Engineer 4 - Site Reliability Engineering
Comcast Chennai, Tamil Nadu, India
Telecommunications · 10,001+ employees
About the role
The role involves ensuring the reliability, scalability, and performance of data systems through monitoring, automation, and incident response. You will collaborate with engineering teams to optimize data pipelines, manage infrastructure, and perform capacity planning for large-scale distributed systems.
What they look for
Requirements
Candidates must have at least 8 years of experience in SRE, DevOps, or Data Operations with proficiency in cloud platforms and big data technologies. A bachelor's degree in Computer Science or a related field is required, along with strong programming skills in languages like Python, Go, or Java.
Benefits
Full description
Comcast brings together the best in media and technology. We drive innovation to create the world's best entertainment and online experiences. As a Fortune 50 leader, we set the pace in a variety of innovative and fascinating businesses and create career opportunities across a wide range of locations and disciplines. We are at the forefront of change and move at an amazing pace, thanks to our remarkable people, who bring cutting-edge products and services to life for millions of customers every day. If you share in our passion for teamwork, our vision to revolutionize industries and our goal to lead the future in media and technology, we want you to fast-forward your career at Comcast.
Job Summary
This job is responsible for the availability, latency, performance, efficiency, change management, monitoring, emergency response, and capacity planning for platforms. It will engage in designing, analyzing and troubleshooting large-scale distributed systems, debugging and optimizing code and automating routine tasks. It will be part of a team consisting of a mix between software and technology infrastructure backgrounds. It will provide subject matter expertise, resolve complex break/fix scenarios and engage broader teams as necessary. It will partner with engineering, vendors and client services to deliver successful technical solutions. Works with limited supervision and direction while executing associated functions and responsibilities. Follows operational practices and independently determines/develops approaches for non-routine solutions. Integrates knowledge of business and functional priorities. Acts as a key contributor in a complex and crucial environment. May lead teams or projects and shares expertise.
Job Description
Freewheel, a Comcast company, provides comprehensive ad platforms for publishers, advertisers, and media buyers. Powered by premium video content, robust data, and advanced technology, we’re making it easier for buyers and sellers to transact across all screens, data types, and sales channels. As a global company, we have offices in nine countries and can insert advertisements around the world.
Position Overview:
FreeWheel is seeking an experienced Data SRE to join Freewheel Data SRE team based in Reston/Chicago. As a member of the Global Operation team, you will be responsible for ensuring the reliability, scalability, and performance of our data systems. Working closely with data engineers and other operation sub-teams, you will manage our data infrastructure, optimize system reliability, automate daily operations, and resolve technical issues that impact our data pipelines and backend data platforms.
Key Responsibilities:
- System Monitoring and Optimization: Design and implement monitoring and alerting systems to ensure the stability, reliability, and performance of data platforms. Quickly respond to and resolve issues impacting data pipelines or storage layers.
- Automation and Tool Development: Develop and maintain automation tools and scripts for deployment, monitoring, backup, recovery, and disaster recovery of data systems.
- Performance Optimization: Analyze and optimize the performance of data storage, query performance, and data flows to ensure efficient processing of large-scale datasets, reduce latency, animprove processing speed.
- Incident Response and Troubleshooting: Respond quickly to data platform failures, perform troubleshooting, and coordinate cross-team efforts to resolve issues and ensure high availability and reliability of data.
- Capacity Planning and Scaling: Work with data engineering teams to analyze and forecast capacity requirements, ensuring the system can handle data growth and scale infrastructure accordingly.
- Documentation and Knowledge Sharing: Document the architecture, configurations, and operational procedures for data platforms, ensuring knowledge is shared across the team and providing relevant training.
- Security and Compliance: Ensure data platforms meet security standards and compliance requirements to prevent data breaches or misuse.
- Cross-Team Collaboration: Collaborate with data science, product, and development teams to support data product design and implementation, solving reliability-related issues.
Qualifications:
- Education: Bachelor’s degree or higher in Computer Science, Software Engineering, or a related field.
- Experience: • At least 8+ years of experience as an SRE,DevOps, or Data Operations Engineer.
- Experience with cloud platforms (e.g.AWS,GCP,Azure)
- Familiarity with modern data architectures and technologies, including big data platforms (e.g., Kafka, Hadoop, Spark), distributed storage (e.g., Cassandra, HDFS, AWS S3), etc.
- Extensive experience in data base management (e.g.,NoSQL databases, MySQL, PostgreSQL).
- Proficiency in automation tools and frameworks (e.g., Ansible, Terraform, Kubernetes, Docker) for automating data system deployment and maintenance.
- Familiarity with modern CICD pipeline.
- Programming Skills: Proficient in at least one programming language, such as Python, Go, Java, or Scala, with the ability to write efficient scripts and automation tools.
- System Monitoring and Log Management: Familiar with using monitoring and log management tools such as Prometheus, Grafana, ELK Stack, or other similar tools.
- Troubleshooting and Debugging: Strong debugging and troubleshooting skills, with the ability to quickly identify and resolve production issues.
- Team Collaboration and Communication: Excellent communication skills with the ability to convey technical information clearly and concisely to both technical and non-technical stakeholders.
Bonus Skills:
- Experience with Aerospike, Kafka, Snowflake and other big data tech stack.
- Familiarity with containerization, micro-services architecture.
- Experience in designing and maintaining large-scale distributed systems.
- Experience in data quality management, data governance, or ETL pipelines.
- Experience in audience targeting and Identity services product support.
Employees at all levels are expected to:
Understand our Operating Principles; make them the guidelines for how you do your job. Own the customer experience - think and act in ways that put our customers first, give them seamless digital options at every touchpoint, and make them promoters of our products and services. Know your stuff - be enthusiastic learners, users and advocates of our game-changing technology, products and services, especially our digital tools and experiences. Win as a team - make big things happen by working together and being open to new ideas. Be an active part of the Net Promoter System - a way of working that brings more employee and customer feedback into the company - by joining huddles, making call backs and helping us elevate opportunities to do better for our customers. Drive results and growth. Support a culture of inclusion in how you work and lead. Do what's right for each other, our customers, investors and our communities.
Disclaimer: This information has been designed to indicate the general nature and level of work performed by employees in this role. It is not designed to contain or be interpreted as a comprehensive inventory of all duties, responsibilities and qualifications.
We believe that benefits should connect you to the support you need when it matters most, and should help you care for those who matter most. That's why we provide an array of options, expert guidance and always-on tools that are personalized to meet the needs of your reality—to help support you physically, financially and emotionally through the big milestones and in your everyday life.
Please visit the benefits summary on our careers site for more details.
Education
Bachelor's Degree
While possessing the stated degree is preferred, Comcast also may consider applicants who hold some combination of coursework and experience, or who have extensive related professional experience.
Certifications (if applicable)
Relevant Work Experience
7-10 Years
Comcast is an equal opportunity workplace. We will consider all qualified applicants for employment without regard to race, color, religion, age, sex, sexual orientation, gender identity, national origin, disability, veteran status, genetic information, or any other basis protected by applicable law.
Similar roles
-
Staff Site Reliability Engineer
NinjaTrader Chicago, Illinois, United States · $160K–$210K/yr
-
Senior Site Reliability Engineer – Network Observability
Blueprint Technologies $104K–$114K/yr
-
Senior Site Reliability Engineer (Cloud Platform)
Salve.Inno Consulting Denver, Colorado, United States
-
Amazon Connect SRE
Miratech Surat, Gujarat, India
-
AWS Site Reliability Engineer
Miratech Ahmedabad, Gujarat, India
-
Senior Data SRE / Cloud Platform Engineer
Flywire Valencia, Valencian Community, Spain · €49K–€62K/yr