Senior Site Reliability Engineer
Publicis Groupe Holdings B.V Bengaluru, Karnataka, India
Information Services · 11-50 employees
About the role
Design and implement reliability strategies for distributed systems while leading incident response and root cause analysis. Collaborate with engineering teams to improve system performance, scalability, and operational readiness through automation and observability.
What they look for
Requirements
Requires a bachelor's degree in Computer Science or equivalent and at least 7 years of experience in SRE, Cloud, or DevOps roles. Candidates must have deep expertise in Linux administration, Kubernetes, cloud platforms, and infrastructure automation.
Full description
Overview
About Business Unit:
At the core of all that Epsilon does is a team that sets the foundation of our IT infrastructure. The team drives innovation and efficiency through pioneering technology across Epsilon's platforms and business verticals. From being the first point of contact for infrastructure needs to final deployment, the team provides end-to-end solutions for our client-facing platforms. ETS supports all aspects of revenue-generating platforms for Epsilon and sets the architectural direction for our enterprise deployments. By adopting the newest technologies, such as Cloud, Automation, and Artificial Intelligence, the team is at the front of redefining our digital business and capturing new opportunities.
Overview:
Publicis Epsilon is seeking a Senior Site Reliability Engineer to help build, operate, and evolve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. As part of a large-scale cloud transformation initiative, you will partner closely with Engineering, DevOps, Platform, and Security teams to establish reliability practices, improve operational excellence, and ensure systems meet performance, availability, and scalability objectives.
This is a hands-on technical leadership role requiring deep expertise in cloud infrastructure, Kubernetes, observability, incident management, and reliability engineering. You will drive technical decisions, influence engineering practices, and help teams design systems that are resilient by design.
Click here to view how Epsilon transforms marketing with 1 View, 1 Vision and 1 Voice.
Responsibilities
Your Impact:
- Design and implement reliability strategies for distributed systems running across AWS and GCP.
- Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics.
- Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
- Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
- Collaborate with engineering teams to improve system performance, resiliency, scalability, and operational readiness.
- Automate operational processes and reduce toil through engineering solutions.
- Guide teams on reliability-focused architecture decisions, capacity planning, and non-functional requirements.
- Experience with configuration management and automation tools (e.g., Ansible, Puppet, Chef).
- Strong experience with Linux system administration including user management, file systems, networking, and performance tuning.
- Experience managing containers and orchestration (Docker, Kubernetes) is a plus.
- Familiarity with monitoring and logging tools (Nagios, Prometheus, Grafana, ELK stack).
- Experience with databases such as MySQL, PostgreSQL, or equivalent tools.
- Develop and enforce system standards, automation practices, and recommend improvements to enhance performance and reliability.
- Install, configure, and maintain commercial and open-source applications on Linux operating systems (e.g., RHEL, CentOS, Ubuntu).
Qualifications
Education, Experience, and Licensing Requirements:
- Bachelor’s degree in Computer Science or related field (or equivalent experience).
- Minimum of 5 years of experience in Linux system administration or IT infrastructure.
- Experience with databases such as MySQL, PostgreSQL, or equivalent tools.
- Experience working across multiple operating systems with strong emphasis on Linux platforms.
- Relevant certifications preferred:
- Red Hat Certified System Administrator (RHCSA) or Engineer (RHCE)
- Linux+ or equivalent
- VMware or cloud certifications (AWS, Azure) are a plus
- Any Core tools mentioned above
Skills & Experience:
- 7+ years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, or Platform Engineering.
- Strong experience supporting production systems in AWS/GCP/Azure environments.
- Deep understanding of SRE principles, including SLIs, SLOs, error budgets, and operational excellence.
- Experience operating and troubleshooting Kubernetes platforms such as EKS and/or GKE.
- Strong knowledge of observability tools such as Prometheus, Grafana, CloudWatch, Cloud Monitoring, Datadog, Splunk, or similar.
- Experience with Infrastructure as Code tools such as Terraform.
- Strong scripting and automation skills using Python, Bash, or comparable languages.
- Solid understanding of networking, distributed systems, cloud security, and performance optimization.
Set Yourself Apart With:
- Experience supporting large-scale cloud migration or modernization programs.
- Expertise in incident management and production operations for high-availability systems.
- Experience implementing chaos engineering or resilience testing practices.
- Knowledge of service mesh technologies such as Istio.
- RHEL/AWS/AZURE/GCP certifications.
- Experience working in Agile, DevOps, or DevSecOps environments.
Additional Information
Epsilon is a global data, technology and services company that powers the marketing and advertising ecosystem. For decades, we’ve provided marketers from the world’s leading brands the data, technology and services they need to engage consumers with 1 View, 1 Vision and 1 Voice. 1 View of their universe of potential buyers. 1 Vision for engaging each individual. And 1 Voice to harmonize engagement across paid, owned and earned channels.
Epsilon’s comprehensive portfolio of capabilities across our suite of digital media, messaging and loyalty solutions bridge the divide between marketing and advertising technology. We process 400+ billion consumer actions each day using advanced AI and hold many patents of proprietary technology, including real-time modeling languages and consumer privacy advancements. Thanks to the work of every employee, Epsilon has been consistently recognized as industry-leading by Forrester, Adweek and the MRC. Epsilon is a global company with more than 9,000 employees around the world.
Our pillars aren't just words. They're how we show up every day.
- People centricity: We focus on employee well-being in an environment where colleagues truly care about each other.
- Collaboration: We work together, support one another, and collectively achieve goals.
- Growth: There are endless opportunities for growth through learning, development and career advancement.
- Innovation: We drive progress through cutting-edge solutions and forward-thinking approaches.
- Flexibility: We’ve created a balance between work and personal life, and we encourage adaptability to solve problems creatively.
Our values guide us to create value for our clients, our people and consumers.
- Act with integrity
- Work together to win together
- Innovate with purpose
- Respect all voices
- Empower with accountability
These pillars and values are our foundation—shaping our culture, guiding our decisions, and uniting us in common purpose.
Epsilon is an Equal Opportunity Employer. Epsilon is committed to promoting diversity, inclusion, and equal employment opportunities by using reasonable efforts to attract, recruit, engage and retain qualified individuals of all ethnicities and backgrounds, including, but not limited to, women, people of color, LGBTQ individuals, people with disabilities and any other underrepresented groups, traits or characteristics.
Similar roles
-
Site Reliability Engineer
Workiy Barrington, Rhode Island, United States
-
Site Reliability Developer 4
Oracle Austin, Texas, United States · $102K–$210K/yr
-
Senior Site Reliability Engineer
Cross River Fort Lee, New Jersey, United States · $160K–$200K/yr
-
Software Dev Senior Engineer – SRE & Cloud Reliability
SonicWall Pune, Maharashtra, India
-
Manager, Site Reliability Engineering (Auth0)
Okta Washington, District of Columbia, United States · $182K–$251K/yr
-
Site Reliability Engineer -Jersey City, NJ & Dallas, TX
StradIT Dallas, Texas, United States