About the role
The SRE Engineer II will implement automation tools and CI/CD pipelines to enhance service reliability and performance. They will also manage risks, troubleshoot large-scale distributed systems, and collaborate with engineering teams to resolve production issues.
What they look for
Requirements
Candidates should possess expertise in infrastructure as code, observability, and incident management. Proficiency in tools such as Terraform, Kubernetes, and various monitoring platforms is essential for maintaining system scalability and efficiency.
Benefits
Full description
Company Description
Entain India is the engineering and delivery powerhouse for Entain, one of the world's leading global sports and gaming groups. Established in Hyderabad in 2001, we've grown from a small tech hub into a dynamic force, delivering modern software solutions and support services that power billions of transactions for millions of users worldwide.
Our focus on quality at scale drives us to create innovative technology that supports Entain’s mission to lead the change in global sports and gaming sector. At Entain India, we make the impossible possible, together.
Job Description
We are seeking a talented and motivated SRE Engineer II to join our dynamic team. In this role, you will execute a range of site reliability activities, ensuring optimal service performance, reliability, and availability. You will collaborate with cross-functional engineering teams to develop scalable, fault-tolerant, and cost-effective cloud services.
If you are passionate about site reliability engineering and ready to make a significant impact, we would love to hear from you!
Key Responsibilities:
● Implement automation tools, frameworks, and CI/CD pipelines, promoting best practices and code reusability.
● Enhance site reliability through process automation, reducing mean time to detection, resolution, and repair.
● Identify and manage risks through regular assessments and proactive mitigation strategies.
● Develop and troubleshoot large-scale distributed systems in both on-prem and cloud environments.
● Deliver infrastructure as code to improve service availability, scalability, latency, and efficiency.
● Monitor support processing for early detection of issues and share knowledge on emerging site reliability trends.
● Analyze data to identify improvement areas and optimize system performance through scale testing.
● Take ownership of production issues within assigned domains, performing initial triage and collaborating closely with engineering teams to ensure timely resolution.
Qualifications
For Site Reliability Engineering (SRE), key skills and tools are essential for maintaining system reliability, scalability, and efficiency. Given your expertise in observability, compliance, and platform stability, here’s a structured breakdown:
Key SRE Skills
- Infrastructure as Code (IaC) – Automating provisioning with Terraform, Ansible, or Kubernetes.
- Observability & Monitoring – Implementing distributed tracing, logging, and metrics for proactive issue detection.
- Security & Compliance – Ensuring privileged access controls, audit logging, and encryption.
- Incident Management & MTTR Optimization – Reducing downtime with automated recovery mechanisms.
- Performance Engineering – Optimizing API latency, P99 response times, and resource utilization.
- Dependency Management – Ensuring resilience in microservices with circuit breakers and retries.
- CI/CD & Release Engineering – Automating deployments while maintaining rollback strategies.
- Capacity Planning & Scalability – Forecasting traffic patterns and optimizing resource allocation.
- Chaos Engineering – Validating system robustness through fault injection testing.
- Cross-Team Collaboration – Aligning SRE practices with DevOps, security, and compliance teams.
Essential SRE Tools
- Monitoring & Observability: Datadog, Prometheus, Grafana, New Relic.
- Incident Response: PagerDuty, OpsGenie.
- Configuration & Automation: Terraform, Ansible, Puppet.
- CI/CD Pipelines: Jenkins, GitHub Actions, ArgoCD.
- Logging & Tracing: ELK Stack, OpenTelemetry, Jaeger.
- Security & Compliance: Vault, AWS IAM, Snyk.
Additional Information
At Entain India, we offer a package and the support people need to make an impact. Join us, and a great compensation package is just the beginning. You can expect to receive benefits like:
- Safe home pickup and home drop.
- A regular bonus and great pension.
- 24 days annual leave.
- Extra paid leave, including wellbeing and development days.
- Life assurance and Income Protection.
- Private healthcare and wellbeing support.
- INR 3,000 per month Communication allowance.
- Up to INR 16,000 per year in Crèche expenses (children under 3).
Equal Opportunities.
If you need any reasonable adjustments at any stage of the recruitment process, please contact us and we'll support you.
We're committed to creating a diverse, equitable and inclusive workplace where everyone feels valued, respected and able to be themselves.
We're an equal opportunities employer. We welcome applications from everyone, and we do not discriminate based on age, disability, gender or gender reassignment, pregnancy or maternity, race, religion or belief, sexual orientation, marriage/civil partnership, or any other basis.
We comply with all applicable recruitment regulations and employment laws in the jurisdictions where we operate, ensuring ethical and compliant hiring practices globally.
- Advertising Department: Technology
Similar roles
-
Site Reliability Engineer (SRE)
Raydar New York, New York, United States · $120K–$180K/yr
-
Senior SRE
Tango Limassol, Cyprus, Cyprus
-
IN-Manager_Site Reliability Architect/Engineer _GCC_Advisory_Pune
PwC pune, Maharashtra, India
-
Site Reliability Engineer
Lloyds Banking Group London, England, United Kingdom · £84K–£93K/yr
-
Observability Engineer - Site Reliability Engineering (SRE)
OCBC Singapore, Singapore, Singapore
-
Site Reliability Engineer
LSEG Port Blair, Andaman and Nicobar Islands, India