Senior Lead Site Reliability Engineer
FIS pune, Maharashtra, India
IT Services and IT Consulting · 10,001+ employees
About the role
The Senior Lead Site Reliability Engineer will design, build, and operate highly available, low-latency payment platforms while implementing reliability standards and observability solutions. They will also lead automation initiatives to eliminate toil and collaborate with cross-functional teams to resolve complex production issues.
What they look for
Requirements
Candidates must have 10 to 15 years of IT experience with deep expertise in distributed systems, cloud platforms, and CI/CD pipelines. Strong proficiency in observability tools, scripting languages, and enterprise database management is required to support mission-critical financial infrastructure.
Benefits
Full description
Senior Lead Site Reliability Engineer
Are you curious, motivated, and forward-thinking? At FIS you’ll have the opportunity to work on some of the most challenging and relevant issues in financial services and technology. Our talented people empower us, and we believe in being part of a team that is open, collaborative, entrepreneurial, passionate and above all fun.
About The Role
- We are hiring a Senior Lead Site Reliability Engineer to define, build, and operate always-on, low-latency, and highly secure payment platforms that power large-scale financial transactions.
- This is a senior hands-on engineering role focused on designing, automating, and supporting highly available, mission-critical platforms. You will work at the intersection of distributed systems, cloud platforms, observability, and reliability engineering, helping to improve platform stability, performance, scalability, and operational efficiency.
- You will collaborate with engineering, development, and infrastructure teams to implement reliable solutions, automate operational processes, and resolve complex production issues while remaining deeply involved in the day-to-day technical work.
- The ideal candidate is a seasoned engineer with strong problem-solving skills, a passion for automation, and the ability to independently drive complex technical initiatives from design through production.
What you be doing
- Design and implement reliability architecture and engineering standards across services, platforms, and infrastructure, ensuring solutions are scalable, resilient, and operationally efficient.
- Design, build, and evolve enterprise-grade observability platforms (metrics, logs, traces, SLOs/SLIs) that provide actionable insights into system health, customer experience, and business impact.
- Implement and continuously improve SRE practices including error budgets, capacity modeling, resilience testing, graceful degradation, and operational readiness.
- Architect, build, and enhance automation and self-service platforms that eliminate toil, reduce operational risk, and enable safe, frequent production releases.
- Design, implement, and optimize CI/CD and release engineering solutions, including secure pipelines, deployment automation, quality gates, and rollback strategies.
- Troubleshoot and resolve complex production issues, performing deep root-cause analysis and implementing long-term reliability improvements.
- Collaborate with engineering, platform, and infrastructure teams to improve system reliability, scalability, performance, and operational excellence.
What You Bring
- 10 to 15 Years of IT Experience.
- Deep software engineering expertise with a proven track record of designing, building, and operating large-scale, distributed, API-driven systems in production.
- Strong CI/CD experience with tools such as Azure DevOps, GitHub Actions, Jenkins, Harness, or equivalent, including pipeline design, deployment automation, and release engineering best practices.
- Expertise in observability, alerting, and reliability engineering, using tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or equivalent ecosystems.
- Strong command of cloud platforms and open systems (AWS, Azure, or GCP), including infrastructure-as-code, platform automation, and cloud-native design patterns.
- Hands-on experience across Linux (RHEL), Windows, databases (e.g. SQL Server, Oracle RDBMS), and complex enterprise technology stacks with strong troubleshooting and debugging skills.
- Strong experience designing and operating high-availability, mission-critical platforms with a focus on reliability, performance, scalability, and automation.
Added bonus if you have:
- Strong automation and scripting skills using Python, Bash, Ansible, PowerShell, or similar tools, with additional experience in C#/.NET considered a strong plus.
- Prior ownership of reliability strategy or platform initiatives spanning multiple teams or business units.
- Experience modernizing legacy financial systems into cloud-native or hybrid architectures with a focus on resilience and compliance.
What we offer you
- A multifaceted job with a high degree of responsibility and a broad spectrum of opportunities
- A modern, international work environment and a dedicated and motivated team
- A broad range of professional education and personal development possibilities – FIS is your final career step!
- A competitive salary and benefits
- A variety of career development tools, resources and opportunities
Privacy Statement
FIS is committed to protecting the privacy and security of all personal information that we process in order to provide services to our clients. For specific information on how FIS protects personal information online, please see the Online Privacy Notice.
Sourcing Model
Recruitment at FIS works primarily on a direct sourcing model; a relatively small portion of our hiring is through recruitment agencies. FIS does not accept resumes from recruitment agencies which are not on the preferred supplier list and is not responsible for any related fees for resumes submitted to job postings, our employees, or any other part of our company.
#pridepass
Similar roles
-
Site Reliability Engineer II
Backblaze External Website Bangalore, Karnataka, India
-
Site Reliability Engineer
Firmus Technologies Singapore, Singapore
-
Site Reliability Engineer I
Backblaze External Website Bangalore, Karnataka, India
-
Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic San Francisco, California, United States · $405K–$485K/yr
-
Linux Site Reliability Engineer (SRE)
OCBC Sepang, Selangor, Malaysia
-
Site Reliability Engineer
LSEG Beijing, Beijing, China