About the role
Serve as the technical SRE reference for Identity and Fraud products, defining and monitoring service objectives while improving availability, performance, observability, and security. Lead incident analysis, automate operational processes, support capacity planning, and strengthen system resilience through architectural improvements and reliable deployment practices.
What they look for
Requirements
Requires substantial SRE, DevOps, or Production Engineering experience in mission-critical environments, with strong skills in Kubernetes, Docker, cloud platforms, Terraform, Ansible, observability, and CI/CD. Candidates should be able to troubleshoot and optimize distributed systems, understand relational and non-relational databases, and collaborate effectively; identity and fraud resilience experience, cloud certifications, chaos engineering, and security knowledge are advantageous.
Benefits
Full description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Especialista de SRE based in Brazil.
As a Site Reliability Engineering specialist, you will play a key role in ensuring the reliability and resilience of critical products in a large-scale technology environment. You will serve as a technical reference for SRE practices, partnering closely with development, product, and operations teams. The role focuses on high availability, performance, observability, automation, and secure infrastructure practices. You will help define and monitor SLIs, SLOs, and SLAs aligned with business objectives. Your expertise will contribute to incident prevention, capacity planning, architectural resilience, and efficient recovery from failures. You will work with modern cloud, Kubernetes, infrastructure-as-code, CI/CD, and observability technologies. This is an opportunity to solve complex reliability challenges while strengthening the resilience of mission-critical systems.
\n
Accountabilities:
- Act as the technical SRE reference for Identity & Fraud products, supporting development and operations teams.
- Define, implement, and monitor SLIs, SLOs, and SLAs aligned with business goals and service expectations.
- Lead incident analysis and implement preventive and corrective actions to reduce recurring issues.
- Automate provisioning, deployment, scaling, and failure-recovery processes to improve operational efficiency and reliability.
- Design and maintain observability solutions covering logs, metrics, traces, and alerts.
- Support capacity and performance engineering to ensure systems can handle demand predictably.
- Contribute to architectural improvements focused on resilience, scalability, and security.
- Promote infrastructure-as-code, CI/CD, version control, and safe change-management practices.
- Troubleshoot and mitigate issues in real time within critical production environments.
Requirements:
- Solid experience in SRE, DevOps, or Production Engineering within mission-critical environments.
- Strong expertise with Kubernetes, Docker, and cloud platforms such as AWS, OCI, Azure, and GCP.
- Advanced knowledge of automation and infrastructure as code, including Terraform and Ansible.
- Experience with monitoring and observability, particularly Datadog, along with familiarity with Prometheus, ELK, and Grafana.
- Hands-on experience with CI/CD pipelines, version control, and reliable deployment practices.
- Strong ability to analyze performance, troubleshoot complex issues, and optimize distributed systems.
- Knowledge of relational and non-relational databases.
- Ability to collaborate effectively with development, product, and operations teams.
- Strong communication, systems thinking, analytical skills, and a problem-solving mindset.
- Experience with resilience engineering in identity and fraud systems is desirable.
- Cloud certifications in AWS, OCI, Azure, or GCP are a plus.
- Experience with chaos engineering and resilience testing is desirable.
- Knowledge of application and infrastructure security is an advantage.
Benefits:
- Opportunity to work on large-scale, mission-critical technology systems.
- Collaborative environment involving development, product, and operations teams.
- People-focused and inclusive workplace culture.
- Environment designed to support career development, professional growth, and personal well-being.
- Work-life balance supported alongside career and personal commitments.
- Opportunities to work with modern cloud, automation, observability, and reliability technologies.
- Exposure to complex challenges in data, technology, identity, and fraud solutions.
\nHow Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Similar roles
-
Principal Site Reliability Engineer, Platform Engineering: Dedicated
GitLab Canada · $223K–$380K/yr
-
Senior Site Reliability Engineer (SRE)
UJET United States · $140K–$180K/yr
-
Senior Site Reliability Engineer
Apple Cupertino, California, United States
-
Site Reliability Engineer (High Performance Computing)
SpaceX Hawthorne, California, United States · $125K–$195K/yr
-
Site Reliability Engineer
N26 Barcelona, Catalonia, Spain
-
SRE
Hitachi Solutions pune, Maharashtra, India