DevOps / Site Reliability Engineer (SRE)
Technology, Information and Media · 11-50 employees
About the role
Design and operate automated CI/CD pipelines, infrastructure-as-code templates, and Kubernetes-based production workloads across cloud platforms. Improve reliability and performance through observability, high-availability and disaster-recovery measures, secure secrets management, on-call support, and incident root-cause analysis.
What they look for
Requirements
Requires 5–9 years of systems engineering or software development experience, including at least 4 years building and operating cloud-native production platforms. Candidates need strong Kubernetes, Terraform, Linux, scripting, networking, and cloud-platform knowledge, plus one of the specified AWS, Azure, or Kubernetes certifications.
Full description
This is a remote position.
DevOps / Site Reliability Engineer (SRE)
Job Details
- Employment Type: Contract
- Work Mode: Remote
- Location: Offshore
- Total Experience Required: 5 to 9 years
- Relevant Experience Required: 4+ years of dedicated experience in infrastructure automation, cloud orchestration, and CI/CD pipelines
- Mandatory Certification: AWS Certified DevOps Engineer - Professional, Microsoft Certified: DevOps Engineer Expert, or Certified Kubernetes Administrator (CKA)
Job Summary
We are seeking an experienced DevOps / Site Reliability Engineer (SRE) to design, automate, and scale our cloud-native infrastructure pipelines. The ideal candidate will bridge the gap between software development and systems operations, building highly available deployment pipelines, writing infrastructure-as-code (IaC), and optimizing cluster scaling to maximize application uptime, system reliability, and performance.
Key Responsibilities
- Design, build, and optimize automated CI/CD pipelines across cloud platforms using industry-standard automation servers (e.g., GitHub Actions, GitLab CI, Jenkins, ArgoCD).
- Architect and manage infrastructure-as-code (IaC) templates using Terraform or OpenTofu to provision secure, modular, and repeatable multi-environment architectures.
- Orchestrate containerized production workloads, configuring cluster scaling, service meshes, network routing policies, and deployment strategies on Kubernetes (EKS/AKS/GKE).
- Implement automated monitoring, logging, and alerting systems utilizing observability tools (e.g., Prometheus, Grafana, Datadog, ELK stack) to actively track platform performance metrics.
- Drive system high-availability and fault tolerance efforts, designing disaster recovery plans, automated load balancing parameters, and self-healing cluster scripts.
- Manage centralized configuration and secret management systems, securely vaulting database credentials, API tokens, and certificate profiles (e.g., HashiCorp Vault, AWS Secrets Manager).
- Participate in on-call rotations and lead incident root-cause analysis (RCA), systematically diagnosing runtime infrastructure failures, performance bottlenecks, and resource leaks.
Requirements
- 5 to 9 years of core systems engineering or software development experience, with 4+ dedicated years actively designing, building, and operating cloud-native production platforms.
- Strong technical mastery of Kubernetes cluster administration, Terraform automation layouts, Linux system internals, shell scripting (Bash, Python, or Go), and network protocols.
- Deep structural understanding of microservices design, caching mechanics, database scaling limits, and cloud provider API governance.
- Mandatory certification: AWS DevOps Professional, Azure DevOps Expert, or CKA.
Preferred Qualifications
- Prior experience implementing DevSecOps controls (e.g., integrating SAST/DAST tools directly into container build phases).
- Experience with GitOps methodologies and progressive delivery mechanisms (e.g., Canary or Blue/Green deployments using Flagger or Istio).
Benefits
- 12+ years of total IT software engineering or operational management background, with 6+ dedicated years acting as a CISO, Director of Security, or Principal Enterprise GRC Advisor.
- Strong visionary mastery of modern security trends, zero-trust target states, risk calculation paradigms, and multi-cloud information landscape parameters.
- Deep communication execution skills, with a proven history of negotiating security budgets, steering board panels, and handling high-pressure public communication events.
- Mandatory certification: CISM or CISSP.
Preferred Qualifications
- Certified in the Governance of Enterprise IT (CGEIT) or Certified in Risk and Information Systems Control (CRISC) credential.
- Prior experience steering complex post-merger information platform integrations or stabilizing security posture profiles during major corporate equity restructurings.
Similar roles
-
Principal Site Reliability Engineer, Platform Engineering: Dedicated
GitLab Canada · $223K–$380K/yr
-
Senior Site Reliability Engineer (SRE)
UJET United States · $140K–$180K/yr
-
Senior Site Reliability Engineer
Apple Cupertino, California, United States
-
Site Reliability Engineer (High Performance Computing)
SpaceX Hawthorne, California, United States · $125K–$195K/yr
-
Site Reliability Engineer
N26 Barcelona, Catalonia, Spain
-
SRE
Hitachi Solutions pune, Maharashtra, India