Jobgether

Site Reliability Engineer - AWS

Jobgether United States · $110K–$140K/yr

Internet Marketplace Platforms · 11-50 employees

19 h ago
Remote sre Senior (5-10 yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The Site Reliability Engineer will ensure the stability, scalability, and performance of cloud-based SaaS products by managing AWS infrastructure and automating operational processes. They will also participate in incident response, troubleshoot system issues, and collaborate with development teams to improve production reliability.

What they look for

AWS Site Reliability Engineering DevOps Kubernetes Terraform CI/CD Python Bash PowerShell Observability Infrastructure as Code GitOps Amazon EKS Docker SQL Networking

Requirements

Candidates must have at least 5 years of experience in DevOps, cloud infrastructure, or site reliability engineering with a strong background in AWS and containerization. Proficiency in automation scripting, infrastructure-as-code tools like Terraform, and CI/CD methodologies is required.

Benefits

Paid time off Paid holidays Medical insurance Dental insurance Vision insurance Life insurance Disability coverage 401(k) retirement plan Discretionary bonus

Full description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer - AWS based in the United States.

This is a fully remote opportunity focused on ensuring the stability, scalability, security, and performance of cloud-based SaaS products. You will support applications recently migrated from on-premises environments to AWS and help establish strong cloud-native operating practices. The role combines site reliability engineering, DevOps, cloud infrastructure, automation, observability, and application support. You will work closely with development, platform, infrastructure, and business teams to improve reliability and enable secure, manageable production releases. A significant part of the role involves automation, infrastructure optimization, incident response, modernization, and continuous operational improvement. You will have broad exposure to AWS services, Kubernetes, GitOps, Terraform, CI/CD, monitoring, databases, networking, and scripting. This is an ideal environment for a hands-on engineer who enjoys solving complex operational challenges and building resilient technology platforms.

\n

Accountabilities:

  • Support and continuously improve production applications running in AWS, ensuring high availability, reliability, performance, and operational stability.
  • Monitor system health, availability, latency, performance, logs, metrics, traces, and alerts, proactively identifying and addressing reliability risks.
  • Participate in incident response, troubleshooting, escalation management, root cause analysis, and post-incident improvement activities.
  • Strengthen operational readiness, resiliency, disaster recovery capabilities, and overall production support processes.
  • Engineer and optimize AWS environments using services such as EC2, ECS/EKS, Lambda, S3, RDS, CloudWatch, IAM, VPC, Elastic Beanstalk, Load Balancers, and related cloud technologies.
  • Apply AWS Well-Architected Framework principles across reliability, security, performance efficiency, cost optimization, and operational excellence.
  • Support cloud-native architecture decisions and contribute to infrastructure modernization and optimization initiatives.
  • Build and enhance end-to-end observability through dashboards, monitoring, alerting, logging, metrics, and tracing.
  • Automate recurring operational processes using Python, Bash, PowerShell, CI/CD pipelines, Terraform, and infrastructure-as-code practices.
  • Support applications across both on-premises and AWS environments, contributing to modernization and migration initiatives.
  • Partner with application development teams to identify performance bottlenecks, infrastructure constraints, security concerns, and reliability risks.
  • Manage and improve containerized environments using Kubernetes, Docker, and Amazon EKS, including GitOps-based deployment approaches.
  • Support CI/CD and release management processes using tools and methodologies such as Git, Azure DevOps, ArgoCD, and FluxCD.
  • Troubleshoot APIs, microservices, network connectivity, HTTP-based applications, and cloud infrastructure to improve application performance and uptime.
  • Manage and optimize database environments, including RDS configuration, schemas, users, performance troubleshooting, data integrity, and storage practices.
  • Collaborate across Development, Business, Platform, and Infrastructure teams in Agile environments using Scrum and Kanban practices.
  • Serve as a technical point of contact for infrastructure, reliability, and operational requirements within product teams.
  • Continuously identify opportunities for reengineering, process improvement, efficiency gains, automation, and cloud optimization.

Requirements

  • Bachelor's degree or an equivalent combination of education and professional experience.
  • 5+ years of relevant industry or technical experience, with substantial hands-on experience in DevOps, cloud infrastructure, site reliability, or related engineering disciplines.
  • Strong DevOps background, ideally with an emphasis on approximately 70% operations and 30% development activities.
  • Extensive experience with AWS cloud infrastructure and cloud-based production environments.
  • Strong experience with containerization and orchestration technologies, particularly Kubernetes, Docker, and Amazon EKS.
  • Hands-on experience building and managing CI/CD pipelines using GitOps methodologies and tools such as ArgoCD or FluxCD.
  • Strong knowledge of Git-based version control and experience with Azure DevOps, including tickets, releases, CI/CD pipelines, and infrastructure-as-code workflows.
  • Strong Terraform experience and knowledge of infrastructure-as-code best practices.
  • Advanced scripting and automation skills using Python, Bash, and/or PowerShell.
  • Experience with monitoring and observability tools, particularly AWS CloudWatch, including alert configuration, troubleshooting, and escalation management.
  • Strong understanding of AWS services including EC2, AMIs, S3, RDS, Lambda, ECS/EKS, IAM, VPC, Elastic Beanstalk, Load Balancers, and Transfer Family.
  • Experience managing large-file transfers and supporting highly available cloud infrastructure.
  • Strong knowledge of APIs and microservices, including configuration, performance tuning, security, and reliability best practices.
  • Solid understanding of HTTP concepts and protocols, with the ability to analyze and optimize web application performance.
  • Experience using diagnostic tools such as curl and wget to troubleshoot connectivity and network performance issues.
  • Strong SQL and database administration skills, including RDS configuration, schemas, catalogs, users/logins, synonyms, performance troubleshooting, and database hygiene.
  • Strong networking and network management capabilities.
  • Experience with service mesh technologies such as Linkerd or Istio is a plus.
  • Excellent written and verbal communication skills, with the ability to collaborate effectively across technical and business teams.
  • Comfortable working independently in a remote environment and adapting to flexible work hours.
  • Strong troubleshooting, analytical, problem-solving, and continuous-improvement mindset.

Benefits

  • Annual full-time base salary range of $110,000–$140,000, with final compensation determined by experience, education, skills, training, geographic location, and market considerations.
  • Potential eligibility for a discretionary bonus based on applicable bonus program guidelines and position eligibility.
  • Fully remote work arrangement.
  • Paid time off and paid holidays in accordance with applicable company policies.
  • Medical, dental, and vision insurance options.
  • Life insurance and disability coverage.
  • 401(k) retirement plan.
  • Opportunity to work on cloud modernization, SaaS reliability, DevOps, automation, and large-scale AWS environments.
  • Exposure to a broad technology stack spanning AWS, Kubernetes, Terraform, GitOps, CI/CD, observability, networking, and database technologies.
  • Collaborative environment with opportunities to work closely with development, platform, infrastructure, and business teams.
  • Flexible work environment designed to support distributed teams and evolving work requirements.

\nHow Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Similar roles