Workiy

Site Reliability Engineer

Workiy · Scottsdale, Arizona, United States

Information Technology & Services · 11-50 employees

Yesterday
Senior (5-10 yrs) Contractor United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The Site Reliability Engineer will support large-scale, high-performance applications within a hybrid cloud and on-premises environment. Responsibilities include building automation scripts, managing transaction journeys, and implementing observability for real-time monitoring and incident resolution.

What they look for

Site Reliability Engineering Kubernetes Cloud Infrastructure Automation Go Python Java Rust GCP Rancher Observability OTEL GraphQL Networking Protocols Database Management Distributed Tracing

Requirements

Candidates must have extensive experience in cloud infrastructure, containerization, and managing high-performance applications. Proficiency in programming languages such as Go, Python, or Java and knowledge of networking protocols and database systems is required.

Full description

We are seeking an experienced Site Reliability Engineer (SRE) to support large-scale, high-performance applications running in a hybrid environment (on-premises and cloud). The ideal candidate will have strong experience in cloud infrastructure, Kubernetes, observability, automation, and production operations.

Requirements

  • Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
  • Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
  • Experience working with Programming languages such as Go, Python, Java, Rust etc.
  • Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any time-series databases
  • Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher
  • Experience maintaining containerized app in GKE/RKE/AKE environments.
  • Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
  • Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
  • Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
  • Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
  • Experience working with Programming languages such as Go, Python, Java, Rust etc.
  • Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any time-series databases
  • Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher
  • Experience maintaining containerized app in GKE/RKE/AKE environments.
  • Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
  • Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
  • Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.