Site Reliability Engineer
Workiy · Scottsdale, Arizona, United States
Information Technology & Services · 11-50 employees
About the role
The Site Reliability Engineer will support large-scale, high-performance applications within a hybrid cloud and on-premises environment. Responsibilities include building automation scripts, managing transaction journeys, and implementing observability for real-time monitoring and incident resolution.
What they look for
Requirements
Candidates must have extensive experience in cloud infrastructure, containerization, and managing high-performance applications. Proficiency in programming languages such as Go, Python, or Java and knowledge of networking protocols and database systems is required.
Full description
We are seeking an experienced Site Reliability Engineer (SRE) to support large-scale, high-performance applications running in a hybrid environment (on-premises and cloud). The ideal candidate will have strong experience in cloud infrastructure, Kubernetes, observability, automation, and production operations.
Requirements
- Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
- Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
- Experience working with Programming languages such as Go, Python, Java, Rust etc.
- Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any time-series databases
- Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher
- Experience maintaining containerized app in GKE/RKE/AKE environments.
- Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
- Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
- Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
- Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
- Experience working with Programming languages such as Go, Python, Java, Rust etc.
- Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any time-series databases
- Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher
- Experience maintaining containerized app in GKE/RKE/AKE environments.
- Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
- Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
- Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.