Site Reliability Engineer
Alvaria Inc Texas, United States
Software Development · 1,001-5,000 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
The Site Reliability Engineer will design, build, and operate highly available, secure, and scalable cloud infrastructure while partnering with teams to define service-level objectives. They will also automate repetitive tasks, improve system observability, and lead incident response efforts to ensure platform resilience.
What they look for
Requirements
Candidates must have a bachelor's degree in a relevant field and 4-5+ years of experience in site reliability, DevOps, or cloud operations. Proficiency in at least one programming language and hands-on experience with AWS and infrastructure-as-code tools are required.
Benefits
Full description
About Aspect Software
Building on more than 50 years of industry experience, Aspect Software is reimagining workforce management through cloud technology, AI, automation, and human-centered innovation. Our Workforce Engagement Management solutions help organizations solve complex workforce challenges, improve operational performance, and deliver better employee and customer experiences.
We foster a collaborative environment where engineers work across teams and geographies to build secure, scalable, and dependable software. Join us as we modernize intelligent workforce systems used by organizations around the world.
Position Overview
We are seeking a Site Reliability Engineer to help build and operate reliable, secure, and efficient cloud services. In this role, you will combine software engineering and systems expertise to improve availability, performance, scalability, observability, and operational readiness across our SaaS platforms.
This person will partner closely with application engineering, platform, security, and product teams to define service-level objectives, automate repetitive work, strengthen incident response, and design systems that remain resilient as they scale. The ideal candidate is a pragmatic problem-solver who measures what matters, learns from failures, and improves systems through automation rather than manual intervention.
Key Responsibilities
- Design, build, and operate highly available, scalable, and secure production services and cloud infrastructure
- Partner with engineering teams to define and maintain service-level indicators (SLIs), service-level objectives (SLOs), error budgets, and actionable alerts
- Build automation and self-service tooling that reduces toil, accelerates safe delivery, and improves operational consistency
- Improve observability through meaningful metrics, logs, distributed tracing, dashboards, synthetic checks, and health monitoring
- Participate in an on-call rotation and lead or support incident response, communication, mitigation, and recovery
- Facilitate blameless post-incident reviews and ensure corrective actions address systemic causes
- Improve CI/CD pipelines using automated validation, deployment health gates, progressive delivery, canary releases, and reliable rollback strategies
- Review system designs for reliability, performance, capacity, security, recoverability, and cost efficiency
- Strengthen resilience through redundancy, autoscaling, fault-tolerant patterns, disaster recovery planning, and appropriate load or chaos testing
- Troubleshoot complex issues across applications, infrastructure, networking, databases, and third-party dependencies
- Develop and maintain runbooks, operational standards, architectural documentation, and support procedures
- Collaborate across teams and time zones while clearly communicating risk, trade-offs, incident status, and reliability priorities
Required Qualifications
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience
- 4–5+ years of experience in site reliability engineering, DevOps, platform engineering, cloud operations, or production software engineering
- Hands-on experience operating production workloads in AWS; comparable experience with Azure or GCP is also valuable
- Proficiency in at least one programming or scripting language such as Python, Go, TypeScript/JavaScript, or C#
- Experience with infrastructure as code and configuration automation, such as AWS CDK, CloudFormation, Terraform, or similar tools
- Experience building or supporting CI/CD pipelines and source-control workflows, preferably with GitHub and GitHub Actions
- Working knowledge of Linux, networking, DNS, load balancing, firewalls, certificates, and common distributed-systems failure modes
- Experience with observability and monitoring tools such as Datadog, Grafana, CloudWatch or equivalent platforms
- Understanding of incident management, root-cause analysis, and blameless post-incident practices
- Strong troubleshooting, documentation, collaboration, and communication skills
Preferred Qualifications
- Experience with containerized or serverless architecture, including Kubernetes, Docker, AWS Lambda, or related technologies
- Experience with AWS services such as API Gateway, CloudFront, S3, IAM, WAF, DynamoDB, Aurora, EventBridge, SQS, Kinesis, Cognito, VPC, Route 53, and Secrets Manager
- Experience designing multi-region systems, disaster recovery strategies, backup and restoration processes, and business-continuity controls
- Experience with performance testing, capacity planning, chaos engineering, or reliability testing
- Familiarity with security, privacy, audit, and compliance requirements for enterprise SaaS products
- Experience supporting data-intensive, real-time, or high-volume enterprise applications
- Experience mentoring engineers or leading cross-team reliability initiatives
Why Join Us?
- Help shape the reliability practices behind modern workforce technology
- Work on meaningful, technically challenging systems used by enterprise customers
- Collaborate with skilled colleagues across engineering, product, security, and operations
- Influence architecture, tooling, and operational standards as our cloud platforms evolve
- Access professional development and career-growth opportunities
- Join a team that values innovation, accountability, inclusion, and measurable customer outcomes
This job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee. Duties, responsibilities and activities may change or new ones may be assigned at any time with or without notice.
Similar roles
-
Senior Manager, Site Reliability Engineering
Oracle Nashville, Tennessee, United States
-
Senior Site Reliability Engineer
RADAR San Diego, California, United States · $170K–$219K/yr
-
Intermediate Site Reliability Engineer
ContactMonkey Toronto, Ontario, Canada · CA$130K–CA$150K/yr
-
Site Reliability Engineer
ACI Worldwide Timișoara, Romania
-
Security Site Reliability Engineer - Apple Service Engineering
Apple Seattle, Washington, United States
-
Software Engineer III, Site Reliability Engineering
Google Fremont, California, United States · $147K–$210K/yr