Site Reliability Engineering (SRE) Director
AARP Washington, District of Columbia, United States
Non-profit Organizations · 1,001-5,000 employees
About the role
The SRE Director leads the strategy and operational discipline to ensure the reliability, availability, and scalability of technical platforms. This role oversees incident management, automation, and observability practices while partnering with digital teams to embed operational excellence.
What they look for
Requirements
Candidates must have a bachelor's degree and at least 8 years of progressive experience in SRE, DevOps, or infrastructure engineering. The role requires 5+ years of experience in cloud-native architectures and modern software delivery practices.
Benefits
Full description
Overview
AARP is the nation's largest nonprofit, nonpartisan organization dedicated to empowering people 50 and older to choose how they live as they age. With a nationwide presence, AARP strengthens communities and advocates for what matters most to the more than 100 million Americans 50-plus and their families: health and financial security, and personal fulfillment. AARP also works for individuals in the marketplace by sparking new solutions and allowing carefully chosen, high-quality products and services to carry the AARP name. As a trusted source for news and information, AARP produces the nation's largest-circulation publications, AARP The Magazine and the AARP Bulletin.
The Site Reliability Engineering (SRE) Director is responsible for leading the strategy, operational discipline, and engineering practices that ensure the reliability, availability, scalability, and performance of technical and digital products and platforms. This role drives automation, observability, incident response maturity, and continuous improvement while building a SRE function that supports both enterprise continuity and innovation.
Responsibilities
- Oversees the reliability and operational readiness of production systems, ensuring platforms can scale effectively and perform consistently under changing demand.
- Implements and continuously improves automated testing (e.g., QA, regression, performance, and reliability testing) within CI/CD pipelines to proactively identify defects, reduce production risk, and ensure consistent, high-quality releases.
- Develops and implements reliability goals, service level objectives, and performance expectations for critical digital products and platforms.
- Establishes and strengthens monitoring, logging, alerting, and observability practices to provide clear visibility into service health and performance.
- Owns the operational framework for incident management, change management, escalation, service restoration, root cause analysis, and post-incident review.
- Partners with digital teams to embed reliability, automation, and operational excellence across the software lifecycle and platforms.
Qualifications
- Bachelor’s degree or equivalent experience in Computer Science, Information Technology, Software Engineering, or a related field; advanced degree preferred.
- Minimum 8 years of progressive experience in Site Reliability Engineering, DevOps, cloud operations, infrastructure engineering, or a related field, including leadership of enterprise-scale reliability and operations functions.
- Demonstrated experience developing and executing SRE strategies, operational frameworks, and reliability programs that improve system availability, scalability, performance, and organizational resiliency.
- 5+ years in cloud-native architectures, distributed systems, infrastructure automation, CI/CD pipelines, and modern software delivery practices supporting mission-critical platforms, with experience establishing and governing Secure Software Development Lifecycle (SSDLC) practices that embed security, compliance, reliability, and risk management throughout the software delivery lifecycle.
AARP will not sponsor an employment visa for this position at this time.
Additional Requirements
- Regular and reliable job attendance
- Effective verbal and written communication skills
- Exhibit respect and understanding of others to maintain professional relationships
- Independent judgement in evaluation options to make sound decisions
- In office/open office environment with the ability to work effectively surrounded by moderate noise
Hybrid Work Environment
AARP observes Mondays and Fridays as remote workdays, except for essential functions. Remote work can only be done within the United States and its territories.
Compensation and Benefits
AARP offers a competitive compensation and benefits package including a 401(k); 100% company-funded pension plan; health, dental, and vision plans; life insurance; paid time off to include company and individual holidays, vacation, sick, caregiving, and parental leave; performance-based and peer-based recognition and tuition reimbursement.
Equal Employment Opportunity
AARP is an equal opportunity employer committed to hiring a diverse workforce and sustaining an inclusive culture. AARP does not discriminate on the basis of race, ethnicity, religion, sex, color, national origin, age, sexual orientation, gender identity or expression, mental or physical disability, genetic information, veteran status, or on any other basis prohibited by applicable law.
Similar roles
-
Site Reliability Engineer, Global Banking & Markets, Frontline Production Engineering
Goldman Sachs New York, New York, United States · $130K–$250K/yr
-
Site Reliability Engineer/ Expert/ Specialist (DevOps)
SITA Switzerland Sarl Amman, Amman, Jordan
-
Senior Site Reliability Engineer (Performance and Scalability)
Digital Zone Poland
-
Senior Site Reliability Engineer
TeamViewer Germany GmbH Austin, Texas, United States
-
Senior Site Reliability Engineer
Duplo Lagos, Lagos State, Nigeria
-
Sr Staff Site Reliability Engineer
Palo Alto Networks Sofia, Sofia-City, Bulgaria