Principal Production Engineer, Database Infrastructure
GitHub, Inc. United States
Software Development · 501-1,000 employees
About the role
The Principal Production Engineer will lead design reviews and shape data modeling to ensure high performance and reliability across GitHub's database platforms. They will also develop SDKs, tooling, and disaster recovery plans while participating in on-call rotations to maintain system health.
What they look for
Requirements
Candidates must have extensive experience in software engineering and operating large-scale distributed systems, with a minimum of 5 to 11 years depending on their educational background. Proficiency in languages such as Go, Python, Ruby, or Rust and experience with cloud-based stateful services are required.
Benefits
Full description
About GitHub
GitHub is the world’s leading platform for agentic software development — powered by Copilot to build, scale, and deliver secure software. Over 180 million developers, including more than 90% of the Fortune 100 companies, use GitHub to collaborate, and more than 77,000 organisations have adopted GitHub Copilot.
Locations
In this role you can work from Remote, United States
Overview
GitHub is looking for a Principal Production Engineer to help scale our data platform to millions of developers. We are software engineers who specialize in reliability, working with technical partners, leading design reviews, writing SDKs and tooling to build against, and shaping how our data platform is used to prevent scaling problems before they reach production.
Responsibilities
- Partner with product and feature teams by leading design reviews and shaping how they model, access, and scale their data so the applications they build are performant, available, and operable at GitHub's scale
- Design and ship SDKs, client libraries, and the tooling applications are built on
- Set reliability strategy for GitHub's database platforms across multiple systems, defining SLOs and operational standards
- Write technical documentation and advocate for the health and quality of the systems the team builds
- Participate in an on-call rotation and respond to incidents as needed
- Develop and design plans for disaster recovery, load shedding, and regional failover
The team is highly distributed across geographies and time zones, and you will thrive in an environment of remote work and asynchronous communication.
Qualifications
Required Qualifications:
- 11+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
- OR Associate's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 10+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
- OR Bachelor's Degree in Computer Science or related field AND 9+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
- OR Master's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 7+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python.
- OR Doctorate in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 5+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python.
- OR equivalent experience.
- 5+ years experience operating large-scale distributed systems in production, including participation in an on-call rotation.
Preferred Qualifications:
- Excitement about building, operating, and maintaining resilient, scalable systems that impact a global community of users with the ability to break down complex systems into manageable components.
- A track record of partnering with product and feature teams and materially changing the reliability and scalability of what they ship.
- Ability to influence engineering decisions and proactively engage in system design conversations.
- Experience running stateful services on managed cloud data stores, specifically Azure Cosmos DB and Azure SQL Database, or equivalents such as DynamoDB, Aurora, or Cloud Spanner with a focus on partition and index design, consistency and isolation tradeoffs, throughput provisioning, and hot-partition diagnosis.
- Experience diagnosing and resolving application-level scalability problems: N+1 query patterns, hot partitions, unbounded fan-out, cache stampedes, and data access patterns that don’t scale.
- A track record of building internal platforms, SDKs, or developer tools adopted across an engineering organization, written in production-grade Go, Python, Ruby, or Rust
- Deep familiarity with the failure modes of large-scale systems, both in the application and in the platform beneath it e.g. cascading failures, retry storms, thundering herds, partial outages, throttling and quota limits, control plane outages, noisy neighbors, and the patterns that mitigate them.
- Experience leading large-scale cloud migrations of live, high-traffic services e.g. dual-write and backfill strategies, traffic shifting, correctness verification, and rollback under load.
- Effective communication skills and willingness to pair on problems, brainstorm in public, and enthusiastically engage with your teammates in group problem solving.
Compensation Range
The base salary range for this job is USD $160,200.00 - USD $425,000.00 /Yr.
These pay ranges are intended to cover roles based across the United States. An individual's base pay depends on various factors including geographical location and review of experience, knowledge, skills, abilities of the applicant. At GitHub certain roles are eligible for benefits and additional rewards, including annual bonus and stock. These rewards are allocated based on individual impact in role. In addition, certain roles also have the opportunity to earn sales incentives based on revenue or utilization, depending on the terms of the plan and the employee's role.
This position will be open for a minimum of 3 days, with applications accepted on an ongoing basis until the position is filled. GitHub values
- Customer-obsessed
- Ship to learn
- Growth mindset
- Own the outcome
- Better together
- Diverse and inclusive
Manager fundamentals
- Model
- Coach
- Care
Leadership principles
- Create clarity
- Generate energy
- Deliver success
Who We Are
GitHub is the world’s leading AI-powered developer platform with 150 million developers and counting. We’re also home to the biggest open-source community on earth (and 99% of the world’s software has open-source code in its DNA). Many of the apps and programs you use every day are built on GitHub. Our teams are dreamers, doers, and pioneers, leading the way in AI, driving humanitarian efforts around the globe, and even sending open source to Mars (and beyond!). At GitHub, our goal is to create the space you need to do your best work. We’re remote-first and offer competitive pay, generous learning and growth opportunities, and excellent benefits to support you, wherever you are—because we know that people flourish when they can work on their own terms. Join us, and let’s change the world, together.
EEO Statement
GitHub is made up of people from a wide variety of backgrounds and lifestyles. We embrace diversity and invite applications from people of all walks of life. We don't discriminate against employees or applicants based on gender identity or expression, sexual orientation, race, religion, age, national origin, citizenship, disability, pregnancy status, veteran status, or any other differences. Also, if you have a disability, please let us know if there's any way we can make the interview process better for you; we're happy to accommodate!