Site Reliability Engineer - rednote
rednote Palo Alto, California, United States · $200K–$400K/yr
Technology, Information and Internet · 11-50 employees
About the role
You will design and implement international infrastructure architecture, focusing on disaster recovery, high availability, and multi-region deployment. Additionally, you will manage overseas technical platforms and lead incident response efforts to ensure system stability.
What they look for
Requirements
Candidates must have extensive experience in SRE, large-scale system stability, and cross-region architecture design. Proficiency in Linux, cloud-native infrastructure, and programming languages like Python, Go, or Java is required.
Full description
What you'll do
1、International Architecture & Disaster Recovery — Participate in the design and implementation of Rednote's international infrastructure architecture. Build and evolve cross-region architecture, disaster recovery, and high-availability capability development. Drive critical services toward multi-region deployment, failover, and fault isolation to improve overall stability of overseas operations. 2、Overseas Infrastructure Platform Development & Operations — Own the deployment, operations, and continuous optimization of core internal technical platforms (release systems, monitoring & alerting, configuration services,service management, traffic scheduling, etc.) in overseas regions. Ensure consistency and availability across overseas and domestic platform environments. 3、Reliability Engineering & Incident Response — Build and continuously improve the reliability framework for overseas business, including observability capabilities, incident response, root cause analysis, and post-mortem mechanisms. Lead cross-functional coordination during major incidents to restore services quickly and drive (long-term)systemic improvements. 4、International Technical Solution Delivery — Develop a deep understanding of overseas business requirements and architecture characteristics. Drive infrastructure capabilities to fit overseas scenarios, including multi-region architecture design, network and data architecture optimization, and adaptation of foundational services. 5、Cross-functional Collaboration & Best Practice Development — Work closely with domestic infrastructure, product engineering teams, and platform teams to align overseas technical standards with domestic architecture standards. Consolidate and promote overseas stability best practices across the organization.
Qualifications
1、Reliability Engineering & SRE Experience — Familiar with large-scale internet system stability frameworks; experienced in high-availability architecture design, fault governance, capacity planning, and incident response. Experience in SRE, platform engineering, or infrastructure engineering is preferred. 2、International Architecture Experience — Familiar with cross-region architecture design and disaster recovery systems (multi-region deployment, traffic scheduling, data sync, failover, etc.). Experience with overseas business architecture or international infrastructure development preferred. 3、Core Technical Skills — Proficient in Linux systems, networking, and common middleware (MySQL, Redis, Kafka, etc.); good understanding of cloud-native infrastructure (Kubernetes, Service Mesh, etc.) and observability stacks (monitoring, logging, tracing). 4、Development & Automation Skills — Proficient in at least one of Python, Go, or Java; experience building automation ops platforms, stability tooling, or infrastructure systems. 5、Problem Solving& Collaboration — Strong troubleshooting and root cause analysis skills in complex system environments; excellent communication and teamwork. 6、Language — Fluent in both English and Chinese (spoken and written)
Similar roles
-
Senior Site Reliability Engineer (SRE)
Sequoia Connect Ciudad de México, Mexico
-
Summer 2027 SRE - Observability Engineering Internship
Tradeweb Jersey City, New Jersey, United States · $73K–$156K/yr
-
Senior Site Reliability Engineering- CTJ- Secret (Cleared Environments)
Microsoft Redmond, Washington, United States · $120K–$261K/yr
-
Senior Site Reliability Engineer
Aptean Alpharetta, Georgia, United States
-
Site Reliability Engineer
Cosm El Segundo, California, United States · $110K–$145K/yr
-
Infrastructure Site Reliability Engineer (w/ active Secret)
Critical Solutions Norfolk, Virginia, United States · $87K–$112K/yr