R

DevOps Platform R&D Engineer – rednote

rednote Palo Alto, California, United States · $200K–$400K/yr

Technology, Information and Internet · 11-50 employees

Aug 14
devops Mid (2-5 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will drive the rollout, integration, and operational support of DevOps platforms across international regions while ensuring high availability. Additionally, you will participate in production operations, incident response, and the maintenance of operational runbooks to improve service reliability.

What they look for

DevOps SRE Kubernetes Terraform Linux Go Java Python Shell CI/CD Infrastructure-as-Code Prometheus Grafana ELK OpenTelemetry Cloud Computing

Requirements

Candidates must hold at least a bachelor's degree in Computer Science or a related field and possess hands-on experience in DevOps, SRE, or infrastructure engineering. Proficiency in Linux, cloud-native technologies like Kubernetes, and at least one programming language such as Go, Java, or Python is required.

Full description

What you'll do

1、Drive the rollout, integration, localization, and operational support of DevOps and high-availability platforms across international regions, ensuring platform capabilities are delivered reliably and on schedule. 2、Evaluate infrastructure differences across international public cloud, private cloud, and data center environments, including compute, networking, storage, Kubernetes, security, and compliance requirements. Design and implement practical adaptation plans. 3、Support the international adoption of internal DevOps platforms, including service release systems, change management platforms, and large-scale operations tooling. Improve CI/CD pipelines, infrastructure automation, and release operations workflows. 4、Implement high-availability capabilities for international environments, including incident response playbooks, failover mechanisms, disaster recovery validation, and chaos engineering practices. Track remediation actions and continuously improve service reliability. 5、Participate in production operations and incident response for international environments. Quickly identify and troubleshoot issues related to releases, infrastructure, and platform integration, and drive root-cause analysis and closed-loop improvements. 6、Maintain deployment guides, operational runbooks, incident response playbooks, and standard operating procedures to improve delivery quality and operational efficiency across international regions. 7、Work closely with central platform R&D teams, as well as international infrastructure, networking, security, and business teams. Manage project progress, dependencies, and risks, and provide feedback to improve platform capabilities for international scenarios.

Qualifications

1、Bachelor’s degree or above in Computer Science, Software Engineering, or a related field. 2、Hands-on experience in DevOps, SRE, cloud platforms, or infrastructure engineering, with strong delivery and troubleshooting capabilities. 3、Solid understanding of Linux, computer networking, and common infrastructure components. Able to independently perform deployment, configuration, and issue diagnosis. 4、Familiar with cloud-native and Infrastructure-as-Code technologies, such as Kubernetes, Helm, and Terraform. Experience with KubeVela is a plus. 5、Proficient in at least one programming or scripting language, such as Go, Java, Python, or Shell. Able to build automation tools and read/debug backend service code. 6、Familiar with at least one major public cloud or enterprise private cloud environment. Understand concepts such as multi-region deployment, network connectivity, access control, and security isolation. 7、Familiar with monitoring, logging, and alerting systems. Experience with Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, or similar tools is preferred. 8、Fluent in both English and Chinese, with the ability to communicate effectively in a cross-region, international team environment. 9、Strong ownership, execution, and communication skills. Able to translate central platform designs into stable, maintainable implementations in local international environments.

Similar roles