Jobgether

Site Reliability / Gitops Engineer

Jobgether Germany

Internet Marketplace Platforms · 11-50 employees

16 h ago
Remote sre Senior (5-10 yrs) Full-time Germany
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will apply Infrastructure as Code and GitOps practices to automate and improve the reliability, scalability, and performance of cloud and container infrastructure. Additionally, you will troubleshoot complex systems, maintain observability solutions, and collaborate with global engineering teams to resolve operational issues.

What they look for

Site Reliability Engineering GitOps Infrastructure as Code Python Linux Cloud Computing CI/CD Prometheus Grafana Elasticsearch Container Infrastructure Automation Observability System Troubleshooting Networking Agile Methodologies

Requirements

The role requires deep experience in IT operations, CI/CD, and modern software engineering practices, along with significant Python development skills. Candidates must have a strong background in Linux administration, cloud architectures, and effective communication skills for a distributed team environment.

Benefits

Remote work flexibility International collaboration Professional development time Mentoring opportunities Company events

Full description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability / Gitops Engineer based in Germany.

Join a global Information Systems team responsible for operating and evolving critical production services at significant scale. In this role, you’ll use automation, Infrastructure as Code, and GitOps practices to make cloud operations more reliable, consistent, and efficient. You’ll work across private and public cloud environments, strengthening infrastructure resilience, scalability, observability, and performance. Your expertise will also influence the evolution of open-source infrastructure technologies through hands-on feedback, bug reporting, and collaboration. You’ll troubleshoot complex systems, improve operational processes, and help eliminate repetitive manual work through thoughtful automation. Working with a distributed team of experienced SREs, you’ll have opportunities to share knowledge, mentor colleagues, and contribute to major engineering initiatives. This is an ideal opportunity for an automation-first technologist who is passionate about Linux, open source, and building robust systems at scale.

\n

Accountabilities:

  • Apply Infrastructure as Code expertise to continuously improve automation practices, processes, and operational consistency.
  • Automate software operations across private and public clouds while accounting for the complexities of distributed systems.
  • Develop new capabilities and improve the resilience, scalability, and reliability of cloud and container infrastructure.
  • Maintain operational responsibility for core services, networks, and infrastructure, ensuring reliable day-to-day performance.
  • Troubleshoot complex systems, perform capacity planning and performance investigations, and develop strong operational expertise.
  • Set up, maintain, and use observability and monitoring solutions such as Prometheus, Grafana, and Elasticsearch.
  • Design and maintain monitoring and alerting for critical systems and services.
  • Collaborate with development teams on service architecture, documentation, playbooks, policies, and operational procedures.
  • Work closely with globally distributed engineering, operations, and support teams to resolve issues and improve services.
  • Dedicate focused development time to larger engineering projects and the automation of repetitive manual processes.
  • Share technical knowledge and best practices through design sessions, mentoring, and collaborative problem-solving.
  • Take final responsibility for resolving time-critical operational escalations.

Requirements:

  • Deep experience defining IT operations through code, using version control, peer review, and CI/CD to deploy application and infrastructure changes.
  • Strong modern software engineering practices, including peer review, unit testing, source control management, CI/CD, and Agile methodologies.
  • Significant Python development experience, including work on large or complex projects.
  • Practical knowledge of Linux networking, routing, firewalls, and related infrastructure concepts.
  • Familiarity with Linux storage technologies, ranging from Ceph to database systems.
  • Hands-on experience administering enterprise Linux servers.
  • Extensive understanding of cloud computing concepts, architectures, and technologies.
  • Bachelor’s degree or higher, preferably in computer science, software engineering, or a related technical discipline.
  • Strong English communication skills across written and spoken channels, including email, chat, video calls, voice calls, and in-person collaboration.
  • Strong troubleshooting abilities, with the curiosity and persistence to investigate issues from the Linux kernel through to the web layer.
  • Ability to collaborate effectively while knowing when to seek input from teammates and subject-matter experts.
  • Adaptability, willingness to learn quickly, and comfort working in fast-changing technical environments.
  • Ability to thrive within globally distributed teams and collaborate across different locations and time zones.
  • Strong interest in open-source technologies, particularly Ubuntu or Debian.

Benefits:

  • Opportunity to work on production infrastructure supporting large-scale global services.
  • Exposure to private and public cloud environments, Infrastructure as Code, GitOps, CI/CD, observability, and open-source technologies.
  • Dedicated development time for impactful automation and larger engineering projects.
  • Collaboration with a highly experienced, globally distributed SRE and engineering community.
  • Opportunities for mentoring, knowledge sharing, and cross-functional technical collaboration.
  • Remote work flexibility, with the role available across time zones.
  • Opportunities to meet colleagues in person 2–4 times per year at internal events, typically lasting 1–2 weeks.
  • International exposure through collaboration with distributed teams and participation in global company events.
  • Compensation and benefits are determined according to the role, location, experience, and applicable company policies.

\nHow Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Similar roles