Lio

Site Reliability Engineer / SRE (all genders)

Lio Munich, Bavaria, Germany

Software Development · 51-200 employees

Aug 06
sre Mid (2-5 yrs) Full-time Germany
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will build and operate reliable cloud infrastructure while improving the performance and scalability of backend services. Additionally, you will define SLOs, manage incident responses, and automate operational processes to support enterprise-scale AI workloads.

What they look for

Python Cloud infrastructure Site reliability engineering Monitoring Observability Incident management Distributed systems Asynchronous processing Scalable architectures Database optimization MongoDB CI/CD pipelines GitHub Actions Automation Root cause analysis

Requirements

Candidates should have experience operating production workloads on major cloud platforms and possess strong Python skills. A solid understanding of distributed systems, monitoring, and CI/CD pipelines is required.

Benefits

Competitive compensation Meaningful equity

Full description

Build and scale the infrastructure behind enterprise AI.

At Lio, we're building the AI workforce for procurement. As a Site Reliability Engineer, you'll work directly with our CTO to build the infrastructure that powers one of Europe's fastest-growing enterprise AI startups. From scaling our platform across regions to optimizing databases for agentic AI, you'll help shape the foundation that enables Lio to serve global enterprise customers with exceptional performance and reliability.

What you'll do

  • Work directly with our CTO on the architecture and scaling of Lio's infrastructure
  • Design and operate highly reliable cloud infrastructure for production AI workloads
  • Scale our platform globally with multi-region deployments and high-availability architectures
  • Optimize databases for agentic AI workloads, including Vector Search, Hybrid RAG, MCP, and high-performance query execution
  • Improve backend performance, latency, throughput, and resource efficiency
  • Build observability, SLOs, alerting, and incident response processes
  • Automate deployments, infrastructure, and developer workflows
  • Partner closely with product engineering teams to build systems that scale with rapid growth

What we're looking for

  • Experience operating production workloads on a major cloud platform
  • Strong Python skills and experience optimizing backend services
  • Solid understanding of distributed systems, asynchronous processing, and scalable architectures
  • Experience with monitoring, observability, and incident management
  • Experience optimizing databases at scale (MongoDB is a plus)
  • Familiarity with CI/CD pipelines (GitHub Actions preferred)
  • Passion for automation, infrastructure, and solving complex scaling challenges

Why Lio?

Work directly with our CTO on company-defining infrastructure decisions. Build the foundation for one of Europe's fastest-growing enterprise AI companies. Solve challenging problems around global scaling, reliability, and agentic AI infrastructure. Competitive compensation, meaningful equity, and exceptional teammates. 100% on-site in our Munich office, where we build together every day

Any questions about the role? Feel free to reach out to Cansel.

Similar roles