Sovereign Engineering Platform SRE (m/f/d)
T-Systems Iberia Granada, Andalusia, Spain
IT Services and IT Consulting · 1,001-5,000 employees
About the role
Build and operate Kubernetes environments to host AI engineering tools, model gateways, and workflow services. Implement GitOps and Infrastructure as Code patterns while ensuring observability and security for engineering workloads.
What they look for
Requirements
Requires 5+ years of experience in SRE, platform engineering, or DevOps with strong Kubernetes and Linux expertise. Candidates must have hands-on experience with Terraform, Ansible, and secure configuration management in high-security environments.
Benefits
Full description
Company Description
T‑Systems is part of the Deutsche Telekom Group, with around 30.000 employees worldwide. We create technology with purpose to generate a positive impact on society. We are looking for curious talent, eager to learn, take on challenges, and contribute ideas that transform our customers’ experience.
We trust people: we offer autonomy, continuous support, and a collaborative environment where you can grow without limits. We are one global team, guided by respect, integrity, and a passion for doing better every day.
Job Description
Key responsibilities
- Build and operate Kubernetes environments that host AI engineering tools, internal model gateways, retrieval components, workflow services, CI/CD runners, and documentation services.
- Implement GitOps and Infrastructure as Code patterns for reproducible provisioning, configuration, policy enforcement, platform upgrades, and disaster recovery readiness.
- Manage private registries, package mirrors, secrets, identity integration, network segmentation, storage classes, backup routines, and controlled connectivity models.
- Provide observability for engineering workloads, including metrics, logs, traces, GPU and CPU utilization, service health, cost signals, and operational runbooks.
- Work with software, security, and architecture teams to ensure the platform supports AI-assisted SDLC workflows without creating uncontrolled data exposure or audit gaps.
Examples of market tools, models, and platform components expected
- Platform tooling such as Kubernetes, Helm, Terraform, Ansible, ArgoCD, Crossplane, GitLab runners, Jenkins agents, private registries, and internal package mirrors.
- AI platform components such as vLLM, Ollama, OpenAI-compatible gateways, Qdrant or similar vector stores, Open WebUI, Continue-compatible endpoints, and workflow services.
- Observability and operations stacks such as Prometheus, Grafana, Loki, OpenTelemetry, ELK/OpenSearch, Alertmanager, SRE runbooks, and incident management tooling.
- Security and governance components such as Vault, Keycloak, network policies, RBAC, admission controls, image scanning, SBOM tooling, and audit logging.
- Infrastructure awareness covering GPU-backed nodes, CPU-only fallback, storage performance, network isolation, proxy patterns, on-premise environments, and dedicated landing zones.
Qualifications
- 5+ years in SRE, platform engineering, DevOps, cloud infrastructure, or operations roles with strong Kubernetes and Linux expertise.
- Proven experience building and operating production-grade engineering platforms with GitOps, Infrastructure as Code, observability, and operational runbooks.
- Hands-on skills in Terraform, Ansible, Helm, Python or shell scripting, CI/CD runners, private registries, and secure configuration management.
- Good understanding of networking, storage, secrets, access control, monitoring, backup, disaster recovery, and operational hardening in high-security environments.
- Comfortable supporting AI-enabled engineering workloads in sovereignty-driven contexts where isolation, controlled data handling, reliability, and auditability are mandatory.
Additional Information
What do we offer you?
Work environment & flexibility
- International, dynamic and collaborative environment.
- T-Social: social initiatives (sports, community, health, ...).
- Hybrid work model (remote/on-site).
- Flexible working hours.
Growth & development
- Customized training: access to Coursera to learn whatever you want, whenever you want.
- Weekly language classes (English & German).
- International Mentoring Sessions & Experience Days.
Compensation & benefits
- Flexible compensation plan (health insurance, meal vouchers, childcare, transport).
- Telemedicine.
- Life and accident insurance.
- Social fund.
Wellbeing & time off
- 26+ working days of vacation per year.
- Free access to specialist services (medical, legal, wellness).
- 100% salary coverage during medical leave.
And many more advantages of being part of T-Systems!
If you are looking for a new challenge, do not hesitate to send us your CV! Please send CV in English. Join our team!
T-Systems Iberia will only process the CVs of candidates who meet the requirements specified for each offer.
Similar roles
-
Senior DevOps & Site Reliability Engineer (GCP)
Burjline Builders Sofia, Sofia-City, Bulgaria
- Platform Engineer (SRE - India)- III
-
Senior Site Reliability Engineer (Cloud Platform)
Salve.Inno Consulting United Kingdom
-
Engineer Lead, Site Reliability
Zensar Pune, Maharashtra, India
-
Azure Site Reliability Engineer (SRE) - SaaS Operations
Zensar Bangalore South, Karnataka, India
-
Site Reliability Engineer
OSTTRA Gurugram, Haryana, India