Jobgether

Site Reliability Engineer in Network Infrastructure

Jobgether · Spain

Internet Marketplace Platforms · 11-50 employees

13 h ago
Remote Senior (5-10 yrs) Full-time Spain
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will be responsible for improving the reliability, performance, and operational maturity of large-scale network systems through automation and engineering excellence. This includes defining reliability objectives, managing incident responses, and building robust observability systems to support cloud and AI platforms.

What they look for

Site Reliability Engineering Network Infrastructure Linux Go Python Automation CI/CD Infrastructure as Code Observability Troubleshooting Networking fundamentals Container platforms Load balancers eBPF Telemetry System operability

Requirements

The ideal candidate possesses strong experience in site reliability engineering, infrastructure operations, and network systems with proficiency in Linux and automation languages like Go or Python. You must have a solid understanding of networking fundamentals and experience operating high-availability systems in complex production environments.

Benefits

Competitive compensation package Career growth and continuous learning opportunities Flexible working environment Impactful cloud and AI infrastructure projects Collaborative international engineering culture Exposure to cutting-edge technologies

Full description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer in Network Infrastructure based in Spain.

This role offers the opportunity to strengthen the reliability and scalability of critical network infrastructure supporting advanced cloud and AI platforms. As a Site Reliability Engineer, you will help design, automate, and operate the systems that enable high-performance digital services at scale. You will focus on building resilient environments, improving observability, and creating safer operational processes across complex network architectures. Working closely with network and platform engineering teams, you will transform operational challenges into long-term solutions through automation and engineering excellence. This position combines software development, infrastructure expertise, and reliability engineering in a fast-moving international environment. You will have significant ownership and the chance to influence the future of large-scale cloud infrastructure.

\n

Accountabilities: As a Site Reliability Engineer in Network Infrastructure, you will be responsible for improving the reliability, performance, and operational maturity of large-scale network systems. You will combine engineering practices with operational expertise to ensure critical infrastructure remains secure, scalable, and efficient.

  • Define and manage reliability objectives for network services, including SLIs, SLOs, availability targets, and error budgets where applicable.
  • Drive reliability improvements across network infrastructure, including services, site readiness, inter-site connectivity, and operational processes.
  • Own incident response activities, lead technical investigations, conduct postmortems, and implement long-term solutions to prevent recurring issues.
  • Build and improve observability systems through meaningful metrics, logs, traces, alerting, and faster troubleshooting workflows.
  • Design safer infrastructure change processes through automation, CI/CD workflows, testing environments, staged deployments, rollback strategies, and auditability.
  • Collaborate closely with network engineers and platform teams to improve system operability and integrate reliability practices into technical designs.
  • Automate operational workflows and continuously improve infrastructure management processes.
  • Contribute to building scalable and resilient network environments that support growing cloud workloads.

Requirements:

The ideal candidate has strong experience in site reliability engineering, infrastructure operations, and network systems. You are comfortable working with complex production environments, debugging challenging technical issues, and developing automation to improve reliability and efficiency.

  • Strong knowledge of production Linux environments and a structured approach to troubleshooting complex systems.
  • Solid understanding of networking fundamentals, including control plane and data plane concepts, latency, packet loss, and failure domains.
  • Experience operating high-availability systems and continuously improving their reliability over time.
  • Ability to write and maintain automation and infrastructure software, with experience in Go preferred and Python welcomed.
  • Experience with modern infrastructure tooling, including Infrastructure as Code, CI/CD systems, and container platforms.
  • Strong engineering mindset with the ability to balance operational stability, automation, and continuous improvement.
  • Experience with high-throughput traffic systems such as load balancers, tunneling, decapsulation, NAT64, or similar technologies is a plus.
  • Knowledge of low-level networking performance optimization, including eBPF/XDP, DPDK, perf/ftrace, or Linux kernel networking internals is advantageous.
  • Experience building safe network delivery pipelines, including testing environments, staged rollouts, automated validation, and drift detection is beneficial.
  • Familiarity with large-scale network observability and telemetry solutions is considered a plus.

Benefits:

  • Competitive compensation package.
  • Career growth and continuous learning opportunities.
  • Flexible working environment with a strong focus on ownership and autonomy.
  • Opportunity to work on impactful cloud and AI infrastructure projects.
  • Collaborative culture with experienced international engineering teams.
  • Exposure to cutting-edge technologies in networking, reliability engineering, and cloud platforms.
  • Opportunity to contribute to systems shaping the future of AI infrastructure.
  • Fast-paced environment focused on innovation, meaningful impact, trust, and professional growth.

\nHow Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1