DRIVENETS

Senior DevOps Engineer — Cybersecurity business unit

DRIVENETS Tel Aviv, Tel-Aviv District, Israel

Software Development · 501-1,000 employees

2 h ago
devops Senior (5-10 yrs) Full-time Israel
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The engineer will own the platform's Helm charts and manage the self-hosted operator stack across various Kubernetes distributions. They are also responsible for building zero-downtime release pipelines, managing observability, and providing on-call support for production environments.

What they look for

Kubernetes Helm GitOps ArgoCD Infrastructure as Code Prometheus Grafana OpenTelemetry Kafka PostgreSQL Elasticsearch KEDA Temporal DevOps Platform Engineering GPU

Requirements

Candidates must have at least 5 years of DevOps or platform engineering experience with deep hands-on expertise in Kubernetes and Helm. Proficiency in GitOps, stateful workload management, and observability fundamentals is required, along with the ability to operate in air-gapped environments.

Full description

Senior DevOps Engineer - Cybersecurity

Hybrid | Israel

About the Company

DriveNets is a leader in large-scale networking solutions for AI infrastructure and service providers. The company's disaggregated networking architecture transforms the economics of large-scale infrastructures while maximizing performance, utilization, and operational efficiency. Its high-performance AI fabric maximizes GPU utilization and accelerates deployments by optimizing the AI stack end-to-end, resulting in higher tokens-per-second and lower cost-per-token. DriveNets' solutions power production networks for global tier-1 operators like AT&T and Comcast, and scale multi-vendor AI infrastructures at foundation model labs, NeoClouds, and enterprises.

DriveNets' Cybersecurity business unit is building an AI-powered security platform for the world's largest service provider networks. The platform combines a cloud-native microservices backend with self-hosted LLM inference, delivered into customers' own Kubernetes environments — often air-gapped, with no data leaving their premises.

Responsibilities

  • Own the platform's Helm charts and make them portable across any Kubernetes distribution, including OpenShift (non-root, restricted SCCs, NetworkPolicies).
  • Operate the self-hosted operator stack: CloudNativePG, Strimzi, ECK, MinIO, Redis, and Temporal.
  • Run the Nx monorepo pipelines and drive the ArgoCD GitOps rollout.
  • Build zero-downtime releases using blue/green and canary deployments, expand/contract migrations, and automated rollback.
  • Manage autoscaling with KEDA/HPA, backup/restore for stateful services, and load-validation of stamp sizing tiers.
  • Provide on-call support for dev/staging environments and customer stamps.
  • Build and maintain observability using Grafana Alloy, OpenTelemetry, and Prometheus/Grafana — dashboards and alerts across services, Kafka, databases, and GPUs.
  • Improve developer experience through a shared dev cluster with local-debug traffic interception (Telepresence / mirrord), targeting sub-15-minute onboarding.

Requirements

Technical Skills

  • 5+ years of DevOps / platform engineering experience in a product company.
  • Deep hands-on experience with Kubernetes and Helm, including shipping software onto clusters you don't control.
  • Production experience with GitOps (ArgoCD/Flux), infrastructure as code, and zero-downtime release engineering.
  • Experience running stateful workloads on Kubernetes via operators (Kafka, PostgreSQL, Elasticsearch).
  • Solid grasp of observability fundamentals: Prometheus, Grafana, OpenTelemetry.
  • Comfortable operating in air-gapped environments without managed cloud services.

Soft Skills

  • Strong cross-functional collaboration — works daily with AI/ML Ops, backend, research, and solutions teams on customer onboarding.
  • Comfortable owning production reliability and on-call responsibilities.

Nice to Have / Advantage

  • Experience with OpenShift, KEDA, and Temporal.
  • Hands-on AI inference infrastructure experience: deploying and operating LLM inference on Kubernetes with GPUs (vLLM, TGI, Triton, or similar).
  • GPU scheduling and node readiness, model packaging for offline delivery.
  • Experience with LLM gateways (LiteLLM) and LLM observability (Langfuse).

If your experience is close but doesn't fulfil all requirements, please submit your application. DriveNets is on a mission to build a special company comprised of individuals with different backgrounds, perspectives, and experiences.

DriveNets is an equal opportunity employer. We do not discriminate based on upon race, religion, national origin, sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with disability, or other applicable legally protected characteristics.

More About DriveNets

Based in Israel with extended teams located in the US, Japan, and Romania, DriveNets operations cover more than twelve countries globally. Powering production networks for global tier-1 operators, DriveNets is a leader in large-scale networking solutions for AI infrastructure and service providers. Visit our website to learn more:

https://drivenets.com/company/

Similar roles