Senior DevOps Engineer, AI Agent Platforms
NVIDIA Yokneam Ilit, North District, Israel
Computer Hardware Manufacturing · 10,001+ employees
About the role
Build and operate AI agents and the underlying platforms to solve engineering problems across the networking stack. Drive responsible AI adoption by establishing best practices, guardrails, and observability pipelines for production workloads.
What they look for
Requirements
Requires a BSc in Computer Science or equivalent and 3+ years of experience in DevOps, platform engineering, or SRE. Candidates must have strong skills in Linux, Kubernetes, CI/CD, and proficiency in Python, Bash, and a compiled language like Go or Rust.
Full description
NVIDIA is spearheading the AI revolution and the creation of powerful accelerated compute platforms for global utilization. Our mission is to put AI at the center of engineering throughout NVIDIA's networking architecture team. We build AI agents that solve real engineering problems, the platforms that make them reliable, and we lead the team's AI adoption, making sure it happens fast and with strong engineering practices.
Powered by the AI revolution we are changing how hardware architecture work happens! We strive to define and implement this new agentic architecture, and we are seeking a DevOps software engineer for our agentic AI team. If you are excited to work at the frontier of AI agents and want your work to multiply the output of hundreds of engineers, we want to hear from you.
What you'll be doing:
- Build AI agents that solve real engineering problems across the networking stack, with a focus on the infrastructure and services that bring them into reliable daily use.
- Design, build, and operate the platforms that make those agents reliable: MCP tools, service APIs, cloud and on-prem services, and the systems behind them.
- Work broadly across networking architecture teams to identify high-impact AI use cases and turn them into working solutions.
- Drive rapid and responsible AI adoption across the organization: establish best practices, guardrails, and evaluation methods so teams embrace agentic workflows with confidence.
- Build evaluation and observability pipelines that measure AI-agent behavior and platform reliability on real production workloads.
What we need to see:
- BSc in Computer Science or a related field, or equivalent experience.
- 3+ years of relevant practical experience in DevOps, platform engineering, or SRE roles ideally including work on a large-scale software product.
- Deep hands-on experience with Linux, cloud and on-prem environments, Kubernetes and container runtimes (Docker/containerd), and infrastructure as code.
- Proven experience designing, building, and maintaining CI/CD pipelines for production services.
- Strong scripting and automation skills in Python/Bash and one compiled language (Go/Rust).
- Experience with observability, monitoring, and security practices.
- Advanced, demonstrable use of AI and agentic workflows in your daily engineering work well beyond standard coding assistant usage.
- A drive for end-to-end ownership and strong communication skills.
Ways to Stand Out from the Crowd:
- Hands-on experience with LLM APIs and agent frameworks, including building tools/MCP integrations.
- Background in computer networking or HPC.
NVIDIA has some of the most forward-thinking and hardworking people in the world working for us. If you are a creative and autonomous engineer with a real passion for technology, we want to hear from you.
Similar roles
-
Senior DevOps Engineer | TS/SCI
Phoenix Operations Group Bexar County, Texas, United States
-
Senior DevOps Engineer (m/f/d)
PROBIS Munich, Bavaria, Germany
-
Senior DevOps Developer
Mappedin Waterloo, Ontario, Canada · CA$120K–CA$150K/yr
-
DevOps Engineer
Evolve Today Bucharest, Romania
-
Private Cloud DevOps Engineer
Vattenfall Oskarshamn, Kalmar County, Sweden
-
Platform Architect | Cloud, DevOps & Automation
knowmad mood Madrid, Community of Madrid, Spain