monday.com

DevOps Tech Lead (BigBrain)

monday.com · Tel Aviv, Tel Aviv, Israel

Software Development · 1,001-5,000 employees

14 h ago
Remote Senior (5-10 yrs) Full-time Israel
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will lead the technical direction and architecture for critical infrastructure domains while managing multi-region Kubernetes clusters and data pipelines. Additionally, you will build and operate AI infrastructure, drive automation with AIOps, and mentor other engineers to raise the technical bar.

What they look for

Kubernetes AWS Terraform Kafka Python TypeScript Go ArgoCD CI/CD GitOps Data Engineering AI Infrastructure System Architecture Observability Security Leadership

Requirements

Candidates must have 6-8+ years of experience in DevOps or infrastructure engineering with deep hands-on expertise in Kubernetes and cloud infrastructure. Strong proficiency in Infrastructure as Code, CI/CD pipelines, and a security-first mindset are essential for this high-ownership role.

Full description

About monday.com

monday.com is the AI work platform powering the most ambitious teams. 250,000+ customers across departments use us to bring people, workflows, and AI agents together on one flexible platform where AI doesn't just assist, it executes. We move fast, build things that matter, and foster an ownership-driven culture where you're empowered to shape how organizations work and outpace their competition.

About the team

We are looking for a Senior DevOps Engineer / Tech Lead to join the BigBrain group.

The team that builds and operates monday.com's Data Platform and leads the company's internal AI innovation.

We own the infrastructure behind billions of daily events - from streaming pipelines and data orchestration to the AI Gateway and ML inference platform that power monday.com's intelligent features. We manage some of the most sensitive data in the company, operate across multiple global regions, and are responsible for keeping it all secure and running at scale.

The BigBrain group consists of diverse teams, including Data Scientists, Full-Stack Engineers, Data Engineers, and BI Engineers. As a Senior DevOps Engineer / Tech Lead on this team, you'll own the technical direction for critical infrastructure domains, mentor other engineers, and partner closely with product, security, data, and platform teams to build infrastructure that's secure, scalable, and increasingly autonomous.

This is a high-ownership role with significant influence over the technical roadmap of our data and AI infrastructure.

This is an exciting time to join: we're building an AI Gateway to govern LLM usage across the company, scaling our ML inference platform, designing AIOps agents that automate infrastructure operations, and evolving our data platform with modern technologies - all while keeping the foundation rock-solid across multiple regions.

You'll join our DevOps team based in our headquarters in Tel Aviv, Israel.

Team podcasts - https://www.startupforstartup.com/110-on-the-operations-behind-the-client-facing-teams-growth/ https://pod.link/1595260676/episode/fa4d99bbf8e536e7a139a1ba679692f0 BigBrain: https://www.startupforstartup.com/on-the-bigbrain-that-makes-kpis-accessible-to-every-employee/ https://www.youtube.com/watch?v=x-m0ag0cty0 https://engineering.monday.com/ai-brain-ai-charged-tool-for-internal-usage/

About The Role

  • Lead technical direction — Own the architecture and technical roadmap for critical infrastructure domains, and make high-impact trade-offs across teams.
  • Own and scale platform infrastructure — Manage multi-region Kubernetes clusters (EKS), streaming pipelines (Kafka/MSK, Debezium CDC), and data orchestration (Airflow) that handles billions of daily events.
  • Build and operate AI infrastructure — Deploy and maintain the AI Gateway (governance, observability, guardrails for LLM usage), ML inference platform, and the tooling that enables AI adoption across the company.
  • Drive infrastructure automation with AIOps — Design and build autonomous agents and intelligent tooling (n8n, LangChain, Claude Code) that automate infrastructure operations, analyze workflows, and reduce manual toil.
  • Drive cross-team initiatives — Lead complex, multi-team infrastructure projects end-to-end, from design through rollout.
  • Mentor and grow engineers — Guide and mentor other DevOps/infra engineers, review designs, and raise the technical bar across the team.
  • Build and maintain CI/CD & GitOps pipelines — Design and operate deployment pipelines using GitHub Actions, ArgoCD, Helm, and Terraform (CDKTF), enabling fast, safe, and reliable releases for product teams.
  • Ensure data security — Protect the company's most sensitive data through access control, data governance, and security-first infrastructure design.
  • Evolve the data platform — Work with modern data technologies (Snowflake, ClickHouse, Apache Iceberg, EMR) and contribute to the next generation of our data infrastructure.
  • Enable developer self-service — Improve our internal platform so engineering teams across the company can deploy, configure, and operate services independently.
  • Provide observability and reliability — Build monitoring, alerting, and debugging tools (Datadog, OpenTelemetry, ClickHouse) across all data and AI processes, and lead incident response for critical, high-scale systems.

Our Stack — AWS, Kubernetes (EKS), Kafka/MSK, Debezium, Airflow, Snowflake, ClickHouse, Apache Iceberg, EMR, ArgoCD, Terraform/CDKTF, Helm, GitHub Actions, Docker, Datadog, OpenTelemetry, MLFlow, n8n, API Gateway, Redis, MySQL, Teleport, TypeScript, Node.js, Python.

Your Experience & Skills

  • 6-8+ years of experience as a DevOps / Infrastructure / Platform Engineer, including experience leading technical initiatives, owning architecture decisions, or mentoring other engineers.
  • Deep, hands-on experience with Kubernetes — cluster management, networking, scaling, and troubleshooting at production scale.
  • Strong experience with Infrastructure as Code (Terraform/CDKTF, Helm) and GitOps workflows (ArgoCD or similar).
  • Proven ownership of CI/CD pipelines — designing, maintaining, and optimizing the full release cycle for multiple teams.
  • Deep experience with cloud infrastructure (AWS preferred) — networking, IAM, security, cost optimization.
  • Security-first mindset — strong understanding of application security, access control, and data protection best practices.
  • Experience with data infrastructure — Kafka, Airflow, EMR, Apache Iceberg, Snowflake, ClickHouse, or similar technologies.
  • Fluent in Linux environments, scripting, and at least one programming language (Python, TypeScript, Go).
  • Proven ability to design and own system architecture at scale, and to make and defend complex technical trade-offs across multiple teams and domains.
  • Experience owning production reliability for critical, high-scale systems — including incident leadership, postmortems, and driving systemic fixes.
  • Strong communicator and cross-functional collaborator, comfortable influencing engineering standards and best practices across an organization.
  • Understanding of products and a passion for building software that impacts millions of users.
  • Comfortable operating with high autonomy and ambiguity.

Big advantage:

  • Experience with AI/ML infrastructure — model serving, LLM deployment, AI gateways, GPU workloads, or AI observability tools (MLFlow).
  • Hands-on experience with AI agents and automation — LangChain, LangGraph, n8n, Claude Code, or similar agentic frameworks.
  • Experience building internal developer platforms or self-service tooling.
  • Passion for pushing the boundaries of what AI agents and autonomous