Conifers.ai

Site Reliability Engineering Manager

Conifers.ai Tel-Aviv, Tel-Aviv District, Israel

Computer and Network Security · 11-50 employees

5 h ago
sre Senior (5-10 yrs) Full-time Israel
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Site Reliability Engineering Manager will own the end-to-end health and reliability of the production platform while leading a team of engineers. They will establish monitoring, observability, and incident management processes to ensure system stability and performance.

What they look for

Site Reliability Engineering Production Engineering Incident Management Observability Monitoring Distributed Systems Automation Cloud Computing Containerization Root Cause Analysis Technical Leadership Software Engineering Scalability Performance Tuning System Resilience

Requirements

Candidates must have at least 4 years of experience in SRE, production, or backend engineering with a strong background in technical leadership. Proficiency in distributed systems, observability, and troubleshooting complex production issues is essential.

Full description

Conifers is transforming security operations centers (SOCs) with CognitiveSOC™, its AI SOC platform, enabling enterprises and MSSPs to achieve SOC excellence.

By leveraging agentic AI, Conifers helps security teams investigate complex, multi-tier incidents with speed, accuracy, and trust.

Led by seasoned cybersecurity leaders and backed by SYN Ventures, PICUS Capital, and others, the company brings deep industry knowledge and innovation to an increasingly AI-driven threat landscape.

We’re building an AI-native security platform that enables autonomous agents to investigate real-world threats at a massive scale.

About The Role :

At Conifers, reliability means more than keeping the infrastructure up. It means making sure our product works correctly, consistently, and predictably for our customers, 24/7.

We are looking for a Site Reliability Engineering Manager to take end-to-end ownership of the health and reliability of our production system.

This is a highly hands-on role. You will lead a small engineering team focused on production reliability and work closely with R&D, Product, and GTM to make sure the system is observable, measurable, stable, and ready for production.

You will own the mechanisms that allow us to understand, at any point in time, whether the system is healthy and behaving as expected, from infrastructure and services to integrations, investigation pipelines, AI agents, and customer-facing functionality.

When something goes wrong, you will help drive it from detection through investigation, resolution, and prevention.

The goal is simple: make sure Conifers works reliably, accurately, and continuously for every customer.

What You’ll Do :

  • Own the end-to-end health and reliability of the Conifers production platform, beyond infrastructure availability.
  • Build and continuously improve the monitoring, observability, alerting, dashboards, health checks, and operational tooling required to understand whether the system is working correctly.
  • Define what "healthy" means across the platform and establish clear reliability, availability, performance, and product health metrics and targets.
  • Proactively identify production issues, abnormal behavior, degradation, and reliability risks before they impact customers.
  • Own and coordinate the response to production incidents and critical bugs, working with the relevant engineering teams until issues are fully resolved.
  • Establish strong incident management, root cause analysis, postmortem, and follow-up processes to ensure we learn from failures and prevent recurrence.
  • Work closely with R&D teams to improve system resilience, error handling, monitoring, scalability, and production readiness.
  • Partner with Product to ensure new capabilities have clear production health indicators, monitoring, failure handling, and operational readiness before release.
  • Work closely with GTM and customer-facing teams when production issues affect customers, helping investigate complex problems and drive them to resolution.
  • When needed, participate directly in technical troubleshooting with customers.
  • Build automation and internal tools that reduce manual operational work and enable faster detection, diagnosis, and recovery.
  • Own production readiness standards and help ensure new features are truly ready before reaching customers.
  • Lead a small team of engineers focused on production reliability while remaining deeply hands-on with the system and day-to-day production challenges.
  • Drive a strong engineering culture of ownership, operational excellence, measurable reliability, and continuous improvement.

What You’ll Need :

  • 4+ years of experience in SRE, production engineering, backend engineering, platform engineering, DevOps, or another highly production-focused engineering role.
  • Experience in technical leadership or engineering management.
  • Strong software engineering background and the ability to investigate complex production problems across multiple services and components.
  • Proven experience operating complex, highly available production systems.
  • Strong understanding of observability, monitoring, alerting, distributed systems, performance, scalability, and incident management.
  • Experience defining and working with production health, availability, performance, and reliability metrics and targets.
  • Strong troubleshooting and root cause analysis skills, with the ability to move from symptoms to the underlying technical problem.
  • Strong automation and programming skills.
  • Experience with modern cloud and containerized environments.
  • Ability to work effectively across Engineering, Product, GTM, and customer-facing teams.
  • Strong sense of ownership and the ability to drive complex production issues from detection through resolution.
  • Ability to lead a small team while remaining highly hands-on.

Nice to Have:

  • Experience in B2B SaaS or cybersecurity.
  • Experience operating AI-driven or agentic production systems.
  • Experience with systems where reliability includes not only availability and performance, but also the correctness and quality of system outputs.
  • Experience supporting enterprise customers and participating in complex production troubleshooting.
  • Experience building an SRE or Production Engineering function from the ground up in a fast-growing company.

If this role resonates with you and you’re excited about shaping how modern SOCs really work, this is your opportunity to join Conifers and build something that truly matters 🚀.

Our Commitment:

We are an equal opportunity employer and value diversity at our company.

All qualified applicants will receive consideration without regard to race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

#LI-AM1 #LI-Hybrid

Similar roles