Calix

Staff Engineer Cloud (Embedded SRE)

Calix · Bengaluru, Karnataka, India

Software Development · 1,001-5,000 employees

13 h ago
Principal (10+ yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design and maintain scalable backend infrastructure while embedding within engineering teams to drive production readiness and service reliability. Own observability, monitoring, and incident response improvements to ensure high availability and performance of the cloud platform.

What they look for

Java SRE Cloud Computing Microservices Kafka Observability AWS GCP Distributed Systems API Design NoSQL Incident Management Data Ingestion Reliability Engineering System Architecture Automation

Requirements

Requires 10+ years of hands-on development experience with expertise in Java and microservices-based architectures. Candidates must have a strong background in cloud platforms, event-based workflows, and reliability engineering practices.

Full description

The Calix platform enables Communication Service Providers (CSPs) of all sizes to transform and future-proof their businesses. Through real-time data, automation, and actionable insights delivered via Calix One — our cloud-first, AI-powered platform — CSPs can simplify operations, collapse cost, and accelerate innovation. Calix One brings together the automation of everything and the experience of one, empowering customers to deliver differentiated subscriber experiences while driving acquisition, loyalty, and revenue growth. This is the Calix mission: to enable CSPs of all sizes to simplify, innovate, and grow, strengthening both their businesses and the communities they serve.

We’re at the forefront of a once in a generational change in the broadband industry. Join us as we innovate, help our customers reach their potential, and connect underserved communities with unrivaled digital experiences.

As part of a high-performing global engineering team, the right candidate will play a critical role in expanding the Calix Cloud solution with a strong focus on reliability, observability, and operational excellence. Calix Cloud empowers customer service organizations with insights and real-time data to maximize support operations efficiency and deliver the best customer experiences.

Responsibilities:

  • Design, develop and maintain backend infrastructure, workflows, and services with a focus on reliability, scalability, and operability (SRE principles)
  • Develop solutions to support onboarding, partner integrations, managing, collecting, and analyzing data from large-scale deployments of home networks
  • Embed within engineering teams to drive production readiness, service health, and continuous improvement of reliability metrics (SLIs/SLOs)
  • Work closely with Cloud product owners to understand, analyze product requirements, provide feedback, and deliver a complete solution
  • Technical leadership in software design to meet requirements of service stability, availability, scalability, and security
  • Drive technical discussions across SDLC phases including requirements, design, peer reviews, and test strategy
  • Own observability, monitoring, alerting, and incident response improvements; partner with TAC and operations teams for faster resolution
  • Support test strategy and automation in both end-to-end solution and functional testing
  • Customer-facing engineering role in debugging and resolving field issues
  • Drive root cause analysis (RCA), post-incident reviews, and ensure systemic fixes and prevention mechanisms

Qualifications:

  • 10+ years of highly technical, hands-on development experience in any programming language
  • Independent, self-driven, and able to work in a team environment
  • Strong problem-solving skills with ability to abstract and communicate effectively
  • Ability to drive technical discussions across cross-functional teams
  • Proficient in design and implementation of microservices-based, API/endpoint architectures
  • Strong background in event-based / pub-sub workflows & data ingestion solutions (Kafka or similar)
  • Good understanding of cloud-based solutions (preferably AWS or GCP)
  • Strong background in transactional databases and experience with NoSQL data stores
  • Experience with monitoring/observability tools (Prometheus, Grafana, OpenTelemetry or similar)
  • Experience with incident management, production support, and reliability engineering practices
  • Understanding of scaling, resiliency patterns, and failure handling in distributed systems
  • Experience with IoT/home gateway protocols (TR-069/TR-369 etc.) a plus
  • Expert in Java; experience in Go/Python/NodeJS a plus
  • Experience with streaming/data platforms (Kafka, Spark, Flink, etc.)
  • Experience building scalable data solutions (Spanner, Elastic, etc.)
  • Practical understanding of AWS/GCP cloud platform

Education:

BS degree in Computer Science, engineering, or equivalent experience

Location:

  • India – (Flexible hybrid work model - work from Bangalore office for 20 days in a quarter)