Senior Platform Software Engineer (Stateful Data, Search and Analytics Platforms)
Medallia Pune, Maharashtra, India
Software Development · 1,001-5,000 employees
About the role
The engineer will operate and maintain stateful data, search, and analytics platforms including Elasticsearch, MongoDB, and ClickHouse. Responsibilities include ensuring system reliability, performing capacity planning, and leading production incident triage and performance tuning.
What they look for
Requirements
Candidates must have at least 5 years of relevant experience, including 3 years operating stateful distributed systems in production. Proficiency in at least two of the specified database platforms and experience with infrastructure automation in Java, Go, or Python are required.
Full description
Overview
Medallia is the pioneer and market leader in Experience Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and actions for candidates, customers, employees, patients, and residents alike.
We believe that every experience is a memory that can last a lifetime. Experiences shape the way people feel about a company. And they greatly influence how likely people are to advocate, contribute, and stay. At Medallia, we are committed to creating a world where organizations are loved by their customers and their employees.
We empower exceptional people to create extraordinary experiences together.
Bring your whole self.
The Role and Team
We are seeking a hands-on Senior Platform Software Engineer with experience operating stateful data, search, and analytics platforms at scale. As a key member of the Platform Services team, you will help ensure the availability, reliability, performance, and scalability of Elasticsearch, MongoDB, and ClickHouse.
This role is based remotely in Pune. Candidates for this position are required to reside within the Pune metropolitan area. Relocation support is not available at this time.
Responsibilities
- Operate and maintain Elasticsearch, MongoDB, and ClickHouse platforms.
- Manage Elasticsearch indices, shards, replicas, allocation, lifecycle policies, upgrades, and performance.
- Operate MongoDB replica sets and sharded clusters, including balancing, backups, restores, and recovery.
- Operate ClickHouse replication, Keeper, distributed tables, storage, ingestion, and query workloads.
- Troubleshoot availability, replication, storage, query performance, and client-integration issues.
- Perform capacity planning based on data growth, ingestion, retention, replication, and query patterns.
- Plan and validate upgrades, failover, backup, recovery, and failure scenarios.
- Lead production incident triage, root-cause analysis, and performance tuning.
- Review data-platform architectures and build operational automation and observability.
- Participate in a periodic on-call rotation to maintain 24/7 reliability and performance of production services.
Qualifications
Minimum Qualifications
- 5 years of software, systems, platform, database, DevOps, or SRE experience.
- 3 years of experience operating stateful distributed systems in production.
- Database Platforms: Hands-on operational experience with at least two of the following platforms in production: Elasticsearch, MongoDB, or ClickHouse.
- Incidents & Troubleshooting: Experience leading production incident triage, root-cause analysis (RCA), and performance tuning for database availability, replication, and storage issues.
- Resiliency & Capacity: Experience executing cluster upgrades, failovers, backups, disaster recovery, and capacity forecasting for data-intensive platforms.
- Software Automation: Experience writing and maintaining infrastructure automation or backend tooling in Java, Go, or Python.
- Technical Collaboration & Review: Experience conducting formal architecture reviews, documenting system design proposals (e.g., RFCs/ADRs), and partnering across engineering teams to standardize infrastructure patterns.
Preferred Qualifications
- Experience with the third platform among Elasticsearch, MongoDB, and ClickHouse.
- Experience operating stateful platforms on Kubernetes or cloud infrastructure.
- Experience designing and validating disaster-recovery and failure scenarios.
- Experience with platform automation, observability, GitOps, or infrastructure as code.
- Experience supporting large-scale, business-critical data platforms.
At Medallia, we celebrate diversity and recognize the value it brings to our customers and employees. Medallia is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age (40 and over), disability, genetic information, veteran status or military service, or any other status protected by state or local law. Individuals with a disability who need an accommodation to apply please contact us at ApplicantAccessibility@medallia.com. For information regarding how Medallia collects and uses personal information, please review our Privacy Policies. Applications will be accepted for 30 days from the date this role was posted or until the role has been filled.