Engineering Manager
DDN · Pune, Maharashtra, India
Software Development · 1,001-5,000 employees
About the role
Lead the design and development of manageability solutions for the DDN Infinia AI Data Platform, focusing on centralized control and automated operations. Manage a high-performing engineering team to deliver cloud-native, API-first, and AI-powered systems at petabyte scale.
What they look for
Requirements
Requires 15+ years of experience in software engineering, distributed systems, or cloud platforms, with at least 5 years in technical leadership roles. Candidates must have a strong background in cloud-native technologies, infrastructure-as-code, and building enterprise-scale systems.
Full description
As an Engineering Manager - Control Plane, you will lead the design and development of Manageability solutions for the DDN Infinia AI Data Platform. This role is responsible for building foundational capabilities that enable centralized control, automated operations, and intelligent support across large-scale hybrid (OnPrem + cloud) environments. You will lead a team delivering cloud-native, API-first, and AI/ML-powered systems that ensure operational excellence, proactive incident management, and seamless user experiences at petabyte scale. This is a ground-up platform leadership role focused on scalability, reliability, automation, and innovation.
Key responsibilities
- Lead a high-performing engineering team across distributed systems, cloud infrastructure, and AI/ML.
- Collaborate with cross-functional teams (product, engineering, SRE, security, and customer
- success) to align platform capabilities with business and customer needs.
- Establish engineering best practices, development standards, and operational excellence
- frameworks.
- Implement policy-driven infrastructure management and Infrastructure-as-Code (IaC)
- frameworks.
- Develop self-service tooling and role-based access control (RBAC) for enterprise customers.
- Design API-first management interfaces for integration with external tools and automation
- workflows.
- Drive proactive capacity planning and performance optimization for large-scale deployments.
- Build self-healing systems that reduce manual intervention and improve system resilience.
- Develop predictive analytics capabilities for capacity planning, performance forecasting, and
- failure prevention.
- Integrate intelligent recommendations and prescriptive insights into operational workflows.
- Define and enforce an API-first, cloud-native architecture across all components.
- Ensure systems are highly scalable, resilient, secure, and capable of operating at petabyte
- scale.
- Promote automation-first principles across development, testing, deployment, and
- operations.
- Oversee the design of distributed systems with high availability and fault tolerance.
Qualifications
- 15+ years of experience in software engineering, distributed systems, or cloud platforms
- 5+ years in technical leadership or management roles
- Proven experience building large-scale platform management or infrastructure systems
- Strong background in distributed systems architecture and cloud-native technologies
- Experience with APIs, microservices, and infrastructure-as-code (IaC)
- Familiarity with AI/ML concepts applied to operational analytics or automation• Experience managing teams delivering production-grade, enterprise-scale systems
- Experience in storage systems, data platforms, or high-performance computing
- environments
- Background in building AI-driven operations or AIOps platforms
- Experience with hybrid cloud and OnPrem deployments
- Knowledge of security, compliance, and enterprise governance requirements
- Familiarity with DevOps, SRE practices, and CI/CD pipelines
Success Metrics:
- Delivery of a unified management platform at scale
- Reduction in incident response and resolution times through automation
- Increased system uptime and reliability (zero or near-zero disruption)
- Adoption of self-service and automated operational workflows by customers
- High customer satisfaction and operational efficiency across deployments