Sr Machine Learning Engineer (MLOps)
Jobgether India
Internet Marketplace Platforms · 11-50 employees
About the role
Design, build, and maintain reliable deployment pipelines and infrastructure for machine learning and generative AI models. Own production monitoring, model governance, and automated retraining processes to ensure system reliability and performance.
What they look for
Requirements
Requires 6+ years of experience in software or machine learning engineering with at least 3 years of hands-on production experience. Proficiency in Python, cloud platforms like AWS, and experience with LLM-based workloads is essential.
Benefits
Full description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Sr Machine Learning Engineer (MLOps) based in India.
This senior engineering role focuses on building the production foundation that enables AI and machine learning systems to operate reliably at scale. You will own the infrastructure, pipelines, monitoring, governance, and deployment practices that take models from experimentation into production. The role spans traditional machine learning as well as modern generative AI, LLM, retrieval-based, and agentic workloads. You will work closely with Data Science and Data Engineering teams to establish repeatable, observable, secure, and cost-effective ML operations. With significant ownership of production AI infrastructure, you will drive reliability, automation, governance, and continuous improvement. You will also collaborate with distributed U.S.-based teams while operating with a high degree of autonomy and technical responsibility.
\n
Accountabilities:
- Design, build, and maintain reliable deployment pipelines that move ML models from development and validation through production.
- Establish model versioning, lineage, registry, and automated model promotion practices.
- Define repeatable production-readiness standards and deployment patterns across machine learning and AI workloads.
- Partner with Data Scientists to make model handoffs efficient, consistent, and production-ready.
- Own production monitoring for model performance, drift, data quality, inference health, latency, availability, and other operational metrics.
- Establish alerts and operational thresholds that identify degradation before it materially impacts products or customers.
- Diagnose production failures, perform root-cause analysis, and implement durable corrective actions.
- Deploy and support production LLM applications, including retrieval-based and agentic architectures.
- Build evaluation frameworks to measure the quality, reliability, and performance of generative AI capabilities.
- Monitor token consumption, inference costs, and cost per interaction to help ensure AI workloads remain economically sustainable.
- Implement appropriate controls around model access, usage, safety, and production behavior.
- Design and operate infrastructure for model training, validation, and retraining.
- Build automated retraining pipelines triggered by relevant performance, data, or business conditions.
- Ensure training environments and workflows are reproducible, scalable, and observable.
- Partner with Data Engineering and Data Science to ensure reliable data movement throughout the ML lifecycle.
- Implement model access controls, auditability, lineage, and governance standards.
- Support model risk classification and establish appropriate controls based on use case and business impact.
- Produce documentation and technical evidence supporting security, compliance, and internal governance requirements.
- Respond to production incidents, troubleshoot failures, and coordinate resolution across teams when required.
- Create runbooks and operational procedures that reduce reliance on undocumented knowledge.
- Identify recurring operational issues and automate them away wherever practical.
- Contribute to continuous improvements in reliability, automation, security, scalability, and operational efficiency.
Requirements:
- 6+ years of experience in software engineering, data engineering, machine learning engineering, or a related technical discipline.
- 3+ years of hands-on experience deploying and operating AI or machine learning systems in production.
- Demonstrated experience supporting both traditional machine learning and LLM-based workloads in production environments.
- Strong understanding of the complete model lifecycle, including development, validation, deployment, monitoring, retraining, and retirement.
- Production experience with LLM-powered applications and agentic frameworks.
- Experience with retrieval architectures, generative AI evaluation methodologies, and production monitoring.
- Understanding of LLM performance, latency, token utilization, and cost-per-interaction management.
- Ability to establish practical operational, security, and governance controls around generative AI systems.
- Deep experience with a major cloud platform and its managed AI/ML services; AWS experience is strongly preferred.
- Hands-on experience with model registries, pipeline orchestration, ML CI/CD, automated retraining, and production monitoring.
- Strong Python engineering skills.
- Experience with containerization and infrastructure-as-code.
- Experience designing reliable, repeatable, automated production environments.
- Proven experience operating production services with meaningful responsibility for reliability and availability.
- Strong incident response, troubleshooting, and root-cause analysis skills.
- Ability to distinguish symptoms from underlying system failures and implement sustainable solutions.
- Comfortable making sound operational decisions independently when immediate support from distributed teams may not be available.
- Strong written technical communication skills.
- Experience creating runbooks, architectural documentation, standards, and operational procedures.
- Proactive communication style suited to distributed and asynchronous teams.
- Ability to collaborate effectively across Data Engineering, Data Science, Product, and other technical functions.
- Bachelor's degree in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience.
- Experience with AI governance, model risk tiering, access-control frameworks, modern data warehouse technologies, or production analytics environments is a plus.
- Experience working successfully with U.S.-based colleagues and distributed global teams is advantageous.
Benefits:
- Competitive, market-aligned compensation for India.
- Full-time employment through an Employer of Record partner.
- Remote working environment within India.
- Competitive local benefits provided through the Employer of Record.
- Ongoing professional development opportunities.
- Opportunity to work directly with U.S.-based Data and Technology teams.
- Meaningful ownership of production systems supporting a growing AI strategy.
- Exposure to traditional ML, generative AI, LLMs, retrieval systems, and emerging agentic technologies.
- Opportunity to solve complex MLOps challenges involving large-scale data and real-world customer applications.
- High-impact senior individual contributor role with significant technical ownership.
- Collaboration with Data Science, Data Engineering, Product, and technology leadership.
- Opportunity to influence AI/ML engineering standards, governance, automation, and production practices.
- Autonomy and responsibility within a distributed, remote-first environment.
\nHow Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Similar roles
-
Machine Learning Engineer - Voice AI & Generative Music (Part-time)
MWDN Kyiv, Ukraine
-
Machine Learning Engineer - Multimodal Intelligence
Apple Sunnyvale, California, United States
-
Senior Machine Learning Engineer (W/M/X)
Ubisoft Saint-Mandé, Ile-de-France, France
-
Staff Machine Learning Scientist/Engineer
Wayve Sunnyvale, California, United States · $370K–$419K/yr
-
Senior Machine Learning Engineer, Ads Response Prediction
Instacart Wasaga Beach, Ontario, Canada · CA$180K–CA$190K/yr
-
AI / Machine Learning Engineer
Lynx Madrid, Community of Madrid, Spain