Machine Learning Systems and Infrastructure Developer
Siemens Kitchener, Ontario, Canada · CA$90K–CA$140K/yr
Automation Machinery Manufacturing · 10,001+ employees
About the role
The developer will build and maintain infrastructure for training robotics AI models, including GPU resource management, experiment tracking, and model rollout pipelines. They will collaborate with cross-functional teams to ensure model reproducibility, scalability, and operational readiness.
What they look for
Requirements
Candidates must hold a Bachelor's or Master's degree in a relevant engineering or science field and possess at least 3 years of experience in ML platform or infrastructure development. Strong proficiency in Python, containerization, and experience with ML frameworks like PyTorch or Ray is required.
Benefits
Full description
Change the future with us.
We are looking for dedicated and talented people who solve ever-changing challenges, customer needs, and questions from colleagues with clever concepts and creativity. We embrace change and work with curious minds re-inventing the future of work. Join us and let us focus together on what’s truly important: making lives better with new ideas and the latest technology around the world.
About the role:
Are you excited by the engineering work that turns model development into repeatable, governed production rollout?
Do you enjoy building the infrastructure behind training jobs, GPU environments, experiment tracking, model registries, evaluation gates, and controlled deployment of robotics AI models?
We are seeking a Machine Learning Systems and Infrastructure Developer to build the training infrastructure and model rollout layer for Physical AI. This developer will provision and operate ML environments, automate training and evaluation pipelines, manage model artifacts and registries, and define the promotion path from trained model to validated runtime candidate. The role connects robotics data, partner VLA/RFM work, model evaluation, release gates, and runtime deployment without owning robotics data semantics or the final runtime integration layer alone.
What will you do?
- Provision and maintain ML training environments for robotics AI workloads across local, edge, on-premise, and cloud-backed infrastructure where applicable.
- Build repeatable training, fine-tuning, evaluation, and model-packaging pipelines for VLA, VLM, perception, and robot-learning workflows.
- Manage model artifacts, model registries, dataset references, experiment metadata, lineage, approvals, promotion states, and rollback paths.
- Automate GPU resource allocation, job scheduling, environment reproducibility, dependency management, and capacity tracking.
- Define model rollout workflows from training output to validated runtime candidate, including evaluation gates, release notes, traceability, and operational handoff.
- Partner with the Senior AI/ML and Robotics Platform Engineer on dataset readiness, evaluation semantics, and model-quality criteria.
- Partner with ML inference, Edge Runtime, AX integration, QA, and DevOps engineers on packaging, deployment readiness, observability, and release automation.
What you will bring:
- Bachelor's or Master's degree in Computer Science, Machine Learning, Robotics, Computer Engineering, Software Engineering, Electrical Engineering, or a related field.
- 3+ years of Strong experience building ML platforms, training infrastructure, MLOps systems, model lifecycle tooling, or production AI infrastructure.
- Strong programming skills in Python and practical experience with containers, Linux, automation, APIs, configuration management, and reproducible environments.
- Experience with training orchestration, job scheduling, GPU provisioning, artifact management, experiment tracking, model registries, or release pipelines.
- Practical experience with ML frameworks and infrastructure such as PyTorch, Hugging Face tooling, Ray, Kubernetes, Slurm, MLflow, Weights & Biases, Docker, CUDA, or comparable systems.
- Understanding of GPU capacity, storage, networking, data locality, dependency isolation, cost controls, and failure recovery for training workloads.
- Experience connecting datasets, metadata, evaluation outputs, model artifacts, and deployment packages through traceable workflows.
- Ability to define model promotion and rollback processes with clear gates for quality, reproducibility, security, and runtime readiness.
- Experience partnering with data engineering, ML research, robotics, DevOps, QA, and runtime teams to make model development operational.
- Strong troubleshooting skills across training jobs, data access, dependencies, GPU infrastructure, model artifacts, and rollout automation.
What sets you apart:
- Experience with VLA, VLM, robotic foundation models, robot learning, imitation learning, reinforcement learning, or multimodal training workflows.
- Experience operating GPU clusters, shared training platforms, high-volume data pipelines, or hybrid on-premise and cloud ML infrastructure.
- Experience with distributed training, multi-node GPU workloads, or training orchestration frameworks such as Ray, Slurm, Kubernetes, Kubeflow, or comparable systems.
- Experience designing portable ML infrastructure across on-premise, cloud, or multicloud environments, including data locality, reproducibility, access control, and cost management.
- Familiarity with dataset versioning, data lineage, model cards, evaluation reports, approval workflows, and audit-ready model release practices.
- Experience with NVIDIA tooling, CUDA environments, distributed training, model export, quantization handoff, or edge-deployment preparation.
- Knowledge of governance concerns for robotics AI models, including reproducibility, traceability, fallback versions, data provenance, and release accountability.
- Ability to separate training infrastructure ownership from robotics data semantics, runtime inference optimization, and industrial platform integration.
- Strong judgment about when a model is ready to leave experimentation and enter controlled runtime validation.
Work Environment:
- This role is primarily focused on software and infrastructure, with close collaboration across robotics data, ML research, QA, DevOps, edge runtime, and platform integration teams.
- Hybrid work is supported, with on-site presence expected for training environment setup, model rollout validation, data pipeline integration, and reference-cell release readiness.
Salary is commensurate with experience, and ranges between $ 90,000 CAD- $ 140,000 CAD, excluding bonus and benefits. In addition to base salary, this role includes eligibility for an annual discretionary bonus of 10% of base salary, based on Company performance metrics.
Why Join Us?
- Impact: Build the founding team and platform capability for a strategic Physical AI initiative in Canada.
- Innovation: Work at the intersection of robotics, AI, industrial software, workflow orchestration, and platform engineering.
- Leadership: Shape the direction, culture, and operating model of a new strategic engineering organization.
- Collaboration: Partner with leading global research, product, and engineering teams.
- Growth: Help define how Physical AI platform capabilities scale from data foundations to runtime-enabled industrial applications.
Why you’ll love working for Siemens!
- Freedom and a healthy work- life balance– Embrace our flexible work environment with flex hours, telecommuting and digital workspaces.
- Solve the world’s most significant problems – Be part of exciting and innovative projects.
- Engaging, challenging, and fast evolving, cutting edge technological environment.
- Opportunities to advance your career and mentorship programs on a local and global scale.
- Competitive total rewards package.
- Profit sharing available.
- Rewarding vacation entitlement with the opportunity to buy and sell your vacation depending on your lifestyle.
- Contribute to our social responsibility initiatives focused on access to education, access to technology and sustaining communities and make a positive impact on the community.
- Participate in our celebrations, social events and offsite business events.
- Opportunities to contribute your innovative ideas and get paid for them!
- Employee perks and discounts.
- Diversity and inclusivity focused.
Siemens is proud to be an eight-time award winner of Canada’s Top 100 Employers, Canada’s Greenest Employers 2025 and Canada’s Top Employers for Young People 2025.
About us.
We share our ideas and champion the people behind them.
Siemens Canada is a leading technology company focused on industry, infrastructure, mobility and healthcare. The company’s purpose is to create technology with purpose, transforming the everyday, for everyone, since 1912. By combining the real and the digital worlds, Siemens empowers its customers to accelerate their digital and sustainability transformations, making factories more efficient, cities more liveable, and transportation more sustainable. Siemens also owns a majority stake in the publicly listed company Siemens Healthineers, a leading global medical technology provider pioneering breakthroughs in healthcare. For everyone. Everywhere. Sustainably.
In fiscal 2024, which ended September 30, 2024, Siemens Canada had revenues of approx. $2.2 billion CAD. The company has approximately 4,400 employees from coast-to-coast and 37 office and production facilities across Canada.
To learn more about Siemens Canada, visit our website at www.siemens.ca
While we appreciate all applications we receive, we advise that only candidates under consideration will be contacted.
Siemens is committed to creating a diverse environment and is proud to be an equal opportunity employer. Upon request, Siemens Canada will provide reasonable accommodation for disabilities to support participation of candidates in all aspects of the recruitment process. All qualified applicants will receive consideration for employment.
By submitting personal information to Siemens Canada Limited or its affiliates, service providers and agents, you consent to our collection, use and disclosure of such information for the purposes described in our Privacy Notice available at www.siemens.ca.
Siemens s’engage à créer un environnement diversifié et est fière d’être un employeur souscrivant au principe de l’égalité d’accès à l’emploi. Sur demande, Siemens Canada prendra des mesures d’accommodement raisonnables pour les personnes handicapées, dans le but de soutenir la participation des candidats dans tous les aspects du processus de recrutement. Tous les candidats qualifiés seront pris en considération pour ce poste.
En transmettant des renseignements personnels à Siemens Canada limitée ou à ses sociétés affiliées, à ses fournisseurs de services ou à ses agents, vous nous autorisez à recueillir, à utiliser et à divulguer ces renseignements aux fins prévues dans notre Déclaration de protection de la confidentialité, que vous pouvez consulter au www.siemens.ca.
Job ID
524099
Posted since
01-Oct-2026
Organization
Digital Industries
Field of work
Research & Development
Company
Siemens Canada Limited
Experience level
Experienced Professional
Job type
Full-time
Work mode
Hybrid (Remote/Office)
Employment type
Permanent
Location(s)
• Kitchener - Ontario - Canada
Similar roles
-
Machine Learning Scientist - Apple Services Engineering, GenAI & ML Frameworks
Apple New York, New York, United States
-
Machine Learning Co-Op (Fall 2027)
Hendrickson Canton, Ohio, United States
-
Machine Learning Engineer (5-8 yrs)
Advanced Space Westminster, Colorado, United States · $124K–$171K/yr
-
Senior Machine Learning Engineer - New Verticals Agentic Foundations
DoorDash USA San Francisco, California, United States · $137K–$299K/yr
-
Senior Machine Learning Engineer, Payments
Airbnb United States · $191K–$223K/yr
-
Artificial Intelligence and Machine Learning Engineer, Senior
Booz Allen Hamilton Wahiawa, Hawaii, United States · $99K–$225K/yr