Nscale

Software Engineering Manager - Fleet Management

Nscale · Seattle, Washington, United States · $300K–$350K/yr

Technology, Information and Internet · 201-500 employees

2 h ago
Remote Principal (10+ yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Lead and grow a team of software engineers to build and maintain the Fleet Manager workflow automation platform for GPU infrastructure. Oversee end-to-end delivery, team health, and operational excellence while partnering with technical leadership on architecture.

What they look for

Python Distributed systems Infrastructure automation Workflow orchestration GPU infrastructure Team leadership Performance management Incident response System architecture Cloud infrastructure Kubernetes Terraform Hardware lifecycle management HPC Networking CI/CD

Requirements

Requires 8+ years of software engineering experience with at least 2 years in a management role, specifically with Python and distributed systems. Candidates must have a proven track record of delivering complex production systems and strong leadership skills in hiring and coaching.

Benefits

Medical Dental Vision Flexible paid time off Parental leave Retirement plan Equity Bonus

Full description

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.

About the Role

We're hiring a Software Engineering Manager to lead the team building Fleet Manager — the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale.

Reporting to the Director of Software Engineering, Fleet Management based in EMEA, you'll lead the engineers who build the Python-based systems managing the entire operational lifecycle of our compute infrastructure: device enrollment, burn-in testing, network configuration, GPU health monitoring, and automated remediation. You'll own delivery and team health end to end — planning, execution, hiring, and career development — while partnering closely with Principal and Staff engineers who drive technical architecture.

We're a hands-on engineering culture: you'll stay close to the systems through design reviews, incident deep-dives, and code, with people leadership as your primary craft.

What you'll be doing

Lead and grow the team

  • Hire, onboard, and develop software engineers who thrive in a high-autonomy, high-accountability environment.
  • Coach performance and career growth, set clear expectations, and handle performance management directly and fairly.
  • Own on-call health, handover quality, and a sustainable operational load for the team.

Own delivery and execution

  • Turn the Fleet Manager roadmap into executable plans: drive prioritization, manage dependencies and risks, and ship on commitments.
  • Run the team's planning, review, and escalation mechanisms, and remove blockers before they become delays.
  • Communicate status, risks, and trade-offs clearly to stakeholders — including when the news is bad.

Drive engineering and operational excellence

  • Uphold engineering standards: code review, testing, CI/CD, incident response, and postmortems that produce durable fixes rather than closed tickets.
  • Own the operational health of the team's services: SLOs, observability, alerting, and incident management.

Stay technical and partner across Nscale

  • Partner with Principal and Staff engineers on architecture and design decisions that balance automation complexity, reliability, and maintainability.
  • Stay hands-on where it counts: review designs and code, dig into incidents, and write code where it unblocks the team.
  • Collaborate with Product, Infrastructure, Platform, SRE, and UI/UX to capture requirements early, align on interfaces, and ship integrations that meet operator needs.
  • Champion AI tools like Claude and Cursor across the team as a fundamental multiplier of what your engineers can build.

About You

  • 8+ years of software engineering experience building and operating production systems, including 2+ years directly managing software engineers.
  • Strong technical foundation in Python and distributed systems: credible in design and code reviews, and able to guide trade-offs in infrastructure automation or workflow tooling.
  • Track record of delivering complex projects from ambiguous requirements to production, with hands-on day-2 operations experience (monitoring, incident response, performance optimization).
  • Proven people leadership: hiring, coaching, performance management, and developing engineers toward senior and staff levels.
  • You are driven by building distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement.
  • You use AI tools like Claude or Cursor as a core part of your workflow and know how to raise a whole team's leverage with them.
  • You stay effective while context-switching between technical depth, delivery judgment calls, and people leadership — reviewing a design, unblocking an engineer, and running a hiring debrief in the same morning.
  • Excellent communication skills to build consensus with stakeholders, both internally and externally.

Strong candidates will have

  • Experience leading teams that build workflow orchestration systems (Temporal, Airflow, Prefect, or similar)
  • Hands-on experience with infrastructure tooling: DCIMs, NetBox, OpenStack, or ERP systems
  • Bare-metal provisioning and automation: MAAS, Ironic, IPMI, PXE boot, or network automation
  • Experience building or managing hardware lifecycle automation: provisioning, validation, testing, or remediation workflows
  • GPU infrastructure experience: health monitoring, burn-in testing, or cluster management
  • HPC and networking: datacenter topology, high-performance interconnects (InfiniBand, RoCE)
  • Working knowledge of Kubernetes, Infrastructure as Code (Terraform, Pulumi), AWS, and GCP
  • Experience scaling a team through rapid growth: hiring pipelines, onboarding, and team processes that hold up under pressure

What we can offer you

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.

  • Highly competitive US compensation package (base + bonus + equity), with performance reviews every 12 months. 🚀
  • Join one of the fastest-growing AI infrastructure companies — your chance to directly shape how global AI capacity is planned and deployed. ✨
  • Expect a dynamic progression plan tailored to your ambitions. Grow by leading critical cross-functional initiatives — always with our full support.
  • Human-First Flexibility: We treat you as humans first. 🫶🏽 Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.

Equal Opportunities Statement

We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there's anything we can do to accommodate your specific situation, please let us know.

The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$300,000—$350,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.