Senior Manager, Network Engineering
Nscale Seattle, Washington, United States · $210K–$360K/yr
Technology, Information and Internet · 201-500 employees
About the role
You will lead and mentor a US-based network engineering team while actively contributing to fabric design, automation, and complex troubleshooting. The role involves owning the delivery of network workstreams and driving operational excellence through incident reduction and infrastructure-as-code practices.
What they look for
Requirements
The ideal candidate possesses extensive experience in people management and deep technical expertise in high-performance Ethernet fabrics and InfiniBand operations. Proficiency in network automation tools like Python, Ansible, and GitOps workflows is essential for this senior-level position.
Benefits
Full description
.
About the Role
The Network Engineering Team is responsible for the design, validation, and ongoing operation of all networking services that underpin both the internal management platform and the customer-facing cloud infrastructure — including high-performance Ethernet fabrics, InfiniBand, WAN connectivity, and DC networking. The team acts as a 3rd/4th line escalation point for the support organisation.
As Senior Network Engineering Manager, you will lead a US-based team of network engineers, anchoring US-hours network engineering capability within our global follow-the-sun operation. This is a player-coach role: you'll spend a meaningful portion of your time as a hands-on senior engineer — designing, automating, researching and troubleshooting alongside your team — while owning its delivery, development, and operational performance.
What You'll be Doing (Responsibilities)
- Lead, coach, and grow a team of four network engineers: setting clear ownership of technical areas, managing performance, unblocking delivery, and raising capability through mentoring and reviews.
- Contribute directly as a senior individual contributor: fabric design, complex troubleshooting, automation development, and design reviews across Ethernet, InfiniBand, and security domains.
- Own delivery of network workstreams for US deployments — scope, timelines, dependencies, and outcomes — from design through validation and into operations.
- Drive operational excellence and incident reduction: lead root-cause analysis for performance and stability issues, act as a senior escalation point during US hours, and reduce reactive load through runbooks, automation, and measurable SLOs.
- IImplement and uphold reference architectures and standards for high-performance Ethernet fabrics (BGP, EVPN, VxLAN, LACP, QoS), delivered consistently through infrastructure-as-code and GitOps practices — version-controlled configuration, peer-reviewed changes, and automated CI/CD pipelines.
- Provide technical leadership on InfiniBand operations for accelerated compute clusters: Subnet Managers, routing, congestion control, QoS, firmware management, and fabric health.
- Ensure the accuracy of source-of-truth network inventory and configuration data, with changes flowing through structured change management.
- Coordinate cross-team dependencies with deployment, DC operations, platform engineering, procurement, and vendors, ensuring clean handovers across regions and shifts.
- Provide clear visibility of team output, operational metrics, and escalations to engineering leadership.
About You
- Proven people-management experience building and leading a technically strong network engineering team through coaching, performance management, and clear ownership of outcomes — while retaining genuine hands-on depth.
- Strong technical depth in high-performance Ethernet fabrics: hands-on production experience with BGP, EVPN, VxLAN overlays, IPv6, LACP, QoS, and troubleshooting at scale (CLI tooling, Wireshark).
- Experience operating InfiniBand and/or RoCE fabrics for accelerated compute (NVIDIA Quantum/Mellanox), including performance tuning and fabric health management.
- Strong network automation experience: Python and Ansible for provisioning, configuration validation, and compliance across multi-vendor environments; comfortable working in a GitOps model with version-controlled network configuration and CI/CD-driven change.
- Familiarity with modern IaC and orchestration tooling (Terraform, Git-based workflows, pipeline tooling such as GitLab CI/GitHub Actions); experience treating the network as code rather than managing devices by hand.
- Depth in large-scale data centre network design (Clos/spine-leaf) and resilient multi-site connectivity.
- Comfortable leading incident response at senior escalation levels and driving measurable improvements in reliability, performance, and change success.
- Nice to have: SONiC-based switching experience; production multi-tenant InfiniBand operations (SR-IOV, isolation); experience standing up a new regional team or function.
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
Salary Range
$210,000—$360,000 USD
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.