Jobgether

Principal Operations Engineer, Network

Jobgether United States

Internet Marketplace Platforms · 11-50 employees

13 h ago
Remote Principal (10+ yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Principal Operations Engineer will serve as the senior technical authority for network infrastructure across a hyperscale AI data center portfolio. Responsibilities include leading site readiness, executing high-risk production changes, and driving fleet-wide incident resolution.

What they look for

Network operations IP networking Optical networking TCP/IP Routing protocols OSPF IS-IS BGP MPLS Change management Incident management Vendor management Technical leadership Root cause analysis Data center infrastructure Linux

Requirements

Candidates must have extensive professional experience operating mission-critical network topologies at scale and deep expertise in IP and optical networking. Strong leadership, vendor management, and the ability to travel 50-75% of the time are essential for this role.

Benefits

Competitive compensation Equity Retirement plan Health insurance Dental insurance Vision insurance Paid time off

Full description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Principal Operations Engineer, Network based in the United States.

This is a senior technical leadership role responsible for the operational reliability of network infrastructure across a hyperscale AI data center portfolio. You will serve as the highest-level technical authority for operational networks spanning switches, routers, copper and fiber cabling, and optical systems. The role combines hands-on network operations with site readiness, technical audits, high-risk change execution, and fleet-wide incident leadership. You will influence how network platforms and integration designs transition from deployment into reliable production operations at unprecedented scale. Working across network engineering, compute operations, facilities, supply chain, vendors, and customer-facing teams, you will act as a technical force multiplier across the organization. You will also establish operational standards and playbooks while helping teams develop the technical readiness required to activate and operate new sites. This opportunity is designed for an accomplished network operations expert who thrives in high-intensity environments, takes end-to-end ownership, and is comfortable traveling extensively to support critical infrastructure.

\n

Accountabilities

  • Network Operations Leadership: Serve as the senior technical authority for the operational network fleet across large-scale AI data centers, covering switches, routers, copper and fiber cabling, optics, and associated physical infrastructure.
  • Site Readiness: Lead network site assessments, operational audits, and readiness activities to ensure network operations teams are technically prepared ahead of new site activations.
  • Operational Design Review: Review network platforms, integration designs, and deployment approaches from an operational perspective and provide feedback to engineering, deployment, and supply chain teams.
  • Production Change Management: Author, approve, and execute high-risk network Methods of Procedure (MOPs) and change records within live production environments.
  • Incident & RCA Leadership: Lead fleet-wide root cause analysis for significant network disruptions, driving investigations through resolution and closure while ensuring lessons learned are captured and applied.
  • Vendor Accountability: Establish and enforce technical and operational standards with OEMs, ODMs, deployment partners, and service vendors, including addressing RMA and integration process deficiencies while maintaining productive relationships.
  • Cross-Functional Leadership: Serve as the connective link between network operations, network engineering, compute operations, facilities, supply chain, and customer-facing teams.
  • Operational Standards: Develop and document the operating standards, processes, health assessments, incident procedures, and technical playbooks required to scale network operations effectively.
  • Technical Readiness: Teach and mentor operational teams, strengthening technical capabilities and ensuring teams are equipped to operate increasingly complex network environments.
  • Build-to-Operations Transition: Help establish new sites from construction or handover through activation and steady-state operations, ensuring operational requirements are incorporated early.
  • Fleet-Scale Improvement: Identify opportunities to improve reliability, efficiency, tooling, and operational processes across a rapidly expanding network fleet.
  • Strategic Technical Influence: Provide clear technical feedback on architecture, deployment, integration, and operational decisions while challenging assumptions and applying first-principles thinking.
  • Travel & Field Support: Travel approximately 50–75% to support site assessments, activations, audits, operational readiness, and critical network initiatives.

Requirements

  • Network Operations Experience: Extensive professional experience operating mission-critical network topologies at scale, including significant experience serving as the senior technical authority for a site, campus, or network fleet.
  • IP Networking: Deep hands-on experience with IP network operations and production network environments.
  • Optical Networking: Strong experience with optical networking technologies and physical network infrastructure.
  • Networking Protocols: Strong understanding of TCP/IP and routing protocols including OSPF, IS-IS, BGP, and MPLS.
  • Physical Infrastructure: Practical knowledge of network devices, cabling, copper and fiber infrastructure, optics, and related physical components.
  • High-Risk Changes: Proven experience authoring and executing high-risk network MOPs and production change records.
  • Incident Management: Demonstrated ability to lead root cause analysis for major network events and drive investigations through complete resolution.
  • Vendor Management: Experience holding OEMs, ODMs, and deployment or service partners accountable to technical standards while maintaining strong working relationships.
  • Technical Communication: Excellent written communication skills, with the ability to produce clear health assessments, RCAs, design feedback, operational documentation, and technical procedures.
  • Leadership & Mentoring: Ability to teach, mentor, and raise the technical readiness of engineers and operations teams.
  • Autonomy: Highly self-directed, comfortable taking ownership of broad technical scope, and capable of making sound decisions without constant oversight.
  • High-Intensity Environment: Comfortable working with urgency and managing complex operational challenges in rapidly scaling infrastructure environments.
  • Hyperscale Experience: Experience operating hyperscale or large HPC fleets supporting thousands of endpoints is highly desirable.
  • Linux & Tooling: Familiarity with Linux and hardware management tooling is a plus.
  • New Site Deployment: Experience standing up new sites from handover through operational steady state is beneficial.
  • Automation: Scripting experience for fleet-scale network operations is advantageous.
  • Travel: Willingness and ability to travel approximately 50–75% as required by the role.

Benefits

  • Competitive Compensation: Competitive total compensation package combining base salary and equity.
  • Equity: Full-time compensation may include equity in the form of restricted stock units.
  • Retirement: Retirement or pension plan aligned with applicable local norms.
  • Healthcare: Health, dental, and vision insurance.
  • Time Off: Generous paid time off policy aligned with local norms.
  • Pay Transparency: Commitment to compensation transparency and pay equity.
  • Remote Work: U.S.-based remote work arrangement with substantial travel requirements.
  • Technical Impact: Opportunity to shape network operations for infrastructure supporting the rapid expansion of AI computing.
  • Scale & Ownership: Significant autonomy and ownership across large-scale, mission-critical infrastructure.
  • Professional Growth: Opportunity to influence operational standards, develop teams, and help establish scalable systems and processes from the ground up.

\nHow Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1