Network Solutions Architect, AI Factory Services
Lenovo Morrisville, North Carolina, United States
IT Services and IT Consulting · 10,001+ employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
The Network Solutions Architect will design, deploy, and optimize high-performance network infrastructure for AI Factory and GigaFactory service offerings. This role involves developing reference architectures, validating GPU cluster fabrics, and providing technical leadership for global professional services teams.
What they look for
Requirements
Candidates must have a Bachelor's or Master's degree in a technical discipline and at least 5 years of experience in high-performance networking for AI, HPC, or GPU environments. Expertise in NVIDIA networking technologies, large-scale fabric design, and customer-facing technical consulting is required.
Full description
Why Work at Lenovo
We are Lenovo. We do what we say. We own what we do. We WOW our customers.
Lenovo is a US$83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY).
This transformation together with Lenovo’s world-changing innovation is building a more inclusive, trustworthy, and smarter future for everyone, everywhere. To find out more visit www.lenovo.com, and read about the latest news via our StoryHub.
Description and Requirements
Job Summary
Lenovo seeks a highly experienced Network Solutions Architect to join the Hybrid Cloud Solutions and AI Offering Engineering team within SSG. This senior-level role designs, deploys, validates, and optimizes high-performance network infrastructure for Lenovo's AI Factory and GigaFactory service offerings, supporting both enterprise-scale and hyperscale GPU environments powered by NVIDIA Spectrum-X Ethernet and InfiniBand fabrics.
The architect will develop network reference architectures, deployment runbooks, performance validation procedures, and field-ready engineering documentation consumed by Lenovo Professional Services and Managed Services teams globally. This includes GPU cluster fabric design, BlueField DPU architectures, multi-tenant network isolation, high-performance RDMA and RoCEv2 deployments, and production-scale AI infrastructure supporting large distributed training and inference workloads.
The ideal candidate brings deep expertise in NVIDIA Spectrum-X, Quantum InfiniBand, and BlueField DPU technologies, along with hands-on experience architecting, deploying, validating, and troubleshooting large-scale GPU clusters. Experience supporting hyperscale AI infrastructure, high-density liquid-cooled environments, GPU cluster bring-up, infrastructure validation, performance tuning, and customer-facing technical engagements is highly desired.
Key Responsibilities
AI Fabric Architecture & Design
- Design GPU cluster network architectures for NVIDIA AI Factory environments utilizing:
- Spectrum-X Ethernet (Spectrum-4 SN5600 + BlueField-3 DPUs) for enterprise deployments
- Quantum XDR InfiniBand and Spectrum Ethernet for rack-scale GigaFactory deployments
- Design and validate large-scale AI fabrics supporting RDMA, RoCEv2, GPUDirect, NCCL collectives, and high-bandwidth GPU-to-GPU communications.
- Develop rail-optimized InfiniBand topologies and Spectrum-X Adaptive Routing, SHARP, congestion management, and performance optimization strategies for AI training and inference environments.
- Contribute to architecture decisions supporting large-scale distributed AI and HPC workloads.
Deployment, Validation & Performance Engineering
- Lead network bring-up, validation, and production-readiness activities for GPU infrastructure.
- Develop and validate RDMA/RoCEv2 configuration guides and deployment runbooks for AI and HPC environments.
- Establish performance baselines for:
- RDMA throughput
- RoCEv2 latency
- GPU communication efficiency
- Storage and network throughput
- End-to-end infrastructure readiness
- Design failure-domain isolation strategies and resilient network architectures for large-scale AI deployments.
- Support network architecture for high-density liquid-cooled GPU environments with power-aware design considerations.
Troubleshooting & Operational Engineering
- Diagnose and resolve complex InfiniBand, Ethernet, RDMA, and GPU workload performance issues.
- Analyze congestion, telemetry, traffic patterns, link utilization, routing behavior, and fabric health to identify bottlenecks and optimize performance.
- Support root cause analysis and remediation of network, storage, and infrastructure issues impacting AI workload performance.
- Develop operational runbooks and troubleshooting procedures consumed directly by Lenovo field teams and customers.
Multi-Tenant & Cloud Architecture
- Design multi-tenant networking architectures supporting AI Factory, NeoCloud, and managed-service provider environments.
- Implement namespace isolation, east-west traffic segmentation, secure tenant separation, and site resiliency aligned with NVIDIA Cloud Partner Reference Architectures.
- Design BlueField-3 DPU solutions utilizing DOCA for infrastructure offload, security services, observability, and service mesh capabilities.
Customer & Cross-Functional Engagement
- Collaborate with Lenovo engineering, product, ISG, and services organizations to validate designs against NVIDIA reference architectures and future hardware roadmaps.
- Work directly with customers, partners, and delivery teams to translate AI workload requirements into production-ready infrastructure solutions.
- Provide technical leadership during customer engagements, infrastructure deployments, escalations, and architecture reviews.
- Support global deployments through approximately 40-50% travel, including customer workshops, implementation support, solution validation, and executive-level technical discussions.
Basic Qualifications
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, Information Technology, or related discipline.
- 5+ years of experience designing and deploying high-performance networking solutions for:
- AI infrastructure
- HPC environments
- GPU clusters
- Hyperscale cloud environments
- Experience designing and validating large-scale network fabrics supporting AI and distributed compute workloads.
- Experience with customer-facing technical consulting, architecture reviews, or deployment leadership.
Preferred Qualifications
NVIDIA AI Infrastructure Expertise
- Expert-level knowledge of:
- NVIDIA Spectrum-X
- Spectrum-4 SN5600 Ethernet
- Quantum InfiniBand (HDR/NDR/XDR)
- BlueField DPUs
- DOCA SDK
- RDMA and RoCEv2
AI Fabric & GPU Cluster Experience
- Experience supporting large-scale GPU cluster deployments and production AI environments.
- Deep understanding of:
- GPUDirect
- NCCL
- AI workload communication patterns
- GPU cluster validation and performance tuning
- Expertise in:
- Adaptive Routing
- SHARP
- Congestion control
- Rail-optimized InfiniBand architectures
Networking & Infrastructure
- Strong experience with:
- BGP
- EVPN
- VXLAN
- MPLS
- OSPF
- IS-IS
- Leaf-Spine architectures
- Experience designing lossless Ethernet environments utilizing:
- PFC
- ECN
- DCQCN
- QoS
- Knowledge of high-performance storage networking and end-to-end infrastructure optimization.
Automation & Observability
- Experience with:
- Python
- Ansible
- Terraform
- REST APIs
- Experience implementing infrastructure observability, telemetry, monitoring, and troubleshooting solutions for large-scale networks.
- Familiarity with data center operations, performance analytics, and capacity planning.
Certifications
Preferred certifications include:
- NVIDIA Certified Networking Professional
- NVIDIA InfiniBand Specialist
- Cisco CCNP Data Center
- Cisco CCIE Data Center
- OCI Networking Specialist
- AWS Solutions Architect Associate
Hybrid Schedule on campus in Morrisville, NC. 3 days in office, 2 days work from home.
We are an Equal Opportunity Employer and do not discriminate against any employee or applicant for employment because of race, color, sex, age, religion, sexual orientation, gender identity, national origin, status as a veteran, and basis of disability or any federal, state, or local protected class.
Additional Locations:
* United States of America - North Carolina - Morrisville
General Information
Req #
WD00099196
Career area:
Artificial Intelligence
Country/Region:
United States of America
State:
North Carolina
City:
Morrisville
Date:
Wednesday, June 3, 2026
Working time:
Full-time
Additional Locations:
* United States of America - North Carolina - Morrisville
Similar roles
-
Oracle Cloud Architect
Booz Allen Hamilton Springfield, Virginia, United States · $87K–$198K/yr
-
Software Architect - Houston, TX - 19562
ManpowerGroup WorkMyWay Houston, Texas, United States
-
Cloud Architect
Booz Allen Hamilton Arlington, Virginia, United States · $62K–$141K/yr
-
Solutions Architect
Booz Allen Hamilton McLean, Virginia, United States · $99K–$225K/yr
-
Defense Solutions Architect
Booz Allen Hamilton McLean, Virginia, United States · $99K–$225K/yr
-
Solutions Architect, Senior
Booz Allen Hamilton Ashburn, Virginia, United States · $113K–$257K/yr