Cornelis Networks, Inc.

Senior System Software Engineer for Kubernetes Fabric Integration

Cornelis Networks, Inc. · San Jose, California, United States

Computer Networking Products · 51-200 employees

Yesterday
Remote Senior (5-10 yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will design, develop, and maintain Kubernetes operators and controllers to integrate high-performance interconnect hardware into cloud-native environments. Additionally, you will lead the architecture of scalable fabric management solutions and contribute to open-source projects within the cloud-native ecosystem.

What they look for

Kubernetes Go Kubernetes Operators Kubebuilder Operator SDK System software engineering High-performance computing HPC Data center networking Multithreaded programming C++ Python Container technologies Docker Containerd Cloud-native orchestration

Requirements

The role requires at least 5 years of professional software development experience with expert proficiency in Go and deep knowledge of Kubernetes architecture. Candidates must have hands-on experience building custom operators and a solid understanding of high-performance networking environments.

Benefits

Performance incentives Equity participation Paid holidays Flexible work arrangements

Full description

At Cornelis we’re building the future of AI and HPC networking with an AI-first approach to silicon and software development. We’re seeking engineers who are energized by working on cutting-edge ASIC design and distributed software systems, and who are motivated to push the boundaries on how AI can transform everything from chip architecture to system performance at scale.

Cornelis Networks delivers the world’s highest performance scale-out networking solutions for AI and HPC datacenters. Our differentiated architecture seamlessly integrates hardware, software and system level technologies to maximize the efficiency of GPU, CPU and accelerator-based compute clusters at any scale. Our solutions drive breakthroughs in AI & HPC workloads, empowering our customers to push the boundaries of innovation. Backed by top-tier venture capital and strategic investors, we are committed to innovation, performance and scalability - solving the world’s most demanding computational challenges with our next-generation networking solutions.   

  

We are a fast-growing, forward-thinking team of architects, engineers, and business professionals with a proven track record of building successful products and companies. As a global organization, our team spans multiple U.S. states and six countries, and we continue to expand with exceptional talent in onsite, hybrid, and fully remote roles.   

 

Cornelis Networks is seeking a talented and experienced Senior System Software Engineer for Kubernetes Fabric Integration to build the bridge between our high-performance interconnect hardware and the cloud-native orchestration ecosystem. In this role, you will be extending the capabilities of Kubernetes to manage highly specialized large scale network fabrics. You will design, develop and test advanced fabric management software that will be deployed in massive AI and HPC clusters. You will design, build, and maintain Kubernetes operators, controllers, and other components necessary to ensure our high-performance interconnect solutions can be seamlessly deployed, managed, and scaled in containerized environments. This is a critical role that will directly impact our customers' ability to leverage Cornelis Networks' technology in large-scale, modern data centers. 

 

Key Responsibilities: 

  • Architect and Design: Lead the design of robust, scalable solutions for integrating Cornelis Networks' platform and fabric management software with Kubernetes. 
  • Develop Kubernetes Operators: Build and maintain custom Kubernetes Operators and Controllers in Go to manage the lifecycle of our software and hardware components within a cluster. 
  • Cloud-Native Integration: Develop solutions that allow for the seamless orchestration of our high-performance fabric services and platform management tools alongside other containerized workloads. 
  • Comprehensive Testing: Develop end-to-end automated validation frameworks to stress-test the product.
  • Cluster Management: Work on extending Kubernetes for managing specialized hardware, scheduling, and networking requirements unique to HPC and AI workloads. 
  • Collaborate: Partner with the core platform, fabric, and hardware teams to ensure a cohesive and performant end-to-end solution. 
  • Upstream Contribution: Engage with the open-source community and contribute to relevant projects within the cloud-native ecosystem. 
  • Documentation and Best Practices: Author high-quality technical documentation and champion best practices for software development in a cloud-native environment. 
  • Leverage AI-powered tools to accelerate software development workflows, including intelligent code generation, refactoring, and performance optimization. 
  • Apply AI-driven techniques for automated code review, testing, and quality assurance to improve reliability and reduce development cycles. 

Minimum Qualifications: 

  • 5+ years of professional software development experience. 
  • Proven experience in designing and developing solutions for Kubernetes, including building custom operators/controllers using tools like the Operator SDK or Kubebuilder. 
  • Expert level proficiency in using Go for systems programming. Experience with C++ or Python is also valuable. 
  • Knowledge of Multithreaded Programming.
  • Deep understanding of Kubernetes architecture, including the control plane, networking (CNI), and storage (CSI) interfaces. 
  • Hands-on experience with container technologies such as Docker or containerd. 
  • Demonstrable experience in integrating existing software platforms or services with Kubernetes. 
  • Solid understanding of high-performance data center networking environment.
  • Bachelor's or Master’s degree in Computer Science, Computer Engineering, or a related technical field.

Preferred Qualifications: 

  • Experience with high-performance computing (HPC) or high-performance networking. 
  • Familiarity with performance-sensitive environments and low-latency application requirements. 
  • Experience with monitoring and observability stacks like Prometheus, Grafana, and Fluentd. 
  • Knowledge of CI/CD principles and experience building deployment pipelines. 
  • Contributions to open-source projects in the Kubernetes or cloud-native ecosystem. 

 

Job Location: This role may work remote from Costa Rica via an EOR (Employer of Record).

  

We offer a competitive compensation package that includes base salary, performance incentives and equity participation.

At Cornelis Networks, your compensation is only one component of your comprehensive total rewards package. Compensation will be determined by factors such as experience, qualifications, skills, and location relative to the hiring range for the position. 

Depending on your geographic location, in addition to your total rewards package, you may also be eligible for additional benefits, paid holidays, and flexible work arrangements.

Cornelis Networks does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. Cornelis Networks is an equal opportunity employer, and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity or expression, pregnancy, age, national origin, disability status, genetic information, protected veteran status, or any other characteristic protected by law.