Senior Manager, Kubernetes Runtime Engineering

NVIDIA

US, CA, Santa ClaraFull-timePosted 1 day ago

Get more Mechanical Engineering openings

A short daily email when similar roles appear in Santa Clara. No account needed.

The NVIDIA Kubernetes Engine (NKE) team is looking for a technical leader to lead the Runtime Engineering team responsible for the full configuration lifecycle of NKE tenant workload clusters. This team is responsible for software components that keep GPU workloads reliable and secure at scale. Their scope includes cluster bootstrapping, node configuration, and the container execution environment, including NVIDIA's AI Container Runtime (AICR). You will work across networking, storage, GPU resource management, and cluster security to deliver a production-grade, multi-tenant Kubernetes platform. Your team's decisions directly shape the runtime foundation that internal and external customers depend on.

What You'll Be Doing:

  • Be responsible for the build, implementation, and operational reliability of cluster configurations for NKE tenant workloads across all supported topologies
  • Manage a team of engineers coordinating the entire container runtime stack: AICR, GPU management operator, DCGM, and related node-level components
  • Drive architecture decisions for cluster networking (CNI), storage (CSI), cluster HA , and GPU resource partitioning (MIG, MPS, time-slicing)
  • Define and implement cluster hardening standards, RBAC models, pod security policies, and multi-tenancy isolation boundaries
  • Partner with NKE platform, infrastructure, and cybersecurity teams to integrate new capabilities and resolve cross-cutting runtime concerns
  • Build and maintain tooling for AICR lifecycle management β€” provisioning, upgrades, configuration drift detection, and remediation
  • Represent the runtime team in architecture reviews, roadmap planning, and customer communications with NVIDIA leadership
  • Contribute to open source communities anywhere NKE has upstream dependencies or influence

What We Need to See:

  • BS/MS degree in Computer Science or related field (or equivalent experience)
  • 12+ overall years of relevant experience designing and delivering large-scale distributed software systems, including 5+ years of people-management experience leading, developing, and scaling high-performing software engineering teams responsible for complex, production-critical software.
  • Experience leading a group of engineers with varying specializations and seniority levels β€” bridging runtime, networking, and security fields is a core part of this role
  • Kubernetes internals knowledge β€” not just usage; you understand how the scheduler, kubelet, API server, and admission controllers interact
  • Cluster lifecycle management experience β€” Cluster API, kubeadm, or equivalent; experience leading fleet-scale cluster provisioning and upgrades
  • Security and compliance posture β€” CIS Kubernetes Benchmark, pod security admission, image signing, supply chain integrity
  • Proven ability to design and implement maintainable APIs for consumers
  • Familiarity with Identity and Access Management approaches
  • Excel in managing up, down, and across organizations
  • Demonstrated ability to reach cross-organization consensus without all the details

Ways to Stand Out from the crowd:

  • Prior experience with NVIDIA GPU Operator, DCGM Exporter, or NVLink-aware scheduling
  • Experience running Kubernetes at hyperscale with GPU node pools
  • Track record of upstream open source contributions in the Kubernetes or any open source runtime ecosystem
  • Experienced, persuasive, and effective interpersonal skills β€” written, verbal, and in front of engineering leadership
  • Demonstrated skills in coaching, analysis, problem solving, and short/long-term technical planning

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction β€” from artificial intelligence to autonomous vehicles. NVIDIA is widely considered one of the technology world's most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. If you're passionate about building the infrastructure that runs AI at scale, we want to hear from you.

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 3, 2026.

This posting is for an existing vacancy.Β 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

The average job posting receives 250 applications.

Stand out by tailoring your resume to this specific role. Our AI resume builder highlights the skills and experience that matter most to this employer.

Frequently asked questions

Who is hiring for Senior Manager, Kubernetes Runtime Engineering at NVIDIA?+
NVIDIA is actively hiring for this Senior Manager, Kubernetes Runtime Engineering role. Click "Apply Now" to submit your application directly on NVIDIA's careers page β€” Careeronaut doesn't charge employers or candidates for referrals.
When was this Senior Manager, Kubernetes Runtime Engineering role posted?+
This listing was first posted on 2026-09-29. We pull the latest copy from the source feed daily, and any role that's taken down gets removed from Careeronaut within seven days so you don't waste time on stale listings.
How should I apply to this Senior Manager, Kubernetes Runtime Engineering role?+
Start by tailoring your resume to the posting β€” most applicants send generic CVs and the first filter recruiters use is keyword relevance. Careeronaut's AI does this automatically: paste the job description, get a matched resume in under a minute, and download as PDF or DOCX.
Where can I find more Mechanical Engineering jobs in Santa Clara?+
Browse all open mechanical engineering roles in Santa Clara on the listing pages linked below. You can filter by salary, remote-friendly, and posting date.