NVIDIA
Posted 1mo ago

Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA
Santa Clara or Illinois or Washington or California or Massachusetts or United States
$168k-$334k/yrHybridFull Time
Responsibilities
  • designing platform
  • operating platform
  • supporting services
Requirements
  • 8+ years building or operating production Kubernetes platforms
  • Proficiency in Go or Python
  • GitOps, CI/CD
  • Production on-call and incident response, and experience with network infrastructure or distributed systems
Technical tools mentioned
KubernetesCluster API (CAPI)Metal3GoPythonGitOpsCI/CD

Job description

Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, and operate NVIDIA’s global network across data centers, colocation facilities, and cloud environments. The team owns the architecture and lifecycle of this platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments.

We are looking for a hands-on senior engineer to own the lifecycle and automation of the Kubernetes platform supporting GNI network systems. You will also provide production support for network services running on the platform, partnering with their engineering owners when issues or changes cross the platform boundary. You will take complex problems from design through production and remain accountable for the outcome. You will bring deep Kubernetes expertise and help establish consistent engineering practices across the US and Bangalore teams. This is a senior individual contributor role with end-to-end ownership and production responsibility.

What You’ll Be Doing:

  • Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.

  • Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.

  • Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps.

  • Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features.

  • Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified resolution.

  • Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery.

  • Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion.

What We Need to See:

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.

  • 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.

  • Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.

  • Proficiency in at least one general-purpose programming language, such as Go or Python.

  • Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.

  • Experience deploying and supporting network automation or telemetry services on Kubernetes.

  • Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion.

Ways to Stand Out From the Crowd:

  • Strong knowledge of IP routing, data center fabrics, and cloud networking is a great plus.

  • Experience designing and operating large, multi-region Kubernetes fleets, including fleet-wide upgrades and recovery.

  • Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades.

  • Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns. Experience designing or operating network automation and telemetry services on Kubernetes at global scale.

  • Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects.

NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these unique opportunities in deep learning cloud solutions. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 264,500 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until July 20, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About NVIDIA

Computing platform for AI and accelerated graphics.

Similar jobs

Platform Engineer roles near Santa Clara, California
10h
Save
Mark Applied
Hide
Staff Platform Engineer
San Francisco or New York City
OnsiteFull Time
Ciridae
Ciridae: N AI transformation firm that builds AI-native operating systems for services businesses.
8+ YOERequires 8+ years of software engineering experience, staff-level technical impact, strong programming skills, cloud-native platform expertise, and hands-on experience with cloud infrastructure, Kubernetes, Terraform, CI/CD, security, and observability.
Go, Python, TypeScript, AWS, Azure, GCP, Kubernetes, Terraform, CI/CD
5d
Save
Mark Applied
Hide
Lead Platform Engineer, Workday Functional (HCM)
McLean or Plano or New York City or San Francisco or Chicago or Richmond or San Jose or Cambridge
$150k-$205k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: A technology-driven bank providing diverse financial services.
6+ YOERequires 6 years of Workday functional configuration, 4 years configuring Workday business processes, and 2 years with calculated fields or custom reports. High school diploma or equivalent required; bachelor's preferred.
Workday, HCM, Recruiting, Compensation, Payroll, Benefits, Workday Security, Workday Reporting, Enterprise Interface Builders (EIB)
5d
Save
Mark Applied
Hide
Lead Platform Engineer, Workday Functional (HCM)
McLean or Plano or New York City or San Francisco or Chicago or Richmond or San Jose or Cambridge
$150k-$205k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: A technology-driven bank providing diverse financial services.
6+ YOERequires 6+ years of Workday functional configuration, 4+ years configuring Workday business processes, and 2+ years with calculated fields or custom reports. High school diploma or equivalent required; bachelor's preferred.
Workday, Human Capital Management (HCM), Enterprise Interface Builders (EIBs)
1w
Save
Mark Applied
Hide
Associate Platform Engineer (College Grad 2027)
Redwood City or United States
OnsiteFull Time
Solace Health
Solace Health: U.S. healthcare technology connecting patients and families with expert advocates for care navigation.
0+ YOENew graduate role requiring curiosity, troubleshooting ability, systems thinking, willingness to learn from failures, and U.S. residence. The position involves cloud infrastructure and Linux systems.
Linux, GCP, AWS, OpenTofu, Kubernetes
1w
Save
Mark Applied
Hide
Platform Engineer
San Francisco or Singapore or Europe
$150k-$250k/yr OnsiteFull Time
Clera
Clera: AI-powered talent agent matching candidates to startup roles.
2+ YOERequires 2–4 years owning production cloud infrastructure, hands-on AWS and containerized systems experience, CI/CD and observability expertise, backend engineering judgment, and production-quality coding skills.
Amazon Web Services (AWS), Terraform, Kubernetes, Amazon EKS, Helm, Docker, Amazon EC2, AWS CodeBuild, Amazon ECR, Amazon S3, AWS IAM, CI/CD
1w
Save
Mark Applied
Hide
IIB and MQ Platform Engineer
Minnesota or United States or Eden Prairie or District of Columbia or San Francisco
$92k-$164k/yr RemoteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Diversified health care helping people live healthier lives.
5+ YOERequires 5+ years of software or systems engineering experience, extensive IIB, ACE, and IBM MQ expertise, Java, Python, React, Ansible, Docker, Kubernetes, security protocols, and messaging infrastructure experience.
IBM App Connect Enterprise (ACE), IBM Integration Bus (IIB), XML, JSON, SOAP, Kafka, Java, Spring Boot, Python, React, Ansible, Docker, GitHub, GitHub Actions, Kubernetes, PingFederate SSO, OIDC, LDAP, CyberArk, HashiCorp Vault, IBM MQ, SSL/TLS, Grafana, Splunk, Prometheus, MySQL, DB2, OAuth, JWT, COBOL, CICS, API Gateway
1w
Save
Mark Applied
Hide
Platform Engineer 4
San Jose, California, United States
$178k-$258k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Empowering everyone to create through innovative digital experiences.
5+ YOERequires 5+ years in public cloud, Linux, and software engineering; Kubernetes, CI/CD, Git, cloud security, automation, architecture, and technical leadership experience; relevant bachelor's degree or equivalent required.
Azure, AWS, Linux, Python, Go, Kubernetes, Git, CI/CD, GitOps
1w
Save
Mark Applied
Hide
Platform Engineer
San Francisco, California, United States
$200k-$220k/yr OnsiteFull Time
Hyperbound
Hyperbound: Private AI sales-coaching platform helping enterprise revenue teams practice conversations and improve performance.
Deep API design and versioning experience, OAuth 2.0 and OpenID Connect fluency, experience integrating third-party APIs, end-to-end ownership, and production AWS and CI/CD experience preferred.
Slack, OAuth 2.0, OpenID Connect, AWS, CI/CD, SOC 2 Type II