NVIDIA
Posted 1mo ago

Senior DevOps Engineer, Platform Engineering

NVIDIA
Santa Clara, California, United States
$176k-$276k/yrOnsiteFull Time
Responsibilities
  • building pipelines
  • managing infrastructure
  • owning observability
Requirements
  • 6+ years industry experience
  • BS/MS or equivalent
  • Strong Python
  • Kubernetes and Helm expertise
  • CI/CD pipeline experience (Jenkins/GitHub/GitLab)
  • Linux administration
  • Observability (Prometheus/Grafana/ELK)
Technical tools mentioned
JenkinsGitHub ActionsGitLab ActionsKubernetesHelmPythonPrometheusGrafanaELKTerraformAnsibleDeepStreamNVIDIA MetropolisAWSGCPAzure

Job description

At NVIDIA, our work is dedicated to a computing model passionate about visual and AI computing. For twenty years, NVIDIA has led the way in visual computing, the science and art of computer graphics, through our invention of the GPU. The GPU has proven extremely effective in solving complex computer science challenges. Today, NVIDIA's GPU powers deep learning algorithms, simulating human intelligence. It serves as the brain for computers, robots, and self-driving cars that perceive and interpret the world. We aim to expand our company and teams with the brightest minds globally, and now is an exciting time to join us!

NVIDIA invites applications for a Senior DevOps Platform Engineer skilled in Platform and Release Engineering to join the Metropolis team. The role involves developing, building, and maintaining foundational infrastructure and CI/CD systems that run AI/Machine Learning video analytics workloads at scale using NVIDIA Data Center GPUs. You will foster engineering rigor by setting up reliable release workflows, automation systems, and developer tools to improve efficiency on the Metropolis platform.

What you'll be doing:

  • Compose, build, and maintain scalable CI/CD pipelines using Jenkins, GitHub/GitLab Actions and Runners for Metropolis software products.

  • Develop and manage Kubernetes-based platform infrastructure supporting AI/ML workloads on NVIDIA Data Center GPUs.

  • Build and implement scaling and performance measurement frameworks within Kubernetes to ensure platform reliability and efficiency under AI/ML workload demands.

  • Define and implement release engineering processes, branching strategies, versioning standards, and gating criteria.

  • Drive developer efficiency by building and maintaining DevOps MCP servers, tooling, and automation frameworks.

  • Own observability and monitoring infrastructure using Prometheus, Grafana, and log aggregation pipelines.

  • Troubleshoot hardware and operating system issues across BareMetal and GPU-accelerated servers to minimize downtime and maintain platform stability.

What we need to see:

  • BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience, with over 6+ years of relevant industry background.

  • Advanced skills in Python for scripting, tooling, and automation.

  • Deep expertise with Kubernetes, Helm, and container orchestration in production environments.

  • Verified background in building and maintaining CI/CD pipelines at scale (Jenkins, GitHub/GitLab Actions and Runners, or similar).

  • Solid understanding of Linux systems administration, networking, and distributed systems.

  • Experience with release engineering practices including semantic versioning, release gating, and change management.

  • Hands-on experience with observability stacks (Prometheus, Grafana, ELK, or similar).

Ways to stand out from the crowd:

  • Experience with GPU infrastructure and AI/ML platform engineering at scale.

  • Background in BareMetal and hybrid cloud (AWS, GCP, Azure) environment management.

  • Familiarity with NVIDIA Metropolis, DeepStream, or similar AI video analytics platforms.

  • Experience with GitOps workflows, Infrastructure as Code (Terraform, Ansible).

  • Track record of driving DevOps culture transformation and developer experience improvements.

NVIDIA is often viewed as one of the most attractive companies to work for in the technology world. We have some of the most progressive and committed professionals in the field on our team. If you are inventive and independent, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD - 276,000 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 17, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About NVIDIA

Designs graphics processing units and artificial intelligence hardware.

Similar jobs

DevOps Engineer roles near Santa Clara, California
2d
Save
Mark Applied
Hide
Senior DevOps Engineer
San Francisco or United States
HybridFull Time
Scanner
Scanner: Cloud-native security data lake for fast threat detection.
5+ YOE5+ years in infrastructure, DevOps, SRE, or platform engineering; deep AWS and infrastructure-as-code experience; CI/CD ownership; security, observability, incident response, and TypeScript proficiency.
AWS, IAM, VPC, ECS, Lambda, S3, Pulumi, Terraform, CDK, TypeScript, Rust, MySQL, Amazon Elastic Container Service (ECS), Svelte
2d
Save
Mark Applied
Hide
DevOps Engineer
Palo Alto, California, United States
$162k-$229k/yr OnsiteFull Time
SAP
SAPFrankfurt Stock Exchange: SAP: Sells enterprise resource planning and business management software solutions.
3+ YOEBachelor's degree plus 5 years or master's degree plus 3 years in a related field; 3 years with CI/CD, Kubernetes, cloud, monitoring, Python, SAP technologies, and testing; 2 years with data applications and microservices.
Jenkins, ArgoCD, GitOps, Helm, Terraform, Kubernetes, Hashi Corp, Google, Azure, AWS, Splunk, Cloud Logging Search (CLS), Dynatrace, Python, SAP BTP, SAP AI Core, HANA Cloud, HANA Spark, HANA Data Lake Filesystem, Java, REST APIs
2d
Save
Mark Applied
Hide
DevOps Engineer
Epalinges or Lausanne or Menlo Park or Boston or Europe
HybridFull Time
Atinary Technologies
Atinary Technologies: AI platform for autonomous materials discovery and R&D optimization.
3+ YOERequires 3+ years in DevOps, site reliability, or cloud infrastructure, or 2+ software engineering and 1+ DevOps years; AWS, CI/CD, containers, Python, Bash, and Infrastructure as Code experience required.
Python, AWS, GitHub Actions, Bash, Terraform, OpenTofu, Infrastructure as Code (IaC)
4d
Save
Mark Applied
Hide
macOS DevOps / Founding Engineer
San Francisco, California, United States
$150k-$200k/yr OnsiteFull Time
Clera
Clera: AI talent agent matching professionals with high-growth startup roles
Requires macOS DevOps experience, strong Swift and bash/zsh scripting skills, CI/CD, monitoring, deployment, macOS administration at scale, networking, Linux, debugging, and systems reliability expertise.
Swift, bash, zsh, CI/CD, Jamf, Rust, Objective-C, Linux
5d
Save
Mark Applied
Hide
Storage & Database DevOps Engineer - rednote
Singapore or Palo Alto
OnsiteFull Time
Rednote
Rednote: A lifestyle-focused social media and e-commerce discovery platform.
Strong storage/database architecture experience; proficiency in Python, Go, or Java; cross-functional project leadership; strong communication, problem-solving, and English and Chinese fluency.
Redis, MySQL, TiDB, MongoDB, S3, Nebula, Python, Go, Java
5d
Save
Mark Applied
Hide
Principal DevOps Engineer
San Jose, California, United States
$147k-$339k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
15+ YOE15+ years in DevOps or SRE for large-scale production systems; expertise in Linux, networking, distributed systems, cloud infrastructure, containerization, CI/CD, reliability, and systems design.
Linux, Terraform, Ansible, Kubernetes, Docker, AWS, OCI, WebRTC, TCP/IP, DNS, CI/CD, SLIs, SLOs, BrightHire
5d
Save
Mark Applied
Hide
Staff Devops Engineer
San Francisco, California, United States
RemoteFull Time
Way2B1
Way2B1: Software platform for family offices and wealth management
Deep AWS production experience, expert Terraform, containers and distributed systems, incident ownership, Linux/networking, Bash or Python scripting, CI/CD, security, and IAM/secrets management.
AWS, Terraform, Aurora PostgreSQL, ECS, Fargate, ECR, Kafka, Auto Scaling Groups, EC2, Buildkite, Vault, AWS Secrets Manager, IAM, Cloudflare, WAF, CDN, OpenTelemetry, LangSmith, LangFuse, Claude Code, Cursor, Linux, Bash, Python
5d
Save
Mark Applied
Hide
DevOps Engineer
Palo Alto, California, United States
OnsiteFull Time
Mind Robotics
Mind Robotics: A robotics developing physical AI systems for industrial deployment.
4+ YOE4+ years in DevOps, platform, SRE, or infrastructure engineering; Linux, cloud, Kubernetes, Docker, Terraform, CI/CD, automation, programming, observability, networking, security, and distributed systems experience.
AWS, GCP, Azure, Terraform, Kubernetes, Docker, GitHub Actions, GitLab CI, Jenkins, Buildkite, Python, Go, Bash, Prometheus, Grafana, Datadog, ELK, OpenTelemetry, Linux, NVIDIA GPUs, CUDA