NVIDIA
Posted 1mo ago

Senior Infrastructure Software Engineer, Deep Learning Libraries

NVIDIA
Santa Clara, California, United States
$152k-$288k/yrOnsiteFull Time
Responsibilities
  • designing software
  • building automation
  • maintaining deployments
Requirements
  • Masters in CS/CE (or equivalent)
  • 3+ years experience
  • Strong Python and C/C++ familiarity
  • CI/CD and Kubernetes experience
  • Web frontend and SCM/build system proficiency
Technical tools mentioned
cuDNNTensorRTCUDAKubernetesJenkinsDockerCMakeGitlabJiraHTMLCSSJavaScriptPythonC/C++GitPerforceMakeBazelGitHub ActionsGitLab pipelinesMicrosoft Azure DevOpsNodeJSReactGroovyUbuntuRedHatWindowsQNX

Job description

We are now looking for a Senior Infrastructure Software Engineer for Deep Learning Libraries!

NVIDIA's Deep Learning Libraries Group is seeking excellent software engineers to enable the next wave of NVIDIA’s highest performing deep learning libraries. The role spans multiple products, including cuDNN, TensorRT, and CUDA kernel libraries. The mission is to design and develop scalable, modular infrastructure that streamlines development, builds, and tests across NVIDIA’s diverse set of platforms, from Drive AGX for autonomous vehicles to DGX servers for datacenters and large language models. Join our technically diverse team of software engineers and infrastructure experts to design the systems that enable NVIDIA to stay ahead of the competition as we deliver the world's fastest deep learning platforms.

What you'll be doing:

  • Designing and developing software for testing and analysis of our codebases

  • Building scalable automation for build, test, integration, and release processes for publicly distributed deep learning libraries

  • Developing throughout the software stack, from the user experience and user interfaces down to the cluster and database layers

  • Configuring, maintaining, and building upon deployments of industry-standard tools (e.g. Kubernetes, Jenkins, Docker, CMake, Gitlab, Jira, etc.)

  • Develop front-end solutions using HTML, CSS, JavaScript, and related web technologies

  • Advancing the state of the art in those industry-standard tools

What we need to see:

  • A Masters Degree in Computer Science or Computer Engineering or equivalent experience.

  • 3+ years of relevant experience

  • Strong programming skills in Python (or similar) and familiarity with C/C++ development

  • Experience setting up, maintaining, and automating continuous integration systems (e.g. Jenkins, GitHub Actions, GitLab pipelines, Azure DevOps)

  • Experience in HTML5, CSS, NodeJS, or React

  • Fluency in SCM (e.g. Git, Perforce) and build systems (e.g. Make, CMake, Bazel)

  • Background with distributed systems and cluster/cloud computing, especially with Kubernetes

Ways to stand out from the crowd:

  • Prior experience designing and developing automation in Jenkins with Groovy (or similar)

  • Track record of identifying useful new technologies and incorporating them into SW development flows

  • A strong understanding of unit and integration test frameworks and experience with crafting them

  • Experience with mobile/embedded platforms and multiple operating systems (Ubuntu, RedHat, Windows, QNX, or similar)

This is an opportunity to have a wide impact at NVIDIA by improving development velocity across our many AI/DL/Compute Software projects. Are you creative, driven, and autonomous? Do you love a challenge? If so, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until June 28, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About NVIDIA

Designs graphics processing units and artificial intelligence hardware.

Similar jobs

Infrastructure Software Engineer roles near Santa Clara, California
12h
Save
Mark Applied
Hide
Infrastructure Software Engineer, Fleet & Automation
Houston or New York City or San Francisco or Seattle
OnsiteFull Time
Nscale
Nscale: Vertically integrated AI infrastructure provider for high-performance computing.
5+ YOEBachelor's degree or equivalent experience, 5+ years building large-scale infrastructure applications, and expertise in Python, Linux, networking, distributed systems, and infrastructure tooling.
C, C++, Java, Python, Linux, TCP/IP, BGP, Ansible, Terraform, DCIMs, NetBox, OpenStack, MAAS, Ironic, IPMI, NVIDIA GPUs, InfiniBand, NCCL, SLURM, Prometheus, Grafana, OpenTelemetry, Kubernetes, Docker
1w
Save
Mark Applied
Hide
Member of Technical Staff - SWE Infrastructure
Palo Alto, California, United States
OnsiteFull Time
Ricursive Intelligence
Ricursive Intelligence: Automates semiconductor chip design using recursive artificial intelligence.
5+ YOERequires 5+ years of software engineering experience, with expertise in production systems, CI/CD, regression testing, cloud infrastructure, distributed systems, and automated deployment.
CI/CD
1w
Save
Mark Applied
Hide
Senior Infrastructure Software Engineer, TensorRT Edge-LLM
Santa Clara or United States
$184k-$288k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
7+ YOEBachelor's degree or equivalent, 7+ years of experience, Python and modern C/C++ skills, CI systems, cloud platforms, Git or Perforce, and build systems. Experience with infrastructure, DevOps, and automation preferred.
Python, C, C++, CMake, GitLab, GitHub Actions, Kubernetes, Docker, Jenkins, AWS, GCP, Azure, Git, Perforce, Make, Bazel, TensorRT, TensorRT-LLM, vLLM, SGLang, Ubuntu, JetPack, QNX
2w
Save
Mark Applied
Hide
Staff Infrastructure Software Engineer
Sunnyvale, California, United States
$210k-$314k/yr OnsiteFull Time
Carbon
Carbon: Manufacturer of industrial 3D printers and advanced polymer materials.
7+ YOE7+ years building and operating production cloud infrastructure with deep AWS or GCP expertise, Terraform, Kubernetes, Istio, CI/CD (Jenkins/GitHub Actions), Linux, and a scripting language (Python/Go/Bash).
AWS, GCP, Terraform, Kubernetes, Istio, Jenkins, GitHub Actions, Linux, Python, Go, Bash, Envoy, Bazel
3w
Save
Mark Applied
Hide
Infrastructure Software Engineer
Campbell, California, United States
$180k-$230k/yr RemoteFull Time
Camus Energy
Camus Energy: Software platform for managing renewable energy grid integration.
3+ YOE3+ years software engineering with infrastructure focus; Python 3, Kubernetes, GCP, CI/CD, observability (Prometheus,Grafana); familiarity with SQL; collaborative incident response and reliability practices.
Python 3, Kubernetes, GCP, CI/CD pipelines, Prometheus, Grafana, Bazel, Go, Node.js, SQL
2mo
Save
Mark Applied
Hide
Constellation Software Engineer, Infrastructure
Redwood Shores or Palo Alto
OnsiteFull Time
WindBorne Systems
WindBorne Systems: Operates smart weather balloons to provide global atmospheric data.
Experience building/shipping full-stack applications (Ruby on Rails, Postgres), operating 24/7 uptime systems with on-call responsibility, and designing low-latency data pipelines and robust infrastructure for flight operations and data processing.
Ruby on Rails, Postgres
2mo
Save
Mark Applied
Hide
Staff Software Engineer, Infrastructure
San Francisco, California, United States
$200k-$300k/yr HybridFull Time
F2
F2: AI-driven financial analysis platform for private market investors
7+ YOE7+ years building and scaling cloud infrastructure on AWS/GCP; expertise with containers, IaC (Terraform/Pulumi/CloudFormation), CI/CD, Postgres/Redis, message systems (Temporal/SQS/Kafka), SOC 2 and security practices.
AWS, GCP, ECS/Fargate, Python, Node, Temporal, Terraform, Pulumi, CloudFormation, GitHub Actions, Postgres, RDS, Supabase, Redis, SQS, Kafka, Kubernetes, GitOps, IaC, IAM, VPC, LLM
2mo
Save
Mark Applied
Hide
Software Infrastructure Engineer (Starlink)
Palo Alto or Redmond
$135k-$185k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
1+ YOEBachelor's in CS/IT/engineering or equivalent experience; 1+ years SRE/DevOps experience; Linux, Terraform/Ansible, Docker/Kubernetes, Bash/Python/C/C++ experience; strong CI/CD, testing, and networking knowledge. Must meet ITAR eligibility.
Linux, Terraform, Ansible, Docker, Kubernetes, Bash, Python, C++, C, Bazel