NVIDIA
Posted 2w ago

Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements

NVIDIA
Santa Clara or Austin or California or Oregon
$272k-$431k/yrHybridFull Time
Responsibilities
  • driving workstreams
  • synthesizing feedback
  • orchestrating updates
Requirements
  • Requires 15+ years in GPU system software
  • Firmware, or accelerator platforms
  • BS/MS in computer science
  • Electrical engineering, or related field
  • Deep GPU architecture and firmware expertise
Technical tools mentioned
NVLinkVBIOSInfoROM

Job description

We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for GPU firmware and GPU system software, working directly with engineering teams of key CSP / hyperscale customers to ensure they can reliably manage, update, and operate NVIDIA GPU firmware at fleet scale. You will drive work streams with engineering teams of key CSPs/hyperscale customers to build shared understanding of GPU firmware and system software integration, incorporate their feedback into NVIDIA's feature roadmap and delivery plan, and ensure customer-side automation and recovery procedures are ready before each firmware release. Your cross-CSP visibility enables you to identify patterns in GPU firmware operational challenges that drive systemic improvements no single customer engagement could surface alone.

What you'll be doing:

  • Drive GPU firmware & siftware work streams with CSP engineering teams — ensuring they understand GPU firmware architecture (VBIOS, InfoROM, microcontroller firmware), update sequencing, recovery procedures, and GPU power management

  • Gather and synthesize CSP feedback on GPU firmware/software — covering manageability, observability, security requirements (e.g., multi-tenancy isolation, secure boot, attestation), and performance — and champion those priorities into NVIDIA's GPU firmware/software feature roadmap and delivery plan

  • Drive GPU firmware update orchestration for large-scale deployments — multi-GPU update sequencing, rollback strategy, failure handling, and validation across hundreds of GPUs per rack

  • Serve as the technical focal point between NVIDIA and CSP firmware/software engineering — ensuring GPU behaviors (error recovery flows, thermal protection, power state transitions) are well-documented and accessible for customer integration

  • Identify cross-CSP GPU SW/FW issue patterns — common update failures, recovery gaps, and configuration problems — and drive documentation, tooling, and test strategy improvements

What we need to see:

  • 15+ years of experience in GPU system software, GPU firmware, or accelerator platform engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience)

  • Deep understanding of GPU architecture internals: streaming multiprocessors, GEMM execution, compute kernels, memory hierarchy, and how firmware/driver decisions impact GPU compute performance

  • Understanding of multi-GPU fabric architectures (NVLink, or similar) and how firmware coordinates across multiple GPUs in a rack-scale system

  • Understanding of GPU firmware architecture: VBIOS, GPU microcontroller firmware, InfoROM, and their interaction with the GPU driver stack

  • Experience with firmware update lifecycle management at scale: multi-device update sequencing, A/B updates, rollback, staged rollout, emergency recovery

  • Understanding of GPU error handling and recovery flows — how firmware-level errors propagate through the driver stack to application-visible failures

  • Experience with GPU health monitoring and telemetry: Xid errors, thermal events, power events, ECC counters, and their significance for firmware/software teams

  • Customer obsession — genuine passion for simplifying GPU firmware integration for fleet-scale customers. Proven success influencing engineering teams to improve quality and fleet manageability

Ways to stand out from the crowd:

  • Direct experience with NVIDIA GPU VBIOS, GPU microcontroller firmware, or GPU driver internals

  • Background in GPU fleet management at 10K+ GPU scale — firmware rollout, health-based remediation, fleet-wide configuration management

  • Experience with GPU error taxonomy (Xid classification, NVLink error counters, ECC events) and building runbooks around GPU firmware behavior

  • Understanding of GPU security: secure boot chain, code signing, attestation, debug authentication, multi-tenancy isolation at the firmware level

  • Familiarity with GPU power management architecture and its impact on workload performance at fleet scale

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative, hardworking and self-motivated, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 16, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About NVIDIA

Computing platform for AI and accelerated graphics.

Similar jobs

Software Engineer roles near Santa Clara, California
7h
Save
Mark Applied
Hide
Software Engineer - Game of Thrones Slots
Austin or San Mateo or Chicago or Toronto
$71k-$127k/yr HybridFull Time
Zynga
Zynga: Mobile video game developer and publisher creating games for players worldwide as a Take-Two subsidiary.
1+ YOEB.Sc. in Computer Science or equivalent; 1–2 years software development experience, or 3+ years relevant engineering experience; Unity, C#, PHP backend, Git, algorithms, and data structures.
Unity, C#, PHP, Git
8h
Save
Mark Applied
Hide
Software Engineer, GTM
San Francisco, California, United States
$160k-$200k/yr OnsiteFull Time
Hyperbound
Hyperbound: Private AI sales-coaching platform helping enterprise revenue teams practice conversations and improve performance.
Build and own production systems for lead routing, pipeline, forecasting, scoring, deal workflows, and revenue data across CRM, product, and billing systems.
10h
Save
Mark Applied
Hide
Senior Software Engineer, AI Ecosystem 
Mountain View, California, United States
$160k-$205k/yr OnsiteFull Time
Aerospike
Aerospike: Aerospike is a private software providing a real-time database platform for enterprise AI and mission-critical applications.
5+ YOERequires 5+ years coding experience in Go, Python, Java, Scala, C, or C++, plus distributed systems integration, Kubernetes, Docker, and DevOps/SRE expertise. NoSQL, vector databases, and open-source experience are bonuses.
Go, Python, Java, Scala, C, C++, Kubernetes, Docker, NoSQL, GitHub, Stack Overflow
10h
Save
Mark Applied
Hide
Software Engineer
San Francisco, California, United States
$150k-$300k/yr OnsiteFull Time
Arini
Arini: AI receptionist and healthcare automation software provider serving dental practices and dental service organizations.
Design, deploy, and iterate AI agents and automation workflows, integrate production code with enterprise systems, and deliver measurable customer outcomes in a fast-moving environment.
EHR
1d
Save
Mark Applied
Hide
Software Engineer
Menlo Park or Seattle
$219k-$301k/yr OnsiteFull Time
Meta
MetaNASDAQ: META: Builds technologies that help people connect, find communities, and grow businesses.
10+ YOEBachelor's degree or equivalent experience, 10+ years in networking or infrastructure software, 4+ years designing production dataplane/control-plane systems, Kubernetes networking, C/C++, scripting, and test automation.
Kubernetes, C, C++, Python, Shell Scripting, DPDK, eBPF/XDP, AF_XDP, io_uring, RDMA/RoCEv2, Linux, TC/iptables/nftables
1d
Save
Mark Applied
Hide
Software Engineer Simulation Infrastructure
Cupertino or North America
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
Strong Python skills, HTTP service and relational database experience, frontend framework knowledge, cloud infrastructure experience, networking and debugging expertise, Linux/macOS background, and a CS degree or equivalent experience.
Python, FastAPI, Flask, Postgres, React, Kubernetes, TCP/IP, HTTP, WebSockets, SSH, Linux, macOS, Swift, C++, OIDC, OAuth
1d
Save
Mark Applied
Hide
Software Engineer
San Francisco, California, United States
OnsiteFull Time
Proximal
Proximal: Private San Francis research lab developing coding data and benchmarks for autonomous coding agents.
Experience designing scalable systems from scratch with strong correctness, reliability, and performance; experimental instincts, systems design judgment, and ability to solve ambiguous technical problems independently.
1d
Save
Mark Applied
Hide
Software Engineer, Simulation Systems
United States or San Francisco
$140k-$185k/yr RemoteFull Time
Aalyria
Aalyria: Private aerospace communications providing network-orchestration software and optical terminals to commercial and government customers.
5+ YOEBachelor's degree in computer science or related technical field or equivalent experience; 5+ years developing production software; proficiency in Golang, C/C++, or Java; distributed or complex backend systems experience.
Golang, C, C++, Java