NVIDIA
Posted 2mo ago

Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements

NVIDIA
Santa Clara or Austin or Oregon or California or United States
$272k-$431k/yrOnsiteFull Time
Responsibilities
  • driving workstreams
  • orchestrating updates
  • identifying patterns
Requirements
  • 15+ years in GPU system software
  • Firmware, or accelerator platform engineering
  • Deep GPU architecture and firmware expertise
  • Experience with multi-GPU fabrics
  • Firmware update lifecycle at scale
  • Telemetry/monitoring, and customer-facing integration work
Technical tools mentioned
NVLinkVBIOSInfoROM

Job description

We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for GPU firmware and GPU system software, working directly with engineering teams of key CSP / hyperscale customers to ensure they can reliably manage, update, and operate NVIDIA GPU firmware at fleet scale. You will drive work streams with engineering teams of key CSPs/hyperscale customers to build shared understanding of GPU firmware and system software integration, incorporate their feedback into NVIDIA's feature roadmap and delivery plan, and ensure customer-side automation and recovery procedures are ready before each firmware release. Your cross-CSP visibility enables you to identify patterns in GPU firmware operational challenges that drive systemic improvements no single customer engagement could surface alone.

What you'll be doing:

  • Drive GPU firmware & siftware work streams with CSP engineering teams — ensuring they understand GPU firmware architecture (VBIOS, InfoROM, microcontroller firmware), update sequencing, recovery procedures, and GPU power management

  • Gather and synthesize CSP feedback on GPU firmware/software — covering manageability, observability, security requirements (e.g., multi-tenancy isolation, secure boot, attestation), and performance — and champion those priorities into NVIDIA's GPU firmware/software feature roadmap and delivery plan

  • Drive GPU firmware update orchestration for large-scale deployments — multi-GPU update sequencing, rollback strategy, failure handling, and validation across hundreds of GPUs per rack

  • Serve as the technical focal point between NVIDIA and CSP firmware/software engineering — ensuring GPU behaviors (error recovery flows, thermal protection, power state transitions) are well-documented and accessible for customer integration

  • Identify cross-CSP GPU SW/FW issue patterns — common update failures, recovery gaps, and configuration problems — and drive documentation, tooling, and test strategy improvements

What we need to see:

  • 15+ years of experience in GPU system software, GPU firmware, or accelerator platform engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience)

  • Deep understanding of GPU architecture internals: streaming multiprocessors, GEMM execution, compute kernels, memory hierarchy, and how firmware/driver decisions impact GPU compute performance

  • Understanding of multi-GPU fabric architectures (NVLink, or similar) and how firmware coordinates across multiple GPUs in a rack-scale system

  • Understanding of GPU firmware architecture: VBIOS, GPU microcontroller firmware, InfoROM, and their interaction with the GPU driver stack

  • Experience with firmware update lifecycle management at scale: multi-device update sequencing, A/B updates, rollback, staged rollout, emergency recovery

  • Understanding of GPU error handling and recovery flows — how firmware-level errors propagate through the driver stack to application-visible failures

  • Experience with GPU health monitoring and telemetry: Xid errors, thermal events, power events, ECC counters, and their significance for firmware/software teams

  • Customer obsession — genuine passion for simplifying GPU firmware integration for fleet-scale customers. Proven success influencing engineering teams to improve quality and fleet manageability

Ways to stand out from the crowd:

  • Direct experience with NVIDIA GPU VBIOS, GPU microcontroller firmware, or GPU driver internals

  • Background in GPU fleet management at 10K+ GPU scale — firmware rollout, health-based remediation, fleet-wide configuration management

  • Experience with GPU error taxonomy (Xid classification, NVLink error counters, ECC events) and building runbooks around GPU firmware behavior

  • Understanding of GPU security: secure boot chain, code signing, attestation, debug authentication, multi-tenancy isolation at the firmware level

  • Familiarity with GPU power management architecture and its impact on workload performance at fleet scale

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative, hardworking and self-motivated, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until June 30, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About NVIDIA

Computing platform for AI and accelerated graphics.

Similar jobs

Software Engineer roles near Santa Clara, California
2h
Save
Mark Applied
Hide
Senior Software Engineer (Python), Mortgage
San Francisco or Denver or Nashville or Santiago
$185k-$218k/yr HybridFull Time
Truework
Truework: AI verification platform providing background, identity, employment, mortgage, tenant, and trust checks for businesses and individuals.
5+ YOE5+ years of software development experience, bachelor's degree or equivalent, Python and JavaScript proficiency, RESTful API development experience, strong communication, documentation, and ownership skills.
Python, JavaScript, RESTful APIs
2h
Save
Mark Applied
Hide
Software Engineer III
Pune or San Francisco or Atlanta or Boston or Chicago or Dubuque or Dallas
OnsiteFull Time
OpenGov
OpenGov: Private American software providing cloud-based finance, permitting, procurement, and public-works software to local and state governments.
6+ YOEBA/BS in computer science or equivalent experience; 6–9 years developing software; proficiency in Python, JavaScript/TypeScript, ReactJS, NodeJS, GraphQL, data structures, databases, algorithms, and observability.
Python, JavaScript, TypeScript, ReactJS, GraphQL, NodeJS, Effect, Drizzle, Kafka, Postgres, RDBMS
3h
Save
Mark Applied
Hide
Sr. Software Engineer
Sunnyvale or Atlanta or Fort Worth or Denver or Huntsville
$150k-$200k/yr OnsiteFull Time
Lynx Software Technologies
Lynx Software Technologies: Lynx, a privately held software, serves aerospace, defense, industrial, and critical-infrastructure customers with secure edge platforms.
5+ YOEBachelor's degree in STEM, 5+ years of C/C++ development, embedded software, digital simulation, integration and testing experience, Linux and Windows proficiency, and active Secret clearance.
C, C++, GitLab, Linux, Microsoft Windows, Atlassian Tools, Confluence, JIRA, Bitbucket, Python, Cameo, AADL
3h
Save
Mark Applied
Hide
Senior Software Engineer
Denver or Menlo Park or United States or Colorado
$195k/yr RemoteFull Time
LeoLabs
LeoLabs: American aerospace providing radar-based orbital intelligence and space situational awareness to government and commercial operators.
5+ YOEActive US TS/SCI clearance, 5+ years of software engineering experience, cloud-scale distributed systems expertise, Python or Go proficiency, databases, message brokers, CI/CD, and on-call availability.
Python, Go, AWS, GCP, Azure, Postgres, MySQL, Kafka, SQS, CI/CD, TypeScript, JavaScript, Kubernetes
3h
Save
Mark Applied
Hide
Staff Software Engineer, Systems Infrastructure - Agent Evaluation
Mountain View, California, United States
$175k-$287k/yr HybridFull Time
LinkedIn: The world's largest professional network.
4+ YOEBachelor's degree or equivalent practical experience; 4+ years building deep learning systems and coding in Java, C++, Python, Go, Rust, C#, Scala, or other relevant languages; distributed systems and production AI agent experience.
Java, C++, Python, Go, Rust, C#, Scala, MLFlow, Kubeflow, Flink, Beam, Spark, PyTorch, TensorFlow, JAX, FLAX
3h
Save
Mark Applied
Hide
Software Engineer, Full Stack
Seattle or San Francisco or New York City or Washington or Raleigh or London or Amsterdam or United States or Canada or United Kingdom or Europe
$176k-$244k/yr OnsiteFull Time
Plaid
Plaid: Fintech data network connecting consumers’ financial accounts to apps and services.
2+ YOERequires 2–4 years of full-stack development experience, HTML, CSS, JavaScript, modern frameworks, a programming language, relational databases, microservices, API design, coding, and testing skills.
HTML, CSS, JavaScript, Python, Java, Go, Node.js, MySQL
4h
Save
Mark Applied
Hide
Software Engineer
San Francisco or San Mateo
HybridFull Time
MintMCP
MintMCP: Private B2B agent-governance platform that gives enterprises governed access to data and tools.
Track record of shipping products or tools to real users; strong fundamentals, fast learning, product judgment, infrastructure expertise, and comfort across TypeScript, Python, and CockroachDB.
Claude, Codex, Slack, GitHub, Jira, Snowflake, TypeScript, Python, CockroachDB
4h
Save
Mark Applied
Hide
Senior Software Engineer, Autonomous Pilot Integration (R5423)
San Mateo, California, United States
FieldFull Time
Shield AI
Shield AI: Developer of AI-powered autonomous systems for defense and aerospace.
2+ YOEBachelor's or master's degree or equivalent experience; 2+ years related experience, C++ and Python, Linux, embedded systems, RTOS, sensor integration, simulation tools, and ability to obtain SECRET clearance.
C++, Python, real-time operating systems (RTOS), Linux, shell scripting, ROS, DDS, AFSIM, NGTS, UCI, OMS Standards