NVIDIA
Posted 3mo ago

Principal Software Engineer - Rack Scale Systems Infrastructure

NVIDIA
Santa Clara, California, United States
$272k-$431k/yrOnsiteFull Time
Responsibilities
  • Define architecture
  • Use Kubernetes
  • Bridge hardware
Requirements
  • BS or MS in Computer Science/Engineering
  • 15+ years in systems software
  • Distributed systems, or infrastructure engineering
  • Go/C++, or Rust
  • Kubernetes
  • Data center networking
  • Open source
  • Security
  • Mentoring
Technical tools mentioned
GoC++RustKubernetesLinuxRedfishIPMI

Job description

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.

What You Will Be Doing:

  • Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.

  • Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments.

  • Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization.

  • Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments.

  • Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development.

  • Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution.

What We Need To See:

  • BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.

  • Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs.

  • Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. Experience with Rust is highly valued.

  • Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.

  • Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems.

  • Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation.

  • Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations.

  • Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation.

  • Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.

Ways To Stand Out from the crowd:

  • Built software supporting multiple adoption models — internal services, CSP-integrated offerings, reusable libraries, and customer-extensible APIs. Strong Rust skills in systems, infrastructure, or hardware-adjacent software.

  • Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI-assisted engineering. Hands-on with fleet-scale provisioning, updates, rollback, observability, health, and remediation.

  • Led across the full data center product lifecycle: inception, pre- and post-silicon, manufacturing, deployment, and operations. Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs.

  • Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain. Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long-term direction.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until May 19, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About NVIDIA

Designs GPU-accelerated computing and artificial intelligence hardware.

Similar jobs

Principal Software Engineer roles near Santa Clara, California
2d
Save
Mark Applied
Hide
(USA) Principal, Software Engineer
Bellevue or Sunnyvale
$132k-$264k/yr OnsiteFull Time
Walmart
WalmartNYSE: WMT: Operates a chain of hypermarkets, discount stores, and grocery stores.
5+ YOEBachelor's degree plus 5 years of software engineering experience, or 7 years of experience. Preferred master's degree plus 3 years. Experience with cloud-native systems, AI/ML, distributed architecture, and technical leadership.
Python, Java, JavaScript, Dart, Rust, C++, GitHub Copilot, CI/CD, SDKs, model runtimes, feature stores
4w
Save
Mark Applied
Hide
Sr Principal Software Engineer (L7 Security - Data Path)
Santa Clara, California, United States
$170k-$277k/yr OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
7+ YOE7+ years in security or networking, proficient in C on Linux, multithreaded and high-performance design, TCP/IP experience, Redis/NoSQL and Go/Python preferred, BS in CS or equivalent required.
C, Linux, Go, Python, Redis, NoSQL, TCP/IP, PANOS
1mo
Save
Mark Applied
Hide
Principal Software Engineer
Santa Clara, California, United States
$221k-$387k/yr OnsiteFull Time
ServiceNow
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
15+ YOE15+ years engineering experience; strong system design and architecture for large-scale distributed systems; technical depth in Java or Python; experience with AI-powered products, observability, and enterprise platforms; mentoring and cross-team influence.
Java, Python, ServiceNow Platform
2mo
Save
Mark Applied
Hide
Senior Principal Software Engineer
Seattle or Santa Clara or United States
$135k-$306k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
10+ YOE10+ years building large-scale distributed systems; strong C/C++; Python or Java; cloud and high-throughput I/O experience; familiarity with RDMA, high-performance networking, AI/ML frameworks, and strong algorithms and troubleshooting skills.
C, C++, Python, Java, OCI, AWS, GCP, Azure, DAOS, SPDK, RMDA, SmartNICs, NVMe/TCP, RoCEv2, Tensorflow/Keras, PyTorch, Scikit-Learn, XGBoost, Caffe, MLOps, Kubernetes
2mo
Save
Mark Applied
Hide
Principal Software Engineer
San Francisco, California, United States
RemoteFull Time
DocuSign
DocuSignNASDAQ: DOCU: Provides electronic signature and agreement management software solutions.
15+ YOE15+ years in distributed systems/infra; Go or Python; Kubernetes and Terraform; incident commander; cloud migrations; OpenTelemetry; Azure/AKS; on-call leadership.
Go, Python, Kubernetes, Terraform, AKS, Azure DevOps, GitHub Actions, OpenTelemetry
2mo
Save
Mark Applied
Hide
Principal Software Engineer, Platform Security
San Francisco, California, United States
$197k-$314k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
10+ YOE10+ years in software development; strong in databases, APIs, data pipelines; experience with large projects and security infrastructure.
API development, REST APIs, SDLC, AWS, GCP, Azure, Python, Ruby, Unix/Linux, Security infrastructure
2mo
Save
Mark Applied
Hide
Principal Software Engineer, Data
Cambridge or San Francisco
$204k-$348k/yr HybridFull Time
Lila Sciences
Lila Sciences: Develops an AI platform for autonomous scientific research and discovery.
8+ YOE8-15 years of backend-focused engineering; full stack experience; AWS, Kubernetes; API design; data pipelines; Python; collaboration with scientists and engineers.
React, TypeScript, Nx, Tailwind, FastAPI, SQL/NoSQL, Python, Pydantic, SQLAlchemy, SQLModel, Django, Airflow, Prefect, Temporal, Dagster, pandas, numpy, scipy, jax, pytorch, AWS, Kubernetes, Terraform, CloudFormation, GitHub Actions
3mo
Save
Mark Applied
Hide
Principal Software Engineer (Platform Team)
Bastrop or Hawthorne or Palo Alto or Redmond or Starbase or Sunnyvale
OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
7+ YOEBachelor’s degree in CS/CE or related field and 7+ years building production platforms; or 9+ years experience without a degree.
Docker, Kubernetes, Cloud, Frontier Models, MLOps, Python