This job has expired

This job posting is no longer active and is not accepting applications. Explore similar roles below!

NVIDIA
Posted 3w ago

Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA
Santa Clara or Texas or United States or Washington or California
$152k-$288k/yrHybridFull Time
Responsibilities
  • designing platform
  • diagnosing assets
  • ensuring reliability
Requirements
  • 5+ years building large-scale production distributed systems
  • Strong systems programming (Go
  • Python)
  • Cluster management and operational experience
  • BS in CS or equivalent
  • Strong communication skills
Technical tools mentioned
GoPythonKubernetesSlurmBase Command Manager

Job description

NVIDIA is hiring experienced software engineers to help scale up its AI Infrastructure. We expect you to have significant software engineering experience with cluster operations, operator development, node health monitoring and working with GPU resource scheduling. We encourage out-of-the-box problem solvers who can provide new ideas with strong execution bias. Expect to be constantly challenged, improving, and evolving for the better. You will help advance NVIDIA's capacity to build and deploy leading infrastructure solutions for a broad range of AI-based applications. If you're creative, passionate about GPUs, and love having fun, please apply today!

For two decades, we have pioneered visual computing, the art and science of computer graphics. With the invention of the GPU - the engine of modern visual computing - the field has expanded to encompass video games, movie production, product design, medical diagnosis and scientific research. Today, we stand at the beginning of the next era, the AI computing era, ignited by a new computing model, GPU deep learning.

What you will be doing:

  • You will be part of an DGX Cloud team responsible for production systems that enable large scalable GPU clusters to be used for a variety of AI workloads.

  • Designing and developing a massively distributed scalable platform which would be used to identify, diagnose and remediate non-performant GPU assets.

  • Working with teams across NVIDIA to ensure production AI clusters run reliability and consistently with maximum performance.  Evaluating system failures and improving services based on a well-defined incident management process.


What we need to see:

  • Direct experience in a software engineering role within a highly technical organization with demonstrable impact from your work.

  • Highly motivated with strong communication skills, you can work successfully with multi-functional teams, principles, and architects and coordinate effectively across organizational boundaries and geographies.

  • 5+ years in similar role and experience on large-scale production systems.  Experience with common software engineering principles, tools and techniques.

  • You possess a BS in Computer Science, Engineering, Physics, Mathematics or a comparable Degree or equivalent experience.

  • Technical knowledge, including a systems programming language (Go, Python) and a solid understanding of data structures and algorithms.


Ways to stand out from the crowd:

  • Technical competency in managing and automating large-scale distributed systems independent of cloud providers. Advanced hands-on experience and deep understanding of cluster management systems (Kubernetes, Slurm, Base Command Manager).

  • Prior experience in asynchronous workflows and/or event driven architecture.

  • Proven operational excellence in maintaining reliable and performant infrastructure.


NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you are creative and autonomous, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 1, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About NVIDIA

Designs graphics processing units and artificial intelligence hardware.

Similar jobs

Software Engineer roles near Santa Clara, California
5h
Save
Mark Applied
Hide
Software Engineer Intern (Fall 2026)
London or San Francisco
HybridFull Time, Internship
Cloudflare
CloudflareNYSE: NET: Provides security and performance services for internet properties.
0+ YOECurrently pursuing a relevant degree, with critical thinking, adaptability, curiosity, and software development passion. Must commit to a full-time 12-week minimum internship in London.
TypeScript, JavaScript, Go, Rust, C, C++, Python, Cloudflare for Students
5h
Save
Mark Applied
Hide
Software Engineer III, Ads Bidding
Mountain View, California, United States
$147k-$210k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
2+ YOEBachelor's degree or equivalent practical experience; 2 years of Python or C++ software development; 1 year testing, maintaining, or launching software products. Master's degree or PhD and data structures experience preferred.
Python, C++
5h
Save
Mark Applied
Hide
Staff+ Software Engineer, Claude Managed Agents
San Francisco or New York City or Seattle or California
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
8+ YOEMinimum 8 years of backend, distributed systems, or infrastructure engineering experience; production systems expertise, API design, product sense, and end-to-end ownership including on-call. Bachelor's degree or equivalent required.
Claude, AI, APIs, SDK, CLI, LLM
5h
Save
Mark Applied
Hide
Software Engineer (Ray Data)
San Francisco, California, United States
$215k-$230k/yr OnsiteFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
3+ YOERequires 3–4 years of relevant experience and a background in scalable, fault-tolerant distributed systems, data processing, database internals, and large-scale AI performance.
Ray, Python
5h
Save
Mark Applied
Hide
Principal Software Engineer
Redmond or United States or Washington or San Francisco or New York City
$143k-$275k/yr HybridFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
6+ YOEBachelor's degree in computer science or related field and 6+ years of coding experience required; preferred qualifications include advanced degrees, Rust, AKS, distributed systems, cloud infrastructure, and technical leadership.
Rust, C, C++, C#, Java, JavaScript, Python, Microsoft Fabric, Azure SQL DB, Azure Cosmos DB, Azure PostgreSQL, Azure Data Factory, Azure Synapse Analytics, Azure Service Bus, Azure Event Grid, Power BI, Azure Monitor, Log Analytics, Application Insights, Container Insights, Hosted Prometheus, Azure Managed Grafana, Microsoft Sentinel, AKS
18h
Save
Mark Applied
Hide
Senior Software Engineer - Marketplace Foundation
San Mateo, California, United States
$197k-$243k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
4+ YOERequires 4+ years of software development experience focused on distributed systems, systems programming, Docker, Kubernetes, Terraform, Helm, big data, distributed databases, monitoring, and technical leadership.
C#, Go, Rust, Java, C++, Python, Docker, Kubernetes, Terraform, Helm, Hadoop, Spark, Kafka, Cassandra, MongoDB, Grafana, Datadog, Prometheus, ELK stack
19h
Save
Mark Applied
Hide
Lead Software Engineer, Payment Platform
San Jose, California, United States
$400k-$500k/yr HybridFull Time
Roku
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
15+ YOERequires 15+ years in large-scale backend services and payment systems, strong architecture and cloud expertise, AI-assisted coding experience, payment-flow knowledge, and familiarity with SOX, PCI, and 3DS compliance.
Claude Code, Cursor, Microsoft Copilot, GitHub Copilot, AWS, GCP, CI/CD, SOX, PCI, 3DS
20h
Save
Mark Applied
Hide
Software Engineer Graduate (TikTok Search Data Infra) - 2027 Start
San Jose or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
$128k-$317k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's degree in computer science or related technical field in progress; strong Java, C++, or Python skills; production data framework experience; distributed systems fundamentals.
Java, C++, Python, Spark, Flink, Hive, Iceberg
This job has expired