NVIDIA
Posted 2w ago

Senior Software Engineer - NVLink Rack Scale Stability and Reliability

NVIDIA
Santa Clara or Colorado or Illinois or Arizona or California or Massachusetts
$152k-$288k/yrHybridFull Time
Responsibilities
  • driving bringup
  • developing diagnostics
  • triaging issues
Requirements
  • BS/MS in a relevant engineering or computer science field or equivalent experience
  • 5+ years in system software
  • Firmware
  • Networking
  • Infrastructure, or distributed systems
  • C/C++, Python, debugging, and networking expertise
Technical tools mentioned
CC++PythonBashShellTCP/IPEthernetInfiniBandRDMARoCENVLinkNVSwitchCUDAPCIeDMACI/CD

Job description

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.

We are looking for highly motivated Senior Software Engineers to join our Fabric Networking team with a targeted focus on NVLink Rack-Scale Systems Stability & Reliability. In this role, you will partner closely with architects and developers building our next-generation NVLink and NVSwitch systems, helping transform first-of-their-kind platforms into stable, reliable, and volume production-ready systems. You will work on complex system-level challenges spanning resiliency, diagnostics, recovery, and large-scale AI infrastructure, contributing directly to the software foundation powering next-generation datacenter deployments. 

What you will be doing:

  • Drive platform bringup, feature enablement, end-to-end software validation, and debug for next-generation NVLink-based GPU and rack-scale systems.

  • Develop tools, diagnostics, automation, and infrastructure for system validation, regression testing, and fleet support.

  • Lead reliability and MTBI validation through stress testing, telemetry analysis, failure injection, and issue resolution.

  • Triage complex software, firmware, networking, and platform issues across validation, deployment, and production environments.

  • Collaborate with architecture, hardware, firmware, software, and Customer engagement teams to improve system quality and reliability.

  • Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and operational readiness.

  • Create automation, dashboards, runbooks, and debug workflows that improve root-cause analysis and operational efficiency.

What we need to see:

  • BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.

  • 5+ years of experience in system software, firmware, networking, platform enablement, data center infrastructure, or distributed systems.

  • Strong programming skills in C/C++ and Python; Bash/Shell scripting experience  is a plus.

  • Strong system-level debugging across software, firmware, hardware, and networking layers.

  • Solid networking fundamentals, including TCP/IP, Ethernet and/or InfiniBand, RDMA/RoCE, routing, switching, and fabric performance analysis.

  • Experience with large-scale AI systems, including platform bringup, validation, reliability engineering, stress testing, telemetry analysis, and root-cause debugging.

  • Ability to triage complex multi-domain issues using logs, telemetry, experiments, and structured debugging methods.

  • Strong communication and collaboration skills across engineering, customer, and operations teams. Passion for building reliable next-generation AI infrastructure and solving complex system-level challenges at scale.

Ways to stand out from the crowd:

  • Experience with NVIDIA GPU systems, NVLink, NVSwitch, CUDA, and large-scale AI/HPC clusters such as NVIDIA GB200 NVL72.

  • Strong understanding of large-scale AI system architecture, including PCIe, memory hierarchy, DMA, high-speed interconnects, and distributed training/inference systems.

  • Experience with server management technologies, data center operations, cluster provisioning, scaling, and fleet monitoring.

  • Proven experience building diagnostics, automation, CI/CD pipelines, dashboards, and reliability tooling.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 21, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About NVIDIA

Computing platform for AI and accelerated graphics.

Similar jobs

Software Engineer roles near Santa Clara, California
1h
Save
Mark Applied
Hide
Staff Software Engineer, Android Agent
Mountain View, California, United States
$207k-$300k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOG, GOOGL: Global technology specializing in internet-related services and products.
8+ YOE3+ MgmtBachelor’s degree or equivalent experience, 8 years of software development, 5 years in machine learning design and infrastructure, software testing and launches, and 3 years in software architecture.
Large Language Model (LLM)
1h
Save
Mark Applied
Hide
Software Engineer: Intern Opportunities for University Students - CoreAI - Boston, Massachusetts
Boston or California or San Francisco or New York City or New York
$6k-$11k/mo HybridInternship
Microsoft
MicrosoftNASDAQ: MSFT: Multinational technology providing software, cloud, and AI solutions.
Currently pursuing a bachelor's or master's degree in computer science or a related technical field with one semester remaining; programming, AI/ML, data structures, and algorithms experience preferred.
TensorFlow, PyTorch, scikit-learn, GitHub Copilot, Claude Code, Microsoft Roo Code
2h
Save
Mark Applied
Hide
Senior Staff Software Engineer – Test Automation Infrastructure (ZIA Core)
San Jose, California, United States
$158k-$225k/yr HybridFull Time
Zscaler
ZscalerNASDAQ: ZS: Cloud-native Zero Trust cybersecurity platform for digital transformation.
12+ YOERequires 12+ years of experience, strong Python programming and test automation framework development, cloud security testing, networking expertise, and advanced debugging skills.
Python, AI/ML, TCP/IP, HTTP(S), TLS, PKI, DNS, DHCP, VPN, HA, IPsec/GRE
5h
Save
Mark Applied
Hide
Founding Software Engineer
San Francisco, California, United States
OnsiteFull Time
Andreessen Horowitz
Andreessen Horowitz: Private venture capital firm investing in startups from seed through growth stages and helping founders build companies.
System design, rapid development with AI coding tools, detailed problem-solving, and ownership of scalable systems; research track requires deep learning, open-model fine-tuning, and ideally publications.
AI, GPU
8h
Save
Mark Applied
Hide
Software Engineer / Sr. Software Engineer, Planning Selection Autonomy
San Jose or Mountain View
$129k-$215k/yr OnsiteFull Time
DiDi Autonomous Driving
DiDi Autonomous Driving: Chinese autonomous-driving developing Level 4 self-driving technology, robotaxis, and autonomous trucking logistics for mobility-fleet applications.
Bachelor's or master's degree in computer science, robotics, or related field; autonomous systems experience; strong C++ and Python proficiency; robotics and motion planning knowledge.
C++, Python
8h
Save
Mark Applied
Hide
Senior Software Engineer - Platform Apps & Technologies
Cupertino or California
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
5+ YOE5+ years' software engineering experience deploying large-scale systems; expertise in distributed, data-centric, reliable cloud systems; applied machine learning or data engineering; bachelor's degree in computer science or related experience.
8h
Save
Mark Applied
Hide
Senior Software Engineer
San Mateo, California, United States
$160k-$180k/yr HybridFull Time
Vynca
Vynca: Private health technology and clinical-services providing palliative care, care navigation, and serious-illness software.
5+ YOEBachelor's degree in computer science or related field, 5–7 years of software engineering experience, and proficiency in Python, Java, distributed systems, event-driven architectures, healthcare data standards, databases, and cloud platforms.
Python, Java, Apache Kafka, Airflow, Flink, DataDog, Kibana, Kafka Streams, Spark Streaming, AWS, Kubernetes, Flask, FastAPI, Spring Boot, Kinesis, SQL, NoSQL, DynamoDB, Redshift, BigQuery, Snowflake, FHIR, CCDA, HL7 V2/V3, DICOM, HIPAA, HITRUST
8h
Save
Mark Applied
Hide
Software Enginee , Data & AI Platform
Menlo Park, California, United States
HybridFull Time
Axiamatic
Axiamatic: Agentic control plane for enterprise transformation.
2+ YOERequires a BS, MS, or PhD in computer science or related field, 2–4 years of backend production experience, strong Python, LLM feature delivery, data structuring, and public cloud experience.
Large Language Models (LLM), Python, Java, AWS