NVIDIA
Posted 1w ago

Senior Software Engineer, Core Infrastructure Services - DGX Cloud

NVIDIA
United States or Texas or Colorado or California or Massachusetts
$168k-$322k/yrRemoteFull Time
Responsibilities
  • building services
  • operating infrastructure
  • automating provisioning
Requirements
  • BS or equivalent with 8+ years experience
  • Strong Python and Go
  • Cloud-native microservices on Kubernetes
  • Infrastructure automation (Terraform
  • Ansible)
  • Workflow orchestration (Temporal)
  • Distributed systems and Linux fundamentals
Technical tools mentioned
PythonGoKubernetesFastAPIgRPCRESTTerraformAnsibleTemporalRedisKafkaNATSSQSDNSNTPRADIUSOAuthPrometheusGrafanaOpenTelemetrygNMIBGPInfiniBandRDMANetBoxNautobotLinux

Job description

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

NVIDIA is seeking an experienced software engineer to join the Cloud Foundations Automation team. Our team builds and operates the core infrastructure services that power NVIDIA's DGX Cloud and SuperPod deployments, delivering secure, reliable, and observable platforms at global scale.

What you'll be doing:

  • Build and operate core infrastructure services that power NVIDIA's global AI infrastructure.

  • Architect and develop secure, scalable and highly available cloud-native platform services.

  • Develop software that enables infrastructure orchestration, self-service workflows, and platform automation.

  • Own integrations with internal and external platforms to automate infrastructure provisioning and lifecycle management.

  • Build observability and security capabilities that improve the reliability and resilience of our infrastructure.

  • Partner with infrastructure and networking teams to deliver production services at scale.

  • Drive operational excellence through automation, monitoring, incident response, and continuous improvement.

What we need to see:

  • BS or equivalent experience with 8+ years of relevant industry experience.

  • Strong proficiency in Python and Go, with experience building production-quality software.

  • Experience building cloud-native microservices and APIs on Kubernetes using frameworks such as FastAPI, gRPC, or REST.

  • Experience with infrastructure automation (Terraform, Ansible), workflow orchestration (Temporal), and distributed systems using databases, Redis, and messaging platforms (Kafka, NATS, SQS).

  • Experience designing, building, and operating production infrastructure services such as DNS, NTP, AAA (RADIUS/OAuth), and observability platforms.

  • Strong Linux fundamentals with experience in observability (Prometheus, Grafana, OpenTelemetry, gNMI), networking (BGP, switching, routing, load balancing), and security (VPNs, firewalls, iptables/nftables).

  • Excellent problem-solving, communication, and collaboration skills.

Ways to stand out from the crowd:

  • Hands-on experience with network infrastructure including switches, routers, and firewalls.

  • Familiarity with InfiniBand, RDMA, and AI/HPC networking.Experience with NetBox, Nautobot, or similar network source of truth platforms.

  • Contributions to open-source software. Experience with public cloud platforms.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD for Level 4, and 200,000 USD - 322,000 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 8, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About NVIDIA

Designs graphics processing units and artificial intelligence hardware.

Similar jobs

Software Engineer roles in Texas
3h
Save
Mark Applied
Hide
Senior Software Engineer (Semiconductor Process and Device Simulation)
San Francisco or Santa Clara or Hsinchu or Sapporo or Tokyo or Seoul
$90k-$162k/yr HybridFull Time
Siemens
SiemensXETRA: SIE: Manufactures industrial automation, infrastructure, and energy technology systems.
5+ YOEPhD or MS with 5+ years of research or industry experience in computer science, applied mathematics, or engineering; expertise in scientific computing, high-performance computing, and C/C++, CUDA, and Python.
C, C++, CUDA, Python, Code Copilot, AI/ML, Message Passing Interface (MPI)
5h
Save
Mark Applied
Hide
Staff Software Engineer, Voice AI Platform
United States or Boston or New York City
$125k-$254k/yr HybridFull Time
Toast
ToastNYSE: TOST: Cloud-based technology platform for the restaurant and hospitality industry.
8+ YOE8+ years with Java or Kotlin, a relevant bachelor's degree, AI-integrated product development, distributed systems, scalable production services, and leadership of complex cross-team projects.
Java, Kotlin, Dropwizard, AWS, Amazon DynamoDB, Amazon RDS, AWS Lambda, PostgreSQL, React, ES6, Claude Code, Cursor, LLMs
6h
Save
Mark Applied
Hide
Staff Software Engineer, Data Engineering
London or San Francisco
HybridFull Time
Ripple
Ripple: Provides blockchain solutions for global payments and liquidity.
10+ YOE10+ years of data engineering experience; mastery of Databricks, Delta Live Tables, Unity Catalog, Delta Lake, and Spark; expertise in SQL, Python, and AWS; experience with AI data engineering tooling.
Databricks, Delta Live Tables, Unity Catalog, Delta Lake, Spark, SQL, Python, AWS, AI, LLMs
7h
Save
Mark Applied
Hide
Senior Software Engineer (Semiconductor Process and Device Simulation)
San Francisco or Hsinchu or Hokkaido or Tokyo or Seoul
$90k-$162k/yr HybridFull Time
Siemens Healthineers
Siemens HealthineersXetra: SHL: Global provider of medical technology and diagnostic imaging solutions.
5+ YOEPhD, or MS with 5+ years of research or industry experience, in computer science, applied mathematics, or engineering. Requires scientific computing, HPC, C/C++, CUDA, Python, and collaborative problem-solving expertise.
C, C++, CUDA, Python, Code Copilot, Message Passing Interface (MPI), AI, ML, Calibre, EDA, CAD, FEM, FD, FVM, BEM, Monte Carlo
9h
Save
Mark Applied
Hide
Staff Software Engineer - Video Performance - (Bay area only)
San Francisco, California, United States
$251k-$329k/yr OnsiteFull Time
Canva
Canva: Cloud-based visual communication and graphic design software platform.
Strong C++ proficiency; systems performance optimization, CPU/GPU architecture, SIMD, graphics APIs, multimedia codecs, profiling, diagnostics, telemetry, and cross-team technical collaboration experience.
C++, Rust, GLSL, HLSL, Metal, Vulkan, WebGPU, OpenGL, Perf, Instruments, Chrome DevTools, Systrace, CMake, Web, iOS, Android, H.264, H.265, VP9, AV1, SIMD
9h
Save
Mark Applied
Hide
Staff Software Engineer - Video Performance - (Bay area only)
San Francisco, California, United States
$251k-$329k/yr OnsiteFull Time
Canva
Canva: Online graphic design platform for creating and editing visual content.
Strong C++ proficiency; experience optimizing multithreaded systems, graphics APIs, multimedia, profiling, telemetry, and CPU/GPU architectures. Systems language experience and performance engineering expertise are valued.
C++, Rust, GLSL, HLSL, Metal, Vulkan, WebGPU, OpenGL, Perf, Instruments, Chrome DevTools, Systrace, H.264, H.265, VP9, AV1, iOS, Android, Web, SIMD
11h
Save
Mark Applied
Hide
Senior Software Engineer, Robotics Tracking and Fusion
Fort Collins, Colorado, United States
$190k-$252k/yr OnsiteFull Time
Anduril Industries
Anduril Industries: Defense technology building autonomous military hardware and software.
Expertise in C/C++, Python, and Matlab; target tracking, sensor fusion, state estimation, applied mathematics, signal processing, machine learning, and production software lifecycles. Eligible for U.S. Top Secret clearance.
C, C++, Python, Matlab, Lattice OS, NoSQL
13h
Save
Mark Applied
Hide
Sr. Software Engineer, Client Backend
Santa Clara or Wisconsin
$99k-$200k/yr OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
10+ YOE10+ years developing cloud security, network, or endpoint solutions; strong Golang, Python, and Node.js skills; scalable API, Kubernetes, cryptography, CI/CD, and distributed-systems experience; BS in Computer Science required.
Golang, Python, Node.js, Kubernetes, RESTful APIs, HTTPS, TLS, Jenkins, Claude Code, Pi, Codex, CASB, Secure Web Gateway (SWG), Zero Trust Network Access (ZTNA), Cloud Firewall (CFW), Endpoint DLP