NVIDIA
Posted 2w ago

Principal Engineer, Cloud Site Reliability Engineering

NVIDIA
Santa Clara, California, United States
$272k-$431k/yrOnsiteFull Time
Responsibilities
  • architecting cloud infrastructure
  • developing software solutions
  • optimizing development workflows
Requirements
  • Requires a BS or MS in Electrical Engineering
  • Computer Science, or equivalent experience
  • 15+ years of systems software development
  • AI experience
  • Cloud infrastructure
  • Programming
  • Distributed systems
  • Databases
  • Containers, and cloud technologies
Technical tools mentioned
JavaPythonShellMySQLCassandraMongoDBElasticsearchDockerOpenStackKubernetesChefPuppetHadoopCephSwiftStackLXCGitPerforceJFrogKafkaREST APIsSQLNoSQLVirtual MachinesWindowsLinuxAndroid

Job description

NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence and Autonomous Vehicles to cater to their infrastructure needs. These cloud services provide almost half a million automated jobs per day on thousands of servers helping with the efficiency of thousands of NVIDIA's software engineers worldwide. The cloud hosts various machines and devices with operating systems like Windows, Linux, and Android. It supports hardware platforms including NVIDIA GPUs and Tegra Processors. It delivers unified CI/CD solutions and cloud-based software development. Are you passionate about distributed infrastructure and looking for sophisticated, critical issues, ready to build the next generation of cloud services, design creative solutions, mine through data to uncover real problems and fix them?

What you'll be doing:

  • Serve as an SRE Architect part of GPU Private Cloud team used by thousands of NVIDIANs globally for interactive development, centralized CI/CD, and QA testing.

  • Evaluating, identifying and developing software solutions to optimize critical software development workflows across various organizations within NVIDIA.

  • Architecting, implementing, and supporting end-to-end CI/CD system using open-source and NVIDIA proprietary software.

  • Customer (NVIDIA Internal development teams) onboarding to Private cloud infrastructure with a good discovery of the use case and available solutions within the cloud.

  • Identify performance bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.

  • Leading software development projects and technically direct a team of brilliant engineers and guide them to provide efficient and impactful solutions.

  • Looking for problems within software systems and resolving the issues

  • Craft and implement critical metrics using various analytics methods and dashboards.

What we need to see:

  • BS or MS in Electrical Engineering, Computer Science, or relevant field (or equivalent experience).

  • 15+ years of systems software development including at least 1 year dedicated to developing/exploring AI.

  • Experience of maintaining cloud infrastructure and highly available production environment.

  • Strong programming and software development skills in JAVA, Python, Shell-script along with good understanding of distributed systems and REST APIs.

  • Experience in working with SQL/NoSQL database systems such as MySQL, Cassandra, MongoDB or Elasticsearch.

  • Excellent knowledge and working experience with Docker containers and Virtual Machines.

  • Good background of Cloud technologies like: OpenStack, Docker, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, Kafka.

  • Ability to work across organizational boundaries effectively to improve alignment and productivity between teams in a multi-national, multi-time-zone corporate environment.

Ways to stand out from the crowd:

  • Depth in AI, Machine Learning and Deep Learning algorithms and techniques.

  • Strong collaborative and interpersonal skills, with a consistent record of guiding and influencing others in dynamic environments.

  • Experience developing large-scale software systems using modular architecture under real-time performance requirements.

  • Background in designing high-performance, scalable software systems with a strong focus on hardware cost optimization.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 9, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About NVIDIA

Designs GPU-accelerated computing and artificial intelligence hardware.

Similar jobs

Principal Engineer roles near Santa Clara, California
13h
Save
Mark Applied
Hide
Principal Engineer, Core - Pinner Journeys
San Francisco or United States or Seattle or Palo Alto
$243k-$450k/yr RemoteFull Time
Pinterest
PinterestNYSE: PINS: Visual discovery engine for finding inspiration and creative ideas.
15+ YOE15+ years in large-scale systems or machine learning; principal-level technical leadership; deep retrieval, ranking, search, representation learning, or generative AI expertise; bachelor's degree or equivalent experience.
AI, Machine Learning, Homefeed, Search, AI Assistant, Large Language Models (LLMs), Foundation Models, Learning-to-Rank, RL, Embeddings, RAG
21h
Save
Mark Applied
Hide
Principal Engineer, Core - Pinner Journeys
San Francisco or United States or Palo Alto or Seattle
$243k-$450k/yr RemoteFull Time
Pinterest
PinterestNYSE: PINS: Visual discovery engine for finding and saving ideas.
15+ YOE15+ years in large-scale systems or machine learning, deep expertise in retrieval/ranking, representation learning, search, or generative AI, technical strategy leadership, and a bachelor's degree or equivalent experience.
AI, machine learning, large language models, foundation models, retrieval, ranking, representation learning, embeddings, multi-modal models, dense and sparse retrieval, search systems, indexing, learning-to-rank, bandits, reinforcement learning (RL), generative AI, prompting, fine-tuning, retrieval-augmented generation (RAG), agents, causal inference
2d
Save
Mark Applied
Hide
Principal Engineer - Land
Walnut Creek, California, United States
$150k-$190k/yr HybridFull Time
Fugro
FugroEuronext Amsterdam: FUR: Provides geotechnical, survey, and geoscience data for infrastructure and energy.
15+ YOEBachelor’s in civil or geotechnical engineering required; 15+ years leading complex geotechnical projects; California PE and GE licenses required; advanced soil and rock mechanics knowledge and U.S. work authorization required.
Microsoft Office Suite, LinkedIn Learning
6d
Save
Mark Applied
Hide
Principal Engineer, Cloud Capacity Planning
Sunnyvale, California, United States
$307k-$427k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
15+ YOEBachelor's degree or equivalent practical experience, 15 years of software engineering experience, and experience delivering large-scale capacity planning, IaaS/PaaS, or fleet management systems.
IaaS, PaaS, TPUs, GPUs
1w
Save
Mark Applied
Hide
Principal Engineer- Optics Manufacturing Integration
Vista or San Diego or Santa Clara or Colorado Springs
$159k-$250k/yr OnsiteFull Time
Samtec
Samtec: Manufacturer of high-speed electronic connectors and cable assemblies.
15+ YOEBachelor's degree in engineering or materials science and 15+ years in optoelectronic or microelectronics manufacturing, including process development, assembly, technology transfer, and automated tools.
DOE, FMEA, Six Sigma
1w
Save
Mark Applied
Hide
Principal Engineer- Optics Manufacturing Integration
Vista or San Diego or Santa Clara or Colorado Springs
$159k-$250k/yr OnsiteFull Time
Samtec
Samtec: Manufactures electronic connectors, cables, and optical interconnect solutions.
15+ YOERequires a bachelor's degree in engineering or related field, 15+ years in optoelectronic or microelectronics manufacturing, process development, technology transfer, and automated manufacturing tools.
DOE, FMEA, Six Sigma
1w
Save
Mark Applied
Hide
Principal Engineer- Optics Manufacturing Integration
Vista or San Diego or Santa Clara or Colorado Springs
$159k-$250k/yr OnsiteFull Time
Samtec
Samtec: Manufactures electronic connectors and high-speed cable assembly systems.
15+ YOEBachelor's degree in engineering or related field, 15+ years in optoelectronic or microelectronics manufacturing, process development, technology transfer, and automated manufacturing tools; US citizenship or green card required.
Design of Experiments (DOE), Failure Mode and Effects Analysis (FMEA), Six Sigma
1w
Save
Mark Applied
Hide
Principal Engineer, Mapping and Localization
Santa Clara, California, United States
$200k-$260k/yr HybridFull Time
PlusAI
PlusAI: AI-based virtual driver software for factory-built autonomous trucks.
7+ YOEM.S. or Ph.D. in a relevant field, 7+ years in robotics, 4+ years in production localization or mapping, expert Modern C++, probabilistic robotics, SLAM, 3D geometry, and sensor-fusion tools.
C++, LiDAR, Radar, Camera, IMU, GNSS, HD Maps, Factor Graphs, EKF, UKF, Ceres, g2o, GTSAM, SLAM, Eigen, PCL, OpenCV, LOAM, lego-LOAM, ORB-SLAM, VINS, CUDA, GPU, ARM