NVIDIA
Posted 2w ago

Senior Solutions Architect, NVIDIA Cloud Partner Operations

NVIDIA
Santa Clara, California, United States
$224k-$357k/yrOnsiteFull Time
Responsibilities
  • solving operations problems
  • improving reliability
  • creating workflows
Requirements
  • Requires BS, MS, or PhD or equivalent experience
  • 12+ years in production infrastructure or related technical roles, or 5+ years of exceptional GPU/AI infrastructure expertise
  • Strong Linux
  • Automation, and distributed systems experience
Technical tools mentioned
DCGMBMCRedfishInfiniBandhigh-speed EthernetNCCLUFMLustreIBM Storage ScaleWEKAVAST DataKubernetesSlurmPrometheusGrafanaOpenTelemetryTerraformAnsibleArgo CDLinuxPythonBashGB200GB300 NVL72Spectrum-XBase Command ManagerMission ControlGPU OperatorsNetwork OperatorsGitOps

Job description

NVIDIA is looking for a hands-on Solutions Architect to raise the Day 2 operations bar across our NVIDIA Cloud Partner ecosystem. Day 2 starts when a cluster is installed and validated: keeping the service healthy, adapting it as technology and customer demand change, and improving performance, stability, efficiency and economics over time. You will work with engineers running AI clouds at scale on the problems that decide whether customers stay and whether the next generation of NVIDIA technology lands successfully.

Our job is to work hand in hand with NCPs to solve real problems and drive real optimizations, prove the answer, and turn it into something the next partner can use! This is not an outsourced operations role. The partner owns its cloud; success means leaving its team more capable, not more dependent on ours.
 

What you'll be doing:

  • Solve hard Day 2 operations problems at scale. Work alongside partner engineers to find the cause, prototype an approach, validate it under representative load, and leave behind a practice their team can operate.

  • Make new technology Day 2 ready. Help partners prepare the operating model for new NVIDIA platforms, capacity, services, and use cases before customers depend on them, and help drive adoption in live environments without degrading service.

  • Improve reliability, performance, and economics together. Use measures such as incident frequency, recovery time, utilization, and cost per token to show where the cloud is losing performance or margin - and whether the fix worked.

  • Raise each partner's Day 2 maturity. Identify and help close the gaps that matter across people, process, tooling, telemetry, security, and incident response.

  • Turn one solution into ecosystem capability. Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows that other NCPs can integrate into their standard operating model.

  • Create the feedback loop only NVIDIA can. Spot patterns across partners early and bring clear field evidence to account teams, support, product, and engineering so repeated problems are fixed at the right level.

What we need to see:

  • BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field - or equivalent experience.

  • 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure.

  • Experience building, operating, or improving distributed infrastructure under real production load - not only designing or deploying it.

  • Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-scale GPU, HPC, or cloud infrastructure. Relevant technologies may include DCGM, BMC/Redfish, and firmware and driver lifecycle; InfiniBand or high-speed Ethernet, NCCL, and UFM; or high-performance storage such as Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms.

  • Working experience across the broader operating platform, including Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling.

  • Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation.

  • A detailed evidence-led approach to troubleshooting across system boundaries, paired with the judgment to make difficult technical findings clear.

  • The ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority or taking ownership away from the operator.

  • Strong communication, prioritization, and time-management skills across multiple partner engagements.

Ways to stand out from the crowd:

  • Real world experience operating a GPU cloud, HPC environment, or large-scale AI platform under customer load.

  • Built or matured a 24/7 operations function, including observability, incident and problem management, coverage, and on-call design.

  • Hands on experience with NVIDIA rack-scale platforms such as GB200 or GB300 NVL72 into production, or have hands-on experience with NVIDIA operations technologies such as Spectrum-X, UFM, Base Command Manager, Mission Control, and the GPU or Network Operators.

  • Driven improved fleet health or unit economics through benchmarking, infrastructure as code, GitOps, automated diagnosis, or agent-based remediation.

Even if your background doesn't match every line above, we'd love to hear how your experience applies.
 

With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. This role presents an opportunity to have a wide impact at NVIDIA by improving the factory planning function. Are you creative, hard-working, dedicated, and determined? Do you love a challenge? If so, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 16, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About NVIDIA

Computing platform for AI and accelerated graphics.

Similar jobs

Solutions Architect roles near Santa Clara, California
2d
Save
Mark Applied
Hide
Senior Solutions Architect, NVIDIA Cloud Partners
Santa Clara or Texas or Virginia or New York or Washington or California
$184k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
10+ YOEBS, MS, PhD, or equivalent experience in engineering, computer science, physics, mathematics, or related fields; 10+ years in solution, sales, cloud engineering, or solution architecture; Python, PyTorch, TensorFlow, AI cloud, and GPU infrastructure expertise.
Python, PyTorch, TensorFlow, NVIDIA Nemotron, NVIDIA NeMo Framework, NVIDIA Dynamo, NeMo Retriever, NVIDIA Triton Inference Server, TensorRT, TensorRT-LLM, NVIDIA CUDA-X, AWS, Azure, GCP, NCCL, DCGM, UFM, Mission Control, Base Command Manager, SLURM, K8s
2d
Save
Mark Applied
Hide
Solutions Architect - NVIDIA Cloud Partners
Santa Clara or California or Texas or Virginia
$184k-$357k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
7+ YOEBS, MS, PhD, or equivalent experience in engineering, computer science, physics, or related fields; 7+ years in solution, sales, or cloud engineering; large-scale clusters, AI/HPC data centers, and customer engagements.
NVIDIA GPUs, NCCL, DCGM, UFM, Mission Control, Base Command Manager, SLURM, K8s
2d
Save
Mark Applied
Hide
Solutions Architect - Communications, Media, Entertainment and Games
San Francisco or California
$180k-$248k/yr RemoteFull Time
Databricks
Databricks: Data and AI software providing a unified platform.
6+ YOERequires 6+ years in solutions architecture, data engineering, technical pre-sales, or senior technical roles; Python and SQL coding proficiency; distributed data systems and public cloud experience; and a relevant bachelor's or master's degree.
Python, SQL, Databricks Platform, AWS, Azure, GCP, Snowflake, Azure Synapse, Databricks, Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, Unity Catalog
2d
Save
Mark Applied
Hide
Senior Solutions Architect, FSI Payments, Financial Services Industry - Payments
San Francisco or Austin or Seattle or East Palo Alto
$154k-$239k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Multinational technology focused on e-commerce and cloud computing.
10+ YOERequires 10+ years of IT development or implementation/consulting experience, including 8+ years in technology domains and 3+ years designing, implementing, or consulting on applications and infrastructure.
AWS
3d
Save
Mark Applied
Hide
Specialist Solutions Architect, Data
San Francisco or New York or United States
HybridFull Time
Stripe
Stripe: Financial infrastructure platform for online businesses and enterprises.
7+ YOE7+ years in technical sales, pre-sales, or solutions architecture; 3+ years with data warehousing, pipelines, and ETL/ELT; expertise in financial reporting, BI platforms, software engineering, and API integrations.
Snowflake, BigQuery, Redshift, Databricks, ETL/ELT, Looker, Tableau, Sigma, Metabase, RESTful APIs, SQL
3d
Save
Mark Applied
Hide
Solutions Architect (Workfront)
Addison or San Francisco or United States
$160k-$205k/yr HybridFull Time
Interpublic Group
Interpublic GroupNYSE: IPG: Global advertising and marketing services holding.
3+ YOERequires 3+ years designing and implementing Adobe Workfront or workflow automation systems, strong Workfront architecture knowledge, technical team leadership, communication skills, and project prioritization.
Adobe Workfront, Adobe Workfront Fusion, AEM Assets, Adobe Creative Suite, Firefly, GenStudio, APIs
3d
Save
Mark Applied
Hide
Solutions Architect - Nordics
Stockholm or London or Toronto or New York City or San Francisco or Montreal or Paris or Berlin or Seoul
RemoteFull Time
Cohere
Cohere: Security-first enterprise AI building foundation models.
3+ YOE3+ years of customer-facing technical pre/post-sales experience; generative AI and foundational ML knowledge; strong communication; Python proficiency; English fluency. NLP/AI/LLM experience preferred.
Python, Assistive Coding tools, Generative LLMs, NLP, AI, LLM
4d
Save
Mark Applied
Hide
Solutions Architect, Expert
Oakland, California, United States
$140k-$238k/yr HybridFull Time
Pacific Gas and Electric Company
Pacific Gas and Electric CompanyNYSE American: PCG-PA: Provides natural gas and electric service.
7+ YOEBachelor's degree or equivalent experience, 7 years in IT, and 4 years in solution architecture and project implementation in asset investment planning, ERP, or GIS. Strong communication and architecture expertise required.
Copperleaf, SAP, PLS-CADD, AutoCAD, AUD, Bentley ProjectWise, AWS, Azure, CI/CD, Microsoft Visio, Lucidchart, TOGAF, Zachman, DODAF, CISSP, ITIL, AI, ML, GenAI, LLM, RAG, APIs