NVIDIA
Posted 6mo ago

Senior Performance Architect - Heterogeneous Workload Optimization

NVIDIA
Santa Clara or Durham or Austin or Westford
$184k-$357k/yrHybridFull Time
Responsibilities
  • architecting frameworks
  • benchmarking applications
  • analyzing performance
Requirements
  • Expertise in CUDA and GPU profiling (Nsight)
  • System-level performance analysis with 8+ years relevant experience (5+ in systems performance)
  • Familiarity with perf/eBPF/VTune/Valgrind and distributed compute (Slurm, LSF
  • Kubernetes)
Technical tools mentioned
CUDANVIDIA Nsight SystemsNVIDIA Nsight ComputeperfeBPFVTuneValgrindSlurmLSFKubernetes

Job description

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

As EDA workloads transition from traditional CPU-bound tasks to massively parallel GPU-accelerated engines, the complexity of identifying bottlenecks has scaled exponentially. We are seeking a Senior Systems Performance Engineer to build our next generation of profiling infrastructure. You will be responsible for measuring, analyzing, and optimizing the interaction between extensive design graphs in system memory and high-throughput kernels on the GPU. Join us to push the boundaries of what's possible in the future of computing!

What you'll be doing:

  • Architecting and maintaining custom profiling frameworks that provide a unified view of execution across CPU (multi-core/multi-socket) and GPU (multi-node/NVLink) environments.

  • Conducting deep-dive benchmarking of EDA applications to characterize memory access patterns, cache hit rates, and instruction-level parallelism.

  • Using GPU profilers to detect GPU-side inefficiencies such as warp divergence, sub-optimal occupancy, and PCIe/NVLink bottlenecks.

  • Developing tools to monitor and attribute high-watermark memory usage in multi-terabyte EDA builds, finding opportunities for data structure compression or smarter memory pooling.

  • Developing predictive models to guide hardware procurement and cloud instance selection based on built gate-count and algorithmic complexity.

What we need to see:

  • A grasp of the CUDA programming model and experience employing GPU profiling tools like NVIDIA Nsight Systems/Compute to address PCIe bottlenecks and kernel stalls.

  • Extensive knowledge of profiling tools such as perf, eBPF, VTune, or Valgrind, along with insight into their internal mechanisms.

  • A passion for meticulous benchmarking and the ability to distill sophisticated performance data into actionable engineering roadmaps.

  • Experience with distributed compute environments (Slurm, LSF, or Kubernetes).

  • A BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field (or equivalent experience) with more than 8+yrs of relevent experience and at least 5 years involved in systems-level performance analysis.

NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most brilliant and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you!

#LI-Hybrid

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until February 16, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About NVIDIA

Designs graphics processing units and artificial intelligence hardware.

Similar jobs

Performance Architect roles near Santa Clara, California
3w
Save
Mark Applied
Hide
Principal Performance Architect
Mountain View or Austin or Raleigh or Hillsboro or Redmond
$143k-$275k/yr HybridFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
3+ YOEAdvanced degree in EE/CE/CS or equivalent experience, 3+ years technical engineering experience (depending on degree), experience with performance modeling and SoC architecture, strong Python/C/C++ skills, ability to pass security and export-control screening.
Python, C, C++
2mo
Save
Mark Applied
Hide
Senior Performance Architect, Nemotron
Santa Clara, California, United States
OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
3+ YOEMaster's in CS/EE or equivalent; strong architecture, roofline modeling, queuing theory; ML fundamentals; Python; 3+ years AI/ML performance analysis; PyTorch, CUDA experience.
Python, C++, PyTorch, CUDA, TRT-LLM, VLLM, SGLang
1w
Save
Mark Applied
Hide
Core Performance Architect
Santa Clara or Hsinchu or Austin or Berkeley
$179k-$219k/yr OnsiteFull Time
SiFive
SiFive: Designs and licenses high-performance RISC-V processor intellectual property.
3+ YOEMS in Computer Science or Computer Architecture, 3+ years in high-performance processor development, hands-on RTL/FPGA experience, EM/FPGA debug, performance analysis, strong communication and leadership.
RISC-V, RTL, FPGA, Integrated Logic Analyzer, FPGA flows
3w
Save
Mark Applied
Hide
CPU Performance Architect
Mountain View, California, United States
$163k-$237k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's in EE/CE/CS or equivalent,8+ years performance analysis and microarchitecture experience,proficiency with C/C++ and Python,EMR? (not mentioned),knowledge of ARM and low-level Linux components preferred.
C, C++, Python, Linux kernel
3w
Save
Mark Applied
Hide
SoC Performance Architect (Server CPU)
Santa Clara or Austin or Hillsboro
$142k-$213k/yr OnsiteFull Time
Qualcomm
QualcommNASDAQ: QCOM: Designs and manufactures semiconductors and wireless telecommunications products.
2+ YOEExpertise in CPU microarchitecture and SoC internals, performance analysis using PMU/profiling, Linux server workload debugging, scripting and data analysis; PhD or advanced degree preferred; 2+ years systems/architecture experience.
Linux, PMU
1mo
Save
Mark Applied
Hide
Performance Modeling Architect – AI Systems
Durham or Santa Clara or Boston or Austin
$200k-$500k/yr OnsiteFull Time
Velaura AI
Velaura AI: Developing ultra-low-power silicon and IP for AI accelerators.
Experienced in computer/system architecture and performance modeling; building simulation or analytical models; strong programming skills (Python, C++); knowledge of CPUs/GPUs/accelerators; ability to analyze system bottlenecks.
Python, C++
1mo
Save
Mark Applied
Hide
GPU Performance Architect
Folsom or Santa Clara
$127k-$217k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Develop performance models and prototypes for GPU/SoC systems; analyze ML/HPC workloads; write and optimize GPU kernels; strong programming in C/C++/Python; knowledge of CUDA/OpenCL and ML frameworks.
TensorFlow, PyTorch, CUDA, OpenCL, C, C++, Python, Hip, Triton, MLIR, LLVM, RTL, System C
1mo
Save
Mark Applied
Hide
Sr. Performance Modeling Architect
Santa Clara or Austin
$100k-$500k/yr HybridFull Time
Tenstorrent
Tenstorrent: Designs and manufactures AI processors and RISC-V CPU solutions.
1+ YOEPhD preferred (MS considered) in a related field; 1+ years industry or research experience in CPU/core or microarchitecture; experience with Gem5, SST, SimpleScalar; programming in C++ and Python; strong processor subsystem knowledge.
Gem5, SST, SimpleScalar, C++, Python, RISC-V