9 gpu performance profiling engineer jobs at 5 companies in Washington

1mo
Save
Mark Applied
Hide
Senior Performance Engineer - DGX Cloud
Santa Clara or Austin or Redmond or Washington or Oregon
$224k-$431k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
12+ YOE12+ years experience, strong C++ and Python programming, foundations in OS/architecture/distributed systems, performance engineering and profiling experience, BS in CS/CE or equivalent.
C++, Python, CUDA, PyTorch, JAX, XLA
1mo
Save
Mark Applied
Hide
Senior Performance Engineer - DGX Cloud
Santa Clara or Austin or Redmond or Oregon or Washington
$224k-$431k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
12+ YOE12+ years experience; BS or higher in CS/CE or equivalent; strong C++ and Python skills; foundation in OS, computer architecture, distributed systems; performance engineering and profiling experience.
C++, Python, CUDA, PyTorch, JAX, XLA
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Kernel Authoring & Optimization
Sunnyvale or Bellevue
$182k-$242k/yr OnsiteFull Time
CoreWeave
CoreWeaveNasdaq: CRWV: Specialized cloud provider for large-scale AI and machine learning.
5+ YOE5+ years building HPC/GPU software, hands-on CUDA kernel authoring and optimization, C++/Python coding, GPU profiling, and experience delivering performance at scale.
CUDA, Nsight Compute, Nsight Systems, C++, Python, Triton, Mojo, CuTe DSL, JAX, HIP, ROCm, NCCL, Kubernetes, SUNK, Slurm, MLPerf, vLLM, TensorRT-LLM, llm-d, SGLang, KNYFE, Pallas, CUTLASS
1mo
Save
Mark Applied
Hide
Fellow Software Engineer — AI Performance & Reliability
San Jose or Bellevue
$235k-$402k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
PhD or equivalent in AI/ML/CS, strong software engineering, experience profiling and optimizing ML models and AI workloads, proficiency in Python/C++, ML frameworks, customer-facing troubleshooting and performance analysis.
Python, C++, PyTorch, TensorFlow, JAX, ROCm, HIP, CUDA, Triton, XLA, MLIR, NCCL
3mo
Save
Mark Applied
Hide
Senior Deep Learning Tools Engineer – CUDA Tile
Santa Clara or Utah or Remote or Remote or Seattle or Redmond or Salt Lake City
$152k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
5+ YOE5+ years software engineering; Python (C++ a plus); CI/CD and automation; GPU/accelerator performance analysis; experience with DL frameworks (PyTorch, TensorFlow, JAX, TensorRT); data analysis and profiling; able to debug complex systems.
Python, C++, PyTorch, TensorFlow, JAX, TensorRT, LLVM, MLIR, CUDA, CI/CD, Telemetry, Profiling
3mo
Save
Mark Applied
Hide
Senior Deep Learning Systems Engineer, Datacenters
Santa Clara or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOEBachelor's in EE/CS or equivalent, 8+ years experience, strong system architecture and performance analysis, programming in C++ and Python, familiarity with Linux, CUDA, DL frameworks, and profiling tools.
Linux, Compilers, CUDA, PyTorch, TensorFlow, C++, Python, Bash, Docker, Slurm, perf, gprof, nvidia-smi, dcgm
1mo
Save
Mark Applied
Hide
Senior Deep Learning Software Engineer, Inference
California or Texas or New York or Washington or Massachusetts
$152k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
5+ YOEMasters/PhD or equivalent,5+ years software development,excellent C/C++ skills,CUDA and GPU programming experience preferred,experience optimizing/deploying DL inference,Python and performance profiling experience helpful.
CUTLASS, OAI Triton, NCCL, CUDA, vLLM, SGLang, FlashInfer, PyTorch, NVSHMEM, C/C++, Python
1w
Save
Mark Applied
Hide
Member of Technical Staff - ML Performance
Toronto or Zürich or Seattle or California
HybridFull Time
Veeda AI
Veeda AI: Private Canadian AI startup building multimodal world models and simulated environments for robotics and Physical AI.
Bachelor's degree or equivalent experience in a related technical field; deep PyTorch and multi-node parallelism experience; Python and C++/CUDA fluency; training-run profiling experience; expertise in low-precision numerics, kernels, or fault diagnosis.
PyTorch, FSDP2, Megatron-Core, TorchTitan, DeepSpeed, Python, C++, CUDA, Nsight Systems, Triton, FlashAttention-4, FlexAttention, torch.compile, NCCL, CUTLASS, CuTe-DSL, ROCm, JAX, XLA
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: AI research developing safe and steerable AI systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD