99 gpu performance profiling engineer jobs at 48 companies in United States

3w
Save
Mark Applied
Hide
GPU Performance Profiling Engineer
Austin or Santa Clara
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
2+ YOEBachelor's in electrical engineering or computer science with 4+ years, master's with 2+ years, or Ph.D.; expertise in C/C++/Python, drivers, GPU architecture, performance analysis, and hardware-software co-design.
C, C++, Python, CUDA, x86, ARM, GPU
3w
Save
Mark Applied
Hide
GPU Performance Profiling Engineer
Austin or Santa Clara
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
2+ YOERequires a BS with 4+ years, MS with 2+ years, or PhD; expertise in C/C++/Python software stacks, user- and kernel-mode drivers, CPU/GPU architecture, hardware pipelines, performance analysis, and hardware-software co-design.
C, C++, Python, CUDA, x86, ARM
3w
Save
Mark Applied
Hide
GPU Performance Profiling Engineer
Austin or Santa Clara
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
2+ YOERequires a bachelor's degree and 4+ years, master's degree and 2+ years, or Ph.D.; expertise in C/C++/Python, drivers, CPU/GPU architecture, hardware pipelines, and performance analysis.
C, C++, Python, CUDA, x86, ARM, GPU
1mo
Save
Mark Applied
Hide
GPU Performance Engineer
New York City, New York, United States
$165k-$300k/yr HybridFull Time
Two Sigma
Two Sigma: Quantitative investment and trading firm using data science to manage institutional capital and trade financial markets.
1+ YOEExpert CUDA and GPU architecture knowledge, C++ and Python skills, BS/MS in STEM, minimum 1 year relevant experience, GPU profiling and optimization experience.
CUDA, Nsight Systems, Nsight Compute, RAPIDS, CUTLASS, cuBLAS, TensorRT, NCCL, MPI, cuDF, C++, Python
1mo
Save
Mark Applied
Hide
Senior GPU Inference Performance Engineer
Santa Clara, California, United States
$144k-$246k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
Hands-on GPU performance engineering experience; proficent in GPU profiling, distributed inference, AI serving frameworks, Python and C/C++; bachelor's degree preferred.
ROCm, rocProfiler, Omniperf, Omnitrace, CUDA, Nsight Systems, Nsight Compute, DCGM, vLLM, SGLang, FP8, MXFP4, INT4, AWQ, GPTQ, RDMA, RoCE, NCCL, RCCL, Pensando, ConnectX, Docker, containerd, Kubernetes, Python, C++, C, HIP, perftest, ib_write_bw, tcpdump
2mo
Save
Mark Applied
Hide
Senior Engineer, GPU Graphic Core Performance Verification
San Jose, California, United States
$139k-$208k/yr OnsiteFull Time
Samsung Electronics
Samsung ElectronicsKorea Exchange: 005930: Global leader in technology, semiconductors, and consumer electronics.
2+ YOE2+ years experience with a Bachelor’s in computer science/engineering (or Master’s). Strong GPU architecture and performance verification experience; proficiency in C++, Python, Linux; familiarity with OpenGL/Vulkan/OpenCL; performance test and profiling experience.
C++, Python, Linux, OpenGL, Vulkan, OpenCL, SystemC, Sparta, RTL
2d
Save
Mark Applied
Hide
GPU Software Engineer (CUDA)
United States
$100k-$175k/yr RemoteFull Time
Bright Vision Technologies
Bright Vision Technologies: AI-powered enterprise automation and software development firm.
6+ YOEBachelor’s or master’s degree in computer science, computer engineering, or related field; 6+ years in GPU programming and performance engineering; expertise in CUDA C/C++, GPU architectures, profiling, optimization, and distributed workloads.
CUDA, CUDA C/C++, Nsight Systems, Nsight Compute, CUDA profilers, NCCL, RDMA, PyTorch, JAX, Triton, MPI, C++, CUTLASS, TensorRT, FasterTransformer, vLLM, LLVM, MLIR
1mo
Save
Mark Applied
Hide
GPU Kernel Engineer
San Francisco, California, United States
$180k-$280k/yr OnsiteFull Time
TypeSafe AI
TypeSafe AI: Private frontier AI lab building reliable, general AI systems for real-world automation and decision-making.
Deep CUDA/GPU kernel expertise, experience building and optimizing training and inference kernels, LLM training experience, profiling and eliminating performance bottlenecks.
CUDA, CuTe DSL
2mo
Save
Mark Applied
Hide
Member of Technical Staff, TPU & AMD GPU Performance Engineering
San Francisco, California, United States
$200k-$400k/yr OnsiteFull Time
Inferact
Inferact: Private AI infrastructure building open-source software that makes LLM inference faster and cheaper.
Bachelor's or equivalent experience in CS/engineering; hands-on AMD GPU/TPU optimization; experience with ROCm/HIP/Triton, XLA/JAX/Pallas; strong profiling, benchmarking, and cross-platform performance skills.
vLLM, ROCm, HIP, Triton, CK, AITER, TPU, XLA, JAX, Pallas, SGLang, TensorRT-LLM, ATOM, MLIR, LLVM, PyTorch
3mo
Save
Mark Applied
Hide
Performance Engineer
Palo Alto, California, United States
OnsiteFull Time
RadixArk
RadixArk: Private AI infrastructure building open training, inference, and post-training systems for developers and research labs.
Strong systems engineering in performance-critical software; GPU/distributed systems; profiling tools; Python and C++; CUDA/Triton/ROCm/XLA familiarity; LLM inference concepts; ability to debug across software, hardware, and infra layers; strong communication.
CUDA, Triton, Pallas, ROCm, XLA, Python, C++
1mo
Save
Mark Applied
Hide
Performance Library Engineer
San Jose or Pittsburgh
$130k-$230k/yr OnsiteFull Time
Efficient Computer
Efficient Computer: Energy-efficient processor building programmable general-purpose chips for AI and edge applications.
5+ YOE5+ years experience developing low-level C/C++ libraries for performance on hardware; experience with at least two RISC/DSP/GPU platforms, CUDA or HIP, profiling/benchmarking, EM/real-time library design, strong communication, and a Bachelor's in a technical field.
CUDA, HIP, PTX, LLVM IR, MLIR, C, C++
3mo
Save
Mark Applied
Hide
HPC & AI Performance Engineer
Bloomington, Minnesota, United States
$63k-$145k/yr HybridFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Global edge-to-cloud advancing how people live and work.
Experience with HPC benchmarking, performance analysis, and HPC software; knowledge of CPU/GPU HPC architectures; proficiency with C/C++, OpenMP, MPI, Python; GPU programming (CUDA/HIP); Linux; English proficiency; Master’s degree preferred.
C++, C, Fortran, OpenMP, MPI, MPI-IO, Python, CUDA, HIP, Linux, Performance profiling tools
2w
Save
Mark Applied
Hide
Helix AI Engineer, Training Performance
San Jose, California, United States
$200k-$400k/yr OnsiteFull Time
Figure
Figure: Develops general-purpose humanoid robots.
3+ YOEBachelor's or Master's in computer science, computer/electrical engineering, or related field; 3+ years AI performance engineering; Python, CUDA/C++, GPU architecture, profiling, NCCL, and networking expertise.
Triton, CUDA, Nsight Systems, Nsight Compute, PyTorch Profiler, HTA, NCCL, RDMA, NVLink, InfiniBand, RoCE, Python, CUDA/C++, PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, FSDP, Gluon
1mo
Save
Mark Applied
Hide
Senior Performance Engineer
San Jose, California, United States
$138k-$206k/yr OnsiteFull Time
Samsung Semiconductor
Samsung SemiconductorKorea Exchange (KRX): 005930: Global leader in semiconductor solutions including memory, system LSI, and foundry services.
0+ YOEAdvanced degree in CS/CE/EE or equivalent experience; strong LLM systems and NVIDIA GPU performance knowledge; experience profiling AI workloads with Nsight tools; proficiency in Python and C++; experience with PyTorch, DeepSpeed, Ray, or similar.
Nsight Systems, Nsight Compute, Python, C++, PyTorch, vLLM, SGLang, TensorRT-LLM, DeepSpeed, Ray, Megatron-LM
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Kernel Authoring & Optimization
Sunnyvale or Bellevue
$182k-$242k/yr OnsiteFull Time
CoreWeave
CoreWeaveNasdaq: CRWV: Specialized cloud provider for large-scale AI and machine learning.
5+ YOE5+ years building HPC/GPU software, hands-on CUDA kernel authoring and optimization, C++/Python coding, GPU profiling, and experience delivering performance at scale.
CUDA, Nsight Compute, Nsight Systems, C++, Python, Triton, Mojo, CuTe DSL, JAX, HIP, ROCm, NCCL, Kubernetes, SUNK, Slurm, MLPerf, vLLM, TensorRT-LLM, llm-d, SGLang, KNYFE, Pallas, CUTLASS
3mo
Save
Mark Applied
Hide
Inference Performance Engineer
New York, New York, United States
HybridFull Time
Material Group
Material Group: Material Group is an Austin-based specialized talent practice recruiting technical workers for companies building artificial general intelligence.
BS in CS/EE or related field; proficiency in Rust/Go/Python/C++; knowledge of concurrency, tail latency; experience with model serving; GPU/ASIC programming; low-precision inference; profiling and benchmarking.
Rust, Go, Python, C++, vLLM, TensorRT-LLM, llama.cpp, CUDA, ROCm, Triton, TGI, SGLang, Nsight, perf
2mo
Save
Mark Applied
Hide
System Performance Engineer, Senior
San Diego, California, United States
$123k-$184k/yr OnsiteFull Time
Qualcomm Technologies, Inc.
Qualcomm Technologies, Inc.: Developing semiconductor, wireless, connectivity, automotive, AI, and computing technologies for device and enterprise customers.
4+ YOE4+ years post‑silicon SoC power and performance analysis experience preferred; proficiency in Python and C/C++; experience profiling on Android/Linux/Windows; strong ARM SoC and memory hierarchy knowledge; leadership and communication skills.
Android, Linux, Windows, Python, C/C++, Vulkan, OpenGL, DX12, CUDA, ARM v8, ARM v9, SMMU, GIC, Coresight-PMU
2w
Save
Mark Applied
Hide
Machine Leaning Performance Engineer (Inference)
New York City, New York, United States
$200k-$300k/yr HybridFull Time
Tower Research Capital
Tower Research Capital: Proprietary quantitative trading firm employing traders, engineers, researchers, and business-support staff to trade global financial markets.
2+ YOERequires 2+ years optimizing deep learning inference, PyTorch or JAX, Python/C++, mixed-precision computation, custom GPU kernels, optimization libraries, compilers, profiling tools, and GPU microarchitecture expertise.
PyTorch, JAX, Python, C++, Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS, Nsight Systems, Nsight Compute, FPGA, ASIC
3mo
Save
Mark Applied
Hide
Member of Technical Staff, Hardware, Performance Engineer
Palo Alto or Austin
$200k-$420k/yr OnsiteFull Time
River AI
River AI: Full-stack AI building personal AI, custom models, training infrastructure, and local hardware for developers and enterprises.
5+ YOEBachelor's in EE/CE/CS and 5+ years experience with advanced process nodes; expert in C/C++ or SystemC; familiarity with compilers (LLVM/GCC/XLA), CUDA/Triton, computer architecture, and profiling/trace analysis.
C, C++, SystemC, LLVM, GCC, XLA, CUDA, Triton, QEMU
3w
Save
Mark Applied
Hide
Performance Tools Intern
San Jose, California, United States
OnsiteFull Time, Internship
Etched
Etched: Private semiconductor startup building AI inference chips, racks, and software for frontier-model customers.
Strong C++ or Rust programming, computer architecture and low-level systems knowledge, interest in performance analysis and profiling, familiarity with profiling tools and OS/driver concepts.
C++, Rust, Python, Nsight, VTune, Xprof, Perfetto, PCIe, Linux, Windows, GPU, TPU