102 inference performance engineer jobs at 54 companies in United States

1w
Save
Mark Applied
Hide
Inference Performance Engineer
San Francisco, California, United States
HybridFull Time
Adaption
Adaption: AI building adaptive intelligence that continually learns for industries, languages, and specialized workflows.
5+ YOE5+ years in ML systems, inference infrastructure, or performance engineering; model-serving expertise; Python and systems-language proficiency; and GPU performance experience with measurable cost or latency improvements.
vLLM, SGLang, TensorRT-LLM, Python, C++, Rust, CUDA, NCCL
1mo
Save
Mark Applied
Hide
Senior GPU Inference Performance Engineer
Santa Clara, California, United States
$144k-$246k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
Hands-on GPU performance engineering experience; proficent in GPU profiling, distributed inference, AI serving frameworks, Python and C/C++; bachelor's degree preferred.
ROCm, rocProfiler, Omniperf, Omnitrace, CUDA, Nsight Systems, Nsight Compute, DCGM, vLLM, SGLang, FP8, MXFP4, INT4, AWQ, GPTQ, RDMA, RoCE, NCCL, RCCL, Pensando, ConnectX, Docker, containerd, Kubernetes, Python, C++, C, HIP, perftest, ib_write_bw, tcpdump
3mo
Save
Mark Applied
Hide
Inference Performance Engineer
New York, New York, United States
HybridFull Time
Material Group
Material Group: Material Group is an Austin-based specialized talent practice recruiting technical workers for companies building artificial general intelligence.
BS in CS/EE or related field; proficiency in Rust/Go/Python/C++; knowledge of concurrency, tail latency; experience with model serving; GPU/ASIC programming; low-precision inference; profiling and benchmarking.
Rust, Go, Python, C++, vLLM, TensorRT-LLM, llama.cpp, CUDA, ROCm, Triton, TGI, SGLang, Nsight, perf
3mo
Save
Mark Applied
Hide
Senior DL Algorithms Engineer - Inference Performance
Santa Clara or California
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
3+ YOEPhD or equivalent experience; 3+ years experience; strong deep learning and inference expertise; performance profiling/optimization for GPU apps; proficient with PyTorch; deep understanding of computer/GPU architecture.
PyTorch, CUDA, OpenCL, TRT-LLM, vLLM, SGLang, FlashInfer
3mo
Save
Mark Applied
Hide
Performance Engineer, Inference Systems
San Francisco or New York City or Seattle
$350k-$850k/yr OnsiteFull Time
Anthropic
Anthropic: AI research developing safe and steerable AI systems.
Hands-on performance engineering with Python, data analysis, and cross-layer investigations; strong communication of quantitative results.
Python, SQL, Pandas
3mo
Save
Mark Applied
Hide
ML Inference Engineer
San Francisco, California, United States
OnsiteFull Time
Reactor
Reactor: Private developer infrastructure helping developers deploy real-time generative media and world-model applications.
Strong expertise in ML engineering, PyTorch, CUDA, and high-performance inference; experience with diffusion models and low-latency systems.
PyTorch, TensorRT, TransformerEngine, Nsight, ONNX Runtime, CUDA
2mo
Save
Mark Applied
Hide
Inference Optimization Engineer
Los Altos or United States or Canada
$198k-$286k/yr HybridFull Time
Modular AI
Modular AINASDAQ: QCOM: AI infrastructure helping developers optimize and deploy AI workloads across heterogeneous computing hardware.
5+ YOE5+ years in distributed systems or performance engineering; experience building reusable tooling; strong technical judgment, communication, and leadership; GPU/kernel, inference engine, Kubernetes, and LLM familiarity helpful.
GPU, ASIC, inference engine, Kubernetes, Modular Cloud, LLM architectures
2w
Save
Mark Applied
Hide
Machine Learning Performance Engineer - Offboard Training & Inference
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$215k-$285k/yr OnsiteFull Time
Applied Intuition
Applied Intuition: Providing digital infrastructure for physical AI and autonomy.
ML performance engineering experience with distributed training, batch inference, GPU or accelerator optimization, Python, and C++ or another systems language; strong debugging and analytical skills required.
FSDP, DeepSpeed, Megatron, NCCL, NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, Python, C++, CUDA, Triton, CUTLASS, Nsight Systems, Nsight Compute, PyTorch Profiler, perf, Kubernetes, Slurm, ROS, OpenCV
2w
Save
Mark Applied
Hide
Inference Systems Performance Architect
San Jose, California, United States
$245k-$325k/yr OnsiteFull Time
SambaNova Systems
SambaNova Systems: AI infrastructure providing a full-stack platform for enterprises, AI labs, service providers, and sovereign AI initiatives.
12+ YOERequires 12+ years in performance engineering, distributed-systems analysis, workload generation, simulation, modeling, technical leadership, cross-functional influence, mentoring, and delivering complex ambiguous projects.
LLM, GPU, RDU, Headspace, Gympass+, One Medical, Employee Assistance Program (EAP)
2w
Save
Mark Applied
Hide
Machine Leaning Performance Engineer (Inference)
New York City, New York, United States
$200k-$300k/yr HybridFull Time
Tower Research Capital
Tower Research Capital: Proprietary quantitative trading firm employing traders, engineers, researchers, and business-support staff to trade global financial markets.
2+ YOERequires 2+ years optimizing deep learning inference, PyTorch or JAX, Python/C++, mixed-precision computation, custom GPU kernels, optimization libraries, compilers, profiling tools, and GPU microarchitecture expertise.
PyTorch, JAX, Python, C++, Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS, Nsight Systems, Nsight Compute, FPGA, ASIC
1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
San Francisco, California, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNYSE: DOCN: The AI-Native Cloud purpose-built for inference and agentic workloads.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
CUDA, ROCm, TensorRT, OpenAI Triton, AITER, FlashAttention
3w
Save
Mark Applied
Hide
Software Engineer, Inference Runtime
New York City, New York, United States
$150k-$350k/yr HybridFull Time
LM Studio
LM Studio: Private U.S. AI software building local and cloud tools for running large language models on personal computers.
Significant production ML, inference runtime, or performance infrastructure experience; strong Python and C++; transformer and inference expertise; CPU/GPU profiling; PyTorch and inference system experience.
Python, C++, PyTorch, llama.cpp, MLX, ExecuTorch, vLLM, SGLang, TensorRT-LLM, CUDA, Metal, Vulkan, ROCm
3mo
Save
Mark Applied
Hide
Senior DL Algorithms Engineer - Inference Performance
Santa Clara or California
$152k-$288k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
3+ YOEPhD in computer science, electrical engineering, or equivalent experience; 3+ years' experience; deep learning, inference, GPU performance optimization, PyTorch, computer architecture, and GPU architecture expertise.
TRT-LLM, vLLM, SGLang, FlashInfer, PyTorch, CUDA, OpenCL
3mo
Save
Mark Applied
Hide
Founding Engineer - ML Performance
San Francisco, California, United States
$250k-$395k/yr RemoteFull Time
uRun
uRun: AI infrastructure helping model labs, builders, and research teams run real-time interactive video and stateful inference.
Hands-on CUDA, GPU optimization, and large-scale model inference experience; strong systems and performance engineering skills.
CUDA, GPU, NCCL, PyTorch, Triton, TensorRT, CUDA kernels
3w
Save
Mark Applied
Hide
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Global technology specializing in AI-powered content platforms.
Bachelor's or master's degree in a technical discipline; proficient in C/C++, Python, CUDA, GPU architecture, deep learning operators, inference compilation, performance analysis, and distributed model inference.
C, C++, Python, CUDA, Nsight, Profiler, vLLM, TensorRT-LLM, SGLang
3mo
Save
Mark Applied
Hide
Senior DL Algorithms Engineer - Inference Performance
Santa Clara, California, United States
$152k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
3+ YOEPhD in CS/EE/CSEE or equivalent; 3+ years in deep learning inference; strong profiling/optimization for GPU; PyTorch experience; understanding of GPU architecture.
PyTorch, CUDA, OpenCL, GPU profiling tools
3mo
Save
Mark Applied
Hide
Performance Engineer
Palo Alto, California, United States
OnsiteFull Time
RadixArk
RadixArk: Private AI infrastructure building open training, inference, and post-training systems for developers and research labs.
Strong systems engineering in performance-critical software; GPU/distributed systems; profiling tools; Python and C++; CUDA/Triton/ROCm/XLA familiarity; LLM inference concepts; ability to debug across software, hardware, and infra layers; strong communication.
CUDA, Triton, Pallas, ROCm, XLA, Python, C++
1mo
Save
Mark Applied
Hide
Sr. Inference Optimization Engineer (local / edge runtime)
Santa Clara or Hillsboro or Folsom or Phoenix
$195k-$361k/yr HybridFull Time
Intel
IntelNASDAQ: INTC: Design and manufacture of semiconductors and computing technology.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
C++, Python, llama.cpp, vLLM, ggml, Vulkan, SYCL, oneAPI, CUDA, Metal, SIMD, Linux, GGUF, AWQ, GPTQ
2mo
Save
Mark Applied
Hide
ML Systems & Performance Engineer
San Francisco, California, United States
OnsiteFull Time
Engram
Engram: Applied AI research lab building cognitive systems that augment human intelligence and accelerate research and engineering.
5+ YOE5+ years building training/inference systems; strong engineering skills; experience with ML frameworks, GPUs, distributed systems; bachelor's degree or equivalent experience.
PyTorch, JAX, GPUs
3w
Save
Mark Applied
Hide
AI Performance Modeling Engineer
Burlingame, California, United States
$180k-$225k/yr HybridFull Time
Quadric
Quadric: Private semiconductor IP licensor providing programmable AI processors for on-device inference to chip designers.
Strong Python and quantitative modeling skills; computer architecture knowledge; technical writing ability; AI inference or performance modeling expertise; BS, MS, PhD, or equivalent practical experience.
Python, C++, CUDA, Triton, gem5, Timeloop, MAESTRO, Accel-Sim