102 inference performance engineer jobs at 54 companies in United States
1w
Save
Mark Applied
Hide
1w
Inference Performance Engineer
San Francisco, California, United States
HybridFull Time
Adaption: AI building adaptive intelligence that continually learns for industries, languages, and specialized workflows.
5+ YOE5+ years in ML systems, inference infrastructure, or performance engineering; model-serving expertise; Python and systems-language proficiency; and GPU performance experience with measurable cost or latency improvements.
Material Group: Material Group is an Austin-based specialized talent practice recruiting technical workers for companies building artificial general intelligence.
BS in CS/EE or related field; proficiency in Rust/Go/Python/C++; knowledge of concurrency, tail latency; experience with model serving; GPU/ASIC programming; low-precision inference; profiling and benchmarking.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
3+ YOEPhD or equivalent experience; 3+ years experience; strong deep learning and inference expertise; performance profiling/optimization for GPU apps; proficient with PyTorch; deep understanding of computer/GPU architecture.
Modular AINASDAQ: QCOM: AI infrastructure helping developers optimize and deploy AI workloads across heterogeneous computing hardware.
5+ YOE5+ years in distributed systems or performance engineering; experience building reusable tooling; strong technical judgment, communication, and leadership; GPU/kernel, inference engine, Kubernetes, and LLM familiarity helpful.
Machine Learning Performance Engineer - Offboard Training & Inference
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$215k-$285k/yrOnsiteFull Time
Applied Intuition: Providing digital infrastructure for physical AI and autonomy.
ML performance engineering experience with distributed training, batch inference, GPU or accelerator optimization, Python, and C++ or another systems language; strong debugging and analytical skills required.
Tower Research Capital: Proprietary quantitative trading firm employing traders, engineers, researchers, and business-support staff to trade global financial markets.
2+ YOERequires 2+ years optimizing deep learning inference, PyTorch or JAX, Python/C++, mixed-precision computation, custom GPU kernels, optimization libraries, compilers, profiling tools, and GPU microarchitecture expertise.
DigitalOceanNYSE: DOCN: The AI-Native Cloud purpose-built for inference and agentic workloads.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
LM Studio: Private U.S. AI software building local and cloud tools for running large language models on personal computers.
Significant production ML, inference runtime, or performance infrastructure experience; strong Python and C++; transformer and inference expertise; CPU/GPU profiling; PyTorch and inference system experience.
ByteDance: Global technology specializing in AI-powered content platforms.
Bachelor's or master's degree in a technical discipline; proficient in C/C++, Python, CUDA, GPU architecture, deep learning operators, inference compilation, performance analysis, and distributed model inference.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
3+ YOEPhD in CS/EE/CSEE or equivalent; 3+ years in deep learning inference; strong profiling/optimization for GPU; PyTorch experience; understanding of GPU architecture.
RadixArk: Private AI infrastructure building open training, inference, and post-training systems for developers and research labs.
Strong systems engineering in performance-critical software; GPU/distributed systems; profiling tools; Python and C++; CUDA/Triton/ROCm/XLA familiarity; LLM inference concepts; ability to debug across software, hardware, and infra layers; strong communication.
IntelNASDAQ: INTC: Design and manufacture of semiconductors and computing technology.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
Engram: Applied AI research lab building cognitive systems that augment human intelligence and accelerate research and engineering.
5+ YOE5+ years building training/inference systems; strong engineering skills; experience with ML frameworks, GPUs, distributed systems; bachelor's degree or equivalent experience.