161 inference optimization engineer jobs at 87 companies in United States

2mo
Save
Mark Applied
Hide
Inference Optimization Engineer
Los Altos or United States or Canada
$198k-$286k/yr HybridFull Time
Modular
Modular: Unified software infrastructure and programming language for AI development.
5+ YOE5+ years in distributed systems or performance engineering; experience building reusable tooling; strong technical judgment, communication, and leadership; GPU/kernel, inference engine, Kubernetes, and LLM familiarity helpful.
GPU, ASIC, inference engine, Kubernetes, Modular Cloud, LLM architectures
3mo
Save
Mark Applied
Hide
Inference Optimization ML Engineer
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
PyTorch, JAX, TensorRT, Triton, CUDA, XLA, TorchServe, vLLM
2w
Save
Mark Applied
Hide
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or California
$124k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
3+ YOEBS, MS, PhD, or equivalent experience; 3+ years engineering experience; AI inference optimization expertise; GPU profiling; Python; C++/CUDA; reproducible benchmarking and strong communication.
TensorRT-LLM, SGLang, vLLM, Dynamo, Nsight Systems, Nsight Compute, CUPTI, PyTorch profiler, Python, C++, CUDA, FlashInfer, NCCL, NIXL, NVSHMEM, MLPerf Inference, SemiAnalysis InferenceX
2w
Save
Mark Applied
Hide
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or United States
$124k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
3+ YOEBachelor's, master's, or doctoral degree in a related field or equivalent experience; 3+ years' engineering experience; GPU profiling, Python, C++/CUDA, and AI inference optimization expertise required.
TensorRT-LLM, SGLang, vLLM, Dynamo, Nsight Systems, Nsight Compute, CUPTI, PyTorch profiler, Python, C++, CUDA, FlashInfer, NCCL, NIXL, NVSHMEM, MLPerf Inference, SemiAnalysis InferenceX
1mo
Save
Mark Applied
Hide
Sr. Inference Optimization Engineer (local / edge runtime)
Santa Clara or Hillsboro or Folsom or Phoenix
$195k-$361k/yr HybridFull Time
Intel
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
C++, Python, llama.cpp, vLLM, ggml, Vulkan, SYCL, oneAPI, CUDA, Metal, SIMD, Linux, GGUF, AWQ, GPTQ
1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Boston, Massachusetts, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and inference optimization expertise, experience with CUDA/Triton/ROCm, attention-layer and kernel-level optimization, strong system design and leadership through influence.
AITER, CUDA, ROCm, TensorRT, Triton, FlashAttention
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Inference
Palo Alto, California, United States
$185k-$250k/yr HybridFull Time
Pika
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
CUDA, NCCL
1mo
Save
Mark Applied
Hide
Senior Inference Engineer - AI
Eagan or Toronto
$110k-$204k/yr HybridFull Time
Thomson Reuters
Thomson ReutersNASDAQ: TRI: Provides professional software, data, and news services globally.
5+ YOE5+ years experience with ML/LLM inference optimization, GPU programming (CUDA), TensorRT/ONNX Runtime, PyTorch/TensorFlow, Python and C++, Kubernetes and multi-cloud deployment (AWS/Azure/GCP).
AWS, Azure, GCP, OCI, Snowflake, OpenSearch, OpenSearch vector search, OpenAI, Anthropic, Vertex AI, Kubernetes, CUDA, TensorRT, ONNX Runtime, PyTorch, TensorFlow, Python, C++
1mo
Save
Mark Applied
Hide
Principal LLM Inference Engineer
Santa Clara, California, United States
$195k-$285k/yr HybridFull Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Python, C, C++, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, CUDA, Triton, JAX
2mo
Save
Mark Applied
Hide
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yr OnsiteFull Time
Nuance Labs
Nuance Labs: A building photorealistic, real-time AI avatars and full-duplex audiovisual systems.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, CUDA, Triton
1w
Save
Mark Applied
Hide
Senior Inference Engineer, AGI
Sunnyvale or Boston or Seattle or Los Angeles County
$193k-$262k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOERequires 5+ years software development, 4+ years systems architecture, a computer science bachelor's degree, 2+ years neural inference optimization, GPU optimization, real-time systems, and technical leadership experience.
Nsight Compute, Nsight Systems, vLLM, PyTorch, TensorRT-LLM, CUTLASS, Triton, CUDA, PTX, FlashAttention, NCCL, NVLink, AWS Neuron, Trainium, Microsoft Excel
1mo
Save
Mark Applied
Hide
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yr OnsiteFull Time
Elorian AI
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
vLLM, TensorRT-LLM, Triton, SGLang, C++, CUDA, Python
1mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer, LLM Inference Optimization
Palo Alto or California
$195k-$262k/yr OnsiteFull Time
Nebius
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
1mo
Save
Mark Applied
Hide
AI Inference Engineer - Speech
Seattle or San Jose
$152k-$332k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Python, shell, C/C++, PyTorch, TensorFlow, CUDA, TensorRT, CUDA Graphs, NVIDIA GPUs, TPU, BrightHire
1mo
Save
Mark Applied
Hide
Applied AI Inference Engineer
San Francisco or Sunnyvale
$250k-$300k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
1w
Save
Mark Applied
Hide
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
San Jose or Mountain View
$170k-$351k/yr OnsiteFull Time
DiDi Autonomous Driving
DiDi Autonomous Driving: Develops Level 4 autonomous driving technology for shared mobility.
3+ YOEMaster's degree or higher in a technical field; 3+ years in HPC, AI infrastructure, model optimization, or embedded deployment; C++, Python, CUDA, OpenMP, inference engines, GPU architectures, and system profiling expertise.
C++, Python, CUDA, OpenMP, TensorRT, ONNX Runtime, vLLM, SGLang, TensorRT-LLM, NVIDIA Hopper, NVIDIA Thor, TGI, LightLLM, PyTorch, INT8, FP8, AWQ, LLaMA, Qwen, GPT
2w
Save
Mark Applied
Hide
Research Engineer, Infrastructure, Inference
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines
Thinking Machines: Building AI systems to extend human will and judgment.
Bachelor's in CS or equivalent, strong engineering skills, experience with deep learning frameworks and inference serving, ability to optimize distributed GPU systems and contribute production-quality code.
PyTorch, JAX, SGLang, vLLM, Kubernetes, Ray, SLURM, Triton, DeepSpeed, XLA
1mo
Save
Mark Applied
Hide
Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
Seattle, Washington, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's in CS/EE/Software, strong C/C++ and Python, experience with PyTorch or TensorFlow, production LLM/VLM inference optimization, GPU familiarity and operator optimization, containerization experience.
C, C++, Python, PyTorch, TensorFlow, CUDA, OpenCL, TensorRT, Triton, CUTLASS, FlashAttention, GEMM, GEMV, Conv2D
4d
Save
Mark Applied
Hide
Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind
Mountain View, California, United States
$207k-$300k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's degree or equivalent experience, 8 years of software development, Python and C++, AI inference optimization expertise, and experience with serving codebases and performance tradeoffs.
Python, C++, vLLM, TensorRT-LLM, SGLang, Dynamo, PyTorch profiler, GPU, TPU
1w
Save
Mark Applied
Hide
Machine Leaning Performance Engineer (Inference)
New York City, New York, United States
$200k-$300k/yr HybridFull Time
Tower Research Capital
Tower Research Capital: Global quantitative trading firm developing automated algorithmic strategies.
2+ YOERequires 2+ years optimizing deep learning inference, PyTorch or JAX, Python/C++, mixed-precision computation, custom GPU kernels, optimization libraries, compilers, profiling tools, and GPU microarchitecture expertise.
PyTorch, JAX, Python, C++, Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS, Nsight Systems, Nsight Compute, FPGA, ASIC