40 quantization engineer jobs at 31 companies in United States

6d
Save
Mark Applied
Hide
Staff Machine Learning Engineer – Model Optimization & Quantization
Santa Clara, California, United States
$161k-$241k/yr OnsiteFull Time
Qualcomm
QualcommNASDAQ: QCOM: Designs and manufactures semiconductors and wireless telecommunications products.
4+ YOEBachelor's and 4+ years, master's and 3+ years, or PhD and 2+ years in engineering or related work. Requires Python, ML frameworks, quantization expertise, and software development experience.
AIMET, PyTorch, ONNX, TensorFlow, GPTQ, AWQ, SmoothQuant, TFLite/LiteRT, git, C++, Python, AI Hub Workbench, Snapdragon
2d
Save
Mark Applied
Hide
Staff Machine Learning Engineer - LLM Quantization & Deployment
Santa Clara or Mountain View
$215k-$364k/yr OnsiteFull Time
XPeng
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
3+ YOEMaster's in CS, CE, or EE with 3–5 years' industry experience; expertise in Transformer architectures, LLM inference, model quantization, PyTorch, inference stacks, Python, and software engineering.
Python, PyTorch, TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, AWQ, GPTQ, SmoothQuant, INT8, FP4
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Inference
Palo Alto, California, United States
$185k-$250k/yr HybridFull Time
Pika
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
CUDA, NCCL
3mo
Save
Mark Applied
Hide
Lead AI Engineer
San Francisco, California, United States
$300k-$400k/yr HybridFull Time
Noon
Noon: A building AI-powered design tooling.
5+ YOE5+ yrs software engineering; 2+ yrs ML model training/deploy; NLP/LLMs; data pipelines; model optimizations; agentic AI systems; end-to-end ownership
NLP, LLMs, data pipelines, quantization, distillation, LoRA, pruning, agentic systems
3mo
Save
Mark Applied
Hide
Inference Optimization ML Engineer
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
PyTorch, JAX, TensorRT, Triton, CUDA, XLA, TorchServe, vLLM
2mo
Save
Mark Applied
Hide
Model Distillation Engineer
San Jose, California, United States
$120k-$300k/yr OnsiteFull Time
Hark
Hark: A building multimodal AI models and next-generation hardware to create natural human-machine interfaces.
3+ YOE3+ years in model compression/distillation/quantization, strong fluency in PyTorch or TensorFlow, experience with PTQ/QAT and int8 conversion, hardware-aware optimization for constrained devices, and familiarity with audio/sequence model architectures.
PyTorch, TensorFlow, TFLite, ONNX Runtime, AIMET, Hexagon DSP, NPUs, Ambiq MCUs
1mo
Save
Mark Applied
Hide
Principal LLM Inference Engineer
Santa Clara, California, United States
$195k-$285k/yr HybridFull Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Python, C, C++, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, CUDA, Triton, JAX
3w
Save
Mark Applied
Hide
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yr OnsiteFull Time
Elorian AI
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
vLLM, TensorRT-LLM, Triton, SGLang, C++, CUDA, Python
2mo
Save
Mark Applied
Hide
Staff ML Engineer, Hardware Software Co-Design
Palo Alto, California, United States
$206k-$258k/yr OnsiteFull Time
Rivian
RivianNASDAQ: RIVN: Designs and manufactures electric vehicles and charging networks.
Ph.D. or M.S. in a related field; hands-on experience deploying quantized models, ML compilers, and code generation for embedded/heterogeneous systems; strong CV model optimization skills; proficiency with PyTorch, TensorFlow, ONNX, C++, Python, and CUDA/OpenCL.
PyTorch, TensorFlow, ONNX, C++, Python, CUDA, OpenCL
1mo
Save
Mark Applied
Hide
Staff AI Inference and Acceleration Engineer
San Jose, California, United States
$180k-$275k/yr OnsiteFull Time
Figure
Figure: Develops autonomous humanoid robots for commercial and residential tasks.
8+ YOEMS/PhD or equivalent, 8+ years in hardware acceleration/ML systems, expertise in inference runtimes, quantization and pruning, profiling and benchmarking, model-to-hardware mapping, and strong C++/Python skills.
ONNX, TFLite, TVM, MLIR, TensorRT, Torch, SNPE/QNN, JAX, CUDA, ROCm, C++, Python
4d
Save
Mark Applied
Hide
Principal Software Engineer - LLM Optimization
Jersey City or London
$204k-$285k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
7+ YOEFormal software engineering training or certification and 7+ years of experience, with hands-on LLM inference, GPU infrastructure, quantization, benchmarking, cloud systems, and agentic AI development expertise.
vLLM, TensorRT-LLM, SGLang, LLM-D, FP8, INT8, INT4, GPTQ, AWQ, AWS, EKS, GuideLLM, DCGM, NVML, XID
2mo
Save
Mark Applied
Hide
Machine Learning Engineer 5 - Globalization
United States
$466k-$750k/yr RemoteFull Time
Netflix
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
Extensive ML engineering experience with LLMs and multimodal models; expertise in training and inference optimization, distributed training, GPU/accelerator optimization, KV cache/batching/quantization; proficient in PyTorch; strong communication and technical leadership.
PyTorch
1mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer
California or San Francisco or Mountain View
$167k-$251k/yr RemoteFull Time
Unity
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
5+ YOE5+ years in software/ML engineering with on-device or performance-critical systems; production deployment of transformer/diffusion models on-device; experience with inference runtimes, quantization, operator fusion, and GPU/compute APIs; strong Python; communication and mentoring skills.
WebGPU, WebNN, WGSL, Metal, Vulkan, SPIR-V, CUDA, ONNX Runtime Web, ONNX Runtime, Transformers.js, WebLLM, TensorFlow.js, CoreML, TFLite, ExecuTorch, Chrome, Dawn, PIX, Instruments, Snapdragon Profiler, Nsight, RenderDoc, Python, TypeScript, JavaScript, MLIR, TVM, IREE, XLA, C++, Objective-C, Swift
1w
Save
Mark Applied
Hide
AI Developer
United States
RemoteFull Time
Salvo Software
Salvo Software: Develops cloud-connected diagnostic tools and software for the automotive industry.
Strong Python, machine learning, LLM, RAG, MCP, database, Docker, Git, GPU optimization, quantization, and offline deployment experience required; AWS and secure-environment experience preferred.
Python, PyTorch, TensorFlow, scikit-learn, pandas, Postgres, MySQL, Docker, Git, Hugging Face Transformers, Hugging Face Datasets, XML, XSD, python-docx, OpenXML, vLLM, TGI, Ollama, GGUF, GPTQ, AWQ, bitsandbytes, CUDA, cuDNN, FAISS, Chroma, Weaviate, pgvector, Azure DevOps, AWS, LLaMA, Mistral, Qwen, LoRA, Q-LoRA, HyDE, MCP
3mo
Save
Mark Applied
Hide
AI/ML Infrastructure Engineer
San Francisco, California, United States
OnsiteFull Time
Zensors
Zensors: AI platform that turns existing cameras into intelligent sensors.
BS/MS/PhD in CS or EE; strong C/C++ and Python; model optimization, quantization, pruning; GPU performance tuning; profiling tools; cross-functional collaboration.
C/C++, Python, Nsight Systems, Nsight Compute, PyTorch, CUDA, TensorRT, NVIDIA DeepStream, DALI, FFmpeg, TVM, MLIR, ONNX Runtime, Triton
1mo
Save
Mark Applied
Hide
R&D Engineer
London or Madrid or Shenzhen or New York City
RemoteFull Time
Ultralytics
Ultralytics: Developing open-source computer vision models and AI platforms.
5+ YOE5+ years in computer vision and deep learning with architecture design experience; expert Python and PyTorch; experience reproducing papers, efficiency research (quantization/pruning/distillation), distributed training and strong research portfolio.
Python, PyTorch, CUDA, GitHub, Ultralytics Python package
1mo
Save
Mark Applied
Hide
Sr. Principal Software Engineer
Burlington or United States or Europe or Asia or North America
$141k-$226k/yr RemoteFull Time
Cerence
CerenceNASDAQ: CRNC: Develops AI-powered voice assistants and software for automotive vehicles.
Proven experience optimizing ML inference in production, deep GPU architecture knowledge, hands-on CUDA kernel development, quantization techniques (INT8/INT4/FP4/FP8/AWQ/GPTQ), and edge/embedded deployment expertise.
vLLM, TensorRT‑LLM, llama.cpp, QAIRT, CUDA, AWQ, GPTQ
2mo
Save
Mark Applied
Hide
Software Engineer, ML Infrastructure, Optimization
Mountain View, California, United States
$160k-$241k/yr OnsiteFull Time
Nuro
Nuro: Builds autonomous driving software and electric delivery robots.
2+ YOE2+ years in ML optimization infrastructure; experience with quantization, pruning, ML compilers and GPU runtimes; proficient in Python, C++, CUDA and deep learning frameworks (PyTorch, JAX, TensorFlow, Keras).
Python, C++, CUDA, PyTorch, JAX, TensorFlow, Keras, FTL
1w
Save
Mark Applied
Hide
Machine Learning Senior Software Engineer
United States
$121k-$219k/yr RemoteFull Time
Akamai
AkamaiNASDAQ: AKAM: Provides content delivery, cybersecurity, and cloud computing services globally.
5+ YOERequires 5 years of relevant experience, a computer science or machine learning degree, ML frameworks, quantization, LLMs, Python, production pipelines, and familiarity with safety, containers, CI/CD, and cloud infrastructure.
PyTorch, TensorFlow, JAX, GPTQ, AWQ, GGUF, Python, CI/CD, cloud infrastructure
2w
Save
Mark Applied
Hide
Senior Machine Learning Engineer
San Diego, California, United States
$180k-$250k/yr OnsiteFull Time
Seasats
Seasats: Builds autonomous surface vessels for persistent maritime intelligence.
7+ YOE7+ years ML experience (5+ on perception), strong Python/PyTorch, edge model optimization (quantization/pruning, TensorRT/ONNX Runtime), experience with sensor data and MLOps; must be U.S. person.
Python, PyTorch, OpenCV, TensorRT, ONNX Runtime