40 quantization engineer jobs at 31 companies in United States
6d
Save
Mark Applied
Hide
6d
Staff Machine Learning Engineer – Model Optimization & Quantization
Santa Clara, California, United States
$161k-$241k/yrOnsiteFull Time
QualcommNASDAQ: QCOM: Designs and manufactures semiconductors and wireless telecommunications products.
4+ YOEBachelor's and 4+ years, master's and 3+ years, or PhD and 2+ years in engineering or related work. Requires Python, ML frameworks, quantization expertise, and software development experience.
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
3+ YOEMaster's in CS, CE, or EE with 3–5 years' industry experience; expertise in Transformer architectures, LLM inference, model quantization, PyTorch, inference stacks, Python, and software engineering.
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
5+ YOE5+ yrs software engineering; 2+ yrs ML model training/deploy; NLP/LLMs; data pipelines; model optimizations; agentic AI systems; end-to-end ownership
NLP, LLMs, data pipelines, quantization, distillation, LoRA, pruning, agentic systems
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
Hark: A building multimodal AI models and next-generation hardware to create natural human-machine interfaces.
3+ YOE3+ years in model compression/distillation/quantization, strong fluency in PyTorch or TensorFlow, experience with PTQ/QAT and int8 conversion, hardware-aware optimization for constrained devices, and familiarity with audio/sequence model architectures.
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
RivianNASDAQ: RIVN: Designs and manufactures electric vehicles and charging networks.
Ph.D. or M.S. in a related field; hands-on experience deploying quantized models, ML compilers, and code generation for embedded/heterogeneous systems; strong CV model optimization skills; proficiency with PyTorch, TensorFlow, ONNX, C++, Python, and CUDA/OpenCL.
Figure: Develops autonomous humanoid robots for commercial and residential tasks.
8+ YOEMS/PhD or equivalent, 8+ years in hardware acceleration/ML systems, expertise in inference runtimes, quantization and pruning, profiling and benchmarking, model-to-hardware mapping, and strong C++/Python skills.
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
7+ YOEFormal software engineering training or certification and 7+ years of experience, with hands-on LLM inference, GPU infrastructure, quantization, benchmarking, cloud systems, and agentic AI development expertise.
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
Extensive ML engineering experience with LLMs and multimodal models; expertise in training and inference optimization, distributed training, GPU/accelerator optimization, KV cache/batching/quantization; proficient in PyTorch; strong communication and technical leadership.
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
5+ YOE5+ years in software/ML engineering with on-device or performance-critical systems; production deployment of transformer/diffusion models on-device; experience with inference runtimes, quantization, operator fusion, and GPU/compute APIs; strong Python; communication and mentoring skills.
Zensors: AI platform that turns existing cameras into intelligent sensors.
BS/MS/PhD in CS or EE; strong C/C++ and Python; model optimization, quantization, pruning; GPU performance tuning; profiling tools; cross-functional collaboration.
Ultralytics: Developing open-source computer vision models and AI platforms.
5+ YOE5+ years in computer vision and deep learning with architecture design experience; expert Python and PyTorch; experience reproducing papers, efficiency research (quantization/pruning/distillation), distributed training and strong research portfolio.
Burlington or United States or Europe or Asia or North America
$141k-$226k/yrRemoteFull Time
CerenceNASDAQ: CRNC: Develops AI-powered voice assistants and software for automotive vehicles.
Proven experience optimizing ML inference in production, deep GPU architecture knowledge, hands-on CUDA kernel development, quantization techniques (INT8/INT4/FP4/FP8/AWQ/GPTQ), and edge/embedded deployment expertise.
Software Engineer, ML Infrastructure, Optimization
Mountain View, California, United States
$160k-$241k/yrOnsiteFull Time
Nuro: Builds autonomous driving software and electric delivery robots.
2+ YOE2+ years in ML optimization infrastructure; experience with quantization, pruning, ML compilers and GPU runtimes; proficient in Python, C++, CUDA and deep learning frameworks (PyTorch, JAX, TensorFlow, Keras).
5+ YOERequires 5 years of relevant experience, a computer science or machine learning degree, ML frameworks, quantization, LLMs, Python, production pipelines, and familiarity with safety, containers, CI/CD, and cloud infrastructure.
Seasats: Builds autonomous surface vessels for persistent maritime intelligence.
7+ YOE7+ years ML experience (5+ on perception), strong Python/PyTorch, edge model optimization (quantization/pruning, TensorRT/ONNX Runtime), experience with sensor data and MLOps; must be U.S. person.