Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or United States
$124k-$242k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
3+ YOEBachelor's, master's, or doctoral degree in a related field or equivalent experience; 3+ years' engineering experience; GPU profiling, Python, C++/CUDA, and AI inference optimization expertise required.
Machine Learning Performance Engineer - Offboard Training & Inference
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$215k-$285k/yrOnsiteFull Time
Applied Intuition: Developing software and simulation infrastructure for autonomous vehicles.
ML performance engineering experience with distributed training, batch inference, GPU or accelerator optimization, Python, and C++ or another systems language; strong debugging and analytical skills required.
Modular: Unified software infrastructure and programming language for AI development.
5+ YOE5+ years in distributed systems or performance engineering; experience building reusable tooling; strong technical judgment, communication, and leadership; GPU/kernel, inference engine, Kubernetes, and LLM familiarity helpful.
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's or master's degree in a technical discipline; proficient in C/C++, Python, CUDA, GPU architecture, deep learning operators, inference compilation, performance analysis, and distributed model inference.
RadixArk: Building scalable open-source infrastructure for AI training and inference.
Strong systems engineering in performance-critical software; GPU/distributed systems; profiling tools; Python and C++; CUDA/Triton/ROCm/XLA familiarity; LLM inference concepts; ability to debug across software, hardware, and infra layers; strong communication.
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
8+ YOEBachelor's degree or equivalent experience, 8 years of software development, Python and C++, AI inference optimization expertise, and experience with serving codebases and performance tradeoffs.
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
8+ YOE4+ Mgmt8+ years software/ML engineering with 4+ years on-device/edge inference; production deployment of transformer/diffusion models; WebGPU/WGSL and GPU API performance tuning; proficiency with TypeScript/JavaScript and Python; leadership experience.
San Jose or Durham or Mexico City or Vancouver or Bengaluru or Pune or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
$171k-$257k/yrHybridFull Time
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
8+ YOE8+ years building distributed, high-performance systems; strong Go/Python, Docker, Kubernetes, CI/CD; knowledge of datacenter, OS internals, virtualization, and ML frameworks.
AI Infra Engineer - Large Model Inference Systems (Multimodal/LLM/VLM)
San Jose, California, United States
$156k-$388k/yrOnsiteFull Time
TikTok: Global short-form video hosting and social media platform.
2+ YOEBachelor's in CS or related,2+ years in high-performance computing or distributed scheduling,experience with large-model inference and system design,knowledge of asynchronous scheduling and resource pooling.
Member of Technical Staff, ML Inference Engineering
Palo Alto, California, United States
OnsiteFull Time
Sanas: Provides real-time speech transformation and accent translation software.
5+ YOERequires 5+ years writing high-performance code, NVIDIA GPU and CUDA expertise, LLM serving knowledge, and research or systems experience in language or speech inference. Production-scale systems experience preferred.
NIONYSE: NIO: Designs and manufactures premium smart electric vehicles and technology
5+ YOE5+ years building and optimizing large-scale LLM/VLM inference systems; strong C/C++ and performance engineering skills; GPU/NPU programming (CUDA), PyTorch/TensorFlow, and BS/MS in CS/CE or related field required.
Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Cupertino, California, United States
$165k-$224k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
3+ YOEBachelor's degree or equivalent; 3+ years professional software development and systems design experience; C++ or Python; machine learning, LLM, performance, memory, parallel computing, debugging, and profiling expertise.
[2026] Senior Machine Learning Engineer (Systems), Embodied AI/NPCs, ML Platform - PhD Early Career
San Mateo, California, United States
$197k-$243k/yrHybridFull Time
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
PhD (pursuing or completed) in a technical field; experience building end-to-end ML pipelines, model inference and deployment, distributed inference systems, Kubernetes and major cloud providers (AWS/Azure/GCP); strong systems and performance optimization skills.
Kubernetes, AWS, Azure, GCP, GPU, LLMs, Roblox Studio IDE