88 model deployment engineer jobs at 55 companies in Live Oak, CA

PromotedHiringCafe
ML Engineer - Inference & Model Deployment
Cupertino, CA, US
$250k-$310k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Turn powerful AI and ML models into fast, reliable production systems. Own inference latency, throughput, model-serving architecture, multi-GPU systems, and production deployment for millions of users.
Python, PyTorch, vLLM, SGLang, TensorRT, LLMs
3mo
Save
Mark Applied
Hide
AI Inference Engineer - Model Optimization & Deployment
Foster City or San Diego or Seattle
$242k-$290k/yr HybridFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
Seeking an engineer to optimize production-ready large-scale models for edge deployment on autonomous vehicle hardware.
CUDA, TensorRT, TensorRT-LLM, PyTorch, Python, C++
1mo
Save
Mark Applied
Hide
ML Engineer - Inference & Model Deployment
Cupertino, California, United States
$250k-$310k/yr OnsiteFull Time
Hiring.Cafe
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
vLLM, TensorRT, SGLang, GPU
1mo
Save
Mark Applied
Hide
Lead AI Engineer (Vision model customization, VLM)
New York or Cambridge or McLean or San Jose
$197k-$246k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
4+ YOEBachelor's in CS/AI/EE/CE + 4 years AI/ML experience (or Master's + 2 years). 4+ years programming with Python/Go/Scala/Java. Experience with LLM inference, similarity search/VectorDBs, cloud deployments, PyTorch, and optimizing training/inference.
AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++, C#, Golang, AWS, Google Cloud, Azure, LLM
2mo
Save
Mark Applied
Hide
Applied Research Scientist / Engineer - Deployment
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
Strong ML research and engineering skills with hands-on experience fine-tuning or adapting large models; translate customer requirements into model adaptations; customer-facing applied research or solutions engineering experience; staff-level/ senior execution expectations.
2mo
Save
Mark Applied
Hide
AI Intern – VLA Deployment
Santa Clara or Mountain View
OnsiteInternship
XPeng
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
0+ YOEIntern to support optimization and deployment of multimodal models onto vehicle-grade compute; strong CS/EE/ robotics background; hands-on C++/Python; DL frameworks; model optimization and edge deployment.
C++, Python, PyTorch, ONNX, TensorRT, CUDA
3d
Save
Mark Applied
Hide
Senior Perception Engineer
Santa Clara, California, United States
$125k-$187k/yr OnsiteFull Time
John Deere
John DeereNYSE: DE: Manufactures agricultural, construction, and forestry machinery and equipment.
3+ YOE3+ years software engineering with modern C++, applied ML for perception, experience with sensor data pipelines, model training/deployment, and system-level debugging.
C++, PyTorch, TensorFlow, ROS 2, clang-tidy, ASAN, TSAN, UBSAN, CMake, Bazel, colcon, Docker
1mo
Save
Mark Applied
Hide
Application Engineer
Santa Clara, California, United States
$100k-$137k/yr OnsiteFull Time
Applied Materials
Applied MaterialsNASDAQ: AMAT: Produces equipment and services for chip and display manufacturing.
5+ YOEBachelor's in engineering/CS/data science; 5+ years (or 2+ with a Master's) in application engineering, algorithm development, or similar; experience in industrial/manufacturing environments; model lifecycle and production deployment; Python/R/C# and AI/ML familiarity.
Python, R, C#
1mo
Save
Mark Applied
Hide
Application Engineer
Santa Clara, California, United States
$100k-$137k/yr OnsiteFull Time
Applied Materials
Applied MaterialsNASDAQ: AMAT: Manufacturers of equipment for semiconductor and display production.
2+ YOEBachelor's in engineering/computer science/data science required; 5+ years relevant experience (or 2+ with a Master’s). Strong analytics, AI/ML familiarity, Python/R/C# programming, model lifecycle and production deployment experience.
Python, R, C#
2w
Save
Mark Applied
Hide
Forward Deployed Engineer
San Mateo or United States or Canada
$170k-$220k/yr RemoteFull Time
P-1 AI
P-1 AI: Developing AI agents for industrial engineering and physical design.
Experience shipping data-driven or AI systems to production (Python preferred); physical engineering background; building integrations, fine-tuning models, customer-facing deployment and troubleshooting experience.
Python
3w
Save
Mark Applied
Hide
Senior Software Engineer, Inference
Palo Alto, California, United States
$185k-$250k/yr HybridFull Time
Pika
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
CUDA, NCCL
6d
Save
Mark Applied
Hide
Senior Staff AI Engineer, Edge AI
Sunnyvale, California, United States
$227k-$300k/yr HybridFull Time
Sonatus
Sonatus: Develops software platforms for AI-enabled software-defined vehicles.
10+ YOE10+ years ML engineering with 3+ years in Edge AI/embedded systems, Bachelor’s in CS/EE/Software Engineering, expert Python, C++14/17, PyTorch/TensorFlow, edge deployment and model optimization experience.
Python, C++14, C++17, PyTorch, TensorFlow, ONNX, TFLite, TVM, scikit-learn, tslearn, statsmodels, NVIDIA TensorRT, Qualcomm SNPE, Gemini, OpenAI, Claude, Linux, QNX, CAN, DBC, UDS, SOME/IP, MQTT, ARM
1mo
Save
Mark Applied
Hide
Infrastructure Engineer
Redwood City, California, United States
HybridFull Time
Vantaca
Vantaca: AI software for community association and HOA management.
8+ YOE8+ years in infrastructure/DevOps/SRE; strong cloud expertise; experience with CI/CD, PostgreSQL, Redis, APM, model serving, vector databases, GPU optimization, and LLM deployment.
PostgreSQL, Redis, APM, CI/CD, vector databases, model serving frameworks, LLM
1mo
Save
Mark Applied
Hide
ML Engineer, End-to-End Autonomy
Santa Clara, California, United States
$135k-$378k/yr HybridFull Time
Blue River Technology
Blue River Technology: Develops robotic computer vision systems for precision agriculture.
Proven ML model development, production deployment for robotics, Python and PyTorch expertise, and cross-functional collaboration.
Python, PyTorch, ROS, CUDA, GPU
1w
Save
Mark Applied
Hide
Computer Vision Engineer
Burlingame, California, United States
$184k-$257k/yr OnsiteFull Time
Meta
MetaNASDAQ: META: Develops social networking platforms and virtual reality technologies.
Bachelor's in CS/CE or equivalent experience; experience shipping ML models to production; background in computer vision/ML with focus on gesture recognition/pose estimation; model optimization for on-device deployment; cross-functional collaboration.
3w
Save
Mark Applied
Hide
Principal Machine Learning Engineer
Mountain View, California, United States
$278k-$417k/yr OnsiteFull Time
Unity
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
8+ YOE4+ Mgmt8+ years software/ML engineering with 4+ years on-device/edge inference; production deployment of transformer/diffusion models; WebGPU/WGSL and GPU API performance tuning; proficiency with TypeScript/JavaScript and Python; leadership experience.
WebGPU, WebNN, WGSL, Metal, Vulkan, SPIR-V, D3D12, CUDA, Chrome, Dawn, PIX, Instruments, Snapdragon Profiler, Nsight, RenderDoc, ONNX Runtime Web, ONNX Runtime, Transformers.js, WebLLM, TensorFlow.js, CoreML, TFLite, ExecuTorch, TypeScript, JavaScript, Python, MLIR, TVM, IREE, XLA, wgpu
1mo
Save
Mark Applied
Hide
Staff Software Engineer/ Tech Lead - Onboard Model Consolidation
Mountain View, California, United States
$251k-$310k/yr OnsiteFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
8+ YOE8+ years professional software development; BS/MS in CS/EE/Robotics/related or equivalent experience; extensive C++ experience building large-scale, high-performance systems; leadership on cross-functional projects; ML deployment and inference expertise.
C++, TPU, GPU
1w
Save
Mark Applied
Hide
ML Engineer – Robotics
San Francisco or Mountain View
$220k-$300k/yr OnsiteFull Time
Clera
Clera: AI talent agent matching professionals with high-growth startup roles
3+ YOE3+ years ML/robotics experience; proficiency in Python and C++; ROS/ROS2 experience; sensor fusion and simulation experience; bachelor's degree or equivalent; proven production deployment of learning-based models.
Python, C++, ROS, ROS2, Gazebo, Isaac Sim, CARLA, MuJoCo, PyBullet
1mo
Save
Mark Applied
Hide
Senior Principal Engineer- MLOps & AI Machinery, ADAS/AV
Sunnyvale, California, United States
$240k-$320k/yr HybridFull Time
Bosch
Bosch: Global manufacturer of automotive and industrial engineering technology.
10+ YOEMaster's or PhD in CS/Robotics/EE/AI, 10+ years in software/system engineering for autonomous driving or ADAS, experience releasing L2+ AI systems, knowledge of training pipelines, model optimization, and embedded deployment.
TensorFlow, PyTorch, Python, C++, MLOps, CICD, SIL, HIL
1w
Save
Mark Applied
Hide
Environmental Test Engineer
Saratoga or Arlington
$140k-$210k/yr OnsiteFull Time
E-Space
E-Space: Builds sustainable LEO satellite networks for global IoT connectivity
Plan and execute vibration, shock, acoustic, and thermal-vacuum tests for large deployable antenna structures; instrumentation, data acquisition, test procedures, anomaly disposition, and model-test correlation.
Nastran, ANSYS
4h
Save
Mark Applied
Hide
Machine Learning Engineer (Staff)
San Francisco or Menlo Park
$220k-$270k/yr HybridFull Time
Sprinter Health
Sprinter Health: Mobile provider of in-home diagnostic and preventive healthcare services.
8+ YOE8+ years building production ML systems and infrastructure; experience with training/serving pipelines, feature pipelines, monitoring, deployment, cloud, containers, CI/CD, and model governance.
CI/CD, APIs, containers, feature stores, MLOps, LLM