64 model deployment engineer jobs at 45 companies in Watsonville, CA

1mo
Save
Mark Applied
Hide
ML Engineer - Inference & Model Deployment
Cupertino, California, United States
$250k-$310k/yr OnsiteFull Time
Hiring.Cafe
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
vLLM, TensorRT, SGLang, GPU
1mo
Save
Mark Applied
Hide
Lead AI Engineer (Vision model customization, VLM)
New York or Cambridge or McLean or San Jose
$197k-$246k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
4+ YOEBachelor's in CS/AI/EE/CE + 4 years AI/ML experience (or Master's + 2 years). 4+ years programming with Python/Go/Scala/Java. Experience with LLM inference, similarity search/VectorDBs, cloud deployments, PyTorch, and optimizing training/inference.
AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++, C#, Golang, AWS, Google Cloud, Azure, LLM
2mo
Save
Mark Applied
Hide
Applied Research Scientist / Engineer - Deployment
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
Strong ML research and engineering skills with hands-on experience fine-tuning or adapting large models; translate customer requirements into model adaptations; customer-facing applied research or solutions engineering experience; staff-level/ senior execution expectations.
2mo
Save
Mark Applied
Hide
AI Intern – VLA Deployment
Santa Clara or Mountain View
OnsiteInternship
XPeng
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
0+ YOEIntern to support optimization and deployment of multimodal models onto vehicle-grade compute; strong CS/EE/ robotics background; hands-on C++/Python; DL frameworks; model optimization and edge deployment.
C++, Python, PyTorch, ONNX, TensorRT, CUDA
1w
Save
Mark Applied
Hide
Staff Software Engineer- Foundation Model Inference
San Francisco or Mountain View
$190k-$265k/yr OnsiteFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
8+ YOE8+ years backend or infrastructure engineering experience; distributed systems, scalable APIs, real-time serving or ML/GPU orchestration experience; familiarity with service-oriented architecture, deployment pipelines, and observability.
OpenAI, Anthropic, Gemini, Qwen, GPT-OSS, Llama, SageMaker, Vertex AI, Azure ML, MLflow, PyTorch, Ray, vLLM, SGLang, Apache Spark, Delta Lake
2w
Save
Mark Applied
Hide
Senior Perception Engineer
Santa Clara, California, United States
$125k-$187k/yr OnsiteFull Time
John Deere
John DeereNYSE: DE: Manufactures agricultural, construction, and forestry machinery and equipment.
3+ YOE3+ years software engineering with modern C++, applied ML for perception, experience with sensor data pipelines, model training/deployment, and system-level debugging.
C++, PyTorch, TensorFlow, ROS 2, clang-tidy, ASAN, TSAN, UBSAN, CMake, Bazel, colcon, Docker
1mo
Save
Mark Applied
Hide
Application Engineer
Santa Clara, California, United States
$100k-$137k/yr OnsiteFull Time
Applied Materials
Applied MaterialsNASDAQ: AMAT: Produces equipment and services for chip and display manufacturing.
5+ YOEBachelor's in engineering/CS/data science; 5+ years (or 2+ with a Master's) in application engineering, algorithm development, or similar; experience in industrial/manufacturing environments; model lifecycle and production deployment; Python/R/C# and AI/ML familiarity.
Python, R, C#
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Inference
Palo Alto, California, United States
$185k-$250k/yr HybridFull Time
Pika
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
CUDA, NCCL
2w
Save
Mark Applied
Hide
Senior Staff AI Engineer, Edge AI
Sunnyvale, California, United States
$227k-$300k/yr HybridFull Time
Sonatus
Sonatus: Develops software platforms for AI-enabled software-defined vehicles.
10+ YOE10+ years ML engineering with 3+ years in Edge AI/embedded systems, Bachelor’s in CS/EE/Software Engineering, expert Python, C++14/17, PyTorch/TensorFlow, edge deployment and model optimization experience.
Python, C++14, C++17, PyTorch, TensorFlow, ONNX, TFLite, TVM, scikit-learn, tslearn, statsmodels, NVIDIA TensorRT, Qualcomm SNPE, Gemini, OpenAI, Claude, Linux, QNX, CAN, DBC, UDS, SOME/IP, MQTT, ARM
1mo
Save
Mark Applied
Hide
Infrastructure Engineer
Redwood City, California, United States
HybridFull Time
Vantaca
Vantaca: AI software for community association and HOA management.
8+ YOE8+ years in infrastructure/DevOps/SRE; strong cloud expertise; experience with CI/CD, PostgreSQL, Redis, APM, model serving, vector databases, GPU optimization, and LLM deployment.
PostgreSQL, Redis, APM, CI/CD, vector databases, model serving frameworks, LLM
2mo
Save
Mark Applied
Hide
ML Engineer, End-to-End Autonomy
Santa Clara, California, United States
$135k-$378k/yr HybridFull Time
Blue River Technology
Blue River Technology: Develops robotic computer vision systems for precision agriculture.
Proven ML model development, production deployment for robotics, Python and PyTorch expertise, and cross-functional collaboration.
Python, PyTorch, ROS, CUDA, GPU
1mo
Save
Mark Applied
Hide
Principal Machine Learning Engineer
Mountain View, California, United States
$278k-$417k/yr OnsiteFull Time
Unity
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
8+ YOE4+ Mgmt8+ years software/ML engineering with 4+ years on-device/edge inference; production deployment of transformer/diffusion models; WebGPU/WGSL and GPU API performance tuning; proficiency with TypeScript/JavaScript and Python; leadership experience.
WebGPU, WebNN, WGSL, Metal, Vulkan, SPIR-V, D3D12, CUDA, Chrome, Dawn, PIX, Instruments, Snapdragon Profiler, Nsight, RenderDoc, ONNX Runtime Web, ONNX Runtime, Transformers.js, WebLLM, TensorFlow.js, CoreML, TFLite, ExecuTorch, TypeScript, JavaScript, Python, MLIR, TVM, IREE, XLA, wgpu
1mo
Save
Mark Applied
Hide
Staff Software Engineer/ Tech Lead - Onboard Model Consolidation
Mountain View, California, United States
$251k-$310k/yr OnsiteFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
8+ YOE8+ years professional software development; BS/MS in CS/EE/Robotics/related or equivalent experience; extensive C++ experience building large-scale, high-performance systems; leadership on cross-functional projects; ML deployment and inference expertise.
C++, TPU, GPU
1mo
Save
Mark Applied
Hide
Senior Principal Engineer- MLOps & AI Machinery, ADAS/AV
Sunnyvale, California, United States
$240k-$320k/yr HybridFull Time
Bosch
Bosch: Global manufacturer of automotive and industrial engineering technology.
10+ YOEMaster's or PhD in CS/Robotics/EE/AI, 10+ years in software/system engineering for autonomous driving or ADAS, experience releasing L2+ AI systems, knowledge of training pipelines, model optimization, and embedded deployment.
TensorFlow, PyTorch, Python, C++, MLOps, CICD, SIL, HIL
3w
Save
Mark Applied
Hide
Environmental Test Engineer
Saratoga or Arlington
$140k-$210k/yr OnsiteFull Time
E-Space
E-Space: Builds sustainable LEO satellite networks for global IoT connectivity
Plan and execute vibration, shock, acoustic, and thermal-vacuum tests for large deployable antenna structures; instrumentation, data acquisition, test procedures, anomaly disposition, and model-test correlation.
Nastran, ANSYS
1w
Save
Mark Applied
Hide
Machine Learning Engineer (Staff)
San Francisco or Menlo Park
$220k-$270k/yr HybridFull Time
Sprinter Health
Sprinter Health: Mobile provider of in-home diagnostic and preventive healthcare services.
8+ YOE8+ years building production ML systems and infrastructure; experience with training/serving pipelines, feature pipelines, monitoring, deployment, cloud, containers, CI/CD, and model governance.
CI/CD, APIs, containers, feature stores, MLOps, LLM
3w
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Model Serving Infrastructure (Multiple Positions)
San Jose, California, United States
$265k-$388k/yr OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOEMaster's (plus 2 years) or Bachelor's (plus 5 years) in a quantitative field; 2+ years coding in Python or C++; Linux development experience; ML, system design, and production deployment experience.
Python, C++, Linux
5d
Save
Mark Applied
Hide
Principal Engineer, Local AI - Agents and Systems
Santa Clara or Redmond
$272k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
15+ YOE3+ Mgmt15+ years software engineering experience, deep Windows internals and security, LLM inference and GPU acceleration experience, proficiency in C++ and Python, experience with agent frameworks and local model deployment.
Nemoclaw, OpenClaw, Nemotron, Ollama, Llama.cpp, vLLM, CUDA, TensorRT, Hermes, LangChain, Windows, C++, Python
1mo
Save
Mark Applied
Hide
Physical AI Engineer – Simulation & Synthetic Data
Fremont or United States
HybridFull Time
TD SYNNEX
TD SYNNEXNYSE: SNX: Distributes information technology products and provides supply chain services.
5+ YOE5+ years in robotics/RL/simulation/applied ML; hands‑on RL/imitation learning, simulation agent training, NVIDIA Omniverse/Isaac Sim/USD experience preferred; foundation model integration; strong Python and deployment experience.
NVIDIA Omniverse, Isaac Sim, USD, GPT, Claude, Opus, Python
3w
Save
Mark Applied
Hide
Machine Learning Engineer
Palo Alto, California, United States
$165k-$185k/yr RemoteFull Time
Allocate
Allocate: A that provides a data-rich platform for discovering, modeling, and managing private market investments.
4+ YOE4+ years applied ML/LLM engineering with strong Python, production model deployment, model fine-tuning and evaluation, Git proficiency, and experience with AI-assisted development tools.
Python, Git, Claude Code, OpenAI's Codex, Cursor

Explore Jobs

Expand Your Job Search