67 model deployment engineer jobs at 38 companies in Prunedale, CA

2mo
Save
Mark Applied
Hide
ML Engineer - Inference & Model Deployment
Cupertino, California, United States
$250k-$310k/yr OnsiteFull Time
Hiring.Cafe
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
vLLM, TensorRT, SGLang, GPU
2mo
Save
Mark Applied
Hide
Lead AI Engineer (Vision model customization, VLM)
New York or Cambridge or McLean or San Jose
$197k-$246k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
4+ YOEBachelor's in CS/AI/EE/CE + 4 years AI/ML experience (or Master's + 2 years). 4+ years programming with Python/Go/Scala/Java. Experience with LLM inference, similarity search/VectorDBs, cloud deployments, PyTorch, and optimizing training/inference.
AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++, C#, Golang, AWS, Google Cloud, Azure, LLM
2w
Save
Mark Applied
Hide
Staff Software Engineer- Foundation Model Inference
San Francisco or Mountain View
$190k-$265k/yr OnsiteFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
8+ YOE8+ years backend or infrastructure engineering experience; distributed systems, scalable APIs, real-time serving or ML/GPU orchestration experience; familiarity with service-oriented architecture, deployment pipelines, and observability.
OpenAI, Anthropic, Gemini, Qwen, GPT-OSS, Llama, SageMaker, Vertex AI, Azure ML, MLflow, PyTorch, Ray, vLLM, SGLang, Apache Spark, Delta Lake
3w
Save
Mark Applied
Hide
Senior Perception Engineer
Santa Clara, California, United States
$125k-$187k/yr OnsiteFull Time
John Deere
John DeereNYSE: DE: Manufactures agricultural, construction, and forestry machinery and equipment.
3+ YOE3+ years software engineering with modern C++, applied ML for perception, experience with sensor data pipelines, model training/deployment, and system-level debugging.
C++, PyTorch, TensorFlow, ROS 2, clang-tidy, ASAN, TSAN, UBSAN, CMake, Bazel, colcon, Docker
2mo
Save
Mark Applied
Hide
Application Engineer
Santa Clara, California, United States
$100k-$137k/yr OnsiteFull Time
Applied Materials
Applied MaterialsNASDAQ: AMAT: Produces equipment and services for chip and display manufacturing.
5+ YOEBachelor's in engineering/CS/data science; 5+ years (or 2+ with a Master's) in application engineering, algorithm development, or similar; experience in industrial/manufacturing environments; model lifecycle and production deployment; Python/R/C# and AI/ML familiarity.
Python, R, C#
2mo
Save
Mark Applied
Hide
Application Engineer
Santa Clara, California, United States
$100k-$137k/yr OnsiteFull Time
Applied Materials
Applied MaterialsNASDAQ: AMAT: Manufacturers of equipment for semiconductor and display production.
2+ YOEBachelor's in engineering/computer science/data science required; 5+ years relevant experience (or 2+ with a Master’s). Strong analytics, AI/ML familiarity, Python/R/C# programming, model lifecycle and production deployment experience.
Python, R, C#
3w
Save
Mark Applied
Hide
Senior Staff AI Engineer, Edge AI
Sunnyvale, California, United States
$227k-$300k/yr HybridFull Time
Sonatus
Sonatus: Develops software platforms for AI-enabled software-defined vehicles.
10+ YOE10+ years ML engineering with 3+ years in Edge AI/embedded systems, Bachelor’s in CS/EE/Software Engineering, expert Python, C++14/17, PyTorch/TensorFlow, edge deployment and model optimization experience.
Python, C++14, C++17, PyTorch, TensorFlow, ONNX, TFLite, TVM, scikit-learn, tslearn, statsmodels, NVIDIA TensorRT, Qualcomm SNPE, Gemini, OpenAI, Claude, Linux, QNX, CAN, DBC, UDS, SOME/IP, MQTT, ARM
2mo
Save
Mark Applied
Hide
ML Engineer, End-to-End Autonomy
Santa Clara, California, United States
$135k-$378k/yr HybridFull Time
Blue River Technology
Blue River Technology: Develops robotic computer vision systems for precision agriculture.
Proven ML model development, production deployment for robotics, Python and PyTorch expertise, and cross-functional collaboration.
Python, PyTorch, ROS, CUDA, GPU
4d
Save
Mark Applied
Hide
Distinguished AI Engineer (Hybrid)
San Jose or Austin or Seattle or New York City
$343k-$433k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
10+ YOEExtensive AI/ML leadership with deep learning, NLP, LLMs, large-scale model development and production deployment; strong communication and mentorship skills.
1mo
Save
Mark Applied
Hide
Principal Machine Learning Engineer
Mountain View, California, United States
$278k-$417k/yr OnsiteFull Time
Unity
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
8+ YOE4+ Mgmt8+ years software/ML engineering with 4+ years on-device/edge inference; production deployment of transformer/diffusion models; WebGPU/WGSL and GPU API performance tuning; proficiency with TypeScript/JavaScript and Python; leadership experience.
WebGPU, WebNN, WGSL, Metal, Vulkan, SPIR-V, D3D12, CUDA, Chrome, Dawn, PIX, Instruments, Snapdragon Profiler, Nsight, RenderDoc, ONNX Runtime Web, ONNX Runtime, Transformers.js, WebLLM, TensorFlow.js, CoreML, TFLite, ExecuTorch, TypeScript, JavaScript, Python, MLIR, TVM, IREE, XLA, wgpu
1mo
Save
Mark Applied
Hide
Staff Software Engineer/ Tech Lead - Onboard Model Consolidation
Mountain View, California, United States
$251k-$310k/yr OnsiteFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
8+ YOE8+ years professional software development; BS/MS in CS/EE/Robotics/related or equivalent experience; extensive C++ experience building large-scale, high-performance systems; leadership on cross-functional projects; ML deployment and inference expertise.
C++, TPU, GPU
1mo
Save
Mark Applied
Hide
Senior Principal Engineer- MLOps & AI Machinery, ADAS/AV
Sunnyvale, California, United States
$240k-$320k/yr HybridFull Time
Bosch
Bosch: Global manufacturer of automotive and industrial engineering technology.
10+ YOEMaster's or PhD in CS/Robotics/EE/AI, 10+ years in software/system engineering for autonomous driving or ADAS, experience releasing L2+ AI systems, knowledge of training pipelines, model optimization, and embedded deployment.
TensorFlow, PyTorch, Python, C++, MLOps, CICD, SIL, HIL
4w
Save
Mark Applied
Hide
Environmental Test Engineer
Saratoga or Arlington
$140k-$210k/yr OnsiteFull Time
E-Space
E-Space: Builds sustainable LEO satellite networks for global IoT connectivity
Plan and execute vibration, shock, acoustic, and thermal-vacuum tests for large deployable antenna structures; instrumentation, data acquisition, test procedures, anomaly disposition, and model-test correlation.
Nastran, ANSYS
3w
Save
Mark Applied
Hide
Lead AI Engineer -- Advanced AI (applied ML, LLMs, agentic AI, ML Ops)
Brooklyn Park or Sunnyvale
$132k-$286k/yr HybridFull Time
Target
TargetNYSE: TGT: General merchandise retailer operating physical stores and e-commerce.
5+ YOEDegree in quantitative field or equivalent experience,5+ years applied ML/AI experience,experience with LLMs,agentic systems,model deployment,software engineering practices and strong communication.
Python, PyTorch, TensorFlow, LangChain, LlamaIndex, Semantic Kernel
1w
Save
Mark Applied
Hide
Principal Engineer, Local AI - Agents and Systems
Santa Clara or Redmond
$272k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
15+ YOE3+ Mgmt15+ years software engineering experience, deep Windows internals and security, LLM inference and GPU acceleration experience, proficiency in C++ and Python, experience with agent frameworks and local model deployment.
Nemoclaw, OpenClaw, Nemotron, Ollama, Llama.cpp, vLLM, CUDA, TensorRT, Hermes, LangChain, Windows, C++, Python
1mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Model Serving Infrastructure (Multiple Positions)
San Jose, California, United States
$265k-$388k/yr OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOEMaster's (plus 2 years) or Bachelor's (plus 5 years) in a quantitative field; 2+ years coding in Python or C++; Linux development experience; ML, system design, and production deployment experience.
Python, C++, Linux
2w
Save
Mark Applied
Hide
Principal Machine Learning Engineer
Seattle or United States or Redwood City or Santa Clara or Austin
$126k-$264k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
6+ YOE6+ years experience building and productionizing ML models; strong software engineering, model deployment, monitoring, data quality and debugging skills; stakeholder collaboration and mentoring experience.
PyTorch, TensorFlow, Keras
1d
Save
Mark Applied
Hide
Builder - Senior Software Engineer, AI
Santa Clara, California, United States
OnsiteFull Time
Reevo
Reevo: AI-native revenue operating system for go-to-market teams.
5+ YOE5+ years of software engineering experience focused on AI/ML systems, with expertise in LLMs, RAG, AI frameworks, model deployment, monitoring, prompt engineering, evaluation, vector databases, and semantic search.
RAG, large language models, AI/ML, ML operations, vector databases, embedding systems, semantic search
2mo
Save
Mark Applied
Hide
Matterport - Senior ML Ops Engineer
Sunnyvale, California, United States
$173k-$253k/yr HybridFull Time
CoStar Group
CoStar GroupNASDAQ: CSGP: Provides global real estate information, analytics, and online marketplaces.
3+ YOEBachelor's in CS/Data Science/Engineering or equivalent, 3+ years ML engineering experience with model optimization and deployment, Python, TensorFlow/PyTorch, cloud (AWS/Azure/GCP), Git, strong communication and problem-solving skills.
Python, TensorFlow, PyTorch, AWS, Azure, GCP, Git, Temporal, Airflow, Kubeflow, Docker, Kubernetes
1mo
Save
Mark Applied
Hide
Principal Perception Engineer, Obstacle Foundation Models - Autonomous Vehicles
Santa Clara, California, United States
$272k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
15+ YOE15+ years developing deep-learning perception systems, technical leadership experience, proficiency in PyTorch, Python and/or C++, EM/deployment experience, strong communication and collaboration, BS/MS/PhD in CS/EE or equivalent.
PyTorch, Python, C++, CUDA, LoRA