75 model deployment engineer jobs at 52 companies in Soquel, CA
1mo
Save
Mark Applied
Hide
1mo
ML Engineer - Inference & Model Deployment
Cupertino, California, United States
$250k-$310k/yrOnsiteFull Time
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
Applied Research Scientist / Engineer - Deployment
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
Strong ML research and engineering skills with hands-on experience fine-tuning or adapting large models; translate customer requirements into model adaptations; customer-facing applied research or solutions engineering experience; staff-level/ senior execution expectations.
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
0+ YOEIntern to support optimization and deployment of multimodal models onto vehicle-grade compute; strong CS/EE/ robotics background; hands-on C++/Python; DL frameworks; model optimization and edge deployment.
John DeereNYSE: DE: Manufactures agricultural, construction, and forestry machinery and equipment.
3+ YOE3+ years software engineering with modern C++, applied ML for perception, experience with sensor data pipelines, model training/deployment, and system-level debugging.
Applied MaterialsNASDAQ: AMAT: Produces equipment and services for chip and display manufacturing.
5+ YOEBachelor's in engineering/CS/data science; 5+ years (or 2+ with a Master's) in application engineering, algorithm development, or similar; experience in industrial/manufacturing environments; model lifecycle and production deployment; Python/R/C# and AI/ML familiarity.
P-1 AI: Developing AI agents for industrial engineering and physical design.
Experience shipping data-driven or AI systems to production (Python preferred); physical engineering background; building integrations, fine-tuning models, customer-facing deployment and troubleshooting experience.
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
Sonatus: Develops software platforms for AI-enabled software-defined vehicles.
10+ YOE10+ years ML engineering with 3+ years in Edge AI/embedded systems, Bachelor’s in CS/EE/Software Engineering, expert Python, C++14/17, PyTorch/TensorFlow, edge deployment and model optimization experience.
Vantaca: AI software for community association and HOA management.
8+ YOE8+ years in infrastructure/DevOps/SRE; strong cloud expertise; experience with CI/CD, PostgreSQL, Redis, APM, model serving, vector databases, GPU optimization, and LLM deployment.
PostgreSQL, Redis, APM, CI/CD, vector databases, model serving frameworks, LLM
MetaNASDAQ: META: Develops social networking platforms and virtual reality technologies.
Bachelor's in CS/CE or equivalent experience; experience shipping ML models to production; background in computer vision/ML with focus on gesture recognition/pose estimation; model optimization for on-device deployment; cross-functional collaboration.
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
8+ YOE4+ Mgmt8+ years software/ML engineering with 4+ years on-device/edge inference; production deployment of transformer/diffusion models; WebGPU/WGSL and GPU API performance tuning; proficiency with TypeScript/JavaScript and Python; leadership experience.
Fireworks AI: Provides high-performance generative AI model inference and deployment infrastructure.
5+ YOE5+ years in customer-facing technical engineering roles, strong Python and Kubernetes skills, experience with LLM inference, model serving and fine-tuning, cloud GPU deployment across major clouds, and exceptional communication.
Python, Kubernetes, vLLM, SGLang, TensorRT-LLM, AWS, Microsoft Azure, GCP, Azure AI Foundry, AWS Bedrock, SageMaker, GCP Vertex
Staff Software Engineer/ Tech Lead - Onboard Model Consolidation
Mountain View, California, United States
$251k-$310k/yrOnsiteFull Time
Waymo: Autonomous driving technology for ride-hailing and logistics.
8+ YOE8+ years professional software development; BS/MS in CS/EE/Robotics/related or equivalent experience; extensive C++ experience building large-scale, high-performance systems; leadership on cross-functional projects; ML deployment and inference expertise.
Senior Principal Engineer- MLOps & AI Machinery, ADAS/AV
Sunnyvale, California, United States
$240k-$320k/yrHybridFull Time
Bosch: Global manufacturer of automotive and industrial engineering technology.
10+ YOEMaster's or PhD in CS/Robotics/EE/AI, 10+ years in software/system engineering for autonomous driving or ADAS, experience releasing L2+ AI systems, knowledge of training pipelines, model optimization, and embedded deployment.
TensorFlow, PyTorch, Python, C++, MLOps, CICD, SIL, HIL
E-Space: Builds sustainable LEO satellite networks for global IoT connectivity
Plan and execute vibration, shock, acoustic, and thermal-vacuum tests for large deployable antenna structures; instrumentation, data acquisition, test procedures, anomaly disposition, and model-test correlation.
Sprinter Health: Mobile provider of in-home diagnostic and preventive healthcare services.
8+ YOE8+ years building production ML systems and infrastructure; experience with training/serving pipelines, feature pipelines, monitoring, deployment, cloud, containers, CI/CD, and model governance.