54 model deployment engineer jobs at 30 companies in Marina, CA
🚀PromotedHiringCafe
ML Engineer - Inference & Model Deployment
Cupertino, CA, US
$250k-$310k/yrOn-SiteFull Time
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Turn powerful AI and ML models into fast, reliable production systems. Own inference latency, throughput, model-serving architecture, multi-GPU systems, and production deployment for millions of users.
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
0+ YOEIntern to support optimization and deployment of multimodal models onto vehicle-grade compute; strong CS/EE/ robotics background; hands-on C++/Python; DL frameworks; model optimization and edge deployment.
John DeereNYSE: DE: Manufactures agricultural, construction, and forestry machinery and equipment.
3+ YOE3+ years software engineering with modern C++, applied ML for perception, experience with sensor data pipelines, model training/deployment, and system-level debugging.
Applied MaterialsNASDAQ: AMAT: Produces equipment and services for chip and display manufacturing.
5+ YOEBachelor's in engineering/CS/data science; 5+ years (or 2+ with a Master's) in application engineering, algorithm development, or similar; experience in industrial/manufacturing environments; model lifecycle and production deployment; Python/R/C# and AI/ML familiarity.
Applied MaterialsNASDAQ: AMAT: Manufacturers of equipment for semiconductor and display production.
2+ YOEBachelor's in engineering/computer science/data science required; 5+ years relevant experience (or 2+ with a Master’s). Strong analytics, AI/ML familiarity, Python/R/C# programming, model lifecycle and production deployment experience.
Sonatus: Develops software platforms for AI-enabled software-defined vehicles.
10+ YOE10+ years ML engineering with 3+ years in Edge AI/embedded systems, Bachelor’s in CS/EE/Software Engineering, expert Python, C++14/17, PyTorch/TensorFlow, edge deployment and model optimization experience.
Senior Principal Engineer- MLOps & AI Machinery, ADAS/AV
Sunnyvale, California, United States
$240k-$320k/yrHybridFull Time
Bosch: Global manufacturer of automotive and industrial engineering technology.
10+ YOEMaster's or PhD in CS/Robotics/EE/AI, 10+ years in software/system engineering for autonomous driving or ADAS, experience releasing L2+ AI systems, knowledge of training pipelines, model optimization, and embedded deployment.
TensorFlow, PyTorch, Python, C++, MLOps, CICD, SIL, HIL
E-Space: Builds sustainable LEO satellite networks for global IoT connectivity
Plan and execute vibration, shock, acoustic, and thermal-vacuum tests for large deployable antenna structures; instrumentation, data acquisition, test procedures, anomaly disposition, and model-test correlation.
Senior Machine Learning Engineer, Model Serving Infrastructure (Multiple Positions)
San Jose, California, United States
$265k-$388k/yrOnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOEMaster's (plus 2 years) or Bachelor's (plus 5 years) in a quantitative field; 2+ years coding in Python or C++; Linux development experience; ML, system design, and production deployment experience.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOE5+ years with LLMs/VLMs and large-scale inference; MSc/PhD or equivalent; strong Python and ML framework skills; experience in benchmarking, optimization, and production deployment.
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years working with LLMs/VLMs and large-scale inference; experience in fine-tuning, benchmarking, evaluation, optimization, and production deployment; strong Python and PyTorch/JAX/TensorFlow skills; familiarity with inference stacks and infrastructure.
Seattle or United States or Redwood City or Santa Clara or Austin
$126k-$264k/yrOnsiteFull Time
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
6+ YOE6+ years experience building and productionizing ML models; strong software engineering, model deployment, monitoring, data quality and debugging skills; stakeholder collaboration and mentoring experience.
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experience with foundation models, LLMs and multimodal LLMs; applied research and engineering in computer vision and machine perception; work across data collection, curation, modeling, evaluation, and deployment; collaborative cross-team communication.
CoStar GroupNASDAQ: CSGP: Provides global real estate information, analytics, and online marketplaces.
3+ YOEBachelor's in CS/Data Science/Engineering or equivalent, 3+ years ML engineering experience with model optimization and deployment, Python, TensorFlow/PyTorch, cloud (AWS/Azure/GCP), Git, strong communication and problem-solving skills.
KnightscopeNASDAQ: KSCP: Manufactures autonomous robots and communication devices for physical security.
10+ YOE10+ years in DevOps/SRE or platform engineering owning CI/CD, cloud (AWS), Docker, Kubernetes, Linux, GitHub Actions/GitLab CI; Python and C++ support; experience with OTA/fleet deployments and ML model rollout preferred.
CoStar GroupNASDAQ: CSGP: Global provider of commercial and residential real estate data.
3+ YOEBachelor's or equivalent experience, 3+ years ML engineering focused on model optimization/deployment, proficiency in Python, ML frameworks (TensorFlow/PyTorch), cloud (AWS/Azure/GCP), Git, and strong communication.
BetterHelp: Provides online therapy and professional mental health counseling services.
3+ YOE3+ years building ML systems; strong Python; experience with NLP, LLMs, PyTorch/TensorFlow, SQL; model evaluation, fine-tuning, inference optimization, and production deployment experience.