142 model deployment engineer jobs at 98 companies in Albany, CA

2mo
Save
Mark Applied
Hide
ML Engineer - Inference & Model Deployment
Cupertino, California, United States
$250k-$310k/yr OnsiteFull Time
Hiring.Cafe
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
vLLM, TensorRT, SGLang, GPU
2mo
Save
Mark Applied
Hide
AI Deployment Strategist
San Francisco, California, United States
HybridFull Time
Pigment
Pigment: AI-powered business planning and performance management platform.
2+ YOEEngineering or computer science degree, 2+ years in technical client-facing/implementation roles, experience in data modeling and AI/ML, proficiency with formulas and structured modeling, programming (Python/SQL), and APIs/data pipelines.
Pigment, Python, SQL
2mo
Save
Mark Applied
Hide
Lead AI Engineer (Vision model customization, VLM)
New York or Cambridge or McLean or San Jose
$197k-$246k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
4+ YOEBachelor's in CS/AI/EE/CE + 4 years AI/ML experience (or Master's + 2 years). 4+ years programming with Python/Go/Scala/Java. Experience with LLM inference, similarity search/VectorDBs, cloud deployments, PyTorch, and optimizing training/inference.
AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++, C#, Golang, AWS, Google Cloud, Azure, LLM
3mo
Save
Mark Applied
Hide
Applied Research Scientist / Engineer - Deployment
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
Strong ML research and engineering skills with hands-on experience fine-tuning or adapting large models; translate customer requirements into model adaptations; customer-facing applied research or solutions engineering experience; staff-level/ senior execution expectations.
1mo
Save
Mark Applied
Hide
Engineering Manager, Model Flywheel
San Francisco, California, United States
$293k-$385k/yr HybridFull Time
OpenAI
OpenAI: Develops artificial intelligence models and generative AI software services.
Proven experience leading engineering teams, shipping production systems at scale, and owning model deployment, experimentation, and measurement tooling; strong communication and cross-functional collaboration skills.
1mo
Save
Mark Applied
Hide
Staff Software Engineer- Foundation Model Inference
San Francisco or Mountain View
$190k-$265k/yr OnsiteFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
8+ YOE8+ years backend or infrastructure engineering experience; distributed systems, scalable APIs, real-time serving or ML/GPU orchestration experience; familiarity with service-oriented architecture, deployment pipelines, and observability.
OpenAI, Anthropic, Gemini, Qwen, GPT-OSS, Llama, SageMaker, Vertex AI, Azure ML, MLflow, PyTorch, Ray, vLLM, SGLang, Apache Spark, Delta Lake
2mo
Save
Mark Applied
Hide
Forward Deployed Robotics Engineer
Freiburg im Breisgau or San Francisco
€120k-€180k/yr HybridFull Time
Black Forest Labs
Black Forest Labs: Developing frontier generative AI models for visual intelligence.
Proven robotics/forward-deployed engineering experience, customer-facing integration skills, experience with action/VLA models, model deployment and latency optimization, strong communication and collaboration.
Latent Diffusion, Stable Diffusion, FLUX, π0/π0.5, LeRobot
2mo
Save
Mark Applied
Hide
Forward Deployed, Robotics Engineer
Freiburg im Breisgau or San Francisco or Europe
€120k-€180k/yr HybridFull Time
Black Forest Labs
Black Forest Labs: Develops generative AI models for image and video creation.
Proven robotics engineering experience; hands-on with action/VLA models (π0/π0.5, LeRobot); customer-facing deployments, model hosting, inference/edge constraints, and strong communication skills.
FLUX, Latent Diffusion, Stable Diffusion, π0, π0.5, LeRobot, VLA, transformer
1mo
Save
Mark Applied
Hide
Senior Perception Engineer
Santa Clara, California, United States
$125k-$187k/yr OnsiteFull Time
John Deere
John DeereNYSE: DE: Manufactures agricultural, construction, and forestry machinery and equipment.
3+ YOE3+ years software engineering with modern C++, applied ML for perception, experience with sensor data pipelines, model training/deployment, and system-level debugging.
C++, PyTorch, TensorFlow, ROS 2, clang-tidy, ASAN, TSAN, UBSAN, CMake, Bazel, colcon, Docker
3w
Save
Mark Applied
Hide
Infrastructure Engineer, Applied AI
San Francisco, California, United States
$250k-$400k/yr OnsiteFull Time
Paradigm
Paradigm: Venture capital firm focused on crypto and frontier technologies.
Experienced engineer with infra, security, and AI model experience; familiarity with durable control planes, runtimes, secrets, observability, and self-hosted deployment.
Reth, Foundry, EVMBench, OpenAI, Centaur, Kubernetes
2mo
Save
Mark Applied
Hide
Application Engineer
Santa Clara, California, United States
$100k-$137k/yr OnsiteFull Time
Applied Materials
Applied MaterialsNASDAQ: AMAT: Produces equipment and services for chip and display manufacturing.
5+ YOEBachelor's in engineering/CS/data science; 5+ years (or 2+ with a Master's) in application engineering, algorithm development, or similar; experience in industrial/manufacturing environments; model lifecycle and production deployment; Python/R/C# and AI/ML familiarity.
Python, R, C#
1w
Save
Mark Applied
Hide
Staff Machine Learning Engineer - LLM Quantization & Deployment
Santa Clara or Mountain View
$215k-$364k/yr OnsiteFull Time
XPeng
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
3+ YOEMaster's in CS, CE, or EE with 3–5 years' industry experience; expertise in Transformer architectures, LLM inference, model quantization, PyTorch, inference stacks, Python, and software engineering.
Python, PyTorch, TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, AWQ, GPTQ, SmoothQuant, INT8, FP4
2d
Save
Mark Applied
Hide
ML/AI Engineer
Santa Clara, California, United States
$175k-$300k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Bachelor’s or master’s degree in computer science, computer engineering, mathematics, or related field; expertise in machine learning, statistical analysis, algorithms, and production model deployment.
Python, JavaScript, C++, PyTorch, TensorFlow, NumPy, Flask, Angular, ASP.NET Core, SQL, NoSQL
2mo
Save
Mark Applied
Hide
Application Engineer
Santa Clara, California, United States
$100k-$137k/yr OnsiteFull Time
Applied Materials
Applied MaterialsNASDAQ: AMAT: Manufacturers of equipment for semiconductor and display production.
2+ YOEBachelor's in engineering/computer science/data science required; 5+ years relevant experience (or 2+ with a Master’s). Strong analytics, AI/ML familiarity, Python/R/C# programming, model lifecycle and production deployment experience.
Python, R, C#
3w
Save
Mark Applied
Hide
Senior Optimization Engineer
San Francisco, California, United States
$150k-$210k/yr HybridFull Time
Verse
Verse: Software to plan, procure, and manage clean energy portfolios.
5+ YOE5+ years building production optimization models, strong Python software engineering, deployment of scalable optimization services, energy industry experience (electricity markets, battery storage), Master's degree in quantitative field.
Python, CI/CD
1mo
Save
Mark Applied
Hide
Forward Deployed Engineer
San Mateo or United States or Canada
$170k-$220k/yr RemoteFull Time
P-1 AI
P-1 AI: Developing AI agents for industrial engineering and physical design.
Experience shipping data-driven or AI systems to production (Python preferred); physical engineering background; building integrations, fine-tuning models, customer-facing deployment and troubleshooting experience.
Python
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Inference
Palo Alto, California, United States
$185k-$250k/yr HybridFull Time
Pika
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
CUDA, NCCL
2mo
Save
Mark Applied
Hide
Lead AI Engineer (Vision model customization, VLM)
New York City or McLean or San Jose or Cambridge
$197k-$246k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: A diversified financial services providing banking and credit products.
4+ YOEBachelor’s degree plus 4 years or master’s degree plus 2 years in AI/ML development; 4 years programming with Python, Go, Scala, or Java; cloud AI deployment experience preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, NeMo Guardrails, PyTorch, Python, Go, Scala, Java, AWS, Google Cloud, Azure, C++, C#, Golang
1mo
Save
Mark Applied
Hide
Senior Staff AI Engineer, Edge AI
Sunnyvale, California, United States
$227k-$300k/yr HybridFull Time
Sonatus
Sonatus: Develops software platforms for AI-enabled software-defined vehicles.
10+ YOE10+ years ML engineering with 3+ years in Edge AI/embedded systems, Bachelor’s in CS/EE/Software Engineering, expert Python, C++14/17, PyTorch/TensorFlow, edge deployment and model optimization experience.
Python, C++14, C++17, PyTorch, TensorFlow, ONNX, TFLite, TVM, scikit-learn, tslearn, statsmodels, NVIDIA TensorRT, Qualcomm SNPE, Gemini, OpenAI, Claude, Linux, QNX, CAN, DBC, UDS, SOME/IP, MQTT, ARM
2mo
Save
Mark Applied
Hide
Applied AI Engineer
San Francisco, California, United States
$171k-$242k/yr HybridFull Time
Artos
Artos: AI-powered document authoring platform for life sciences R&D.
2+ YOE2+ years building and deploying AI/ML applications, hands-on LLM experience, backend engineering with Python, API development (FastAPI/Django), cloud container deployment, R&D on model capabilities, and evaluation tooling experience.
Langfuse, LangSmith, FastAPI, Django, AWS, GCP, Azure, Terraform, Pulumi, GitHub Actions, React, Python, OpenAI, Anthropic, Google