141 optimization modeling jobs at 46 companies in Marina, CA

1mo
Save
Mark Applied
Hide
Inference Optimization Intern – Performance Modeling
Sunnyvale, California, United States
OnsiteInternship
Institute of Foundation Models
Institute of Foundation Models: Develops open-source frontier-class AI foundation models and research.
Currently pursuing a quantitative degree; experience or coursework in CUDA, GPU kernel development, performance modeling, Nsight profiling, C++, Python, and deep learning frameworks preferred.
CUDA, Nsight Systems, Nsight Compute, PTX, SASS, PyTorch, TensorFlow, C++, Python
3mo
Save
Mark Applied
Hide
Senior Optimization Engineer
Santa Clara, California, United States
HybridFull Time
Hitachi
HitachiTokyo Stock Exchange: 6501: Global provider of digital systems and social infrastructure solutions.
5+ YOEMinimum 5 years in optimization modeling tools; 3+ years programming; 3+ years Linux experience; 2+ years developing power system software; English proficiency; US work authorization.
AMPL, AIMMS, CPLEX, Gurobi, Linux, AIX
2mo
Save
Mark Applied
Hide
Principal AI Performance Modeling Architect
Santa Clara or Austin
$203k-$348k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Extensive, senior experience optimizing large-scale ML systems and GPU architectures; CUDA programming; memory hierarchies; distributed training; transformer models.
PyTorch, CUDA, TensorRT, OpenAI Triton, Ray, Megatron-LM, NSight Compute, nvprof, PyTorch Profiler, KV cache optimization, Flash Attention, InfiniBand, RDMA, NVLink
2w
Save
Mark Applied
Hide
Multimodal Model Training and Inference Optimization Engineer
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
M.S. or PhD in CS/EE/AI, experience optimizing model training and inference, proficiency in Python, C++, CUDA, PyTorch, Megatron, Deepspeed, distributed training, and knowledge of transformers/diffusion models.
Python, C++, CUDA, PyTorch, Megatron, Deepspeed
2w
Save
Mark Applied
Hide
Multimodal Model Training and Inference Optimization Engineer
San Jose, California, United States
$156k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
MS/PhD in CS/EE/AI or related, experience optimizing model training and inference, distributed training, Python/C++/CUDA, PyTorch/Megatron/Deepspeed, knowledge of transformers and diffusion models.
Python, C++, CUDA, PyTorch, Megatron, Deepspeed
3mo
Save
Mark Applied
Hide
SoC Power Analysis and Optimization Engineer
Beaverton or Cupertino or San Diego
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
3+ YOEBachelor's degree and 3+ years in SOC power analysis, optimization, and modeling; ML, Python, Verilog/SystemVerilog; ASIC/SOC design experience.
Python, Verilog, SystemVerilog, Machine Learning, Power modeling, Power analysis
2w
Save
Mark Applied
Hide
Solutions Architect, Agentic Optimization
Santa Clara or United States
$152k-$242k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years AI/software engineering experience, strong Python/C++ coding, GPU model training/optimization, communication and customer-facing skills, BS/MS/PhD or equivalent.
Python, C++, GitHub, Kubernetes, containers
2w
Save
Mark Applied
Hide
Solutions Architect, Agentic Optimization
Santa Clara or United States
$152k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOEDegree in CS/EE/Physics/Math or equivalent experience,5+ years AI/software engineering,proficient in Python/C++,experience profiling and optimizing models on GPUs and developing GPU kernels,strong communication skills.
Python, C++, GitHub, Kubernetes, containers, GPUs
1mo
Save
Mark Applied
Hide
Machine Learning Engineer 5 - Decisioning & Optimization
New York City or Seattle or Los Angeles or Los Gatos
$466k-$750k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
7+ YOE7+ years software engineering experience with 3+ years on ML infrastructure or model serving; proficiency in Java, Python, or Scala; experience building high‑QPS, low‑latency model serving, feature serving, and model monitoring.
Java, Python, Scala, Chronon, Signal Service, JVM
1mo
Save
Mark Applied
Hide
Machine Learning Engineer 5 - Decisioning & Optimization
New York or Los Angeles or Los Gatos or Seattle
$466k-$750k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global video streaming and media production service.
7+ YOE7+ years software engineering; 3+ years ML infrastructure, model serving, or ML platform experience in ads/real-time decisioning; real-time model serving with sub-20ms latency; proficiency in Java, Python, or Scala; experience with ML serving frameworks and real-time feature pipelines; strong model monitoring and production readiness.
Java, Python, Scala, ML serving frameworks, feature stores, model registries
1mo
Save
Mark Applied
Hide
Technical Lead, Front-End Power Optimization
San Jose, California, United States
$184k-$264k/yr OnsiteFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
8+ YOEBachelor’s or Master’s in Electrical/Computer Engineering with 6+ years (Master) or 8+ years (Bachelor); PhD with 3+ years; RTL design, power modeling; PrimePower RTL; scripting (Tcl/Python).
PrimePower RTL, RTL design tools, Tcl, Python
2w
Save
Mark Applied
Hide
Senior Principal Engineer (PE) / Subject Matter Expert (SME) – Quartus Timing Analysis & Optimization
San Jose, California, United States
$266k-$392k/yr OnsiteFull Time
Altera
Altera: Manufacturer of field-programmable gate arrays and programmable logic devices.
15+ YOEMS or PhD in CS/CE/EE,15+ years EDA/timing analysis experience,expertise in STA,timing closure/modeling,physical design optimization,and large-scale C++ development.
Quartus, Timing Analyzer, Fitter, Routing, C++
2w
Save
Mark Applied
Hide
Senior Staff / Principal Machine Learning Scientist, AI Inference & Optimization
Santa Clara, California, United States
$183k-$261k/yr OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
10+ YOE10+ years industry experience with 4+ years hands-on ML; experience in model fine-tuning, quantization, inference runtimes, transformer internals; MS required, PhD preferred.
LoRA, QLoRA, GGUF, AWQ, GPTQ, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp, MLX, CoreML, Python, C++
1mo
Save
Mark Applied
Hide
ML Engineer - Inference & Model Deployment
Cupertino, California, United States
$250k-$310k/yr OnsiteFull Time
Hiring.Cafe
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
vLLM, TensorRT, SGLang, GPU
6d
Save
Mark Applied
Hide
Software Development Manager, LLM Inference Model Enablement, Neuron SDK
Cupertino, California, United States
$213k-$288k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
7+ YOE3+ MgmtManage engineering team to onboard and optimize LLMs for inference on Trainium; strong background in LLM architectures, model performance optimization, and inference techniques; experience with PyTorch and Neuron stack.
PyTorch, AWS Neuron, Neuron compiler, Neuron runtime
1mo
Save
Mark Applied
Hide
Model Distillation Engineer
San Jose, California, United States
$120k-$300k/yr OnsiteFull Time
Hark
Hark: A building multimodal AI models and next-generation hardware to create natural human-machine interfaces.
3+ YOE3+ years in model compression/distillation/quantization, strong fluency in PyTorch or TensorFlow, experience with PTQ/QAT and int8 conversion, hardware-aware optimization for constrained devices, and familiarity with audio/sequence model architectures.
PyTorch, TensorFlow, TFLite, ONNX Runtime, AIMET, Hexagon DSP, NPUs, Ambiq MCUs
2mo
Save
Mark Applied
Hide
AI Intern – VLA Deployment
Santa Clara or Mountain View
OnsiteInternship
XPeng
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
0+ YOEIntern to support optimization and deployment of multimodal models onto vehicle-grade compute; strong CS/EE/ robotics background; hands-on C++/Python; DL frameworks; model optimization and edge deployment.
C++, Python, PyTorch, ONNX, TensorRT, CUDA
3mo
Save
Mark Applied
Hide
3D Game Artist
Hong Kong or San Jose
HybridFull Time
Nex
Nex: Gaming system that turns body movement into interactive play.
3D modeling/rigging experience; Unity proficiency; optimize assets for performance; strong communication.
Unity, Shader Graph, 3D Modeling, Rigging, Particle Systems, In-Engine Implementation
1mo
Save
Mark Applied
Hide
Power Systems Research Scientist
Cupertino, California, United States
$175k-$235k/yr OnsiteFull Time
Gridmatic
Gridmatic: AI-powered platform for optimizing energy trading and battery storage.
Advanced degree in EE/power systems, strong power systems modeling and optimization background, experience with power flow, transmission/congestion analysis, large-scale optimization, and Python programming.
Python, PSS/E, PowerWorld, PSLF, PyTorch, JAX
3mo
Save
Mark Applied
Hide
Sr. Applied Scientist
San Jose or Seattle
$164k-$313k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
Master’s or Ph.D. in CS/ML; track record in mid-training of multimodal models; diffusion architectures; Vision-Language Models; scalable data pipelines; optimize inference.
Python, PyTorch, TensorFlow, Distributed Training, Vision-Language Models