Institute of Foundation Models: Develops open-source frontier-class AI foundation models and research.
Currently pursuing a quantitative degree; experience or coursework in CUDA, GPU kernel development, performance modeling, Nsight profiling, C++, Python, and deep learning frameworks preferred.
HitachiTokyo Stock Exchange: 6501: Global provider of digital systems and social infrastructure solutions.
5+ YOEMinimum 5 years in optimization modeling tools; 3+ years programming; 3+ years Linux experience; 2+ years developing power system software; English proficiency; US work authorization.
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Extensive, senior experience optimizing large-scale ML systems and GPU architectures; CUDA programming; memory hierarchies; distributed training; transformer models.
Multimodal Model Training and Inference Optimization Engineer
San Jose, California, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
M.S. or PhD in CS/EE/AI, experience optimizing model training and inference, proficiency in Python, C++, CUDA, PyTorch, Megatron, Deepspeed, distributed training, and knowledge of transformers/diffusion models.
Multimodal Model Training and Inference Optimization Engineer
San Jose, California, United States
$156k-$388k/yrOnsiteFull Time
TikTok: Global short-form video hosting and social media platform.
MS/PhD in CS/EE/AI or related, experience optimizing model training and inference, distributed training, Python/C++/CUDA, PyTorch/Megatron/Deepspeed, knowledge of transformers and diffusion models.
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years AI/software engineering experience, strong Python/C++ coding, GPU model training/optimization, communication and customer-facing skills, BS/MS/PhD or equivalent.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOEDegree in CS/EE/Physics/Math or equivalent experience,5+ years AI/software engineering,proficient in Python/C++,experience profiling and optimizing models on GPUs and developing GPU kernels,strong communication skills.
New York City or Seattle or Los Angeles or Los Gatos
$466k-$750k/yrOnsiteFull Time
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
7+ YOE7+ years software engineering experience with 3+ years on ML infrastructure or model serving; proficiency in Java, Python, or Scala; experience building high‑QPS, low‑latency model serving, feature serving, and model monitoring.
NetflixNASDAQ: NFLX: Global video streaming and media production service.
7+ YOE7+ years software engineering; 3+ years ML infrastructure, model serving, or ML platform experience in ads/real-time decisioning; real-time model serving with sub-20ms latency; proficiency in Java, Python, or Scala; experience with ML serving frameworks and real-time feature pipelines; strong model monitoring and production readiness.
Java, Python, Scala, ML serving frameworks, feature stores, model registries
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
8+ YOEBachelor’s or Master’s in Electrical/Computer Engineering with 6+ years (Master) or 8+ years (Bachelor); PhD with 3+ years; RTL design, power modeling; PrimePower RTL; scripting (Tcl/Python).
Altera: Manufacturer of field-programmable gate arrays and programmable logic devices.
15+ YOEMS or PhD in CS/CE/EE,15+ years EDA/timing analysis experience,expertise in STA,timing closure/modeling,physical design optimization,and large-scale C++ development.
Senior Staff / Principal Machine Learning Scientist, AI Inference & Optimization
Santa Clara, California, United States
$183k-$261k/yrOnsiteFull Time
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
10+ YOE10+ years industry experience with 4+ years hands-on ML; experience in model fine-tuning, quantization, inference runtimes, transformer internals; MS required, PhD preferred.
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
Software Development Manager, LLM Inference Model Enablement, Neuron SDK
Cupertino, California, United States
$213k-$288k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
7+ YOE3+ MgmtManage engineering team to onboard and optimize LLMs for inference on Trainium; strong background in LLM architectures, model performance optimization, and inference techniques; experience with PyTorch and Neuron stack.
Hark: A building multimodal AI models and next-generation hardware to create natural human-machine interfaces.
3+ YOE3+ years in model compression/distillation/quantization, strong fluency in PyTorch or TensorFlow, experience with PTQ/QAT and int8 conversion, hardware-aware optimization for constrained devices, and familiarity with audio/sequence model architectures.
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
0+ YOEIntern to support optimization and deployment of multimodal models onto vehicle-grade compute; strong CS/EE/ robotics background; hands-on C++/Python; DL frameworks; model optimization and edge deployment.
Gridmatic: AI-powered platform for optimizing energy trading and battery storage.
Advanced degree in EE/power systems, strong power systems modeling and optimization background, experience with power flow, transmission/congestion analysis, large-scale optimization, and Python programming.
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
Master’s or Ph.D. in CS/ML; track record in mid-training of multimodal models; diffusion architectures; Vision-Language Models; scalable data pipelines; optimize inference.