113 reinforcement learning jobs at 58 companies in San Francisco, CA
6d
Save
Mark Applied
Hide
6d
Senior Reinforcement Learning Engineer
Sunnyvale or Austin
OnsiteFull Time
Apptronik: Designs and manufactures humanoid robots for industrial automation.
5+ YOERequires 5+ years with reinforcement learning frameworks and physics simulators, Python and C++, distributed training, robotics, and policy deployment; PhD or MS preferred and mentoring experience expected.
PyTorch, JAX, MuJoCo, IsaacGym, Python, C++, CoRL, RSS, ICRA
Bugcrowd: Provides a crowdsourced platform for security vulnerability testing.
Experience with reinforcement learning workflows, Linux ML environments, vulnerability research/binary exploitation, proficiency in Python and C, DevOps pipelines and reproducible builds, and low-level debugging.
Mayhem, GitHub Actions, docker, buildkit, nix, Python, C, Rust, Linux
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
1+ YOEAdvanced degree preferred but open to fresh graduates; proficiency in Python; 1+ years experience with deep learning frameworks such as PyTorch; strong RL and imitation learning knowledge; experience with PPO/DQN/SAC; C++ and hardware experience preferred.
Pony.aiNASDAQ: PONY: Develops autonomous driving technology and operates robotaxi services.
3+ YOEM.S./Ph.D. or equivalent experience; 3+ years building production ML with strong RL experience; depth in deep learning and generative models; distributed training and large-scale data processing; strong communication.
DeepRoute.ai: Develops full-stack autonomous driving software and VLA models.
Proficiency in modern RL and RLHF algorithms, experience with reward model training and LLM/VLM fine-tuning, distributed RL training and massively parallel simulation, sim-to-real transfer, Python and C++, PyTorch, and distributed training frameworks.
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years distributed systems experience, strong Python and PyTorch or JAX, multi-node GPU orchestration (Ray, SLURM, Kubernetes), experience with actor-learner architectures and RL training pipelines.
Hunyuan Multimodal Reinforcement Learning Research Intern
Palo Alto, California, United States
$80k-$125k/yrOnsiteFull Time
TencentHong Kong Stock Exchange: 0700: Multinational technology conglomerate providing internet and entertainment services.
PhD student in computer science required; strong research record with top-tier publications; hands-on deep learning implementation, training, CPU/GPU acceleration, and distributed training experience.
DoorDashNYSE: DASH: Local food delivery and on-demand logistics platform.
BS, MS, or PhD in CS, EE, Robotics, or related field; deep RL and deep learning expertise; large-scale RL training; JAX or similar framework; GPU simulation and software development experience.
JAX, Claude Code, Codex, Cursor, NeurIPS, ICML, ICLR, CoRL, RSS, ICRA
Lead AI Infrastructure Engineer, Reinforcement Learning
Santa Clara, California, United States
$179k-$306k/yrHybridFull Time
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Design and operate distributed RL training infrastructure at scale; strong systems experience in ML platforms; proficiency with PyTorch/JAX, NCCL/MPI-style distributed training, C++/Python performance tuning; Bachelor's degree required.
Pony.aiNASDAQ: PONY: Develops autonomous driving technology for passenger and freight transportation.
MS/PhD in CS/ML/AI or related field, or equivalent experience; RL/production ML experience; deep learning, sequence modeling, generative models; publish or ship impactful ML systems; large-scale training and data processing; lead ambiguous work.
Member of Technical Staff – Senior Engineer, Reinforcement Learning – Policy Post-Training
Cambridge or San Francisco
$255k-$340k/yrOnsiteFull Time
Walden Robotics: Builds general-purpose robots and advances robot manipulation through research and development.
Hands-on RL experience for manipulation, sim-to-real transfer, reward and curriculum design, strong software engineering, and ability to run large-scale training and experiments.
Research Engineer – Reinforcement Learning (RL) Systems & Infrastructure (Seed Infra)
San Jose, California, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
Design and build scalable RL systems and infrastructure for large-scale model training; expertise in distributed systems, GPU optimization, Python/C++, and RL workflows.
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Redwood City, California, United States
HybridFull Time
Luma AI: Develops multimodal AI for video generation and creative production.
Experience operating post-trained LLMs with reinforcement learning at scale, distributed PyTorch training, building RL environments and reward infrastructure, and debugging large asynchronous rollout pipelines.
Applied Deep Learning PhD Research Intern, Reinforcement Learning for LLMs - Fall 2026
Santa Clara, California, United States
$30-$94/hrOnsiteFull Time, Internship
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
Pursuing a PhD with strong background in reinforcement learning and NLP, excellent Python and PyTorch skills, experience with large-scale training and experimental research.
1X: Manufacturing safe, general-purpose humanoid robots for home and work.
Expertise in RL (PPO/SAC/TD-MPC), sim-to-real transfer, PyTorch, simulation platforms, Python/C++ experience in large codebases, and deploying RL policies to physical robots.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOEBachelor's in CS or related field; 5+ years exp including 3+ in engineering; strong communication; experience presenting to technical audiences; RL pipelines; PyTorch/JAX/NeMo-RL; able to work on training pipelines and demos.
Hikinex: Provider of outsourced sales and administrative business solutions.
PhD or equivalent; deep experience building robot learning systems for real-world manipulation, expertise in imitation/reinforcement learning, experience with physical robots, and technical leadership experience.
Mariana Minerals: Building software-first infrastructure to produce and refine critical minerals.
0+ YOE0–4 years ML or scientific computing experience; strong ML fundamentals and deep learning exposure; reinforcement learning a plus; proficiency in Python; ability to debug codebases and collaborate with process/chemistry experts.
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOERequires 5+ years building production ML systems, Python, PyTorch, Spark, Airflow, modern ML infrastructure, and expertise in deep learning, reinforcement learning, optimization, LLMs, or VLMs.
PyTorch, Spark, Airflow, Python, Claude Code, Codex, Cursor, large language models (LLMs), vision-language models (VLMs)