34 reinforcement learning engineer jobs at 19 companies in Novato, CA
4w
Save
Mark Applied
Hide
4w
Reinforcement Learning Engineer (Cybersecurity)
United States or San Francisco or New Hampshire
$176k-$243k/yrRemoteFull Time
Bugcrowd: Provides a crowdsourced platform for security vulnerability testing.
Experience with reinforcement learning workflows, Linux ML environments, vulnerability research/binary exploitation, proficiency in Python and C, DevOps pipelines and reproducible builds, and low-level debugging.
Mayhem, GitHub Actions, docker, buildkit, nix, Python, C, Rust, Linux
Pony.aiNASDAQ: PONY: Develops autonomous driving technology and operates robotaxi services.
3+ YOEM.S./Ph.D. or equivalent experience; 3+ years building production ML with strong RL experience; depth in deep learning and generative models; distributed training and large-scale data processing; strong communication.
DoorDashNYSE: DASH: Local food delivery and on-demand logistics platform.
BS, MS, or PhD in CS, EE, Robotics, or related field; deep RL and deep learning expertise; large-scale RL training; JAX or similar framework; GPU simulation and software development experience.
JAX, Claude Code, Codex, Cursor, NeurIPS, ICML, ICLR, CoRL, RSS, ICRA
Member of Technical Staff – Senior Engineer, Reinforcement Learning – Policy Post-Training
Cambridge or San Francisco
$255k-$340k/yrOnsiteFull Time
Walden Robotics: Builds general-purpose robots and advances robot manipulation through research and development.
Hands-on RL experience for manipulation, sim-to-real transfer, reward and curriculum design, strong software engineering, and ability to run large-scale training and experiments.
Pony.aiNASDAQ: PONY: Develops autonomous driving technology for passenger and freight transportation.
MS/PhD in CS/ML/AI or related field, or equivalent experience; RL/production ML experience; deep learning, sequence modeling, generative models; publish or ship impactful ML systems; large-scale training and data processing; lead ambiguous work.
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Redwood City, California, United States
HybridFull Time
Luma AI: Develops multimodal AI for video generation and creative production.
Experience operating post-trained LLMs with reinforcement learning at scale, distributed PyTorch training, building RL environments and reward infrastructure, and debugging large asynchronous rollout pipelines.
Mariana Minerals: Building software-first infrastructure to produce and refine critical minerals.
0+ YOE0–4 years ML or scientific computing experience; strong ML fundamentals and deep learning exposure; reinforcement learning a plus; proficiency in Python; ability to debug codebases and collaborate with process/chemistry experts.
Research Engineer, Chip Design RL (Reinforcement Learning)
San Francisco or New York City
$500k-$850k/yrHybridFull Time
Anthropic: Developing safe and reliable artificial intelligence systems.
Bachelor's or equivalent, expertise in ASIC/FPGA design and EDA tools, experience with RTL, verification (UVM, formal methods), physical design and tapeout experience; RL experience and tooling experience preferred.
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOERequires 5+ years building production ML systems, Python, PyTorch, Spark, Airflow, modern ML infrastructure, and expertise in deep learning, reinforcement learning, optimization, LLMs, or VLMs.
PyTorch, Spark, Airflow, Python, Claude Code, Codex, Cursor, large language models (LLMs), vision-language models (VLMs)
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
Master's or PhD in CS/robotics/ML, deep reinforcement learning experience, strong C++ and Python skills, hands-on with JAX or PyTorch, experience bridging sim-to-real fidelity gaps and working with GPU/CUDA environments.
Macroscope: Provides AI-powered codebase analysis and automated code reviews.
3+ YOE3+ years applied ML experience; experience building, training, fine-tuning, or evaluating ML models; dataset creation and evaluation; experiment design; familiarity with LLMs and reinforcement learning techniques.
Autonomique: Developing hardware-agnostic physical AI software for industrial robotics.
3+ YOEMS or PhD in Robotics/CS, 3+ years shipping code for physical robots, expertise in manipulation planning, imitation or reinforcement learning, production Python and C++ experience, familiarity with sim-to-real.
Research Engineer / Research Scientist - RE / RS - Proactivity
San Francisco, California, United States
$295k-$555k/yrHybridFull Time
OpenAI: Develops artificial intelligence models and generative AI software services.
Strong ML engineering and research experience with LLM post-training, reinforcement learning, dataset creation, evaluations, and ability to work in a large ML codebase.
Mountain View or New York City or Kirkland or San Francisco
$251k-$310k/yrHybridFull Time
Waymo: Autonomous driving technology for ride-hailing and logistics.
4+ YOEMaster's or PhD in a technical field and 4+ years of industry or postdoctoral research experience in reinforcement learning or foundation models. Requires distributed training, JAX, Flax, and Transformer optimization expertise.
Databricks: A unified platform for data analytics and artificial intelligence.
2+ YOEBS/MS/PhD in CS or related; 2+ years applied research with shipped prototypes; experience with LLMs, agents, reinforcement learning, and post-training workflows; strong communication and cross-functional collaboration.
Backbone: AI-powered infrastructure for healthcare payments and authorizations.
0+ YOEExperience with ML research, language models, NLP, evals, reinforcement learning, agentic systems, and model improvement; demonstrated technical depth via papers/projects/internships; new grads through ~5-6 years experience considered.
Abundant: Building simulation infrastructure for AI model training.
Deep technical fluency in evaluations, reinforcement learning, LLMs and agents; research-oriented with ability to read SOTA papers; highly proficient coding agents (e.g., Codex, Claude Code); ownership mindset and experience building scalable systems.
Member of Technical Staff, Post-Training, RL Environments
San Francisco, California, United States
$350k-$500k/yrOnsiteFull Time
Mirendil: Developing frontier artificial intelligence models to accelerate scientific research.
Build and own data systems and execution environments for reinforcement learning; experience with long-horizon RL tasks, scalable sandboxed environments, and preventing reward hacking; strong collaboration and system-design skills.
Moonlake: Generates interactive 3D world simulations using artificial intelligence.
Experience training large language, vision-language, multimodal, or code models; strong reinforcement learning and post-training expertise; distributed training and high-throughput systems experience; Python and PyTorch/JAX proficiency.
Xterra AI: A stealth-stage AI building agents and foundation models to solve complex scientific problems.
Hands-on ML experience training large models, reinforcement learning (RLHF/RLAIF), reward modeling, and ability to work across research and engineering.