Applied Intuition
Posted 2d ago

Machine Learning Performance Engineer - Offboard Training & Inference

Applied Intuition
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$215k-$285k/yrOnsiteFull Time
Responsibilities
  • optimizing training
  • improving scaling
  • building tooling
Requirements
  • ML performance engineering experience with distributed training
  • Batch inference
  • GPU or accelerator optimization
  • Python, and C++ or another systems language
  • Strong debugging and analytical skills required
Technical tools mentioned
FSDPDeepSpeedMegatronNCCLNVIDIA Triton Inference ServerTensorRTONNX RuntimeRayPythonC++CUDATritonCUTLASSNsight SystemsNsight ComputePyTorch ProfilerperfKubernetesSlurmROSOpenCV

Job description

About Applied Intuition

Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon Valley company is creating the digital infrastructure needed to bring intelligence to every moving machine on the planet. Applied Intuition services the automotive, defense, trucking, construction, mining and agriculture industries in three core areas: tools and infrastructure, operating systems, and autonomy. Eighteen of the top 20 global automakers, as well as the United States military and its allies, trust the company’s solutions to deliver physical intelligence. Applied Intuition is headquartered in Sunnyvale, California, with offices in Washington, D.C.; San Diego; Ft. Walton Beach, Florida; Ann Arbor, Michigan; London; Stuttgart; Munich; Stockholm; Bangalore; Seoul; and Tokyo. Learn more at applied.co.

We are an in-office company, and our expectation is that employees primarily work from their Applied Intuition office 5 days a week. However, we also recognize the importance of flexibility and trust our employees to manage their schedules responsibly. This may include occasional remote work, starting the day with morning meetings from home before heading to the office, or leaving earlier when needed to accommodate family commitments.

About the Role

We are looking for a performance engineer who specializes in making large-scale machine learning workloads fast and cost-efficient in the datacenter. This role is focused on distributed training runs spanning many nodes, and high-throughput batch inference sweeping petabytes of real-world autonomy logs for auto-labeling, data mining, ground-truth generation, and evaluation.

The optimization target here is not tail latency on a vehicle - it is throughput, cluster goodput, and cost per unit of data processed. A training run that wastes 30% of its GPU-hours on stalled data loaders, or an offline inference sweep that takes a week instead of a day, directly slows down how fast the whole company can iterate. You will own the gap between what our fleet of accelerators is theoretically capable of and what our workloads actually achieve: profiling across the stack, finding where the compute and the wall-clock time actually go, and closing the difference.

You will work at the intersection of accelerators, ML frameworks, and large-scale data infrastructure, partnering with the teams who own each layer to land wins that show up in training time-to-result and offline processing cost. At Applied, we encourage all engineers to take ownership over technical and product decisions, closely interact with users to collect feedback, and contribute to a thoughtful, dynamic team culture.

At Applied, you will:

  • Profile and optimize distributed training end to end - data loading and preprocessing, augmentation, kernel execution, gradient communication, and checkpointing

  • Optimize large-scale offline and batch inference over petabyte-scale sensor logs: batching and scheduling strategies, quantization and low-precision execution, graph optimization, and accelerator saturation across long-running sweeps

  • Establish roofline and performance models for our workloads, quantify the gap between achieved and theoretical performance, and stack-rank optimization opportunities by impact and effort

  • Improve multi-node scaling efficiency: sharding and parallelism strategies, collective communication, interconnect utilization, and memory-bandwidth and kernel-fusion bottlenecks

  • Drive cluster goodput - reduce GPU idle time from input pipeline stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery on long-running jobs

  • Build the benchmarking, observability, and regression-detection tooling that keeps performance from silently degrading as models and code evolve

  • Collaborate with engineers across functions to solve complex data and compute problems at scale

  • Contribute to a team culture that values effective collaboration, technical excellence, and innovation

We're looking for someone who has:

  • Hands-on ML performance engineering experience: profiling, roofline analysis, throughput optimization, and root-cause investigation in production systems

  • Experience with distributed multi-node training at scale (FSDP, DeepSpeed, Megatron, NCCL, or equivalent), including diagnosing scaling inefficiency as node count grows

  • Deep familiarity with GPU or accelerator performance concepts - memory bandwidth, kernel launch overhead, occupancy, quantization, collective communication

  • Experience with high-throughput or batch inference systems (NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar)

  • Fluency in Python and proficiency in C++ or another systems language

  • Excellent debugging, analytical, and problem-solving skills

  • A deep understanding of machine learning foundations, and the ability to develop technical solutions for problems with no established playbook

Nice to have:

  • GPU kernel development experience: CUDA, Triton, CUTLASS, or hand-tuned attention implementations

  • Experience with profiling toolchains such as Nsight Systems/Compute, PyTorch Profiler, or perf

  • Experience with GPU scheduling and orchestration on Kubernetes, Slurm, or Ray, including multi-tenant cluster utilization

  • Experience with fault tolerance and elastic training for long-running jobs - checkpointing strategy, straggler mitigation, preemption recovery

  • Familiarity with autonomy or robotics data (ROS, OpenCV, multi-sensor log formats)

Don’t meet every single requirement? If you’re excited about this role but your past experience doesn’t align perfectly with every qualification in the job description, we encourage you to apply anyway. You may be just the right candidate for this or other roles.

Applied Intuition is an equal opportunity employer and federal contractor or subcontractor. Consequently, the parties agree that, as applicable, they will abide by the requirements of 41 CFR 60-1.4(a), 41 CFR 60-300.5(a) and 41 CFR 60-741.5(a) and that these laws are incorporated herein by reference. These regulations prohibit discrimination against qualified individuals based on their status as protected veterans or individuals with disabilities, and prohibit discrimination against all individuals based on their race, color, religion, sex, sexual orientation, gender identity or national origin. These regulations require that covered prime contractors and subcontractors take affirmative action to employ and advance in employment individuals without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status or disability. The parties also agree that, as applicable, they will abide by the requirements of Executive Order 13496 (29 CFR Part 471, Appendix A to Subpart A), relating to the notice of employee rights under federal labor laws.

About Applied Intuition

Developing software and simulation infrastructure for autonomous vehicles.

Year founded
2017
Employees
1300
Organization type
Private
Latest investment
Raised $600.00M Series F (2025) — led by BlackRock, Kleiner Perkins
Headquarters
US

Similar jobs

Performance Engineer roles near Sunnyvale, California
2w
Save
Mark Applied
Hide
Senior Performance Engineer - DGX Cloud
Santa Clara or Austin or Redmond or Oregon or Washington
$224k-$431k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE12+ years experience; BS or higher in CS/CE or equivalent; strong C++ and Python skills; foundation in OS, computer architecture, distributed systems; performance engineering and profiling experience.
C++, Python, CUDA, PyTorch, JAX, XLA
2w
Save
Mark Applied
Hide
Senior Performance Engineer - DGX Cloud
Santa Clara or Austin or Redmond or Washington or Oregon
$224k-$431k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
12+ YOE12+ years experience, strong C++ and Python programming, foundations in OS/architecture/distributed systems, performance engineering and profiling experience, BS in CS/CE or equivalent.
C++, Python, CUDA, PyTorch, JAX, XLA
3w
Save
Mark Applied
Hide
Performance Engineer
San Jose, California, United States
$135k-$170k/yr OnsiteFull Time
Astera Labs
Astera LabsNASDAQ: ALAB: Designs connectivity solutions for cloud and AI infrastructure.
2+ YOEBachelor's in engineering/computer science, experience benchmarking and characterizing GPU cluster performance, proficiency with performance tooling and scripting (Python); strong systems and networking knowledge.
NVBandwidth, NCCL, Confluence, CUDA, MPI, Python, COSMOS, NVLink, PCIe, Ethernet, UALink, UEC
3w
Save
Mark Applied
Hide
Senior Performance Engineer
San Francisco, California, United States
$170k-$205k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
Expertise in Linux kernel systems, low-level optimization, performance benchmarking, and high-performance coding in Go/C/C++/Rust. Bachelor's degree or equivalent experience and cross-functional collaboration skills.
Go, C, C++, Rust, Linux
1mo
Save
Mark Applied
Hide
CoDesign & NextGen Performance Engineer
Sunnyvale or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
3+ YOE3+ years experience in computer architecture/performance, strong analytical skills, low-level deep learning math exposure, experience with simulators, kernel optimization, and performance profiling; comfortable with C++ and Python.
C++, Python
1mo
Save
Mark Applied
Hide
Senior Performance Engineer - Apple Frameworks
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Analyze system and UI architecture, establish performance baselines, diagnose and resolve complex performance issues, and build regression-prevention mechanisms.
1mo
Save
Mark Applied
Hide
Senior Performance Engineer
San Jose, California, United States
$138k-$206k/yr OnsiteFull Time
Samsung Semiconductor
Samsung SemiconductorKorea Exchange: 005930: Designs and manufactures memory chips, processors, and sensors.
0+ YOEAdvanced degree in CS/CE/EE or equivalent experience; strong LLM systems and NVIDIA GPU performance knowledge; experience profiling AI workloads with Nsight tools; proficiency in Python and C++; experience with PyTorch, DeepSpeed, Ray, or similar.
Nsight Systems, Nsight Compute, Python, C++, PyTorch, vLLM, SGLang, TensorRT-LLM, DeepSpeed, Ray, Megatron-LM
1mo
Save
Mark Applied
Hide
Staff Performance Engineer, Data Protection & Security Platform Engineering
Santa Clara, California, United States
$203k-$226k/yr HybridFull Time
Cohesity
Cohesity: AI-powered data security and management software provider.
10+ YOE10+ years experience, BS/MS/PhD in Electrical Engineering or Computer Science (or equivalent), expertise in performance analysis and modeling, strong OS/hardware knowledge, and scripting skills in GoLang and Python.
GoLang, Python