84 inference engineer jobs at 31 companies in Seattle, WA
1mo
Save
Mark Applied
Hide
1mo
Staff Engineer, Inference Optimizations
Seattle, Washington, United States
$191k-$239k/yrHybridFull Time
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
8+ YOE8+ years building/operating multi-tenant distributed systems, 2+ years Go/Golang, 2+ years Kubernetes, strong SRE/observability and production debugging skills, experience optimizing inference and GPU utilization.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
AI Inference Infrastructure Software Engineer (Kubernetes / Cloud)
Seattle, Washington, United States
HybridFull Time
ElastixAI: A software startup focused on AI inference infrastructure and high-performance cloud platforms.
3+ YOE3–5 years in Kubernetes and cloud deployments; BS in CS/Software Eng; strong Python/Bash/Go; experience with Terraform/Pulumi, Ansible/Puppet, GitOps; production ML/inference on Kubernetes.
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yrOnsiteFull Time
Nuance Labs: A building photorealistic, real-time AI avatars and full-duplex audiovisual systems.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
ByteDance: Developing AI-driven content platforms and mobile applications.
PhD in CS or related discipline, strong knowledge of large-model inference, distributed systems, container orchestration, and proficiency in Go/Rust/Python/C++ with cloud/ML infrastructure experience.
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
8+ YOE8–12+ years in distributed systems; leadership of cross-team initiatives; proficient in Go, Python or C++; production Kubernetes expertise; low-latency, high-throughput system design and optimization.
Constellation Space: AI-powered operating system for satellite network management.
BS/MS in CS or Engineering (or equivalent), proven experience deploying ML models to production, strong software engineering in Python and C++, experience with Docker, cloud platforms and MLOps, and optimizing low-latency inference.
8+ YOEBachelor's or equivalent; 8+ years software development with C++, Java, Python, Kotlin, or Go; 4+ years technical leadership; 5+ years testing/launch experience; EMR not applicable; strong communication and critical thinking.
San Francisco or New York City or Los Angeles or Seattle
$245k-$345k/yrHybridFull Time
Whatnot: Social marketplace for buying and selling via live streams
4+ YOE4+ years building ML systems, 3+ years engineering production systems, 1+ year Python, experience with distributed training/inference, databases, monitoring, and cloud services.
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Harell Data: A managed platform that enables secure sharing of scientific datasets and provides compute and tooling for model training and deployment.
5+ YOE5+ years in customer-facing technical roles; familiar with ML training/inference, cloud infrastructure (AWS/GCP), strong written/verbal communication, ability to create docs and runbooks.
Senior Software Engineer, Machine Learning Infrastructure - Generative AI
San Francisco or Sunnyvale or Seattle
$137k-$202k/yrOnsiteFull Time
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
6+ YOE6+ years software engineering experience; BS/MS/PhD in CS or equivalent; deep backend fundamentals in Python and distributed systems; experience with LLM inference/fine-tuning, production reliability, observability, and technical leadership.
Chicago or New York City or San Francisco or Seattle or Sunnyvale
$182k-$202k/yrOnsiteFull Time
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
4+ YOE4+ years building ML models; BS in CS/CE or related; experience with PyTorch, causal inference or constrained optimization preferred; product and marketplace experience a plus.
Member of Technical Staff - Machine Learning Infrastructure Engineer
San Francisco or Toronto or Seattle
$180k-$300k/yrOnsiteFull Time
Preference Model: Building reinforcement learning environments to train frontier AI models.
Experienced software engineer with production ML/data infrastructure skills, proficiency with PyTorch or JAX, distributed systems, AWS/GCP, Kubernetes, data pipelines, and familiarity with transformers and inference libraries like vLLM.
Wherobots: Cloud-native platform for geospatial data analytics and AI.
5+ YOE5+ years building distributed data or ML systems; experience with Ray/Spark/Dask, GPU inference, Python and scientific stack; geospatial and object-storage expertise.
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
San Francisco or Seattle or New York City or United States
$250k-$485k/yrOnsiteFull Time
Perplexity: AI-powered search engine providing conversational answers with citations.
Deep Kubernetes and GPU cluster experience, multi-cloud orchestration, strong distributed systems fundamentals, systems-level coding in Go/Rust/C++, and experience with training and inference workloads.