84 inference engineer jobs at 32 companies in Lake Stevens, WA
1mo
Save
Mark Applied
Hide
1mo
Staff Engineer, Inference Optimizations
Seattle, Washington, United States
$191k-$239k/yrHybridFull Time
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
8+ YOE8+ years building/operating multi-tenant distributed systems, 2+ years Go/Golang, 2+ years Kubernetes, strong SRE/observability and production debugging skills, experience optimizing inference and GPU utilization.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
AI Inference Infrastructure Software Engineer (Kubernetes / Cloud)
Seattle, Washington, United States
HybridFull Time
ElastixAI: A software startup focused on AI inference infrastructure and high-performance cloud platforms.
3+ YOE3–5 years in Kubernetes and cloud deployments; BS in CS/Software Eng; strong Python/Bash/Go; experience with Terraform/Pulumi, Ansible/Puppet, GitOps; production ML/inference on Kubernetes.
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yrOnsiteFull Time
Nuance Labs: A building photorealistic, real-time AI avatars and full-duplex audiovisual systems.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
ByteDance: Developing AI-driven content platforms and mobile applications.
PhD in CS or related discipline, strong knowledge of large-model inference, distributed systems, container orchestration, and proficiency in Go/Rust/Python/C++ with cloud/ML infrastructure experience.
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
8+ YOE8–12+ years in distributed systems; leadership of cross-team initiatives; proficient in Go, Python or C++; production Kubernetes expertise; low-latency, high-throughput system design and optimization.
Constellation Space: AI-powered operating system for satellite network management.
BS/MS in CS or Engineering (or equivalent), proven experience deploying ML models to production, strong software engineering in Python and C++, experience with Docker, cloud platforms and MLOps, and optimizing low-latency inference.
8+ YOEBachelor's or equivalent; 8+ years software development with C++, Java, Python, Kotlin, or Go; 4+ years technical leadership; 5+ years testing/launch experience; EMR not applicable; strong communication and critical thinking.
San Francisco or New York City or Los Angeles or Seattle
$245k-$345k/yrHybridFull Time
Whatnot: Social marketplace for buying and selling via live streams
4+ YOE4+ years building ML systems, 3+ years engineering production systems, 1+ year Python, experience with distributed training/inference, databases, monitoring, and cloud services.
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Harell Data: A managed platform that enables secure sharing of scientific datasets and provides compute and tooling for model training and deployment.
5+ YOE5+ years in customer-facing technical roles; familiar with ML training/inference, cloud infrastructure (AWS/GCP), strong written/verbal communication, ability to create docs and runbooks.
Senior Software Engineer, Machine Learning Infrastructure - Generative AI
San Francisco or Sunnyvale or Seattle
$137k-$202k/yrOnsiteFull Time
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
6+ YOE6+ years software engineering experience; BS/MS/PhD in CS or equivalent; deep backend fundamentals in Python and distributed systems; experience with LLM inference/fine-tuning, production reliability, observability, and technical leadership.
Chicago or New York City or San Francisco or Seattle or Sunnyvale
$182k-$202k/yrOnsiteFull Time
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
4+ YOE4+ years building ML models; BS in CS/CE or related; experience with PyTorch, causal inference or constrained optimization preferred; product and marketplace experience a plus.
Member of Technical Staff - Machine Learning Infrastructure Engineer
San Francisco or Toronto or Seattle
$180k-$300k/yrOnsiteFull Time
Preference Model: Building reinforcement learning environments to train frontier AI models.
Experienced software engineer with production ML/data infrastructure skills, proficiency with PyTorch or JAX, distributed systems, AWS/GCP, Kubernetes, data pipelines, and familiarity with transformers and inference libraries like vLLM.
Staff Machine Learning Engineer, Generative AI Modeling and Inference
Los Angeles or Seattle or Palo Alto or New York City or Bellevue
$195k-$343k/yrOnsiteFull Time
SnapNYSE: SNAP: Develops social media applications and augmented reality technology.
8+ YOEBachelor's degree or equivalent experience and 8+ years of post-bachelor's ML experience, or advanced degree with equivalent experience. Requires computer vision or generative modeling and ML framework experience.