84 inference engineer jobs at 31 companies in Everett, WA

1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Seattle, Washington, United States
$191k-$239k/yr HybridFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
8+ YOE8+ years building/operating multi-tenant distributed systems, 2+ years Go/Golang, 2+ years Kubernetes, strong SRE/observability and production debugging skills, experience optimizing inference and GPU utilization.
Go, Golang, Kubernetes, vLLM, Triton, TensorRT-LLM
1w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
1w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
3w
Save
Mark Applied
Hide
AI Inference Engineer - Speech
Seattle or San Jose
$152k-$332k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Python, shell, C/C++, PyTorch, TensorFlow, CUDA, TensorRT, CUDA Graphs, NVIDIA GPUs, TPU, BrightHire
3mo
Save
Mark Applied
Hide
AI Inference Infrastructure Software Engineer (Kubernetes / Cloud)
Seattle, Washington, United States
HybridFull Time
ElastixAI
ElastixAI: A software startup focused on AI inference infrastructure and high-performance cloud platforms.
3+ YOE3–5 years in Kubernetes and cloud deployments; BS in CS/Software Eng; strong Python/Bash/Go; experience with Terraform/Pulumi, Ansible/Puppet, GitOps; production ML/inference on Kubernetes.
Kubernetes, AWS, GCP, Terraform, Pulumi, Ansible, GitOps, Argo CD, Flux, Python, Bash, Go, Linux
2mo
Save
Mark Applied
Hide
Performance Engineer, Inference Systems
San Francisco or New York City or Seattle
$350k-$850k/yr OnsiteFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Hands-on performance engineering with Python, data analysis, and cross-layer investigations; strong communication of quantitative results.
Python, SQL, Pandas
2mo
Save
Mark Applied
Hide
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yr OnsiteFull Time
Nuance Labs
Nuance Labs: A building photorealistic, real-time AI avatars and full-duplex audiovisual systems.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, CUDA, Triton
3w
Save
Mark Applied
Hide
Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PhD)
Seattle, Washington, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
PhD in CS or related discipline, strong knowledge of large-model inference, distributed systems, container orchestration, and proficiency in Go/Rust/Python/C++ with cloud/ML infrastructure experience.
AIBrix, Kubernetes, vLLM, SGLang, TensorRT-LLM, Docker, Go, Rust, Python, C++, Ray, CUDA, AWS, Azure, GCP, SageMaker, Azure ML, Vertex AI, DeepSpeed, PyTorch
3mo
Save
Mark Applied
Hide
Staff Software Engineer, Inference
Sunnyvale or Bellevue
$188k-$275k/yr HybridFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
8+ YOE8–12+ years in distributed systems; leadership of cross-team initiatives; proficient in Go, Python or C++; production Kubernetes expertise; low-latency, high-throughput system design and optimization.
Go, Python, C++, Kubernetes, CUDA, NCCL, RDMA, NUMA, TensorRT-LLM, Ray Serve, TorchServe
1mo
Save
Mark Applied
Hide
Machine Learning Engineer
Seattle, Washington, United States
$120k-$180k/yr OnsiteFull Time
Constellation Space
Constellation Space: AI-powered operating system for satellite network management.
BS/MS in CS or Engineering (or equivalent), proven experience deploying ML models to production, strong software engineering in Python and C++, experience with Docker, cloud platforms and MLOps, and optimizing low-latency inference.
Python, C++, Docker
3w
Save
Mark Applied
Hide
Senior Staff Engineer, GDC AI Inference Platform
Sunnyvale or Kirkland
$262k-$365k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's or equivalent; 8+ years software development with C++, Java, Python, Kotlin, or Go; 4+ years technical leadership; 5+ years testing/launch experience; EMR not applicable; strong communication and critical thinking.
C++, Java, Python, Kotlin, Go, Kubernetes, Docker
4w
Save
Mark Applied
Hide
Machine Learning Platform Engineer
San Francisco or New York City or Los Angeles or Seattle
$245k-$345k/yr HybridFull Time
Whatnot
Whatnot: Social marketplace for buying and selling via live streams
4+ YOE4+ years building ML systems, 3+ years engineering production systems, 1+ year Python, experience with distributed training/inference, databases, monitoring, and cloud services.
Python, PostgreSQL, DynamoDB, Elasticsearch, Redis, DataDog, Grafana, AWS Sagemaker, Lambda, Kinesis, S3, EC2, EKS, ECS, Apache Kafka, Flink
5d
Save
Mark Applied
Hide
Staff Software Engineer - Data Cloud Applied ML
San Francisco or Seattle or New York City
$189k-$315k/yr OnsiteFull Time
Rippling
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Python, Go, Java, Kubernetes
2mo
Save
Mark Applied
Hide
Senior Software Engineer - AI/ML, AWS Neuron Inference
Seattle, Washington, United States
$168k-$227k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOE5+ years in full software development; BS in CS; 5+ years in Java/C++/C# OO; ML basics and optimization.
Java, C++, C#, PyTorch, JAX, AWS Neuron, ML optimization
3mo
Save
Mark Applied
Hide
Senior Solutions Engineer
Bellevue or Palo Alto
OnsiteFull Time
Harell Data
Harell Data: A managed platform that enables secure sharing of scientific datasets and provides compute and tooling for model training and deployment.
5+ YOE5+ years in customer-facing technical roles; familiar with ML training/inference, cloud infrastructure (AWS/GCP), strong written/verbal communication, ability to create docs and runbooks.
PyTorch, Hugging Face, AWS, GCP
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Machine Learning Infrastructure - Generative AI
San Francisco or Sunnyvale or Seattle
$137k-$202k/yr OnsiteFull Time
DoorDash
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
6+ YOE6+ years software engineering experience; BS/MS/PhD in CS or equivalent; deep backend fundamentals in Python and distributed systems; experience with LLM inference/fine-tuning, production reliability, observability, and technical leadership.
Python, Claude Code, Codex, Cursor, vLLM, SGLang, TensorRT-LLM, Kubernetes, AWS, GCP, Modal
6d
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Chicago or New York City or San Francisco or Seattle or Sunnyvale
$182k-$202k/yr OnsiteFull Time
Uber
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
4+ YOE4+ years building ML models; BS in CS/CE or related; experience with PyTorch, causal inference or constrained optimization preferred; product and marketplace experience a plus.
PyTorch
2w
Save
Mark Applied
Hide
Member of Technical Staff - Machine Learning Infrastructure Engineer
San Francisco or Toronto or Seattle
$180k-$300k/yr OnsiteFull Time
Preference Model
Preference Model: Building reinforcement learning environments to train frontier AI models.
Experienced software engineer with production ML/data infrastructure skills, proficiency with PyTorch or JAX, distributed systems, AWS/GCP, Kubernetes, data pipelines, and familiarity with transformers and inference libraries like vLLM.
PyTorch, JAX, AWS, GCP, Kubernetes, transformers, vLLM, SGLang
2mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer – GeoAI Platform
San Francisco or Bellevue
$185k-$275k/yr HybridFull Time
Wherobots
Wherobots: Cloud-native platform for geospatial data analytics and AI.
5+ YOE5+ years building distributed data or ML systems; experience with Ray/Spark/Dask, GPU inference, Python and scientific stack; geospatial and object-storage expertise.
Ray, Spark, Dask, Python, PyTorch, PyArrow, NumPy, Xarray, Zarr, Cloud-Optimized GeoTIFF (COG), GeoParquet, Parquet, S3, CUDA
1d
Save
Mark Applied
Hide
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
San Francisco or Seattle or New York City or United States
$250k-$485k/yr OnsiteFull Time
Perplexity
Perplexity: AI-powered search engine providing conversational answers with citations.
Deep Kubernetes and GPU cluster experience, multi-cloud orchestration, strong distributed systems fundamentals, systems-level coding in Go/Rust/C++, and experience with training and inference workloads.
Kubernetes, kubectl, NVIDIA, CUDA, InfiniBand, RoCE, CoreWeave, AWS, GCP, Go, Rust, C++, vLLM, SGLang, TensorRT-LLM, Slurm, Triton, RDMA, Prometheus, Grafana, Weights & Biases