NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOEMaster's or PhD in computer science, computer engineering, or related field, 6+ years' experience, strong Python and C++, GPU profiling, LLM inference frameworks, and CUDA kernel optimization expertise.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
DigitalOceanNYSE: DOCN: The AI-Native Cloud purpose-built for inference and agentic workloads.
8+ YOE8+ years building/operating multi-tenant distributed systems, 2+ years Go/Golang, 2+ years Kubernetes, strong SRE/observability and production debugging skills, experience optimizing inference and GPU utilization.
Sunnyvale or Boston or Seattle or Los Angeles County
$167k-$260k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Multinational technology focused on e-commerce and cloud computing.
3+ YOERequires 3+ years building machine learning models, 2+ years optimizing neural-model inference, production real-time inference experience, GPU optimization expertise, and Java, C++, or Python programming.
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yrOnsiteFull Time
Nuance Labs: Private AI research building real-time audiovisual foundation models for face-to-face conversational AI.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
F5NASDAQ: FFIV: Delivering and securing applications across any multi-cloud environment.
Proficiency in Python, C++, Rust, or Golang; experience with inference tools and cloud infrastructure; expertise in GPU and AI hardware optimization; scalable AI serving and monitoring experience.
vLLM, TGI (Text Generation Inference), NVIDIA Triton, NVIDIA GPUs, CUDA, TensorRT, Apple Silicon, CoreML, TPUs, LPUs, Kubernetes, Python, C++, Rust, Golang, Llama.cpp, Ollama, Docker, AWS, GCP, Azure, Speculative Decoding, PagedAttention, Triton kernels, MLOps, SRE, Time to First Token (TTFT), SLAs
ZoomNasdaq Global Select Market: ZM: American publicly traded communications platform serving businesses and individuals with AI-assisted video, voice, chat, and phone services.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
Seattle, Washington, United States
OnsiteFull Time
ByteDance: Global technology specializing in AI-powered content platforms.
Bachelor's in CS/EE/Software, strong C/C++ and Python, experience with PyTorch or TensorFlow, production LLM/VLM inference optimization, GPU familiarity and operator optimization, containerization experience.
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
15+ YOE15+ years software development with 5+ years technical leadership; deep expertise in AI frameworks and ROCm; mastery of performance profiling and distributed training/inference optimization; PhD/Master's or equivalent experience.
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Seattle, Washington, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
Build and optimize inference frameworks, services, and tools for large-scale foundation models, supporting low-latency AI services across Apple products.
Anthropic: AI research developing safe and steerable AI systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
CoreWeaveNasdaq: CRWV: Specialized cloud provider for large-scale AI and machine learning.
8+ YOE8–12+ years in distributed systems; leadership of cross-team initiatives; proficient in Go, Python or C++; production Kubernetes expertise; low-latency, high-throughput system design and optimization.
Constellation Space: Private Seattle software building AI-driven satellite-network operations infrastructure for space companies.
BS/MS in CS or Engineering (or equivalent), proven experience deploying ML models to production, strong software engineering in Python and C++, experience with Docker, cloud platforms and MLOps, and optimizing low-latency inference.
Rippling: Unified workforce management platform for HR, IT, and Finance.
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Chicago or New York City or San Francisco or Seattle or Sunnyvale
$182k-$202k/yrOnsiteFull Time
Uber FreightNYSE: UBER: Global technology platform for ride-hailing, delivery, and freight logistics.
4+ YOE4+ years building ML models; BS in CS/CE or related; experience with PyTorch, causal inference or constrained optimization preferred; product and marketplace experience a plus.
AdobeNASDAQ: ADBE: Empowering everyone to create through innovative digital experiences.
10+ YOE10+ years in data engineering/ML infrastructure, distributed systems expertise, Python and a systems language, experience with Ray or Spark, GPU inference optimization, large-scale databases and data curation for model training.
MicrosoftNASDAQ: MSFT: Multinational technology providing software, cloud, and AI solutions.
6+ YOEBachelor's in CS or related + 6+ years engineering experience (or equivalent); coding experience in C, C++, C#, Java, JavaScript, or Python; experience with inference stacks, performance optimization, open-source code; ability to pass Microsoft security screening.