Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
Seattle, Washington, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's in CS/EE/Software, strong C/C++ and Python, experience with PyTorch or TensorFlow, production LLM/VLM inference optimization, GPU familiarity and operator optimization, containerization experience.
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
4+ YOEPhD in OR/applied math or related field; or Masters with 4+ years of modeling/optimization; strong optimization theory and software experience; proficiency in C/C++, Python/Julia; SQL and large data handling; ability to communicate and mentor.
CPLEX, Gurobi, XPRESS, SQL, MySQL, ETL, C++, Java, Python, Julia
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
15+ YOE15+ years software development with 5+ years technical leadership; deep expertise in AI frameworks and ROCm; mastery of performance profiling and distributed training/inference optimization; PhD/Master's or equivalent experience.
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOE8+ years in compiler optimizations, strong C++ skills, knowledge of processor ISA, experience with MLIR/LLVM/Clang preferred; BS/MS/PhD in CS/CE or equivalent experience.
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure with GPU architecture expertise, experience optimizing attention layers and distributed GPU kernels, and strong low-level systems design and open-source contributions.
Sr Electricity Market Optimization Software Engineer
Bellevue, Washington, United States
$128k-$192k/yrOnsiteFull Time
GE VernovaNYSE: GEV: Designs and services technologies for global power generation and electrification.
2+ YOEMaster's or Ph.D. in electrical engineering with power systems/optimization, 2+ years power systems experience, MIP knowledge, GitHub experience, testing and automation, strong communication and leadership.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
8+ YOE8+ years in compiler optimizations, strong C++ skills, BS/MS/PhD in CS or related, experience with MLIR/LLVM/Clang or CUDA preferred, excellent communication and software engineering skills.
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
5+ YOE5+ years building HPC/GPU software, hands-on CUDA kernel authoring and optimization, C++/Python coding, GPU profiling, and experience delivering performance at scale.
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yrOnsiteFull Time
Nuance Labs: A building photorealistic, real-time AI avatars and full-duplex audiovisual systems.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
SnowflakeNYSE: SNOW: Cloud-based platform for data storage, processing, and analytics.
10+ YOE10+ years building/optimizing large-scale data systems; expertise in query optimization, incremental/stream processing, or materialized view maintenance; strong CS fundamentals; proficient in C++ or Java; ability to lead multi-engineer initiatives.
Sr. Multimodal Model Training and Inference Optimization Engineer
Seattle, Washington, United States
$233k-$428k/yrOnsiteFull Time
TikTok: Global short-form video hosting and social media platform.
3+ YOEMS or PhD in CS/EE/AI, 3+ years optimizing AI model training and inference, strong software engineering, proficiency in Python,C++,CUDA, PyTorch, Megatron, Deepspeed, distributed training, transformers and diffusion models.