Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or United States
$124k-$242k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
3+ YOEBachelor's, master's, or doctoral degree in a related field or equivalent experience; 3+ years' engineering experience; GPU profiling, Python, C++/CUDA, and AI inference optimization expertise required.
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
3+ YOEMaster's degree or higher in a technical field; 3+ years in HPC, AI infrastructure, model optimization, or embedded deployment; C++, Python, CUDA, OpenMP, inference engines, GPU architectures, and system profiling expertise.
Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
San Jose, California, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's in CS/EE/Software or related; strong C/C++ and Python; experience with PyTorch or TensorFlow; production LLM/VLM inference optimization experience; familiarity with GPU architecture and containerized server debugging.
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Santa Clara, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build and optimize inference frameworks, services, and tools for large-scale foundation models across cloud infrastructure, including language, vision, and speech models.
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
6+ YOEBachelor's plus 6 years or master's plus 4 years developing AI/ML technologies, and 6 years programming with Python, Go, Scala, or Java. Cloud AI deployment and team leadership are preferred.
Senior Lead AI Engineer (FM Hosting, LLM Inference)
New York or McLean or Cambridge or San Jose
$251k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
6+ YOEBachelor's in CS/AI/EE/CE or related with 6+ years (or Master's with 4+ years); 6+ years programming with Python, Go, Scala, or Java; experience deploying scalable AI systems, LLM inference, similarity search, and optimization of training/inference.
Sr. Software Development Engineer, Inference Team - AWS Neuron
Seattle or Cupertino or Los Angeles County
$168k-$262k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOERequires 5+ years of software development and programming experience, a computer science degree or equivalent, machine learning optimization experience, and knowledge of inference frameworks and accelerator hardware.
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: A diversified financial services providing banking and credit products.
6+ YOEBachelor's degree plus 6 years or master's degree plus 4 years developing AI/ML technologies; 6 years programming with Python, Go, Scala, or Java; cloud AI deployment experience preferred.
Senior AI Infra Engineer - Large Model Inference Systems (Multimodal/LLM/VLM)
San Jose, California, United States
$213k-$450k/yrOnsiteFull Time
TikTok: Global short-form video hosting and social media platform.
4+ YOEBachelor's degree,4+ years in high-performance computing or distributed scheduling, familiarity with large-model architectures, strong system design and performance-optimization skills, experience with CUDA/Triton/Cutlass and inference frameworks.
Distinguished Engineer - AI (San Jose, CA, US, 95128)
San Jose, California, United States
$266k-$396k/yrOnsiteFull Time
NetAppNASDAQ: NTAP: Sells enterprise data storage and cloud management software.
15+ YOE15+ years building low-latency, fault-tolerant distributed systems and AI/ML inference platforms; expertise with inference engines, model optimization, storage for AI, RDMA/DPDK, and Kubernetes-based orchestration.
NetAppNasdaq: NTAP: Provides intelligent data infrastructure for hybrid cloud environments.
15+ YOEExpert in AI inferencing and distributed systems at scale with 15+ years experience; hands-on with inference engines, model optimization, GPU/TPU orchestration, Kubernetes, RDMA/DPDK; strong architecture, communication, and mentorship skills.
Alpha Design AI: AI-native EDA platform for semiconductor design and verification.
Experience with large-scale ML systems and GPU computing; strong Python and C++/CUDA skills; familiarity with vLLM, PyTorch, SGLang, Ray; experience deploying and optimizing LLMs, profiling and benchmarking inference.
NioNYSE: NIO: Designs and manufactures smart premium electric vehicles.
Master's degree or higher in CS/AI/EE or related; deep learning framework proficiency (PyTorch, TensorFlow, ONNX); model optimization and edge deployment experience; embedded Linux/RTOS/QNX and C/C++ skills; experience with heterogeneous platforms (NPU/DSP/GPU) preferred.
PyTorch, TensorFlow, ONNX, TensorRT, llama.cpp, MNN, TNN, C/C++, QNX, ARM
NIONYSE: NIO: Designs and manufactures premium smart electric vehicles and technology
5+ YOE5+ years building and optimizing large-scale LLM/VLM inference systems; strong C/C++ and performance engineering skills; GPU/NPU programming (CUDA), PyTorch/TensorFlow, and BS/MS in CS/CE or related field required.
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
10+ YOE10+ years in data engineering/ML infrastructure, distributed systems expertise, Python and a systems language, experience with Ray or Spark, GPU inference optimization, large-scale databases and data curation for model training.