Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or United States
$124k-$242k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
3+ YOEBachelor's, master's, or doctoral degree in a related field or equivalent experience; 3+ years' engineering experience; GPU profiling, Python, C++/CUDA, and AI inference optimization expertise required.
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's or master's degree in a technical discipline; proficient in C/C++, Python, CUDA, GPU architecture, deep learning operators, inference compilation, performance analysis, and distributed model inference.
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
San Jose or Durham or Mexico City or Vancouver or Bengaluru or Pune or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
$171k-$257k/yrHybridFull Time
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
8+ YOE8+ years building distributed, high-performance systems; strong Go/Python, Docker, Kubernetes, CI/CD; knowledge of datacenter, OS internals, virtualization, and ML frameworks.
AI Infra Engineer - Large Model Inference Systems (Multimodal/LLM/VLM)
San Jose, California, United States
$156k-$388k/yrOnsiteFull Time
TikTok: Global short-form video hosting and social media platform.
2+ YOEBachelor's in CS or related,2+ years in high-performance computing or distributed scheduling,experience with large-model inference and system design,knowledge of asynchronous scheduling and resource pooling.
Open Source Software Engineer — ML Systems & AMD Hardware
San Jose, California, United States
$179k-$306k/yrHybridFull Time
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Systems software engineering experience with LLM inference, GPU kernel optimization, open source development, performance analysis, and GPU programming environments; bachelor's or master's degree preferred.
NIONYSE: NIO: Designs and manufactures premium smart electric vehicles and technology
5+ YOE5+ years building and optimizing large-scale LLM/VLM inference systems; strong C/C++ and performance engineering skills; GPU/NPU programming (CUDA), PyTorch/TensorFlow, and BS/MS in CS/CE or related field required.
Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Cupertino, California, United States
$165k-$224k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
3+ YOEBachelor's degree or equivalent; 3+ years professional software development and systems design experience; C++ or Python; machine learning, LLM, performance, memory, parallel computing, debugging, and profiling expertise.