7 gpu inference performance engineer jobs at 6 companies in New York
3mo
Save
Mark Applied
Hide
3mo
Inference Performance Engineer
New York, New York, United States
HybridFull Time
Material: Specialized inference cloud platform for high-performance AI workloads.
BS in CS/EE or related field; proficiency in Rust/Go/Python/C++; knowledge of concurrency, tail latency; experience with model serving; GPU/ASIC programming; low-precision inference; profiling and benchmarking.
LTIMindtreeNational Stock Exchange of India: LTIM: Global technology consulting and digital solutions.
8+ YOERequires 8+ years in infrastructure or ML engineering, hands-on NVIDIA GPU operations, Kubernetes GPU workloads, model serving, GPU scheduling and partitioning, KEDA autoscaling, and inference performance optimization.
Amazon Web Services (AWS), Amazon Elastic Kubernetes Service (EKS), AWS CloudFormation, NVIDIA AI Enterprise (NVAIE), NVIDIA GPU Operator, CUDA, Data Center GPU Manager (DCGM), NIM, NVIDIA Dynamo, OpenAI-compatible API, RunAI, Kubernetes, KEDA, NVIDIA Triton Inference Server, TensorRT-LLM, vLLM, NVIDIA Multi-Instance GPU (MIG), Amazon Outposts
LM Studio: Desktop software for running large language models locally and privately.
Significant production ML, inference runtime, or performance infrastructure experience; strong Python and C++; transformer and inference expertise; CPU/GPU profiling; PyTorch and inference system experience.
Santa Clara or Washington or Texas or New York or Washington or Massachusetts
$184k-$357k/yrRemoteFull Time
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOE3+ MgmtRequires MS, PhD, or equivalent experience; 6+ years software development; 3+ years technical leadership or engineering management; C/C++, GPU programming, performance optimization, and production deep learning deployment.
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.