7 gpu inference performance engineer jobs at 6 companies in New York

3mo
Save
Mark Applied
Hide
Inference Performance Engineer
New York, New York, United States
HybridFull Time
Material
Material: Specialized inference cloud platform for high-performance AI workloads.
BS in CS/EE or related field; proficiency in Rust/Go/Python/C++; knowledge of concurrency, tail latency; experience with model serving; GPU/ASIC programming; low-precision inference; profiling and benchmarking.
Rust, Go, Python, C++, vLLM, TensorRT-LLM, llama.cpp, CUDA, ROCm, Triton, TGI, SGLang, Nsight, perf
2w
Save
Mark Applied
Hide
Machine Leaning Performance Engineer (Inference)
New York City, New York, United States
$200k-$300k/yr HybridFull Time
Tower Research Capital
Tower Research Capital: Global quantitative trading firm developing automated algorithmic strategies.
2+ YOERequires 2+ years optimizing deep learning inference, PyTorch or JAX, Python/C++, mixed-precision computation, custom GPU kernels, optimization libraries, compilers, profiling tools, and GPU microarchitecture expertise.
PyTorch, JAX, Python, C++, Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS, Nsight Systems, Nsight Compute, FPGA, ASIC
5d
Save
Mark Applied
Hide
NVIDIA GPU AI SME
Albany, New York, United States
$150k-$164k/yr RemoteFull Time
LTIMindtree
LTIMindtreeNational Stock Exchange of India: LTIM: Global technology consulting and digital solutions.
8+ YOERequires 8+ years in infrastructure or ML engineering, hands-on NVIDIA GPU operations, Kubernetes GPU workloads, model serving, GPU scheduling and partitioning, KEDA autoscaling, and inference performance optimization.
Amazon Web Services (AWS), Amazon Elastic Kubernetes Service (EKS), AWS CloudFormation, NVIDIA AI Enterprise (NVAIE), NVIDIA GPU Operator, CUDA, Data Center GPU Manager (DCGM), NIM, NVIDIA Dynamo, OpenAI-compatible API, RunAI, Kubernetes, KEDA, NVIDIA Triton Inference Server, TensorRT-LLM, vLLM, NVIDIA Multi-Instance GPU (MIG), Amazon Outposts
2w
Save
Mark Applied
Hide
Software Engineer, Inference Runtime
New York City, New York, United States
$150k-$350k/yr HybridFull Time
LM Studio
LM Studio: Desktop software for running large language models locally and privately.
Significant production ML, inference runtime, or performance infrastructure experience; strong Python and C++; transformer and inference expertise; CPU/GPU profiling; PyTorch and inference system experience.
Python, C++, PyTorch, llama.cpp, MLX, ExecuTorch, vLLM, SGLang, TensorRT-LLM, CUDA, Metal, Vulkan, ROCm
1mo
Save
Mark Applied
Hide
Senior Deep Learning Software Engineer, Inference
California or Texas or New York or Washington or Massachusetts
$152k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOEMasters/PhD or equivalent,5+ years software development,excellent C/C++ skills,CUDA and GPU programming experience preferred,experience optimizing/deploying DL inference,Python and performance profiling experience helpful.
CUTLASS, OAI Triton, NCCL, CUDA, vLLM, SGLang, FlashInfer, PyTorch, NVSHMEM, C/C++, Python
2w
Save
Mark Applied
Hide
Engineering Manager, Deep Learning Inference
Santa Clara or Washington or Texas or New York or Washington or Massachusetts
$184k-$357k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOE3+ MgmtRequires MS, PhD, or equivalent experience; 6+ years software development; 3+ years technical leadership or engineering management; C/C++, GPU programming, performance optimization, and production deep learning deployment.
vLLM, SGLang, FlashInfer, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, C/C++, Python, PyTorch, TensorRT-LLM, Agile
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD