6 gpu performance profiling engineer jobs at 6 companies in New York City, NY

1mo
Save
Mark Applied
Hide
GPU Performance Engineer
New York City, New York, United States
$165k-$300k/yr HybridFull Time
Two Sigma
Two Sigma: Quantitative investment and trading firm using data science to manage institutional capital and trade financial markets.
1+ YOEExpert CUDA and GPU architecture knowledge, C++ and Python skills, BS/MS in STEM, minimum 1 year relevant experience, GPU profiling and optimization experience.
CUDA, Nsight Systems, Nsight Compute, RAPIDS, CUTLASS, cuBLAS, TensorRT, NCCL, MPI, cuDF, C++, Python
3mo
Save
Mark Applied
Hide
Inference Performance Engineer
New York, New York, United States
HybridFull Time
Material Group
Material Group: Material Group is an Austin-based specialized talent practice recruiting technical workers for companies building artificial general intelligence.
BS in CS/EE or related field; proficiency in Rust/Go/Python/C++; knowledge of concurrency, tail latency; experience with model serving; GPU/ASIC programming; low-precision inference; profiling and benchmarking.
Rust, Go, Python, C++, vLLM, TensorRT-LLM, llama.cpp, CUDA, ROCm, Triton, TGI, SGLang, Nsight, perf
2w
Save
Mark Applied
Hide
Machine Leaning Performance Engineer (Inference)
New York City, New York, United States
$200k-$300k/yr HybridFull Time
Tower Research Capital
Tower Research Capital: Proprietary quantitative trading firm employing traders, engineers, researchers, and business-support staff to trade global financial markets.
2+ YOERequires 2+ years optimizing deep learning inference, PyTorch or JAX, Python/C++, mixed-precision computation, custom GPU kernels, optimization libraries, compilers, profiling tools, and GPU microarchitecture expertise.
PyTorch, JAX, Python, C++, Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS, Nsight Systems, Nsight Compute, FPGA, ASIC
2mo
Save
Mark Applied
Hide
Staff ML Engineer, Generative Model Performance & Efficiency
Mountain View or New York City
$251k-$310k/yr OnsiteFull Time
Waymo
Waymo: Autonomous driving technology and robotaxi service provider.
5+ YOEMS/PhD in CS/ML/Robotics, 5+ years deep learning experience (Transformers, Diffusion, MoEs), proficiency with JAX/Flax and ML tooling, experience with model compression and profiling, strong Python and C++ skills.
JAX, Flax, TensorFlow, PyTorch, XLA, xprof, Perfetto, NVIDIA Nsight, Python, C++, Gemax, XManager, TPUs, GPUs
3w
Save
Mark Applied
Hide
Software Engineer, Inference Runtime
New York City, New York, United States
$150k-$350k/yr HybridFull Time
LM Studio
LM Studio: Private U.S. AI software building local and cloud tools for running large language models on personal computers.
Significant production ML, inference runtime, or performance infrastructure experience; strong Python and C++; transformer and inference expertise; CPU/GPU profiling; PyTorch and inference system experience.
Python, C++, PyTorch, llama.cpp, MLX, ExecuTorch, vLLM, SGLang, TensorRT-LLM, CUDA, Metal, Vulkan, ROCm
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: AI research developing safe and steerable AI systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD