6 gpu performance profiling engineer jobs at 6 companies in New York City, NY
1mo
Save
Mark Applied
Hide
1mo
GPU Performance Engineer
New York City, New York, United States
$165k-$300k/yrHybridFull Time
Two Sigma: Quantitative investment and trading firm using data science to manage institutional capital and trade financial markets.
1+ YOEExpert CUDA and GPU architecture knowledge, C++ and Python skills, BS/MS in STEM, minimum 1 year relevant experience, GPU profiling and optimization experience.
Material Group: Material Group is an Austin-based specialized talent practice recruiting technical workers for companies building artificial general intelligence.
BS in CS/EE or related field; proficiency in Rust/Go/Python/C++; knowledge of concurrency, tail latency; experience with model serving; GPU/ASIC programming; low-precision inference; profiling and benchmarking.
Tower Research Capital: Proprietary quantitative trading firm employing traders, engineers, researchers, and business-support staff to trade global financial markets.
2+ YOERequires 2+ years optimizing deep learning inference, PyTorch or JAX, Python/C++, mixed-precision computation, custom GPU kernels, optimization libraries, compilers, profiling tools, and GPU microarchitecture expertise.
Staff ML Engineer, Generative Model Performance & Efficiency
Mountain View or New York City
$251k-$310k/yrOnsiteFull Time
Waymo: Autonomous driving technology and robotaxi service provider.
5+ YOEMS/PhD in CS/ML/Robotics, 5+ years deep learning experience (Transformers, Diffusion, MoEs), proficiency with JAX/Flax and ML tooling, experience with model compression and profiling, strong Python and C++ skills.
LM Studio: Private U.S. AI software building local and cloud tools for running large language models on personal computers.
Significant production ML, inference runtime, or performance infrastructure experience; strong Python and C++; transformer and inference expertise; CPU/GPU profiling; PyTorch and inference system experience.
Anthropic: AI research developing safe and steerable AI systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.