196 inference jobs at 108 companies in New York City, NY

1w
Save
Mark Applied
Hide
Software Engineer, Inference Runtime
New York City, New York, United States
$150k-$350k/yr HybridFull Time
LM Studio
LM Studio: Desktop software for running large language models locally and privately.
Significant production ML, inference runtime, or performance infrastructure experience; strong Python and C++; transformer and inference expertise; CPU/GPU profiling; PyTorch and inference system experience.
Python, C++, PyTorch, llama.cpp, MLX, ExecuTorch, vLLM, SGLang, TensorRT-LLM, CUDA, Metal, Vulkan, ROCm
3mo
Save
Mark Applied
Hide
Inference Performance Engineer
New York, New York, United States
HybridFull Time
Material
Material: Specialized inference cloud platform for high-performance AI workloads.
BS in CS/EE or related field; proficiency in Rust/Go/Python/C++; knowledge of concurrency, tail latency; experience with model serving; GPU/ASIC programming; low-precision inference; profiling and benchmarking.
Rust, Go, Python, C++, vLLM, TensorRT-LLM, llama.cpp, CUDA, ROCm, Triton, TGI, SGLang, Nsight, perf
3w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
2w
Save
Mark Applied
Hide
AI Inference Engineer
San Francisco or United States or Toronto or New York City or Montreal
$165k-$330k/yr HybridFull Time
Baseten
Baseten: Scalable infrastructure platform for deploying and serving AI models.
2+ YOEDegree in CS/Engineering/Math,2+ years experience,production programming (Python preferred),familiarity with ML model lifecycle,strong communication and customer-facing skills.
Python, Docker, Whisper, ComfyUI
3mo
Save
Mark Applied
Hide
Head of Inference, Stealth Edge AI Co
New York City, New York, United States
HybridFull Time
Montauk Capital
Montauk Capital: Investment firm building and funding climate technology companies.
Hands-on technical leader with production inference systems experience, architecture definition, distributed multi-GPU optimization, and startup–level execution, strong C++/CUDA/Rust skills, and ability to ship fast.
vLLM, TensorRT-LLM, Triton Inference Server, CUDA, C++, Rust, Kubernetes, Ray, NVidia, KV-cache, observability tooling
1mo
Save
Mark Applied
Hide
Member of Technical Staff - Inference Research
New York City or San Francisco
$150k-$350k/yr OnsiteFull Time
Modal
Modal: Serverless cloud platform for running AI and data workloads
Research-leaning or systems background in LLM inference; experience with kernels, quantization, schedulers, and autoscaling; record of shipping research/systems; able to take research bets end-to-end and work onsite in NYC or San Francisco.
Flash Attention 4, Python
1mo
Save
Mark Applied
Hide
Machine Learning Engineer - Inference Maintainer & Developer Experience
New York City or San Francisco or United States or Europe
RemoteFull Time
Roboflow
Roboflow: Platform for building and deploying custom computer vision models.
5+ YOE5+ years building and operating production ML systems, strong CV/inference foundation, CI/CD and test infra experience, proficiency with PyTorch/TensorFlow/ONNX/TensorRT/vLLM, and experience with image/video processing tools.
inference, PyTorch, TensorFlow, ONNX, TensorRT, vLLM, OpenCV, DeepStream, Pillow, PyAV, GitHub, CI/CD
2mo
Save
Mark Applied
Hide
Senior Data Scientist, Causal Inference
New York City, New York, United States
$148k-$185k/yr HybridFull Time
Lyft
LyftNASDAQ: LYFT: Provides an on-demand ride-hailing and multimodal transportation platform.
4+ YOEAdvanced degree or equivalent experience in statistics/economics/math, 4+ years in causal inference or data science, expertise in causal inference and MMM, SQL and Python proficiency, strong communication and critical thinking.
SQL, Python
2mo
Save
Mark Applied
Hide
AI Research Engineer, Inference
New York City, New York, United States
$250k-$300k/yr OnsiteFull Time
Hudson River Trading
Hudson River Trading: A quantitative firm using technology to trade global financial markets.
2+ YOETwo+ years of experience building deep learning systems; strong low-level engineering in CUDA/Triton/CuTe or PyTorch/JAX; experience with inference optimization; familiarity with multiple domains.
CUDA, PyTorch, JAX, CuTe DSL, FPGA, ASIC, CUDA Graphs
1mo
Save
Mark Applied
Hide
Technical Program Manager, Inference
Livingston or New York or Sunnyvale or Bellevue
$198k-$264k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
8+ YOEBachelor's degree in a technical field or equivalent experience, 8+ years TPM experience in distributed systems/cloud/AI platforms, strong technical fluency in inference systems and GPU compute, program delivery and launch readiness experience.
1w
Save
Mark Applied
Hide
Machine Leaning Performance Engineer (Inference)
New York City, New York, United States
$200k-$300k/yr HybridFull Time
Tower Research Capital
Tower Research Capital: Global quantitative trading firm developing automated algorithmic strategies.
2+ YOERequires 2+ years optimizing deep learning inference, PyTorch or JAX, Python/C++, mixed-precision computation, custom GPU kernels, optimization libraries, compilers, profiling tools, and GPU microarchitecture expertise.
PyTorch, JAX, Python, C++, Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS, Nsight Systems, Nsight Compute, FPGA, ASIC
1mo
Save
Mark Applied
Hide
Staff Backend Engineer, ML Inference Systems
Mountain View or Montreal or Texas or New York
$245k-$318k/yr RemoteFull Time
Unity
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
5+ YOE5+ years building and operating distributed systems; expertise in Golang, cloud (GCP), Kubernetes, monitoring with Prometheus/Grafana; experience with high-throughput, low-latency inference systems.
Golang, GCP, Kubernetes, Prometheus, Grafana, Docker, NVIDIA Triton Inference Server
1mo
Save
Mark Applied
Hide
Senior Customer Success Manager, Managed Inference
New York, New York, United States
$190k-$215k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
3+ YOE3+ years supporting enterprise cloud/AI/ML customers; Bachelor’s degree; strong technical foundation with Kubernetes, containers, APIs, and GPU inference; excellent communication and customer relationship skills.
Kubernetes, APIs
1mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (FM Hosting, LLM Inference)
New York or McLean or Cambridge or San Jose
$251k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
6+ YOEBachelor's in CS/AI/EE/CE or related with 6+ years (or Master's with 4+ years); 6+ years programming with Python, Go, Scala, or Java; experience deploying scalable AI systems, LLM inference, similarity search, and optimization of training/inference.
AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++, C#, Golang, AWS, Google Cloud, Azure
1w
Save
Mark Applied
Hide
Engineering Lead, Inference
New York City, New York, United States
HybridFull Time
Mistral AI
Mistral AI: Developing frontier artificial intelligence models and enterprise AI solutions.
Experience building and managing engineering teams; proficiency in Python, C#, Golang, TypeScript, and a web framework; distributed systems, databases, caching, messaging, cross-functional collaboration, and end-to-end delivery.
Python, FastAPI, Django, Flask, C#, Golang, TypeScript
1mo
Save
Mark Applied
Hide
Machine Learning Engineer, Causal Inference, Level 5
Santa Monica or Los Angeles or San Francisco or Palo Alto or New York City
$178k-$313k/yr OnsiteFull Time
Snap
SnapNYSE: SNAP: Develops social media applications and augmented reality technology.
5+ YOE5+ years post-Bachelor's ML experience with expertise in causal inference, experimentation, uplift modeling, and productionizing models; proficient in Python, pandas, NumPy, scikit-learn; strong communication and mentorship skills.
Python, pandas, NumPy, scikit-learn, CausalM, CausalML, EconML, DoWhy
1mo
Save
Mark Applied
Hide
Machine Learning Engineer, Causal Inference, Level 5
Los Angeles or San Francisco or Palo Alto or New York or Santa Monica
$178k-$313k/yr OnsiteFull Time
Snap
SnapNYSE: SNAP: Provides visual messaging software and augmented reality wearable devices.
5+ YOE5+ years post-Bachelor's ML experience (or Master's/PhD with reduced experience), strong causal inference and experimentation experience, proficiency in Python and ML libraries, production ML experience, strong communication and mentorship skills.
Python, pandas, NumPy, scikit-learn, CausalM, CausalML, EconML, DoWhy
17h
Save
Mark Applied
Hide
Member of Research Staff, Causal Inference, Voleon Securities
New York City or United States or Berkeley
$250k-$275k/yr HybridFull Time
The Voleon Group
The Voleon Group: Quantitative investment management firm using machine learning strategies.
Ph.D.-level coursework, causal inference and statistics expertise, top-tier research publications, strong mathematical ability, Python production coding willingness, and interest in financial applications.
Python
2mo
Save
Mark Applied
Hide
Performance Engineer, Inference Systems
San Francisco or New York City or Seattle
$350k-$850k/yr OnsiteFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Hands-on performance engineering with Python, data analysis, and cross-layer investigations; strong communication of quantitative results.
Python, SQL, Pandas
2mo
Save
Mark Applied
Hide
AI Research Engineer, Inference
London or New York City
$250k-$300k/yr OnsiteFull Time
Hudson River Trading
Hudson River Trading: Proprietary quantitative trading firm specializing in automated market making.
2+ YOEStrong engineering skills and 2+ years building deep learning systems; experience with GPU kernels, PyTorch, JAX, XLA, CUDA Graphs, FPGA, or ASICs preferred.
CUDA, Triton, Pallas, CuTe DSL, PyTorch, JAX, XLA, CUDA Graphs, FPGA, ASICs