196 inference jobs at 108 companies in New York City, NY
1w
Save
Mark Applied
Hide
1w
Software Engineer, Inference Runtime
New York City, New York, United States
$150k-$350k/yrHybridFull Time
LM Studio: Desktop software for running large language models locally and privately.
Significant production ML, inference runtime, or performance infrastructure experience; strong Python and C++; transformer and inference expertise; CPU/GPU profiling; PyTorch and inference system experience.
Material: Specialized inference cloud platform for high-performance AI workloads.
BS in CS/EE or related field; proficiency in Rust/Go/Python/C++; knowledge of concurrency, tail latency; experience with model serving; GPU/ASIC programming; low-precision inference; profiling and benchmarking.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
San Francisco or United States or Toronto or New York City or Montreal
$165k-$330k/yrHybridFull Time
Baseten: Scalable infrastructure platform for deploying and serving AI models.
2+ YOEDegree in CS/Engineering/Math,2+ years experience,production programming (Python preferred),familiarity with ML model lifecycle,strong communication and customer-facing skills.
Montauk Capital: Investment firm building and funding climate technology companies.
Hands-on technical leader with production inference systems experience, architecture definition, distributed multi-GPU optimization, and startup–level execution, strong C++/CUDA/Rust skills, and ability to ship fast.
Modal: Serverless cloud platform for running AI and data workloads
Research-leaning or systems background in LLM inference; experience with kernels, quantization, schedulers, and autoscaling; record of shipping research/systems; able to take research bets end-to-end and work onsite in NYC or San Francisco.
New York City or San Francisco or United States or Europe
RemoteFull Time
Roboflow: Platform for building and deploying custom computer vision models.
5+ YOE5+ years building and operating production ML systems, strong CV/inference foundation, CI/CD and test infra experience, proficiency with PyTorch/TensorFlow/ONNX/TensorRT/vLLM, and experience with image/video processing tools.
LyftNASDAQ: LYFT: Provides an on-demand ride-hailing and multimodal transportation platform.
4+ YOEAdvanced degree or equivalent experience in statistics/economics/math, 4+ years in causal inference or data science, expertise in causal inference and MMM, SQL and Python proficiency, strong communication and critical thinking.
Hudson River Trading: A quantitative firm using technology to trade global financial markets.
2+ YOETwo+ years of experience building deep learning systems; strong low-level engineering in CUDA/Triton/CuTe or PyTorch/JAX; experience with inference optimization; familiarity with multiple domains.
CUDA, PyTorch, JAX, CuTe DSL, FPGA, ASIC, CUDA Graphs
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
8+ YOEBachelor's degree in a technical field or equivalent experience, 8+ years TPM experience in distributed systems/cloud/AI platforms, strong technical fluency in inference systems and GPU compute, program delivery and launch readiness experience.
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
5+ YOE5+ years building and operating distributed systems; expertise in Golang, cloud (GCP), Kubernetes, monitoring with Prometheus/Grafana; experience with high-throughput, low-latency inference systems.
Golang, GCP, Kubernetes, Prometheus, Grafana, Docker, NVIDIA Triton Inference Server
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
3+ YOE3+ years supporting enterprise cloud/AI/ML customers; Bachelor’s degree; strong technical foundation with Kubernetes, containers, APIs, and GPU inference; excellent communication and customer relationship skills.
Senior Lead AI Engineer (FM Hosting, LLM Inference)
New York or McLean or Cambridge or San Jose
$251k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
6+ YOEBachelor's in CS/AI/EE/CE or related with 6+ years (or Master's with 4+ years); 6+ years programming with Python, Go, Scala, or Java; experience deploying scalable AI systems, LLM inference, similarity search, and optimization of training/inference.
Mistral AI: Developing frontier artificial intelligence models and enterprise AI solutions.
Experience building and managing engineering teams; proficiency in Python, C#, Golang, TypeScript, and a web framework; distributed systems, databases, caching, messaging, cross-functional collaboration, and end-to-end delivery.
Santa Monica or Los Angeles or San Francisco or Palo Alto or New York City
$178k-$313k/yrOnsiteFull Time
SnapNYSE: SNAP: Develops social media applications and augmented reality technology.
5+ YOE5+ years post-Bachelor's ML experience with expertise in causal inference, experimentation, uplift modeling, and productionizing models; proficient in Python, pandas, NumPy, scikit-learn; strong communication and mentorship skills.
5+ YOE5+ years post-Bachelor's ML experience (or Master's/PhD with reduced experience), strong causal inference and experimentation experience, proficiency in Python and ML libraries, production ML experience, strong communication and mentorship skills.
Member of Research Staff, Causal Inference, Voleon Securities
New York City or United States or Berkeley
$250k-$275k/yrHybridFull Time
The Voleon Group: Quantitative investment management firm using machine learning strategies.
Ph.D.-level coursework, causal inference and statistics expertise, top-tier research publications, strong mathematical ability, Python production coding willingness, and interest in financial applications.
Hudson River Trading: Proprietary quantitative trading firm specializing in automated market making.
2+ YOEStrong engineering skills and 2+ years building deep learning systems; experience with GPU kernels, PyTorch, JAX, XLA, CUDA Graphs, FPGA, or ASICs preferred.