21 inference optimization engineer jobs at 11 companies in Passaic, NJ

4w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
4w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
2w
Save
Mark Applied
Hide
Machine Leaning Performance Engineer (Inference)
New York City, New York, United States
$200k-$300k/yr HybridFull Time
Tower Research Capital
Tower Research Capital: Global quantitative trading firm developing automated algorithmic strategies.
2+ YOERequires 2+ years optimizing deep learning inference, PyTorch or JAX, Python/C++, mixed-precision computation, custom GPU kernels, optimization libraries, compilers, profiling tools, and GPU microarchitecture expertise.
PyTorch, JAX, Python, C++, Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS, Nsight Systems, Nsight Compute, FPGA, ASIC
3mo
Save
Mark Applied
Hide
Head of Inference, Stealth Edge AI Co
New York City, New York, United States
HybridFull Time
Montauk Capital
Montauk Capital: Investment firm building and funding climate technology companies.
Hands-on technical leader with production inference systems experience, architecture definition, distributed multi-GPU optimization, and startup–level execution, strong C++/CUDA/Rust skills, and ability to ship fast.
vLLM, TensorRT-LLM, Triton Inference Server, CUDA, C++, Rust, Kubernetes, Ray, NVidia, KV-cache, observability tooling
2mo
Save
Mark Applied
Hide
AI Research Engineer, Inference
New York City, New York, United States
$250k-$300k/yr OnsiteFull Time
Hudson River Trading
Hudson River Trading: A quantitative firm using technology to trade global financial markets.
2+ YOETwo+ years of experience building deep learning systems; strong low-level engineering in CUDA/Triton/CuTe or PyTorch/JAX; experience with inference optimization; familiarity with multiple domains.
CUDA, PyTorch, JAX, CuTe DSL, FPGA, ASIC, CUDA Graphs
2mo
Save
Mark Applied
Hide
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
6+ YOEBachelor's plus 6 years or master's plus 4 years developing AI/ML technologies, and 6 years programming with Python, Go, Scala, or Java. Cloud AI deployment and team leadership are preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, NeMo Guardrails, PyTorch, Python, Go, Scala, Java, AWS, Google Cloud, Azure, C++, C#, Golang
1mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (FM Hosting, LLM Inference)
New York or McLean or Cambridge or San Jose
$251k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
6+ YOEBachelor's in CS/AI/EE/CE or related with 6+ years (or Master's with 4+ years); 6+ years programming with Python, Go, Scala, or Java; experience deploying scalable AI systems, LLM inference, similarity search, and optimization of training/inference.
AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++, C#, Golang, AWS, Google Cloud, Azure
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD
2mo
Save
Mark Applied
Hide
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: A diversified financial services providing banking and credit products.
6+ YOEBachelor's degree plus 6 years or master's degree plus 4 years developing AI/ML technologies; 6 years programming with Python, Go, Scala, or Java; cloud AI deployment experience preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, Google Cloud, Azure, C++, C#, Golang
2w
Save
Mark Applied
Hide
Senior AI/ML Engineer
New York City, New York, United States
RemoteFull Time
Pypestream
Pypestream: Enterprise-grade conversational AI agents for customer service automation.
Build inference infrastructure, own model evaluation pipelines, optimize latency, and integrate orchestration engines with frontier large language models.
4w
Save
Mark Applied
Hide
AI Foundational Model Engineer
Jersey City, New Jersey, United States
$140k-$210k/yr OnsiteFull Time
NTT DATA
NTT DATA: Global provider of IT and business consulting services.
7+ YOE7+ years in AI/ML or platform engineering; hands-on LLM, RAG, embeddings, PyTorch/TensorFlow, Python, Terraform, CI/CD, cloud-native deployment, model evaluation, inference optimization, and secure data handling.
Python, PyTorch, TensorFlow, Hugging Face, LangChain, LlamaIndex, Semantic Kernel, Terraform, CI/CD, AWS Bedrock, Amazon SageMaker, OpenSearch, Kendra, AWS Lambda, EKS, ECS, Azure OpenAI, Vertex AI, Databricks, vLLM, Triton, MLflow, Kubeflow, Kubernetes
3w
Save
Mark Applied
Hide
Staff Software Engineer - Data Cloud Applied ML
San Francisco or Seattle or New York City
$189k-$315k/yr OnsiteFull Time
Rippling
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Python, Go, Java, Kubernetes
3w
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Chicago or New York City or San Francisco or Seattle or Sunnyvale
$182k-$202k/yr OnsiteFull Time
Uber
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
4+ YOE4+ years building ML models; BS in CS/CE or related; experience with PyTorch, causal inference or constrained optimization preferred; product and marketplace experience a plus.
PyTorch
2mo
Save
Mark Applied
Hide
Senior AI Engineer (US)
Boston or New York City
HybridFull Time
Assail
Assail: Autonomous AI platform for offensive security testing.
5+ YOE5+ years building production ML/AI systems with 2+ years on LLMs/agents; deep Python; fine-tuning (SFT, DPO/GRPO, RLHF/RLAIF); transformer expertise; PyTorch/Hugging Face/DeepSpeed stack; inference optimization; retrieval/vector pipelines; Kubernetes.
Python, PyTorch, Hugging Face, DeepSpeed, FSDP, accelerate, vLLM, TensorRT-LLM, Kubernetes