20 inference optimization engineer jobs at 12 companies in Peabody, MA

1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Boston, Massachusetts, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and inference optimization expertise, experience with CUDA/Triton/ROCm, attention-layer and kernel-level optimization, strong system design and leadership through influence.
AITER, CUDA, ROCm, TensorRT, Triton, FlashAttention
1w
Save
Mark Applied
Hide
Senior Inference Engineer, AGI
Sunnyvale or Boston or Seattle or Los Angeles County
$193k-$262k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOERequires 5+ years software development, 4+ years systems architecture, a computer science bachelor's degree, 2+ years neural inference optimization, GPU optimization, real-time systems, and technical leadership experience.
Nsight Compute, Nsight Systems, vLLM, PyTorch, TensorRT-LLM, CUTLASS, Triton, CUDA, PTX, FlashAttention, NCCL, NVLink, AWS Neuron, Trainium, Microsoft Excel
2mo
Save
Mark Applied
Hide
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
6+ YOEBachelor's plus 6 years or master's plus 4 years developing AI/ML technologies, and 6 years programming with Python, Go, Scala, or Java. Cloud AI deployment and team leadership are preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, NeMo Guardrails, PyTorch, Python, Go, Scala, Java, AWS, Google Cloud, Azure, C++, C#, Golang
3w
Save
Mark Applied
Hide
Staff/Principal DevOps Engineer, AI Inference
Cambridge, Massachusetts, United States
$192k-$272k/yr OnsiteFull Time
Lila Sciences
Lila Sciences: Develops an AI platform for autonomous scientific research and discovery.
Expertise operating GPU/accelerator infrastructure for ML, Kubernetes and AWS deployment experience, infrastructure-as-code (Terraform, Helm), Python proficiency, networking and performance optimization for low-latency inference.
Kubernetes, vLLM, Triton Inference Server, TGI, Terraform, Helm, EKS, EC2, S3, EFA, IAM, NCCL, Python, Rust, Go, CUDA
1mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (FM Hosting, LLM Inference)
New York or McLean or Cambridge or San Jose
$251k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
6+ YOEBachelor's in CS/AI/EE/CE or related with 6+ years (or Master's with 4+ years); 6+ years programming with Python, Go, Scala, or Java; experience deploying scalable AI systems, LLM inference, similarity search, and optimization of training/inference.
AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++, C#, Golang, AWS, Google Cloud, Azure
2mo
Save
Mark Applied
Hide
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: A diversified financial services providing banking and credit products.
6+ YOEBachelor's degree plus 6 years or master's degree plus 4 years developing AI/ML technologies; 6 years programming with Python, Go, Scala, or Java; cloud AI deployment experience preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, Google Cloud, Azure, C++, C#, Golang
1mo
Save
Mark Applied
Hide
Staff Embedded ML Engineer, Edge AI
Boston, Massachusetts, United States
$186k-$245k/yr HybridFull Time
SimpliSafe
SimpliSafe: Provides wireless home security systems and professional monitoring services.
8+ YOE8+ years embedded/performance engineering experience, strong C/C++ skills, experience optimizing on-device ML inference and runtimes (TFLite/ONNX/TensorRT), profiling and debugging performance under constrained hardware.
C, C++, TFLite, ONNX Runtime, TensorRT, XNNPACK, QNNPACK, oneDNN, CMSIS-NN, perf, flame graphs
1mo
Save
Mark Applied
Hide
Systems ML Engineer (Member of the Technical Staff)
Cambridge, Massachusetts, United States
OnsiteFull Time
Transfyr
Transfyr: Building physical AI infrastructure for scientific research and automation.
Experience optimizing and deploying large-scale ML models for training and inference, profiling and custom GPU kernel development, distributed training, cloud and edge deployment, and infrastructure automation.
Nsight, PyTorch Profiler, PyTorch Distributed, PyTorch, JAX, Triton, CUDA, Kubernetes, Terraform, AWS, NCCL
2mo
Save
Mark Applied
Hide
AI Systems Engineer
Boston, Massachusetts, United States
OnsiteFull Time
Glia AI
Glia AI: Automated AI systems engineering for high-performance infrastructure.
Experience with AI inference and distributed serving, strong systems and GPU optimization skills, proficiency in Python and PyTorch, familiarity with inference frameworks (vLLM, Triton, Ray Serve); PhD or equivalent experience preferred.
vLLM, PyTorch, Triton inference server, Ray Serve, CUDA, Python
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Westford or Austin or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE12+ years software engineering experience in GPU computing or ML systems, strong Python or C++ skills, GPU kernel optimization (CUDA/Triton), container engineering, and LLM inference knowledge.
Python, C++, CUDA, Triton, Docker, OCI, NCCL, RCCL, NVIDIA Container Toolkit
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Austin or Westford or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
12+ YOE12+ years software engineering experience with GPU computing or ML systems, strong Python or C++ skills, GPU kernel optimization (CUDA/Triton), container engineering, and LLM inference knowledge.
Python, C++, CUDA, Triton, Docker, OCI, NVIDIA Container Toolkit, NCCL, RCCL
1mo
Save
Mark Applied
Hide
Sr. Principal Software Engineer
Burlington or United States or Europe or Asia or North America
$141k-$226k/yr RemoteFull Time
Cerence
CerenceNASDAQ: CRNC: Develops AI-powered voice assistants and software for automotive vehicles.
Proven experience optimizing ML inference in production, deep GPU architecture knowledge, hands-on CUDA kernel development, quantization techniques (INT8/INT4/FP4/FP8/AWQ/GPTQ), and edge/embedded deployment expertise.
vLLM, TensorRT‑LLM, llama.cpp, QAIRT, CUDA, AWQ, GPTQ
2mo
Save
Mark Applied
Hide
Senior AI Engineer (US)
Boston or New York City
HybridFull Time
Assail
Assail: Autonomous AI platform for offensive security testing.
5+ YOE5+ years building production ML/AI systems with 2+ years on LLMs/agents; deep Python; fine-tuning (SFT, DPO/GRPO, RLHF/RLAIF); transformer expertise; PyTorch/Hugging Face/DeepSpeed stack; inference optimization; retrieval/vector pipelines; Kubernetes.
Python, PyTorch, Hugging Face, DeepSpeed, FSDP, accelerate, vLLM, TensorRT-LLM, Kubernetes
3mo
Save
Mark Applied
Hide
Software Solutions Architect
Austin or Boxborough or Markham
$212k-$318k/yr OnsiteFull Time
AMD
AMDNasdaq: AMD: Designs and sells microprocessors and graphics hardware for computers.
Architect and deliver enterprise software solutions leveraging AMD GPUs/APUs; engage customers and partners; strong software engineering, AI inference stack and performance optimization experience.
ROCm, ONNX Runtime, ONNX, PyTorch, TensorRT, CUDA, Docker, Kubernetes, C, C++, Python
1w
Save
Mark Applied
Hide
Member of Technical Staff, Performance & Capacity
Boston, Massachusetts, United States
HybridFull Time
Physical Superintelligence
Physical Superintelligence: Building AI systems to discover new physics at scale.
5+ YOERequires 5+ years with GPU and large-scale multi-node compute workloads, distributed training performance, networking, parallel file systems, AI training and inference optimization, and capacity decisions.
GPU, H100, B200, InfiniBand, RDMA, parallel file systems