36 inference engineer jobs at 19 companies in Massachusetts

1w
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Boston, Massachusetts, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and inference optimization expertise, experience with CUDA/Triton/ROCm, attention-layer and kernel-level optimization, strong system design and leadership through influence.
AITER, CUDA, ROCm, TensorRT, Triton, FlashAttention
1w
Save
Mark Applied
Hide
Staff/Principal DevOps Engineer, AI Inference
Cambridge, Massachusetts, United States
$192k-$272k/yr OnsiteFull Time
Lila Sciences
Lila Sciences: Develops an AI platform for autonomous scientific research and discovery.
Expertise operating GPU/accelerator infrastructure for ML, Kubernetes and AWS deployment experience, infrastructure-as-code (Terraform, Helm), Python proficiency, networking and performance optimization for low-latency inference.
Kubernetes, vLLM, Triton Inference Server, TGI, Terraform, Helm, EKS, EC2, S3, EFA, IAM, NCCL, Python, Rust, Go, CUDA
3w
Save
Mark Applied
Hide
Senior Deep Learning Software Engineer, Inference
California or Texas or New York or Washington or Massachusetts
$152k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOEMaster's/PhD or equivalent experience, 5+ years software development, strong C/C++ skills, GPU/CUDA experience, DL inference optimization and production deployment experience.
CUTLASS, OAI TRITON, NCCL, CUDA, Python, C, C++, vLLM, SGLang, FlashInfer, PyTorch, NVSHMEM
1mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (FM Hosting, LLM Inference)
New York or McLean or Cambridge or San Jose
$251k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
6+ YOEBachelor's in CS/AI/EE/CE or related with 6+ years (or Master's with 4+ years); 6+ years programming with Python, Go, Scala, or Java; experience deploying scalable AI systems, LLM inference, similarity search, and optimization of training/inference.
AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++, C#, Golang, AWS, Google Cloud, Azure
3w
Save
Mark Applied
Hide
Senior Deep Learning Software Engineer, Inference
California or Texas or New York or Washington or Massachusetts
$152k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOEMasters/PhD or equivalent,5+ years software development,excellent C/C++ skills,CUDA and GPU programming experience preferred,experience optimizing/deploying DL inference,Python and performance profiling experience helpful.
CUTLASS, OAI Triton, NCCL, CUDA, vLLM, SGLang, FlashInfer, PyTorch, NVSHMEM, C/C++, Python
2w
Save
Mark Applied
Hide
Staff Embedded ML Engineer, Edge AI
Boston, Massachusetts, United States
$186k-$245k/yr HybridFull Time
SimpliSafe
SimpliSafe: Provides wireless home security systems and professional monitoring services.
8+ YOE8+ years embedded/performance engineering experience, strong C/C++ skills, experience optimizing on-device ML inference and runtimes (TFLite/ONNX/TensorRT), profiling and debugging performance under constrained hardware.
C, C++, TFLite, ONNX Runtime, TensorRT, XNNPACK, QNNPACK, oneDNN, CMSIS-NN, perf, flame graphs
5d
Save
Mark Applied
Hide
Applied AI Research Engineer - Model Cost Optimization
Cambridge, Massachusetts, United States
$210k-$260k/yr OnsiteFull Time
Blitzy
Blitzy: Autonomous AI platform for enterprise software development.
Applied ML/AI research experience with hands-on LLM evaluation or fine-tuning, strong software engineering, familiarity with LLM evaluation and benchmarking, and understanding of inference cost/latency tradeoffs.
LLM
2mo
Save
Mark Applied
Hide
Principal Audio/DSP Systems Engineer
Boston, Massachusetts, United States
$200k-$275k/yr OnsiteFull Time
Analog Devices
Analog DevicesNASDAQ: ADI: Designs and manufactures semiconductors for signal processing and power management.
10+ YOEPhD in Electrical Engineering or related field; 10+ years in audio/signal processing; deployment on DSP/NPU; fixed-point, PTQ/QAT; C/Python/MATLAB; RTL verification; embedded AI inference.
C, Python, MATLAB, TensorFlow/TFLite, PyTorch/ONNX, DSP, RTL verification, SIMD, NPU, Fixed-point programming, Model quantization, Beamforming
2w
Save
Mark Applied
Hide
Systems ML Engineer (Member of the Technical Staff)
Cambridge, Massachusetts, United States
OnsiteFull Time
Transfyr
Transfyr: Building physical AI infrastructure for scientific research and automation.
Experience optimizing and deploying large-scale ML models for training and inference, profiling and custom GPU kernel development, distributed training, cloud and edge deployment, and infrastructure automation.
Nsight, PyTorch Profiler, PyTorch Distributed, PyTorch, JAX, Triton, CUDA, Kubernetes, Terraform, AWS, NCCL
4d
Save
Mark Applied
Hide
Member of Technical Staff, ML Engineer
Boston, Massachusetts, United States
OnsiteFull Time
Physical Superintelligence
Physical Superintelligence: Building AI systems to discover new physics at scale.
3+ YOE3+ years building and operating ML training or inference infrastructure; distributed training and model-serving experience; strong software engineering and ML fluency to debug training and inference; cloud and IaC experience preferred.
PyTorch, Ray, vLLM, SGLang, Triton, CUDA, GCP, AWS, Terraform
1mo
Save
Mark Applied
Hide
AI Systems Engineer
Boston, Massachusetts, United States
OnsiteFull Time
Glia AI
Glia AI: Automated AI systems engineering for high-performance infrastructure.
Experience with AI inference and distributed serving, strong systems and GPU optimization skills, proficiency in Python and PyTorch, familiarity with inference frameworks (vLLM, Triton, Ray Serve); PhD or equivalent experience preferred.
vLLM, PyTorch, Triton inference server, Ray Serve, CUDA, Python
2mo
Save
Mark Applied
Hide
Machine Learning Systems Engineer
Boston or Las Vegas or Pittsburgh or Remote
$144k-$192k/yr HybridFull Time
Motional
Motional: Develops autonomous vehicle technology for driverless ride-hailing and delivery.
Bachelor’s/Master’s/PhD in CS/CE or related; strong Python; PyTorch; ML training/inference optimization; problem solving.
Python, PyTorch, Nsight, PyTorch Profiler, Triton, CUDA
2mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer (Data Science Algorithms)
Boston, Massachusetts, United States
$150k-$210k/yr OnsiteFull Time
WHOOP
WHOOP: Wearable technology for personalized health and performance tracking.
4+ YOE4+ years ML engineer/related experience; BS in CS/Math/Data Science; strong Python; ML inference at scale; cloud/CI/CD; time-series experience preferred.
Python, cloud platforms (AWS or GCP), MLOps, CI/CD, time series data
1mo
Save
Mark Applied
Hide
Sr. Principal Software Engineer
Burlington or United States or Europe or Asia or North America
$141k-$226k/yr RemoteFull Time
Cerence
CerenceNASDAQ: CRNC: Develops AI-powered voice assistants and software for automotive vehicles.
Proven experience optimizing ML inference in production, deep GPU architecture knowledge, hands-on CUDA kernel development, quantization techniques (INT8/INT4/FP4/FP8/AWQ/GPTQ), and edge/embedded deployment expertise.
vLLM, TensorRT‑LLM, llama.cpp, QAIRT, CUDA, AWQ, GPTQ
2mo
Save
Mark Applied
Hide
Senior AI Engineer (US)
Boston or New York City
HybridFull Time
Assail
Assail: Autonomous AI platform for offensive security testing.
5+ YOE5+ years building production ML/AI systems with 2+ years on LLMs/agents; deep Python; fine-tuning (SFT, DPO/GRPO, RLHF/RLAIF); transformer expertise; PyTorch/Hugging Face/DeepSpeed stack; inference optimization; retrieval/vector pipelines; Kubernetes.
Python, PyTorch, Hugging Face, DeepSpeed, FSDP, accelerate, vLLM, TensorRT-LLM, Kubernetes
1mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Wellesley or New York City
$111k-$222k/yr HybridFull Time
CVS Health
CVS HealthNYSE: CVS: Provides retail pharmacy, health insurance, and pharmacy benefit management services.
5+ YOE5+ years building ML systems in production, 5+ years Python, familiarity with PyTorch/TensorFlow, MLOps (Kubeflow, Vertex), 2+ years SQL with Snowflake/BigQuery, 2+ years system design leadership, experience in reinforcement learning/causal inference/LLMs, mentoring experience, strong communication.
Python, PyTorch, TensorFlow, Kubeflow, Vertex, SQL, Snowflake, BigQuery, Node.js, React, Spark, Kafka, Airflow
2mo
Save
Mark Applied
Hide
Software Solutions Architect
Austin or Boxborough or Markham
$212k-$318k/yr OnsiteFull Time
AMD
AMDNasdaq: AMD: Designs and sells microprocessors and graphics hardware for computers.
Architect and deliver enterprise software solutions leveraging AMD GPUs/APUs; engage customers and partners; strong software engineering, AI inference stack and performance optimization experience.
ROCm, ONNX Runtime, ONNX, PyTorch, TensorRT, CUDA, Docker, Kubernetes, C, C++, Python
3w
Save
Mark Applied
Hide
AI Systems Architect (Models & Hardware Co-Design)
Santa Clara or Austin or Boston
$200k-$500k/yr OnsiteFull Time
Velaura AI
Velaura AI: Developing ultra-low-power silicon and IP for AI accelerators.
Deep knowledge of modern ML architectures (e.g., transformers), strong mathematical foundations, experience with ML training/inference, and familiarity with ML frameworks and hardware co-design.
PyTorch, JAX, TensorFlow
2w
Save
Mark Applied
Hide
Manager III, Applied Science, PXT Central Science
Boston or Bellevue or Seattle or Arlington or San Francisco
$184k-$286k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
3+ Mgmt3+ years managing scientists/ML engineers; deep ML, NLP, IR knowledge; experience with causal inference, building production ML systems, and leading hiring and talent strategy.
machine learning, NLP, Information Retrieval, computer vision, deep learning, large language models, LLMs, Generative AI
1mo
Save
Mark Applied
Hide
Director of Product Management – AI Essentials
Spring or San Jose or Durham or Fort Collins or Andover
$170k-$413k/yr HybridFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Provides global edge-to-cloud technology solutions and IT infrastructure services.
15+ YOE5+ MgmtBachelor's in CS/engineering required; 15+ years product management experience with 5+ years AI/ML product leadership; experience building AI/ML platforms, inference/model serving, GPU ecosystem, executive communication, and GTM strategy.
GreenLake, GenAI, GPU, agent frameworks, AI tools, MLOps, DevOps, SaaS