24 inference engineer jobs at 15 companies in Warwick, RI
1w
Save
Mark Applied
Hide
1w
Staff Engineer, Inference Optimizations
Boston, Massachusetts, United States
$191k-$239k/yrRemoteFull Time
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and inference optimization expertise, experience with CUDA/Triton/ROCm, attention-layer and kernel-level optimization, strong system design and leadership through influence.
Senior Lead AI Engineer (FM Hosting, LLM Inference)
New York or McLean or Cambridge or San Jose
$251k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
6+ YOEBachelor's in CS/AI/EE/CE or related with 6+ years (or Master's with 4+ years); 6+ years programming with Python, Go, Scala, or Java; experience deploying scalable AI systems, LLM inference, similarity search, and optimization of training/inference.
SimpliSafe: Provides wireless home security systems and professional monitoring services.
8+ YOE8+ years embedded/performance engineering experience, strong C/C++ skills, experience optimizing on-device ML inference and runtimes (TFLite/ONNX/TensorRT), profiling and debugging performance under constrained hardware.
Applied AI Research Engineer - Model Cost Optimization
Cambridge, Massachusetts, United States
$210k-$260k/yrOnsiteFull Time
Blitzy: Autonomous AI platform for enterprise software development.
Applied ML/AI research experience with hands-on LLM evaluation or fine-tuning, strong software engineering, familiarity with LLM evaluation and benchmarking, and understanding of inference cost/latency tradeoffs.
Analog DevicesNASDAQ: ADI: Designs and manufactures semiconductors for signal processing and power management.
10+ YOEPhD in Electrical Engineering or related field; 10+ years in audio/signal processing; deployment on DSP/NPU; fixed-point, PTQ/QAT; C/Python/MATLAB; RTL verification; embedded AI inference.
Systems ML Engineer (Member of the Technical Staff)
Cambridge, Massachusetts, United States
OnsiteFull Time
Transfyr: Building physical AI infrastructure for scientific research and automation.
Experience optimizing and deploying large-scale ML models for training and inference, profiling and custom GPU kernel development, distributed training, cloud and edge deployment, and infrastructure automation.
Physical Superintelligence: Building AI systems to discover new physics at scale.
3+ YOE3+ years building and operating ML training or inference infrastructure; distributed training and model-serving experience; strong software engineering and ML fluency to debug training and inference; cloud and IaC experience preferred.
Glia AI: Automated AI systems engineering for high-performance infrastructure.
Experience with AI inference and distributed serving, strong systems and GPU optimization skills, proficiency in Python and PyTorch, familiarity with inference frameworks (vLLM, Triton, Ray Serve); PhD or equivalent experience preferred.
vLLM, PyTorch, Triton inference server, Ray Serve, CUDA, Python
WHOOP: Wearable technology for personalized health and performance tracking.
4+ YOE4+ years ML engineer/related experience; BS in CS/Math/Data Science; strong Python; ML inference at scale; cloud/CI/CD; time-series experience preferred.
Python, cloud platforms (AWS or GCP), MLOps, CI/CD, time series data
Assail: Autonomous AI platform for offensive security testing.
5+ YOE5+ years building production ML/AI systems with 2+ years on LLMs/agents; deep Python; fine-tuning (SFT, DPO/GRPO, RLHF/RLAIF); transformer expertise; PyTorch/Hugging Face/DeepSpeed stack; inference optimization; retrieval/vector pipelines; Kubernetes.
CVS HealthNYSE: CVS: Provides retail pharmacy, health insurance, and pharmacy benefit management services.
5+ YOE5+ years building ML systems in production, 5+ years Python, familiarity with PyTorch/TensorFlow, MLOps (Kubeflow, Vertex), 2+ years SQL with Snowflake/BigQuery, 2+ years system design leadership, experience in reinforcement learning/causal inference/LLMs, mentoring experience, strong communication.
AI Systems Architect (Models & Hardware Co-Design)
Santa Clara or Austin or Boston
$200k-$500k/yrOnsiteFull Time
Velaura AI: Developing ultra-low-power silicon and IP for AI accelerators.
Deep knowledge of modern ML architectures (e.g., transformers), strong mathematical foundations, experience with ML training/inference, and familiarity with ML frameworks and hardware co-design.
Boston or Bellevue or Seattle or Arlington or San Francisco
$184k-$286k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
3+ Mgmt3+ years managing scientists/ML engineers; deep ML, NLP, IR knowledge; experience with causal inference, building production ML systems, and leading hiring and talent strategy.
machine learning, NLP, Information Retrieval, computer vision, deep learning, large language models, LLMs, Generative AI