MakerMaker: Small San Francis-based AI startup focused on autonomous agents and production ML systems.
3+ YOESenior ML systems engineer with 3+ years building production-grade, large-scale serving infrastructure; strong distributed systems; GPU-accelerated inference; fluent Python and systems languages (C++, CUDA, ROCm or Triton).
San Francisco or United States or Toronto or New York City or Montreal
$165k-$330k/yrHybridFull Time
Baseten: Scalable infrastructure platform for deploying and serving AI models.
2+ YOEDegree in CS/Engineering/Math,2+ years experience,production programming (Python preferred),familiarity with ML model lifecycle,strong communication and customer-facing skills.
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
Thinking Machines: Building AI systems to extend human will and judgment.
Bachelor's in CS or equivalent, strong engineering skills, experience with deep learning frameworks and inference serving, ability to optimize distributed GPU systems and contribute production-quality code.
New York City or San Francisco or United States or Europe
RemoteFull Time
Roboflow: Platform for building and deploying custom computer vision models.
5+ YOE5+ years building and operating production ML systems, strong CV/inference foundation, CI/CD and test infra experience, proficiency with PyTorch/TensorFlow/ONNX/TensorRT/vLLM, and experience with image/video processing tools.
Radical Numerics: Building general biological intelligence models for scientific discovery.
Deep expertise in large-model inference, GPU performance engineering, kernel development (CUDA/Triton), Python and PyTorch, distributed systems, and production model deployment.
Together AI: Cloud platform for training and deploying artificial intelligence models.
5+ YOE5+ years experience with inference systems, open-source LLM deployment, and post-training pipelines; expert with inference engines; strong Python skills; Mandarin and English proficiency.
Modal: Serverless cloud platform for running AI and data workloads
Research-leaning or systems background in LLM inference; experience with kernels, quantization, schedulers, and autoscaling; record of shipping research/systems; able to take research bets end-to-end and work onsite in NYC or San Francisco.
Saviynt: Provides AI-powered identity governance and cloud security platforms.
ML platform or MLOps engineer with production Ray experience; LLM serving, distributed training, Python and PyTorch; MLflow/Flyte; Bachelor's degree in CS/Engineering.
Ray Train, Ray Serve, Ray Core, Ray Data, vLLM, SGLang, NVIDIA Triton, TorchTrainer, DDP, NCCL, PPO, RLlib, Flyte, MLflow, Qdrant, Pgvector, PyTorch, Python
Abridge: Automates medical documentation through AI-powered speech analysis
5+ YOE1+ Mgmt5+ years engineering with 1+ year in technical leadership; ML systems and inference experience; GPU, latency, throughput expertise; strong people leadership and collaboration.
Artificial Analysis: Independent AI benchmarking and performance analysis platform.
3+ YOE3+ years professional experience (min 2 years with inference providers/neoclouds), proficiency in Python and data analysis, hands-on familiarity with vLLM,SGLang,TensorRT-LLM and inference APIs, strong analytical skills, and fluency with inference performance metrics.
Santa Monica or Los Angeles or San Francisco or Palo Alto or New York City
$178k-$313k/yrOnsiteFull Time
SnapNYSE: SNAP: Develops social media applications and augmented reality technology.
5+ YOE5+ years post-Bachelor's ML experience with expertise in causal inference, experimentation, uplift modeling, and productionizing models; proficient in Python, pandas, NumPy, scikit-learn; strong communication and mentorship skills.
5+ YOE5+ years post-Bachelor's ML experience (or Master's/PhD with reduced experience), strong causal inference and experimentation experience, proficiency in Python and ML libraries, production ML experience, strong communication and mentorship skills.
Sail Research: Infrastructure platform for long-horizon agentic AI workloads.
Deep knowledge of LLM mechanics and MLSys research, experience with inference engines and GPU profiling, familiarity with tile-based GPU programming (Triton/CUTLASS/ThunderKittens), strong communication and critical thinking.