Featherless AI
Posted 6mo ago

Machine Learning Engineer — Inference Optimization

Featherless AI
North America
RemoteFull Time
Responsibilities
  • Optimize latency
  • Profile pipelines
  • Benchmark performance
Requirements
  • Experience in ML inference optimization
  • Deep learning internals
  • PyTorch
  • GPU performance tuning, and scaling inference in production
Technical tools mentioned
PyTorchCUDAROCmTritonTensorRTONNX RuntimevLLMKernels

Job description

About the Role

We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale. You’ll work at the intersection of research and production—turning cutting-edge models into fast, reliable, and cost-efficient systems that serve real users.

This role is ideal for someone who enjoys deep technical work, profiling systems down to the kernel/GPU level, and translating research ideas into production-grade performance gains.

What You’ll Do

  • Optimize inference latency, throughput, and cost for large-scale ML models in production

  • Profile and bottleneck GPU/CPU inference pipelines (memory, kernels, batching, IO)

  • Implement and tune techniques such as:

    • Quantization (fp16, bf16, int8, fp8)

    • KV-cache optimization & reuse

    • Speculative decoding, batching, and streaming

    • Model pruning or architectural simplifications for inference

  • Collaborate with research engineers to productionize new model architectures

  • Build and maintain inference-serving systems (e.g. Triton, custom runtimes, or bespoke stacks)

  • Benchmark performance across hardware (NVIDIA / AMD GPUs, CPUs) and cloud setups

  • Improve system reliability, observability, and cost efficiency under real workloads

What We’re Looking For

  • Strong experience in ML inference optimization or high-performance ML systems

  • Solid understanding of deep learning internals (attention, memory layout, compute graphs)

  • Hands-on experience with PyTorch (or similar) and model deployment

  • Familiarity with GPU performance tuning (CUDA, ROCm, Triton, or kernel-level optimizations)

  • Experience scaling inference for real users (not just research benchmarks)

  • Comfortable working in fast-moving startup environments with ownership and ambiguity

Nice to Have

  • Experience with LLM or long-context model inference

  • Knowledge of inference frameworks (TensorRT, ONNX Runtime, vLLM, Triton)

  • Experience optimizing across different hardware vendors

  • Open-source contributions in ML systems or inference tooling

  • Background in distributed systems or low-latency services

Why Join Us

  • Real ownership over performance-critical systems

  • Direct impact on product reliability and unit economics

  • Close collaboration with research, infra, and product

  • Competitive compensation + meaningful equity at Series A

  • A team that cares about engineering quality, not hype

About Featherless AI

Provides serverless infrastructure for deploying open-source AI models.

Year founded
2023
Employees
15
Organization type
Private
Latest investment
Raised $5.00M Seed (2025) — led by Airbus Ventures
Headquarters
US

Similar jobs

Machine Learning Engineer roles
3h
Save
Mark Applied
Hide
Machine Learning Engineer Graduate (E-Commerce Governance) - 2027 Start
San Jose, California, United States
$128k-$317k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's degree in software development, computer science, computer engineering, or related field; strong Python, Go, or C++ skills; Linux development; data structures, algorithms, machine learning, and deep learning knowledge.
Linux, Python, Go, C++, natural language processing, computer vision, multimodal, graph algorithms, search algorithms, text mining, data mining, LLM, reinforcement learning, operational research
5h
Save
Mark Applied
Hide
Consultant | Machine Learning Engineer | Public Sector
Amsterdam, North Holland, Netherlands
HybridFull Time
Amsterdam Data Collective
Amsterdam Data Collective: Data and AI consultancy helping organizations turn data into impact.
3+ YOERequires 3+ years in machine learning or related engineering, Python, production ML lifecycle experience, cloud and MLOps expertise, regulated-environment experience, and fluency in English and Dutch.
Python, Azure, AWS, Databricks, MLflow, Docker, Airflow, CI/CD, Azure Machine Learning, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG)
5h
Save
Mark Applied
Hide
Staff Machine Learning Engineer for AI Product
Paris, Île-de-France, France
OnsiteFull Time
Qonto
Qonto: European business finance workspace for SMEs and freelancers.
6+ YOE6+ years of ML engineering experience with MLOps, end-to-end customer-facing model deployment, Python, FastAPI or similar, production integrations, and fluent English.
Python, FastAPI, Snowflake, Kafka, Kibana, PostgreSQL, Airflow, AWS, Prometheus, ArgoCD, GitHub, Cursor
7h
Save
Mark Applied
Hide
Senior Machine Learning Engineer, MLE
Beijing, Beijing, China
OnsiteFull Time
Grab
GrabNASDAQ: GRAB: Superapp providing transportation, food delivery, and digital financial services.
2+ YOERequires 2+ years in machine learning engineering, strong algorithms and data structures, large-scale microservices, production batch and real-time pipelines, evaluation metrics, and willingness to work in Golang.
Golang, Kafka, Flink, Spark Streaming, RAG, AWS, Microsoft Azure
13h
Save
Mark Applied
Hide
2027 - Full Time Assistant Vice President - Corporate Functions, Analytics Lab (Machine Learning Engineer)
Jersey City or Montreal or Paris or Lisbon or London or New York or Chesterbrook or San Francisco or Boston or Chicago or Denver or Miami or Washington
$150k/yr OnsiteFull Time
BNP Paribas
BNP ParibasEuronext Paris: BNP: Provides global banking, financial, and investment services.
Master’s degree in a quantitative field; December 2026–June 2027 graduation; Python or C, SQL, statistical modeling, data analysis, LLM/agentic workflows, ML frameworks, RAG systems, Git, and communication skills.
Python, C, SQL, LangChain, CrewAI, Google ADK, OpenAI Agent SDK, TensorFlow, PyTorch, scikit-learn, Git, DVC, MLflow, Jenkins, Docker
16h
Save
Mark Applied
Hide
Lead Machine Learning Engineer
Ann Arbor, Michigan, United States
OnsiteFull Time
Domino's
Domino'sNASDAQ: DPZ: Global pizza chain providing delivery and carryout services.
5+ YOERequires 5–8 years of machine learning engineering experience, a bachelor's degree in a technical discipline, advanced ML skills, Python, ML libraries, scalable systems, and production deployment experience.
Microsoft Azure, Azure Container Apps, Azure AI Services, Azure Storage, Azure Key Vault, Azure Monitor, Event Hubs, Azure Machine Learning, Databricks, GitHub, GitHub Enterprise, GitHub Actions, Python, TensorFlow, PyTorch, scikit-learn, Pandas, NumPy, Docker, Kubernetes, REST APIs
19h
Save
Mark Applied
Hide
Senior Geospatial Machine Learning Engineer
Canada or United States or Europe
RemoteFull Time
Clera
Clera: AI talent agent matching professionals with high-growth startup roles
5+ YOE5+ years building and deploying production machine learning or deep learning models; satellite or aerial imagery experience; Python geospatial libraries; ML frameworks; pipeline orchestration; QGIS; and production model monitoring.
Python, rasterio, geopandas, shapely, GDAL, TensorFlow, PyTorch, scikit-learn, Dagster, Airflow, dbt, QGIS, Grafana, Sentry, Prometheus
21h
Save
Mark Applied
Hide
Senior Principal Machine Learning Engineer
San Jose, California, United States
$220k-$300k/yr OnsiteFull Time
SambaNova Systems
SambaNova Systems: Develops custom AI hardware and software for enterprise computing.
8+ YOEBachelor's degree in computer science, electrical engineering, or related field; 8+ years of ML engineering experience; deep LLM expertise; technical leadership and complex project delivery experience.
Large Language Models (LLMs), RDU, SambaStack, SambaCloud, speculative decoding, reinforcement learning, mixture-of-experts