Whatnot
Posted 1mo ago

Machine Learning Infrastructure Engineer

Whatnot
San Francisco or New York or Los Angeles or Seattle
$200k-$345k/yrHybridFull Time
Responsibilities
  • designing infrastructure
  • serving models
  • scaling systems
Requirements
  • 4+ years building ML systems,3+ years software engineering,1+ year Python,experience with databases,monitoring,cloud services and production ML deployments
Technical tools mentioned
PythonPostgreSQLDynamoDBElasticsearchRedisDataDogGrafanaAWS SagemakerLambdaKinesisS3EC2EKSECSApache KafkaFlink

Job description

🚀 Join the Future of Commerce with Whatnot!

Whatnot is the largest live shopping platform in North America and Europe to buy, sell, and discover the things you love. Whether it's trading cards, fashion, electronics, or live plants, our sellers are building real businesses across hundreds of categories. We're building live commerce at a scale that's never been done in the West, and there's no playbook to copy. The people here are shaping how an entirely new industry develops.

As a remote co-located team, we're inspired by our values and anchored in hubs across the US, UK, Ireland, Poland, Germany, and Australia. We move fast, stay close to our users, and focus on the work that drives the most impact.

We're one of the fastest growing marketplaces and were recently named the #1 Best Startup Employer in America by Forbes. Check out the latest Whatnot updates on our news and engineering blogs and join us as we enable anyone to turn their passion into a business and bring people together through commerce.

💻 Role

We’re looking for builders–intellectually curious, highly entrepreneurial engineers eager to shape the future of AI and ML at Whatnot. You’ll design and scale the core infrastructure that powers machine learning and self-hosted large language model applications across the company, working side by side with machine learning scientists to bring cutting-edge models into production and unlock entirely new product experiences. This means building systems that make advanced ML dependable and fast at scale–from low-latency, large model serving to distributed training & high-throughput GPU inference.

What you'll do:

  • Own the infrastructure powering AI and ML models across critical business surfaces–supporting growth, recommendations, trust and safety, fraud, seller tooling, and more.

  • Prototype, deploy, and productionalize novel ML architectures that directly shape user experience and marketplace dynamics.

  • Design and scale inference infrastructure capable of serving large models with low latency and high throughput.

  • Build distributed training and inference pipelines leveraging GPUs and both model and data parallelism.

  • Stretch beyond your comfort zone to take on new technical challenges as we scale AI across Whatnot’s ecosystem.

US Based: We offer flexibility to work from home or from one of our global office hubs, and we value in-person time for planning, problem-solving, and connection. Team members in this role must live within commuting distance of our New York, Seattle, Los Angeles, and San Francisco hubs.

👋 You

People who do well at Whatnot tend to be comfortable figuring things out as they go, biased toward action, and genuinely curious about what they're building. They care more about outcomes than credit and stay close to the product and the people using it.

As our next AI/ML Platform Engineer you should have 4+ years of professional experience developing machine learning systems and algorithms, plus:

  • Bachelor’s degree in Computer Science, Statistics, Applied Mathematics or a related technical field, or equivalent work experience.

  • 3+ years of software engineering experience building and maintaining production systems for consumer-scale loads.

  • 1+ years of professional experience developing software in Python

  • Ability to work autonomously and drive initiatives across multiple product areas and communicate findings with leadership and product teams.

  • Experience with operational, search, and key-value databases such as PostgreSQL, DynamoDB, Elasticsearch, Redis.

  • Firm grasp of visualization tools for monitoring and logging e.g. DataDog, Grafana.

  • Familiarity with cloud computing platforms and managed services such as AWS Sagemaker, Lambda, Kinesis, S3, EC2, EKS/ECS, Apache Kafka, Flink.

  • Professionalism around collaborating in a remote working environment and well tested, reproducible work.

  • Exceptional documentation and communication skills.

🎁 Benefits

  • Flexible Time off Policy and Company-wide Holidays (including a spring and winter break)

  • Health Insurance options including Medical, Dental, Vision

  • Work From Home Support

    • Home office setup allowance

    • Monthly allowance for cell phone and internet

  • Care benefits

    • Monthly allowance for wellness

    • Annual allowance towards Childcare

    • Lifetime benefit for family planning, such as adoption or fertility expenses

  • Retirement; 401k offering for Traditional and Roth accounts in the US (employer match up to 4% of base salary) and Pension plans internationally

  • Monthly allowance to dogfood the app

    • All Whatnauts are expected to develop a deep understanding of our product. We're passionate about building the best user experience, and all employees are expected to use Whatnot as both a buyer and a seller as part of their job (our dogfooding budget makes this fun and easy!).

  • Parental Leave

    • 16 weeks of paid parental leave + one month gradual return to work *company leave allowances run concurrently with country leave requirements which take precedence.

💛 EOE

Whatnot is proud to be an Equal Opportunity Employer. We value diversity, and we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, parental status, disability status, or any other status protected by local law. We believe that our work is better and our company culture is improved when we encourage, support, and respect the different skills and experiences represented within our workforce.

About Whatnot

Social marketplace for buying and selling via live streams

Year founded
2019
Employees
1600
Organization type
Private
Latest investment
Raised $225.00M Series F (2025) — led by DST Global, CapitalG
Subsidiaries
Headquarters
US

Similar jobs

Machine Learning Infrastructure Engineer roles near San Francisco, California
1d
Save
Mark Applied
Hide
Senior Machine Learning Infrastructure Engineer, Recommendations and Search
San Jose, California, United States
$187k-$360k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
4+ YOEBachelor’s or master’s degree in a technical discipline; 4+ years software engineering, 2+ years enterprise ML infrastructure, and expertise in Python, C++/Java, PyTorch/TensorFlow, GPUs, CUDA, and distributed systems.
Python, C++, Java, PyTorch, TensorFlow, CUDA, InfiniBand, RoCE, Megatron, DeepSpeed, Triton, Cutlass, TensorRT, Triton Inference Server
2w
Save
Mark Applied
Hide
Staff ML Infra Engineer
Mountain View, California, United States
$174k-$299k/yr OnsiteFull Time
Coupang
CoupangNYSE: CPNG: Provides an end-to-end e-commerce and logistics network.
5+ YOEBachelor's in CS/EE/math,5+ years applied ML experience,proficiency in Python/Java,experience with big data pipelines,ML frameworks,cloud platforms,and building scalable low-latency services.
Python, Java, Hadoop, Hive, Presto, Spark, Scala, Apache Airflow, TensorFlow, PyTorch, AWS, GCP
4w
Save
Mark Applied
Hide
Staff Engineer - ML Infra / MLOps
Palo Alto, California, United States
$218k-$285k/yr OnsiteFull Time
Quince
Quince: Sells high-quality apparel and home goods at accessible prices.
8+ YOE8+ years industry experience with 4+ years in ML infrastructure/MLOps. Experience designing production ML platforms, cloud-native infra (AWS), Kubernetes, IaC, distributed training, feature stores, and cost/compute optimization.
AWS, Kubernetes (EKS), Docker, Terraform, Pulumi, PyTorch, TensorFlow, Kubeflow, SageMaker, Spark, Flink, Kafka
1mo
Save
Mark Applied
Hide
Machine Learning Infrastructure Engineer, Safeguards Research
San Francisco or New York City
$350k-$500k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Proven software engineering with Python, experience building and operating data-intensive or distributed systems, tooling for researcher workflows, debugging performance and correctness, and strong communication skills.
Python, transformers, GPU
1mo
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Francisco, California, United States
$180k-$230k/yr OnsiteFull Time
Echo Neurotechnologies
Echo Neurotechnologies: Developing brain-computer interface technologies to improve patient autonomy.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Python, PyTorch, FSDP, DeepSpeed, Megatron-LM, Ray, C++, Go, CUDA, Rust, Java, Kubernetes, Docker
1mo
Save
Mark Applied
Hide
Machine Learning Infrastructure Engineer
Redwood City, California, United States
$140k-$240k/yr HybridFull Time
WindBorne Systems
WindBorne Systems: Operates smart weather balloons to provide global atmospheric data.
Experience running production ML systems, working with large datasets, PyTorch/Docker experience, managing GPU/cluster/job-scheduler environments, and building reliable data and training pipelines.
PyTorch, Docker
1mo
Save
Mark Applied
Hide
Machine Learning Infrastructure Tech Lead
San Francisco, California, United States
$200k-$300k/yr OnsiteFull Time
Reducto
Reducto: AI platform extracting structured data from unstructured documents
5+ YOE5+ years building production infrastructure with significant ML systems experience; strong Python, Kubernetes, GPU training/serving, distributed systems, and performance optimization skills.
Python, Kubernetes, CUDA, Triton, vLLM, SGLang, PyTorch, TensorRT-LLM, Ray
1mo
Save
Mark Applied
Hide
Member of Technical Staff - Machine Learning Infrastructure Engineer
San Francisco or Toronto or Seattle
$180k-$300k/yr OnsiteFull Time
Preference Model
Preference Model: Building reinforcement learning environments to train frontier AI models.
Experienced software engineer with production ML/data infrastructure skills, proficiency with PyTorch or JAX, distributed systems, AWS/GCP, Kubernetes, data pipelines, and familiarity with transformers and inference libraries like vLLM.
PyTorch, JAX, AWS, GCP, Kubernetes, transformers, vLLM, SGLang