Reddit
Posted 5d ago

Senior Machine Learning Infrastructure Engineer, Embedding Platform

Reddit
United States
$191k-$267k/yrRemoteFull Time
Responsibilities
  • designing ML platforms
  • building ML pipelines
  • optimizing inference
Requirements
  • Requires 5+ years in machine learning engineering
  • Deep learning and distributed ML expertise
  • Python and modern ML frameworks
  • Production-scale infrastructure
  • System design
  • Testing
  • Optimization
  • Evaluation, and strong communication
Technical tools mentioned
PythonPyTorchTensorFlow

Job description

Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com.

The LS Embedding Machine Learning Platform team is at the forefront of building highly expressive, machine learning models that power Reddit’s recommendation systems. We go beyond standard retrieval and ranking architectures, leveraging modern deep learning approaches and scalable model designs to enhance personalization across Reddit’s ecosystem. Our work impacts content discovery, user engagement, and platform growth at a massive scale.

About the Role

As a Senior Machine Learning Infrastructure Engineer, you will work across both model development and ML platform to build large-scale learning systems that improve recommendation and personalization on Reddit. At the senior level, you will own major technical components end to end: designing models, implementing training and evaluation pipelines, and driving production deployment in close partnership with ML platform, product, and cross-functional ML teams.

Responsibilities

  • Design, train, and improve large-scale machine learning platforms for recommendation or personalization systems.
  • Own and deliver major ML systems components end to end, from problem framing through production rollout.
  • Build and optimize end-to-end ML pipelines spanning data preparation, feature generation, training, evaluation, and deployment.
  • Improve distributed training, model efficiency, and online inference performance.
  • Apply modern modeling approaches including sequence modeling and related foundation-model techniques to Reddit use cases.
  • Develop reliable serving and monitoring patterns for low-latency, high-throughput production ML systems.
  • Work with cross-functional partners across product, relevance, ads, and core ML teams to deliver measurable improvements in user experience and business impact.
  • Drive rigorous offline and online evaluation, including experimentation, model diagnostics, and feedback-loop improvement.
  • Contribute to engineering quality through strong code, design reviews, documentation, and operational excellence.

Qualifications

  • 5+ years of experience in machine learning engineering, with a strong focus on large-scale ML infrastructure and recommendation or personalization systems.
  • Expertise in modern deep learning architectures, including sequence models and foundational models.
  • Experience building or scaling ML platform for large datasets and high-traffic production environments.
  • Demonstrated ability to independently scope and execute ambiguous technical work, while owning high-quality implementation details.
  • Solid understanding of distributed training and inference concepts, such as data parallelism, model parallelism, pipeline parallelism, or related optimization techniques.
  • Proficiency in Python and experience with modern ML frameworks such as PyTorch, TensorFlow, or similar.
  • Strong software engineering fundamentals, including system design, debugging, testing, and performance optimization.
  • Experience with A/B testing, model evaluation frameworks, and real-time feedback loops in large-scale production systems.
  • Excellent communication skills, with the ability to effectively present complex ML concepts to technical and non-technical stakeholders.

Benefits:

  • Comprehensive Healthcare Benefits and Income Replacement Programs
  • 401k with Employer Match
  • Global Benefit programs that fit your lifestyle, from workspace to professional development to caregiving support
  • Family Planning Support
  • Gender-Affirming Care
  • Mental Health & Coaching Benefits
  • Flexible Vacation & Paid Volunteer Time Off
  • Generous Paid Parental Leave 

#LI-Remote

Pay Transparency:

This job posting may span more than one career level.

In addition to base salary, this job is eligible to receive equity in the form of restricted stock units, and depending on the position offered, it may also be eligible to receive a commission. Additionally, Reddit offers a wide range of benefits to U.S.-based employees, including medical, dental, and vision insurance, 401(k) program with employer match, generous time off for vacation, and parental leave. To learn more, please visit https://www.redditinc.com/careers/.

To provide greater transparency to candidates, we share base salary ranges for all US-based job postings regardless of state. We set standard base pay ranges for all roles based on function, level, and country location, benchmarked against similar stage growth companies. Final offer amounts are determined by multiple factors including, skills, depth of work experience and relevant licenses/credentials, and may vary from the amounts listed below.

The base salary range for this position is:
$190,800$267,100 USD

In select roles and locations, the interviews will be recorded, transcribed and summarized by artificial intelligence (AI). You will have the opportunity to opt out of recording, transcription and summarization prior to any scheduled interviews.

During the interview, we will collect the following categories of personal information: Identifiers, Professional and Employment-Related Information, Sensory Information (audio/video recording), and any other categories of personal information you choose to share with us. We will use this information to evaluate your application for employment or an independent contractor role, as applicable.  We will not sell your personal information or disclose it to any third party for their marketing purposes.  We will delete any recording of your interview promptly after making a hiring decision.  For more information about how we will handle your personal information, including our retention of it, please refer to our Candidate Privacy Policy for Potential Employees and Contractors.

Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve.  Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.

About Reddit

Social platform for community-driven discussion and content sharing.

Similar jobs

Machine Learning Infrastructure Engineer roles
6d
Save
Mark Applied
Hide
Machine Learning Infrastructure Engineer, Technology
New York City, New York, United States
$185k-$300k/yr OnsiteFull Time
Point72
Point72: Global alternative investment firm managing capital and venture investments.
3+ YOEBachelor’s or master’s degree in a technical field; 3–7 years building scalable compute or ML infrastructure; expertise in distributed systems, Kubernetes, cloud platforms, MLOps, Python, and systems programming.
Kubernetes, AWS, Google Cloud Platform, Azure, MLflow, Ray, Airflow, Kubeflow, Terraform, Python, Go, C++, Rust
6d
Save
Mark Applied
Hide
Tech Lead, Machine Learning Infrastructure Engineer
Seattle or Bellevue
$198k-$416k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
5+ YOEBachelor’s or master’s degree in a technical discipline; 5+ years software engineering and 3+ years ML infrastructure experience. Requires Python, C++ or Java, distributed systems, GPU clusters, Kubernetes or Slurm, PyTorch or TensorFlow, CUDA, and networking expertise.
Python, C++, Java, Kubernetes, Slurm, PyTorch, TensorFlow, CUDA, InfiniBand, RoCE, Megatron, DeepSpeed, Triton, Cutlass, TensorRT, Triton Inference Server
1w
Save
Mark Applied
Hide
Staff ML Infra Engineer
Mountain View, California, United States
$174k-$299k/yr OnsiteFull Time
Coupang
CoupangNYSE: CPNG: Provides an end-to-end e-commerce and logistics network.
5+ YOEBachelor's in CS/EE/math,5+ years applied ML experience,proficiency in Python/Java,experience with big data pipelines,ML frameworks,cloud platforms,and building scalable low-latency services.
Python, Java, Hadoop, Hive, Presto, Spark, Scala, Apache Airflow, TensorFlow, PyTorch, AWS, GCP
3w
Save
Mark Applied
Hide
Staff Engineer - ML Infra / MLOps
Palo Alto, California, United States
$218k-$285k/yr OnsiteFull Time
Quince
Quince: Sells high-quality apparel and home goods at accessible prices.
8+ YOE8+ years industry experience with 4+ years in ML infrastructure/MLOps. Experience designing production ML platforms, cloud-native infra (AWS), Kubernetes, IaC, distributed training, feature stores, and cost/compute optimization.
AWS, Kubernetes (EKS), Docker, Terraform, Pulumi, PyTorch, TensorFlow, Kubeflow, SageMaker, Spark, Flink, Kafka
3w
Save
Mark Applied
Hide
Machine Learning Infrastructure Engineer, Safeguards Research
San Francisco or New York City
$350k-$500k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Proven software engineering with Python, experience building and operating data-intensive or distributed systems, tooling for researcher workflows, debugging performance and correctness, and strong communication skills.
Python, transformers, GPU
4w
Save
Mark Applied
Hide
Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)
Toronto or Zürich or California
HybridFull Time
Veeda AI
Veeda AI: Building multimodal foundation world models for physical AI.
Bachelor's degree or equivalent experience; deep experience with distributed deep-learning training, numerical precision (BF16/FP8), hardware/software root-cause analysis, and strong Python and C++/CUDA skills.
PyTorch, FSDP, Megatron-LM, DeepSpeed, Tensor Parallelism, Pipeline Parallelism, BF16, FP8, Python, C++/CUDA, NCCL, ROCm, JAX/XLA, CUDA, Triton
1mo
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Francisco, California, United States
$180k-$230k/yr OnsiteFull Time
Echo Neurotechnologies
Echo Neurotechnologies: Developing brain-computer interface technologies to improve patient autonomy.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Python, PyTorch, FSDP, DeepSpeed, Megatron-LM, Ray, C++, Go, CUDA, Rust, Java, Kubernetes, Docker
1mo
Save
Mark Applied
Hide
Machine Learning Infrastructure Engineer
Redwood City, California, United States
$140k-$240k/yr HybridFull Time
WindBorne Systems
WindBorne Systems: Operates smart weather balloons to provide global atmospheric data.
Experience running production ML systems, working with large datasets, PyTorch/Docker experience, managing GPU/cluster/job-scheduler environments, and building reliable data and training pipelines.
PyTorch, Docker