Reddit
Posted 1w ago

Staff Machine Learning Infrastructure Engineer, Embedding Platform

Reddit
United States
$253k-$355k/yrRemoteFull Time
Responsibilities
  • architecting ML systems
  • defining ML strategy
  • mentoring engineers
Requirements
  • Requires 8+ years in machine learning engineering
  • Expertise in deep learning and scalable ML systems
  • Python or C++
  • Distributed training
  • Real-time inference
  • Cloud ML pipelines
  • A/B testing, and technical leadership
Technical tools mentioned
PythonC++

Job description

Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com.

The LS Embedding Machine Learning Platform team is at the forefront of building highly expressive machine learning models that power Reddit’s recommendation systems. We go beyond standard retrieval and ranking architectures, leveraging modern deep learning approaches and scalable model designs to enhance personalization across Reddit’s ecosystem. Our work impacts content discovery, user engagement, and platform growth at a massive scale.

How You'll Have Impact

As a Staff Machine Learning Infrastructure Engineer, you will own the technical direction for large-scale machine learning platform, guiding the development of advanced deep learning architectures and high-impact ML systems. You will partner with leadership to define ML roadmaps, drive innovation in scalable model design and training approaches, and ensure efficient, reliable deployment of ML models in production. This role offers an opportunity to influence key AI-driven systems across Reddit while mentoring and uplifting the team’s technical capabilities.

What You’ll Do

  • Architect and lead the development of next-generation, large-scale machine learning techniques.
  • Define and execute the ML strategy, identifying opportunities to enhance personalization and recommendation quality across Reddit.
  • Lead research initiatives on scalable machine learning systems and real-time model adaptation, bringing cutting-edge advancements into production.
  • Partner with ML infrastructure teams to build high-performance, distributed training systems that efficiently scale across multiple GPUs and cloud environments.
  • Establish and optimize real-time serving architectures for large-scale embeddings, ensuring low-latency inference and high throughput.
  • Collaborate cross-functionally with teams in Feed Ranking, Ads, Content Understanding, and Core ML to integrate ML models into Reddit’s key AI-driven systems.
  • Mentor and guide senior and mid-level ML engineers, fostering a culture of excellence, innovation, and knowledge sharing.
  • Stay at the forefront of AI research, evaluating and introducing new modeling paradigms to keep Reddit’s ML ecosystem cutting-edge.
  • Drive technical discussions, present findings to leadership, and contribute to long-term ML planning and decision-making.

Who You Might Be:

  • 8+ years of experience in machine learning engineering, with a strong focus on large-scale ML systems and recommendation or personalization systems.
  • Expertise in modern deep learning architectures, including sequence models and foundational models.
  • Deep understanding of complex multi-entity relationships in machine learning applications and how they are modeled in large-scale systems.
  • Proven ability to design, implement, and optimize scalable ML architectures, from distributed training to real-time inference.
  • Strong software engineering skills in Python, C++, or similar languages, with experience in ML infrastructure, high-performance computing, and cloud-based ML pipelines.
  • Demonstrated leadership in driving ML strategy, mentoring engineers, and influencing cross-functional teams.
  • Experience with A/B testing, model evaluation frameworks, and real-time feedback loops in large-scale production systems.
  • Excellent communication skills, with the ability to effectively present complex ML concepts to technical and non-technical stakeholders. 

Benefits:

  • Comprehensive Healthcare Benefits and Income Replacement Programs
  • 401k with Employer Match
  • Global Benefit programs that fit your lifestyle, from workspace to professional development to caregiving support
  • Family Planning Support
  • Gender-Affirming Care
  • Mental Health & Coaching Benefits
  • Flexible Vacation & Paid Volunteer Time Off
  • Generous Paid Parental Leave 

#LI-Remote

Pay Transparency:

This job posting may span more than one career level.

In addition to base salary, this job is eligible to receive equity in the form of restricted stock units, and depending on the position offered, it may also be eligible to receive a commission. Additionally, Reddit offers a wide range of benefits to U.S.-based employees, including medical, dental, and vision insurance, 401(k) program with employer match, generous time off for vacation, and parental leave. To learn more, please visit https://www.redditinc.com/careers/.

To provide greater transparency to candidates, we share base salary ranges for all US-based job postings regardless of state. We set standard base pay ranges for all roles based on function, level, and country location, benchmarked against similar stage growth companies. Final offer amounts are determined by multiple factors including, skills, depth of work experience and relevant licenses/credentials, and may vary from the amounts listed below.

The base salary range for this position is:
$253,300$354,600 USD

In select roles and locations, the interviews will be recorded, transcribed and summarized by artificial intelligence (AI). You will have the opportunity to opt out of recording, transcription and summarization prior to any scheduled interviews.

During the interview, we will collect the following categories of personal information: Identifiers, Professional and Employment-Related Information, Sensory Information (audio/video recording), and any other categories of personal information you choose to share with us. We will use this information to evaluate your application for employment or an independent contractor role, as applicable.  We will not sell your personal information or disclose it to any third party for their marketing purposes.  We will delete any recording of your interview promptly after making a hiring decision.  For more information about how we will handle your personal information, including our retention of it, please refer to our Candidate Privacy Policy for Potential Employees and Contractors.

Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve.  Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.

About Reddit

Social platform for community-driven discussion and content sharing.

Similar jobs

Machine Learning Infrastructure Engineer roles
1w
Save
Mark Applied
Hide
Machine Learning Infrastructure Engineer, Technology
New York City, New York, United States
$185k-$300k/yr OnsiteFull Time
Point72
Point72: Global alternative investment firm managing capital and venture investments.
3+ YOEBachelor’s or master’s degree in a technical field; 3–7 years building scalable compute or ML infrastructure; expertise in distributed systems, Kubernetes, cloud platforms, MLOps, Python, and systems programming.
Kubernetes, AWS, Google Cloud Platform, Azure, MLflow, Ray, Airflow, Kubeflow, Terraform, Python, Go, C++, Rust
1w
Save
Mark Applied
Hide
Tech Lead, Machine Learning Infrastructure Engineer
Seattle or Bellevue
$198k-$416k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
5+ YOEBachelor’s or master’s degree in a technical discipline; 5+ years software engineering and 3+ years ML infrastructure experience. Requires Python, C++ or Java, distributed systems, GPU clusters, Kubernetes or Slurm, PyTorch or TensorFlow, CUDA, and networking expertise.
Python, C++, Java, Kubernetes, Slurm, PyTorch, TensorFlow, CUDA, InfiniBand, RoCE, Megatron, DeepSpeed, Triton, Cutlass, TensorRT, Triton Inference Server
2w
Save
Mark Applied
Hide
Staff ML Infra Engineer
Mountain View, California, United States
$174k-$299k/yr OnsiteFull Time
Coupang
CoupangNYSE: CPNG: Provides an end-to-end e-commerce and logistics network.
5+ YOEBachelor's in CS/EE/math,5+ years applied ML experience,proficiency in Python/Java,experience with big data pipelines,ML frameworks,cloud platforms,and building scalable low-latency services.
Python, Java, Hadoop, Hive, Presto, Spark, Scala, Apache Airflow, TensorFlow, PyTorch, AWS, GCP
3w
Save
Mark Applied
Hide
Staff Engineer - ML Infra / MLOps
Palo Alto, California, United States
$218k-$285k/yr OnsiteFull Time
Quince
Quince: Sells high-quality apparel and home goods at accessible prices.
8+ YOE8+ years industry experience with 4+ years in ML infrastructure/MLOps. Experience designing production ML platforms, cloud-native infra (AWS), Kubernetes, IaC, distributed training, feature stores, and cost/compute optimization.
AWS, Kubernetes (EKS), Docker, Terraform, Pulumi, PyTorch, TensorFlow, Kubeflow, SageMaker, Spark, Flink, Kafka
4w
Save
Mark Applied
Hide
Machine Learning Infrastructure Engineer, Safeguards Research
San Francisco or New York City
$350k-$500k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Proven software engineering with Python, experience building and operating data-intensive or distributed systems, tooling for researcher workflows, debugging performance and correctness, and strong communication skills.
Python, transformers, GPU
1mo
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Francisco, California, United States
$180k-$230k/yr OnsiteFull Time
Echo Neurotechnologies
Echo Neurotechnologies: Developing brain-computer interface technologies to improve patient autonomy.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Python, PyTorch, FSDP, DeepSpeed, Megatron-LM, Ray, C++, Go, CUDA, Rust, Java, Kubernetes, Docker
1mo
Save
Mark Applied
Hide
Machine Learning Infrastructure Engineer
Redwood City, California, United States
$140k-$240k/yr HybridFull Time
WindBorne Systems
WindBorne Systems: Operates smart weather balloons to provide global atmospheric data.
Experience running production ML systems, working with large datasets, PyTorch/Docker experience, managing GPU/cluster/job-scheduler environments, and building reliable data and training pipelines.
PyTorch, Docker
1mo
Save
Mark Applied
Hide
Machine Learning Infrastructure Tech Lead
San Francisco, California, United States
$200k-$300k/yr OnsiteFull Time
Reducto
Reducto: AI platform extracting structured data from unstructured documents
5+ YOE5+ years building production infrastructure with significant ML systems experience; strong Python, Kubernetes, GPU training/serving, distributed systems, and performance optimization skills.
Python, Kubernetes, CUDA, Triton, vLLM, SGLang, PyTorch, TensorRT-LLM, Ray