Pluralis Research
Posted 4mo ago

Machine Learning Engineer - Distributed ML Systems

Pluralis Research
San Francisco or Melbourne
RemoteFull Time
Responsibilities
  • design training
  • optimize parallelism
  • build monitoring
Requirements
  • 5+ years in distributed systems and large-scale ML training
  • Senior/staff engineer level
Technical tools mentioned
PythonDeepSpeedMegatronFSDPgRPC

Job description

Overview

Pluralis Research carries out foundational research on Protocol Learning: multi-participant training of foundation models where no single participant has, or can ever obtain, a full copy of the model. The purpose of Protocol Learning is to facilitate the creation of community-trained and community-owned frontier models with self-sustaining economics.

We're looking for Senior/Staff engineers with 5+ years of experience in distributed systems and ML large-scale training. You'll be implementing a novel substrate for training distributed ML models that work under consumer grade internet connection.

Responsibilities

Distributed Training Architecture & Optimization

  • Design and implement large-scale distributed training systems optimized for heterogeneous hardware operating under low-bandwidth, high-latency conditions.

  • Develop and optimize model-parallel training strategies (data, tensor, pipeline parallelism) with custom sharding techniques that minimize communication overhead.

  • Optimize GPU utilization, memory efficiency, and compute performance across distributed nodes.

  • Implement robust checkpointing, state synchronization, and recovery mechanisms for long-running, fault-prone training jobs.

  • Build monitoring and metrics systems to track training progress, model quality, and system bottlenecks.

Decentralized Networking & Resilience

  • Architect resilient training systems where nodes can fail, networks can partition, and participants can dynamically join or leave.

  • Design and optimize peer-to-peer topologies for decentralized coordination across non-co-located nodes.

  • Implement NAT traversal, peer discovery, dynamic routing, and connection lifecycle management.

  • Profile and optimize communication patterns to reduce latency and bandwidth overhead in multi-participant environments.

What You’ll Bring

  • Strong experience building and operating distributed systems in production.

  • Hands-on expertise with distributed training frameworks (FSDP, DeepSpeed, Megatron, or similar).

  • Deep understanding of model parallelism (data, tensor, pipeline parallelism).

  • Expert-level Python with production experience (concurrency, error handling, retry logic, clean architecture).

  • Strong networking fundamentals: P2P systems, gRPC, routing, NAT traversal, distributed coordination.

  • Experience optimizing GPU workloads, memory management, and large-scale compute efficiency.

What We Offer

  • Equity-heavy compensation with meaningful ownership in a mission-driven company

  • Competitive base salary for senior engineering roles in Australia

  • Visa sponsorship available for exceptional candidates

  • Remote-first with optional access to our Melbourne hub

  • World-class team — team mates were previously at at Google, Amazon, Microsoft, and leading startups

Backed by Union Square Ventures and other tier-1 investors, we're a world-class, deeply technical team of ML researchers and engineers. Pluralis is unapologetically ideological. We view the world as a better place if we are able to implement what we are attempting, and Protocol Learning as the only plausible approach to preventing a handful of massive corporations monopolising model development, access and release, and achieving massive economic capture. If this resonates, please apply.

About Pluralis Research

Decentralized protocol for collaborative and open-source AI model training.

Year founded
2024
Employees
10
Organization type
Private
Latest investment
Raised $7.60M Seed (2025) — led by Union Square Ventures, CoinFund
Headquarters
AU

Similar jobs

Machine Learning Engineer roles near San Francisco, California
3h
Save
Mark Applied
Hide
Machine Learning Engineer Graduate (E-Commerce Supply Chain & Logistics - LLM/Agent) - 2027 Start (PhD)
San Jose or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo or Los Angeles County
$162k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
PhD candidate or recent PhD graduate in AI, computer science, machine learning, NLP, data mining, software engineering, or related field; experience with LLMs or agents; strong Python and production-language programming skills.
Python, Java, C++, Go, TypeScript, PyTorch, TensorFlow, JAX, vLLM, Hugging Face, LangChain, LlamaIndex, RAG, GRPO, PPO, CPT, SFT, DPO, RLHF, RLAIF, AutoResearch, Harness, Clone & Adapt
4h
Save
Mark Applied
Hide
Staff Machine Learning, GAI Search Relevance - Moveworks
Mountain View, California, United States
$176k-$308k/yr OnsiteFull Time
ServiceNow
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
Experience with information retrieval or machine learning ranking models, query log analysis, search quality improvement, and text processing, natural language understanding, or machine learning.
Golang
4h
Save
Mark Applied
Hide
Staff Machine Learning, GAI Search Relevance - Moveworks
Mountain View, California, United States
$176k-$308k/yr OnsiteFull Time
ServiceNow
ServiceNowNYSE: NOW: Enterprise cloud platform for digital workflow automation.
Experience with information retrieval or machine-learning ranking models, query-log analysis, text processing, natural language understanding, or machine learning; Golang experience preferred.
Golang
6h
Save
Mark Applied
Hide
Senior Machine Learning Engineer
San Jose, California, United States
$230k-$360k/yr HybridFull Time
Roku
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
5+ YOE5+ years building software solutions; strong computer science fundamentals; proficiency in Java, Scala, Kotlin, or Python; big data and ML platform experience; AI literacy; and an MS in computer science or related field.
Java, Scala, Kotlin, Python, Spark, Kafka, Flink, Amazon S3, Airflow, Ray, PyTorch, Hugging Face, AWS SageMaker, Claude Code, Cursor, MCP servers
8h
Save
Mark Applied
Hide
Senior ML Engineer
San Francisco, California, United States
$170k-$220k/yr OnsiteFull Time
Echo Neurotechnologies
Echo Neurotechnologies: Developing brain-computer interface technologies to improve patient autonomy.
5+ YOEBachelor’s or master’s degree in a quantitative field, 5+ years of combined ML experience, Python and TensorFlow or PyTorch proficiency, and experience with GPUs and GPU software toolkits.
Python, TensorFlow, PyTorch, GPUs, GPU software toolkits, LLMs
9h
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Vision Models
Sunnyvale, California, United States
$312k-$370k/yr HybridFull Time
Wayve
Wayve: Develops end-to-end artificial intelligence for autonomous driving systems.
4+ YOERequires 4+ years in ML engineering, production deep learning, computer vision and foundation models, Python, PyTorch, large-scale training, model evaluation, and strong ownership. Autonomous vehicle experience is desirable.
Python, PyTorch
9h
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Vision Models
Sunnyvale, California, United States
$312k-$370k/yr HybridFull Time
Wayve
Wayve: Develops AI software for autonomous vehicle navigation.
4+ YOE4+ years of ML engineering experience shipping deep learning models, with computer vision, foundation model fine-tuning, Python, PyTorch, large-scale training, software engineering, model measurement, ownership, collaboration, and mentoring skills.
Python, PyTorch
1d
Save
Mark Applied
Hide
Machine Learning Engineer Graduate (E-Commerce Risk Control) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's degree in computer science, computer engineering, or related field; master's preferred. Requires machine learning knowledge, Python, SQL/Hive, Hadoop, and strong data structures, algorithms, and problem-solving skills.
Python, SQL, Hive, Hadoop, TikTok