T

Senior/Staff Machine Learning Engineer - Distributed Systems

torentify United State
Remote Visa Sponsorship
Apply
AI Summary

Design and implement large-scale distributed machine learning training infrastructure for decentralized, community-owned models. Optimize model parallelism, GPU utilization, and peer-to-peer networking for heterogeneous, low-bandwidth environments. Requires 5+ years of experience in distributed systems, Python, and large-scale ML frameworks like FSDP or DeepSpeed.

Key Highlights
Build novel infrastructure for distributed ML training across consumer-grade internet connections.
Optimize for heterogeneous hardware, low-bandwidth, and high-latency environments.
Develop resilient peer-to-peer communication and fault-tolerant training systems.
Key Responsibilities
Design and implement large-scale distributed machine learning training systems.
Optimize distributed training for heterogeneous hardware operating in low-bandwidth and high-latency environments.
Develop model-parallel training strategies, including data, tensor, and pipeline parallelism.
Implement custom sharding techniques to reduce communication overhead.
Optimize GPU utilization, memory efficiency, and overall compute performance.
Develop reliable checkpointing and state synchronization mechanisms.
Build recovery systems for long-running and fault-prone training workloads.
Develop monitoring and metrics systems to track training progress, model quality, and infrastructure bottlenecks.
Architect resilient distributed training systems capable of handling node failures and network partitions.
Design systems that support participants dynamically joining or leaving the network.
Develop peer-to-peer communication and coordination mechanisms.
Implement peer discovery, NAT traversal, dynamic routing, and connection lifecycle management.
Analyze and optimize communication patterns across distributed nodes.
Reduce network latency and bandwidth requirements in multi-participant training environments.
Build systems capable of operating reliably across non-co-located infrastructure.
Technical Skills Required
Python Distributed Systems Machine Learning Infrastructure
Benefits & Perks
Equity-heavy compensation
Competitive base salary
Visa sponsorship available
Remote-first working environment
Optional access to Melbourne hub

Job Description



About the Company


Pluralis Research conducts foundational research in Protocol Learning, an approach to training foundation models across multiple participants without requiring any single participant to hold a complete copy of the model.

The company is developing technology designed to enable community-trained and community-owned frontier models through decentralized machine learning infrastructure and sustainable economic models.

Pluralis Research is backed by leading investors and brings together experienced machine learning researchers and distributed systems engineers from major technology companies and innovative startups.


About the Role


Pluralis Research is seeking Senior and Staff-level Machine Learning Engineers with strong experience in distributed systems and large-scale machine learning training.

In this role, you will help build a novel infrastructure layer for distributed ML training designed to operate across consumer-grade internet connections and heterogeneous computing environments.

You will work at the intersection of distributed systems, machine learning infrastructure, networking, GPU optimization, and decentralized computing.


Responsibilities


Distributed Training Architecture & Optimization

  • Design and implement large-scale distributed machine learning training systems.
  • Optimize distributed training for heterogeneous hardware operating in low-bandwidth and high-latency environments.
  • Develop model-parallel training strategies, including data, tensor, and pipeline parallelism.
  • Implement custom sharding techniques to reduce communication overhead.
  • Optimize GPU utilization, memory efficiency, and overall compute performance.
  • Develop reliable checkpointing and state synchronization mechanisms.
  • Build recovery systems for long-running and fault-prone training workloads.
  • Develop monitoring and metrics systems to track training progress, model quality, and infrastructure bottlenecks.

Decentralized Networking & Resilience

  • Architect resilient distributed training systems capable of handling node failures and network partitions.
  • Design systems that support participants dynamically joining or leaving the network.
  • Develop peer-to-peer communication and coordination mechanisms.
  • Implement peer discovery, NAT traversal, dynamic routing, and connection lifecycle management.
  • Analyze and optimize communication patterns across distributed nodes.
  • Reduce network latency and bandwidth requirements in multi-participant training environments.
  • Build systems capable of operating reliably across non-co-located infrastructure.

Required Qualifications

  • 5+ years of professional experience building and operating distributed systems.
  • Strong production experience with large-scale machine learning systems.
  • Hands-on experience with distributed training frameworks such as FSDP, DeepSpeed, Megatron, or comparable technologies.
  • Deep understanding of model parallelism, including:
  • Data parallelism
  • Tensor parallelism
  • Pipeline parallelism
  • Expert-level Python programming skills.
  • Production experience with Python concurrency, error handling, retry mechanisms, and clean software architecture.
  • Strong understanding of networking and distributed systems fundamentals.
  • Experience with peer-to-peer systems, gRPC, routing, NAT traversal, and distributed coordination.
  • Experience optimizing GPU workloads and memory management.
  • Strong understanding of large-scale compute performance and resource optimization.
  • Ability to troubleshoot complex distributed systems and performance issues.


Technical Skills


  • Python
  • Distributed Systems
  • Machine Learning Infrastructure
  • FSDP
  • DeepSpeed
  • Megatron
  • Model Parallelism
  • GPU Optimization
  • Memory Management
  • P2P Networking
  • gRPC
  • NAT Traversal
  • Distributed Coordination
  • Checkpointing
  • State Synchronization
  • Fault Tolerance
  • Performance Optimization


What We Offer


  • Equity-heavy compensation with meaningful ownership opportunities.
  • Competitive base salary for senior engineering positions.
  • Visa sponsorship available for exceptional candidates.
  • Remote-first working environment.
  • Optional access to the company's Melbourne hub.
  • Opportunity to work with an experienced team of ML researchers and distributed systems engineers.
  • Exposure to cutting-edge distributed machine learning research and infrastructure.
  • Opportunity to contribute to an ambitious approach to decentralized AI development.


Work Environment


This is a remote-first engineering position focused on highly technical machine learning infrastructure. Engineers will work on challenging problems involving distributed training, decentralized networking, GPU optimization, and fault-tolerant systems.

The role is particularly suited to engineers who enjoy working on complex infrastructure problems and developing new approaches to large-scale machine learning.

Equal Opportunity

Pluralis Research is committed to building a diverse and inclusive team. Employment decisions are based on qualifications, skills, experience, and organizational needs, without discrimination based on legally protected characteristics.


Similar Jobs

Explore other opportunities that match your interests

Machine Learning Research Scientist

Machine Learning
6h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

Harnham

United State

Senior Machine Learning Engineer

Machine Learning
1d ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

hired

United State

Researcher in Self-Improving Machine Learning

Machine Learning
1d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

chemanager international

United State

Subscribe our newsletter

New Things Will Always Update Regularly