T

Senior/Staff Machine Learning Engineer, Distributed Systems

torentify United State
Remote Visa Sponsorship
Apply
AI Summary

Pluralis Research seeks experienced ML Engineers to build novel infrastructure for distributed ML model training across consumer-grade internet. Responsibilities include designing large-scale distributed training systems, optimizing for low-bandwidth networks, and ensuring system resilience. Requires 5+ years in distributed systems, large-scale ML training, and expert Python skills.

Key Highlights
Build novel infrastructure for distributed ML training.
Optimize training for low-bandwidth and high-latency networks.
Ensure system resilience and fault tolerance in decentralized environments.
Key Responsibilities
Design and implement large-scale distributed training systems for heterogeneous hardware.
Optimize training systems for low-bandwidth and high-latency network conditions.
Develop model-parallel training strategies, including data, tensor, and pipeline parallelism.
Implement custom sharding techniques to minimize communication overhead.
Optimize GPU utilization, memory efficiency, and compute performance across distributed nodes.
Build robust checkpointing, state synchronization, and recovery mechanisms for long-running training workloads.
Develop monitoring and metrics systems to track training progress, model quality, and system bottlenecks.
Architect resilient distributed training systems capable of handling node failures and network partitions.
Support dynamic participant joining and leaving within distributed environments.
Design and optimize peer-to-peer network topologies for decentralized coordination.
Implement NAT traversal, peer discovery, dynamic routing, and connection lifecycle management.
Profile and optimize communication patterns to reduce latency and bandwidth requirements.
Build systems capable of operating reliably across non-co-located infrastructure.
Technical Skills Required
Distributed Systems Large-scale Machine Learning Training Python
Benefits & Perks
Equity-heavy compensation
Competitive base salary
Visa sponsorship available
Remote-first working environment
Nice to Have
FSDP, DeepSpeed, Megatron, or comparable frameworks
Data, tensor, and pipeline parallelism
GPU optimization and memory efficiency
Concurrent systems
Peer-to-peer networking
gRPC and distributed coordination
NAT traversal and peer discovery
Fault-tolerant distributed infrastructure

Job Description




## About the Company


Pluralis Research conducts foundational research into Protocol Learning, a distributed approach to training foundation models where no single participant holds or can obtain a complete copy of the model.


The company is focused on enabling community-trained and community-owned frontier models through decentralized infrastructure and sustainable economic models. Its deeply technical team includes ML researchers and engineers with backgrounds at major technology companies and leading startups.


Pluralis Research is backed by Union Square Ventures and other tier-1 investors and is building novel infrastructure for distributed machine learning.


## About the Role


Pluralis Research is seeking Senior and Staff-level Machine Learning Engineers with 5+ years of experience in distributed systems and large-scale machine learning training.


In this role, you will help build a novel technical substrate for training distributed ML models across consumer-grade internet connections. You will work on large-scale distributed training, model parallelism, GPU optimization, decentralized networking, fault tolerance, and communication efficiency.


The ideal candidate will have deep production experience with distributed systems and hands-on expertise in modern distributed training frameworks, along with strong Python and networking skills.


### Key Responsibilities


#### Distributed Training Architecture and Optimization


* Design and implement large-scale distributed training systems for heterogeneous hardware.

* Optimize training systems for low-bandwidth and high-latency network conditions.

* Develop model-parallel training strategies, including data, tensor, and pipeline parallelism.

* Implement custom sharding techniques to minimize communication overhead.

* Optimize GPU utilization, memory efficiency, and compute performance across distributed nodes.

* Build robust checkpointing, state synchronization, and recovery mechanisms for long-running training workloads.

* Develop monitoring and metrics systems to track training progress, model quality, and system bottlenecks.


#### Decentralized Networking and Resilience


* Architect resilient distributed training systems capable of handling node failures and network partitions.

* Support dynamic participant joining and leaving within distributed environments.

* Design and optimize peer-to-peer network topologies for decentralized coordination.

* Implement NAT traversal, peer discovery, dynamic routing, and connection lifecycle management.

* Profile and optimize communication patterns to reduce latency and bandwidth requirements.

* Build systems capable of operating reliably across non-co-located infrastructure.


### Required Qualifications


* 5+ years of professional experience in distributed systems and large-scale machine learning training.

* Strong experience building and operating distributed systems in production.

* Hands-on experience with distributed training frameworks such as FSDP, DeepSpeed, Megatron, or similar technologies.

* Deep understanding of data, tensor, and pipeline parallelism.

* Expert-level Python development experience.

* Production experience with concurrency, error handling, retry mechanisms, and clean software architecture.

* Strong networking fundamentals.

* Experience with peer-to-peer systems, gRPC, routing, NAT traversal, and distributed coordination.

* Experience optimizing GPU workloads, memory management, and large-scale compute efficiency.

* Strong problem-solving and systems engineering abilities.


### Technical Focus


* Distributed machine learning.

* Large-scale model training.

* FSDP, DeepSpeed, Megatron, or comparable frameworks.

* Data, tensor, and pipeline parallelism.

* GPU optimization and memory efficiency.

* Python and concurrent systems.

* Peer-to-peer networking.

* gRPC and distributed coordination.

* NAT traversal and peer discovery.

* Fault-tolerant distributed infrastructure.


### Compensation and Benefits


* Equity-heavy compensation with meaningful ownership.

* Competitive base salary for senior engineering roles in Australia.

* Visa sponsorship available for exceptional candidates.

* Remote-first working environment.

* Optional access to the Melbourne hub.

* Opportunity to work alongside experienced ML researchers and engineers.

* Career opportunity within a mission-driven, technically focused organization.


### Work Arrangement


* Remote-first position.

* Optional access to the Melbourne hub.

* Role supports distributed collaboration across geographically separated infrastructure and teams.


## Equal Opportunity


Pluralis Research is committed to providing equal employment opportunities to qualified professionals. Employment decisions are made without regard to race, color, religion, sex, sexual orientation, gender identity, pregnancy, age, national origin, disability, veteran status, or any other status protected by applicable law. The company supports equality, inclusion, and a respectful working environment for all qualified candidates.



Similar Jobs

Explore other opportunities that match your interests

Machine Learning Engineering Manager

Machine Learning
13h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Harnham

United State

AI/ML Engineer - Remote

Machine Learning
2d ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

sundayy

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Deloitte

United State

Subscribe our newsletter

New Things Will Always Update Regularly