--- title: 'Research Engineer - Post-Training at Pluralis Research' canonical: 'https://feeny.ai/job/research-engineer-post-training-pluralis-research-usa-or-dvzaxxpvrv0x' type: 'job' last_seen: '2026-09-07' --- # Research Engineer - Post-Training at Pluralis Research - **Company:** Pluralis Research - **Location:** Usa OR, Australia - **Employment:** full-time - **Work type:** remote - **Posted:** 2026-08-31 - **Last confirmed live:** 2026-09-07 - **Apply:** https://jobs.ashbyhq.com/pluralis-research/37a272e7-f55c-42e6-8db8-5c94f9934d86 ## Job description Pluralis Research works on Protocol Learning: training and serving large models in a fully decentralized way on small consumer-grade devices connected via the internet. Despite being dismissed as infeasible, we have made significant advances on this problem, most recently Agora, a permissionless run that pretrained an 8B model from scratch on consumer GPUs spread over the internet, with no single participant ever holding the full weights (tech report https://arxiv.org/abs/2607.13332). While many of the core research problems have been solved, Protocol Learning unlocks a series of new challenges. For the mission in full, read A Third Path: Protocol Learning https://pluralis.ai/blog/a-third-path-protocol-learning/. Agora gave us a pretrained 8B model. Post-training is how we make it useful for agentic use-cases. But every post-training stack you've seen assumes a datacenter — synchronous rollouts, fast interconnects, trusted workers. Ours gets none of that. It has to run on consumer GPUs, and Macs spread across the public internet, training a model whose weights no single participant ever holds, with rollouts arriving from a geo-distributed inference pipeline at high latencies. Your primary role is to make RL post-training work here anyway — the algorithms and the system, end-to-end. ## KEY RESPONSIBILITIES - Build the post-training stack: You build the RL training loop end-to-end: rollout ingestion from the geo-distributed inference pipeline, reward computation, policy updates, and getting updated weights back out to the network. You set the direction, and you make things happen. - Invent the algorithms: Standard RL recipes assume on-policy rollouts from fast, trusted hardware. You adapt them to asynchronous, high-latency, partially trusted generation: staleness tolerance, off-policy corrections, and communication-efficient policy updates. - Ship first post-trained models: You build the evals that show the models are improving, and you take the first decentralized post-trained release from run to public artifact. ## WHAT WE'RE LOOKING FOR - Hands-on RL post-training: You've run RL post-training on large language models — RLHF, RLVR, or reasoning-focused RL — and touched the systems layer yourself: rollout generation, async training loops, weight synchronization. Not just launched jobs on someone else's stack. - Strong engineering: Production-quality Python and PyTorch: concurrency, failure handling, profiling before optimizing. - Research ability: Publications in RL post-training, asynchronous or distributed RL, or nearby fields are a strong signal. So is unpublished work you can defend in detail. - Mission alignment: You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI. ## NICE TO HAVE - Experience training over slow networks, or with decentralized or federated setups. - Familiarity with serving-engine internals such as vLLM or SGLang — our rollout pipeline is a serving system. - Experience with reward modeling or building verifiable-reward datasets. - Experience with P2P networking and NAT traversal. - Experience at proprietary, open-weight and open-source AI labs ## COMPENSATION & BENEFITS - Equity-Heavy Package: We offer significant ownership for key technical contributors in addition to a high base salary. - Remote-First Culture: Flexible work environment with team members distributed globally. - Visa Sponsorship: Optional full visa sponsorship and relocation support to either Australia or the US. - Open Problems: Training and serving frontier models on hardware you don't control, over networks you don't own, mostly has no published answers yet. You'll write some of the first ones. ## FYI'S - We work remotely across the world, with the main teams in Australia and North America. You'll need to be comfortable working across timezones. - Applicants must have professional-level English proficiency (written and spoken). - Recruiters: we aren't looking for agency support at this time. We'll reach out if we need help. We are backed by Union Square Ventures https://www.usv.com/ and other tier-1 investors, and we are a world-class, deeply technical team of ML researchers. Pluralis is unapologetically ideological. We believe AI, and the world, end up on a better path if we succeed in implementing the protocol for intelligence. If this resonates, please apply. ## About Pluralis Research ## Company Overview - **One-liner**: Pluralis Research is developing protocol learning to enable decentralized, collectively-owned and collaboratively-trained frontier AI models that are unextractable and operate over open, permissionless networks. - **Entity Type**: Private (Seed stage) - **Headquarters**: San Francisco, California, United States - **Founded**: 2024 - **Founders**: Alexander Long (Founder and CEO); Founding Scientists include Ajanthan Thalaiyasingam, Gil Avraham, and Sameera Ramasinghe. ## Core Business - **Primary Industry**: Foundational AI Research (Decentralized AI / Protocol Learning) - **Target Customers**: B2B (research institutions, compute providers, open-source community contributors); B2C indirectly (future users of collectively owned models) - **Mission/Purpose**: To create a third path for AI development beyond closed-source models and open-weight models—a self-sustaining, collectively owned ecosystem where models are trained and governed by a community, not a single corporation. ## Products & Services - **Protocol Learning (Core Protocol)**: A novel training paradigm where large AI models are split across geographically separated devices connected only by standard internet links. - **Type**: Research Protocol / Decentralized Infrastructure - **Key Innovation**: Models exist only within the protocol, not on any single device. They are trainable and usable but **unextractable**, preventing any single entity from capturing the model weights. - **Node0**: A collaborative event powered by Protocol Learning, enabling decentralized AI training runs. - **Type**: Open-source tool / Event Framework - **Agora**: A collaborative training library for multi-participant model training. - **Type**: Open-source Library (GitHub) - **AsyncPP & AsyncMesh**: Communication-efficient pipeline and mesh parallelism techniques for scaling decentralized models over low-bandwidth networks (achieving >95% compression and matching centralized wall-clock convergence). - **Type**: Open-source Algorithms (GitHub) ## Market Standing - **Valuation**: Not publicly disclosed - **Total Funding**: $8.6 million - **Non-Equity Assistance** (Nov 2025): $1.0M led by Amazon Web Services - **Pre-Seed Round** (Jan 2025): Led by CoinFund - **Seed Round** (Mar 2025): $7.6M co-led by Union Square Ventures and CoinFund, with participation from Topology, Variant, Eden Block, and Bodhi Ventures. Angel investors include Balaji Srinivasan and Clem Delangue (co-founder of Hugging Face). - **Notable Investors/Partners**: Union Square Ventures, CoinFund, Amazon Web Services, Balaji Srinivasan, Clem Delangue (Hugging Face) - **Growth Signals**: - Headcount grew **+100% YoY** (from ~7 to 17 employees). - Published multiple papers at **NeurIPS 2025** on decentralized training (e.g., Subspace Networks, Mixtures of Subspaces). - 10 active job openings across Research, ML Engineering, and Distributed Systems. - Team has deep FAANG/Anthropic pedigree (Google, Amazon, Oracle, Atlassian alumni). ## Competitive Advantages - **Technical Moat (Unextractable Models)**: Their core IP ensures model weights are never materialized on any participant's device, preventing theft or centralization even during training. - **Economic Moat (Self-Sustaining Economics)**: The protocol allows value to flow programmatically to contributors, solving the financial sustainability problem that plagues purely open-weight models. - **Published Research**: Validated by acceptance at NeurIPS 2025, demonstrating academic credibility and technical rigor. - **Talent Density**: The team comprises PhD researchers who previously worked together at top AI labs (Anthropic, Google, Amazon), with deep expertise in distributed systems and ML. ## Strategic Focus - **Hiring in All Areas**: Currently recruiting across Research, ML Training Platforms, and Distributed ML Systems to accelerate protocol scaling. - **Scaling Decentralized Networks**: Aiming to prove that model-parallel training over low-bandwidth (30Mbps) internet is feasible for billion-parameter models. - **Building Community**: Opening participation in training runs to external contributors, moving toward a truly permissionless network. ## Why Work Here - **Culture of Radical Openness**: The mission is to democratize AI ownership. Work here is published openly, and researchers contribute to a public good. - **High Impact, Small Team**: With only 17 people, every hire has an outsized influence on shaping the protocol and company direction. - **Research First**: The team is PhD-heavy and publishes at top-tier conferences (NeurIPS). The environment is scholarly and technically deep. - **Remote-Flexible (Hybrid)**: Presence in San Francisco (US) and Australia (largest cohort of 12 employees). The job postings do not mandate 5 days in-office; distributed collaboration is core to the product itself. - **Top-Tier Investor Backing**: Backed by USV and CoinFund, providing strong financial runway and network effects in both the AI and crypto/Web3 ecosystems. - **Founding Team Pedigree**: Work alongside former researchers from Anthropic, Google, and Amazon, creating a steep learning curve for ML engineers and scientists. ## Sources 1. [pluralis.ai](https://pluralis.ai/) 2. [LinkedIn - Pluralis Research](https://www.linkedin.com/company/pluralis-research) 3. [GlobeNewswire - Seed Round Announcement](https://www.globenewswire.com/news-release/2025/03/19/3045635/0/en/Pluralis-Research-Pioneers-Protocol-Learning-to-Scale-Decentralized-AI-Announces-7-6M-Seed-Round-Led-by-USV-and-CoinFund.html) 4. [GitHub - Pluralis Research](https://github.com/PluralisResearch) 5. [Jobs.ashbyhq.com - Pluralis Research Careers](https://jobs.ashbyhq.com/pluralis-research) ## Other roles at Pluralis Research - [Machine Learning Engineer - ML Training Platform](https://feeny.ai/job/machine-learning-engineer-ml-training-platform-pluralis-research-usa-or-cb0b1g58vsjh) — Usa OR, Australia - [Research Engineer Intern](https://feeny.ai/job/research-engineer-intern-pluralis-research-australia-62bya07kzppe) — Australia - [Research Scientist Intern](https://feeny.ai/job/research-scientist-intern-pluralis-research-usa-or-gc7q96xt6kh1) — Usa OR, Australia - [Research Engineer - Decentralized Training and Inference Verification](https://feeny.ai/job/research-engineer-decentralized-training-and-inference-verification-pluralis-7gw1639gefck) — Usa OR, Australia - [Research Engineer - Geo-Distributed Inference](https://feeny.ai/job/research-engineer-geo-distributed-inference-pluralis-research-usa-or-gk4y7fqhyk80) — Usa OR, Australia - [Research Engineer - Pre-training](https://feeny.ai/job/research-engineer-pre-training-pluralis-research-usa-or-55zxjrp2f7s5) — Usa OR, Australia - [Research Scientist](https://feeny.ai/job/research-scientist-pluralis-research-usa-or-v7z1rqs3sas3) — Usa OR, Australia - [Founders Associate (Business Operations)](https://feeny.ai/job/founders-associate-business-operations-pluralis-research-sydney-z2b3q8c92pwz) — Sydney, Australia - [Events & Operations Associate](https://feeny.ai/job/events-operations-associate-pluralis-research-san-francisco-51ht4svfj5tx) — San Francisco, CA - [Academic Collaboration](https://feeny.ai/job/academic-collaboration-pluralis-research-sydney-kwpagzyta862) — Sydney, Australia