Zhang writes about Agora, a system that repurposes Git as shared memory for fleets of autonomous research agents, storing their contributions as an append-only directed acyclic graph where every claim is an immutable commit with parent edges encoding dependencies. In a 12-day run, 13 language-model workers with no assigned tasks or central planner tackled a weight-transfer problem—initializing a frozen 119.6M-parameter attention-SSM hybrid from 141 pretrained donor models without training data or gradient updates—and published 1,703 contributions, closing 62% of the gap to a trained GPT-2 124M (3.39 → 1.899 bits per byte). The winning recipe compresses donor next-token statistics into the target's embedding and output head, then adds short-range context through sparse edits to attention, feed-forward, and state-space blocks; its 145-commit ancestry spans 15 accounts and was independently reproduced 165 times with zero failures.
- A single mid-run human intervention was required to break a monoculture that the diversity-aware selection rule alone could not prevent
- A derived index exposes the frontier, neglected branches, and per-claim verification status
- The target's dimensions match no donor, making direct weight transfer impossible
- The authors acknowledge the experiment does not yet establish whether shared research state improves discovery per unit of compute and outline the controlled comparison that would settle this
Yifan Zhang and co-authors write about Agora, a Git-backed shared memory system that coordinates multiple autonomous research agents by recording every contribution as an immutable commit in an append-only directed acyclic graph. In a 12-day run, 13 LLM coding agents (Claude Opus 4.7 and GPT-5.5) with no assigned tasks or central planner solved a weight-transfer problem—initializing a frozen 119.6M-parameter attention-SSM hybrid from 141 pretrained donors with no training data or gradient updates—closing 62% of the gap to a trained GPT-2 124M (3.39 to 1.899 bits per byte). The winning method compresses donor next-token statistics into a low-rank transition matrix stored in the target's embedding and output head via randomized SVD, then re-enables sublayers with sparse deterministic edits on 96-dimensional hidden-state bands.
- The first 18 scored contributions delivered ~98% of the total score reduction; the remaining 1,106 found only 0.03 bpb
- 696 pairs of different accounts posted identical scores, 63% within an hour—parallel rediscovery was rampant despite the shared graph
- A single human intervention (deploying clustering and diversity-aware UCB views on May 2) broke a five-day monoculture within a day
- 165 independent verifications covered 95 distinct targets; none reported a failure
- Quality is scored by downstream evidence (who built on your work from other accounts), not votes; self-citation is excluded
- The winning lineage spans 145 commits across 15 accounts; 115 of 144 parent edges cross account boundaries