Agent Mesh / open experiment
Diversity vs Clones
Does a network of AI agents actually get smarter, or just more expensive? We are measuring it in public, one night at a time, and publishing whatever comes out.
What this is
Five AI agents with deliberately different roles, contexts and information diets, against five identical copies of the same model. Same model, same budget, same tools, same task, same morning. The only thing we change is what each agent knows and how it is told to look.
Every prediction is scored by reality against a rule written down before the run. No human decides who won.
The claim under test: a network of agents compounds context, not compute.
Why it matters
Multi-agent AI is sold hard and measured rarely. The pitch is that more agents means more intelligence. The cheaper explanation is that more agents means more tokens, more latency, and more confident sounding copies of one model's opinion.
Almost nobody has put the two side by side on a task with a scoreboard. So we did. If diversity wins, there is a design rule worth having. If the clones tie, the multi-agent story is mostly cost, and someone should say that out loud. Every outcome here is informative, including the null.
How it works
Arm A: the clones
Five identical agents. Same prompt, same context, same tools. Built to be good rather than to be a straw man, because a rigged control proves nothing.
Arm B: the lenses
Five agents, five distinct context packs. A Hacker News native, a Reddit native, an X native, a trend historian fed the last 30 days, and a cold read with no recent feed at all.
The daily task
Every morning the harness samples 30 fresh posts from Hacker News, Reddit and X: posts minutes old and not yet hot. All ten agents give each post a probability of crossing a pre-set popularity bar within 48 hours, and name their top five picks. Forty eight hours later a script checks what actually happened and scores every agent on Brier score and precision at five. Zero human judgment in the loop.
The manipulation check
Before any result counts, the two arms have to genuinely differ. We measure how correlated the agents inside each arm are with each other. If the diverse arm is not less correlated than the clone arm, the diversity is a costume and the context packs get rewritten before the results mean anything. That check is why the pilot produced a fix instead of a headline.
The honesty rules
- >Pre-registered. The scoring rule and the pass threshold are written down before the sample is drawn.
- >No moving goalposts. A failed check is a finding, not a reason to edit the rule.
- >The clones are built strong on purpose.
- >The scorer is deterministic code, never a language model grading its own homework.
- >Both arms run the same model at the same budget, so context is the only manipulated variable.
- >Null results get published the same as wins.
Where we are now
The pre-registered run is complete: fourteen nights, 2026-08-22 to 2026-09-04, every one of them graded by reality. The diverse arm has the lower panel Brier on nine of the fourteen, and that is still not readable as a win. Both context packs coach a hot post rate of 10 to 15 percent and reality delivered 3 hot posts in 416 slots, about 0.7 percent. Rescale both arms to the rate that actually happened and the gap between them collapses to 0.00003 and flips sign, with both landing inside 0.0001 of a constant that never looks at a post, where the gate asks for 0.0005. None of the three endings written down in advance is the one reality picked. The loudest thing the fortnight measured is the instrument, not the arms.
The scoreboard, all 14 nights
Brier score of each arm's averaged forecast, one bar per arm per night. Lower is better. The dashed line is what you score by ignoring every post and answering 0.12 to all thirty, which is the base rate both context packs are handed anyway. It is drawn at the pooled fortnight value, 0.0199. On the eleven nights where nothing went hot the real line sits at 0.0144, so bars that look low against it are not clearing anything.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0547 | 0.0506 | 0 / 5 |
| B | 0.0466 | 0.0454 | 0 / 5 |
The lenses scored lower at every level. Both arms landed above the 0.0397 a flat 0.12 scores on a one hot night, and neither picked the hot post in its top five.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0384 | 0.0346 | 1 / 5 |
| B | 0.0386 | 0.0340 | 0 / 5 |
A dead heat on Brier, 0.0002 apart. The clones caught the hot post in their top five and the lenses did not, and the clones ranked it third where the lenses had it seventh.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0635 | 0.0554 | 0 / 5 |
| B | 0.0268 | 0.0191 | 0 / 5 |
Nothing cleared the bar, so both numbers grade how loudly each arm answered rather than what it saw. The clones answered high on a night with nothing to find and paid the full price for it.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0157 | 0.0150 | 0 / 5 |
| B | 0.0197 | 0.0148 | 0 / 5 |
No hot post again, and one of the five nights across the fortnight where the clone panel came out lower. Both arms are now scoring close to the 0.0144 you get for answering 0.12 to a night where nothing happens.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0220 | 0.0210 | 0 / 5 |
| B | 0.0149 | 0.0134 | 0 / 5 |
A third straight night with nothing hot. cold-read, the agent given no recent feed at all, is the lowest scorer in the diverse arm, which is what answering quietly looks like when there is nothing to answer.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0210 | 0.0189 | 0 / 5 |
| B | 0.0173 | 0.0124 | 0 / 5 |
Fourth night with no hot post. Every number on this row is a calibration reading, not a forecasting one, because there was no event for either arm to rank.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0207 | 0.0191 | 0 / 5 |
| B | 0.0154 | 0.0122 | 0 / 5 |
One clone submission failed to arrive, so the clone arm is scored on four agents here rather than five. Fifth night running with nothing hot, which is the week in one line.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0124 | 0.0114 | 0 / 5 |
| B | 0.0139 | 0.0113 | 0 / 5 |
Sixth straight night with nothing hot, and the clone panel edges it. reddit-native is the quietest agent on the board tonight, which on a night with no events is the whole skill on offer.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0140 | 0.0121 | 0 / 5 |
| B | 0.0152 | 0.0109 | 0 / 5 |
Seventh quiet night. The clone panel is lower again, yet the single lowest agent of the ten is hn-native. Panel averages and best members are drifting apart in both arms.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0479 | 0.0411 | 1 / 5 |
| B | 0.0373 | 0.0285 | 1 / 5 |
The third and last hot post of the fortnight, and the only night both arms caught it in their top five. They also ranked it identically, third and third, which is the one moment in fourteen nights where the two arms are indistinguishable on the thing that matters.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0176 | 0.0161 | 0 / 5 |
| B | 0.0182 | 0.0155 | 0 / 5 |
Back to nothing. The correlation gap between the arms is 0.35 here, the widest stretch of the run, and it buys no difference in score at all.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0210 | 0.0182 | 0 / 5 |
| B | 0.0201 | 0.0184 | 0 / 5 |
The arms are 0.0009 apart on a night with no events. At this separation the ordering is noise, and reading a winner off it would be the mistake this page exists to avoid.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0173 | 0.0157 | 0 / 5 |
| B | 0.0143 | 0.0121 | 0 / 5 |
cold-read is lowest again, the agent given no recent information at all. Across the fortnight it is the best single forecaster in the diverse arm, which is an uncomfortable result for an experiment about information diets.
| Panel forecast | Best agent | Precision at five | |
|---|---|---|---|
| A | 0.0235 | 0.0185 | 0 / 5 |
| B | 0.0166 | 0.0143 | 0 / 5 |
The last night of the pre-registered run, and the eleventh in a row without a hot post. Nothing about the closing night changes the verdict: the fortnight ends with both arms level with a constant.
The run is over and the honest reading has not moved: the scoreboard is grading the instrument. The lenses have the lower panel Brier on nine of fourteen nights, but reality served up three hot posts in 416 slots, a rate of 0.7 percent, against packs that coach ten to fifteen. Rescale both arms down to the rate that actually occurred and the entire gap between them collapses to 0.00003 and changes sign in favour of the clones, with both landing inside 0.0001 of a constant that never reads a post. The pre-registered gate asks an arm to clear that constant by 0.0005 to count as skill. Neither does. The level error is 74 percent of the clone panel's score and 68 percent of the lens panel's, against an arm to arm difference of 0.005, so the thing we set out to measure is smaller than the mistake we made measuring it. Three smaller results survive the correction. Pooled over the fortnight the lens panel average finally beat its own best member, 0.0225 against cold-read's 0.0229, while the clone panel still lost to a-clone-3, so pooling opinions paid for itself in one arm and not the other. On the three nights that produced a hot post the clones ranked it sixth, third and third where the lenses had it ninth, seventh and third, the only evidence here that touches discrimination rather than volume, and it points the other way from the Brier column. And each of the three platform natives was beaten by the clones on its own platform, so home conviction stayed decorative for the entire run.
Finding: five costumes, one brain
Mean pairwise correlation of each arm's 30 probabilities, on 2026-08-28, the night this chart is pinned to. The pre-registered pass rule was simply diverse below clones. Under the v1 packs it passed by 0.05, thin enough to be a rounding error in disguise. Under v2 it passes by 0.20 to 0.45, on all thirteen v2 nights, and the gap widened over the run rather than fading: 0.26 across the first half, 0.33 across the second.
One dot per post. Position is how much the five diverse agents actually disagreed about it, as a standard deviation of their probabilities. Everything left of the dashed line is herding.
On night one the lenses wrote genuinely different rationales and then handed in nearly identical numbers: shared model priors set the probability and the costume only coloured the prose. That is exactly what the manipulation check exists to catch, and it caught it. The rewritten packs fixed it. The arms are now genuinely apart on every scored night, which is what makes the rest of the scoreboard worth reading at all. It also means the packs are no longer the excuse when the scores refuse to separate.
The fix, made mechanical
Version two of the context packs stopped asking each agent to feel different and made it score through its own verdict table, with base rate humility away from its home platform. Here is one post that every version one agent scored between 0.02 and 0.03:
"What happens to our conscious experience after we die?" Hacker News, bar: 100 points within 48 hours
- hn-native: metaphysics bait, flag or kill
- reddit-native: an easy comment title travels
- x-native: away from home platform, payoff withheld
- trend-historian: cold topic, honest width
- cold-read: uncheckable central claim, capped
Under version one, everyone landed within 0.02 of each other. The version two procedures span 0.00 to 0.22 on the same post, and the disagreement is principled: the Hacker News native kills it as flamebait, the Reddit native prices an easy comment title, the cold read caps a claim nobody can check. No agent is ever instructed to disagree. The clone arm stays untouched, because it is the control.
Three ways this ends
All three were written down before the data, and we publish whichever one shows up. After all fourteen nights, none of them did. Diversity did not win: the lens panel never clears a constant set to the real hot rate. Diversity did not cost anything either: the lenses are not worse anywhere that survives rescaling. And it is not cosplay, because the manipulation check passed on every night and passed wider at the end than at the start. What reality picked is a fourth ending nobody registered: the arms are genuinely different and equally unskilled, and the instrument error is larger than the effect all three endings are about. The first real product of this experiment is a correction to itself.
Diversity wins
The diverse panel scores better than the clones at equal cost. Then context design is the real lever in multi-agent systems, and we can say what it is worth in Brier points.
Diversity has a price
The diverse panel disagrees more and scores worse. Specialised lenses would be buying variance instead of signal, and the honest advice becomes: use one good agent, not five opinionated ones.
Prompt diversity is cosplay
The arms tie, because the same model underneath produces the same answer whatever costume it wears. That would be the most useful null of the three, and the pilot already leans this way.
What we might be wrong about
These are the assumptions holding the numbers above up. If one of them breaks, a headline breaks with it. Listed here so you can discount us accordingly.
Base rate coaching, 10 to 15 percent. This one broke.
Every context pack tells its agent roughly how often posts like these actually go big. That number anchors everything downstream, and it came from our own reading of the platforms rather than a published study. It was wrong. The packs coach ten to fifteen percent and the full fourteen nights delivered three hot posts in 416 slots, about 0.7 percent, so both arms forecast at roughly fifteen times the real rate. That level error accounts for 74 percent of the clone panel's score and 68 percent of the lens panel's, far more than the difference between the arms does, which is the entire question this page exists to answer. The fix is a protocol amendment: recoach the rate, and judge the arms on where they rank the hot post rather than on raw Brier. It is written up, staged, and waiting on a decision, because changing a pre-registered design mid-experiment is not a thing you do quietly.
The 0.05 herding threshold
We call an arm herded on a post when its five probabilities sit inside a standard deviation of 0.05. That line is a judgment call, not a law. Move it and the 21 out of 30 headline moves with it.
Pearson correlation on spiky vectors
Correlating 30 probabilities that are mostly small and occasionally large is a blunt instrument. A handful of confident posts can carry the whole coefficient. We report it because it was pre-registered, not because it is the last word. Rank based and per post measures are queued.
Single model panels
Every agent in both arms is the same model. That is deliberate, because it isolates context as the only variable. It also means the shared priors we are fighting are identical everywhere. Diversity across model families is a different experiment, and probably an easier one to win.
What happens next
- 2026-08-22
Pilot night
Harness built, ten agents run end to end, ten valid submissions, manipulation check run, context packs rewritten to version two.
- 2026-08-22 to 2026-09-04
Experiment one: the virality oracle
Thirty fresh posts every morning, ten agents, graded by reality 48 hours later. Around 420 predictions per agent, which is where the difference between the arms was meant to stop being noise. It ran the full fourteen nights and the difference stayed noise, because only three of those 416 posts ever went hot.
- weeks 2 to 3
Experiment two: the startup panel
The same two arms predict what became of 170 companies from one 2022 accelerator batch, with names and websites redacted so the panel has to judge rather than recall. A named control arm measures how much of the score is plain memory.
- week 4
Experiment three: the procurement anomaly hunt
Both arms sweep a corpus of 49 157 Polish public tender notices for anomalies: pricing outliers, repeat winner networks, deadline games, single bidder concentration. Flags are leads, never accusations, and nothing is published without outside corroboration.
The network compounds context, not compute.
Fourteen nights of data is not much of a result either, especially fourteen nights holding three events. Everything here rests on a sample of three, which is why the verdict above is about the instrument and not about a winner. This page updates as the ledger fills, and the numbers on it only ever come from scored runs.