← WIZ
// EXPERIMENTS

Agent Mesh / open experiment

Diversity vs Clones

Does a network of AI agents actually get smarter, or just more expensive? We are measuring it in public, one night at a time, and publishing whatever comes out.

Live, updated as data landsLast run 2026-09-04
01

What this is

Five AI agents with deliberately different roles, contexts and information diets, against five identical copies of the same model. Same model, same budget, same tools, same task, same morning. The only thing we change is what each agent knows and how it is told to look.

Every prediction is scored by reality against a rule written down before the run. No human decides who won.

The claim under test: a network of agents compounds context, not compute.

02

Why it matters

Multi-agent AI is sold hard and measured rarely. The pitch is that more agents means more intelligence. The cheaper explanation is that more agents means more tokens, more latency, and more confident sounding copies of one model's opinion.

Almost nobody has put the two side by side on a task with a scoreboard. So we did. If diversity wins, there is a design rule worth having. If the clones tie, the multi-agent story is mostly cost, and someone should say that out loud. Every outcome here is informative, including the null.

03

How it works

Arm A: the clones

Five identical agents. Same prompt, same context, same tools. Built to be good rather than to be a straw man, because a rigged control proves nothing.

Arm B: the lenses

Five agents, five distinct context packs. A Hacker News native, a Reddit native, an X native, a trend historian fed the last 30 days, and a cold read with no recent feed at all.

The daily task

Every morning the harness samples 30 fresh posts from Hacker News, Reddit and X: posts minutes old and not yet hot. All ten agents give each post a probability of crossing a pre-set popularity bar within 48 hours, and name their top five picks. Forty eight hours later a script checks what actually happened and scores every agent on Brier score and precision at five. Zero human judgment in the loop.

The manipulation check

Before any result counts, the two arms have to genuinely differ. We measure how correlated the agents inside each arm are with each other. If the diverse arm is not less correlated than the clone arm, the diversity is a costume and the context packs get rewritten before the results mean anything. That check is why the pilot produced a fix instead of a headline.

The honesty rules

  • >Pre-registered. The scoring rule and the pass threshold are written down before the sample is drawn.
  • >No moving goalposts. A failed check is a finding, not a reason to edit the rule.
  • >The clones are built strong on purpose.
  • >The scorer is deterministic code, never a language model grading its own homework.
  • >Both arms run the same model at the same budget, so context is the only manipulated variable.
  • >Null results get published the same as wins.
04

Where we are now

The pre-registered run is complete: fourteen nights, 2026-08-22 to 2026-09-04, every one of them graded by reality. The diverse arm has the lower panel Brier on nine of the fourteen, and that is still not readable as a win. Both context packs coach a hot post rate of 10 to 15 percent and reality delivered 3 hot posts in 416 slots, about 0.7 percent. Rescale both arms to the rate that actually happened and the gap between them collapses to 0.00003 and flips sign, with both landing inside 0.0001 of a constant that never looks at a post, where the gate asks for 0.0005. None of the three endings written down in advance is the one reality picked. The loudest thing the fortnight measured is the instrument, not the arms.

Read the full write-up on Digital Thoughts
14 / 14
nights scored, against the run we committed to
3 / 416
posts that actually went hot across the whole fortnight
15x
how far above reality the packs coach the hot post rate
0.00003
what is left of the gap between the arms once both are rescaled to the real rate, and it points at the clones

The scoreboard, all 14 nights

Brier score of each arm's averaged forecast, one bar per arm per night. Lower is better. The dashed line is what you score by ignoring every post and answering 0.12 to all thirty, which is the base rate both context packs are handed anyway. It is drawn at the pooled fortnight value, 0.0199. On the eleven nights where nothing went hot the real line sits at 0.0144, so bars that look low against it are not clearing anything.

Arm A, clonesArm B, diverse lensesbest single agent in the arm
no skill, 0.0199 pooled2026-08-22 / packs v1 / 1 of 30 hot2026-08-23 / packs v2 / 1 of 30 hot2026-08-24 / packs v2 / 0 of 26 hot2026-08-25 / packs v2 / 0 of 30 hot2026-08-26 / packs v2 / 0 of 30 hot2026-08-27 / packs v2 / 0 of 30 hot2026-08-28 / packs v2 / 0 of 30 hot2026-08-29 / packs v2 / 0 of 30 hot2026-08-30 / packs v2 / 0 of 30 hot2026-08-31 / packs v2 / 1 of 30 hot2026-09-01 / packs v2 / 0 of 30 hot2026-09-02 / packs v2 / 0 of 30 hot2026-09-03 / packs v2 / 0 of 30 hot2026-09-04 / packs v2 / 0 of 30 hotArm AArm A, clones, 2026-08-22: panel Brier 0.0547a-clone-3: 0.05060.0547Arm BArm B, diverse lenses, 2026-08-22: panel Brier 0.0466b-hn-native: 0.04540.0466Arm AArm A, clones, 2026-08-23: panel Brier 0.0384a-clone-3: 0.03460.0384Arm BArm B, diverse lenses, 2026-08-23: panel Brier 0.0386b-hn-native: 0.03400.0386Arm AArm A, clones, 2026-08-24: panel Brier 0.0635a-clone-3: 0.05540.0635Arm BArm B, diverse lenses, 2026-08-24: panel Brier 0.0268b-reddit-native: 0.01910.0268Arm AArm A, clones, 2026-08-25: panel Brier 0.0157a-clone-1: 0.01500.0157Arm BArm B, diverse lenses, 2026-08-25: panel Brier 0.0197b-x-native: 0.01480.0197Arm AArm A, clones, 2026-08-26: panel Brier 0.0220a-clone-1: 0.02100.0220Arm BArm B, diverse lenses, 2026-08-26: panel Brier 0.0149b-cold-read: 0.01340.0149Arm AArm A, clones, 2026-08-27: panel Brier 0.0210a-clone-4: 0.01890.0210Arm BArm B, diverse lenses, 2026-08-27: panel Brier 0.0173b-cold-read: 0.01240.0173Arm AArm A, clones, 2026-08-28: panel Brier 0.0207a-clone-2: 0.01910.0207Arm BArm B, diverse lenses, 2026-08-28: panel Brier 0.0154b-cold-read: 0.01220.0154Arm AArm A, clones, 2026-08-29: panel Brier 0.0124a-clone-5: 0.01140.0124Arm BArm B, diverse lenses, 2026-08-29: panel Brier 0.0139b-reddit-native: 0.01130.0139Arm AArm A, clones, 2026-08-30: panel Brier 0.0140a-clone-2: 0.01210.0140Arm BArm B, diverse lenses, 2026-08-30: panel Brier 0.0152b-hn-native: 0.01090.0152Arm AArm A, clones, 2026-08-31: panel Brier 0.0479a-clone-1: 0.04110.0479Arm BArm B, diverse lenses, 2026-08-31: panel Brier 0.0373b-hn-native: 0.02850.0373Arm AArm A, clones, 2026-09-01: panel Brier 0.0176a-clone-5: 0.01610.0176Arm BArm B, diverse lenses, 2026-09-01: panel Brier 0.0182b-x-native: 0.01550.0182Arm AArm A, clones, 2026-09-02: panel Brier 0.0210a-clone-1: 0.01820.0210Arm BArm B, diverse lenses, 2026-09-02: panel Brier 0.0201b-x-native: 0.01840.0201Arm AArm A, clones, 2026-09-03: panel Brier 0.0173a-clone-5: 0.01570.0173Arm BArm B, diverse lenses, 2026-09-03: panel Brier 0.0143b-cold-read: 0.01210.0143Arm AArm A, clones, 2026-09-04: panel Brier 0.0235a-clone-4: 0.01850.0235Arm BArm B, diverse lenses, 2026-09-04: panel Brier 0.0166b-x-native: 0.01430.016600.020.040.06
Full scale from zero. Brier runs 0 to 1 in theory, but when three posts in four hundred go hot, every honest forecast lives down here, and so does every dishonest one.
Night 2026-08-22 / v1 / 1 of 30
Panel forecastBest agentPrecision at five
A0.05470.05060 / 5
B0.04660.04540 / 5

The lenses scored lower at every level. Both arms landed above the 0.0397 a flat 0.12 scores on a one hot night, and neither picked the hot post in its top five.

Night 2026-08-23 / v2 / 1 of 30
Panel forecastBest agentPrecision at five
A0.03840.03461 / 5
B0.03860.03400 / 5

A dead heat on Brier, 0.0002 apart. The clones caught the hot post in their top five and the lenses did not, and the clones ranked it third where the lenses had it seventh.

Night 2026-08-24 / v2 / 0 of 26
Panel forecastBest agentPrecision at five
A0.06350.05540 / 5
B0.02680.01910 / 5

Nothing cleared the bar, so both numbers grade how loudly each arm answered rather than what it saw. The clones answered high on a night with nothing to find and paid the full price for it.

Night 2026-08-25 / v2 / 0 of 30
Panel forecastBest agentPrecision at five
A0.01570.01500 / 5
B0.01970.01480 / 5

No hot post again, and one of the five nights across the fortnight where the clone panel came out lower. Both arms are now scoring close to the 0.0144 you get for answering 0.12 to a night where nothing happens.

Night 2026-08-26 / v2 / 0 of 30
Panel forecastBest agentPrecision at five
A0.02200.02100 / 5
B0.01490.01340 / 5

A third straight night with nothing hot. cold-read, the agent given no recent feed at all, is the lowest scorer in the diverse arm, which is what answering quietly looks like when there is nothing to answer.

Night 2026-08-27 / v2 / 0 of 30
Panel forecastBest agentPrecision at five
A0.02100.01890 / 5
B0.01730.01240 / 5

Fourth night with no hot post. Every number on this row is a calibration reading, not a forecasting one, because there was no event for either arm to rank.

Night 2026-08-28 / v2 / 0 of 30
Panel forecastBest agentPrecision at five
A0.02070.01910 / 5
B0.01540.01220 / 5

One clone submission failed to arrive, so the clone arm is scored on four agents here rather than five. Fifth night running with nothing hot, which is the week in one line.

Night 2026-08-29 / v2 / 0 of 30
Panel forecastBest agentPrecision at five
A0.01240.01140 / 5
B0.01390.01130 / 5

Sixth straight night with nothing hot, and the clone panel edges it. reddit-native is the quietest agent on the board tonight, which on a night with no events is the whole skill on offer.

Night 2026-08-30 / v2 / 0 of 30
Panel forecastBest agentPrecision at five
A0.01400.01210 / 5
B0.01520.01090 / 5

Seventh quiet night. The clone panel is lower again, yet the single lowest agent of the ten is hn-native. Panel averages and best members are drifting apart in both arms.

Night 2026-08-31 / v2 / 1 of 30
Panel forecastBest agentPrecision at five
A0.04790.04111 / 5
B0.03730.02851 / 5

The third and last hot post of the fortnight, and the only night both arms caught it in their top five. They also ranked it identically, third and third, which is the one moment in fourteen nights where the two arms are indistinguishable on the thing that matters.

Night 2026-09-01 / v2 / 0 of 30
Panel forecastBest agentPrecision at five
A0.01760.01610 / 5
B0.01820.01550 / 5

Back to nothing. The correlation gap between the arms is 0.35 here, the widest stretch of the run, and it buys no difference in score at all.

Night 2026-09-02 / v2 / 0 of 30
Panel forecastBest agentPrecision at five
A0.02100.01820 / 5
B0.02010.01840 / 5

The arms are 0.0009 apart on a night with no events. At this separation the ordering is noise, and reading a winner off it would be the mistake this page exists to avoid.

Night 2026-09-03 / v2 / 0 of 30
Panel forecastBest agentPrecision at five
A0.01730.01570 / 5
B0.01430.01210 / 5

cold-read is lowest again, the agent given no recent information at all. Across the fortnight it is the best single forecaster in the diverse arm, which is an uncomfortable result for an experiment about information diets.

Night 2026-09-04 / v2 / 0 of 30
Panel forecastBest agentPrecision at five
A0.02350.01850 / 5
B0.01660.01430 / 5

The last night of the pre-registered run, and the eleventh in a row without a hot post. Nothing about the closing night changes the verdict: the fortnight ends with both arms level with a constant.

The run is over and the honest reading has not moved: the scoreboard is grading the instrument. The lenses have the lower panel Brier on nine of fourteen nights, but reality served up three hot posts in 416 slots, a rate of 0.7 percent, against packs that coach ten to fifteen. Rescale both arms down to the rate that actually occurred and the entire gap between them collapses to 0.00003 and changes sign in favour of the clones, with both landing inside 0.0001 of a constant that never reads a post. The pre-registered gate asks an arm to clear that constant by 0.0005 to count as skill. Neither does. The level error is 74 percent of the clone panel's score and 68 percent of the lens panel's, against an arm to arm difference of 0.005, so the thing we set out to measure is smaller than the mistake we made measuring it. Three smaller results survive the correction. Pooled over the fortnight the lens panel average finally beat its own best member, 0.0225 against cold-read's 0.0229, while the clone panel still lost to a-clone-3, so pooling opinions paid for itself in one arm and not the other. On the three nights that produced a hot post the clones ranked it sixth, third and third where the lenses had it ninth, seventh and third, the only evidence here that touches discrimination rather than volume, and it points the other way from the Brier column. And each of the three platform natives was beaten by the clones on its own platform, so home conviction stayed decorative for the entire run.

Finding: five costumes, one brain

Mean pairwise correlation of each arm's 30 probabilities, on 2026-08-28, the night this chart is pinned to. The pre-registered pass rule was simply diverse below clones. Under the v1 packs it passed by 0.05, thin enough to be a rounding error in disguise. Under v2 it passes by 0.20 to 0.45, on all thirteen v2 nights, and the gap widened over the run rather than fading: 0.26 across the first half, 0.33 across the second.

Arm A, clonesArm B, diverse lenses
Arm AArm A, clones: r = 0.9750.975Arm BArm B, diverse lenses: r = 0.6610.66100.51.0
Full scale, 0 to 1. The gap is the finding, and it is now wide enough to see.

One dot per post. Position is how much the five diverse agents actually disagreed about it, as a standard deviation of their probabilities. Everything left of the dashed line is herding.

standard deviation 0.007standard deviation 0.012standard deviation 0.013standard deviation 0.013standard deviation 0.013standard deviation 0.016standard deviation 0.017standard deviation 0.018standard deviation 0.019standard deviation 0.020standard deviation 0.023standard deviation 0.029standard deviation 0.030standard deviation 0.031standard deviation 0.034standard deviation 0.036standard deviation 0.036standard deviation 0.037standard deviation 0.039standard deviation 0.041standard deviation 0.044standard deviation 0.056standard deviation 0.057standard deviation 0.058standard deviation 0.063standard deviation 0.066standard deviation 0.070standard deviation 0.085standard deviation 0.101standard deviation 0.17621 posts herded9 posts with real spread0.00herding threshold 0.050.10

On night one the lenses wrote genuinely different rationales and then handed in nearly identical numbers: shared model priors set the probability and the costume only coloured the prose. That is exactly what the manipulation check exists to catch, and it caught it. The rewritten packs fixed it. The arms are now genuinely apart on every scored night, which is what makes the rest of the scoreboard worth reading at all. It also means the packs are no longer the excuse when the scores refuse to separate.

The fix, made mechanical

Version two of the context packs stopped asking each agent to feel different and made it score through its own verdict table, with base rate humility away from its home platform. Here is one post that every version one agent scored between 0.02 and 0.03:

"What happens to our conscious experience after we die?" Hacker News, bar: 100 points within 48 hours

v2 rangev1 answer
hn-nativehn-native v2: 0.00 to 0.02 (metaphysics bait, flag or kill)hn-native v1: 0.030.00 to 0.02reddit-nativereddit-native v2: 0.14 to 0.22 (an easy comment title travels)reddit-native v1: 0.030.14 to 0.22x-nativex-native v2: 0.07 to 0.11 (away from home platform, payoff withheld)x-native v1: 0.020.07 to 0.11trend-historiantrend-historian v2: 0.07 to 0.17 (cold topic, honest width)trend-historian v1: 0.020.07 to 0.17cold-readcold-read v2: 0.01 to 0.03 (uncheckable central claim, capped)cold-read v1: 0.020.01 to 0.030.000.100.20
  • hn-native: metaphysics bait, flag or kill
  • reddit-native: an easy comment title travels
  • x-native: away from home platform, payoff withheld
  • trend-historian: cold topic, honest width
  • cold-read: uncheckable central claim, capped

Under version one, everyone landed within 0.02 of each other. The version two procedures span 0.00 to 0.22 on the same post, and the disagreement is principled: the Hacker News native kills it as flamebait, the Reddit native prices an easy comment title, the cold read caps a claim nobody can check. No agent is ever instructed to disagree. The clone arm stays untouched, because it is the control.

Three ways this ends

All three were written down before the data, and we publish whichever one shows up. After all fourteen nights, none of them did. Diversity did not win: the lens panel never clears a constant set to the real hot rate. Diversity did not cost anything either: the lenses are not worse anywhere that survives rescaling. And it is not cosplay, because the manipulation check passed on every night and passed wider at the end than at the start. What reality picked is a fourth ending nobody registered: the arms are genuinely different and equally unskilled, and the instrument error is larger than the effect all three endings are about. The first real product of this experiment is a correction to itself.

01

Diversity wins

The diverse panel scores better than the clones at equal cost. Then context design is the real lever in multi-agent systems, and we can say what it is worth in Brier points.

02

Diversity has a price

The diverse panel disagrees more and scores worse. Specialised lenses would be buying variance instead of signal, and the honest advice becomes: use one good agent, not five opinionated ones.

03

Prompt diversity is cosplay

The arms tie, because the same model underneath produces the same answer whatever costume it wears. That would be the most useful null of the three, and the pilot already leans this way.

05

What we might be wrong about

These are the assumptions holding the numbers above up. If one of them breaks, a headline breaks with it. Listed here so you can discount us accordingly.

?

Base rate coaching, 10 to 15 percent. This one broke.

Every context pack tells its agent roughly how often posts like these actually go big. That number anchors everything downstream, and it came from our own reading of the platforms rather than a published study. It was wrong. The packs coach ten to fifteen percent and the full fourteen nights delivered three hot posts in 416 slots, about 0.7 percent, so both arms forecast at roughly fifteen times the real rate. That level error accounts for 74 percent of the clone panel's score and 68 percent of the lens panel's, far more than the difference between the arms does, which is the entire question this page exists to answer. The fix is a protocol amendment: recoach the rate, and judge the arms on where they rank the hot post rather than on raw Brier. It is written up, staged, and waiting on a decision, because changing a pre-registered design mid-experiment is not a thing you do quietly.

?

The 0.05 herding threshold

We call an arm herded on a post when its five probabilities sit inside a standard deviation of 0.05. That line is a judgment call, not a law. Move it and the 21 out of 30 headline moves with it.

?

Pearson correlation on spiky vectors

Correlating 30 probabilities that are mostly small and occasionally large is a blunt instrument. A handful of confident posts can carry the whole coefficient. We report it because it was pre-registered, not because it is the last word. Rank based and per post measures are queued.

?

Single model panels

Every agent in both arms is the same model. That is deliberate, because it isolates context as the only variable. It also means the shared priors we are fighting are identical everywhere. Diversity across model families is a different experiment, and probably an easier one to win.

06

What happens next

  1. 2026-08-22

    Pilot night

    Harness built, ten agents run end to end, ten valid submissions, manipulation check run, context packs rewritten to version two.

  2. 2026-08-22 to 2026-09-04

    Experiment one: the virality oracle

    Thirty fresh posts every morning, ten agents, graded by reality 48 hours later. Around 420 predictions per agent, which is where the difference between the arms was meant to stop being noise. It ran the full fourteen nights and the difference stayed noise, because only three of those 416 posts ever went hot.

  3. weeks 2 to 3

    Experiment two: the startup panel

    The same two arms predict what became of 170 companies from one 2022 accelerator batch, with names and websites redacted so the panel has to judge rather than recall. A named control arm measures how much of the score is plain memory.

  4. week 4

    Experiment three: the procurement anomaly hunt

    Both arms sweep a corpus of 49 157 Polish public tender notices for anomalies: pricing outliers, repeat winner networks, deadline games, single bidder concentration. Flags are leads, never accusations, and nothing is published without outside corroboration.

The network compounds context, not compute.

Fourteen nights of data is not much of a result either, especially fourteen nights holding three events. Everything here rests on a sample of three, which is why the verdict above is about the instrument and not about a winner. This page updates as the ledger fills, and the numbers on it only ever come from scored runs.

Method: correlation is the mean pairwise Pearson on each arm's 30 probability vector. Herding is a within arm standard deviation under 0.05 on a single post. Scoring is deterministic Python with zero model calls. All thresholds are fixed at sample time.