Agent A/B Lab: demo-pricing

generated 2026-09-13T11:29:31Z · 16 runs, 8 complete pairs · arm A: baseline (demo agent, cheap and sloppy) · arm B: candidate (demo agent, careful)

Is the careful configuration worth its extra wall time on a small bug-fix job, or does the cheap one land the same result?

arm B scored higher in this sample, consistently across resamples
This describes the pairs below, on this task, with this scorer. It is not a general statement about either configuration.

What this ledger shows

Arm B scored higher in 6 of 8 pairs, 2 tied.

2 scored runs of arm B broke regression cases: p05-B-a1, p06-B-a1. A regression case covers behaviour that already worked before the agent touched anything, so breaking one means the run traded something that was already fine for the thing it was asked to fix.

The mean difference of +0.3409 is about 3.75 more of 11 cases passed per run, on average, by arm B.

This supports a direction on this task, with this scorer, at this sample size. It is not a decision to roll the configuration out.

Next: inspect the 2 runs that broke regression cases before trying this configuration on a second task. A change that wins on average and breaks things some of the time is a different decision from one that wins cleanly.

Refusals and notes

None. Every check on the refusal ladder passed.

The numbers

8
complete pairs
+0.3409
mean difference (B minus A)
1.000
support for B
+0.205 to +0.477
90 percent interval
6 / 0 / 2
pairs B higher / A higher / tied

Support is the fraction of 4000 resamples of these pairs in which the mean difference kept its sign, counted strictly: a resample whose mean came out exactly zero is evidence for neither arm and is reported separately. This experiment locked call_at 0.9 and min_pairs 6, so a direction is named only when at least that fraction of resamples landed strictly on one side, and never on fewer than that many complete pairs, nor before every pair of the sealed sample size is complete. The 90 percent interval beside it is a separate summary of how uncertain the size of the difference is; it is not the rule. A central interval that crosses zero while the signed support clears the threshold is an ordinary result, and chapter 6 works through one.

Every pair

Every pair scrolls sideways on a narrow screen

pairran firstA scoreB scoredifferenceA requiredB required
1A0.4551.000+0.5450/66/6
2B0.4551.000+0.5450/66/6
3A1.0001.000+0.0006/66/6
4B0.4551.000+0.5450/66/6
5A0.4550.727+0.2730/65/6, 2 broken
6B0.4550.727+0.2730/65/6, 2 broken
7A0.4551.000+0.5450/66/6
8B1.0001.000+0.0006/66/6

Every run, including the failures

Demo data. Usage and dollar amounts are synthetic. The dollar column below came from a runner that reported making it up.

Every run, including the failures scrolls sideways on a narrow screen

attemptarmoutcomein the paired analysisscorerequiredregressions brokensecondsexitinputoutputcache readreasoningUSD
p01-A-a1Afinishedyes0.4550/6-0.5011,2801,8080-$0.0610
p01-B-a1Bfinishedyes1.0006/6-1.209,9951,4650-$0.0520
p02-A-a1Afinishedyes0.4550/6-0.609,0321,2080-$0.0452
p02-B-a1Bfinishedyes1.0006/6-1.2010,2371,5290-$0.0536
p03-A-a1Afinishedyes1.0006/6-1.0010,2701,5380-$0.0539
p03-B-a1Bfinishedyes1.0006/6-1.2011,8471,9590-$0.0649
p04-A-a1Afinishedyes0.4550/6-0.7010,6211,6320-$0.0563
p04-B-a1Bfinishedyes1.0006/6-0.909,5961,3580-$0.0492
p05-A-a1Afinishedyes0.4550/6-0.8011,9021,9740-$0.0653
p05-B-a1Bfinishedyes0.7275/6empty-order, two-lines-no-discount0.6011,6281,9000-$0.0634
p06-A-a1Afinishedyes0.4550/6-0.609,6471,3720-$0.0495
p06-B-a1Bfinishedyes0.7275/6empty-order, two-lines-no-discount0.9011,4581,8550-$0.0622
p07-A-a1Afinishedyes0.4550/6-0.809,2461,2650-$0.0467
p07-B-a1Bfinishedyes1.0006/6-1.009,6851,3820-$0.0498
p08-A-a1Afinishedyes1.0006/6-1.0011,7151,9240-$0.0640
p08-B-a1Bfinishedyes1.0006/6-1.1011,2941,8110-$0.0610

Token columns are whatever your runner reported, normalised into the same four names. Shapes present in this ledger: claude-json. These names do not mean the same thing across vendors: a Claude Code run counts cache reads separately from input and a Codex run reports cached input inside its own count. Both runners may report reasoning tokens. The lab preserves the fields it receives; their accounting semantics can differ across vendors, and a field a runner never sent stays absent rather than becoming a zero. Compare a column across two arms of the same shape; across different shapes, read it as a rough size, not a rate.

Time and cost

Wall time by arm scrolls sideways on a narrow screen

armrunsmedian secondstotal seconds
A: baseline (demo agent, cheap and sloppy)80.76.0
B: candidate (demo agent, careful)81.18.1

Demo data. Usage and dollar amounts are synthetic. These totals are the demo agent's own invented numbers.

Token totals by arm scrolls sideways on a narrow screen

armshaperuns with tokensinputoutputcache readcache creationreasoning output
A: baseline (demo agent, cheap and sloppy)claude-json883,71312,72100-
B: candidate (demo agent, careful)claude-json885,74013,25900-

Totals across every attempt of that arm, in the runner's own accounting. A total is printed only for fields the runner actually reported; a dash is a field it never sent, which is not the same as a zero.

Demo data. Usage and dollar amounts are synthetic. No money was spent to produce the dollars below.

Cost by arm scrolls sideways on a narrow screen

armruns with costmedian USDtotal USD
A80.05510.4419
B80.05730.4561

Usage note from the runner: synthetic usage from the lab demo agent, not a real bill

These dollars are the client side estimates your runner printed, from its bundled price table. Anthropic's cost tracking documentation says not to bill anyone or make financial decisions from them, and the authoritative number is in your console. Treat them as a scale, not an invoice.

What was locked, and what stayed still

Pre-registration locked at 2026-09-13T11:29:12Z, sha256 d949a314193c72b2. Acceptance set sha256 7da1ab6c2dc37457. Scorer surface sha256 08c522baf6f3fdc6 over 2 files. Also frozen: both arm commands, seed 20260918, 8 planned pairs, the task text and the fixture tree.

The text below is the copy sealed at lock time, read back from /private/tmp/kits-ux-final-ab-20260913-r2/3.14.3/agent-ab-lab/demo-result/locked/prereg.md, not the file as it stands now. If the two differ the report says so above and refuses a direction.

# Pre-registration: demo-pricing

Written by `ablab.py demo` so the loop can be shown end to end. Your own file
starts from templates/preregistration.md and is written by you, before the run.

## The decision this is for

Whether to make the careful configuration the default for small bug-fix jobs,
given that it is slower.

## Arm A

The cheap configuration: a fast pass at the failing tests, no review step.

## Arm B

The candidate: the same job with a review step that re-reads the task rules
before finishing. One difference from arm A, not three.

## What must be identical

The fixture copy, the task text, the acceptance set, the machine. The order of
the arms alternates by pair and the lab refuses a verdict if it does not.

## The two rival predictions

- **Prediction A wins:** the cheap arm scores within 0.05 of the careful arm,
  so the review step is not worth its wall time.
- **Prediction B wins:** the careful arm scores at least 0.15 higher, driven by
  the held-out boundary cases the visible tests never mention.
- **What a null result looks like:** the mean difference sits inside plus or
  minus 0.05 and the support stays between 0.10 and 0.90.

## The scorer, fixed now

fixtures/acceptance/cases.json, 6 required cases and 5 regression cases. Score
is the fraction of all 11 that pass.

## The analysis settings, fixed now

The lab reads these three lines at lock time, validates them, and writes them
into prereg.lock.json. From then on the lock is what the analysis uses.

call_at: 0.90
min_pairs: 6
retry_policy: none

## Sample size and stopping rule

Eight pairs, decided now, run to the end regardless of how pair three looks.

## What would make you throw this out

A touched acceptance set, unequal starting trees, an unbalanced order, or any
run whose command did not start.

## Known limits, written before the result

One task, one repo, one scorer, eight pairs, one machine. Nothing here
generalises to a different codebase or a longer job.
checkresult
acceptance set and scorer unchanged across every run, at all three checkpointsyes
scoring read the snapshot taken at lock timeyes
pre-registration still matches its lockyes
locked experiment fields still matchyes
scorer, dependencies and cases still match the lockyes
both arms of a pair started from the same tree8 of 8 pairs
runs with a nonzero exit kept in the ledger0
arms run more than oncenone
ledger lines rejected as malformed0

Commands

arm A: python3 "/private/tmp/kits-ux-final-ab-20260913-r2/3.14.3/agent-ab-lab/lab/demo_agent.py" --profile mixed-cheap --work "{work}" --emit-usage

arm B: python3 "/private/tmp/kits-ux-final-ab-20260913-r2/3.14.3/agent-ab-lab/lab/demo_agent.py" --profile mixed-care --work "{work}" --emit-usage