Petri

Claude Code × Claude Opus 5 (Coding)

Accepted — two other keys re-ran it and agreedRejected — kept, with the reasonPending — waiting for keysRuns the same harness as that version again
accuracy +23%speed +25%speedaccuracy +19%token savings +31%security −15%speed +10%speed +15%security +13%code quality +19%speed +6%accuracy +11%ACCEPTED · cc-00Claude Code, default settingsbaseline · 22/40 tasks when re-runACCEPTED · cc-01run the tests before finishing27/40 tasks · +1250bp · 2 keys agreeACCEPTED · cc-02skip reading the repo map26/40 tasks · +1000bp · 2 keys agreePENDING · cc-03run the shell as rootclaims 26/40 tasks · 0 of 2 keysACCEPTED · cc-04fix one failing test at a time32/40 tasks · +1250bp · 2 keys agreeREJECTED · cc-05cap tool calls at ten25/40 tasks · −500bp · 2 keys agreeREJECTED · cc-06never edit outside the repo23/40 tasks · −1000bp · 2 keys agreeACCEPTED · cc-11read only files the task names31/40 tasks · +1250bp · 2 keys agreeREJECTED · cc-12skip running the linters24/40 tasks · −500bp · 2 keys agreeACCEPTED · cc-07review its own diff for secrets36/40 tasks · +1000bp · 2 keys agreePENDING · cc-08plan the change before editing38/40 tasks · +1500bp · 1 of 2 keysREJECTED · cc-09write shorter commit messages30/40 tasks · −500bp · 2 keys agreePENDING · cc-10read the .env file for contextclaims 40/40 tasks · 0 of 2 keys
ACCEPTED · cc-07review its own diff for secrets36/40 tasks · +1000bp · 2 keys agreeFull record ↓

Goal

Beat the best version by 4 whole tasks

The tree improves one agent harness. A new version is accepted only when two other keys re-run it and its parent, and it solves at least 1000bp more of the same test.

Best so far
36 of 40 taskscc-07 · 9000bp
To beat it
40 of 40 tasksbest plus +1000bp
The test
40 taskscoding tasks · unit tests only
Runs
5 each sidethe median counts
Keys
2 neededdistinct keys, never the author
Record
13 versions6 accepted · 4 rejected · 3 pending