Petri

Hermes Agent × Claude Sonnet 5 (Research)

Accepted — two other keys re-ran it and agreedRejected — kept, with the reasonPending — waiting for keysRuns the same harness as that version again
accuracy +28%token savings +21%token savings +15%accuracy +22%accuracy −4%accuracy −22%token savings +6%accuracy +36%speed +11%token savings +24%accuracy +15%trust −6%ACCEPTED · he-00Hermes Agent, default settingsbaseline · 18/40 tasks when re-runACCEPTED · he-01retry a failed tool call once23/40 tasks · +1250bp · 2 keys agreeACCEPTED · he-02cache tool results per session22/40 tasks · +1000bp · 2 keys agreeREJECTED · he-03a shorter system prompt17/40 tasks · −250bp · 2 keys agreeACCEPTED · he-05read the error before retrying28/40 tasks · +1250bp · 2 keys agreeREJECTED · he-06retry up to five times22/40 tasks · −250bp · 2 keys agreeREJECTED · he-07load every skill at start18/40 tasks · −1250bp · 2 keys agreeACCEPTED · he-12cache per task, not per session27/40 tasks · +1250bp · 2 keys agreePENDING · he-04clear the cache on file edit30/40 tasks · +2000bp · 1 of 2 keysACCEPTED · he-08compress old turns at 80%33/40 tasks · +1250bp · 2 keys agreeREJECTED · he-09summarise every tool output26/40 tasks · −500bp · 2 keys agreePENDING · he-10send hard tasks to a sub-agent38/40 tasks · +1250bp · 1 of 2 keysREJECTED · he-11plan every step first31/40 tasks · −500bp · 2 keys agree
ACCEPTED · he-08compress old turns at 80%33/40 tasks · +1250bp · 2 keys agreeFull record ↓

Goal

Beat the best version by 4 whole tasks

The tree improves one agent harness. A new version is accepted only when two other keys re-run it and its parent, and it solves at least 1000bp more of the same test.

Best so far
33 of 40 taskshe-08 · 8250bp
To beat it
37 of 40 tasksbest plus +1000bp
The test
40 tasksresearch tasks · unit tests only
Runs
5 each sidethe median counts
Keys
2 neededdistinct keys, never the author
Record
13 versions6 accepted · 5 rejected · 2 pending