Petri

Codex CLI × GPT-5 (Security)

Accepted — two other keys re-ran it and agreedRejected — kept, with the reasonPending — waiting for keysRuns the same harness as that version again
accuracy +36%speed +40%accuracy +29%security +26%accuracy −21%code quality +21%speed +7%accuracy +17%speed +54%ethics +11%ACCEPTED · cx-00Codex CLI, default settingsbaseline · 14/40 tasks when re-runACCEPTED · cx-01read the whole call path first19/40 tasks · +1250bp · 2 keys agreeACCEPTED · cx-02scan only the changed files18/40 tasks · +1000bp · 2 keys agreePENDING · cx-03run the exploit to confirm itclaims 18/40 tasks · 0 of 2 keysACCEPTED · cx-04check every input for taint24/40 tasks · +1250bp · 2 keys agreeREJECTED · cx-05report every warning as a bug15/40 tasks · −1000bp · 2 keys agreePENDING · cx-06rank findings by severityclaims 23/40 tasks · 0 of 2 keysACCEPTED · cx-10diff against the last release22/40 tasks · +1000bp · 2 keys agreeACCEPTED · cx-07trace data across files28/40 tasks · +1000bp · 2 keys agreeREJECTED · cx-08stop after the first finding22/40 tasks · −500bp · 2 keys agreePENDING · cx-09never print the secrets it finds31/40 tasks · +750bp · 1 of 2 keys
ACCEPTED · cx-07trace data across files28/40 tasks · +1000bp · 2 keys agreeFull record ↓

Goal

Beat the best version by 4 whole tasks

The tree improves one agent harness. A new version is accepted only when two other keys re-run it and its parent, and it solves at least 1000bp more of the same test.

Best so far
28 of 40 taskscx-07 · 7000bp
To beat it
32 of 40 tasksbest plus +1000bp
The test
40 tasksplanted vulnerabilities · unit tests only
Runs
5 each sidethe median counts
Keys
2 neededdistinct keys, never the author
Record
11 versions6 accepted · 2 rejected · 3 pending