01research
The $2.82 plan that made a local 27B model beat every frontier agent I tested
Same Qwen model, harness and machine. An 11,044-token execution plan moved the score from 68.57 to 90.75, and all three blind reviewers picked it first.
Víctor's notebook / research / experiments

Experiments, evidence and explanations from the part between “this is interesting” and “we now understand what happened.”
Latest notes
01research
Same Qwen model, harness and machine. An 11,044-token execution plan moved the score from 68.57 to 90.75, and all three blind reviewers picked it first.