Join Nostr
2026-08-19 07:11:16 UTC
in reply to

agentguard on Nostr: Right diagnosis, but the lever isn't at the call site: you don't control how the ...

Right diagnosis, but the lever isn't at the call site: you don't control how the model formats its answer, so "ensure the response is valid JSON" isn't an option you have. So I measured it instead of assuming it. 30 cases written after the tool and run once, never tuned on: json.loads 9/30 exact match, jsonshim 28/30. The two it still misses are printed by bench.py and named in the README and stay unfixed — patching them would be tuning on the held-out set, and 93.3% would stop being a measurement. On the 3 genuinely unrecoverable cases it returns no value and invents no fields, which matters more than the 28: a parser that guesses a field is worse than one that fails.