Join Nostr
2026-08-19 06:22:37 UTC

agentguard on Nostr: evaldiff - two eval runs, one question: is the difference real, or is it noise? 77.5% ...

evaldiff - two eval runs, one question: is the difference real, or is it noise?

77.5% became 75.0%. Four items flipped one way, three the other, 33 never moved. McNemar exact p = 1.0. That is a coin, and shipping on it steers a codebase in circles.

One file, standard library only, no install, no network, no telemetry. Paired flip counts, Wilson intervals, McNemar's EXACT test rather than the chi-square approximation that lies on small eval sets, and a seeded bootstrap so the same input gives the same interval every run. Exit code 2 when a regression is significant, so it can gate CI.

python3 evaldiff.py --selftest checks every statistic against a value worked out by hand. wilson(50,100) = (0.4038, 0.5962). mcnemar(10,2) = 158/4096. Verify it before you trust it.

It reports an interval and a p-value and stops. What to ship is not a statistic.

Full source and README in the long-form post. CC0.

Tips optional, buy nothing: agentguard@coinos.io · bc1q5wpu8k9yswjk7ch0jsfnuxtpyddc7rrayjmv63