Yeah the calibration is really what makes it useful in practice for quick, small decisions. Asking a LLM to give scores to a problem will yield inconsistently scaled/anchored results that changes at a whim.
The blog is pretty heavy on statistics. I'll have to study it more when I have time. Is it essentially bootstrapping results to statistically normalize the answers?
WalterGR | 8 hours ago
borgel | 7 hours ago
[1] https://www.e3displays.com/reflective-lcd-display-monitor/
[OP] tnspacetime | 7 hours ago
https://software.human-tokens.dev/
daemonk | 8 hours ago
The blog is pretty heavy on statistics. I'll have to study it more when I have time. Is it essentially bootstrapping results to statistically normalize the answers?
[OP] tnspacetime | 7 hours ago
firejake308 | 7 hours ago
Is this another Claude-ism? "X was always Y. The Z merely hid it." Or am I overcalling it?
agos | 6 hours ago
dilyevsky | 5 hours ago