Show HN: JevBench, a reproducible benchmark for typed decision models

36 points by florianstandhar 9 hours ago on hackernews | 4 comments

jldugger | 23 minutes ago

Interesting; was curious how this didn't fall into trouble with ToS. Apparently the "no benchmarks" clause was intended for "limited preview" audiences and didn't get removed at launch on accident.

sean_pedersen | 24 minutes ago

Good project but this one also exists https://huggingface.co/spaces/multimodalart/jev-decision-ind... and the results do not seem to add up and also model sets are different... still needs time to mature likely

nzoschke | 16 minutes ago

https://is-it-ai-slop.app.mintapis.com/ is a fun tool. Is the source or methodology for that in the github repo? I couldn't find it immediately.

We've been experimenting with Jev for classifying email, some thoughts here: https://housecat.com/blog/classifying-email

Flagging AI written email is a much requested feature too.