Great feedback. After all this time I just realized a lot of people may not even know that menu exists. It's been like that for so long I'm literally blind to it.
Just a few high level points about how this works. It works in multiple stages:
1. I apply a domain and keyword filter to the feed
2. The content of the remaining articles are run twice daily through a Modern Bert-based classifier fine-tuned to detect AI-related content (~8000 training examples)[0].
It also filters out Github repos that contain AI authorship. My backend scans:
- Commit messages for agent attribution
- The contributor graph for agents
- Repo files for instructions/configs
0: The original workflow for this was a little different. For the better part of the past year, I had an AI agent detect AI-related stories and raise ~20 stories to me to make a judgement call on. After a while, a workflow like this just doesn't make sense when small models can do it equally as well. The training data is based on the machine and my labeling.
I think you should make it clear, that when you say "without AI" you mean AI-related content, because some people in the comments seem to think that this is about AI-generated content.
I don't personally use those apps so I'm not sure, but if they're open source I'm confident you could easily update them to work with my site. In case it's useful, here's the source code https://github.com/leiDnedyA/hn-without-ai
Huh, perhaps there is a project to be done to consume the HN firehose and serve it AT Proto style from the firehose. HN mobile clients can then pick the feed they want. AI? No AI? Politics? No politics? Frameworks? No frameworks? Choose your own HN post feed adventure. Read replica esq consumption, writes (comments, posts, votes) go to the primary data store (HN).
At the bottom of the linked site, it says "79 of 179 submissions survived the filter" so if only ~44% survived I'd say that counts as "prevalent" yeah.
* It filters out things that are clearly not AI related (like the one titled ' Mind-altering drugs played key role in rise of Andean civilization ') - or maybe it's just slow to update (that post was 3 hours ago)
* No ability to go to the next page
* Order is often different than on HN, even account for the missing posts. It's mostly the same, but not quite, which always throws me for a loop to see if something else is missing.
People were actually complaining specifically about OpenAI/Anthropic glazing and the countless influencer comments and posts here every hour.
And it's getting twisted into "People want HN without AI."
But that was never the case. I mean some surely do. But the blow-up this morning was specific to Anthropic ads to the point there's basically no information on this website that isn't related to Anthropic or OpenAI.
If it was announced some rando kid made a LLM in PHP or something, sure share it! If it was announced someone actually made something cool with AI (lol) then share it!
But Claude 24/7 aint it. It's at least extremely boring!
There is also the hide button under every post, which I’ve started using generously. It only takes a few seconds to clean up the front page to your taste
I'm mixed on that. Definitely topics I hide when I think people can't discuss them like adults, but sometimes even bad stories have some great discussions in them.
Depends on what you want "without AI" to mean... is it just stories about AI, or anything that an AI has touched at all? Or purely vibe-coded projects?
Instead of filtering keywords this ranks the content using Pangram. But it's also pretty effective at filtering content about AI. Turns out that a lot of LLM tooling projects have LLM written project descriptions/blog posts.
It's amazing to see how many of the "AI hurts my brain" posts are apparently written by an LLM.
Great point. I think the key is differentiating between "prompt generated" content vs "AI as a Topic" content and then being able to toggle them at will.
I am not averse to prompt-generated content, but it should be marked and quoted as such, just like any other reference. If it's not, it's like the author is double-cheating.
In fact there are no shortage of comments here on HN, especially open-ended questions that could've been plugged into a bot for an answer. Avoiding LLMs altogether is foolhardy. It's a matter of balance, and being able to write a good prompt is a skill - no mental atrophy there. Content derived from non-English sources using English-based LLMs is also an interesting area. [1]
It's possible to obtain AIaaT from https://aibriefs.news/ and then prune it elsewhere if need be. But the best aggregators like that may be machine-generated with a lot of human moderation. I wonder if that is 100% machine generated.
Nice! I was considering doing this but the Pangram API is very expensive. Then I considered training my own model and I fortunately stopped at the edge of that rabbit hole.
Yea very expensive (hence me siphoning another page haha).
Im trying to get something cheaper to work, Pangram has some nice docs on how to build something like their service https://github.com/pangramlabs/EditLens, they even have training data online.
The harm a missclassification carries is much lower for a hn post than a masters thesis, so we might be fine with a worse model.
Yeah agree, Pangram puts out interesting material. I recently came across their v4 technical report and they shared a lot more than I would have expected them to.
But see, here I am back at the edge of the rabbit hole, and you're trying to pull me in. I refuse!
If you do end up training a model, send me an email and maybe I can tie it into the site.
Will do, thanks for hcker.news btw, has been my main mobile client for couple of months. Of all things i like the changelog most, its nice to have a quick view about new stuff, especially if the website changes somewhat faster than usual.
I quite like this, particularly that the filter is just a preset! Will be back.
Minor UI feedback - when navigating comments (and maybe other stuff) with the keyboard, the actual focus change is blocked behind some automated scroll behavior, and the effect is that someone quickly hitting keybinds will collapse the wrong thread. Would suggest making the focus change happen concurrently.
Thanks for the feedback. Comment drawer nav could be a lot better. I started working on the keyboard nav but got distracted with other things. I'll fix it.
It looks nice but I can't load any comments, and I believe it's because I have Firefox's Enhanced Tracking Protection turned on to the strictest level.
I know you guys hate X/Twitter but the muted words feature is really awesome for this kind of thing. You never have to hear about a topic you don't give a crap about ever again.
Would if be possible to set up push notifications for new stories matching a filter - one that's separate to the main filter? Like say, I want to get notified every time an article about "Linux" is posted. But if I go to main page, I want to see all stories (or whatever my main filter is set to).
Yeah, all possible, the infra for all of that is set up. I think I understand your request but can u send me an email so I can talk it through with you a little more?
loldidntwork | 12 hours ago
[OP] postalcoder | 12 hours ago
TouchGrass | 12 hours ago
Checking the connection Checking the proxy and the firewall ERR_CONNECTION_RESET
[OP] postalcoder | 12 hours ago
Gualdrapo | 12 hours ago
niyonx | 12 hours ago
chilipepperhott | 12 hours ago
[OP] postalcoder | 12 hours ago
vanschelven | 12 hours ago
basch | 12 hours ago
its incredibly easy to immediately tell when something was posted, and how popular it was. there s no automagic rearranging or movement.
similar to https://techmeme.com/river
[OP] postalcoder | 12 hours ago
basch | 10 hours ago
[OP] postalcoder | 10 hours ago
dorkrawk | 12 hours ago
[OP] postalcoder | 11 hours ago
ant6n | 10 hours ago
[OP] postalcoder | 12 hours ago
HN Frontpage minus AI: https://hcker.news/?view=frontpage&ai=exclude
You may also like Small Web HN: https://hcker.news/?view=frontpage&smallweb=include
0: The original workflow for this was a little different. For the better part of the past year, I had an AI agent detect AI-related stories and raise ~20 stories to me to make a judgement call on. After a while, a workflow like this just doesn't make sense when small models can do it equally as well. The training data is based on the machine and my labeling.
bananaflag | 11 hours ago
[OP] postalcoder | 11 hours ago
NeedNewForums | 11 hours ago
vova_hn2 | 7 hours ago
RIMR | 12 hours ago
account42 | 12 hours ago
amelius | 11 hours ago
busymom0 | 11 hours ago
https://imgur.com/a/P1qMIok
9 out of 30 posts have been removed.
Steps:
Install uBlock Origin Lite. Then create a custom filter:
Any improvements to above filter are welcome.layer8 | 12 hours ago
otherayden | 12 hours ago
layer8 | 11 hours ago
gonzalohm | 11 hours ago
otherayden | 11 hours ago
toomuchtodo | 11 hours ago
[OP] postalcoder | 7 hours ago
xvinci | 10 hours ago
otherayden | 6 hours ago
WarcrimeActual | 10 hours ago
buggy6257 | 8 hours ago
Night_Thastus | 6 hours ago
* It filters out things that are clearly not AI related (like the one titled ' Mind-altering drugs played key role in rise of Andean civilization ') - or maybe it's just slow to update (that post was 3 hours ago)
* No ability to go to the next page
* Order is often different than on HN, even account for the missing posts. It's mostly the same, but not quite, which always throws me for a loop to see if something else is missing.
visarga | 11 hours ago
Edit: ah yes, sure it exists
[OP] postalcoder | 11 hours ago
jrm4 | 11 hours ago
NeedNewForums | 11 hours ago
And it's getting twisted into "People want HN without AI."
But that was never the case. I mean some surely do. But the blow-up this morning was specific to Anthropic ads to the point there's basically no information on this website that isn't related to Anthropic or OpenAI.
If it was announced some rando kid made a LLM in PHP or something, sure share it! If it was announced someone actually made something cool with AI (lol) then share it!
But Claude 24/7 aint it. It's at least extremely boring!
jessetemp | 11 hours ago
datakan | 11 hours ago
gdulli | 11 hours ago
sejje | 11 hours ago
bronlund | 11 hours ago
armchairhacker | 11 hours ago
TIL it was rare in the actual wartime (WWII) and only became popular in 2000: https://en.wikipedia.org/wiki/Keep_Calm_and_Carry_On
Tepix | 11 hours ago
Looking at the iPod 6G QEMU emulation story:
> "With the help of Claude Code, I was able to create code to extract it using emCORE."
ranger_danger | 10 hours ago
nonsensical_ | 11 hours ago
ChrisArchitect | 11 hours ago
Stick to sharing your thing in the current thread today:
Ask HN: Can we please limit the AI news flood?
https://news.ycombinator.com/item?id=49657850
tcfhgj | 11 hours ago
nilsherzig | 11 hours ago
nilsherzig | 10 hours ago
[OP] postalcoder | 10 hours ago
nilsherzig | 11 hours ago
Instead of filtering keywords this ranks the content using Pangram. But it's also pretty effective at filtering content about AI. Turns out that a lot of LLM tooling projects have LLM written project descriptions/blog posts.
It's amazing to see how many of the "AI hurts my brain" posts are apparently written by an LLM.
nilsherzig | 11 hours ago
OPs solution does more than keyword filtering. But its still about filtering out "content about ai" not filtering out "ai written content".
adrianwaj | 3 hours ago
I am not averse to prompt-generated content, but it should be marked and quoted as such, just like any other reference. If it's not, it's like the author is double-cheating.
In fact there are no shortage of comments here on HN, especially open-ended questions that could've been plugged into a bot for an answer. Avoiding LLMs altogether is foolhardy. It's a matter of balance, and being able to write a good prompt is a skill - no mental atrophy there. Content derived from non-English sources using English-based LLMs is also an interesting area. [1]
It's possible to obtain AIaaT from https://aibriefs.news/ and then prune it elsewhere if need be. But the best aggregators like that may be machine-generated with a lot of human moderation. I wonder if that is 100% machine generated.
Likewise, HN-AI could well be a decent site.
[1] How much content can be derived from non-English sources using LLMs? https://share.gemini.google/ObjMQHCj8JWj"
[OP] postalcoder | 11 hours ago
nilsherzig | 11 hours ago
Im trying to get something cheaper to work, Pangram has some nice docs on how to build something like their service https://github.com/pangramlabs/EditLens, they even have training data online.
The harm a missclassification carries is much lower for a hn post than a masters thesis, so we might be fine with a worse model.
[OP] postalcoder | 10 hours ago
But see, here I am back at the edge of the rabbit hole, and you're trying to pull me in. I refuse!
If you do end up training a model, send me an email and maybe I can tie it into the site.
nilsherzig | 10 hours ago
[OP] postalcoder | 10 hours ago
verdverm | 11 hours ago
I suppose it depends on what point of view we answer from. It is very hacker-newsery to make something like this
mpalmer | 11 hours ago
Minor UI feedback - when navigating comments (and maybe other stuff) with the keyboard, the actual focus change is blocked behind some automated scroll behavior, and the effect is that someone quickly hitting keybinds will collapse the wrong thread. Would suggest making the focus change happen concurrently.
[OP] postalcoder | 11 hours ago
sochowski | 11 hours ago
grandwizardmarv | 11 hours ago
flashu | 11 hours ago
thisisauserid | 11 hours ago
dr_kiszonka | 9 hours ago
1vuio0pswjnm7 | 7 hours ago
https://hcker.news/feeds/atom?period=day&ai=exclude
https://hcker.news/feeds/json?period=week&ai=exclude
More info about parameters
https://hcker.news/feeds
dsiegel2275 | 7 hours ago
sajithdilshan | 6 hours ago
[OP] postalcoder | 6 hours ago
ares623 | 6 hours ago
ragall | 6 hours ago
[OP] postalcoder | 6 hours ago
ragall | 4 hours ago
Perhaps it's Cloudflare that's doing the blocking.
[OP] postalcoder | 8 minutes ago
narrator | 6 hours ago
[OP] postalcoder | 6 hours ago
d3Xt3r | 6 hours ago
[OP] postalcoder | 6 hours ago