Same here. I would seriously doubt the scraping activity would drop, because I've got a few URLs/domains that have never been indexed by any search engine and are constantly hit by scrapers and bots attempting to find something (they look for .env files and .php files though there's no PHP running in that server).
I believe the list of publicly certified domain certificates from Let's Encrypt is the most popular source for bots now, since I've had new domains get hit within 1h of existing and the only "thing" that knew about them at that point was their service.
I believe the list of publicly certified domain certificates from Let's Encrypt is the most popular source for bots now, since I've had new domains get hit within 1h of existing and the only "thing" that knew about them at that point was their service.
Solution could be to do wildcard certs and then put stuff you don't want found on a subdomain.
It would be expected, since they're blocking all bots. What's surprising is the explicit need to allow google bot in order to read the global disallow.
At my previous job we've had issues with Huawei bot trying to scrape absolutely everything from our server - we've done everything by the book (including their specifoc recommendations) and at the end of thr day we had to forcefully block its UA, as they claimed they update cached robots.txt every SIX MONTHS, so 3 months are really good!
I'm not sure I understand the end goal perfectly here.
The idea, basically, is that if you are not available via Google, you are harder to find by the scrapers? Because those are the annoying scrapper you don't want?
ryan-duve | 14 hours ago
I was not expecting this to be a happy post just based on the title. Glad the author got what they wanted.
I am curious whether there is a drop in other server activity (LLM scrapers, other search engines, SSH login attempts, etc) after the Google index.
[OP] brn | 12 hours ago
Same here. I would seriously doubt the scraping activity would drop, because I've got a few URLs/domains that have never been indexed by any search engine and are constantly hit by scrapers and bots attempting to find something (they look for .env files and .php files though there's no PHP running in that server).
I believe the list of publicly certified domain certificates from Let's Encrypt is the most popular source for bots now, since I've had new domains get hit within 1h of existing and the only "thing" that knew about them at that point was their service.
marginalia | 12 hours ago
Solution could be to do wildcard certs and then put stuff you don't want found on a subdomain.
[OP] brn | 12 hours ago
That might be a good incentive to use subdomains more, instead of domains. Thanks!
ploum | 9 hours ago
Cannot find you on Kagi either.
[OP] brn | 4 hours ago
It would be expected, since they're blocking all bots. What's surprising is the explicit need to allow google bot in order to read the global disallow.
hwj | 7 hours ago
Interestingly, 6% of my nobody-cares-website are coming from Kagi, whereas only <1% are from Google.
AndrewStephens | 7 hours ago
I accidentally blocked GoogleBot a few months ago and thought long and hard before white-listing it again. I get very few referrals from Google.
Long gone are the days where I would go out of my way to make my site easy to spider.
patryk | 6 hours ago
At my previous job we've had issues with Huawei bot trying to scrape absolutely everything from our server - we've done everything by the book (including their specifoc recommendations) and at the end of thr day we had to forcefully block its UA, as they claimed they update cached robots.txt every SIX MONTHS, so 3 months are really good!
laurentbroy | 4 hours ago
I'm not sure I understand the end goal perfectly here.
The idea, basically, is that if you are not available via Google, you are harder to find by the scrapers? Because those are the annoying scrapper you don't want?
motet-a | 4 hours ago
It’s really a shame that Google is so slow to update low-traffic websites. It should be way faster!
yharnam | an hour ago
The only spiders I allow on my little web project are from the Internet Archive and from Marginalia Search.
I've been meaning to allow Kagi for a while but haven't gotten around to it.