Meta's Muse is fantastic for web scraping

59 points by STRiDEX 17 hours ago on hackernews | 74 comments
I built an AI "web harness" running on a sandboxed Chromium (using a custom side-loaded plugin that talks over websockets to a "driver") to basically do anything a normal user could do in a browser. It totally bypasses any and all bot measures and only gets the ones you yourself would get as well (and passes those successfully, e.g. Cloudflare checkbox or those annoying OCR puzzles).

Not sure if I should release it, but I'm sure more people are catching onto the power of agentic browsing.

ares623 | 16 hours ago

By describing it here you've already released it no?
FYI Muse smashes through those captchas natively without prompting.

frabcus | 16 hours ago

Right, but what happens when everyone uses that at scale? Without agreements and standards it isn't pretty.
It's sad, but this train has left the station a quarter of a century ago imo. Many people have since become billionaires scraping the web without anyone agreeing (Google, Yahoo, and now possibly Anthropic and OpenAI).

csnover | 15 hours ago

There is a huge difference between a clearly identifiable, robots.txt-respecting, largely symbiotic search engine spider versus this new AI-driven casual sociopathy of firing shotguns of deliberately masked bots at sites in ways which benefit no one except for, maybe, the bot-owner.

I agree with you, though, that the future has almost certainly already been decided, and see the eventual outcome of all this selfishness to be mandatory device attestation to access most services on the internet—which will of course never be allowed on any open platform. AI engineers who believed in freedom to compute (or just individual freedom more generally) have committed perhaps the greatest self-own in the history of our species to date.

bayindirh | 16 hours ago

Thanks for letting us know that we need a new layer of detection systems.

Also it’s great(!) to see that we’re going from “but ethics” to “I got mine, who cares”.

Humans are interesting creatures.

Edit: Please before assuming that I'm assuming things, this is an observation I'm making over time. It's possible that I'm in a bubble, but it's not a sample size of 1 (i.e. The comment I replied only).

pcthrowaway | 16 hours ago

There is no detection method that will prevent AI from accessing systems without also blocking humans. The only thing we can do at this point is throttling.

kees99 | 15 hours ago

Agreed on inevitable collateral blockage of real people using real "headed" browser. I'm getting a ton of that already, personally.

Throttling is poor help though. Mass scrapers are using "residential proxy" loophole + rotating UA and other attributes. You can't throttle somebody without identifying them. Unless you're talking about a global rate-limit.

pcthrowaway | 15 hours ago

throttling based on sessions kind of works; we're headed in the direction that sites like Reddit will probably prevent logged-out users from viewing threads (as they already do with mobile devices)

Once the LLMs create sockpuppets to get around that, the web services will need to resort to profiling users more aggressively so that they know which actual human an account corresponds to.

If someone has a malicious browser extension that uses their session to scrape Reddit then, they're probably going to see significant usage obstacles.

We are headed to a very user-hostile place.

mvt67 | 13 hours ago

Only if your life revolves around reddit.

cyanydeez | 8 hours ago

I assume most people have see the photos of various "far east" people sitting at a bench with an array of 100 phones.

This predates AI as a _capitalism_ problem.

afro88 | 16 hours ago

I'm becoming more and more convinced that a big source of outrage on the internet is caused by people assuming that all other people are a homogenous blob.

It's not that "we" are going from one thing to another. It's that these are two different people, with different ethical boundaries.

Cakez0r | 15 hours ago

You're framing this as if people are deliberately making decisions that they believe are unethical. The reality is that people have different ethical frameworks. For example, I believe that there is no ethical distinction between whether a web request originates from a browser or from an LLM on my behalf.

dd8601fn | 15 hours ago

It’s a big question mark in the conversation.

Is defeating captchas unethical? I don’t think so. Not on its own.

Is scraping unethical? I don’t think so. Not on its own.

Are there tons of uses for both of those that are sketchy or outright wrong? Yeah, absolutely.

Kepten-Hook | 14 hours ago

Another person replied already and I agree with them, but also I think you need to see this from the perspective of the people operating the service you are accessing.

As an example, imagine a really small community maintaining a small site/wiki/cms/forum that has the ultimate goal of promoting human relationships around a common interest (let's say retro computing as an example, but it could be anything). There are many many many such communities on the internet.

It's very improbable you will specifically instruct your LLM to access their site, it's way more probable it will happen without you even knowing, as a result of you doing some /deep-research or something. And not only that, but your LLM will probably spawn a ton of agents to gather as much information as possible in as little time as possible. A torrent of requests will go at this community's site, effectively killing it. They are a small community, they use their spare time and money to maintain something to serve them, they don't have the resources to serve your LLM and until you showed up they probably never even had to think about Cloudflare. They are certainly not against you getting the content, but they don't want you causing them issues either and you just did.

End result? Your LLM (effectively you) DoSed a small community's site. You caused harm. Could you have caused similar harm if you were doing it on your own? Sure. But you would have done it on purpose, not accidentally while instructing your LLM to do something else.

This isn't a made up story, it has happened already more than 1 times.

So the question is, now that you know your LLM can cause harm without you even knowing it, how does this change your stance?

mvt67 | 13 hours ago

Already have whatsapp and signal for this. Hardly need a website for these thing anymore.

bayindirh | 13 hours ago

So shall we build more walls around our content and make it anti-FAIR?

We can stop putting information in the open in any form, as well.

FAIR: Findable, Accessible, Interoperable, Reusable.

mvt67 | 13 hours ago

Why do you build a wall around your house genius?

bayindirh | 13 hours ago

A public website is not a house, it's something between a big board in the city square or a community center where people can meet an interact.

Considering that, shall we paint the boards black or lock the doors to community spaces?

Also, while I assume this is not your first account here, these are worthy of reminding:

> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.

> Throwaway accounts are ok for sensitive information, but please don't create accounts routinely. HN is a community—users should have an identity that others can relate to.

For more, please refer to https://news.ycombinator.com/newsguidelines.html

murderfs | 12 hours ago

The answer is that no one gives a shit: just look at people's views on adblock.

watwut | 14 hours ago

> The reality is that people have different ethical frameworks.

Sure, Nazi, Hitler or Stalin all considered themselves to be the good guys. Kushner, Trump, etc are also acting within their ethical system (if I am getting money or power or fame it is ok to do it). Just about the only exception is Thiel who openly frames himself as evil, but is proud of it.

That does not mean we cant criticize their crappy actions or "ethical frameworks".

And yes, all the above are making deliberate decisions to be unethical assholes.

cyanydeez | 8 hours ago

There's two types of cares:

1. People like other people.

2. Businesses need to sell product

The internet mixes those people, and an Agent basically pushes the signal to noise ratio that businesses have relied on the intenet to basically zero.

If the internet just allowed indescriminate traffic, neither #1 nor #2 survives, and while it's interesting to think the value is, it certainly isn't anything we understand.

Maybe you think the value will still exist, but some how agents will replace all value with equivelents. That's an argument, but I don't think it will.

jareklupinski | 9 hours ago

humans also change over time

in batman, bruce wayne shuts does the massive surveillance network, seeing it was problematic (2008)

in spiderman, peter parker just shrugs off being able to know where anything is happening anytime (2026)

bayindirh | 9 hours ago

> humans also change over time

Yes, that's what I'm referring to. Citing myself:

> Also it’s great(!) to see that we’re going from “but ethics” to “I got mine, who cares”.

This is exactly how I observe a shift in general.

jareklupinski | 9 hours ago

maybe it'll go backwards? or the direction is chaotic / tick-tock

bayindirh | 8 hours ago

I believe that the movement is sinusoidal,since every trend feeds its counterbalance.

OTOH, stabilizing at "0" or any point is impossible since the system is 2nd order and the response of the system is also an input to itself.

jareklupinski | 7 hours ago

> sinusoidal

> OTOH, stabilizing

any function is stable; im just dissppointed if its predictable / manipulable / only has 2d

Interesting, how is the chromium sandboxed? profile dir cli param? or deeper like chromium engine framework embeded in the application?
Just simply a separate Chromium binary (not your usual browser).
anyone can spin these kind of side projects and do easy talk, but the moment you actually try to use this on signed-in Linkedin or Amazon its going to fail

the only solution is to drive your regular browser with all your sessions/cookies via an extension

> the only solution is to drive your regular browser with all your sessions/cookies via an extension

This is exactly what I'm doing, but mocking/randomizing all the sessions/cookies/params (like resolution, OS, WebGL , etc.) in a separate Chromium binary. It's popular these days, but imo using your normal browser for agentic stuff is a very bad idea. These models do dumb stuff all the time.

carsoon | 16 hours ago

no it doesn't fail. Agents are very good at comparing traffic characteristics from a real browser and a headless/automation browser and getting it to behave in the same manner. It's a cat and mouse game for sites stopping unauthorized access but right now llm agents are ahead.

For testing proxies should be used to avoid IP ban issues but given enough time modern agents can figure out how to bypass most of the modern scraping/automation prevention mechanisms.

I have built and used a lot of different automations and web scraping implementations for my business and it's never got permanently stuck yet, some take a bit longer, some shorter, but all within a reasonable time with little external help they have succeeded in their tasks.

kurisufag | 16 hours ago

the 4get dev did the same thing a little while ago for his metasearch engine: https://git.lolcat.ca/lolcat/4play
Yep, it's super similar to this! I think his is a bit overengineered, but tbh I haven't built Firefox plugins in forever so that might be the way you have to do stuff on FF. All you need is the websocket plugin to have "full access" to all visited webpages, and then your driver just hooks into the DOM like normal (with the extra ws layer—and even the websocket layer can be simplified away I think).

pprotas | 16 hours ago

Camoufox bypasses most blocks with no problems https://camoufox.com/

onion2k | 16 hours ago

(using a custom side-loaded plugin that talks over websockets to a "driver")

Is that necessary? You could start Chromium with an open debug port and use Chrome Devtools Protocol to send commands.

It is, because `--remote-debugging-port` is detected via the root DOM object (and that can't be changed unless you want to recompile Chromium), so for example, if you try doing a Google search with debugging enabled, you'll get blocked (usually just by being served a blank page).

LunaSea | 15 hours ago

Interesting, do you know what exactly changes on said root DOM node in case the remote debugging feature is enabled?
`navigator.webdriver` afaik, but I think there's a couple more, I did some research on it a few months ago. If using Playwright/Puppeteer, they also have a few special things injected that you have to strip.

Cakez0r | 15 hours ago

What model did you use to build this? I recently looked in to building something similar and got refusals.
I built the POC by hand (just had like 3 things it could do: scroll, click, type), and then Codex used my established patterns to make it slightly more general and cleaned it up.

LunaSea | 13 hours ago

Ah right, there are stealth plugins to remediate these types of hints I believe.

raffraffraff | 15 hours ago

Does remote debugging port itself get detected, or does the presence of a webdriver connection to the port get detected? I believe it's the latter.

I've had success launching the browser and using dumb dumb methods to get around the captcha before attaching playwright.

Dumb dumb methods = wmctrl, xdotool, bash (work fine for Cloudflare's "are you human" check)

I think it's the latter, and that's why I think using a "normal" non-debug browser is the best idea because it's essentially undetectable. And for the record, you can technically control Chromium using IPC if you feel like adding that feature & fully rebuilding it from scratch.

kees99 | 15 hours ago

Some webshops (e.g. Aliexpress) already soft-block Chromium (at least Chromium-on-Linux). I.e. such users see larger-than-usual share of captchas.

Assuming "more people catching onto" this, expect Cloudflare and most everyone else to follow the suite.

Are you saying they're blocking the `User-Agent`? Because that's trivial to bypass. I'm not exactly sure what you mean by "Chromium" because that's a whole class of browsers (Edge, Brave, Chrome, etc. are all forked "Chromium" browsers).

kees99 | 15 hours ago

Aliexpress specifically uses a javascript blob to do detection, and quite a sophisticated one. There was a recent story about that blob messing with bluetooth headphones.

> not sure what you mean by "Chromium"

I meant that literally. This thing: https://www.chromium.org/getting-involved/download-chromium/...

Here's my agentic Chromium browser browsing Aliexpress[1]. Usually I'd run it in headless mode, but wanted to test and make sure I'm getting the website proper.

It searched for "Aliexpress" on Google and clicked the link, hence the Google UTM params.

[1] https://imgur.com/a/eNulg11

kotaKat | 12 hours ago

And the worst part is humans just getting more cognitively demanding captchas.

Some days I'm literally fatigued enough I can't clear the overly complex hCaptchas these days I keep getting forced upon me and bots sail right through this crap.

Cakez0r | 15 hours ago

I think it's inevitable that eventually one of the big AI companies makes something like this. It's such an obvious consumer win. "I am not a bot" checkboxes are a string and peg tethering the elephant.

busssard | 13 hours ago

please release it and call it SilverSurfer

cyanydeez | 8 hours ago

I've got a plugin in development that'll do the same: via an extension, connect to a standard coding bot, so it can drive.

It can be used for testing software and using it as a human support.

arjunchint | 16 hours ago

their static ip's were initially good and didn't get flagged, but now most sites are recognizing their ip ranges and blocking.

Muse's utility has significantly dropped with the blockages.

To become truly useful again they will need to use residential proxies, but I can't see them use those due to the risks and reputational damage.

sejje | 16 hours ago

they can just use the user ip. i think grok already does this.

koolala | 15 hours ago

How can it do this? Wouldn't you see a hundred fetch requests in your browser network tab?

arjunchint | 14 hours ago

bruh have you heard of CSP?

Gareth321 | 16 hours ago

Most residential proxies are already far more blocked and rate limited than any Meta IP. The internet is becoming a very weird place, where individual and "trusted" personal IPs are becoming a kind of commodity. Some sites are already scoring IPs based on usage activity - like a credit score. It's only a matter of time until this data is collated and commoditised. AI analysis is turning this up to 11.

iamacyborg | 15 hours ago

That doesn’t track with reality as far as I can tell.

simoncion | 15 hours ago

> Some sites are already scoring IPs based on usage activity - like a credit score. It's only a matter of time until this data is collated and commoditised.

Spamhaus is nearly thirty years old and the notion of electronic distribution of IP and domain "reputation" lists is at least that old.

I'll bet my hat that the Internet "advertising" [0] industry has been calculating and determining the reputation of individual households (if not individual users) for at least a decade.

[0] The scare quotes are because its primary purpose these days is for dragnet private-sector surveillance.

wraptile | 15 hours ago

Every time a new tool launches there's a good window where it can act as a scraping proxy. Back in the 2010s I used Google Translate for years to scrape hard targets like LinkedIn but these windows are much shorter these days as scraping is so much bigger.

One thing with Muse though is that you can scrape Meta's own sites which are currently all going under login walls and restricting discovery/search entirely.

[OP] STRiDEX | 15 hours ago

similarly, gemini via google ai studio will happily run workloads over youtube that would be very annoying to run at scale. Especially if you needed to download the video.

memcg | 8 hours ago

"risks and reputational damage"

Good one, I can't stop laughing. Thanks!

kevmo314 | 16 hours ago

Muse ran into a captcha and asked me if I wanted it to solve it.

So of course I clicked yes and it dutifully convinced the site that it was not a bot.

xnickb | 16 hours ago

For 2(3?) decades we've been training the robots to tell traffic lights from fire hydrants. It's finally paying off.

Gareth321 | 15 hours ago

I think this is the end of sites where it is expected that only humans may interact with them. It's been cat and mouse for a while, and some places like Reddit sell API access, but these agents are for all intents and purposes, humans interacting with the site. They're going to need to figure out new business models.

mejutoco | 15 hours ago

Or the captchas will evolve.

therein | 15 hours ago

This will be used as an excuse to normalize "scan QR code with your phone to verify you are human" along with Web Environment Integrity stuff.

wiether | 15 hours ago

As someone having to fight Meta's bots everyday to keep websites accessible to actual customers, I'm not surprised to read that it can be seen as something positive on the other side of the fence.

But I'm wondering: at what cost?

[OP] STRiDEX | 15 hours ago

I think for my own side projects i would require the user to login if they were making requests from those ip addresses or block.

Cakez0r | 15 hours ago

The end game will be that either your site is fully open to bots and humans alike, or your site is open to humans only(×) and requires Airport security style identity verification.

(×) and their AI delegates

gunalx | 15 hours ago

Im guessing more of the internet will be login walled from now on.

sixtyj | 13 hours ago

> I don’t know where this leads, but we’ll likely see more websites block Muse unless Meta prevents abuse.

Multiply it by 1,000,000 access attempts - daily.

It is question of time when even normal browsers would need some allowed fingerprint to go to sites otherwise caught by clouflare or similar wall.

lm411 | 5 hours ago

There is at least one very large derivatives marketplace that will require basic recent price information (OHLC) to be behind user registration by next year. Enjoy what you have done scrapers.

"I'd sure hate if someone did this to my site, but at least the agent did it for me. I don't have to feel guilty now."