Cloudflare announces AI Labyrinth, which uses AI-generated content to confuse and waste the resources of AI Crawlers and bots that ignore “no crawl” directives.

[email protected]

So the world is now wasting energy and resources to generate AI content in order to combat AI crawlers, by making them waste more energy and resources. Great!

[email protected]

It's the consequences of the MIT and Apache licenses showing up in real time.

GPL your software, people!

[email protected]

Not exactly how I expected the AI wars to go, but I guess since we're in a cyberpunk world, we take what we get

[email protected]

Next step is an AI that detects AI labyrinth.

It gets trained on labyrinths generated by another AI.

So you have an AI generating labyrinths to train an AI to detect labyrinths which are generated by another AI so that your original AI crawler doesn't get lost.

It's gonna be AI all the way down.

? Offline

yeah. it's pretty fucked. hopefully it's temporary.

so do we make everything inaccessible to everyone, or just inaccessible to disabled people? we don't have a way to include them yet. we should work on it, but we are not the ones who fucked accessibility.

yeah. search engine web crawlers are a public service. they are responsible. but we are in a conflict. we must struggle tooth and nail against capital for every nice thing.

[email protected]

Geez, that's a lot of requests!

[email protected]

I swear someone released this exact thing a few weeks ago

[email protected]

It sure is. Needless to say, I noticed it happening.

[email protected]

and try to slam your site with like 200+ requests per second

Your solution would do nothing to stop the crawlers that are operating 10ish rps. There's ones out there operating at a mere 2rps but when multiple companies are doing it at the same time 24x7x365 it adds up.

Some incredibly talented people have been battling this since last year and your solution has been tried multiple times. It's not effective in all instances and can require a LOT of manual intervention and SysAdmin time.

https://thelibre.news/foss-infrastructure-is-under-attack-by-ai-companies/

[email protected]

How can authority not exist? That's staggeringly broad

[email protected]

Cloudflare offers that too, but you can't always tell

[email protected]

Cloudflare is providing the service, not libraries

[email protected]

Damned ~~Arasaka~~Cloudflare ice walls are such a pain

[email protected]

the only problem with that solution being applied to generic websites is schools and institutions can have many legitimate users from one IP address and many sites don't want a chance to accidentally block one.

[email protected]

It’s difficult to imagine a group of people voluntarily amassing and then using the resources necessary for “AI” absent the desire to cash in on their investment.

No imagination necessary.

I mean Dmitry Pospelov was arguing for AI control in the Soviet Union clear back in the 70s.

[email protected]

everyone remembers tomogatchi, they were like a digital houseplant.

[email protected]

It's worked alright for me. Your mileage may vary.

If someone is scraping my site at a low crawl rate I honestly don't care so long as it doesn't impact my performance for everyone else. If I hosted anything that was not just public knowledge or copy regurgitated verbatim from the bumf provided by the vendors of the brands I sell, I might oppose to it ideologically. But I don't. So I don't.

If parallel crawling from multiple organizations legitimately becomes a concern for us I will have to get more creative. But thus far it hasn't, and honestly just wholesale blocking Amazon from our shit instantly solved 90% of the problem.

[email protected]

This is fair in those applications. I only run an ecommerce web site, though, so that doesn't come into play.

? Offline

given what domains we're hosted on; i think we've both had a version of this conversation about a thousand times, and both ended up where we ended up. do you want us to explain hypothetically-at-but-mostly-past each other again? I can do it while un-sober, if you like.

? Offline

cool, but where do you get them? you can't, right? because they were stupid?

agnos.is Forums

Cloudflare announces AI Labyrinth, which uses AI-generated content to confuse and waste the resources of AI Crawlers and bots that ignore “no crawl” directives.