[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


4chan will be down for scheduled maintenance starting at 12 p.m. Eastern Time.


[Advertise on 4chan]


File: 2026-08-16_15-48-32.png (396 KB, 1920x1080)
396 KB PNG
My website isn't even 10 days old and everyday i get hammered by thousands of AI requests.
i have failed2ban and ip blocks set up with all sorts of regex now....but this can't be sustainable no?
mind you this is a low end vps and no one knows about the website except for my close friends yet i find myself in this situation i can't even imagine how bad it is for more popular sites or projects.
First they came for your hardware, then they came for your software and now they don't even want you to self host your own online space...You guys can call us luddites or whatever as much as you want but something has to be done about altman and his jewish peers raping everybody on the web
>>
I only get anthropic's crawler on mine, surprisingly never any google at all
>>
I know it looks bad, but does it really matter? It's just background noise
>>
>>109571537
why are you even responding to the requests? just drop the packets via firewall or close the connection on your server.
>>109571639
yes, it drains your resources.
>>
>>109571698
>why are you even responding to the requests? just drop the packets via firewall or close the connection on your server.
not him but that's what tools like fail2ban are for. fail2ban can be setup for many kinds of services, but the point is that repeated unsuccessful attempts at things like ssh logins result in dynamically added firewall rules to reject or blackhole further attempts by the offending ip
not op but i also have a low end vps for personal use and while occasional automated/bot attempts has been a thing as long as i can remember, the sheer volume of them these days is alarming, and only getting worse. thing about networking is that even if you blackhole an attempt (don't even respond to it), you still have to download it at an absolute minimum before your machine can even decide to not respond to it, so there's no way to completely eliminate wasting resources handling bogus requests.
>>
>>109571610
Too much work
>>109571639
Talking about copyright issues and intellectual property is a dead horse at this point, but the grip they have on the modern web is so exhausting and dystopian...and im pretty sure these are regular mindless scrappers i could imagine it would be a nightmare if they were more directed like with popular projects and sites
>>109571698
im not a web dev nor do i have alot of experience with network so i thought this was the way to go because thats what most people do...but i may look into your solution too
>>
>>109571793
>im not a web dev nor do i have alot of experience with network so i thought this was the way to go because thats what most people do...but i may look into your solution too
in networking/firewalling terms, the difference between "reject" and "drop" is that reject sends a packet back to the sender to tell them the request was denied, while drop just drops (ignores/deletes) the packet without responding to it.
dropping is usually better if you suspect you're dealing with undesirable or malicious traffic, since it makes your server come off like it's unresponsive or just gone, which may make bots slow down or stop, while rejected packets instead tell the bot what they can "get away with" and may try different things instead.

it's like the difference between answering a call to say "not interested" and just not picking up the call at all
>>
>>109571537
Am I the only person left in the world who knows how to setup fail2ban?
>>
>>109571854
not op but if you have some neat tricks then do tell
>>
As Cloudfare said, "humans will soon be a rounding error on the Internet", bot traffic is going to takes over bro it is over, average normie in 2026 just searches on chatgpt.
>>
>>109571639
Retards don't know how to host things properly so the hosting price scales exponentially with the number of requests they get.
>>
>>109571859
>setup fail2ban with default settings
>???
>virtually 0 ai traffic
>very limited bot traffic, a few non-abusive crawlers sometimes
There is literally 0 to do. Just setup fail2ban and done. Nothing special needs to be organized there.
>>
>>109571911
oh, you just meant knowing to use fail2ban at all
>>
>>109571969
Yes. It seems like nobody's heard of it before and then everyone goes pikachu face when they're getting bot hits, as if AI invented web botting.
>>
>>109571639
It's just icky - they're /touching/ my /site/!
>>
>>109572030
nigga i already have it setup if you look closely at the screenshot they are all returning 403 requests, its just annoying how nobodies like most people here have ai crawlers sniffing around their websites
>>
>>109572065
An ip that gets banned does not get 403 and you should have a repeated 403 -> ban rule enabled. You're doing something wrong.
>>
>>109571610
This gave me an idea to setup a local llm redirect to poison the data they receive
>>
>>109571537
>Is this the state of the web
Unfortunately yes. AI theft is rampant. Check this out for more methods to block the thieves: https://nochan.net/b/Internet-Crap/20260606-How-To-Block-Some-Of-The-Bots/
>>
>>109571537
Curiousity:
What's on the site?
What (non-"yours") DNS's have previously or (even currently) pointed to that IP?

Incidently, the FSF 'recently' had similar issues... when they hit the limits of fai2ban they moved to reaction: https://reaction.ppom.me/

>>109571610
>content is just word nigger
Evidence that the racial slur will cause any reaction from the unfeeling machine as it absorbs and catalogs the data as instructed?

>>109571911
>Nothing special needs to be organized there.
Don't worry anon. Maybe one day you'll be interesting enough for all the winxp botnet on residential comcast IP's to start taking a swing.
>>
>>109571537
Why are you running a website? Just replace your own usage with AI like everyone else is doing. AI permanently replaces all content—the entire web.
>>
>>109572483
Botnets I don't think I've seen, but bots definitely. As I said they just get one-shot banned, no problem. Most of the bot traffic I get is in china, the philipines, and cloud providers.
>>
File: Linked_attacks.png (879 KB, 1366x768)
879 KB PNG
>>109572602
>Most of the bot traffic I get is in china, the philipines, and cloud providers.
You actually expecting legit traffic from these places?
Strikes me as /8 /32 fodder...

>philipines
Ahh. The Seychells islands. Remarkably lax in the way of legal restriction. All the datacenter space is already leased to alphabet soups tho. =(

>As I said they just get one-shot banned, no problem
Rate at which I've occured 'em I'm legit shocked. Sans firewalling, I can 10k+ attempts on one fing, each from a unique IP.

Picrel. What that demonstrates is independant sources attempting *the same things*. Three seperate services surrounded by their hostiles. As the hostiles create bonds to the other hostiles they clump together...

And that's a snapshot of not even a full month back in like 2014? 2016? that ballpark...
>>
>>109572750
What kind of site are you running? For me it's just a professional blog that brings in inbounds.
>>
>>109571537
it was always like this, crawlers were always a thing
>>
>>109571909
huh interesting, so it's yet another nail in the road approach just like virus industry
>>
You hear about AI scraping sites a lot but why? Why would hundreds of AI agents be constantly hitting a random blog on a topic of generally low interest? Storing a copy of site would take one pass, and you have for pay for storage remember.

I think this meme is meant to sell solutions and not a real thing.
>>
>>109572990
AI don't cache any web use, they just use e.g. google or bing api to get results, then read the page in full, every single time.
>>
>>109573014
Why?
>>
>>109571537
pretty much. i think a read Cloudfalre said that over half of all web traffic is just bots.
>>
>>109573023
Because "AI" "developers" are profoundly retarded. None of them have any AI background, they're just webshitters, often fresh webshitters.
>>
>>109572833
That big blob to the left, that was a blog. Long dead nao.

The two blobs on the right is email n ssh IIRC, both on the same system, seperate system to the blog. Seem to think something of a forum, but that could of been somewhere else.

All them yellow dots is inbounds... IPv4's specifically. Each one 'did something' to get in that list. The 'what' they did is what clumps them together....

These isn't crawlers, screwglebot n that ilk. Or even shodan et al. In the example of the blog, the light blue/cyan dots are usernames/emails used - typically to throw various spam. But some actively attempted to leverage vulnerabiities... And before the end of the month started drawing lines to blocks associated with spam... And assaults on SSH on a seperate system...
>>
File: Linked_attacks1.png (1.63 MB, 1366x768)
1.63 MB PNG
>>109573106
Forgot, picrel
Same graph, greater zoom...
>>
>>109573023
>Why?

Could it be, that you had previously answered your own question?

>>109572990
>and you have for pay for storage remember.

Also:
>>109573049
>Because "AI" "developers" are profoundly retarded.
>they're just webshitters, often fresh webshitters
It's highly likely they consider 'internet' to be a 'long term storage medium'. I mean, it works for their photos in icloud, right?
>>
>>109571698
> just drop the packets via firewall
That’s the wrong thing to do. Back in the 80s we rewarded spammers with uuencoded coredumps.

The right solution is to poison the results with reasonable looking nonsense. Let the “claudebot” agent through but give it some slightly randomized realistic looking nonsense—just like AI responses are.

If you just drop them, they will just scan them using other, harder to detect methods.
>>
>>109571610
serve them a zip (or GZ in web case) bomb of niggers, small to store locally and most web browsers and shit can accept gzip compressed documents
>>
>>109572990
I am more interested in knowing how the fuck will they parse a react blob, since modern websites are all spa slop, detect the json api or something? but what if the javascript blob is obfuscated
>>
>>109574601
They use headless browsers in containers.
>>
>>109571754
>>109572056
>>109573435
i would image the ip address ranges are overwhelmingly 3rd-world-ish.

surely the simplest and most straightforward solution is to ban any packet not from the US/EU? this would eliminate 95% of malicious login attempts at probably no more than 5% cost to readership.

i mean, think about it. what could a user from nairobi or thailand conceivably contribute to you or your environment by visiting your website? infact, we would probably be doing them a courtesy by not bothering their lived experience with web content that is, at best, of no consequence to them.
>>
File: opnbsdpf.png (230 KB, 2428x1360)
230 KB PNG
>>109574970
>i would image the ip address ranges are overwhelmingly 3rd-world-ish.

Not really, those days are long gone. Maybe you get some Chinese/Hong Kong ips once in a while, but nowadays it's always AWS/Azure/Cloudflare. Everybody runs their bots from the cloud now.
>>
>>109571537
not sure why but people are incredibly excited about the already consolidated net to become even more consolidated and controlled
>>
>>109575196
that is sad if true. is that really the case? i would have thought nobody would be harsher than clouds on misuse, given they need to protect their ip range reputation.

then again, i suppose they have a strong bargaining position given the vast amount of legitimate services they overwhelmingly host, so in terms of game theory they could get away with offloading the costs of misuse to the rest of the internet ("not my problem", oil draining into sink jpg)
>>
its not "AI" traffic, its just regular bots
it always happened way before AI
AI scrapes websites to get data there is no need for them to repeatedly make login attempts on your shitty vps
>>
>>109575259
>no need for them to repeatedly make login attempts on your shitty vps
Huh? who said anything about that?
>>
>>109574970
you again ameritranny, fuck off
>>
>>109571537
If you have a TLS certificate, your domain is immediately published on the transparency log and thus everyone knows about it. You have to write performant webservers now. You can't get away with some garbage that can barely do 2 or 3 requests a second, so your shitty Django crap or cgit instead of some proper forge.
>>
>>109572065
Why does your server even respond to pings?
>>
>>109571537
Disable all ipv4 in your hosting, your site should ipv6-only. That bots scan the entire ipv4 space in 1 hour, this why they found your website
>>
>>109574970
> i would image the ip address ranges are overwhelmingly 3rd-world-ish.
The fsf and gnu folks are were blocking over 5 million IPs and ranges at one point.
That database should be published and we can take the appropriate action.
One action is to go to the upstream providers and get their service throttled.
>>
>>109571537
My website only like 10 people know about is basically under a constant DDOS for ages, plus SSH login requests onto another IP that's not public.
It's truly ghastly and unsustainable.
>>
>>109576795
They’ve been trying to prevent individual/independent hosting for quite some time.

Why not just use Twitter?

Now they’ve moved on to try and prevent computer ownership.

I remember in the early 2000’s our internet provider blocked everybody’s inward access to port 80 and 443 to prevent hosting at your house. It competed with their web host service where they gave you a few hundred K of static files you could serve from their webhost infrastructure.

Of course, that’s long dead now (the Internet provider company is now a nontechnical money extraction entity of some sort) but everything is still blocked.
>>
>>109574970
true but ai companies are overwhelmingly chinese or american
>>
>>109574970
>>109575236
They all use cloud provider like Google Cloud, Microsoft Azure, etc. The IP range are from US/EU
https://www.abuseipdb.com/check/34.123.33.125
>>
>>109571859
NTA, but disabling password login and changing ssh ports (22 -> 48220) is a good start. Search for guides on "hardening vps".
>>
>>109578605
changing port doesn't do all that much, bots are smarter than they used to be. you absolutely should disable password login though, even if your password is excellent you can reject attempts earlier by rejecting password attempts outright
>>
>>109578624
>>109578605
-- if you want to do something fancy with ports, port knocking is probably more useful. i haven't tried that myself but i'd be surprised if it doesn't cut down/cut out background noise
>>
>>109571639
By that metaphor it may as well be the digital equivalent of tinnitus.
>>
>>109573036
>Cloudfalre
those scumbags have fingers in both pies - hosting bots and then selling bot protection to sites.
>>
>>109571854
fail2ban does nothing to prevent people hammering your servers though, if anything it causes more resources to be used. same goes for any redirect tricks that people do
>>
>>109578754
fail2ban will cause people hammering your server to get blocked at the request level, before any handling happens. You can easily handle billions of requests at that level without breaking a sweat. Very different than having to go through routing and response and then deciding to reply unauthorized or something, let alone providing an actual response.
>>
>>109579029
only if those recurring requests come from the same ip, which is not true most of the time
>>
>>109572990
Little known fact but these AI scrapers have been raping even Wikipedia's servers, despite Wikipedia regularly publishing data dumps they could just download.
>>
>>109579254
>only if those recurring requests come from the same ip
Yes
> which is not true most of the time
My experience has been that though there is provenance from many IPs, they all repeat from the same ip and thus get caught.
>>
>>109579306
Because they're not looking at using data from wikipedia. The user is asking "what is a cow" and the ai is using wikipedia as a source by scraping the page.
>>
>>109571537
serve up a custom page for bots that is just the word "NIGGER"
>>
>>109571537
>Is this the state of the web in 2026?
pretty much. before that it was chinese bots doing similar things and trying to break into your ssh

>>109578754
>fail2ban does nothing to prevent people hammering your servers though,
anon is right
>>
>>109579371
this is true. since google, etc. are all useless and want to keep up this farce of all knowing ai systems, they want to keep this free stream of real time data alive, leeched from other websites. the industry is kind of rotten and its intelligence is way oversold



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.