[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: 1791649487028.png (51 KB, 944x867)
51 KB PNG
>blacklists planet
>whitelists north america

Thousands of rotating ip addresses were trying to download images from my website. Its a 4chan archive, so all the images are on the internet archive, or easily accessible from the 4chan api itself

They were using an old url pattern for images, which is how I could do anything about this. Once the faggots wipe the drool off their faces and switch to the new pattern, I'm not sure how to combat this

Why do they perform the most inefficient data collection strategy possible?
>>
Goodnight sweet prince
>>
>>110029894
>I'm not sure how to combat this
Put all images in a hidden folder and set up a htaccess url rewrite that serves them from another link, like files are in /hiddenfolderasdfg/*.* and accessing /lalalakatamaridamacy/*.* will read the ones in the other folder.
Change the site links so the image folder is served by one function, that function reading the media url prefix from say a config file (a non public one obviously).
Set up a cron job to a function that changes the media directory name randomly and regenerates the htaccess url rewrite using it and call it every I dunno 30 minutes or something.

So basically you are still just switching patterns, but you switch to a random one every x minutes.

alternatively also block all datacenter IPs and ASNs in your firewall.
>>
>>110030177
both very good ideas
that will save a lot of bandwidth
>>
>>110029894
They perform the most efficient work for themselves—if you don't like people using your API then don't host it: you have only yourself to blame.
>>
>>110030353
no they don't, if you could comprehend my post stating easier solutions, you'd know it
>>
>>110030177
This is smart but are there better solutions than a cron?
>>
>>110030699
Better in what fucking way?
>>
>>110030722
Idk, cons seem disconnected from the codebase. Maybe it doesn't matter.
>>
>>110030760
>cons seem disconnected from the codebase.
if you mean cronjob, if you can monitor bandwidth you should have access to the cron.

But you could also just do something like "if cfg file was last changed >30 minutes ago, generate random key again", before you load said config. Or even store the last change time in Redis or something and look it up on page load.
>>
>>110030760
can insert some logic in the media route, which I plan to do
>>
wait, having a map in nginx would be even cleaner
>>
Zump
>>
>>110029894
Look up the ASN of the offending IP and block it, the entire ASN, their IP ranges are public.
>>
>>110029894
> thousands of users want to use your website
Whats the issue?
>>
I saw this analogy for AI scrapers on Reddit:

It’s like how a roomba just cleans in random directions — it doesn’t need to be efficient because it’s automated.
>>
>>110029894
blacklist north america.

Especially all the corpo subnets.
You did blacklist Google, AWS, Cloudflare and Digital Ocean, right?
>>
>>110030177
Back in the day, people just hid honeypot URLs, that bots click on, but no human would ever find, and if they trigger it, they get auto-blocked for a day.
At best, this honeypot URL would break the connection, to let it timeout, so the automated bot rolls for a different IP to try it again, getting blocked again, and so on.

Weird how this is still the most effective method. No need to break URLs.
>but they may have ten thousand IPs
ten thousand requests should be able to survive
>>
>>110035462
Yeah that sounds interesting but I wouldn't know how to link that up to the firewall, and nowadays even googlebot ignores things like nofollow links and I don't want to ban googlebot server wide.
>>
>>110037736
You link it to the firewall by making a fail2ban jail. Say the trigger is a certain response code or request or user agent from your server log. Fail2ban can monitor that, and block their IPs automatically.
>>
File: 1771609332379628.jpg (16 KB, 294x349)
16 KB JPG
>>110035346
>on Reddit
>—
>’



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.