>blacklists planet>whitelists north americaThousands of rotating ip addresses were trying to download images from my website. Its a 4chan archive, so all the images are on the internet archive, or easily accessible from the 4chan api itselfThey were using an old url pattern for images, which is how I could do anything about this. Once the faggots wipe the drool off their faces and switch to the new pattern, I'm not sure how to combat thisWhy do they perform the most inefficient data collection strategy possible?
Goodnight sweet prince
>>110029894>I'm not sure how to combat thisPut all images in a hidden folder and set up a htaccess url rewrite that serves them from another link, like files are in /hiddenfolderasdfg/*.* and accessing /lalalakatamaridamacy/*.* will read the ones in the other folder.Change the site links so the image folder is served by one function, that function reading the media url prefix from say a config file (a non public one obviously).Set up a cron job to a function that changes the media directory name randomly and regenerates the htaccess url rewrite using it and call it every I dunno 30 minutes or something.So basically you are still just switching patterns, but you switch to a random one every x minutes.alternatively also block all datacenter IPs and ASNs in your firewall.
>>110030177both very good ideasthat will save a lot of bandwidth
>>110029894They perform the most efficient work for themselves—if you don't like people using your API then don't host it: you have only yourself to blame.
>>110030353no they don't, if you could comprehend my post stating easier solutions, you'd know it
>>110030177This is smart but are there better solutions than a cron?
>>110030699Better in what fucking way?
>>110030722Idk, cons seem disconnected from the codebase. Maybe it doesn't matter.
>>110030760>cons seem disconnected from the codebase.if you mean cronjob, if you can monitor bandwidth you should have access to the cron.But you could also just do something like "if cfg file was last changed >30 minutes ago, generate random key again", before you load said config. Or even store the last change time in Redis or something and look it up on page load.
>>110030760can insert some logic in the media route, which I plan to do
wait, having a map in nginx would be even cleaner
Zump
>>110029894Look up the ASN of the offending IP and block it, the entire ASN, their IP ranges are public.
>>110029894> thousands of users want to use your websiteWhats the issue?
I saw this analogy for AI scrapers on Reddit:It’s like how a roomba just cleans in random directions — it doesn’t need to be efficient because it’s automated.
>>110029894blacklist north america.Especially all the corpo subnets.You did blacklist Google, AWS, Cloudflare and Digital Ocean, right?
>>110030177Back in the day, people just hid honeypot URLs, that bots click on, but no human would ever find, and if they trigger it, they get auto-blocked for a day.At best, this honeypot URL would break the connection, to let it timeout, so the automated bot rolls for a different IP to try it again, getting blocked again, and so on.Weird how this is still the most effective method. No need to break URLs.>but they may have ten thousand IPsten thousand requests should be able to survive
>>110035462Yeah that sounds interesting but I wouldn't know how to link that up to the firewall, and nowadays even googlebot ignores things like nofollow links and I don't want to ban googlebot server wide.
>>110037736You link it to the firewall by making a fail2ban jail. Say the trigger is a certain response code or request or user agent from your server log. Fail2ban can monitor that, and block their IPs automatically.
>>110035346>on Reddit>—>’