The entire fucking point of the Web was to make information as easily-accessible as possible, structured and semantically tagged, and consumable by humans and further machine transformation alike. “Scraping” is facilitated by design!
Using Javascript to deliberately break that is evil and every programmer who participates it is a piece of shit. No exceptions.
That’s true but they probably didn’t account for AI data scrapers ramfucking your server so they could steal all the value you assembled for general consumption and serve it themselves for profit.
The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API. The “ramfucking” is caused by the attempt to block bots; it is entirely self-inflicted.
Remember, it’s all our content to begin with and Reddit does not have any right to try to lock it up for itself.
That doesn’t mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.
so why was I getting hit with over 1,400,000 request a day to the web URI and not the API by some bot farm in China the other week. They were also hitting other lemmy instances.
I blocked the fuckers, no qualms at all.
Even if they were using the API they were not being nice about their shit.
Well they had an API, but…
The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API
That’s simply not true. These bots are essentially DDOSing the entire internet, API or not.
Okay, if efficient APIs existed and they weren’t incompetently failing to use them, it wouldn’t be a problem. Happy now?
(I should’ve addressed that in my previous comment, as I was aware of how one of the Lemmy instances was taken down by scrapers the other day despite the fact that they could easily get all the content simply by consuming ActivityPub directly. But I was naively hoping it wouldn’t be necessary because, as you can see from this text, it would’ve cluttered up my writing with double the words.)
Reddit does have RSS feeds
But we all know AI companies are unethically scraping and selling shit back to us right? I just really need people to acknowledge that.
It is the “selling shit back to us” specifically, not the “scraping,” that’s the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue “no.”
Hmffh, anti-copyright. After all, every view is a copy to your machine. Just information being free.
even logged in I can’t access old reddit anymore.
if you use the reddit enhanced suite extension you can get redirected
If I wasn’t permabanned I might bother trying work arounds.
I just checked my old.reddit login on the desktop… still working just fine today. I don’t think you need work-arounds, you just need to access it without any work-arounds.
working for me
How dare you question King Steven the Turd, Greediest of Pigboys? If he proclaims HTML to be unsafe, it must be so. That’s a King’s job, to tell the Landed Gentry how to behave…
I’m not forgetting or forgiving that shit either.
I get blocked and asked to login to reddit no matter if it’s old or new. safereddit still works though so I’m using that for now.
afaik supposedly it was because a large chunk of the bot network went through the old.reddit portal over the standard reddit portal.
Only because the bots were already set up to do it that way. It’s quicker to use what exists when it works.
I am so glad I’m off reddit. Even though I could use their help on a great number of subjects. Neverheless, because Isreal I can’t use reddit apparently. Not permabanned yet but they are on my shit with fake violations like right away now. I abandon them after a 2nd violation, go through one every 3 to 6 months when using it.
Every single company that goes public gets worse, reddit will be no exception. Those craven amoral cockscum are not to be trusted, nor patronized with our words they can use for their own benefit.






