So Reddit has decided that plain HTML is unsafe

So Reddit has decided that plain HTML is unsafe

Reddit 决定:纯 HTML 是不安全的

Reddit-The-Company If you don’t know Reddit, it basically is the host of many popular forums. And like any company which encourages you to “come for the cats [and] stay for the empathy,” Reddit seems to be in the business of extracting as much value as it can from said forums without completely destroying them. After all, simply fostering community is not a noble enough goal for the New Tech, and fortunately for Reddit, genuinely human-generated data is now gold in the LLM Age. You are welcome to read about the last time they decided to pluck the metaphorical liver from their communities.

Reddit 公司:如果你不了解 Reddit,它基本上是许多热门论坛的宿主。像任何鼓励你“为猫而来,为共情而留”的公司一样,Reddit 的业务似乎是在不彻底摧毁论坛的前提下,尽可能地从中榨取价值。毕竟,对于“新科技”而言,单纯培育社区并不是一个足够崇高的目标;幸运的是,对于 Reddit 来说,真正由人类生成的数据在 LLM(大语言模型)时代已成了黄金。欢迎阅读他们上次决定从社区中“掏心挖肺”时的往事。

Reddit-The-Search-Results While I no longer wish to engage with Reddit, I still visit it occasionally, especially in the LLM Age. This is because appending site: reddit.com to a search query is basically a surefire way to find results written by genuine humans. Which, just to be extremely clear, I still find desirable. Now behind a login… sort of After doing such a query yesterday, to my absolute delight I was greeted with I guess this makes me old, but I use the original frontend for Reddit, old.reddit.com. I’m going to try really hard not to preach about why it’s a better frontend, but that’s all it is! A design for Reddit.

Reddit 搜索结果:虽然我不再想参与 Reddit 的互动,但我偶尔还是会访问它,尤其是在 LLM 时代。这是因为在搜索查询中加上 site:reddit.com 基本上是找到由真实人类撰写的结果的万无一失的方法。需要明确的是,我仍然认为这很有价值。现在它被放在了登录墙之后……某种程度上。昨天进行这样的查询后,我非常“惊喜”地收到了登录提示。我想这说明我老了,但我确实在使用 Reddit 的原始前端 old.reddit.com。我会尽量不去说教为什么它是一个更好的前端,但事实就是如此!它只是 Reddit 的一种设计。

Look, I know I’m a fringe user. I use Firefox. I noscript! I am no stranger to being forced off a product that worked just fine because someone decides to no longer support the two people who still use it. It happened to my phone of 8 years, which works fine by the way, but is on too outdated of an OS. May it rest in peace in its tiny glory. It happened to my tablet of 10 years, which works even better than my phone and holds a charge like champ, for the same reason. It happened to the API for Stack Exchange used by my copy of their outdated app long removed from the app store. This one hurt me the most. Where I feel like things get personal here is that Reddit is saying that this is all To keep Reddit safe Keep Reddit safe from whom exactly? Me? My desire for knowledge?? Safety is what exactly?

听着,我知道我是个边缘用户。我用 Firefox,我用 NoScript!我对于那种“因为某人决定不再支持仅存的两个用户,而被迫放弃一个运行良好的产品”的情况并不陌生。我用了 8 年的手机就遭遇了这种情况,顺便说一句,它运行得很好,但操作系统太老了。愿它在小巧的荣光中安息。我用了 10 年的平板电脑也因为同样的原因遭遇了这种情况,它比我的手机运行得更好,电池续航依然强劲。我那份早已从应用商店下架的旧版 Stack Exchange 应用所使用的 API 也遭遇了同样的情况。这最让我心痛。我觉得这件事之所以让我感到被针对,是因为 Reddit 声称这一切都是为了“保持 Reddit 安全”。到底是为了防谁?我吗?我那对知识的渴望吗??所谓的“安全”到底是什么?

So let’s see what they have to say on this matter by going to the announcement, which of course I didn’t see because I don’t read Reddit anymore: https://old.reddit.com/r/modnews/comments/1ujtebf/logging_in_to_use_old_reddit/. Hope you’re logged in. Old Reddit’s logged-out experience is a significant source of abusive scraping and automated traffic on the platform. Hmmm, OK. But then why is New Reddit still accessible logged out? Oh, someone asked that. [Question]: What’s so different about new reddit that people don’t try to scrape that? Seems to me like it would be better to just implement that on old reddit too. Besides, won’t this just cause people to try and scrape new reddit? [Admin reply]: I was about to type an answer but just saw u/Nestramutat- gave a really eloquent answer in another comment! [The comment (snipped)]: … To your first question, the shape of malicious traffic is always changing. It’s going to be a constant cat and mouse game as you ban one method, a new one gets developed. It’s easy to see abusive traffic in hindsight, but it’s harder to pre-emptively block it. Given that they’re claiming Old Reddit doesn’t have the modern security stack, this is likely proving to be an even greater challenge…

让我们去看看他们的公告是怎么说的,当然我之前没看到,因为我不再阅读 Reddit 了:https://old.reddit.com/r/modnews/comments/1ujtebf/logging_in_to_use_old_reddit/。希望你已经登录了。公告称:“旧版 Reddit 的未登录体验是平台上滥用抓取和自动化流量的重要来源。” 嗯,好吧。但为什么新版 Reddit 在未登录状态下仍然可以访问?哦,有人问过这个问题。[提问]:新版 Reddit 有什么不同,以至于人们不去抓取它?在我看来,直接在旧版 Reddit 上实现同样的机制会更好。此外,这难道不会导致人们转而去抓取新版 Reddit 吗?[管理员回复]:我正要打字回答,但刚看到 u/Nestramutat- 在另一条评论中给出了非常精彩的回答![评论节选]:……对于你的第一个问题,恶意流量的形态总是在变化。这是一个持续的猫鼠游戏,当你封禁一种方法时,新的方法就会被开发出来。事后发现滥用流量很容易,但预先阻止它却很难。鉴于他们声称旧版 Reddit 没有现代安全栈,这很可能是一个更大的挑战……

So it doesn’t have “the modern security stack.” Now I may not have really earned my Full Stack stripes, but I can right click and select Inspect Element so I’d say I’d say I’m qualified enough to see why. Old versus New Old Reddit Let’s start by — begrudgingly — logging in to see what is so insecure about old.reddit.com. I’ll use their announcement thread to test. Ahhh, so much nicer. Let’s check what Old Reddit is doing that makes it so insecure. The best I can guess is that their precious, precious user-created content is available in plain HTML, since that’s basically all Old Reddit does: you don’t even need JS unless you want to load more comments (ask me how I know). It DLs about 1 megabyte and sends about half a megabyte. It’s not shown there, but the page’s HTML itself comprises most of the response. I’ve certainly seen worse, but what’s with the load time? GitHub loads its massive payload about 4x as fast (relative to size). Oh… 2 whole seconds of waiting for a reply. Smells of rate-limiting. Or was Reddit always this slow?

所以它没有“现代安全栈”。虽然我可能还没拿到全栈开发的“勋章”,但我会右键点击并选择“检查元素”,所以我认为我有足够的资格去看看原因。旧版与新版:旧版 Reddit。让我们先——不情愿地——登录进去,看看 old.reddit.com 到底哪里不安全。我将用他们的公告帖来测试。啊,好多了。让我们看看旧版 Reddit 到底做了什么让它变得如此“不安全”。我能猜到的最好解释是,他们那宝贵的用户生成内容是以纯 HTML 形式提供的,因为这基本上就是旧版 Reddit 的全部工作方式:你甚至不需要 JavaScript,除非你想加载更多评论(问我怎么知道的)。它下载约 1MB 数据,上传约 0.5MB。虽然图中没显示,但页面本身的 HTML 占据了响应的大部分。我当然见过更糟的,但加载时间是怎么回事?GitHub 加载其庞大的内容速度大约是它的 4 倍(相对于大小)。哦……整整 2 秒的等待响应。闻起来像是速率限制。还是说 Reddit 一直都这么慢?

Let’s load some more comments. (Reddit never loads all of the comments initially) Well that was nice and lean, and pretty snappy. I don’t see my secrets being sniffed. All I got here is that Old Reddit is a pretty normal webpage, which I guess makes it insecure in comparison to… New Reddit Let’s see why New Reddit is so much better. In case you have unrealistic expectations, let me right them: New Reddit will not load anything more than the post itself without Javascript (JS). That’s probably what makes it more secure. There’s a lot loading here (about 5x Old Reddit), and this is why my analysis gets rather unscientific. Rather than try to get around the “security things” (whatever that means), I instead tried to do the bare minimum necessary to fetch the content. In browser — I did not feel like writing a scraper. This led to me basically blocking all requests to domains (including reddit.com) except for www.redditstatic.com/js/concat www.reddit.com/svc/shreddit/more-comments/ www.reddit.com/svc/shreddit/comment/ When you do this, the page loads a lot less, but it does load. When you click to load more comments, it spins forever, but I inspected the request fired off and it did get a response with comment text. So as far as I can gather, simply running the JS on the page is sufficient to get enough information to get comments. So I guess that’s what’s stopping the scrapers? Executing Javascript? Just for the heck of it, let’s load some more comments without the request filter. Well, that’s certainly less lean than Old Reddit. In my (again, unscientific) experimenting, I reloaded the page several times and tried to load comments and replies and didn’t get a…

让我们加载更多评论。(Reddit 从来不会一开始就加载所有评论)嗯,这很简洁,而且相当快。我没看到我的秘密被窃取。我在这里得到的一切结论是,旧版 Reddit 就是一个非常普通的网页,我想这使得它在对比之下显得“不安全”……新版 Reddit。让我们看看为什么新版 Reddit 好得多。以防你有不切实际的期望,让我来纠正一下:没有 JavaScript,新版 Reddit 除了帖子本身什么都不会加载。这可能就是它更安全的原因。这里加载的东西很多(大约是旧版 Reddit 的 5 倍),这就是为什么我的分析变得相当不科学。与其试图绕过那些“安全机制”(不管那意味着什么),我尝试只做获取内容所需的最低限度操作。在浏览器中——我不想写爬虫。这导致我基本上屏蔽了除 www.redditstatic.com/js/concatwww.reddit.com/svc/shreddit/more-comments/www.reddit.com/svc/shreddit/comment/ 之外的所有域名请求。当你这样做时,页面加载的内容少了很多,但它确实能加载。当你点击加载更多评论时,它会一直转圈,但我检查了发出的请求,它确实得到了包含评论文本的响应。所以据我所知,仅仅运行页面上的 JS 就足以获取评论信息。所以我想这就是阻止爬虫的原因?执行 JavaScript?为了好玩,让我们在没有请求过滤器的情况下加载更多评论。嗯,这肯定比旧版 Reddit 臃肿得多。在我的(再次声明,不科学的)实验中,我重新加载了几次页面并尝试加载评论和回复,结果并没有得到……