Bug Blindness

Bug Blindness

I used to wonder why I see so many more bugs than most people. I easily observe hundreds to thousands of bugs per week and nothing seems to work, but most people I talk to don’t see anything like this. For a long time, I thought this had something to do with how I use computers but, over time, I’ve realized that it’s mostly that people are hitting the same bugs and don’t notice. 我曾经纳闷,为什么我发现的程序错误(Bug)比大多数人多得多。我每周能轻松观察到成百上千个 Bug,感觉什么东西都运行不正常,但我交流过的大多数人却对此毫无察觉。很长一段时间里,我以为这与我使用电脑的方式有关,但随着时间的推移,我意识到其实是因为大家都在遇到同样的 Bug,只是没注意到而已。

If you’re not a programmer, that’s probably a better way to see the world, but I think curing quality/bug blindness is helpful for programmers. I’ve done this with a lot of friends and acquaintances (just by pointing out bugs). After a few weeks, people who are so inclined tend to start noticing bugs as well. 如果你不是程序员,这或许是看待世界的一种更好的方式,但我认为,治愈“质量/Bug 盲区”对程序员来说很有帮助。我曾带着许多朋友和熟人做过练习(仅仅是通过指出 Bug)。几周后,那些有心的人往往也开始能注意到 Bug 了。

Because I notice these kinds of things, I’ve had multiple jobs where directors/VPs/execs/etc. sometimes ask me to evaluate something when they want an actual opinion from someone who is relatively likely to notice issues (and fix them or drive fixes for them if necessary). Sometimes I won’t find any issues (there are likely issues that just aren’t the kind I notice). More often, I find issues that fall somewhere from “mild” to “moderate”. And, sometimes, the issues are severe, to the point where one might even say the thing actually doesn’t work. 因为我能注意到这些问题,我曾多次在工作中被总监、副总裁或高管等要求评估产品,因为他们想听听那些更有可能发现问题(并在必要时修复或推动修复)的人的真实意见。有时我找不到任何问题(可能确实存在问题,只是不属于我能注意到的类型)。更多时候,我发现的问题处于“轻微”到“中等”之间。有时,问题甚至严重到可以说产品根本无法使用。

I find this last category a bit mysterious, as when I look up discussions on how the thing got into this state, there’s usually a stream of internal comments indicating that the thing is great, it works well, etc., but when I open up the thing and try it, it’s in a state where the thing only works if you do quite a few non-intuitive workarounds. More likely than not, not only would a normal user not be able to use the thing, they’d have such a hilariously/infuriatingly bad experience that they’d tell their friends. 我发现最后这一类情况有点神秘,因为当我查看关于产品为何会变成这样的讨论时,内部评论往往充斥着“产品很棒”、“运行良好”之类的溢美之词。但当我打开并试用时,却发现它必须通过许多反直觉的变通方法才能勉强运行。通常情况下,普通用户不仅无法使用,甚至会因为体验极其糟糕(既可笑又令人愤怒)而向朋友吐槽。

I’ve had this post in mind for maybe a decade or so, but I was hesitant to write it up because, in the back of my mind, I always wondered if I’m somehow triggering weird corner case behavior most users don’t hit without realizing it. But after seeing more and more cases where the product launches and falls flat on its face because users run into the exact same issues I saw, I don’t think that, in general, I’m hitting bugs because I’m doing unusual things a normal user wouldn’t do. If a product seems severely flawed when I use it, it probably is. And with the magic of LLMs, nowadays, I can even have LLMs act like normal users in a lot of ways and show that the issues reproduce across many different scenarios. 我构思这篇文章大概有十年了,但一直犹豫要不要写,因为我内心深处总在怀疑:是不是我无意中触发了大多数用户遇不到的奇怪边界情况?但看到越来越多的产品在发布后因为用户遇到了和我一模一样的问题而惨败,我认为我并不是因为做了普通用户不会做的异常操作才遇到 Bug。如果一个产品在我使用时显得漏洞百出,那它很可能就是真的有问题。如今,借助大语言模型(LLM)的魔力,我甚至可以让 LLM 在很多方面模拟普通用户,并证明这些问题在多种不同场景下都会复现。

A few examples I don’t want to give any specific examples where it was my job to see how well the thing worked because, even if the internal examples are meant in a constructive, blameless, way, they may not always read that way when re-posted externally, so I’ll give a few less interesting and less well supported “random” examples. 关于例子,我不想列举任何我曾受雇评估其性能的具体案例,因为即使内部讨论是出于建设性且不指责的目的,一旦被发布到外部,读起来可能就不是那么回事了。因此,我将提供几个不太引人注目、证据也不那么充分的“随机”例子。

A while ago, I wrote up the results of some web search queries and found poor results from Google, Bing and Kagi. In general, the major search engines failed to return good results for the queries and returned pages full of low-quality SEO spam as well as some sites that were actually scams. BTW, on the scale mentioned above, I would consider this “moderate” and not “severe” (severe would be something like, the search engine returns 500 errors half the time, the majority of results are scams, etc; my bar for severe is that a normal user likely won’t be able to use the thing at all, not that they have a bad experience). 前段时间,我记录了一些网络搜索查询的结果,发现 Google、Bing 和 Kagi 的表现都很差。总的来说,主流搜索引擎未能针对查询返回高质量结果,反而充斥着低质量的 SEO 垃圾信息,甚至还有一些诈骗网站。顺便提一下,按照上面提到的标准,我认为这属于“中等”而非“严重”(严重的情况应该是:搜索引擎有一半时间返回 500 错误,或者大部分结果都是诈骗等;我定义的“严重”门槛是普通用户根本无法使用,而不仅仅是体验糟糕)。

Almost nobody objected to my characterization of Google and Bing search results, but people told me that I was wrong about Kagi. In some cases, people sent me their actual search results. In every such case, the search results did not contain a good result that I could see (e.g., for the seasonal forecast query, the search failed to return an up-to-date seasonal forecast) and was full of SEO spam. In one case, a person passed me both their list of Kagi filters as well as the search results they got without making claims that the results were good or bad, but people generally insisted the results were good even though the results both failed to link to a useful result and were full of spam except in cases where the user did something like pin GitHub to the top of their results, which worked for the queries where the goal was to download software that’s hosted on GitHub, but of course completely fails for the other queries from the post. 几乎没有人反对我对 Google 和 Bing 搜索结果的评价,但有人说我对 Kagi 的看法错了。在某些情况下,人们发给我他们实际的搜索结果。在每一个案例中,我都没能从中看到好的结果(例如,针对季节性预报的查询,搜索未能返回最新的预报),而且结果中充满了 SEO 垃圾信息。有一次,一个人发给我他的 Kagi 过滤器列表以及他得到的搜索结果,并没有断言结果好坏,但人们普遍坚持认为结果很好——尽管这些结果既没有指向有用的信息,又充满了垃圾内容。除非用户采取了诸如将 GitHub 置顶之类的操作,这对于下载托管在 GitHub 上的软件的查询确实有效,但对于文章中提到的其他查询则完全失效。

In the abstract, I get that people who are fans of things tend to be blind to the thing’s faults. For example, since I bought a Volvo after seeing how they do in out-of-sample crash tests, I sometimes search for answers to my questions on Volvo car forums. For well over a decade, the reliability data that exists (and I think this is backed up by the anecdotal experience that mechanics who work on Volvos have) is that Volvo reliability is mediocre to poor, but of course Volvo forums are full of people who insist that Volvos are among the most reliable cars and that the data are all wrong. 抽象地讲,我理解人们作为某事物的粉丝,往往会对该事物的缺陷视而不见。例如,我因为看过沃尔沃在样本外碰撞测试中的表现而购买了它,所以我有时会在沃尔沃汽车论坛上搜索问题的答案。十多年来,现有的可靠性数据(我认为这得到了维修沃尔沃的技师们的经验支持)显示,沃尔沃的可靠性处于中等偏下水平,但沃尔沃论坛里当然充斥着坚称沃尔沃是最可靠汽车之一、并认为所有数据都是错误的人。

An example that might be more central to the topic is Blackboard (the course management software). Back when it was the most widely used software by universities for coursework, the software was widely disliked by both students and professors. I think it would be fair to say that it was the most widely disliked software in my social circles (there was more strongly disliked software, like Visual Source Safe, but any more strongly disliked software wasn’t widely used enough to be the most widely disliked overall). The Wikipedia page notes Blackboard had become “one of the most disliked — even detested — companies in education.” as well as In December 2011, Fast Company reported that 93% of respondents to the Amplicate customer opinion survey “hate” the company. Back when I was much younger and had less of a filter, I ran into someone who worked at Blackboard and, without thinking, I stupidly blurted out something like “what’s it like to work on this software that so many people dislike?”. 一个可能更贴近主题的例子是 Blackboard(课程管理软件)。在它被大学广泛用于课程作业时,学生和教授们都非常讨厌这款软件。我认为可以公平地说,它是我社交圈中最不受欢迎的软件(虽然有更令人讨厌的软件,比如 Visual Source Safe,但那些软件的使用范围不够广,不足以成为“最广泛被讨厌”的软件)。维基百科页面指出,Blackboard 已成为“教育界最不受欢迎,甚至最令人憎恶的公司之一”。2011 年 12 月,《快公司》(Fast Company)报道称,在 Amplicate 客户意见调查中,93% 的受访者“讨厌”该公司。在我年轻得多、说话还没那么顾忌的时候,我遇到过一个在 Blackboard 工作的人,我不假思索地脱口而出:“为这款被这么多人讨厌的软件工作是什么感觉?”