My experience writing automated tests for a SPA
My experience writing automated tests for a SPA
我编写单页应用(SPA)自动化测试的经验
I think I managed to build quite a nice and interesting test suite recently; I’ll do my best to describe it in this post. 我认为我最近成功构建了一套相当不错且有趣的测试套件;我将尽力在这篇文章中对其进行描述。
It’s basically just a bunch of notes, and the code is not open-source, but I think these explanations can have more value than raw source code, especially if you want to adapt some of these ideas for one of your own projects. 这基本上只是一些笔记,代码也没有开源,但我认为这些解释比原始源代码更有价值,特别是如果你想将其中一些想法应用到你自己的项目中时。
The application
应用程序
Let’s start with a quick description of what we want to actually test, because as you can imagine, this is crucial for everything else. 让我们先快速描述一下我们实际要测试的内容,因为正如你所想象的那样,这对其他一切都至关重要。
Réécoute is a single-page web application (SPA), i.e., a website rendered with client-side JavaScript. It’s mainly an audio player, optimized for long recordings (typically 2 or 3 hours), with quite a few interactive features that couldn’t work with server-side rendering alone. It uses React, and the client-side JavaScript communicates with a single server by sending JSON over HTTP. Nothing special. Réécoute 是一个单页应用(SPA),即通过客户端 JavaScript 渲染的网站。它主要是一个音频播放器,针对长录音(通常为 2 或 3 小时)进行了优化,并具有许多无法仅靠服务器端渲染实现的交互功能。它使用 React,客户端 JavaScript 通过 HTTP 发送 JSON 与单个服务器进行通信。没什么特别的。
Now, how can we test that? Unlike a classic server-side rendered website, the complexity is split into two roughly equal parts between the backend and the client-side JavaScript. Ideally, we should test both together in a realistic fashion to exercise all the chatter between the client and the server. I’ve made the extreme choice of testing the app as a whole, using a real web browser. 那么,我们该如何测试它呢?与传统的服务器端渲染网站不同,其复杂性大致平分在后端和客户端 JavaScript 之间。理想情况下,我们应该以真实的方式将两者结合起来测试,以验证客户端和服务器之间的所有交互。我做出了一个极端的选择:使用真实的 Web 浏览器对整个应用程序进行测试。
The project also has a few backend-only tests that I won’t discuss here because there is really nothing special about them. 该项目还有一些仅针对后端的测试,我不会在这里讨论,因为它们确实没什么特别之处。
The main test suite
主测试套件
The main test suite is written with Playwright, running against a real web browser. It consists of about 20 files, each containing between 1 and 4 test cases. 主测试套件是使用 Playwright 编写的,在真实的 Web 浏览器中运行。它由大约 20 个文件组成,每个文件包含 1 到 4 个测试用例。
Regarding my personal preferences: I tend to write rather lengthy test cases that describe full user journeys, rather than small tests for individual steps. For an e-commerce website, for example, I would likely write a test that adds an item to the cart, signs up, goes to the checkout page, and actually purchases the item: it’s the most critical user journey for the business, and you do not want it to break. Of course, I also write smaller, specialized tests for things like sign-up, but IMO these tend to be somewhat less critical than the end-to-end flows. 关于我的个人偏好:我倾向于编写描述完整用户旅程的较长测试用例,而不是针对单个步骤的小型测试。例如,对于电子商务网站,我可能会编写一个测试,涵盖添加商品到购物车、注册、进入结账页面并实际购买商品的过程:这是业务中最关键的用户旅程,你绝对不希望它出错。当然,我也会为注册等功能编写较小、专门的测试,但在我看来,这些测试的重要性往往不如端到端流程。
Data isolation between tests
测试之间的数据隔离
Tests are not jailed in isolated environments, because: 测试并没有被限制在隔离的环境中,原因如下:
- When using something like Playwright, this is very complicated to achieve with database transactions;
- 当使用 Playwright 之类的工具时,通过数据库事务来实现这一点非常复杂;
- I could spawn an instance of the backend for each test, but it would be much slower, so I’m not going to do that;
- 我可以为每个测试启动一个后端实例,但这会慢得多,所以我不会这样做;
- Running each test on a tiny subset of the dataset does not help catch database queries that only slow down when there’s a lot of data;
- 在极小的数据集子集上运行每个测试,无法帮助发现那些只有在数据量大时才会变慢的数据库查询;
- It’s simply more complicated and less realistic than writing tests that run against the same database without disturbing other tests.
- 与编写在同一个数据库上运行且互不干扰的测试相比,这种做法不仅更复杂,而且真实性更低。
Basically, I write tests just like anyone would use the app in production: each test creates its own objects without relying on any existing data, never touches data it did not create, and never cleans up anything. Data just accumulates. This strategy works really well for apps like Réécoute, where nothing is actually public. 基本上,我编写测试的方式就像任何人在生产环境中使用该应用一样:每个测试都会创建自己的对象,而不依赖任何现有数据,从不触碰非自己创建的数据,也从不清理任何东西。数据只是在不断累积。这种策略对于像 Réécoute 这样没有任何公开内容的应用程序非常有效。
I use a few helper functions to create data (createUser, createBand, createSession, etc.). Note that I do not use before/after hooks at all. 我使用一些辅助函数来创建数据(createUser、createBand、createSession 等)。请注意,我完全不使用 before/after 钩子。
Mocks
模拟(Mocks)
The test suite uses two kinds of mocks: 测试套件使用两种类型的模拟:
-
Each external service has its own global mock: things like S3, Stripe, Twilio, etc. I tend to write one large, realistic mock for each of them. It’s much faster and more reliable than using actual third-party services, and it allows running the tests without an internet connection. These mocks are enabled by default and used across all tests.
-
每个外部服务都有自己的全局模拟:例如 S3、Stripe、Twilio 等。我倾向于为它们中的每一个编写一个大型、真实的模拟。这比使用实际的第三方服务要快得多且更可靠,并且允许在没有互联网连接的情况下运行测试。这些模拟默认启用,并用于所有测试。
-
For some complicated cases (emails and passkeys, especially), I have a few (2 or 3?) custom code paths enabled by test-only parameters/HTTP headers in API queries. These parameters are ignored by the backend in production builds.
-
对于一些复杂的情况(特别是电子邮件和通行密钥),我有一些(2 或 3 个?)自定义代码路径,通过 API 查询中的测试专用参数/HTTP 标头启用。这些参数在生产构建中会被后端忽略。
(I really hate when a test suite forces you to write custom mocks for every single test…) (我真的很讨厌测试套件强迫你为每一个测试编写自定义模拟……)
The most complex mock I wrote for this project is probably the one for passkeys: I couldn’t get actual passkeys to work in headless Chromium, so I hacked together a fake client around the passkey crate. But it is very specific and I am not very proud of it, so I won’t go into details here! 我为这个项目编写的最复杂的模拟可能是通行密钥(passkeys)的模拟:我无法让实际的通行密钥在无头 Chromium 中工作,所以我围绕 passkey crate 拼凑了一个伪客户端。但它非常具体,我对此并不感到自豪,所以这里就不详细介绍了!
Speed
速度
As you can imagine, browser automation is much slower than simply parsing HTTP response bodies, so without parallelism it can quickly become unmanageable. This is why Playwright runs test files in parallel by default. With Réécoute, I went a step further by enabling fullyParallel in the Playwright config, so tests within the same file also run concurrently. However, the most important factor here is the app itself, since a test suite can’t be more efficient than the app being tested! To give you an idea, the Playwright suite currently completes in just over 20 seconds on my fanless M3 MacBook Air. 正如你所想象的那样,浏览器自动化比简单地解析 HTTP 响应体要慢得多,因此如果没有并行处理,它很快就会变得难以管理。这就是为什么 Playwright 默认并行运行测试文件的原因。对于 Réécoute,我更进一步,在 Playwright 配置中启用了 fullyParallel,因此同一文件内的测试也会并发运行。然而,这里最重要的因素是应用程序本身,因为测试套件的效率不可能超过被测应用程序!为了让你有个概念,Playwright 套件目前在我无风扇的 M3 MacBook Air 上只需 20 多秒即可完成。
Also, Playwright supports all major web browsers and runs your tests across 3 or 4 of them by default. I changed the settings to only use Chromium: modern browsers behave very similarly, this makes the suite 3 to 4 times faster to run, and it is nearly as effective. 此外,Playwright 支持所有主流 Web 浏览器,并默认在 3 或 4 个浏览器中运行测试。我更改了设置,仅使用 Chromium:现代浏览器的行为非常相似,这使得套件的运行速度提高了 3 到 4 倍,而且效果几乎一样好。
Reliability
可靠性
Here’s the main downside to browser testing, especially for SPAs: because we are testing an entire app and an entire browser, it’s difficult to make tests perfectly reliable. Yet with a large test suite, you must have high reliability, because the more tests you have, the less reliable the overall suite becomes, and re-running failed suites is expensive. 这是浏览器测试的主要缺点,尤其是对于 SPA:因为我们测试的是整个应用程序和整个浏览器,所以很难使测试达到完美的可靠性。然而,对于大型测试套件,你必须具备高可靠性,因为测试越多,整体套件就越不可靠,而且重新运行失败的套件成本很高。
There is a trick here—it’s not pretty, but it works well: Playwright has a retries option, which I set to 2 in CI. When a test fails, it is retried individually up to 2 times. In practice, tests in Réécoute’s suite rarely fail and retry. I could probably eliminate flakes entirely if I spent a few hours on it, but I’m not sure it’s worth the effort right now. 这里有一个技巧——虽然不优雅,但效果很好:Playwright 有一个重试选项,我在 CI 中将其设置为 2。当测试失败时,它会单独重试最多 2 次。在实践中,Réécoute 套件中的测试很少失败和重试。如果我花几个小时,可能完全可以消除不稳定性(flakes),但我目前不确定这是否值得。
In fact, the main issue I faced with reliability was related to dual server-side/client-side rendering, in… 事实上,我在可靠性方面遇到的主要问题与双重服务器端/客户端渲染有关,在……