Remove Watermark Without Uploading: How We Built an In-Browser AI Editor
Remove Watermark Without Uploading: How We Built an In-Browser AI Editor
无需上传即可去除水印:我们是如何构建浏览器端 AI 编辑器的
Most online watermark removers ask you to hand over your photo first and ask questions later. We built ClearPix the other way around: the AI models come to your device, your files never leave it, and inference runs entirely in your browser tab. This post is the architecture overview of how that actually works in production — the ONNX Runtime Web + WebGPU inference stack, the model distribution pipeline (R2 CDN, multi-mirror fallback, SHA-256 self-healing cache), the places it broke on us, and how we keep the privacy claim structural instead of rhetorical. No marketing math; every number below comes from our own engineering logs. 大多数在线去水印工具都会要求你先交出照片,然后再谈其他。我们构建 ClearPix 的思路恰恰相反:AI 模型被发送到你的设备上,你的文件永远不会离开设备,推理过程完全在你的浏览器标签页中运行。本文概述了其在生产环境中的架构实现——包括 ONNX Runtime Web + WebGPU 推理栈、模型分发流水线(R2 CDN、多镜像回退、SHA-256 自愈缓存)、我们遇到的故障点,以及我们如何确保隐私承诺是结构性的而非口头上的。文中没有营销辞令,以下所有数据均来自我们的工程日志。
Does a watermark remover upload your photos? Usually, yes. The upload question is not paranoia — it is the default architecture of the category. Nearly every “free online watermark remover” is a thin frontend over a server-side inference API. Your image goes up, gets processed on someone else’s GPU, and comes back down. Whether the privacy policy says the file is deleted afterward, you have no way to verify it, and neither do we. The structural fact is that your media crossed the wire. 去水印工具会上传你的照片吗?通常会。关于上传的担忧并非多疑——这是该类产品的默认架构。几乎每一个“免费在线去水印工具”都只是服务器端推理 API 的一个轻量级前端。你的图像被上传,在别人的 GPU 上处理,然后再传回来。无论隐私政策是否声称文件随后会被删除,你都无法验证,我们也无法验证。结构性的事实是,你的媒体文件确实经过了网络传输。
We kept running into this as a genuine user anxiety, and we shared it. Photos with watermarks are often personal — screenshots, ID documents, family pictures, client work. “Trust us, we delete it” is a weak foundation for a tool people use on sensitive files. So we asked the harder question: can the inference just happen on the user’s machine? The answer, in 2026, is yes — if you are willing to solve the distribution and reliability problems that server-side tools offload to a data center. That trade is the whole story of ClearPix. 我们不断遇到用户对此表达真实的焦虑,我们也感同身受。带有水印的照片通常涉及个人隐私——截图、身份证件、家庭照片、客户作品等。“请相信我们,我们会删除它”对于处理敏感文件的工具来说,是一个脆弱的信任基础。因此,我们提出了一个更具挑战性的问题:推理能否直接在用户的机器上完成?在 2026 年,答案是肯定的——前提是你愿意解决那些服务器端工具通过数据中心规避掉的分发和可靠性问题。这种权衡正是 ClearPix 的核心故事。
The architecture in one paragraph: ClearPix is a static site. There is no processing backend. When you open a tool page, your browser downloads ONNX model weights (once, then cached forever), loads them into ONNX Runtime Web, and runs inference on WebGPU where available with a WASM fallback. Your image or video is decoded locally, processed locally, re-encoded locally, and never serialized into a network request. The only bytes that travel are the models coming down to you, and a small set of aggregate analytics events that contractually cannot contain your media. That single paragraph hides about 90% of the engineering. Let’s unpack it. 架构概览:ClearPix 是一个静态网站,没有后端处理逻辑。当你打开工具页面时,浏览器会下载 ONNX 模型权重(仅下载一次,随后永久缓存),将其加载到 ONNX Runtime Web 中,并在支持 WebGPU 的情况下运行推理,否则回退到 WASM。你的图像或视频在本地解码、处理、重新编码,绝不会被序列化为网络请求。唯一传输的数据是下载到你本地的模型文件,以及一小部分在协议上绝不包含你媒体内容的聚合分析事件。这一段话隐藏了约 90% 的工程细节,让我们来拆解它。
In-browser AI image processing: the runtime layer. We run all neural inference through ONNX Runtime Web (ort 1.30.x), loaded via dynamic import so its ~1 MB parsed weight stays out of the initial page bundle — Core Web Vitals matter even for tools. Session creation prefers the WebGPU execution provider and falls back to WASM. WebGPU is worth real money here: when we enabled the WebGPU EP (ort 1.30+), our inpainting step went from 2100–3500 ms to roughly 400 ms per region in our benchmarks. But “session created” does not mean “inference works” — we have seen WebGPU sessions deadlock silently. So the session runs behind a watchdog, and a dead WebGPU session falls back to WASM rather than hanging the tab. 浏览器端 AI 图像处理:运行时层。我们通过 ONNX Runtime Web (ort 1.30.x) 运行所有神经网络推理,并通过动态导入加载,使其约 1 MB 的解析权重不包含在初始页面包中——即使对于工具类产品,核心网页指标(Core Web Vitals)也很重要。会话创建优先选择 WebGPU 执行提供程序,并回退到 WASM。WebGPU 在这里价值巨大:在我们的基准测试中,启用 WebGPU EP (ort 1.30+) 后,我们的修复(inpainting)步骤从每个区域 2100–3500 毫秒缩短至约 400 毫秒。但“会话创建成功”并不意味着“推理正常工作”——我们曾遇到 WebGPU 会话静默死锁的情况。因此,会话运行在看门狗(watchdog)机制之后,一旦 WebGPU 会话死锁,系统会回退到 WASM,而不是让标签页卡死。
The fallback chain is not an afterthought; every production stall we have ever shipped was closed by one. Even the runtime binaries need their own distribution strategy. The jsep.wasm file (28 MB) is a hard dependency of the WebGPU EP — if it fails to load, ORT silently drops back to CPU inference and the user just experiences a mysteriously slow tool. We serve the ORT runtime from jsDelivr first (measured ~7× the throughput of our own R2 custom domain for these files) with a 5-second HEAD probe, falling back to our self-hosted R2 copy when the CDN is unreachable — which happens in some network environments. 回退链并非事后补救;我们发布过的每一个生产环境故障都是通过它解决的。即使是运行时二进制文件也需要自己的分发策略。jsep.wasm 文件(28 MB)是 WebGPU EP 的强依赖——如果加载失败,ORT 会静默回退到 CPU 推理,用户只会感到工具莫名其妙地变慢。我们优先从 jsDelivr 提供 ORT 运行时(经测量,其吞吐量约为我们 R2 自定义域名的 7 倍),并配合 5 秒的 HEAD 探测,当 CDN 不可达时(在某些网络环境下会发生),则回退到我们自托管的 R2 副本。
A client-side watermark remover lives or dies by its model pipeline. Here is the uncomfortable truth of client-side AI: you are now a software distribution company. A server-side tool updates weights by deploying a container. We have to push tens of megabytes of model files to arbitrary browsers on arbitrary networks, keep them intact, and never strand a user. Distribution. Models live in a Cloudflare R2 bucket behind a custom domain on our zone. This was a measured decision, not a default: direct r2.dev URLs gave us 75–110 KB/s, while the same-zone custom domain delivered 2.3 MB/s — a 20–30× difference that turns a model download from a coffee break into a pause. 客户端去水印工具的成败取决于其模型流水线。客户端 AI 有一个令人不安的事实:你现在实际上是一家软件分发公司。服务器端工具通过部署容器来更新权重,而我们必须将数十兆字节的模型文件推送到各种网络环境下的各种浏览器中,确保文件完整,且绝不能让用户掉队。关于分发:模型存储在 Cloudflare R2 存储桶中,并配置了我们区域下的自定义域名。这是一个经过深思熟虑的决定,而非默认设置:直接使用 r2.dev URL 的速度仅为 75–110 KB/s,而同区域的自定义域名可达 2.3 MB/s——20 到 30 倍的性能差异,让模型下载从“喝杯咖啡的时间”缩短为“短暂的停顿”。
Permanent caching. Once downloaded, model bytes go into IndexedDB and are reused forever — repeat visits load from the device, and the tool works offline. Two details cost us real debugging time. First, big models are written in 32 MB chunks (above a 64 MB threshold), because a single structured clone of a 198 MB model pushed the renderer process into OOM. Second, the cache is self-healing: we verify the SHA-256 of a cached model on every hit (~100 ms for a 67 MB file), because a download interrupted mid-write would otherwise poison a cache that never expires — permanently bricking that user’s tool with no visible error. 永久缓存。一旦下载,模型字节就会进入 IndexedDB 并永久复用——再次访问时直接从设备加载,工具甚至可以离线工作。有两个细节花费了我们大量的调试时间。首先,大模型被拆分为 32 MB 的块进行写入(超过 64 MB 阈值),因为对 198 MB 模型进行单次结构化克隆会导致渲染进程内存溢出(OOM)。其次,缓存具有自愈功能:我们在每次命中时都会验证缓存模型的 SHA-256(67 MB 文件耗时约 100 毫秒),因为如果下载在中途被中断,可能会污染永不过期的缓存,从而导致用户的工具永久失效且没有任何可见错误。
Mirrors and fallback specs. Each model spec carries an ordered mirror list, a pinned SHA-256, and an optional fallback spec with a different weight format (fp16 → fp32) and an independent cache key. The first mirror whose bytes hash-match wins; a stalled connection (no bytes for 30 seconds) is aborted and handed to the next mirror; if every mirror fails, the fallback spec takes over. 镜像与回退规范。每个模型规范都包含一个有序的镜像列表、一个固定的 SHA-256 哈希值,以及一个可选的、具有不同权重格式(fp16 → fp32)和独立缓存键的回退规范。第一个哈希匹配成功的镜像胜出;连接停滞(30 秒无数据传输)会被中止并切换到下一个镜像;如果所有镜像都失败,则启用回退规范。