We Open Sourced Our Web Crawler. Here's How to Run a Node.
We Open Sourced Our Web Crawler. Here’s How to Run a Node.
我们开源了网页爬虫,以下是运行节点的方法
The crawler behind cl0q.com is now public. Anyone can read the code, run a node, and help index the open web. https://codeberg.org/cl0qsearch/crawler-node cl0q.com 背后的爬虫现已公开。任何人都可以阅读代码、运行节点,并帮助索引开放网络。https://codeberg.org/cl0qsearch/crawler-node
Why open source it? Search shouldn’t be a black box. The big engines don’t tell you what they crawl, how they rank, or what they do with the data. We think that’s wrong. cl0q is an independent search engine with its own crawler and index. We’ve mapped 38.5M domains so far, working toward 100M. Now the crawler itself is open for anyone to inspect, run, or improve. If you’ve ever wondered “what does a real web crawler look like under the hood,” this is your answer. Every line is readable. No magic. 为什么要开源?搜索不应该是一个黑箱。大型搜索引擎不会告诉你它们抓取了什么、如何排名,或者如何处理数据。我们认为这是不对的。cl0q 是一个拥有独立爬虫和索引的搜索引擎。到目前为止,我们已经映射了 3850 万个域名,并正朝着 1 亿个的目标努力。现在,爬虫本身已向所有人开放,供大家检查、运行或改进。如果你曾经好奇“真正的网页爬虫内部是什么样子的”,这就是你的答案。每一行代码都清晰可读,没有任何魔法。
What it does: The node software discovers websites, fetches their homepages politely, and sends what it finds to cl0q’s index. That’s it. It doesn’t scrape personal data, doesn’t follow you around, doesn’t build profiles. Five services, one Postgres database, zero public ports. Everything runs locally on your machine. 它的功能:该节点软件负责发现网站,以礼貌的方式抓取其主页,并将发现的内容发送到 cl0q 的索引中。仅此而已。它不会抓取个人数据,不会跟踪你的行踪,也不会建立用户画像。它包含五个服务、一个 Postgres 数据库,且没有开放任何公共端口。一切都在你的本地机器上运行。
It’s a good bot (by design, not by promise): Most crawlers say they respect robots.txt. Ours can’t not respect it. The politeness rules are hardcoded: One honest User-Agent that names your node. It never pretends to be a browser. Public internet only. It physically cannot connect to private networks, localhost, or cloud metadata endpoints. robots.txt is checked first, every time. Crawl-delay is obeyed. At least 1 second between requests to any server. No hammering. Only fetches homepages, at most once every 120 days per site. These aren’t settings you can turn off. They’re in the code. 它是一个“好机器人”(基于设计,而非承诺):大多数爬虫声称它们遵守 robots.txt,而我们的爬虫是“不得不”遵守。礼貌规则被硬编码在程序中:使用一个诚实的 User-Agent 来标识你的节点,从不伪装成浏览器。仅限公共互联网,物理上无法连接到私有网络、本地主机或云元数据端点。每次抓取前都会先检查 robots.txt。严格遵守抓取延迟,对任何服务器的请求间隔至少 1 秒,绝不进行轰炸式抓取。仅抓取主页,每个站点最多每 120 天抓取一次。这些不是可以关闭的设置,而是直接写在代码里的规则。
Run your own node: You need Docker and a machine that’s online. That’s it. 运行你自己的节点:你只需要 Docker 和一台联网的机器。就是这样。
git clone https://codeberg.org/cl0qsearch/crawler-node.git
cd crawler-node
cp .env.example .env
docker compose build
docker compose run --rm cl0q-node enroll --contact you@example.com
docker compose run --rm preflight
docker compose up -d
The enroll step registers your node with cl0q and gives you a token. The preflight check makes sure everything works before you start crawling. “enroll”步骤会将你的节点注册到 cl0q 并为你提供一个令牌。“preflight”检查则确保在开始抓取之前一切运行正常。
What’s next: This is v0.1.0. Coming up: more repos open sourced, a GitHub mirror, and performance work. The code is MIT licensed. Fork it, break it, improve it, send PRs. Run a node. Help index the open web. 下一步计划:这是 v0.1.0 版本。接下来我们将开源更多仓库、建立 GitHub 镜像,并进行性能优化。代码采用 MIT 许可证。欢迎 Fork、测试、改进并提交 PR。运行一个节点,帮助索引开放网络。