Changes to SourceHut's terms of service regarding LLMs

Changes to SourceHut’s terms of service regarding LLMs

关于 SourceHut 服务条款中涉及大语言模型(LLM)的变更

In late July, Codeberg announced that they would limit the use of LLMs on their platform, something that SourceHut has also been considering for some time. Last week, we discussed with the SourceHut community the possibility of limiting or prohibiting the use of LLMs (large language models) and other generative AI technologies on SourceHut. Thank you to everyone who participated, shared a variety of views, and ensured we had a productive and civil conversation on the subject. Following this conversation and more discussions internally, we have elected to move forward with the necessary changes to our terms of service.

7 月下旬,Codeberg 宣布将限制在其平台上使用大语言模型(LLM),这也是 SourceHut 一段时间以来一直在考虑的事情。上周,我们与 SourceHut 社区讨论了在 SourceHut 上限制或禁止使用 LLM 及其他生成式 AI 技术的可能性。感谢每一位参与讨论、分享不同观点并确保我们进行了一场富有成效且文明对话的朋友。在这次对话及内部进一步讨论后,我们决定推进服务条款的必要变更。

I encourage you to read through this thread to see the varied and thoughtful opinions of the SourceHut community on this subject, both for and against the change. In summary, the SourceHut community broadly has strong support for restrictions on the use of AI on this platform. However, ultimately this decision was made by staff, so we’ll focus on our own rationale for it, give examples of how the new policies will be interpreted and enforced, the expected timeline for the change, and options for people affected by the change and how they can move their projects forward afterwards.

我鼓励大家阅读这个讨论帖,看看 SourceHut 社区对此事各种深思熟虑的观点,无论是支持还是反对。总之,SourceHut 社区普遍强烈支持对该平台上的 AI 使用进行限制。然而,最终决定权在于员工,因此我们将重点阐述我们的决策依据,举例说明新政策将如何解读和执行,变更的预期时间表,以及受影响用户在变更后的选择和如何继续推进其项目。

There are material incentives for us to curtail “vibe coded” projects on SourceHut. They tend to, as a class, be outliers in terms of resource usage, as the practice tends to invent complex CI configurations that waste build minutes, produce larger codebases, and push commits often, all of which causes excessive resource usage and can lead to outages. Often these codebases are abandoned after burning through all of these resources for naught. And, of course, the scourge of LLM crawlers scouring the internet causes more outages and headaches for our sysadmins, and AI is causing skyrocketing costs for the hardware we need to run the platform and the hardware you need to use it.

我们有实质性的动机去限制 SourceHut 上的“AI 生成式(vibe coded)”项目。这类项目在资源使用方面往往表现异常,因为这种做法倾向于创建复杂的 CI 配置,浪费构建时长,产生庞大的代码库,并频繁推送提交,所有这些都会导致资源过度消耗,甚至引发宕机。通常,这些代码库在耗尽资源却一无所获后就被废弃了。当然,LLM 爬虫在互联网上肆虐也给我们的系统管理员带来了更多的宕机和麻烦,而且 AI 正在导致我们运行平台所需的硬件成本以及你们使用平台所需的硬件成本飙升。

But, in truth, the rationale is more about politics and ethics than anything else. Deciding which “hills to die on” is tricky. We are an organization which is transparent about politics, both the fact that we, like all others, have politics, and also about what our politics are. But, although the world is a turbulent place, and this calls all of us to action, there are many issues we believe firmly in but which need not necessarily steer SourceHut policy. If we were to shape our policies in terms of our values on, for example, Palestine, it would serve to stoke controversy and increase our workload without meaningfully contributing anything to the problem. We lack the cultural and intellectual context to know what to do about that situation, and we have access to few resources knowledgeable enough to steer our policies with wisdom and care.

但事实上,这一决策背后的理由更多是出于政治和伦理考量。决定在哪些问题上“坚持到底”是很棘手的。我们是一个在政治立场上保持透明的组织,既承认我们像其他人一样拥有政治立场,也明确我们的立场是什么。但是,尽管世界动荡不安,呼吁我们采取行动,但我们坚信的许多问题并不一定非要指导 SourceHut 的政策。如果我们根据自己在例如巴勒斯坦问题上的价值观来制定政策,只会引发争议并增加我们的工作负担,而对解决问题毫无实质性贡献。我们缺乏相关的文化和知识背景来判断如何处理该情况,也没有足够的资源来明智且审慎地引导我们的政策。

In the case of AI, though it touches on complex social, economic, and political factors, many of which are beyond the realm of our expertise, it overlaps considerably with matters on which we are experts – software development, open source, and open culture, for example. As experts, we can more easily see that we have an obligation to evaluate their impact and to weigh in and make decisions accordingly. If we focus narrowly on how LLMs affect our community, we can consult our mission statement for guidance: We are here to make free software better. We will be honest, transparent, and empathetic. We care for our users, and we will not exploit them, and we hope that they will reward our care and diligence with success.

在 AI 的问题上,虽然它涉及复杂的社会、经济和政治因素,其中许多超出了我们的专业领域,但它与我们擅长的领域——例如软件开发、开源和开放文化——有相当大的重叠。作为专家,我们更容易看到我们有义务评估其影响,并据此权衡和做出决定。如果我们狭隘地关注 LLM 如何影响我们的社区,我们可以参考我们的使命宣言来寻求指导:我们致力于让自由软件变得更好。我们将保持诚实、透明和同理心。我们关心我们的用户,不会剥削他们,我们也希望他们能以成功来回报我们的关怀与勤勉。

It’s not easy to argue that LLMs make free software better. The means by which they operate involves taking, by force and without regard for the wishes of the software authors, the obligations of software licenses, or the platforms on which our work relies, vast swaths of open source software, disregarding the licensing requirements, and charging rent to produce code which is a plausible reinterpretation of its training data applied to new problems. The models which are produced in this process are proprietary, owned and controlled by private corporations, and give little back to the open source community. Therefore, the relationship between LLMs and open source seems to exploit the users we are responsible for caring for.

很难论证 LLM 让自由软件变得更好了。它们的运作方式涉及强行获取大量开源软件——无视软件作者的意愿、软件许可的义务或我们工作所依赖的平台——无视许可要求,并收取费用来生成代码,而这些代码不过是其训练数据在面对新问题时的一种貌似合理的重新诠释。在此过程中产生的模型是专有的,由私营公司拥有和控制,几乎没有回馈给开源社区。因此,LLM 与开源之间的关系似乎是在剥削我们有责任关心的用户。

We have also seen some other troubling incidents, for example the use of LLMs to rewrite open source software to circumvent copyleft obligations, such as in the case of the chardet library, a practice which seems to straightforwardly work against the interests of open source software authors. However, LLMs have shown some utility. They find real bugs, including security bugs, and open source software has become more robust as a consequence. Some programmers find them useful for code reviews and other creative applications, and indeed many programmers report good results in using them to write or assist in writing code.

我们也看到了一些其他令人不安的事件,例如使用 LLM 重写开源软件以规避 Copyleft(著佐权)义务,如 chardet 库的案例,这种做法似乎直接损害了开源软件作者的利益。然而,LLM 也确实展现出了一些实用性。它们能发现真正的漏洞,包括安全漏洞,从而使开源软件变得更加稳健。一些程序员发现它们在代码审查和其他创造性应用中很有用,事实上,许多程序员报告称在编写或辅助编写代码时使用它们效果良好。

The more we zoom out, though, the more alarming it gets. The biggest concern is the impact on the climate. Cars fueled by fossil fuels are also useful, but most people support efficiency improvements and ultimately a complete replacement with electric vehicles to mitigate their effects. The transformation of society to address the climate emergency must be comprehensive, and each of us, in our respective industries and expertise, have a moral obligation to work to reduce our environmental footprint, particularly in systemic and institutional ways. The roll-out of AI is a massive trend in the wrong direction. LLMs and similar tools are an extremely energy-inefficient way of solving the problems they address, such as authoring software, and are fundamentally at odds with mitigating the climate disaster. Planned developments in AI infrastructure are expected to generate an electricity demand comparable to the entire country of India, a country which is experiencing tens of thousands of deaths due to heatwaves exacerbated by climate change. This build-out consumes a large portion.

然而,当我们看得更远时,情况就越发令人担忧。最大的担忧是对气候的影响。化石燃料驱动的汽车也很有用,但大多数人支持提高效率,并最终用电动汽车完全替代它们以减轻其影响。社会为应对气候紧急状况所做的转型必须是全面的,我们每个人在各自的行业和专业领域,都有道德义务致力于减少我们的环境足迹,特别是在系统性和制度层面。AI 的推广是一个巨大的错误方向。LLM 和类似工具在解决它们所处理的问题(如编写软件)时,是一种极其低效的能源使用方式,从根本上与缓解气候灾难背道而驰。计划中的 AI 基础设施开发预计将产生相当于整个印度全国的电力需求,而印度正因气候变化加剧的热浪而经历数万人的死亡。这种建设消耗了很大一部分……