The model that didn't exist, so you made it yourself

本文为原文前 6,000 字符的节选翻译,完整内容请查看原文。

Back to Articles The model that didn’t exist, so you made it yourself Published October 8, 2026 Update on GitHub Upvote 1 yuvraj sharma ysharma Follow Abubakar Abid abidlabs Follow Last week, I wanted a small version of the prompt rewriter that ships with Qwen-Image 2.1. The official one is a 9B model that needs about 20 GB of memory and thinks for thousands of tokens before writing a single paragraph. On the Hub, I found only compressed copies of that same 9B model. So I described what I wanted to ML Intern, and the next day I had a 0.8B version that runs on a CPU. It returns valid output 99.7% of the time and uses about a quarter of the teacher’s tokens. The compute for the whole project, including having the 9B model label 8,797 example requests, came to USD 16.

回到文章:那个不存在的模型,于是你亲手打造了它。发布于 2026 年 10 月 8 日。上周,我想要一个 Qwen-Image 2.1 自带的提示词重写器的小型版本。官方版本是一个 9B 模型,需要约 20 GB 内存,并且在写出一段话之前要思考数千个 token。在 Hugging Face Hub 上,我只找到了该 9B 模型的压缩版本。于是我向 ML Intern 描述了我的需求,第二天我就得到了一个可以在 CPU 上运行的 0.8B 版本。它有 99.7% 的概率返回有效输出,且使用的 token 数量仅为教师模型的四分之一。整个项目的计算成本,包括让 9B 模型标注 8,797 个示例请求,总计 16 美元。

Over the course of the next few days, I made five more models the same way. Each one started as a message in HuggingChat with ML-intern switched on, and each one ended as a public model on the Hub with its evaluation in the model card. ML-intern plans the work, asks me for a budget before it spends anything, runs a small test before the real job, then trains, evaluates and publishes on Hugging Face hardware. How I prompt ML Intern The first message is where I spend my effort. My first prompt, for the citrus model shared below, was about 450 words. By my 6th project it was closer to 2,000, because each project taught me something I wanted in the next one. All seven prompts are on GitHub at yvrjsharma/ml-intern-prompts, exactly as I wrote them.

在接下来的几天里,我用同样的方法又制作了五个模型。每一个模型都始于开启 ML-intern 的 HuggingChat 消息,并最终成为 Hub 上的公开模型,且在模型卡中附带了评估结果。ML-intern 会规划工作,在花费任何资金前向我询问预算,在正式任务前运行小规模测试,然后进行训练、评估,并发布在 Hugging Face 的硬件上。我是如何向 ML Intern 提示的:第一条消息是我投入精力最多的地方。我为下面分享的柑橘模型编写的第一个提示词大约有 450 字。到了第六个项目时,字数已接近 2,000 字,因为每个项目都教会了我一些想在下一个项目中实现的东西。所有七个提示词都已发布在 GitHub 的 yvrjsharma/ml-intern-prompts 上,保持了我编写时的原样。

A prompt starts with the idea in one line and why I want it. Then it names the exact pieces: the dataset, the base model, the training script. Anything I have already checked goes under a heading that literally says “Verified facts, do not re-derive”, so the agent spends its budget on the work instead of rediscovering what I know. For the camera-angle LoRA that section listed which trainer had just added transparent-image support, and which open GitHub issues made the fallback trainer risky. Two lines in the prompt are critical. The first asks for a baseline before any training. For example, the citrus prompt says: “Also report the base model’s zero-shot score on the same metric before training so we can see the gain.” Without it you get a trained model and no idea whether it is better than what you started with.

提示词以一行核心想法及其原因开头。接着明确指出具体组件:数据集、基础模型、训练脚本。任何我已经核实过的内容都会放在一个标题为“已验证事实,无需重新推导”的下方,这样代理就会把预算花在工作上,而不是重新发现我已经知道的东西。对于相机角度 LoRA,该部分列出了哪个训练器刚刚添加了透明图像支持,以及哪些 GitHub 开放问题导致备用训练器存在风险。提示词中有两行至关重要。第一行要求在训练前提供基准测试。例如,柑橘模型的提示词写道:“同时报告基础模型在训练前相同指标下的零样本得分,以便我们观察提升效果。”没有这一步,你只会得到一个训练好的模型,却不知道它是否比初始模型更好。

The second is a smoke test with a check attached. For the image LoRAs I asked for 50 training steps, then a check that the saved weights had actually changed, before paying for the full run. At the end of the prompt, I lay out the expected deliverables and limit the cost. I define what belongs in the model card and include a instruction like: “Cap total spend at USD 12 and ask me before exceeding it.” Because ML-intern begins every task with zero dollar budget and needs permission before executing paid jobs, this spending limit stays strictly enforced. When you leave out a budget, the agent suggests a couple of paths depending on project size and asks which one you prefer. You don’t necessarily need all of that on your first attempt. For example, I didn’t have the verified-facts section in my citrus brief and ML Intern still produced a model that more than tripled the accuracy of the Qwen3.5-2B model.

第二行是一个带有检查机制的冒烟测试。对于图像 LoRA,我要求先进行 50 步训练,然后检查保存的权重是否确实发生了变化,之后才支付完整运行的费用。在提示词末尾,我列出了预期的交付成果并限制了成本。我定义了模型卡中应包含的内容,并加入了一条指令:“总支出上限为 12 美元,超过前请询问我。”由于 ML-intern 在开始每项任务时预算均为零,且在执行付费任务前需要许可,因此该支出限制得到了严格执行。如果你没有设定预算,代理会根据项目规模建议几种路径,并询问你的偏好。你不需要在第一次尝试时就做到面面俱到。例如,我的柑橘项目简报中没有“已验证事实”部分,但 ML Intern 依然产出了一个准确率比 Qwen3.5-2B 模型高出三倍多的模型。

Let me walk you through 6 things I built with Ml-Intern in just a couple of days. 1. A model that knows your field A general vision model can describe a yellowing citrus leaf. However, telling you whether it is a mite problem or a magnesium deficiency, and the bio and non-bio remedies to treat the plant is very hard. Using Claude, I put together a training dataset merged from three sources hosted by the Project-AgML organization on the Hub. The resulting citrus-disease-vlm-instruct is a dataset containing 3,017 annotated images across 21 distinct pests, illnesses, nutritional gaps, and treatment approaches. ML-intern handled the fine-tuning of Qwen3.5-2B using these examples, making sure to benchmark the foundation model beforehand. On the 335 test photos, the base model named the right problem 14.9% of the time. After two epochs on one A10G, the fine-tuned model got 52.8%. Compute cost, about USD 1.90. Check out: Model · Dataset · Citrus Doctor App

让我带你看看我在短短几天内用 ML-Intern 构建的 6 个项目。1. 一个了解你所在领域的模型。通用的视觉模型可以描述一片发黄的柑橘叶,但要判断它是螨虫问题还是镁缺乏症,并给出生物或非生物的治疗方案是非常困难的。我利用 Claude 整理了一个训练数据集,该数据集合并了 Hub 上 Project-AgML 组织托管的三个来源。最终生成的 citrus-disease-vlm-instruct 数据集包含 3,017 张标注图像,涵盖 21 种不同的害虫、疾病、营养缺失及治疗方法。ML-intern 使用这些示例对 Qwen3.5-2B 进行了微调,并确保事先对基础模型进行了基准测试。在 335 张测试照片中,基础模型识别出正确问题的概率为 14.9%。在单张 A10G 上运行两个 epoch 后,微调后的模型准确率达到了 52.8%。计算成本约为 1.90 美元。查看:模型 · 数据集 · 柑橘医生应用。

  1. A model that draws your character Image models know plenty of characters. Huggy, drawn in the flat style of the Hugging Face brand assets, was not one of them. I asked ML-Intern for a LoRA on FLUX.2 klein base 4B, trained on 84 captioned drawings from Chunte/huggy_for_training dataset. The agent saved a checkpoint every 100 steps and drew the same set of prompts with each one, which made choosing easy. Step 200 was the first where Huggy was fully on-model. From step 500 on, Huggy’s style started bleeding into prompts that had nothing to do with Huggy! The trained LoRA also works on the distilled klein model at 4 steps. Compute cost, about USD 7.60. Check out: Model · Dataset · Huggy Generator App

  2. 一个绘制你角色的模型。图像模型认识很多角色,但以 Hugging Face 品牌资产的扁平风格绘制的 Huggy 并不在其中。我要求 ML-Intern 在 FLUX.2 klein base 4B 上训练一个 LoRA,训练数据来自 Chunte/huggy_for_training 数据集中的 84 张带标题的绘图。代理每 100 步保存一个检查点,并用每个检查点绘制同一组提示词,这使得选择变得很容易。第 200 步是 Huggy 完全符合模型风格的起点。从第 500 步开始,Huggy 的风格开始渗透到与 Huggy 无关的提示词中!训练好的 LoRA 也可以在 4 步蒸馏的 klein 模型上运行。计算成本约为 7.60 美元。查看:模型 · 数据集 · Huggy 生成器应用。

  3. A model that does a new trick Camera-angle LoRAs are among the most-liked community add-ons for earlier Qwen-Image models. You can give the model a picture of an object and ask to see it 45 degrees from the left. When I checked a few days after the Qwen-Image 2.1 model release, nobody had made one, so I tasked ML-intern to build it. ML-intern rendered 1,030 scanned household objects from Google Scanned Objects at 24 angles each, 24,722 transparent images, on a CPU job that cost a few cents. It later finalised 461 objects for training and 40 held out for testing, and 1,844 before-and-after training pairs spread evenly over 23 camera instructions. Training ran 2,000 steps in about 90 minutes on one A100 (~USD 3.75). The whole project took about half a day and 48 jobs, counting the ones that failed on missing packages or wrong paths and had to be resubmitted by ML-Intern. Total compute cost, about USD 16. Check out: Model · Dataset · Viewpoint Orbit App Doodle-in LoRA is another cool idea. Upload a photo with a magenta scribble on it and add

  4. 一个掌握新技巧的模型。相机角度 LoRA 是早期 Qwen-Image 模型中最受欢迎的社区插件之一。你可以给模型一张物体图片,并要求从左侧 45 度角查看它。在 Qwen-Image 2.1 发布几天后,我查看时发现还没人制作过,于是我指派 ML-intern 来构建它。ML-intern 将 Google Scanned Objects 中的 1,030 个扫描家居物体渲染为每个 24 个角度,共 24,722 张透明图像,这项 CPU 任务仅花费了几美分。随后,它确定了 461 个用于训练的物体和 40 个用于测试的保留物体,并生成了 1,844 对训练前后对比图,均匀分布在 23 个相机指令中。训练在单张 A100 上运行了 2,000 步,耗时约 90 分钟(约 3.75 美元)。整个项目耗时约半天,共 48 个任务,包括那些因缺少包或路径错误而失败并由 ML-Intern 重新提交的任务。总计算成本约为 16 美元。查看:模型 · 数据集 · 视角轨道应用。Doodle-in LoRA 是另一个很酷的想法。上传一张带有洋红色涂鸦的照片并添加……