Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
在消费级硬件(RTX 4090)上以 100T/s 的速度运行 Qwen 3.8 Flash Next (125B)
Run a 125-billion-parameter AI model on your own gaming PC. NVIDIA or AMD graphics card (12 GB or more) · Windows or Linux · free and open source. 在你的个人游戏电脑上运行 1250 亿参数的 AI 模型。支持 NVIDIA 或 AMD 显卡(12 GB 或以上)· Windows 或 Linux 系统 · 免费且开源。
A voxel pagoda garden, 1 shot prompt running on an RTX 5070 with Strata (IQ3_S, 128K context) · full video (49 s). 一个体素宝塔花园,通过 Strata 在 RTX 5070 上运行的单次提示词生成(IQ3_S,128K 上下文)· 完整视频(49 秒)。
Strata runs Qwen3.8-Flash-Next on a normal PC. This is a large, smart AI model that usually needs a server. It chats, writes code, reads pictures and works with your apps and coding agents. Nothing leaves your PC. Strata 可以在普通电脑上运行 Qwen3.8-Flash-Next。这是一个通常需要服务器才能运行的大型智能 AI 模型。它可以聊天、编写代码、读取图片,并与你的应用程序和编程代理协同工作。所有数据都不会离开你的电脑。
How fast is it? We measured it on two ordinary gaming PCs. A token is about ¾ of a word. 它有多快?我们在两台普通游戏电脑上进行了测试。一个 Token 大约是 ¾ 个单词。
Writes answers: how fast the reply appears in a short chat. 60 tokens per second is faster than you can read. 写入答案:在简短对话中回复出现的速度。每秒 60 个 Token 的速度比你的阅读速度还要快。
Reads your prompt: how fast it takes in what you send (here a 32K-token document, code or chat history). 读取提示词:处理你发送内容的速度(此处为 32K Token 的文档、代码或聊天记录)。
NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAM
NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB 内存
| Size | Writes answers | Reads your prompt |
|---|---|---|
| Q2_0 | 94 tokens/s | 2,650 tokens/s |
| IQ2_XS | 79 tokens/s | 2,090 tokens/s |
| IQ3_XXS | 62 tokens/s | 1,750 tokens/s |
| IQ3_S | 53 tokens/s | 1,620 tokens/s |
| Coder | 55 tokens/s | 2,180 tokens/s |
AMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM
AMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB 内存
| Size | Writes answers | Reads your prompt |
|---|---|---|
| Q2_0 | 60 tokens/s | 1,160 tokens/s |
| IQ2_XS | 52 tokens/s | 1,110 tokens/s |
| Coder | 44 tokens/s | 1,420 tokens/s |
NVIDIA: Q2_0 with engine 0.1.36, the other rows with 0.1.26 (4K answers, 32K prompts). The full tables are in DETAILS.md. A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second. Long chats and other cards: speed of each model, community results. Strata is free. If it runs well on your PC, a coffee keeps the work on it going. NVIDIA:Q2_0 使用 0.1.36 引擎,其他行使用 0.1.26(4K 答案,32K 提示词)。完整表格请见 DETAILS.md。显存更大的显卡速度更快:RTX 3090 (24 GB) 的写入速度应在每秒 100-140 个 Token 左右。长对话及其他显卡:各模型速度、社区测试结果。Strata 是免费的。如果它在你的电脑上运行良好,请买杯咖啡支持后续开发。
What you need
你需要什么
Graphics card: NVIDIA GeForce RTX 20, 30, 40 or 50 series, or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700 or RX 6800 / 6900 series. It needs 12 GB of VRAM or more. 显卡: NVIDIA GeForce RTX 20、30、40 或 50 系列,或 AMD Radeon RX 7900 XT / XTX、RX 7800 XT / 7700 XT、RX 9060 XT、RX 9070 / 9070 XT、Radeon AI PRO R9700 或 RX 6800 / 6900 系列。需要 12 GB 或以上的显存。
RAM: 32 GB or more. Your RAM decides which model fits. 64 GB runs every size. 内存: 32 GB 或以上。你的内存决定了能运行哪个模型。64 GB 可以运行所有尺寸的模型。
Disk: About 80 GB free. Use an SSD if you can: the first start is much faster. 磁盘: 约 80 GB 可用空间。如果可以,请使用 SSD:首次启动会快得多。
System: Windows 10 / 11 or Linux, and a current graphics driver from NVIDIA or AMD. The installer sets up everything else. Two or three cards can share the model (multi-GPU). 系统: Windows 10 / 11 或 Linux,以及最新的 NVIDIA 或 AMD 显卡驱动。安装程序会处理其余所有设置。两到三张显卡可以共享模型(多 GPU)。
Experimental, written and tested by community members on their own machines: Older graphics cards (Tesla P40 / V100, GTX 10, Radeon VII / MI50, RX 6700 XT, RX 5500 XT): Older GPUs. Intel Arc, built from source on Linux: Intel Arc. Older processors without AVX2: they work, but slowly. Older CPUs. The full list: docs/INSTALL.md. 实验性支持,由社区成员在各自机器上编写和测试:旧款显卡(Tesla P40 / V100, GTX 10, Radeon VII / MI50, RX 6700 XT, RX 5500 XT):旧款 GPU。Intel Arc,在 Linux 上从源码构建:Intel Arc。不支持 AVX2 的旧处理器:可以运行,但速度较慢。旧款 CPU。完整列表:docs/INSTALL.md。
Install
安装
Let your AI set it up: Do you use an AI coding assistant (Claude Code, Cursor, Codex, GitHub Copilot, …)? Paste this into it: “Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.” It checks your graphics card, RAM and disk and picks the model that fits. Then it installs and starts it and tells you how to connect your apps. AI tools can also install, start and stop Strata through its MCP server. 让你的 AI 来设置: 你在使用 AI 编程助手(Claude Code, Cursor, Codex, GitHub Copilot 等)吗?将这段话粘贴进去:“Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.” 它会检查你的显卡、内存和磁盘,并选择合适的模型。然后它会完成安装、启动,并告诉你如何连接你的应用程序。AI 工具还可以通过其 MCP 服务器安装、启动和停止 Strata。
Or do it yourself: Download Strata and unzip it (or git clone it). Windows: double-click START-HERE.bat. Linux: run ./setup.sh in the Strata folder. The steps are the same for NVIDIA and AMD. The installer finds your card and sets up the right engine for it. It asks you a few questions: which model and which size, how much context (how much text the model keeps in mind), whether it should read pictures. Press Enter each time for the recommended answer. Then it downloads the model (about 70 GB) and starts it. If the download stops, run it again: it continues where it left off. 或者手动操作: 下载 Strata 并解压(或 git clone)。Windows:双击 START-HERE.bat。Linux:在 Strata 文件夹中运行 ./setup.sh。NVIDIA 和 AMD 的步骤相同。安装程序会自动识别你的显卡并配置合适的引擎。它会问你几个问题:模型及其尺寸、上下文长度(模型能记住多少文本)、是否需要读取图片。每次按回车键即可选择推荐答案。然后它会下载模型(约 70 GB)并启动。如果下载中断,再次运行即可:它会从断点处继续。
Your browser opens the Strata app at http://127.0.0.1:8080. While the model starts, your PC can be slow or stop responding for 1-3 minutes (longest the first time). Strata loads 35-55 GB into your RAM and locks part of it for the graphics card. This is normal. Wait, and don’t close the window. The window shows what Strata is doing. Next time, run START-HERE.bat (or ./setup.sh) again. It starts right away and downloads nothing twice. Close its window to stop the model. UPDATE.bat (./update.sh) updates Strata without starting it. Updating, Docker, several cards, where the files go and every option: docs/INSTALL.md. 你的浏览器会在 http://127.0.0.1:8080 打开 Strata 应用。模型启动时,你的电脑可能会变慢或停止响应 1-3 分钟(首次启动最久)。Strata 会将 35-55 GB 数据加载到内存中,并为显卡锁定一部分内存。这是正常的。请等待,不要关闭窗口。窗口会显示 Strata 的运行状态。下次运行时,再次运行 START-HERE.bat(或 ./setup.sh)。它会立即启动,不会重复下载。关闭窗口即可停止模型。UPDATE.bat(./update.sh)可以在不启动的情况下更新 Strata。关于更新、Docker、多显卡、文件路径及所有选项:请见 docs/INSTALL.md。
Which model should I pick?
我该选哪个模型?
The installer recommends one for your RAM. The same model comes in several sizes, compressed more or less. Smaller sizes are faster. Larger sizes are a bit smarter. 安装程序会根据你的内存推荐模型。同一模型有多种尺寸,压缩程度不同。尺寸越小速度越快,尺寸越大越聪明。
| Your RAM | Take | Why |
|---|---|---|
| 32 GB | Coder | 它适合 32 GB 内存,且专为代码优化(24 GB 显存的卡也可运行 Q2_0 和 IQ2_XS) |
| 48 GB | IQ2_XS (or Q2_0) | 最快,更大的尺寸装不下 |
| 64 GB | IQ2_XS (recommended), or IQ3_XXS / IQ3_S | 所有尺寸均可运行;IQ3_S 效果最好但速度最慢 |
| 96 GB+ | IQ3_S, or Unsloth’s UD-IQ4_XS (~4-bit) | 有空间运行最大尺寸,同时还能运行其他程序 |
Coder: a coding version with half of the experts removed. It reaches 91% of the full model’s SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM. It is weaker outside code, including Chinese and other CJK text (#438). For those, take Q2_0, IQ2_XS or IQ3_S, which keep every expert. Coder: 一个移除了半数专家的编程版本。它达到了完整模型 SWE-bench Verified 分数的 91%(由作者测定),且适配 32 GB 内存。它在代码之外的能力较弱,包括中文和其他 CJK 文本(#438)。如果需要处理这些,请选择保留了所有专家的 Q2_0、IQ2_XS 或 IQ3_S。
Swift 1.5: a fine-tune that thinks for a much shorter time before it answers. You get the answer sooner, at about the same quality. Swift 1.5: 一个微调版本,在回答前思考时间更短。你能更快得到答案,且质量基本相同。
Unsloth UD-IQ4_XS: Unsloth’s ~4-bit version, between IQ3_S and UD-Q4_K_XL in quality. A 94 GB download. With less than ~80 GB of RAM, Strata reads part of it from the SSD while it answers, so it is slower there (an NVMe SSD helps). Unsloth UD-IQ4_XS: Unsloth 的约 4-bit 版本,质量介于 IQ3_S 和 UD-Q4_K_XL 之间。下载量为 94 GB。如果内存小于约 80 GB,Strata 在回答时会从 SSD 读取部分数据,因此速度会变慢(NVMe SSD 有帮助)。
Unsloth UD-Q4_K_XL (experimental): the closest to the full model. But Strata reads most of it from the SSD while it answers, so it writes only 7-8.5 tokens/s on a 64 GB PC. Unsloth UD-Q4_K_XL(实验性): 最接近完整模型的版本。但由于 Strata 在回答时会从 SSD 读取大部分数据,在 64 GB 内存的电脑上写入速度仅为每秒 7-8.5 个 Token。
OrcaRouter’s Uncensored IQ3_XXS: you set it up by hand. It is not in the installer’s menu. OrcaRouter 的 Uncensored IQ3_XXS: 需要手动设置,不在安装程序的菜单中。
Sizes, downloads and what fits where: docs/MODELS.md. To add another model later, run SETUP.bat (Linux: ./setup.sh —setup). 尺寸、下载及适配情况:docs/MODELS.md。若后续要添加其他模型,请运行 SETUP.bat(Linux:./setup.sh —setup)。
Using it
使用方法
The Strata app’s Monitor (left) while a coding agent writes the pagoda garden from the video (right). In the browser: open http://127.0.0.1:8080. Strata 应用的监控界面(左),此时编程代理正在编写视频中的宝塔花园(右)。在浏览器中打开:http://127.0.0.1:8080。