HandNotes: teaching an open model to read my friend's handwriting, on his own laptop

HandNotes: teaching an open model to read my friend’s handwriting, on his own laptop

HandNotes:教导开源模型识别朋友的手写笔记,并在其个人笔记本电脑上运行

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝 This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend. Hacktoberfest 周末挑战:为朋友而建(Build for a Friend)参赛作品 🤝 这是我为 Hacktoberfest 周末挑战赛提交的作品。

What I Built: My friend writes down everything in class. Every lecture, every assignment, page after page. The problem shows up the night before an exam: he photographs his notes, uploads them to an AI chatbot, asks for a summary, and gets nothing useful back. It can’t read his handwriting. 我构建了什么:我的朋友在课堂上会记下所有内容。每一场讲座、每一项作业,一页又一页。问题出现在考试前夜:他拍下笔记照片,上传到 AI 聊天机器人,要求总结,却得不到任何有用的结果。因为它读不懂他的字迹。

When I ran his pages through a vision model myself, I found something worse than “can’t read it”. The model read most of it, and the mistakes looked right: The Perennial Student (the book he was writing about) became The Peripheral Student; “satirical” became “spatialical”; “his reputation” became “her reputation”; a date in a letter disappeared, and “5:00 to 7:00 pm” became “8:00 to 8:00”. 当我亲自用视觉模型处理他的笔记时,我发现情况比“读不懂”更糟。模型读出了大部分内容,但错误看起来却“很合理”:他笔记中提到的书名《The Perennial Student》(常青学生)变成了《The Peripheral Student》(边缘学生);“satirical”(讽刺的)变成了“spatialical”;“his reputation”(他的名声)变成了“her reputation”(她的名声);信中的日期消失了,而“下午 5:00 到 7:00”变成了“8:00 到 8:00”。

And then there’s his shorthand. He writes “acc” for according and “gitq” for given in the question, which no model is ever going to guess. A tired student skimming a summary at 2 AM would never catch any of that. So I built HandNotes: a handwriting reader that runs entirely on his laptop, lets you correct what it got wrong, and remembers those corrections for the next page. 此外还有他的速记。他用“acc”代表 according(根据),用“gitq”代表 given in the question(题目中给出的),这是任何模型都无法猜到的。一个凌晨两点在浏览总结的疲惫学生绝不可能发现这些错误。因此,我构建了 HandNotes:一个完全在他笔记本电脑上运行的手写识别器,允许用户纠正错误,并为下一页笔记记住这些修正。

What it does: 它能做什么:

  • Reads a photo of his notes live. Text streams in word by word from an open vision model running locally.
  • 实时读取笔记照片。 文字通过本地运行的开源视觉模型逐词流式输出。
  • Highlights words it isn’t sure about, with one-click jumps to each one so checking a page is fast.
  • 高亮显示不确定的单词,并支持一键跳转,使页面检查变得快捷。
  • Learns from corrections. Every fix is compared word by word with what the model wrote. His vocabulary and past misreads go into the prompt for the next page, and a misread you’ve corrected twice gets fixed automatically.
  • 从修正中学习。 每次修正都会与模型原始输出进行逐词对比。他的词汇和过去的误读会被加入到下一页的提示词(Prompt)中,如果同一个误读被修正两次,系统会自动修复。
  • Knows his short forms. Teach it “gitq = given in the question” once, either by adding it or just by expanding it while correcting a page, and it expands it on every page after that.
  • 识别速记。 只需教它一次“gitq = given in the question”,无论是通过手动添加还是在纠错时展开,之后每一页它都会自动展开该缩写。
  • Shows its work. An accuracy chart per page, a list of his usual misreads, and a “compare with plain model” view that shows, word by word, what the memory changed.
  • 展示工作成果。 提供每页的准确率图表、常见误读列表,以及“与原始模型对比”视图,逐词展示记忆功能带来的改变。
  • Turns notes into revision material: a summary with likely exam questions, flip-to-reveal flashcards, and export to PDF, Word or text.
  • 将笔记转化为复习资料: 生成包含潜在考题的总结、翻转式抽认卡,并支持导出为 PDF、Word 或文本格式。

What he said: when I showed it to him, he really liked it, and he told me he’s going to use it on a regular basis. Coming from the person who photographs his notes because typing them up is too much effort, that’s the review I was hoping for. 他的评价:当我向他展示时,他非常喜欢,并告诉我他会经常使用。对于那个因为觉得打字太麻烦而选择拍照片的人来说,这就是我所期待的评价。

Demo: It runs locally (that’s the point), so the GIF at the top is a real run recorded on my laptop, sped up 4x: adding his short forms, reading one of his pages live, comparing against the plain model, and making flashcards. 演示:它在本地运行(这是重点),所以顶部的 GIF 是在我笔记本电脑上录制的真实运行过程,加速了 4 倍:添加速记、实时读取页面、与原始模型对比以及制作抽认卡。

Code: jemankalita / Hacktoberfest-Build-for-a-Friend. A handwriting reader that learns one friend’s handwriting from corrections. Gemma 3 via Ollama, runs fully offline on a laptop. 代码:jemankalita / Hacktoberfest-Build-for-a-Friend。一个通过修正学习特定朋友字迹的手写识别器。基于 Ollama 运行 Gemma 3,完全离线运行于笔记本电脑。

How I Built It: The model: Gemma 3 4B, an open-weight vision model, served locally by Ollama. On my laptop (GTX 1650 Ti, 4 GB) Ollama splits it about half GPU, half CPU, and a page takes 30 to 40 seconds. 我是如何构建的:模型采用 Gemma 3 4B,这是一个开源权重的视觉模型,由 Ollama 在本地提供服务。在我的笔记本电脑(GTX 1650 Ti, 4GB 显存)上,Ollama 将任务分配给 GPU 和 CPU 各一半,处理一页大约需要 30 到 40 秒。

The app: a small FastAPI server bound to 127.0.0.1, with a hand-written HTML/CSS/JS front end. Fonts and icons are bundled, so even the UI makes no network requests. 应用:一个绑定在 127.0.0.1 的小型 FastAPI 服务器,配合手写的 HTML/CSS/JS 前端。字体和图标均已打包,因此即使是 UI 界面也不会发起任何网络请求。

The prompt matters more than I expected. My first prompt let the model “tidy up” his notes: it dropped a whole line with a date in it. Telling it to copy literally, skip nothing, and mark unclear words with [?] brought the date back. 提示词(Prompt)的重要性超出了我的预期。我最初的提示词让模型“整理”笔记,结果它删掉了一整行包含日期的内容。告诉它逐字复制、不要跳过任何内容,并用 [?] 标记不清楚的单词后,日期又回来了。

The learning is a memory, not retraining. Fine-tuning a vision model in a weekend on a 4 GB GPU wasn’t realistic, so HandNotes keeps a small per-writer memory file. 这种学习是基于记忆的,而非重新训练。在一个周末内用 4GB 显存的 GPU 微调视觉模型是不现实的,因此 HandNotes 为每个用户维护了一个小型记忆文件。