Hacktoberfest Challenge Week 1: Touch Grass - Nature Quest

Hacktoberfest Challenge Week 1: Touch Grass - Nature Quest

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿 Hacktoberfest 开源 AI 挑战赛第一周:Touch Grass(亲近自然)参赛作品 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass. 这是我为 Hacktoberfest 开源 AI 挑战赛第一周“Touch Grass”主题提交的作品。

What I Built 我的作品

Nature Quest is an offline scavenger hunt for families and kids. You pick a place (park, forest, garden, urban walk), how many things to find (5, 8 or 12) and the age of the youngest player. Then an open-source AI model running on the phone writes a hunt list for that place, season and age: “a red leaf”, “something with a spiral shape”, “bark with moss on it”. The app reads the list aloud and asks everyone to put the phone away. When someone finds something, they tap it on the list and take one photo. The same model checks it, answers with a short, encouraging message (plus a friendly hint if it is not a match), and reads that aloud too. When the hunt ends, there is a bronze, silver or gold medal and the time spent outside. Nature Quest 是一款专为家庭和儿童设计的离线寻宝游戏。你可以选择地点(公园、森林、花园、城市步道)、寻找物品的数量(5、8 或 12 个)以及最小玩家的年龄。随后,运行在手机上的开源 AI 模型会根据地点、季节和年龄生成一份寻宝清单,例如:“一片红叶”、“螺旋形状的东西”、“长满苔藓的树皮”。应用会朗读清单,并提醒大家收起手机。当有人找到目标时,只需在清单上点击并拍一张照片。同一个 AI 模型会进行核对,并以简短、鼓励的话语(如果未匹配成功,还会给出友好提示)进行回复,同样通过语音朗读出来。寻宝结束后,系统会颁发铜牌、银牌或金牌,并记录户外活动时间。

The theme was “Touch Grass”, so every design decision started from one question: how do I make the screen the shortest part of the experience? Setup is one screen with everything pre-selected, so starting a hunt is one tap. The list and every piece of feedback are read aloud, so nobody needs to look at the phone. Big tap targets and a high-contrast palette, because the phone is used outdoors in the sun, for a few seconds at a time. One photo per find. No camera roll, no feed, nothing to scroll. It is for families and kids on a walk, with an adult, in English and Spanish. The first screen always says: hunt with an adult, only look and take photos, never pick or touch anything. 本次主题是“Touch Grass”(亲近自然),因此每一个设计决策都始于一个问题:如何让屏幕使用时间降到最低?设置界面仅需一屏,所有选项预先选定,因此只需点击一下即可开始寻宝。清单和所有反馈均通过语音朗读,无需盯着手机看。界面采用大点击目标和高对比度配色,因为手机是在户外阳光下使用的,且每次仅需操作几秒钟。每个目标仅限一张照片。没有相册、没有信息流、无需滚动。该应用旨在供家庭和儿童在成人陪同下散步时使用,支持英语和西班牙语。首屏始终显示:请在成人陪同下寻宝,仅限观察和拍照,严禁采摘或触摸任何东西。

Setup 设置

The hunt, written on the phone. Medal and time outside. 手机上生成的寻宝清单。奖牌和户外活动时间。

Demo 演示

There is no web demo, because the whole point is that it runs on a phone. Install it: Nature Quest 1.0.1 on GitHub Releases (signed APK, about 55 MB, Android 12+, 64-bit ARM). On first launch the app offers to download the AI model (2.59 GB, Wi-Fi recommended). That is the only time it uses the internet. 本项目没有网页演示,因为其核心意义在于完全运行在手机上。安装方式:在 GitHub Releases 下载 Nature Quest 1.0.1(已签名 APK,约 55 MB,支持 Android 12+,64 位 ARM 架构)。首次启动时,应用会提示下载 AI 模型(2.59 GB,建议在 Wi-Fi 下进行)。这是该应用唯一需要联网的时刻。

How I Built It 开发过程

The open pieces. The model is Gemma 4 E2B (open weights, Apache-2.0) in LiteRT-LM’s .litertlm format. It runs through LiteRT-LM (Apache-2.0), Google’s on-device runtime, with no network and no server. The rest is a normal Android stack: Kotlin, Jetpack Compose, Hilt, CameraX and Android’s own text-to-speech. Nothing closed runs in the app. 开源组件:模型采用 Gemma 4 E2B(开放权重,Apache-2.0 协议),格式为 LiteRT-LM 的 .litertlm。它通过 Google 的端侧运行时 LiteRT-LM(Apache-2.0 协议)运行,无需网络,无需服务器。其余部分是标准的 Android 技术栈:Kotlin、Jetpack Compose、Hilt、CameraX 以及 Android 原生语音合成(TTS)。应用内不包含任何闭源组件。

The shape of it. The app talks to the model through one interface, InferenceEngine, so ViewModels and use cases never touch LiteRT-LM and tests use a fake. The model answers in strict JSON, which I parse with kotlinx.serialization. If the JSON is bad I retry once, then fall back to something safe: “I’m not sure, try another photo”, or a built-in hunt list. Prompts are versioned text files, not strings in Kotlin. 架构设计:应用通过 InferenceEngine 接口与模型交互,因此 ViewModels 和用例层无需直接接触 LiteRT-LM,测试时也可使用模拟对象。模型以严格的 JSON 格式回答,我使用 kotlinx.serialization 进行解析。如果 JSON 解析失败,我会重试一次,若仍失败则回退到安全方案:“我不确定,请尝试拍另一张照片”,或使用内置的寻宝清单。提示词(Prompts)以版本化文本文件形式存储,而非硬编码在 Kotlin 字符串中。

I benchmarked before I built. The first real step was a debug-only screen that loaded the model and timed it on my Pixel 10. That changed the design: The CPU backend was about twice as fast as the GPU: a hunt list took 5.7 s on average on CPU and 11.9 s on GPU, and the first GPU load took 25 s against 3 s. Google’s published numbers for a flagship phone did not carry over to mine. Decoding is slow (7 to 23 tokens per second), so every answer is kept short. “Twelve words at most” in a prompt is a performance feature. 我在开发前进行了基准测试。第一步是创建一个仅用于调试的界面,加载模型并在我的 Pixel 10 上进行计时。这改变了我的设计:CPU 后端的运行速度大约是 GPU 的两倍:生成寻宝清单在 CPU 上平均耗时 5.7 秒,而在 GPU 上则需 11.9 秒;首次加载 GPU 模型需 25 秒,而 CPU 仅需 3 秒。Google 公布的旗舰机数据在我的设备上并不适用。解码速度较慢(每秒 7 到 23 个 token),因此所有回答都保持简短。提示词中要求的“最多十二个词”实际上是一种性能优化手段。

The phone slows about 1.6x after 25 minutes of back-to-back runs, even though the OS reported no thermal throttling. So I picked settings that hit my targets in the slow state: 140 visual tokens and a 640 px photo. A photo check takes 4.5 to 5 s on a cool phone and 5.7 to 6.5 s on a warm one. Peak memory is 2.2 to 2.6 GB, so the app loads the model for each use and releases it right after. A warm reload is about half a second, and memory drops to about 240 MB. 在连续运行 25 分钟后,手机性能下降了约 1.6 倍,尽管系统并未报告热节流。因此,我选择了在性能下降状态下也能达标的设置:140 个视觉 token 和 640 像素的照片。在手机温度正常时,照片核对耗时 4.5 到 5 秒,发热时则为 5.7 到 6.5 秒。峰值内存占用为 2.2 到 2.6 GB,因此应用在每次使用时加载模型,用完即释放。热重载(Warm reload)耗时约 0.5 秒,内存占用会降至约 240 MB。

Safety cannot live in the prompt. The model still produced “a coiled snake”, “orange fungus” and “a bird’s nest hidden in a hollow”, and later “a cloud shaped like a deer”. For a kids’ app that is not acceptable, so every item now passes a deterministic validator in code (English and Spanish word lists: no picking or touching, no animals, insects or mushrooms, nothing near water or roads, no climbing, nothing to eat, no photos of people). Rejected items are regenerated, and if the model still falls short, a built-in safe list fills the gap. The model’s feedback is checked too, because it is read aloud to children. 安全性不能仅依赖提示词。模型曾生成过“盘绕的蛇”、“橙色真菌”、“藏在树洞里的鸟巢”,甚至“像鹿一样的云”。对于儿童应用来说,这是不可接受的。因此,现在每个项目都必须通过代码中的确定性验证器(包含英西双语词库:禁止采摘或触摸、禁止动物、昆虫或蘑菇、禁止靠近水边或道路、禁止攀爬、禁止食用、禁止拍摄人物)。被拒绝的项目会重新生成,如果模型仍无法满足要求,则由内置的安全清单补齐。模型的反馈也会经过检查,因为这些内容会朗读给孩子听。

Bugs the phone found that tests did not. The same settings produced the identical hunt three times in a row. LiteRT-LM’s sampler seed defaults to 0, which makes sampling deterministic. Each request now gets a random seed. A regex that passed every JVM test crashed on Android, because Android’s regex engine rejects a bare }. 手机运行中发现而测试未发现的 Bug:相同的设置连续三次生成了完全相同的寻宝清单。LiteRT-LM 的采样器种子默认为 0,导致采样具有确定性。现在每个请求都会分配一个随机种子。一个在所有 JVM 测试中都能通过的正则表达式在 Android 上崩溃了,因为 Android 的正则表达式引擎拒绝接受单独的 }。