Show HN: AI search for every photo and every frame of video on macOS

Show HN: AI search for every photo and every frame of video on macOS

SCM — Screen Memories Deep AI search for every photo and every frame of video in any folder on macOS. Local-first — no accounts, no cloud, no uploads. Inference runs on your Mac. SCM (Screen Memories) 是一款适用于 macOS 的深度 AI 搜索工具,可搜索任何文件夹中的每一张照片和视频的每一帧。它采用本地优先原则——无需账户、无需云端、无需上传。所有推理过程均在你的 Mac 上运行。

What makes it different 它有何不同?

Search like you think — describe a memory in plain language; a local vision model does the rest. 像思考一样搜索——用通俗的语言描述一段记忆,剩下的交给本地视觉模型处理。

Video, down to the moment — scenes are segmented and embedded, so you land on the shot, not just the file. 视频搜索,精确到瞬间——场景会被分割并嵌入,因此你可以直接定位到具体的镜头,而不仅仅是文件本身。

Text and dialogue too — OCR over visible text; exact spoken-line search via Whisper, each as its own mode. 支持文本和对话——通过 OCR 识别可见文本;通过 Whisper 进行精确的语音行搜索,每种功能都有独立的模式。

Your tabs, your prompts — save any query as a tab; Screenshots and Email tabs are toggleable. 标签页与提示词——将任何查询保存为标签页;截图和电子邮件标签页均可切换。

Self-maintaining library — watched folders auto-import, content hashes dedupe renames, and model switches re-embed in the background without blocking search. 自动维护的库——自动导入受监控文件夹,通过内容哈希处理重命名后的去重,模型切换时会在后台重新嵌入,且不会阻塞搜索。

Truly private — your media never leaves the machine. Weights download once; everything after that is offline. 真正的隐私保护——你的媒体文件永远不会离开设备。权重仅需下载一次,之后所有操作均处于离线状态。

Five ways to search / 五种搜索方式

ModeFinds
FilesWhole photos/videos by meaning — vision rank with filename and phrase boosts
ScenesMoments inside video — search a shot, jump to its timecode
OCRText visible in images and frames, matched literally (Tesseract; eng + 35 language toggles)
DialogueExact spoken words in videos (Whisper), tiered exactness
LLMs (opt-in)Local chat over the dialogue, OCR, and filenames your Mac already extracted — cited answers
模式查找内容
文件 (Files)按含义查找整个照片/视频——通过文件名和短语加权进行视觉排序
场景 (Scenes)视频内的瞬间——搜索镜头,跳转至对应时间码
OCR图像和帧中可见的文本,进行字面匹配(Tesseract;支持英语 + 35 种语言切换)
对话 (Dialogue)视频中精确的口述词汇(Whisper),分级精确度
LLM (可选)基于 Mac 已提取的对话、OCR 和文件名进行本地聊天——提供带引用的回答

Requirements / 系统要求

macOS (packaged with electron-builder; menu-bar/tray features are macOS-only) macOS(使用 electron-builder 打包;菜单栏/托盘功能仅限 macOS)

Bun — the project uses bun as package manager and runner Bun — 该项目使用 bun 作为包管理器和运行环境

Node modules installed: bun install 安装 Node 模块:bun install

First use of a model downloads its weights (~435MB for the default CLIP); after that, fully offline. 首次使用模型时会下载权重(默认 CLIP 约为 435MB);之后完全离线。

Download / 下载

The easiest install is via Homebrew (Apple Silicon, macOS 12+). The tap’s cask clears the macOS quarantine flag automatically on every install and upgrade, so the app launches with no manual Gatekeeper steps: 最简单的安装方式是通过 Homebrew(适用于 Apple Silicon,macOS 12+)。该 tap 的 cask 会在每次安装和升级时自动清除 macOS 的隔离标志,因此启动应用时无需手动执行 Gatekeeper 步骤:

brew tap allenv0/scm
brew trust allenv0/scm
brew install --cask allenv0/scm/scm

Upgrades keep the same behavior: brew upgrade --cask allenv0/scm/scm 升级时保持同样的操作:brew upgrade --cask allenv0/scm/scm

Prefer least privilege? Trust just the cask instead of the whole tap: 更倾向于最小权限原则?仅信任 cask 而非整个 tap:

brew tap allenv0/scm
brew trust --cask allenv0/scm/scm
brew install --cask scm

The tap lives at allenv0/homebrew-scm. 该 tap 位于 allenv0/homebrew-scm。

Development / 开发

bun run dev # build the renderer bundle, then launch the Electron app
bun start # launch the Electron app without rebuilding
bun run build # just rebuild the renderer bundle into dist/

Building the app package (DMG / ZIP) 构建应用包 (DMG / ZIP)

bun run dist # signed if an identity is in the keychain
bun run dist:unsigned # skip code-sign discovery

This runs two steps in sequence: vite build — compiles the React renderer into dist/ (picks up all changes under src/). electron-builder --mac — packages the app. It bundles the fresh dist/ bundle together with main.js, preload.js, main-lib/, and indexer/ (the file list is configured under build.files in package.json), then produces the installers. 此操作按顺序执行两个步骤:vite build — 将 React 渲染器编译到 dist/(获取 src/ 下的所有更改)。electron-builder --mac — 打包应用。它将最新的 dist/ 包与 main.js、preload.js、main-lib/ 和 indexer/ 一起打包(文件列表在 package.json 的 build.files 下配置),然后生成安装程序。

Output: the installers land in dist-app/ (see build.directories.output in package.json) — look for SCM-0.2.4.dmg and SCM-0.2.4.zip. 输出:安装程序位于 dist-app/(参见 package.json 中的 build.directories.output)— 请查找 SCM-0.2.4.dmg 和 SCM-0.2.4.zip。

Features / 功能特性

Files Typing starts an instant filename-keyword pre-pass, then the vision model takes over: results are scored by cosine similarity against image embeddings, with gated phrase and filename boosts, an honesty floor calibrated per model, and a near-duplicate diversity filter. Every tile carries a “why it matched” badge (Visual match / Filename match / …) and a hover tooltip with the per-component score breakdown. CJK queries search as overlapping bigrams (“台北車站” also matches 台北, 車站). 文件 输入时会立即进行文件名关键词预筛选,随后视觉模型接管:结果通过图像嵌入的余弦相似度进行评分,并结合短语和文件名加权、针对各模型校准的置信度下限以及近重复多样性过滤器。每个磁贴都带有“匹配原因”徽章(视觉匹配/文件名匹配等),悬停时会显示各组件的得分明细。中日韩(CJK)查询以重叠双字词组方式搜索(“台北車站”也会匹配“台北”和“車站”)。

Scenes Every scene segment across all videos is scored, so a hit lands on the exact shot: tiles show the scene poster with a timecode badge, and opening the video jumps straight to that moment. A noise gate returns “no scene match” instead of flooding the grid with gibberish, and each video contributes at most 3 scenes. 场景 所有视频中的每个场景片段都会被评分,因此搜索结果会精确到镜头:磁贴显示带有时间码徽章的场景海报,打开视频即可直接跳转到该时刻。噪声门控会返回“无场景匹配”,而不是用乱码填满网格,且每个视频最多贡献 3 个场景。

OCR Matches the fraction of query tokens literally visible in each image’s OCR text — the filename is ignored and no vision model is involved, so it works even while the AI engine is warming up or offline. Matched words are boxed in amber on tiles and in the lightbox. OCR 匹配图像 OCR 文本中字面可见的查询标记部分——忽略文件名且不涉及视觉模型,因此即使在 AI 引擎预热或离线时也能工作。匹配的单词在磁贴和灯箱视图中会用琥珀色框标出。

Dialogue Exact literal retrieval over Whisper transcripts — no embeddings, no thresholds, works with the AI engine down. Results come in three tiers: Exact line (contiguous phrase in one utterance), Exact words (all words in one utterance or an ≤8s window), and Words spoken (all words in the same video). Matching words are highlighted in a speech snippet; opening the result seeks straight to the line. 对话 基于 Whisper 转录的精确字面检索——无需嵌入,无需阈值,即使 AI 引擎关闭也能工作。结果分为三个层级:精确行(单次发言中的连续短语)、精确词(单次发言或 ≤8 秒窗口内的所有词)以及口述词(同一视频中的所有词)。匹配的词汇会在语音片段中高亮显示;打开结果可直接跳转到该行。

LLMs (Ask) Opt-in — nothing downloads or runs until enabled in Settings → LLMs Chat. A llama.cpp sidecar bound to loopback answers your question from evidence the app already extracted — dialogue lines, OCR text, and filename keyword hits — with numbered citations you can click, streamed token-by-token with a live tok/s readout. Leading /screenshots, /videos, /email narrow the corpus; Stop keeps the partial answer; empty evidence short-circuits before the model ever runs. LLM (询问) 可选功能——在“设置 → LLMs Chat”中启用前不会下载或运行任何内容。一个绑定到本地回环的 llama.cpp 侧边程序会根据应用已提取的证据(对话行、OCR 文本和文件名关键词)回答你的问题,并提供可点击的带编号引用,以逐 token 流式传输并实时显示 tok/s。使用 /screenshots、/videos、/email 前缀可缩小搜索范围;“停止”可保留部分回答;若无证据,模型在运行前会直接跳过。

Chat modelSizeNotes
Qwen3 1.7B (default)~1.1GBFast everyday chat; fits 8GB Macs
Llama 3.2 3B~2GBStronger long answers; needs headroom
聊天模型大小备注
Qwen3 1.7B (默认)~1.1GB快速日常聊天;适合 8GB 内存的 Mac
Llama 3.2 3B~2GB更强的长文本回答;需要更多内存空间

Tabs & library views Built-in browse tabs: All, Videos, plus Screenshots and Email — the latter two toggleable in Settings → Smart Tabs. Selecting Videos auto-enables Scenes mode. Save any query as a tab: the pin pill under the search bar saves the current prompt with its mode (Files/Scenes/OCR/Dialogue) — up to 20 tabs, renameable, each restored exactly as saved. Five semantic views behind remappable shortcuts (⌘1–⌘5 by default), plus ⌘I import / AI Insights, ⌘, for Settings. Search stays in the selected tab: pick Screenshots, Email, Videos, or any saved tab and results are filtered to it — scope first, then search. In LLMs chat the same idea is explicit: leading /screenshots, /videos, /email narrow the corpus before the model ever runs. 标签页与库视图 内置浏览标签页:全部、视频,以及截图和电子邮件——后两者可在“设置 → 智能标签页”中切换。选择“视频”会自动启用“场景”模式。可将任何查询保存为标签页:搜索栏下方的固定按钮可保存当前提示词及其模式(文件/场景/OCR/对话)——最多支持 20 个标签页,可重命名,每个标签页都能精确恢复。五个语义视图可通过可重映射的快捷键(默认 ⌘1–⌘5)访问,另有 ⌘I 用于导入/AI 洞察,⌘, 用于设置。搜索会保留在选定的标签页中:选择截图、电子邮件、视频或任何已保存的标签页,结果将过滤至该范围——先确定范围,再进行搜索。在 LLM 聊天中,这一逻辑同样明确:使用 /screenshots、/videos、/email 前缀可在模型运行前缩小语料库范围。

Email tab Surfaces photos whose visible OCR text contains an email address — an overlapping view (a photo keeps its category too). Detection is OCR-tolerant: 电子邮件标签页 展示可见 OCR 文本中包含电子邮件地址的照片——这是一个重叠视图(照片仍保留其原有类别)。检测具有 OCR 容错性: