Git's Staging Area — What Actually Happens Between `git add` and `git commit`
Git 的暂存区——git add 和 git commit 之间到底发生了什么?
Edit a file, run git add ., then git commit -m ”…” — for most people these two commands feel like a single operation performed in two keystrokes. But why are they separate commands at all? What happens if you edit the file again after git add but before git commit? Many developers who use Git daily can’t answer either question cleanly. This article walks through the staging area, the mechanism sitting between those two commands.
编辑文件,运行 git add .,然后执行 git commit -m "..." —— 对大多数人来说,这两个命令感觉就像是分两次按键完成的单一操作。但为什么它们被设计成两个独立的命令呢?如果在 git add 之后但在 git commit 之前再次编辑文件会发生什么?许多每天使用 Git 的开发者都无法清晰地回答这两个问题。本文将带你深入了解暂存区(staging area),即位于这两个命令之间的核心机制。
The three-area mental model
三层模型思维模式
Note: A common way to understand Git is the three-layer model — the working directory, the staging area (also called the index), and the repository. A change has to pass through all three, in order, before it becomes part of the commit history.
注:理解 Git 的一种常用方式是“三层模型”——工作目录(working directory)、暂存区(staging area,也称为索引 index)和仓库(repository)。一项更改必须按顺序通过这三层,才能成为提交历史的一部分。
The working directory is where you actually edit files in your editor. From Git’s point of view, changes here are just filesystem differences — they have no effect on history at all until something moves them forward.
工作目录是你实际在编辑器中修改文件的地方。从 Git 的角度来看,这里的更改仅仅是文件系统的差异——在有操作将其推进之前,它们对历史记录没有任何影响。
The repository is the committed history you see with git log. Once a commit exists, its content is fixed unless you explicitly rewrite it with something like git commit —amend or git reset.
仓库是你通过 git log 看到的已提交历史。一旦提交存在,其内容就是固定的,除非你使用 git commit --amend 或 git reset 等命令显式地重写它。
Sitting between those two is the staging area. It’s a temporary holding place for exactly the parts of your working-directory changes that you want included in the next commit — nothing more, nothing less. git add copies changes from the working directory into the staging area; git commit takes a snapshot of whatever is currently in the staging area and permanently records it in the repository. That’s the entire division of labor.
位于这两者之间的是暂存区。它是一个临时存放处,专门存放你希望包含在下一次提交中的工作目录更改部分——不多也不少。git add 将更改从工作目录复制到暂存区;git commit 则对暂存区当前的内容进行快照,并将其永久记录到仓库中。这就是它们全部分工。
What git add actually does
git add 到底做了什么
Under the hood, git add does two things. First, it takes the content of the changed file and stores it as a blob object under .git/objects/, keyed by a content hash. Second, it writes a mapping into a binary file called .git/index: “this path now points to the blob with this hash.”
在底层,git add 做了两件事。首先,它获取更改后的文件内容,并将其作为 blob 对象存储在 .git/objects/ 下,以内容哈希作为键。其次,它向一个名为 .git/index 的二进制文件写入映射:“此路径现在指向具有该哈希值的 blob”。
The key detail is that .git/index records a snapshot of the file’s content at the exact moment git add ran — nothing more current than that. If you edit the file again after staging it, the working directory changes, but the staging area still points to the old blob. Run git status in that state and the same file shows up under both “changes to be committed” and “changes not staged for commit” at once. That’s not a bug — it’s the direct consequence of the staging area being a frozen copy rather than a live reference.
关键细节在于,.git/index 记录的是 git add 运行那一刻的文件内容快照——不会比那一刻更新。如果你在暂存后再次编辑文件,工作目录会发生变化,但暂存区仍然指向旧的 blob。在这种状态下运行 git status,同一个文件会同时出现在“要提交的变更”(changes to be committed)和“未暂存以备提交的变更”(changes not staged for commit)中。这不是 Bug,而是因为暂存区是一个冻结的副本,而不是实时引用,这正是其必然结果。
$ echo "v1" > file.txt
$ git add file.txt
$ echo "v2" > file.txt
$ git status
Changes to be committed:
modified: file.txt ← the index still holds v1
Changes not staged for commit:
modified: file.txt ← the working directory now has v2
What git commit actually does
git commit 到底做了什么
git commit builds a tree object out of whatever is currently in the staging area (the index), then wraps that tree together with a pointer to the parent commit and metadata (author, timestamp, message) into a commit object. What it reads from is always the index — never the live state of the working directory. Anything you never staged, or staged and then edited again afterward, is simply not part of the commit.
git commit 根据暂存区(索引)中当前的内容构建一个树对象(tree object),然后将该树与指向父提交的指针以及元数据(作者、时间戳、提交信息)打包成一个提交对象。它读取的始终是索引,绝不是工作目录的实时状态。任何你未暂存的内容,或者暂存后又再次修改的内容,都不会包含在提交中。
git commit -a is a shortcut that makes this look like a single step: internally, it silently runs git add first, but only for files Git is already tracking (newly created files are excluded). If you don’t know that automatic staging is happening behind the scenes, you can be caught off guard by a commit that’s missing a brand-new file you clearly just created.
git commit -a 是一个让操作看起来像一步完成的快捷方式:在内部,它会静默地先运行 git add,但仅针对 Git 已经跟踪的文件(新创建的文件会被排除在外)。如果你不知道后台正在进行自动暂存,可能会因为提交中缺少了你明明刚创建的新文件而感到措手不及。
Why split it into two steps at all
为什么要分成两步?
A version-control system that committed the working directory directly, with no intermediate step, would still work in principle. The reason Git inserts a staging layer is to let you control what goes into a single commit independently of how you actually edited the files.
原则上,一个直接提交工作目录而没有中间步骤的版本控制系统也是可以工作的。Git 引入暂存层的原因是让你能够独立于实际编辑文件的方式,来控制哪些内容进入单次提交。
A common scenario: you make two unrelated changes to the same file in one editing session — say, a bug fix and an incidental formatting cleanup. git add -p file.txt lets you stage changes at the granularity of individual hunks rather than whole files, so you can pick “just this part” interactively. That makes it possible to commit the bug fix on its own and leave the formatting change for a separate commit — a kind of after-the-fact reorganization that would otherwise require manually copying and diffing files by hand.
一个常见的场景是:你在一次编辑会话中对同一个文件做了两项不相关的更改——比如一个 Bug 修复和一个顺手的格式清理。git add -p file.txt 允许你以代码块(hunk)为粒度而非整个文件来暂存更改,因此你可以交互式地选择“仅这一部分”。这使得你可以单独提交 Bug 修复,而将格式更改留到另一次提交中——这是一种事后的重组,否则就需要手动复制和对比文件。
The other practical use is reviewing exactly what’s about to be committed with git diff —cached (equivalently git diff —staged). While plain git diff shows the difference between the working directory and the index, git diff —cached shows the difference between the index and the last commit — in other words, the exact content that’s about to become the next commit. Making a habit of running this after git add and before the actual commit gives you one last checkpoint to catch a stray debug print statement, or a file that shouldn’t be part of the commit at all, before it becomes permanent history.
另一个实际用途是使用 git diff --cached(等同于 git diff --staged)来精确审查即将提交的内容。普通的 git diff 显示的是工作目录与索引之间的差异,而 git diff --cached 显示的是索引与上一次提交之间的差异——换句话说,就是即将成为下一次提交的确切内容。养成在 git add 之后、实际提交之前运行此命令的习惯,能为你提供最后一道检查点,以便在内容成为永久历史之前,捕获遗漏的调试打印语句或不应包含在提交中的文件。
A small habit worth building
一个值得养成的小习惯
In workflows where multiple people — or an automated agent — are creating commits, staging specific files by name tends to be safer than a blanket git add -A, followed by a git status check before committing. git add -A stages every change under the current directory indiscriminately, which means it can just as easily pick up a stray temp file, or a credentials file someone accidentally left in the working tree. Having a staging area — a deliberate holding pen between “edited” and “committed” — is what makes it possible to insert that kind of mechanical check at the one moment it can still do any good: after git add, before git commit.
在多人协作或有自动化代理创建提交的工作流中,按名称暂存特定文件通常比盲目使用 git add -A 并随后在提交前进行 git status 检查更安全。git add -A 会不加区分地暂存当前目录下的所有更改,这意味着它很容易把遗留的临时文件或某人不小心留在工作树中的凭据文件也一并暂存。拥有一个暂存区——即在“已编辑”和“已提交”之间的一个刻意的缓冲地带——使得我们能够在最关键的时刻插入这种机械检查:即在 git add 之后、git commit 之前。
Summary
总结
git add and git commit being separate commands isn’t a historical accident — it’s a deliberate design choice to decouple “what you’ve edited” from “what goes into the next commit.” The staging area is the temporary snapshot that makes that decoupling possible, and both hunk-level staging with git add -p and pre-commit review with git diff —cached only exist because of this intermediate layer.
git add 和 git commit 作为独立命令并非历史偶然,而是一种刻意的设计选择,旨在将“你编辑了什么”与“什么进入下一次提交”解耦。暂存区是实现这种解耦的临时快照,而使用 git add -p 进行代码块级暂存以及使用 git diff --cached 进行提交前审查,都是因为有了这一中间层才得以存在。