Here (e.g., git switch``git restoreSparse checkout) are unavailable in older versions.
Git’s design is the product of several deliberate trade-offs, each motivated by the Linux kernel Workflow:
Every repository clone contains the complete object database, every commit, every tree, every Blob. This means:
- Offline operation:
git log``git diff``git blame``git show all work without network access. You can commit, branch, and merge entirely offline. - Speed: Local operations read from the filesystem, not the network.
git log on a cold repository scans the local object store. - Resilience: No single point of failure. If the remote server burns down, any clone can recreate it entirely with
git push --mirror.
The cost is disk space, a full clone of the Linux kernel is ∼5 GB. Mitigations exist (shallow clones, sparse checkout, partial clone), but the default is to replicate everything.
Most VCS (CVS, Subversion, Perforce) store a series of deltas: file v2 is expressed as “file v1 with these lines changed.” Git instead stores full snapshots of the entire project tree at each commit. If a file has not changed between two commits, Git does not store it again, it stores A pointer to the identical blob object.
This design choice has deep implications:
- Content-addressable storage: Every object is identified by the SHA-1 hash (or SHA-256, as of Git 2.29) of its content. Two identical files at different paths or in different commits produce the same blob object. This deduplication is automatic and transparent.
- Fast branching: Creating a branch is a O(1) operation, it writes a 41-byte reference file. There is no copying of file data.
- Merge correctness: Three-way merge compares full tree snapshots, not a chain of deltas, which makes it robust against complex history topologies.
The cost is that Git’s object store can appear larger than a delta-based store for repositories with very large files that change frequently. This is why Git added the packfile format (see Internals: Packing and Garbage Collection) to Compress objects using delta compression between similar objects.
Every Git object (blob, tree, commit, tag) is identified by a cryptographic hash of its content plus header. This means:
- Tamper detection: If a single byte in any object is modified, its hash changes, and all objects referencing it become invalid.
git fsck can detect this. - Deterministic builds: Given the same source tree and the same commit hash, you are guaranteed the same content. This is foundational for reproducible builds and supply-chain security.
- No ambiguity: A commit hash uniquely identifies a snapshot of the entire project. Two developers referring to
a3f2b1c are guaranteed to be referring to the same state.
With the exception of git fetch``git pull``git push``git cloneAnd git ls-remoteEvery Git operation works on local data. This was a hard requirement for the Linux kernel workflow, where Contributors on dial-up connections needed to work efficiently.
| Feature | Git | Mercurial (Hg) | Subversion (SVN) | Perforce (P4) |
|---|
| Architecture | Distributed | Distributed | Centralized | Centralized |
| Storage model | Content-addressable snapshots | Content-addressable snapshots | Delta-based | Delta-based (server-side) |
| Branching model | Pointer-based (O(1)) | Bookmark-based (O(1)) | Directory copy (O(n)) | Streams (server-side) |
| Offline commits | Full | Full | No | Limited (shelving) |
| Performance at scale | Excellent (Linux kernel, Chromium) | Good (Facebook used it) | Degrades with large trees | Excellent with Helix Core |
| Learning curve | Steep | Moderate | Shallow | Steep |
| Binary file handling | Poor (use Git LFS) | Poor (use Largefiles) | Good | Good |