Snapshots & epoch coordination¶
Code:
src/Arius.Core/Shared/Snapshot/(SnapshotService.cs,SnapshotManifest.cs,SnapshotSerializer.cs,FileTreeHashJsonConverter.cs) · Decisions: ADR-0002 · ADR-0016 · ADR-0017 · Terms: snapshot · epoch · filetree
Purpose¶
A snapshot is the immutable, point-in-time manifest that names one complete repository state: the root filetree hash plus its totals. SnapshotService creates, lists, and resolves snapshots against snapshots/<timestamp> blobs with a local plain-JSON cache. The latest snapshot is also the repository's commit point and the epoch marker the rest of the cache stack coordinates against.
How it works¶
The manifest¶
SnapshotManifest is a tiny immutable record — it does not embed the tree, only its root:
| Field | Meaning |
|---|---|
Timestamp |
UTC creation time; doubles as the blob name and version id |
RootHash |
root FileTreeHash (SHA-256 hex) produced by the tree builder |
FileCount / OriginalSize |
snapshot totals: file count and the logical size — summed original (uncompressed) bytes of all files, counting duplicates once per file (the size you would restore). Not deduplicated or compressed. |
AriusVersion |
tool version that wrote it — shared AriusVersion.Informational, stamped from the git tag at publish |
The manifest is the root of the Merkle structure: RootHash points at the top filetree blob, which references the rest. FileTreeHashJsonConverter persists RootHash as its canonical lowercase-hex string.
Two storage representations of the same manifest:
- Disk (SnapshotService.SerializerOptions, indented): plain JSON at ~/.arius/{account}-{container}/snapshots/<timestamp> via a RelativeFileSystem rooted at RepositoryLocalStatePaths.GetSnapshotCacheRoot.
- Azure wire (SnapshotSerializer, compact): JSON → compress → optional encrypt, uploaded at BlobPaths.SnapshotPath(timestamp) on the Cool tier with content type ContentTypes.SnapshotPlaintext (application/zstd) or SnapshotGcmEncrypted (application/aes256gcm+zstd).
Create (write-through) and resolve (disk-first)¶
flowchart TD
subgraph create[CreateAsync — write-through]
C1[Build manifest at ts] --> C2[WriteToDiskAsync: plain JSON to local snapshots/]
C2 --> C3[SnapshotSerializer.SerializeAsync: zstd + optional encrypt]
C3 --> C4["UploadAsync snapshots/<ts>, overwrite:false (default), Cool tier"]
end
subgraph resolve[ResolveAsync — disk-first]
R1[ListBlobNamesAsync: list snapshots/, sort by ParseTimestamp] --> R2{version?}
R2 -->|null| R3["pick names[^1] (latest)"]
R2 -->|prefix| R4[StartsWith match, last wins]
R3 --> R5{local JSON exists?}
R4 --> R5
R5 -->|yes| R6[read + deserialize local]
R5 -->|no| R7[LoadFromAzureAsync] --> R8[WriteToDiskAsync, then return]
end
CreateAsync writes the disk copy before the upload. The local file is the per-machine "I wrote the last snapshot" marker (see invariants); writing it before the upload would be wrong, so the order is disk-then-Azure but both must succeed for the marker to be meaningful — the marker is only authoritative once the matching remote blob also exists.
ResolveAsync always lists snapshots/ first to pick the target name (latest, or the last name whose filename StartsWith the requested version), then prefers the local JSON cache and only falls back to Azure on a miss — populating the disk cache on the way out. ListBlobNamesAsync returns names sorted oldest→newest by ParseTimestamp.
Snapshot as the epoch / commit point¶
The latest snapshot name is the single coordination point for the whole cache stack. Two consumers read it:
FileTreeService.ValidateAsynccompares the latest local snapshot name to the latest remote one once per archive run. Equal ⇒ fast path (this machine was last writer, local caches fully trusted, nofiletrees/listing). Different ⇒ slow path (another machine archived; materialize markers and signal a mismatch, which makesArchiveCommandHandlercallChunkIndexService.InvalidateCaches()before flush). See ADR-0016 and chunk-index / filetree for the consuming side.ArchiveCommandHandlerstage 6d is the only place a snapshot is created, and it is the last durability step — every chunk, thin chunk, filetree blob, and chunk-index shard it references is already durable. AfterCreateAsyncit callsChunkIndexService.PromoteToSnapshotVersionAsync(<snapshot name>)so validated coverage claims carry forward to the new epoch.
sequenceDiagram
participant H as ArchiveCommandHandler (stage 6d)
participant S as SnapshotService
participant CI as ChunkIndexService
H->>S: ResolveAsync() — latest
alt rootHash == latest.RootHash (no-op)
H-->>H: reuse latest Timestamp + RootHash, no write
else changed
H->>S: CreateAsync(rootHash, fileCount, originalSize)
S->>S: WriteToDisk (epoch marker) then Upload
H->>CI: PromoteToSnapshotVersionAsync(snapshot name)
H-->>H: publish SnapshotCreatedEvent
end
A re-archive of unchanged data rebuilds the same rootHash; stage 6d resolves the latest snapshot, sees latestSnapshot.RootHash == rootHash, and reuses it instead of publishing a new manifest — no-op archives create no history. See ADR-0002.
Key invariants¶
- Manifests are immutable and content-rooted. A snapshot is never edited;
RootHashis the SHA-256 root of the filetree, so two snapshots with the same root describe the same state.CreateAsyncdefaults tooverwrite: false(each normal run gets a freshnowtimestamp, so manifests are append-only). The one exception is a caller that creates a snapshot at a deterministictimestampand needs re-runs to be idempotent — the v5→v7 migration passes an explicittimestamp(the source state's version) withoverwrite: trueso replaying it rewrites the same blob rather than failing on a conflict. - Timestamp lexicographic order == chronological order.
TimestampFormat = "yyyy-MM-ddTHHmmss.fffZ"(UTC, zero-padded) makes "latest" a plain string sort on both disk filenames and blob names —ResolveAsyncandValidateAsyncrely on this to pick the epoch without parsing every name. - Disk-write-before-upload, and "latest local == latest remote" ⇒ this machine wrote last. The local
snapshots/marker is written only when this machine completes an archive, so name equality is exactly the fast-path proposition inFileTreeService.ValidateAsync(ADR-0016). - Snapshot last. The manifest is published only after all data it references is durable; a resolvable snapshot is therefore always fully restorable (ADR-0017).
- No-op archives reuse the latest snapshot. Equal root hash ⇒ no new manifest (ADR-0002); callers that need to know whether a new snapshot was published must compare versions, not just archive success.
- The version id is the timestamp filename.
GetVersion/ParseTimestamp/ResolveAsync(version)all key off the baresnapshots/<name>filename;GetSnapshotFileNamerejects any blob name not directly under root orSnapshotsPrefix.
Why this shape¶
- Snapshot as epoch instead of a distributed lock. The snapshot already exists as the repository's single commit point, so it doubles as a free coherence marker — the same-machine repeat-archive path stays nearly listing-free. Rationale and the accepted last-writer-wins tradeoff: ADR-0016.
- Snapshot-last commit, metadata-presence recovery. Publishing the manifest only after referenced data is durable makes re-running the command the entire recovery procedure, with no coordinator: ADR-0017.
- Skip no-op snapshots. Snapshot history records state changes, not command invocations: ADR-0002.
- Two serializers (disk plain, Azure compressed+encrypted). The local cache is human-readable plain JSON for inspection; the wire format mirrors how every other blob is stored. The
FileTreeHashconverter keeps both byte-identical in their JSON shape.
Open seams / future¶
- Concurrent multi-machine writes are tolerated but not coordinated. Two machines publishing snapshots interleave by timestamp; the chunk-index shard rewrites underneath them are last-writer-wins, recovered only by
RepairAsync. ETag-conditional shard writes are the noted future hardening (ADR-0016). - Crash between snapshot upload and local marker write costs a spurious slow path (a re-list + shard revalidation) on the next run — a performance cost, not a correctness bug.
versionresolution is a prefixStartsWithover the full listing. Fine at human snapshot counts; if a single repository ever accumulates very many snapshots, bothResolveAsyncandListBlobNamesAsyncmaterialize and sort the entiresnapshots/listing per call.- No retention / pruning. Snapshots accumulate indefinitely; there is no expiry or GC of old manifests today.