Expand description
Reader for Apache Jackrabbit Oak segment-tar (TarMK) repositories.
TarMK is the storage engine used by Apache Jackrabbit Oak and Adobe
Experience Manager: content is stored as immutable segments packed into
tar archives, and a journal records the sequence of repository head
states. This crate opens such a repository directly from disk, resolves
the current head state, and exposes the content tree for traversal and
extraction — without a running Oak instance.
The reading API (Repository, store, content, tooling) is
read-only by design: it never takes the repository lock and never
modifies any file, so it is safe to point at a live repository. Like Oak,
the reader memory-maps archives and relies on the store’s
never-modify-in-place file protocol; an external process that truncates or
rewrites an archive would disturb both froe and a running Oak instance.
The mutating writing API (writer) covers commits, checkpoints, applying
cleanup, compaction, backup, restore, and journal recovery. It takes the
exclusive repository lock first, so it cannot race a cooperating running
instance, and produces stores byte-for-byte compatible with Oak (apart from
the documented extreme-subnormal rendering residue; see
content::property::double_to_text). Planning cleanup is the read-only
exception and never takes the lock. Run mutations only against a stopped
repository. The writer requires a Unix operating-system entropy source and
therefore refuses to open on Windows.
The writing API is verified against a real Oak instance: the
workspace interoperability suite round-trips it through Apache
Jackrabbit Oak oak-segment-tar 1.90.0 — Oak writes the store, froe
commits, checkpoints, compacts, cleans up, backs up, restores and
recovers the journal, and Oak then boots against each result and serves
a byte-identical content tree without logging any of its own repair
messages. Still unverified against a live instance: store.version=1
stores, external blob stores, native macOS or Windows execution, and
Adobe AEM itself, which ships its own Oak build. Writing still requires
a stopped repository, and keeping a copy before a destructive operation
on irreplaceable data remains ordinary prudence.
§Example
use froe::store::Repository;
fn main() -> froe::Result<()> {
let repository = Repository::open(std::path::Path::new("/path/to/segmentstore"))?;
if let Some(node) = repository.node_at_path("/content")? {
for property in node.properties()? {
println!("{} = {:?}", property.name, property.values);
}
for (name, child) in node.child_node_entries()? {
println!("{name}: {} children", child.child_node_count()?);
}
}
Ok(())
}§Layers
Each layer is usable on its own:
tar_archive— archives, indexes, segment graphs, binary reference catalogs;segment— segment parsing and record addressing;content— decoding records into nodes, properties, and values;journal— the head revision log;gc_journal— optional garbage-collection history;store— the assembled read-only repository.
Long-running operations — opening a large store, planning a cleanup,
compacting, checking consistency — have a _with_progress twin that
reports what they are doing to a progress::ProgressObserver, so a
caller need not guess whether a silent minute means work or a hang.
Custom backends (in-memory fixtures, remote stores) implement
SegmentProvider and reuse the whole content layer unchanged.
Re-exports§
pub use content::BinaryStream;pub use content::BinaryValue;pub use content::NodeState;pub use content::PropertyState;pub use content::PropertyType;pub use content::PropertyValue;pub use content::PropertyValues;pub use content::SegmentProvider;pub use content::Template;pub use content::read_binary_stream;pub use error::Error;pub use error::Result;pub use gc_journal::GarbageCollectionJournalEntry;pub use journal::JournalEntry;pub use progress::DiscardedProgress;pub use progress::ProgressObserver;pub use progress::Step;pub use progress::WorkUnit;pub use segment::GarbageCollectionGeneration;pub use segment::RecordIdentifier;pub use segment::RecordType;pub use segment::SegmentIdentifier;pub use segment::SegmentKind;pub use store::Repository;pub use writer::CleanupAction;pub use writer::CleanupDeletionFailure;pub use writer::CleanupOptions;pub use writer::CleanupOutcome;pub use writer::CleanupPlan;pub use writer::CleanupTask;pub use writer::CompactionKind;pub use writer::CompactionOutcome;pub use writer::JournalLineRemoval;pub use writer::JournalRemovalReason;pub use writer::PreparedCleanup;pub use writer::RecoveryBackupPolicy;pub use writer::RecoveryOutcome;pub use writer::StaleArchiveReason;pub use writer::WritableRepository;pub use writer::backup;pub use writer::backup;pub use writer::backup_with_progress;pub use writer::cleanup;pub use writer::cleanup;pub use writer::cleanup_with_progress;pub use writer::compact;pub use writer::compact_with_progress;pub use writer::plan_cleanup;pub use writer::plan_cleanup_with_progress;pub use writer::recover_journal;pub use writer::recover_journal_with_progress;pub use writer::restore;pub use writer::restore_with_progress;
Modules§
- checksum
- CRC32 checksum as used by the segment-tar format.
- content
- The content layer: decoding records into the content tree.
- error
- Error and result types shared across the crate.
- gc_
journal - The garbage-collection journal (
gc.log). - hashing
- Hash functions that must match their Java counterparts bit for bit.
- journal
- The journal: the repository’s sequence of head states.
- progress
- Progress observation for long-running operations.
- segment
- Segments: the unit of storage below the tar archive layer.
- store
- The repository: opening a segment store directory read-only.
- tar_
archive - The tar archive layer: segment archives on disk.
- tooling
- Read-only diagnostic tooling: consistency checking, revision differencing, node history, node search, segment dumps, and archive path attribution.
- writer
- The write path: building segments, archives, and repository state.