Expand description
Reader for Apache Jackrabbit Oak segment-tar (TarMK) repositories.
TarMK is the storage engine used by Apache Jackrabbit Oak and Adobe
Experience Manager: content is stored as immutable segments packed into
tar archives, and a journal records the sequence of repository head
states. This crate opens such a repository directly from disk, resolves
the current head state, and exposes the content tree for traversal and
extraction — without a running Oak instance.
The reading API (Repository, store, content, tooling) is
read-only by design: it never takes the repository lock and never
modifies any file, so it is safe to point at a live repository. Like Oak,
the reader memory-maps archives and relies on the store’s
never-modify-in-place file protocol; an external process that truncates or
rewrites an archive would disturb both froe and a running Oak instance.
The mutating writing API (writer) covers commits, checkpoints,
compaction and the reclamation it performs, backup, restore, and journal
recovery. It takes the exclusive repository lock first, so it cannot race a
cooperating running instance, and produces stores byte-for-byte compatible
with Oak (apart from the documented extreme-subnormal rendering residue;
see content::property::double_to_text). Planning a compaction is the
read-only exception and never takes the lock. Run mutations only against a stopped
repository. The writer requires a Unix operating-system entropy source and
therefore refuses to open on Windows.
The writing API is verified against a real Oak instance: the
workspace interoperability suite round-trips it through Apache
Jackrabbit Oak oak-segment-tar 1.90.0 — Oak writes the store, froe
commits, checkpoints, compacts, cleans up, backs up, restores and
recovers the journal, and Oak then boots against each result and serves
a byte-identical content tree without logging any of its own repair
messages. Still unverified against a live instance: store.version=1
stores, external blob stores, native macOS or Windows execution, and
Adobe AEM itself, which ships its own Oak build. Writing still requires
a stopped repository, and keeping a copy before a destructive operation
on irreplaceable data remains ordinary prudence.
§Example
use froe::store::Repository;
fn main() -> froe::Result<()> {
let repository = Repository::open(std::path::Path::new("/path/to/segmentstore"))?;
if let Some(node) = repository.node_at_path("/content")? {
for property in node.properties()? {
println!("{} = {:?}", property.name, property.values);
}
for (name, child) in node.child_node_entries()? {
println!("{name}: {} children", child.child_node_count()?);
}
}
Ok(())
}§Layers
Each layer is usable on its own:
tar_archive— archives, indexes, segment graphs, binary reference catalogs;segment— segment parsing and record addressing;content— decoding records into nodes, properties, and values;journal— the head revision log;gc_journal— optional garbage-collection history;store— the assembled read-only repository;index— Oak’s query-index definitions, lanes and storage.
Long-running operations — opening a large store, planning a compaction,
compacting, checking consistency — have a _with_progress twin that
reports what they are doing to a progress::ProgressObserver, so a
caller need not guess whether a silent minute means work or a hang.
Custom backends (in-memory fixtures, remote stores) implement
SegmentProvider and reuse the whole content layer unchanged.
Re-exports§
pub use content::BinaryStream;pub use content::BinaryValue;pub use content::NodeState;pub use content::PropertyState;pub use content::PropertyType;pub use content::PropertyValue;pub use content::PropertyValues;pub use content::SegmentProvider;pub use content::Template;pub use content::read_binary_stream;pub use error::Error;pub use error::Result;pub use gc_journal::GarbageCollectionJournalEntry;pub use index::lucene::writer::DocValue;pub use index::lucene::writer::Document;pub use index::lucene::writer::Field as LuceneField;pub use index::lucene::writer::IndexWriterStatistics;pub use index::lucene::writer::LuceneIndexWriter;pub use index::lucene::writer::MAXIMUM_TERM_LENGTH;pub use index::lucene::writer::StoredValue;pub use index::lucene::writer::Token;pub use index::lucene::writer::WrittenIndex;pub use index::lucene::writer::index_writer;pub use journal::JournalEntry;pub use progress::DiscardedProgress;pub use progress::ProgressObserver;pub use progress::Step;pub use progress::WorkUnit;pub use segment::GarbageCollectionGeneration;pub use segment::RecordIdentifier;pub use segment::RecordType;pub use segment::SegmentIdentifier;pub use segment::SegmentKind;pub use store::Repository;pub use units::format_byte_size;pub use units::format_count;pub use writer::ArchiveIndexSurvey;pub use writer::ArchiveRewritePolicy;pub use writer::CompactedGeneration;pub use writer::CompactionAction;pub use writer::CompactionKind;pub use writer::CompactionOptions;pub use writer::CompactionOutcome;pub use writer::CompactionPlan;pub use writer::DefinitionReport;pub use writer::ExternalBinaryFootprint;pub use writer::FileDeletionFailure;pub use writer::ImportedIndex;pub use writer::JournalLineRemoval;pub use writer::JournalRemovalReason;pub use writer::LuceneImportOptions;pub use writer::LuceneImportOutcome;pub use writer::LuceneImportPlan;pub use writer::NoWorkReason;pub use writer::OrphanedVersionHistoryReport;pub use writer::PlannedImport;pub use writer::PreparedCompaction;pub use writer::PreparedLuceneImport;pub use writer::PreparedReindex;pub use writer::RecoveryBackupPolicy;pub use writer::RecoveryBackupSurvey;pub use writer::RecoveryOutcome;pub use writer::ReindexAction;pub use writer::ReindexOptions;pub use writer::ReindexOutcome;pub use writer::ReindexPlan;pub use writer::ReindexWarning;pub use writer::StaleArchiveReason;pub use writer::WorkDirectory;pub use writer::WritableRepository;pub use writer::backup;pub use writer::backup;pub use writer::backup_with_progress;pub use writer::compact;pub use writer::compact_with_progress;pub use writer::lucene_import;pub use writer::lucene_import_with_progress;pub use writer::plan_compaction;pub use writer::plan_compaction_with_progress;pub use writer::plan_lucene_import;pub use writer::plan_reindex;pub use writer::plan_reindex_with_progress;pub use writer::recover_journal;pub use writer::recover_journal_with_progress;pub use writer::reindex;pub use writer::reindex_with_progress;pub use writer::restore;pub use writer::restore_with_progress;pub use writer::survey_archive_indexes;pub use writer::survey_recovery_backups;
Modules§
- checksum
- CRC32 checksum as used by the segment-tar format.
- content
- The content layer: decoding records into the content tree.
- error
- Error and result types shared across the crate.
- gc_
journal - The garbage-collection journal (
gc.log). - hashing
- Hash functions that must match their Java counterparts bit for bit.
- index
- Oak’s query indexes: definitions, asynchronous lanes, status nodes and the structures each index type stores in the repository.
- journal
- The journal: the repository’s sequence of head states.
- progress
- Progress observation for long-running operations.
- segment
- Segments: the unit of storage below the tar archive layer.
- store
- The repository: opening a segment store directory read-only.
- tar_
archive - The tar archive layer: segment archives on disk.
- tooling
- Read-only diagnostic tooling: consistency checking, revision differencing, node history, node search, segment dumps, and archive path attribution.
- units
- Rendering byte counts for people to read.
- writer
- The write path: building segments, archives, and repository state.
Structs§
- RunLocation
- What a caller of the Lucene index writer needs from the external sort.
- Sort
Budget - What a caller of the Lucene index writer needs from the external sort.
- Sorted
Pass - What a caller of the Lucene index writer needs from the external sort.
- Sorted
Passes - What a caller of the Lucene index writer needs from the external sort.
Traits§
- Spill
Record - What a caller of the Lucene index writer needs from the external sort.