rd-rds 0.5.0

Read-only reader for the subset of R's RDS serialization format used by installed-package help databases
Documentation

rd-rds

rd-rds is a scoped, read-only reader for installed-R-package information and selected CRAN-like repository indexes. It is not a general R serialization library; strict decoding rejects unknown SEXPs. Its compression support is implemented in pure Rust, so the crate links no C library. See the workspace README for repository status and crate relationships.

The API provides these entry points:

  • parse reads a decompressed XDR serialization stream only.
  • inspection::inspect observes a decompressed stream's root kind and closure formals without constructing an object or validating its complete payload.
  • file::from_bytes and file::read accept the complete envelope and apply bounded decompression. Supported envelopes are raw X\n XDR, gzip, xz, bzip2, and zstd (when the corresponding feature is enabled).
  • package provides validated convenience views for Meta/package.rds and CRAN-like PACKAGES.rds matrices.
  • matrix::CharacterMatrix provides a validated, owned view of general R character matrices, including matrices without dimnames.

With the opt-in lazyload feature, [lazyload::LazyLoadIndex] reads an .rdx index without requiring its companion .rdb data file. Its open and open_with_options constructors expose stored variables, persistence-reference descriptors, and record compression. Variables and references retain their index order and duplicate names; name lookups return the last matching entry. For example, callers can enumerate code or lazy-data names with LazyLoadIndex::open(path)?.variables() without loading any records.

Options::max_index_bytes bounds both stored and decompressed index bytes, each to 256 MiB by default. The index reader also applies the default RDS decoder limits. Record-size options apply only to record reads. Index parsing validates descriptors; it does not check their ranges against a data file. The returned Compression describes records, independently of the index file's compression envelope.

[lazyload::LazyLoadDb] uses the same index reader and opens the .rdx/.rdb pair to read direct records without an R session. The record layer recognizes uncompressed and zlib records. bzip2 and xz records are recognized in the index, but reading them reports an explicit unsupported error. Compound persistence references are described by their eager and lazy record locations but are not resolved. Unrecognized persistence descriptors are retained as RecordReference::Unsupported. Record reads return both the exact addressed bytes and the decompressed payload, with independent 256 MiB default bounds. Raw records are already XDR bytes and do not carry a length prefix; zlib records carry a four-byte declared length.

With the same feature, [package::InstalledCodeDb] provides the package-level API for that pair. The caller supplies the installed package directory; the reader selects R/<basename>.rdx and .rdb and does not discover libraries or scan runtime exports. [InstalledCodeDb::stored_bindings] is the complete variables map in index order, including duplicate names. Consequently this view's completeness domain is CodeDatabaseVariables, not exports or runtime bindings. [InstalledCodeDb::inspect_stored_binding] returns structured unknown/ambiguous errors instead of choosing among duplicate names, reads only a unique direct record, and returns bounded closure prefix metadata. Closure bodies are reported as BodyValidation::NotValidated. The selected record is fully loaded, decompressed, and checked by the LazyLoadDb container layer first; stored/decompressed size limits, framing, compression, and trailing-stream corruption therefore fail before prefix inspection. InstalledCodeOptions::max_bytes_visited applies only to the semantic prefix walker, which stops before reading or validating a closure's decompressed body payload. The provenance includes both selected paths, compression, and an opaque CodeDbGeneration derived from best-effort file metadata. It is an identity hint rather than a content hash or a transaction guarantee.

The lazyload feature includes the gzip feature because installed package .rdx indexes use the standalone gzip envelope in the normal package profile. Callers that enable lazyload therefore also get gzip .rds handling; the other standalone codecs remain independently selectable. The low-level lazyload::decode_stored_record helper applies the same bounded record decoder to an already isolated (offset, length) byte slice.

Record reads take a metadata snapshot when the database opens and compare it before and after each read. On Unix this includes device and inode, and on other platforms it uses file length and modification time when available. This is best-effort detection of replacement or concurrent modification and cannot provide a transaction guarantee against races after the final check. Neither the generation nor replacement detection is a content hash or a portable guarantee: a replacement that preserves all observed metadata may not be distinguishable on every platform.

The rd-helpdb crate uses the file layer for standalone help-database RDS files, and rd-ast can lower supported decoded documentation objects into the common document model.

Namespace metadata

[package::NamespaceMetadata] provides an owned view of declarations from a decoded Meta/nsInfo.rds object. Its exports, imports, S3 registrations, and S4 declarations are static metadata: they are not runtime namespace exports, stored lazy-load bindings, evaluated export patterns, or .onLoad results. Each known field is independently represented by [package::MetadataField], so a malformed S3 schema does not hide valid declared exports. Declared exports are returned as [package::NamespaceExport] values: the source binding and namespace-facing name are preserved separately, so an assignment-shaped export such as export(public = internal) is not confused with an ordinary export(name). Empty export-name attributes use the source name. The installed-package fixture also exercises the R-written list("utils", except = ...) import shape and an aliased importFrom.

A consumer such as a mini-roxygen provider can retain its existing policy boundary while replacing ad-hoc S3 extraction with positive evidence:

use std::collections::BTreeSet;
use rd_rds::package::{MetadataField, NamespaceMetadata};

fn generic_evidence(object: &rd_rds::RObject) -> BTreeSet<String> {
    match NamespaceMetadata::from_object(object)
        .ok()
        .map(|metadata| metadata.s3_generic_evidence().clone())
    {
        Some(MetadataField::Present(generics)) => generics.into_iter().collect(),
        Some(MetadataField::Missing)
        | Some(MetadataField::Invalid(_))
        | Some(MetadataField::UnsupportedSchema { .. })
        | None => BTreeSet::new(),
        _ => BTreeSet::new(),
    }
}

The caller still decides how missing metadata, invalid schemas, base-generic catalogs, shadowing, and library precedence should affect its provider.

Installed metadata paths

Consumers that already have an installed package directory can use [package::NamespaceMetadata::read_installed] and [package::PackageMeta::read_installed] to read the canonical Meta/nsInfo.rds and Meta/package.rds artifacts. These helpers only join the canonical path, perform the bounded file read, and construct the typed view; they do not discover packages, resolve exports, or model runtime state. read_installed_with_options is available on both views when a consumer needs explicit [file::ReadOptions] bounds. [package::InstalledMetadataError] keeps file-layer failures distinct from typed-view validation failures and retains the selected artifact path in either case.

PackageMeta::built().r_version() and PackageMeta::description_field("Priority") expose the validated Built.R and DESCRIPTION metadata used by consumers such as base-package catalogs. An absent DESCRIPTION field is not the same as a present NA field, while an absent Built element is represented by PackageMeta::built() == None.

Closure prefix inspection

inspection::inspect and inspection::inspect_with_options expose the bounded inspector without requiring optional features. They accept decompressed XDR bytes starting with the serialization header. A consumer that already loaded a lazy-load record can inspect and decode the same bytes without reopening the database or decompressing twice:

let record = database.read(binding_name)?;
let observed = rd_rds::inspection::inspect(record.decompressed_bytes())?;
// A consumer can retain its decoder for expressions the inspector omits.
let object = consumer_decode(record.decompressed_bytes())?;

The caller retains path selection and duplicate-name policy. Inspection does not look up stored bindings, evaluate promises, or resolve persistent references. InstalledCodeDb::inspect_stored_binding delegates to the same inspector after its own lookup and container checks. Existing inspection types at package::* remain available with lazyload and are the same types as inspection::*.

For a closure, inspection walks attributes, environment fields, and the formal/default pairlist while sharing strict decoding's reference registration/resolution, encoding, depth, and element accounting. It also applies inspection-specific byte and formal-count limits through InspectionOptions: 256 MiB and one million formals by default. These bounds do not limit the caller's file reads or decompression. It reports formal names in wire order (including ..., duplicate names, and non-syntactic UTF-8 names), distinguishes missing defaults from present defaults including NULL, and stops immediately after observing the body tag. The body payload is intentionally not validated. Prefix failures retain their phase and byte offset in an unavailable result after the root kind is known; failures before the root flags remain top-level errors. Non-closures return InspectionExtent::RootTagOnly, including unknown root kinds and records whose payloads are missing or invalid. The root kind describes the stored wire tag; it does not necessarily describe a materialized value, establish runtime callability, or prove that the object can be decoded. In particular, ALTREP values still require decoding to determine the materialized value's kind, and the shared decoder currently rejects those values as unsupported. Unvisited body payloads and trailing bytes remain unchecked. Use parse when full decoding is required for a supported object. Deterministic plain and compiler-produced format-2/format-3 fixtures, along with generated diagnostic, ALTREP, S4, namespace, and persisted-reference cases, are produced in the source repository by tests/fixtures/generate_closure_inspection_fixture.R.

Runnable examples

cargo run -p rd-rds --example inspect_packages -- /path/to/PACKAGES.rds
cargo run -p rd-rds --example inspect_rds -- /path/to/archive.rds
cargo run -p rd-rds --features lazyload --example inspect_installed_package -- /path/to/installed/package [binding]

inspect_packages demonstrates the typed, stable package-index view. inspect_rds provides a bounded advanced inspection of unfamiliar decoded objects, including shapes that are not package matrices. inspect_installed_package reads Meta/nsInfo.rds with file::read, reports declared S3 generic evidence, and lists or inspects stored code-database bindings from the explicit package directory. Its output keeps three domains separate: declared namespace metadata, code-database variables, and runtime namespace state. It never starts or inspects an R runtime. Consumers should derive a closure-formals signature only from FormalsInspection::Available; NotApplicable (including callable built-ins and specials) and Unavailable remain distinct states. A missing index is reported as NoCodeDatabase, while an existing index with no variables is a valid empty database. An unknown binding is reported as UnknownStoredBinding, and a duplicate name as AmbiguousStoredBinding; consumers should keep these database states distinct rather than treating them as an empty or last-wins lookup. The stored-code view is not a runtime namespace: a stored binding is evidence about the installed code database only. DefaultPresence::Absent means that the serialized closure has no default expression; it does not by itself prove that the argument is required (an R function may use missing()). Available formals likewise do not imply that the closure body was validated; BodyValidation::NotValidated remains explicit.

Repository-index interoperability

The supported contract is the tested decoding behaviour described in the workspace stability policy, not the continued availability or unchanged schema of files hosted by third parties. Deterministic fixtures cover these CRAN profiles:

  • src/contrib/PACKAGES.rds: xz envelope, serialization format 2, and the 17-column main-index schema.
  • src/contrib/Archive/<package>/PACKAGES.rds: gzip envelope, serialization format 3, and the 15-column package-archive schema.
  • src/contrib/Meta/archive.rds: gzip envelope, serialization format 3, and a named list of file.info()-shaped data frames.

Real CRAN examples were compared cell-for-cell with R 4.6.1 readRDS() on 2026-08-04; decoded cell values matched in all three profiles. R-universe was manually verified on 2026-08-05: source indexes used gzip, the Windows and macOS binary-repository indexes used zstd, and the observed schema had 15 columns with SHA256 in place of CRAN's MD5sum. These observations fall within the reader's general matrix, encoding, and compression behaviour, but the test suite contains no R-universe-specific fixture.

These statements describe observed interoperability at the stated dates. They do not guarantee that an external service retains the same paths, schemas, compression, or serialization behaviour.

Upstream archive semantics

These are upstream semantics, not reader guarantees. As observed on 2026-08-04, CRAN's per-package Archive/<package>/PACKAGES.rds excludes the current package version, and its rows are in archival rather than semantic-version order. Consumers must not infer inclusion of the current release or version precedence from row position. This is a recently introduced and undocumented CRAN facility and may change or disappear independently of rd-rds.

String encoding metadata

RStr::encoding() reports the CHARSXP encoding flag stored in the serialized data. R-universe files are generated by a JavaScript serializer rather than by R, and currently flag every string as UTF-8, including ASCII-only strings, while R's Encoding() reports those strings as "unknown" after readRDS(). Encoding labels can therefore differ even when decoded string contents are identical; this is not a decoding incompatibility.

R serialization format 2 has no native-encoding field in its header. A non-ASCII CHARSXP marked Native is therefore ambiguous. The reader preserves the bytes lazily for retained RStr values rather than guessing a locale: with the default policy, conversion by RStr::as_str() or a typed view rejects the value. A SYMSXP print name is converted during parsing instead, so a symbol name that cannot be decoded fails with Error::InvalidSymbolName at parse time under either policy. A caller with an independent UTF-8 contract may opt in with ReadOptions::native_encoding_policy(NativeEncodingPolicy::AssumeUtf8) (or the corresponding parse_with_options API with ParseOptions). The opt-in still validates retained RStr bytes with str::from_utf8 when conversion occurs and never performs lossy replacement. The policy applies only when the header field is absent, which means format 2; a format-3 header value is always authoritative. RStrValue::native_encoding_source() distinguishes a header-declared encoding from a caller assumption, while RStrValue::header_native_encoding() reports header evidence only.

Unknown or unsupported SEXP values reachable from the decoded result are hard decode errors; they are never silently converted to a known value. The one exception is environment internals: environments are collapsed to opaque handles, and a limited set of verified value shapes inside them (for example complex, raw, and S4 objects) is wire-consumed and discarded rather than rejected. Decoder defaults are a depth limit of 5,000, a vector limit of 8,000,000 elements, and a total-element limit of 16,000,000, plus a reference-table limit of 16,000,000 entries. The reference-table cap can be tightened independently with Limits::max_references; it is checked before a decoded reference is registered. The file layer defaults to 256 MiB compressed and decompressed input caps.

RObject and RValue access is a supported advanced API. Their fields are encapsulated and accessed through constructors and accessors. Enum variants may be added in minor releases, so consumers must use wildcard match arms; the public enums are non-exhaustive. The typed package and matrix views are the stable convenience surface for ordinary consumers.

Installed-code corpus observation

The scheduled Installed Code Corpus workflow runs an R-side oracle and the Rust scanner as separate processes against base and stats from setup-r plus the latest versions available in the fixed 2026-09-01 P3M snapshot for Matrix, dplyr, rlang, and R6. The manifest records the expected package version and the provenance records the observed version, build, R version, platform, locale, and installed path. It is intentionally limited to scheduled and manually dispatched runs; it is not a pull-request gate. The R 4.6.1 profile has a measured baseline containing stored-entry, root-kind, formals-availability, and oracle comparison counts. The R 4.5.3 profile has a measured compatibility baseline with the same aggregate counts. Both fixed-release profiles are blocking; R-devel is observational.

An eligible oracle comparison requires a same-named, non-active runtime closure. Stored and declared names, runtime-only names, active bindings, and runtime kind differences are reported as independent diagnostic buckets. A bounded Rust inspection that cannot obtain formals remains an explicit unavailable classification; it is not treated as a guessed signature.

Stability

Typed package-metadata views are the recommended supported surface. The RObject/RValue object model is supported as an advanced surface, with variants subject to addition; unsupported SEXPs are hard errors except for selected environment internals consumed as opaque or discarded wire data. See the workspace stability policy.

License

MIT; see the workspace license.