Skip to main content

RestStreamConfig

Struct RestStreamConfig 

Source
pub struct RestStreamConfig {
Show 41 fields pub base_url: String, pub path: String, pub method: Method, pub auth: AuthSpec<Auth>, pub headers: HashMap<String, String>, pub query_params: HashMap<String, String>, pub query_params_multi: HashMap<String, Vec<String>>, pub body: Option<Value>, pub pagination: PaginationStyle, pub records_path: Option<String>, pub max_pages: Option<usize>, pub request_delay: Option<Duration>, pub timeout: Option<Duration>, pub max_retries: u32, pub retry_backoff: Duration, pub tolerated_http_errors: Vec<u16>, pub replication_method: ReplicationMethod, pub replication_key: Option<String>, pub start_replication_value: Option<Value>, pub state_key: Option<String>, pub name: Option<String>, pub primary_keys: Vec<String>, pub schema: Option<Value>, pub schema_sample_size: usize, pub partitions: Vec<HashMap<String, Value>>, pub partition_concurrency: Option<usize>, pub tls: Option<TlsClientConfig>, pub response_format: ResponseFormat, pub csv_delimiter: u8, pub csv_has_headers: bool, pub excel_sheet: Option<String>, pub excel_header_row: usize, pub replication_bind: Option<ReplicationBind>, pub odata: Option<ODataConfig>, pub decode: Vec<DecodeStep>, pub async_job: Option<AsyncJobConfig>, pub window: Option<WindowSpec>, pub record_ancestors: Option<HashMap<String, String>>, pub records_multi: Vec<RecordsMultiSpec>, pub op_field: Option<String>, pub persist_cursor: bool,
}
Expand description

Configuration for a RestStream.

Fields§

§base_url: String§path: String

URL path, relative to base_url. May contain {key} placeholders that are substituted per-partition (e.g. "/orgs/{org_id}/users").

§method: Method§auth: AuthSpec<Auth>

Authentication: either inline ({ type, config }) or a { ref: <name> } pointer to a shared provider in the CLI’s top-level auth: catalog.

§headers: HashMap<String, String>

Static request headers sent on every request (data pages, async-job submit/poll/fetch requests, and OData $metadata discovery probes). Applied before the auth provider’s header placements, so an auth header of the same name always wins on a clash. Values honor ${env:} / ${param.*} load-time interpolation and pass through the secrets/redaction boundary like other config strings. Invalid header names/values are rejected at config load (FaucetError::Config), never a mid-run panic.

headers:
  Prefer: transient
  Accept: application/json
§query_params: HashMap<String, String>§query_params_multi: HashMap<String, Vec<String>>

Repeated / array-valued query params (#536), rendered as repeated keys — e.g. { "group_by[]": ["api_key_id", "model"] }?group_by[]=api_key_id&group_by[]=model. Applied alongside (in addition to) query_params; use this for APIs that need a key to appear more than once (group_by[], repeated expand/fields). Values honor {placeholder} context substitution for child sources, like query_params. Empty by default.

§body: Option<Value>§pagination: PaginationStyle§records_path: Option<String>§max_pages: Option<usize>§request_delay: Option<Duration>§timeout: Option<Duration>§max_retries: u32

Number of retries (after the first attempt) for transient request failures. Default 3.

Precedence note: the REST source predates the unified pipeline resilience: policy. When this field (or retry_backoff) is left at its default, an injected RetryPolicy (e.g. from a pipeline-level resilience: block, via RestStream::with_retry_policy) governs the retry budget. Setting this field away from its default makes it win — an explicit per-connector value is never silently overridden by a pipeline-wide default.

§retry_backoff: Duration

Base exponential-backoff delay between retries. Default 1s. Shares the legacy-field precedence rule documented on max_retries.

§tolerated_http_errors: Vec<u16>

HTTP status codes that should not cause an error. Responses with these codes are treated as empty pages (no records, no further pages).

§replication_method: ReplicationMethod§replication_key: Option<String>

Field name (not a JSONPath) used for incremental replication bookmarking.

§start_replication_value: Option<Value>

Bookmark value: records where record[replication_key] <= start_replication_value are filtered out when replication_method is Incremental.

§state_key: Option<String>

Opt-in identifier used by Pipeline::with_state_store to persist this stream’s bookmark across runs. When set, the pipeline will load any previously-stored bookmark before fetching and write the new bookmark only after the sink confirms the batch.

Keys must satisfy faucet_core::state::validate_state_key.

§name: Option<String>

Human-readable stream name (used in logging and Singer SCHEMA messages).

§primary_keys: Vec<String>

Field names that uniquely identify a record (Singer key_properties).

§schema: Option<Value>

JSON Schema describing the structure of each record.

§schema_sample_size: usize

Maximum number of records to sample when inferring the schema via crate::stream::RestStream::infer_schema. 0 means sample all available records (up to max_pages). Defaults to 100.

§partitions: Vec<HashMap<String, Value>>

Each entry is a context map whose values are substituted into path placeholders. The stream is executed once per partition and results are concatenated. Empty means run once with no substitution.

§partition_concurrency: Option<usize>

Maximum number of partitions to fetch concurrently. None means sequential processing (backward compatible default).

§tls: Option<TlsClientConfig>

Optional client-certificate (mutual TLS) config. When set, the source presents a client certificate on every request — data requests and any inline auth token request (both go through the same HTTP client). Requires the crate’s mtls feature; a tls block on a build without it is a load-time error rather than being silently ignored.

§response_format: ResponseFormat

How to parse the response body. json (default) uses JSONPath extraction (records_path); csv / excel parse a tabular file body into records — for authenticated file endpoints such as a Microsoft Graph / OneDrive / SharePoint …/content download or any signed export URL. In file mode a single response is fetched (pagination must be none) and records_path does not apply. excel requires the crate’s excel feature.

§csv_delimiter: u8

CSV field delimiter byte (default ,). Used only when response_format: csv.

§csv_has_headers: bool

Whether the first CSV row is a header row supplying field names (default true). When false, fields are named column_0, column_1, …

§excel_sheet: Option<String>

Excel worksheet to read: a sheet name, or a 0-based index as a string. When omitted, the first worksheet is used. response_format: excel only.

§excel_header_row: usize

0-based index of the Excel header row (default 0). Rows above it are skipped; the header row supplies field names. response_format: excel only.

§replication_bind: Option<ReplicationBind>

Bind the stored bookmark into the outgoing request (query param / header / body field / path) so the server returns only new rows. Composes with the existing replication_key client-side filter, which stays active as a safety net. Requires replication_method: incremental + replication_key.

§odata: Option<ODataConfig>

Speak the OData protocol: @odata.nextLink paging, the $.value envelope, $select/$filter/$expand/$orderby sugar, and $metadata (EDMX) → schema discovery. When set, it derives the pagination, records_path, query params, and Prefer header at load time (explicit values still win). See ODataConfig.

§decode: Vec<DecodeStep>

Decode the response body before record extraction: a chain of extract (JSONPath) / base64 / gunzip / unzip / parse (json|csv|xlsx|xml) steps. Lets a source consume base64/compressed/file payloads (e.g. a base64 XLSX inside a SOAP body, or a gzipped-CSV export). When set, it replaces the response_format body parsing, and pagination must be none. See DecodeStep.

§async_job: Option<AsyncJobConfig>

Run a submit→poll→fetch job lifecycle instead of a single GET, for bulk/export/report-run APIs (Salesforce Bulk, Stripe Reporting, …). The fetched result flows through decode: / response_format. When set, pagination must be none. See AsyncJobConfig.

§window: Option<WindowSpec>

Bound each request to a rolling [start, end) window between the stored bookmark and now, iterating the windows within one run (each step wide) with per-window bookmark durability. For APIs that require — or cap — a bounded date range (analytics/ads/reporting feeds). Parity with Airbyte’s DatetimeBasedCursor. Requires replication_method: incremental + replication_key, and a start bookmark (from state, or start_replication_value). See WindowSpec.

§record_ancestors: Option<HashMap<String, String>>

When records_path selects a nested array element (e.g. $.data[*].data.object), copy fields from the enclosing [*] array-element ancestor onto each emitted record. The map is dest_field: ancestor_relative_path — for each matched leaf, the source walks up to the array-element ancestor and copies the named path onto the record under dest_field. Absent ⇒ records are emitted unchanged.

records_path: "$.data[*].data.object"
record_ancestors: { event_id: "id", event_created: "created" }

Requires records_path to contain an array wildcard [*]; mutually exclusive with records_multi.

§records_multi: Vec<RecordsMultiSpec>

Emit several record arrays from one response in a single page (sharing one pagination advance), each stamped with a user-defined op marker under op_field. Composes with a downstream write_mode: upsert + delete_marker so added/modified/removed feeds route correctly. Mutually exclusive with records_path / record_ancestors, and requires response_format: json with no decode: pipeline.

records_multi:
  - { path: "$.added[*]",    op: upsert }
  - { path: "$.modified[*]", op: upsert }
  - { path: "$.removed[*]",  op: delete }
op_field: _op
§op_field: Option<String>

Field name each records_multi record is stamped with its spec’s op value. Defaults to _op when omitted.

§persist_cursor: bool

Persist the terminal pagination cursor as this run’s bookmark (riding the existing StreamPage.bookmark / StateStore path — no core trait change) and, on resume, seed the stored bookmark back into the first request (query param for cursor, request body field for cursor_in_body) before paging. Only meaningful with pagination: cursor / cursor_in_body; mutually exclusive with window slicing. Default false.

Implementations§

Source§

impl RestStreamConfig

Source

pub fn validate(&self) -> Result<(), FaucetError>

Validate cross-field invariants that serde alone can’t express.

File response formats (csv / excel) fetch a single response and parse the whole body, so paginated / JSONPath-extracted requests are rejected rather than silently ignored.

Source

pub fn apply_odata_defaults(&mut self)

Derive request defaults from the odata: block (paging, $.value envelope, $select/$filter/$expand/$orderby params, and the Prefer page-size header). Explicit config always wins — a field the user already set is never overwritten. Idempotent.

Source

pub fn new(base_url: &str, path: &str) -> Self

Source

pub fn method(self, m: Method) -> Self

Source

pub fn auth(self, a: Auth) -> Self

Source

pub fn header(self, k: &str, v: &str) -> Self

Add a static request header. Validation is deferred to RestStream::new (via validate), so an invalid name/value surfaces as a typed FaucetError::Config rather than panicking here.

Source

pub fn query(self, k: &str, v: &str) -> Self

Source

pub fn body(self, b: Value) -> Self

Source

pub fn tls(self, tls: TlsClientConfig) -> Self

Attach a mutual-TLS client identity (requires the mtls feature at build time; otherwise RestStream::new errors).

Source

pub fn pagination(self, p: PaginationStyle) -> Self

Source

pub fn records_path(self, p: &str) -> Self

Source

pub fn max_pages(self, n: usize) -> Self

Source

pub fn request_delay(self, d: Duration) -> Self

Source

pub fn timeout(self, d: Duration) -> Self

Source

pub fn max_retries(self, n: u32) -> Self

Source

pub fn retry_backoff(self, d: Duration) -> Self

Source

pub fn tolerate_http_error(self, status: u16) -> Self

HTTP status codes that should be silently ignored (treated as empty pages).

Source

pub fn replication_method(self, m: ReplicationMethod) -> Self

Source

pub fn replication_key(self, key: &str) -> Self

Field name (not JSONPath) used as the incremental replication bookmark.

Source

pub fn start_replication_value(self, v: Value) -> Self

Bookmark start value: records at or before this value are filtered out when using ReplicationMethod::Incremental.

Source

pub fn state_key(self, key: &str) -> Self

Opt the stream into resumable runs by giving it a stable state key. When this is set and the Pipeline is configured with a state store, the previously persisted bookmark is applied to the stream before fetching.

Source

pub fn replication_bind(self, bind: ReplicationBind) -> Self

Bind the stored bookmark into the outgoing request (#513).

Source

pub fn window(self, window: WindowSpec) -> Self

Slice the run into rolling [start, end) datetime windows (#527).

Source

pub fn odata(self, odata: ODataConfig) -> Self

Speak OData: derive paging, the $.value envelope, the query-option sugar, and $metadata discovery from the block (#512).

Source

pub fn decode(self, steps: Vec<DecodeStep>) -> Self

Set the response-decode pipeline (#515).

Source

pub fn name(self, n: &str) -> Self

Human-readable stream name.

Source

pub fn primary_keys(self, keys: Vec<String>) -> Self

Field names that uniquely identify a record (Singer key_properties).

Source

pub fn schema(self, s: Value) -> Self

JSON Schema for the stream’s records.

Source

pub fn schema_sample_size(self, n: usize) -> Self

Maximum records to sample for schema inference (0 = unlimited).

Source

pub fn add_partition(self, ctx: HashMap<String, Value>) -> Self

Add a partition context. The stream will execute once for each partition, substituting {key} placeholders in path with values from the context.

Source

pub fn add_query_param_multi(self, key: &str, values: Vec<String>) -> Self

Add a repeated / array-valued query parameter (#536): key is emitted once per value (?key=v0&key=v1). Chainable.

Source

pub fn partition_concurrency(self, concurrency: Option<usize>) -> Self

Set the maximum number of partitions to fetch concurrently. None (default) means sequential processing.

Source

pub fn record_ancestors(self, map: HashMap<String, String>) -> Self

Copy enclosing [*] ancestor fields onto each nested record (#549).

Source

pub fn records_multi(self, specs: Vec<RecordsMultiSpec>) -> Self

Emit several op-stamped record arrays from one response in one page (#548).

Source

pub fn op_field(self, field: &str) -> Self

Field name each records_multi record is stamped with its op value (default _op).

Source

pub fn persist_cursor(self, enabled: bool) -> Self

Persist the terminal pagination cursor as the run’s bookmark and seed it on resume (#547).

Trait Implementations§

Source§

impl Clone for RestStreamConfig

Source§

fn clone(&self) -> RestStreamConfig

Returns a duplicate of the value. Read more
1.0.0 (const: unstable) · Source§

fn clone_from(&mut self, source: &Self)

Performs copy-assignment from source. Read more
Source§

impl Debug for RestStreamConfig

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more
Source§

impl Default for RestStreamConfig

Source§

fn default() -> Self

Returns the “default value” for a type. Read more
Source§

impl<'de> Deserialize<'de> for RestStreamConfig

Source§

fn deserialize<__D>(__deserializer: __D) -> Result<Self, __D::Error>
where __D: Deserializer<'de>,

Deserialize this value from the given Serde deserializer. Read more
Source§

impl JsonSchema for RestStreamConfig

Source§

fn schema_name() -> Cow<'static, str>

The name of the generated JSON Schema. Read more
Source§

fn schema_id() -> Cow<'static, str>

Returns a string that uniquely identifies the schema produced by this type. Read more
Source§

fn json_schema(generator: &mut SchemaGenerator) -> Schema

Generates a JSON Schema for this type. Read more
Source§

fn inline_schema() -> bool

Whether JSON Schemas generated for this type should be included directly in parent schemas, rather than being re-used where possible using the $ref keyword. Read more
Source§

impl Serialize for RestStreamConfig

Source§

fn serialize<__S>(&self, __serializer: __S) -> Result<__S::Ok, __S::Error>
where __S: Serializer,

Serialize this value into the given Serde serializer. Read more

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> CloneToUninit for T
where T: Clone,

Source§

unsafe fn clone_to_uninit(&self, dest: *mut u8)

🔬This is a nightly-only experimental API. (clone_to_uninit)
Performs copy-assignment from self to dest. Read more
Source§

impl<T> DeserializeOwned for T
where T: for<'de> Deserialize<'de>,

Source§

impl<T> DynClone for T
where T: Clone,

Source§

fn __clone_box(&self, _: Private) -> *mut ()

Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T> Instrument for T

Source§

fn instrument(self, span: Span) -> Instrumented<Self>

Instruments this type with the provided Span, returning an Instrumented wrapper. Read more
Source§

fn in_current_span(self) -> Instrumented<Self>

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> PolicyExt for T
where T: ?Sized,

Source§

fn and<P, B, E>(self, other: P) -> And<T, P>
where T: Sized + Policy<B, E>, P: Policy<B, E>,

Create a new Policy that returns Action::Follow only if self and other return Action::Follow. Read more
Source§

fn or<P, B, E>(self, other: P) -> Or<T, P>
where T: Sized + Policy<B, E>, P: Policy<B, E>,

Create a new Policy that returns Action::Follow if either self or other returns Action::Follow. Read more
Source§

impl<T> Same for T

Source§

type Output = T

Should always be Self
Source§

impl<T> ToOwned for T
where T: Clone,

Source§

type Owned = T

The resulting type after obtaining ownership.
Source§

fn to_owned(&self) -> T

Creates owned data from borrowed data, usually by cloning. Read more
Source§

fn clone_into(&self, target: &mut T)

Uses borrowed data to replace owned data, usually by cloning. Read more
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, <T as TryFrom<U>>::Error>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<T> WithSubscriber for T

Source§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self>
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a WithDispatch wrapper. Read more
Source§

fn with_current_subscriber(self) -> WithDispatch<Self>

Attaches the current default Subscriber to this type, returning a WithDispatch wrapper. Read more