pub struct TableSchema { /* private fields */ }Expand description
The overall schema for potentially partitioned data sources.
When reading partitioned data (such as Hive-style partitioning), a TableSchema
consists of up to three parts:
- File schema: The schema of the actual data files on disk
- Partition columns: Columns whose values are encoded in the directory structure, but not stored in the files themselves
- Virtual columns: Columns produced by the file reader (e.g. Parquet
row_number) that are not stored in the files
The full table schema is composed in that order: file columns, then
partition columns, then virtual columns. Consumers that need a different
output ordering should use a projection on top of
TableSchema::table_schema.
§Example: Partitioned Table
Consider a table with the following directory structure:
/data/date=2025-10-10/region=us-west/data.parquet
/data/date=2025-10-11/region=us-east/data.parquetIn this case:
- File schema: The schema of
data.parquetfiles (e.g.,[user_id, amount]) - Partition columns:
[date, region]extracted from the directory path - Table schema: The full schema combining both (e.g.,
[user_id, amount, date, region])
§When to Use
Use TableSchema when:
- Reading partitioned data sources (Parquet, CSV, etc. with Hive-style partitioning)
- You need to efficiently access different schema representations without reconstructing them
- You want to avoid repeatedly concatenating file and partition schemas
For non-partitioned data or when working with a single schema representation,
working directly with Arrow’s Schema or SchemaRef is simpler.
§Performance
This struct pre-computes and caches the full table schema, allowing cheap references to any representation without repeated allocations or reconstructions.
Implementations§
Source§impl TableSchema
impl TableSchema
Sourcepub fn builder(file_schema: SchemaRef) -> TableSchemaBuilder
pub fn builder(file_schema: SchemaRef) -> TableSchemaBuilder
Start building a TableSchema from its (required) file schema.
Partition columns are optional and added with
TableSchemaBuilder::with_table_partition_cols; the full table schema
is computed once by TableSchemaBuilder::build. This is the preferred
way to construct a TableSchema.
§Example
let file_schema = Arc::new(Schema::new(vec![
Field::new("user_id", DataType::Int64, false),
Field::new("amount", DataType::Float64, false),
]));
let table_schema = TableSchema::builder(file_schema)
.with_table_partition_cols(vec![
Arc::new(Field::new("date", DataType::Utf8, false)),
Arc::new(Field::new("region", DataType::Utf8, false)),
])
.build();
// Table schema will have 4 columns: user_id, amount, date, region
assert_eq!(table_schema.table_schema().fields().len(), 4);Sourcepub fn new(file_schema: SchemaRef, table_partition_cols: Vec<FieldRef>) -> Self
👎Deprecated since 55.0.0: use TableSchema::builder(file_schema).with_table_partition_cols(cols).build() (or TableSchema::from(file_schema) for no partition columns)
pub fn new(file_schema: SchemaRef, table_partition_cols: Vec<FieldRef>) -> Self
use TableSchema::builder(file_schema).with_table_partition_cols(cols).build() (or TableSchema::from(file_schema) for no partition columns)
Create a new TableSchema from a file schema and partition columns.
This is a convenience for
TableSchema::builder(file_schema).with_table_partition_cols(cols).build().
Sourcepub fn from_file_schema(file_schema: SchemaRef) -> Self
👎Deprecated since 55.0.0: use TableSchema::from(file_schema) / file_schema.into()
pub fn from_file_schema(file_schema: SchemaRef) -> Self
use TableSchema::from(file_schema) / file_schema.into()
Create a new TableSchema with no partition columns.
Sourcepub fn with_table_partition_cols(self, partition_cols: Vec<FieldRef>) -> Self
👎Deprecated since 55.0.0: use TableSchema::builder(file_schema).with_table_partition_cols(cols).build()
pub fn with_table_partition_cols(self, partition_cols: Vec<FieldRef>) -> Self
use TableSchema::builder(file_schema).with_table_partition_cols(cols).build()
Return a new TableSchema with partition_cols as its partition columns,
replacing any existing ones. Existing virtual columns are preserved.
Sourcepub fn file_schema(&self) -> &SchemaRef
pub fn file_schema(&self) -> &SchemaRef
Get the file schema (without partition columns).
This is the schema of the actual data files on disk.
Sourcepub fn table_partition_cols(&self) -> &Fields
pub fn table_partition_cols(&self) -> &Fields
Get the table partition columns.
These are the columns derived from the directory structure that will be appended to each row during query execution.
Sourcepub fn virtual_columns(&self) -> &Fields
pub fn virtual_columns(&self) -> &Fields
Get the virtual columns.
Virtual columns are produced by the file reader (e.g. Parquet
row_number) and are not stored in the data files or derived from
partition paths.
Sourcepub fn table_schema(&self) -> &SchemaRef
pub fn table_schema(&self) -> &SchemaRef
Get the full table schema (file schema + partition columns + virtual columns).
This is the complete schema that will be seen by queries. Fields appear in the order: file columns, partition columns, virtual columns.
Sourcepub fn schema_without_virtual_columns(&self) -> &SchemaRef
pub fn schema_without_virtual_columns(&self) -> &SchemaRef
Schema of columns that can be referenced by predicates pushed into the file reader: file columns plus partition columns, excluding virtual columns.
Virtual columns are produced by the reader itself (e.g. Parquet
row_number) and cannot be referenced inside the reader’s row filter,
so predicates that reference them must stay above the scan. Callers
deciding which filters to push down should check against this schema
rather than Self::table_schema.
When there are no virtual columns this returns the same schema as
Self::table_schema.
Trait Implementations§
Source§impl Clone for TableSchema
impl Clone for TableSchema
Source§fn clone(&self) -> TableSchema
fn clone(&self) -> TableSchema
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read moreSource§impl Debug for TableSchema
impl Debug for TableSchema
Auto Trait Implementations§
impl Freeze for TableSchema
impl RefUnwindSafe for TableSchema
impl Send for TableSchema
impl Sync for TableSchema
impl Unpin for TableSchema
impl UnsafeUnpin for TableSchema
impl UnwindSafe for TableSchema
Blanket Implementations§
impl<T> Allocation for T
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
Source§impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> CloneToUninit for Twhere
T: Clone,
Source§impl<T> IntoEither for T
impl<T> IntoEither for T
Source§fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left is true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left(&self) returns true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read more