Skip to main content

Module csv_dialect

Module csv_dialect 

Source
Expand description

The parts of a CSV’s dialect that Polars’ reader does not have, named as the Frictionless Table Dialect names them.

FrictionlessHereDone by
commentChar--commentPolars’ comment_prefix, before the header and in the data
headerRows, headerJoin--header-rows, header_join[head] reads those lines; Polars reads the rest without a header
skipInitialSpace--skip-initial-spaceskip_initial_space, lazy expressions over the text columns

Header names are trimmed whatever the dialect: shown_names.

Structs§

NoHeader
A file that ends before the header line a read needs. blank when all it holds is white space, or nothing.

Constants§

DEFAULT_HEADER_JOIN
What header_join is when nothing sets it: Frictionless’ headerJoin default.

Functions§

check_comment_char
Why c cannot mark comment lines, if it cannot: it must be something, and on one line. --comment, its config key and the Python option share it.
header_fields
Header line row’s fields, trimmed: without a byte-order mark on line 1, its line break, or comment’s prefix, split on separator.
is_blank_file
Whether e says the file holds nothing but white space where its header should be.
name_columns
lf with its columns named as shown_names says. Nothing is read: Polars has the schema from the scan’s own inference.
named_lines
The lines rows names (1-based, from the top of the file), in the order rows gives them, each with its line break. Only those lines are held, each up to a bound; a file that ends before the last of them is an error: it has no header there.
names_of
The names of the columns, from the lines rows names (1-based, counted from the top of the file before anything is skipped), each split on separator and trimmed; lines are the lines read names as named_lines read them, and one not read is blank.
read_after_header
An eager read of the lines after header’s, with nothing there read as no columns, which name_columns makes the header’s columns with no rows.
shown_names
The names a read shows for the columns Polars named raw: the header lines’ names by position when --header-rows gave them, else Polars’ own, trimmed. A blank name is column_N, as Polars names a headerless file’s; a name already taken gets _duplicated_K, as Polars marks a repeated header.
skip_initial_space
skipInitialSpace: the spaces after a delimiter are not part of a text value, so " 152.6" is "152.6" and a cell of spaces is empty, which reads as null like an empty cell. nulls are the null values this column takes, matched after the padding is gone, as they would be against the unpadded file.
skip_lines
Pass over n lines of source.
window
The first rows data lines of source, read on from where its header lines ended, each split on separator and trimmed: the lines a scan infers its types from. A line that starts with comment, or is blank, is not one. Stops at MAX_WINDOW_BYTES.
window_of
window, and whether a line it read holds bytes that are not UTF-8.