Expand description
The parts of a CSV’s dialect that Polars’ reader does not have, named as the Frictionless Table Dialect names them.
| Frictionless | Here | Done by |
|---|---|---|
commentChar | --comment | Polars’ comment_prefix, before the header and in the data |
headerRows, headerJoin | --header-rows, header_join | [head] reads those lines; Polars reads the rest without a header |
skipInitialSpace | --skip-initial-space | skip_initial_space, lazy expressions over the text columns |
Header names are trimmed whatever the dialect: shown_names.
Structs§
- NoHeader
- A file that ends before the header line a read needs.
blankwhen all it holds is white space, or nothing.
Constants§
- DEFAULT_
HEADER_ JOIN - What
header_joinis when nothing sets it: Frictionless’headerJoindefault.
Functions§
- check_
comment_ char - Why
ccannot mark comment lines, if it cannot: it must be something, and on one line.--comment, its config key and the Python option share it. - header_
fields - Header line
row’s fields, trimmed: without a byte-order mark on line 1, its line break, orcomment’s prefix, split onseparator. - is_
blank_ file - Whether
esays the file holds nothing but white space where its header should be. - name_
columns lfwith its columns named asshown_namessays. Nothing is read: Polars has the schema from the scan’s own inference.- named_
lines - The lines
rowsnames (1-based, from the top of the file), in the orderrowsgives them, each with its line break. Only those lines are held, each up to a bound; a file that ends before the last of them is an error: it has no header there. - names_
of - The names of the columns, from the lines
rowsnames (1-based, counted from the top of the file before anything is skipped), each split onseparatorand trimmed;linesare the linesreadnames asnamed_linesread them, and one not read is blank. - read_
after_ header - An eager read of the lines after
header’s, with nothing there read as no columns, whichname_columnsmakes the header’s columns with no rows. - shown_
names - The names a read shows for the columns Polars named
raw: the header lines’ names by position when--header-rowsgave them, else Polars’ own, trimmed. A blank name iscolumn_N, as Polars names a headerless file’s; a name already taken gets_duplicated_K, as Polars marks a repeated header. - skip_
initial_ space skipInitialSpace: the spaces after a delimiter are not part of a text value, so" 152.6"is"152.6"and a cell of spaces is empty, which reads as null like an empty cell.nullsare the null values this column takes, matched after the padding is gone, as they would be against the unpadded file.- skip_
lines - Pass over
nlines ofsource. - window
- The first
rowsdata lines ofsource, read on from where its header lines ended, each split onseparatorand trimmed: the lines a scan infers its types from. A line that starts withcomment, or is blank, is not one. Stops atMAX_WINDOW_BYTES. - window_
of window, and whether a line it read holds bytes that are not UTF-8.