Expand description
CSV, TSV and other delimited text for deser.
use deser::{Deserialize, Serialize};
#[derive(Debug, Deserialize, Serialize)]
struct City {
name: String,
country: String,
population: Option<u64>,
}
let input =
"name,country,population\nVienna,Austria,1897000\nAtlantis,,\n";
let cities: Vec<City> = deser_csv::from_str(input).unwrap();
assert_eq!(cities[0].population, Some(1897000));
assert_eq!(cities[1].population, None);
assert_eq!(deser_csv::to_string(&cities).unwrap(), input);§Data Model
A document is a sequence of records. By default the first record holds
the names of the columns (see Headers) and the other records are
maps of the names to their fields, without names records are sequences
(for instance tuples):
| input | deser |
|---|---|
a,b\n1,2\n3,4\n | [{"a": "1", "b": "2"}, {"a": "3", "b": "4"}] |
1,2\n3,4\n (no names) | [["1", "2"], ["3", "4"]] |
Everything in a CSV file is text, only the type a field is deserialized
into knows what it means. Fields are therefore passed on as lexical
atoms which are parsed by the types they are
delivered to: numbers parse them, strings take them as they are. They
are interpreted with the lenient
rules: yes, on and 1 are
booleans too and empty fields are None for optional numbers (other
LexicalRules can be given in the
Context). This also works when values are buffered, so flattened structs and
internally tagged and untagged enums work:
use deser::Deserialize;
#[derive(Debug, Deserialize, PartialEq)]
#[deser(tag = "kind", rename_all = "lowercase")]
enum Shape {
Circle { radius: f64 },
Rect { width: f64, height: f64 },
}
#[derive(Debug, Deserialize, PartialEq)]
struct Row {
id: u32,
#[deser(flatten)]
shape: Shape,
}
let input = "id,kind,radius,width,height\n1,circle,2,,\n2,rect,,3,4\n";
let config =
deser_csv::DeserializerConfig::builder().nulls(deser_csv::Nulls::Empty).build();
let rows: Vec<Row> = config.from_str(input).unwrap();
assert_eq!(rows[1].shape, Shape::Rect { width: 3.0, height: 4.0 });Empty fields are None for optionals if the type does not accept them
(an empty field is None for an Option<u32> and Some("") for an
Option<String>). Which fields are null can be configured (see
Nulls). Fields cannot hold maps or sequences, but a list in a
field can be read with the Separated
adapter (#[deser(as = Separated<';'>)] reads a;b;c).
Serializing works the other way around (see SerializerConfig): the
value is a sequence of records which are maps (the keys of the first
record are the names of the columns) or sequences.
§Dialects
There is no single CSV format, the configurations can be adjusted to what the other side writes and reads:
- The delimiter (
,,;,\t,|, the ASCII unit separator, …), the quote character (or no quotes at all) and if quotes are doubled or escaped (seeEscape). - The line ending (see
Terminator), by default\n,\r\nand\rall end records and\nis written. - Comment lines, blank lines, whitespace around fields (see
Trim), records with a different number of fields (flexible) and fields with quotes that do not follow the rules (lenient_quotes). - Tab separated values as written by databases (with backslash escapes
and
\Nfor null), seeDeserializerConfig::tsv.
Input is UTF-8, a byte order mark at the start is skipped (and UTF-16
input is reported as such). Fields which are not UTF-8 are passed on as
bytes. The sep=; line Excel writes can be enabled with
DeserializerConfig::set_sep_line.
§Errors
Errors point at the position in the input and (with deser-path) at
the path of the value, for instance [3].price. By default quotes
that do not follow the rules and records with the wrong number of fields
are errors.
§Streams
A stream of records is read with a reader of deser::io which the
configuration creates (DeserializerConfig::reader): every read
returns the next record. Errors of a record (like a field that does
not fit the type) only discard the record. The names of the columns
are known to the stream deserializer (see
StreamDeserializer::headers).
use deser_csv::DeserializerConfig;
#[derive(deser::Deserialize)]
struct Row {
name: String,
age: u32,
}
let input = &b"name,age\njane,42\njohn,x\nmax,7\n"[..];
let mut reader = DeserializerConfig::new().reader(input);
let mut ages = Vec::new();
let mut errors = Vec::new();
loop {
match reader.read::<Row>() {
Ok(Some(row)) => ages.push(row.age),
Ok(None) => break,
Err(err) => errors.push(err.line()),
}
}
assert_eq!(ages, [42, 7]);
assert_eq!(errors, [Some(3)]);
assert_eq!(reader.deserializer().headers().unwrap(), ["name", "age"]);Likewise every value written with a writer created by
SerializerConfig::writer is a record, the names are written before
the first one. The functions that read and write a single value
(from_reader and to_writer) read and write all records as a
sequence, like from_str and to_string.
The stream serializer (Serializer) and the stream deserializer
(StreamDeserializer) do not do IO themselves (see
deser::stream), they also work with other kinds
of IO and without the standard library.
§Features
Structs§
- Deserializer
- Deserializes delimited text.
- Deserializer
Config - Configures how delimited text is deserialized.
- Deserializer
Config Builder - Builds a
DeserializerConfig. - Records
- An iterator over the records of a
Deserializer. - Serializer
- Serializes records into delimited text.
- Serializer
Config - Configures how values are serialized into delimited text.
- Serializer
Config Builder - Builds a
SerializerConfig. - Stream
Deserializer - Reads the records of a stream (see
deser::stream).
Enums§
- Escape
- How characters are escaped.
- Headers
- Where the names of the columns come from.
- Nulls
- Which fields are null.
- Quote
Style - When fields are quoted.
- Terminator
- What ends records.
- Trim
- Which whitespace is removed.
Functions§
- from_
reader - Deserializes the records of a reader.
- from_
slice - Deserializes the records of a byte slice.
- from_
str - Deserializes the records of a string.
- to_
string - Serializes the records of a value to delimited text.
- to_
writer - Serializes the records of a value to a writer.