1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
//! Turning chunks into rows and rows back into chunks.
//!
//! Every pipeline breaker here holds its input as `Vec<Vec<Value>>`. That is the slowest reasonable
//! layout and it is chosen knowingly: a boxed value per field per row is what section 7.4 replaces
//! with a fixed width row prefix and a payload beside it, and the reason to write the slow one
//! first is that the fast one has to produce the same answer and there has to be something to
//! compare it against. Sorting, grouping and joining are the three places a database is most often
//! subtly wrong, and none of those bugs are about the row layout.
//!
//! It is also why the memory limit is charged here. Turning chunks into rows and rows back into
//! chunks is where an unbounded amount of memory is taken, so this is where the limit has to be
//! asked, and a limit charged in one place per crate is one somebody can believe rather than one
//! that has to be re-audited every time an operator is added. The reading half of it lives in
//! `gather` now, because that is the sink every operator that holds its input ends in, and it
//! charges through the helpers here rather than counting anything itself.
use ;
use ;
/// What one buffered row costs, counting the values and the vector holding them.
pub
/// What one buffered row owns away from itself, which is everything [`footprint`] counts except
/// the three words of the vector's own header.
///
/// Separate because a row sitting inside a container has those three words counted already, by
/// [`capacity`] over the container that holds it, and counting them again here would charge them
/// twice. Per #227 the containers are now charged for what they took rather than for what they are
/// using, and the two have to divide the row between them without overlapping.
///
/// A string is counted as the room it has and not as the bytes in it, for the same reason, so the
/// row this is asked about has to be the copy something kept rather than a buffer that is about to
/// be filled again. Grouping reads a row into one buffer and copies it out only when the group is
/// new, and that buffer keeps whatever the longest string it has held needed, so asking this about
/// it charges every row for the longest row.
pub
/// What one value owns away from itself, which is its footprint without its own bytes.
///
/// For a value going into room that has been charged already, where charging the footprint would
/// charge that room twice. A group by asks for the room its aggregate results need before it has
/// them, because by the time it has them it has taken the memory, and then it has only what each
/// result owns left to charge.
pub
/// [`ALLOCATION`] in the type the sizes around it are in.
const ALLOCATION_USIZE: usize = ALLOCATION as usize;
/// Charges whatever a set of containers has grown by since it was last charged.
///
/// `now` is what they take today and `charged` is what they were last charged, which this updates.
///
/// A container's capacity is what it took from the allocator and its length is what it is using,
/// and the difference is not small. A `Vec` doubles, so it is between half empty and full, and a
/// `HashMap` fills to seven eighths and then doubles as well. Charging the entries alone charges
/// the used part of a structure that paid for all of it, which is the larger half of #227.
///
/// Asked once per input chunk, the same granularity everything else in here charges at, so the
/// overshoot before a query is told is bounded by one chunk of insertions.
///
/// # Errors
///
/// [`rudb_common::ErrorCode::OutOfMemory`] when the growth passes the limit.
pub
/// How many buckets a hash table has to have to hold `entries` without growing again.
///
/// Both `HashMap` and `HashSet` are open addressed and refuse to fill past seven eighths, and they
/// report the seven rather than the eight, so the table on the heap is a bucket and a control byte
/// for each of eight sevenths of what `capacity` says.
pub
/// Rows back into chunks of at most [`VECTOR_SIZE`], in the order given.
///
/// An operator with no columns keeps its row count, which is the `SELECT count(*)` case and the
/// reason this takes the types rather than deriving the width from the first row.
///
/// What the chunks take is charged as they are built, because this is the second copy of the data:
/// the rows that went in are still alive while it runs, and an operator that is about to hand out
/// the chunks is holding both.
///
/// # Errors
///
/// If a value does not belong in the column it was placed in, if a row is not as wide as the type
/// list, or if the chunks pass the limit the database was opened with.
pub