1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
//! Built-in dataset loaders.
//!
//! Every module here wraps one data source in a [`Dataset`](dataset_core::Dataset).
//! Each module is a worked example of the same four steps:
//!
//! 1. Download from a URL.
//! 2. Verify a SHA-256 hash.
//! 3. Parse the source: CSV records, raw documents from an archive, or binary IDX images.
//! 4. Return a [`Table`](crate::table::Table) of named, typed columns.
//!
//! The crate root also re-exports every loader struct, so
//! [`dataset_ml::Iris`](crate::Iris) and
//! [`dataset_ml::dataset::iris::Iris`](crate::dataset::iris::Iris) name the same
//! type. Use whichever path reads better.
//!
//! This module needs the `dataset` feature, which is on by default.
//!
//! # Datasets
//!
//! | Module | Samples | Features | Task Type |
//! |--------|---------|----------|-----------|
//! | [`abalone`](crate::dataset::abalone) | 4,177 | 8 | Regression |
//! | [`adult`](crate::dataset::adult) | 32,561 | 14 | Classification |
//! | [`bank_marketing`](crate::dataset::bank_marketing) | 45,211 | 16 | Classification |
//! | [`banknote_authentication`](crate::dataset::banknote_authentication) | 1,372 | 4 | Classification |
//! | [`bike_sharing_hourly`](crate::dataset::bike_sharing::bike_sharing_hourly) | 17,379 | 12 | Regression (multi-output) |
//! | [`bike_sharing_daily`](crate::dataset::bike_sharing::bike_sharing_daily) | 731 | 11 | Regression (multi-output) |
//! | [`iris`](crate::dataset::iris) | 150 | 4 | Classification |
//! | [`breast_cancer`](crate::dataset::breast_cancer) | 569 | 30 | Classification |
//! | [`boston_housing`](crate::dataset::boston_housing) | 506 | 13 | Regression |
//! | [`california_housing`](crate::dataset::california_housing) | 20,640 | 8 | Regression |
//! | [`car_evaluation`](crate::dataset::car_evaluation) | 1,728 | 6 | Classification |
//! | [`covtype`](crate::dataset::covtype) | 581,012 | 54 | Classification |
//! | [`diabetes`](crate::dataset::diabetes) | 442 | 10 | Regression |
//! | [`digits`](crate::dataset::digits) | 1,797 | 64 | Classification |
//! | [`fashion_mnist`](crate::dataset::fashion_mnist) | 60,000 / 10,000 / 70,000 | 784 (28×28 pixels) | Classification (10 classes) |
//! | [`heart_disease`](crate::dataset::heart_disease) | 303 | 13 | Classification |
//! | [`ionosphere`](crate::dataset::ionosphere) | 351 | 34 | Classification |
//! | [`kddcup99`](crate::dataset::kddcup99) | 494,021 / 4,898,431 | 41 | Classification |
//! | [`letter_recognition`](crate::dataset::letter_recognition) | 20,000 | 16 | Classification (26 classes) |
//! | [`linnerud`](crate::dataset::linnerud) | 20 | 3 | Regression (multi-output) |
//! | [`mnist`](crate::dataset::mnist) | 60,000 / 10,000 / 70,000 | 784 (28×28 pixels) | Classification (10 classes) |
//! | [`movielens_100k`](crate::dataset::movielens_100k) | 100,000 ratings | 943 users × 1,682 movies | Recommendation |
//! | [`mushroom`](crate::dataset::mushroom) | 8,124 | 22 | Classification |
//! | [`spambase`](crate::dataset::spambase) | 4,601 | 57 | Classification |
//! | [`titanic`](crate::dataset::titanic) | 891 | 11 | Classification |
//! | [`palmer_penguins`](crate::dataset::palmer_penguins) | 344 | 7 | Classification |
//! | [`sms_spam`](crate::dataset::sms_spam) | 5,574 | text | Classification |
//! | [`wholesale_customers`](crate::dataset::wholesale_customers) | 440 | 8 | Clustering (no target) |
//! | [`wine_recognition`](crate::dataset::wine_recognition) | 178 | 13 | Classification |
//! | [`red_wine_quality`](crate::dataset::wine_quality::red_wine_quality) | 1,599 | 11 | Regression |
//! | [`white_wine_quality`](crate::dataset::wine_quality::white_wine_quality) | 4,898 | 11 | Regression |
//! | [`youtube_spam`](crate::dataset::youtube_spam) | 1,956 | text | Classification |
//! | [`sentiment_sentences`](crate::dataset::sentiment_sentences) | 3,000 | text | Classification |
//! | [`newsgroups20`](crate::dataset::newsgroups20) | 11,314 / 18,846 | text | Classification |
//! | [`movie_review_polarity`](crate::dataset::movie_review_polarity) | 2,000 | text | Classification |
//!
//! Each module documents its own source and column layout.
/// Reader for the IDX binary format, shared by [`mnist`] and [`fashion_mnist`].
///
/// The two datasets ship the same four-file layout and the same 28×28 image
/// shape. This module is internal to the crate, unlike every other module here.
/// Abalone dataset module.
///
/// Contains the Abalone dataset (UCI, Nash et al. 1994) for **regression**. It
/// predicts an abalone's `rings` (age in years is `rings + 1.5`) from 8 mixed
/// features: 1 categorical `sex` feature and 7 numeric physical measurements.
/// `Abalone::FEATURE_NAMES` names the 8 inputs. Unlike the other mixed-type
/// loaders, which are classification tasks, its target column
/// `Abalone::TARGET` holds numeric values.
/// Adult / Census Income dataset module.
///
/// Contains the Adult dataset (also called "Census Income") for binary
/// classification. It predicts whether a person earns over $50K/year from 14
/// mixed features: 8 categorical and 6 numeric, covering demographic and
/// employment attributes. Extracted from the 1994 US Census. Uses the canonical
/// `adult.data` training partition.
/// Bank Marketing dataset module.
///
/// Contains the Bank Marketing dataset for binary classification. It predicts
/// whether a client subscribes to a term deposit from 16 mixed features: 9
/// categorical and 7 numeric, covering client, contact, and campaign attributes.
/// Recorded from a Portuguese bank's phone campaigns. Uses the full
/// `bank-full.csv` partition. Sourced from a ZIP archive (like `digits`).
/// Banknote Authentication dataset module.
///
/// Contains the Banknote Authentication dataset (UCI, Lohweg 2012) for binary
/// classification. It tells genuine banknote specimens from forged ones, using 4
/// continuous statistics (variance, skewness, curtosis, entropy) of
/// Wavelet-transformed banknote images. This is the crate's most compact
/// pure-numeric benchmark. Its target column `BanknoteAuthentication::TARGET`
/// holds the source's raw `0`/`1` code, because UCI does not document which code
/// means which.
/// Bike Sharing dataset module.
///
/// Contains the Bike Sharing dataset (UCI, Fanaee-T 2013) for **regression**:
/// predicting the rental count of the Capital Bikeshare system in Washington,
/// D.C., over 2011 and 2012 from the calendar attributes and the weather. Each
/// sample carries its calendar date in the `dteday` column, and the rows stay in
/// chronological order, so a split by time is possible. One ZIP archive holds
/// two aggregations of the same rental log, and each one has its own loader:
/// `bike_sharing_hourly::BikeSharingHourly` (17,379 records) and
/// `bike_sharing_daily::BikeSharingDaily` (731 records). Both use the
/// multi-output target `(casual, registered, cnt)`, which each loader names in
/// its own `TARGET_NAMES` constant.
/// Boston Housing dataset module.
///
/// Contains the Boston Housing dataset for predicting median house values
/// in Boston suburbs, based on features like crime rate, room count,
/// and accessibility to highways.
/// Breast Cancer Wisconsin (Diagnostic) dataset module.
///
/// Contains the Breast Cancer Wisconsin dataset for binary classification of
/// tumors as malignant or benign. It uses 30 features computed from digitized
/// images of cell nuclei.
/// California Housing dataset module.
///
/// Contains the California Housing dataset for predicting median house values
/// in California districts. Reproduces the eight derived features of
/// scikit-learn's `fetch_california_housing`.
/// Car Evaluation dataset module.
///
/// Contains the Car Evaluation dataset (UCI, Bohanec 1988) for multi-class
/// classification. It predicts a car's overall acceptability (`unacc`, `acc`,
/// `good`, `vgood`) from 6 categorical price and technical attributes. Like
/// [`mushroom`], it is **all-categorical**: every feature column is
/// `ColumnData::String`.
/// Forest Cover Type dataset module.
///
/// Contains the scikit-learn Forest CoverType dataset (`fetch_covtype`) for
/// multi-class classification: predicting one of seven forest cover types from 54
/// cartographic features of 30×30 meter cells. Its source is a gzip-compressed
/// file. The `Cover_Type` column holds the source's codes `1` to `7`, and the
/// `Covtype::CLASS_NAMES` constant names each one.
/// Diabetes dataset module.
///
/// Contains the scikit-learn diabetes dataset (`load_diabetes`) for regression:
/// predicting disease progression from 10 standardized physiological features.
/// Optical Recognition of Handwritten Digits dataset module.
///
/// Contains the scikit-learn digits dataset (`load_digits`) for multi-class
/// classification: recognizing handwritten digits (`0`–`9`) from 8×8 grayscale
/// images flattened into 64 integer pixel intensities.
/// Fashion-MNIST dataset module.
///
/// Contains the Fashion-MNIST dataset (Zalando, Xiao et al. 2017) for
/// multi-class classification: recognizing one of 10 garment classes from a
/// 28×28 grayscale image. It matches [`mnist`] in image size, class count,
/// partition sizes, and IDX file format, but its classes overlap more. They are
/// exactly balanced: 6,000 images per class for training and 1,000 for test.
/// Like `mnist`, it holds one `pixels` column of `ColumnData::Bytes`, plus a
/// `FashionMnist::CLASS_NAMES` constant that names each label code. It offers
/// `new`/`new_test`/`new_all` subset constructors.
/// Heart Disease (Cleveland) dataset module.
///
/// Contains the Cleveland Heart Disease dataset (UCI, Janosi et al. 1988) for
/// classification: predicting the presence of heart disease (`num`, `0`–`4`) from
/// 13 clinical features. The loader maps the `?` missing values in `ca`/`thal` to
/// `NaN` (like [`titanic`]/[`palmer_penguins`]). `HeartDisease::TARGET` names the
/// `num` column, which holds one integer code per patient.
/// Ionosphere dataset module.
///
/// Contains the Ionosphere dataset (UCI, Sigillito et al. 1989) for binary
/// classification. It predicts whether a radar return shows structure in the
/// ionosphere (`good`) or passes through it (`bad`), from 34 continuous
/// autocorrelation features. A compact pure-numeric benchmark like
/// [`breast_cancer`].
/// Iris flower dataset module.
///
/// Contains the classic Iris dataset for classifying iris flowers into
/// three species (setosa, versicolor, virginica) based on sepal and petal
/// measurements.
/// KDD Cup 1999 network-intrusion dataset module.
///
/// Contains the scikit-learn KDD Cup 1999 dataset (`fetch_kddcup99`) for
/// multi-class classification: detecting network intrusions from 41 mixed
/// (3 categorical + 38 numeric) connection features. `Kddcup99::new` loads the
/// default 10% subset (494,021 samples) and `Kddcup99::new_full` the full set
/// (4,898,431 samples). Like `covtype`, the loader downloads a gzip-compressed
/// file and decompresses it with `gunzip`.
/// Letter Recognition dataset module.
///
/// Contains the Letter Recognition dataset (UCI, Slate 1991) for multi-class
/// classification. It identifies which of the 26 capital letters a distorted
/// glyph shows, from 16 integer statistics of its pixel image. This is the
/// crate's widest classification problem by class count. Its label column
/// `LetterRecognition::TARGET` holds one capital letter per sample, as a
/// one-character string, so it needs no lookup table.
/// Linnerud dataset module.
///
/// Contains the scikit-learn Linnerud dataset (`load_linnerud`) for multi-output
/// regression. It predicts three physiological variables (`Weight`, `Waist`,
/// `Pulse`) from three exercise variables (`Chins`, `Situps`, `Jumps`), measured
/// on 20 middle-aged men.
/// Movie Review Polarity dataset module.
///
/// Contains the Cornell Movie Review Polarity dataset (Pang and Lee 2004,
/// polarity dataset v2.0) for binary **text** classification. It labels 2,000
/// full IMDb movie reviews as `positive` or `negative` (1,000 each). Like
/// [`sms_spam`], its document column is `ColumnData::String`. It complements the
/// sentence-level [`sentiment_sentences`] with full-document reviews. Sourced
/// from a `.tar.gz` archive (decompressed with `untar_gz`).
/// MNIST handwritten digits dataset module.
///
/// Contains the MNIST database (LeCun et al. 1998) for multi-class
/// classification: recognizing a handwritten digit (`0`-`9`) from a 28×28
/// grayscale image. Its source is the binary IDX format, read from four
/// gzip-compressed files. One `pixels` column holds `ColumnData::Bytes` of shape
/// `(n_samples, 784)`, which a `(n_samples, 28, 28)` view reads at no copy. It
/// offers `new`/`new_test`/`new_all` subset constructors.
/// MovieLens 100K dataset module.
///
/// Contains the MovieLens 100K dataset (GroupLens, Harper & Konstan 2015) for
/// **recommendation**: 100,000 ratings that 943 users gave to 1,682 movies
/// between September 1997 and April 1998. One sample is one rating. The table
/// holds the `user_id` and `item_id` columns, the `rating` column
/// (`MovieLens100k::TARGET`), and a `timestamp` column of Unix seconds.
/// `MovieLens100k::N_USERS` and `MovieLens100k::N_ITEMS` give the two identifier
/// ranges. GroupLens permits research use under conditions that this crate's MIT
/// license does not cover, so read the struct docs before you use the data.
/// Mushroom dataset module.
///
/// Contains the Mushroom dataset (UCI `agaricus-lepiota`) for binary
/// classification: predicting whether a mushroom is edible or poisonous from 22
/// categorical attributes. It is **all-categorical**: every feature is a
/// single-letter string code in a `ColumnData::String` column.
/// 20 Newsgroups dataset module.
///
/// Contains the classic 20 Newsgroups dataset (Lang 1995, the `bydate` version)
/// for multi-class **text** classification: labeling ~18,846 Usenet posts with
/// one of 20 newsgroups. It is the framework-agnostic analogue of scikit-learn's
/// `fetch_20newsgroups`. It is the **multi-class** text corpus, with 20 classes.
/// Like [`sms_spam`], its document column is `ColumnData::String`.
/// `new`/`new_test`/`new_all` mirror scikit-learn's train/test/all subsets.
/// Sourced from a `.tar.gz` archive (decompressed with `untar_gz`).
/// Palmer Penguins dataset module.
///
/// Contains the Palmer Penguins dataset for classifying penguins into three
/// species (Adelie, Chinstrap, Gentoo). It uses bill and flipper measurements,
/// body mass, and categorical island and sex features.
/// Sentiment Labelled Sentences dataset module.
///
/// Contains the Sentiment Labelled Sentences dataset (UCI, Kotzias et al. 2015)
/// for binary **text** classification. It labels 3,000 review sentences from
/// three sites (Amazon, IMDb, Yelp) as `positive` or `negative`. Like
/// [`sms_spam`] and [`youtube_spam`], it is a text-modality loader: its document
/// column is `ColumnData::String`. A third column, `source`, names the review
/// site of each sentence. Sourced from a ZIP archive of three per-site files.
/// SMS Spam Collection dataset module.
///
/// Contains the SMS Spam Collection dataset (UCI, Almeida and Hidalgo 2011) for
/// binary **text** classification: labeling 5,574 SMS messages as `ham` or
/// `spam`. Its document column is `ColumnData::String`, one raw message per
/// sample. Sourced from a ZIP archive.
/// Spambase dataset module.
///
/// Contains the Spambase dataset (UCI, Hopkins et al. 1999) for binary
/// classification. It labels 4,601 emails as `ham` or `spam` from 57 numeric
/// features: word and character frequencies, plus capital-run-length statistics.
/// This is the feature-engineered counterpart to the crate's raw-text spam
/// corpora ([`sms_spam`], [`youtube_spam`]). Those loaders leave vectorization to
/// you, but Spambase already does it, so it drops straight into a numeric model.
/// Titanic dataset module.
///
/// Contains data about Titanic passengers for predicting survival based
/// on features like passenger class, sex, age, and fare.
/// Wholesale Customers dataset module.
///
/// Contains the Wholesale Customers dataset (UCI, Cardoso 2013) for
/// **clustering**: segmenting 440 clients of a Portuguese wholesale distributor
/// by their annual spending across six product categories, plus a sales-channel
/// code and a region code. The dataset has **no target column**, so the loader
/// has no target constant: all 8 columns are numeric features.
/// `WholesaleCustomers::COLUMN_NAMES` names them in source order.
/// Wine Quality dataset module.
///
/// Contains wine quality assessment data for predicting quality scores
/// based on physicochemical properties like acidity, sugar content, and
/// alcohol percentage.
/// Wine Recognition dataset module.
///
/// Contains the scikit-learn Wine recognition dataset for classifying wines
/// into three cultivars based on 13 chemical constituents. Distinct from
/// [`wine_quality`], which is a regression task on quality scores.
/// YouTube Spam Collection dataset module.
///
/// Contains the YouTube Spam Collection dataset (UCI, Alberto, Lochter, and
/// Almeida 2017) for binary **text** classification. It labels 1,956 comments
/// from five popular music videos as `ham` or `spam`. Like [`sms_spam`] (a
/// sibling by the same authors), it is a text-modality loader: its document
/// column is `ColumnData::String`, one raw comment per sample. Sourced from a ZIP
/// archive of five per-video CSVs.