falcon_mdf 0.6.0

High-performance Rust library for reading ASAM MDF v4 (MF4) measurement data files
Documentation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
<p align="center">
  <img src="https://raw.githubusercontent.com/mohammad-albarham/falcon_mdf/main/assets/logo.jpg" alt="falcon_mdf logo" width="200">
</p>

# Falcon MDF

A high-performance Rust library for reading ASAM MDF (Measurement Data Format) v2.14, v3.x and v4.x files.

[![Crates.io](https://img.shields.io/crates/v/falcon_mdf.svg)](https://crates.io/crates/falcon_mdf)
[![Documentation](https://docs.rs/falcon_mdf/badge.svg)](https://docs.rs/falcon_mdf)
[![License](https://img.shields.io/badge/License-MIT%20OR%20Apache--2.0-blue.svg)](LICENSE-MIT)
[![Demo](https://img.shields.io/badge/demo-in%20your%20browser-4f8cff.svg)](https://mohammad-albarham.github.io/falcon_mdf/)

**Try it without installing anything:** the [browser demo](https://mohammad-albarham.github.io/falcon_mdf/)
opens an `.mf4` file and plots its channels entirely client-side via
[WebAssembly](wasm/) — the file never leaves your machine.

## Overview

**falcon_mdf** reads MDF measurement files, the format automotive and industrial
acquisition tools record to. It aims at three things in this order:

- **Correct, or it says so.** A channel decodes to the right values, or reading
  it fails with a reason. It never returns part of the data, or a raw value in
  place of a converted one, dressed up as a measurement. Decoded output is
  checked against an independent reference implementation over a corpus of CAN,
  LIN and GPS/IMU logs.
- **Safe on files you did not write.** Malformed input comes back as an error —
  not a panic, an aborted process, or a loop that never ends. An audit over
  1,049 mutated files found three hard failures, against 71 for asammdf and 61
  for mdfreader; those three are fixed, and a fresh sweep of 1,200 mutated
  files — truncations, corrupted lengths and links, bad block IDs, zeroed
  fields — produces no panic, no abort and no hang. That sweep is only worth
  the paper it is written on because a deliberately crashing build was fed
  through the same harness first, to prove it reports a crash when one happens.
- **Fast.** Usually the faster reader, and often by a large margin: on the
  reference OBD2 CANedge log (326,623 samples) roughly 3.9× for decoding and
  4.8× for a whole read, and 3.1× to 31.9× across other uncompressed files —
  though at parity or slower on some vendor-compressed files. The spread is
  real; see Performance.

## Features

- Read MDF 4.x files (4.0, 4.1, 4.2), sorted and unsorted, finished and unfinished
- Read MDF 3.x files (2.14, 3.20, 3.30) — structure, samples and conversions
  (`mdf3` feature, off by default)
- Memory-mapped and buffered I/O
- HD, DG, CG, CN, DT, DZ, DL, HL, TX, MD, CC and SI blocks
- `##LD` linked-data list blocks, `##DV` data-value blocks and `##DI` data-invalidation blocks
- Compressed data blocks, in every form asammdf writes: plain and transposed
  deflate, and — behind the `zstd` and `lz4` features, both off by default —
  zstd, transposed zstd, LZ4 and transposed LZ4. A block whose compression is
  not compiled in reports itself by name rather than returning wrong bytes
- Typed samples: an integer channel decodes to an integer of its own width, a
  frame payload to bytes, a text table to text
- Variable-length signal data, in both storage forms
- Array (CA) channels whose elements sit in the record, decoded to flat values
  with the per-dimension shape available — including look-up arrays composed of
  nested CA blocks, and arrays whose length varies per sample
- The acquisition source behind a channel or group: which ECU, bus or tool it
  came from
- Conversion rules: identity, linear, rational, algebraic formulas, value and
  range tables, value-to-text, text-keyed and bitfield tables
- CAN frames out of bus-logged groups: timestamp, identifier, extended flag,
  bus channel and a payload trimmed to the logged length — with no database
  needed, and no interpretation of the payload
- Streamed reading of a channel in bounded windows, so peak memory does not
  scale with the largest data group — including unsorted groups, which are
  demultiplexed per window, and bus-log payload channels
- CAN payloads decoded against a CAN database into named physical signals: bit
  position and width, Intel and Motorola byte order, signedness, scaling,
  multiplexed signals and `VAL_` tables decoded to text — from DBC files (`dbc`
  feature) or AUTOSAR ECU extracts (`arxml` feature), both through one decoder
- DBC extended multiplexing (`SG_MUL_VAL_`) with range selectors
- DBC global value tables (`VAL_TABLE_`)
- J1939 parameter-group matching (`IdMatching::J1939Pgn`), matching frames from
  any source address against a database written for one
- J1939 parameter-group and source-address matching (`IdMatching::J1939PgnAndSource`),
  matching by exact `(PGN, source)` pair before falling back to PGN alone
- ARXML dynamic multiplexed PDUs, with selector-field resolution
- LIN frames out of bus-logged groups: timestamp, the six-bit identifier, bus
  channel and a payload trimmed to the logged length — frames only, with no
  database and no interpretation of the payload
- LIN Description File (LDF) database decoding: parse `.ldf` files and decode
  LIN frames into physical signals with units and value tables
- Per-sample validity from invalidation bits
- Metadata as a comment plus named properties, rather than raw XML
- Attachments (embedded data only) and events
- Channel hierarchy (CH blocks): full tree traversal resolving nodes to their
  member channels and element paths (asammdf stores the root link without traversing it)
- Sample reduction (SR blocks): reading reduction level descriptors and decoding
  condensed value series (mean, min, max) via `reduced_signal` (asammdf does not implement SR blocks)
- The file as a file: `block_map` walks the block graph and returns every block
  in address order — length, links under the format's own names, referrers, and
  a line describing its fields — plus the bytes no block covers, and `read_raw`
  hands back the bytes at any offset
- Time-domain operations: `cut` a channel to a time window, `resample` onto a
  fixed raster with step-hold or linear interpolation
- Multi-channel operations: `filter` picks named channels out of a file,
  `concatenate` joins measurements end to end, `stack` overlays them aligned by
  start time or recorded time
- Batched reading: `signals()` decodes any set of channels in one pass, assembling
  each channel group's records once
- Signal algebra: `SignalSeries` implements `Add`, `Sub`, `Mul`, `Div` against
  both other series and scalars, with automatic resampling onto the union of
  their timestamps
- Channel search: `find_channels` for exact name matches, `search_channels` for
  substring, wildcard or regex queries
- Anonymisation: `scramble_file` replaces every piece of identifying text with
  random bytes of the same length, leaving sample data and decoding formulas
  untouched
- Export decoded channels to CSV, Apache Parquet (`parquet` feature) or MATLAB
  level 5 MAT-files (`mat` feature)
- Writing MF4 files from scratch: typed channels, per-sample validity, conversion
  rules, and optional deflate compression into `##DZ` blocks

### Not supported

Named so you can tell before you depend on it:

- **MDF 4.20 `##LD` chains carrying separate invalidation with incompatible
  channel-group layouts.** LD/DV reading and compatible split invalidation
  layouts are supported. Layouts that cannot be interleaved unambiguously are
  rejected during opening.
- **Big-endian MDF 3.x files.** Little-endian 3.x reads correctly; big-endian
  is reported by name.
- **Arrays stored one channel group or data group per element**
  (CG- and DG-template `ca_storage`), and arrays with more than one
  dynamically-sized dimension. asammdf does not support them either (it logs
  "Only CN template arrays are supported"); only CN-template arrays (elements
  stored contiguously in the record) are supported in both readers.
- **Sync channels** (`cn_type` 4), which index a media stream rather than
  measure something.
- **Streamed reading of a variable-length channel whose payloads sit in its own
  signal-data block.** `signal_chunks` refuses it by name rather than reading it
  wrongly; `signal` reads it, materialising the group. The companion-group form
  that bus loggers write *is* streamed.
- **Lossless editing of arbitrary existing files.** `Mf4Writer::from_file`
  creates an editable representation of supported channels; it can skip
  unreadable or unrepresentable channels and does not preserve every metadata
  block. Fixed CN-template arrays, VLSD strings/bytes and multiple channel
  groups per data group can be written and have dedicated conformance tests.

Of the channel-level items above, each reports itself by name through
`Mf4Error::Unsupported` when you read such a channel, and the rest of the file
still opens and decodes. Whole-file limits are different: a big-endian MDF 3.x
file, and an MDF 4.20 LD chain whose separate invalidation requires incompatible
record layouts, fail during opening.

### Tested against

Every claim above is exercised by the test suite. Two areas are implemented but
have no file available to test them: **big-endian channels** are covered by
synthetic tests only, and the vendor reference corpus is primarily **MDF 4.11**.
An asammdf-generated **MDF 4.20** LD/DV file is checked by
`tests/asammdf_ld_conformance.rs`; separate invalidation is covered by synthetic
fixtures in `tests/synthetic_blocks.rs`. These cases do not establish complete
MDF 4.20 compatibility across vendors. See the limitations above and
[the format review](docs/mf4-review.md) for remaining priorities.

## Installation

```toml
[dependencies]
falcon_mdf = "0.6"
```

Memory mapping is on by default. For a file another process may be writing, or
one on a network share, open it buffered instead — the buffered backend does
not require the file to stay unmodified. Its memory footprint relative to the
mapped backend is not a fixed saving; it varies with file size (see Memory).

```toml
[dependencies]
falcon_mdf = { version = "0.6", default-features = false }
```

Decoding CAN payloads against a database needs the `dbc` feature (DBC files) or
`arxml` (AUTOSAR ECU extracts). Reading a data block compressed with something
other than deflate needs `zstd` or `lz4`. Reading MDF 3.x files needs `mdf3`.
Exporting to Parquet needs `parquet`; exporting to MATLAB MAT needs `mat`. All
six are off by default, so that reading a plain, deflate-compressed measurement
file pulls in neither a database parser nor a second decompressor.

```toml
[dependencies]
falcon_mdf = { version = "0.6", features = ["dbc", "arxml", "zstd", "lz4", "mdf3", "parquet", "mat"] }
```

The crate's MSRV is **1.89**, and it covers every feature: CI builds
`--all-features` on 1.89 on every push. The floor is set by `autosar-data`,
which the `arxml` feature pulls in; without that feature the crate builds on
considerably less, but a declared MSRV that only holds for some feature
combinations is not a number anyone can rely on.

### Feature flags

| Flag | Default | What it pulls in |
|---|---|---|
| `mmap` | on | Memory-mapped I/O backend (`memmap2`) |
| `dbc` | off | DBC file parsing via `can-dbc` 10.x |
| `arxml` | off | AUTOSAR ARXML parsing via `autosar-data` 0.22 |
| `zstd` | off | Zstandard decompression via `ruzstd` 0.7 |
| `lz4` | off | LZ4 frame decompression via `lz4_flex` 0.11 |
| `mdf3` | off | MDF 3.x reader (`src/mdf3/`) |
| `parquet` | off | Apache Parquet export via `parquet` 59 + Arrow 59 |
| `mat` | off | MATLAB level 5 MAT-file export (no extra dependency) |

## Quickstart

### Opening an MDF4 File

```rust
use falcon_mdf::Mf4File;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Open with memory-mapped I/O (recommended for large files)
    let file = Mf4File::open("measurement.mf4")?;

    // Print file version and statistics
    println!("Version: {}", file.version());
    println!("Data groups: {}", file.data_group_count());
    println!("Total channels: {}", file.channel_count());
    println!("Start time: {:?}", file.start_time());

    Ok(())
}
```

### Opening an MDF3 File

```rust
use falcon_mdf::mdf3::Mdf3File;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let file = Mdf3File::open("measurement.mdf")?;
    println!("Version: {}", file.version());
    println!("Channels: {}", file.channel_count());
    Ok(())
}
```

### Reading a channel

Samples come back in the channel's own type. A 29-bit CAN identifier is a
`u32`, a two-bit bus number a `u8`, a frame payload bytes — nothing is forced
through `f64` unless you ask for it.

A group that recorded no samples decodes to an empty channel — a bus logger
writes one such group for every bus that carried no traffic — so index with
`first()` rather than `[0]`.

```rust
use falcon_mdf::{Mf4File, SignalValues};

let file = Mf4File::open("measurement.mf4")?;

if let Some(channel) = file.find_channel("VehicleSpeed") {
    let signal = file.signal(channel)?;

    match signal.values()? {
        SignalValues::F64(v) => println!("first: {:?} {}", v.first(), signal.unit()),
        SignalValues::U32(v) => println!("first: {:?}", v.first()),
        other => println!("{} samples of {}", other.len(), other.kind().name()),
    }

    // Or a uniform numeric view, lossy for wide integers and byte channels.
    let as_f64 = signal.values_f64()?;
    println!("{} samples", as_f64.len());
}
# Ok::<(), falcon_mdf::error::Mf4Error>(())
```

### Samples the file marks invalid

A channel may carry an invalidation bit. `values()` does not filter those out —
dropping them would break alignment with the master channel — so check validity
alongside the data.

```rust
# use falcon_mdf::Mf4File;
# let file = Mf4File::open("measurement.mf4")?;
# let channel = file.find_channel("VehicleSpeed").unwrap();
let signal = file.signal(channel)?;
let values = signal.values_f64()?;

match signal.validity() {
    Some(valid) => {
        for (value, ok) in values.iter().zip(&valid) {
            if *ok {
                println!("{value}");
            }
        }
    }
    // No invalidation bit: every sample is valid.
    None => println!("{} valid samples", values.len()),
}
# Ok::<(), falcon_mdf::error::Mf4Error>(())
```

### Channels this build cannot decode

Reading fails rather than returning something plausible. Handle it explicitly if
you process files you have not seen.

```rust
# use falcon_mdf::{Mf4File, error::Mf4Error};
# let file = Mf4File::open("measurement.mf4")?;
for channel in file.channels() {
    match file.signal(channel).and_then(|s| s.values()) {
        Ok(values) => println!("{}: {} samples", channel.name, values.len()),
        Err(Mf4Error::Unsupported { feature, .. }) => {
            println!("{}: skipped, needs {feature}", channel.name)
        }
        Err(e) => return Err(e),
    }
}
# Ok::<(), falcon_mdf::error::Mf4Error>(())
```

### File metadata

```rust
# use falcon_mdf::Mf4File;
# let file = Mf4File::open("measurement.mf4")?;
println!("{}", file.comment());

if let Some(serial) = file.metadata().get("Device Information/serial number") {
    println!("recorded by {serial}");
}
# Ok::<(), falcon_mdf::error::Mf4Error>(())
```

### Writing a file

`Mf4Writer` creates MF4 files from scratch: one data group per channel group,
an implicit `Time` master per group, records sorted by time. Channels are
written in their own type — an integer of its own width and signedness, a
32- or 64-bit float, a fixed-length string, or a fixed-width byte run. A
channel may also carry a conversion rule so that raw counts read back as the
physical quantity they stand for. Validity can be carried over per sample, so
an export keeps the gaps the source declared. Optional deflate compression
writes each group's records as a `##DZ` block behind the `##HL`/`##DL` pair
the standard requires.

```rust
# use falcon_mdf::Mf4Writer;
let mut writer = Mf4Writer::new();
let group = writer.add_group(&[0.0, 0.1, 0.2])?;
group.add_channel("Speed", "km/h", &[0.0, 5.0, 10.0])?;
group.add_channel_with_validity(
    "Boost", "psi", &[1.0, 2.0, 3.0], Some(&[true, false, true]),
)?;
writer.write_to_file("out.mf4")?;
# Ok::<(), falcon_mdf::error::Mf4Error>(())
```

### Decoding a bus log

A bus logger records raw frames; what the payload bytes mean lives in a DBC or
ARXML database you supply. `decode_bus` reads the whole file against one and
hands back each signal as a time series.

```rust,no_run
use falcon_mdf::{CanDatabase, IdMatching, Mf4File};

let file = Mf4File::open("truck.mf4")?;
let database = CanDatabase::from_dbc_path("j1939.dbc")?
    // A J1939 database keys messages by parameter group, while the identifier
    // on the wire also carries the sending ECU's address. Without this, a real
    // heavy-duty log decodes to nothing.
    .with_matching(IdMatching::J1939Pgn);

for signal in file.decode_bus(&database)?.iter() {
    println!(
        "{}.{}: {} readings [{}]",
        signal.message, signal.name, signal.len(), signal.unit
    );
    // Enum-valued signals carry their `VAL_` label as well as the number.
    if let Some(text) = signal.text_at(0) {
        println!("  first reading: {text}");
    }
}
# Ok::<(), falcon_mdf::error::Mf4Error>(())
```

Decoded signals are a namespace of their own — they do **not** appear in
`channels()`, because they are derived from a database the file does not
contain. A series is identified by bus, message and name together: two
messages may spell one signal name, and the same identifier on two buses is two
different signals.

For frame-level access with no database at all, `can_frame_groups` and
`can_frames` give the timestamp, identifier and payload directly. See
`examples/decode_bus.rs` for the whole path in one file.

### LIN database decoding

```rust,no_run
use falcon_mdf::{CanDatabase, Mf4File};

let file = Mf4File::open("lin_log.mf4")?;
let database = CanDatabase::from_ldf_path("lin_database.ldf")?;

for signal in file.decode_lin(&database)?.iter() {
    println!(
        "{}.{}: {} readings [{}]",
        signal.message, signal.name, signal.len(), signal.unit
    );
}
# Ok::<(), falcon_mdf::error::Mf4Error>(())
```

### Signal algebra

`filter` hands back `SignalSeries`, which implement the arithmetic operators,
resampling onto the union of their timestamps automatically:

```rust
# use falcon_mdf::Mf4File;
# let file = Mf4File::open("measurement.mf4")?;
let mut series = file.filter(&["Speed".into(), "RPM".into()])?;
let rpm = series.pop().unwrap();
let speed = series.pop().unwrap();
let sum = &speed + &rpm;
let scaled = &speed * 2.0;
# Ok::<(), falcon_mdf::error::Mf4Error>(())
```

### Export

`filter` picks the channels to export by name — a name several channels share,
such as every group's master, is rejected rather than guessed at; the
`ChannelSelector` enum disambiguates by group or position.

```rust
# use falcon_mdf::{Mf4File, export::write_parquet};
# let file = Mf4File::open("measurement.mf4")?;
let series = file.filter(&["VehicleSpeed".into(), "EngineRPM".into()])?;
let mut out = std::fs::File::create("measurement.parquet")?;
write_parquet(&series, &mut out)?;
# Ok::<(), falcon_mdf::error::Mf4Error>(())
```

CSV is always available. Parquet needs the `parquet` feature; MATLAB MAT needs
the `mat` feature.

## Architecture (Click on the image)

[![Runtime architecture map](docs/architecture.png)](https://mohammad-albarham.github.io/falcon_mdf/architecture.html)

The image links to an interactive version of this map (`docs/architecture.html`,
served by GitHub Pages): one self-contained HTML file with dark and light
themes, pan and zoom, guided tours, and per-node links back into this
repository's sources.

The library is organized in layers, each with a clear responsibility:

```text
┌─────────────────────────────────────────────┐
│                  file.rs                     │  User-facing API
│            (Mf4File, high-level)             │
├─────────────────────────────────────────────┤
│                  model/                      │  Domain types
│      (DataGroup, Channel, Signal, etc.)     │
├─────────────────────────────────────────────┤
│                 parser/                      │  Version-aware parsing
│    (Mf4Version, LinkedBlockIterator)        │
├─────────────────────────────────────────────┤
│                 blocks/                      │  Low-level block types
│  (HdBlock, DgBlock, CnBlock, DtBlock, etc.) │
├─────────────────────────────────────────────┤
│                   io/                        │  File access abstraction
│     (ByteSource, MmapSource, Buffered)      │
└─────────────────────────────────────────────┘
```

### Module Descriptions

| Module | Description |
|--------|-------------|
| `io/` | File I/O abstraction with mmap and buffered backends |
| `blocks/` | Low-level MDF4 and MDF3 block parsers following the ASAM spec |
| `parser/` | Version detection and block traversal utilities |
| `model/` | High-level types representing channels, signals, and metadata |
| `file.rs` | Main `Mf4File` API for opening and reading MDF4 files |
| `mdf3/` | `Mdf3File` API for opening and reading MDF3 files (`mdf3` feature) |
| `stream.rs` | `signal_chunks` — reading a channel in bounded windows |
| `cache.rs` | Parsed-block cache, shared by file offset |
| `bus.rs` | CAN frame extraction and bus signal decoding against a database |
| `candb.rs` | Format-neutral CAN database model and signal decoder |
| `dbc.rs` | `CanDatabase::from_dbc` — reading DBC files (`dbc` feature) |
| `arxml.rs` | `CanDatabase::from_arxml_path` — reading AUTOSAR ECU extracts (`arxml` feature) |
| `ldf.rs` | `CanDatabase::from_ldf_path` — reading LIN Description Files |
| `lin.rs` | `lin_frames` — LIN frames out of a bus-logged group |
| `write.rs` | `Mf4Writer` — creating MF4 files from scratch |
| `export/` | `write_csv`, `write_parquet`, `write_mat` — decoded channels to other formats |
| `inspect.rs` | `block_map` — every block in a file, in address order |
| `time_ops.rs` | `SignalSeries` — cut, resample, and signal arithmetic |
| `multi_ops.rs` | `filter`, `concatenate`, `stack` — operations spanning channels and files |
| `scramble.rs` | `scramble_file` — anonymise text in a measurement file |
| `error.rs` | Comprehensive error types with `thiserror` |

## Performance

Medians, decoding 326,623 samples from the reference OBD2 CANedge log against
asammdf. As the Overview says, the result depends on the file and on the
asammdf entry point you compare against:

The full comparison is tracked in this repository:
[`benchmarks/COMPARISON.md`](benchmarks/COMPARISON.md) curates per-file
timings across an 81-file corpus, size-bucket aggregates, memory measurements,
and 122 MB / 480 MB fixtures. The
[performance review](benchmarks/performance-review.html) includes paired
before/after measurements and verification results. Raw generated reports
are available in [`benchmarks/`](benchmarks/).

The table below records earlier measurements, before the September 2026
reader optimizations; use the linked comparison for current results.

| Scene | Speedup over asammdf | Measured |
|---|---|---|
| OBD2 CANedge log, decoding only | 3.9× | yes |
| Same file, whole read | 4.8× | yes |
| Uncompressed, 13 other files | 3.1×–31.9× | yes |
| DZ-compressed, per-channel via `mdf.get` | 6.7×–9.1× | yes, 4 files |
| DZ-compressed, per-channel via `mdf.select` | 5.6×–7.6× | yes, 4 files |
| DZ blocks written by native vendor tools | 0.85×–1.01× | **no — see below** |
| Compressed, 126 MB file | 0.81× — slower | **no — see below** |

Medians of three to five runs each, warm cache. The entry point matters: the
compressed figures above are against asammdf's per-channel `mdf.get`, and drop
by roughly a fifth against `mdf.select`, which amortises its setup across
channels.

The last two rows are the ones you should weigh most and we can least support.
They come from an audit whose corpus included a 126 MB file and vendor-written
DZ blocks that this repository's fixtures do not contain, so nothing here
reproduces them — and they are precisely the cases where falcon_mdf stops
winning. The measured rows all come from files of 5 MB or less; a reader that
is several times faster on those may well converge toward parity as the file
outgrows cache, which is what that audit reports and what the "at parity or
slower" clause in the Overview refers to.

Opening a file — parsing its structure without reading samples — is quicker than
decoding samples, which matters when you only want to know what a file contains.

### Choosing a backend

| Backend | When | Trade-off |
|---|---|---|
| Memory-mapped (default) | Files that are finished being written | Fastest reads. The file must not be modified while open: another process truncating it raises `SIGBUS`, which is not a catchable Rust error. |
| Buffered | Files still being written, on a network share, or that another user can replace. Also large files. | Copies what it reads, so it carries no such requirement: the file may be modified or replaced while it is open. How its memory compares to the mapped backend varies with file size (see Memory). |

### Memory

A data group's records are assembled into one buffer before they are read, so
peak memory scales with the **largest data group**, not with the file. Under the
memory-mapped backend the data is resident twice — once as mapped pages, once as
that buffer.

Measured reading a 416 MB file:

| Backend | Peak resident |
|---|---|
| Memory-mapped | 826 MB |
| Buffered | 434 MB |

That single measurement is not a general result, and the 416 MB file is no
longer available to repeat it on. Re-measuring peak resident size across every
file that *is* available — 1.6 KB to 5.2 MB, median of three runs each — the
buffered backend shows no consistent saving:

| File size | Buffered ÷ mapped |
|---|---|
| 1.6 KB – 70 KB | 0.98–1.01 — parity; peak is the process baseline, not the file |
| 1.2 MB | 0.80 — the largest saving seen anywhere |
| 5.2 MB | **1.13 — buffered costs more**, on four separate files |

Whether the halving above holds at 416 MB is untested: the mapped backend's
resident pages grow with the file in a way five-megabyte samples cannot show,
so these numbers neither confirm nor refute it. What they do rule out is
reading it as a saving you can count on at any size.

Prefer `Mf4File::open_buffered` for the reasons in the table above — a file
another process may modify — not for a memory ratio.
Decoding block by block, which would make memory independent of group size, is
planned but not implemented.

### Build settings

The release profile in this repository already sets these; if you vendor the
crate, they are worth keeping:

```toml
[profile.release]
lto = true
codegen-units = 1
opt-level = 3
```

### Reproducing the benchmarks

To reproduce the benchmark comparisons against asammdf:

1. Fetch the vendor reference MF4 measurement files:
   ```bash
   scripts/fetch_reference_files.sh
   ```
   This downloads the public reference corpus into `test_data/reference/`.

2. Run the comparison benchmark using a Python environment with `asammdf` installed (such as `.venv/bin/python`):
   ```bash
   .venv/bin/python scripts/bench_vs_asammdf.py --limit 10
   ```

Running the benchmark requires Python 3 with `asammdf` (e.g. `pip install asammdf` in a virtual environment, or `.venv/bin/python`). The script recursively scans `test_data/` (or `--data-dir <path>`), automatically compiles the release build of `examples/bench.rs` if needed, and measures the median of warm-cache runs for each file. Use `--limit N` (default 10, or `0` for all files) to control the number of files benchmarked. If `asammdf` is not installed or the data directory contains no `.mf4` files, the script reports `skipped: <reason>` and exits with status 0.

The results of the full comparison suite are committed under
[`benchmarks/`](benchmarks/): `COMPARISON.md` is the curated summary, and the
`latest_*` and `large_*` files beside it are the raw generated reports
(per-file timings, size buckets, sample-count agreement, memory) that the
summary is built from.

## MDF4 Block Reference

| Block | Description |
|-------|-------------|
| **ID** | File identification (MDF signature, version) |
| **HD** | Header block (file metadata, timestamps) |
| **DG** | Data group (logical grouping of channels) |
| **CG** | Channel group (channels with shared time axis) |
| **CN** | Channel block (signal definition) |
| **DT** | Data block (raw measurement data) |
| **DZ** | Zipped data block (compressed) |
| **DL** | Data list (linked data blocks) |
| **LD** | Linked data list block |
| **DV** | Data-value block |
| **DI** | Data-invalidation block |
| **HL** | Hierarchy list (nested block structure) |
| **TX** | Text block (strings) |
| **MD** | Metadata block (XML content) |
| **CC** | Conversion block (value transformation) |
| **SI** | Source information (ECU, tool metadata) |

## Error Handling

All operations return `Result<T, Mf4Error>`:

```rust
use falcon_mdf::{Mf4File, Mf4Error};

fn open_file(path: &str) -> Result<(), Mf4Error> {
    let file = Mf4File::open(path)?;
    // ... use file
    Ok(())
}
```

Error types include:
- `Mf4Error::Io` - File I/O errors
- `Mf4Error::InvalidSignature` - Not a valid MDF file
- `Mf4Error::UnsupportedVersion` - Unknown MDF version
- `Mf4Error::InvalidBlockId` / `InvalidBlockSize` - Malformed block structure
- `Mf4Error::Decompression` - Zlib/zstd/LZ4 decompression failure
- `Mf4Error::ChannelNotFound` - Channel lookup failed
- `Mf4Error::Unsupported` - A channel or feature this build cannot decode, named by feature

## Examples

The `examples/` directory contains:

- **list_channels.rs** - Enumerate all channels in a file
- **export_to_csv.rs** - Export a channel to CSV format
- **write_mf4.rs** - Create an MF4 file with `Mf4Writer`
- **decode_bus.rs** - Decode a bus-logged file against a DBC, printing the
  named physical signals the frames carry (requires `--features dbc`)
- **block_map.rs** - Print every block in a file, in the order they sit on disk,
  with the gaps and anything the walk could not make sense of

Run examples with:

```bash
cargo run --example list_channels -- measurement.mf4
cargo run --example export_to_csv -- measurement.mf4 VehicleSpeed output.csv
cargo run --example write_mf4 -- out.mf4
cargo run --features dbc --example decode_bus measurement.mf4 database.dbc [--j1939]
cargo run --example block_map -- measurement.mf4 [--summary]
```

## Testing

Run the test suite:

```bash
cargo test
```

Run with logging:

```bash
RUST_LOG=debug cargo test
```

## GUI

> **The viewer is not stable.** It is pre-1.0, and the least settled part of
> this project. The interface, the CSV and MF4 it exports, and the session
> state it saves may all change between versions, and its coverage of what
> vendors actually emit is uneven — a file that opens elsewhere may still fail
> or plot as nothing here. Use it to look at measurements, not as something to
> build a process on, and check anything that matters against the source.
> [gui/RUNNING.md](gui/RUNNING.md#status-not-stable) sets out what is unstable.

`gui/` contains `falcon`, a desktop viewer built as a library (`falcon_mdf_gui`)
plus a thin binary, so its search, decimation, session, formatting and loader
logic are testable without a window. The window is two panes: on the left the
file — a structure tree with Expand/Collapse and group filtering, the
address-ordered block list, and the searchable channel list (matching names,
units, comments and acquisition names via substring, wildcard or regex, with
filter toggles for arrays, unreadable channels and masters) — and on the right
whatever is selected there across six tabs. That content is a details view (a
block shows its links as buttons and its bytes as a hex dump), overlay and
stacked plots with min-max decimation, placeable measurement cursors with
region statistics, per-signal display styling, and relative or absolute UTC
time axes, a numeric view answering instantaneous values across plotted
channels, a sample table with column sorting and row filtering, a CAN or LIN
frame list (CAN frames optionally decoded against a DBC), or per-channel
statistics with a distribution. Attachments, events, file history and the
channel hierarchy are nodes in the tree; CSV and MF4 export sit above the plot
and sample table. The core library itself stays GUI-free.
[gui/RUNNING.md](gui/RUNNING.md) covers running it and opening a measurement;
build and packaging notes live in [gui/PACKAGING.md](gui/PACKAGING.md).

## License

Licensed under either of:

- Apache License, Version 2.0 ([LICENSE-APACHE](LICENSE-APACHE) or http://www.apache.org/licenses/LICENSE-2.0)
- MIT license ([LICENSE-MIT](LICENSE-MIT) or http://opensource.org/licenses/MIT)

at your option.

## Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

## References

- [ASAM MDF Standard](https://www.asam.net/standards/detail/mdf/)

## Acknowledgments

This library was designed to provide a robust, performant foundation for working with measurement data in Rust. Special thanks to the ASAM organization for the MDF specification.