rullama-datasets 0.12.0

Training data pipelines for rullama: JSONL I/O, tokenization, dedup, format conversion (Alpaca / ChatML / OpenAI / ShareGPT / Together), and quality stats. Shared by cloud (`rullama-finetune`) and local (`rullama-lora`, in the rullama workspace) fine-tuning.
Documentation

rullama-datasets

There is very little structured metadata to build this page from currently. You should check the main library docs, readme, or Cargo.toml in case the author documented the features in them.

This version has 4 feature flags, 0 of them enabled by default.

default

This feature flag does not enable additional features.

datasets-dedup

datasets-full

datasets-hf-tokenizer

datasets-tiktoken