Русский: README.ru.md · 中文: README.zh.md
CMF — Cortiq Model Format
CMF is an auditable model container: one file can hold weights, tokenizer, chat metadata, task masks and skill overlays for inference without a large framework runtime.
Status
CMF v2 is the current on-disk format. Readers validate the envelope, section bounds, tensor metadata and hashes; incompatible changes require a feature bit or version bump. The Rust crate APIs are still pre-1.0 and may change.
The project is usable for local inference and format experiments. Treat model quality and speed numbers as workload-specific measurements, not guarantees; see the focused guides for the test setup behind each claim.
Quick start
Install the CLI and convert a small public checkpoint:
Already have GGUF? Import it directly:
Native Qwen Image Edit
Version 0.6.9 adds the native Qwen-Image-Edit-2509 path and the explicit q2tp_affine Prism profile. The ready default is
one self-contained CMF, while the standalone component files remain available
when users want to stage them independently. Build the bundle from a directory
that contains transformer.cmf, text_encoder.cmf, vae.cmf, and
scheduler_config.json:
The bundle embeds the transformer, Qwen2.5-VL text/vision encoder, tokenizer/processor/configuration, VAE, and scheduler. Run it directly with no companion files:
The command requires at least one reference image; repeat --image for
additional references. PNG, JPEG, and PPM output are supported. CMF_GPU=0
forces the portable CPU path. The documented profile keeps
--reference-size 1024: this is the square root of the VAE reference area and
is independent of the requested output dimensions. On a capable Vulkan device,
cortiq imagine automatically selects the resident Qwen transformer forward:
hidden state stays on the device across transformer blocks and is read back once
at the end of each forward. CPU and Metal fallback paths remain available, and
the memory-budgeted admission chooses a compatible path when needed. See the
Qwen Image guide for pinned acquisition, standalone staging,
and bundle packing commands.
The standalone layout remains available for explicit component selection:
For pinned sources and component packing, see the Qwen Image guide.
The CLI also exposes info, bench, ppl, serve, skill, moe-mask,
moe-defrag, requant, compact, sign, imagine, animate and
ltx-video. Run cortiq <command> --help for flags and current limitations.
What a CMF file contains
- A fixed 128-byte envelope addressing every section.
- Header JSON with architecture, quantization defaults, chat metadata and provenance.
- A binary tensor directory with dtype, shape, offsets, lengths and
hash64. - A page-aligned weight blob that can be memory-mapped and read in place.
- Optional task masks, skill replacement tensors, tokenizer bytes and sparse indexes.
cortiq verify fails on malformed bounds or a hash mismatch; python/cmf_reader.py
is a small independent reader for inspection. The normative layout is in the
CMF v2 specification.
Quantization
Quantization is selected per tensor, so sensitive tensors can stay at a higher precision while large matrix blocks use a compact codec.
| Codec | Typical use |
|---|---|
f16, f32 |
norms, embeddings and exact control tensors |
q8, q8_2f |
high-fidelity weights; q8_2f adds input-channel scales |
q4, q4t, q4tp |
general dense and MoE weights |
q2tp |
existing dtype-16 midrise four-level codec; zero is not an individual code |
q2tp_affine |
dtype-16 Prism profile with explicit (c - 1) * s affine operator; requires signed-Hadamard and affine metadata |
vbit, vbit_ro |
size-constrained or mixed-bit profiles |
q1, q1t, q1s |
trained binary or experimental ternary/PTQ paths |
See Q1T/PTQ for the experimental low-bit path and the quantization coverage matrix for codec support by execution path.
The q2tp_affine profile is an explicit Prism/Bonsai operator, not a
reinterpretation of ordinary q2tp: its signed Hadamard descriptor and affine
feature metadata are required at load time, and older readers must reject it.
It preserves the public q2tp codec and its existing semantics unchanged.
Runtime capabilities
- CPU inference: portable Rust implementation with memory-mapped weights.
- GPU inference: native Metal on macOS and wgpu backends (Vulkan/DX12)
where the device and build support them. Use
CMF_GPU=1to request GPU. - Long context: optional
--o1attention uses fixed-size state and trades memory growth for a measured quality delta; measure on your model. - Skills: one backbone can carry task masks and replacement-tensor overlays; inactive overlays do not need a second full model copy.
- MoE: task masks and physical defragmentation can reduce the active expert set. Validate perplexity on held-out data before shipping a restriction.
- Speculative decode: supported MTP/draft paths are enabled only where the model metadata and measured acceptance make them useful.
- Serving:
cortiq serveprovides an OpenAI-compatible local HTTP API. - Media:
imagine,animateandltx-videopack model assets and run image, video or audio-capable pipelines when their model guides apply.
Platforms and model families
| Target | Support |
|---|---|
| Linux/macOS/Windows | CPU builds; GPU through the available wgpu backend |
| Apple Silicon | CPU and Metal paths; memory is shared with the system |
| NVIDIA/AMD GPUs | Vulkan or DX12 through wgpu where tested |
| Android/iOS integrations | See the companion mobile project and split guide |
Native conversion covers Qwen, Llama, Mistral, Gemma, Phi, DeepSeek, Kimi
Linear, MiniCPM and several MoE/video families. Model-specific constraints and
published files are indexed in the model guides. For an
unsupported checkpoint, try import-gguf and file a reproducible issue if it
fails.
Focused documentation
- CMF v2 specification — normative file layout.
- Format comparison — criteria and evidence.
- Skills — masks, overlays and routing workflows.
- GPU kernel recipes — backend measurements.
- Multi-GPU execution and mobile split.
- FCD restoration and low-bit PTQ.
- Qwen Image Edit — one-file native image editing and standalone component staging.
- Model cards and conversion notes.
- Cortiq Spectra — deterministic CPU streaming colorization for dual-energy X-ray captures, scanner profiles, refusal masks, and measured limits (Russian).
Build from source
The workspace is Apache-2.0. Read LICENSE, PATENTS.md, CONTRIBUTING.md and SECURITY.md before redistributing or reporting a vulnerability. Releases and checksums are listed on the GitHub releases page.