cuttlefish 0.0.2

Native tooling for agents: a local wasm runtime that runs delegated jobs against local models
Documentation

Cuttlefish VM

(API Docs)

Cuttlefish is native tooling for AI agents. A coding agent hands off a job — summarize this, classify that, extract these fields — and Cuttlefish runs it locally against a local model, returning a structured result.

Two things make that worth doing:

  • Your data stays on your machine. For jobs marked local-only, the calling agent passes paths, not contents. Cuttlefish reads the files itself, so proprietary source and personal data never enter a frontier model's context at all.
  • You stop paying frontier prices for grunt work. Bulk, repetitive, and mechanical subtasks don't need a trillion-parameter model. Offloading them keeps tokens and context for the work that does.

The tradeoff is honest and worth stating up front: you cannot run a frontier-class model on a laptop. Cuttlefish is for the large subset of agent work that doesn't need frontier reasoning — not a replacement for it.

How it works

A job is described by a .cuttlefish spec: which model, what it's allowed to touch, and which processing block runs it. The block compiles to WebAssembly and runs sandboxed inside the cuttlefishd daemon, which serves inference from a local model and streams results back.

spec summarize_docs = {
  description = "Use when the agent needs a summary of a local file
                  and content must not leave the machine.";
  model = Path "../models/qwen2.5-7b-instruct-q4_k_m.gguf";
  data_policy = Local_only;
  capabilities = [ Read "./docs" ];
  block = "../blocks/echo-summarize";
}

Three design decisions shape everything else:

The guest orchestrates; the host does the work. A block is a state machine the daemon drives — it returns commands (Infer, Open, Slice, Done) and the host executes them. Nothing blocks waiting on a callback, so cancellation is simply the host declining to take the next step, and every inference iteration is observable and metered.

Bulk data never enters guest memory. Blocks receive a handle and a length, then pull bounded windows. Guest memory stays flat whether the input is a README or a corpus.

Capabilities are deny-by-default and enforced twice. A block gets no filesystem or network access unless its spec grants it. The grant is checked when the spec compiles and again by the sandbox at runtime — the compile-time check is a convenience, not the security boundary.

Status

Early, but it runs end to end. A .cuttlefish spec drives a sandboxed wasm block through the daemon and returns a structured result, with capabilities enforced and output streamed:

$ cuttlefish run --spec summarize_docs --input '{"path": "examples/docs/a.txt"}'
{
  "result": { "path": "examples/docs/a.txt", "summary": "a stub summary" },
  "status": "completed",
  "usage": { "duration_ms": 196, "model": "stub", "tokens_in": 12, "tokens_out": 3 }
}

Inference is still a deterministic stub rather than llama.cpp — everything around it is real. Also still to come: the typed DSL with block signatures, multi-block pipelines, the model pool, the block registry, and the agent harness.

Design rationale lives in the code, not in a separate design document — each crate's module docs explain what it is responsible for and why it looks the way it does. Start at the API docs, or read lib.rs of crates/cuttlefish-host for the core of the system. See AGENTS.md for why the project is organized that way.

Development

The toolchain is pinned with Nix, which supplies the exact Rust (host and wasm32-unknown-unknown targets) and Python versions the project expects:

$ nix develop                              # drops you into a shell with everything
$ nix develop --command cargo test --workspace

To run the example job end to end:

$ nix develop --command bash -c '
    cargo build -p cf-block-echo-summarize --target wasm32-unknown-unknown &&
    cargo build -p cuttlefishd -p cuttlefish &&
    ./target/debug/cuttlefishd examples/summarize.cuttlefish \
      target/wasm32-unknown-unknown/debug/cf_block_echo_summarize.wasm /tmp/cf.sock &
    until [ -S /tmp/cf.sock ]; do sleep 0.1; done
    ./target/debug/cuttlefish run --socket /tmp/cf.sock --spec summarize_docs \
      --input "{\"path\": \"examples/docs/a.txt\"}"
  '

The daemon serves over a unix domain socket, so it is unix-only for now; the library crates are cross-platform and tested on Windows too.

Running cargo outside that shell will pick up whatever toolchain happens to be on your PATH, which is a reliable source of confusing errors. If you'd rather not use Nix, check flake.nix for the pinned versions and match them yourself.

Test coverage is measured with cargo-llvm-cov:

$ nix develop --command cargo llvm-cov --workspace          # summary
$ nix develop --command cargo llvm-cov --workspace --html   # browsable report

That is the same command CI runs. There is no coverage service and no upload token — CI publishes the resulting percentage as a small JSON file alongside the API docs, and the badge above renders from it. Coverage data never leaves the build.

Contributing

Contributions are welcome. Please read CONTRIBUTING.md for how the project is built and tested, and CODE_OF_CONDUCT.md for the standards expected of participants. Security issues should follow SECURITY.md rather than being filed as public issues.

Two conventions worth knowing before you start, both covered in AGENTS.md:

  • Documentation lives in the code, as rustdoc. There is no docs/ tree — it is gitignored. Explanations belong next to what they explain, where review catches them going stale.
  • Comments explain why, never what. Several of this codebase's choices exist to avoid failures that are invisible from reading the result. Preserve and extend those rather than tidying them away.

Prior art and acknowledgements

Cuttlefish stands on work that came before it:

  • rune (hotg-ai) — the "declarative spec compiles to a single portable wasm binary" model, typed dataflow pipelines, and versioned addressable processing blocks all come from rune. Cuttlefish is a successor in that spirit, retargeted from edge ML inference to agent job delegation.
  • llama.cpp — local model inference.
  • Wasmtime and the Bytecode Alliance — the WebAssembly runtime and the sandboxing model.
  • superpowers (obra) — the agent-harness patterns: discovery metadata that states when to use a tool rather than summarizing how it works, and a hard line between what is enforced by code and what is merely suggested in prose.

License

This project is licensed under either of

at your option.

It is recommended to always use cargo crev to verify the trustworthiness of each of your dependencies, including this one.

Contribution

Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.

The intent of this crate is to be free of soundness bugs. The developers will do their best to avoid them, and welcome help in analysing and fixing them.