Skip to main content

Crate llm_token_visualizer

Crate llm_token_visualizer 

Source
Expand description

Flag the words an LLM answer was unsure about, from the token log probabilities (logprobs) that OpenAI-compatible APIs return.

Each token’s probability is exp(logprob). A word is flagged when any of its tokens with letters or digits falls below the threshold, neighbouring flagged words merge into one span, and each span keeps the alternatives the model was weighing at its weakest token.

use llm_token_visualizer::detect::{detect, parse_logprobs};

// Three real tokens from a Llama 3.1 8B answer ("... in Dordrecht").
let json = r#"[
  {"token": " D", "logprob": -0.0153},
  {"token": "ord", "logprob": -0.5620, "top_logprobs": [
    {"token": "ord", "logprob": -0.5620},
    {"token": "üsseldorf", "logprob": -0.9370},
    {"token": "elf", "logprob": -3.5620}]},
  {"token": "recht", "logprob": -0.0004}
]"#;

let tokens = parse_logprobs(json)?;
let report = detect(&tokens, 0.6);

assert_eq!(report.spans.len(), 1);
assert_eq!(report.spans[0].text, " Dordrecht");
assert_eq!(
    report.spans[0].describe(),
    r#"p=0.57 at "ord"; model also considered "üsseldorf" (0.39), "elf" (0.03)"#
);

report renders the same result as a terminal heatmap, a self-contained HTML page or Markdown. The command-line tool is llm-token-visualizer (cargo install llm-token-visualizer).

Re-exports§

pub use data::ConfidenceLevel;
pub use data::FlagType;
pub use data::TokenAnalysis;
pub use data::TokenFlag;
pub use data::TokenInfo;
pub use data::VisualizationConfig;
pub use renderer::HtmlRenderer;
pub use renderer::MarkdownRenderer;
pub use renderer::Renderer;
pub use renderer::TerminalRenderer;
pub use utils::create_mock_analysis;Deprecated
pub use utils::detect_issues;
pub use utils::simple_tokenize;
pub use utils::AnalysisMetrics;

Modules§

data
Input types for the visualizer renderers.
detect
Hallucination-risk detection from token log probabilities.
interop
Adapters for other crates, each behind its own optional feature.
livelive
Fetch a completion with token logprobs from an OpenAI-compatible API.
renderer
Terminal, HTML and Markdown renderers for TokenAnalysis input (visualizer mode).
report
Reports for detect mode: a confidence heatmap of the answer, the flagged spans, and the alternatives the model weighed at each weak token.
utils
Helpers for the visualizer: a simple tokenizer, metrics and issue detection.

Functions§

analyze_with_issues
Comprehensive analysis with issue detection
quick_analyzeDeprecated
Renders made-up scores (see utils::create_mock_analysis); the confidence values are a hash of each word, not anything a model said.
visualize_tokens
Main visualization function that can be used by other applications