Skip to main content

Crate llm_token_visualizer

Crate llm_token_visualizer 

Source
Expand description

Flag the words an LLM answer was unsure about, from the token log probabilities (logprobs) that OpenAI-compatible APIs return.

Each token’s probability is exp(logprob). A word is flagged when any of its tokens with letters or digits falls below the threshold, neighbouring flagged words merge into one span, and each span keeps the alternatives the model was weighing at its weakest token.

use llm_token_visualizer::detect::{detect, parse_logprobs};

// Three real tokens from a Llama 3.1 8B answer ("... in Dordrecht").
let json = r#"[
  {"token": " D", "logprob": -0.0153},
  {"token": "ord", "logprob": -0.5620, "top_logprobs": [
    {"token": "ord", "logprob": -0.5620},
    {"token": "üsseldorf", "logprob": -0.9370},
    {"token": "elf", "logprob": -3.5620}]},
  {"token": "recht", "logprob": -0.0004}
]"#;

let tokens = parse_logprobs(json)?;
let report = detect(&tokens, 0.6);

assert_eq!(report.spans.len(), 1);
assert_eq!(report.spans[0].text, " Dordrecht");
assert_eq!(
    report.spans[0].describe(),
    r#"p=0.57 at "ord"; model also considered "üsseldorf" (0.39), "elf" (0.03)"#
);

report renders the same result as a terminal heatmap, a self-contained HTML page or Markdown. The command-line tool is llm-token-visualizer (cargo install llm-token-visualizer).

Re-exports§

pub use data::ConfidenceLevel;
pub use data::FlagType;
pub use data::TokenAnalysis;
pub use data::TokenFlag;
pub use data::TokenInfo;
pub use data::VisualizationConfig;
pub use renderer::HtmlRenderer;
pub use renderer::MarkdownRenderer;
pub use renderer::Renderer;
pub use renderer::TerminalRenderer;
pub use utils::create_mock_analysis;
pub use utils::detect_issues;
pub use utils::simple_tokenize;
pub use utils::AnalysisMetrics;

Modules§

data
detect
Hallucination-risk detection from token log probabilities.
live
Fetch a completion with token logprobs from an OpenAI-compatible API.
renderer
report
Reports for detect mode: a confidence heatmap of the answer, the flagged spans, and the alternatives the model weighed at each weak token.
utils

Functions§

analyze_with_issues
Comprehensive analysis with issue detection
quick_analyze
Quick analysis function for testing - creates mock data and visualizes
visualize_tokens
Main visualization function that can be used by other applications