Expand description
Flag the words an LLM answer was unsure about, from the token log probabilities (logprobs) that OpenAI-compatible APIs return.
Each token’s probability is exp(logprob). A word is flagged when any of
its tokens with letters or digits falls below the threshold, neighbouring
flagged words merge into one span, and each span keeps the alternatives the
model was weighing at its weakest token.
use llm_token_visualizer::detect::{detect, parse_logprobs};
// Three real tokens from a Llama 3.1 8B answer ("... in Dordrecht").
let json = r#"[
{"token": " D", "logprob": -0.0153},
{"token": "ord", "logprob": -0.5620, "top_logprobs": [
{"token": "ord", "logprob": -0.5620},
{"token": "üsseldorf", "logprob": -0.9370},
{"token": "elf", "logprob": -3.5620}]},
{"token": "recht", "logprob": -0.0004}
]"#;
let tokens = parse_logprobs(json)?;
let report = detect(&tokens, 0.6);
assert_eq!(report.spans.len(), 1);
assert_eq!(report.spans[0].text, " Dordrecht");
assert_eq!(
report.spans[0].describe(),
r#"p=0.57 at "ord"; model also considered "üsseldorf" (0.39), "elf" (0.03)"#
);report renders the same result as a terminal heatmap, a self-contained
HTML page or Markdown. The command-line tool is llm-token-visualizer
(cargo install llm-token-visualizer).
Re-exports§
pub use data::ConfidenceLevel;pub use data::FlagType;pub use data::TokenAnalysis;pub use data::TokenFlag;pub use data::TokenInfo;pub use data::VisualizationConfig;pub use renderer::HtmlRenderer;pub use renderer::MarkdownRenderer;pub use renderer::Renderer;pub use renderer::TerminalRenderer;pub use utils::create_mock_analysis;Deprecated pub use utils::detect_issues;pub use utils::simple_tokenize;pub use utils::AnalysisMetrics;
Modules§
- data
- Input types for the visualizer renderers.
- detect
- Hallucination-risk detection from token log probabilities.
- interop
- Adapters for other crates, each behind its own optional feature.
- live
live - Fetch a completion with token logprobs from an OpenAI-compatible API.
- renderer
- Terminal, HTML and Markdown renderers for
TokenAnalysisinput (visualizer mode). - report
- Reports for detect mode: a confidence heatmap of the answer, the flagged spans, and the alternatives the model weighed at each weak token.
- utils
- Helpers for the visualizer: a simple tokenizer, metrics and issue detection.
Functions§
- analyze_
with_ issues - Comprehensive analysis with issue detection
- quick_
analyze Deprecated - Renders made-up scores (see
utils::create_mock_analysis); the confidence values are a hash of each word, not anything a model said. - visualize_
tokens - Main visualization function that can be used by other applications