pub fn measure(
web: &Web,
entry: &ProviderEntry,
model: &str,
) -> Result<Measured, AskError>Expand description
Measures the window model really gets on this server, by sending prompts of growing size
and watching for the first one the server does not read whole.
This is the number the page shows beside what the model claims, and the two are often not the same: the endpoint a harness uses can pin the window far below the model’s own capacity and throw the front of every larger prompt away in silence. Probing stops at the first prompt that was cut, so a server with a small window is asked twice rather than four times.
Every probe carries its own opening, and run makes this measurement’s openings unlike any
other’s. That is not decoration. A server that remembers the front of a prompt it has already
read counts only what it had to read afresh, so probes that shared an opening were answered
with numbers far below the prompts they were sent — on the owner’s own server a window of
16 384 was reported as 9 512, and asking twice reported 4. Nothing shares an opening now, so
the count is of the prompt rather than of the part that was new.
§Errors
When the provider could not be reached, refused, or answered without a token count.