Expand description
Vision support — screenshot capture/downscale/encode for visual asserts. Vision support — screenshot capture, tiling, downscaling, and JPEG encoding for LLM visual assertions.
When an assert step sets screenshot = true, the runner captures the
full scrollable page (via the CDP captureBeyondViewport flag —
nothing below the fold is skipped) and splits it into viewport-tall
bands (“tiles”) from the top, covering at most
scenario::ScenarioConfig::screenshot_max_height (default "20x" =
four viewports). Each tile is downscaled independently so its longest
edge is at most [default_max_dimension] (the page’s [config] screenshot_max_dimension wins when set), JPEG-encoded, and sent as its
own OpenAI-compatible image_url content part next to the text prompt —
detail is preserved at every depth, and the tile count (hence token
cost) is bounded by the height cap and [max_tiles].
Downscaling and cropping happen in Rust (no page JS, no fragile
canvas evaluation): the PNG bytes are decoded with the image crate,
banded, resized with Lanczos filtering, and re-encoded as quality-85
JPEG in a single capture call.
Constants§
- DEFAULT_
MAX_ DIMENSION - Default longest edge (px) of screenshots sent to vision endpoints.
- MAX_
TILES - Hard ceiling on the number of tiles attached to a single vision assert,
so a huge
screenshot_max_heightcan never produce an unbounded request (or token bill).
Functions§
- capture_
screenshot_ data_ urls - Captures the full page and returns one JPEG data URL per viewport-tall
band, ordered from the top of the page down — suitable for the
OpenAI-compatible
image_urlcontent-part array. - tile_
count - Number of viewport-tall bands covering
coveragepixels, clamped to [max_tiles] so a single assert can never send an unbounded number of image parts.