Skip to main content

Module vision

Module vision 

Source
Expand description

Vision support — screenshot capture/downscale/encode for visual asserts. Vision support — screenshot capture, tiling, downscaling, and JPEG encoding for LLM visual assertions.

When an assert step sets screenshot = true, the runner captures the full scrollable page (via the CDP captureBeyondViewport flag — nothing below the fold is skipped) and splits it into viewport-tall bands (“tiles”) from the top, covering at most scenario::ScenarioConfig::screenshot_max_height (default "20x" = four viewports). Each tile is downscaled independently so its longest edge is at most [default_max_dimension] (the page’s [config] screenshot_max_dimension wins when set), JPEG-encoded, and sent as its own OpenAI-compatible image_url content part next to the text prompt — detail is preserved at every depth, and the tile count (hence token cost) is bounded by the height cap and [max_tiles].

Downscaling and cropping happen in Rust (no page JS, no fragile canvas evaluation): the PNG bytes are decoded with the image crate, banded, resized with Lanczos filtering, and re-encoded as quality-85 JPEG in a single capture call.

Constants§

DEFAULT_MAX_DIMENSION
Default longest edge (px) of screenshots sent to vision endpoints.
MAX_TILES
Hard ceiling on the number of tiles attached to a single vision assert, so a huge screenshot_max_height can never produce an unbounded request (or token bill).

Functions§

capture_screenshot_data_urls
Captures the full page and returns one JPEG data URL per viewport-tall band, ordered from the top of the page down — suitable for the OpenAI-compatible image_url content-part array.
tile_count
Number of viewport-tall bands covering coverage pixels, clamped to [max_tiles] so a single assert can never send an unbounded number of image parts.