Skip to main content

Module vision

Module vision 

Source
Expand description

Vision support — screenshot capture/downscale/encode for visual asserts. Vision support — screenshot capture, downscaling, and JPEG encoding for LLM visual assertions.

When an assert step sets screenshot = true, the runner captures the current viewport as PNG, downscales it so its longest edge is at most [default_max_dimension] (the page’s [config] screenshot_max_dimension wins when set), and encodes it as a JPEG data URL. The image is sent to the vision endpoint as an OpenAI-compatible image_url content part alongside the text prompt.

Downscaling happens in Rust (no page JS, no fragile canvas evaluation): the PNG bytes are decoded with the image crate, resized with Lanczos filtering, and re-encoded as quality-85 JPEG.

Constants§

DEFAULT_MAX_DIMENSION
Default longest edge (px) of screenshots sent to vision endpoints.

Functions§

capture_screenshot_data_url
Captures the current viewport and returns a JPEG data URL suitable for the OpenAI-compatible image_url content part.