Expand description
Vision support — screenshot capture/downscale/encode for visual asserts. Vision support — screenshot capture, downscaling, and JPEG encoding for LLM visual assertions.
When an assert step sets screenshot = true, the runner captures the
current viewport as PNG, downscales it so its longest edge is at most
[default_max_dimension] (the page’s [config] screenshot_max_dimension wins when set), and encodes it as a JPEG data
URL. The image is sent to the vision endpoint as an OpenAI-compatible
image_url content part alongside the text prompt.
Downscaling happens in Rust (no page JS, no fragile canvas evaluation):
the PNG bytes are decoded with the image crate, resized with Lanczos
filtering, and re-encoded as quality-85 JPEG.
Constants§
- DEFAULT_
MAX_ DIMENSION - Default longest edge (px) of screenshots sent to vision endpoints.
Functions§
- capture_
screenshot_ data_ url - Captures the current viewport and returns a JPEG data URL suitable for
the OpenAI-compatible
image_urlcontent part.