auv-daemon 0.0.28

Server-side SDK for hosting an AUV daemon
docs.rs failed to build auv-daemon-0.0.28
Please check the build logs for more information.
See Builds for ideas on how to fix a failed build, or Metadata for how to configure docs.rs builds.
If you believe this is docs.rs' fault, open an issue.

License

AUV means Application Use Via ....

Think of it as a programmable computer use, without agents.

Table of Contents

Getting Started

Install

Install a prebuilt release:

macOS

brew install moeru-ai/tap/auv
auv --version

Alternatively, without Homebrew:

curl -fsSL https://raw.githubusercontent.com/moeru-ai/auv/main/install/install.sh | sh

Linux

curl -fsSL https://raw.githubusercontent.com/moeru-ai/auv/main/install/install.sh | sh
auv --version

[!NOTE]

Set AUV_VERSION or AUV_INSTALL_DIR to change the release version or the install directory (default: ~/.local/bin).

Windows

Scoop
scoop bucket add auv https://github.com/moeru-ai/auv
scoop install auv/auv
auv --version
Manual installation

Download the archive for your architecture:

Extract the archive to a permanent directory. Add that directory to your user PATH. The archive contains a single auv.exe; the Windows helper is embedded.

Install with proto

Install and configure proto first. Then add the AUV plugin and install the latest release:

proto plugin add auv "https://raw.githubusercontent.com/moeru-ai/auv/main/toolchain/proto/auv.toml" --to global
proto install auv latest --config-mode global --pin global
auv --version

[!NOTE]

AUV Helper.app for macOS is included in the proto installation. On Windows, auv-helper.exe is embedded in auv.exe and extracted only by the elevated helper setup command.

[!WARNING]

Linux musl is not supported. (But PRs are welcomed!)

Install with Nix

Install Nix 2.27 or later and enable the nix-command and flakes experimental features. On macOS, install Apple's build tools first:

xcode-select --install

Then install the default AUV package from this repository:

nix profile install 'git+https://github.com/moeru-ai/auv#default'
auv --version

The git+https transport is required so Nix fetches AUV's Git submodules. The flake defines source-built packages for Apple Silicon and Intel macOS and for x86-64 and ARM64 Linux. The package does not support Windows or Linux musl.

The Nix package does not embed the signed AUV Helper.app. On macOS, use Homebrew, proto, or a direct release download if you need to run auv setup macos-helper install with the official helper.

Install with Cargo

Prerequisites: Rust and the platform build tools below. AUV includes the Protobuf sources, so Buf is not required.

[!WARNING] cargo install does not include AUV Helper. Use an install method above if you need it.

macOS

Install the Xcode Command Line Tools:

xcode-select --install
cargo install --git https://github.com/moeru-ai/auv auv-cli --bin auv
auv --version

Linux

On Ubuntu or Debian, install the native build dependencies:

sudo apt-get update
sudo apt-get install -y \
  pkg-config libclang-dev libxcb1-dev libxrandr-dev libdbus-1-dev \
  libpipewire-0.3-dev libwayland-dev libxkbcommon-dev libegl-dev \
  libleptonica-dev libtesseract-dev
cargo install --git https://github.com/moeru-ai/auv auv-cli --bin auv
auv --version

[!NOTE] Other Linux distributions can use different package names.

Windows

Install Rust with the MSVC toolchain, Visual Studio Build Tools, and the Windows SDK.

cargo install --git https://github.com/moeru-ai/auv auv-cli --bin auv
auv --version

Setup

macOS

Official macOS releases include the signed AUV Helper.app. On macOS 13 or later, install it for the current user:

auv setup macos-helper install
auv setup macos-helper status

[!TIP] The installation does not require sudo or an administrator password.

If macOS requests approval, open the Background Items and Accessibility settings:

auv setup macos-helper open-background-items-settings
auv setup macos-helper open-accessibility-settings

[!NOTE] Projects that integrate AUV can rebrand AUV Helper.app. They can change its name, icon, bundle identifier, and Apple Developer signing identity. See Shipped helper identity for packaging options.

Grant these permissions to the application that starts AUV, usually your terminal application:

Permission Needed for
Accessibility AX tree reads, focused element control, keyboard/pointer automation.
Screen Recording Screenshots, OCR, visual inspection, and evidence capture.
Automation AppleScript/System Events app activation and foreground fallback paths.

After you change the permissions, restart the terminal. Then run:

auv doctor
auv invoke app.probePermissions

Windows

[!IMPORTANT] The Windows setup commands require an elevated PowerShell.

auv setup windows-helper install
auv setup windows-helper status

[!NOTE] The setup command installs auv.exe and extracts its embedded auv-helper.exe into %ProgramFiles%\AUV. The helper is not a standalone command and does not need to be downloaded or placed beside auv.exe.

Uninstall

Use the instructions that match your installation method. If a platform Helper is installed, remove it first.

macOS

auv setup macos-helper uninstall
brew uninstall auv

[!NOTE] Helper removal keeps the enrollment data in the login Keychain. It also keeps other AUV data in the Application Support directory.

Linux

rm "$HOME/.local/bin/auv"

Windows

Run these commands from an elevated PowerShell:

auv setup windows-helper uninstall
scoop uninstall auv

[!NOTE] Helper removal keeps the device data in %ProgramData%.

If you installed the ZIP manually, remove its directory from the file system. Then remove that directory from PATH.

Cargo

cargo uninstall auv-cli

Understand AUV

For Cua, agent-browser, and similar computer-use projects, it is common to execute screenshot, read image, click, type, wait, and follow-up verification steps in sequence, then ask LLMs or agents to judge the next move.

flowchart LR
  A[Agent] --> B[screenshot]
  B --> C[read image]
  C --> D[decide next step]
  D --> E[click]
  E --> F[wait]
  F --> G[type]
  G --> H[verify]
  H --> D

Many of those repeated sequences can be squashed into reusable GUI operations. Opening an app, waiting for readiness, filling a form, and checking the result should be callable as one command instead of spending tokens on the same step-by-step loop every time.

Modern agents often use skills or project instructions to orchestrate tool calls, CLIs, and scripts. But built-in computer-use surfaces, such as OpenAI Computer Use or Claude Computer Use, are still primarily interactive model-tool loops, not scriptable GUI automation libraries.

Similar to Playwright, what if we could organize those actions into executable scripts, reusable?

• Ran screenshot
  └ saved screen.png
• Ran read image screen.png
  └ form is visible
• Ran click "Email"
  └ clicked
• Ran type "user@example.com"
  └ typed
• Ran screenshot
  └ saved after.png
• Ran verify form state
  └ ready
pub fn open_and_fill_form(
  app: &mut AppSession,
  data: FormData,
) -> AuvResult<OperationResult> {
  app.open()?;
  app.wait_for_ready()?;
  app.fill(data)?;
  app.verify_submitted()
}
• Ran screenshot
  └ saved page-1.png
• Ran OCR visible rows
  └ 12 rows
• Ran scroll
  └ scrolled down
• Ran OCR visible rows
  └ 10 rows, 4 repeated
• Ran guess when to stop
  └ uncertain
pub fn scan_visible_rows(
  region: &mut WindowRegion,
) -> AuvResult<ScrollScanArtifact> {
  region.scan_rows_until_stop()
}
• Ran click target
  └ clicked
• Ran screenshot
  └ saved after-click.png
• Ran semantic check
  └ mismatch
• Ran retry manually
  └ repeated tool loop
pub fn verify_and_retry<F>(
  mut operation: F,
) -> AuvResult<OperationResult>
where
  F: FnMut() -> AuvResult<OperationResult>,
{
  retry_until_verified(&mut operation)
}

AUV expects agents to write, test, and improve reusable GUI automation for E2E tests and rapid application actions.

In fact, AUV is not a computer-use agent. It does not ship an agent or harness. It offers tools, CLIs, drivers, and verifiable observable results so agents can build reusable GUI operations.

AUV is meant to work with coding agents and agent products such as:

That means:

  • If your agent can call a CLI, AUV can be used as computer use.
  • If your agent can write code, AUV can move repeated GUI work into reusable Rust or JavaScript/TypeScript operations. Once a GUI flow is finalized as an operation, repeated execution can approach zero reasoning-token cost.
  • AUV's daemon and extension APIs use versioned Protobuf/gRPC contracts. A language with compatible Protobuf/gRPC generators can generate a client for those contracts without AUV inventing another language-specific protocol. First-party SDK quality, packaging, and documentation are still separate support claims: Rust and JavaScript/TypeScript are available today, while a first-party Python SDK remains planned.

The reusable pieces are split by responsibility, but they use one execution model instead of becoming unrelated wrappers:

flowchart LR
  A[CLI / MCP / Rust / JS / generated clients] --> B[typed operation]
  B --> C[local or remote Device / Runner]
  C --> D[capability Driver]
  D --> E[direct result]
  D --> F[Run trace and artifacts]
  E --> G[separate semantic verification]

Drivers own platform capabilities, operation crates own reusable workflows, and auv-tracing owns Run evidence and artifacts. The visual overlay remains a separate trust and debugging surface; drawing a cursor never stands in for input delivery or semantic verification. This package structure lets another frontend or generated language client reuse the same operations rather than reimplementing them around the CLI.

Why even build AUV?

AUV born from the grounding knowledge of building general gaming agents for Project AIRI, since 2024, we tried to build agents to allow LLMs to play the following games, you can find how we implement the agents in the following repos:

There are more games we implemented where you can find in Project AIRI organization, but these four requires YOLO, OCR, screen understanding, and computer-use capabilities.

Now you have the framework to build for any applications, games.

Since Vercel published the agent-browser, we fell in love with it and have it assisted agents to build many web projects, but we found that the loop it requires for agents to call agent-browser CLI to execute the commands is too slow and inefficient, while in computer use world, many operations can be repeated thousands of times, just like how Playwright/Vitest would allow us to write E2E test for applications, why don't we expand this idea of writing code to control application to computer use world?

Capability Matrix

What AUV can do, compared to other computer-use projects.

  • ✅: yes.
  • ❌: no.
  • ⚠️: partial support. The cell states the limit.
  • ⏳: planned.
  • —: not assessed.

Platform support comes from the Native desktop drivers row. Other rows name a platform only when their support is different.

Capability AUV Cua @oai/sky[^sky]bundled OpenBridge (KWWK core) Playwright
Agent model 💡 BYOA 💡 BYOA 💡 agent-free API 💡 OpenBridge built-in agentKWWK is agent-free 💡 BYOA + built-in Test Agents
Language-agnostic API ✅ Protobuf/gRPC ✅ HTTP/WebSocket ❌ ❌ ❌
Scriptable (Rust) ✅ ✅ ❌ ❌ ❌
Scriptable (TypeScript) ✅ ✅ ✅ ❌ ✅
Scriptable (Python) ⏳ first-party SDK ✅ ❌ ❌ ✅
Native desktop drivers ✅ macOS/Linux/Windows⏳ Android/iOS ✅ macOS/Linux/Windows ✅ macOS/Linux/Windows ✅ macOS❌ Linux/Windows ❌ browser only
CLI ✅ ✅ ❌ ❌ ✅
MCP ✅ ✅ ❌ ❌ ✅ browser MCP
REPL / Codemode ⏳ planned ❌ ✅ Node REPL ❌ ❌
Screen Lock/Unlock ✅[^device-entry] ❌ ❌ ❌ ❌
Trace ✅ Runs, artifacts, OpenTelemetry ✅ trajectories ❌ ❌ ✅ test traces
Screenshot ✅ ✅ ✅ ✅ ✅
OCR ✅ macOS Vision/Linux Tesseract/Windows OCR ⚠️ requires an external model key ❌ ❌ ❌
Template Matching ❌ locator✅ result contract ❌ ❌ ❌ ❌
Accessibility tree ✅ ✅ ✅ ✅ ✅
Accessibility actions ⚠️ focus and selection ✅ ✅ ✅ ✅
Mouse Click ✅ ✅ ✅ ✅ ✅
Mouse Move ✅ ✅ ✅ Linux❌ macOS/Windows — ✅
Background pointer input ✅ macOS❌ Linux/Windows ⚠️ some apps require foreground ✅ Linux window target❌ macOS/Windows ✅ ✅ browser context
Foreground pointer input ✅ ✅ ✅ ✅ ✅
Keyboard Hold ✅ ✅ ✅ Linux timed hold❌ macOS/Windows — ✅
Keyboard Input ✅ ✅ ✅ ✅ ✅
Scroll ✅ ✅ ✅ ✅ ✅
Ghost Cursor ✅ macOS: multiple named cursors[^ghost-cursor]❌ Linux/Windows ⚠️ one agent cursor ❌ ❌ ❌
Customizable Cursor ✅ macOS: colors, SVG, shadow⚠️ Windows: colors only❌ Linux ❌ ❌ ❌ ❌
Scroll-to-list ✅ library and app integrations❌ generic CLI ❌ ❌ ❌ ✅ browser lists❌ desktop lists
Feedback ✅ attempts, fallback, disturbance, verification ✅ outputs and trajectories ⚠️ state read after action ⚠️ metadata only ⚠️ assertions and traces
YOLO / Custom Models ✅ ✅ ❌ ❌ ❌
  • Scroll scan is a major reason AUV exists. Most desktop automation stacks can scroll and capture a screenshot. They do not make page records, row candidates, crop artifacts, OCR fragments, or clear stop reasons. The current scroll-scan implementation is contract work. The old scan window-region CLI will return when the reusable API is clear.
  • Feedback is machine-readable evidence for an action. It records the input path, changes, artifacts, fallbacks, and verification result. This evidence tells an operation when to retry, stop, or fail.

[^device-entry]: Evidence level: configuration-specific installed-host test. This API locks and unlocks an existing login session. It does not sign in a user from the signed-out screen. The 2026-10-01 test ran 300 normal-use lock/unlock cycles. The API passed 298 cycles on the first attempt (99.33%). The requested OS state occurred on the first attempt in 299 cycles (99.67%). All 300 cycles ended in the USABLE state. The test used dwell times of 15, 20, 25, and 30 seconds. A separate stress test used delays near zero. It measured OS transition readiness, not normal-use reliability. Read the Device lock contract and platform evidence for the typed contract, native mechanisms, and configuration limits. The raw logs remain local. They are not in a durable evidence pack.

[^ghost-cursor]: AUV does not define a numeric cursor limit. Host memory and WindowServer resources limit the actual count. Ghost cursors are visual overlays. They do not deliver input or prove an action result.

[^sky]: Evidence level: installed package documentation and TypeScript declarations. The inspected package is @oai/sky 0.7.1 from the ChatGPT app. It is not available from the public npm registry. No native execution or native binary inspection supports this column. See the local Sky API research and the background-delivery comparison.

Development

auv

cargo fmt --check
cargo check
cargo test

To update vendored Protobuf dependencies, see the Protobuf source distribution reference.

@auv-js/sdk

Prerequisites

[!NOTE]

If you use proto, then

proto install buf
proto install node
proto install pnpm

, this should help you install necessary tools.

pnpm install
pnpm generate:proto
pnpm exec playwright install chromium
pnpm build
pnpm test:run
pnpm lint
pnpm typecheck

Documentation

After you change headings in the root or package READMEs, run pnpm docs:update. This command updates all three tables of contents.

Useful entrypoints:

auv doctor
auv invoke <command-id> --help
auv serve --help
auv devices list
auv runner --help
auv run --help
auv mcp serve
auv plugin list

Use docs/TERMS_AND_CONCEPTS.md for shared vocabulary. Durable design and evidence notes live under docs/ai/references/.

Related

[!NOTE]

This project is part of the Project AIRI ecosystem.

Acknowledgements

Special Thanks

Special thanks to all contributors for their contributions to auv ❤️

Star History

License

Apache License 2.0