Overview
agent-otel-bridge is a high-performance, native OpenTelemetry sidecar bridge designed to solve the hot-path lifecycle hook latency bottleneck in autonomous AI CLI agent harnesses (Google Antigravity, Claude Code, OpenAI Codex CLI, xAI Grok, Inflection Pi, and custom LLM developer agents).
Modern AI coding agents invoke synchronous lifecycle hooks (PostToolUse, PreInvocation, PostInvocation, Stop) dozens to hundreds of times per session. Traditional script-based hooks (Python, Node.js, PowerShell) impose 150ms to 1,200ms of cold-start latency per tool call, accumulating up to several minutes of wasted developer time in a single pairing session.
agent-otel-bridge decouples hot-path event emission into:
agent-hook.exe: An ultra-fast, microscopic static PE32 binary (242 KB) that executes in < 1ms, drops the event into a local Win32 Named Pipe via Overlapped I/O, and enforces a 3ms watchdog fail-open guarantee (never blocks or breaks the agent's workflow).agent-otel-bridge daemon: A Tokio async background daemon maintaining connection pooling with the OpenTelemetry Collector (localhost:4318), performing micro-batching (50 spans / 200ms), and serializing canonical OTLP Protobuf traces and quota metrics without requiringprotocbuild-time dependencies.
Architecture
flowchart TD
subgraph AgentHarness ["AI CLI Agent Harnesses (Antigravity / Claude / Codex / Grok / Pi)"]
HookEvent["Synchronous Lifecycle Hook (e.g. PostToolUse)"]
end
subgraph ClientPath ["Ultra-Fast Client (< 1ms execution)"]
HookBin["agent-hook.exe (242 KB Native PE32)"]
Watchdog["Hard Watchdog Thread (3ms Deadline)"]
PipeClient["Win32 Overlapped Client (\\.\\pipe\\agent-otel)"]
end
subgraph DaemonProcess ["agent-otel-bridge daemon (Tokio Background Process)"]
PipeServer["Named Pipe Server (64KB Ring Buffer)"]
Parser["ProtoJSON Parser & SemConv Mapper"]
Batcher["Micro-Batcher (Configurable Buffer)"]
QuotaEngine["Quota & Heartbeat Engine"]
OtlpExporter["OTLP HTTP/Protobuf Exporter (Keep-Alive Pool)"]
end
subgraph OTelPipeline ["OpenTelemetry Collector Pipeline"]
Collector["OTel Collector (:4318 / :4317)"]
Processor["Batch / Transform Processor (OTTL)"]
end
subgraph ObservabilityBackends ["Pluggable Observability Ecosystem"]
Traces["Distributed Traces (Jaeger / Tempo / SigNoz / Datadog)"]
Metrics["Metrics & Gauges (Prometheus / Mimir / VictoriaMetrics)"]
Dashboards["Dashboards (Grafana / SigNoz / Honeycomb)"]
end
HookEvent -->|"stdin (ProtoJSON)"| HookBin
HookBin -->|"Spawn watchdog"| Watchdog
HookBin -->|"Non-blocking write"| PipeClient
HookBin -->|"stdout: {} & exit 0"| HookEvent
PipeClient -->|"Binary Wire Framing [AG:v1]"| PipeServer
PipeServer -->|"Tokio mpsc channel"| Parser
Parser --> Batcher
QuotaEngine -->|"Metrics"| OtlpExporter
Batcher -->|"Traces"| OtlpExporter
OtlpExporter -->|"POST /v1/traces (Protobuf)"| Collector
OtlpExporter -->|"POST /v1/metrics (Protobuf)"| Collector
Collector --> Processor
Processor --> Traces
Processor --> Metrics
Processor --> Dashboards
Benchmarks
Measured on Windows 11 (AMD Ryzen / NVMe) via the dedicated benchmark suite (agent-otel-bench):
| Measurement Stage | p50 (median) | p90 | p95 | p99 | p99.9 | SLA Target | Status |
|---|---|---|---|---|---|---|---|
| Win32 Named Pipe RTT | 22 µs | 38 µs | 57 µs | 101 µs | 277 µs | < 3,000 µs | ✅ PASS |
| ProtoJSON Parse + Span Build | 1 µs | 1 µs | 2 µs | 2 µs | 5 µs | < 3,000 µs | ✅ PASS |
| Real OS Process Lifecycle | 3.56 ms | 4.25 ms | 4.70 ms | 5.59 ms | 8.36 ms | < 10.0 ms | ✅ PASS |
- OTLP Span Processing Throughput: > 546,000 spans/second.
- Hook Cost across 200 Tool Invocations:
- PowerShell Hook Script: 90s – 240s lost.
- Python Hook Script: 30s – 56s lost.
agent-hook.exe(Rust): 0.7s total accumulated time.
- Full methodology and data in
docs/BENCHMARKS.md.
OpenTelemetry Semantic Conventions
Fully aligned with canonical CNCF OpenTelemetry GenAI & Agent Semantic Conventions (v1.28+ / v1.36):
Spans & Canonical Attributes
- Span Names:
execute_tool {gen_ai.tool.name}(SpanKind::Internal)invoke_agent {gen_ai.agent.name}(SpanKind::Internal)agent.stop(SpanKind::Internal)
- GenAI Attributes:
gen_ai.operation.name:execute_tool,invoke_agentgen_ai.provider.name: dynamically inferred ("anthropic","google","openai","xai","inflection")gen_ai.agent.name: dynamically inferred ("antigravity","claude-code","codex","grok","pi","ai-agent")gen_ai.conversation.id: correlated session identifier (mapped fromconversationIdor Claude Codesession_id)gen_ai.tool.name&gen_ai.tool.call.idgen_ai.request.model: e.g.claude-3-5-sonnet,gemini-2.5-pro,gpt-4o,grok-2
- Canonical Agent Attributes:
agent.hook.event:PostToolUse,PreToolUse,PostInvocation,PreInvocation,Stopagent.step.index&gen_ai.agent.step_index: integer turn indexagent.execution.num: execution sequence numberagent.fully_idle: boolean quiescence indicatoragent.termination_reason: stop / completion reason (e.g.NO_TOOL_CALL,model_stop)
- Agent Quota Gauges:
agent.quota.remaining_fraction: Gauge (unit"1",0.0..=1.0).agent.quota.seconds_to_reset: Gauge (unit"s",>= 0.0).- Attributes:
bucket,group.
- Universal Agent Intelligence & Context (v0.3):
- Behavioral CLI Archetypes (Preference-Agnostic): Commands are categorized into 6 functional archetypes (
filter_compressor,structured_parser,inspector_diff,search_retrieval,build_test_verify,generic_exec) instead of hardcoding developer-specific tools (rtk,jq,bat). Tracks compression ratio, tokens saved, and pipeline depth. - Universal MCP & Skills Taxonomy + Waste Tracking: Standardizes tool calls across
mcpservers,skillbundles, andnativetools. Quantifies Schema Tax (dead-weight prompt tokens spent on inactive tools), retry thrashing, and response payload bloat. - Multi-Tier Zero-Subprocess Context Harvester (< 150 µs): Extracts workspace path, project root, project type (
rust,node,python,go), and direct file-based Git metadata (.git/HEAD, sanitized remote origin, worktrees) with zero external process calls (git.exeis never spawned). - Cross-Agent Distributed Tracing (W3C
traceparent): Injects and propagates W3C trace context via process environment variables ($env:TRACEPARENT). Heterogeneous subagents (e.g. Antigravity spawning Claude Code, which delegates to Codex) are unified into a single distributed trace DAG in SigNoz.
- Behavioral CLI Archetypes (Preference-Agnostic): Commands are categorized into 6 functional archetypes (
- OpenTelemetry Layering & Opt-In Compatibility:
- Emits purely canonical, vendor-neutral attributes by default so all AI agents (Antigravity, Claude Code, Codex, Grok, Pi) share a clean data model without vendor pollution.
- Optional
AGENT_OTEL_LEGACY_ATTRIBUTES=trueenables legacyagy.*aliases for older dashboards, or use an OpenTelemetry Collectortransformprocessor(OTTL) to alias attributes in the collection tier.
Crates in the Workspace
| Crate | Binary / Role | Description |
|---|---|---|
agent-otel-core |
Library | Domain models, SemConv constants, deterministic 16-byte trace_id / 8-byte span_id generators, OTLP Protobuf builders. |
agent-otel-ipc |
Library | Binary wire protocol (MAGIC: AG), Win32 Overlapped client (windows-sys), Tokio Named Pipe server (\\.\pipe\agent-otel). |
agent-otel-client |
agent-hook.exe |
242 KB static PE32 binary with 3ms watchdog fail-open backstop. |
agent-otel-daemon |
Library | Tokio micro-batcher, reqwest keep-alive exporter, quota ticker, 30-min idle shutdown. |
agent-otel-bridge |
agent-otel-bridge.exe |
Unified CLI (hook, daemon, doctor, hooks, install-hooks, emit-quota, stop). |
agent-otel-bench |
agent-otel-bench.exe |
High-precision microsecond benchmark suite. |
Installation & Setup
1. Install via Cargo or Pre-built Binaries
From crates.io:
cargo install agent-otel-bridge
cargo install agent-otel-client
Or download pre-compiled release archives directly from GitHub Releases.
Or build locally from repository:
cargo install --path crates/agent-otel-client --force
cargo install --path crates/agent-otel-cli --force
2. Automated Hook Installation (install-hooks)
agent-otel-bridge automatically detects and configures lifecycle hooks for your AI agents:
# Automatically detects installed agents (Antigravity, Claude Code, Codex, Grok, Pi) and registers hooks
agent-otel-bridge install-hooks
# Or target a specific client
agent-otel-bridge install-hooks --client antigravity
agent-otel-bridge install-hooks --client claude
agent-otel-bridge install-hooks --client codex
agent-otel-bridge install-hooks --client grok
agent-otel-bridge install-hooks --client pi
# Or register hooks for all supported clients
agent-otel-bridge install-hooks --client all
# Or install at project level (.claude/settings.json or .gemini/hooks.json, etc.)
agent-otel-bridge install-hooks --client claude --project
Check hook status anytime:
agent-otel-bridge hooks status
3. Verify Installation (doctor)
agent-otel-bridge doctor
Output:
=== agent-otel-bridge doctor ===
Diagnosing station telemetry pipeline & OpenTelemetry invariants
[1/5] Checking environment contract variables...
[ok] OTEL_EXPORTER_OTLP_ENDPOINT = http://127.0.0.1:4318
[info] OTEL_SERVICE_NAME is unset (using default 'agent-otel-bridge')
[ok] OTEL_RESOURCE_ATTRIBUTES = deployment.environment=homelab
[2/5] Checking named pipe IPC (\\.\pipe\agent-otel)...
[ok] Daemon is RUNNING and responding on named pipe!
[3/5] Checking OTLP Collector HTTP endpoint...
[ok] OTLP Collector reachable at http://127.0.0.1:4318/v1/traces
[4/5] Checking Observability UI reachability (http://localhost:8080)...
[ok] Observability UI reachable at http://localhost:8080
[5/5] Checking Client Hook Registrations...
agent-hook in PATH: [ok] present
Google Antigravity: [ok] registered (C:\Users\samue\.gemini\config\hooks.json)
Claude Code: [ok] registered (C:\Users\samue\.claude\settings.json)
OpenAI Codex: [info] not registered (C:\Users\samue\.codex\hooks.json)
xAI Grok: [info] not registered (C:\Users\samue\.grok\hooks.json)
Inflection Pi: [info] not registered (C:\Users\samue\.pi\hooks.json)
4. Run CLI Commands
# Start background daemon
agent-otel-bridge start
# Send manual quota probe to collector
agent-otel-bridge emit-quota --ping
# Gracefully stop daemon
agent-otel-bridge stop
# Run benchmarks and submit results to community leaderboard
cargo run --release -p agent-otel-bench -- --submit --open-browser
# Or via unified CLI:
agent-otel-bridge benchmark --submit --open-browser
Documentation
| Guide | Description |
|---|---|
| AI Agent Harness Integration | Setup & hook configs for Antigravity, Claude Code, Codex, Grok, Pi, and Custom Agents. |
| AI Agent Architectural Invariants & SLAs | Enforced design constraints, sub-millisecond SLAs, and AI agent operating instructions. |
| Contributing & Branching Policy | Branching strategy (feat/*, fix/*), PR workflow, and automated guardrail verification. |
| Configuration & Environment | Complete environment variable catalog and OTel Collector recipes (config.yaml). |
| Telemetry Taxonomy & Variable Dictionary | Complete captured attribute dictionary, 6-dimension taxonomy, and Codex dashboard evaluation. |
| OpenTelemetry Dashboard & Visualization Guide | Panel queries and visualization recipes across Grafana, SigNoz, Jaeger, and Prometheus. |
| SigNoz Contrib Dashboard Template | Ready-to-import agent-native JSON template (contrib/dashboards/signoz/ai-agent-observability.json). |
| Troubleshooting & Diagnostic Runbook | Step-by-step resolution for common issues and fail-open verification. |
| Performance Benchmark Report | Detailed microsecond measurement methodology and hardware SLA evaluation. |
| Community Benchmark Matrix | Living leaderboard comparing benchmark results across community hardware. |
| Benchmark Submission Guide | Instructions on running automated benchmarks and contributing hardware results. |
| Architecture & Sequence Diagrams | Zero-copy IPC wire protocol and Win32 Named Pipe architecture. |
| Agent Observability Skill | Built-in pair programming skill for inspecting and diagnosing telemetry pipelines. |
License
Copyright (c) 2026 Samuel Mota. Licensed under the Apache License, Version 2.0.