# Symbiont Agent Runtime API Reference
Complete API documentation for the Symbiont Agent Runtime System.
## Core Types
### Identifiers
```rust
// Unique identifiers for various entities
pub struct AgentId(Uuid);
pub struct TaskId(Uuid);
pub struct MessageId(Uuid);
pub struct RequestId(Uuid);
pub struct AuditId(Uuid);
pub struct SandboxId(Uuid);
pub struct SnapshotId(Uuid);
impl AgentId {
pub fn new() -> Self;
}
// Similar for all ID types
```
### Agent Types
```rust
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum AgentState {
Created,
Initializing,
Ready,
Running,
Suspended,
Waiting,
Completed,
Failed,
Terminating,
Terminated,
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum ExecutionMode {
Persistent,
Ephemeral,
Scheduled { interval: Duration },
EventDriven,
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum Priority {
Critical = 4,
High = 3,
Normal = 2,
Low = 1,
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum Capability {
FileSystem,
Network,
Database,
Custom(String),
}
pub struct AgentConfig {
pub id: AgentId,
pub name: String,
pub dsl_source: String,
pub execution_mode: ExecutionMode,
pub security_tier: SecurityTier,
pub resource_limits: ResourceLimits,
pub capabilities: Vec<Capability>,
pub policies: Vec<Policy>,
pub metadata: HashMap<String, String>,
pub priority: Priority,
}
pub struct AgentInstance {
pub id: AgentId,
pub config: AgentConfig,
pub state: AgentState,
pub created_at: SystemTime,
pub last_updated: SystemTime,
pub execution_count: u64,
pub error_count: u32,
pub restart_count: u32,
}
```
### Resource Types
```rust
pub struct ResourceLimits {
pub memory_mb: u64,
pub cpu_cores: f64,
pub disk_io_mbps: u64,
pub network_io_mbps: u64,
pub execution_timeout: Duration,
pub idle_timeout: Duration,
}
pub struct ResourceUsage {
pub memory_used: u64,
pub cpu_usage: f64,
pub disk_io_rate: u64,
pub network_io_rate: u64,
pub uptime: Duration,
}
pub struct ResourceAllocation {
pub agent_id: AgentId,
pub allocated_at: SystemTime,
pub limits: ResourceLimits,
pub current_usage: ResourceUsage,
}
```
### Security Types
```rust
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord)]
pub enum SecurityTier {
Tier1 = 1, // Docker
Tier2 = 2, // gVisor
Tier3 = 3, // Firecracker
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum IsolationLevel {
None,
Low,
Medium,
High,
Maximum,
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum EncryptionAlgorithm {
Aes256Gcm,
ChaCha20Poly1305,
}
pub struct SecurityContext {
pub tier: SecurityTier,
pub isolation_level: IsolationLevel,
pub encryption_algorithm: EncryptionAlgorithm,
pub signing_key: Option<Vec<u8>>,
pub encryption_key: Option<Vec<u8>>,
}
```
### Communication Types
```rust
pub struct Message {
pub id: MessageId,
pub from: AgentId,
pub to: AgentId,
pub topic: String,
pub payload: Vec<u8>,
pub priority: Priority,
pub ttl: Duration,
}
pub struct SecureMessage {
pub message: Message,
pub signature: Vec<u8>,
pub encrypted_payload: Vec<u8>,
pub timestamp: SystemTime,
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum DeliveryStatus {
Pending,
Delivered,
Failed,
Expired,
}
```
### Error Types
```rust
#[derive(Debug, Clone)]
pub enum RuntimeError {
Resource(ResourceError),
Communication(CommunicationError),
Security(SecurityError),
Scheduler(SchedulerError),
Lifecycle(LifecycleError),
ErrorHandler(ErrorHandlerError),
Configuration(ConfigurationError),
Policy(PolicyError),
Sandbox(SandboxError),
Audit(AuditError),
Internal(String),
}
#[derive(Debug, Clone)]
pub enum ResourceError {
InsufficientResources { requirements: String },
AllocationFailed { agent_id: AgentId },
DeallocationFailed { agent_id: AgentId },
UsageExceeded { agent_id: AgentId, resource: String },
MonitoringFailed { reason: String },
}
#[derive(Debug, Clone)]
pub enum LifecycleError {
AgentNotFound { agent_id: AgentId },
InvalidStateTransition { from: AgentState, to: AgentState },
InitializationFailed { agent_id: AgentId, reason: String },
TerminationFailed { agent_id: AgentId, reason: String },
ConfigurationInvalid { reason: String },
}
#[derive(Debug, Clone)]
pub enum CommunicationError {
AgentNotRegistered { agent_id: AgentId },
MessageTooLarge { size: usize, max_size: usize },
DeliveryFailed { message_id: Option<MessageId>, reason: String },
EncryptionFailed { reason: String },
TopicNotFound { topic: String },
}
```
## Core Interfaces
### 1. Lifecycle Controller
```rust
#[async_trait]
pub trait LifecycleController {
async fn create_agent(&self, config: AgentConfig) -> Result<AgentInstance, LifecycleError>;
async fn initialize_agent(&self, agent_id: AgentId) -> Result<(), LifecycleError>;
async fn start_agent(&self, agent_id: AgentId) -> Result<(), LifecycleError>;
async fn stop_agent(&self, agent_id: AgentId) -> Result<(), LifecycleError>;
async fn suspend_agent(&self, agent_id: AgentId) -> Result<(), LifecycleError>;
async fn resume_agent(&self, agent_id: AgentId) -> Result<(), LifecycleError>;
async fn terminate_agent(&self, agent_id: AgentId) -> Result<(), LifecycleError>;
async fn get_agent_state(&self, agent_id: AgentId) -> Result<AgentState, LifecycleError>;
async fn list_agents(&self) -> Vec<AgentInstance>;
async fn get_agent(&self, agent_id: AgentId) -> Result<AgentInstance, LifecycleError>;
async fn update_agent_config(&self, agent_id: AgentId, config: AgentConfig) -> Result<(), LifecycleError>;
async fn restart_agent(&self, agent_id: AgentId) -> Result<(), LifecycleError>;
async fn get_system_status(&self) -> SystemStatus;
async fn shutdown(&self) -> Result<(), LifecycleError>;
}
pub struct LifecycleConfig {
pub initialization_timeout: Duration,
pub termination_timeout: Duration,
pub state_check_interval: Duration,
pub enable_auto_recovery: bool,
pub max_restart_attempts: u32,
pub max_agents: usize,
}
```
### 2. Resource Manager
```rust
#[async_trait]
pub trait ResourceManager {
async fn allocate_resources(&self, agent_id: AgentId, limits: ResourceLimits) -> Result<ResourceAllocation, ResourceError>;
async fn deallocate_resources(&self, agent_id: AgentId) -> Result<(), ResourceError>;
async fn update_resource_limits(&self, agent_id: AgentId, limits: ResourceLimits) -> Result<(), ResourceError>;
async fn get_resource_usage(&self, agent_id: AgentId) -> Result<ResourceUsage, ResourceError>;
async fn get_system_resources(&self) -> SystemResourceStatus;
async fn check_resource_violations(&self) -> Vec<ResourceViolation>;
async fn set_resource_alerts(&self, agent_id: AgentId, thresholds: ResourceThresholds) -> Result<(), ResourceError>;
async fn get_resource_history(&self, agent_id: AgentId, duration: Duration) -> Result<Vec<ResourceSnapshot>, ResourceError>;
async fn shutdown(&self) -> Result<(), ResourceError>;
}
pub struct ResourceManagerConfig {
pub total_memory: usize,
pub total_cpu_cores: u32,
pub total_disk_space: usize,
pub total_network_bandwidth: usize,
pub enforcement_enabled: bool,
pub auto_scaling_enabled: bool,
pub resource_reservation_percentage: f32,
pub monitoring_interval: Duration,
}
```
### 3. Scheduler
```rust
#[async_trait]
pub trait Scheduler {
async fn schedule_task(&self, task: ScheduledTask) -> Result<(), SchedulerError>;
async fn cancel_task(&self, task_id: TaskId) -> Result<(), SchedulerError>;
async fn get_task_status(&self, task_id: TaskId) -> Result<TaskStatus, SchedulerError>;
async fn list_pending_tasks(&self) -> Vec<ScheduledTask>;
async fn list_running_tasks(&self) -> Vec<RunningTask>;
async fn get_scheduler_metrics(&self) -> SchedulerMetrics;
async fn update_task_priority(&self, task_id: TaskId, priority: Priority) -> Result<(), SchedulerError>;
async fn pause_scheduling(&self) -> Result<(), SchedulerError>;
async fn resume_scheduling(&self) -> Result<(), SchedulerError>;
async fn shutdown(&self) -> Result<(), SchedulerError>;
}
pub struct ScheduledTask {
pub id: TaskId,
pub agent_id: AgentId,
pub priority: Priority,
pub scheduled_time: SystemTime,
pub timeout: Duration,
pub retry_count: u32,
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum LoadBalancingStrategy {
RoundRobin,
LeastConnections,
ResourceBased,
WeightedRoundRobin,
}
pub struct SchedulerConfig {
pub max_concurrent_tasks: usize,
pub task_timeout: Duration,
pub retry_attempts: u32,
pub load_balancing_strategy: LoadBalancingStrategy,
pub enable_priority_scheduling: bool,
pub task_queue_size: usize,
pub worker_threads: usize,
pub health_check_interval: Duration,
}
```
### 4. Communication Bus
```rust
#[async_trait]
pub trait CommunicationBus {
async fn register_agent(&self, agent_id: AgentId, capabilities: Vec<Capability>) -> Result<(), CommunicationError>;
async fn unregister_agent(&self, agent_id: AgentId) -> Result<(), CommunicationError>;
async fn send_message(&self, message: Message) -> Result<MessageId, CommunicationError>;
async fn receive_messages(&self, agent_id: AgentId) -> Result<Vec<SecureMessage>, CommunicationError>;
async fn subscribe_to_topic(&self, agent_id: AgentId, topic: String) -> Result<(), CommunicationError>;
async fn unsubscribe_from_topic(&self, agent_id: AgentId, topic: String) -> Result<(), CommunicationError>;
async fn broadcast_message(&self, topic: String, message: Message) -> Result<Vec<MessageId>, CommunicationError>;
async fn get_message_status(&self, message_id: MessageId) -> Result<DeliveryStatus, CommunicationError>;
async fn get_agent_topics(&self, agent_id: AgentId) -> Result<Vec<String>, CommunicationError>;
async fn shutdown(&self) -> Result<(), CommunicationError>;
}
pub struct CommunicationConfig {
pub message_ttl: Duration,
pub max_queue_size: usize,
pub delivery_timeout: Duration,
pub retry_attempts: u32,
pub enable_encryption: bool,
pub enable_compression: bool,
pub max_message_size: usize,
pub dead_letter_queue_size: usize,
}
```
### 5. Error Handler
```rust
#[async_trait]
pub trait ErrorHandler {
async fn handle_error(&self, agent_id: AgentId, error: RuntimeError) -> Result<ErrorAction, ErrorHandlerError>;
async fn register_strategy(&self, error_type: ErrorType, strategy: RecoveryStrategy) -> Result<(), ErrorHandlerError>;
async fn get_error_stats(&self, agent_id: AgentId) -> Result<ErrorStatistics, ErrorHandlerError>;
async fn get_system_error_stats(&self) -> SystemErrorStatistics;
async fn set_error_thresholds(&self, agent_id: AgentId, thresholds: ErrorThresholds) -> Result<(), ErrorHandlerError>;
async fn clear_error_history(&self, agent_id: AgentId) -> Result<(), ErrorHandlerError>;
async fn shutdown(&self) -> Result<(), ErrorHandlerError>;
}
#[derive(Debug, Clone)]
pub enum ErrorAction {
Retry { max_attempts: u32, backoff: Duration },
Restart,
Suspend,
Terminate,
Failover,
}
#[derive(Debug, Clone)]
pub enum RecoveryStrategy {
Retry { max_attempts: u32, backoff: Duration },
Restart { preserve_state: bool },
Failover { backup_agent: Option<AgentId> },
Terminate { cleanup: bool },
Manual { reason: String },
None,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
pub enum ErrorType {
ResourceExhaustion,
NetworkError,
SecurityViolation,
PolicyViolation,
SystemError,
ValidationError,
}
pub struct ErrorHandlerConfig {
pub max_error_history: usize,
pub error_aggregation_window: Duration,
pub escalation_threshold: u32,
pub circuit_breaker_threshold: u32,
pub circuit_breaker_timeout: Duration,
pub enable_auto_recovery: bool,
pub max_recovery_attempts: u32,
pub recovery_backoff_multiplier: f32,
}
```
## External Integrations
### 1. Policy Engine
```rust
#[async_trait]
pub trait PolicyEngine {
async fn validate_agent_config(&self, config: &AgentConfig) -> Result<PolicyValidationResult, PolicyError>;
async fn check_operation_allowed(&self, agent_id: AgentId, operation: &str, context: &PolicyContext) -> Result<bool, PolicyError>;
async fn get_agent_policies(&self, agent_id: AgentId) -> Result<Vec<Policy>, PolicyError>;
async fn update_policy(&self, policy: Policy) -> Result<(), PolicyError>;
async fn delete_policy(&self, policy_id: String) -> Result<(), PolicyError>;
async fn evaluate_policy(&self, policy_id: String, context: &PolicyContext) -> Result<PolicyDecision, PolicyError>;
}
pub struct Policy {
pub id: String,
pub name: String,
pub description: String,
pub rules: Vec<PolicyRule>,
pub priority: u32,
pub enabled: bool,
}
pub struct PolicyContext {
pub agent_id: AgentId,
pub operation: String,
pub resource_requirements: Option<ResourceLimits>,
pub security_context: SecurityContext,
pub metadata: HashMap<String, String>,
}
```
### 2. Sandbox Orchestrator
```rust
#[async_trait]
pub trait SandboxOrchestrator {
async fn create_sandbox(&self, config: SandboxConfig) -> Result<SandboxId, SandboxError>;
async fn start_sandbox(&self, sandbox_id: SandboxId) -> Result<(), SandboxError>;
async fn stop_sandbox(&self, sandbox_id: SandboxId) -> Result<(), SandboxError>;
async fn destroy_sandbox(&self, sandbox_id: SandboxId) -> Result<(), SandboxError>;
async fn get_sandbox_status(&self, sandbox_id: SandboxId) -> Result<SandboxStatus, SandboxError>;
async fn execute_command(&self, sandbox_id: SandboxId, command: &str, args: Vec<String>) -> Result<CommandResult, SandboxError>;
async fn upload_file(&self, sandbox_id: SandboxId, local_path: &str, remote_path: &str) -> Result<(), SandboxError>;
async fn download_file(&self, sandbox_id: SandboxId, remote_path: &str, local_path: &str) -> Result<(), SandboxError>;
}
pub struct SandboxConfig {
pub agent_id: AgentId,
pub security_tier: SecurityTier,
pub resource_limits: ResourceLimits,
pub network_config: NetworkConfig,
pub filesystem_config: FilesystemConfig,
pub environment_variables: HashMap<String, String>,
}
```
### 3. Audit Trail
```rust
#[async_trait]
pub trait AuditTrail {
async fn record_event(&self, event: AuditEvent) -> Result<AuditId, AuditError>;
async fn query_events(&self, query: AuditQuery) -> Result<Vec<AuditEvent>, AuditError>;
async fn verify_integrity(&self, from_time: SystemTime, to_time: SystemTime) -> Result<IntegrityReport, AuditError>;
async fn get_event(&self, audit_id: AuditId) -> Result<AuditEvent, AuditError>;
async fn export_events(&self, query: AuditQuery, format: ExportFormat) -> Result<Vec<u8>, AuditError>;
}
pub struct AuditEvent {
pub id: AuditId,
pub timestamp: SystemTime,
pub event_type: AuditEventType,
pub agent_id: Option<AgentId>,
pub details: String,
pub metadata: HashMap<String, String>,
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum AuditEventType {
AgentCreated,
AgentStarted,
AgentStopped,
AgentTerminated,
ResourceAllocated,
ResourceDeallocated,
MessageSent,
MessageReceived,
ErrorOccurred,
PolicyViolation,
SecurityEvent,
}
```
## Usage Examples
### Complete Agent Lifecycle
```rust
use symbiont_runtime::*;
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
// Initialize components
let lifecycle_controller = DefaultLifecycleController::new(LifecycleConfig::default()).await?;
let resource_manager = DefaultResourceManager::new(ResourceManagerConfig::default()).await?;
let scheduler = DefaultScheduler::new(SchedulerConfig::default()).await?;
let comm_bus = DefaultCommunicationBus::new(CommunicationConfig::default()).await?;
let error_handler = DefaultErrorHandler::new(ErrorHandlerConfig::default()).await?;
// Create agent configuration
let agent_config = AgentConfig {
id: AgentId::new(),
name: "example_agent".to_string(),
dsl_source: "agent logic".to_string(),
execution_mode: ExecutionMode::Persistent,
security_tier: SecurityTier::Tier2,
resource_limits: ResourceLimits {
memory_mb: 512,
cpu_cores: 1.0,
disk_io_mbps: 50,
network_io_mbps: 10,
execution_timeout: Duration::from_secs(3600),
idle_timeout: Duration::from_secs(300),
},
capabilities: vec![Capability::FileSystem, Capability::Network],
policies: vec![],
metadata: HashMap::new(),
priority: Priority::Normal,
};
// Create and manage agent
let agent = lifecycle_controller.create_agent(agent_config.clone()).await?;
println!("Created agent: {}", agent.id);
// Allocate resources
let allocation = resource_manager.allocate_resources(agent.id, agent_config.resource_limits).await?;
println!("Allocated resources for agent: {}", agent.id);
// Register with communication bus
comm_bus.register_agent(agent.id, agent_config.capabilities).await?;
println!("Registered agent with communication bus");
// Initialize and start agent
lifecycle_controller.initialize_agent(agent.id).await?;
lifecycle_controller.start_agent(agent.id).await?;
println!("Agent started successfully");
// Schedule a task
let task = ScheduledTask {
id: TaskId::new(),
agent_id: agent.id,
priority: Priority::Normal,
scheduled_time: SystemTime::now(),
timeout: Duration::from_secs(60),
retry_count: 0,
};
scheduler.schedule_task(task).await?;
// Send a message
let message = Message {
id: MessageId::new(),
from: agent.id,
to: agent.id, // Self-message for demo
topic: "test_topic".to_string(),
payload: b"Hello, world!".to_vec(),
priority: Priority::Normal,
ttl: Duration::from_secs(300),
};
comm_bus.send_message(message).await?;
// Monitor and cleanup
tokio::time::sleep(Duration::from_secs(5)).await;
let state = lifecycle_controller.get_agent_state(agent.id).await?;
println!("Agent state: {:?}", state);
let usage = resource_manager.get_resource_usage(agent.id).await?;
println!("Resource usage: {:?}", usage);
// Shutdown
lifecycle_controller.terminate_agent(agent.id).await?;
resource_manager.deallocate_resources(agent.id).await?;
comm_bus.unregister_agent(agent.id).await?;
Ok(())
}
```
This API reference provides complete type definitions and interface specifications for all components of the Symbiont Agent Runtime System, including the optional HTTP API.
## HTTP API Reference
### Overview
The HTTP API provides RESTful endpoints for external system integration. This API is optional and requires the `http-api` feature flag to be enabled.
#### Feature Activation
```toml
[dependencies]
symbiont-runtime = { version = "0.1.0", features = ["http-api"] }
```
#### Configuration
```rust
#[cfg(feature = "http-api")]
use symbiont_runtime::api::{HttpApiServer, HttpApiConfig};
let config = HttpApiConfig {
bind_address: "127.0.0.1".to_string(),
port: 8080,
enable_cors: true,
enable_tracing: true,
};
let server = HttpApiServer::new(config);
server.start().await?;
```
### HTTP API Types
#### Request/Response Types
```rust
#[cfg(feature = "http-api")]
pub struct WorkflowExecutionRequest {
pub workflow_id: String,
pub parameters: serde_json::Value,
pub agent_id: Option<AgentId>,
}
#[cfg(feature = "http-api")]
pub struct AgentStatusResponse {
pub agent_id: AgentId,
pub state: AgentState,
pub last_activity: chrono::DateTime<chrono::Utc>,
pub resource_usage: ResourceUsage,
}
#[cfg(feature = "http-api")]
pub struct HealthResponse {
pub status: String,
pub uptime_seconds: u64,
pub timestamp: chrono::DateTime<chrono::Utc>,
pub version: String,
}
#[cfg(feature = "http-api")]
pub struct ErrorResponse {
pub error: String,
pub code: String,
pub details: Option<serde_json::Value>,
}
#[cfg(feature = "http-api")]
pub struct ResourceUsage {
pub memory_bytes: Option<u64>,
pub cpu_percent: Option<f64>,
pub active_tasks: u32,
}
```
### Endpoints
#### Health Check
**Endpoint:** `GET /api/v1/health`
**Description:** Returns system health status and version information.
**Response:**
```json
{
"status": "healthy",
"uptime_seconds": 3600,
"timestamp": "2025-07-18T06:45:00Z",
"version": "0.1.0"
}
```
**Example:**
```bash
curl http://localhost:8080/api/v1/health
```
#### List Agents
**Endpoint:** `GET /api/v1/agents`
**Description:** Returns a list of all active agent IDs.
**Response:**
```json
[
"agent-id-1",
"agent-id-2",
"agent-id-3"
]
```
**Example:**
```bash
curl http://localhost:8080/api/v1/agents
```
#### Get Agent Status
**Endpoint:** `GET /api/v1/agents/{id}/status`
**Description:** Returns detailed status information for a specific agent.
**Parameters:**
- `id` (path): Agent ID
**Response:**
```json
{
"agent_id": "agent-id-1",
"state": "Running",
"last_activity": "2025-07-18T06:45:00Z",
"resource_usage": {
"memory_bytes": null,
"cpu_percent": null,
"active_tasks": 3
}
}
```
**Example:**
```bash
curl http://localhost:8080/api/v1/agents/agent-id-1/status
```
CPU and memory are nullable measurements. The scheduler currently has no
per-agent sampler and returns `null` for both internal and external agents;
`last_activity` is not a resource sample timestamp. `active_tasks` counts tasks
owned by this scheduler, so it remains a known number. Clients must accept null
and must not convert unavailable measurements to zero. Fleet Overview displays
**Not sampled**. Administrative worker measurements and reservations are available
separately through [worker capacity inspection](#worker-capacity-inspection).
#### Execute Workflow
**Endpoint:** `POST /api/v1/workflows/execute`
**Description:** An administrator submits raw DSL source for a scheduled invocation.
`workflow_id` contains the source, up to 1 MiB. The request uses the same governed
scheduler as registered-agent execution. `parameters` becomes its JSON input.
Authentication is required. Agent-scoped keys receive `403 ADMIN_REQUIRED`, even
when `agent_id` matches their scope. Scoped callers invoke registered source through
`POST /api/v1/agents/{id}/execute` with `{"input": ...}`. The legacy
`SYMBIONT_API_TOKEN`, where enabled, carries administrative authority.
Both execution endpoints require an `Idempotency-Key` UUID retained for retries.
Omit `agent_id` to derive a registration ID from that UUID. Supplying an existing UUID replaces its
source for subsequent invocations. The first declared agent is selected by default;
`metadata { name = "selected" }` can select another declaration. The registered
name is the selected declaration's actual name. Parsing, selection and supported
inline policies are validated before registration. Unsupported policies are refused.
This endpoint retains its Docker/Tier1 registration default; conflicting sandbox
selection fails during governed execution. It does not interpret arbitrary executable
DSL bodies. See [inline policies](../../docs/inline-policies.md).
**Request Body:**
```json
{
"workflow_id": "agent report() { with sandbox = \"docker\" {} }",
"parameters": {
"input_file": "input.csv",
"output_format": "json"
},
"agent_id": "19b183f7-97c4-4e42-9c62-5e9c940bfae3"
}
```
**Fresh response (200):**
```json
{
"status": "queued",
"invocation_id": "0949b393-a3b1-4564-a206-6dbe3e19dd57",
"agent_id": "19b183f7-97c4-4e42-9c62-5e9c940bfae3",
"execution_id": "c7022f13-7140-4a09-8e30-b1941e0cbb32",
"replayed": false,
"audit": {"run_id": "c7022f13-7140-4a09-8e30-b1941e0cbb32", "path": "/project/.symbiont/governed/agent.run.jsonl", "public_key": "hex"}
}
```
Repeat the same authenticated submission with the same UUID to retrieve the
saved completion without execution. HTTP 409 distinguishes `in_progress`,
`unresolved` and `conflict`; HTTP 422 returns a saved known failure. Changed caller,
source, agent configuration or input cannot reuse an ID. Invalid source returns
400; an unavailable registered agent returns 404; admission/storage refusal returns
503. See [scheduler idempotency](../../docs/scheduler-idempotency.md) for the complete
contract, restart behavior and limits.
Agent history records `queued`, then `Completed`, `Failed`, `TimedOut`,
`Terminated` or `Unresolved` under `execution_id`. It remains a bounded in-memory
view; the durable claim and signed journal preserve retry evidence across restart.
Completed saved results expose output, usage, shared budget and audit information.
The focused `workflow_execution` target requires Unix, `http-api` and `cedar`.
It uses actual HTTP authentication and protected execution to test authority,
source selection, parameter delivery and caller isolation. The shipping
`scripts/test-scheduler-invocations.py` adds actual Docker effects, concurrent
admission and crash/restart checks. `scripts/test-workflow-execution.py` retains
source-selection and file-grant coverage.
**Example:**
```bash
# Generate once and retain for retries.
INVOCATION_ID=$(cat /proc/sys/kernel/random/uuid)
curl -X POST http://localhost:8080/api/v1/workflows/execute \
-H "Idempotency-Key: $INVOCATION_ID" \
-H "Authorization: Bearer $SYMBIONT_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"workflow_id":"agent report() { with sandbox = \"docker\" {} }","parameters":{"input_file":"input.csv"}}'
```
#### Get System Metrics
**Endpoint:** `GET /api/v1/metrics`
**Description:** Returns system performance metrics and statistics.
**Response:**
```json
{
"agents": {
"total": 10,
"running": 8,
"stopped": 2
},
"system": {
"memory_usage": 85.5,
"cpu_usage": 45.2,
"uptime_seconds": 7200
},
"performance": {
"messages_per_second": 1250,
"avg_response_time_ms": 15
}
}
```
**Example:**
```bash
curl http://localhost:8080/api/v1/metrics
```
### Error Handling
All endpoints return consistent error responses:
```json
{
"error": "Agent not found",
"code": "AGENT_NOT_FOUND",
"details": {
"agent_id": "invalid-agent-id"
}
}
```
**HTTP Status Codes:**
- `200 OK` - Successful operation
- `400 Bad Request` - Invalid request parameters
- `404 Not Found` - Resource not found
- `500 Internal Server Error` - Server error
### Authentication & Security
The HTTP API includes middleware for:
- CORS handling (configurable)
- Request tracing and logging
- Configurable rate limiting
- Bearer authentication on protected routes
- Security response headers
Administrative endpoints additionally reject agent-scoped keys. Health is public;
callback routes use their route-specific verification. Swagger is opt-in and
protected by bearer authentication.
This API reference provides complete type definitions and interface specifications for all components of the Symbiont Agent Runtime System, including the optional HTTP API.
### Persistent cron triggers
`POST /api/v1/schedules/{id}/trigger` requires an administrative bearer token and
one `Idempotency-Key` UUID. It returns the same queued/saved/in-progress/unresolved/
conflict contract as agent execution. Reuse the UUID after a lost response;
changing it requests new work. Schedule history persists occurrence IDs, protected
admission audit references and final execution results. Unknown occurrences block
resume and further execution of that job. See [cron recovery](../../docs/cron-recovery.md)
for project storage migration, timer identities and restart behavior.
## Coordinator chat admission
Coordinator WebSocket messages at `/ws/chat` require a retained non-nil UUID in
`ChatSend.id`. The server durably admits that caller/content identity before
queueing and correlates replies through the same `request_id`. `ChatInspect`
with the same UUID/content performs read-only lookup. Completed replies include
`ChatChunk.replayed`; active, unresolved, reconciled and conflicting IDs never
repeat execution. See [chat recovery](../../docs/chat-recovery.md) for the protocol,
error codes, queue contract and console migration from arbitrary string IDs.
## Verified run inspection
`GET /api/v1/audit/runs/{agent_uuid}/{run_uuid}?public_key={64_hex_characters}`
requires an administrative bearer key. It returns a bounded signed-journal
snapshot from the trusted project, including parent/child links, recorded shared
budgets, per-action permissions, original unknown/incomplete outcomes and a
separate invocation assessment when available. It never executes or retries work.
Scoped keys receive 403; malformed keys receive 400; invalid or unavailable
evidence receives 422. See [operator run inspection](../../docs/run-inspector.md)
for trust inputs, limits and response interpretation.
The run view's `recovery.recovered_budget` reconstructs signed root-family
reservation accounting, with unknown requests retaining their full charge even
without a final budget snapshot. Older or child journals without root history
return null. A child's `budget_root` links to the root history when available.
These fields are read-only evidence and never authorize execution resumption or
provider-request replay. See [inference budget recovery](../../docs/provider-budget-recovery.md).
## Worker capacity inspection
`GET /api/v1/sandbox/capacity` returns the running supervisor's retained worker,
memory and CPU reservations, pool limits, remaining capacity and unknown resource
metadata count. Attributed workers include the originating run, signing public
key, tool, iteration, dispatch ID and call fingerprint in nullable `origin`.
The Inspector verifies that run separately; retained metadata is not signed proof.
On Unix, nullable `staging` reports snapshot slots/bytes, remaining capacity and
ownership references. A staging inspection error leaves worker totals available
and populates `staging_error` without initializing or reconciling staging state.
Landlock host workers currently bypass this supervisor and are absent from totals.
`GET /api/v1/sandbox/workers/{lease_uuid}/usage` separately samples
one retained worker. Both require an administrative bearer key and never start a
missing helper or release capacity. Scoped keys receive 403; unavailable or busy
inspection receives 503 with `CAPACITY_UNAVAILABLE`. Unknown values remain null,
and successful/handler-generated error responses are not cacheable. These are
live observations, not signed run snapshots. See
[worker capacity](../../docs/worker-capacity.md) for response fields, units,
backend semantics and deployment scope.