onlyne-client
One workspace, one role, one daemon. The role runs many sessions at once.
Verbs
| verb | one line |
|---|---|
run --workspace <dir> |
Foreground role runtime: connect, handshake, pull, dispatch, report. Backgrounding is the operator's job, never the client's. |
status --workspace <dir> |
Print uptime, socket path, recorded fault count, and whether the server link is up. |
doctor |
Print host-detection JSON. No workspace, no socket. Exit 0. |
init --workspace <dir> --role <r> --server-root <dir> |
Build the minimal role workspace and print the [[client]] spec fragment. |
roles --workspace <dir> |
Answer role prose from the local cache. |
sessions --workspace <dir> |
Reserved for the live role runtime. |
watch --workspace <dir> |
Reserved for the live role runtime. |
history --workspace <dir> |
Reserved for the live role runtime. |
run is the only launch verb, and it stays in the foreground. --workspace takes a relative path and resolves it to an absolute path before use, so the daemon and its generated sessions all read one location. A supervisor that wants the client in the background owns that decision — a visible terminal tab, launchd, nohup — so the client never detaches, writes no pid file, and nothing signals it by number. A run whose adapter socket cannot be bound ends there with exit 1 and names the failure on stderr; an accept error after a successful bind logs at error level (adapter socket accept failed; retrying) and retries every 100 ms with the listener held.
status prints onlyne: client running uptime <n>s socket <path> faults <n>. The <path> is the served socket in the machine-level runtime directory, /tmp/onlyne-<uid>/<digest>.sock ($ONLYNE_RUNTIME_DIR overriding the directory), where <digest> is the first 16 hex characters of sha256 over the workspace's canonical root. The uptime is the age of the <digest>.json registration that client published, and a client counts as running only when that socket answers an admin hello, so a socket an unclean exit left behind reads as not running. When the answering client holds no server link it adds onlyne: client not connected on stderr.
The printed [[client]] fragment is a complete role entry: it carries role, key, admin, max_sessions, the ACL lists, prose, and [client.runtime] drive and command. Paste it into spec.toml and reload; the client can then spawn sessions for that role.
Workspace layout
Both init and run create these paths under --workspace:
| path | mode | content |
|---|---|---|
.onlyne/config.toml |
role, cert_pin, key_path, plugins = [...], [server] host and port, [orca] worktree |
|
.onlyne/client.db |
SQLite: intents, sessions, faults, prose_cache, config_cache, events |
|
.onlyne/keys/role.key |
0600 |
32 raw ed25519 bytes, generated once |
.onlyne/run/ |
0700 |
runtime directory |
.onlyne/logs/client.log |
stdout and stderr, when the operator starts run under a shell that redirects them |
|
.onlyne/agent/<id>/ |
installed plugin package with plugin.toml |
|
.onlyne/cache/orca-tabs.jsonl |
append-only Orca tab to session map: a supervisor/display side-channel, not the identity (the adapter protocol owns that) |
The adapter socket and its registration live outside the workspace, in the
machine-level runtime directory — /tmp/onlyne-<uid>/, $ONLYNE_RUNTIME_DIR
overriding it, 0700. <workspace>/.onlyne/run/s stays the canonical spelling
operators read; nothing binds there.
| path | mode | content |
|---|---|---|
<runtime>/<digest>.sock |
0600 |
the adapter socket, bound for the life of the run |
<runtime>/<digest>.json |
0600 |
the registration: kind (client), role, root, pid, version, and the runtime hosting the role's sessions |
run binds the socket, then publishes the registration, and the run's end
removes it: a registration outliving its surface is what an external runtime's
plugin reads as a live client.
init never writes spec.toml. A workspace holding the pre-v1 layout is refused before any write: exit 6, and one line naming the marker that decided it plus the way forward. There is no migrate command — point the client at a workspace init has written and keep the old one for reference.
Three config values take a $NAME spelling: cert_pin, key_path, and [server] host. At startup run reads the environment variable named after the $, then puts its value where the config line sits. The gateway plugins use that same idiom for platform tokens. A name the environment carries no value for — absent, or present and blank — stops the launch with exit 1 and one line on stderr naming both the field and the variable: onlyne-client: missing secret $ONLYNE_CERT for cert_pin; set the environment variable. A value with no leading $ travels verbatim, so a literal $ inside a value stays part of the string.
Exit codes
| code | meaning |
|---|---|
| 0 | the verb finished |
| 1 | the verb failed; the reason is one line on stderr |
| 1 | run could not bind the adapter socket; stderr is onlyne-client: bind the workspace socket <canonical path>: <detail>, the detail naming the served path, both byte lengths, and the OS reason |
| 2 | status found no client answering its socket, printed as onlyne: client not running |
| 2 | status found a client with no server link, printed as onlyne: client not connected |
| 2 | the workspace holds the legacy layout |
| 5 | run received an ONLYNE_BACKEND value that is not a placement |
status exits 0 only for a client that is up and connected to its server. That is the fact a script reads.
doctor exits 0 for every placement-detection result, including placement: null.
Backends
Selection is env ONLYNE_BACKEND (nonempty) > workspace config.toml placement > probe > headless fallback.
| name | parse aliases | how it is chosen | notes |
|---|---|---|---|
orca |
env, config, or auto probe (first) | tab host | |
zellij |
env, config, or auto probe | pane host; probe maps EXITED / exit_status |
|
headless |
env or config only | machine placement for exec and ACP drives | |
external |
env or config only | externally managed placement | |
fake |
env only | in-process test runtime |
A nonempty value that names a placement selects it. The spec's [client.runtime] drive names plugin, acp, or exec; acp requires headless. An unknown explicit ONLYNE_BACKEND value makes onlyne-client run exit 5.
fake runs sessions in-process and needs no external tool; the end-to-end scripts set ONLYNE_BACKEND=fake. The exec drive spawns the role's [client.runtime] command as a child of the client, holds stdin open, and appends the child's output to .onlyne/logs/session-<task>.log. On child exit, probe may fill detail.output_tail (at most 200 lines / 16 KiB). crates/onlyne-testkit/e2e/pi-live.sh and exec-headless.sh set this path. Windows close uses CREATE_NEW_PROCESS_GROUP plus CTRL_BREAK, then kill; a process with no console terminates the child directly. Operator-facing graceful stop of the daemons is onlyne shutdown.
A role at max_sessions keeps pulling with control_only, which is the path that lets focus, recycle, and cancel reach the session occupying the last free slot.
An agent that mounts naming no session — the always-running plugin — parks as the connection for the next staged session. The claim binds that socket to the session it takes, and a mount that arrives after a work item hands that item over on the spot. A connection that named no session releases only the transports sharing its socket.
doctor
onlyne-client doctor is a read-only verb. It prints one JSON object and exits 0. Fields:
| field | meaning |
|---|---|
placement |
selected placement name, or null |
placement_selection |
explicit, config, probe, or fallback |
explicit |
raw ONLYNE_BACKEND when nonempty |
binary |
CLI path or name for orca/zellij; null for exec, fake, and no host |
refusal |
present only when an explicit placement name is unknown |
A missing placement yields placement: null and exit 0. refusal is present only for an unknown explicit name. The verb is a pre-deploy check.
[orca] worktree in config.toml sets which Orca tab list a session tab joins. Three states:
host(the default) readsORCA_WORKTREE_ID, the worktree id Orca exports to the tab the supervisor started the client in and which the daemon inherits. Every session tab lands flat in that worktree's tab list, beside the supervisor's own tabs. Start the client outside an Orca tab and the variable is absent, so the policy behaves likeinherit.inheritpasses no selector and leaves the choice to Orca's active worktree.- Any other value is used verbatim as an Orca worktree selector (
id:<…>,path:<abs>,name:<…>,branch:<…>).
Tab ownership and working directory are independent. The selector decides which tab list the tab joins; the spawned command's own cd decides where the agent runs. The role workspace therefore never has to exist in Orca: it is not registered, not opened, and not cleaned up. That is the whole reason a generated (non-git) role workspace works at all. Orca's public registration command accepts git checkouts only, so a path:<workspace> selector would fail for exactly the directories this client hands out.
Sessions
max_sessions from the role's spec entry caps how many sessions a role runs at once. A session whose stored lifecycle reads exited spends none of that cap: the rows of sessions the role has ended stay in client.db as its history and stay queryable. The client keeps pulling while fewer than max_sessions sessions have not exited. Each task gets its own session and its own spawn. A session that has finished one task takes no further task; its slot releases, its host resource closes, and it stops counting against max_sessions.
The host resource retires with the session: a pane, tab, zellij session, or exec child closes once that session holds no task and no plugin transport is attached. Three paths do the closing — a graceful plugin detach closes each idle session that connection served, a settle with no attached agent closes at settle time, and the 250 ms readiness tick closes any tracked session whose stored lifecycle projects exited with a stored outcome while its agent is gone, taking the reason from that outcome (Completed, Fault, or Cancelled). One case keeps the resource: a connection that dropped without a detach, where that agent may reconnect. Past [client] reconnect_grace_secs that agent is gone, and the sweep settles the task the session still owed failed and refuses that task's delivery with reason session_dead: the row leaves in_flight, so the ledger carries the ending an operator reads and repair retry is what brings the work back. The same pass publishes the session's own projection — the heartbeat report every ordinary ending travels — so the server's mirrored row for it reads exited at once, instead of reading working until the server's stale observer records a stale_working or heartbeat_missing fault. A retirement with the stored resource still open refreshes a stale backend_ref through attach, projects resource_closed, logs retiring idle session resource with task, backend, resource, and reason, then closes the resource and drops the slot; a close that fails is a warning.
Session scopes
[client.session] scope decides how long a session lives and what it serves. The scope is
the client's own rule: the server delivers by role and knows nothing about it.
oneshot(the default) — one delivery per session. The delivery settles, the session ends, and a delivery that was in flight when the client or the runtime restarted is requeued.task— one session per task family, keyed on the delivery's causality family (causality.family, falling back to the delivery's own task id). In planner → builder → reviewer → builder the second delivery to the builder enters the session the builder used the first time, which is the point of the scope: the conversation keeps its context, and no summary is assembled from the journal to fake it.role— a standing pool for the role: deliveries reuse the sessions the role already holds, and one pastmax_sessionswaits for a session instead of opening another one (it is not refused; it is dispatched when a session frees up).
A scoped session that has finished a delivery goes idle rather than retiring: its binding is released, its host resource stays, and the next delivery of its scope enters it again.
idle_close bounds how long an idle task or role session waits — "2h" by default,
0 (or Duration::ZERO) meaning it is never closed for idleness.
Suspension
An idle task or role session may suspend: the runtime saves the conversation, and the
client releases the process and the slot, so a suspended session spends none of
max_sessions. The next delivery bound to that session resumes it instead of opening a new
one.
Suspend depends on the runtime's own capability, declared as resume in its spec entry's
capabilities. A runtime without it degrades to "process alive, session alive": the session
is never suspended, its process and slot stay, and the family's next delivery still enters
it. The client never writes a history summary of the journal to stand in for a resumed
conversation; a runtime that cannot restore one keeps its process.
The server link
A session the client holds survives a lost server link: the link drops, intents queue,
and every session and the delivery it is serving stay exactly as they were. Only the loss
of the plugin connection — the agent gone, past [client] reconnect_grace_secs — settles
work.
Every hello, first and after each redial, carries live_sessions: each session the client
holds, the delivery that session is bound to (none when it is between deliveries), and
whether it is suspended. That is what keeps the server from requeuing work a live session
is still serving.
Server link
The client reconnects on a ladder of 1, 2, 4, 8, 16, 32, 60 seconds; 60 seconds repeats for every later attempt. After a reconnect the order is handshake, welcome, intent flush, pull resume.
A link failure or a bye frame sets accept_new = false. Queued deliveries wait on the server, running sessions continue to their terminal state, and the completions those sessions produce enter the intent queue. accept_new = false blocks new session spawns and new pulls. The gate follows the connection rather than any one frame: the runloop sets it from the link's readiness, and a frame that could not leave — a request past its own deadline with the link still up — goes to the intent queue alone.
A delivery the pull already had in hand when the gate shut is left unanswered: the row stays in flight and the next hello requeues it. A refusal would settle that row rejected, which is terminal, so the work would come back only through an operator's repair retry.
Intent queue
Every outbound envelope lands in client.db intents before the first socket write, keyed by op_id.
| state | meaning | next state |
|---|---|---|
pending |
enqueued, first attempt owed | accepted, retrying, exhausted |
retrying |
waiting for a later attempt | accepted, retrying, exhausted |
accepted |
receipt stored | terminal |
exhausted |
attempt ceiling reached with a recorded fault | terminal |
The role spec sets the attempt ceiling in intent.attempts and the delay between attempts in intent.backoff_ms. The queue lives in the client database, so a restart resumes it.
| answer | rule |
|---|---|
ok = true |
store the receipt, mark accepted |
duplicate |
replay the stored receipt from the first attempt |
acl_denied, invalid, conflict, forbidden, unknown_role, not_admin, bad_frame, frame_too_large, protocol_version |
drop the row without a retry |
internal, connection loss |
count the attempt and retry after the ladder |
Hitting the ceiling records fault kind intent_exhausted and sends report{kind:"fault"} once the server link exists. An exhausted row is never dropped silently.
Role prose
welcome.prose is cached in prose_cache, keyed by role with spec_hash. roles and the cluster prose export read that record.
onlyne-client(中文版 / Chinese Mirror)
一个工作区、一个角色、一个守护进程。该角色可同时运行多个会话。
动词
| 动词 | 一句话说明 |
|---|---|
run --workspace <dir> |
前台角色运行时:连接、握手、拉取、分派、报告。后台运行由操作者负责,客户端自身不负责。 |
status --workspace <dir> |
打印运行时长、socket 路径、已记录故障数以及服务器链接是否可用。 |
doctor |
打印主机检测 JSON。无需工作区或 socket。退出码为 0。 |
init --workspace <dir> --role <r> --server-root <dir> |
构建最小角色工作区并打印 [[client]] spec 片段。 |
roles --workspace <dir> |
从本地缓存回答角色 prose。 |
sessions --workspace <dir> |
为实时角色运行时保留。 |
watch --workspace <dir> |
为实时角色运行时保留。 |
history --workspace <dir> |
为实时角色运行时保留。 |
run 是唯一的启动动词,并始终留在前台。--workspace 接受相对路径,并在使用前将其解析为绝对路径,因此守护进程与它所生成的会话都读取同一位置。需要让客户端在后台运行的管理器负责这一决定——可见终端标签页、launchd、nohup——客户端自身不会分离、不写 pid 文件,也没有东西按编号向其发送信号。无法绑定适配器 socket 的 run 会就此结束,退出码为 1,并在 stderr 指明失败原因;成功绑定后出现 accept 错误时,会以 error 级别记录(adapter socket accept failed; retrying),并保持监听器、每 100 ms 重试一次。
status 打印 onlyne: client running uptime <n>s socket <path> faults <n>。<path> 是机器级运行时目录中的已提供服务 socket,即 /tmp/onlyne-<uid>/<digest>.sock($ONLYNE_RUNTIME_DIR 可覆盖该目录),其中 <digest> 是工作区规范根路径 sha256 的前 16 个十六进制字符。运行时长取自该 client 发布的 <digest>.json 注册文件的存续时间;只有该 socket 回应 admin hello 时,客户端才计为运行中,因此异常退出遗留的 socket 文件会显示为未运行。作出响应的客户端若没有服务器链接,会在 stderr 附加 onlyne: client not connected。
打印出的 [[client]] 片段是完整的角色条目:其中包含 role、key、admin、max_sessions、ACL 列表、prose 和 session_command。将其粘贴到 spec.toml 并重新加载后,客户端便可为该角色生成会话。
工作区布局
init 和 run 都会在 --workspace 下创建以下路径:
| 路径 | 模式 | 内容 |
|---|---|---|
.onlyne/config.toml |
角色、cert_pin、key_path、plugins = [...]、[server] 主机和端口、[orca] 工作树 |
|
.onlyne/client.db |
SQLite:intents、sessions、faults、prose_cache、config_cache、events |
|
.onlyne/keys/role.key |
0600 |
32 个原始 ed25519 字节,只生成一次 |
.onlyne/run/ |
0700 |
运行时目录 |
.onlyne/logs/client.log |
操作者通过会重定向输出的 shell 启动 run 时的 stdout 和 stderr |
|
.onlyne/agent/<id>/ |
包含 plugin.toml 的已安装插件包 |
|
.onlyne/cache/orca-tabs.jsonl |
仅追加的 Orca 标签页到会话映射:供管理器/显示使用的旁路信息,不是身份来源(身份由适配器协议管理) |
适配器 socket 及其注册文件位于工作区之外,位于机器级运行时目录——/tmp/onlyne-<uid>/,$ONLYNE_RUNTIME_DIR 可覆盖它,权限为 0700。<workspace>/.onlyne/run/s 只作为操作者阅读的规范拼写保留;那里不绑定任何东西,也不存在路径长度规则。
| 路径 | 模式 | 内容 |
|---|---|---|
<runtime>/<digest>.sock |
0600 |
适配器 socket,在整个 run 期间保持绑定 |
<runtime>/<digest>.json |
0600 |
注册文件:kind(client)、role、root、pid、version,以及承载该角色会话的 runtime |
run 先绑定 socket,再发布注册文件,运行结束时将其删除:一份比其服务面活得更久的注册文件,正是外部运行时插件读作"客户端在线"的依据。
init 绝不会写入 spec.toml。如果工作区采用 v1 之前的布局,程序会在任何写入之前拒绝处理:退出码为 6,并输出一行,指明决定它的标志以及下一步该做什么。没有 migrate 子命令——把 client 指向 init 写出的工作区,旧的那份留作参考。
三个配置值支持 $NAME 写法:cert_pin、key_path 和 [server] host。启动时,run 读取以 $ 之后名称命名的环境变量,再把其值填入配置行所在位置。网关插件也以相同方式处理平台令牌。如果某个名称对应的环境变量没有值——不存在,或存在但为空白——启动会停止,退出码为 1,并在 stderr 输出一行,同时指明字段和变量:onlyne-client: missing secret $ONLYNE_CERT for cert_pin; set the environment variable。不以 $ 开头的值会原样传递,因此值中的字面量 $ 仍是字符串的一部分。
退出码
| 代码 | 含义 |
|---|---|
| 0 | 动词已完成 |
| 1 | 动词失败;原因以一行形式写入 stderr |
| 1 | run 无法绑定适配器 socket;stderr 为 onlyne-client: bind the workspace socket <canonical path>: <detail>,详情指明所服务的路径、两个字节长度和 OS 原因 |
| 2 | status 未发现客户端回应其 socket,打印 onlyne: client not running |
| 2 | status 发现客户端没有服务器链接,打印 onlyne: client not connected |
| 2 | 工作区采用旧版布局 |
| 5 | run 收到不是 placement 的 ONLYNE_BACKEND 值 |
只有客户端正在运行且已连接到服务器时,status 才退出 0。脚本读取的就是这一事实。
doctor 对每一种主机检测结果(包括 host: null)都退出 0。
后端
选择优先级为环境变量 ONLYNE_BACKEND(非空)> 工作区 config.toml 的 backend > auto。
| 名称 | 解析别名 | 选择方式 | 备注 |
|---|---|---|---|
orca |
环境变量、配置或 auto 探测(首个) | 标签页主机 | |
zellij |
环境变量、配置或 auto 探测 | 窗格主机;探测会映射 EXITED / exit_status |
|
exec |
headless |
仅环境变量或配置 | 投影将后端记录为 exec |
fake |
仅环境变量或配置 | 进程内运行,用于测试 | |
auto |
空字符串 | 环境变量和配置均为空时的默认值 | 依次探测 orca、zellij |
非空值若为 orca、zellij、exec/headless 或 fake,就会选择相应后端。auto 从不发现 exec 和 fake。没有匹配项时,onlyne-client run 退出 5,并写入 onlyne: no supported host detected; run inside orca or zellij, or set ONLYNE_BACKEND。
fake 在进程内运行会话,无需外部工具;端到端脚本会设置 ONLYNE_BACKEND=fake。exec 将角色的 session_command 作为客户端的子进程生成,保持 stdin 打开,并把子进程输出追加到 .onlyne/logs/session-<task>.log。子进程退出时,probe 可以填充 detail.output_tail(最多 200 行 / 16 KiB)。crates/onlyne-testkit/e2e/pi-live.sh 和 exec-headless.sh 会设置此路径。在 Windows 上关闭时,使用 CREATE_NEW_PROCESS_GROUP 加 CTRL_BREAK,然后 kill;没有控制台的进程会直接终止子进程。面向操作者、用于优雅停止守护进程的命令是 onlyne shutdown。
角色达到 max_sessions 后,仍会通过 control_only 拉取;这条路径可让 focus、recycle 和 cancel 到达占用最后一个空槽位的会话。
挂载时未指定会话的代理——即始终运行的插件——会作为下一个待定会话的连接等待。认领会把该 socket 绑定到所接管的会话;工作项之后才到达的挂载会立即移交给该工作项。未指定会话名称的连接只释放共享其 socket 的传输。
doctor
onlyne-client doctor 是只读动词。它打印一个 JSON 对象并退出 0。字段如下:
| 字段 | 含义 |
|---|---|
host |
所选后端名称,或 null |
backend_selection |
explicit、env 或 none |
explicit |
非空时的原始 ONLYNE_BACKEND |
binary |
orca/zellij 的 CLI 路径或名称;对于 exec、fake 和无主机情况为 null |
refusal |
NO_SUPPORTED_HOST 行;当 host 为 null 时存在 |
未检测到主机会得到 host: null 和 refusal,并退出 0。该动词用于部署前检查。
config.toml 中的 [orca] worktree 决定会话标签页加入哪个 Orca 标签页列表。三种状态为:
host(默认)读取ORCA_WORKTREE_ID,即 Orca 导出到管理器启动客户端所在标签页、并由守护进程继承的工作树 id。每个会话标签页都会平铺到该工作树的标签页列表中,与管理器自身的标签页并列。如果在 Orca 标签页外启动客户端,该变量不存在,因此此策略的行为与inherit相同。inherit不传选择器,将选择交给 Orca 的活动工作树。- 任何其他值都会原样用作 Orca 工作树选择器(
id:<…>、path:<abs>、name:<…>、branch:<…>)。
标签页归属与工作目录彼此独立。选择器决定标签页加入哪个标签页列表;所生成命令自身的 cd 决定代理的运行位置。因此,角色工作区从不需要存在于 Orca:它不会被注册、打开或清理。这正是生成式(非 git)角色工作区能够工作的原因。Orca 的公共注册命令只接受 git 检出,所以 path:<workspace> 选择器恰好会因本客户端分发的目录而失败。
会话
角色 spec 条目中的 max_sessions 限制该角色同时运行的会话数量。存储的 lifecycle 状态为 exited 的会话不占该配额:角色已结束会话对应的行作为历史保留在 client.db 中,仍可查询。只要尚未退出的会话少于 max_sessions,客户端就会继续拉取。每个任务都有自己的会话和自己的生成过程。完成一个任务的会话不再接收任务;其槽位释放,主机资源关闭,并且不再计入 max_sessions。
主机资源随会话退役:当会话不再持有任务且没有插件传输连接时,窗格、标签页、zellij 会话或 exec 子进程就会关闭。共有三条关闭路径——插件正常 detach 会关闭该连接服务过的每个空闲会话;结算时没有已连接代理,则在结算时关闭;250 ms 就绪检查则在代理消失且存储的 lifecycle 投影为 exited、并带有存储 outcome 时,关闭任何被跟踪的会话,原因取自该 outcome(Completed、Fault 或 Cancelled)。有一种情况会保留资源:连接在未执行 detach 的情况下中断,因为该代理可能重连。超过 [client] reconnect_grace_secs 后,即认为该代理已经消失;清理流程会将会话仍欠下的任务结算为 failed,并以 session_dead 为原因拒绝对应任务交付:该行会离开 in_flight,因此账本保留了操作者可见的结束结果,而工作只能由 repair retry 重新带回。同一次处理还会发布会话自身的投影——每次正常结束都会发送的心跳报告——因此服务器上镜像的行会立即读取为 exited,无需在服务器的过期观察器记录 stale_working 或 heartbeat_missing 故障之前一直读取为 working。如果退役时存储的资源仍处于打开状态,会通过 attach 刷新过期的 backend_ref,投影 resource_closed,以任务、后端、资源和原因为由记录 retiring idle session resource,然后关闭资源并释放槽位;关闭失败只产生警告。
会话作用域
[client.session] scope 决定一个会话活多久、为谁服务。作用域完全是客户端自己的规则:服务器只按角色投递,对此一无所知。
oneshot(默认)——一个会话只服务一次投递。投递结算后会话结束;客户端或运行时重启时仍在途的投递会被重新排队。task——每个任务家族一个会话,家族取自投递的因果关系家族(causality.family,缺失时回退到投递自身的任务 id)。在 planner → builder → reviewer → builder 中,第二次投递给 builder 的投递会进入 builder 第一次使用的会话,这正是该作用域的意义:对话保留自己的上下文,绝不从日志里拼出一份摘要来顶替它。role——角色级的常驻会话池:投递优先复用角色已持有的会话,超出max_sessions的那一个会等待(不是被拒绝),等到有会话空出来再派发。
已完成一次投递的带作用域会话进入空闲而不是退役:绑定释放,宿主资源保留,下一个属于同一作用域的投递会再次进入它。
idle_close 限定空闲的 task 或 role 会话等待多久——默认 "2h",0(即 Duration::ZERO)表示永不因空闲而关闭。
挂起
空闲的 task 或 role 会话可以挂起:运行时保存对话,客户端释放进程与槽位,因此挂起的会话不占用 max_sessions。下一个绑定到该会话的投递会恢复它,而不是新开一个。
挂起取决于运行时自己的能力,即 spec 条目能力列表中的 resume。没有该能力的运行时退化为“进程在,会话就在”:会话永不挂起,进程与槽位保留,该家族的后续投递仍然进入它。客户端绝不写一份日志摘要来顶替被恢复的对话;无法恢复对话的运行时只能继续持有进程。
服务器链接
客户端持有的会话能挺过服务器链路断开:链路断了,intent 排队,每个会话以及它正在服务的投递都原样保留。只有插件连接断开——代理消失且超过 [client] reconnect_grace_secs——才会结算工作。
每一次 hello(首次以及每次重连后)都携带 live_sessions:客户端持有的每个会话、该会话绑定的投递(处于两次投递之间时为空),以及它是否处于挂起。服务器正是靠它避免把仍在被服务的投递重新排队。
服务器链接
客户端按 1、2、4、8、16、32、60 秒的阶梯重连;之后每次尝试都重复 60 秒间隔。重连后的顺序为握手、welcome、intent 刷新、恢复拉取。
链接失败或收到 bye 帧会设置 accept_new = false。队列中的交付会在服务器等待,正在运行的会话继续到达终态,这些会话产生的完成事件则进入 intent 队列。accept_new = false 会阻止生成新会话和执行新拉取。此开关跟随连接,而非某一个帧:运行循环根据链接就绪状态设置它;无法发出的帧——请求超过自身期限但链接仍在时——只进入 intent 队列。
开关关闭时,拉取已经持有的交付不会得到回答:该行保持飞行状态,下一次 hello 会重新入队。拒绝会将该行结算为 rejected,这是终态,因此工作只能通过操作者执行 repair retry 再次返回。
Intent 队列
每个出站 envelope 都会以 op_id 为键,在首次 socket 写入之前落入 client.db 的 intents。
| 状态 | 含义 | 下一状态 |
|---|---|---|
pending |
已入队,尚需首次尝试 | accepted、retrying、exhausted |
retrying |
等待后续尝试 | accepted、retrying、exhausted |
accepted |
已存储回执 | 终态 |
exhausted |
达到尝试次数上限并记录故障 | 终态 |
角色 spec 通过 intent.attempts 设置尝试次数上限,通过 intent.backoff_ms 设置尝试之间的延迟。队列位于客户端数据库中,因此重启后会恢复。
| 回答 | 规则 |
|---|---|
ok = true |
存储回执,标记为 accepted |
duplicate |
重放首次尝试时存储的回执 |
acl_denied、invalid、conflict、forbidden、unknown_role、not_admin、bad_frame、frame_too_large、protocol_version |
删除该行,不重试 |
internal、连接丢失 |
计入此次尝试,并按阶梯延迟后重试 |
达到上限会记录 intent_exhausted 故障类型,并在服务器链接存在后发送一次 report{kind:"fault"}。已耗尽的行绝不会被静默丢弃。
角色 prose
welcome.prose 会缓存在 prose_cache 中,以角色为键,并带有 spec_hash。roles 和集群 prose 导出会读取该记录。