# CPU Hotspot Profiling Workflow
Maintainer workflow for issue #48 / PRD-0056. Default path is credential-free and offline: `scripts/profile_cpu.py` runs ignored Rust harness tests with a temporary synthetic `MC_HOME`, no `auth.json`, credential-looking environment variables removed from child processes, local SSE fixtures, local session/cache/skill data, and no live model-catalog refresh.
## Requirements
- All platforms: Rust/Cargo, `python3`.
- macOS: built-in `sample` and `spindump`; optional Instruments.
- Linux: `perf`; optional FlameGraph and `hyperfine`.
## Default credential-free run
```sh
python3 scripts/profile_cpu.py --repo . --all --min-seconds 3
```
Automatic wrapper artifacts:
```text
target/profiling/issue-48/profile-results.json
target/profiling/issue-48/profile-summary.md
target/profiling/issue-48/<scenario>.log
target/profiling/issue-48/tui-render-pipeline.json
```
The wrapper does not collect `.sample.txt` files or idle PTY samples. Collect those separately using the manual [macOS sampling](#macos-sampling) workflow.
Run one scenario:
```sh
python3 scripts/profile_cpu.py --repo . --scenario provider_sse_parser --min-seconds 10
cargo test --release profile_cpu_sse_parser -- --ignored --nocapture
```
Self-test:
```sh
python3 scripts/profile_cpu.py --self-test
```
## Mission Control render pipeline
Default render profiling is credential-free and uses Ratatui `TestBackend`; it does not enter raw mode, alternate screen, provider transport, network, or real stdout ANSI output. Historical baseline results live in [`docs/archive/cpu-benchmark-history.md`](../archive/cpu-benchmark-history.md#render-pipeline-history).
Run full harness through Python wrapper:
```sh
python3 scripts/profile_cpu.py --repo . --scenario tui_render_pipeline --min-seconds 5
```
Run all render-pipeline scenarios once through the exact aggregate test (a partial filter also selects six individual tests):
```sh
cargo test --release --lib profiling::profile_harness::render::profile_tui_render_pipeline -- --exact --ignored --nocapture
```
Run one scenario directly:
```sh
cargo test --release profile_tui_render_pipeline_idle -- --ignored --nocapture
cargo test --release profile_tui_render_pipeline_streaming -- --ignored --nocapture
cargo test --release profile_tui_render_pipeline_heavy_transcript -- --ignored --nocapture
cargo test --release profile_tui_render_pipeline_scroll -- --ignored --nocapture
cargo test --release profile_tui_render_pipeline_resize -- --ignored --nocapture
cargo test --release profile_tui_render_pipeline_modal -- --ignored --nocapture
```
Direct Cargo runs print results to stdout; they do not write a JSON artifact. The Python wrapper writes:
```text
target/profiling/issue-48/tui-render-pipeline.json
```
Backend modes are selected by `PROFILE_TUI_RENDER_BACKEND`:
```sh
PROFILE_TUI_RENDER_BACKEND=test cargo test --release --lib profiling::profile_harness::render::profile_tui_render_pipeline -- --exact --ignored --nocapture
PROFILE_TUI_RENDER_BACKEND=crossterm-sink cargo test --release --lib profiling::profile_harness::render::profile_tui_render_pipeline -- --exact --ignored --nocapture
PROFILE_TUI_RENDER_BACKEND=crossterm-real cargo test --release --lib profiling::profile_harness::render::profile_tui_render_pipeline -- --exact --ignored --nocapture
```
| backend mode | purpose | caveat |
| --- | --- | --- |
| `test` | default Ratatui `TestBackend`; measures process-side render and diff cost | records 0 for stdout/write/flush; no terminal emulator paint |
| `crossterm-sink` | measures generated ANSI bytes and crossterm write/flush work against an in-memory sink | still does not measure terminal emulator paint |
| `crossterm-real` | writes to real stdout for manual terminal-path checks | env-gated; use only in disposable terminal; manual real-TTY validation has not been recorded in baseline |
Metrics:
| metric | meaning | interpretation |
| --- | --- | --- |
| `terminal_draw_ms` | end-to-end `Terminal::draw` wall time in default `TestBackend` mode | closest process-side total draw number for baseline tables |
| `controlled_draw_ms` | phase-instrumented manual frame path wall time | use when comparing `render_draw_ms`, `diff_ms`, backend draw, write, and flush phases |
| `render_draw_ms` | Mission Control `render::draw` CPU time | historical baseline bottleneck: 71-82% of p50 frame time; rerun to assess current cost |
| `diff_ms` | Ratatui buffer diff duration | not the historical baseline bottleneck: 0.049-0.111 ms |
| `diff_cells` | changed-cell count emitted by Ratatui diff | high values can raise terminal output, but baseline resize still diffed 7995 cells in 0.111 ms p50 |
| `changed_cell_ratio` | changed cells divided by terminal area | resize/modal first frames should be high; steady idle should be near zero |
| `backend_draw_ms` | backend `draw` call duration inside controlled path | zero/near-zero for `TestBackend`; useful for crossterm modes |
| `stdout_bytes` | ANSI/output bytes generated by crossterm backend | zero for `TestBackend`; terminal I/O ceiling needs crossterm modes |
| `write_calls` / `write_ms` | writer call count and write duration | zero for `TestBackend`; sink/real backends isolate output overhead |
| `flush_calls` / `flush_ms` | flush count and flush duration | process flush time is not terminal emulator paint telemetry |
| `frames_per_sec` | scripted harness throughput | not display refresh rate; compare same machine/profile/backend/area only |
Interpretation rules:
1. Compare same backend, terminal area, Rust profile, scenario mix, and machine.
2. If `render_draw_ms` dominates p50 frame time, optimize Mission Control render/state projection first.
3. If `diff_ms` stays sub-millisecond while `diff_cells` is high, Ratatui diff is not blocking frame throughput.
4. If `stdout_bytes`, `write_ms`, or `flush_ms` rise in crossterm modes, terminal I/O may cap real-world FPS even when `TestBackend` is fast.
5. Treat `TestBackend` results as process-side render ceiling only. Real-world ceiling depends on terminal emulator write/flush/paint behavior.
## Typing while output is queued
```sh
cargo test --release --locked --lib profile_typing_under_streaming_load \
-- --ignored --nocapture --test-threads=1
```
This offline diagnostic uses existing controller drain, input-classification, input-reduction, and draw components with `TestBackend` at 120 × 36. It compares idle typing, a 1,024-event streaming queue, and the same queue with 120 prior synthetic turns. Each scenario excludes 10 warmup samples and emits 100 raw rows as `typing_under_load` JSON. Setup and assertions stay outside timing; assertions verify the edited draft appears in the drawn prompt while streaming output remains queued.
Columns: `drain_ms`, `edit_to_draw_ms`, `drain_then_edit_to_draw_ms`, `events_drained`, `backlog_after_draw`. Edit timing starts immediately after the drain, before applying its scheduling state and classifying input. The draft resets between samples; the streaming response grows throughout the run. These are correlated component-workflow observations, not independent whole-session trials.
The diagnostic scripts component order, not the complete event loop. It excludes input capture/queue wait, timer pacing, concurrent producer contention, terminal writes and physical display. No latency threshold gates ordinary tests. Repeat separate invocations when comparing changes; retain raw rows and machine/build context.
Initial measurements and interpretation limits: [interaction baseline](../archive/interaction-baseline-2026-09-27.md).
## Real application PTY workflows
```sh
cargo build --release --locked --bin magi-code
python3 scripts/profile_interaction.py --cycles 60 --idle-seconds 10 \
--out-dir target/profiling/pty-interaction/new-run
```
`profile_interaction.py` launches a private snapshot of the actual CLI at 160 × 48 cells with disposable working directory, `HOME`, and `MC_HOME`. The report hashes the executed copy rather than a shared build path. An allowlisted environment and nonforwarding loopback proxy keep configured provider traffic local without supplying credentials. The scripted provider exercises text replies, shell-tool continuation, streaming and process-tree cancellation, and recovery prompts. This is not an OS network sandbox.
Reports include draft-write-to-output timing, submission/request/output boundaries, cancellation feedback and observed cleanup, startup model-label observation, application RSS, idle CPU/output, and shutdown checks. Application CPU and RSS exclude the observer and tool children. Separate observer-thread CPU counters cover turns and idle windows; they are not disjoint wall-time stages. Plain content with delayed completion is separate from streaming content prefixed by `</think>`, matching the current Chat Completions reasoning parser.
Limits: Unix PTYs and `ps`, verified on macOS only; no new dependencies. `--cycles` accepts 1-200 (default 20), `--idle-seconds` accepts 1-300 (default 10). A unique output directory is generated when omitted; explicitly named directories must not exist. Each cancellation wait and process fixture is bounded. Fixture cleanup never signals reported callback PIDs.
Require command exit 0 and `valid: true` in `results.json`. Missing workflow evidence, unexpected process exit, fixture failures, surviving reported helpers, forced shutdown, or failed terminal restoration invalidate the run. Raw failure output is capped at 2 MiB. The fixed-grid ANSI observer waits for complete synchronized updates but approximates Unicode widths: these are reconstructed-output observations, not physical pixels or terminal-compatibility results. Small draft/cancellation sample counts and bounded turn loops do not establish tail latency or leak freedom.
Optional diagnostics on macOS:
```sh
python3 scripts/profile_interaction.py --cycles 60 --idle-seconds 10 --diagnostics \
--out-dir target/profiling/pty-interaction/new-profiled
```
`--diagnostics` adds `application.sample.txt` (10 ms sampling interval, 120-second maximum), `observer.prof` (main Python thread only), and bounded raw/screen captures for both idle windows. Captures record truncation; native sample files can contain local paths. A failed or missing native sample invalidates the run. Sampling starts after initial startup and may cover only a prefix of long workloads. Profiles perturb timing: keep diagnostic runs separate from ordinary baselines, and do not sum nested cumulative profile times. Open the observer profile with `python3 -m pstats <run>/observer.prof`.
Add `--reduced-motion` to disable animation in the disposable fixture only. Compare both orders when attributing output to animation; this does not change user settings or measure physical-terminal cost.
For observer-side delay attribution, add `--trace-timing` to ordinary or diagnostic runs. It records monotonic select/read/processing timestamps, processing-thread CPU, byte counts, input-write completion and first literal marker bytes. Coverage includes initial scheduled cycles plus lifecycle turns, `/new`, compaction and session switching. Each trace ends when reconstructed output confirms its marker; pacing, idle, cancellation and resource counters stay outside. Differential ANSI output can split a marker with escapes, so absent literal bytes do not mean absent feedback. Select waits include scheduling; processing wall time can include descheduling or cursor-query writes. These traces alone cannot distinguish application waiting, kernel delivery and host contention.
Trace storage is capped at 20,000 pump rows across the run; truncation or missing operation traces invalidate the run. Keep traced and untraced repetitions separate and compare both orders.
On macOS, add `--stall-threshold-ms 250 --trace-timing` for concurrent application and observer native samples when a traced operation crosses the threshold. Accepted threshold: 25-5000 ms. A watchdog checks every 10 ms and starts at most four capture pairs per run, each requesting one second at 1 ms intervals. Both samplers start before either is awaited; samples and logs use `stall-N-{application,observer}` filenames. Reports retain operation, trigger, launch and reap timestamps, target PIDs, commands, exit codes and file sizes. Sampler failures invalidate the run; child waits share a 15-second deadline, followed by kill/reap cleanup. This mode cannot combine with continuous `--diagnostics` sampling.
Native `sample` output is an aggregate, not timestamped stacks. Launch time precedes attachment/collection; sampling may begin after a short stall ends and extend into later work. The watchdog shares Python's process and can itself be delayed by scheduling or the interpreter lock. Captures pause no application thread intentionally, but attachment and sampling perturb timing. Keep low-threshold capture-validation runs separate from natural-stall evidence. A run without captures does not establish absence of stalls beyond traced operations or the four-pair limit.
Optional paced session lifecycle/resource measurement on macOS:
```sh
python3 scripts/profile_interaction.py --cycles 20 --idle-seconds 10 \
--lifecycle-rounds 6 --lifecycle-pause 20 \
--out-dir target/profiling/pty-interaction/new-lifecycle
```
`--lifecycle-rounds` accepts 0-12 (default 0). After the ordinary workflows, it creates a second session and repeats 20 turns, `/compact`, a continuation prompt, session-picker switching, and another prompt. It verifies durable inputs/replies in the selected session, the compaction checkpoint, and summary use in subsequent requests. Summaries are synthetic: correctness checks do not assess semantic quality or archival recovery.
`--lifecycle-pause` accepts 1-60 seconds (default 30) after each compaction and switch. Lifecycle mode also waits one second before provider submissions, outside latency timing; it does not test rapid resubmission. macOS `ps -M` and `lsof` provide thread and numeric file-descriptor counts at workflow boundaries and approximately every five seconds during these pauses. Each counter command has a five-second timeout. Counts are sequential, application-only samples and can miss brief resources; mapped-file and working-directory entries are not numeric descriptors. Active JSONL byte counts exclude archives. Synthetic session inspection is capped at 32 MiB per active file.
Measured results and limitations: [PTY interaction baseline](../archive/pty-interaction-baseline-2026-09-27.md), [paired profiling and 200-turn follow-up](../archive/pty-profiling-followup-2026-09-27.md), [eight-minute session lifecycle runs](../archive/session-lifecycle-resources-2026-09-27.md), [first-display/resize, observer tracing and 19.5-minute investigation](../archive/performance-investigation-2026-09-27.md).
## Large-response and text-modal comparisons
These ignored, credential-free workflows compare current render paths with local reference algorithms, not historical binaries or terminal-emulator paint:
```sh
cargo test --release --locked --lib streaming_viewport_large_response_workload -- --ignored --nocapture
cargo test --release --locked --lib text_modal_large_document_workload -- --ignored --nocapture
```
Streaming uses independent states with identical synthetic events and alternates measurement order. It reports initial indexing separately, then medians across 30 appends. `rebuild` disables only completed-prefix reuse; `full_reference` materializes and measures every line. Text-modal results include `Terminal::draw` with `TestBackend`, across 21 top/middle/end draws.
For repeated first-display and resize observations:
```sh
cargo test --release --locked --lib profile_large_response_first_display_and_resize \
-- --ignored --nocapture --test-threads=1
```
The diagnostic rotates three fixture types across ten fresh display states each. For every state it measures widths `80, 80, 120, 120, 40, 80` at 24 rows: first projection, cached projection, two newly encountered widths, and return to retained width. Ingestion, assertions and result cleanup are outside timing. It verifies visible content survives resize and returns unchanged at width 80. These are projection costs, not complete frame/input latency; no performance threshold gates tests. Repeat fresh process invocations and report their medians separately from single samples.
Representative local results on macOS with Rust 1.98.1, release profile, September 2026:
| Synthetic workload | Viewport median | Sparse rebuild median | Full-reference median |
| --- | ---: | ---: | ---: |
| 390,000-byte multiline ASCII, 80×24 | 0.89 ms | 8.00 ms | 36.53 ms |
| 430,000-byte multiline Unicode, 80×24 | 1.30 ms | 45.57 ms | 38.33 ms |
| 276,000-byte unbroken Unicode, 80×24 | 9.96 ms | 9.96 ms | 10.08 ms |
| Large system-prompt text, 120×40 | 5.66 ms | — | 16.29 ms |
The multiline Unicode cold index took 45.52 ms versus 38.24 ms for full formatting. Completed-prefix reuse—not faster cold Unicode wrapping—produces the steady append improvement. Reuse requires exact equality of the fully sanitized completed prefix and compatible row/layout identity. Unbroken responses retain full formatting because sparse rebuilding measured slower. Whole-source sanitization, cold indexing, final Markdown, explicit copying, and modal height/prefix scans remain full-work paths. Allocations and physical-terminal latency were not measured.
## Scenarios
| scenario | harness | exercises |
| --- | --- | --- |
| startup_discovery | `profile_cpu_startup_discovery` | `src/config`, `src/instructions`, `src/skills`, `src/sessions`, `src/model_catalog` local cache reads |
| provider_sse_parser | `profile_cpu_sse_parser` | `src/providers/stream.rs::StreamParser` using Responses and Chat Completions SSE fixtures |
| rendering_heavy_transcript | `profile_cpu_rendering_heavy_transcript` | `src/tui/state`, `src/tui/render/transcript.rs`, `src/rendering/markup.rs`, `src/rendering/highlight.rs` |
| tui_render_pipeline | `profile_tui_render_pipeline` | full Mission Control `Terminal::draw`, controlled `render::draw`, Ratatui diff, TestBackend/crossterm byte and flush accounting |
| tui_streaming_simulation | `profile_cpu_tui_streaming_simulation` | assistant deltas, activity events, transcript/activity caches |
| tool_timeout_cleanup | `profile_cpu_tool_timeout_cleanup` | `src/tools/process.rs::terminate_child_tree_and_wait` process-group cleanup |
| idle Mission Control TUI | manual real terminal or non-interactive PTY | static `--no-session` wakeups/CPU while no prompt is submitted |
## Issue #381 skill discovery benchmark
The dedicated release benchmark uses 55 selected skills plus three deterministic diagnostics across small, empty-body, representative-body, near-1-MiB, disabled, malformed, colliding, and oversized fixtures. The fixture contains 5,439,395 aggregate body bytes. That count describes generated fixture content; bytes read and allocations were not instrumented and are not claimed.
Both runs used `PROFILE_CPU_ITERATIONS=25 cargo test --release profile_cpu_discover_skills -- --ignored --nocapture` on a MacBook Pro `Mac17,8` with Apple M5 Pro, 64 GB RAM, macOS 27.0 arm64, and `rustc 1.98.0`, release profile.
| run | exact commit | elapsed / iterations | first `read skill://near-0` | saved artifact |
| --- | --- | ---: | ---: | --- |
| metadata-body baseline | `4085f5f4ed51f0b219b7ca03e4e86ae97baea283` | 325 ms / 25 | 687 µs | `target/profiling/issue-381/before.log` |
| metadata-only discovery | `40d6f1410c6097e12b352afd88bb34d87402165c` | 74 ms / 25 | 723 µs | `target/profiling/issue-381/after-final.log` |
The dedicated result improved from 13.0 ms to 3.0 ms per iteration (about 77%). This is a same-machine synthetic comparison, not evidence of a user-visible startup regression or a hard budget. As secondary context only, the aggregate `startup_discovery` harness completed 25 iterations in 42 ms at the optimized commit (`target/profiling/issue-381/startup-after.log`); it also measures settings, instructions, sessions, and model-cache work, and the committed Issue #48 profile identifies session listing as dominant and skill discovery as secondary.
## macOS sampling
Harness sample:
```sh
PROFILE_CPU_MIN_SECONDS=20 PROFILE_CPU_ITERATIONS=1 \
cargo test --release profile_cpu_rendering_heavy_transcript -- --ignored --nocapture &
cargo_pid=$!
sleep 1
pgrep -P "$cargo_pid" -fl magi_code
test_pid=<pid-from-pgrep>
sample "$test_pid" 10 -mayDie -file target/profiling/issue-48/rendering-heavy-transcript.sample.txt
wait "$cargo_pid"
```
Idle Mission Control TUI real-terminal command:
```sh
MC_HOME="$(mktemp -d)" cargo run --bin magi-code -- --no-session
# In another terminal:
pgrep -fl "magi-code.*--no-session"
sample <pid> 30 -file target/profiling/issue-48/idle-tui.sample.txt
spindump <pid> 10 -file target/profiling/issue-48/idle-tui.spindump.txt
```
Idle rules: do not submit a prompt, do not run `/login`, do not configure provider credentials, do not refresh model catalogs, exit with `/quit` after sampling. A non-interactive `/usr/bin/script` PTY baseline is acceptable for issue #48 if it launches `magi-code --no-session`, keeps stdin open, uses isolated `MC_HOME`, captures `sample`, and is labeled as PTY rather than physical terminal evidence. Recorded baselines below retain the launch flags used at the time.
## Linux perf parity
The Linux commands provide an equivalent workflow; they have not been run unless a Linux host is named in the baseline notes:
```sh
PROFILE_CPU_MIN_SECONDS=20 PROFILE_CPU_ITERATIONS=1 \
perf record -F 99 -g -- cargo test --release profile_cpu_tui_streaming_simulation -- --ignored --nocapture
perf report --stdio > target/profiling/issue-48/tui-streaming.perf-report.txt
perf script > target/profiling/issue-48/tui-streaming.perf-script.txt
```
Optional FlameGraph:
```sh
perf script | stackcollapse-perf.pl > target/profiling/issue-48/tui-streaming.folded
flamegraph.pl target/profiling/issue-48/tui-streaming.folded > target/profiling/issue-48/tui-streaming.svg
```
## Issue #48 historical baseline
Machine baseline from `target/profiling/issue-48/profile-results.json` generated on 2026-05-31T07:32:23Z. macOS stack samples were collected with `/usr/bin/sample` for all synthetic harness scenarios plus idle TUI under a non-interactive `/usr/bin/script` PTY. The Linux profiler was not run on this Darwin host.
| scenario | date/time | machine/OS/arch | Rust version/profile | command | duration/iterations | primary metric | result | artifact path | notes |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| startup_discovery | 2026-05-31T07:32:23Z | Darwin 25.5.0 arm64 | rustc 1.95.0 release | `cargo test --release profile_cpu_startup_discovery -- --ignored --nocapture` | 3000 ms / 4401 | discovered_items | 563618466 | `target/profiling/issue-48/startup_discovery.log`; `target/profiling/issue-48/startup_discovery.sample.txt` | synthetic settings, sessions, skills, model cache |
| provider_sse_parser | 2026-05-31T07:32:23Z | Darwin 25.5.0 arm64 | rustc 1.95.0 release | `cargo test --release profile_cpu_sse_parser -- --ignored --nocapture` | 3000 ms / 102104 | events_parsed | 1735768 | `target/profiling/issue-48/provider_sse_parser.log`; `target/profiling/issue-48/provider_sse_parser.sample.txt` | local Responses + Chat Completions fixtures |
| rendering_heavy_transcript | 2026-05-31T07:32:23Z | Darwin 25.5.0 arm64 | rustc 1.95.0 release | `cargo test --release profile_cpu_rendering_heavy_transcript -- --ignored --nocapture` | 3001 ms / 263 | rendered_units | 956268 | `target/profiling/issue-48/rendering_heavy_transcript.log`; `target/profiling/issue-48/rendering_heavy_transcript.sample.txt` | markdown/code/diff transcript fixture |
| tui_streaming_simulation | 2026-05-31T07:32:23Z | Darwin 25.5.0 arm64 | rustc 1.95.0 release | `cargo test --release profile_cpu_tui_streaming_simulation -- --ignored --nocapture` | 3000 ms / 1059 | state_units | 277458 | `target/profiling/issue-48/tui_streaming_simulation.log`; `target/profiling/issue-48/tui_streaming_simulation.sample.txt` | synthetic assistant deltas and activity deltas |
| tool_timeout_cleanup | 2026-05-31T07:32:23Z | Darwin 25.5.0 arm64 | rustc 1.95.0 release | `cargo test --release profile_cpu_tool_timeout_cleanup -- --ignored --nocapture` | 3020 ms / 127 | cleanup_units | 254 | `target/profiling/issue-48/tool_timeout_cleanup.log`; `target/profiling/issue-48/tool_timeout_cleanup.sample.txt` | Unix process-group termination path |
| idle Mission Control TUI | 2026-05-31T07:24:36Z | Darwin 25.5.0 arm64 | release binary | `/usr/bin/script -q target/profiling/issue-48/idle-tui-pty.script.txt target/release/magi-code --tui --no-session` + `/usr/bin/sample <pid> 8 -mayDie` | 8 s sample / PTY idle | ps CPU + stack sample | 0.4% CPU before sample, 0.1% CPU after sample; 6832/6853 samples in `kevent` | `target/profiling/issue-48/idle-tui-pty-summary.json`; `target/profiling/issue-48/idle-tui-pty.sample.txt` | isolated temp `MC_HOME`; no prompt submitted; no provider request; no `auth.json` created; PTY baseline, not physical terminal |
## Hotspot report
Historical findings from the Issue #48 baseline above. Source paths, symbols, and priorities reflect that run, not a current bottleneck assessment.
| scenario | source file | function/symbol | metric type | metric value | profiler artifact | impact | confidence | follow-up priority |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| idle Mission Control TUI | `src/tui.rs` + crossterm event path | `MissionControlApp::run` / `crossterm::event::poll` / `kevent` | macOS `sample` stack + `ps` CPU | 6832/6853 samples blocked in `kevent`; 0.4% -> 0.1% CPU around 8 s sample | `target/profiling/issue-48/idle-tui-pty.sample.txt`; `target/profiling/issue-48/idle-tui-pty-summary.json` | idle loop appears mostly blocked on terminal events under PTY; low idle CPU on this machine | High for PTY evidence; physical-terminal wakeups still optional confirmation | P2 |
| rendering_heavy_transcript | `src/tui/render/transcript.rs` | `cached_transcript_lines` / `append_transcript_entry_lines` | macOS `sample` stack | 3225 samples in `MissionControlState::cached_transcript_lines`; 3191 in `cached_transcript_lines`; 3163 in `append_transcript_entry_lines` | `target/profiling/issue-48/rendering_heavy_transcript.sample.txt` | heavy transcript redraw cost; markdown projection repeated across many assistant turns | High; stack sample plus harness throughput | P1 |
| rendering_heavy_transcript | `src/rendering/markup.rs` + `src/rendering/highlight.rs` | `render_markdown` / `highlight_code` | macOS `sample` stack | 3055 samples in `render_markdown`; 2391 samples in `highlight_code` | `target/profiling/issue-48/rendering_heavy_transcript.sample.txt` | code fence highlighting and Markdown parsing dominate long transcript rendering | High | P1 |
| tui_streaming_simulation | `src/tui/state.rs` | `apply_activity_event` / `cached_visible_nodes` / `collect_visible` | macOS `sample` stack | 1164 samples in `apply_activity_event`; 1137 in `cached_visible_nodes`; 447 in `collect_visible` | `target/profiling/issue-48/tui_streaming_simulation.sample.txt` | many small activity deltas clone and rebuild visible activity state | High | P1 |
| provider_sse_parser | `src/providers/stream.rs` | `StreamParser::push_chunk_outcome` / `serde_json::de::from_trait` | macOS `sample` stack | 376 samples in `push_chunk_outcome`; 369 in serde JSON parse beneath it | `target/profiling/issue-48/provider_sse_parser.sample.txt` | parser overhead is real but smaller than rendering/state hotspots in synthetic run | High | P2 |
| startup_discovery | `src/sessions.rs` + `src/skills.rs` | `SessionManager::list` / `discover_skills` / `discover_agents` | macOS `sample` stack | 5750 samples in `SessionManager::list`; 303 + 301 in `discover_skills`; 83 + 79 in `discover_agents` | `target/profiling/issue-48/startup_discovery.sample.txt` | startup scales with local session count first, then skill/instruction discovery | High | P2 |
| tool_timeout_cleanup | `src/tools/process.rs` | `terminate_child_tree_and_wait` / `poll_process_group_exit` / `signal_process_group` | macOS `sample` stack | 4525 samples in `terminate_child_tree_and_wait`; 3772 in `poll_process_group_exit`; 693 in `signal_process_group` | `target/profiling/issue-48/tool_timeout_cleanup.sample.txt` | timeout cleanup cost dominated by sleep/process polling and subprocess probes | High | P3 |
## Historical CPU-reduction priority order
1. P1: reduce repeated transcript markdown/render projection in heavy transcript and streaming paths. Expected user impact: smoother long-answer TUI and lower redraw CPU.
2. P1: reduce activity visible-node clone/rebuild churn in streaming paths. Expected user impact: lower CPU during tool-heavy runs.
3. P2: inspect startup session listing cost when many JSONL files exist; keep session semantics unchanged.
4. P2: keep idle loop under observation with optional physical-terminal wakeup sample; PTY evidence shows low CPU and no P0 idle blocker.
5. P3: process cleanup polling changes are riskier because they touch safety semantics; optimize only with process-tree regression coverage.
## Before/after comparison method
1. Use the same machine, branch base, and Rust profile, and the same terminal size for TUI checks.
2. Run:
```sh
python3 scripts/profile_cpu.py --repo . --all --min-seconds 10
cargo test --release profile_cpu -- --ignored --nocapture
```
3. Profile the changed area with a platform profiler:
```sh
sample <pid> 10 -file target/profiling/issue-48/<scenario>.sample.txt
# or Linux:
perf record -F 99 -g -- cargo test --release <scenario> -- --ignored --nocapture
```
4. Compare baseline table fields: `scenario`, `date/time`, `machine/OS/arch`, `Rust version/profile`, `command`, `duration/iterations`, `primary metric`, `result`, `artifact path`, `notes`.
5. Compare hotspot table fields: `scenario`, `source file`, `function/symbol`, `metric type`, `metric value`, `profiler artifact`, `impact`, `confidence`, `follow-up priority`.
6. Report delta as `(after - before) / before`, plus whether profiler top symbols moved away from target function.
## Optional live-provider profiling
Live provider profiling is not part of the default workflow. Skip it unless credentials and network are intentionally available. Do not store provider captures, bearer-style headers, provider secret strings, account identifiers, authorization codes, or secret payloads in fixtures or artifacts.