pub struct RunFlags {Show 14 fields
pub modified_files: Vec<String>,
pub modified_file_count: usize,
pub empty_output: bool,
pub no_output_tools: bool,
pub searches_run: usize,
pub searches_empty: usize,
pub max_iterations_hit: usize,
pub gates_forced: usize,
pub required_regions_abandoned: Vec<String>,
pub workspace_lost: bool,
pub produced_output: bool,
pub output_forced: usize,
pub splits_degraded: usize,
pub broken_scripts: Vec<String>,
}Expand description
Post-hoc diagnosis of a run’s productivity, persisted in meta.json so a
harness (or the dashboard) can tell an empty run from a successful one
without inspecting the workspace or parsing logs.
The motivating failure: 13/300 SWE-bench runs completed their whole stage pipeline and produced no file changes at all. Nothing on disk said so, or said why.
Fields§
§modified_files: Vec<String>Paths passed to file-modifying tools that succeeded, in first-touch
order. Capped at MAX_TRACKED_MODIFIED_FILES; modified_file_count
keeps the true total.
modified_file_count: usizeTotal successful file-modifying tool calls across the run (uncapped).
empty_output: boolThe run reached a terminal status having modified nothing, and its
blueprint gave it a way to modify something. See Self::no_output_tools.
no_output_tools: boolNo stage of the blueprint advertised a file-modifying tool, so this run
could never have produced the file changes empty_output looks for.
Recorded because “modified no files” only diagnoses an agent that was supposed to modify files. A router that spawns sub-agents, or an agent whose answer is its text, would otherwise report itself empty on every successful run. The framework has no basis to judge such a run, so it says nothing rather than accusing.
This mirrors the escape the runtime’s gate_blocks already applies per
stage: a require_modifications gate on a stage that advertises no
modifying tool is skipped, because it could never pass.
Phrased negatively so the false that Default and serde(default)
produce means “was capable” - the behavior every meta.json written
before this field had.
searches_run: usizeweb_search calls this run made, across every stage.
Zero for an agent that never had the tool, which is most of them - the
pair says nothing about a run that could not have searched, the same
escape Self::no_output_tools applies to file modifications.
searches_empty: usizeOf those, how many came back with nothing usable: no results, or a diagnostic saying the search could not run.
A research run whose searches all came back empty still finishes
complete and still writes a confident, fully cited report, because a
model handed an empty result set fills the gap from its training data
and cites what it remembers. One did exactly that across 47 consecutive
failed searches - no engine was configured - and nothing on disk
recorded that the report rested on nothing.
searches_empty == searches_run with searches_run > 0 is that run.
max_iterations_hit: usizeHow many stages exhausted their max_iterations.
gates_forced: usizeHow many transitions proceeded past an unsatisfied gate because the gate’s re-run budget ran out.
required_regions_abandoned: Vec<String>Regions declared required that a stage gave up on and that are still
empty, in the order they were abandoned.
The mechanism re-runs the stage a bounded number of times and then
proceeds with a log line, which nothing downstream reads: a run whose
agent wrote its plan and a run where we asked twice and moved on both
finished complete, with the second silently missing the artifact every
later stage’s prompt says to work from. Names rather than a count
because knowing which region was abandoned is what makes it
actionable, and a run cannot abandon many.
A later stage can fill a region an earlier one gave up on, and when that
happens the name is dropped from here - the artifact exists, so a reader
told it is missing would be told something false. A deep-researcher run
abandoned sources_index in gather, analyze wrote it, and the run
finished with a fifty-citation bibliography while still reporting the
region as never written. The moment is kept in the log; this field
answers “what is actually missing”, which is the question a consumer is
asking when it renders a warning.
workspace_lost: boolThe working directory disappeared mid-run.
produced_output: boolThe run submitted a final output.
Counts as having produced something, alongside file modifications.
Without this an agent whose whole deliverable is its answer - a
researcher, a reviewer, a router - reported itself empty on every
successful run, which is the same mistake Self::no_output_tools was
added to correct from the other direction.
output_forced: usizeHow many stages transitioned without the final output they required, because the re-run budget ran out.
The counterpart to Self::gates_forced: the run finished, and this
says the answer it hands back may be missing.
splits_degraded: usizeHow many fan-out splits were unusable and were degraded to an empty
fan-out because the blueprint declared no error and no dead_end
escape from the stage.
Ending the run on a split that cannot be parsed throws away everything the parent has already done, workers finished and later stages still pending. The stage moves on instead, and this is what says the fan-out it moved on from produced nothing. Non-zero means the merge stage worked from less than it was meant to.
broken_scripts: Vec<String>Rhai scripts this run needed that could not be used, by name.
A script that will not compile, or that throws where the runtime has to
carry on regardless, is skipped rather than fatal - which is the right
call and used to be completely silent; this list is the trace. An
output validator that cannot run is the exception: by default the
submission is rejected and the script’s own error goes back to the
model as retry feedback, while on_validator_error = "accept" records
the submission unchecked and leaves this list as the only trace of an
answer nothing checked. The script is named here in both modes.
Named rather than counted, because the useful question is which one - the answer tells you which file to open.
Implementations§
Source§impl RunFlags
impl RunFlags
Sourcepub fn record_modification(&mut self, path: &str)
pub fn record_modification(&mut self, path: &str)
Record a successful modifying tool call on path.
Sourcepub fn note_modified_path(&mut self, path: &str) -> bool
pub fn note_modified_path(&mut self, path: &str) -> bool
Note a path that changed on disk without a modifying tool naming it -
a file a shell command created or rewrote, found by scanning the
working directory. It joins the list (deduped, capped) but does not
touch modified_file_count, which counts modifying tool calls: a
shell call is not one, and a scan that ran twice must not double-count
the same file. Returns whether the path was newly added.