# Round 5 — 2026-09-05: the methods themselves
_Score family: forecast consensus IoU and Brier · base = the E31 recommended configuration · all six fires, nothing chosen per fire · pre-registered TEST_PLAN v1.7 · terms: [GLOSSARY.md](GLOSSARY.md)_
## What we knew before
Rounds 1–4 asked what the *fire model* gets wrong. The recommended
ensemble (assimilating, β 10, σ 0.2, immigrants 0.2, containment only, 32
members) forecasts within 0.03–0.07 of the Circle on five fires. Its
knobs (members, operators, prior) were chosen from one run each, and E31
had just shown that one run can differ from another by 0.02–0.03. Round 5
asks what the *methods* do: the Monte Carlo ensemble, the particle
filter's genetic-algorithm operators, and the offline genetic algorithm.
The fire model is not touched.
**The bar.** E33 measured the run-to-run noise of the base configuration
from five seeds per fire. A difference smaller than that fire's standard
deviation (0.02 on Bear and Chimney, 0.01 on Brattain, Ferguson and Pier,
0.04 on Buck) is a tie; a gain has to clear the bar on several fires at
once.
## What we ran
| [E33](34-e33-noise-floor.md) | How much do two runs of the same configuration differ? | sd ≤ 0.015 on five fires, 0.039 on Buck (one locked-in seed). FINDING, the bar |
| [E32](35-e32-ensemble-size.md) | Does skill keep rising with more members? | 32 is the knee; 8 too few; Brier keeps improving to 128. FINDING |
| [E34](36-e34-operator-ablation.md) | Which operator carries the gain; is crossover worth having? | Plateau: 38 of 54 cells tie; lock-in decides the big cells; crossover 0.5 worth replicating. FINDING |
| [E35](37-e35-prior-width.md) | Does the prior still decide the answer after learning? | Narrow hurts (Bear −0.037); very broad is free. FINDING |
| [E36](38-e36-offline-fit-versus-filter.md) | Learn as it burns, or fit the start and extrapolate? | The filter wins on every fire by 0.025–0.104. FINDING |
| [E37](39-e37-illuminate-the-fire-model.md) | What shapes can the fire model make at all? | A wedge: elongated only while small; Brattain, Ferguson, Pier unreachable. FINDING |
| [E38](40-e38-immigrant-reset.md) | Can a fully contained population recover if immigrants start fresh? | Yes: +0.105 on the locked seed, ties elsewhere, small Brier cost. KEPT as option |
Tooling note: the runner binary was rebuilt (crossover option, driver
re-apply, new modes) while E32 was running, so E32/E35 jobs used two
builds. One E32 job and one E35 job were re-run with the final build and
matched the stored reports to the last digit: the additions are inert in
`assim` mode, and every Round 5 number is reproducible from one binary.
`summarize_r5.py` prints the round's tables with tie marks from the E33
sd. Each experiment file carries one figure, generated by
`scripts/experiments/figures_r5.py` from the result JSON.
## What we know now
1. **The noise floor is small and now known** (E33): sd ≤ 0.015 on five
fires, 0.039 on Buck. Round 4's single-run comparisons were mostly
safe; E28's Bear gain was not (E31), and this round's bar catches that
class of claim.
2. **32 members is the knee** (E32). Eight is clearly too few; past 32
only Buck's IoU and everyone's Brier keep improving. Use 64–128 when
the probability map is the product.
3. **The filter is on an operator plateau** (E34). Immigrants, σ, β and
crossover within a factor of two are ties on most fires; only
Ferguson, whose answer sits at the prior's edge, punishes added genome
noise. Crossover 0.5 is the one setting worth replicating. Keep the
defaults.
4. **Never narrow the prior** (E35). An "informed" ±25 % prior lost 0.037
on Bear; a very broad prior was a tie with better calibration. The
filter recovers from wide, not from narrow.
5. **Learning as it burns beats fitting the start** (E36), on every fire
by 0.025–0.104. Three-day fits drive knobs to the box edges. Offline
fitting is retired as a forecaster; a fitted *start* for the filter is
the one hybrid worth a test (Ferguson's first week).
6. **Three of six fires are outside what the model can draw** (E37). The
reachable growth × elongation region is a wedge: elongated only while
small. Brattain, Ferguson and Pier are big and elongated; wind in this
kernel changes speed, not shape. This converts E12/E19 into an
acceptance test for the E30 kernel refit: after it, the observed dots
must fall inside the shaded region.
7. **The filter's one real failure mode is lock-in** (E33, E34, E38):
every run ends with the population contained and a flat tail, and one
seed in ~30 freezes early enough to lose 0.06–0.10. `immigrant_reset`
repairs it (+0.105 on the locked seed, ties elsewhere) at a small
Brier cost where the fire really stopped. Kept as an option.
Method lessons for `cella_lib::explore` users: measure the noise floor
first (five seeds is enough). Judge every delta against it. Put the prior
wide. Prefer the filter to an offline fit when observations arrive over
time. Use MAP-Elites with no objective as a reachability test before
tuning anything. Watch for state the population cannot leave.
## Still open after this round
In order:
1. **E30** kernel wind-rate refit, with E37's map as the acceptance test.
2. **E39** area-ratio-gated immigrant reset (reset only while the
consensus under-predicts the observed area).
3. Replicate crossover 0.5 over five seeds.
4. The ICS-209 containment check and a fuel term in the containment
operator (from Round 4).
5. A fitted start for the filter on Ferguson-like fires (E36).
## Configuration after this round
Unchanged from Round 4: 32 members, β 10, σ 0.2, immigrants 0.2,
containment only, E25 broad prior. New option `immigrant_reset`, default
off; switch it on for fires still growing after their first week. Five-
seed means: 0.479 / 0.416 / 0.590 / 0.434 / 0.344 / 0.535 (E33).