The v1.5.0 re-baseline campaign — plan¶
Status: the campaign is COMPLETE (2026-08-14). All three measurement phases
have run and are published. The default is flipped
(#411), the six-cell grid
is re-baselined at a1daa7a and republished on grid.md, the
cap × detector set ran at 7471447 and is published on
reinjection.md — the detector A/B and the ≥ 59 % claim both
confirmed at 20 seeds, the cap sweep landing on its pre-registered "maintainer
decides" branch with no default change — and phase 3's oracle control is
measured at 40b434d across all six grid cells plus the satellite torus, with
its three a-priori assertions passing at 20 seeds and its 480-row baseline
byte-identity control clean. The headline it existed to produce:
on two-ray the channel costs nothing, so 100 % of AntHocNet's 7.0–9.9 pp
delivery shortfall is protocol overhead; under fading the channel costs
0.45–0.59 pp and the routing share is still 95.5–95.8 %
(grid.md). This page
exists so the campaign is designed once, in the open, before the runtime is
spent — and so the order of operations is a written artefact rather than a
sequence of decisions made under the pressure of a running job.
← Benchmark index · Methodology · Roadmap
Why a plan, and not just a dispatch¶
Four independent workstreams now want to measure the same scenario set, at the same time, at defaults that are themselves about to change:
| workstream | issue | what it needs measured |
|---|---|---|
| hold-cap default flip (1 s → 200 ms) | #371 | the whole published corpus, at the new default |
| re-injection cap default | #386 | cap ∈ {1, 2, ∞} on the fading cells |
| publishable detector A/B | #386 | detector ON/OFF at 20 seeds (#293 floor) |
| oracle control | #296 | a new routing arm in both suites |
Run naively — each workstream dispatching what it needs when it is ready — the same cells get measured three or four times, at three or four different sets of defaults, and the results are not comparable with each other. That is the #365 measure-twice failure, prospectively. This page sequences them into one campaign so every number lands at one configuration, once.
The corpus this campaign replaces is v1.4.0's. Per the
provenance rule,
the merge that flips a shipped default invalidates the published corpus and must
say so — this campaign is what re-establishes it.
Phase order, and why it is this order¶
The dependencies are hard: each phase changes what the next phase measures, so no two phases may be swapped or run concurrently.
Phase 0 — prerequisites that are not campaign runs¶
Cheap work that must land first, because it changes what the campaign can read. All four items are complete (2026-08-11) — the record of each is appended in place below.
- #402 — cap-aware drop-cause attribution. Until it lands,
drop_mac/drop_chan/sumare unreadable on any capped arm (scenario_checkFAILs them, correctly: +8.50 pp residue at cap=1 against +3.32 pp uncapped). FlowMonitor metrics are unaffected in every arm, so this blocks only the drop-cause columns of phase 2 — but it blocks them absolutely, and diagnosing it needs no dispatch (the archived cap=1 block carries the books). Done — fixed by #407 (cap-aware mac numeratormacTerminal − skipsDelivEvents, in event units): the cap=1 sum went 108.49 → 101.57, and capped arms are readable for every column. - Anchors review (#59). The regression floors were set against the 1 s hold cap. If the default flips, floors calibrated on the old default will either mis-fire or silently pass. Decide per anchor: re-derive, or state why it is default-independent. Done — per-anchor verdict recorded on #371: all calibrated floors are AODV-gated and proven default-independent (SHA-256-identical across cap arms). Because anchors also cannot catch a botched flip, the phase-1 pre-merge A/B expectations are pre-registered there.
- #229 — DSDV's
dense-regime queue drops land in no bucket (−9.84 pp at
dense-smallin the last committed measurement, and the fix the accounting rework never touched: DSDV's internalPacketQueuesheds packets above every book). A P1 identity defect in a published-page scenario cannot ride into the re-baseline unfixed — phase 1's grid would republish a known-broken DSDV breakdown. The settle measurement is cheap (rescue the per-mergebenchmark-resultsartifact, or one preflighted dense-small dispatch); the fix is a DSDV drop-trace hook with the standing byte-identity probe. Done — #409 landed the loopback-Tx conservation hook (mergede1014fe), gated by the standing byte-identity probe. The settle measurement came from exactly the predicted rescue (effectiveness read): dense-small dsdvdrop_queue_pct0.000 → 21.235 with sum 78.18 → 99.41, and aodv's deferred-queue sheds reclassified route → queue (52.758 → 0.064, queue 0.000 → 55.018, sum 97.64 → 99.96), on byte-identical non-queue columns. - Sweep pre-registration written on #386, stating readability per arm explicitly (which columns each arm may be read for) and the refutation conditions, before phase 2 dispatches. Done — posted as #386 (comment): per-arm readability, the Q1 cap-adoption / rejection criteria, and the Q2/Q3 withdrawal conditions for the +6.4 pp and ≥59 % claims.
Phase 1 — flip the hold-cap default, then re-baseline¶
The #371 decision is already taken: 200 ms is the measured-better operating point and the flip was deferred to this campaign precisely so the re-baseline measures it once.
- Land the default change (
ReconvHoldCap1 s → 200 ms) with the standing pre-merge A/B on identical seeds — a protocol-behaviour change, so the AGENTS.md rule applies. Done — merged as #411 (a1daa7a), all botched-flip gates PASS (reconvMax 192.8–199.9 ms, setupMax 1105–2924 ms untouched, aodv/olsr/dsdv byte-identical across arms). Measured effect at paper-base/disk, 20 seeds: paired ΔPDR −4.38 [−4.80, −3.96] pp, delay99 862 → 517 ms (−40.1 %), NRL down — inside the phase-0 pre-registered envelope, smaller than the ablation-derived prediction in the pre-caveated direction (the ablation arm movedRepairHoldCaptoo). Full A/B verdict on the #411 PR thread. - Re-baseline at the new default: the six-cell grid
({
rwp,ssrwp,gaussmarkov} × {tworay,nakagami}) at 20 seeds, which is the corpus every headline claim rests on. Done — six cells × 20 seeds dispatched onmainata1daa7a, runs 31618105814 (rwp-tworay), 31618108070 (rwp-nakagami), 31618110426 (ssrwp-tworay), 31618114426 (ssrwp-nakagami), 31618116286 (gaussmarkov-tworay), 31618118283 (gaussmarkov-nakagami). Baseline attribution control: 0/18 aodv/olsr/dsdv rows moved vs the4cdfb96corpus, so every delta is the flip's. Grid-wide trade: tail −20 % to −48 %, delivery −2.3 to −6.9 pp, overhead flat-to-down. Two of grid.md's ranking sentences were obsoleted and rewritten (two-ray "…anthocnet last" and "aodv wins the fading tail"); the corpus is re-established and grid.md republished ata1daa7a. Full readout with per-cell tables and re-checked headline claims: #371 (comment). - The three sweeps (pause / area / scale) are deliberately not re-run here.
#365 was closed by accepting the
v1.3.0pins; nothing in this campaign changes that trade. A sweep gets re-measured when a claim needs its shape at the new default — and then it is a phase of its own, not a rider on this one.
Why first: every later phase's numbers are quoted against the shipped default. Measuring the cap sweep at 1 s and then flipping to 200 ms would invalidate the cap decision the day it was made.
Phase 2 — the re-injection cap sweep and the detector A/B, together¶
One dispatch set answers both, because they share arms:
| arm | MaxReinjectPerPacket |
EnableMacFailureDetector |
|---|---|---|
| A | 0 (unlimited) | true — the ∞ arm and the A/B's ON arm are the same arm |
| B | 1 | true |
| C | 2 | true |
| D | — | false — the A/B's OFF arm |
× {rwp, gaussmarkov} × nakagami × 20 seeds (four-way firstRun splitting
per cell, #126) — the
fading cells only, since the mechanism does not engage under two-ray
(reorder ratio 0.0004; reading 2).
Sharing arm A between the sweep and the A/B is what makes this affordable: the publishable detector comparison (A vs D) falls out of the sweep at no extra runtime, at the post-flip defaults it should be measured under anyway.
Readability per arm (the thing the pre-registration must state, and the
reason phase 0 item 1 exists): arms A and D may be read for every column; arms
B and C may be read for FlowMonitor metrics and the ##REINJ## books, and for
drop-cause columns only if #402 has landed.
Done (2026-08-13) — all 8 cells ran on main @ 7471447, time=900,
runs=20 firstRun=1, nakagami, gaussmarkov at pause=0; every cell's
##PROV## reads commit=7471447 with ReconvHoldCap=+2e+08ns. Published as
reinjection.md; full readout on
#386 (comment).
- What ran, and the corrected run-ID table.* The dispatch comment's
run-ID → cell table was wrong for 4 of 8 cells* — it assigned run IDs by
assuming the two mobilities interleaved, when all four
rwparms went first and then all fourgaussmarkovarms. The design is complete and balanced (4 arms × 2 mobilities × 20 seeds, all present); only the labelling was wrong, and it was caught because every cell identity was re-derived from its own##CONFIG##rather than from dispatch order. The corrected table is the one on reinjection.md § Provenance:
| arm | cap | detector | rwp | gaussmarkov |
|---|---|---|---|---|
| A | 0 (unlimited) | true |
31662269404 | 31662277724 |
| B | 1 | true |
31662271976 | 31662280223 |
| C | 2 | true |
31662274139 | 31662282155 |
| D | — | false |
31662275527 | 31662284121 |
- Deviation from the plan above, recorded rather than smoothed over.* This
section planned 32 dispatches via #126
four-way
firstRunsplits (and the affordability table below still counts them that way). What ran was 8 dispatches ofruns=20 firstRun=1*. Seeds 1–20 are present in every cell, so validity is unaffected and the samples are the intended ones — but plan and artefact disagree, and the artefact is what the tables are read from. - Validation.
scenario_check.py resultson all 8: zero anthocnet FAILs anywhere; every FAIL is the known non-blocking #230 path-diversity FAIL on the three baselines. Baselines byte-identical across arms A/B/C/D within each mobility (all 63 rows), so the #51-class control passes and every anthocnet delta is attributable to the knob. The cap's self-check passes by construction:skipsexactly 0 in arm A (unlimited) and arm D (no events), non-zero in B/C with cap=1 > cap=2. - The phase-0 item 1 readability gate: PASS.* Anthocnet drop-cause
sumresidues — A-rwp +1.29, B-rwp +0.69, C-rwp +0.89, D-rwp +0.00; A-gm +2.22, B-gm +1.22, C-gm +1.60, D-gm +0.00. The capped arms are now tighter than the uncapped ones (cap=1 rwp: 108.49 pre-#407 → 101.57 post-#407 → 100.69 here at 20 seeds), so #402/#407 hold at 20 seeds and no arm's drop columns are withheld* — which is exactly what phase 0 item 1 was for. - Q2 — the publishable detector A/B: confirmed.* Paired, n=20: ΔPDR
+5.55 [+4.78, +6.32] pp (rwp, p=9.54e-05) and +6.54 [+5.79, +7.29] pp
(gaussmarkov, p=1.91e-06), 20/20 sign-consistent in both. The n=5 "+6.4 pp"
finding replicates and sits inside the gaussmarkov CI — the pre-registered
withdrawal condition did not fire. Its reorder corroboration survived its own
direct test: detector-off collapses the fading reorder ratio from
0.1745 / 0.2301 to 0.0010 / 0.0009* (175× / 256×), with the cap arms
interpolating monotonically, so duplicates — not multipath — drive
AntHocNet's fading reordering and the #399
metrics.mdgloss correction does not reopen. - Q3 — the ≥ 59 % claim: stands, and the bound's caveat is retired.* The
NotifyTxErrorcounter measures the overlap directly at 20 seeds: 0.6519 [0.6405, 0.6633] (A-rwp) and 0.6683 [0.6536, 0.6830]* (A-gm). Both CIs lie entirely above 0.59 — the direct measurement reproduces the inclusion–exclusion floor and clears it by ~6 pp, as a conservative floor should behave. - Q1 — the cap sweep: neither auto-criterion fires; the pre-registered
"maintainer decides" branch TRIGGERED.* Auto-adopt cap=1 requires dominance
and the ΔPDR (A−B) upper CI bound must be < 1 pp; it is +2.62 (rwp) / +1.80
(gm), so it does not fire (the other two conditions pass:
postTxcut 48.20 % / 49.61 % ≥ 30 %, delay99 and NRL both improve). Auto-reject-all also does not fire: worst PDR loss 1.89 pp (< 2 pp), C-rwp's CI includes zero (p=0.0668), and there are compensating d99/NRL/reorder gains. So the frontier table went to the maintainer as the deliverable, with the recommendation keepMaxReinjectPerPacket=0(unlimited) on the issue: the cap converts delivered-after-only re-injections into never (pktsNever2.66 % → 13.41 % at cap=1 on rwp), which is the whole PDR loss; the overhead saving does not land in NRL (−2.4 % / −4.8 % for a ~48–50 %postTxcut); and the capped arms' delay gains are survivorship-confounded. "No default chosen in advance" therefore resolved as no default change* — measured, not assumed.
Phase 3 — the oracle control¶
#296 item 1: a global-knowledge shortest-path arm replayed against the ground-truth topology — the upper bound both suites currently lack, and the largest unknown on the v1.5.0 path (it is harness work, not a sweep).
It runs after phases 1–2 rather than before, for one reason: an oracle arm is only meaningful against a fixed protocol configuration, and phases 1–2 are what fix it. It reuses the phase-1 grid cells, so its incremental cost is one arm per cell rather than a new campaign.
The arm is built and verified (2026-08-13); the campaign cells are what
remains. contrib/oracle is a self-contained ns-3 module wired into both
anthocnet-compare and isl-grid behind --protocols=…,oracle (off by
default, no existing arm or RNG stream touched). Design, the exactness
limitation per propagation model and the recompute cadence are in
ns3/oracle/README.md; the framing that says how
to read it sits with the other baselines in
methodology.md.
What phase 3 dispatches must know before it spends runtime:
- The oracle is exact on the satellite torus and on the
rangedisk, and approximate ontworay/nakagami— a fading channel has no crisp adjacency, so the control is held to the scenario's--rangeand every such row is flaggedapprox=1. The phase-1 grid is exactly the six{tworay, nakagami}cells, so every phase-3 grid cell is an approximate oracle: read it as a reference point, not as a proven upper bound, and expectscenario_check.pyto WARN on all six. The direction of the error is known (the 300 m radius is conservative under two-ray, so the control uses a subset of the links that actually work), which is why the WARN is the right level rather than a refusal. - The a-priori assertions are wired, not left to review: NRL exactly 0
(
NS_ABORTin-harness plus ascenario_checkrule), no protocol routing the same packet in fewer hops, and no arm beating full knowledge on a static lossless topology. The hop-count rule carries a survivorship guard —path_hops_meanis a mean over delivered packets, and a low-PDR arm skims the short flows (measured: olsr 1.90 hops at 75.4 % against the oracle's 2.09 at 95.9 %) — so it fires only when the other arm delivered at least as much. - Preflight first, as always:
scenario_check.py preflight --protocols=…,oracle --propagation=nakagamiFAILs the combination that would abort the run and WARNs that--range, inert for every other arm on that channel, is what pins the control's adjacency.
Smoke cells behind those statements (ns-3.36, paper base, single-seed unless
noted): range oracle 95.9 % PDR / 4.8 ms / NRL 0.00 against aodv 80.2 % /
45.4 ms (2 seeds, 300 s); tworay 100.0 % vs aodv 69.6 %; nakagami 99.3 % vs
aodv 55.2 %; 4×4 ISL torus 100.0 % PDR, mode=wired approx=0, one Dijkstra
solve for the whole run.
Phase 3 is where the v1.5.0 exit criteria actually begin. The AOMDV and GPSR spikes (#296 items 2–3) are independent of this campaign — they are feasibility work whose deliverable may well be a written infeasibility verdict, and they should proceed in parallel with phases 0–2 rather than queue behind them.
Done (2026-08-14) — six grid cells × 20 seeds on main @ 40b434d (the
#419 merge commit), image
ghcr.io/danieljoppi/ns3:3.42-opt, time=900, runs=20,
protocols=anthocnet,aodv,olsr,dsdv,oracle, gaussmarkov at pause=0; every
cell's identity read from its own ##CONFIG## rather than from dispatch order.
Published as the oracle sections of grid.md;
full readout on
#415 (comment),
which also closed #415.
| mobility | channel | run ID |
|---|---|---|
| rwp | tworay | 31807666381 |
| ssrwp | tworay | 31807668353 |
| gaussmarkov | tworay | 31807670848 |
| rwp | nakagami | 31807672820 |
| ssrwp | nakagami | 31807676290 |
| gaussmarkov | nakagami | 31807678924 |
- The satellite cell. The exact (
approx=0) half of phase 3 ran on the ISL suite as run 31830581632 and is reported on #216 (comment), published on satellite/isl-grid.md. It is the cell where the control's adjacency is the wiring rather than an assumed disk, so it — not the grid — is where an oracle number is a proven bound. Read it there; it is a different regime and its numbers are not comparable with the MANET grid's (network-regimes.md). - The three a-priori assertions, all PASS at 20 seeds. These were pre-registered above, before any cell was dispatched.
- NRL == 0.00 exactly* — not merely as a mean. All 120 per-seed oracle
rows* carry
nrlmin = max = 0.00 andnrl_bytesmax = 0.0000, corroborated independently by the harness's own# stddev oracle nrl=0.000and# drops oracle route=0.00lines. This is an asserted invariant (NS_ABORTin-harness plus ascenario_checkrule), so the correct reading is "the assertion held", not "overhead measured very low" — a distinction now recorded in metrics.md. - Oracle PDR ≥ every arm, in all six cells* — margins +7.05 to
+35.42 pp*. The sub-100 % oracle PDR under fading (99.41–99.55 %) is the
expected shape and the harness attributes it correctly:
# drops oracle route=0.00with the loss in the channel bucket. On two-ray the oracle is 100.00 % with every drop bucket at zero. - Hop rule — PASS, but vacuously, and that is the interesting result.*
No arm ties the oracle's PDR in any cell, so
#419's survivorship
guard suppresses the rule in all six. A naive version of the same rule
would have FAILED: on all three two-ray cells the oracle's
hopsMean(2.12 / 2.07 / 2.43) is the highest of the five arms, above anthocnet (1.97 / 1.91 / 2.26) and every baseline. The guard — added because a low-PDR arm skims short flows and scores a deceptively low mean — did real work here for a second, unanticipated reason: underapprox=1the control's 300 m disk misses real links, so it genuinely routes longer. Both readings are recorded because the guard would have been credited with catching only the first.* - Baseline byte-identity control — 480
##RUN##rows.* Every per-seed row of all four original arms (6 cells × 4 protocols × 20 seeds) is byte-identical to the phase-1 corpus, as are the##BENCH##,# stddev,# paths,# dropsand# energylines — anthocnet included*. Adding a fifth arm perturbed nothing, which is what licenses reading the phase-3 oracle column inside the phase-1 grid tables despite the commit gap (a1daa7a→40b434d). Without this control the composition would be a provenance-rule violation; with it, the comparison is attributable. Same control class as phase 1's 0/18 attribution check, applied to an added arm instead of a changed default. - Validation.*
scenario_check.py resultson all six cells: exit 0, zero FAILs*, 25 checks per cell — the #418 gate rewrite confirmed working on real campaign data. Column-mapping self-check OK; the oracle's positional##RUN##mapping was validated against the harness's own# stddev oracleline rather than assumed. - The #421
consequence: these cells ran pre-fix, so their
##BENCH##blocks have four rows, not five.*paper-benchmark.yml's compact-block step hardcoded the four original protocols in an awk alternation, so every arm added since v1.4.0 —oracle,gpsr,aomdv— was silently dropped from the##BENCH##re-emit, with no warning and no empty-table failure (the #28 gate fires only when zero rows match, and the four originals always match). All six phase-3 blocks confirm it:##BENCH##rows = 4,##BENCH## oracle= 0,##RUN## … oracle= 20. Consequence for these numbers: every oracle mean, CI and hop count published from this phase was computed from the per-seed##RUN##rows, not from the##BENCH##summary line* — which is possible only because the##RUN##rows are appended by an unfilteredgrep '^##RUN##'and therefore survived. Severity was bounded by luck, not by design: had the filter been applied one step later, phase 3 would have been a silent six-cell loss. #421 is fixed and closed; the phase-3 blocks themselves are not re-emitted, so anyone re-reading them must take the oracle from##RUN##. A related omission —##ORACLE##not being re-emitted in the compact block, which costs a ~2400-line tail instead of ~900 to read — is filed separately as the same class. - Anomalies carried forward, not silently dropped. (i) The oracle's
delay99is pinned at ~2.01 s on all three Nakagami cells with a per-seed sd of 2.3–3.4 ms — a cliff, not a distribution — which generalised into a metric rule rather than a footnote:delay99is not comparable across arms with materially different PDR. (ii) OraclenoRoutetotals over 20 seeds reach 225 on gaussmarkov-nakagami (seed 9 alone contributes 181), against 44/25/12/8/2 elsewhere — yet# drops oracle route=0.00in every cell; whether retries mask genuine partition time is an open follow-up. (iii) One##ORACLE##line was corrupted by an interleaved##RSS##write from the CI memory monitor — a log-capture artefact; that seed's data row is intact.
Affordability¶
#121 governs. Rough shape, at the measured ~2–2.5 h per 5-seed 900 s fading cell:
| phase | cells | 20-seed cells (4 × 5-seed jobs each) |
|---|---|---|
| 1 (grid re-baseline) | 6 | 24 jobs |
| 2 (cap × detector) | 8 (4 arms × 2 mobility) | 32 jobs |
| 3 (oracle) | 6 (one arm added per grid cell) | 24 jobs |
Concurrency is the lever that makes this tractable — the TCP arm's four 5-seed cells finished a 20-seed campaign in 2 h 34 m of wall clock. The number worth watching is not total CPU-hours but how many campaigns this replaces: three, if run in this order; zero, if run out of it.
What this plan deliberately does not do¶
- No sweep re-measure. #365's disposition stands.
- No
##DROPID##un-suppression under TCP. That is #389's follow-up and needs its own identity validation; it is not a campaign question. - No RL baseline. #296 item 5 is explicitly out of scope until this campaign and the metrics epic land — a leaky comparison is worse than none.
- No default chosen in advance. The cap default is whatever the phase-2
frontier says, including "unlimited" if the +6.4 pp PDR does not survive
capping. Phase 2 is a measurement, not a confirmation.
Resolved (2026-08-13): the +6.4 pp survived at 20 seeds (+5.55 / +6.54 pp),
every cap arm cost PDR, and neither auto-criterion fired — so the frontier
went to the maintainer and the default stayed
MaxReinjectPerPacket=0. A measured no-change, not a deferral.