Methodology
Coverage gate
Every interval this engine produces is simulation-tested at known ground truth: nominal 95% intervals must cover at least 93% empirically or the build fails. Generated by ppi_core.simulate, seed 20260816, byte-identical on rerun. Status: ✓ PASS (zero failures)
| Scenario · method | Coverage | Mean width | Reps | Status |
|---|---|---|---|---|
| active_ipw_mean.classical_ipw | 0.9590 | 0.8713 | 1000 | ✓ pass |
| active_ipw_mean.ppi_ipw | 0.9560 | 0.5439 | 1000 | ✓ pass |
| active_ipw_mean.ppi_naive | 0.0000 | 0.3713 | 1000 | ungated |
| active_loop_asym.classical | 0.9525 | 0.4228 | 400 | ✓ pass |
| active_loop_asym.ppi | 0.9600 | 0.2405 | 400 | ✓ pass |
| active_loop_asym.ppi_power_tuned | 0.9600 | 0.2405 | 400 | ✓ pass |
| active_loop_mean.classical | 0.9500 | 0.4687 | 400 | ✓ pass |
| active_loop_mean.ppi | 0.9450 | 0.2961 | 400 | ✓ pass |
| active_loop_mean.ppi_power_tuned | 0.9450 | 0.2961 | 400 | ✓ pass |
| active_loop_mean_div.classical | 0.9725 | 0.4799 | 400 | ✓ pass |
| active_loop_mean_div.ppi | 0.9625 | 0.3032 | 400 | ✓ pass |
| active_loop_mean_div.ppi_power_tuned | 0.9625 | 0.3032 | 400 | ✓ pass |
| active_loop_mean_vr.classical | 0.9600 | 0.4672 | 400 | ✓ pass |
| active_loop_mean_vr.ppi | 0.9550 | 0.2952 | 400 | ✓ pass |
| active_loop_mean_vr.ppi_power_tuned | 0.9550 | 0.2952 | 400 | ✓ pass |
| bootstrap_mean.analytic_ppi | 0.9333 | 0.1991 | 300 | ✓ pass |
| bootstrap_mean.bootstrap_ppi | 0.9300 | 0.1958 | 300 | ✓ pass |
| logistic.classical | 0.9500 (per-coord min of [0.95, 0.95]) | 0.6390 | 300 | ✓ pass |
| logistic.ppi | 0.9500 (per-coord min of [0.95, 0.95]) | 0.6476 | 300 | ✓ pass |
| mean.classical | 0.9420 | 0.4888 | 2000 | ✓ pass |
| mean.ppi | 0.9455 | 0.2391 | 2000 | ✓ pass |
| mean.ppi_power_tuned | 0.9405 | 0.2371 | 2000 | ✓ pass |
| ols.classical | 0.9300 (per-coord min of [0.952, 0.938, 0.93]) | 0.2481 | 500 | ✓ pass |
| ols.ppi | 0.9420 (per-coord min of [0.958, 0.958, 0.942]) | 0.2728 | 500 | ✓ pass |
| ols.ppi_power_tuned | 0.9300 (per-coord min of [0.956, 0.94, 0.93]) | 0.2439 | 500 | ✓ pass |
| policy_gain.diversity | 0.9333 | 0.2517 | 60 | ungated |
| policy_gain.random | 0.9333 | 0.2581 | 60 | ungated |
| policy_gain.uncertainty | 0.9333 | 0.2451 | 60 | ungated |
| policy_gain.variance_reduction | 0.9333 | 0.2451 | 60 | ungated |
| quantile.classical_p0.5 | 0.9640 | 0.8513 | 500 | ✓ pass |
| quantile.classical_p0.9 | 0.9480 | 1.2752 | 500 | ✓ pass |
| quantile.ppi_p0.5 | 0.9800 | 0.4463 | 500 | ✓ pass |
| quantile.ppi_p0.9 | 0.9680 | 0.6691 | 500 | ✓ pass |
| runner_bootstrap.bootstrap_pool_mean | 0.9500 | 0.2651 | 300 | ✓ pass |
| runner_logistic.classical | 0.9450 (per-coord min of [0.98, 0.945]) | 0.6559 | 200 | ✓ pass |
| runner_logistic.ppi | 0.9450 (per-coord min of [0.985, 0.945]) | 0.6534 | 200 | ✓ pass |
| runner_ols.classical | 0.9700 (per-coord min of [0.98, 0.97]) | 0.2510 | 200 | ✓ pass |
| runner_ols.ppi | 0.9800 (per-coord min of [0.985, 0.98]) | 0.2702 | 200 | ✓ pass |
| runner_ols.ppi_power_tuned | 0.9700 (per-coord min of [0.98, 0.97]) | 0.2471 | 200 | ✓ pass |
| runner_quantile.classical | 1.0000 | 2.0353 | 200 | ✓ pass |
| runner_quantile.ppi | 0.9900 | 0.6161 | 200 | ✓ pass |
| weight_skew_mean.classical_ipw | 0.9375 | 1.4274 | 800 | ✓ pass |
| weight_skew_mean.ppi_ipw | 0.9625 | 0.8118 | 800 | ✓ pass |
active_ipw_mean.ppi_naive is a deliberate demonstration of the selection-bias trap (0% coverage without IPW correction); policy_gain rows are a width-efficiency experiment whose coverage is gated separately at higher replication counts.
Agency research reports
The research agent scores each agency's published feed against the versioned rubric agency-data-quality v1; the verification agent independently recomputes every finding and refuses anything it cannot ground. Fixture snapshots of 2026-08-16.
Mississippi DOT
✓ verified79.7 / 100 · 145 records
| date-verification | 145/145 = 100.0% |
| impact-completeness | 27/145 = 18.6% |
| geometry-completeness | 145/145 = 100.0% |
| description-quality | 145/145 = 100.0% |
| duration-derivability | 145/145 = 100.0% |
Utah DOT
✓ verified60.9 / 100 · 744 records
| date-verification | 744/744 = 100.0% |
| impact-completeness | 0/744 = 0.0% |
| geometry-completeness | 744/744 = 100.0% |
| description-quality | 46/744 = 6.2% |
| duration-derivability | 744/744 = 100.0% |
Missouri DOT
✓ verified79.9 / 100 · 615 records
| date-verification | 615/615 = 100.0% |
| impact-completeness | 489/615 = 79.5% |
| geometry-completeness | 615/615 = 100.0% |
| description-quality | 0/615 = 0.0% |
| duration-derivability | 615/615 = 100.0% |
Kentucky Transportation Cabinet
✓ verified71.4 / 100 · 298 records
| date-verification | 298/298 = 100.0% |
| impact-completeness | 0/298 = 0.0% |
| geometry-completeness | 298/298 = 100.0% |
| description-quality | 228/298 = 76.5% |
| duration-derivability | 296/298 = 99.3% |
Known limitations (stated, not hidden)
- Plug-in selection weights. Multi-round adaptive selection uses realized cumulative inclusion probabilities; exact Horvitz–Thompson unbiasedness does not hold under adaptive designs. The scheme is certified empirically by gated end-to-end simulations, including asymmetric-uncertainty designs built to stress it (docs/gauntlet/statistical-core.md).
- Two interval targets. Analytic CIs target the superpopulation; the runner's bootstrap targets the mean of the ingested pool (selection uncertainty only). They are labeled distinctly everywhere and answer different questions.
- Policy gains depend on the oracle. Active policies beat random by ~6% width in the gated experiment because oracle uncertainty is informative there; with an uninformative oracle they degrade to random. The random baseline is always available for comparison.
- Label budget semantics. Ground-truth labels derive from authoritative feed fields via the verification agent; the budget simulates real-world acquisition cost rather than paying it (docs/02-feeds.md).
- The heuristic oracle is weak by design and labeled as such (
heuristic:v1); live Anthropic labeling requires an API key and is tagged with its model identity. - Determinism is certified single-platform (byte-identical reruns on the build machine and CI image; cross-platform BLAS bit-identity is not claimed).
- lane_restricted covers MS + MO only — Utah and Kentucky publish vehicle_impact as "unknown" on every record in our snapshots, so the verifier refuses them for that estimand (516 of 1,802 records eligible).