PPI EngineDOT prototype

Methodology

Coverage gate

Every interval this engine produces is simulation-tested at known ground truth: nominal 95% intervals must cover at least 93% empirically or the build fails. Generated by ppi_core.simulate, seed 20260816, byte-identical on rerun. Status: ✓ PASS (zero failures)

Scenario · methodCoverageMean widthRepsStatus
active_ipw_mean.classical_ipw0.95900.87131000✓ pass
active_ipw_mean.ppi_ipw0.95600.54391000✓ pass
active_ipw_mean.ppi_naive0.00000.37131000ungated
active_loop_asym.classical0.95250.4228400✓ pass
active_loop_asym.ppi0.96000.2405400✓ pass
active_loop_asym.ppi_power_tuned0.96000.2405400✓ pass
active_loop_mean.classical0.95000.4687400✓ pass
active_loop_mean.ppi0.94500.2961400✓ pass
active_loop_mean.ppi_power_tuned0.94500.2961400✓ pass
active_loop_mean_div.classical0.97250.4799400✓ pass
active_loop_mean_div.ppi0.96250.3032400✓ pass
active_loop_mean_div.ppi_power_tuned0.96250.3032400✓ pass
active_loop_mean_vr.classical0.96000.4672400✓ pass
active_loop_mean_vr.ppi0.95500.2952400✓ pass
active_loop_mean_vr.ppi_power_tuned0.95500.2952400✓ pass
bootstrap_mean.analytic_ppi0.93330.1991300✓ pass
bootstrap_mean.bootstrap_ppi0.93000.1958300✓ pass
logistic.classical0.9500 (per-coord min of [0.95, 0.95])0.6390300✓ pass
logistic.ppi0.9500 (per-coord min of [0.95, 0.95])0.6476300✓ pass
mean.classical0.94200.48882000✓ pass
mean.ppi0.94550.23912000✓ pass
mean.ppi_power_tuned0.94050.23712000✓ pass
ols.classical0.9300 (per-coord min of [0.952, 0.938, 0.93])0.2481500✓ pass
ols.ppi0.9420 (per-coord min of [0.958, 0.958, 0.942])0.2728500✓ pass
ols.ppi_power_tuned0.9300 (per-coord min of [0.956, 0.94, 0.93])0.2439500✓ pass
policy_gain.diversity0.93330.251760ungated
policy_gain.random0.93330.258160ungated
policy_gain.uncertainty0.93330.245160ungated
policy_gain.variance_reduction0.93330.245160ungated
quantile.classical_p0.50.96400.8513500✓ pass
quantile.classical_p0.90.94801.2752500✓ pass
quantile.ppi_p0.50.98000.4463500✓ pass
quantile.ppi_p0.90.96800.6691500✓ pass
runner_bootstrap.bootstrap_pool_mean0.95000.2651300✓ pass
runner_logistic.classical0.9450 (per-coord min of [0.98, 0.945])0.6559200✓ pass
runner_logistic.ppi0.9450 (per-coord min of [0.985, 0.945])0.6534200✓ pass
runner_ols.classical0.9700 (per-coord min of [0.98, 0.97])0.2510200✓ pass
runner_ols.ppi0.9800 (per-coord min of [0.985, 0.98])0.2702200✓ pass
runner_ols.ppi_power_tuned0.9700 (per-coord min of [0.98, 0.97])0.2471200✓ pass
runner_quantile.classical1.00002.0353200✓ pass
runner_quantile.ppi0.99000.6161200✓ pass
weight_skew_mean.classical_ipw0.93751.4274800✓ pass
weight_skew_mean.ppi_ipw0.96250.8118800✓ pass

active_ipw_mean.ppi_naive is a deliberate demonstration of the selection-bias trap (0% coverage without IPW correction); policy_gain rows are a width-efficiency experiment whose coverage is gated separately at higher replication counts.

Agency research reports

The research agent scores each agency's published feed against the versioned rubric agency-data-quality v1; the verification agent independently recomputes every finding and refuses anything it cannot ground. Fixture snapshots of 2026-08-16.

Mississippi DOT

✓ verified
79.7 / 100 · 145 records
date-verification145/145 = 100.0%
impact-completeness27/145 = 18.6%
geometry-completeness145/145 = 100.0%
description-quality145/145 = 100.0%
duration-derivability145/145 = 100.0%

Utah DOT

✓ verified
60.9 / 100 · 744 records
date-verification744/744 = 100.0%
impact-completeness0/744 = 0.0%
geometry-completeness744/744 = 100.0%
description-quality46/744 = 6.2%
duration-derivability744/744 = 100.0%

Missouri DOT

✓ verified
79.9 / 100 · 615 records
date-verification615/615 = 100.0%
impact-completeness489/615 = 79.5%
geometry-completeness615/615 = 100.0%
description-quality0/615 = 0.0%
duration-derivability615/615 = 100.0%

Kentucky Transportation Cabinet

✓ verified
71.4 / 100 · 298 records
date-verification298/298 = 100.0%
impact-completeness0/298 = 0.0%
geometry-completeness298/298 = 100.0%
description-quality228/298 = 76.5%
duration-derivability296/298 = 99.3%

Known limitations (stated, not hidden)

  • Plug-in selection weights. Multi-round adaptive selection uses realized cumulative inclusion probabilities; exact Horvitz–Thompson unbiasedness does not hold under adaptive designs. The scheme is certified empirically by gated end-to-end simulations, including asymmetric-uncertainty designs built to stress it (docs/gauntlet/statistical-core.md).
  • Two interval targets. Analytic CIs target the superpopulation; the runner's bootstrap targets the mean of the ingested pool (selection uncertainty only). They are labeled distinctly everywhere and answer different questions.
  • Policy gains depend on the oracle. Active policies beat random by ~6% width in the gated experiment because oracle uncertainty is informative there; with an uninformative oracle they degrade to random. The random baseline is always available for comparison.
  • Label budget semantics. Ground-truth labels derive from authoritative feed fields via the verification agent; the budget simulates real-world acquisition cost rather than paying it (docs/02-feeds.md).
  • The heuristic oracle is weak by design and labeled as such (heuristic:v1); live Anthropic labeling requires an API key and is tagged with its model identity.
  • Determinism is certified single-platform (byte-identical reruns on the build machine and CI image; cross-platform BLAS bit-identity is not claimed).
  • lane_restricted covers MS + MO only — Utah and Kentucky publish vehicle_impact as "unknown" on every record in our snapshots, so the verifier refuses them for that estimand (516 of 1,802 records eligible).