SummerProject2026/docs/modeling_protocol.md

5.9 KiB

Leakage-safe modeling protocol

Reporting boundary

This protocol governs the private, page-sampled 10,000-vehicle development cohort. All resulting counts, rates, comparisons, and performance metrics are sample results, not population estimates. County/source reporting is for coverage and drift context only; it must not be turned into population county rankings.

Prediction unit

Estimate the probability that the first attempt of a returning vehicle's next inspection episode is a non-pass, immediately before the episode begins. Consecutive attempts for the same private vehicle token form one episode when they are no more than 30 days apart. Rapid retests are not separate targets.

This is the sole research target for the completed prototype. Cold-start, four-class, fail-only, failure-to-pass, and individualized prediction tasks are out of scope.

Eligibility

  • Use 2010-2015 events only as historical context; supervised targets begin in 2016.
  • Exclude the four known 1990 date outliers.
  • Require a recognized first-attempt binary outcome and at least one prior completed episode.
  • Keep first-observed vehicles outside the modeled returning-vehicle cohort.
  • Normalize slc and slco to Salt Lake County for geography while preserving source era for diagnostics.
  • Deduplicate exact uploads and quarantine shared, test, or placeholder identifiers under preregistered conflict/activity rules.

Every exclusion appears in the private cohort-flow audit.

Binary label contract

The target is 0 for a recognized pass and 1 for recognized fail, reject, or abort. Blank, null, and unrecognized values remain unlabeled rather than being treated as passes.

The normalized target prefers a recognized overall_result. The older Utah County feed is the sole exception: when its overall result is blank, obd_result may supply the binary label only when source is utah, program is obd, test type is OBD, and the controlled value is pass, fail, reject, or abort. B, blank, TSI, other/C, and unknown values remain unlabeled. Private label provenance is retained for audit and is never a predictor or public field.

This exception is a binary pass/non-pass proxy. It does not justify interpreting fail, reject, and abort as interchangeable mechanical conditions.

Point-in-time features

Every feature window ends strictly before the target episode. The final inspection-history feature set includes:

  • vehicle age;
  • prior episode and attempt counts;
  • days since the prior episode and prior adverse outcome;
  • prior non-pass rate;
  • prior episode first and final outcomes;
  • previously observed make and model;
  • public county context; and
  • target season.

The final prototype does not use DMV enrichment or rich current-test OBD data.

Leakage and privacy exclusions

The model and public product exclude:

  • the target attempt's result, reason, DTC count, MIL/readiness state, PIDs, visual checks, measurements, certificate, or calibration fields;
  • later attempts or eventual outcome from the target episode;
  • station, technician, or station-level outcome information;
  • raw VIN, private vehicle token, plate, ZIP, exact address, or another direct or pseudonymous identifier as a feature;
  • future records, full-history aggregates, or target-row information that was not available before the episode;
  • preprocessing, category mappings, imputation, or calibration learned from a later partition; and
  • raw JSON, credentials, operational records, and row-level predictions in any public artifact.

Timestamp ties are resolved before lag/window calculations. Cumulative windows end at the preceding event or episode.

Frozen chronology

Partition Target dates Role in the private 10,000-vehicle development sample
Historical context 2010-2015 Lag features only
Train 2016-2022 Fit preprocessing and model parameters
Tune 2023 Select logistic regularization; compare fixed candidates
Calibrate 2024 Fit Platt scaling and inspect calibration behavior
One-time holdout 2025 Development comparison after specifications were frozen
Shadow 2026-01-01 to 2026-06-22 Partial-period monitoring context only

The 2025 gate was opened once during live verification. It is no longer a pristine unseen test, no 2025-informed tuning is permitted, and every reported 2025 value must say “one-time development-sample holdout.” A later complete-data 2025 analysis would be confirmatory, not a new final test.

Repeated vehicles may cross chronological partitions because returning-vehicle prediction is the intended scenario. Separately held-out private vehicle buckets form a never-fit stress test. Random row splitting is prohibited.

Models and decision rule

  1. Training prevalence: probability baseline.
  2. Previous episode's first outcome: literal history baseline.
  3. Regularized logistic regression: selected using the 2023 sample partition, then Platt-calibrated on the 2024 sample partition; the final model.
  4. Histogram gradient boosting with Platt calibration: fixed nonlinear benchmark only.

The benchmark does not power a separate dashboard experience. Added complexity would require a clear, predeclared improvement to displace the logistic model; the one-time 2025 sample comparison did not show that improvement.

Metrics and interpretation

Headline metrics are non-pass PR-AUC and Brier score. Log loss, ROC-AUC, and precision at the highest-scored 10% are supporting diagnostics. Calibration values computed on the 2024 partition used to fit Platt scaling are calibration diagnostics, not independent final performance.

All tables and charts must state that they use the private 10,000-vehicle development sample and are not population estimates. Report sample support with every subgroup result. Do not describe observed differences as causal, do not claim generalization to all Utah vehicles, and do not publish county rankings.