SummerProject2026/docs/project_options.md

169 lines
8.2 KiB
Markdown

# Project options based on the countydata inventory
Inventory date: 2026-07-15
> **Archived planning context:** This shortlist records ideas considered during
> discovery; it is not the scope of the finished prototype. DMV enrichment,
> multiclass outcomes, station analysis, forecasting, and population/county
> rankings below are not current deliverables. The authoritative scope is the
> [project charter](project_charter.md), and the completed development story is
> in the [final report](final_report.md).
> **Selected:** Option 1, Utah Vehicle Health. The implementation contract is in
> [project_charter.md](project_charter.md).
## What is actually available
The useful analytical source is the `countydata` PostgreSQL database, which is
about 93.4 GB. Its normalized tables are projections of two large logical
datasets, so their row counts should not be added together:
- 18,009,278 DMV registration/tax records with county, registration date,
make, model, model year, fuel type, registration type/place, expiration and
emissions dates, ZIP codes, and temporary-registration status.
- 19,357,287 inspection records with county/source, timestamp, make, model,
model year, station, result, OBD result, test/program type, and OBD summary.
- Raw inspection JSON adds odometer, fuel, engine, transmission, vehicle type,
readiness monitors, visual checks, DTCs, PIDs, and communication protocol for
newer county/source feeds.
The vehicle histories link well without exposing identifiers: 88.2% of a
10,000-VIN inspection sample had DMV history, and the sampled median was eight
inspection visits per vehicle. In the other direction, 74.2% of sampled DMV
vehicles had an inspection match.
Important limitations:
- DMV coverage is strong in 2016-2020 and 2022 through March 2024, but only 117
rows are dated 2021. Treat 2024 as partial.
- Inspection coverage is strong from 2010 through June 2026, but the four rows
dated 1990 are clear date outliers.
- Raw inspection `overall_result` labels are imbalanced: 71.06% pass, 3.61%
fail, 3.56% reject, 2.31% abort, and 19.45% blank/null. Most missing values
come from the older Utah County feed; its separately encoded OBD result can
support a provenance-tagged binary target, but not the four-class analysis.
- Rich OBD/odometer fields are source-dependent. They are essentially absent
from the older `slc`, `utah`, and `weber` feeds and concentrated in `slco`,
`davis`, and `cache` records from late 2024 onward.
- Categories require cleaning: fuel values differ by case and punctuation, and
inspection make values include aliases such as `TOYOTA`/`TOYOT` and
`CHEVROLET`/`CHEVR`.
## Ranked ideas
### 1. Utah Vehicle Health and Reliability Observatory
Build a model that estimates the chance a vehicle will fail, reject, or abort
its next emissions inspection using only information known before that test:
vehicle age, canonical make/model, county, season, fuel, prior test outcomes,
time since the previous test, and longitudinal DMV history.
The dashboard could include:
- make/model/year reliability scorecards with uncertainty intervals;
- risk-versus-age curves and county comparisons;
- prior-failure and repeat-test patterns;
- calibrated individual what-if estimates without accepting or displaying a
VIN; and
- a methodology/data-quality page showing drift and missingness.
This is the strongest overall option because it combines SQL/data engineering,
entity resolution, longitudinal feature engineering, classification or
survival analysis, model explainability, and a compelling public dashboard.
Guardrail: do not use same-test OBD result, DTC count, or overall result as
features when claiming to predict an outcome before the test. That would be
target leakage. Use a time-based holdout and report PR-AUC, calibration, and
Brier score rather than accuracy alone.
### 2. Failure-to-Pass Journey Analyzer
Follow vehicles forward after a fail or reject and model how many attempts and
how much time it takes to achieve a pass. Compare journeys by vehicle age,
canonical make/model, program, county/source, and failure history using
time-to-event or competing-risk methods.
The dashboard could use journey funnels, transition diagrams, survival curves,
and a cohort comparison tool. This is an unusually good fit for the repeated
histories and can become a major page within option 1 instead of a separate app.
### 3. Utah EV and Hybrid Transition Atlas
Deduplicate repeated registrations into vehicle-by-period snapshots, estimate
electric and hybrid share by county and ZIP, measure transitions in the vehicle
fleet, and forecast adoption scenarios.
The dashboard could use a choropleth, adoption curves, county rankings, and a
make/model explorer. This is visually strong and easy to explain, but the DMV
2021 gap and March 2024 endpoint make honest uncertainty and partial-period
handling essential.
### 4. OBD Early-Warning Lab
Scope the analysis to the richer `slco`, `davis`, and `cache` feeds. Use
odometer, vehicle age, fuel, engine, readiness monitors, communication protocol,
MIL state, and prior history to distinguish likely pass, fail, reject, and
abort outcomes and discover common failure signatures.
The dashboard could show diagnostic pathways, readiness-monitor patterns, and
failure signatures by vehicle cohort. It is technically novel, but it is not a
statewide study because rich-feature coverage begins mainly in 2024-2025.
### 5. Risk-Adjusted Inspection Station Consistency
Fit a hierarchical model of station outcomes after controlling for vehicle age,
make/model, program, county, and time. Use empirical-Bayes shrinkage or control
charts to flag unusual reject, abort, or failure rates and detect process drift.
This is excellent applied statistics and anomaly detection. Public results
should anonymize station identifiers, suppress small groups, show uncertainty,
and describe anomalies as review signals rather than evidence of misconduct.
### 6. AutoClean: Government Vehicle Data Entity Resolution
Use the high VIN match rate as weak supervision to learn canonical make/model
mappings between messy inspection strings and cleaner DMV values. Compare
rules, fuzzy matching, and a supervised ranking model, then quantify how much
normalization improves downstream analytics.
The dashboard could show before/after category fragmentation, match confidence,
and correction examples using non-identifying values. This is a particularly
strong data-engineering project and could also serve as the first pipeline stage
for option 1.
### 7. Inspection Demand Forecast and Operations Dashboard
Forecast daily or weekly test volume by county and station using timestamps,
seasonality, holidays, long-term trend, and recent history. Add change-point and
volume-anomaly detection.
The dashboard could show workload forecasts, day/hour heatmaps, forecast
intervals, and historical disruptions. It is feasible and useful, although it
has less machine-learning depth than the reliability project unless forecasting
and anomaly evaluation are developed carefully.
### 8. Vehicle Longevity and Survival Index
Treat the repeated DMV and inspection histories as censored longitudinal data.
Estimate how long makes, models, fuel types, and model-year cohorts remain
active, using Kaplan-Meier curves and a Cox or gradient-boosted survival model.
The dashboard could answer "Which vehicles stay on Utah roads the longest?"
The main methodological challenge is that disappearance from the data can mean
sale, relocation, incomplete coverage, or retirement, so the result must be
described as observed-system retention rather than mechanical lifespan.
## Recommended project shape
Use option 1 as the main story and option 6 as its data-engineering foundation:
1. Build an incremental, read-only SQL extraction and canonical vehicle model.
2. Create one row per eligible upcoming inspection using only prior information.
3. Compare a transparent logistic baseline with a tree model.
4. Validate on a later time period and separately by county/source.
5. Export only aggregate scorecards, curves, and de-identified model outputs.
6. Build the static dashboard over those safe outputs.
This gives the project a coherent end-to-end narrative across data engineering,
analytics, machine learning, responsible validation, and product design.