174 lines
7.1 KiB
Markdown
174 lines
7.1 KiB
Markdown
# Utah Vehicle Health: 10-minute presentation outline
|
|
|
|
## Presentation rule
|
|
|
|
Keep this sentence visible on every results or dashboard slide:
|
|
|
|
> Private 10,000-vehicle development sample; sample results, not population
|
|
> estimates.
|
|
|
|
Do not show county rankings, imply statewide coverage, or describe non-pass as
|
|
mechanical failure. Refer to calibrated logistic regression as the **final
|
|
model** and histogram gradient boosting as the **benchmark**.
|
|
|
|
## Timed outline
|
|
|
|
| Time | Slide / action | Core message |
|
|
| --- | --- | --- |
|
|
| 0:00-0:40 | 1. Title and boundary | One focused prediction question; development prototype, not a population study or decision tool |
|
|
| 0:40-1:25 | 2. Research question | Returning vehicle, next episode, first attempt, binary non-pass |
|
|
| 1:25-2:15 | 3. Development data | Private page-sampled 10,000-vehicle histories; complete sampled histories but nonrepresentative selection |
|
|
| 2:15-3:10 | 4. Leakage-safe design | 30-day episodes, only pre-episode history, chronological split, never random rows |
|
|
| 3:10-4:20 | 5. Model comparison | Logistic final versus prevalence/previous-outcome baselines and boosted-tree benchmark |
|
|
| 4:20-4:50 | 6. Model decision | Calibrated logistic wins the declared sample comparison and is easier to explain |
|
|
| 4:50-7:35 | Live dashboard demo | Overview, Sample cohorts, Model & benchmark, Data & methods |
|
|
| 7:35-8:30 | 7. Privacy architecture | Private pipeline to suppressed aggregate JSON to an allowlisted same-network demo server |
|
|
| 8:30-9:20 | 8. Limitations | Sampling, feed coverage, source/time confounding, heterogeneous non-pass, one-time holdout |
|
|
| 9:20-10:00 | 9. Close | Prototype is complete and honest about what it can—and cannot—claim |
|
|
|
|
Total: **10:00**
|
|
|
|
## Slide notes
|
|
|
|
### 1. Title and boundary — 40 seconds
|
|
|
|
**Title:** Utah Vehicle Health
|
|
|
|
**Subtitle:** Predicting a returning vehicle's next-episode first-attempt
|
|
non-pass outcome
|
|
|
|
Say:
|
|
|
|
> This is a presentation-ready development prototype built from a private,
|
|
> page-sampled 10,000-vehicle cohort. Every result I show is sample-based, not a
|
|
> population estimate. The project does not rank counties or diagnose vehicles.
|
|
|
|
### 2. Research question — 45 seconds
|
|
|
|
Show the question:
|
|
|
|
> Using only information available before an inspection episode begins, how
|
|
> well can a calibrated logistic regression estimate whether a returning
|
|
> vehicle's next episode will have a first-attempt non-pass outcome?
|
|
|
|
Explain that attempts separated by no more than 30 days form one episode. This
|
|
keeps rapid fail/retest sequences from becoming repeated target rows. Non-pass
|
|
combines fail, reject, and abort; blanks remain unlabeled.
|
|
|
|
### 3. Development data — 50 seconds
|
|
|
|
Show a compact flow:
|
|
|
|
```text
|
|
sample up to 10,000 vehicles -> retrieve their histories -> build episodes
|
|
-> create 44,659 eligible returning targets
|
|
```
|
|
|
|
Say that 9,996 private tokens remained after validation, producing 83,552 source
|
|
events and 69,588 episodes in the private sample. Immediately repeat that these
|
|
are reconciliation counts, not Utah totals. Page sampling over-represents
|
|
vehicles with more inspection records, and participating feeds do not cover all
|
|
29 counties.
|
|
|
|
### 4. Leakage-safe design — 55 seconds
|
|
|
|
Show the chronology:
|
|
|
|
```text
|
|
2010-15 context | 2016-22 train | 2023 tune | 2024 calibrate
|
|
| 2025 one-time holdout | 2026 partial shadow
|
|
```
|
|
|
|
Explain that every feature window ends before the target episode. The model uses
|
|
prior outcomes, history depth, timing, age, previously observed make/model,
|
|
public county context, and season. It excludes current-test diagnostics, future
|
|
records, identifiers, stations, technicians, and later attempts.
|
|
|
|
Note that the 2025 sample gate was opened once after specifications were frozen;
|
|
it is a development holdout, not a pristine future test.
|
|
|
|
### 5. Model comparison — 70 seconds
|
|
|
|
Use a two-metric table. Label it “Private 10,000-vehicle development sample; not
|
|
population performance.”
|
|
|
|
| 2025 one-time sample holdout | Role | PR-AUC | Brier |
|
|
| --- | --- | ---: | ---: |
|
|
| Training prevalence | Baseline | 0.123 | 0.1076 |
|
|
| Previous episode | Baseline | 0.158 | 0.1731 |
|
|
| Logistic + Platt | **Final** | **0.261** | **0.1011** |
|
|
| Boosted tree + Platt | Benchmark | 0.238 | 0.1024 |
|
|
|
|
Say that higher PR-AUC and lower Brier are better. Logistic beat both baselines
|
|
and the boosted-tree benchmark in this one-time sample comparison. Avoid the
|
|
word “accuracy,” which obscures class imbalance and calibration.
|
|
|
|
### 6. Model decision — 30 seconds
|
|
|
|
Say:
|
|
|
|
> The final model is calibrated logistic regression. The boosted tree is a
|
|
> benchmark only. It did not improve the declared sample metrics enough to
|
|
> justify greater complexity, while logistic regression is easier to audit and
|
|
> explain.
|
|
|
|
### Live dashboard demo — 2 minutes 45 seconds
|
|
|
|
Follow [demo_script.md](demo_script.md). Keep the development-preview banner in
|
|
view at the start. Explicitly show calibrated logistic as final, the tree as
|
|
benchmark-only, and the absence of a vehicle-level prediction service.
|
|
|
|
### 7. Privacy architecture — 55 seconds
|
|
|
|
Show:
|
|
|
|
```text
|
|
private read-only data -> local modeling -> suppressed aggregate JSON
|
|
-> allowlisted local static server -> same-network demo browser
|
|
```
|
|
|
|
State that the public bundle has no VINs, plates, ZIPs, stations, technician
|
|
identifiers, private tokens, raw JSON, credentials, operational records, or
|
|
row-level predictions. The browser has no database connection and fails closed
|
|
if its aggregate contract does not validate.
|
|
|
|
### 8. Limitations — 50 seconds
|
|
|
|
Name the limitations directly:
|
|
|
|
- page-sampled cohort is not population-representative;
|
|
- feeds are partial and change over time;
|
|
- geography and source era are confounded;
|
|
- reject/abort do not necessarily mean mechanical failure;
|
|
- 2025 is a one-time development holdout, not external validation; and
|
|
- sample associations do not support county rankings or causal conclusions.
|
|
|
|
### 9. Close — 40 seconds
|
|
|
|
End with:
|
|
|
|
> The prototype is complete at the development stage: one precise question, a
|
|
> leakage-safe timeline, calibrated logistic regression as the final model, a
|
|
> boosted-tree benchmark, and a privacy-reviewed static dashboard. Its value is
|
|
> not just the sample performance; it is the discipline of making only the
|
|
> claims the data and release boundary can support.
|
|
|
|
Pause, then invite questions about episode construction, calibration, privacy,
|
|
or the dashboard—not about county “winners” and “losers,” which this sample
|
|
cannot establish.
|
|
|
|
## Presentation checklist
|
|
|
|
- Put both computers on the same trusted network, stop any older dashboard
|
|
process, and run `node dashboard/server.mjs --lan`.
|
|
- On the demo computer, open the printed **Other devices** URL for the shared
|
|
Wi-Fi or Ethernet interface with `/#overview` appended. Close unrelated
|
|
browser tabs.
|
|
- Use a fresh page load to confirm aggregate validation succeeds.
|
|
- Keep a screenshot or short recording as a backup, with the sample warning
|
|
visible.
|
|
- Never open private data, artifacts, terminal environment variables, database
|
|
tools, or browser developer tools during the talk.
|
|
- Do not type or display a VIN, plate, ZIP, station, technician identifier, raw
|
|
record, or row-level prediction.
|