UC Berkeley · ZoneZero

Pre-Eaton Fire Zone Zero Risks

Model evaluation

Fixed validation results

Final fine-tuned model performance on the fixed seed-42 validation split: 400 buildings for Stage 1 and 298 human-QC-pass buildings for Stage 2. The six hierarchical outcomes are reported only as Boolean Zone 0 risk labels.

Validation buildings400
Stage 2 buildings298
Stage 1 F10.915
Stage 2 macro F10.734

Fine-tuned model

Per-label validation performance

FIXED VALIDATION
Per-label positive-class F1, precision, and recall for the fine-tuned model on the fixed validation split
Direct labels are reported once. For the six hierarchical object outcomes, risk is derived per building when the parent object is present and inside Zone 0. Metrics use the final fixed seed-42 validation split, with uncertain human labels excluded per outcome.
Evaluation scopeAll 400 validation buildings400 evaluated · 0 excluded for this outcome

Confusion matrix

Stage 1 visual QC

F1 0.915
Model predictionGround truthNegativePositiveNegativePositive

All outcomes

Final validation metrics

OutcomeF1PrecisionRecallMax validation F1Precision / RecallEvaluatedExcluded
0.9150.9050.9260.9160.919 / 0.9134000
0.8310.8650.8000.8330.795 / 0.875197101
0.6370.7130.5760.6700.636 / 0.70728414
0.6460.7370.5750.6890.579 / 0.84928810
0.7250.7840.6740.7470.775 / 0.72126335
0.8930.9260.8620.9090.888 / 0.9312962
0.9530.9210.9870.9640.970 / 0.95828117
0.6340.4920.8890.7370.700 / 0.77823365
0.9360.9080.9650.9560.948 / 0.96528711
0.4940.4150.6110.5640.524 / 0.61124553
0.6490.5620.7690.7420.780 / 0.70828513
0.6730.5380.8990.7400.760 / 0.72223959

Threshold-tuned values maximize positive-class F1 on this same fixed validation split; the associated precision and recall appear below each F1. They are apparent validation results, not independent test estimates. Derived risk rows jointly tune the parent and conditional Zone 0 margins under the prompt-v3 scoring contract. Hover a value to see its selected threshold.

Final-model diagnostic

Trainval Diagnostic Atlas

IN-SAMPLE DIAGNOSTIC

Full confusion matrices for the 2,000-building labeled trainval pool, with representative high-confidence examples from each cell. This is an in-sample error-analysis view for model diagnosis and prompt refinement; it is not a validation or held-out performance estimate.

Stage 1 buildings2,000
Stage 2 QC-pass1,493
Selection seed42
Image views3 per sample
Loading the audited trainval diagnostic atlas…