UC Berkeley · ZoneZero

Pre-Eaton Fire Zone Zero Risks

Pre-Eaton analysis

Held-out test performance and population inference summary

Locked test-set comparisons for four model settings and audited inference summaries for 55,608 study buildings. Building-level results are available on the interactive map.

Study coverage

Building and parcel populations

Research unitStudy totalQC passedNo SVINo building
Building55,60823,200 (41.7%)10,438 (18.8%)
Parcel34,94017,984 (51.5%)4,365 (12.5%)2,998 (8.6%)

Model evaluation

Held-out test performance

LOCKED TEST SET
ModelSettingStage 1 F1Stage 2 macro-F1Stage 2 macro precisionStage 2 macro recall
Qwen3.5-9BFine-tuned0.8530.7450.8430.701
Qwen3.5-9BZero-shot0.8310.4140.6720.345
GPT-5.6 SolZero-shot0.4860.6950.450
GPT-5.6 LunaZero-shot0.5070.7650.440

Stage 1 uses all 1,000 held-out buildings. Stage 2 uses the locked cohort of 526 human-QC-pass buildings; uncertain human labels are excluded separately by outcome and the public CSV reports the scored denominator for each model-outcome pair. Qwen fine-tuned uses the frozen threshold profile selected on validation only; zero-shot models use their native binary decisions. GPT-5.6 Sol and Luna were evaluated only for conditional Stage 2, so Stage 1 is not reported. Stage 2 F1, precision, and recall are equal-weight positive-class macros over the 11 final outcomes shown below. A dash means not reported, not zero. The best value in each metric column is bold. Sol produced five invalid structured responses, which remain penalized.

Open validation results and diagnostic atlas →

Four-model comparison

Per-label held-out performance

526 TEST BUILDINGS
Per-label positive-class F1 for Qwen3.5-9B fine-tuned, Qwen3.5-9B zero-shot, GPT-5.6 Sol zero-shot, and GPT-5.6 Luna zero-shot on the held-out test set
Positive-class F1 is shown for all 11 final Stage 2 outcomes in the locked cohort. Qwen fine-tuned uses the frozen validation-selected threshold profile; zero-shot models use native binary decisions. Uncertain human labels are excluded independently by outcome, and the public CSV reports the scored denominator for every model-outcome pair. Missing or invalid model outputs are penalized. The source table is available as a public CSV.

Formal Eaton inference

Population coverage and predicted label prevalence

AUDITED MODEL OUTPUT
Stage 1 QC coverage and selected panorama rank
Rank 2 and Rank 3 were evaluated only after earlier panorama candidates failed Stage 1 visual QC.
Predicted prevalence of 11 final Stage 2 outcomes across QC-pass Eaton buildings
Direct labels are shown once. For each hierarchical outcome, risk is positive only when the parent object is present and inside Zone 0; all rates use the 23,200-building Stage 2 denominator.