Experimental research artifact. Not an alerting service. Do not use for safety-critical decisions. For authoritative information see weather.gov, nhc.noaa.gov, or your local emergency management authority. Read more.

Agent

Walk through a forecaster detection pass, then inspect critic fitness and proposals. Critic fire stays operator-gated.

Forecaster

A walkthrough of a detection pass — reasoning over signals, running detectors, and the forecasts that result.

Transcript

No steps yet.

Loading map…

Critic

Transcript over skill fitness and lineage — no map. Targets the mutator and generator; proposals still need human review.

Transcript

No steps yet.

14d Brier
SkillBrierHits
wildfire_rapid_growth0.4980
wildfire_risk_elevated0.5840
Lineage
Loading lineage…

System

Curator
Enabled
Active forecasts
15013
Evaluations
1000
last Oct 5, 7:00 AM UTC
LLM (10m)
No calls

Approval queue7 pending

  • wildfire_wind_fire_weather_fusionv0Jul 21, 4:00 AM UTC

    This skill targets the two highest-priority uncovered signal types — `aifs/high_wind_corridor` and `open_meteo/fire_weather` — which are the most direct atmospheric precursors to wildfire spread and ignition. High winds drive fire behavior (rate of spread, spotting distance) while FWI is the globally accepted composite index encoding fuel moisture, wind, and atmospheric dryness. Fusing them spatially provides compound-risk detection that neither signal achieves alone. Cross-referencing with FIRMS MODIS/VIIRS hotspots grounds the atmospheric risk in observed fire activity, reducing false-positive pressure. The two-parameter run(now, db) signature matches the required interface, and all probabilities are clamped at 0.85 to satisfy the DB CHECK constraint.

  • Wildfire Risk Elevatedv1Jul 19, 4:00 AM UTC

    **Root cause analysis from failure traces:** All three worst-Brier forecasts (brier=0.7225, p=0.85, outcome=false_positive) share the same run timestamp and the same 24798 hotspot input. The culprits are: 1. **Clusters too small to be credible signals** — cluster sizes were 14 and 20 (traces ec2c16c4, 9a3a4025). With MIN_SAMPLES=5 and EPS_KM=10, DBSCAN was forming tiny, spatially tight clusters from noise. These intersected broad ECMWF `fire_weather_grid` polygons (which cover large regions), resulting in false positives. The third worst trace (b25d54f0, size=86) also false-positived, driven entirely by the over-inflated base probability. 2. **Probability formula was too aggressive** — `base=0.40 + cluster_size_factor=0.30 + polygon_overlap_factor=0.30` pegged forecasts at the 0.85 cap far too easily. Any cluster ≥20 pts with ≥2 polygon intersections hits 0.85 regardless of cluster quality. 3. **Only LIKE 'firms%' used** — the inventory specifies explicit sources `firms_modis` and `firms_viirs`; switching to `IN (...)` is cleaner and avoids potential pattern-match leakage. **Changes made:** | Parameter | v1 | v2 | Rationale | |---|---|---|---| | `EPS_KM` | 10.0 | 15.0 | Merges nearby sparse noise into fewer, denser clusters | | `MIN_SAMPLES` | 5 | 12 | Suppresses formation of tiny noisy clusters | | `MIN_CLUSTER_SIZE` | — | 25 | Hard gate: clusters 14 and 20 (both FP) would be dropped entirely | | `base` probability | 0.40 | 0.30 | Reduce systematic over-confidence | | `cluster_size_factor` cap | 0.30 | 0.25 | Less headroom for size inflation | | `cluster_size_factor` rate | 0.02/hotspot | 0.015/hotspot | Slower ramp-up | | `polygon_overlap_factor` cap | 0.30 | 0.20 | Tighter cap; broad ECMWF polygons intersect too easily | | `polygon_overlap_factor` rate | 0.15/match | 0.10/match | Less aggressive per-match bonus | | Internal probability cap | 0.85 | 0.80 | Stay below DB cap with margin | | Hotspot source filter | `LIKE 'firms%'` | `IN ('firms_modis', 'firms_viirs')` | Explicit inventory-correct sources | | Added polygon source | — | `open_meteo/fire_weather` | Additional true-positive coverage from inventory | **Brier impact estimate:** The three worst cases (all p=0.85, outcome=FP) contribute brier=(0.85)²=0.7225 each. With MIN_CLUSTER_SIZE=25, clusters of size 14 and 20 are dropped entirely (brier→0). The size-86 cluster at p=0.85 would now score p≈0.30 + 0.015*(86-12)=0.30+0.111≈0.411 + min(0.20,0.10*1)=0.10 → p≈0.511, reducing its brier from 0.7225 to ≈0.261. Overall mean Brier should drop substantially.

  • Wildfire Rapid Growthv2Jul 18, 4:00 AM UTC

    **Root-cause analysis of the three worst FP forecasts (Brier=0.7225, p=0.85):** All three share identical `inputs`: `hotspot_count_last_24h=15975`, `hotspot_count_prior_24h=24891`. This means the **global** hotspot field was sharply contracting (ratio ≈ 0.64) while the skill was still emitting high-confidence rapid-growth forecasts. That global-decline context was completely ignored in v2, making local cell growth ratios (3.5×–12.6×) grossly misleading. **Change 1 — Global hotspot-trend gate (primary fix)** `GLOBAL_DECLINE_RATIO_THRESHOLD = 0.85`: if `last_24h / prior_24h < 0.85`, the overall hotspot field is contracting and any local growth is almost certainly noise or artifact. Return `[]` immediately. This single change would have suppressed all three worst FP traces and many others with similar global-decline signatures. **Change 2 — Noisy-signal-field cap** `threshold_met_count=26` in all three traces means 26 cells passed the SQL filter simultaneously — a strong sign of spatially diffuse noise rather than a genuine fire front. Added `MAX_CANDIDATE_CELLS=15`: if more than 15 cells pass, cap probability at `NOISY_SIGNAL_PROB_CAP=0.45`. **Change 3 — Raise `MIN_BASELINE_COUNT` to 5, add `MIN_DAY_T_ABS=20`** The SQL filter now also requires `day_t >= 20` (= 5 × 2²), enforcing that the current window must be substantive in absolute terms, not just a large ratio over a tiny base. **Change 4 — Sigmoid-rescaled growth factor** Replaced the linear `0.04*(compound-4)` formula with a sigmoid centred at 8.0. Extreme compound ratios (e.g. 12.5×3.8 ≈ 47) were inflating `growth_factor` toward its cap and pushing probability to ceiling — the sigmoid saturates gracefully instead. **Change 5 — Lower `PROB_CEIL` to 0.65, base to 0.28** All three FP had p=0.85 (v1) and v2 claimed 0.75 ceiling but scoring still approached it. Floor/ceiling tightened and base reduced so the scoring headroom is smaller. **Change 6 — High-wind-corridor bonus (+0.05)** `aifs/high_wind_corridor` is in the signal inventory and is highly correlated with true rapid-growth events. Added a small positive bonus so genuinely dangerous cells can still score meaningfully, compensating somewhat for the tighter ceiling. **Change 7 — Broadened fire-weather gate signal types** Gate now explicitly matches `(signal_type, source)` pairs verbatim from the inventory: `fire_weather_grid`/`aifs`, `fire_weather_grid`/`ecmwf_open_data`, `fire_weather`/`open_meteo`, `fire_warning`/`nws_alerts`, `high_wind_corridor`/`aifs` — ensuring no valid forecast is silently dropped due to a missing source string.

  • Wildfire Rapid Growthv2Jul 17, 4:00 AM UTC

    ## What changed and why ### Root-cause diagnosis from failure traces All three worst forecasts (brier=0.7225, p=0.85, outcome=false_positive) share a telltale signature: - **Global hotspot count was declining**: `last_24h=21012 < prior_24h=24981` — the overall fire environment was cooling off globally, yet the skill still fired on local cell growth. - **p=0.85** was the output — the old PROB_CEIL of 0.85 was the exact ceiling hit, meaning the scoring saturated in the wrong direction. - **growth_ratios of 1.65, 2.538, 3.25** — the 1.65 cell in trace 09b353bc would have been blocked by raising GROWTH_THRESHOLD to 2.5; the 2.538 cell would be blocked in a declining-global environment by the +0.5 bump to 3.0. ### Specific changes (v2 → v3) 1. **Global-decline guard (new)**: If `last_24h < prior_24h` globally, `effective_threshold` is raised by +0.5 (2.5 → 3.0). This directly addresses all three worst-case traces, which all shared `last_24h=21012 < prior_24h=24981`. In a globally-declining hotspot environment, local cell growth is more likely to be noise. 2. **GROWTH_THRESHOLD 2.0 → 2.5**: The v2 threshold of 2.0 still admitted growth_ratio=1.65 (trace 09b353bc). Raising to 2.5 blocks low-confidence growers. In a declining environment this further rises to 3.0. 3. **MIN_BASELINE_COUNT 3 → 5**: Sparse cells with very few baseline hotspots generate noisy ratios; requiring 5 hotspots in day_t2 filters these out. 4. **Stricter fire-weather gate**: When globally declining AND `last_24h > 15000`, require at least **2** supporting fire-weather/warning signals per cell (not just 1). High-volume globally-declining fire days need stronger corroboration to emit. 5. **Probability scoring tightened**: - `base` 0.35 → 0.30 - `persistence_factor` cap 0.15 → 0.12, coefficient 0.005 → 0.004 - `growth_factor` trigger compound threshold raised from 4.0 → 6.25, cap 0.25 → 0.20 - New `ratio_bonus`: contributes only when both day-over-day ratios substantially exceed threshold (≥2.5), capped at 0.15 — rewards genuinely explosive growth without inflating moderate cases - Net max achievable probability = 0.30 + 0.12 + 0.20 + 0.15 = 0.77, but capped at PROB_CEIL 6. **PROB_CEIL 0.75 → 0.70**: The saturated p=0.85 on false positives argues for a lower ceiling. Even in v2 the 0.75 cap was too generous given the base+factors achievable. 7. **PROB_FLOOR 0.10 → 0.15**: Marginally raise floor to avoid noisy near-zero emissions. 8. **SKILL_VERSION bumped to 3**. These changes are targeted directly at the failure mode: a globally-declining fire environment where local cells showed moderate growth ratios (1.65–3.25) and produced overconfident false-positive forecasts.

  • Typhoon Landfall Imminentv1Jul 7, 4:00 AM UTC

    ## Root-cause analysis of false-positive failures All three worst evaluations (Brier 0.534, 0.490, 0.479) shared the same structural failure: 1. **Probability over-inflation**: All three produced p ≈ 0.69–0.73 despite being false positives. The v1 formula `base=0.45 + n_bonus + pop_bonus` could reach ~0.72 for any storm with 5+ cities in the cone and moderate population (1–4M), regardless of how distant those cities were from the storm's actual position. 2. **Distant cities inflating population signal**: The cone traces show top cities at 340–548 km from the storm centre. These cities were only captured because the 72h buffer radius (200 km) was added to a ~350–500 km forward displacement. For slow storms (3–10 km/h as seen in all three FPs), the cone area was 143k–255k km² — largely uncertainty buffer, not actual track. Yet the raw intersected population of 700k–4M drove the probability bonus without any distance penalty. 3. **Flat base too high**: `base=0.45` gives a 45% prior before any city data is considered. NHC 72h landfall climatology for any single storm is much lower; a prior of 0.30 is more calibrated. 4. **Cone buffer too generous for slow storms**: All three FPs had speed ≤ 10 km/h. At 72h, a 10 km/h storm only travels 720 km, but the 200 km buffer expands the cone width dramatically for lateral cities. v1 applied the same buffer regardless of speed. ## Changes made - **BASE_PROBABILITY**: 0.45 → 0.30 (directly reduces systematic over-forecast) - **Distance-weighted effective population** (`effective_population()`): exponential decay with 300 km half-life. At 500 km (the worst FP case), weight is ~0.19×, so a 428k-population city contributes only ~81k to the signal rather than the full 428k. The three FP cones would yield effective_pop well below the new gate. - **MIN_EFFECTIVE_POP gate**: 50,000 — forecasts are suppressed if the distance-attenuated population is too low, preventing distant-buffer false positives. - **Speed-adaptive cone narrowing** (`effective_cone_steps()`): for slow storms (< 15 km/h), buffer radii beyond 24h are scaled by a factor in [0.65, 1.0] proportional to speed. At 5 km/h (FP average), factor ≈ 0.72 — shrinking the 72h buffer from 185 km to ~133 km, reducing lateral city capture. - **Cone buffer base values** slightly tightened (200→185 km at 72h, etc.) to better match NHC 5-year average track error growth. - **n_bonus cap** reduced (0.20 → 0.10) and per-city rate reduced (0.02 → 0.015). - **pop_bonus** now computed on log10(effective_pop / 100k) not raw population, with scale 0.08 vs old 0.05, but effective_pop is typically 5–10× smaller than raw total_pop for distant-buffer cases. - **LATEST_BULLETIN_WINDOW_HOURS**: 6 → 12 to avoid advisory-cycle coverage gaps. - Signal source remains `nhc` / `cyclone_advisory` per inventory.

  • Wildfire Rapid Growthv2Jun 22, 1:24 AM UTC

    **Root cause from traces:** All three worst forecasts (brier=0.7225, p=0.85, outcome=false_positive) share identical inputs (32,833 hotspots last 24h vs 27,643 prior 24h — only ~19% global growth) and were emitted for cells with moderate growth ratios: 2.414, 3.862, and 6.316. The v2 GROWTH_THRESHOLD of 2.0 allowed the 2.414-ratio cell through, and the probability model still reached the PROB_CEIL of 0.75 → 0.85 in v1. The fire-weather gate requiring only 1 matching signal was insufficient to block them. **Changes made and why:** 1. **GROWTH_THRESHOLD 2.0 → 2.5**: The 2.414-ratio cell (forecast 7756fa52) would be pruned entirely at the SQL level — it barely exceeded 2.0 but doesn't reach 2.5. This directly eliminates one of the three worst forecasts before any downstream processing. 2. **MIN_BASELINE_COUNT 3 → 5**: The worst trace cells triggered at day_t2 counts that, while ≥3, were still modest. Requiring ≥5 hotspots in the oldest window ensures only established cluster growth is detected, not transient flares. 3. **MIN_DAY_T_ABS = 10 (new absolute floor)**: Added a direct SQL filter `day_t >= %(min_day_t)s` requiring at least 10 hotspots in the current 24h window. Purely relative growth (e.g., 2→5→13) can still be noise; an absolute minimum tethers significance to real fire activity volume. 4. **Fire-weather gate: count ≥ 2 required when compound < 8.0**: The v2 gate required only 1 matching signal. For borderline compound ratios below 8.0 (which covered cells 1 and 2 in the worst traces with compound ~5.8 and ~3.7), we now require 2 corroborating signals. High-confidence cells (compound ≥ 8.0, like cell 0 with ratio 6.316 giving compound ~39) still only need 1 signal. 5. **Broadened fire-weather signal sources**: Added `high_wind_corridor` (from `aifs`) to the gate query — this signal type is in the inventory and is physically correlated with rapid fire spread, giving the gate more opportunities to confirm real events while still blocking false positives. 6. **Probability base 0.35 → 0.30, growth_factor compound threshold 4.0 → 6.0, persistence coefficient 0.005 → 0.004**: The three worst forecasts all hit p=0.85 in v1 (PROB_CEIL). Even in v2 with PROB_CEIL=0.75, high compound ratios easily reached the cap. Raising the growth_factor compound threshold to 6.0 means cells with compound < 6.0 contribute zero growth_factor; only genuinely explosive growth adds to the score. 7. **PROB_CEIL 0.75 → 0.70**: All three worst FPs were at p=0.85. A lower ceiling limits the maximum Brier penalty for any single overconfident false positive that slips through. 8. **Cell deduplication by IoU > 0.60**: When a large fire spans multiple adjacent 50km grid cells, they can generate near-duplicate forecasts that all fail together. Keeping only the highest day_t cell per overlapping cluster reduces correlated false positives and unnecessary spread of probability mass.

  • Wildfire Risk Elevatedv1Jun 8, 12:58 AM UTC

    fixture mutant

Attribution: FIRMS, NWS, Open-Meteo, NHC, JTWC, ECMWF, AIFS, GDACS — see How it works. 4 skills with recent activity.