They are judgements, not measurements
Nobody measured an event’s systemic importance with an instrument. A language model read the coverage, answered a fixed set of questions against a published rubric, and a fixed formula turned those answers into a number. That is more auditable than a score with no stated method — you can read every input and disagree with a specific one — but it is still judgement.They do not predict anything
propagation_potential is not a probability that something will spread. market_sensitivity is not a
forecast price move or a direction. Both score conditions present in the event as reported.
The frameworks are borrowed structure, not endorsement
We adapt the shape of established methods — Basel’s equal-weighted G-SIB categories, the ERCS barrier model, Richardson’s log-fatality scale, the reasonable-investor materiality test. That structure is what makes these scores legible to an analyst who already knows them. It does not mean the results carry those institutions’ authority.Coverage against hard registries is thin
Across 42,099 events, 2.2% touch a node matchable to a hard registry and 1.5% attach to an identifiable traded instrument. Treat all four as ordinal ranking signals — good for sorting and triage, not for asserting one event is 1.4× another.CAMEO+ only
The four metrics exist only on CAMEO+ events. Conflict events omit them rather than returning0 — a
0 would assert “measured, and it is the minimum”, which is false. A filter like
min_magnitude=6&event_family=conflict therefore returns an empty 200.
The error on each scale
We coded the same 59 events three times and measured how far each score moved when nothing changed.
Magnitude is the least stable, and nearly all of that sits on the
verbal path — verbal events swung up to
4 points between identical runs, while hazard averaged σ 0.14.
Reliability is not validity
Three runs agreeing proves the instrument is consistent, not that it is right. A scorer that was reliably wrong would look identical in the table above. Establishing correctness needs a rubric-blind human gold set graded by domain experts, which we have not built.propagation_potential in particular has no outcome validation — we do not currently measure whether
high-scoring events are followed by more downstream events than low-scoring ones.
Effective date and backfill
metric_version is stored but not currently returned by the API. Use the presence of
metrics.metric_inputs as the boundary marker — it appears only on v2 events.
