magnitude, systemic_importance, propagation_potential or
market_sensitivity in a model, an alert threshold, or a customer-facing product. What each one
individually does not claim is on its own page under Event metrics; this page
is what applies to all four.
Rubric scores, not measurements. For each metric, a model reads a handful of concrete sub-factors off
the source text — a fatality count, how substitutable a supplier is, whether a barrier that was holding got
breached, whether a traded claim is exposed — and states a reason for each one. Fixed, published formulas
then turn those sub-factors into the score. Nothing is fitted, estimated from price history, or forecast, and
nothing is a black box: any value is reconstructable by a third party from the sub-factors and the
published formula.
The frameworks give the rubric its structure — they do not make the values empirical.
magnitude follows domain severity scales (Richardson log-deaths, anchored MEPV-style on 0–10; an
EM-DAT-style realized-impact tier for hazards; the CAMEO coercion ladder for speech acts).
systemic_importance follows the BCBS G-SIB / ECB O-SII equal-weight indicator practice.
propagation_potential follows the ERCS barrier model and the ESRB systemic-risk shape.
market_sensitivity follows the reasonable-investor materiality test. We adapt these frameworks for
their vocabulary and structure. None of these institutions endorses, reviews, or is connected to this
work, and their use does not make our values measured.
Coverage — why “measured” is the wrong word. Over 42,099 events in a 45-day window: a hard registry
attribute — the observable that would make systemic_importance measured rather than judged — resolves for
about 2.2% of events, and a traded instrument, behind market_sensitivity’s exposure gate, resolves
for about 1.5%. For everything else the score is a judged rubric reading. Treat all four as ordinal
ranking signals: use them to sort, filter and triage, not to assert a quantity about a single event.
Reliability — the smallest difference worth reading. The same 59 events, coded three times with identical
prompts and formulas, so everything that moved is noise:
Never read a single-event difference smaller than the last column. Distribution-level comparisons are far
steadier than individual events, so aggregate claims hold well below these thresholds — but a claim that one
event moved does not.
magnitude is within-domain only. Each domain has its own anchored ladder, so a conflict magnitude of 8
and an economic magnitude of 8 are not the same “size.” Use significance — the family-scoped composite —
whenever you rank events across domains. And magnitude is null when no severity observable was found:
null means unknown, never zero.
The four are not fully independent axes. Under the ESRB framing, market_sensitivity is a propagation
channel — common exposure and confidence effects expressed in prices — so it is closer to a special case of
propagation_potential than an orthogonal dimension. Both are served because market_sensitivity carries
substantial independent variance and fires domain-appropriately, but do not treat the set as four independent
dimensions in a model or a weighted score.
Magnitude is the least stable, and nearly all of that sits on the verbal path — verbal events swung
up to 4 points between identical runs, while hazard averaged σ 0.14.
Reliability is not validity
Three runs agreeing proves the instrument is consistent, not that it is right. A scorer that was reliably wrong would look identical in the table above. Establishing correctness needs a rubric-blind human gold set graded by domain experts, which we have not built.propagation_potential in particular has no outcome validation — we do not currently measure
whether high-scoring events are followed by more downstream events than low-scoring ones.
CAMEO+ only
The four metrics exist only on CAMEO+ events. Conflict events omit them rather than returning0 — a 0 would assert “measured, and it is the minimum”, which is false. A filter like
magnitude_min=6&event_family=conflict therefore returns an empty 200.
Effective date and backfill
metric_version is returned on the event card whenever it is set, so it is the direct boundary
marker: v2-2026-07-24 means these formulas produced the scores, and an absent value means the
event predates them. The presence of metrics.metric_inputs tells you the same thing — it appears
only on v2 events — and additionally tells you the per-sub-factor why is available.
