Skip to main content
Read this before putting magnitude, systemic_importance, propagation_potential or market_sensitivity in a model, an alert threshold, or a customer-facing product. What each one individually does not claim is on its own page under Event metrics; this page is what applies to all four.
How to read magnitude, systemic_importance, propagation_potential and market_sensitivity. These four are rubric scores, not measurements. Read this before using any of them in a threshold, a model feature, or a customer-facing claim.
Rubric scores, not measurements. For each metric, a model reads a handful of concrete sub-factors off the source text — a fatality count, how substitutable a supplier is, whether a barrier that was holding got breached, whether a traded claim is exposed — and states a reason for each one. Fixed, published formulas then turn those sub-factors into the score. Nothing is fitted, estimated from price history, or forecast, and nothing is a black box: any value is reconstructable by a third party from the sub-factors and the published formula. The frameworks give the rubric its structure — they do not make the values empirical. magnitude follows domain severity scales (Richardson log-deaths, anchored MEPV-style on 0–10; an EM-DAT-style realized-impact tier for hazards; the CAMEO coercion ladder for speech acts). systemic_importance follows the BCBS G-SIB / ECB O-SII equal-weight indicator practice. propagation_potential follows the ERCS barrier model and the ESRB systemic-risk shape. market_sensitivity follows the reasonable-investor materiality test. We adapt these frameworks for their vocabulary and structure. None of these institutions endorses, reviews, or is connected to this work, and their use does not make our values measured. Coverage — why “measured” is the wrong word. Over 42,099 events in a 45-day window: a hard registry attribute — the observable that would make systemic_importance measured rather than judged — resolves for about 2.2% of events, and a traded instrument, behind market_sensitivity’s exposure gate, resolves for about 1.5%. For everything else the score is a judged rubric reading. Treat all four as ordinal ranking signals: use them to sort, filter and triage, not to assert a quantity about a single event. Reliability — the smallest difference worth reading. The same 59 events, coded three times with identical prompts and formulas, so everything that moved is noise: Never read a single-event difference smaller than the last column. Distribution-level comparisons are far steadier than individual events, so aggregate claims hold well below these thresholds — but a claim that one event moved does not. magnitude is within-domain only. Each domain has its own anchored ladder, so a conflict magnitude of 8 and an economic magnitude of 8 are not the same “size.” Use significance — the family-scoped composite — whenever you rank events across domains. And magnitude is null when no severity observable was found: null means unknown, never zero. The four are not fully independent axes. Under the ESRB framing, market_sensitivity is a propagation channel — common exposure and confidence effects expressed in prices — so it is closer to a special case of propagation_potential than an orthogonal dimension. Both are served because market_sensitivity carries substantial independent variance and fires domain-appropriately, but do not treat the set as four independent dimensions in a model or a weighted score.
Never read a single-event difference smaller than the last column above. It is indistinguishable from the instrument talking to itself.
Magnitude is the least stable, and nearly all of that sits on the verbal path — verbal events swung up to 4 points between identical runs, while hazard averaged σ 0.14.

Reliability is not validity

Three runs agreeing proves the instrument is consistent, not that it is right. A scorer that was reliably wrong would look identical in the table above. Establishing correctness needs a rubric-blind human gold set graded by domain experts, which we have not built. propagation_potential in particular has no outcome validation — we do not currently measure whether high-scoring events are followed by more downstream events than low-scoring ones.

CAMEO+ only

The four metrics exist only on CAMEO+ events. Conflict events omit them rather than returning 0 — a 0 would assert “measured, and it is the minimum”, which is false. A filter like magnitude_min=6&event_family=conflict therefore returns an empty 200.

Effective date and backfill

metric_version is returned on the event card whenever it is set, so it is the direct boundary marker: v2-2026-07-24 means these formulas produced the scores, and an absent value means the event predates them. The presence of metrics.metric_inputs tells you the same thing — it appears only on v2 events — and additionally tells you the per-sub-factor why is available.