Public beta: Methods and coverage remain under review.

Methodology

The monitor distinguishes publisher facts, tracker arithmetic and tracker judgement. Missing evidence does not become a negative assessment.

Indicator selection

Indicators are selected when they help answer a public question, have a clear unit and denominator, and can be linked to a stable official or reputable series. Comparable definitions take priority over the newest isolated number.

Collection, periods and revisions

Data are transcribed or imported from exact dated publications. Calendar years, financial years, months and rolling quarters remain separate. Provisional values are labelled. When a publisher revises a series, the current value and dependent calculations are updated together and material changes enter the correction log.

Calculations and missing data

Calculations retain their input sources, units and formula. Per-person values divide a matching-period amount by the matching population. Real values use the stated deflator or official chained-volume series. Missing values remain null and display “Data not yet integrated” or “Insufficient evidence to assess progress”; zero is used only when the source reports a genuine zero.

Commitment assessments

Each assessment requires the original promise, a target, baseline, latest comparable evidence, reporting period, primary source and an explanation.

Three separate fields

The controlled status vocabulary is:

Delivered
The measurable target is met.
On track
Current evidence is consistent with a stated milestone or required pace.
In progress
Verified delivery activity exists, but evidence does not support a stronger conclusion.
At risk
Evidence indicates a material threat to the target or deadline.
Delayed
A published milestone or deadline has been missed or formally moved.
Partially delivered
Some, but not all, measurable elements are delivered.
Not started
Reliable evidence shows delivery has not begun.
Unable to verify
Comparable evidence is insufficient; this is not failure.
Not yet measurable
The target cannot yet be observed on its stated timetable.
Commitment changed or withdrawn
The original promise has been formally altered or withdrawn.

Tracker assessment — not an official government classification. Status, interpretation and phase are reviewed independently and are never automatically copied into one another.

Confidence

High uses recent primary evidence, a measurable target and independently reproducible progress. Medium has a useful official source but material scope, timing or revision limitations. Low relies on incomplete, old, indirect or non-comparable evidence.

Context, correlation and causation

Event dates and correlations identify patterns worth investigating. They do not establish that a policy or event caused an outcome. A causal claim also requires a plausible mechanism, correct timing, comparison with alternatives and evidence robust to competing explanations.

Relationship direction summaries

“Generally positive”, “Generally negative”, “Mixed”, “Little meaningful change”, “Insufficient evidence” and “Direction depends on perspective” are tracker interpretations based on explicit indicator metadata. Each indicator records a normally preferred direction, a small-change threshold where available and an interpretation limitation. Context-dependent measures do not receive an automatic positive or negative direction. The summary describes the latest loaded movement; it is not a causal conclusion or political recommendation.

When two series use different units, the levels view defaults to an index where the first common period equals 100. Dual axes are avoided. Same-unit series may be displayed on their shared original scale. Differences in frequency or calendar versus financial-year labelling remain visible.

Relationship annotations and usefulness feedback

Annotations are manually reviewed records with dates, types, relevant indicator IDs and source links. They provide timeline context only. They are not scraped headlines and timing alone does not show an effect. Usefulness ratings are kept separate from public evidence. Raw feedback is not public and aggregates should be suppressed until a disclosure-safe sample exists.

Public Priorities accounts and aggregation

Survey options use stable IDs so wording can change without breaking historic comparisons. Authenticated responses are validated against a versioned option registry and limited to one contribution per account and month. Public results are produced from grouped server-side counts; raw identifiable responses are not public. Demonstration figures remain clearly labelled and must not be presented as genuine submissions.

Review cycle

Sources display a retrieval or review date and an expected frequency where known. Headline datasets are reviewed when their source publishes; the broader register receives a documented editorial review before each public release.