Methodology
The monitor distinguishes publisher facts, tracker arithmetic and tracker judgement. Missing evidence does not become a negative assessment.
Indicator selection
Indicators are selected when they help answer a public question, have a clear unit and denominator, and can be linked to a stable official or reputable series. Comparable definitions take priority over the newest isolated number.
Collection, periods and revisions
Data are transcribed or imported from exact dated publications. Calendar years, financial years, months and rolling quarters remain separate. Provisional values are labelled. When a publisher revises a series, the current value and dependent calculations are updated together and material changes enter the correction log.
Calculations and missing data
Calculations retain their input sources, units and formula. Per-person values divide a matching-period amount by the matching population. Real values use the stated deflator or official chained-volume series. Missing values remain null and display “Data not yet integrated” or “Insufficient evidence to assess progress”; zero is used only when the source reports a genuine zero.
Commitment assessments
Each assessment requires the original promise, a target, baseline, latest comparable evidence, reporting period, primary source and an explanation.
Three separate fields
- Status is the controlled methodology classification.
- Tracker interpretation is the balanced evidence conclusion: Positive, Mixed, At risk or Unable to verify.
- Delivery phase describes the implementation stage without implying success: Established delivery, Early delivery, Early implementation or Multi-wave delivery.
The controlled status vocabulary is:
The measurable target is met.
Current evidence is consistent with a stated milestone or required pace.
Verified delivery activity exists, but evidence does not support a stronger conclusion.
Evidence indicates a material threat to the target or deadline.
A published milestone or deadline has been missed or formally moved.
Some, but not all, measurable elements are delivered.
Reliable evidence shows delivery has not begun.
Comparable evidence is insufficient; this is not failure.
The target cannot yet be observed on its stated timetable.
The original promise has been formally altered or withdrawn.
Tracker assessment — not an official government classification. Status, interpretation and phase are reviewed independently and are never automatically copied into one another.
Confidence
High uses recent primary evidence, a measurable target and independently reproducible progress. Medium has a useful official source but material scope, timing or revision limitations. Low relies on incomplete, old, indirect or non-comparable evidence.
Context, correlation and causation
Event dates and correlations identify patterns worth investigating. They do not establish that a policy or event caused an outcome. A causal claim also requires a plausible mechanism, correct timing, comparison with alternatives and evidence robust to competing explanations.
Relationship direction summaries
“Generally positive”, “Generally negative”, “Mixed”, “Little meaningful change”, “Insufficient evidence” and “Direction depends on perspective” are tracker interpretations based on explicit indicator metadata. Each indicator records a normally preferred direction, a small-change threshold where available and an interpretation limitation. Context-dependent measures do not receive an automatic positive or negative direction. The summary describes the latest loaded movement; it is not a causal conclusion or political recommendation.
When two series use different units, the levels view defaults to an index where the first common period equals 100. Dual axes are avoided. Same-unit series may be displayed on their shared original scale. Differences in frequency or calendar versus financial-year labelling remain visible.
Relationship annotations and usefulness feedback
Annotations are manually reviewed records with dates, types, relevant indicator IDs and source links. They provide timeline context only. They are not scraped headlines and timing alone does not show an effect. Usefulness ratings are kept separate from public evidence. Raw feedback is not public and aggregates should be suppressed until a disclosure-safe sample exists.
Public Priorities accounts and aggregation
Survey options use stable IDs so wording can change without breaking historic comparisons. Authenticated responses are validated against a versioned option registry and limited to one contribution per account and month. Public results are produced from grouped server-side counts; raw identifiable responses are not public. Demonstration figures remain clearly labelled and must not be presented as genuine submissions.
Review cycle
Sources display a retrieval or review date and an expected frequency where known. Headline datasets are reviewed when their source publishes; the broader register receives a documented editorial review before each public release.