PolicyEngine / PolicyEngine/microcosm

UK reform-validation suite: CenTax/HMRC/OBR published costings + the #462 backtest families as shipped payload

Open
#365 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
0
Forks
4
Avg merge
1d 3h
Merged PRs (30d)
94

Description

The UK validation surface is thin next to the US (policyengine.py#462's UK run had ~17 candidate metrics vs the US's state-legislative suite #319, JCT/OBBBA rows, and SPM backtests #348). uk-v2's #133 shows the pattern that fills it: its CenTax replication cross-checked an engine-computed reform cost against a published analysis to the pound (£28.55bn) on a defined subset.

Proposal: a UK reform-validation suite in the reform-validation payload, seeded from published, citable costings —

  • CenTax published analyses (with the same attribution discipline #133 used),
  • HMRC policy costings / OBR Economic and Fiscal Outlook policy tables,
  • the #462 out-of-sample families (HBAI poverty, DWP caseloads/expenditure, HMRC liabilities) as baseline backtests — promoting the adjudication pack's benchmark set into the shipped payload the way #348 did for US poverty.

Out-of-sample discipline per #348/#302: diagnostics and promotion-gate evidence, never calibration targets.

Refs: policyengine.py#462, #348, #319 (US pattern), #302, policyengine-uk-v2#133.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the reform-validation payload and the patterns referenced in policyengine.py#462, #348, and #302, then compare the UK proposal with policyengine-uk-v2#133. Done means a shipped UK validation suite with citable CenTax, HMRC, and OBR costings plus the specified out-of-sample backtest families, kept as diagnostics and promotion-gate evidence rather than calibration targets.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
testing-qa
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.