PolicyEngine / PolicyEngine/microcosm
UK reform-validation suite: CenTax/HMRC/OBR published costings + the #462 backtest families as shipped payload
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 0
- Forks
- 4
- Avg merge
- 1d 3h
- Merged PRs (30d)
- 94
Description
The UK validation surface is thin next to the US (policyengine.py#462's UK run had ~17 candidate metrics vs the US's state-legislative suite #319, JCT/OBBBA rows, and SPM backtests #348). uk-v2's #133 shows the pattern that fills it: its CenTax replication cross-checked an engine-computed reform cost against a published analysis to the pound (£28.55bn) on a defined subset.
Proposal: a UK reform-validation suite in the reform-validation payload, seeded from published, citable costings —
- CenTax published analyses (with the same attribution discipline #133 used),
- HMRC policy costings / OBR Economic and Fiscal Outlook policy tables,
- the #462 out-of-sample families (HBAI poverty, DWP caseloads/expenditure, HMRC liabilities) as baseline backtests — promoting the adjudication pack's benchmark set into the shipped payload the way #348 did for US poverty.
Out-of-sample discipline per #348/#302: diagnostics and promotion-gate evidence, never calibration targets.
Refs: policyengine.py#462, #348, #319 (US pattern), #302, policyengine-uk-v2#133.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the reform-validation payload and the patterns referenced in policyengine.py#462, #348, and #302, then compare the UK proposal with policyengine-uk-v2#133. Done means a shipped UK validation suite with citable CenTax, HMRC, and OBR costings plus the specified out-of-sample backtest families, kept as diagnostics and promotion-gate evidence rather than calibration targets.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- testing-qa
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100