AllenNeuralDynamics / AllenNeuralDynamics/aind-dynamic-foraging-bfm-dispatcher
Study 09: top-seven plus species-balanced external transfer panel
- 主要語言
- Python
- 星號
- 0
- 分支
- 0
- 平均合併
- 2 小時 9 分鐘
- 30 天內合併 PR
- 31
描述
## Context
The first three Study 09 targets showed promising transfer. Stage A expanded the benchmark only when a complete released cohort fit the existing v1 or v2 split contract. Draft PR: #139.
## Cohort audit
### Admitted
- [x] Lebedeva — mouse, v1
- [x] Beron — mouse, v1
- [x] Kwak — mouse, v1
- [x] Miller — rat, v1
- [x] Findling — human, v1
- [x] Tang — macaque, v1
- [x] Alsiö cohorts II–V — rat, v1
- [x] Eckstein — human, v2
- [x] Costa — macaque, v1
- [x] López-Yépez mouse — mouse, v1
Grossman, Chen, and Zid are reused from the completed first round.
### Skipped without schema v3
- [x] López-Yépez human — mixed 1–4 sessions per subject; complete cohort is neither v1 nor v2.
- [x] Shin — released sessions cannot be mapped back to subject identity.
- [x] Alsiö cohort VI — chronological real-session identity is unavailable.
- [x] Hattori — pinned public behavior derivative is blocked by authorization/WAF responses.
- [x] Samejima — no stable public trial-level choice/reward release located.
- [x] Kubanek — excluded as non-binary.
## Validation and execution
- [x] Pin release identities, terms, checksums, inclusion rules, exact counts, and deterministic manifests.
- [x] Pass canonical binary choice/reward and v1/v2 contract validation.
- [x] Pass local real-subject CPU smoke tests.
- [x] Pass one D=614, seed-0 GRU GPU smoke for all ten admitted expansion cohorts.
- [x] Complete and freeze full GRU D={10,30,100,300,614} × three-seed matrices on Beaker GPU.
- [x] Complete all common-Q fits on Allen HPC CPU.
- [x] Freeze exact ordered GRU/Q trial-key parity and publish Results 2 and 3.
- [ ] Review the Stage-A decision report before selecting any new author model.
Draft PR #139 contains the completed common-Q fits, frozen parity evidence, and reports; those acceptance boxes remain open until the PR is merged.
## Stage-A decision result
At D=614, GRU wins four unadjusted subject-paired comparisons, common Q wins five, and four are unresolved. All 13 cohorts improve in trial-pooled GRU likelihood from D=10 to D=614, but the largest D is not always the curve maximum. The embedding report shows every external cohort farther from the source distribution than held-out AIND mice in every seed.
No CPU-only work was submitted to Beaker. No new split abstraction or author-model family was introduced.
貢獻指南
這個儲存庫沒有索引到貢獻指南
研究方向
Start by reviewing draft PR #139, its frozen parity evidence, completed common-Q fits, and Results 2 and 3, then read the Stage-A decision report. Confirm the reported comparisons and embedding findings, complete the open review item, and decide whether a new author model should be selected.
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- python
- 領域
- machine-learning
- Issue 類型
- 功能
- 難度
- 5/5
- 預估耗時
- 一週以上
- 活躍度
- 停滯
- 描述清晰度
- 需要釐清
- 新手友好度
- 20/100