AllenNeuralDynamics / AllenNeuralDynamics/aind-dynamic-foraging-bfm-dispatcher

Study 09: top-seven plus species-balanced external transfer panel

未關閉
#134 2 則留言 0 個 reaction 已指派 1 人 已被 @hanhou 認領 在 GitHub 檢視
evaluation extension priority:P1
主要語言
Python
星號
0
分支
0
平均合併
2 小時 9 分鐘
30 天內合併 PR
31

描述

## Context

The first three Study 09 targets showed promising transfer. Stage A expanded the benchmark only when a complete released cohort fit the existing v1 or v2 split contract. Draft PR: #139.

## Cohort audit

### Admitted

- [x] Lebedeva — mouse, v1
- [x] Beron — mouse, v1
- [x] Kwak — mouse, v1
- [x] Miller — rat, v1
- [x] Findling — human, v1
- [x] Tang — macaque, v1
- [x] Alsiö cohorts II–V — rat, v1
- [x] Eckstein — human, v2
- [x] Costa — macaque, v1
- [x] López-Yépez mouse — mouse, v1

Grossman, Chen, and Zid are reused from the completed first round.

### Skipped without schema v3

- [x] López-Yépez human — mixed 1–4 sessions per subject; complete cohort is neither v1 nor v2.
- [x] Shin — released sessions cannot be mapped back to subject identity.
- [x] Alsiö cohort VI — chronological real-session identity is unavailable.
- [x] Hattori — pinned public behavior derivative is blocked by authorization/WAF responses.
- [x] Samejima — no stable public trial-level choice/reward release located.
- [x] Kubanek — excluded as non-binary.

## Validation and execution

- [x] Pin release identities, terms, checksums, inclusion rules, exact counts, and deterministic manifests.
- [x] Pass canonical binary choice/reward and v1/v2 contract validation.
- [x] Pass local real-subject CPU smoke tests.
- [x] Pass one D=614, seed-0 GRU GPU smoke for all ten admitted expansion cohorts.
- [x] Complete and freeze full GRU D={10,30,100,300,614} × three-seed matrices on Beaker GPU.
- [x] Complete all common-Q fits on Allen HPC CPU.
- [x] Freeze exact ordered GRU/Q trial-key parity and publish Results 2 and 3.
- [ ] Review the Stage-A decision report before selecting any new author model.

Draft PR #139 contains the completed common-Q fits, frozen parity evidence, and reports; those acceptance boxes remain open until the PR is merged.

## Stage-A decision result

At D=614, GRU wins four unadjusted subject-paired comparisons, common Q wins five, and four are unresolved. All 13 cohorts improve in trial-pooled GRU likelihood from D=10 to D=614, but the largest D is not always the curve maximum. The embedding report shows every external cohort farther from the source distribution than held-out AIND mice in every seed.

No CPU-only work was submitted to Beaker. No new split abstraction or author-model family was introduced.

貢獻指南

這個儲存庫沒有索引到貢獻指南

研究方向

Start by reviewing draft PR #139, its frozen parity evidence, completed common-Q fits, and Results 2 and 3, then read the Stage-A decision report. Confirm the reported comparisons and embedding findings, complete the open review item, and decide whether a new author model should be selected.

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
machine-learning
Issue 類型
功能
難度
5/5
預估耗時
一週以上
活躍度
停滯
描述清晰度
需要釐清
新手友好度
20/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。