acl-org / acl-org/acl-anthology
Appetite for Subject and References in MODs XML
- 主要语言
- Python
- 星标
- 797
- 派生
- 408
- 平均合并
- 3 天 19 小时
- 30 天内合并 PR
- 36
描述
Greetings,
I am unfamiliar with the MODs creation pipeline in use at ACL. But as a librarian I know that it can hold references and subjects. As far as I know there is no overt way to look at lang-o-metrics across the ACL corpus. That is, which papers are _about_ which languages, or use evidence from certain languages. This information could be captured in and communicated through the MODs XML file. However, to make a meaningful systemic change at ACL it needs to be captured at the paper submission time from the authors. This implies a larger workflow change. I'm wondering if there is an appetite for this or not. If there is I could follow up with a broader plan where we might need to make some adjustments and provide some code around the MODs structure.
In a similar vein, I am wondering if there is appetite for including outbound references in the MODs XML. Again, MODs is designed to support this. But it would imply some ingest workflows, and data stewardship architecture. I know that it has been of interest to some in the past, however, I believe the coverage has been less than universal, and the work hasn't initiated any supporting architecture. If there are open issues on this please link them in the reply. I didn't see any specifically addressing MODs, outbound references, or subject languages.
贡献指南
这个仓库没有索引到贡献指南
调研方向
The issue names no files or tests. Start by reviewing the existing MODS XML creation, paper-submission, and ingest workflows, then check for related issues about subjects, subject languages, or outbound references; done would first require an agreed scope and implementation plan for capturing and exposing this metadata.
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- xml
- 领域
- data
- Issue 类型
- 功能
- 难度
- 5/5
- 预计耗时
- 一周以上
- 活跃度
- 冷清
- 描述清晰度
- 需要澄清
- 新手友好度
- 25/100