acl-org / acl-org/acl-anthology

Appetite for Subject and References in MODs XML

未关闭
#7,923 2 条评论 0 个 reaction 已指派 1 人 已被 @mbollmann 认领 在 GitHub 查看
enhancement
主要语言
Python
星标
797
派生
408
平均合并
3 天 19 小时
30 天内合并 PR
36

描述

Greetings,

I am unfamiliar with the MODs creation pipeline in use at ACL. But as a librarian I know that it can hold references and subjects. As far as I know there is no overt way to look at lang-o-metrics across the ACL corpus. That is, which papers are _about_ which languages, or use evidence from certain languages. This information could be captured in and communicated through the MODs XML file. However, to make a meaningful systemic change at ACL it needs to be captured at the paper submission time from the authors. This implies a larger workflow change. I'm wondering if there is an appetite for this or not. If there is I could follow up with a broader plan where we might need to make some adjustments and provide some code around the MODs structure.

In a similar vein, I am wondering if there is appetite for including outbound references in the MODs XML. Again, MODs is designed to support this. But it would imply some ingest workflows, and data stewardship architecture. I know that it has been of interest to some in the past, however, I believe the coverage has been less than universal, and the work hasn't initiated any supporting architecture. If there are open issues on this please link them in the reply. I didn't see any specifically addressing MODs, outbound references, or subject languages.

贡献指南

这个仓库没有索引到贡献指南

调研方向

The issue names no files or tests. Start by reviewing the existing MODS XML creation, paper-submission, and ingest workflows, then check for related issues about subjects, subject languages, or outbound references; done would first require an agreed scope and implementation plan for capturing and exposing this metadata.

由索引模型根据 Issue 内容生成。

评估

技术栈
xml
领域
data
Issue 类型
功能
难度
5/5
预计耗时
一周以上
活跃度
冷清
描述清晰度
需要澄清
新手友好度
25/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。