catalyst-cooperative / catalyst-cooperative/mozilla-sec-eia
Running list of Ex. 21 extraction model improvement ideas
- Dominant language
- Jupyter Notebook
- Stars
- 2
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
### Overview
This is a running list of improvements that we could try implementing to improve the performance of the Ex. 21 extraction model. I moved any "nice to have" straggler items from #78 into this issue. These items can be experimented with after record linkage, when we have a better idea of remaining budget and performance needs.
* Example of a filing with a "footnotes" section that can be excluded:
* 103872-0001193125-13-444053
```[tasklist]
### Next steps
- [ ] Use Corpwatch dataset for further validation
- [ ] Nice to have: breakout `layoutlm-finetune` into ops
- [ ] Try clustering the final hidden states instead of using heuristic based table extractor
- [ ] Exclude anything below "Footnotes" or similar keywords
- [ ] Create threshold for entity classification failure based on logits returned by LayoutLM
```
Contributor guide
Research direction
No file, test, or entry point is named. Start by reviewing the Ex. 21 extraction model and the existing items moved from #78, then evaluate the Corpwatch dataset, LayoutLM outputs, and table-extraction approach against the remaining budget and performance needs. Done requires choosing a scoped experiment and validating its effect.
Written by the indexing model from the issue text.
Assessment
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100