catalyst-cooperative / catalyst-cooperative/mozilla-sec-eia

Running list of Ex. 21 extraction model improvement ideas

Open
#88 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
2
Forks
0
PR merge metrics
No merged PRs in 30d

Description

### Overview

This is a running list of improvements that we could try implementing to improve the performance of the Ex. 21 extraction model. I moved any "nice to have" straggler items from #78 into this issue. These items can be experimented with after record linkage, when we have a better idea of remaining budget and performance needs.

* Example of a filing with a "footnotes" section that can be excluded:
* 103872-0001193125-13-444053

```[tasklist]
### Next steps
- [ ] Use Corpwatch dataset for further validation
- [ ] Nice to have: breakout `layoutlm-finetune` into ops
- [ ] Try clustering the final hidden states instead of using heuristic based table extractor
- [ ] Exclude anything below "Footnotes" or similar keywords
- [ ] Create threshold for entity classification failure based on logits returned by LayoutLM
```

Contributor guide

Open the contributing guide

Research direction

No file, test, or entry point is named. Start by reviewing the Ex. 21 extraction model and the existing items moved from #78, then evaluate the Corpwatch dataset, LayoutLM outputs, and table-extraction approach against the remaining budget and performance needs. Done requires choosing a scoped experiment and validating its effect.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.