3.0.6 (pypi) not fully extracting?

Open
#346 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
35/100
Issue type
Bug
Clarity
Needs clarification
Activity status
Stale
Tech stack
python

Research direction

The report concerns PyPI 3.0.6 processing XML dumps for enwiki and ruwiki. Reproduce the extraction with those inputs and compare the decompressed dump sizes with the reported 18 GB and 8 GB outputs; done when the expected size relationship is established or a specific extraction defect is identified.

Written by the indexing model from the issue text.

Description

I'm using the xml dumps for the latest enwiki and ruwiki and i'm not sure it's fully extracting the articles. my enwiki text goes to GY and my ruwiki text goes to AI, they decompress to 100gb and 30gb respectively and my wikiextractor text output is only 18gb and 8gb respectively. Is this correct?

Dominant language
Python
Stars
4k
Forks
1k
Avg merge
1m
Merged PRs (30d)
10

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from WikiExtractor/wikiextractor

All issues in WikiExtractor/wikiextractor

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.