jsumners / jsumners/feedparser

Don't use chardet unless you have to

Open
#421 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

auto-migrated
Dominant language
Python
Stars
0
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Automatic character set detection can be CPU intensive, especially with large 
feeds. It seems a waste to run it on every feed, considering that the result 
will be discarded most of the time.

The attached patch makes the chardet invocation lazy.

Original issue reported on code.google.com by itsa...@gmail.com on 20 Jan 2014 at 8:40

Attachments:

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Review the attached lazy_chardet.patch and locate the feed-processing path where character-set detection is invoked. Confirm when detection is needed and verify that large feeds no longer trigger unnecessary chardet work; add or update tests if the relevant test location is identified.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
performance
Issue type
Refactor
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.