maniac-tech / maniac-tech/Web-Crawling-using-Python
get_text error in webCrawling.py
Nobody has claimed this yet.
- Dominant language
- HTML
- Stars
- 0
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Description
Traceback (most recent call last):
File "webCrawling.py", line 42, in
blog_posts = get_blog_posts(fp)
File "webCrawling.py", line 35, in get_blog_posts
'content': cleanHtml(content.value),
File "webCrawling.py", line 11, in cleanHtml
return BeautifulStoneSoup(get_text(html),
NameError: global name 'get_text' is not defined
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start at webCrawling.py lines 11 and 35, then inspect how HTML text extraction is intended to work and whether get_text is defined or imported. Re-run the crawler with the same input and verify it completes without the NameError and produces the content field.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- web-dev
- Issue type
- Bug
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 58/100