aboutcode-org / aboutcode-org/scancode-toolkit
Using BeautifulSoup4 and html.parser to parse XML FIles.
- Linguagem predominante
- Python
- Estrelas
- 2.6k
- Forks
- 791
- Merge médio
- 1d 12h
- PRs com merge (30d)
- 5
Descrição
In order to parse XML documents we will be using BeautifulSoup4 and `html.parser`.
Now we are using this option instead of the Python's built in XML Parser is because at times the XML that has to be parsed is malformed and the `html.parser` is linient in parsing, whereas the standard library only handle well-formed XML.
There is an issue with this approach:
- Parsing the document as HTML, all the tags were converted to lower case by default, and there is no option to change that.
In order to deal with this:
- We create a mapping of all the tags to another another tags and replace those tags to the mappings.
- We then parse the document normally as HTML.
- We then use the mapping to replace the mapped tags to the original tags.
Example:
```
org.jboss.seam
root
https://github.com/
https://github.com/35C4n0r
```
The Mapping:
```
{'groupId': 'TAG0', 'parent': 'TAG1', 'url': 'TAG2', 'artifactId': 'TAG3', 'Url': 'TAG4'}
```
The new XML
```
org.jboss.seam
root
https://github.com/
https://github.com/35C4n0r
```
After parsing it with BeautifullSoup
```
org.jboss.seam
root
https://github.com/
https://github.com/35C4n0r
```
After using the map to convert the tags back
```
org.jboss.seam
root
https://github.com/
https://github.com/35C4n0r
```
Guia de contribuição
Avaliação
Esta issue ainda não foi avaliada.