aboutcode-org / aboutcode-org/scancode-toolkit

Using BeautifulSoup4 and html.parser to parse XML FIles.

Ouverte
#3,486 4 commentaires 0 réactions 1 personne assignée Réclamée par @35C4n0r Voir sur GitHub
GSoC
Langage dominant
Python
Étoiles
2.6k
Forks
791
Merge moyen
1 j 12 h
PR mergées (30 j)
5

Description

In order to parse XML documents we will be using BeautifulSoup4 and `html.parser`.
Now we are using this option instead of the Python's built in XML Parser is because at times the XML that has to be parsed is malformed and the `html.parser` is linient in parsing, whereas the standard library only handle well-formed XML.

There is an issue with this approach:
- Parsing the document as HTML, all the tags were converted to lower case by default, and there is no option to change that.

In order to deal with this:
- We create a mapping of all the tags to another another tags and replace those tags to the mappings.
- We then parse the document normally as HTML.
- We then use the mapping to replace the mapped tags to the original tags.

Example:
```

org.jboss.seam
root
https://github.com/
https://github.com/35C4n0r

```

The Mapping:
```
{'groupId': 'TAG0', 'parent': 'TAG1', 'url': 'TAG2', 'artifactId': 'TAG3', 'Url': 'TAG4'}
```

The new XML
```

org.jboss.seam
root
https://github.com/
https://github.com/35C4n0r

```

After parsing it with BeautifullSoup
```

org.jboss.seam
root
https://github.com/
https://github.com/35C4n0r

```

After using the map to convert the tags back
```

org.jboss.seam
root
https://github.com/
https://github.com/35C4n0r

```

Guide de contribution

Ouvrir le guide de contribution

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.