aboutcode-org / aboutcode-org/scancode-toolkit

Using BeautifulSoup4 and html.parser to parse XML FIles.

未关闭
#3,486 4 条评论 0 个 reaction 已指派 1 人 已被 @35C4n0r 认领 在 GitHub 查看
GSoC
主要语言
Python
星标
2.6k
派生
791
平均合并
1 天 12 小时
30 天内合并 PR
5

描述

In order to parse XML documents we will be using BeautifulSoup4 and `html.parser`.
Now we are using this option instead of the Python's built in XML Parser is because at times the XML that has to be parsed is malformed and the `html.parser` is linient in parsing, whereas the standard library only handle well-formed XML.

There is an issue with this approach:
- Parsing the document as HTML, all the tags were converted to lower case by default, and there is no option to change that.

In order to deal with this:
- We create a mapping of all the tags to another another tags and replace those tags to the mappings.
- We then parse the document normally as HTML.
- We then use the mapping to replace the mapped tags to the original tags.

Example:
```

org.jboss.seam
root
https://github.com/
https://github.com/35C4n0r

```

The Mapping:
```
{'groupId': 'TAG0', 'parent': 'TAG1', 'url': 'TAG2', 'artifactId': 'TAG3', 'Url': 'TAG4'}
```

The new XML
```

org.jboss.seam
root
https://github.com/
https://github.com/35C4n0r

```

After parsing it with BeautifullSoup
```

org.jboss.seam
root
https://github.com/
https://github.com/35C4n0r

```

After using the map to convert the tags back
```

org.jboss.seam
root
https://github.com/
https://github.com/35C4n0r

```

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。