aboutcode-org / aboutcode-org/scancode-toolkit

Integration of sanexml fails.

Offen
#3,483 6 Kommentare 0 Reaktionen 1 zugewiesene Person Beansprucht von @35C4n0r Auf GitHub ansehen
GSoC
Vorherrschende Sprache
Python
Sterne
2.6k
Forks
791
Ø Merge
1 T. 12 Std.
Gemergte PRs (30 T.)
5

Beschreibung

The integration of [sanexml](https://github.com/nexB/sanexml) with Pymaven and SCTK are failing.
Here is the [log](https://gist.github.com/35C4n0r/c44ec67e80e7bb276847bdb9809d1172).
I think a lot of these are failing because the xmls are not in a proper format and i didn't implement any logic corresponding to the `recover=True` for the XMLParser. What this does that it Parses XML liniently.

I see two possibilitoes to solve this issue:

- Use `BeautifulSoup4` and use it's `html.parser` for linient parsing to get a correct xml string and then parse that string normally using python's default XML parsers. (BS4 [Docs](https://www.crummy.com/software/BeautifulSoup/bs4/doc/))
- Create a custom parser :(

If we go with the first approch we can either use this process ``html.parser -> get correct xml -> DefaultXMLParser`` as the default or we can do:

try: `DefaultXMLParser`
except:
try: `html.parser` -> get correct xml -> `DefaultXMLParser` # Slower #Lininent
except:
try: `html5lib` -> get correct xml -> `DefaultXMLParser` # Slowest #MoreLinient

@JonoYang @pombredanne @AyanSinhaMahapatra please share your opinions on this.

Beitragsleitfaden

Beitragsleitfaden öffnen

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.