aboutcode-org / aboutcode-org/source-inspector

xgettext: multiple starting lines for a string are not well supported

Open
#13 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
C
Stars
3
Forks
2
PR merge metrics
No merged PRs in 30d

Description

In the current xgettext implementation I can see at line 115 https://github.com/nexB/source-inspector/blob/9511f56b44ac7c5644b34d413146d58dd9fa7ea0/src/source_inpector/strings_xgettext.py#L115 the following:

```
_, _, start_line = line.rpartition(":")
```

This is likely leading to the wrong results, as a line can have multiple instances of `start_line`, which you aren't catching. As an example, I used `xgettext` with the same parameters as you did on `libbb/lineedit.c` from BusyBox:

```
$ xgettext --omit-header --extract-all --no-wrap lineedit.c
```

Some of the result lines:

```
#: lineedit.c:834 lineedit.c:890 lineedit.c:893
msgid "."
msgstr ""
```

As you can see there are multiple file/line number entries there. It seems that at some point the authors of `xgettext` decided to combine these. Your code does not correctly process these lines:

```
>>> line = '#: lineedit.c:834 lineedit.c:890 lineedit.c:893'
>>> _, _, start_line = line.rpartition(":")
>>> start_line
'893'
```

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.