PyCQA / PyCQA/docformatter

Multiline summary is not split correctly

Open
#315 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

C: convention P: bug U: high
Dominant language
Python
Stars
598
Forks
93
PR merge metrics
No merged PRs in 30d

Description

I saw the error in CI because I used master (that was needed for a while to use with pre-commit version 4 and above).

Some multi lines sentences are split incorrectly. This is the behavior of the function split_summary

>>> split_summary(["First sentence is here. Second sentence is long", "and split in two. I even have a third sentence here.", "", "And other text here."])
['First sentence is here.',
 'and split in two. I even have a third sentence here.',
 'Second sentence is long',
 '',
 'And other text here.']

I think this comes from the split_summary function that does this:

    lines[0] = first_sentence
    if rest_text:
        lines.insert(2, rest_text)

and inserts at the wrong place if we have sentences that are too long.

I am not familiar with the code base, but maybe something along those lines could work?

def split_summary(lines) -> List[str]:
    """Split multi-sentence summary into the first sentence and the rest."""
    if not lines or not lines[0].strip():
        return lines

    text = lines[0].strip()

    tokens = re.split(r"(\s+)", text)  # Keep whitespace for accurate rejoining
    sentence = []
    rest = []
    i = 0

    while i < len(tokens):
        token = tokens[i]
        sentence.append(token)

        if token.endswith(".") and not any(
            "".join(sentence).strip().endswith(abbr) for abbr in ABBREVIATIONS
        ):
            i += 1
            break

        i += 1

    rest = tokens[i:]
    first_sentence = "".join(sentence).strip()
    rest_text = "".join(rest).strip()

    new_lines = [first_sentence, ""]
    if rest_text:
        new_lines.append(rest_text)
        new_lines.extend(line for line in  lines[1:] if line)

    return new_lines

This gives:

>>> split_summary(["First sentence is here. Second sentence is long", "and split in two. I even have a third sentence here.", "", "And other text here."])
['First sentence is here.',
 '',
 'Second sentence is long',
 'and split in two. I even have a third sentence here.',
 'And other text here.']

I do not know if the result should be processed more before returning or if it is something that is taken into account elsewhere in the codebase.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the split_summary function and reproduce the multiline example from the issue. Inspect how its returned lines are consumed, then add regression coverage for the shown ordering and confirm the corrected output preserves subsequent lines.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
tooling
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
55/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.