trentm / trentm/python-markdown2

2.5.5 regression: strong emphasis adjacent to word/CJK chars fails when content starts or ends with punctuation

Open
#688 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Bug
Dominant language
Python
Stars
2.8k
Forks
459
Avg merge
2d 19h
Merged PRs (30d)
4

Description

Summary

After upgrading from markdown2 2.5.4 to 2.5.5, strong emphasis sometimes stops parsing when all of these are true:

  • the **...** span is adjacent to alphanumeric or CJK text without surrounding spaces
  • the emphasized content starts or ends with punctuation, for example Chinese quotes or parentheses

On 2.5.4 these cases render correctly. On 2.5.5 the parser leaves literal ** in the output, and in longer inputs it can also emit malformed mixed HTML.

This reproduces with the default parser configuration, no extras required.

Related to #679 because it also looks like an emphasis-regression in the recent parser changes, but this one reproduces without middle-word-em and with plain markdown2.markdown(...).

Minimal repro

import markdown2

cases = [
    'a**“b”**c',
    '**“b”**c',
    'a**(b)**c',
]

print('version:', markdown2.__version__)
for text in cases:
    print('INPUT :', text)
    print('OUTPUT:', markdown2.markdown(text).strip())
    print()
2.5.4
<p>a<strong>“b”</strong>c</p>
<p><strong>“b”</strong>c</p>
<p>a<strong>(b)</strong>c</p>
2.5.5
<p>a**“b”**c</p>
<p>**“b”**c</p>
<p>a**(b)**c</p>

Longer repro that produces malformed mixed HTML

import markdown2

text = '*   **示例**:系统会使用**(方案A)**或**(方案B)**进行处理。'
print(markdown2.markdown(text))
2.5.4
<ul>
<li><strong>示例</strong>:系统会使用<strong>(方案A)</strong>或<strong>(方案B)</strong>进行处理。</li>
</ul>
2.5.5
<ul>
<li><strong>示例</strong>:系统会使用**(方案A)<strong>或</strong>(方案B)**进行处理。</li>
</ul>

Suspected regression window

Looking at the source, Markdown._do_italics_and_bold() changed between these versions:

2.5.4
text = self._strong_re.sub(r"<strong>\\2</strong>", text)
text = self._em_re.sub(r"<em>\\2</em>", text)
2.5.5
if not self._iab_processor:
    self._iab_processor = GFMItalicAndBoldProcessor(self, None)
if self._iab_processor.test(text):
    text = self._iab_processor.run(text)

So this looks related to the switch to GFMItalicAndBoldProcessor as the default implementation for _do_italics_and_bold().

Environment

  • Python: 3.14.0
  • markdown2: 2.5.4 vs 2.5.5
  • OS: Windows 11

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with Markdown._do_italics_and_bold() and GFMItalicAndBoldProcessor, comparing the 2.5.4 regular-expression path with the 2.5.5 processor path. Run the minimal and longer reproductions against both versions, then add regression coverage showing that adjacent alphanumeric or CJK text with punctuation renders the same strong-emphasis HTML as 2.5.4.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
58/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.