Kludex / Kludex/python-multipart

Skip header area?

Open
#59 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
532
Forks
93
PR merge metrics
No merged PRs in 30d

Description

RFC 1341 says about multipart payloads:

> This [the multipart content-type] indicates that the entity consists of several parts, each itself with a structure that is syntactically identical to an RFC 822 message, except that the header area might be completely empty...

I read that as: "a multipart payload *may* have a header area". As a matter of fact, when you construct a multipart payload using email.mime.multipart.MIMEMultipart and then calling as_string(), you get:

```
MIME-Version: 1.0
Content-Type: multipart/form-data; boundary="========== bounda r y 930"

--========== bounda r y 930\nMIME-Version: 1.0
Content-Type: application/octet-stream
...
```

The multipart parser from the old cgi module would correctly grok that when I uploaded it (provided I arranged for the content-type to be in the HTTP header). python-multipart (tried 0.0.5) does not and fails when it sees the M of the MIME-Version.

I appreciate the "as browsers send it" part of multipart's rationale; however, being somewhat friendly to machine-generated multipart messages would, I think, be a friendly gesture, and it might prevent breakage as people move from cgi's FieldStorage (used, e.g., in twisted) to python-multipart.

What I think should be done: in MultipartParser's _internal_write, we should blindly skip everything until the MIME header area is consumed, i.e., until we have found a CRLFCRLF sequence. Would you consider such a PR?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start at MultipartParser._internal_write and reproduce the reported input generated by email.mime.multipart.MIMEMultipart and as_string(). Trace how the parser handles the initial MIME-Version header and the CRLFCRLF boundary, then verify that multipart payloads with this header area are accepted without breaking browser-style uploads.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.