Inconsistency in Content-Disposition format
Nobody has claimed this yet.
- Dominant language
- Java
- Stars
- 1.1k
- Forks
- 564
- Avg merge
- 2d 2h
- Merged PRs (30d)
- 29
Description
Dataverse returns Content-Disposition header for a file in two different formats depending on where the file is physically stored. E.g.:
import requests
url = "https://dataverse.harvard.edu/api/access/datafile/3315157"
req = requests.get(url, allow_redirects=True)
print(f"File (id:3315157) - tab version not on S3: {req.headers['Content-Disposition']}")
req = requests.get(url + "?format=original", allow_redirects=True)
print(f"File (id:3315157) - original version on S3: {req.headers['Content-Disposition']}")
yields:
File (id:3315157) - tab version not on S3: attachment; filename="Karnataka_DD%26FS_Data-1.tab"
File (id:3315157) - original version on S3: attachment; filename*=UTF-8''Karnataka_DD%26FS_Data-1.xlsx
Note that in both cases filename has escaped and encoded characters per RFC 5987. However, only the latter format of the header is the correct ("filename*" indicates that name is encoded, see e.g. Content-Disposition def)
Related PRs: #4542 and #7503
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Reproduce the two responses with the Python requests example and compare the Content-Disposition headers for the default and original formats. Trace the API response handling for files stored in different locations, then verify that both paths use the RFC 5987 filename* form consistently.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- api, backend
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100