MIT-LCP / MIT-LCP/wfdb-python

Authentication handling in wfdb.io.dl_files, error 403

Offen
#568 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Vorherrschende Sprache
Jupyter Notebook
Sterne
853
Forks
322
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

Context

I am trying to download specific files (subsets) from credentialed databases on PhysioNet using wfdb helper functions.

What I am trying to do is downloading selected files (not the entire dataset), e.g. :

from wfdb.io import dl_files
dl_files('mimiciv','data\\tmp',['hosp/patients.csv.gz'])

Stack trace

Created local base download directory: data\tmp
Downloading files...
Traceback (most recent call last):
  File "c:\Users\pa10198\Documents\_code\zarrow\dev.py", line 12, in <module>
    dl_files('mimiciv','data\\tmp',['hosp/patients.csv.gz'])
  File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\download.py", line 544, in dl_files
    pool.map(dl_pn_file, dl_inputs)
  File "C:\Users\pa10198\AppData\Local\Programs\Python\Python312\Lib\multiprocessing\pool.py", line 367, in map
    return self._map_async(func, iterable, mapstar, chunksize).get()
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\pa10198\AppData\Local\Programs\Python\Python312\Lib\multiprocessing\pool.py", line 774, in get
    raise self._value
  File "C:\Users\pa10198\AppData\Local\Programs\Python\Python312\Lib\multiprocessing\pool.py", line 125, in worker
    result = (True, func(*args, **kwds))
                    ^^^^^^^^^^^^^^^^^^^
  File "C:\Users\pa10198\AppData\Local\Programs\Python\Python312\Lib\multiprocessing\pool.py", line 48, in mapstar
    return list(map(*args))
           ^^^^^^^^^^^^^^^^
  File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\download.py", line 450, in dl_pn_file
    dl_full_file(url, local_file)
  File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\download.py", line 472, in dl_full_file
    content = readfile.read()
              ^^^^^^^^^^^^^^^
  File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\_url.py", line 580, in read
    result = b"".join(self._read_range(start, end))
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\_url.py", line 473, in _read_range
    with RangeTransfer(self._current_url, req_start, req_end) as xfer:
         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\_url.py", line 167, in __init__
    self._parse_headers(method, self._response)
  File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\_url.py", line 213, in _parse_headers
    raise cls(
wfdb.io._url.NetFilePermissionError: 403 Error: Forbidden for url: https://physionet.org/files/mimiciv/3.1/hosp/patients.csv.gz

Remarks

  • I have the credentialed physionet account and access to the databases I request.
  • The same URL works in a browser after login and the file downloads successfully.

Questions

  • Is there any supported way to pass authentification when using dl_files or related functions, e.g. through env vars, config, etc.
  • If not, are there recommended workarounds (e.g., using cookies, tokens, or another library)? Ideally not leaving python and avoiding the multiplication of libraries within projects.

Feature idea

It would be very useful if wfdb supported:

  • authentication for credentialed MIMIC-IV (and similar datasets), and
  • partial downloads (which already works on the open datasets and seem supported structurally, but fail due to auth)

This would allow to avoid downloading entire large datasets and keep workflows within a single library.

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Beginne in wfdb/io/download.py bei dl_files, dl_pn_file und dl_full_file und untersuche anschließend in wfdb/io/_url.py RangeTransfer und _parse_headers. Reproduziere die 403-Anfrage gegen die authentifizierte PhysioNet-URL und verfolge, wie der Zugriff gehandhabt wird; als erledigt gilt die Aufgabe, wenn es einen dokumentierten, unterstützten Authentifizierungspfad gibt, über den dl_files ausgewählte zugangsgeschützte Dateien abrufen kann.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
authentication, data
Issue-Typ
Feature
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Ruhig
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
45/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.