Authentication handling in wfdb.io.dl_files, error 403
Nessuno ha ancora preso questa issue.
- Lingua principale
- Jupyter Notebook
- Stelle
- 853
- Fork
- 322
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
Context
I am trying to download specific files (subsets) from credentialed databases on PhysioNet using wfdb helper functions.
What I am trying to do is downloading selected files (not the entire dataset), e.g. :
from wfdb.io import dl_files
dl_files('mimiciv','data\\tmp',['hosp/patients.csv.gz'])
Stack trace
Created local base download directory: data\tmp
Downloading files...
Traceback (most recent call last):
File "c:\Users\pa10198\Documents\_code\zarrow\dev.py", line 12, in <module>
dl_files('mimiciv','data\\tmp',['hosp/patients.csv.gz'])
File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\download.py", line 544, in dl_files
pool.map(dl_pn_file, dl_inputs)
File "C:\Users\pa10198\AppData\Local\Programs\Python\Python312\Lib\multiprocessing\pool.py", line 367, in map
return self._map_async(func, iterable, mapstar, chunksize).get()
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\Users\pa10198\AppData\Local\Programs\Python\Python312\Lib\multiprocessing\pool.py", line 774, in get
raise self._value
File "C:\Users\pa10198\AppData\Local\Programs\Python\Python312\Lib\multiprocessing\pool.py", line 125, in worker
result = (True, func(*args, **kwds))
^^^^^^^^^^^^^^^^^^^
File "C:\Users\pa10198\AppData\Local\Programs\Python\Python312\Lib\multiprocessing\pool.py", line 48, in mapstar
return list(map(*args))
^^^^^^^^^^^^^^^^
File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\download.py", line 450, in dl_pn_file
dl_full_file(url, local_file)
File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\download.py", line 472, in dl_full_file
content = readfile.read()
^^^^^^^^^^^^^^^
File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\_url.py", line 580, in read
result = b"".join(self._read_range(start, end))
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\_url.py", line 473, in _read_range
with RangeTransfer(self._current_url, req_start, req_end) as xfer:
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\_url.py", line 167, in __init__
self._parse_headers(method, self._response)
File "C:\Users\pa10198\Documents\_code\zarrow\.venv\Lib\site-packages\wfdb\io\_url.py", line 213, in _parse_headers
raise cls(
wfdb.io._url.NetFilePermissionError: 403 Error: Forbidden for url: https://physionet.org/files/mimiciv/3.1/hosp/patients.csv.gz
Remarks
- I have the credentialed physionet account and access to the databases I request.
- The same URL works in a browser after login and the file downloads successfully.
Questions
- Is there any supported way to pass authentification when using dl_files or related functions, e.g. through env vars, config, etc.
- If not, are there recommended workarounds (e.g., using cookies, tokens, or another library)? Ideally not leaving python and avoiding the multiplication of libraries within projects.
Feature idea
It would be very useful if wfdb supported:
- authentication for credentialed MIMIC-IV (and similar datasets), and
- partial downloads (which already works on the open datasets and seem supported structurally, but fail due to auth)
This would allow to avoid downloading entire large datasets and keep workflows within a single library.
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Direzione di ricerca
Inizia in wfdb/io/download.py con dl_files, dl_pn_file e dl_full_file, quindi esamina RangeTransfer e _parse_headers in wfdb/io/_url.py. Riproduci la richiesta 403 verso l’URL autenticata di PhysioNet e traccia la gestione dell’accesso; il lavoro è completato quando esiste un percorso di autenticazione documentato e supportato che consenta a dl_files di recuperare file selezionati che richiedono credenziali.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python
- Ambito
- authentication, data
- Tipo di issue
- Funzionalità
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Tranquilla
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 45/100