dwh-scripts/extract_logged_cda_to_csv_by_zip.sh does not work sufficiently
- Lenguaje dominante
- Shell
- Estrellas
- 0
- Forks
- 0
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Descripción
Bug: extract_logged_cda_to_csv_by_zip.sh does not populate several CSV columns
Summary
Running the Bash script extract_logged_cda_to_csv_by_zip.sh against the 14-day Wolfsburg sample placed on the fileshare produces a CSV whose header lists all required AKTIN / DIVI fields, yet several columns are always NA.
Screenshots provided by Dr. Jennifer Kreklow show the empty fields highlighted in yellow (see cedis_mts_transfer.png & docid_case_diag.png).
This blocks QA and prevents starting the full export.
Affected columns
CSV column | Expected CDA element / CodeSystem | Current value | Notes
-- | -- | -- | --
CEDIS Code | | NA | all rows
MTS Score | | NA | all rows
Transfer | | NA | all rows
Discharge Code | | NA | all rows
Document ID | | NA | all rows
Internal Case Nr. | | NA | all rows
Diagnosis | | NA | all rows
Steps to reproduce
-
Request two real-world test CDAs from IT (see Contact below) that contain non-empty values for the affected fields.
-
Place the CDAs in a directory, e.g.
/share/cda_test/. -
Run the script:
./extract_logged_cda_to_csv_by_zip.sh /path/to/aktin.properties /share/cda_test/ -
Open the generated
cda_data_<timestamp>.csvand verify that the columns listed above are stillNA.
Environment
-
Script version 1.1 (10 Oct 2024)
-
Bash 5.2, Debian 12 container (same behaviour on RHEL 9)
-
aktin.propertiesfrom Wolfsburg deployment (encounter & billing roots verified)
Contact
-
Tomasz Babiuch – (please provide two representative CDA files covering the missing fields)
-
CC: Robert Dietrichs, Dr. Jennifer Kreklow
Analysis / Suspected root causes
-
Incorrect root IDs – The
encounter_rootandbilling_rootvalues inaktin.propertiesmay not match the IDs actually used in the test CDAs. -
Regex too strict – Current grep patterns require an explicit
xsi:typeprefix and may not allow namespace variations. -
Multiple occurrences – Using
head -1might pick the wrong section (e.g. administrative vs. clinical) where the field is genuinely empty.
Proposed fix
-
Generalise regex patterns to accept optional namespace prefixes and whitespace.
-
Validate root IDs by echoing them in the log; abort with a clear error if not found.
-
Add a helper script / unit test that parses a mini-CDA snippet for each field.
-
Emit a warning whenever a header field resolves to
NAbut the tag exists elsewhere in the file (helps QA).
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Evaluación
Este issue todavía no se ha evaluado.