dfe-analytical-services / dfe-analytical-services/dfe-published-data-qa

Add tracking for files screened

Open
#135 0 comments 0 reactions 0 assignees View on GitHub
new feature
Dominant language
R
Stars
4
Forks
1
PR merge metrics
No merged PRs in 30d

Description

We should set up a database table that we populate with the following columns:

- environment - e.g. 'local', `development`, 'production', `shinyapps.io` (can set environment variables to control)
- data_filename - pulled from the csv data file
- file_size - we should just track the raw bytes so it's easier to analyse, rather than the `dfeR::pretty_filesize()` output
- rows_ count - same thing we put in the UI
- cols_count - same thing we put in the UI
- stage - what stage of the checks the file got to
- pass - boolean, did it pass
- warnings - string of the checks that had a warning (take the name from the test col of the screening output table, e.g. "ethnicity_values, total, ob_unit_meta" (we can then parse this if we want to splitting by commas if we want to analyse what warnings are happening a lot)
- time_started - time recorded at start of screening
- time_ended - time recorded at end of screening
- screening_time - count of time taken to screen in raw seconds calculated from above two cols, I think we'd want a raw value like seconds that we can ees-ily average / aggregate

We then add code into the server side of the app that will write a new row into the database table after each screening.

Example UI output now
![image](https://github.com/user-attachments/assets/e55650cf-f15f-4506-b5cd-9d28975b2dfc)

Example screening output table (where we can use the test column to pull from and populate with what warnings are present (as will be interesting thinking of API standards), also shows the 'stages' we have
![image](https://github.com/user-attachments/assets/5acc595f-e299-4180-924d-7d05675e53e2)

## Probable tasks
### Easy-ish first tasks
- [x] Add time tracking into the screener as is (log time at start of screening, log time at end, present in UI with `dfeR::pretty_time_taken()`
- [ ] Check environment variables exist
- [ ] Check all data above exists in app

### Main tasks
- [ ] Create new database table
- [ ] Add a connection into the app to point to our existing SQL databases (can migrate to databricks at a later point)
- [ ] Add code into server file that gathers all of this and writes a new row into the database

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.