ISISComputingGroup / ISISComputingGroup/DataStreaming
`filewriter`: file archiving
- 主要語言
- 沒有語言資料
- 星號
- 0
- 分支
- 0
- PR 合併指標
- 30 天內沒有已合併 PR
描述
The file-writing infrastructure must be able to trigger file-archiving workflows to get the file from wherever it is generated, to places where it can be read by downstream consumers and stored according to the ISIS data policy.
This includes:
- Generating file checksums
* See [comment from @ChrisM-S on ADR to drop .raw files](https://github.com/ISISComputingGroup/ibex_developers_manual/pull/236#issuecomment-5192402543) for further details about how file checksumming is currently done
- Marking files with read-only attribute [for existing windows ISIS archive]
- Moving files from the location where they were generated to the archive
* In the immediate term this is assumed to be the existing ISIS archive, however the chosen approach should be flexible enough to cope with longer-term architectural direction in this area (for example see EPAC 'data pipeline' slides).
* Including sufficiently robust processes to "catch up" and recover gracefully later if the archive is offline or unreachable at the time a file is generated, or the copy onto the archive fails.
- Any other steps which are required to go from some `.nxs` bytes on disk, to a file that a user can download from ISIS data gateway, that mantid can read from its expected locations, and gets archived to SCD long-term tape store.
This functionality does not necessarily have to be in the filewriter itself; it may be better to live *outside* the filewriter process, as long as the filewriter provides adequate 'hooks' for triggering archiving (e.g. `wrdn`)
## Questions
- Do files, or a copy, go _via_ the instrument (NDX) data area at all? Which downstream processes depend on getting data from there - IDAAAAAS, Mantid, AutoReduction, ...?
## Potential differences from existing files
- File checksums are currently written into a Windows alternate file stream in the `.raw` file. This almost certainly is not the best approach for a linux-based filewriter. Do any downstream consumers other than the archival process currently use these checksums embedded in alternate file streams?
貢獻指南
這個儲存庫沒有索引到貢獻指南
研究方向
Start by mapping the filewriter and its possible archiving hooks, including `wrdn`, against the existing ISIS archive and the EPAC data-pipeline direction. Investigate whether data passes through the instrument NDX area and which consumers depend on it, including IDAAAAAS, Mantid, and AutoReduction. Done means an agreed, recoverable path from generated `.nxs` bytes through checksum and archive handling to the data gateway and long-term tape store.
由索引模型根據 Issue 內容生成。
評估
- 領域
- data, infrastructure
- Issue 類型
- 功能
- 難度
- 5/5
- 預估耗時
- 一週以上
- 活躍度
- 活躍
- 描述清晰度
- 需要釐清
- 新手友好度
- 25/100