Remove Pixels name, repo and path columns
Nobody has claimed this yet.
- Dominant language
- Java
- Stars
- 222
- Forks
- 105
- Avg merge
- 2h 21m
- Merged PRs (30d)
- 1
Description
Background
The OMERO 4.2.0 release in July 2010 included an alteration to the database schema to add name, path and repo columns to the Pixels table with a similar meaning as the columns in the OriginalFile table. This change was part of the initial work adding native file format support in OMERO via Bio-Formats also known as FS lite. For a subset of file formats (primarily single file and with large XY dimensions), the original file was uploaded to the binary repository and linked from the Pixels object. This allowed the server to perform certain operations including the generation of OMERO pyramids.
Full support for native file format support in OMERO, also known as OMERO.fs, was introduced in OMERO 5.0.0 in February 2014 with the introduction of the Fileset table linked to the Image. Each Fileset row is linked to an ordered set of FilesetEntry rows each of these being themselves associated with a single OriginalFile entry. This change effectively superseded the FS Lite concept allowing native support for single and multi-file formats as well as multi-image formats. In OMERO 5.1, the series column was also introduced to the Image table to store the mapping between an image and the underlying Bio-Formats series.
Current API
Despite OMERO 5 actually deprecating their usage, the Pixels.name, Pixels.path and Pixels.repo columns are still currently heavily used server-side as of OMERO 5.6.x:
- the pixeldata phase of the importer calls OMEROMetadataStoreClient.setPixelsFile which populates these fields via a PostgresSQLAction
- the PixelsService retrieves the filepath associated with an image using the FilePathResolver.getOriginalFilePath API. The OMERO implementation of this interface OmeroFilePathResolver.getOriginalFilePath also uses an SQL action to retrieve the name, path and repo attributes of the Pixels
Challenges
The current logic is problematic for several reasons:
- for each FS image, with an associated Fileset, identical metadata is maintained in several locations: the
Pixelsobject as well as theOriginalFilelinked to theFilesetEntry - the requirements to update the
Pixelsmetadata in the importer can cause substantial DB operations especially in the case of high-content screening - while the API and the UIs have been updated to expose the
Fileset/FilesetEntry, thePixelsname,repoandpathattributes are still hidden from an API perspective. - a client like Glencoe's image-region micro-service still requires direct access to the database and the
omero.dbconfigurations properly set in its configuration in order to execute the PostgresSQLAction - the IDR team explored the replacements of Filesets e.g. following a conversion and ran into the requirements to update the
Pixelsobject via SQL script - see https://github.com/IDR/idr-metadata/issues/656
As an additional related complication, a historical bug has been reported in the image.sc forum where the OriginalFile are incorrectly linked to FilesetEntry for some multi-file filesets.
Proposal
- Fix the ordering of the FilesetEntry so that it can be used as the single source of truth
a. Fix the mapping between OriginalFile and FilesetEntry at import time - see https://github.com/ome/omero-blitz/pull/148
b. Create an upgrade script allowing to fix all existing FilesetEntry/OriginalFile links in existing OMERO databases
c. Review the API and technical documentation of Bio-Formats and OMERO and if needed clarify and enforce that the first file inIFormatReader.getUsedFiles, the output ofImportCandidatesand the firstFilesetEntryis the file that should be passed toIFormatReader.setId
d. Optionally, create an upgrade script allowing to convert FS lite imports intoFileset - Remove all legacy FS lite API and use the Fileset API consistently
a. Create a new version of the OMERO database schema dropping thePixels.name,Pixels.pathandPixels.repocolumns
b. Update all the server APIs to use the OriginalFile from the first FilesetEntry as the source of truth
/cc @joshmoore @jburel @kkoz @chris-allan @will-moore @dominikl @Tom-TBT
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by tracing the legacy Pixels fields through ManagedImportRequestI, PostgresSqlAction, PixelsService, and OmeroFilePathResolver. Review the FilesetEntry/OriginalFile ordering and the linked import fix before examining the database upgrade requirements. Done means the Fileset API is the consistent source of truth and the legacy Pixels columns and APIs can be removed.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java, postgresql
- Domain
- backend-api-design, databases
- Issue type
- Refactor
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100