Models Inconsistency

Open
#566 4 comments 0 reactions 0 assignees View on GitHub

@shaneahmed is already working on this.

Since Sep 20, 2024.

  • #578 by @shaneahmed — merged
  • #867 by @shaneahmed — open

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
25/100
Issue type
Refactor
Clarity
Mostly clear
Activity status
Stale
Tech stack
python, pytorch

Research direction

Start with ModelABC and compare the model implementations mentioned in vanilla.py, unet.py, hovernet.py, and micronet.py, focusing on forward, infer_batch, _transform, preproc, and _preproc. Review how the existing methods are documented before reorganizing the custom models. Done means the pipeline is consistently structured, ModelABC documents normalization, activation, inference, training, and postprocessing, and a documentation page explains evaluation and training.

Written by the indexing model from the issue text.

Description

code readability documentation enhancement refactoring
  • TIA Toolbox version: 1.3.3
  • Python version: 3.10.8
Description

Tiatoolbox has several pre-trained models helpful for data processing. However, models differ in how they handle input and output, making them confusing to use (especially when customizing):

  • Activation functions apply at different steps for different models. Sometimes in forward method (e.g. CNNModel), sometimes forward returns a raw layer output and the transformation applies in infer_batch (e.g. UNetModel).
  • Moreover, activation functions are hardcoded. To customize, you can't simply change an attribute; you must overwrite the whole method (different for each model).
  • Data normalization distributes across methods: HoVerNet uses it in forward, UNetModel in _transform, MicroNet in preproc, and vanilla models rely on the user to do so.
  • Data preprocessing also lacks consistency. Even though it should happen in preproc_func/_preproc functions, UNetModel uses its own _transform, unrelated to the standard methods. Yet, its behavior could implement in _preproc.
What to do

Refactoring the code will significantly improve readability:

  • Decompose the pipeline into small granular methods in ModelABC: one method for normalization, activation function as an attribute, etc.
  • Explain ModelABC methods in their documentation: does infer_batch rely on postproc_func? Can infer_batch be used for training? How?
  • Reorganize custom model methods to match the new ModelABC structure.
  • Add a new page to the documentation explaining the Tiatoolbox models pipeline: how is it related to the PyTorch pipeline? How to evaluate a model? How to train a model?
Dominant language
Python
Stars
551
Forks
112
Avg merge
1d 6h
Merged PRs (30d)
13

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from TissueImageAnalytics/tiatoolbox

All issues in TissueImageAnalytics/tiatoolbox

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.