astronomy-commons / astronomy-commons/lsdb
Create example notebook for tokenizing catalog data
- Dominant language
- Python
- Stars
- 55
- Forks
- 26
- Avg merge
- 4d 1h
- Merged PRs (30d)
- 8
Description
We may choose to just do this for images for simplicity, or we could extend to multiple modalities if we wanted to overachieve. But largely this should just provide a template for the typical machinery needed to perform tokenization, which itself is an inference style workflow rather than a pure map_partitions style operation.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by locating existing example notebooks and the catalog-processing entry points, then trace how tokenization fits an inference-style workflow rather than a map_partitions operation. Create a notebook that serves as a reusable tokenization template, using images as the initial scope unless the project guidance expands it to multiple modalities; done means the workflow is clear and runnable.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- jupyter-notebook, python
- Domain
- data, machine-learning
- Issue type
- Documentation
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 62/100