astronomy-commons / astronomy-commons/lsdb

Create example notebook for tokenizing catalog data

Open
#1,604 0 comments 0 reactions 0 assignees View on GitHub
documentation modalities-support
Dominant language
Python
Stars
55
Forks
26
Avg merge
4d 1h
Merged PRs (30d)
8

Description

We may choose to just do this for images for simplicity, or we could extend to multiple modalities if we wanted to overachieve. But largely this should just provide a template for the typical machinery needed to perform tokenization, which itself is an inference style workflow rather than a pure map_partitions style operation.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by locating existing example notebooks and the catalog-processing entry points, then trace how tokenization fits an inference-style workflow rather than a map_partitions operation. Create a notebook that serves as a reusable tokenization template, using images as the initial scope unless the project guidance expands it to multiple modalities; done means the workflow is clear and runnable.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook, python
Domain
data, machine-learning
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
62/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.