4paradigm / 4paradigm/OpenMLDB
too many type systems in source code but serves same purpose
- Dominant language
- C++
- Stars
- 1.7k
- Forks
- 331
- Avg merge
- 12d 12h
- Merged PRs (30d)
- 1
Description
- https://github.com/4paradigm/OpenMLDB/blob/ff7e8acf21ead1f9734eef59ac3521adb05dff2e/hybridse/include/node/node_enum.h#L154-L165
- https://github.com/4paradigm/OpenMLDB/blob/ff7e8acf21ead1f9734eef59ac3521adb05dff2e/hybridse/src/proto/fe_type.proto#L21-L27
- https://github.com/4paradigm/OpenMLDB/blob/ff7e8acf21ead1f9734eef59ac3521adb05dff2e/src/proto/type.proto#L25-L32
- `base::StringRef, base::Timestamp` .. etc in udf system
- https://github.com/4paradigm/OpenMLDB/blob/ff7e8acf21ead1f9734eef59ac3521adb05dff2e/hybridse/include/sdk/base_schema.h#L27-L32
And also different implementations build upon those type systems,
Contributor guide
Research direction
The issue points to multiple type definitions across hybridse/include/node/node_enum.h, hybridse/src/proto/fe_type.proto, src/proto/type.proto, and base_schema.h. Start by examining each linked file to understand the overlapping type systems. Look for how these types are used in the UDF system and other implementations. Determine a strategy to unify or reduce duplication, which requires deep knowledge of the codebase's architecture.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- backend, databases, machine-learning
- Issue type
- Refactor
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 20/100