4paradigm / 4paradigm/OpenMLDB

disk table supports duplicate key?

Aperta
#2,318 0 commenti 0 reazioni 1 assegnatario Rivendicata da @zhanghaohit Vedi su GitHub
enhancement storage-engine
Lingua principale
C++
Stelle
1.7k
Fork
331
Merge medio
12g 12h
PR unite (30g)
1

Descrizione

**Describe the feature you'd like**

One big difference between disk table and mem table is that disk table only allows one unique key per index.
For example, if the index is . the rows we inserted are:

|id |col1 | col2|
|-------| ------ | ----- |
1 | k1 | 1
2 | k1 | 1
3 | k2 | 1
4 | k2 | 2

In mem table, it will be:
|id |col1 | col2|
|-------| ------ | ----- |
1 | k1 | 1
2 | k1 | 1
3 | k2 | 1
4 | k2 | 2

However, in disk table, it will be:
|id |col1 | col2|
|-------| ------ | ----- |
1 | k1 | 1
3 | k2 | 1
4 | k2 | 2

because row 1 and row 2 have the same value of .

We should discuss whether we should support the same semantics for both disk table and memory table.

**Additional context**
One possible solution is to add a sequence number to the combined key for disk table.

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Look at the disk table and memory table implementations to understand how keys are handled. The issue suggests adding a sequence number to the combined key for disk tables to allow duplicates. Start by examining the index structures and insertion logic, then design a change that maintains performance while supporting duplicate keys. Testing will involve verifying that duplicate rows are preserved in disk tables as they are in memory tables.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
sql
Ambito
backend, databases
Tipo di issue
Funzionalità
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Ferma
Chiarezza
Specificata chiaramente
Idoneità per principianti
45/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.