google / google/leveldb

Iterator::Seek() gets extremely slow

Open
#777 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
39.4k
Forks
8.2k
PR merge metrics
No merged PRs in 30d

Description

I am running a leveldb instance with 30GB of data. Now it happens that Iterator::Seek() gets extremely slow, as slow as it takes several hours to finish.

The leveldb.stats is:

```
Compactions
Level Files Size(MB) Time(sec) Read(MB) Write(MB)
--------------------------------------------------
0 1946 29505 0 0 0
```

I debug it by adding some printf() codes, and find out `DBIter::FindNextUserEntry()` is the problem.

Here is my code added:

```
printf("%s %d %s %s\n", __FILE__, __LINE__, ikey.DebugString().c_str(), EscapeString(*skip).c_str());
```

after

https://github.com/google/leveldb/blob/5903e7a1125cacaa1d44367b5b84fe9208e42884/db/db_iter.cc#L188
https://github.com/google/leveldb/blob/5903e7a1125cacaa1d44367b5b84fe9208e42884/db/db_iter.cc#L193

Here is some outputs:

```
db/db_iter.cc 189 '\x01\x00\x00\x00\x00\x00f\x99@' @ 98350051 : 0 \x01\x00\x00\x00\x00\x00f\x99@
db/db_iter.cc 195 '\x01\x00\x00\x00\x00\x00f\x99@' @ 21077915 : 1 \x01\x00\x00\x00\x00\x00f\x99@
db/db_iter.cc 189 '\x01\x00\x00\x00\x00\x00f\x99A' @ 98350052 : 0 \x01\x00\x00\x00\x00\x00f\x99A
db/db_iter.cc 195 '\x01\x00\x00\x00\x00\x00f\x99A' @ 21077918 : 1 \x01\x00\x00\x00\x00\x00f\x99A
db/db_iter.cc 189 '\x01\x00\x00\x00\x00\x00f\x99B' @ 98350053 : 0 \x01\x00\x00\x00\x00\x00f\x99B
db/db_iter.cc 195 '\x01\x00\x00\x00\x00\x00f\x99B' @ 21077922 : 1 \x01\x00\x00\x00\x00\x00f\x99B
db/db_iter.cc 189 '\x01\x00\x00\x00\x00\x00f\x99C' @ 98350054 : 0 \x01\x00\x00\x00\x00\x00f\x99C
db/db_iter.cc 195 '\x01\x00\x00\x00\x00\x00f\x99C' @ 21077925 : 1 \x01\x00\x00\x00\x00\x00f\x99C
db/db_iter.cc 189 '\x01\x00\x00\x00\x00\x00f\x99D' @ 98350055 : 0 \x01\x00\x00\x00\x00\x00f\x99D
db/db_iter.cc 195 '\x01\x00\x00\x00\x00\x00f\x99D' @ 21077929 : 1 \x01\x00\x00\x00\x00\x00f\x99D
```

As the output shows, the keys are added and then deleted, but both remain in sst files, so DBIter skips them one by one before reaching at the target. For my use case, I deleted millions of keys, that's why it gets so slow, it iterates over millions of obsoleted keys on a hard disk!

So my question is, why shouldn't these keys be really "deleted"? They do not exist in the database scope, they are marked deleted in level-0 sst file, and they do not exist in higher level sst files, they should be really deleted after compaction.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.