docling-project / docling-project/docling

Integrate docling with RapidDoc (Table and Formula models)

Open
#3,147 1 comment 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
66.4k
Forks
4.8k
Avg merge
2d 21h
Merged PRs (30d)
84

Description

This is a feature request to enable docling to use table and formula models from RapidDoc (https://github.com/RapidAI/RapidDoc). They are generally faster and more accurate than the ones in docling_ibm_models.

Below is a test example. RapidDoc performs better for tables with columns spanning other columns and also runs faster than both default docling table model and also version 2 table model.

Image

### rapiddoc result (took 11 seconds)

Image

### docling result (took 18 seconds)

## TABLE I

RATiO (RAT.) of THE RMSE oBTAINED By kNN WITHoUT PS oVER THE RMSE oF kNN WITH PS. STRICT IMPROVEMENTS oF THE USE OF HML+&NN ovEr kNN aRE PrinTED In BoLD. THe coluMN Red. PRESENTS tHE REDUCTION AFTER PS. COLUMN TIME LISTS tHE EXECUTION TIME IN SECONDS OF PS+kNN.

| Dataset | % noise | Red.(%) | k = 1 Rat. | Time | Red.(%) | k = 10 Rat. | Time |
|-----------|-----------|-----------|--------------|---------|-----------|---------------|---------|
| SID | 0 | 99.9963 | 1.2872 | 1092.18 | 70.2196 | 1.0006 | 1147.18 |
| SID | 10 | 99.9963 | 1,2983 | 1065.59 | 98.7768 | 1.0055 | 1133.59 |
| SID | 20 | 99.9963 | 1.3093 | 1059.96 | 98.7768 | 1.008 | 1099.96 |
| SID | 30 | 99.9963 | 1.3099 | 1089.28 | 99.9913 | 1.0097 | 1097.28 |
| SID | 40 | 99.9959 | 1.3324 | 1853.81 | 46.1739 | 1.0004 | 1885.81 |
| SID | 50 | 99.8817 | 1.2583 | 1084.65 | 98.7768 | 1.0125 | 1104.65 |
| MEPS | 0 | 99.9929 | 1.1682 | 39.74 | 0.0142 | 1.0032 | 40.74 |
| MEPS | 10 | 99.9929 | 1.1948 | 32.17 | 0.0142 | 1.0028 | 33.17 |
| MEPS | 20 | 99.9929 | 1.2235 | 31.94 | 0.0142 | 1.0025 | 33.94 |
| MEPS | 30 | 99.9929 | 1.2565 | 32.12 | 0.0142 | 1.0021 | 34.12 |
| MEPS | 40 | 99.6367 | 1.0348 | 31.96 | 0.0142 | 1.0017 | 33.96 |
| MEPS | 50 | 99.9929 | 1.2966 | 32.26 | 99.9929 | 1.0127 | 33.26 |

ments. For 10NN, the reduction is still quite high for the SID data, but small for most MEPS datasets. The lower reduction rate is explained by the lower susceptibility of kNN to noise for higher values of k. As a result, fewer instances are removed to increase the performance of the regression method. Only for the version in which 50% of the instances were perturbed, does the reduction on the MEPS data attain a high level.

### docling result (using TableStructureV2Options - took 15 seconds)

## TABLE I

RATiO (RAT.) of THE RMSE oBTAINED By kNN WITHoUT PS oVER THE RMSE oF kNN WITH PS. STRICT IMPROVEMENTS oF THE USE OF HML+&NN ovEr kNN aRE PrinTED In BoLD. THe coluMN Red. PRESENTS tHE REDUCTION AFTER PS. COLUMN TIME LISTS tHE EXECUTION TIME IN SECONDS OF PS+kNN.

| Dataset | % noise | Red.(%) | k = 1 Rat. | Time | Red.(%) | k = 10 Rat. | Time |
|-----------|-----------|-----------|--------------|---------|-----------|---------------|---------|
| SID | 0 | 99.9963 | 1.2872 | 1092.18 | 70.2196 | 1.0006 | 1147.18 |
| SID | 10 | 99.9963 | 1,2983 | 1065.59 | 98.7768 | 1.0055 | 1133.59 |
| SID | 20 | 99.9963 | 1.3093 | 1059.96 | 98.7768 | 1.008 | 1099.96 |
| SID | 30 | 99.9963 | 1.3099 | 1089.28 | 99.9913 | 1.0097 | 1097.28 |
| SID | 40 | 99.9959 | 1.3324 | 1853.81 | 46.1739 | 1.0004 | 1885.81 |
| SID | 50 | 99.8817 | 1.2583 | 1084.65 | 98.7768 | 1.0125 | 1104.65 |
| MEPS | 0 | 99.9929 | 1.1682 | 39.74 | 0.0142 | 1.0032 | 40.74 |
| MEPS | 10 | 99.9929 | 1.1948 | 32.17 | 0.0142 | 1.0028 | 33.17 |
| MEPS | 20 | 99.9929 | 1.2235 | 31.94 | 0.0142 | 1.0025 | 33.94 |
| MEPS | 30 | 99.9929 | 1.2565 | 32.12 | 0.0142 | 1.0021 | 34.12 |
| MEPS | 40 | 99.6367 | 1.0348 | 31.96 | 0.0142 | 1.0017 | 33.96 |
| MEPS | 50 | 99.9929 | 1.2966 | 32.26 | 99.9929 | 1.0127 | 33.26 |

ments. For 10NN, the reduction is still quite high for the SID data, but small for most MEPS datasets. The lower reduction rate is explained by the lower susceptibility of kNN to noise for higher values of k. As a result, fewer instances are removed to increase the performance of the regression method. Only for the version in which 50% of the instances were perturbed, does the reduction on the MEPS data attain a high level.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.