catboost / catboost/catboost

Slow single prediction in python

Open
#1,346 3 comments 0 reactions 0 assignees View on GitHub
in progress performance planned python
Dominant language
C++
Stars
9.1k
Forks
1.3k
PR merge metrics
No merged PRs in 30d

Description

Problem: Prediction on a single row works slowly in python
catboost version: 0.23.2
Operating System: Linux
CPU: Intel

# Catboost prediction works slower than XGBoost/lightGBM

Evironment : Colab

```import numpy as np
from catboost import CatBoostClassifier
from lightgbm import LGBMClassifier
from xgboost import XGBClassifier
import time

train_size = 1000
train_features = 30
X_train = np.random.randn(train_size, train_features)
y_train = np.random.randint(2, size=train_size)
nrow = np.random.randn(1, train_features)

n_esimators = 200

cb = CatBoostClassifier(verbose=0, iterations=n_esimators, thread_count=2)
cb.fit(X_train, y_train)

xgb = XGBClassifier(n_jobs=2, n_estimators=n_esimators)
xgb.fit(X_train, y_train)

lgb = LGBMClassifier(n_jobs=2, n_estimators=n_esimators)
lgb.fit(X_train, y_train)

c0 = time.time()
for _ in range(1000):
__ = cb.predict_proba(nrow, thread_count=1)
print(time.time() - c0) # 0.4030110836029053

c0 = time.time()
for _ in range(1000):
_ = xgb.predict_proba(nrow)
print(time.time() - c0) # 0.0883171558380127

c0 = time.time()
for _ in range(1000):
_ = lgb.predict_proba(nrow)
print(time.time() - c0) # 0.14197039604187012
```

# Catboost prediction works slower with thread_count > 1

Operating System: Ubuntu 18.04
CPU Cores/Thread : 8/16

```
for thread_count in [1, 2, 8, 16]:
c0 = time.time()
for _ in range(1000):
__ = cb.predict_proba(nrow, thread_count=thread_count)
print(time.time() - c0)

0.30289459228515625
0.32805943489074707
0.5098705291748047
0.9639191627502441
```

# Simple cpp model export to python works 30+ times faster

Environment: Colab

For this experiment i used cppyy module that can easily export functions from cpp

```
import cppyy

cb = CatBoost(params=dict(verbose=0, iterations=n_esimators, thread_count=2))
cb.fit(X_train, y_train)
cb.save_model("cpp_model.cpp", format="cpp")

with open('cpp_model.cpp', 'r') as f:
cf = f.read()

cppyy.cppdef(cf)

nrow_py = list(nrow[0])
rounds = 1000
c0 = time.time()
for _ in range(rounds):
__ = cppyy.gbl.ApplyCatboostModel(nrow_py)
cpptime = time.time() - c0
print(cpptime)

c0 = time.time()
for _ in range(rounds):
__ = cb.predict(nrow_py, prediction_type='RawFormulaVal')
pytime = time.time() - c0
print(pytime)
print(pytime/cpptime)

0.013201713562011719
0.46095943450927734
34.91663656721809
```

I'm sure you can increase speed for single prediction by many times

Contributor guide

Open the contributing guide

Research direction

Start by running the issue's Colab benchmark around CatBoostClassifier.predict_proba with a one-row nrow input and compare thread_count values against the shown timings. Then inspect the Python prediction path and the exported ApplyCatboostModel comparison to identify where single-row overhead occurs. Done means reproducing the behavior and demonstrating a justified improvement without regressing normal prediction.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, python
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.