Slow engine creation - only when using FastAPI

Open
#422 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
25/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Stale
Tech stack
fastapi, python, sql, sqlalchemy
Domain
api, databases

Research direction

Start with the FastAPI reproduction in main.py and test_route.py, then compare its create_engine path with the standalone databricks.sql connection example. Profile the import and execution of databricks/sql/thrift_api/TCLIService/ttypes.py under the pinned requirements.txt versions. Done means the cause of the FastAPI-only delay is established and a supported remedy is demonstrated without manually editing the generated file.

Written by the indexing model from the issue text.

Description

Hello,

I am currently working on a FastAPI application which calls Databricks using databricks-sql-connector, however the code appears to slow down massively on the engine creation stage.

I've narrowed the issue down to .venv\Lib\site-packages\databricks\sql\thrift_api\TCLIService\ttypes.py, where when running in debug mode the script is loaded incredibly slowly (around 5-10 minutes). However when the exact same code is run outside of a FastAPI function (ie, in a standalone script, notebook, etc) it runs almost instantly and the engine is created in under 1 second.

The details of my test application are:

main.py:

from fastapi import FastAPI
from fastapi.responses import RedirectResponse
import uvicorn
from routes import test_route


app = FastAPI(
     title="TestAPI",
     version="1.0",
     description="Test API",
     )

@app.get("/", include_in_schema=False)
async def docs_redirect():
     return RedirectResponse(url='/docs')

app.include_router(test_route.router)

if __name__ == '__main__':
     uvicorn.run('main:app', host='0.0.0.0', port=8000)

test_route.py:

import fastapi
from fastapi import APIRouter

from sqlalchemy import create_engine
from sqlalchemy import text

router = APIRouter(tags=["Test"])

@router.post("/test_route", description="Test")
async def test():
    import time

    print("Connecting to engine....")
    startT = time.time()
    engine = create_engine(
        url = f"databricks://token:<<TOKEN>>@<<HOST>>?http_path=/sql/1.0/warehouses/<<WAREHOUSE>>&catalog=<<CATALOG>>&schema=<<SCHEMA>>"
    )
    print("Engine Connected!")
    print(f"Time taken: {time.time() - startT} seconds.")

    with engine.connect() as connection:
        result = connection.execute(text("SELECT * from range(10)"))
        for row in result:
            print(row)

    print(f"Execution time: {time.time() - startT} seconds.")
    pass

    return "Ok."

requirements.txt:

fastapi==0.111.1
uvicorn==0.30.3
databricks-sql-connector[SQLAlchemy]==3.3.0

The same issue occurs when using:

from databricks import sql

connection = sql.connect(
                        server_hostname = <HOST>,
                        http_path = <HTTP>
                        access_token = <TOKEN>)

cursor = connection.cursor()

cursor.execute("SELECT * from range(10)")
print(cursor.fetchall())

cursor.close()
connection.close()

but given it all goes to create_engine under the hood I was trying to simplify my test case.

As mentioned, the call stack points to .venv\Lib\site-packages\databricks\sql\thrift_api\TCLIService\ttypes.py taking all of this extra time. This can be fixed through editing ttypes.py directly and removing all rows which begin None, #. However this directly violates the warning given at the top of ttypes.py about not editing. Removing these rows reduces the file from ~100,000 lines to ~10,000, and completely solves the time delay issue when using FastAPI.

So my main questions are:

  • Why does this delay in engine creation only occur in a FastAPI app (is it something to do with it being asynchronous?).
  • How can this delay be remedied?
  • Can ttypes.py be safely edited to remove all "None" rows or is this not a viable solution?

Thanks!

Dominant language
Python
Stars
233
Forks
152
Avg merge
21h 5m
Merged PRs (30d)
10

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from databricks/databricks-sql-python

All issues in databricks/databricks-sql-python

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.