Slow engine creation - only when using FastAPI
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 25/100
Hướng nghiên cứu
Bắt đầu với bản tái hiện FastAPI trong main.py và test_route.py, sau đó so sánh đường dẫn create_engine của nó với ví dụ kết nối databricks.sql độc lập. Phân tích hiệu năng của việc import và thực thi databricks/sql/thrift_api/TCLIService/ttypes.py trong các phiên bản được cố định trong requirements.txt. Công việc được xem là hoàn tất khi đã xác định được nguyên nhân của độ trễ chỉ xảy ra với FastAPI và đã chứng minh được một cách khắc phục được hỗ trợ mà không chỉnh sửa thủ công tệp được sinh ra.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Hello,
I am currently working on a FastAPI application which calls Databricks using databricks-sql-connector, however the code appears to slow down massively on the engine creation stage.
I've narrowed the issue down to .venv\Lib\site-packages\databricks\sql\thrift_api\TCLIService\ttypes.py, where when running in debug mode the script is loaded incredibly slowly (around 5-10 minutes). However when the exact same code is run outside of a FastAPI function (ie, in a standalone script, notebook, etc) it runs almost instantly and the engine is created in under 1 second.
The details of my test application are:
main.py:
from fastapi import FastAPI
from fastapi.responses import RedirectResponse
import uvicorn
from routes import test_route
app = FastAPI(
title="TestAPI",
version="1.0",
description="Test API",
)
@app.get("/", include_in_schema=False)
async def docs_redirect():
return RedirectResponse(url='/docs')
app.include_router(test_route.router)
if __name__ == '__main__':
uvicorn.run('main:app', host='0.0.0.0', port=8000)
test_route.py:
import fastapi
from fastapi import APIRouter
from sqlalchemy import create_engine
from sqlalchemy import text
router = APIRouter(tags=["Test"])
@router.post("/test_route", description="Test")
async def test():
import time
print("Connecting to engine....")
startT = time.time()
engine = create_engine(
url = f"databricks://token:<<TOKEN>>@<<HOST>>?http_path=/sql/1.0/warehouses/<<WAREHOUSE>>&catalog=<<CATALOG>>&schema=<<SCHEMA>>"
)
print("Engine Connected!")
print(f"Time taken: {time.time() - startT} seconds.")
with engine.connect() as connection:
result = connection.execute(text("SELECT * from range(10)"))
for row in result:
print(row)
print(f"Execution time: {time.time() - startT} seconds.")
pass
return "Ok."
requirements.txt:
fastapi==0.111.1
uvicorn==0.30.3
databricks-sql-connector[SQLAlchemy]==3.3.0
The same issue occurs when using:
from databricks import sql
connection = sql.connect(
server_hostname = <HOST>,
http_path = <HTTP>
access_token = <TOKEN>)
cursor = connection.cursor()
cursor.execute("SELECT * from range(10)")
print(cursor.fetchall())
cursor.close()
connection.close()
but given it all goes to create_engine under the hood I was trying to simplify my test case.
As mentioned, the call stack points to .venv\Lib\site-packages\databricks\sql\thrift_api\TCLIService\ttypes.py taking all of this extra time. This can be fixed through editing ttypes.py directly and removing all rows which begin None, #. However this directly violates the warning given at the top of ttypes.py about not editing. Removing these rows reduces the file from ~100,000 lines to ~10,000, and completely solves the time delay issue when using FastAPI.
So my main questions are:
- Why does this delay in engine creation only occur in a FastAPI app (is it something to do with it being asynchronous?).
- How can this delay be remedied?
- Can
ttypes.pybe safely edited to remove all "None" rows or is this not a viable solution?
Thanks!
- Ngôn ngữ chính
- Python
- Star
- 233
- Fork
- 152
- Merge trung bình
- 21 giờ 5 phút
- Pull request đã merge (30 ngày)
- 10
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của databricks/databricks-sql-python
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Tất cả issue của databricks/databricks-sql-python
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
PolicyEngine/policyengine-us#9559 ·
-
priority: p3
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
googleapis/librarian#7636 ·
-
from:qa priority:P2 reliability tech-debt
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
spec-kitty/spec-kitty#4874 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100