graphql-python / graphql-python/graphql-core

performance issues due to breadth first execution of grapqhl queries in case of async resolvers during calls burst

未關閉
#200 4 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
主要語言
Python
星號
531
分支
146
PR 合併指標
30 天內沒有已合併 PR

描述

When we receive calls burst we found that the calls "wait each other" (i.e. the first call of the burst waits the last one).
This result in degradation of performances in both time of execution and memory consumption in the server because we have to keep many calls in fly.

This is particularly evident in big graphql queries where the users request many fields and we have several depth in the queries where each level has many async fields.

Just to be TLDR, looking the implementations of graphql and asyncio we understood that this is due to the following:
- graphql breadth first way to schedule and resolve the fields
- asyncio internal FIFO queue of tasks to be executed

As an example lets have queries like that, where data may be like beer vendors and we want for each beer vendor many fields that describes that vendor, a1...a100, b1...b100, ...:

```
query {
data {
a1 {
b1 {
c1
...
c100
}
...
b100 {
c1
...
c100
}
}
...
a100 { ... }
}
}
```

If we have n of this calls coming in burst when we arrive to the depth of the c fields we have many many task scheduled in the asyncio queue.

If we check the of order of execution we have that the first query, on each level, "waits" the other queries, because all the queries schedules a lot of tasks.

In the proof of concept, that you may find at the end of the post, you can verify the order of execution of the resolvers.

It could be very nice to have some sort of priority in the order to let the first query not wait the scheduling and resolve of all the queries before ending.
I understand that this is something between graphql and asyncio but i think it could affect the use of graphql in environments receiving many calls.
Fixes, helps and hints in how to improve this would be very appreciated.

```python
import asyncio

from graphene import ObjectType, Schema, String, Field

FIELD_NUMBER = 2
CONCURRENT_QUERIES = 10

def make_resolver(i, j=None):
async def resolver(self, info):
print(f"START query {info.context['query_number']} | a{i} | b{j}")
await asyncio.sleep(0.001)
print(f"END query {info.context['query_number']} | a{i} | b{j}")
return i

return resolver

def create_fields():
fields = {}
for i in range(FIELD_NUMBER):
inner_fields = {}
for j in range(FIELD_NUMBER):
inner_fields[f"b{j}"] = String()
inner_fields[f"resolve_b{j}"] = make_resolver(i, j)

MyType = type(
f"MyType",
(ObjectType,),
inner_fields,
)

fields[f"a{i}"] = Field(MyType)
fields[f"resolve_a{i}"] = make_resolver(i)

return fields

async def make_query(schema, query_number):
inner_query_values = [f"b{i}" for i in range(FIELD_NUMBER)]
query_values = [
"a%s {%s}" % (i, " ".join(inner_query_values)) for i in range(FIELD_NUMBER)
]
query_string = "{ %s }" % (" ".join(query_values),)

await schema.execute_async(
query_string, context_value=dict(query_number=query_number)
)

async def main():
Query = type("Query", (ObjectType,), create_fields())
schema = Schema(query=Query)

await asyncio.gather(*[make_query(schema, i) for i in range(CONCURRENT_QUERIES)])

asyncio.run(main())

```

貢獻指南

這個儲存庫沒有索引到貢獻指南

研究方向

從 issue 的 asyncio 概念驗證和 schema.execute_async 進入點開始,然後追蹤非同步 GraphQL 欄位如何被排程和解析。將該行為與並行查詢下 asyncio 的工作佇列進行比較。完成這項工作需要達成一致的排程或優先順序方案,並透過報告中的突發情境加以驗證,同時不降低執行時間或增加記憶體使用量。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
api, backend
Issue 類型
功能
難度
5/5
預估耗時
一週以上
活躍度
停滯
描述清晰度
需要釐清
新手友好度
25/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。