spring-cloud / spring-cloud/spring-cloud-function

AWS Lambda SnapStart with priming hangs during restore phase when using Spring Cloud Function for AWS

オープン
#1,450 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

主要言語
Java
スター
1.1k
フォーク
641
平均マージ
11時間 2分
マージ済み PR(30日)
8

説明

Here is the application https://github.com/Vadym79/aws-lambda-java-25-spring-boot-4/tree/main/aws-spring-cloud-function-dynamodb. You can deploy it with SAM within minutes. SnapStart is on for all Lambdas. I use there this SnapStart priming resource https://github.com/Vadym79/aws-lambda-java-25-spring-boot-4/blob/main/aws-spring-cloud-function-dynamodb/src/main/java/software/amazonaws/example/product/handler/FullPrimingResource.java which causes the problem as SnapStart restore doesn't work (it times out after 10 seconds)

If you then invoke the Lambda function with the name GetProductByIdWithJava25SpringBoot40AWSSCFDynamoDB with some product id (like 1, it doesn't matter, it shouldn't even exist) via API Gateway, you'll see the issue.

SnapStart works for another exisiting priming resource https://github.com/Vadym79/aws-lambda-java-25-spring-boot-4/blob/main/aws-spring-cloud-function-dynamodb/src/main/java/software/amazonaws/example/product/handler/DynamoDBPrimingResource.java which is now deactivated (@Configuration annotation removed), but I'd like to have them both work.

SnapStart worked for nearly the same application https://github.com/Vadym79/AWSLambdaJavaWithSpringBoot/tree/master/spring-boot-3.4-with-spring-cloud-function but it used the older versions: Java 21 (now 25), Spring Boot 3.4 (now 4.0) and Spring Cloud Function 4.2.0 (now 5.0.1).

I contacted AWS Serverless team via email and asked for the investigation and after quite some time, they responded to me the following:

After investigating, this is not a Lambda SnapStart platform bug, the snapshot and restore are working correctly.

The issue is with Spring Cloud Function 5.x's Netty dependency not being CRaC-aware. We'd recommend seeking guidance from the upstream projects:

After I asked for a bit more details, they wrote to me the following:

CRaC-aware" means a library properly handles the checkpoint/restore lifecycle — closing and reopening OS-level resources (sockets, file descriptors, threads) around the snapshot boundary. Libraries that don't do this can cause restores to hang or timeout, which is what customer is experiencing.

For specifics on what needs to change in Spring Cloud Function 5.x and its dependencies, we'd recommend having a conversation with the Spring team.

From the Lambda side, SnapStart's checkpoint/restore mechanism is working correctly — the issue is in the application-level framework behavior during restore. Unfortunately there isn’t any Lambda-side configuration knob to "ignore stale sockets" or "force-close resources on restore".

I can establish contact to the AWS folks, who investigated the problem if required.

I tested now with Spring cloud Fucntion 5.0.4 but the problem persists.

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

リンクされた aws-spring-cloud-function-dynamodb アプリケーションを SAM でデプロイし、GetProductByIdWithJava25SpringBoot40AWSSCFDynamoDB を呼び出して restore timeout を再現します。FullPrimingResource.java と DynamoDBPrimingResource.java、および以前の Spring Boot 3.4 の例を比較し、Spring Cloud Function 5.x の依存関係の挙動に注目します。両方の priming リソースを有効にした状態で SnapStart restore が完了すれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
aws, java, spring
領域
backend, cloud
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
活発
明瞭さ
おおむね明確
初心者へのやさしさ
45/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。