spring-cloud / spring-cloud/spring-cloud-function

AWS Lambda SnapStart with priming hangs during restore phase when using Spring Cloud Function for AWS

Offen
#1,450 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Vorherrschende Sprache
Java
Sterne
1.1k
Forks
641
Ø Merge
11 Std. 2 Min.
Gemergte PRs (30 T.)
8

Beschreibung

Here is the application https://github.com/Vadym79/aws-lambda-java-25-spring-boot-4/tree/main/aws-spring-cloud-function-dynamodb. You can deploy it with SAM within minutes. SnapStart is on for all Lambdas. I use there this SnapStart priming resource https://github.com/Vadym79/aws-lambda-java-25-spring-boot-4/blob/main/aws-spring-cloud-function-dynamodb/src/main/java/software/amazonaws/example/product/handler/FullPrimingResource.java which causes the problem as SnapStart restore doesn't work (it times out after 10 seconds)

If you then invoke the Lambda function with the name GetProductByIdWithJava25SpringBoot40AWSSCFDynamoDB with some product id (like 1, it doesn't matter, it shouldn't even exist) via API Gateway, you'll see the issue.

SnapStart works for another exisiting priming resource https://github.com/Vadym79/aws-lambda-java-25-spring-boot-4/blob/main/aws-spring-cloud-function-dynamodb/src/main/java/software/amazonaws/example/product/handler/DynamoDBPrimingResource.java which is now deactivated (@Configuration annotation removed), but I'd like to have them both work.

SnapStart worked for nearly the same application https://github.com/Vadym79/AWSLambdaJavaWithSpringBoot/tree/master/spring-boot-3.4-with-spring-cloud-function but it used the older versions: Java 21 (now 25), Spring Boot 3.4 (now 4.0) and Spring Cloud Function 4.2.0 (now 5.0.1).

I contacted AWS Serverless team via email and asked for the investigation and after quite some time, they responded to me the following:

After investigating, this is not a Lambda SnapStart platform bug, the snapshot and restore are working correctly.

The issue is with Spring Cloud Function 5.x's Netty dependency not being CRaC-aware. We'd recommend seeking guidance from the upstream projects:

After I asked for a bit more details, they wrote to me the following:

CRaC-aware" means a library properly handles the checkpoint/restore lifecycle — closing and reopening OS-level resources (sockets, file descriptors, threads) around the snapshot boundary. Libraries that don't do this can cause restores to hang or timeout, which is what customer is experiencing.

For specifics on what needs to change in Spring Cloud Function 5.x and its dependencies, we'd recommend having a conversation with the Spring team.

From the Lambda side, SnapStart's checkpoint/restore mechanism is working correctly — the issue is in the application-level framework behavior during restore. Unfortunately there isn’t any Lambda-side configuration knob to "ignore stale sockets" or "force-close resources on restore".

I can establish contact to the AWS folks, who investigated the problem if required.

I tested now with Spring cloud Fucntion 5.0.4 but the problem persists.

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Stelle die verknüpfte aws-spring-cloud-function-dynamodb-Anwendung mit SAM bereit und rufe GetProductByIdWithJava25SpringBoot40AWSSCFDynamoDB auf, um das Restore-Timeout zu reproduzieren. Vergleiche FullPrimingResource.java mit DynamoDBPrimingResource.java und dem älteren Spring Boot 3.4-Beispiel und konzentriere dich dabei auf das Abhängigkeitsverhalten von Spring Cloud Function 5.x. Als abgeschlossen gilt die Aufgabe, wenn der SnapStart-Restore abgeschlossen wird, sobald beide Priming-Ressourcen aktiviert sind.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
aws, java, spring
Bereich
backend, cloud
Issue-Typ
Bug
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Aktiv
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
45/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.