aws / aws/bedrock-agentcore-sdk-python

[Feature] @kb_transformation decorator for Knowledge Base custom transformation Lambdas

Abierto
#488 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
761
Forks
147
Merge medio
1 d 23 h
PR fusionados (30 d)
7

Descripción

## Summary

Knowledge Bases supports [custom transformation Lambda functions](https://docs.aws.amazon.com/bedrock/latest/userguide/kb-custom-transformation.html) that run during ingestion to apply custom chunking or metadata enrichment. Today, customers must manually parse the KB event format, read/write content batches from S3, and return the exact expected response structure — significant boilerplate that obscures the actual transformation logic.

The SDK should provide a `@kb_transformation` decorator that handles all the plumbing (S3 I/O, event parsing, batch iteration, and response formatting) so the customer only writes their transformation logic.

## Proposed API

```python
from bedrock_agentcore.knowledge_base import kb_transformation

@kb_transformation
def my_chunker(content: str, metadata: dict) -> list[dict]:
"""Custom chunking — split on headings and enrich metadata."""
chunks = content.split("\n# ")
return [
{"contentBody": chunk, "contentMetadata": {"section": i}}
for i, chunk in enumerate(chunks)
]
```

The decorator would:

1. **Parse the Lambda event** — extract the S3 bucket/key for the input content batch and the output location
2. **Handle S3 I/O** — read content objects from S3, write transformed output back to S3
3. **Iterate over batches** — call the user function once per document/content item in the batch
4. **Format the response** — return the exact structure the Knowledge Bases ingestion pipeline expects

## Current pain points (without the decorator)

- Customers must understand and implement the undocumented event schema for the transformation Lambda
- S3 read/write logic (including error handling, content-type detection) must be implemented from scratch
- Batch iteration and response assembly is repetitive boilerplate across every transformation Lambda
- Small mistakes in the response structure cause silent ingestion failures that are hard to debug

## Expected behavior

- The decorator should accept a function with signature `(content: str, metadata: dict) -> list[dict]` at minimum
- Each dict in the return list represents a chunk with at least `contentBody` (str) and optionally `contentMetadata` (dict)
- The decorator should handle all S3 operations, event parsing, and response formatting transparently
- Errors in the user function should surface clearly rather than being swallowed by the S3/response plumbing
- Should work as a standard AWS Lambda handler (compatible with `lambda_handler` entry point patterns)

## Additional considerations

- Should the decorator support async transformation functions?
- Should there be a lower-level variant that gives access to the raw batch (for cases where per-document iteration isn't desired)?
- Consider exposing the source document metadata (filename, content type, data source ID) to the user function for context-aware transformations

## References

- [Knowledge Bases Custom Transformation User Guide](https://docs.aws.amazon.com/bedrock/latest/userguide/kb-custom-transformation.html)
- [AgentCore SDK Knowledge Bases design doc](https://quip-amazon.com/aMbfABnehnH7#LPc9AAhEqQj)

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

Comienza con la guía de transformación personalizada de AWS Knowledge Bases y el documento de diseño de AgentCore SDK enlazado; después, inspecciona los patrones existentes del punto de entrada lambda_handler. Se considerará terminado cuando exista una API @kb_transformation documentada que gestione el análisis del evento indicado, la E/S de S3, la iteración por lotes, el formateo de la respuesta y los errores claros de la función del usuario, y se hayan tomado decisiones sobre las cuestiones async, raw-batch y metadata enumeradas.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
aws, python
Área
backend, cloud
Tipo de issue
Nueva funcionalidad
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Tranquilo
Claridad
Bastante claro
Aptitud para principiantes
38/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.