awslabs / awslabs/synthetically_engineered_evaluation_data
Generated highly-visual documents (e.g. driver's licenses) don't look realistic
- Dominant language
- Python
- Stars
- 9
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Description
## Description
When generating highly-visual identity documents such as US state driver's licenses (e.g. Florida DLs) from a text description — including with low-quality scan/fax augmentation — the resulting outputs do not visually resemble real driver's licenses.
## Context
SEED's current generation capability is primarily focused on **letter-sized PDF documents** (text-heavy, form-style layouts). It was not designed to produce **highly visual documents** with photo-ID-style layouts, graphics, holograms/security features, portrait photos, and compact card dimensions. As a result, visual ID documents come out looking unrealistic.
## Proposed Enhancement
Add support for generating realistic visual documents, for example by integrating an **image-generation tool (e.g. Amazon Nova)** to synthesize realistic imagery — driver's license layouts, portrait photos, and other fake-but-realistic image data — rather than relying solely on the existing PDF/text-oriented pipeline.
## Steps to Reproduce
1. Request generation of a US (e.g. Florida) driver's license from a description.
2. Apply low-quality scan/fax augmentation.
3. Observe that the output does not resemble an actual driver's license.
[doc_0001.pdf](https://github.com/user-attachments/files/30557478/doc_0001.pdf)
[doc_0002.pdf](https://github.com/user-attachments/files/30557475/doc_0002.pdf)
[doc_0003.pdf](https://github.com/user-attachments/files/30557474/doc_0003.pdf)
[doc_0004.pdf](https://github.com/user-attachments/files/30557477/doc_0004.pdf)
[doc_0005.pdf](https://github.com/user-attachments/files/30557476/doc_0005.pdf)
Contributor guide
Research direction
The issue does not name files, tests, or an entry point. Start by tracing the current PDF/text-oriented generation pipeline, then determine how a proposed image-generation integration such as Amazon Nova would fit; done means generated visual identity documents resemble realistic card layouts, imagery, and augmentations.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, python
- Domain
- ai, data
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100