awslabs / awslabs/synthetically_engineered_evaluation_data

Generated highly-visual documents (e.g. driver's licenses) don't look realistic

Open
#16 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
9
Forks
1
PR merge metrics
No merged PRs in 30d

Description

## Description

When generating highly-visual identity documents such as US state driver's licenses (e.g. Florida DLs) from a text description — including with low-quality scan/fax augmentation — the resulting outputs do not visually resemble real driver's licenses.

## Context

SEED's current generation capability is primarily focused on **letter-sized PDF documents** (text-heavy, form-style layouts). It was not designed to produce **highly visual documents** with photo-ID-style layouts, graphics, holograms/security features, portrait photos, and compact card dimensions. As a result, visual ID documents come out looking unrealistic.

## Proposed Enhancement

Add support for generating realistic visual documents, for example by integrating an **image-generation tool (e.g. Amazon Nova)** to synthesize realistic imagery — driver's license layouts, portrait photos, and other fake-but-realistic image data — rather than relying solely on the existing PDF/text-oriented pipeline.

## Steps to Reproduce

1. Request generation of a US (e.g. Florida) driver's license from a description.
2. Apply low-quality scan/fax augmentation.
3. Observe that the output does not resemble an actual driver's license.

[doc_0001.pdf](https://github.com/user-attachments/files/30557478/doc_0001.pdf)
[doc_0002.pdf](https://github.com/user-attachments/files/30557475/doc_0002.pdf)
[doc_0003.pdf](https://github.com/user-attachments/files/30557474/doc_0003.pdf)
[doc_0004.pdf](https://github.com/user-attachments/files/30557477/doc_0004.pdf)
[doc_0005.pdf](https://github.com/user-attachments/files/30557476/doc_0005.pdf)

Contributor guide

Open the contributing guide

Research direction

The issue does not name files, tests, or an entry point. Start by tracing the current PDF/text-oriented generation pipeline, then determine how a proposed image-generation integration such as Amazon Nova would fit; done means generated visual identity documents resemble realistic card layouts, imagery, and augmentations.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, python
Domain
ai, data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.