mattpocock / mattpocock/evalite

Synthetic Test Data Set Generation

Open
#292 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
TypeScript
Stars
1.7k
Forks
102
PR merge metrics
No merged PRs in 30d

Description

Problem

LLM evaluations depend on abundant, diverse, and regularly refreshed test data, yet manual dataset creation is slow, costly, and fails to capture complex reasoning, multi‑hop queries, edge cases, and adversarial prompts. This leads to incomplete coverage, inconsistent regression testing, and limited confidence in real world performance.

Solution

I propose a synthetic test data pipeline that generates domain grounded evaluations using a knowledge graph to structure entities and relations, personas to vary intents and language styles, and both single‑hop and multi‑hop synthesizers to control reasoning depth. The pipeline produces balanced test sets across difficulty and scenarios for continuous regression, with traceable provenance from graph nodes, persona templates, and hop strategies.

Knowledge Graph
Image

We ingest source content, use chunking strategies to split it into chunks, then run extractors to label entities and topics per chunk. Each chunk becomes a node, and we build edges with topic similarity and entity overlap to connect related nodes. This yields a compact, navigable graph that mirrors domain structure and preserves provenance to the original content.

Scenario Generation
Image

We sample graph nodes to assemble context passages, then use single‑hop and multi‑hop synthesizers to generate queries that require shallow or compositional reasoning. Personas vary intent, tone, and language style, while configurable question lengths and formats ensure balanced coverage. Each scenario outputs the user query, the source context used, and a golden/reference answer with provenance to the contributing nodes.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No files, tests, or entry points are identified in the issue, so first locate the existing evaluation and data-generation surfaces. Define completion around a reproducible pipeline that produces balanced single- and multi-hop scenarios with persona variation, reference answers, source context, and provenance.

Written by the indexing model from the issue text.

Assessment

Tech stack
typescript
Domain
ai, data, testing-qa
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.