CusmaLinux / CusmaLinux/svu-ai

Review the simula framework to create synthetic datasets

Open
#1 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
No language data
Stars
1
Forks
0
PR merge metrics
No merged PRs in 30d

Description

### Objective
Review the simula frameowrk of google to create synthetic datsets

### Context
The first step to create the pipeline for classify pqrs to the responsible office is the dataset, but even if we obtaion 500-1000 real records, we need to create more to reach a maximum of 3000 records, this limit can change in base reommendations of the differents approaches in the ML world.

### Checklist

- [ ] Review the [simula framework of google](https://research.google/blog/designing-synthetic-datasets-for-the-real-world-mechanism-design-and-reasoning-from-first-principles/)

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading the linked Google Research article about synthetic datasets and relate its approach to the planned PQRS-to-office classification dataset. The issue is done when the framework review, recommended dataset size and approach, and criteria for using synthetic records are documented; no repository file, test, or entry point is named.

Written by the indexing model from the issue text.

Assessment

Tech stack
machine-learning
Domain
data, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.