CusmaLinux / CusmaLinux/svu-ai
Review the simula framework to create synthetic datasets
- Dominant language
- No language data
- Stars
- 1
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
### Objective
Review the simula frameowrk of google to create synthetic datsets
### Context
The first step to create the pipeline for classify pqrs to the responsible office is the dataset, but even if we obtaion 500-1000 real records, we need to create more to reach a maximum of 3000 records, this limit can change in base reommendations of the differents approaches in the ML world.
### Checklist
- [ ] Review the [simula framework of google](https://research.google/blog/designing-synthetic-datasets-for-the-real-world-mechanism-design-and-reasoning-from-first-principles/)
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading the linked Google Research article about synthetic datasets and relate its approach to the planned PQRS-to-office classification dataset. The issue is done when the framework review, recommended dataset size and approach, and criteria for using synthetic records are documented; no repository file, test, or entry point is named.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- machine-learning
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100