codeforboston / codeforboston/maple
Test Out New AssemblyAI Model (Universal 3.5-Pro)
- Dominant language
- TypeScript
- Stars
- 56
- Forks
- 175
- Avg merge
- 2d 5h
- Merged PRs (30d)
- 13
Description
## Summary
The new AssemblyAI model supports up to 30 speakers with speaker diarization (up from our original ~10).
This is a substantial lift in terms of our ability to potentially map voice lines with speakers in our hearing transcripts - we should test out the new model to determine:
* Does the speaker diarization now work properly (First level: Does it distinguish the speakers correctly? Second level: Can it name the legislators (whose names are known via committee/agenda)?)
* We can provide a list of known names to help improve accuracy - we should send in the list of all legislators on the Committee holding the hearing at minimum (possibly there's more specific data in the agenda)
* Cost Difference (if any)
* Compare to the `transcriptions//utterances` to see how different the diarization actually is between versions?
*
We would have to re-process all of our hearings with the new model to get this if it proves useful - so we should determine how much of a lift it really is in terms of data quality.
Contributor guide
Research direction
Start by reviewing the existing AssemblyAI transcription flow and the transcriptions//utterances data. Test Universal 3.5-Pro on a hearing using committee or agenda names, then compare speaker diarization, legislator naming, and cost with the current results. Done means documenting the data-quality and cost differences well enough to decide whether reprocessing hearings is worthwhile.
Written by the indexing model from the issue text.
Assessment
- Domain
- audio-video-rtc
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100