[AI Triage] Spike: Crawler and evaluator for MS Docs sources
- Dominant language
- C#
- Stars
- 135
- Forks
- 260
- Avg merge
- 3d 1h
- Merged PRs (30d)
- 144
Description
The goal for this spike is to complete within the agreed upon timebox and to check difficulty and feasibility of creating a crawler for MS Docs articles which can:
- Determine the type of content and identify general articles related to the Azure SDK packages and/or Azure topics
- Use an AI model to evaluate the quality and relevance of an article to recommend whether or not it should be included in the knowledge base for a repository
- Ideally, perform results-based testing for known representative issues in the area relevant to the article to determine if suggestion quality is positively or negatively impacted by including the article.
At the end of the spike, we should:
- Have a clear answer on "is this feasible" and "should we consider this as a feature?"
- Understand a high-level effort for what it would take to implement this as a feature.
- If this is a potential feature, have notes or other assets that capture the approach the spike used.
- If this is a potential feature, have the reference implementation available for future inspiration.
Contributor guide
Research direction
Start by reviewing the crawler, article-type detection, AI quality evaluation, and results-based testing goals described in the issue, then establish the agreed timebox. Done means documenting feasibility, a high-level feature effort, the proposed approach, and a reference implementation if the feature is viable.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- azure, csharp
- Domain
- ai, documentation, tooling
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100