microsoft / microsoft/sql-ai-promptathon

Mission: Data Scientist - Zava Semantic Retrieval Audit

Open
#8 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Shell
Stars
49
Forks
132
PR merge metrics
No merged PRs in 30d

Description

Mission/open goal Description

Built a semantic theme-discovery and retrieval-quality audit over Zava's
multilingual customer feedback (English, Spanish, and French reviews and
support chats). The goal was to use precomputed embeddings and the
find_similar_docs_by_doc_id vector tool to surface semantically related
feedback across languages, then audit how trustworthy those matches
actually are — producing an analysis grounded in real documents rather
than a polished demo.

Harness and model

GitHub Copilot Chat (Agent mode) in VS Code Codespaces, powered by GPT-5.1

Turn-by-turn journey
  1. Prompt: Explore the SQL database to find tables related to customer reviews and support chats.
    Agent response or action: Ran describe_entities to list all entities, then read_records to sample rows from SupportChat, SupportTicket, and Doc tables.
    Result: Identified Docs as the main feedback corpus containing reviews and support-chat text with precomputed embeddings.

  2. Prompt: Pick diverse seed documents and run the vector tool to find similar documents.
    Agent response or action: Selected 6 seeds (reviews + support chats, spanning English/Spanish/French) and ran find_similar_docs_by_doc_id for each.
    Result: Returned top-5 nearest neighbors with cosine distances; smart-fabric connectivity issues matched correctly across all three languages.

  3. Prompt: Label each neighbor as true match / near miss / false positive and build a precision@k audit notebook.
    Agent response or action: Created zava_semantic_retrieval_audit.ipynb, labeled all 30 neighbor pairs, and computed precision@k per seed.
    Result: Overall precision@k of 0.89 across 6 seeds, with a summary noting the 83-document corpus is too small to support broad claims.

Completion
  • Yes, the agent completed the mission or goal.
  • No, the agent did not complete the mission or goal.
Bonus work

Computed an honest precision@k metric (0.89 average) and explicitly
documented the limitations of the 83-document corpus rather than
overclaiming retrieval quality — included a model-card style note on
what the system is good for and where it falls short.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the issue's described SQL exploration and the zava_semantic_retrieval_audit.ipynb notebook. Reproduce the six multilingual seed searches with find_similar_docs_by_doc_id, inspect the 30 labeled neighbor pairs, and verify the precision@k calculation. Done means the audit and its limitations are reproducible against the 83-document corpus.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook, sql
Domain
data, databases, machine-learning, search, testing-qa
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.