google-gemini / google-gemini/gemini-cli
Long-Context & Complex Reasoning Coding Evaluation Dataset
- Dominant language
- TypeScript
- Stars
- 107k
- Forks
- 14.6k
- Avg merge
- 2d 3h
- Merged PRs (30d)
- 45
Description
### What would you like to be added?
**Long-Context & Complex Reasoning Coding Evaluation Dataset**
**Difficulty**: Medium | **Size**: 175 hours | **Area**: Innovation
**Description**
As current agent evaluation benchmarks (such as SWE-bench Pro and TerminalBench) saturate, they are becoming less effective at measuring true enterprise-level capabilities due to their relatively simple tasks and isolated codebases. To rigorously stress-test the Gemini CLI, we must develop a benchmark that accurately mirrors the complexity of real-world production environments.
This project focuses on architecting a highly challenging, novel dataset comprising 30-50 massive, multi-language repositories. The core objective is to curate these extensive codebases and strategically extract real-world engineering problems that explicitly demand complex, multi-step reasoning across expansive context windows. Ultimately, this dataset will act as the definitive proving ground for our agent's intelligence, verifying its ability to seamlessly navigate, understand, and resolve intricate architectural dependencies spanning thousands of lines of code.
**Expected Outcomes**
Repository Curation: Identify and onboard 30-50 large-scale, highly active open-source repositories spanning a diverse array of modern programming languages and enterprise frameworks.
**Task Formulation**: Extract and define complex, multi-file engineering tasks (e.g., deep architectural bug fixes, sweeping cross-component feature integrations) that strictly require long-context comprehension to solve.
**Schema Design**: Develop a standardized, robust dataset schema optimized for automated, reproducible agent evaluation.
**Pipeline Integration**: Seamlessly integrate this novel dataset into the Gemini CLI's existing evaluation and testing pipelines.
**Baseline Analysis**: Deliver a comprehensive baseline performance report detailing the Gemini CLI's current success rates, while critically analyzing and categorizing specific failure modes in its long-context reasoning.
**Skills Required / Preferred**
Required: Strong proficiency in at least two major programming languages (e.g., Python, Node.js), coupled with deep, practical experience using Git to navigate and manipulate large, complex codebases.
**Preferred**: Familiarity with the landscape of LLM evaluation benchmarks, applied prompt engineering, autonomous AI agents, or building automated data curation pipelines.
### Why is this needed?
High Quality large dataset curation.
### Additional context
_No response_
Contributor guide
Assessment
This issue has not been assessed yet.