google-gemini / google-gemini/gemini-cli

Long-Context & Complex Reasoning Coding Evaluation Dataset

Open
#23,316 80 comments 22 reactions 0 assignees View on GitHub
🔒 maintainer only area/platform kind/enhancement priority/p2 status/bot-triaged
Dominant language
TypeScript
Stars
107k
Forks
14.6k
Avg merge
2d 3h
Merged PRs (30d)
45

Description

### What would you like to be added?

**Long-Context & Complex Reasoning Coding Evaluation Dataset**
**Difficulty**: Medium | **Size**: 175 hours | **Area**: Innovation

**Description**
As current agent evaluation benchmarks (such as SWE-bench Pro and TerminalBench) saturate, they are becoming less effective at measuring true enterprise-level capabilities due to their relatively simple tasks and isolated codebases. To rigorously stress-test the Gemini CLI, we must develop a benchmark that accurately mirrors the complexity of real-world production environments.

This project focuses on architecting a highly challenging, novel dataset comprising 30-50 massive, multi-language repositories. The core objective is to curate these extensive codebases and strategically extract real-world engineering problems that explicitly demand complex, multi-step reasoning across expansive context windows. Ultimately, this dataset will act as the definitive proving ground for our agent's intelligence, verifying its ability to seamlessly navigate, understand, and resolve intricate architectural dependencies spanning thousands of lines of code.

**Expected Outcomes**
Repository Curation: Identify and onboard 30-50 large-scale, highly active open-source repositories spanning a diverse array of modern programming languages and enterprise frameworks.

**Task Formulation**: Extract and define complex, multi-file engineering tasks (e.g., deep architectural bug fixes, sweeping cross-component feature integrations) that strictly require long-context comprehension to solve.

**Schema Design**: Develop a standardized, robust dataset schema optimized for automated, reproducible agent evaluation.

**Pipeline Integration**: Seamlessly integrate this novel dataset into the Gemini CLI's existing evaluation and testing pipelines.

**Baseline Analysis**: Deliver a comprehensive baseline performance report detailing the Gemini CLI's current success rates, while critically analyzing and categorizing specific failure modes in its long-context reasoning.

**Skills Required / Preferred**
Required: Strong proficiency in at least two major programming languages (e.g., Python, Node.js), coupled with deep, practical experience using Git to navigate and manipulate large, complex codebases.

**Preferred**: Familiarity with the landscape of LLM evaluation benchmarks, applied prompt engineering, autonomous AI agents, or building automated data curation pipelines.

### Why is this needed?

High Quality large dataset curation.

### Additional context

_No response_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.