Use `Terminal-Bench 4.0` instead of `Terminal-Bench 2.1` in the `eval` results
- Dominant language
- TypeScript
- Stars
- 5.4k
- Forks
- 502
- Avg merge
- 1d 2h
- Merged PRs (30d)
- 715
Description
### Problem
- Currently, https://maka.apache.org/en/ directs users to https://github.com/apache/maka/blob/main/docs/eval/terminal-bench-2.1-deepseek-v4-flash-nine-arm.md . The relevant tests utilize `Terminal-Bench 2.1`, which is certainly fine.
- However, with the release of `Terminal-Bench 4.0`, using a benchmark that `DeepSeek V4 Flash` has not encountered during training to evaluate the model offers greater discriminative power.
- See https://www.tbench.ai/run and https://hub.harborframework.com/datasets/terminal-bench/terminal-bench/4?leaderboard=4-0-0&tab=leaderboard .
- Of course, the lack of tokens and GPUs could be an issue. ASF does not provide tokens for `DeepSeek V4 Flash`, and the free tier of GitHub Actions Runners lacks GPUs entirely—yet running `Terminal-Bench` requires GPUs.
### Desired outcome
- Use `Terminal-Bench 4.0` instead of `Terminal-Bench 2.1` in the `eval` results.
### Alternatives or workarounds
- See https://maka.apache.org/en/ and https://github.com/apache/maka/blob/main/docs/eval/terminal-bench-2.1-deepseek-v4-flash-nine-arm.md .
Contributor guide
Research direction
Start with docs/eval/terminal-bench-2.1-deepseek-v4-flash-nine-arm.md and the link from https://maka.apache.org/en/. Check how the existing Terminal-Bench 2.1 results were produced and whether the available GitHub Actions runners can run Terminal-Bench 4.0. Done means the eval results use Terminal-Bench 4.0 and the public link points to the updated results.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- github-actions, markdown
- Domain
- documentation, machine-learning
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 48/100