MoonshotAI / MoonshotAI/PerceptionBench
PBench & VisRes
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 208
- Forks
- 12
- PR merge metrics
- No merged PRs in 30d
Description
Thanks for the nice contribution. We've been doing some work in this area and would appreciate if you could test Kimi K3 on a few benchmarks we introduced:
VisRes https://arxiv.org/pdf/2512.21194
Image-only visual reasoning (completion, rules, composition); good check on whether K3 relies on language priors vs. actual visual abstraction.
SalBench https://arxiv.org/abs/2507.04741
Low-level saliency / odd-one-out; shows if K3 catches obvious pop-out features humans spot instantly -> a blind spot for many frontier VLMs.
PBench (counting / compositional grounding) -> https://arxiv.org/abs/2603.27365 https://huggingface.co/datasets/tiiuae/PBench
Compositional counting under dense, crowded scenes; would stress-test K3 beyond PerceptionBench's Count slice, especially with OCR + spatial constraints.
Bests,
Yasser
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No file, test, or entry point is named. Start by reviewing the repository's existing benchmark evaluation workflow, then determine how VisRes, SalBench, and PBench should be tested with Kimi K3. Done means producing comparable results for the requested benchmarks and documenting the evaluation output.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning, testing-qa
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100