OpenEuroLLM / OpenEuroLLM/Taskboard

Benchmark Qwen3.5-9B Base and post-trained Qwen3.5-9B on BFCL V4 non-Web

Open
#384 0 comments 0 reactions 1 assignee View on GitHub

@ReinforcedKnowledge is already working on this.

Since Sep 14, 2026.

4.6 post-training
Dominant language
No language data
Stars
3
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Goal

Establish comparable BFCL V4 non-Web baselines for Qwen/Qwen3.5-9B-Base and the post-trained Qwen/Qwen3.5-9B, and validate the evaluation path used for function-calling experiments.

Scope

  • Run both checkpoints on the same BFCL V4 non-Web categories and configuration.
  • Validate serving, chat-template handling, parsing, expected row coverage, scoring, aggregation, and server logs.
  • Report the exact model and harness revisions, configuration, overall score, and category scores.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.