open-compass / open-compass/opencompass

[Codellama assessment score with meta prompt is half lower than codellama assessment score without meta prompt]

Open
#768 1 comment 0 reactions 1 assignee View on GitHub

@kennymckormick is already working on this.

Since Jan 4, 2024.

Dominant language
Python
Stars
7.5k
Forks
869
Avg merge
17h 52m
Merged PRs (30d)
13

Description

Prerequisite
Type

I'm evaluating with the officially supported tasks/models/datasets.

Environment
{'CUDA available': True,
 'CUDA_HOME': '/usr/local/cuda-12.2',
 'GCC': 'gcc (GCC) 7.3.0',
 'GPU 0,1,2,3,4,5,6,7': 'NVIDIA A100-SXM4-40GB',
 'MMEngine': '0.10.1',
 'NVCC': 'Cuda compilation tools, release 12.2, V12.2.128',
 'OpenCV': '4.8.1',
 'PyTorch': '2.1.2',
 'PyTorch compiling details': 'PyTorch built with:\n'
                              '  - GCC 9.3\n'
                              '  - C++ Version: 201703\n'
                              '  - Intel(R) oneAPI Math Kernel Library Version '
                              '2023.1-Product Build 20230303 for Intel(R) 64 '
                              'architecture applications\n'
                              '  - Intel(R) MKL-DNN v3.1.1 (Git Hash '
                              '64f6bcbcbab628e96f33a62c3e975f8535a7bde4)\n'
                              '  - OpenMP 201511 (a.k.a. OpenMP 4.5)\n'
                              '  - LAPACK is enabled (usually provided by '
                              'MKL)\n'
                              '  - NNPACK is enabled\n'
                              '  - CPU capability usage: AVX2\n'
                              '  - CUDA Runtime 12.1\n'
                              '  - NVCC architecture flags: '
                              '-gencode;arch=compute_50,code=sm_50;-gencode;arch=compute_60,code=sm_60;-gencode;arch=compute_61,code=sm_61;-gencode;arch=compute_70,code=sm_70;-gencode;arch=compute_75,code=sm_75;-gencode;arch=compute_80,code=sm_80;-gencode;arch=compute_86,code=sm_86;-gencode;arch=compute_90,code=sm_90\n'
                              '  - CuDNN 8.9.2\n'
                              '  - Magma 2.6.1\n'
                              '  - Build settings: BLAS_INFO=mkl, '
                              'BUILD_TYPE=Release, CUDA_VERSION=12.1, '
                              'CUDNN_VERSION=8.9.2, '
                              'CXX_COMPILER=/opt/rh/devtoolset-9/root/usr/bin/c++, '
                              'CXX_FLAGS= -D_GLIBCXX_USE_CXX11_ABI=0 '
                              '-fabi-version=11 -fvisibility-inlines-hidden '
                              '-DUSE_PTHREADPOOL -DNDEBUG -DUSE_KINETO '
                              '-DLIBKINETO_NOROCTRACER -DUSE_FBGEMM '
                              '-DUSE_QNNPACK -DUSE_PYTORCH_QNNPACK '
                              '-DUSE_XNNPACK -DSYMBOLICATE_MOBILE_DEBUG_HANDLE '
                              '-O2 -fPIC -Wall -Wextra -Werror=return-type '
                              '-Werror=non-virtual-dtor -Werror=bool-operation '
                              '-Wnarrowing -Wno-missing-field-initializers '
                              '-Wno-type-limits -Wno-array-bounds '
                              '-Wno-unknown-pragmas -Wno-unused-parameter '
                              '-Wno-unused-function -Wno-unused-result '
                              '-Wno-strict-overflow -Wno-strict-aliasing '
                              '-Wno-stringop-overflow -Wno-psabi '
                              '-Wno-error=pedantic -Wno-error=old-style-cast '
                              '-Wno-invalid-partial-specialization '
                              '-Wno-unused-private-field '
                              '-Wno-aligned-allocation-unavailable '
                              '-Wno-missing-braces -fdiagnostics-color=always '
                              '-faligned-new -Wno-unused-but-set-variable '
                              '-Wno-maybe-uninitialized -fno-math-errno '
                              '-fno-trapping-math -Werror=format '
                              '-Werror=cast-function-type '
                              '-Wno-stringop-overflow, LAPACK_INFO=mkl, '
                              'PERF_WITH_AVX=1, PERF_WITH_AVX2=1, '
                              'PERF_WITH_AVX512=1, '
                              'TORCH_DISABLE_GPU_ASSERTS=ON, '
                              'TORCH_VERSION=2.1.2, USE_CUDA=ON, USE_CUDNN=ON, '
                              'USE_EXCEPTION_PTR=1, USE_GFLAGS=OFF, '
                              'USE_GLOG=OFF, USE_MKL=ON, USE_MKLDNN=ON, '
                              'USE_MPI=OFF, USE_NCCL=ON, USE_NNPACK=ON, '
                              'USE_OPENMP=ON, USE_ROCM=OFF, \n',
 'Python': '3.10.13 (main, Sep 11 2023, 13:44:35) [GCC 11.2.0]',
 'TorchVision': '0.16.2',
 'numpy_random_seed': 2147483648,
 'opencompass': '0.2.0+c3e0fcf',
 'sys.platform': 'linux'}
Reproduces the problem - code/configuration sample

The code of Codellama assessment process with meta prompt is as follows:

#eval_codellama_with_metaprompt.py
from mmengine.config import read_base
from opencompass.models import HuggingFaceCausalLM

with read_base():
    from .datasets.humaneval.humaneval_gen_8e312c import humaneval_datasets  # noqa: F401, F403

llama2_meta_template = dict(
    round=[
        dict(role='HUMAN', begin='<s>[INST] ', end=' [/INST] '),
        dict(role='BOT', begin='', end=' </s>', generate=True),
    ],
    eos_token_id=2)

models = [
    dict(
        type=HuggingFaceCausalLM,
        abbr='CodeLlama-13b-Instruct-hf',
        path="the local path of CodeLlama-13b-Instruct-hf",
        tokenizer_path='the local path of CodeLlama-13b-Instruct-hf',
        tokenizer_kwargs=dict(padding_side='left',
                              truncation_side='left',
                              use_fast=False,
                              ),
        max_out_len=100,
        max_seq_len=2048,
        batch_size=8,
        model_kwargs=dict(device_map='auto'),
        batch_padding=False,  # if false, inference with for-loop without batch padding
        meta_template=llama2_meta_template,
        run_cfg=dict(num_gpus=8, num_procs=1),
    )
]

datasets = [*humaneval_datasets]

The code of Codellama assessment process without meta prompt is as follows:

#eval_codellama_without_metaprompt.py
from mmengine.config import read_base
from opencompass.models import HuggingFaceCausalLM

with read_base():
    from .datasets.humaneval.humaneval_gen_8e312c import humaneval_datasets  # noqa: F401, F403

models = [
    dict(
        type=HuggingFaceCausalLM,
        abbr='CodeLlama-13b-Instruct-hf',
        path="the local path of CodeLlama-13b-Instruct-hf",
        tokenizer_path='the local path of CodeLlama-13b-Instruct-hf',
        tokenizer_kwargs=dict(padding_side='left',
                              truncation_side='left',
                              use_fast=False,
                              ),
        max_out_len=100,
        max_seq_len=2048,
        batch_size=8,
        model_kwargs=dict(device_map='auto'),
        batch_padding=False,  # if false, inference with for-loop without batch padding
        run_cfg=dict(num_gpus=8, num_procs=1),
    )
]

datasets = [*humaneval_datasets]
Reproduces the problem - command or script
python run.py configs/eval_codellama_with_metaprompt.py
python run.py configs/eval_codellama_without_metaprompt.py
Reproduces the problem - error message

eval_codellama_with_metaprompt.py

dataset           version    metric            mode      CodeLlama-13b-Instruct-hf
----------------  ---------  ----------------  ------  ---------------------------
openai_humaneval  8e312c     humaneval_pass@1  gen                            12.8

eval_codellama_without_metaprompt.py

dataset           version    metric            mode      CodeLlama-13b-Instruct-hf
----------------  ---------  ----------------  ------  ---------------------------
openai_humaneval  8e312c     humaneval_pass@1  gen                            28.05
Other information

As the website say: https://opencompass.readthedocs.io/en/latest/prompt/meta_template.html
I think codellama assessment score with meta prompt can be high than codellama assessment score without meta prompt

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.