epfl-dlab / epfl-dlab/transformers-CFG

Bus error when running single string as input on MacOS

Open
#84 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
139
Forks
23
PR merge metrics
No merged PRs in 30d

Description

This bug is very wired....

I can not reproduce it on linux machine but only on my local macbook....

Bascially, with a batch size of 1, I get this wired error

```python
input_ids = tokenizer(
[prefix1], add_special_tokens=False, return_tensors="pt", padding=True
)["input_ids"] # 2415 bus error python examples/generate_json.py
```

->
`[1] 2415 bus error python examples/generate_json.py`

But with batch size = 2, it disappears...

```python
input_ids = tokenizer(
[prefix1, prefix1], add_special_tokens=False, return_tensors="pt", padding=True
)["input_ids"]
```

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from transformers_cfg.grammar_utils import IncrementalGrammarConstraint
from transformers_cfg.generation.logits_process import GrammarConstrainedLogitsProcessor

if __name__ == "__main__":

model_id = "gpt2"

# Load model and tokenizer
tokenizer = AutoTokenizer.from_pretrained(model_id)
tokenizer.pad_token = tokenizer.eos_token

model = AutoModelForCausalLM.from_pretrained(model_id)
model.generation_config.pad_token_id = model.generation_config.eos_token_id

# Load grammar
with open("examples/grammars/json.ebnf", "r") as file:
grammar_str = file.read()
grammar = IncrementalGrammarConstraint(grammar_str, "root", tokenizer)
grammar_processor = GrammarConstrainedLogitsProcessor(grammar)

# Generate
prefix1 = "This is a valid json string for http request:"
prefix2 = "This is a valid json string for shopping cart:"
input_ids = tokenizer(
[prefix1], add_special_tokens=False, return_tensors="pt", padding=True
)["input_ids"] # 2415 bus error python examples/generate_json.py

# input_ids = tokenizer(
# [prefix1, prefix2], add_special_tokens=False, return_tensors="pt", padding=True
# )["input_ids"] # this works fine

output = model.generate(
input_ids,
do_sample=False,
max_new_tokens=60,
logits_processor=[grammar_processor],
repetition_penalty=1.1,
num_return_sequences=1,
)
# decode output
generations = tokenizer.batch_decode(output, skip_special_tokens=True)
print(generations)

"""
'This is a valid json string for http request:{ "request": { "method": "GET", "headers": [], "content": "Content","type": "application" }}
'This is a valid json string for shopping cart:This is a valid json string for shopping cart:{ "name": "MyCart", "price": 0, "value": 1 }
"""

```

Contributor guide

No contributing guide indexed for this repository

Research direction

The reproduction is in examples/generate_json.py; start by running the tokenizer call with the single-item batch on macOS, then compare it with the two-item batch. Trace whether the bus error occurs during tokenization or later in the grammar-constrained generation path. Done means the single-item example completes without a bus error.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning, operating-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.