DjangoPeng / DjangoPeng/word2vec

freebase vector file in python gensim

Open
#15 0 comments 0 reactions 0 assignees View on GitHub
auto-migrated Priority-Medium Type-Defect
Dominant language
C
Stars
3
Forks
5
PR merge metrics
No merged PRs in 30d

Description

```
What steps will reproduce the problem?
1. Load freebase.bin files into a word2vec model on freebase
2. attempy .most_similar function
3. error returned

What is the expected output? What do you see instead?
see beloe

What version of the product are you using? On what operating system?
mac osx anaconda python

Please provide any additional information below.

I’m trying to get started by loading the pretrained .bin files from the
google word2vec site ( freebase-vectors-skipgram1000.bin.gz) into the gensim
implementation of word2vec. The model loads fine,

using ..

model = word2vec.Word2Vec.load_word2vec_format('...../free....-en.bin', binary=
True)
and creates a

>>> print model

but when I run the most similar function. It cant find the words in the
vocabulary. My error code is below.

Any ideas where I’m going wrong?

>>> model.most_similar(['girl', 'father'], ['boy'], topn=3)
2013-10-11 10:22:00,562 : WARNING : word ‘girl’ not in vocabulary; ignoring
it
2013-10-11 10:22:00,562 : WARNING : word ‘father’ not in vocabulary;
ignoring it
2013-10-11 10:22:00,563 : WARNING : word ‘boy’ not in vocabulary; ignoring
it
Traceback (most recent call last):
File “”, line 1, in
File
“/....../anaconda/python.app/Contents/lib/python2.7/site-packages/gensim-0.8.7
/py2.7.egg/gensim/models/word2vec.py”, line 312, in most_similar
raise ValueError(“cannot compute similarity with no input”)
ValueError: cannot compute similarity with no input

any ideas welcome?

```

Original issue reported on code.google.com by `Mark.d.g...@gmail.com` on 11 Oct 2013 at 3:41

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the reported load and model.most_similar calls with the freebase-vectors-skipgram1000.bin.gz file and gensim 0.8.7 on the stated Anaconda/macOS setup. Inspect gensim/models/word2vec.py around most_similar and load_word2vec_format; done means identifying why the loaded vocabulary does not contain girl, father, or boy and documenting or correcting the supported loading behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.