lyst / lyst/lightfm

Add new user/item ids or features

Open
#667 1 comment 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
5.1k
Forks
724
PR merge metrics
No merged PRs in 30d

Description

Hello,

Thanks for this nice work, very efficient and designed for real wrold use cases.

Concerning the cold-start issue, you indicate in the documentation to call fit_partial method of the lightfm.data.Dataset class, and to "resize your LightFM model to be able to use the new features".
What does "resize your LightFM model to be able to use the new features" really means ?

First I train the model

from lightfm import LightFM
from lightfm.data import Dataset
from lightfm.evaluation import auc_score, precision_at_k, recall_at_k, reciprocal_rank

dataset = Dataset(user_identity_features=False, item_identity_features=True)
dataset.fit(users=train_users_df.index.unique(), 
            items=train_items_df.index, 
            item_features=train_tag_labels)

train_item_features = dataset.build_item_features(train_item_features_)
train_interactions, train_weights = dataset.build_interactions(train_users_df["MatchId"].items())

recommender = LightFM(loss='warp')

recommender = recommender.fit(interactions=train_interactions,
                              item_features=train_item_features,
                              sample_weight=train_weights,
                              epochs=NUM_EPOCHS,
                              num_threads=NUM_THREADS)

After, I would like to predict using this model but for new items, with new features, unseen during this 1st fit.
I fit again partially my Dataset without issue ...

dataset.fit_partial(users=test_users_df.index.unique(), 
                    items=test_items_df.index, 
                    item_features=test_tag_labels)

test_item_features_ = build_item_features(test_items_df)
test_item_features = dataset.build_item_features(test_item_features_)

... but I get an error when predicting with the model

recommender.predict_rank(test_interactions=next_items,
                         train_interactions=past_items,
                         item_features=test_item_features)

and I get the following error:
"ValueError: The item feature matrix specifies more features than there are estimated feature embeddings: 3623 vs 4985"
where 3623 is the number of item features saw during the fit and 4985 is the number of features after adding new items with new features.

Then, is there a way to "resize the model" as suggested in the documentation ?

Thanks

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the documentation for Dataset.fit_partial and the LightFM prediction entry point predict_rank, then trace how newly added item features are represented after fitting. Clarify what “resize your LightFM model” means and document the supported workflow or limitation, including the reported feature-count error and a way to verify the result.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.