LAION-AI / LAION-AI/CLAP

Loss Implementation

Open
#109 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
2.3k
Forks
213
PR merge metrics
No merged PRs in 30d

Description

I'm a little confused by the exact loss used in the experiments. From #100 it is clear that you only use the projection layers which is basically Linear > ReLU > Linear > Normalise. So MLP loss is not used. Then in the loss function, assuming world_size == 1 we end up with the calculation:

logits_per_audio = logit_scale_a * audio_features @ text_features.T
logits_per_text = logit_scale_a * text_features @ audio_features.T

However, in the paper, after equation 3 you say:

Where τ is a learnable temperature parameter for scaling the loss.
Two logarithmic terms consider either audio-to-text logits or text-to-audio logits.

Only the audio-to-text logit parameter is used in the code if you don't use the mlp layer loss. Should the logits_per_text actually use the logit_scale_t parameter instead of logit_scale_a parameter? Currently only the mlp_loss seems to use the logit_scale_t parameter, but you say that isn't used in #100. The paper says both logit_scale parameters are used but that doesn't seem to line up with the code. Then there is the weighted_loss as well which I'm not sure is used or not.

Could you please clarify?

Thank you.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Begin by tracing the loss function for world_size == 1 and comparing logits_per_audio, logits_per_text, logit_scale_a, logit_scale_t, mlp_loss, and weighted_loss with equation 3 in the paper and issue #100. Done means the discrepancy is resolved and the intended implementation or documentation clearly explains which loss terms and temperature parameters are used.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.