huggingface / huggingface/sentence-transformers

Further fine-tuned the language model with custom corpus

Open
#101 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
19.1k
Forks
2.9k
Avg merge
1d 19h
Merged PRs (30d)
61

Description

Hi, the repository is amazing and the architecture is fantastic stuff.
Along with this, I have below query:
I have my own **text file** which contains sentences per line, how can I use this my custom corpus to further fine-tuned the trained language model, as I want to perform **domain-specific text similarity task**.
Could you suggest any way to fine-tune the model with the **unsupervised** way as I do not have a labeled dataset?

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no file, test, or entry point. First clarify the expected unsupervised fine-tuning workflow for a text file and domain-specific similarity task; completion would require an agreed implementation or documented supported approach.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.