Accenture / Accenture/AmpliGraph

Ampligraph Embedding for Protein Sequences

Open
#275 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
2.2k
Forks
257
PR merge metrics
No merged PRs in 30d

Description

Hi

I am new to AmpliGraph Embedding. I want to know, will AmpliGraph Embedding accurately generate embedding of a protein sequence? I was exploring the ScoringBasedEmbeddingModel used in Ampligraph Embedding and uses the TransE scoring method. TransE is developed on the text-based knowledge base. So the question arises, how any text based embedding model can generate embedding accurately for a protein or kinase enzyme in the knowledge graph data like ?

I was trying to get the embedding of the protein sequences of Protein Knowledge graph data https://www.zjukg.org/project/ProteinKG25/

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the documentation for ScoringBasedEmbeddingModel and the TransE scoring method, then review the ProteinKG25 dataset linked in the issue. The issue names no source files or tests, and a completion condition is not defined because it asks whether protein-sequence embeddings would be accurate rather than requesting a specific change.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.