NVIDIA / NVIDIA/NeMo-Retriever

[FEA]: Consistent ids when connecting llamaIndex to a Milvus VDB populated by NV-Ingest

Open
#205 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

feature request
Dominant language
Python
Stars
3k
Forks
349
Avg merge
1d 23h
Merged PRs (30d)
116

Description

Is this a new feature, an improvement, or a change to existing functionality?

New Feature

How would you describe the priority of this feature request

Significant improvement

Please provide a clear description of problem this feature solves

Currently, when I perform an extraction, upload the results to the Milvus VDB, and then query the VDB with LlamaIndex, the node Ids of the retrieved results change every retrieval, even if the content is the same. For example if I upload a document to the VDB:

from nv_ingest_client.client import NvIngestClient
from nv_ingest_client.primitives import JobSpec
from nv_ingest_client.primitives.tasks import ExtractTask
from nv_ingest_client.primitives.tasks import SplitTask
from nv_ingest_client.primitives.tasks import EmbedTask
from nv_ingest_client.primitives.tasks import VdbUploadTask
from nv_ingest_client.util.file_processing.extract import extract_file_content
import logging, time

logger = logging.getLogger("nv_ingest_client")

file_name = "data/multimodal_test.pdf"
file_content, file_type = extract_file_content(file_name)

job_spec = JobSpec(
    document_type=file_type,
    payload=file_content,
    source_id=file_name,
    source_name=file_name,
    extended_options={"tracing_options": {"trace": True, "ts_send": time.time_ns()}},
)

extract_task = ExtractTask(
    document_type=file_type,
    extract_text=True,
    extract_images=False,
    extract_tables=True,
)

embed_task = EmbedTask(
    text=True,
    tables=True,
)

vdb_upload_task = VdbUploadTask()

job_spec.add_task(extract_task)
job_spec.add_task(embed_task)
job_spec.add_task(vdb_upload_task)


client = NvIngestClient()
job_id = client.add_job(job_spec)

client.submit_job(job_id, "morpheus_task_queue")

result = client.fetch_job_result(job_id, timeout=60)

And then connect to the VDB with LlamaIndex:

embed_model = NVIDIAEmbedding(base_url="http://localhost:8012/v1", model="nvidia/nv-embedqa-e5-v5")

vector_store = MilvusVectorStore(
    uri="http://localhost:19530",
    collection_name="nv_ingest_collection",
    doc_id_field="pk",
    embedding_field="vector",
    text_key="text",
    dim=1024,
    overwrite=False
)
index = VectorStoreIndex.from_vector_store(vector_store=vector_store, embed_model=embed_model)
retriever = index.as_retriever(similarity_top_k=1)

And then retrieve a document:

res = retriever.retrieve("What was the dog doing?")

And get the id:

res[0].id_

I get:

'd87a82ea-c968-42b7-84ae-f628b759eac6'

However If i do it again:

res = retriever.retrieve("What was the dog doing?")
res[0].id_

I get:

'ab450386-7d9b-4d1f-8b58-2735d4cacd76'

Despite that the text is the same in both cases:

'locations. Animal Activity Place Giraffe Driving a car. At the beach Lion Putting on sunscreen At the park. Cat Jumping onto a laptop In a home office Dog Chasing a squirrel In the front yard'

This might be a LlamaIndex issue but when I upload documents to Milvus through LlamaIndex and set the Id with LLamaIndex I get a stable Id when retrieving.

Describe the feature, and optionally a solution or implementation and any alternatives

Ideally I would like the id to be consistent and mapped to the pk field in the nv_ingest_collection

Additional context

No response

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the issue using VdbUploadTask, the nv_ingest_collection Milvus collection, and LlamaIndex's MilvusVectorStore with doc_id_field set to pk. Trace how the pk value becomes res[0].id_ across repeated retrievals; done means identical content returns a stable id mapped to the collection's pk field, with the behavior verified by repeating the provided retrieval example.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
databases
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.