awslabs / awslabs/graphrag-toolkit

[FEATURE] Add support for Cohere Embed v4 (configurable output dimensions)

Open
#438 0 comments 0 reactions 1 assignee Claimed by @oussamahansal View on GitHub
enhancement
Dominant language
Python
Stars
442
Forks
106
Avg merge
2d 3h
Merged PRs (30d)
41

Description

### Package

lexical-graph

### Problem statement

Add support for Amazon Bedrock's Cohere Embed v4 (`cohere.embed-v4:0`) in the lexical-graph embedding configuration. Embed v4 supports multiple output dimensions (256, 512, 1024, 1536) via the `output_dimension` request parameter (Bedrock docs (https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-embed-v4.html)). The toolkit currently has no way to pass model-specific request parameters through to the underlying BedrockEmbedding.

### Current behavior

`GraphRAGConfig.to_embedding_model()` (`lexical-graph/.../config.py:952`) accepts a `BaseEmbedding`, a model-name string, or a JSON string. The JSON path only reads `model_name`, `region_name`, and `profile_name`; there is no mechanism to supply model-specific parameters such as Embed v4's `output_dimension`. The output dimension is also coupled to `embed_dimensions` (default 1024), which sizes the vector index.

### Proposed solution

Extend the embedding model configuration to accept a supplemental/kwargs field that is forwarded to `BedrockEmbedding(additional_kwargs=...)`, so callers can select the Embed v4 output dimension. Keep `embed_dimensions` aligned with the chosen output_dimension (either derive it from the kwarg or validate that they match).

### Sample configuration

JSON-string form (the shape to_embedding_model() already parses, extended with additional_kwargs):

```
from graphrag_toolkit.lexical_graph import GraphRAGConfig
import json

GraphRAGConfig.embed_model = json.dumps({
"model_name": "cohere.embed-v4:0",
"region_name": "us-east-1",
"additional_kwargs": {
"output_dimension": 512 # one of 256 / 512 / 1024 / 1536
}
})
# Keep the index size in sync with the chosen output dimension
GraphRAGConfig.embed_dimensions = 512
```

Escape hatch available today (Option 3) — pass a fully-built instance:

```
from llama_index.embeddings.bedrock import BedrockEmbedding
from graphrag_toolkit.lexical_graph import GraphRAGConfig

GraphRAGConfig.embed_model = BedrockEmbedding(
model_name="cohere.embed-v4:0",
region_name="us-east-1",
additional_kwargs={"output_dimension": 512},
)
GraphRAGConfig.embed_dimensions = 512
```

### Implementation notes / risks

- Forward `additional_kwargs` from both the JSON-string branch and the plain-string branch of to_embedding_model() (config.py:973–989).
- Upstream dependency: the current `llama-index-embeddings-bedrock` Cohere request-body builder (`_get_request_body`) sends only texts + input_type and ignores `additional_kwargs` for Cohere (only Titan V2 reads them). To actually emit `output_dimension`, either bump to an upstream version that supports it, contribute the fix upstream, or ship a thin `BedrockEmbedding` subclass that injects `output_dimension` into the Cohere body. Verify the minimum `llama-index-embeddings-bedrock` version and pin it.
- Add validation so `embed_dimensions` and `output_dimension` cannot silently diverge (mismatched index vs. vector size).
- Add unit tests for the JSON config forwarding and dimension validation; update docs/README embedding-config section.

### Alternatives considered

1. Supplemental `additional_kwargs` in the JSON model config (recommended). Add an optional `additional_kwargs` key to the JSON-string form (and the equivalent object form) that is passed straight to BedrockEmbedding. Most flexible, forward-compatible with future model params (e.g., `embedding_types`), and requires no new config surface beyond one key. Also derive/validate `embed_dimensions` from `additional_kwargs["output_dimension"]`.
2. Dedicated top-level `EMBEDDINGS_OUTPUT_DIMENSION` setting. Simpler for the single v4 use case but narrow — every new model parameter needs another setting, and it duplicates intent with `embed_dimensions`.
3. Require callers to construct and pass a `BedrockEmbedding` instance themselves. Already works today (the instance passes through untouched) but pushes Bedrock wiring onto every user and bypasses the string/JSON config convenience the toolkit is built around. Good as a documented escape hatch, not as the primary path.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.