dotnet / dotnet/extensions

[API Proposal]: IDocumentExtractionClient — a document-extraction capability as its own peer library (Microsoft.Extensions.DocumentExtraction)

Open
#7,587 2 comments 3 reactions 1 assignee Claimed by @luisquintanilla View on GitHub
api-suggestion area-ai
Dominant language
C#
Stars
3.2k
Forks
894
Avg merge
1d 12h
Merged PRs (30d)
23

Description

# [API Proposal]: `IDocumentExtractionClient` — a document-extraction capability as its own peer library (`Microsoft.Extensions.DocumentExtraction`)

> [!IMPORTANT]
> **Review status:** This proposal describes the current 32-type `[Experimental]` surface implemented
> by #7588. The prototype, provider surveys, and implementation validate that the shape is feasible;
> they do not replace reviewer approval.
>
> **Confirmation requested in this review:** the standalone
> `Microsoft.Extensions.DocumentExtraction(.Abstractions)` home and name; the neutral `Document*`
> model living there for v1; the `GetService` / `RawRepresentation` / `AdditionalProperties` posture;
> the v1 middleware boundary; and the next formal approval gate.
>
> **Implementation status (2026-08-10):** #7588 is open and non-draft at `a215825ae2`; its Ubuntu,
> Windows, coverage, aggregate CI, and CLA checks are green.

Proposal change history and superseded intermediate states

> **Updates**
> - _2026-07-30_ — **Moved out of `Microsoft.Extensions.AI` into its own peer library, `Microsoft.Extensions.DocumentExtraction`.** Per review steer, document extraction is its own domain (like `VectorData` / `DataIngestion`), not an `IChatClient` sibling inside M.E.AI: it pulls structured content *out* of documents, complementing `DataIngestion`, which feeds content *into* RAG. The capability keeps the M.E.AI building-block **shape** (abstraction + delegating base + builder + logging/OpenTelemetry/configure-options middleware + DI) and references `Microsoft.Extensions.AI.Abstractions` internally (e.g. `DataContent`) without carrying the "AI" brand. This **absorbs the shared-model + family-rename tracks** below (both now resolved by the move) and lets a provider team (e.g. Azure AI Document Intelligence) own an `IDocumentExtractionClient` impl without an "AI" branding dependency. Naming flips from the M.E.AI-internal verb scheme (`SpeechToText`) to the peers' domain-noun convention: a neutral **`Document*`** content model + **`DocumentExtraction*`** operation/client types. New diagnostic id **`MEDE0001`** (peer precedent: VectorData `MEVD9001`). The extraction builds green as a straw-man (both new packages + M.E.AI(.Abstractions) with OCR removed + the `DataIngestion` consumer + both new test projects, 60 tests). The speclet below reflects the new `Document*` content model and `DocumentExtraction*` operation/client names, with public-API baselines regenerated across all five TFMs (netstandard2.0, net462, net8/9/10). Methods (`ExtractAsync`/`ExtractPagesAsync`/`GetService`/`AsBuilder`) are unchanged. The surface stays one `[Experimental]` unit so a review-driven name change remains a single mechanical re-rename.
> - _2026-07-23_ — **Pre-review spikes: per-page coordinate model, selective promotes, geometry primitives and `GetService` kept.** SPIKE-07/08 moved `DocumentCoordinateUnit` / `DocumentCoordinateOrigin` back onto `DocumentPage` (reported **per page**, not per document — engines emit different units for different pages in mixed image/PDF batches, per Google `Page.Dimension` and Azure DI `DocumentPage.Unit`), and grouped `DocumentPage.Width` / `Height` into a new `DocumentPageDimensions` value type. SPIKE-06's 14-engine raw-output inventory kept `RawRepresentation` / `AdditionalProperties` and promoted per-cell `BoundingRegion` / `Confidence` / `RawRepresentation` / `AdditionalProperties` onto `DocumentTableCell` plus `RawRepresentation` onto `DocumentPage`. SPIKE-09 kept the three OCR-owned geometry primitives (no cross-platform BCL type carries the page-scoped, rotation-capable `DocumentBoundingRegion` polygon). SPIKE-04/05 kept `IDocumentExtractionClient.GetService` (one optional seam for provider metadata, provider-SDK escape, and adapter unwrap). Surface is now **32 public types** (adds `DocumentPageDimensions`).
> - _2026-07-22_ — **Superseded intermediate state: API-review reshape with document-level coordinates.** Collapsed the parallel `DocumentPage.Blocks`/`Tables`/`Images` into one reading-order `DocumentPage.Elements` over a new polymorphic `DocumentElement` base (`DocumentBlock`/`DocumentTable`/`DocumentImage` derive; project with `OfType()`), and added optional nested `DocumentTableCell.Elements`. The coordinate metadata was temporarily moved to the document result; the July 23 provider evidence superseded that placement and restored it per page. Renamed `ExtractStreamingAsync` → `ExtractPagesAsync` and `OcrResponseUpdate` → `DocumentExtractionPageResult` (non-null `Page`, dropped `Status`); `Markdown` → `Text`; added `DocumentExtractionUsage` token counts; added `DocumentTableCellKind.RowHeader`/`RowSection`; dropped response `ModelId`, `DocumentPage.Confidence`, and `DocumentExtractionOptions.IncludeImages`; `DocumentBoundingRegion.FromRectangle` now takes `float`. Surface was **31 public types** at this point.
> - _2026-07-15_ — **API-review alignment (family symmetry).** Replaced the unary-only shape with the family's unary + streaming pair: added `IAsyncEnumerable ExtractStreamingAsync(...)` (the `IChatClient.GetStreamingResponseAsync` twin), `OcrResponseUpdate`, and an `OcrResponseUpdateExtensions.ToDocumentExtractionResult`/`ToDocumentExtractionResultAsync` reducer, and **removed `IProgress` and `DocumentExtractionProgress`** (progress now rides on the streamed update). `DocumentBlock.Kind` / `DocumentTableCell.Kind` are now `ChatRole`-style open structs (`DocumentBlockKind` / `DocumentTableCellKind`), not raw strings. Added `DocumentPage.Width` / `Height` + `DocumentCoordinateUnit` so bounding-box coordinates are interpretable across engines. Removed the leaky `DocumentExtractionResult.OcrSource` (`ModelId` + `DocumentExtractionClientMetadata.ProviderName` carry provenance). Unsealed the result / data types to match `ChatResponse` / `ChatOptions`. Surface is now **29 public types**.
> - _2026-07-14_ — Reformatted to the API-proposal template and refreshed the **API Proposal** speclet to match the surface implemented in PR #7588: `GetTextAsync` → `ExtractAsync`; `IDocumentExtractionClient : IDisposable`; typed geometry `DocumentPoint` / `DocumentBoundingBox` with `DocumentBoundingRegion.Polygon` as `IReadOnlyList`; 1-based `DocumentPage.PageNumber`; `[Experimental("MEDE0001")]`; added `DocumentImage`, `DocumentPage.Images`, `DocumentExtractionOptions.Clone()`, `DocumentExtractionClientMetadata`, the `DocumentExtractionClientExtensions` surface (incl. opt-in `ExtractFromUriAsync`), and the full builder / middleware / DI types. The sibling `IDocumentAnalysisClient` is scoped out to its own future proposal.

## Background and motivation

Document parsing is a core RAG and ingestion building block, but there is no provider-neutral
document-extraction capability in the `Microsoft.Extensions.*` stack. Today, you either wire provider SDKs
directly or route OCR through `IChatClient`, which loses native document structure such as tables,
bounding boxes, confidence, polygons, and reading order.

This proposal adds `IDocumentExtractionClient` as a provider-neutral capability in its **own peer
library**, `Microsoft.Extensions.DocumentExtraction` — a sibling to `Microsoft.Extensions.VectorData` and
`Microsoft.Extensions.DataIngestion`, not a capability inside `Microsoft.Extensions.AI`. It adopts the
same builder, middleware, and DI shape developers already use across the Microsoft.Extensions.AI
capability family, and references `Microsoft.Extensions.AI.Abstractions` internally (e.g. `DataContent`
inputs) without carrying the "AI" brand in its own namespace or types. Extraction pulls structured content
*out* of documents; it complements `DataIngestion`, which feeds content *into* RAG.

### Architecture at a glance

```mermaid
flowchart LR
subgraph Providers
DI[Azure Document Intelligence]
M[Mistral OCR]
CU[Content Understanding]
V[Vision-capable IChatClient adapter]
L[Local document model]
end

DI --> C
M --> C
CU --> C
V --> C
L --> C

C[IDocumentExtractionClient] --> D[Document pages, elements, tables, images, and geometry]
D --> MEDI[Microsoft.Extensions.DataIngestion]
D --> RAG[RAG and indexing pipelines]
D --> APP[Direct application consumers]
```

The client is the provider-neutral seam. The `Document*` types are the structured exchange model;
ingestion is one consumer of that model, not its owner or a required runtime.

### What changed after the July review

| Review concern | Evidence gathered | Current proposed design | Review status |
|---|---|---|---|
| OCR was too narrow and the M.E.AI assembly might be the wrong home | Shared-model and Azure Document Intelligence implementation spikes | Standalone `Microsoft.Extensions.DocumentExtraction(.Abstractions)` peer library | **Confirmation requested** |
| Parallel blocks/tables/images lose reading order | 12-engine output-model survey and MEDI consumer spike | One ordered `DocumentPage.Elements` list over `DocumentElement` | Implemented and evidence-validated |
| Close extensibility where the domain is bounded | Provider taxonomy and coordinate survey | Closed unit/origin enums; open block/cell kind structs | Implemented and evidence-validated |
| A document probably uses one coordinate unit | Google and Azure DI model units per page, including mixed image/PDF inputs | Dimensions, unit, and origin live on each `DocumentPage` | Implemented and evidence-validated |
| A shared document model needs an ownership story | The current MEDI bridge drops tables and typed geometry | Neutral `Document*` model lives here for v1; a later hoist remains additive | **Confirmation requested** |
| Escape hatches must earn their surface | Vision-adapter, metadata, SDK-escape, and 14-engine raw-output investigations | Keep `GetService`, `RawRepresentation`, and `AdditionalProperties` while experimental | **Confirmation requested** |

### What is OCR / document AI?

OCR is the process of extracting text from documents and images. Document AI goes further: it keeps
structure around that text, including pages, tables, blocks, regions, confidence scores, and reading
order.

For RAG and ingestion pipelines, that structure matters. A document reader should not only produce
markdown. It should also preserve enough page, region, table, confidence, and source metadata for
downstream chunking, retrieval, grounding, and evaluation.

### Why an abstraction?

`Microsoft.Extensions.AI.Abstractions` ships a family of capability interfaces: `IChatClient`,
`IEmbeddingGenerator`, `ISpeechToTextClient`, `ITextToSpeechClient`, `IImageGenerator`,
`IRealtimeClient`, and `IHostedFileClient`. There is **no OCR / document-extraction capability**, even
though `Microsoft.Extensions.DataIngestion` (MEDI) already depends on document parsing. Its
`IngestionDocumentReader` roadmap, per MS Learn, includes LlamaParse and Azure Document Intelligence,
both hosted document-AI services that need a provider-agnostic seam.

Today, if you want to use document-AI models, you must:

1. Couple ingestion code directly to a provider SDK.
2. Model OCR as a chat prompt against `IChatClient`.
3. Add reader-mode flags for specific engines or hosts.
4. Rebuild retry, logging, DI, middleware, and test seams per provider.
5. Give up native structure when the abstraction cannot represent it.

[CommunityToolkit/AI #3](https://github.com/CommunityToolkit/AI/issues/3) is a representative example.
It introduced a `PdfReadingMode.VisionOnly` flag that routes whole-document transcription through a
vision LLM (`IChatClient`). That is a layer leak: a *model choice* hardened into a *reader-mode flag*,
with temporal coupling because the reader emits placeholders that are useless unless a specific enricher
runs. The cleaner shape is a capability **client** the reader composes, exactly how MEDI's enrichers
already compose an injected `IChatClient`.

**OCR is not chat.** Most OCR / document-AI engines emit structured output: tables, bounding boxes,
confidence, polygons, reading order. That does not fit `ChatResponse`. A vision LLM *can* transcribe by
prompt, but it is the lowest-fidelity path and loses native structure. Purpose-built engines (Mistral
OCR, Azure Document Intelligence, Azure AI Content Understanding) and local document VLMs
(granite-docling, PaddleOCR) beat it. Modeling OCR as "call `IChatClient` with an image" makes those
engines unrepresentable without discarding their value. That points to a **separate capability
interface**, independent of `IChatClient`.

### Prototype validation

This proposal is **not a sketch**. It describes a working prototype spiked across **four real engine
providers** (three using no `IChatClient` at all) plus one vision-LLM adapter, composed with an
`IChatClient`-style builder pipeline, and validated end-to-end through a MEDI ingestion pipeline:

- `FoundryMistralDocumentExtractionClient`: Azure AI Foundry `mistral-ocr-4-0`, keyless Entra (verified HTTP 200).
- `MistralDocumentExtractionClient`: Mistral-direct, API key.
- `AzureDocumentIntelligenceClient`: `Azure.AI.DocumentIntelligence` (`AnalyzeResult`, native polygons + table cells).
- `ContentUnderstandingClient`: `Azure.AI.ContentUnderstanding` 1.1.0, keyless Entra (markdown path).
- `VisionLlmDocumentExtractionClient`: the **one** adapter over `IChatClient` (gpt-4o / Gemini / local Ollama GLM-OCR), the lowest-fidelity path.

Every claim below ("one pipeline wraps all engines", "the polygon flows losslessly from DI and Mistral",
"streaming yields pages as they finish") is backed by code that **builds and runs**, not by assertion. The demo
is a public, runnable proof: one `IDocumentExtractionClient` in front of four OCR engines, bridged into a MEDI RAG
pipeline, with the identical consumer loop across both provider archetypes (document-native and
image-per-page).

The design goal is **provider-neutrality**: one small set of composable primitives that every provider
maps onto equally, judged on **interoperability, reusability, composition, extensibility**. No provider
is privileged. Providers form a **coverage matrix, not a hierarchy**.

The same precedent that justifies splitting `OpenAIClient` / `AzureOpenAIClient` behind `IChatClient`
applies here. The **interface** is the portability guarantee; concrete classes split by **engine** and by
**host** where credential, route, or provider behavior leaks. The model/deployment id is a parameter,
never a type or a boolean flag.

Building-block symmetry with Microsoft.Extensions.AI

The goal is not a one-off OCR helper, and not a new pattern to learn. `Microsoft.Extensions.DocumentExtraction`
is a **peer library** that mirrors the M.E.AI building-block shape — exactly as `VectorData` and
`DataIngestion` do — rather than a capability living inside M.E.AI.

If you know one Microsoft.Extensions.AI capability, you should know this one: abstraction,
options/result types, delegating base, builder/middleware, provider implementation, and DI registration.
The row below shows the shape it adopts.

| Capability | Abstraction | Options / result types | Delegating base | Builder / middleware | Example implementation | DI registration shape |
|---|---|---|---|---|---|---|
| Chat | `IChatClient` | `ChatOptions`, `ChatResponse`, `ChatResponseUpdate` | `DelegatingChatClient` | `ChatClientBuilder`, `.Use(...)`, logging, OpenTelemetry, caching, function invocation | `OpenAIChatClient` | `AddChatClient`, `AddKeyedChatClient` |
| Embeddings | `IEmbeddingGenerator` | `EmbeddingGenerationOptions`, `GeneratedEmbeddings` | `DelegatingEmbeddingGenerator` | `EmbeddingGeneratorBuilder`, `.Use(...)`, logging, OpenTelemetry, caching | `OpenAIEmbeddingGenerator` | `AddEmbeddingGenerator`, `AddKeyedEmbeddingGenerator` |
| Speech-to-text | `ISpeechToTextClient` | `SpeechToTextOptions`, `SpeechToTextResponse`, response updates | `DelegatingSpeechToTextClient` | `SpeechToTextClientBuilder`, `.Use(...)`, logging, OpenTelemetry, options | `OpenAISpeechToTextClient` | `AddSpeechToTextClient`, `AddKeyedSpeechToTextClient` |
| Text-to-speech | `ITextToSpeechClient` | `TextToSpeechOptions`, `TextToSpeechResponse`, response updates | `DelegatingTextToSpeechClient` | `TextToSpeechClientBuilder`, `.Use(...)`, logging, OpenTelemetry, options | `OpenAITextToSpeechClient` | `AddTextToSpeechClient`, `AddKeyedTextToSpeechClient` |
| Images | `IImageGenerator` | `ImageGenerationOptions`, `ImageGenerationRequest`, `ImageGenerationResponse` | `DelegatingImageGenerator` | `ImageGeneratorBuilder`, `.Use(...)`, logging, options | `OpenAIImageGenerator` | `AddImageGenerator`, `AddKeyedImageGenerator` |
| Realtime | `IRealtimeClient` | `RealtimeSessionOptions`, client/server messages, sessions | `DelegatingRealtimeClient` | `RealtimeClientBuilder`, `.Use(...)`, logging, OpenTelemetry, function invocation | `OpenAIRealtimeClient` | Register `IRealtimeClient` / keyed clients through DI |
| Hosted files | `IHostedFileClient` | `HostedFileClientOptions`, `HostedFileDownloadStream` | `DelegatingHostedFileClient` | `HostedFileClientBuilder`, `.Use(...)`, logging, OpenTelemetry | `OpenAIHostedFileClient` | Register `IHostedFileClient` / keyed clients through DI |
| OCR / document extraction | `IDocumentExtractionClient` | `DocumentExtractionOptions`, `DocumentExtractionResult`, `DocumentExtractionPageResult`, `DocumentPage`, `DocumentBlock`, `DocumentTable`, `DocumentImage`, `DocumentExtractionUsage` | `DelegatingDocumentExtractionClient` | `DocumentExtractionClientBuilder`, `.Use(...)`, logging, OpenTelemetry, configure-options | `FoundryMistralDocumentExtractionClient`, `MistralDocumentExtractionClient`, `AzureDocumentIntelligenceClient`, `ContentUnderstandingClient`, `VisionLlmDocumentExtractionClient` | `AddDocumentExtractionClient`, `AddKeyedDocumentExtractionClient` |

That symmetry is the main API shape. `IDocumentExtractionClient` should feel like a natural next capability, not a
separate pattern you have to relearn.

### Provider coverage

Providers are peers behind a provider-neutral contract, not a tier or hierarchy.

| Provider | `IDocumentExtractionClient` (markdown/structure) | `IDocumentAnalysisClient` (typed fields + grounding) — *future sibling proposal, not in this PR* |
|---|---|---|
| Foundry Mistral OCR | yes | — |
| Azure Document Intelligence | yes | yes (`Documents[].Fields`) |
| Content Understanding | yes (markdown path) | yes (`fields{}` + grounding) |
| Vision-LLM adapter | yes (lowest fidelity) | — |
| Local ONNX / Ollama (roadmap) | yes | — |

No row is privileged. Some providers implement more of the family than others; the **family** is the
design, and **coverage** is a matrix. Content Understanding is the *widest-surface conformance test* (one
service exercises **both** interfaces with the same primitives), **not** an apex; it validates
provider-neutrality because the same polygon / confidence / builder primitives serve its two shapes,
Mistral OCR, Azure DI, and a vision LLM. The second column previews a future sibling capability
(`IDocumentAnalysisClient`, see *Related and future work*) and is shown only to illustrate that the same
region / confidence / builder primitives generalize; **it is not part of this PR**.

## API Proposal

The surface below is the current proposed public API on PR #7588 (32 public types), reshaped after the
2026-07-22 API review and a 12-engine provider survey. It is implemented and evidence-validated, but
remains subject to API review. Signatures only, no method bodies.

| Surface area | Purpose |
|---|---|
| `IDocumentExtractionClient`, options, result, page result, usage | Unary and page-streaming extraction contract |
| `DocumentPage` and `DocumentElement` hierarchy | Provider-neutral reading-order document model |
| Tables, cells, images, regions, points, dimensions, units, origins | Structured output and interpretable page geometry |
| Delegating client, builder, extensions, logging, OpenTelemetry, options, DI | Standard composable capability plumbing |

### Core abstraction, options, and result types (`Microsoft.Extensions.DocumentExtraction.Abstractions`)

```csharp
namespace Microsoft.Extensions.DocumentExtraction;

///
/// A capability for OCR / document-extraction engines. Independent of :
/// engines emit structured output (tables, bounding boxes, confidence, reading order) that does not
/// fit a chat response. One contract, many engines (Mistral OCR, Azure Document Intelligence, Content
/// Understanding, a local ONNX model, or a vision LLM behind an adapter).
///
[Experimental("MEDE0001")]
public interface IDocumentExtractionClient : IDisposable
{
/// Runs OCR / document parsing over a document stream and returns structured text + pages.
Task ExtractAsync(
Stream document, string mediaType, DocumentExtractionOptions? options = null,
CancellationToken cancellationToken = default);

///
/// Streams OCR / document parsing as values — one per page as it finishes
/// (the twin). Reassemble into an
/// via . Lets
/// large-document RAG chunk/embed early pages while later pages are still being parsed.
///
IAsyncEnumerable ExtractPagesAsync(
Stream document, string mediaType, DocumentExtractionOptions? options = null,
CancellationToken cancellationToken = default);

/// Provider escape hatch (the pattern).
object? GetService(Type serviceType, object? serviceKey = null);
}

/// Normalized OCR result — "normalize the common, preserve the raw" (the ChatResponse pattern).
[Experimental("MEDE0001")]
public class DocumentExtractionResult
{
public DocumentExtractionResult(IReadOnlyList pages);

public IReadOnlyList Pages { get; }
public string Text { get; } // derived: page Text joined with blank lines
public DocumentExtractionUsage? Usage { get; set; }
public object? RawRepresentation { get; set; } // provider-native object — nothing is lost
public AdditionalPropertiesDictionary? AdditionalProperties { get; set; }
}

[Experimental("MEDE0001")]
public class DocumentPage
{
public DocumentPage(int pageNumber, string text);

public int PageNumber { get; } // 1-based
public string Text { get; }
public IReadOnlyList Elements { get; set; } // READING ORDER; OfType() to project. default: empty
public DocumentPageDimensions? Dimensions { get; set; } // page extent (width+height), when the engine reports it
public DocumentCoordinateUnit? CoordinateUnit { get; set; } // per page — engines emit different units per page (image vs PDF batches)
public DocumentCoordinateOrigin? CoordinateOrigin { get; set; }
[JsonIgnore] public object? RawRepresentation { get; set; } // provider-native page object; survives ToDocumentExtractionResult reduction (SPIKE-06)
public AdditionalPropertiesDictionary? AdditionalProperties { get; set; }
}

///
/// The reading-order element base — a polymorphic ($type) shape (the AIContent pattern), designed to be
/// promotable to a future shared document-element type. DocumentBlock / DocumentTable / DocumentImage derive from it.
///
[Experimental("MEDE0001")]
[JsonPolymorphic(TypeDiscriminatorPropertyName = "$type")]
[JsonDerivedType(typeof(DocumentBlock), "block")]
[JsonDerivedType(typeof(DocumentTable), "table")]
[JsonDerivedType(typeof(DocumentImage), "image")]
public abstract class DocumentElement
{
protected DocumentElement();

public DocumentBoundingRegion? BoundingRegion { get; set; }
public double? Confidence { get; set; }
[JsonIgnore] public object? RawRepresentation { get; set; }
public AdditionalPropertiesDictionary? AdditionalProperties { get; set; }
}

[Experimental("MEDE0001")]
public class DocumentBlock : DocumentElement
{
public DocumentBlock(string text);

public string Text { get; }
public DocumentBlockKind? Kind { get; set; } // open struct: Paragraph / Title / Figure / ...
// BoundingRegion, Confidence inherited from DocumentElement
}

/// An image or figure extracted from a page (emitted normally; no request flag).
[Experimental("MEDE0001")]
public class DocumentImage : DocumentElement
{
public DataContent? Content { get; set; }
public string? Caption { get; set; }
// BoundingRegion, Confidence inherited from DocumentElement
}

/// A single 2-D point in page coordinates (a polygon vertex).
[Experimental("MEDE0001")]
public readonly record struct DocumentPoint(float X, float Y);

/// An axis-aligned bounding box, for coarse filters / hit-testing.
[Experimental("MEDE0001")]
public readonly record struct DocumentBoundingBox(float Left, float Top, float Right, float Bottom);

///
/// The SHARED, provider-neutral geometry primitive — a polygon of DocumentPoint vertices (populated natively by
/// Azure DI, via FromRectangle for Mistral's rect, and reused for field grounding). GetBounds() gives a coarse box.
///
[Experimental("MEDE0001")]
public class DocumentBoundingRegion
{
public DocumentBoundingRegion(int pageNumber, IReadOnlyList polygon);

public int PageNumber { get; }
public IReadOnlyList Polygon { get; }
public static DocumentBoundingRegion FromRectangle(
int pageNumber, float left, float top, float right, float bottom); // float (was double)
public DocumentBoundingBox? GetBounds();
}

/// Cells are the primary structured representation; MarkdownRepresentation is the fallback (Mistral).
[Experimental("MEDE0001")]
public class DocumentTable : DocumentElement
{
public DocumentTable(
int rowCount, int columnCount,
IReadOnlyList? cells = null, string? markdownRepresentation = null);

public int RowCount { get; }
public int ColumnCount { get; }
public IReadOnlyList? Cells { get; }
public string? MarkdownRepresentation { get; }
// BoundingRegion, Confidence inherited from DocumentElement
}

[Experimental("MEDE0001")]
public class DocumentTableCell
{
public DocumentTableCell(int rowIndex, int columnIndex, string content);

public DocumentTableCellKind? Kind { get; set; } // open struct: ColumnHeader / Content / RowHeader / RowSection
public int RowIndex { get; }
public int ColumnIndex { get; }
public int RowSpan { get; set; } // default 1
public int ColumnSpan { get; set; } // default 1
public string Content { get; } // flat-text convenience (kept)
public IReadOnlyList? Elements { get; set; } // optional NESTED content (structured cells)
// Positioned-node facet mirrored from DocumentElement (NOT inheritance; reversible way-station, SPIKE-06):
public DocumentBoundingRegion? BoundingRegion { get; set; } // per-cell geometry (5 engines: Textract/Google/DI/Adobe/Docling)
public double? Confidence { get; set; }
[JsonIgnore] public object? RawRepresentation { get; set; }
public AdditionalPropertiesDictionary? AdditionalProperties { get; set; }
}

///
/// A streamed OCR update — one completed page (the ChatResponseUpdate pattern). Reduce a sequence of these
/// into an DocumentExtractionResult with DocumentExtractionPageResultExtensions. Page is non-null (no sentinel update).
///
[Experimental("MEDE0001")]
public class DocumentExtractionPageResult
{
[JsonConstructor] public DocumentExtractionPageResult(DocumentPage page); // Page is non-null (no sentinel update)

public DocumentPage Page { get; }
public int? PagesProcessed { get; set; } // progress (absorbs the retired DocumentExtractionProgress)
public int? TotalPages { get; set; }
public DocumentExtractionUsage? Usage { get; set; }
[JsonIgnore] public object? RawRepresentation { get; set; }
public AdditionalPropertiesDictionary? AdditionalProperties { get; set; }
}

/// Reducers that assemble streamed page results back into one DocumentExtractionResult (the ToChatResponseAsync pattern).
[Experimental("MEDE0001")]
public static class DocumentExtractionPageResultExtensions
{
public static DocumentExtractionResult ToDocumentExtractionResult(this IEnumerable updates);
public static Task ToDocumentExtractionResultAsync(
this IAsyncEnumerable updates, CancellationToken cancellationToken = default);
}

[Experimental("MEDE0001")]
public class DocumentExtractionUsage
{
public int? PagesProcessed { get; set; }
public int? InputTokenCount { get; set; } // vision-LLM path; classic OCR leaves these null
public int? OutputTokenCount { get; set; }
public int? TotalTokenCount { get; set; }
public AdditionalPropertiesDictionary? AdditionalProperties { get; set; }
}

/// The kind of a text block — a ChatRole-style OPEN set (well-knowns + provider-specific kinds).
[Experimental("MEDE0001")]
public readonly struct DocumentBlockKind : IEquatable
{
public DocumentBlockKind(string value); // throws on null/whitespace
public static DocumentBlockKind Paragraph { get; } // "paragraph"
public static DocumentBlockKind Title { get; } // "title"
public static DocumentBlockKind Figure { get; } // "figure"
public string Value { get; }
// == / != / IEquatable / GetHashCode / ToString + a JsonConverter (the ChatRole shape)
}

/// The kind of a table cell — a ChatRole-style OPEN set.
[Experimental("MEDE0001")]
public readonly struct DocumentTableCellKind : IEquatable
{
public DocumentTableCellKind(string value);
public static DocumentTableCellKind ColumnHeader { get; } // "columnHeader"
public static DocumentTableCellKind Content { get; } // "content"
public static DocumentTableCellKind RowHeader { get; } // "rowHeader" (added)
public static DocumentTableCellKind RowSection { get; } // "rowSection" (added)
public string Value { get; }
}

/// The unit for page dimensions + region coordinates — a CLOSED enum (units are physically bounded).
[Experimental("MEDE0001")]
public enum DocumentCoordinateUnit { Pixel, Point, Inch, Normalized }

/// Origin corner + y-axis direction of the coordinate space — a CLOSED enum.
[Experimental("MEDE0001")]
public enum DocumentCoordinateOrigin { TopLeft, BottomLeft }

/// Page extent (width + height), expressed in the page's DocumentCoordinateUnit — a readonly record struct (atomic pair).
[Experimental("MEDE0001")]
public readonly record struct DocumentPageDimensions(float Width, float Height);

/// Request knobs — the ChatOptions pattern.
[Experimental("MEDE0001")]
public class DocumentExtractionOptions
{
public string? ModelId { get; set; } // "GetChatClient(model)" analog
public AdditionalPropertiesDictionary? AdditionalProperties { get; set; }
public DocumentExtractionOptions Clone(); // shallow clone (the ChatOptions.Clone pattern)
}

/// Metadata about an (the *ClientMetadata pattern).
[Experimental("MEDE0001")]
public class DocumentExtractionClientMetadata
{
public DocumentExtractionClientMetadata(string? providerName = null, Uri? providerUri = null, string? defaultModelId = null);
public string? ProviderName { get; }
public Uri? ProviderUri { get; }
public string? DefaultModelId { get; }
}
```

### Extension methods (`DocumentExtractionClientExtensions`)

```csharp
/// Convenience helpers over .
[Experimental("MEDE0001")]
public static class DocumentExtractionClientExtensions
{
public static TService? GetService(this IDocumentExtractionClient client, object? serviceKey = null);

// Extract from an in-memory DataContent (unary + streaming twin).
public static Task ExtractAsync(
this IDocumentExtractionClient client, DataContent document, DocumentExtractionOptions? options = null,
CancellationToken cancellationToken = default);
public static IAsyncEnumerable ExtractPagesAsync(
this IDocumentExtractionClient client, DataContent document, DocumentExtractionOptions? options = null,
CancellationToken cancellationToken = default);

// Extract from a UriContent. Handles self-contained data: URIs; throws NotSupportedException for
// file:/http(s) (whether to download vs. hand the URL to the engine is a deliberate non-decision).
public static Task ExtractAsync(
this IDocumentExtractionClient client, UriContent document, DocumentExtractionOptions? options = null,
CancellationToken cancellationToken = default);
public static IAsyncEnumerable ExtractPagesAsync(
this IDocumentExtractionClient client, UriContent document, DocumentExtractionOptions? options = null,
CancellationToken cancellationToken = default);

// Explicit, opt-in remote downloader: fetches http(s) bytes with a caller-supplied HttpClient
// (caller owns handlers/auth/timeouts/lifetime), inlines data: URIs, then extracts.
public static Task ExtractFromUriAsync(
this IDocumentExtractionClient client, UriContent document, HttpClient httpClient, DocumentExtractionOptions? options = null,
CancellationToken cancellationToken = default);
}
```

### Delegating base, builder, middleware, and DI

All Microsoft.Extensions.AI capabilities ship the same five-layer shape:
`interface` → `Delegating` → `Builder` → `Add` (returns the builder) → `.Use*()`
middleware, with one composition primitive, `Builder Use(Func)`. `IDocumentExtractionClient`
mirrors it exactly.

```csharp
// ---- Microsoft.Extensions.DocumentExtraction.Abstractions ----

/// Optional base for an that passes calls through to an inner instance.
[Experimental("MEDE0001")]
public class DelegatingDocumentExtractionClient : IDocumentExtractionClient
{
protected DelegatingDocumentExtractionClient(IDocumentExtractionClient innerClient);
protected IDocumentExtractionClient InnerClient { get; }

public void Dispose();
public virtual Task ExtractAsync(
Stream document, string mediaType, DocumentExtractionOptions? options = null,
CancellationToken cancellationToken = default);
public virtual IAsyncEnumerable ExtractPagesAsync(
Stream document, string mediaType, DocumentExtractionOptions? options = null,
CancellationToken cancellationToken = default);
public virtual object? GetService(Type serviceType, object? serviceKey = null);
protected virtual void Dispose(bool disposing);
}

// ---- Microsoft.Extensions.DocumentExtraction ----

[Experimental("MEDE0001")]
public sealed class DocumentExtractionClientBuilder
{
public DocumentExtractionClientBuilder(IDocumentExtractionClient innerClient);
public DocumentExtractionClientBuilder(Func innerClientFactory);
public IDocumentExtractionClient Build(IServiceProvider? services = null); // first .Use is outermost
public DocumentExtractionClientBuilder Use(Func clientFactory);
public DocumentExtractionClientBuilder Use(Func clientFactory); // THE primitive
}

[Experimental("MEDE0001")] public class LoggingDocumentExtractionClient : DelegatingDocumentExtractionClient { } // logging middleware
[Experimental("MEDE0001")] public sealed class OpenTelemetryDocumentExtractionClient : DelegatingDocumentExtractionClient { } // OTel middleware
[Experimental("MEDE0001")] public sealed class ConfigureOptionsDocumentExtractionClient : DelegatingDocumentExtractionClient { } // options middleware

[Experimental("MEDE0001")]
public static class DocumentExtractionClientBuilderExtensions
{
public static DocumentExtractionClientBuilder AsBuilder(this IDocumentExtractionClient innerClient);
}

[Experimental("MEDE0001")]
public static class LoggingDocumentExtractionClientBuilderExtensions
{
public static DocumentExtractionClientBuilder UseLogging(
this DocumentExtractionClientBuilder builder, ILoggerFactory? loggerFactory = null, Action? configure = null);
}

[Experimental("MEDE0001")]
public static class OpenTelemetryDocumentExtractionClientBuilderExtensions
{
public static DocumentExtractionClientBuilder UseOpenTelemetry(
this DocumentExtractionClientBuilder builder, ILoggerFactory? loggerFactory = null, string? sourceName = null,
Action? configure = null);
}

[Experimental("MEDE0001")]
public static class ConfigureOptionsDocumentExtractionClientBuilderExtensions
{
public static DocumentExtractionClientBuilder ConfigureOptions(this DocumentExtractionClientBuilder builder, Action configure);
}

[Experimental("MEDE0001")]
public static class DocumentExtractionClientBuilderServiceCollectionExtensions
{
public static DocumentExtractionClientBuilder AddDocumentExtractionClient(
this IServiceCollection serviceCollection, IDocumentExtractionClient innerClient,
ServiceLifetime lifetime = ServiceLifetime.Singleton);
public static DocumentExtractionClientBuilder AddDocumentExtractionClient(
this IServiceCollection serviceCollection, Func innerClientFactory,
ServiceLifetime lifetime = ServiceLifetime.Singleton);
public static DocumentExtractionClientBuilder AddKeyedDocumentExtractionClient(
this IServiceCollection serviceCollection, object? serviceKey, IDocumentExtractionClient innerClient,
ServiceLifetime lifetime = ServiceLifetime.Singleton);
public static DocumentExtractionClientBuilder AddKeyedDocumentExtractionClient(
this IServiceCollection serviceCollection, object? serviceKey, Func innerClientFactory,
ServiceLifetime lifetime = ServiceLifetime.Singleton);
}
```

`IDocumentExtractionClient` ships logging, OpenTelemetry, and configure-options middleware in v1, matching the
`ISpeechToTextClient` template. It does not ship a built-in retry client: no `Microsoft.Extensions.AI`
capability does, because resilience belongs in the HTTP pipeline
(`Microsoft.Extensions.Http.Resilience`). The `.Use(...)` primitive still lets a consumer wrap a custom
retry or cache decorator when they want one.

## API Usage

Every snippet below is drawn from the runnable
[`iocrclient-demo`](https://github.com/luisquintanilla/iocrclient-demo), which exercises this exact
surface across four engines and a MEDI RAG pipeline.

**One interface, four engines — the payoff.** Vision LLM, Mistral OCR, Azure Document Intelligence, and
Azure Content Understanding each speak a completely different wire protocol. Behind `IDocumentExtractionClient` they are
`IDocumentExtractionClient`; the consumer loop never changes when you add or swap a provider.

```csharp
var clients = new (string Name, IDocumentExtractionClient Client)[]
{
("vision-llm", new VisionLlmDocumentExtractionClient(chatClient)),
("mistral-ocr", new FoundryMistralDocumentExtractionClient(foundryEndpoint, cred)),
("azure-document-intelligence", new AzureDocumentIntelligenceClient(diEndpoint, cred)),
("azure-content-understanding", new ContentUnderstandingClient(cuEndpoint, cred)),
};

byte[] bytes = await File.ReadAllBytesAsync("report.pdf");
foreach (var (name, client) in clients)
{
using (client)
{
using var stream = new MemoryStream(bytes, writable: false);
DocumentExtractionResult r = await client.ExtractAsync(stream, "application/pdf"); // identical for every engine

int tables = r.Pages.Sum(p => p.Elements.OfType().Count());
Console.WriteLine($"{name}: {r.Pages.Count} pages, {tables} tables, {r.Text.Length} chars");
}
}
```

**Stream pages as they finish** — the `IChatClient.GetStreamingResponseAsync` twin. Each
`DocumentExtractionPageResult` carries one completed page (plus progress + usage), so a RAG pipeline can chunk and
embed early pages while later pages are still being OCR'd; `ToDocumentExtractionResultAsync` reduces the stream back to
the same `DocumentExtractionResult` the unary call would return.

```csharp
await foreach (DocumentExtractionPageResult update in client.ExtractPagesAsync(stream, "application/pdf"))
{
DocumentPage page = update.Page;
Console.WriteLine($"page {page.PageNumber}/{update.TotalPages}: {page.Text.Length} chars");
}

// …or reduce the whole stream back into one DocumentExtractionResult (the ToChatResponseAsync pattern):
DocumentExtractionResult full = await client.ExtractPagesAsync(stream, "application/pdf").ToDocumentExtractionResultAsync();
```

**Consume elements in reading order.** A page has one ordered stream rather than parallel collections
that require consumers to reconstruct order from geometry.

```csharp
DocumentExtractionResult result = await client.ExtractAsync(stream, "application/pdf");
DocumentPage page = result.Pages[0];

foreach (DocumentElement element in page.Elements)
{
switch (element)
{
case DocumentBlock block:
Console.WriteLine(block.Text);
break;
case DocumentTable table:
RenderTable(table);
break;
case DocumentImage image:
RenderImage(image);
break;
}
}

IEnumerable tables = page.Elements.OfType();
```

**Interpret geometry in the page's own coordinate system.** Unit and origin are page-scoped because
one request can contain image pages measured in pixels and PDF pages measured in inches or points.

```csharp
DocumentExtractionResult result = await client.ExtractAsync(stream, "application/pdf");

foreach (DocumentPage page in result.Pages)
{
if (page.Dimensions is { } size &&
page.CoordinateUnit is { } unit &&
page.CoordinateOrigin is { } origin)
{
Console.WriteLine(
$"Page {page.PageNumber}: {size.Width} x {size.Height} {unit}, origin {origin}");

foreach (DocumentElement element in page.Elements)
{
DocumentBoundingBox? bounds = element.BoundingRegion?.GetBounds();
// Normalize or transform bounds using this page's size, unit, and origin.
}
}
}
```

**Compose middleware with the builder** — the same shape as `ChatClientBuilder`: you compose a client,
you don't set flags.

```csharp
IDocumentExtractionClient ocr = new FoundryMistralDocumentExtractionClient(endpoint, cred)
.AsBuilder()
.UseOpenTelemetry(loggerFactory)
.UseLogging(loggerFactory)
.Build();
```

**Discover optional services without widening the core contract.** Middleware can inspect provider
metadata, and an opaque adapter can expose the client it wraps.

```csharp
static void Inspect(IDocumentExtractionClient client)
{
DocumentExtractionClientMetadata? metadata =
client.GetService();

IChatClient? innerChatClient =
client.GetService(); // non-null when an adapter chooses to expose it
}
```

**Register with dependency injection** — the consumer depends only on `IDocumentExtractionClient`; swap the engine
without touching downstream code.

```csharp
services.AddDocumentExtractionClient(sp => new FoundryMistralDocumentExtractionClient(endpoint, new DefaultAzureCredential()))
.UseOpenTelemetry()
.UseLogging();

// Later, swap the engine on one line — nothing downstream changes:
services.AddDocumentExtractionClient(sp => new AzureDocumentIntelligenceClient(diEndpoint, cred)).UseLogging();
```

**Configure per-request options through the pipeline** (for example, a local Ollama GLM-OCR engine that
needs a task-prefix prompt injected even when the caller passes no options):

```csharp
IDocumentExtractionClient ocr = engine
.AsBuilder()
.ConfigureOptions(o => o.ModelId ??= "mistral-ocr-4-0")
.Build();
```

**Bridge into a MEDI ingestion / RAG pipeline.** A hosted document-AI service is a *reader* (the
LlamaParse-as-reader shape). One provider-agnostic reader composes any `IDocumentExtractionClient`:

```csharp
public sealed class OcrDocumentReader(IDocumentExtractionClient ocr, DocumentExtractionOptions? options = null) : IngestionDocumentReader
{
public override async Task ReadAsync(
Stream source, string identifier, string mediaType, CancellationToken ct = default)
{
DocumentExtractionResult r = await ocr.ExtractAsync(source, mediaType, options, ct);
// map r.Pages -> IngestionDocument elements, stamping page/region/confidence/model metadata
}
}

// Usage: one section per OCR page, each stamped with its 1-based PageNumber for downstream chunking.
var reader = new OcrDocumentReader(ocr);
IngestionDocument doc = await reader.ReadAsync(fileStream, "report.pdf", "application/pdf");
```

**Extract from a remote URI (opt-in download).** `ExtractAsync(UriContent)` never touches the network;
`ExtractFromUriAsync` is the explicit counterpart that fetches http(s) bytes with a caller-owned
`HttpClient`:

```csharp
using var http = new HttpClient();
DocumentExtractionResult r = await ocr.ExtractFromUriAsync(new UriContent(url, "application/pdf"), http);
```

## Alternative Designs

1. **Reuse `IChatClient` with multimodal content + a prompt.** Rejected: loses native
tables/bbox/confidence, is nondeterministic and token-expensive, and cannot represent non-chat engines
(Azure DI, CU, local ONNX) at all. A vision LLM is supported as *one provider behind* `IDocumentExtractionClient`
(`VisionLlmDocumentExtractionClient`), not as the contract. **The prototype demonstrates this**: three of four real
engines use no `IChatClient`.
2. **A flag on the reader (the `VisionOnly` approach).** Rejected: a model choice hardened into a reader
mode; temporal coupling; not decoratable (no middleware); does not generalize across engines.
3. **A transport-parameterized single class (`new MistralDocumentExtractionClient(isFoundry: true)`).** Rejected:
smuggles host branching into the type; breaks composition. Follow the `OpenAIClient`/`AzureOpenAIClient`
precedent: the interface is portable; concrete classes split by host where the host leaks.
4. **One interface for OCR *and* field extraction (a flag/overload).** Rejected: different output
contract; modeled as the sibling `IDocumentAnalysisClient` peer (the STT/TTS precedent). See *Future sibling: `IDocumentAnalysisClient`* below.

### Validated design decisions and evidence

The table distinguishes implementation and evidence from review approval. “Implemented and
evidence-validated” means the shape exists in #7588 and survived the provider/prototype work; it does
not mean the API has already been approved.

| Decision | Evidence | Current proposed resolution | Status |
|---|---|---|---|
| Structured geometry | Azure DI emits rotation/skew-capable polygons; Mistral emits rectangles | Typed `DocumentPoint` polygon with `GetBounds()` for coarse boxes | Implemented and evidence-validated |
| Structured tables | Azure DI exposes cells while Mistral can provide a markdown fallback | `DocumentTable` with cells plus optional markdown representation | Implemented and evidence-validated |
| Page identity | Documents, Azure DI, and downstream provenance use one-based pages | `DocumentPage.PageNumber` is one-based | Implemented and evidence-validated |
| Unary and streaming pair | Large documents benefit from early-page processing; target frameworks prevent adding a default interface member later | `ExtractAsync` plus page-at-a-time `ExtractPagesAsync`; reducers rebuild a full result | Implemented and evidence-validated |
| Reading order | Surveyed engines emit a heterogeneous ordered stream; MEDI consumers otherwise reconstruct order from geometry | One `DocumentPage.Elements` list over `DocumentElement`, including nested cell elements | Implemented and evidence-validated |
| Open versus closed vocabularies | Coordinate units/origins are physically bounded; provider block/cell taxonomies are large and evolving | Closed coordinate enums; open `DocumentBlockKind` and `DocumentTableCellKind` structs | Implemented and evidence-validated |
| Coordinate ownership | Google and Azure DI report dimensions/unit per page; mixed image/PDF inputs can use different units | `Dimensions`, `CoordinateUnit`, and `CoordinateOrigin` live on each page | Implemented and evidence-validated |
| Text naming | Many engines return plain text rather than genuine Markdown | `Text` is the common property; Markdown can be additive later | Implemented and evidence-validated |
| Usage | Vision adapters report tokens; classic engines generally report page progress | Optional token counts in `DocumentExtractionUsage`; page progress stays on page results | Implemented and evidence-validated |
| Package and naming family | Shared-model and Azure DI spikes both pointed away from the M.E.AI assembly | Standalone `DocumentExtraction` peer library; neutral `Document*` model and `DocumentExtraction*` operations | **Confirmation requested** |
| Shared-model home | The existing MEDI bridge drops tables and typed geometry | Keep the neutral model here for v1; preserve an additive future hoist | **Confirmation requested** |
| Escape hatches | Adapter unwrapping, metadata discovery, provider SDK access, and a 14-engine raw-output inventory | Keep `GetService`, `RawRepresentation`, and `AdditionalProperties` while experimental | **Confirmation requested** |

### Decisions requested in this review

1. **Library boundary and name:** Confirm
`Microsoft.Extensions.DocumentExtraction(.Abstractions)` as the standalone peer library, or name
the specific replacement.
2. **Model home:** Confirm that the neutral `Document*` model can live in this package for v1, with a
later namespace/package hoist remaining additive.
3. **Escape hatches:** Confirm `GetService`, `RawRepresentation`, and `AdditionalProperties`, or
identify the unsupported scenario or replacement API for each removed seam.
4. **V1 composition surface:** Confirm builder, delegating client, DI, logging, OpenTelemetry, and
configure-options; the recommendation is to defer built-in caching and retry policy.
5. **Collection mutability:** Confirm settable `IReadOnlyList` properties for provider population, or
request a constructor/init-only alternative compatible with the supported target frameworks.
6. **Landing gate:** Identify any remaining blocking evidence or signature change before formal
approval and #7588 review.

### Deferred follow-ups, not part of this proposal

These items are intentionally separated from the current API decision. Reviewers can promote one to a
v1 blocker, but the proposal does not assume that on their behalf.

| Follow-up | Current disposition |
|---|---|
| `IChatClient` vision-capability metadata | Cross-team M.E.AI discussion; the document-extraction contract does not depend on the outcome |
| New C# extension-member syntax | Verify repository language, analyzer, and public-API baseline support before adoption |
| Distributed caching | Proposed defer; compose later through the existing builder seam when a concrete policy is ready |
| Retry | Provider/transport concern handled through HTTP resilience or a custom delegating client |
| Multi-image or batch-page input | Additive follow-up; callers can render/split upstream and invoke once per image today |
| Consumer-ergonomics sample beyond this issue | Add a runnable reading-order consumer, but do not block the public contract on another sample |
| Shared model hoist | Revisit when another ready capability and owner require a common package |
| Azure Document Intelligence provider package | Provider-team-owned satellite implementation; direct implementation is already proven feasible |
| Page versus chunk | Page is the v1 streaming unit; revisit only if a future shared model establishes a different stable unit |
| Region primitive name | Keep `DocumentBoundingRegion` for v1; reconsider a more neutral name only with a second capability |
| `IDocumentAnalysisClient` | Separate proposal because schema-based field extraction has a different output contract |

### Future sibling: `IDocumentAnalysisClient`

Field extraction takes a document plus a schema/analyzer and returns typed fields with confidence and
grounding. That is categorically different from extracting a provider-neutral document structure, so
it should be a peer interface rather than a mode, flag, or overload on `IDocumentExtractionClient`.

```csharp
[Experimental("MEDE0001")]
public interface IDocumentAnalysisClient
{
Task AnalyzeAsync(
Stream document,
string mediaType,
DocumentAnalysisOptions options,
CancellationToken cancellationToken = default);

object? GetService(Type serviceType, object? serviceKey = null);
}
```

The future sibling may reuse `DocumentBoundingRegion` and the builder pattern, but its schema,
typed-field value model, and lifecycle require their own proposal and provider matrix.

### Review provenance

How the July review and follow-up spikes produced the current shape

| Review or spike | Outcome in the current proposal |
|---|---|
| Rename streaming operation and page result | `ExtractPagesAsync` and `DocumentExtractionPageResult` |
| Replace parallel content lists | Ordered `DocumentPage.Elements` hierarchy |
| Shared document-model investigation | Standalone peer library with a neutral `Document*` model |
| Nested table-cell content | Optional `DocumentTableCell.Elements` |
| Vision adapter and `GetService` investigation | Keep one optional service/unwrap seam |
| Raw-output investigation across 14 engines | Keep raw/properties; selectively promote repeated common fields |
| Page dimensions and coordinate survey | Per-page dimensions, unit, and origin |
| BCL geometry investigation | Keep document-owned polygon primitives; no cross-platform BCL polygon fits |
| Azure Document Intelligence adapter spike | Direct implementation proven; provider package remains separate |
| Family rename | `Ocr*` becomes `Document*` / `DocumentExtraction*` after leaving M.E.AI |

The full SPIKE-01..13 ledger, provider survey, raw-output inventory, and ADRs remain the durable
provenance record. They should be linked from the review handoff where those artifacts are published,
rather than reproduced line-by-line in this issue.

### References

- [CommunityToolkit/AI #3](https://github.com/CommunityToolkit/AI/issues/3): document-processing
packages and the `VisionOnly` reader flag this proposal replaces.
- [`area-data-ingestion` IngestionPipeline (#7488)](https://github.com/dotnet/extensions/issues/7488):
a downstream consumer of a provider-neutral document reader.
- [STT/TTS split (#5838)](https://github.com/dotnet/extensions/issues/5838) and
[`IVideoGenerator` proposal (#7420)](https://github.com/dotnet/extensions/issues/7420):
capability-family precedents.
- #7588: the implementation proof for this proposal.
- [`iocrclient-demo`](https://github.com/luisquintanilla/iocrclient-demo): runnable provider and MEDI
coverage prototype.

## Risks

| Risk | Mitigation / review question |
|---|---|
| **Permanent surface area:** 32 public types are a meaningful compatibility commitment | The surface is `[Experimental]`; retain a type only when a demonstrated cross-provider scenario requires it |
| **Shape creep:** extraction could absorb analysis, conversion, batching, caching, or provider policy | Keep field analysis in a sibling proposal and keep orchestration/policy at the edges |
| **Package/model ownership:** evidence supports the standalone home, but governance has not been explicitly confirmed | Make package name and v1 model home explicit review decisions |
| **Escape-hatch misuse:** raw objects and additional properties can become an alternative untyped API | Normalize repeated common concepts; reserve the hatches for provider-specific data and document that guidance |
| **Mutable result collections:** settable `IReadOnlyList` properties aid provider construction but may surprise consumers | Decide in review whether setters, init-only properties, or constructor population best fit the target frameworks |
| **Provider input differences:** document-native and image-per-page engines accept different input strategies | Keep one document stream as the narrow contract; leave rendering, batching, and windowing to additive edge APIs |
| **Middleware scope:** caching or resilience can expand the first release without proving the core capability | Ship standard composition seams; defer policy-specific middleware until a concrete scenario and design are ready |

### Landing criterion

The proposal is ready for the next gate when each confirmation request has an explicit disposition,
every blocking change has an owner, and deferred ideas are not treated as hidden prerequisites.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.