opensearch-project / opensearch-project/sql

[FEATURE] Define a schema-conflict and mismatched-data handling policy for multi-index queries

Open
#5,610 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Java
Stars
176
Forks
229
Avg merge
2d 21h
Merged PRs (30d)
43

Description

Is your feature request related to a problem?

Querying across multiple indices (e.g. a wildcard index pattern like source=logs-*) has no consistent policy for handling two related but distinct problems:

  1. Schema-merge-time conflicts: the same field name has genuinely incompatible mapping types across backing indices (e.g. object in some, scalar text in others). MergeRuleHelper currently resolves conflicts via an ordered rule chain:
    • DeepMergeRule only merges struct-vs-struct sub-properties.
    • TextKeywordConflictRule only resolves scalar-vs-scalar (text/keyword) conflicts.
    • Any other mismatch falls through to LatestRule, which silently picks whichever backing index's mapping was processed last, with no validation, error or coercion.
  2. Execution-time mismatched data: even once a unified schema is settled on, an individual document's actual value can still fail to match the declared type for that field — not because the document is malformed, but because it came from a differently-mapped index (e.g. dynamic mapping drift across backing indices, or one index legitimately storing the field as a plain string while another stores it as a nested object). OpenSearchExprValueFactory blindly casts each document's raw value to the declared type with no defensive handling, so a document from a mismatched-mapping shard throws:
java.lang.ClassCastException: class java.lang.String cannot be cast to class java.util.Map
  at org.opensearch.sql.opensearch.data.utils.ObjectContent.map(ObjectContent.java:81)
  at org.opensearch.sql.opensearch.data.value.OpenSearchExprValueFactory.parseStruct(OpenSearchExprValueFactory.java:376)
  at org.opensearch.sql.opensearch.data.value.OpenSearchExprValueFactory.parse(OpenSearchExprValueFactory.java:218)
  at org.opensearch.sql.opensearch.data.value.OpenSearchExprValueFactory.construct(OpenSearchExprValueFactory.java:189)
  at org.opensearch.sql.opensearch.response.OpenSearchResponse.lambda$handleAggregationResponse$2(OpenSearchResponse.java:263)
  at org.opensearch.sql.opensearch.storage.scan.OpenSearchIndexEnumerator.moveNext(OpenSearchIndexEnumerator.java:139)

What solution would you like?

Solutions generally fall into four strategies. Some apply only at execution time (per document, during row deserialization); others apply at both schema-merge time (before query execution, once per index-pattern resolution) and execution time, depending on where they're implemented. Rather than hardcoding a single strategy, these could be exposed as a configurable mode, similar to Spark's CSV/JSON reader mode option (PERMISSIVE, FAILFAST, DROPMALFORMED), letting users choose the behavior that fits their use case:

Strategy Applies to Mechanism Reference model
Fail-fast Schema-merge time & execution time Reject an unresolvable conflict (schema-merge time) or throw on a per-row mismatch (execution time), instead of silently picking one or crashing uncontrolled Spark (AnalysisException on schema merge; FAILFAST mode for per-record parsing), Trino (fails at split planning), Iceberg (rejects writes with mismatched schema unless mergeSchema is explicitly enabled)
Coercion Schema-merge time & execution time Pick one common declared type across indices (schema-merge time), or convert a mismatched value into the expected type per row (execution time) Spark's Parquet mergeSchema read option (widens compatible types across files); Iceberg's mergeSchema write option (adds missing columns at ingestion time, opt-in via write.spark.accept-any-schema)
Null on mismatch Execution time only When a value doesn't match the expected type, return null for that field/row instead of throwing ES|QL's unsupported column type behavior; Spark's PERMISSIVE mode (default) for malformed field values
Drop mismatched row Execution time only Silently exclude the offending row/document from results entirely Spark's DROPMALFORMED mode

What alternatives have you considered?

  • Schema-on-read with no unified type: the most permissive, but a much larger architectural change inconsistent with OpenSearch SQL/PPL's current Calcite-based single-row-schema design.
  • A VARIANT-like semi-structured type for conflicting fields: would require a new core data type and query syntax (field:path::type), a larger change than extending the existing merge/parsing layers.

Do you have any additional context?

Related work

This gap has been patched twice before, each time narrowly for one specific type pairing rather than as a general policy:

Issue / PR Conflict type Layer fixed Fix
#3625 / #3653 struct vs. struct (differing sub-properties) Schema-merge time Added DeepMergeRule to recursively merge sub-properties from both indices
#4659 text vs. keyword (scalar vs. scalar) Schema-merge time Added TextKeywordConflictRule, forcing _source-based retrieval instead of doc_values
Examples

Minimal reproduction (OpenSearch 3.8.0-SNAPSHOT, main HEAD):

PUT /log-repro-object
{
  "mappings": {
    "properties": {
      "log": { "properties": { "ts": { "type": "date" } } }
    }
  }
}

PUT /log-repro-scalar
{
  "mappings": {
    "properties": {
      "log": { "type": "text" }
    }
  }
}

POST /log-repro-object/_doc
{ "log": { "ts": "2026-07-06T10:00:00Z" }, "path": "/a" }

POST /log-repro-scalar/_doc
{ "log": "plain string log line", "path": "/b" }
POST /_plugins/_ppl
{ "query": "source = log-repro-* | dedup path" }

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading MergeRuleHelper and its DeepMergeRule, TextKeywordConflictRule, and LatestRule implementations, then trace OpenSearchExprValueFactory and ObjectContent through the reported parsing path. Compare how schema conflicts and per-document mismatches are handled today. Done requires an agreed policy and defined behavior for both layers, including how any configurable modes should be specified and tested.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
backend, databases
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.