fix(indexing): unblock chunk indexing and gate reranker model load
Three defects surfaced during the first live ingest:
1. OpenSearch index mapping is dynamic:"strict", so the first chunk
carrying metadata.section_heading was rejected with
strict_dynamic_mapping_exception and zero chunks were indexed. The
metadata and quality_flags sub-objects now opt into
dynamic:true so chunker-specific keys insert without a fixed schema,
while the top-level mapping stays strict.
2. hybrid_search.run_search called get_reranker() unconditionally, so
every search - even search_mode=lexical with RERANKER_ENABLED=false
- triggered a multi-GB FlagReranker download and hung the request.
The loader is now lazy behind the reranker_enabled flag.
3. docker-compose x-common-env never forwarded RERANKER_ENABLED, so the
container always defaulted to true regardless of .env. Added
RERANKER_ENABLED: ${RERANKER_ENABLED:-true} to common-env.
After these fixes POST /search (lexical) returns HTTP 200 in ~50ms with
a ranked hit and citation; OpenSearch _count reports the indexed chunk.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -72,14 +72,18 @@ def run_search(req: SearchRequest) -> SearchResponse:
|
||||
merged = _merge(lexical, semantic, mode)
|
||||
candidates = merged[: settings.rerank_candidates]
|
||||
|
||||
reranker = get_reranker()
|
||||
reranked_flag = False
|
||||
if settings.reranker_enabled and reranker.available and candidates:
|
||||
scores = reranker.score(req.query, [c.text for c in candidates])
|
||||
for c, s in zip(candidates, scores, strict=True):
|
||||
c.dense_score = s
|
||||
candidates.sort(key=lambda c: (c.dense_score or 0.0), reverse=True)
|
||||
reranked_flag = True
|
||||
if settings.reranker_enabled and candidates:
|
||||
# Lazy-load only when the flag is on; FlagReranker construction
|
||||
# downloads the model and we do not want that on the cold path when
|
||||
# the operator has disabled reranking.
|
||||
reranker = get_reranker()
|
||||
if reranker.available:
|
||||
scores = reranker.score(req.query, [c.text for c in candidates])
|
||||
for c, s in zip(candidates, scores, strict=True):
|
||||
c.dense_score = s
|
||||
candidates.sort(key=lambda c: (c.dense_score or 0.0), reverse=True)
|
||||
reranked_flag = True
|
||||
|
||||
final = candidates[: req.limit]
|
||||
|
||||
|
||||
@@ -72,8 +72,12 @@ INDEX_SETTINGS: dict[str, Any] = {
|
||||
},
|
||||
"ocr_confidence": {"type": "float"},
|
||||
"language_hint": {"type": "keyword"},
|
||||
"metadata": {"type": "object", "enabled": True},
|
||||
"quality_flags": {"type": "object", "enabled": True},
|
||||
# metadata + quality_flags carry chunker-specific keys
|
||||
# (section_heading, table_index, low_ocr_confidence, ...) that we do
|
||||
# not want to enumerate up-front. Top-level mapping is dynamic:
|
||||
# "strict", but these two sub-objects opt-in to dynamic insertion.
|
||||
"metadata": {"type": "object", "enabled": True, "dynamic": True},
|
||||
"quality_flags": {"type": "object", "enabled": True, "dynamic": True},
|
||||
"created_at": {"type": "date"},
|
||||
},
|
||||
},
|
||||
|
||||
Reference in New Issue
Block a user