test(core): add reranker real-model smoke and quality harness - #1233
Merged
Conversation
Signed-off-by: phernandez <paul@basicmachines.co>
Signed-off-by: phernandez <paul@basicmachines.co>
Signed-off-by: phernandez <paul@basicmachines.co>
14 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phases 2-3 of #1231 — the remaining automated items. Test-only: six files under
test-int/semantic/, zero production changes.What's added
Real-model smoke (
test_real_fastembed_reranker.py, semantic-marked): loads theactual
jinaai/jina-reranker-v1-tiny-encross-encoder through the productionprovider/factory path (no stubs), asserts the relevant doc wins with calibrated
[0,1]scores and deterministic ordering, plus one end-to-end hybrid search through an indexed
repository with reranking enabled.
Quality regression harness (
test_reranker_quality.py+ ranking corpus/metricsextensions): golden ranking-sensitive queries measured reranker-off vs reranker-on with
real embeddings + real cross-encoder. Asserts on >= off for hit@1/MRR plus an absolute
hit@5 floor so a silent rerank no-op cannot pass. Measured on the curated set:
No queries needed exclusion; the real tiny model degraded nothing.
Latency benchmark (
test_reranker_latency.py, semantic+benchmark-marked, report-focused with a generous catastrophic-regression ceiling only). Warmed local numbers,
20 iterations:
~90 ms P50 overhead for the local tiny cross-encoder — comfortably inside the latency
headroom that motivated Add a rerank stage to search: ~half of LoCoMo benchmark misses are ranking failures, and there is ~20x latency headroom #950.
Verification
cached).
just typecheck,just lint: clean.test_semantic_quality[asyncio-postgres-fastembed]can exceed the 120s per-testtimeout on a loaded local machine — reproduced identically on
main, pre-existing andunrelated to this PR.
Ticks the Phase 2 fastembed-smoke and both Phase 3 boxes on #1231. Remaining there: the
deferred live LiteLLM run and the optional live smoke. These numbers also set the baseline
for revisiting #951 (entity boosting) and closing #618.
🤖 Generated with Claude Code
https://claude.ai/code/session_01CPdSXDbYyhyZ1TwgFnpEv8