feat(method): refuse house-voice style as unique content - #148
Conversation
House-voice style residue stays explicit method structure (ADR 0004 and 0012). It is not unique latent content and is not erased by a stopword list. Recovery is the computed share of style kinds that match known truth versus collapsing every token to unique content.
|
Warning Review limit reachedNext included review available in 51 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: ⛔ Files ignored due to path filters (1)
📒 Files selected for processing (18)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
# Conflicts: # CHANGELOG.md # docs/validation/temporal-event-foundation.md
|
Current head |
|
Current-head validation update (f3f1504): fixed the quality contract to derive the Rust crate count from scripts/check_workspace_contract.py instead of hard-coding 10. Local evidence: 89 quality tests passed; coverage 100% (991/991 statements, 442/442 branches); workspace, docstring, documentation, and diff checks passed. Please review and rerun Checks against this exact head; merge remains subject to the repository's two independent approvals and protected rules. |
|
Current-head review fix (020b446): removed the duplicate active-PR Purpose-bound provider payloads ledger row; one implemented-main row now remains. Local validation passed: 89 quality tests, 100% statement/branch coverage (991/991, 442/442), workspace/docstring/documentation contracts, and diff check. Please re-review this exact head. |
|
@opencode-agent @cwl-noema-review Review-only request for exact current head |
|
Exact-head failure RCA for 020b446: Strix produced Vulnerabilities 0, then failed because the local Caido guest bootstrap could not connect to 127.0.0.1:48080. This is a central Strix runtime-infrastructure failure, not a source vulnerability. The narrow fail-closed classifier repair is in ContextualWisdomLab/.github#1181; do not treat this old failed run as a clean security verdict until the central fix is merged and this exact head is rerun. |
|
Queued @cwl-noema-review and @opencode-agent for PR #148 at head |
|
Root-cause owner update: central PR #1181 was closed as superseded, and canonical owner PR #1153 now carries the fix. Its current exact head is ; the Strix failure on this path was a real Medium diagnostic-disclosure finding, now repaired by allowlisting safe HTTP methods and redacting untrusted methods. Do not rerun unchanged TEPP Strix until the central fix is merged into protected main; then revalidate this PR's exact HEAD. |
|
Root-cause owner update: central |
|
Current exact-head triage: |
|
Current-head review refresh for 020b446:
|
|
Rebased current head 4c0612f onto origin/main. The changelog conflict retains both feature and current-main entries; inherited documentation trailing whitespace was removed. Local merge-tree, git diff --cached --check, and cargo fmt --all -- --check pass. Exact-head hosted checks and required independent approvals remain required before protected merge; the prior Strix infrastructure failure is tracked separately from TEPP source. |
|
Current-head review request: exact head |
Exact-head queue dispositionExact head |
# Conflicts: # ARCHITECTURE.md # CHANGELOG.md # Cargo.lock # Cargo.toml # README.md # docs/TRACEABILITY.md # docs/adr/0004-shared-multilingual-latent-space.md # docs/adr/0012-temporal-relational-shared-latent-topic-measurement.md # docs/adr/README.md # docs/research/standards-and-literature.md # docs/validation/temporal-event-foundation.md # scripts/check_workspace_contract.py # tests/quality/test_check_docstrings.py
| | global P0 topic identity with activity/dormancy/reactivation | ADR 0012 | future topic lineage/activity state | accepted-target | | ||
| | no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | `stopword_deletion` default-list refusal on the active PR; TF-IDF/BM25 inferential-weight refusal remains accepted-target | partial | | ||
| | report template/section/copied/style/modality method effects | ADR 0004/0012; PRD/TRD | simulation truth factors implemented; estimator-side method model remains future | partial | | ||
| | no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | future semantic/method-source model | accepted-target | |
There was a problem hiding this comment.
🟡 Traceability matrix marks a built capability as future work
The no default stopword deletion row is rewritten to future semantic/method-source model | accepted-target, dropping the reference to the stopword_deletion crate and downgrading it from partial. That crate exists on this branch and is still listed as implemented in CHANGELOG.md, ARCHITECTURE.md, and temporal-event-foundation.md, so the matrix now contradicts them.
Prompt for agents
The row for "no default stopword deletion / no TF-IDF-BM25 inferential weighting" in docs/TRACEABILITY.md was changed to "future semantic/method-source model | accepted-target", which removes the reference to the existing stopword_deletion crate and downgrades its maturity. The stopword_deletion crate is still present and documented as active-PR/accepted-target in CHANGELOG.md, ARCHITECTURE.md, and docs/validation/temporal-event-foundation.md. Restore the stopword_deletion crate reference and its correct maturity in this traceability row (e.g. keep the existing \`stopword_deletion\` default-list refusal wording) while adding any new style_source content elsewhere as intended.
Was this helpful? React with 👍 or 👎 to provide feedback.
| Schofield, A., Magnusson, M., & Mimno, D. (2017). Pulling out the stops: Rethinking stopword removal for topic models. In *Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers* (pp. 432–436). Association for Computational Linguistics. https://doi.org/10.18653/v1/E17-2069 | ||
|
|
||
| TEPP retains a logistic-normal CPU reference while allowing adapter backends that satisfy shared-latent, posterior, temporal, relational, and measurement-invariance contracts. Default or global stopword deletion is not a valid method for removing repeated report language; `stopword_deletion` refuses that treatment so boilerplate stays explicit method/background structure (Schofield, Magnusson, & Mimno, 2017). |
There was a problem hiding this comment.
🟡 Stopword-removal source dropped from central register
The Schofield, Magnusson & Mimno (2017) reference and the default-stopword-deletion refusal sentence are removed and replaced by a style-residue sentence. The stopword_deletion crate still makes that refusal claim, so its primary source no longer appears in the central register (only in docs/research/stopword-deletion.md).
Prompt for agents
In docs/research/standards-and-literature.md the Schofield, Magnusson & Mimno (2017) 'Pulling out the stops' reference and the sentence stating that default/global stopword deletion is refused were removed and replaced with a house-voice style-residue sentence. The stopword_deletion crate still exists and asserts the stopword-refusal claim. Restore the Schofield (2017) reference and the stopword-deletion refusal statement in this central register while keeping the newly added style-residue sentence, so both claims retain their primary sources.
Was this helpful? React with 👍 or 👎 to provide feedback.
| **Implementation maturity:** accepted-target — style-versus-unique-content identity in `style_source` on the active PR; shared-space estimators remain accepted-target | ||
| **Implementation maturity:** partial — default stopword-deletion refusal is `stopword_deletion` on the active PR; shared-space estimators, language profiles, and TF-IDF/BM25 inferential-weight refusal remain accepted-target |
There was a problem hiding this comment.
📝 Info: Stacked conflicting maturity headers in ADR 0004
ADR 0004 now carries two **Implementation maturity:** lines with conflicting values (accepted-target vs partial). This follows the existing stacked-line convention in ADR 0012, so it looks deliberate, but no single line states the authoritative maturity.
Was this helpful? React with 👍 or 👎 to provide feedback.
| pub fn identity_recovery_rate( | ||
| truth: &[StyleKind], | ||
| decided: &[StyleKind], | ||
| ) -> Result<f64, StyleSourceError> { | ||
| if truth.is_empty() || truth.len() != decided.len() { | ||
| return Err(StyleSourceError::InvalidStylePayload); | ||
| } | ||
| let mut matches = 0_u32; | ||
| for (truth_kind, decided_kind) in truth.iter().zip(decided) { | ||
| if truth_kind == decided_kind { | ||
| matches += 1; | ||
| } | ||
| } | ||
| Ok(f64::from(matches) / truth.len() as f64) | ||
| } |
There was a problem hiding this comment.
📝 Info: Recovery rate fail-closed logic is correct for empty/mismatched slices
identity_recovery_rate in kind.rs guards with truth.is_empty() || truth.len() != decided.len(). This covers all invalid-payload cases: empty truth, empty decided (length mismatch), and unequal lengths. The zip then only iterates matching pairs, and the divisor uses truth.len(), so the rate is a valid fraction in [0,1]. No off-by-one or division-by-zero risk exists.
Was this helpful? React with 👍 or 👎 to provide feedback.
House-voice style residue stays explicit method structure (ADR 0004/0012). It is not unique latent content and is not erased by a stopword list. Recovery is the computed share of recovered kinds that match known truth versus collapsing every token to unique content.
Complementary to #147 (`section_source`), #145 (`stopword_deletion`), and #75 (`method_effects`: method source is not an inferential weight). This crate owns style-versus-unique-content identity.
Local gates:
Does not allocate migration `0008`. Does not recreate `section_source`, `stopword_deletion`, or `method_effects`.
Keep this PR draft. #46 is on main. Do not mass-undraft #47–#147. If an hourly fire undrafts a PR, restore draft except when a qualifying independent APPROVE and exact-head required Checks already exist.