fix(e2e): restrict readiness polling to loopback destinations - #1314
fix(e2e): restrict readiness polling to loopback destinations#1314cursor[bot] wants to merge 11 commits into
Conversation
Land the #1244 loopback boundary on current main without mixing Strix classifier changes. Accept literal localhost and standard-library loopback addresses, unwrap IPv4-mapped IPv6, and reject userinfo, missing hosts, unspecified binds, and metadata destinations before any request is opened. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
A poisoned hosts file could map the literal name to a public A/AAAA record. Resolve localhost before opening the readiness URL and fail closed unless every answer is loopback after IPv4-mapped unwrap. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
There was a problem hiding this comment.
Follow-up 3f0c3ac5 addresses the Devin localhost-without-resolution note.
Literal localhost is now resolved. Every A/AAAA answer must be loopback after IPv4-mapped unwrap. A poisoned hosts file, empty answer set, resolver error, or non-IP answer fails closed before urlopen. Numeric loopback literals are unchanged.
Local evidence on this head: 28 focused tests, 187 statements / 58 branches at 100% on sandboxed_web_e2e.py. Hosted Checks must rerun on 3f0c3ac5; cancelled or queued predecessor runs are not merge evidence.
Next action: wait for current-head OpenCode, Strix, and required Checks on 3f0c3ac5. Do not treat 9a47c10e results as this head.
Sent by Cursor Automation: Fix Issues
There was a problem hiding this comment.
Revalidated exact head 3f0c3ac5c018770c32a80663a28c1f499ddfa1ba.
Hosted success on this SHA: Python 3.14 exact contract and complete coverage; Python 3.10 contract; quality; exact-head-policy; exact-head-path-policy; noema-review; close-empty; generate-sbom; dependency-review; osv-scan; pip-audit; Semgrep; trivy-fs; Scorecard; gitleaks; CodeQL compatibility/merge-preview (actions and python); latest scan-pr-queue (run 32792998005). Devin Review status is success. The Devin TOCTOU note (validate-then-urlopen re-resolve) is Info only and is accepted for the documented static-hosts threat. Do not push a successor SHA for that note.
Not merge evidence:
- Required-workflow
opencode-reviewsuccess on run32792844690is the pull_request_target bootstrap stub. It is not an OpenCode GitHub App APPROVE on this head. - No
OpenCode Review Dispatchrun exists for #1314@3f0c3ac5. The live dispatch queue is.github#789(in progress) thennaruon#1195(queued). OriginWeave#40 failed after a 2h dispatch. - The #1314 review-event scheduler (
32792998005) ran withTRIGGER_REVIEWS=false. The laterworkflow_runscheduler (32793386272) hadTRIGGER_REVIEWS=truebut inspected only #897. - Org-queue-sweep
32791443215consumed the shared dispatch budget on naruon#1195, not this PR. - Strix
32792844719is still in progress on this SHA. Cancelled predecessorscan-pr-queue/9a47c10eruns are not this head.
Do not treat the stub check, queued/cancelled runs, or Devin Info as OpenCode approval. Wait for a current-head OpenCode App APPROVE plus Strix success on 3f0c3ac5 before the mechanical merge scheduler can land this PR. No new commit from this pass.
Sent by Cursor Automation: Fix Issues
There was a problem hiding this comment.
Strix required check now failed on exact head 3f0c3ac5c018770c32a80663a28c1f499ddfa1ba (run 32792844719). This is not a loopback-readiness vulnerability finding.
Observed on this SHA:
- Primary
nvidia_nim/nvidia/nemotron-3-super-120b-a12b:litellm.RateLimitError429 Too Many Requests, no report artifact, fail-closed. - Fallback
nvidia_nim/nvidia/llama-3.3-nemotron-super-49b-v1.5: penetration test completed with Vulnerabilities 0, then the gate classified the staticMODEL QUALITY WARNINGbanner as provider infrastructure/failure-signal output and failed closed. - Next fallback
openai-direct/gpt-5.6-luna:404 page not found. - Final annotation: provider/backend unavailable (#466, #467). Zero mapped vulnerability reports.
Do not treat this as merge evidence, and do not treat it as a defect in require_loopback_readiness_url. Do not push a successor SHA on #1314 to retry Strix or to carry #1311 banner sanitization — that would invalidate the green quality/Noema/security Checks and mix Strix classifier ownership into this SSRF slice.
Still not merge evidence: required-workflow opencode-review stub success; no OpenCode Review Dispatch for this PR@SHA. Devin TOCTOU remains Info and is accepted. Wait for a same-head Strix success (retry/dispatch, not a new commit) plus OpenCode App APPROVE on 3f0c3ac5.
Sent by Cursor Automation: Fix Issues
|
Exact-head evidence correction for current head
Therefore this PR is not merge-ready. Do not aggregate Semgrep/Trivy association status, the |
|
Live scheduler evidence from the
This comment is not OpenCode approval and not merge evidence. Do not push a successor SHA here to retry Strix or to absorb |
|
Revalidated 2026-08-25T01:24Z. Head still Local this hour: Trusted-base Strix still cannot pass until #1311 is the gate on Org-queue-sweep |
|
Follow-up 2026-08-25T01:32Z, not merge evidence: #1311 auto-merges cleanly onto |
|
Correction 2026-08-25T01:34Z, not merge evidence. Org-queue-sweep The trusted workflow excludes Hub-repo PRs are only scanned by same-repo Land path is unchanged: operator lands #1311 plus the |
|
Follow-up 2026-08-25T01:36Z, not merge evidence. The |
|
Land-order correction 2026-08-25T01:39Z, not merge evidence. Do not push a successor SHA here. This head still overlays Live contrast: #1311 same-head re-run #1316 is required for current- Operator path for this exact head
|
|
Correction after #1318 merged at 01:42:35Z. Not merge evidence for this PR. Protected This inverts the same-head Strix advice I posted at 01:39Z:
#1316@ Operator path: land #1311 onto |
|
Follow-up 02:13Z from #1316@ #1318 smoke is live: that current- This does not change the instruction for |
|
Follow-up 02:24Z from #1311 same-head Strix That 42-minute run produced Vulnerabilities 0 on both the primary 120b (with 429) and the 49b fallback. The trusted pre-sanitizer gate discarded both as infra because report/console logs contain the Do not push |
|
Queue note 02:37Z, not merge evidence. #1297@ Landing #1297's concurrency change onto |
|
Exact-head update: This is not OpenCode approval and not merge evidence. #1311@ Local: loopback/path-policy pytest 27 passed; filtered Strix cases Do not retry required Strix on this SHA (luna overlay vs trusted |
…heading Deleting a whole MODEL QUALITY WARNING box hid a Fatal or provider Warning emitted on the next line of that same box. Strip only the cosmetic heading line, keep the classified console copy so last-attempt stays raw, and cover the same-box fail-closed path. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
A trailing-space fullmatch left MODEL QUALITY WARNING in the classified copy, so a clean 0-finding scan still fail-closed. Search the heading after stripping ANSI so same-box failures stay visible. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
Heading-only stripping left fixture or model text such as -warning- in the leftover banner, which the generic Warn matcher fail-closed. Delete a box that is only MODEL QUALITY WARNING. If the same box also has Fatal, Denied, Timeout, or Provider WARNING, strip the heading and keep the failure line. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
|
Revalidated 2026-08-25T04:17Z. Land SHA is now Protected |
…dence The PR overlay still named the removed gpt-5.6-luna fallback, so the trusted required-path smoke fail-closed before the sanitizer could matter. Match the protected-main gpt-5.4 contract. Also keep RateLimitError and related infra tokens when they share a MODEL QUALITY WARNING box, so deleting a cosmetic-only banner cannot hide them. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
Path-policy still pinned the removed gpt-5.6-luna secret-absent default after the overlay matched protected main. RateLimitError in a MODEL QUALITY WARNING box is retryable, so the keep-path must expect the three-model Vertex sequence; one call would mean the box was deleted. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
|
Revalidated on Path-policy on Local Protected |
|
Hosted path-policy Required Strix |
|
Required Strix Observed, not invented:
The llama-3.3 0-finding attempt is the catch-22 this sanitizer exists to close. RateLimitError and LLM CONNECTION FAILED stay fail-closed even after the sanitizer is trusted. Mechanical merge cannot treat this failed required check as merge evidence. Do not merge |
|
Revalidated 2026-08-25T05:50Z. Do not invent merge evidence.
|
|
Pushed This is not merge evidence. Required Strix still executes the pre-sanitizer gate from Previous head |
Merging 8fd471a brought #1318's global luna-to-gpt-5.4 replace in test_strix_quick_gate.sh. Three OpenCode dispatch-pool assertions now expected openai/gpt-5.4, but opencode-review-dispatch.yml still lists openai/gpt-5.6-luna. Restore those needles so the harness matches the unchanged dispatch workflow. Keep the Strix quota fixtures on gpt-5.4. Do not fold the dispatch-pool model change from #1316. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
|
Follow-up on
|
|
Please review exact current head |
|
Required Strix Observed, not invented:
The two NIM 0-finding attempts are the catch-22 this sanitizer exists to close. |
|
Please review exact head |
|
Revalidated 2026-08-25T08:18Z. This is not an approval and not merge evidence.
|
|
Revalidated 2026-08-25T10:05Z. This is not an approval and not merge evidence.
|
There was a problem hiding this comment.
Pull request overview
OpenCode could not approve from deterministic current-head evidence because GitHub Checks have failed.
Findings
1. HIGH Current-head GitHub Checks - Fix failed required checks before approval
- Problem: Failed same-head checks remain for
d4ea752830df1c6e25bc38147fe08a9b3480f058. - Root cause: The model-unavailable evidence fallback is allowed only when peer GitHub Checks are complete and clean.
- Fix: Read and fix the failed check logs below, then rerun the current-head checks.
- Regression test: Keep the model-unavailable fallback gated on an empty failed-check rollup.
Failed checks:
- Strix Security Scan/strix: FAILURE (https://github.com/ContextualWisdomLab/.github/actions/runs/32814998745/job/97701390309)
- Strix Security Scan/strix: failure (https://github.com/ContextualWisdomLab/.github/actions/runs/32814998745/job/97701390309)
Changed-File Evidence Map
flowchart LR
PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
Evidence --> S1["Changed file (2 files)"]
S1 --> I1["repository behavior"]
I1 --> Conflict["Merge conflict blocks this path"]
Conflict --> V1["required checks"]
Evidence --> S2["Docs (2 files)"]
S2 --> I2["operator or user guidance"]
I2 --> Conflict["Merge conflict blocks this path"]
Conflict --> V2["docs review"]
Evidence --> S3["CI script (3 files)"]
S3 --> I3["review and security gate shell path"]
I3 --> Conflict["Merge conflict blocks this path"]
Conflict --> V3["bash -n plus Strix self-test"]
Evidence --> S4["Test (3 files)"]
S4 --> I4["regression suite"]
I4 --> Conflict["Merge conflict blocks this path"]
Conflict --> V4["targeted test run"]
OpenCode Review Overview
Pull request overviewOpenCode could not approve from deterministic current-head evidence because GitHub Checks have failed. Findings1. HIGH Current-head GitHub Checks - Fix failed required checks before approval
Failed checks:
Changed-File Evidence Mapflowchart LR
PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
Evidence --> S1["Changed file (2 files)"]
S1 --> I1["repository behavior"]
I1 --> Conflict["Merge conflict blocks this path"]
Conflict --> V1["required checks"]
Evidence --> S2["Docs (2 files)"]
S2 --> I2["operator or user guidance"]
I2 --> Conflict["Merge conflict blocks this path"]
Conflict --> V2["docs review"]
Evidence --> S3["CI script (3 files)"]
S3 --> I3["review and security gate shell path"]
I3 --> Conflict["Merge conflict blocks this path"]
Conflict --> V3["bash -n plus Strix self-test"]
Evidence --> S4["Test (3 files)"]
S4 --> I4["regression suite"]
I4 --> Conflict["Merge conflict blocks this path"]
Conflict --> V4["targeted test run"]
Merge Conflict Guidance
gh pr checkout 1314 --repo ContextualWisdomLab/.github
git fetch origin main
git merge --no-ff origin/main # or: git rebase origin/main
git status --short
# resolve files, then git add <resolved-files>
# merge path: git commit
# rebase path: git rebase --continue
git push origin HEAD:cursor/bc-deb78137-73b8-43fe-bfe5-af265490f12a-4f61
# rebase path only: git push --force-with-lease origin HEAD:cursor/bc-deb78137-73b8-43fe-bfe5-af265490f12a-4f61 |
Keep the classified-copy box sanitizer so gate-last-attempt stays raw. Take #1320's HF Hub notice filter and event-visibility/fallback contracts. Point OpenCode dispatch-pool harness needles at openai/gpt-5.4 now that opencode-review-dispatch.yml on main lists that model. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
|
Revalidated 2026-08-25T11:45Z. This is not an approval and not merge evidence.
|
|
Please review exact current head Hosted quality, exact-head-path-policy, exact-head-policy, and both Python contracts are terminal GitHub-success on this SHA. Unresolved review threads are 0. The prior |


Why
Sandboxed web E2E still accepted any HTTP(S) hostname on
--backend-ready-url/--frontend-ready-url. A review run could poll a public host or a link-local metadata endpoint instead of the sandboxed app. #1244 already specified the repair but is behind currentmain. #1313 only allowslocalhost/127.0.0.1, drops IPv6::1and127.0.0.0/8, and mixes Strix classifier ownership into the SSRF slice.This is the current-
mainsuccessor for that security boundary. It also carries a Strix banner sanitizer that deletes a cosmetic-onlyMODEL QUALITY WARNINGbox and keeps Fatal/Denied/Timeout, Provider WARNING, RateLimitError, and related infra tokens when those appear in the same box. #1311 still mutates$STRIX_LOGin place, sogate-last-attempt.logis not byte-faithful. Do not land #1311.Repair
localhost(trailing FQDN dot stripped) or an address Pythonipaddressclassifies as loopback.localhostand require every A/AAAA answer to be loopback after IPv4-mapped unwrap, so a poisoned hosts file cannot smuggle a public address.::ffff:8.8.8.8cannot bypass the rule while::ffff:127.0.0.1still works.0.0.0.0,::),.localhostsubdomains, and169.254.169.254.MODEL QUALITY WARNINGbox. If the same box also carries Fatal, Denied, Timeout, Provider WARNING, RateLimitError, Nvidia_nimException, LLM CONNECTION FAILED, APIConnectionError, or Too Many Requests, strip only the heading and keep the failure line. A preceding Fatal/Denied box stays.gate-last-attempt.logandgate-attempts/stay raw.gpt-5.4default and NVIDIA fallback. Do not treatopenai_direct/gpt-5.6-lunaquota-fixture names as live workflow pins.Verification (exact current head)
d4ea752830df1c6e25bc38147fe08a9b3480f058.main@8fd471a31399a914d9cb22a840f4a4c68e010ea6; the head is a non-force descendant of this base.32815000721checked out exact headd4ea7528…, passed 1,408 tests (1 skipped, 16 subtests), and completed fulltest_strix_quick_gate: PASS.5406826622.gpt-5.4overlay while preserving the raw last-attempt artifact and fail-closed infrastructure/vulnerability classification.Operator next action
Wait for an exact-current-head substantive formal verdict and the #1222 exact-head SAST/Trivy provenance owner repair. Do not admin-merge, self-approve, or count GitHub-success scan labels as authoritative binding. If the verdict identifies a current-source defect, repair it test-first on this owner branch; otherwise merge only under live governance after all evidence is exact-head and independently approved.
Closes nothing. Supersedes the live-head intent of #1244 on current
main. Successor for the sanitizer intent of #1311 without copying its in-place last-attempt mutation.