Skip to content

fix: keep runtime alive during channel recovery - #25

Draft
tylerslaton wants to merge 1 commit into
mainfrom
fix/channel-startup-recovery
Draft

fix: keep runtime alive during channel recovery#25
tylerslaton wants to merge 1 commit into
mainfrom
fix/channel-startup-recovery

Conversation

@tylerslaton

Copy link
Copy Markdown
Contributor

What changed

  • If the initial 15-second Channel readiness wait expires while the runtime is still connecting or reconnecting, start the HTTP server so Railway health checks keep the process alive.
  • Continue waiting for Channel readiness in the background.
  • Keep permanent initial failures terminal, and shut down if background recovery later becomes terminal.

Why

A runtime that starts during a realtime gateway outage currently exits before a retry can connect after the gateway recovers. Keeping the HTTP health route alive lets the CopilotKit retry loop eventually restore the managed Channel without a process restart.

Dependency

Depends on CopilotKit/CopilotKit#6347 and a CopilotKit release that includes it. That change retries transient initial gateway activation with exponential backoff.

Validation

  • pnpm exec vitest run app/server-recovery.test.ts app/server.test.ts --reporter=verbose — 6 tests passed
  • git diff --check

Existing repository failures

  • pnpm check-types fails unchanged on the clean base checkout because @copilotkit/channels 0.6.1 and the runtime transitive channels-core 0.6.0 expose incompatible types.
  • The full test suite also has existing failures from the same package skew. The two server test files changed or exercised here pass.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant