Skip to content

feat(ha): follow-ups to consistent-hash session placement (#3661) #4223

Description

@EmilienM

User Story

I work on gateway HA and run multi-replica OpenShell gateways on Kubernetes with an external PostgreSQL database. #3661 places each sandbox's supervisor session on a replica picked by a consistent hash, and with replica routing on, lets the CLI send SSH, exec and forward traffic straight to that replica. I need operators to be able to see that placement, CI to test it, and every client to use it, not only the CLI.

Problem Statement

Follow-ups to #3661, part of #3528. #3661 ships the placement. These gaps come with it:

  • Nothing reports placement. There are no metrics for redirects, sessions served off their ring owner, or ring size.
  • No e2e checks that a session ends up on its ring owner. A failed redirect dial (a TLS error, say) falls back to the gateway Service, so the current HA lanes still pass.
  • Replica routing (grpcRoute.replicaRouting) has Helm unit tests but no e2e. The unrouted retry (retry_unrouted in crates/openshell-cli/src/tls.rs) only matches Envoy Gateway's error wording.
  • Only openshell-cli asks for the owner (x-openshell-want-owner) and sends x-openshell-replica. With replica routing on, the Rust, Go, TypeScript and Python SDKs still go through the shared route, so their streams relay through a peer whenever it picks a replica other than the owner.
  • Token-mode ssh-proxy (CLI sandbox connect, sync and forward, plus the TUI's shell, exec and forwards) calls GetSandbox again on every connection. The CLI already has the owner in SshSessionConfig but never passes it to build_proxy_command, and the TUI never asks for it.
  • Replica routing requires a StatefulSet, which makes an existing chart gap matter more: with server.externalDbSecret set, the StatefulSet still renders a 1Gi openshell-data PVC per replica that only the SQLite path uses, and it stays Pending on clusters without a default StorageClass.
  • The membership worker has no timeout around register() and ring(), so a database call that hangs keeps a stale ring past the TTL. Ring membership also compares the reader's clock with the writer's timestamp, so clock skew between replicas changes who counts as live.
  • The HA guide and chart text tell multi-replica users to run a Deployment, while replica routing needs a StatefulSet. The guide also says a stopping owner redirects its supervisors, but sandboxes created on a release before v0.1.3 run a supervisor that can't follow a redirect.

Impact / Why This Matters

Operators who enable replica routing can't confirm it works, and CI won't catch a placement regression. With replica routing on, SDK users still pay an extra peer hop for the whole life of most exec, SSH and forward streams. Some clusters can't run the StatefulSet that replica routing needs. A hung database call or a replica clock running ahead can keep a dead replica in the ring, so supervisors redirected to it waste a dial. Operators following the HA guide get conflicting advice on which workload to run, and expect a shutdown handoff that older sandboxes don't get. None of this is a correctness bug: unhinted traffic still reaches the owner through the peer relay, and a failed redirect falls back to the gateway Service.

Proposed Design

  • /metrics shows placement next to the session and peer-relay series from feat(server): add gateway capacity metrics and optional HPA #3978. Notional output (names not final):
    openshell_server_session_redirects_total{reason="connect"} 12
    openshell_server_session_redirects_total{reason="shutdown"} 40
    openshell_server_session_offring_admissions_total 1
    openshell_server_ring_members 3
    
  • The HA e2e checks placement with TLS on between peers, and a replica-routing variant covers direct routing and the unrouted retry.
  • Every first-party client asks for the owner and sends the replica hint, with the same single unrouted retry as the CLI.
  • The chart skips the data PVC when the database is external.
  • Ring membership holds up when the database hangs or replica clocks drift.
  • The HA docs say which workload to run when, and which sandboxes get the shutdown handoff.

Not covered here: measuring the shutdown handoff and pacing it if needed (#3661 redirects every session at once when a replica stops) and rebalancing established sessions after a scale-up or rollout, both in the #3528 checklist; closing admission on already-accepted HTTP/2 connections at shutdown (#3551); and the supervisor reconnect backoff never resetting after an accepted session, which predates #3661 and is in the same checklist.

Acceptance Criteria

Each line is one PR (the SDK line is one PR per SDK).

  • Placement metrics in crates/openshell-server/src/gateway_metrics.rs: redirects actually sent, by reason (connect, shutdown), redirected hellos served by a replica its own ring doesn't name (off-ring admissions), and live ring members. Count at the two SessionRedirect sends in supervisor_session.rs (handle_connect_supervisor, and run_session_loop, where a failed try_send must not count), not in preferred_peer_redirect. Update the ring gauge wherever the ring changes in gateway_members.rs (refresh, clear_ring_if_stale, shutdown). A hello after a failed redirect dial looks the same as one after a successful redirect (redirected=true in both), so counting true fallbacks needs a new SupervisorHello field. Leave that out unless the off-ring count proves too noisy.
  • e2e/rust/tests/kubernetes_ha_operations.rs (ha-tls lane from feat(server)!: harden gateway peer transport and expand HA conformance #3825) asserts that a sandbox created (or reconnected) after every replica has joined the ring is held by its ring owner (x-openshell-owner on a GetSandbox that sends x-openshell-want-owner), and stays there after its supervisor connection drops with no replica change. Sessions left in place by a scale-up or a forced owner loss are expected (moving them is feat(ha): add production scaling signals and graceful gateway redistribution #3528) and are not checked. GatewayRing is private to openshell-server, so the test mirrors it or the PR exposes the expected owner. Needs feat(server)!: harden gateway peer transport and expand HA conformance #3825.
  • A replica-routing e2e variant (StatefulSet plus grpcRoute.replicaRouting on Envoy Gateway) checks that a routed exec and SSH land on the owner (openshell_server_routed_request_attempts_total{route="local"} grows there, not route="peer"), and that the unrouted retry fires once when the owner's replica Service loses its endpoints while the owner pod stays up (for example, by patching its selector). Once feat(helm): add agentgateway ingress support #3714 merges, run it on agentgateway too and extend retry_unrouted if the wording differs.
  • Each SDK sends x-openshell-want-owner on GetSandbox and x-openshell-replica on whichever of ExecSandbox, ExecSandboxInteractive and ForwardTcp it exposes (SSH runs over ForwardTcp), with one unrouted retry like the CLI: Rust (crates/openshell-sdk/src/client.rs), Go (sdk/go/openshell/v1/exec_client.go, ssh_client.go, tcp_client.go), TypeScript (sdk/typescript/src/client.ts), Python (python/openshell/sandbox.py).
  • build_proxy_command (crates/openshell-core/src/forward.rs, shared with the TUI) takes an optional owner and passes it to ssh-proxy as a new flag. sandbox_ssh_proxy in crates/openshell-cli/src/ssh.rs skips its per-connection GetSandbox only when that flag is set, and keeps the lookup otherwise. The CLI passes the owner from ssh_session_config. The TUI (crates/openshell-tui/src/lib.rs) keeps the lookup until it asks for the owner itself, since it doesn't link the CLI's hint helpers.
  • With server.externalDbSecret set, templates/statefulset.yaml renders no openshell-data volumeClaimTemplate and _gateway-workload.tpl drops the mount. volumeClaimTemplates can't change on an existing StatefulSet, so this needs an upgrade note or an opt-in value. feat(helm): make gateway PVC size and StorageClass configurable #3854 edits the same volumeClaimTemplates block, so land one after the other.
  • Membership hardening in crates/openshell-server/src/gateway_members.rs. Put a timeout around register() and ring() in spawn_membership_worker, so a hung store call errors and clear_ring_if_stale can drop the ring after MEMBER_TTL. live_members compares the reader's clock with the writer's updated_at_ms, and the .max(0) clamp only turns a negative age into zero: a replica whose clock runs ahead stays in its peers' rings past the TTL after it dies, and one running well behind drops out. Judge freshness by database time, or at least treat rows dated well ahead of the reader's clock as stale (that only fixes the first case), and update the clamp's comment to match. Also fix the register() doc comment: another replica's sweep (delete_if) can cause the conflict too, not only a duplicate replica ID.
  • Docs: the HA guide (docs/kubernetes/high-availability.mdx) lists workload.kind: deployment as a requirement, and setup.mdx, the workload comments in values.yaml and the multi-replica StatefulSet error in _helpers.tpl also point to a Deployment, while Replica Routing in ingress.mdx needs a StatefulSet with workload.allowMultiReplicaStatefulSet. Keep Deployment as the recommended HA workload, say when a StatefulSet is needed (replica routing), and have the HA guide's peer routing section mention ring placement and link Replica Routing. Also say that placement and the shutdown handoff only apply to supervisors that support redirects: sandboxes created on a release before v0.1.3 keep their supervisor and reconnect through the gateway Service until they're recreated.

Alternatives Considered

Checklist

  • I've reviewed existing issues and the published docs
  • This is a design proposal, not a "please build this" request

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions