stackabletech / stackabletech/cockpit

Shared persistent server-side state

Open
#310 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
TypeScript
Stars
1
Forks
0
Avg merge
22h 6m
Merged PRs (30d)
12

Description

Epic: Shared persistent server-side state

All server-side state is held in process-local Maps, so the deployment is capped at one replica and a restart drops everything.
Part of this is already tackled in the ongoing FileBrowser PRs.

[!NOTE]
This text is AI generated against our normal policy because it is migrated from another place, when we tackle this for real it needs to be properly rewritten

Four places:

  • Trino connection configsrc/lib/server/trino/user-clients.ts:18,
    new Map<string, UserEntry>(). Per-user URL and credentials.
  • Sessions and userssrc/lib/server/auth.ts. betterAuth() is called without a
    database option, so it uses the built-in memory adapter.
  • Query statesrc/lib/server/trino/queries.ts:76,79. Progress, rows, status and
    per-tab access times.
  • Readinesssrc/routes/healthz/+server.ts. Liveness and readiness both point at
    /healthz, which always returns 200.

What this costs

  • replicaCount is a plain user-settable chart value (values.yaml:4, documented as
    "Number of replicas") with nothing stopping someone setting it to 2. At two replicas
    users are logged out at random, a query started on one pod is invisible to the other,
    and a saved connection only exists on the pod that received it.
  • A restart orphans running queries in Trino, loses completed results, and logs everyone out.
  • A readiness probe has nothing meaningful to check, so /readyz would be a placeholder
    until a backing store exists.

Scope

  • Pick a backing store and wire all four consumers to it. #41 proposes PostgreSQL for
    better-auth; whatever is chosen should cover all four rather than solving them separately.
  • Split /healthz (liveness, trivial) and /readyz (readiness, checks the store), and
    update the Helm probes.
  • Decide what replicaCount > 1 does until then: document the constraint, or make the
    chart reject it.

Out of scope

  • Persisting completed query results beyond their TTL (STACKABLE_COCKPIT_QUERY_TTL, default 1800s). Separate question.
  • S3 credentials in localStorage (#306). Client-side, different problem.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading src/lib/server/trino/user-clients.ts, src/lib/server/auth.ts, src/lib/server/trino/queries.ts, and src/routes/healthz/+server.ts, then inspect values.yaml and the Helm probes. Decide on a backing store that covers the four consumers and define the replicaCount behavior. Done means shared state survives restarts and replicas, /healthz and /readyz have distinct checks, and the probes are updated.

Written by the indexing model from the issue text.

Assessment

Tech stack
helm, postgresql, typescript
Domain
backend, databases, devops
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.