Site

These procedures apply to the selected Supabase/Neon database and Supabase Auth/Better Auth composition. Record the exact provider, branch/project, region, pool mode, deployment version and compatibility evidence before every operation. Never switch the provider of an existing installation in place.

Credential rotation

  1. Enable maintenance mode and pause the native or external scheduler. Keep health, Auth recovery and the admin surface reachable.
  2. Confirm no destructive GDPR or billing command is in processing without a valid lease. Do not replay an ambiguous external mutation.
  3. Using a short-lived administrator connection, attest the existing role attributes and grants. Refuse rotation if the role is owner, superuser, BYPASSRLS, replication-capable, or can SET ROLE to an owner.
  4. Generate a new independent password for exactly one login: app_runtime, app_maintenance, app_financial_worker, app_platform_admin, or auth_runtime_login_v1.
  5. Update the matching secret in the deployment platform, start new instances, and verify /api/health plus the role preflight.
  6. Drain old instances and their pools. Only after the old credential has no clients may the previous password be revoked.
  7. Resume the scheduler and inspect terminal failures/backlog.

Never rotate multiple database roles to the same password. Init preflights existing credentials but does not silently rotate them.

Auth signing secrets

Rotate BETTER_AUTH_SECRET and AUTH_SESSION_FINALIZATION_SECRET in a maintenance window. Drain web instances, expire/revoke affected live sessions and pending credentials according to the selected provider, deploy the new values everywhere, then resume admission. These secrets do not have a public runtime bypass or an implicit previous-key fallback.

The Better Auth local drill must restart the real Auth runtime with the new BETTER_AUTH_SECRET, prove that a cookie signed by the old secret is rejected while its database session still exists, and only then revoke that session and drain the old runtime. Rotating only an application-side marker is not evidence of a Better Auth signing-secret rotation.

For Supabase Auth keys, rotate them in the provider dashboard, update server secrets first and publishable keys at build time, rebuild client assets, then revoke the old key after all instances have drained.

Pool drainage

Stop new traffic before closing pools. Pause job claims, wait for active requests and job leases to settle, then terminate the application process normally so the shared Prisma, maintenance, financial-worker, platform-admin and Better Auth pools call their disconnect paths. Do not kill a worker during an unfenced external mutation. A hard stop requires reconciliation of every ambiguous run before retry.

Before scaling or changing a pool maximum, update the declared budget and verify DATABASE_APPLICATION_INSTANCES * (DATABASE_POOL_MAX + 1 maintenance + 1 financial-worker + 2 platform-admin + Better Auth pool when enabled) + DATABASE_RESERVED_CONNECTIONS < DATABASE_BACKEND_CONNECTION_LIMIT. The backend limit comes from the active Supabase/Neon tier. Init refuses to issue credentials and runtime refuses startup when the final margin would be below one connection.

Treat sustained non-zero pool waiters, pool error events, statement timeouts and transaction rollbacks as alerts. The public health response is intentionally small; detailed counts and redacted failures belong in operational logs.

Controlled outage drill

Run outage drills only against an isolated target authorized for destructive QA. Pause admission and the scheduler, record the selected composition, then interrupt one boundary at a time: database reachability first, Auth reachability in a separate run. Do not terminate the database or change a managed endpoint merely to satisfy the repository checker.

For the database drill, verify the public route returns its documented bounded failure instead of provider detail, one redacted incident becomes visible only to a platform administrator, retention purge still removes an eligible fixture, and durable jobs recover after reachability returns. For the Auth drill, verify temporary failure does not become a login redirect or delete an otherwise valid session cookie, the public response is non-cacheable and controlled, the log is redacted and platform-admin-only, and the same provider adapter recovers.

Reconcile every ambiguous external write before resuming claims. A drill that cannot prove log visibility, purge, queue recovery or cookie behavior is failed or skipped; it is never inferred from a unit test.

Backups and restores

Enable the provider's managed backups and point-in-time recovery before production data exists. A backup is not accepted until a restore has been exercised into a separate empty project/branch.

For a restore drill:

  1. Pause writes and record the recovery point.
  2. Restore into an isolated target; never overwrite the only healthy copy.
  3. Keep all runtime logins disabled until schema catalogue comparison and role attestation pass.
  4. Reconfigure external Supabase Auth/Storage only to the restored database identity contract; never import provider subjects by email.
  5. Verify Account isolation, credits ledger consistency, billing queues, deletion latches and job leases.
  6. Issue new runtime credentials, deploy to the restored target, then resume jobs in bounded batches.

Supabase and Neon backup retention/PITR settings are provider-managed. Exported SQL from db:schema:build is installation DDL, not a data backup.

Supabase backup restores do not restore passwords for custom database roles. After a managed restore, reset and redeploy every dedicated runtime credential before admission, then repeat the role and connection-budget preflights. For a repeatable Supabase QA environment, use an isolated persistent branch; branches have independent credentials and start without production data unless seeding is configured explicitly.

The Auth restore scope is deliberately provider-specific. Better Auth backup and restore includes the complete authn schema, including its users, sessions and verification records. Supabase Auth restoration is narrower: export and restore only the stable identity rows in auth.users and auth.identities. Do not carry over auth.sessions, refresh tokens, MFA enrollment/challenge state or SSO/SAML configuration through this archive. Users must sign in again after recovery; re-enroll MFA factors and reconfigure SSO/SAML where needed. Provider subjects are restored by their stable identifiers, never reconstructed by email.

The auth_restore_scope_attested backup metric means the drill proved exactly that boundary; it does not claim uninterrupted Auth sessions or a complete Supabase Auth restore. For neon:supabase, the database drill explicitly attests that no provider-owned GoTrue state was restored into Neon. Run the separate Supabase Auth recovery for auth.users and auth.identities before admitting users. For supabase:supabase, those two tables use the separate data-only Auth archive alongside the application archive.

Durable job recovery

The scheduler only wakes the common runner. After an outage, first restore database reachability, then inspect stale processing leases, pending retry timestamps, failed runs and notification delivery state. Allow the runner's lease recovery to reclaim eligible work. Do not manually duplicate queue rows or re-run provider writes whose outcome is unknown.

Resume with one runner, confirm settlement and backlog movement, then restore normal concurrency. Terminal failures and notification failures stay visible until an operator resolves their cause.

Observability and performance drill

Measure cold and warm starts separately on each exact disposable composition. The Q3 profile records p50/p95/p99 latency, error rate, p95 pool wait, peak client and backend connections, prepared-statement behavior, idle transactions and a mixed interactive/worker load. For Supabase Auth, the Better Auth pool counts are zero; for Better Auth, measure the isolated auth_runtime budget alongside app_runtime, app_maintenance, app_financial_worker and app_platform_admin.

Use the declared provider-tier connection limit and configured pool budgets. Do not raise a threshold to make a run pass. Raw plans and measurements are sanitized artefacts; the checked-in receipt contains only bounded booleans, numbers and SHA-256 bindings. The local-only Phase 1 orchestrator now owns the exact 24-cell matrix, validates dataset/workload/budget/RLS inputs, rejects non-loopback targets, publishes atomically and binds every cell artefact with SHA-256. The repository-owned CLI binds its own bytes and any supplied measurement adapter by repository-relative paths and exact SHA-256 digests. This proves executable integrity, not measurements. The canonical contract remains not_ready: a real measurement adapter, the reviewed dataset, workload parameters, budgets, real RLS-plan captures and all 24 cell runs still need to be supplied and run.

Q3 operational evidence

tests/manifests/database-auth-operational-evidence.json is the fail-closed registry for controlled Database/Auth failures, database and Auth-secret rotation, pool drainage, backup/restore, durable-job recovery, observability redaction and the performance profile. It has one local report for each of the four Database/Auth compositions. The checked-in baseline is passed for all four reports: every one binds all nine registered scenarios without skips. Q3 operational success is required by Q4, but never activates a composition by itself.

The Q4 compatibility checker verifies the matching composition's receipts, raw artefacts, SHA-256 bindings and runner provenance before accepting local certification or activation. Evidence from another composition cannot satisfy the lock. Each composition remains gated independently.

For a local composition whose database provider is Supabase, the runner owns a complete Supabase CLI stack on an internal isolated Docker network and verifies the exact containers, image and managed schemas/roles. Database and Kong ports must not be published directly. Only the digest-pinned, run-owned proxy may publish the API and database ports, both on IPv4 loopback. A generic PostgreSQL/pgvector container is accepted only for the disposable Neon-local profile and must never be reported as Supabase Database evidence.

Run the nine drills only against a disposable target owned by the runner. Replace <exact-id> with one of supabase:supabase, supabase:better_auth, neon:supabase or neon:better_auth:

bash
pnpm run qa:operations:controlled-database-failure -- --composition=<exact-id> --write-local-evidence
pnpm run qa:operations:controlled-auth-failure -- --composition=<exact-id> --write-local-evidence
pnpm run qa:operations:database-credential-rotation -- --composition=<exact-id> --write-local-evidence
pnpm run qa:operations:auth-secret-rotation -- --composition=<exact-id> --write-local-evidence
pnpm run qa:operations:pool-drainage -- --composition=<exact-id> --write-local-evidence
pnpm run qa:operations:backup-restore -- --composition=<exact-id> --write-local-evidence
pnpm run qa:operations:durable-job-recovery -- --composition=<exact-id> --write-local-evidence
pnpm run qa:operations:observability-redaction -- --composition=<exact-id> --write-local-evidence
pnpm run qa:operations:performance -- --composition=<exact-id> --write-local-evidence

Run one composition at a time. The lifecycle lock rejects overlapping runs, and every command must stop and attest cleanup of the exact target it created.

After an authorized run, write one sanitized receipt per scenario. Its exact top-level fields are version, environment, provider, scenario, status, observedAt, metrics and artifactDigests. environment is exactly local; provider is the exact composition id; status is passed, failed or skipped. Every digest binds a repository-relative raw measurement artefact. Receipts cannot contain URLs, DSNs, credentials, bearer tokens, provider payloads or free-form errors.

Register the receipt path and SHA-256 in the matching report, then run:

bash
pnpm run qa:operations:check

The checker verifies the fixed scenario matrix, aggregate outcome, file paths, digests, receipt binding and scenario-specific metrics. It does not execute a drill, assert that a self-reported observation is true, sign evidence, modify the compatibility manifest or authorize runtime activation. A failed or skipped scenario keeps the local report incomplete.

Certification boundary

Database/Auth certification is local_only. Operational success does not certify a composition by itself: keep activation blocked until the final compatibility suite records exact versions, the disposable execution topology, connection budget and all required local passing evidence in the repository-owned compatibility manifest. This certificate does not attest provider-hosted control planes, service tiers, regional latency, autoscaling or backup guarantees.

Evidence producer and review format

The owner of an isolated Q2/Q3 run is the evidence producer. A static audit, init success, copied console output or an unexecuted test name is not a run. The producer records only outcomes actually observed for one exact composition and scope. There is intentionally no command that converts missing results into a passing report.

For each run, write one sanitized JSON artifact below ops/provider-certification/<database>-<auth>/. The verification artifact has exactly this shape; the values below describe the schema and are not evidence:

json
{
  "schemaVersion": 1,
  "selectionId": "supabase:supabase",
  "scope": "local",
  "recordedAt": "<ISO-8601 UTC timestamp>",
  "runId": "<bounded non-sensitive run id>",
  "testResults": [
    { "id": "transaction-context", "status": "passed" },
    { "id": "rollback-timeout", "status": "failed" },
    { "id": "prepared-statements", "status": "skipped" }
  ]
}

scope is exactly local; every test id comes from requiredCertificationTests. Report failed and skipped without rewriting them as missing. The matching verification.local entry records the aggregate status, timestamp, exact outcomes, repository-relative artifact path and SHA-256 digest. not_run owns no timestamp, result, path or digest. partial contains an incomplete set or a skip and no failure; failed contains at least one failure; passed requires all required tests to pass.

Run node scripts/check-database-auth-provider-compatibility.mjs after editing the report. The checker validates paths, real files, digests, safe fields and the binding to composition, scope and outcomes. SHA-256 binds reviewed bytes; it is not a provider signature and must never be described as one.

Partial verification never populates the final evidence field and never changes certificationStatus, activation or initStatus. Remove only a blocking gate whose condition is actually resolved. The compatibility suite remains incomplete until the local scope passes every required test.

The release verification workspace must include every repository-relative JSON file named by evidencePaths. The release checker verifies their exact SHA-256 digests; missing, moved or altered evidence fails pnpm run qa:release. Application startup does not read these files or Git metadata. Standalone Docker images retain the runtime provider, credential, SQL-role and Account/RLS checks.

Only after the local scope passes may the Q4 owner prepare the separate final certification artifact. It contains the exact versions, deployment topology, connection budget, a certified runtime pooling mode, server-revalidated Auth and all passing test results. Its exact top-level fields are schemaVersion, selectionId, recordedAt, runId, versions, deployment, connectionBudget, runtimeMode and testResults; the checker defines and validates every nested field. Direct migration_operator connections cannot be used as runtime evidence. Certification also requires every registered threat to be reviewed and marked mitigated. A reviewed change then sets the final artifact path/digest, clears blockers and changes the composition to certified, allowed and available. The offline release verifier verifies that future local path: Q4 receipt bytes and SHA-256, canonical runtime digest, complete Auth revocation, the certified 24/24 performance matrix, and all nine Q3 receipts for the same composition. Incomplete reports keep release approval blocked; they do not prevent an implemented runtime from starting with valid configuration. The compatibility manifest remains the certification record.

Threat review is also fail-closed. The typed local-only registry at tests/manifests/database-auth-threat-evidence.json defines the exact claims for T01-T09. Its checker dereferences threat-scoped repository paths, verifies the artifact and runner bytes against SHA-256, binds local provenance and time bounds, and requires the compatibility status to match. Read current completion from that registry and its checker, rather than a historical audit count. T05 binds the actual build files and must be renewed after they change. T09 requires an externally signed supply-chain snapshot and an independently approved trust anchor. Missing or stale evidence keeps release blocked. Flipping a threat to mitigated manually cannot satisfy either the checker or the synchronous runtime consumer. Passing these threat proofs does not automatically certify the four Database/Auth compositions.

Reconciling an interrupted AI stream

New chat requests allow one unresolved operation per Account. A pending/failed debit or unknown provider outcome blocks further requests; topping up alone does not clear failed settlement. The output cap is aiConfig.maxOutputTokens. The settle-ai-usage job reports stalled admissions after aiSettlementConfig.stalledAdmissionSeconds. Inspect the operation with your provider before retrying. Never infer a bill from generated text.

For an admission with no committed receipt, use the administrative job runner with operation_id and reconcile_admission: input_tokens, output_tokens, provider_reference_digest (SHA-256 of the provider evidence), and provider_usage_confirmed: true. The financial capability requires a stale admission, records the evidence digest, and stages the exact total for settlement. The same values replay safely; conflicting recovery values fail. Confirmed zero usage releases the fence without a debit. Do not use this operation merely to unblock a customer without checking the provider outcome. For an existing failed receipt, correct the cause and run retry_failed: true for that operation instead. Account deletion cascades admissions; user deletion removes actor attribution.