Site

The fresh-schema bootstrap checks SQL assembly before starting its disposable stack. Edit the owning source listed in database/manifest.json and validate with pnpm run db:schema:check. Bootstrap, guardrails, RLS/type audits and Neon diagnostics all assemble SQL in memory; they require no exported files. The SQL workflow guide explains source ownership, optional manual export and installer error mapping.

The QA system is a hermetic, account-aware release gate. It tests the real Next.js application against disposable local services, provisions explicit account and role fixtures, replaces only remote provider boundaries, and removes the generated state when the run ends.

Use this page when you need to run the gates, understand what they prove, or add coverage for a new route. For the lower-level Vitest, Playwright, Lighthouse, and audit catalog, see Testing. The complete implementation ledger and remaining work live in QATests.md at the repository root.

Current baseline — 15 August 2026

Hermetic local QA Phases 0–6 are complete. Phase 7, which adds representative browser-driven admin CRUD journeys, is in progress. Jobs, handlers, blog categories, and localized tags are covered; CMS/blog pages, media, and cron UI journeys remain open. Real-provider and remote-runner canaries are optional and outside the disposable local gate.

Start here

The repository requires its pinned Node 24.18.1 and pnpm 11.20.0 versions. Docker must be running because the main gate starts a local Supabase stack.

bash
# Recommended comprehensive local gate
pnpm run qa:local:gate

# Inspect or stop the local QA stack when troubleshooting
pnpm run qa:local:status
pnpm run qa:local:stop

qa:local:gate is self-contained. Focused qa:local:* profiles use the same guarded lifecycle and are faster when you are working in one domain. Direct test:e2e:* commands generally expect an already running application and, for authenticated suites, pre-provisioned validated personas.

Fresh SQL installation proof

Run pnpm run qa:local:bootstrap to assemble and install the core and CMS batches from database/manifest.json into a separate, empty local PostgreSQL cluster as Supabase's non-superuser postgres administrator, using the same whole-batch driver calls and schema acceptance queries as initialization. It then checks organization/document and notification erasure, subscription reads, public changelog and runtime role isolation, and provisions and preflights a real shared app_runtime credential using the initialization helper. Credentials stay in memory. The completion summary counts executed and skipped assertions from validated TAP results. Catalogue tests also compare registered files with their SQL plans.

After the SQL assertions, bounded child-process batches run the repository proof files registered in PRISMA_PROOF_BATCHES in scripts/qa/fresh-schema-bootstrap.mjs. The catalogue and configured deadlines are authoritative; each batch reports its file count and elapsed time. Coverage includes public/CMS reads and administration, jobs, analytics, billing, Account/identity paths and privileged workers. It checks published visibility, projection, ordering, limits, denied writes and admin RPCs, read-only transactions, actor clearing, actual role identity, and timeout rollback followed by reuse of the same database connection. Block fixtures also cover nested localized JSON, nullable timestamps, key batches and missing keys; existing seeded blocks are preserved and only run-owned fixtures are deleted. Taxonomy checks cover category/tag ordering, limits, nullable metadata, forbidden association metadata access, and repeatable installation of the reviewed index. CMS page checks cover fully unpublished rows, exact-language false overriding same-language fallback, cross-language publication refusal, and full SEO maps. Paginated blog checks cover category/tag filters, counts, stable ties, locale publication before pagination, and published-only association keys. More than 100 draft and other-language fixtures exercise filtering before limits. The repository proofs run sequentially to isolate shared fixtures. Blog detail checks additionally cover full content/SEO maps, category and tag projections, null taxonomy metadata, stable tag order and overflow rejection. Concurrent unpublishing verifies that the detail transaction keeps one coherent snapshot and the next transaction sees the new publication state. Sitemap checks cover the five-field projection, publication and SEO exclusions, independent CMS/blog quotas, stable traversal, and transaction/role isolation. Admin block read checks cover list/page bounds, nullable timestamps, stable ties, ID lookups and count/page consistency under concurrent changes. Blocks remain global public content; this proof does not certify private admin access.

The block mutation proof exercises create/update/delete through the same Prisma runtime and private SQL capabilities. It checks active platform-admin authority, rejected actors and direct writes, atomic previous-key receipts, concurrent edits, rollback, timeouts and connection reuse. The function owner cannot change identity or profile records: its narrow lock privilege is fenced by restrictive RLS.

Admin page reads verify draft access for active platform administrators through private read-only functions. Public draft access stays denied. The proof checks full page details, bounded pagination, stable ordering, count consistency, snapshot behavior under concurrent changes and revocation on a new transaction. The block/page mutation and page read proofs share the same owned Auth/consent fixture setup and exact restoration of temporary role grants.

The page mutation proof exercises page/tag writes in one transaction, reference failures without partial changes, missing/conflicting slugs and actual previous slugs under concurrent renames. Independent language publication edits preserve each other. It checks rejected actors, direct writes and owner escalation, authority revocation under locks, malformed-receipt rollback, real SQL timeouts and pool reuse. Input checks include localized values, tag limits and the UTF-8 payload bound before JSON conversion. These SQL fixture checks do not claim a browser sign-in ceremony or final composition certification.

Administrative taxonomy checks cover more than 1,000 associations, including draft blog pages and ordinary CMS pages. Category counts include blog pages; tag counts include all associations. They verify the 200-row list limits, stable ordering, complete nullable metadata, private draft tag IDs and rejection of a 201-tag selection. Each list's metadata and counts share a read snapshot; the proof also checks several repository reads within one transaction during concurrent changes. An administrator revocation blocks the next transaction. The page and taxonomy proofs use the same actor read-only context, including real timeout rollback and connection reuse checks.

The taxonomy mutation proof checks category/tag CRUD, preserved omitted fields, missing and conflicting slugs, and actual previous slugs under concurrent edits. Category deletion clears page references and tag deletion removes associations; both preserve pages. It verifies active administrator authority and revocation under locks, denied direct writes and role escalation, and rollback after a malformed receipt, exception or real database timeout. The writer has no direct page or association privileges; deletion effects use the existing foreign keys.

Changelog administration checks private drafts, eight-field bounded lists and nine-field details, repeated version labels, partial updates and publication timestamps. Concurrent publication and content edits must preserve both changes. Read-only snapshots keep their observed state; revoked administrators are refused on the next transaction. Write authority locks, invalid receipts, callback errors, real SQL timeout rollback and pool reuse are exercised. The previous maintenance functions and owner must be absent, and neither new owner can modify identity or profile data. Public publication restrictions remain in force.

The maintenance proof checks the public scalar and administrator writes through the same Prisma pool. It verifies that each successful command changes the flag and records exactly one audit row, including repeated values and concurrent toggles. Invalid receipts, callback failures and real SQL timeouts roll back both records. It covers authority revocation under locks, actor clearing, denied direct writes and role escalation, and removal of the old public/private RPCs. Runtime preflight must reject disabled RLS on either settings or audit records. The original flag value and timestamp are restored, and only fixture-owned audit rows are removed. The website cache and proxy behavior have separate unit tests.

Administrative settings reads verify the complete 200-key bound, explicit overflow failure, UTF-8 ordering and existing text values longer than the mutation limit. The settings list and ten fixed SQL counts share one read-only snapshot, including under concurrent changes and administrator revocation. Counts match independent SQL queries; the two displayed AI counters use the same result. The proof checks minimal owner privileges, actor isolation, real timeouts and pool reuse, and preflight refusal when RLS is disabled on any of the eleven participating public tables. Existing settings are preserved; only fixture-owned keys are removed.

The administrative settings mutation proof checks both setting changes with their audit, canonical administrator authority, rejected actors and direct writes, malformed-receipt and audit-failure rollback, real timeouts and pool reuse. Global purge checks run inside rollback transactions to preserve existing rows. They verify message/session counts, the audit retained after an admin-log purge, and concurrency under the reviewed table-lock order. Setting fixtures restore the original values and timestamps, and cleanup removes only owned records.

The cron URL proof exercises the startup repository with a fixture Vault secret and jobs scheduled outside the test window. It commits the unchanged-URL path; all changes to global schedules run inside forced-rollback transactions. It checks bounded reconciliation, rejected native scheduling, concurrent replicas, lock and statement deadlines, receipt failure, pool reuse and denied access to secrets, cron commands and privileged roles. Original schedules are compared by hash without exposing command text; original settings and timestamps are restored. This proof does not send HTTP requests or certify an external scheduler.

The cron administration read proof uses canonical actor fixtures and distant, non-network cron entries. It verifies complete lists at the 200-entry limit, overflow rejection, masked commands for every job, real run statistics and RepeatableRead behavior across concurrent changes and administrator revocation. SQL coverage rejects ambiguous job links and inspects the deployed guards for missing native tables. Actual removal of extension tables is skipped because the non-superuser installer does not own them; repository tests verify the unsupported result mapping. The database portability work owns this gap until certification on a target without the native scheduler. Privilege and timeout tests verify that the runtime cannot read raw commands or secrets, change schedules or assume an owner role. Only owned fixtures are removed; existing schedules and execution history remain intact.

The cron mutation proof commits toggle and unschedule only on owned distant fixtures. Every global purge is rolled back, including its audit. It checks exact audit attribution, missing targets, overflow, invalid receipts, callback errors, native/audit rollback, concurrent commands and authority revocation after lock waits. SQL also verifies that an audit trigger silently returning NULL cannot commit a mutation without its audit. Real database timeouts are followed by connection reuse checks. An actual foreign-owned native target is skipped because creating that fixture requires unavailable native privileges; the deployed guard and unsupported mapping remain covered. The portability work owns that missing evidence until target certification, without widening grants.

The jobs observability proof creates 1,005 owned runs to verify complete SQL averages beyond the application page size. It checks the seven-column metadata projection, exact pagination totals, status filtering, indexed ordering and microsecond timestamps, including ties and nulls. Signed durations and the exact 24-hour boundary follow the database schema and transaction clock. Concurrent run changes and administrator revocation verify snapshot isolation; denied payload access, real timeouts and connection reuse exercise the runtime boundary. Fixtures have no cron expression and do not invoke handlers or HTTP requests.

The jobs list proof checks empty, 200-row and overflowing 201-row results both globally and after filters. Owned fixtures exercise nullable enabled status, Unicode ordering that differs between UTF-8 and UTF-16, exact UTC timestamps and invalid dates. The eight-field projection excludes configuration, native cron identifiers and unused settings. Authority denials, concurrent job changes and admin revocation, forbidden role/column access and real timeout/pool reuse verify the shared runtime. Original jobs and native schedules remain unchanged.

The handler list proof checks the eight-field display projection without headers, authentication values or timestamps, enabled-only filtering, Unicode ordering, nullable fields and executor array shapes. It verifies the 200-row and two-MiB receipt boundaries, active administrator authority, concurrent snapshots, revocation on the next transaction, real timeout and pool reuse. Fixtures do not execute handlers or contact webhook URLs. The job selector passes only its four required fields to the client; handler editing and credential updates have their own proof below.

The overview usage proof verifies five aggregate metrics in one canonical administrator read-only snapshot: null token values, inclusive timestamp boundaries, totals beyond int32, future rows, role denials, snapshot consistency, revocation and actual SQL cancellation followed by pooled connection reuse. Malformed and unsafe numeric receipts fail closed. The bounded fixture does not claim to aggregate millions of rows to cross the JavaScript safe-integer limit; that boundary is covered through receipt validation. Revenue queries remain a separate implementation and are not certified by this usage proof.

The daily Analytics proof verifies selected UTC chart days through the shared Prisma repository: leap days, microsecond cutoffs, offset equivalence, sparse empty results, excluded unselected dates and fractional costs in units of 1/10,000 cent. It covers token sums beyond int32, invalid dates and numeric receipts, actor and column denials, concurrent snapshot reads and revocation, actual SQL timeout and reuse of the same connection. Oversized JavaScript totals are tested by corrupting the receipt after a real query, not by claiming a multi-million-row aggregate fixture. Other Analytics aggregates remain separate.

The model/agent Analytics proof reads both categories in one snapshot. Real fixtures cover each category at 200 groups and refusal at 201, names of 256 and 257 UTF-8 bytes, Unicode ordering, empty versus unknown names, fractional costs, wide token totals, microsecond cutoffs and future requests. It also checks canonical authority, minimal column access, concurrent writes and revocation, real timeout, pool reuse and fixture cleanup. Corrupted receipts after real queries verify unsafe numeric totals and missing groups. The JSON byte ceiling is additionally covered by repository tests and the exact SQL attestation; the local database fixture does not exercise an overflow caused by escaping control characters in otherwise valid-length names.

The user/Account rankings proof reads both top-ten lists in one snapshot. It checks deterministic UUID ties, cross-Account attribution, rejected unattributed writes, all Account types, exact microsecond cutoffs and fractional costs. Labels are selected after ranking: oversized excluded labels are ignored, while selected emails and names enforce 320/512 UTF-8 bytes. A real control-character fixture exceeds the 32-KiB escaped JSON receipt limit. The proof also checks canonical authority, minimal email/name access, concurrent label changes and revocation, actual timeout and pool reuse. Unsafe numeric receipts are injected after real queries. Defensive missing-label fallbacks and userless read results have separate unit coverage: schema constraints and erasure guards prevent constructing those rows through current application writes.

The administrative payment-list proof checks stable pagination, exact totals on empty pages, normalized payment states and amounts, and Account owner labels read after pagination. It exercises canonical administrator denials, forbidden columns and role escalation, concurrent changes and revocation, rollback, real SQL cancellation and reuse of the same connection. Only selected labels count toward field limits; unsafe selected money and oversized results fail explicitly. Parser-only receipt corruption is labelled separately from real SQL fixtures. The nullable owner email is exercised through an owned Account without an owner. The defensive null Account-name result has unit coverage; the current schema requires a non-null Account name and does not permit that physical fixture.

The administrative subscription-list proof covers the current-subscription filter, stable pagination and totals, Account labels and safe credit balances. It uses separate owned Accounts for the 200-row page because the schema permits only one current subscription per Account. Superseded rows remain excluded. Administrator denials, field and result limits, snapshots, revocation, rollback, actual SQL cancellation and connection reuse use the common Prisma repository. Malformed receipt injection is distinguished from physical database fixtures.

The single-subscription proof reads current and historical IDs through both methods of the shared repository. It verifies absence after authorization, exact requested IDs, bounded Account decoration for GET and label-independent command preparation. Canonical identity and privileged-role denials, snapshots, revocation, rollback and a real five-second SQL cancellation for each method use the same bounded connection pool. Unsafe or oversized receipt injection is identified separately from physical fixtures; real credit changes use the existing RPC.

The membership-role proof reads exact actor/Account pairs through shared Prisma, including owner, admin and member roles, custom and nullable slugs where the schema permits them, and missing memberships. It checks canonical revocation, absence of platform-admin bypass, strict projections, the existing UTF-16 slug limit, snapshots, concurrent actors, rollback and real SQL cancellation with connection reuse. Billing-role policy helpers retain their current behavior.

The subscription-control creation proof replaces the standalone audit proof. It checks canonical admin and Account billing authority, configuration-owned role rejection before commit, exact command attribution and atomic audit, sequential and genuinely concurrent replay, conflicting intents and plan changes. Source timestamps retain microsecond precision and their NOT NULL constraint. End-trial commands freeze the active commercial identity and future trial boundary. Invalid receipts, explicit rollback and real SQL cancellation leave neither command nor audit; connection reuse and foreign-row hashes are checked. Subscription state, payment proof and credits remain unchanged. Final composition certification remains separate.

The billing statistics proof exercises subscription and one-time payment aggregates through the shared Prisma repositories. It covers canonical administrator authority, cross-Account fixtures, currency separation, commercial version matching, fractional annual-plan MRR, refunds and disputed payments. It also checks calendar boundaries, minimal column privileges, denied role escalation, consistent snapshots, revocation, rollback, actual SQL cancellation and connection reuse. A real unsafe payment aggregate fails explicitly; additional malformed or oversized receipts are injected after real queries and are reported separately from physical financial fixtures.

The summary proof covers historical counts and both comparison periods in one administrator snapshot. It checks exact period boundaries, positive-only rounded latencies, status counts, wide token totals and fractional costs, along with cross-Account fixtures, denied actors and sensitive-column access. Concurrent changes must remain invisible inside the snapshot and appear on the next read; revocation must deny the next transaction. A pool configured for thirty seconds still cancels the summary query at the five-second SQL bound and reuses the connection. The jobs period-statistics RPC remains outside this Analytics proof.

The credit totals proof exercises sign-based ledger aggregates across Accounts, using credit RPCs for fixtures. It covers totals beyond int32, inclusive microsecond cutoffs, offsets, future and null dates, all credit sources, denied actors and sensitive columns, coherent snapshots and revocation. The pool starts with a 30-second SQL timeout; the read helper applies the five-second bound before attestation and verifies actual cancellation followed by reuse of the same connection. Unsafe or oversized receipts are injected after real SQL, not represented as physical large-ledger database fixtures.

The handler deletion proof exercises the real shared Prisma transaction: references and idempotent absence, serialized concurrent job writes, active canonical authority, mandatory audit and rollback, malformed receipts, deadlines and pooled connection reuse. It checks the minimal owner/bridge privileges and retirement of all old jobs-admin transport capabilities. Job function names remain open to code handlers; no future-reference foreign key is claimed.

The handler detail/write proof verifies a twelve-field editor projection with presence flags instead of credentials or headers, plus atomic creation and updates with immutable names. Omitted values remain unchanged, including SQL NULL and JSON null headers; replacement and explicit clearing are covered. It checks strict inputs, complete receipt bounds, concurrent writes, canonical authority under locks, mandatory audit rollback, invalid receipts, actual timeout and pool reuse. SQL forbids direct authenticated credential reads and creation/updates. No handler or outbound webhook is executed. UI interaction tests and an isolated Chromium preview complement this proof; they do not certify a full application or real TOTP journey.

The job detail proof verifies a coherent job/run snapshot, explicit absence, 20/50 recent-history limits and exact UTF-8 byte boundaries for the two-MiB receipt ceiling. Stored JSON scalar, array, object and null values are preserved. Oversized or invalid unselected runs cannot block a bounded history sample. Canonical authority, cross-job isolation, concurrent changes, timeout and pool reuse are exercised with real runtime credentials. The retired owner cannot read run rows; operator fixture deletion still verifies the run cascade. Fixtures never schedule or execute a handler.

The job deletion proof verifies native cleanup before the job/run cascade, an idempotent absent result and exactly one minimal audit per actual deletion. It covers stored and canonical-name schedules, ambiguous links, rollback when audit or receipt validation fails, concurrent deletion, run-statistics updates and authority revocation around lock waits. Read-only refusal, real timeouts and connection reuse exercise the runtime boundary. Direct table deletion and the retired transport function are unavailable. Native fixtures are owned, distant and non-networking; they never execute a handler. The existing lack of a foreign-owned native fixture remains an explicit certification limit.

The job creation/update proof checks exact configuration JSON, preserved omitted fields, duplicate-name races and disjoint concurrent patches. Native scheduling and audit must commit with the job or roll back together; metadata-only updates leave schedules untouched. Two-candidate fixtures verify obsolete-reference cleanup and complete unscheduling; shared references fail without effects and an independent schedule is preserved. It verifies that only the inert writer bypasses the old automatic trigger, with fixed native reconciliation instead. Audit failures, silent audit suppression, malformed receipts, authority revocation, run-statistics contention and real timeout/pool reuse are exercised. Owned Vault and URL fixtures are restored, schedules are distant and no handler or network request executes.

The command requires local Docker and reserves project boilerplate-stack-bootstrap-qa with database port 55522. Existing resources for that project cause refusal; it never resets the ordinary QA stack. It uses a temporary configuration without application migrations, seed data, linked project metadata, or inherited provider credentials, and removes its own project and temporary files on completion or handled failure. An abrupt process termination can leave resources or its temporary lock behind; inspect those exact bootstrap resources before cleanup and retry.

This is a standalone SQL/runtime proof. It does not execute the interactive wizard, browser onboarding, or satisfy final composition certification.

How the disposable gate works

  1. Validate the target before creating privileged clients. Zod-validated QA configuration rejects production or indexable targets, hosted Supabase origins under local intent, mismatched approved origins, privileged publishable keys, live Stripe keys, and unsafe email/provider settings.

  2. Start clean local infrastructure. The runner starts the pinned Supabase CLI stack on loopback, installs the assembled core and CMS batches, loads deterministic seed data, and starts the Next.js application on port 3777 with test-only credentials.

    The disposable profile exercises the supabase:supabase composition without certifying it. It installs the same assembled fresh schema used by project initialization, then provisions and preflights distinct app_runtime, app_maintenance, app_financial_worker, and app_platform_admin credentials. Those credentials stay in memory and replace inherited values. Session finalization stays enforced with v2 markers signed by a temporary per-run key. The compatibility matrix remains blocked until the complete four-composition local_only gate records final evidence.

    Fresh registrations remain pending until the real post-auth finalizer verifies consent and personal Account evidence and activates canonical application identity. Business reads use the direct app_runtime login; GDPR operations use the separate maintenance capability. Do not weaken RLS or substitute an operator read to hide a finalization failure.

    The shared composition provisioner checks role attributes and exact capability memberships while every login is still disabled, then enables credentials and preflights each connection. It does not request superuser-only options or silently normalize unexpected privileges. Provisioning diagnostics expose only closed stage names, never SQL errors or credentials.

    The runner supplies a per-run local provider proof to validate ownership of its disposable QA target: matching proof/readiness UUIDs, non-indexable loopback application and Auth endpoints, loopback PostgreSQL, exact run-owned login names, distinct generated passwords, and one database. These safeguards belong to the QA lifecycle. Ordinary runtime validates configured providers, credentials and SQL authority without reading release evidence; pnpm run qa:release separately enforces certification. QA flags do not turn a hosted target into a disposable local target.

  3. Prove the database contract first. pgTAP and database lint verify RLS, role and capability ACLs, constraints, function permissions, Account isolation, credit-ledger behavior, selected scheduler objects, and source/installed-schema consistency.

  4. Provision account-centric personas. Setup creates explicit profiles, Accounts, Memberships, roles, access states, and isolated browser storage states. Mutable tests receive worker-owned users and Accounts so parallel projects cannot share state accidentally.

    The four baseline personas pass through canonical session finalization only after consent and personal Account evidence are ready. Preparation uses the validated disposable database, then verifies each profile through the application runtime before committing. It does not write an active status directly or bypass application session checks.

  5. Run application contracts. Vitest checks domain behavior; Playwright drives API and browser journeys through real routes; axe-core, viewport, visual, cross-browser, PWA, performance, and reliability checks run in their assigned profiles.

  6. Collect bounded diagnostics and clean exactly what the run owns. Failures include a reproduction command and secret-safe evidence. Privileged runs do not persist traces, screenshots, videos, raw URLs, headers, bodies, or credentials. Cleanup removes exact worker/run resources, browser state, temporary email/provider files, and the disposable local volume—even after handled failure or interruption.

    After stopping app processes, cleanup destroys the exact run-owned stack and its ephemeral operational credentials. A provisioning, preflight, or cleanup failure keeps the gate failed; no credential is written to an environment file or recovery journal.

Command map

Self-contained local profiles

CommandUse it for
pnpm run qa:local:gateThe default disposable gate. Runs every retained local QA profile and is the best pre-merge proof.
pnpm run qa:local:auth-deepPublic registration/onboarding, admin-created setup and delivered login by OTP and link. Each spec gets a fresh app process and real persona setup; email is recorded only in a disposable local directory. Other Auth flows use their dedicated suites or the broader gate.
pnpm run qa:local:privatePrivate overview, actor-scoped GDPR export with own/peer/foreign checkout fixtures, profile/consent changes with three atomic preference events, unchanged signup terms and exported evidence, personal deletion scheduling/cancellation, plus deterministic chat/SSE and documents/RAG: accounting, resilience, Storage lifecycle and tenant isolation, using the member project and real persona setup. Cancellation verifies retained history and an active user, not terminal worker erasure.
pnpm run qa:local:orgOrganization owner/admin/member routes, APIs, mutations, billing boundaries, accessibility, and responsive behavior.
pnpm run qa:local:adminPlatform-admin pages and APIs, CRUD/actions, validation, denials, audit rows, and error-log redaction.
pnpm run qa:local:billingB2C/B2B billing, checkout/portal, signed Stripe webhooks, replay/order/idempotency, refunds, disputes, trials, and credit accounting.
pnpm run qa:local:billing:lemonEquivalent Lemon Squeezy checkout, portal, signed webhook, lifecycle, refund/risk, replay, and convergence coverage through a secret-free deterministic boundary.
pnpm run qa:local:billing:paddleEquivalent Paddle checkout/Paddle.js, portal, signed webhook, lifecycle, adjustment/risk, replay, and convergence coverage through a secret-free deterministic boundary.
pnpm run qa:local:phase6Production-build non-functional gate: accessibility, responsive, visual, cross-browser, Lighthouse, PWA/offline, privacy, query budgets, and bounded load.
pnpm run qa:local:phase7Focused browser-driven admin CRUD journeys with persisted database and Storage postconditions.

The registration and admin-setup journeys also verify the current admin deletion contract: preserve the personal Account, receive 202, and prove a matching durable pending request plus profile suspension while the native Auth identity still exists. Their exact run-owned cleanup is separate from that assertion. They do not run the global deletion worker or certify terminal erasure.

Visual baselines require deliberate review

Run pnpm run qa:local:phase6:update-visuals only when you intend to inspect and approve every changed screenshot. The normal gate compares committed baselines and never auto-approves a visual diff.

Lower-level and already-provisioned commands

V2 has no identity-upgrade, equal-UUID or backfill profile. Provider evidence starts from an empty disposable database and exercises the selected fresh composition. Historical migration runners must not be used to claim V2 compatibility or final local-only certification.

The billing profile collects the existing guest-checkout and guest-claim refusal contracts for member, administrator and owner personas. These API-boundary tests do not certify a completed paid guest transfer or the continuation worker.

pnpm run qa:local:auth-user-profile-boundary requires the disposable stack from qa:local:start and qa:local:reset. It checks the production profile port and SDK with one owned pending Auth fixture: pre-cancellation and wrong issuer make no request; editable metadata cannot confirm email; native confirmation yields only the reduced profile; deletion yields typed absence. Fixture cleanup runs in finally. Run qa:local:test-db and qa:local:lint-db for SQL contracts and always qa:local:stop afterwards, including on failure. No environment file or remote project is used. This is not a whole-worker guest-claim or alternate- provider certification.

pnpm run qa:local:auth-user-erasure-boundary requires a running, freshly reset disposable stack from qa:local:start and qa:local:reset. It exercises the real Auth erasure port, configuration, service client and SDK with an owned pending fixture: pre-cancelled and wrong-issuer calls perform no I/O; deletion confirms the exact subject; replay confirms typed absence without another DELETE. The fixture is cleaned in finally; always run qa:local:stop afterwards, including on failure. No environment file, remote target, SQL bypass or external OAuth provider is used. This does not execute the whole GDPR worker or certify another database/Auth combination.

Command familyPurpose
pnpm run testAll Vitest unit and domain contracts.
pnpm run qa:local:test-db / qa:local:lint-dbpgTAP and database lint against the local QA stack.
pnpm run qa:local:start / qa:local:resetStart or rebuild the owned disposable stack using complete in-memory source batches, acceptance checks and the QA seed. CLI migration/seed loading is disabled; no pending-migration command is supported.
pnpm run qa:local:subscription-benchmarkMeasure the shared Prisma subscription repository under bounded synthetic load on the running QA stack. It provisions and cleans run-owned fixtures, writes aggregate local measurements after cleanup, and does not alone certify a Database/Auth composition.
pnpm run test:e2e:auth / test:e2e:oauth / test:e2e:apiAnonymous Auth, deterministic OAuth, and protected-API authorization contracts.
pnpm run test:e2e:private / org / admin / billing / personasFocused Playwright suites against a validated, provisioned target.
pnpm run test:a11y / test:responsivePublic accessibility and required viewport coverage.
pnpm run test:cross-browser / test:visualBrowser compatibility and committed snapshot comparison.
pnpm run check:payment-catalogSelected-provider credentials, mode coherence, and local paid-catalogue completeness/format/collision preflight.
pnpm run smoke:payment-catalog / smoke:payment-webhookSigned, partial Lemon/Paddle sandbox evidence. These commands never replace the complete staging gate.
pnpm run test:stagingBounded readiness probes for the exact non-indexable origin approved by STAGING_APPROVED_ORIGIN.

What the system proves

The automation is layered so a failure is caught as close as possible to its source:

LayerToolingMain proof
Domain and configurationVitestValidation, redirects, permissions, calculations, serializers, provider parsing, and failure semantics.
DatabaseSupabase CLI + pgTAPRLS, grants, constraints, triggers, RPCs, migrations, account isolation, and atomic ledger behavior.
APIPlaywright request/API suitesAuth, role and account scope, CSRF, rate limits, Zod failures, sanitization, idempotency, headers, and safe errors.
Browser journeysPlaywrightReal user behavior across public, private, organization, admin, Auth, billing, CMS, and account flows.
Inclusive UIaxe-core + PlaywrightWCAG 2.2 AA, keyboard behavior, focus, 200% reflow, reduced motion, touch targets, and five required viewport widths.
Release qualityPlaywright + Lighthouse CIVisual regression, Chromium/Firefox/WebKit, PWA/offline, consent boundaries, response and query budgets, and bounded load.

Every configured locale receives navigation/render smoke coverage. Functional journeys use a deterministic default locale, with locale-sensitive assertions repeated across French and English regional variants.

Accounts, personas, and isolation

The fixtures follow the product architecture: domain data attaches to an Account, not directly to a User. A persona is valid only when its profile, Account, Membership, role, billing/access state, and active-account selection are explicit.

The core local matrix includes ordinary members, workspace admins, workspace owners, and a platform admin, plus anonymous and specialized states such as zero credits, disabled access, pending invitations, unrelated workspaces, active subscriptions, and licenses. Tests use run and worker identifiers for mutable records, seed dynamic resource IDs instead of hardcoding UUIDs, and assert both wrong-role denial and cross-account denial.

This is why the gate can run state-changing projects in parallel without weakening tenant-isolation assertions or allowing one test to depend on another test's mutation.

Route manifests prevent silent coverage drift

Two typed manifests classify every page and route handler:

  • tests/manifests/pages.ts records the surface, persona, locale, feature gate, and expected page outcome.
  • tests/manifests/api-routes.ts records exported/effective methods, domain, security wrapper, access class, mutation/CSRF behavior, account scope, validation, idempotency, side effects, and expected errors.

The route-drift guard scans app/**/page.tsx and app/**/route.ts. A new or changed route fails until its manifest entry says how it is tested, redirected, feature-gated, or intentionally excluded with a reason. This turns missing QA classification into a review-time failure instead of a future production surprise.

Third-party boundaries stay deterministic

Core browser journeys do not stub the application's own API. They keep the real route, domain logic, Supabase writes, Storage behavior, security middleware, and browser UI, then replace only the remote dependency:

IntegrationLocal strategy
Supabase Auth, Better Auth, and emailExact-composition Auth profiles plus Mailpit and a run-scoped application-email recorder.
OAuthDeterministic provider-boundary contracts for provider, locale, PKCE, consent, cancellation, and callback failures.
StripeReal SDK/signature/application flow against a secret-free loopback HTTP boundary and signed fixtures.
Lemon SqueezyReal application/domain flow against a deterministic secret-free provider boundary and signed Lemon fixtures.
PaddleReal application/domain and Paddle.js handoff flow against a deterministic secret-free provider boundary and signed Paddle fixtures.
AI and embeddingsLoopback OpenRouter-compatible streaming/embedding provider with fixed usage/error frames; rejects missing ZDR, data-collection-deny, strict-parameter, attribution, or gateway-model routing.
Push, analytics, and TurnstileBrowser/service-worker/network contracts or official test adapters.

Optional staging checks use dedicated test-mode projects and pre-provisioned identities. Production permits only read-only synthetic checks such as health, public availability, TLS, headers, and a non-mutating login render.

Reading failures and reports

QA output identifies the affected surface, route or domain, persona/profile, pass/fail/skip/flaky count, and—where relevant—accessibility impact, performance delta, or visual diff. Stable prefixes such as AUTH-*, ACL-*, ORG-*, ADM-*, API-*, DB-*, BILL-*, AI-*, A11Y-*, and PERF-* make failures searchable.

Retries collect evidence; they do not convert flaky behavior into a pass. Start with the focused reproduction command printed by the failing gate, then use pnpm run qa:local:status to confirm the local stack if the failure occurred during setup. Callback setup reports only a fixed stage, optional HTTP status and closed redirect classification, not the URL, cookies or provider response. Do not copy credentials, callback URLs, browser state, or provider payloads into an issue. Run the full Vitest suite separately from a live local gate: the runner's cleanup unit test removes generated persona state.

Browser console findings include a normalized sourceRoute: a known API/page pattern or a fixed placeholder such as /[framework-resource]. This distinguishes resource origins without retaining console text, arguments, raw URLs, or stacks. The source label does not identify the underlying cause or grant an exception; existing page/kind allowance budgets and diagnostic limits remain blocking.

Adding QA for a change

  1. Update the page or API manifest whenever a route is added, removed, or changes its effective methods or access behavior.
  2. Put pure domain behavior in Vitest and database authorization/invariants in pgTAP.
  3. Add API contracts for validation, auth, CSRF, role, account isolation, idempotency, and documented failures.
  4. Add a browser journey for user-visible behavior; use accessible-name locators and assert persisted outcomes rather than implementation details.
  5. Give state-changing tests worker-owned Accounts and resources, with exact fail-closed cleanup.
  6. Run the smallest relevant focused profile while developing, then pnpm run qa:local:gate before merge.

No test should use fixed sleeps, depend on another test's mutation, auto-approve a snapshot, weaken a rate limit, or target production. Expected skips and allowlists need a reason, owner, and expiry instead of becoming permanent gaps.

Audit regression proofs

node scripts/qa/prove-ai-usage-settlement.mjs verifies stream admission, debt refusal, settlement and recovery on disposable local composition targets. node scripts/qa/prove-member-capacity.mjs checks commercial quotas and concurrent invitation acceptance; an optional composition ID limits either command.

After an offline Webpack build, phase6-cms in scripts/qa/run-local-gate.mjs reuses the exact QA_PHASE6_REUSE_BUILD_ID and checks first CMS reads plus actual administrative mutations on pricing and blog in every locale. Prepare standalone assets with the same runner helpers before reuse. This gate owns its disposable local stack and is never a command for a shared or hosted database.

The Webpack client budget gate reads client-entry-manifest.json, emitted by the build, to count the selected page and ancestor boundaries, then resolve their client reference IDs in the final Flight manifest. This includes chunks reassigned to neighboring entries during manifest merging without counting unrelated modules. Unknown references and unsupported parallel-route topologies fail the gate. Budgets remain in config/bundle-budgets.json; lazy chunks are excluded from the initial-file budget and still need interaction/performance testing.