Site

RAG Document Chat

Upload documents (TXT, MD, PDF) and chat with them using the Knowledge Base AI agent. Documents are chunked and embedded in bounded batches through the shared LlmClient using either direct OpenAI or OpenRouter, and vector similarity search runs inside Postgres through the shared Prisma repository with Account-scoped actor RLS over a pgvector HNSW index. The match cutoff is configurable via embeddingConfig.ragMatchThreshold (default 0.1 in config/ai.ts). Credits are deducted 1:1 with provider-reported embedding tokens. PDF text extraction uses unpdf.

📄

Document Upload

Upload TXT, MD, PDF files (10MB max). MIME, matching extension, and magic bytes are validated. The API acknowledges accepted work with 202 plus a Location header while processing continues.

✂️

Text Chunking

Automatic splitting into 1000-char chunks with 200-char overlap. Max 500 chunks per document to prevent DoS.

🔍

Similarity Search

Vector similarity runs inside Postgres through the shared Prisma repository with Account-scoped actor RLS over a pgvector HNSW index (<=> operator). Threshold is configurable via embeddingConfig.ragMatchThreshold (default 0.1 in config/ai.ts). The shared transaction installs the canonical actor context; queries enforce Account scope and only match clean, ready documents at their current ingestion revision.

🤖

Knowledge Base Agent

Dedicated RAG agent that retrieves relevant document chunks and injects them into the user message (not system prompt). Uses a pre-fetch pattern via prepareRAGContext() called from the stream route. Cites sources as [Document N].

RAG Agent Integration

The Knowledge Base agent (lib/ai/agents/rag/index.ts) is fully integrated into the AI chat system. It uses a pre-fetch pattern — the stream route calls prepareRAGContext() before streaming, which caches the document context. Then buildMessages() injects it into the user message.

  1. Before streaming, the agent checks the canonical actor's Account membership, embeds the query and confirms its token debit before retrieving matches
  2. Generates the query embedding through the selected direct/OpenRouter route without fetching every chunk into the application
  3. Runs vector similarity search through the shared Prisma repository (pgvector HNSW index, bounded results and actor-scoped RLS)
  4. Filters by threshold (configurable via embeddingConfig.ragMatchThreshold, default 0.1) and keeps top 5 matches
  5. buildMessages() injects cached context into the user message (not system prompt) wrapped in XML tags (<document_chunk>) for prompt injection protection
  6. The LLM responds using the document context, citing sources like [Document 1]

The agent is registered in lib/ai/agents/index.ts and appears in the chat agent selector dropdown with a rose/pink color scheme.

PDF support: unpdf extracts pages in a fixed Node child process, without a shell or application secrets. New ingestion checks credits before downloading or parsing. The parser enforces page, extracted-text, time and V8 heap limits; timeout or cancellation kills the process and waits for its exit. TXT and MD files share the byte, text and chunk limits. Production hosting must support Node subprocesses; the job route's output trace includes the worker and unpdf assets. The V8 heap limit does not bound total RSS or native buffers: use an OS/container memory quota when a strict total-memory ceiling is required.

Configurable Embedding Models

The embedding model used for document processing is configurable in config/ai.ts via the embeddingConfig object:

ModelDimensionsCost / 1K tokensNotes
text-embedding-3-small1536$0.002Default. Best cost/quality ratio.
text-embedding-3-large3072$0.013Higher quality, higher cost.
text-embedding-ada-0021536$0.01Legacy model.

To change the default model, update embeddingConfig.defaultModel in config/ai.ts. The model, model provider, transport, reported upstream provider, and exact token count are stored in document/credit metadata for traceability. Embedding calls use batches of at most 96 inputs, bounded concurrency, schema/dimension/index validation, and the selected AI_LLM_TRANSPORT. OpenRouter requests carry the same mandatory ZDR/data-collection-deny routing as chat.

Important: If you switch embedding models, existing document chunks will use the old model's vectors. You should re-process affected documents to ensure consistent similarity search results.

Credit Deduction on Upload

Ingestion stages validated embeddings and exact provider token usage in the private ai_document_ingestions table. This derived document data is unavailable to public, authenticated and application-actor reads. A document becomes ready only after the ledger confirms the Account, semantic ingestion key and token amount. Retries reuse the staged result and the same debit key; they do not repeat a completed, persisted embedding call. A process failure between the provider response and the first durable staging commit can still repeat that provider call. Insufficient credits produce an explicit failure, while transient accounting failures use the existing bounded job retries.

Staging is deleted on successful publication and cascades when the document or its Account is erased. Failed staging stays attached to the failed document for controlled recovery or document deletion. It is not a separate public document or a source of search results. The subject data export includes pending ingestion content and usage, scoped to documents uploaded by that subject in their authorized Accounts. It refuses partial output above 1,000 pending records or 10 MiB of serialized pending data. Query embedding debit failures abort RAG preparation and return no matches.

1 credit = 1 LLM token. Credits are deducted from the account based on the actual usage.total_tokens reported by the selected direct/OpenRouter embeddings route across all chunks of the document, summed and decremented in a single decrement_credits RPC after the loop completes.

Total credits = sum of usage.total_tokens for each chunk embedding call

For example, a document with 12,500 provider-reported embedding tokens costs exactly 12,500 credits. New ingestion requires at least aiConfig.minCreditsRequired (default: 100) before downloading or parsing. A resumed durable receipt skips that preflight and reuses its existing embeddings. If concurrent usage exhausts the balance, publication waits for confirmed settlement; the document is not exposed as ready with unpaid usage.

Config KeyDefaultDescription
embeddingConfig.defaultModeltext-embedding-3-smallDefault embedding model for new documents
aiConfig.minCreditsRequired100Minimum balance for the pre-flight check before any LLM call (chat or RAG)
aiConfig.embeddingRequestTimeoutMs60000Maximum duration of one embedding-provider request
documentConfig.pdfMaxPages200Maximum PDF pages before extraction
documentConfig.maxExtractedTextCharacters400000Maximum extracted characters for every document type
documentConfig.pdfParseTimeoutMs15000Deadline that terminates the parser process
documentConfig.pdfHeapLimitMb128V8 heap ceiling in MiB; not a total RSS ceiling
documentConfig.maxChunks500Chunk construction stops at this cap

Knowledge Base Page

The document manager is accessible at /private-dashboard/documents and appears in the sidebar as "Knowledge Base" (right below "AI Agents"). The page provides:

  • Drag-and-drop file upload with progress tracking
  • Real-time status polling (pending → processing → ready)
  • Document list with status badges, chunk counts, file sizes, and exact provider-token balance refresh
  • Delete with confirmation and full cleanup (chunks + storage)
Database tables: documents (metadata + status + embedding model) has Account-membership RLS; document_chunks (content + vector embeddings + model metadata) is service-only. Server routes establish Account membership before returning chunk summaries.

API Routes

All four handlers use the security wrapper's active application identity for Account membership and upload attribution. Reading and deleting remain shared among Account members, not restricted to the uploader. Client-supplied actor fields cannot choose that identity. Core document helpers are server-only, not public Server Actions. A missing actor returns a private 503; database/storage failures return generic 500 responses with synthetic diagnostics, and malformed multipart or delete JSON returns 400.

Upload keeps its MIME, extension, magic-byte and size checks, durably queues ingestion, and returns 202 without awaiting completion. Deletion durably queues removal of the Account-scoped storage object and its database rows, including chunks and private ingestion staging. Storage and database changes are not atomic; jobs use revision fences and idempotent completion. Document objects use the selected server-only Supabase Storage, Neon Object Storage, or AWS S3 adapter; metadata and authorization remain Account-scoped in the selected PostgreSQL database.

API RouteMethodDescription
/api/documentsGETList documents for account
/api/documentsPOSTUpload document (multipart, 10MB, authenticated standard rate limit). Returns 202 + Location and tracks async processing with exact token deduction.
/api/documentsDELETEDelete document + chunks + storage file
/api/documents/[id]GETDocument details + chunk summaries

File Structure

lib/rag/
├── index.ts                 # bounded chunks, durable document jobs, searchDocuments
├── pdf-parser.ts            # fixed child process, abort/timeout, exit confirmation
└── pdf-parser-worker.mjs    # page-by-page unpdf extraction with resource limits

lib/ai/agents/rag/
└── index.ts                 # RAGAgent class (extends BaseAgent)
                             # prepareRAGContext() + buildMessages() pattern

config/ai.ts
├── embeddingModels          # Available embedding models with dimensions + cost
└── embeddingConfig          # Default embedding model (credits are 1:1 with tokens, no flat rate)

config/ai-server.ts          # Validated server-only transport, credentials, and base URLs
lib/ai/llm-client.ts         # Sole chat/embedding provider boundary

app/api/documents/
├── route.ts                 # GET (list), POST (upload), DELETE (remove)
└── [id]/route.ts            # GET (details + chunk summaries)

app/[locale]/(private)/private-dashboard/documents/
└── page.tsx                 # Knowledge Base page (server component)

components/documents/
└── document-manager.tsx     # Upload UI with drag-and-drop + status polling