Development

Beyond the Prompt: Building Your Backend's AI Knowledge Base

Beyond the Prompt: Building Your Backend's AI Knowledge Base

Most AI features begin as a prompt and end as a production incident. The missing piece is rarely a more clever sentence. It is a backend knowledge base that can supply trustworthy, scoped, current context when a request arrives.

For a PHP application, that knowledge base is not a chatbot transcript pasted into a table. It is a system: content ingestion, normalization, storage, retrieval, authorization, observability, and a clear contract between your application and the model. Treating it as backend infrastructure changes the questions you ask. Instead of “What should the prompt say?” ask “Which records may this user see, how fresh are they, and can we explain why they were retrieved?”

Start with the job, not the model

An AI assistant should have a narrow, useful job. “Answer questions about our product” sounds reasonable but hides important architectural choices. A better definition identifies the audience, the permitted knowledge, the expected action, and the boundary of responsibility.

For example, an internal support assistant may answer from approved help articles and product-release notes. It may summarize an account’s current configuration only when the requesting employee is authorized to view that account. It should not infer billing status from unrelated text or present an uncertain answer as a confirmed fact.

This definition becomes a retrieval policy. It tells the backend which sources to index, which metadata matters, and when the application must decline to answer. A model can produce fluent text from weak context. Your backend must prevent weak context from becoming a confident-looking response.

Model knowledge as documents plus metadata

Keep the original source of truth where it belongs: a CMS, database, ticketing system, repository, or object store. The AI knowledge base should be a derived read model optimized for retrieval. That separation makes re-indexing, auditing, and deletion manageable.

Each indexed chunk needs more than content. Metadata is how a backend turns retrieval into a controlled operation.

  • Source identity: a stable document ID and source URL or internal reference.
  • Versioning: a content hash, source update time, and indexing time.
  • Scope: tenant, product, locale, audience, and visibility level.
  • Lifecycle: published, superseded, archived, or deleted.
  • Provenance: the section or page from which a chunk was created.

Chunking deserves deliberate design. Splitting purely by character count can separate a warning from the procedure it qualifies. Splitting only by headings can create chunks that are too large and unfocused. Start with semantic sections, preserve heading context, and use modest overlap when a section must be split. Store the heading with each chunk so retrieved text remains intelligible outside its original page.

Build ingestion as a repeatable pipeline

Indexing should be an idempotent background workflow, not an HTTP request that happens when an editor clicks Publish. A reliable pipeline fetches source content, converts it to normalized text, validates it, chunks it, enriches it with metadata, creates searchable representations, and writes the resulting records.

In PHP, the request handler can enqueue work and return promptly. A worker can then process the document with retry rules appropriate to the failure. Transient network errors may be retried with backoff. Invalid source markup should be recorded for correction, not retried forever. A deleted source should remove or deactivate its indexed chunks.

$document = $sourceRepository->find($documentId);

if ($document === null || $document->isArchived()) {
    $knowledgeRepository->removeBySourceId($documentId);
    return;
}

$text = $normalizer->toPlainText($document->body());
$chunks = $chunker->split($text, $document->title());

$knowledgeRepository->replaceSource(
    sourceId: $document->id(),
    contentHash: hash('sha256', $text),
    chunks: $chunks,
    metadata: [
        'tenant_id' => $document->tenantId(),
        'visibility' => $document->visibility(),
        'updated_at' => $document->updatedAt()->format(DATE_ATOM),
    ],
);

The important detail is replaceSource, not the exact method name. Reprocessing a document should converge on one current set of chunks. Avoid append-only indexing unless you also have an explicit, tested strategy for superseding old content.

Retrieval is an authorization problem first

Semantic search, keyword search, and hybrid ranking can all be useful. None is safe if filtering happens after retrieval. Apply tenant and visibility constraints in the retrieval query itself. A result that must be discarded after it was fetched has already crossed a boundary it should not have crossed.

A practical retrieval flow looks like this:

  1. Authenticate the caller and construct a permission-aware retrieval scope.
  2. Search only active chunks within that scope.
  3. Rank a small candidate set using lexical, semantic, or hybrid signals.
  4. Apply a relevance threshold and deduplicate near-identical passages.
  5. Send the selected passages, their citations, and task instructions to the model.
  6. Return the answer with source references the client can display.

Do not force an answer when retrieval is poor. A useful fallback is direct: say that the available knowledge does not support a reliable answer, then offer a narrower query or a human escalation path. This is not a failure of the model. It is a correct response to insufficient evidence.

Keep prompts thin and contracts explicit

The prompt should express behavior, not contain your entire knowledge base. Give the model a concise role, the user’s request, retrieved context, and output rules. Require it to distinguish supported facts from uncertainty, and require citations that map to the supplied chunks.

Use structured output where your application needs predictable handling. For instance, an answer endpoint may expect answer, citations, and needs_escalation. Validate the model response before returning it. If JSON parsing fails or citations reference unknown chunks, log the event and use a safe fallback rather than passing malformed output to the client.

Also treat retrieved content as untrusted data. A document may contain text that looks like instructions to the model. Delimit context clearly and tell the model that the context is evidence, not authority over system behavior. This will not solve every prompt-injection risk, but it establishes the correct trust boundary in your application design.

Operate it like any other backend dependency

AI quality is difficult to improve when requests are opaque. Record the retrieval query, selected document IDs, ranking scores where available, model request identifiers, latency, errors, and the final response status. Protect sensitive content in logs; often identifiers and measurements are enough for debugging.

Measure the pipeline separately from the chat endpoint. Indexing lag, failed jobs, stale chunks, empty retrievals, authorization denials, model timeouts, and fallback rates each point to different problems. A slow answer may be caused by database filtering, retrieval, model inference, or a retrying upstream call. One aggregate “AI latency” metric is rarely actionable.

Cache carefully. Cache stable retrieval results only when the permission scope is part of the cache key. Cache generated answers only when the question, relevant source versions, locale, and authorization context match. An answer cache that ignores document versions quietly serves stale guidance after an update.

The durable advantage is trust

The most valuable AI feature is not the one that sounds the most human. It is the one that retrieves the right information for the right person, shows where it came from, admits when evidence is missing, and stays correct as the underlying system changes.

Prompts will evolve, models will change, and ranking methods will improve. A well-designed knowledge base gives those changes a stable foundation. Build that foundation as a maintainable backend system, and your AI layer becomes a useful product capability rather than a fragile demonstration.

Blog author portrait

Mihajlo

I’m Mihajlo — a developer driven by curiosity, discipline, and the constant urge to create something meaningful. I share insights, tutorials, and free services to help others simplify their work and grow in the ever-evolving world of software and AI.