Razvoj

Beyond the Prompt: Architecting Your Backend for AI Companionship

Iza upita: Projektiranje vašeg pozadinskog sustava za AI druženje

An AI companion is not a clever prompt wrapped in a chat window. The prompt may define tone, boundaries, and a first impression, but companionship emerges from what happens after the first message: remembered context, reliable responses, consent-aware personalization, and graceful behavior when systems fail.

That makes backend architecture central to the product. If a companion forgets an important preference, repeats itself, leaks context between users, or becomes unavailable during a vulnerable conversation, no amount of prompt polish will repair the experience. The backend must treat conversation as durable, evolving application state rather than a stream of disposable API calls.

Model the product before modeling the messages

Start by separating the concepts that tend to get collapsed into a single chat table. A user has identity and account settings. A companion has a configured persona and policy. A conversation is a bounded interaction space. A message is an immutable event within that space. Memory is a curated set of facts or summaries that may influence future responses.

Those distinctions make practical features easier to build: multiple companions per user, temporary chats, data deletion, memory review, and model changes without rewriting old transcripts. They also give engineers clear ownership boundaries when a feature request arrives.

A relational database is often the right source of truth for this layer. PostgreSQL or MySQL can enforce ownership relationships, transactions, and retention workflows that become awkward when everything starts in a vector store. Keep the original user and assistant messages intact where product policy permits; derived summaries, embeddings, and classifications should be explicitly marked as derived data.

CREATE TABLE conversation_messages (
  id UUID PRIMARY KEY,
  conversation_id UUID NOT NULL,
  author_role VARCHAR(20) NOT NULL,
  body TEXT NOT NULL,
  created_at TIMESTAMP NOT NULL,
  model_request_id UUID NULL
);

The exact schema will vary, but immutability is a useful default. Corrections, redactions, and moderation actions can be represented as separate records or carefully audited updates. Avoid silently rewriting history: it complicates debugging, trust, and compliance work later.

Build a memory pipeline, not a bigger prompt

Long conversations quickly exceed a model’s useful context window, even when the underlying model accepts very large inputs. Sending every prior message is expensive, slow, and often counterproductive. Old small talk can drown out the one detail that matters now.

A better approach is layered retrieval. Assemble the request context from a few deliberate sources:

  • The system and companion instructions, including non-negotiable safety rules.
  • The recent conversational turn window for local coherence.
  • A compact conversation summary for continuity.
  • Retrieved long-term memories relevant to the current message.
  • User-controlled preferences, such as preferred name or communication style.

Do not let an embedding score alone decide what the model “knows.” A stored memory should have provenance, scope, confidence, creation time, and an expiration or review policy. “The user prefers concise replies” has a different lifetime and sensitivity level from “The user is planning a trip next month.”

Memory extraction should run asynchronously after a turn completes. A worker can propose candidate memories, apply deterministic filters, and store them for review or retrieval. This keeps the visible response path fast and prevents a memory-writing failure from blocking the chat.

Most importantly, make memory inspectable and controllable. Users should be able to see, edit, and delete remembered items. From an engineering perspective, that means a delete request must propagate beyond the primary row: remove derived embeddings, invalidate caches, and ensure future retrieval excludes the item.

Design the request path for failure

An LLM call is a remote dependency with variable latency, capacity limits, malformed responses, and occasional outages. Treat it accordingly. A synchronous PHP endpoint should validate input, authorize the conversation, persist the user message, create a request record, and hand work to a queue when the experience permits it.

For streaming chat, a worker or application process can relay model output to the client while accumulating the final answer. Only mark the assistant message complete after the stream is successfully finalized. If the connection drops, preserve enough state to show an honest “response interrupted” status instead of inventing a completed message.

Use idempotency keys for message submission. Mobile clients retry; browsers resend after flaky connections; queue systems can deliver jobs more than once. Without an idempotency key, one user action can create duplicate messages and duplicate model charges.

$key = $request->header('Idempotency-Key');

$message = DB::transaction(function () use ($conversation, $body, $key) {
    return Message::firstOrCreate(
        ['conversation_id' => $conversation->id, 'idempotency_key' => $key],
        ['author_role' => 'user', 'body' => $body]
    );
});

The surrounding code still needs a database uniqueness constraint; application-level checks alone are vulnerable to races. Queue jobs should likewise be safe to retry, with clear states such as queued, running, completed, failed, and cancelled.

Retries need boundaries

Retry transient network failures and rate limits with bounded exponential backoff. Do not blindly retry validation failures, policy refusals, or requests that may already have succeeded but whose response was lost. Record provider request identifiers when available, but do not assume every provider exposes the same fields or semantics.

A circuit breaker or temporary degradation mode can protect the rest of the application when the model provider is unhealthy. The useful fallback is often simple: acknowledge the delay, retain the user’s message, and invite them to retry later. A fabricated answer is never a reliable fallback.

Privacy and safety are architectural concerns

Companion products invite disclosure. That raises the bar for access control, retention, and operational hygiene. Authorize every conversation lookup by tenant or user ownership; never accept a conversation identifier as proof of access. Encrypt data in transit, limit staff access, and keep secrets outside source control and application images.

Log operational metadata sparingly. Request timing, error class, and token usage may be useful; raw conversation bodies are often not necessary for routine logs. When content must be retained for debugging or quality review, define access controls and retention periods before the data accumulates.

Moderation should be a pipeline with explicit outcomes, not a vague instruction hidden in a prompt. Input and output checks, escalation paths, and user-facing responses need to be observable and testable. The goal is not merely to block content; it is to behave predictably when the conversation crosses a product boundary.

Keep the PHP service boring on purpose

PHP is well suited to this kind of backend when the design respects its strengths: stateless HTTP workers, explicit queues, a relational database, and cache-backed coordination. Put model-provider code behind a small interface so application services depend on capabilities rather than one vendor’s response shape.

Containerize the web process and worker separately, even if they share the same codebase. They scale differently and fail differently. A Docker deployment should inject configuration through environment variables or a secret manager, run database migrations as a controlled release step, and expose health checks that distinguish “the process is alive” from “the application can serve traffic.”

Measure the full turn: queue delay, retrieval time, provider latency, stream duration, completion rate, retries, and failures by reason. These signals reveal whether a slow reply comes from the database, the retrieval layer, a saturated worker pool, or the model call itself.

The durable advantage is trustworthy continuity

The most compelling AI companion experiences feel present, but presence is engineered. It is the result of careful state modeling, selective memory, resilient delivery, clear privacy boundaries, and operational discipline.

Prompts will evolve, models will change, and providers will occasionally fail. A backend built around durable records, explicit workflows, and user control can absorb those changes. That is how a chat feature becomes a companion product worth trusting: not by pretending the system is human, but by making its behavior dependable when the conversation matters.

Portret autora bloga

Mihajlo

Ja sam Mihajlo — programer vođen znatiželjom, disciplinom i stalnom željom da stvorim nešto smisleno. Dijelim uvide, tutorijale i besplatne usluge kako bih pomogao drugima da pojednostave svoj rad i rastu u svijetu softvera i umjetne inteligencije koji se neprestano razvija.