Razvoj

Ditch the Prompt: Architecting Systems AI Learns From

Odbacite upit: Dizajniranje sustava iz kojih AI uči

The most valuable AI feature in a product is rarely the clever prompt. Prompts are interfaces: useful, visible, and easy to revise. The harder work sits underneath them—the system that gives a model the right context, captures outcomes, protects data, and steadily improves from real use.

That distinction matters because prompt-centric prototypes often fail in production for ordinary engineering reasons. Context is stale. Inputs are inconsistent. Retrieval returns irrelevant records. A retry duplicates a side effect. Nobody can explain why an answer was produced. The model may be capable, but the surrounding system has not earned the right to trust it.

Backend engineers are well placed to solve this. The durable advantage is not a paragraph of instructions hidden in application code. It is a well-designed learning loop built from contracts, data quality, evaluation, and operational discipline.

Start with a decision, not a chatbot

“Add AI” is too vague to design. Identify a specific decision or transformation the system should improve: classify an incoming support request, propose metadata for a document, summarize a case for review, or retrieve the most relevant policy for an operator.

Then define the boundary. What inputs are authoritative? What output shape can downstream code safely consume? Which actions require a person to approve them? A model should not be allowed to turn uncertain language into irreversible database writes merely because its response sounds confident.

A useful rule is simple: let AI propose; let deterministic systems validate, authorize, and execute.

For example, a model may suggest a ticket category and priority, but the application should verify that both values belong to known enumerations before persisting them. This makes failure explicit and keeps the model from becoming an undocumented source of truth.

final class TicketSuggestion
{
    public function __construct(
        public readonly string $category,
        public readonly string $priority,
        public readonly string $reason,
    ) {}
}

function isValidSuggestion(TicketSuggestion $suggestion): bool
{
    return in_array($suggestion->category, ['billing', 'technical', 'account'], true)
        && in_array($suggestion->priority, ['low', 'normal', 'high'], true);
}

The exact class is less important than the boundary it represents. Your application owns the contract. The model is one input to that contract.

Build a context pipeline

Models do not learn from your production database by osmosis. They learn only from what you deliberately provide during a request or what you later use in a controlled training or evaluation process. Treat context assembly as a first-class backend pipeline.

For each use case, decide what information belongs in the request:

  • Current facts: the records needed to answer this request.
  • Business rules: concise, versioned instructions that explain constraints.
  • Relevant history: selected prior events, not an unbounded conversation dump.
  • Reference material: retrieved documents with clear provenance.
  • Output schema: a structure your application can parse and validate.

This is where many systems become expensive and unreliable. Passing every available field feels safer, but excess context can obscure the signal, increase latency, and leak information across boundaries. Fetch data using the same discipline you would apply to an API response: minimum necessary fields, clear ownership rules, and explicit authorization checks.

Retrieval deserves similar care. Store document chunks with stable identifiers, source metadata, and a version or timestamp. When a response uses retrieved material, record which chunks were included. Without that record, debugging an incorrect answer turns into guesswork: was the problem the source document, indexing, ranking, prompt instructions, or model behavior?

Make feedback an event stream

A system becomes capable of improvement when outcomes are captured as structured events. “Thumbs up” and “thumbs down” can help, but they are weak on their own. Record the task, the context version, the model output, the final human or system action, and the eventual outcome where available.

For a classification workflow, that might mean storing the suggested category, the accepted or corrected category, whether the ticket was later reassigned, and the versions of the prompt and retrieval strategy used. Do not make the production request wait for all of this bookkeeping. Emit an event and process it asynchronously.

$event = [
    'type' => 'ticket.classification.reviewed',
    'ticket_id' => $ticketId,
    'suggestion' => $suggestion->category,
    'final_category' => $finalCategory,
    'prompt_version' => $promptVersion,
    'context_version' => $contextVersion,
    'reviewed_at' => gmdate('c'),
];

$eventBus->publish($event);

Use an outbox pattern when the event must remain consistent with a database transaction. Write the business change and an unsent event record in the same transaction; a worker can publish it later with retries. Consumers must be idempotent, because delivery can happen more than once. This is familiar distributed-systems engineering, and it matters just as much for AI feedback as it does for payments or notifications.

Version everything that changes behavior

Prompt text should be versioned, but stopping there is a mistake. A response is produced by a configuration set: model choice, system instructions, tool definitions, retrieval settings, document index version, output schema, and application code.

Give that set a traceable identity. When quality changes, you need to compare runs that differ in one intentional dimension. Otherwise, a prompt edit may appear to improve results while a simultaneous reindex silently caused the change.

Versioning also enables safe rollout. Route a small, suitable slice of traffic to a candidate configuration, observe structured outcomes, and retain a rollback path. For higher-risk workflows, run the candidate in shadow mode first: generate its output, record it, but do not expose it or let it act.

Evaluate before optimizing

Production logs are valuable, but they are not a complete evaluation set. Curate representative cases, including ambiguous inputs, missing fields, outdated reference material, hostile text, and cases where the correct answer is “I do not know.” Keep expected outputs or reviewer criteria alongside them.

Measure what the product actually needs. A summarizer may need factual coverage and clear uncertainty. A classifier may need correct routing and calibrated escalation. A retrieval workflow may need the right source document in the context before answer quality is even meaningful.

Also measure operational behavior: validation failures, fallback frequency, latency, queue depth, retry count, and cost per completed task. A system that produces excellent results only when a dependency is healthy and response time is unlimited is not ready for routine work.

Design graceful failure paths

External model calls can time out, reject requests, or return unusable output. Plan for each case. Set bounded timeouts. Retry only transient failures, with backoff and a retry limit. Attach idempotency keys to requests that might trigger actions. Validate every structured response before it crosses into core business logic.

Most importantly, define the fallback before launch. It may be a rules-based route, a queued human review, a cached response, or an explicit unavailable state. Silent degradation is dangerous when users assume an AI-generated answer was checked against current information.

The prompt is still important—just not alone

Good prompts clarify roles, constraints, and output expectations. They are worth testing and refining. But they should be the smallest visible layer of a system designed for evidence, correction, and accountability.

The enduring engineering question is not “What should we ask the model?” It is “What system will tell us whether this decision was useful, safe, and correct enough—and help us improve it tomorrow?” Build that system, and prompts become replaceable components rather than the foundation of your product.

Portret autora bloga

Mihajlo

Ja sam Mihajlo — programer vođen znatiželjom, disciplinom i stalnom željom da stvorim nešto smisleno. Dijelim uvide, tutorijale i besplatne usluge kako bih pomogao drugima da pojednostave svoj rad i rastu u svijetu softvera i umjetne inteligencije koji se neprestano razvija.