Razvoj

Beyond the Prompt: Architecting Systems for AI's Unpredictable Future

Iza prompta: projektiranje sustava za nepredvidivu budućnost umjetne inteligencije

AI features rarely fail because the prompt was not clever enough. They fail because the surrounding system assumes a clean answer, a fast response, a stable format, and a cost profile that reality does not guarantee.

That is the uncomfortable shift for backend teams: an AI model is not a deterministic library call. It is a capable but variable dependency. Treating it as one changes how we design APIs, databases, queues, observability, and even user expectations.

The prompt still matters. It defines intent and constrains output. But production quality comes from the architecture around it: validation, retries, fallbacks, versioning, and a deliberate answer to the question, “What happens when this response is late, malformed, expensive, or simply wrong?”

Design for uncertainty at the boundary

A conventional backend endpoint often maps structured input to structured output. Given the same request, the system is expected to follow predictable rules. An AI-backed endpoint has more moving parts: model behavior can vary, upstream capacity can fluctuate, and generated text may not respect a requested schema.

Do not let that uncertainty leak unchecked into the rest of the application. Put an explicit boundary around the model integration.

For example, instead of allowing controllers to call an AI provider directly, create a service that receives a domain-level request and returns a validated domain-level result. The controller should not know about prompt templates, model identifiers, token limits, or provider-specific response shapes.

final class ProductSummaryService
{
    public function __construct(
        private AiClient $client,
        private SummaryValidator $validator,
    ) {
    }

    public function summarize(Product $product): ProductSummary
    {
        $response = $this->client->generateSummary(
            title: $product->title,
            description: $product->description,
        );

        return $this->validator->toProductSummary($response);
    }
}

The point is not the interface itself. The point is containment. A later provider change, a revised prompt, or a second validation pass belongs inside this boundary rather than across every controller, job, and command.

Require structure, then verify it

If downstream code needs fields such as a title, summary, category, and confidence flag, free-form prose is an unreliable contract. Ask for structured output where the chosen integration supports it, but never mistake a requested schema for a guaranteed schema.

Validate every generated result before persisting it or exposing it through an API. Check required fields, types, allowed values, string lengths, and business rules. A valid JSON document can still be invalid application data.

Consider a support-triage workflow. A model may classify a ticket as “billing,” but the application may only allow a fixed set of categories. It may suggest a priority that conflicts with an account’s contractual service level. Those decisions belong to deterministic application rules.

  • Parse the response at the integration boundary.
  • Validate it against an explicit application schema.
  • Reject or repair invalid results through a bounded retry path.
  • Store the final normalized representation, not an opaque provider payload as the primary record.
  • Retain carefully selected raw metadata for debugging, subject to privacy requirements.

A bounded retry is important. Retrying indefinitely turns malformed output into a queue backlog and a cost problem. One retry with a corrective instruction may be reasonable; beyond that, route the item to a review state or use a safe fallback.

Make asynchronous the default for non-interactive work

Users notice latency long before they admire model sophistication. If an AI operation does not need to finish before the current HTTP response, move it to a queue.

Document enrichment, classification, translation, tagging, report preparation, and background analysis are natural candidates. Persist a job request, return a clear pending state, and let a worker perform the AI call with controlled retry behavior.

In a PHP application, that may mean placing a small, serializable job on the existing queue rather than embedding a long-running request in a web worker. The job should carry stable identifiers, not a large object graph or mutable request context.

final class GenerateArticleTags
{
    public function __construct(
        public readonly int $articleId,
    ) {
    }

    public function handle(ArticleRepository $articles, TaggingService $tagging): void
    {
        $article = $articles->findOrFail($this->articleId);

        if ($article->tagsGeneratedAt !== null) {
            return;
        }

        $tags = $tagging->generate($article);
        $articles->saveGeneratedTags($article->id, $tags);
    }
}

This job is intentionally idempotent. Queue systems can retry work after timeouts or worker failures. If a retry creates duplicate records, sends duplicate notifications, or bills a customer twice, the architecture has converted a normal operational event into a product defect.

Model state belongs in the data model

AI output is often treated as disposable text. That is fine for a temporary chat response, but it is insufficient when output influences a customer-facing page, a workflow decision, or a search index.

Store enough context to understand what produced a result. At minimum, that usually includes an input fingerprint, a prompt or policy version identifier, a model configuration identifier, generation status, timestamps, and the validated result. Avoid assuming that “model name” alone captures meaningful behavior; application-side instructions and validation rules are part of the effective system.

Versioning makes change safe. When a better prompt or different provider becomes available, you can decide whether existing content should remain as-is, be regenerated gradually, or be compared before replacement. Without versioning, every generated record becomes an archaeological puzzle.

Separate source data from generated derivatives

Keep human-authored source content distinct from generated summaries, labels, embeddings, or recommendations. The source is authoritative. Generated data is a derivative that can be regenerated, invalidated, or reviewed.

This distinction also helps database design. A separate table for generated artifacts can hold status, version fields, error details, and audit timestamps without overloading the core domain table. It supports a practical question that every team eventually faces: which records are stale after a change?

Build reliable failure paths before polished happy paths

An AI provider may be unavailable. A response may exceed a deadline. Input may contain sensitive data that should not leave your environment. A request may become too large. These are not edge cases to postpone; they are normal conditions for an external dependency.

Define what the product does in each case. A product description page may show editor-written content without its generated summary. A search feature may fall back to traditional keyword search. A workflow may pause for review rather than automatically taking an irreversible action.

Use timeouts that reflect the user experience, not vague optimism. Apply retries only to transient failures, with limited attempts and increasing delay. Track permanent failures separately from temporary ones. A failed job that silently disappears is worse than a visible pending state because it creates false confidence.

Also treat cost as a failure mode. Set input limits, constrain output length, cache results where the input is unchanged, and avoid repeated generation caused by page refreshes or duplicate jobs. An architecture that performs well in a test environment can become fragile when ordinary usage multiplies calls.

Observe behavior, not just outages

Traditional monitoring asks whether an endpoint is up and how long it takes. AI systems need additional questions: How often does validation fail? How many retries are needed? Which prompt version produces the most review work? Are fallback paths becoming common? Are generated results being replaced by users?

Log request identifiers and versions so a support report can be traced through the system. Be deliberate about content logging: prompts and outputs can contain customer data, credentials, or proprietary material. Observability should improve diagnosis without becoming an ungoverned copy of sensitive data.

Metrics should lead to action. If a particular task has high validation failure, simplify the requested structure or split the task. If latency is inconsistent, move the operation off the request path. If users regularly edit an output, consider whether the model is being asked to make a judgment that belongs in product rules or a human review step.

Keep the system replaceable

The most durable AI architecture does not bet the application on one prompt, one provider, or one fashionable interaction pattern. It gives the model a narrow role, makes its inputs and outputs explicit, and preserves deterministic control where correctness matters.

That is not a retreat from AI. It is how useful AI survives contact with production. The prompt may open the door, but architecture decides whether the feature remains trustworthy when the future behaves unpredictably.

Portret autora bloga

Mihajlo

Ja sam Mihajlo — programer vođen znatiželjom, disciplinom i stalnom željom da stvorim nešto smisleno. Dijelim uvide, tutorijale i besplatne usluge kako bih pomogao drugima da pojednostave svoj rad i rastu u svijetu softvera i umjetne inteligencije koji se neprestano razvija.