Umjetna inteligencija (UI)

Beyond the Prompt: Architecting Software AI Needs to Evolve With

Iza upita: Arhitektura softvera s kojom se AI treba razvijati

Most AI failures in software do not begin with a bad prompt. They begin with an architecture that assumes intelligence is a feature call: send text in, receive text out, ship it. That works for a demo. It rarely works for a system expected to survive changing models, ambiguous requests, unreliable external tools, evolving policies, and the ordinary messiness of production data.

The useful question is not “Which model should we use?” It is “What kind of software can safely improve as its AI capabilities change?” Treating AI as a durable system component, rather than a clever endpoint, leads to better products and calmer engineering teams.

Put the model behind a boundary

A model provider, model version, prompt format, and response shape are all likely to change. Do not let those choices leak through the application. Create a narrow interface that expresses the business capability you need.

type Classification = {
  category: "billing" | "technical" | "account";
  confidence: number;
  rationale?: string;
};

interface TicketClassifier {
  classify(input: {
    subject: string;
    body: string;
  }): Promise<Classification>;
}

The rest of the application should depend on TicketClassifier, not on a particular chat-completions request or provider-specific response object. This makes it possible to replace a prompt with a fine-tuned model, introduce a rules-based fallback, or route different requests to different implementations without rewriting every caller.

This boundary should also normalize failures. A timeout, a malformed response, a rate limit, and a policy refusal are not interchangeable, but they should arrive at the application in a predictable form. When every feature invents its own retry behavior, observability and reliability become accidental.

Design for uncertainty, not theatrical certainty

Language models can produce useful answers without providing a dependable guarantee that every answer is correct. Good architecture accepts that uncertainty explicitly. It does not force a probabilistic output into a deterministic workflow without checks.

For an AI-assisted support triage system, confidence can drive the next action:

  • High-confidence, low-risk classifications can be routed automatically.
  • Ambiguous classifications can be shown to an agent with the supporting context.
  • High-impact actions, such as account changes or payments, should require deterministic validation and appropriate human approval.

Confidence values deserve care. A number emitted by a model is not automatically calibrated. Use it as one signal among others: request completeness, retrieval quality, policy checks, tool results, and the cost of being wrong. The practical goal is not to eliminate uncertainty. It is to make uncertainty visible and manageable.

Separate reasoning from action

An agent becomes much more consequential when it can act on external systems. Reading a document and updating a customer record may appear adjacent in a workflow, but they carry very different operational risks.

Keep tool access explicit. Give the system small, well-defined operations with constrained inputs, authorization checks, and auditable results. Prefer create_draft_invoice over a broad database-writing tool. Prefer a scoped search operation over unrestricted access to an entire knowledge store.

Validate every tool call outside the model. The model can propose structured arguments, but conventional code should verify required fields, types, allowed values, permissions, and business rules before side effects occur. This is not a limitation of AI; it is the same principle that keeps any untrusted input from becoming a production command.

Make irreversible work deliberate

For actions that cannot be easily undone, introduce a confirmation boundary. An agent can prepare a change, explain the intended effect, and request approval. The human reviewer should see the actual action, relevant context, and expected outcome rather than a vague promise that the system has “handled it.”

Idempotency matters too. Network retries and partial failures are normal. If an action is retried after a timeout, the system needs a stable request identifier so it can determine whether the action already succeeded. Otherwise, a helpful agent can create duplicate tickets, send repeated messages, or apply the same change twice.

Build retrieval as a product capability

Many useful AI features depend less on model cleverness than on access to current, relevant information. Retrieval should be treated as a designed subsystem: content ingestion, permissions, chunking, indexing, ranking, citations where appropriate, and lifecycle management.

A support assistant that retrieves outdated policy text can be fluent and still be harmful. A code assistant that sees only a fragment of a repository can make locally plausible changes that violate system boundaries. The solution is not simply to add more context. More context can introduce irrelevant instructions, stale material, privacy exposure, and higher cost.

Define what sources are authoritative, preserve access controls from the original systems, and record which material informed an answer when the use case requires traceability. Evaluation should test retrieval quality separately from answer quality. If the right document never reaches the model, prompt refinement will not solve the real problem.

Evaluate workflows, not just prompts

A prompt can look excellent against a handful of examples and fail when users phrase requests differently, tools return incomplete data, or the model chooses an unexpected path. Evaluation needs representative tasks and explicit success criteria.

Start with a small test set drawn from real workflow patterns after removing or protecting sensitive information. Include straightforward examples, ambiguous cases, missing data, adversarial instructions, and expected failures. Then assess the full chain: input handling, retrieval, model output, validation, tool execution, and user-facing result.

Useful measures vary by product, but they often include task completion, escalation quality, invalid action attempts, latency, cost, and operator correction rate. A change that improves answer style while increasing unsafe tool proposals is not an improvement.

Keep versions of prompts, tool definitions, model settings, and evaluation cases together. Reproducibility is essential when behavior shifts. Without it, teams end up debating impressions instead of identifying which change altered an outcome.

Operate AI features like production systems

AI workloads need ordinary engineering discipline: timeouts, bounded retries, queues where appropriate, rate limiting, access control, logging, alerts, and rollback paths. They also need AI-specific visibility. Capture enough structured information to understand a request’s route, retrieved context identifiers, tool attempts, validation outcomes, and final disposition. Avoid casually logging sensitive prompts or retrieved documents; observability must respect the same privacy expectations as the application itself.

Graceful degradation is often the difference between a useful feature and a fragile dependency. If classification is unavailable, send work to a general queue. If retrieval fails, say that the system cannot verify the answer rather than inventing one. If a model response cannot pass schema validation, retry only when a retry is meaningful, then escalate or fail safely.

Let the architecture evolve with the capability

The strongest AI products are not built around the assumption that models will become perfect. They are built so that better models improve the experience, while imperfect models remain contained by good interfaces, validation, retrieval, evaluation, and human judgment.

That is the work beyond the prompt. A prompt may unlock an impressive first interaction. Architecture determines whether that interaction can become dependable software: software that learns from feedback, adapts to changing capabilities, and earns the right to take on more responsibility over time.

Portret autora bloga

Mihajlo

Ja sam Mihajlo — programer vođen znatiželjom, disciplinom i stalnom željom da stvorim nešto smisleno. Dijelim uvide, tutorijale i besplatne usluge kako bih pomogao drugima da pojednostave svoj rad i rastu u svijetu softvera i umjetne inteligencije koji se neprestano razvija.