Architect Software AI Needs to Understand Before It's Built
AI is often introduced into software architecture as a feature: add a chat box, call a model, ship an assistant. That framing is too small. Once an AI system can interpret requests, choose tools, retrieve company information, or trigger workflows, it becomes part of the architecture’s decision-making surface.
The important question is not “Which model should we use?” It is “What kind of system are we creating, and what must remain dependable when the model is uncertain, slow, unavailable, or wrong?” Good answers lead to useful AI products. Weak answers produce impressive demos with unclear boundaries.
Start with the decision, not the model
An AI capability should exist to improve a specific decision or reduce a specific kind of work. “Summarize support tickets” is a task. “Help support leads identify recurring issues before they become incidents” is an outcome. The second statement gives architects something to design for: inputs, users, review points, latency expectations, and measures of usefulness.
Separate work that requires language judgment from work that requires exact rules. A model can classify a vague customer message, draft a response, or extract likely themes from documents. It should not be the sole authority for applying a discount policy, changing account permissions, or calculating a financial total.
That distinction is foundational. Use deterministic software for deterministic obligations. Use AI where ambiguity, variation, or unstructured information makes traditional programming expensive or brittle.
Design the AI as an unreliable collaborator
Language models are productive because they can produce plausible output from incomplete context. That is also their central engineering risk. They may misunderstand an instruction, invent a detail, select the wrong tool, or express uncertainty with undeserved confidence.
An architecture should therefore treat model output as untrusted until the surrounding system has validated it. This does not mean AI is unusable. It means the system needs layers that convert useful suggestions into safe actions.
- Constrain inputs. Define the task, available context, and allowed actions clearly.
- Validate outputs. Check formats, required fields, ranges, identifiers, and permissions in ordinary application code.
- Limit authority. Give an agent the minimum tools and data needed for its job.
- Require confirmation. Put a human approval step before consequential external actions.
- Record decisions. Preserve relevant inputs, tool calls, outputs, and outcomes for debugging and review.
For example, an AI assistant may turn a natural-language request into a proposed database query. The application should still enforce read-only access, allow only approved tables, cap returned rows, and reject unsafe query structures. The model can help translate intent; the application remains responsible for enforcement.
Keep business rules outside the prompt
Prompts are useful interfaces, but they are poor substitutes for product rules. If a policy matters enough to affect money, access, compliance, or customer treatment, encode it where software can test and enforce it.
A prompt can tell an assistant to avoid revealing confidential information. A permission layer should determine what information the current user may retrieve. A prompt can ask an agent to issue refunds only within policy. A service should verify the refund amount, reason, account state, and approval requirement before executing the transaction.
This separation improves more than safety. It makes policy changes easier to review, test, and deploy without trying to infer whether a revised sentence will change model behavior in every edge case.
Use structured boundaries
When an AI system passes information to another system, prefer a schema over free-form prose. Ask the model for a bounded structure, validate it, and handle validation failures explicitly. A model response that cannot satisfy the schema is not a reason to guess; it is a reason to retry with a clearer instruction, request human input, or stop the workflow.
{
"action": "create_draft",
"customer_id": "string",
"summary": "string",
"requires_review": true
}
The schema does not make the output true. It makes the contract inspectable. Your service still needs to confirm that the customer exists and that the caller is authorized to create the draft.
Retrieval is a product and data problem
Many AI applications need current, organization-specific knowledge. Retrieval can provide relevant documents at request time, but it is not simply “connect a vector database.” The quality of the answer depends on document quality, chunking, metadata, access controls, search behavior, and the way retrieved material is presented to the model.
Architects should ask practical questions early. Which sources are authoritative? How quickly must updates become available? What happens when a document is missing or contradictory? Can a user retrieve material they would not otherwise be allowed to see?
Metadata is especially important. A policy document without an owner, effective date, product area, or access classification is difficult to retrieve reliably and dangerous to use blindly. Build ingestion workflows that preserve provenance. An answer should be able to point back to the material used to support it, even if the product chooses a different presentation than a formal citation.
Plan for failure as a normal state
AI systems add failure modes beyond ordinary application failures: rate limits, timeouts, malformed output, unavailable tools, stale retrieved content, and model behavior that changes as providers update systems. A resilient design assumes these will happen.
Set time budgets and fallback behavior. If an assistant cannot generate a useful summary within the interaction window, show the original content or offer a queued result. If a tool call fails, do not silently claim the action succeeded. If retrieval returns weak evidence, let the system say it does not have enough information.
Retries need care. Retrying a text-generation request may be reasonable. Retrying an action that created a ticket or sent a message can duplicate work unless the operation is idempotent. Attach a stable request identifier to side-effecting operations and let the receiving service recognize duplicates.
Evaluate the whole workflow
Traditional software tests often ask whether a function returns an expected value. AI evaluation also needs to ask whether the workflow is useful, grounded, safe, and appropriately cautious across realistic variation.
Create a small, representative evaluation set before broad deployment. Include straightforward requests, ambiguous requests, missing information, conflicting documents, adversarial instructions, and requests that should be refused or escalated. Review not only final answers but also tool choices, retrieved context, and whether the system followed required controls.
Production feedback matters too. Monitor failure rates, validation rejections, abandoned interactions, escalations, latency, and correction patterns. A high volume of fluent answers is not evidence of value if users regularly repair them afterward.
Give people meaningful control
Human oversight works only when it is designed into the workflow. Asking someone to approve hundreds of opaque AI decisions is not a safeguard; it is an alert fatigue machine. Reviewers need concise context: what the system proposes, the evidence it used, the confidence or uncertainty signals available, and the exact consequence of approval.
Choose review points according to reversibility. Drafting a meeting summary may need lightweight correction. Sending a customer-facing contract change deserves stronger review. Deleting data, changing permissions, or initiating payments should have explicit authorization and clear audit records.
Architecture is where trust becomes real
The most durable AI systems will not be those that give models the most freedom. They will be the ones that pair capable models with clear boundaries, trustworthy data, observable workflows, and deliberate human control.
Think of the model as a powerful component, not the application itself. Let it interpret, generate, rank, and assist. Let the surrounding architecture define truth, permission, policy, and accountability. That is how AI becomes more than a compelling interface: it becomes a system people can depend on.