Architecting Software That Teaches AI What It Needs to Know
Most AI failures are not model failures. They are software design failures.
A capable model can summarize, classify, reason over instructions, and generate useful drafts. But it cannot reliably act on information that is scattered across unreadable documents, hidden behind inconsistent APIs, or mixed with data it should never see. The real work is architectural: deciding what the system should know, when it should know it, and what it is allowed to do with that knowledge.
That is the difference between adding a chat box and building an AI-enabled product. The first exposes a model. The second gives a model a carefully designed operating environment.
Start with the job, not the model
“Add AI” is not a requirement. A useful requirement describes a decision, workflow, or outcome that can improve with assistance. For example: help support staff identify the right policy, help an engineer investigate a deployment failure, or help finance teams extract fields from submitted documents.
Each of those jobs has a different risk profile. A support assistant may need to cite approved policy text. A deployment assistant may need current system status but no ability to change production. A document workflow may need structured extraction followed by human review.
Before selecting a model or building a prompt, define four things:
- User: who receives the result and what expertise do they have?
- Decision: what will they do differently after seeing it?
- Evidence: which sources are authoritative enough to support the result?
- Boundary: what must the system refuse, escalate, or leave to a human?
This framing keeps a team from treating fluent output as the product. The product is the improved workflow around that output.
Make knowledge available in layers
AI systems need context, but “give it all the data” is almost always the wrong design. It increases cost, creates privacy and access-control problems, and makes it harder for the model to distinguish relevant information from noise.
A stronger pattern is progressive disclosure. Give the model stable instructions first, then retrieve only the documents relevant to the current request, then call narrowly scoped tools only when needed.
Separate instructions, facts, and permissions
These three categories should not be blended into one large prompt.
- Instructions define behavior: tone, task boundaries, required checks, and escalation rules.
- Facts come from product documentation, records, tickets, knowledge bases, or other sources that may change.
- Permissions determine which facts and tools a particular user or workflow may access.
This separation makes systems easier to update and audit. A policy change should update a policy source, not require editing a hidden paragraph in application code. A user’s permissions should filter retrieval before context reaches the model, not merely ask the model to ignore confidential records.
Retrieval also needs product thinking. Chunking a document into passages is not enough. Preserve titles, source links, effective dates, ownership, and access metadata. If a response matters, users should be able to inspect the evidence that informed it.
Give agents tools with narrow contracts
An AI agent becomes more useful when it can query systems, create records, or trigger workflows. It also becomes more dangerous. The answer is not to avoid tools; it is to design them as carefully as public APIs.
A good tool has a small purpose, explicit inputs, predictable outputs, and a clear authorization model. “Search customer orders” is safer and easier to evaluate than a general database query interface. “Create a draft refund request” is safer than “issue a refund.”
For consequential actions, split planning from execution. Let the model assemble a proposed action, show the important parameters, and require confirmation or an approval step before committing the change.
{
"action": "create_refund_request",
"order_id": "ORD-1042",
"amount": 49.00,
"reason": "duplicate_charge",
"requires_approval": true
}
The application, not the model, should enforce that approval requirement. Models can suggest and explain; deterministic software should validate schemas, apply authorization, handle retries, and record the result.
Design for uncertainty and failure
AI output is probabilistic. External systems are not always available. Retrieved documents can be incomplete or stale. A production design assumes all three conditions will happen.
Build explicit paths for uncertainty. If the system lacks sufficient evidence, it should say so and identify what it needs next. If a tool call fails, it should report the failure without pretending that an action succeeded. If an action times out, the system should check its final state before retrying, especially where a retry could create duplicate work.
Idempotency matters whenever agents initiate operations. A request identifier can allow a downstream service to recognize that a retry represents the same intended action rather than a new one. This is ordinary distributed-systems discipline, and it matters just as much when an AI is involved.
Also distinguish between a helpful draft and a verified result. A model may draft a change summary; a test suite verifies behavior. It may propose an incident diagnosis; telemetry and logs establish whether that diagnosis is supported. The system should make that distinction visible to users.
Evaluate the workflow, not just the answer
Teams often evaluate AI with a handful of impressive examples. That is useful for exploration, but insufficient for deployment. A reliable evaluation set includes ordinary requests, ambiguous requests, incomplete data, adversarial instructions, permission boundaries, and failure scenarios.
Measure what actually matters for the job. For a support assistant, that might include whether it selected approved guidance, cited the relevant source, avoided unsupported claims, and handed off sensitive cases. For a coding assistant, it might include whether suggested changes pass tests, respect repository conventions, and avoid altering unrelated behavior.
Keep representative examples as regression tests. When prompts, retrieval logic, tools, or models change, rerun them. This turns AI quality from a subjective demo-room debate into an engineering practice with observable tradeoffs.
Human oversight should be intentional
Human review is not a single switch. It belongs where the cost of being wrong exceeds the cost of pausing: irreversible actions, financial commitments, regulated decisions, sensitive communications, and low-confidence cases.
Conversely, forcing approval for harmless, reversible work can erase the value of automation. The goal is calibrated autonomy: let the system prepare, sort, summarize, and draft broadly; require stronger controls as the potential impact rises.
Good interfaces help reviewers act quickly. Show the proposed action, the evidence, the affected records, and the reason it was chosen. Do not make a reviewer reverse-engineer a chain of model messages to understand what will happen.
Teach the system through architecture
Software teaches an AI system what matters. Data models tell it what entities exist. Retrieval tells it which evidence is relevant. Tool contracts tell it what actions are possible. Permissions tell it where the boundaries are. Evaluation tells the team what “good” means.
The model is important, but it is only one component in that lesson plan. The most durable AI systems will be built by teams that treat context, controls, and feedback as first-class product architecture. When the system knows the right things, can verify what it claims, and acts only within well-designed boundaries, intelligence becomes useful rather than merely impressive.