Beyond Prompts: Architecting Software AI Can't Live Without
Most AI conversations start with prompts. The more consequential work starts after them.
A useful prompt can produce code, summarize a document, classify a ticket, or draft a response. But a prompt alone is not a dependable software capability. It has no durable state, no permission boundary, no retry policy, no audit trail, and no clear definition of success. The moment an AI output affects a customer, a production system, or a business decision, the surrounding architecture matters more than the wording of the request.
The durable opportunity is not simply to add a model call to an application. It is to build software that makes AI dependable enough to be useful.
Think of the model as a probabilistic component
Traditional software components are usually expected to behave consistently for the same input. A model is different. It can misunderstand ambiguity, produce incomplete structured output, choose an inappropriate tool, or confidently state something unsupported by available data.
That does not make models unsuitable for production. It means they should be placed where probabilistic judgment creates value and surrounded by deterministic systems that constrain risk.
A practical AI feature typically has at least four layers:
- Context assembly: gather the relevant data, instructions, policies, and user intent.
- Model inference: ask the model to reason, draft, classify, extract, or select an action.
- Validation: check format, permissions, business rules, and confidence conditions.
- Execution and observation: perform approved actions, record outcomes, and make failures visible.
The model belongs in the second layer. It should not silently replace the other three.
Build workflows, not chatbot-shaped features
A chat interface can be useful, but it often hides the real job to be done. Start instead with a workflow: what enters the system, what decisions must be made, what tools are involved, and where a human must remain accountable.
Consider an internal support assistant that helps resolve access requests. A weak implementation gives the model broad administrative access and asks it to “handle the request.” A stronger design asks the model to extract the requested system, identify the requester, summarize the stated justification, and route the request through existing approval rules.
The model may propose a structured result such as this:
{
"request_type": "access_request",
"system": "analytics",
"requested_role": "viewer",
"justification": "Needs dashboard access for quarterly reporting",
"needs_human_review": true
}
Your application should validate that structure before acting on it. It should confirm that the system exists, the requested role is valid, the requester is eligible, and the approval path is known. Only then should it create a ticket or notify an approver.
This distinction is crucial: the model interprets language; the application enforces policy.
Context is a product decision
Many disappointing AI features are not model failures. They are context failures. The model was asked to answer a question without the current policy, the relevant customer record, the correct product documentation, or an explanation of what it must not assume.
Useful context is not the same as maximum context. Sending every available document creates cost, latency, privacy exposure, and distraction. It can also make an answer worse when outdated or conflicting material overwhelms the important facts.
Design context deliberately. For each workflow, define:
- Which data sources are authoritative.
- How current the information must be.
- Which fields are relevant to the decision.
- What information the model must never receive.
- How the system handles missing or contradictory context.
For example, an AI assistant drafting a customer reply may need the open case, the customer’s plan, approved help content, and previous messages in that case. It probably does not need unrestricted access to every customer conversation or every internal document.
Good retrieval is therefore less about clever search alone and more about provenance. When a response matters, the system should know what information supported it and whether that information was appropriate to use.
Give agents narrow tools and explicit boundaries
Agents become interesting when they can use tools: search a system, create a record, run a report, schedule work, or trigger an automation. They also become risky at exactly that moment.
Tool access should be specific, typed, and permissioned. Avoid a single command that lets a model execute arbitrary actions. Prefer focused operations such as find_customer_by_email, create_draft_invoice, or submit_access_request.
Each operation should validate its own inputs. The application, not the model, should determine whether the current user is allowed to perform the action. The application should also decide whether an action is reversible, requires confirmation, or needs approval.
A useful rule is simple: if an action would normally require a person to pause and verify it, the AI workflow should pause too.
Make failure a first-class path
Production systems do not succeed because nothing fails. They succeed because failures are contained and recoverable.
Model calls can time out. External tools can be unavailable. A result can fail schema validation. A request can be duplicated after a retry. An agent can reach a point where available information is insufficient. These are expected conditions, not exceptional signs that the feature was poorly designed.
Plan for them explicitly:
- Use timeouts and bounded retries for transient failures.
- Make write operations idempotent so a retry does not create duplicate records.
- Validate structured outputs before they reach downstream systems.
- Provide a safe fallback, such as a draft, queue, or human handoff.
- Log enough context to diagnose problems without storing sensitive material unnecessarily.
An AI system earns trust when it can say, in effect, “I could not safely complete this, so I preserved the work and routed it correctly.”
Evaluate the workflow, not just the response
Teams often test prompts with a handful of impressive examples. That is useful for exploration, but it is not evaluation.
A meaningful evaluation set reflects real variation: incomplete requests, ambiguous wording, unsupported questions, unusual formats, conflicting records, and adversarial attempts to bypass instructions. Test the entire workflow, including retrieval, permissions, validation, tool use, and fallback behavior.
Define success in operational terms. For a document extraction workflow, it may be field accuracy and the rate at which uncertain cases are correctly routed for review. For a support drafting tool, it may be policy compliance, factual grounding, and edit effort before sending. For an agent that changes records, it may be the absence of unauthorized or duplicate actions.
Evaluation should continue after launch. Production feedback reveals how users phrase requests, where context is missing, and which cases deserve automation versus escalation.
The software around AI is the moat
Models will improve, prices will change, and interfaces will evolve. The durable work is understanding a real workflow deeply enough to encode its rules, connect its data safely, observe its outcomes, and improve it over time.
That is why the most valuable AI systems rarely feel magical after the novelty fades. They feel dependable. They remove repetitive friction, surface the right information at the right moment, and leave people in control of consequential decisions.
Prompts matter. But architecture determines whether a promising demonstration becomes software that people cannot live without.