Build Software That Teaches AI the Unwritten Rules of Your Business
Most businesses do not run on the rules written in policy documents. They run on the judgment people apply between the lines: when a customer deserves an exception, which fields matter in a messy request, how to interpret an incomplete handoff, and when a process that looks correct is actually heading toward trouble.
That is the opportunity behind useful AI software. The goal is not to build a chatbot that can recite a knowledge base. It is to build systems that help capture, apply, and improve the unwritten rules that experienced people use every day.
Start with decisions, not documents
A document repository is valuable, but it is rarely enough. Many operational rules only become visible at the moment someone makes a decision. A support specialist knows which billing issue can be resolved immediately. A sales operations team knows why two seemingly identical deals should be routed differently. A finance reviewer recognizes the combination of signals that calls for a second look.
Before choosing a model or designing an agent, identify the recurring decisions where expertise is currently implicit. Good candidates tend to have three properties:
- The decision happens often enough to matter.
- People can explain at least part of their reasoning after the fact.
- The cost of a wrong recommendation can be contained through review, limits, or escalation.
“Answer employee questions” is usually too broad. “Draft a response to a vendor onboarding question using approved policy, then flag missing information” is specific enough to design, test, and govern.
Turn tacit knowledge into usable context
Unwritten rules should not be treated as mysterious intuition that an AI model will somehow absorb. They need to be translated into evidence the system can use: examples, decision criteria, policy references, workflow state, and feedback from the people responsible for outcomes.
A practical pattern is to collect a small set of representative cases. For each one, record the input, the decision, the rationale, the relevant source material, and the outcome. Include edge cases deliberately. The ordinary cases teach the happy path; exceptions reveal where judgment actually lives.
Separate facts, rules, and recommendations
This separation is essential. Facts might come from a customer record, ticket, contract, or inventory system. Rules may come from formal policy or approved operational guidance. A recommendation is the model’s proposed interpretation of those facts and rules.
When these layers are mixed together, it becomes difficult to tell whether an answer is grounded, outdated, or simply wrong. A well-designed interface should make the distinction visible. For example, an agent reviewing a refund request might show the order details, cite the applicable policy section, and then propose an action with a confidence note and escalation option.
Build narrow agents with clear boundaries
The most reliable business AI systems usually begin as narrow tools. They receive a defined task, access a limited set of data and actions, and operate within a workflow that people already understand.
Consider an internal procurement assistant. Its first version should not negotiate with suppliers, approve spending, and modify contracts. It might instead classify an incoming request, identify required information, retrieve the relevant purchasing policy, and prepare a checklist for a human reviewer.
That scope makes the system easier to evaluate. It also reduces the temptation to grant broad permissions before the organization understands its failure modes.
- Give the agent a precise objective and stopping condition.
- Expose only the tools required for that objective.
- Require confirmation before irreversible or externally visible actions.
- Define escalation paths for low-confidence, high-impact, or conflicting cases.
- Log inputs, retrieved context, proposed actions, and final outcomes.
An agent does not become trustworthy because its prompts are lengthy. Trust comes from a combination of grounded context, constrained actions, observable behavior, and appropriate human control.
Use retrieval to ground answers, but do not mistake it for governance
Retrieval can help a model find relevant internal knowledge at the time of a request. This is often more practical than trying to encode every policy into a prompt or retrain a model whenever a procedure changes.
But retrieval quality is a product problem, not merely a search problem. Documents need ownership, dates, clear titles, sensible sections, and access controls. If the source material is contradictory, obsolete, or inaccessible, adding AI will surface those weaknesses faster.
Design answers so users can inspect why the system made a recommendation. A concise citation or link to the governing source is often more valuable than a polished paragraph. It gives the user a way to verify the result and gives content owners a concrete signal when guidance needs repair.
Make feedback part of the product
The first deployment is not the end of the learning process. It is when the system begins encountering the real variations that were absent from design discussions.
Capture feedback in a form that supports improvement. A simple thumbs-up or thumbs-down signal can indicate general satisfaction, but it rarely explains what changed. Better feedback options include “wrong policy,” “missing context,” “incorrect classification,” “needed escalation,” and “recommendation was useful but incomplete.”
Review this feedback with domain experts on a regular cadence. Look for recurring patterns rather than isolated failures. If users repeatedly override a recommendation for the same reason, that is not just a model issue. It may indicate a missing rule, an ambiguous workflow, or a policy that exists only in someone’s memory.
Measure business behavior, not conversational polish
Fluent output can create a false impression of competence. Evaluate the system against the decisions it is meant to support.
For a support triage agent, useful measures may include correct routing, unnecessary escalations, missed escalations, time to resolution, and reviewer acceptance. For a document assistant, evaluate factual grounding, citation relevance, completeness, and whether the response stayed within authorized scope.
Use a test set that includes realistic ambiguity, conflicting signals, incomplete records, and adversarial instructions. Test what happens when a required system is unavailable, when retrieval returns nothing useful, and when the agent lacks permission to complete an action. In mature systems, a safe refusal or escalation is often a better result than a confident guess.
Protect the people behind the process
AI systems that encode operational knowledge can affect jobs, authority, privacy, and accountability. Those effects should be discussed directly. Employees are more likely to contribute valuable expertise when they understand how it will be used, who can access it, and where human judgment remains essential.
Be especially careful with systems that influence hiring, performance evaluation, credit, pricing, access, healthcare, legal matters, or other consequential decisions. The right design may be decision support rather than automated decision-making. Technical capability does not remove the need for responsible ownership.
Make the unwritten visible without making it rigid
The strongest AI systems do not replace organizational judgment with a brittle rulebook. They make useful judgment easier to find, apply, question, and improve.
Start with one consequential but contained workflow. Preserve the evidence behind recommendations. Let experts correct the system. Treat exceptions as learning material, not inconvenience. Over time, the software becomes more than an interface to a model: it becomes a living map of how the business makes good decisions.
That is the durable value of AI in software work. It is not simply faster output. It is the chance to turn hard-won, scattered knowledge into a system that helps more people act with clarity when the written rules run out.