Iza upita: Projektiranje AI sustava koji zaista djeluju
A prompt can produce an impressive answer in seconds. That is useful, but it is not yet a system that acts. The difference matters when an AI feature must retrieve the right information, choose a safe next step, call a tool, handle failure, and leave behind a result that people can trust.
The most valuable AI work is increasingly less about finding a clever prompt and more about designing the surrounding machinery. Models are one component in a larger product: they need inputs, boundaries, tools, state, evaluation, observability, and a clear definition of success. Treating them as autonomous magic is an efficient way to create fragile automation.
Start with a narrow job, not an “agent”
“Build an AI agent” is usually too vague to guide a technical decision. A better starting point is a specific workflow with a visible outcome. For example: classify inbound support requests, draft a response using approved documentation, and route uncertain cases to a human.
This framing immediately raises productive questions. What information is required? Which actions are allowed? What should happen when the model is unsure? Who owns the final decision? How will the team know whether the workflow improved anything?
A useful first version often has fewer moving parts than people expect. It may be a conventional application flow with one model call in the middle, rather than an open-ended loop that decides its own goals. Reliable systems earn autonomy gradually.
Separate reasoning from authority
A language model is good at interpreting unstructured input, generating alternatives, and translating between human language and structured data. It should not automatically be the authority that executes irreversible operations.
Consider an assistant that helps operations staff process refund requests. The model can extract an order number, summarize the customer’s issue, and propose an action. Your application should still validate the order, check policy rules, verify permissions, and create the refund through an audited service.
The distinction is simple: the model may recommend an action, while deterministic software decides whether that action is permitted. This makes behavior easier to test and reduces the chance that a persuasive but incorrect model output becomes a production incident.
Make tool contracts explicit
When an AI system can use tools, each tool should have a small, well-defined contract. Give it a name that reflects a business action, a typed input shape, an authorization boundary, and a predictable result. Avoid handing a model broad database access or a general-purpose shell and hoping the prompt will contain the risk.
For a ticketing workflow, an action might look conceptually like this:
{
"tool": "create_support_ticket",
"arguments": {
"customer_id": "string",
"summary": "string",
"priority": "low | normal | high"
}
}
The application should reject invalid values before calling the ticketing service. It should also record who or what initiated the action, the inputs supplied, and the resulting ticket identifier. The model does not need to know the implementation details; it needs a constrained interface that is hard to misuse.
Design for uncertainty as a normal state
Models can sound certain even when the available context is incomplete or misleading. Good AI systems do not try to eliminate uncertainty through wording alone. They make uncertainty operational.
One practical pattern is to classify outcomes into three paths:
- Proceed: the request is well understood and the action is low risk.
- Ask: required information is missing, ambiguous, or contradictory.
- Escalate: the action has material consequences, falls outside policy, or cannot be verified.
These paths should be part of the product flow, not an apology at the end of a generated answer. If an assistant cannot match a request to a known account with sufficient confidence, it should request an identifier or hand the case to a person. It should not improvise a match.
Human review is particularly valuable for actions involving money, legal commitments, sensitive personal data, security permissions, or public communication. The goal is not to keep people in every loop forever. It is to place review where mistakes are expensive and automation where repetition is safe.
Ground answers in controlled context
Many disappointing AI experiences are actually information-retrieval problems. A model cannot reliably answer questions about an organization’s current policies, product behavior, or internal decisions unless relevant context is provided at the time of the request.
That context needs governance. Decide which documents are authoritative, how they are updated, which users may access them, and what the assistant should do when no relevant source is available. A system that retrieves the wrong document confidently is not much better than one that retrieves nothing.
For user-facing answers, it is often useful to preserve links or identifiers for the materials used. This supports review, makes corrections easier, and gives users a way to inspect the basis for an important answer. The system should clearly distinguish sourced facts from generated recommendations.
Build the boring engineering around the model
The model call is rarely the hardest production concern. The surrounding software needs the same discipline as any other dependency: timeouts, retries where safe, rate-limit handling, input validation, logging, access control, and deployment controls.
Retries deserve special care. Retrying a failed text-generation request may be harmless. Retrying an action such as sending a message or creating a payment can produce duplicates unless the downstream operation supports idempotency. Give action requests an idempotency key, persist the result, and return the existing result if the same request is received again.
Observability should capture enough to diagnose behavior without casually collecting sensitive content. Track latency, tool failures, completion rates, escalation rates, and user corrections. Review samples under appropriate access controls. The most useful signal is often not whether the response sounded polished, but whether it led to a correct, completed outcome.
Evaluate workflows, not demos
A demo answers one hand-picked question. An evaluation tests a representative set of real tasks, awkward inputs, missing data, conflicting instructions, and failure conditions. Build a small test set before broad rollout, then expand it as new edge cases appear.
Measure the result against the job being done. A document assistant might be evaluated on answer accuracy, citation quality, and whether it appropriately says it cannot find an answer. A coding assistant might be evaluated on test outcomes, maintainability, and whether it respects repository conventions. An automation assistant might be evaluated on successful completion, reversibility, and the rate of human intervention.
Version prompts, tool schemas, retrieval settings, and model choices together with the evaluation results. Otherwise, a seemingly minor change can alter system behavior with no reliable way to explain why.
Adoption is a product and organizational design problem
AI changes work most effectively when it removes friction from an existing process rather than imposing a theatrical new one. Invite the people closest to the workflow to define failure cases, review early outputs, and decide what “useful” means. They usually know where the exceptions live.
Training should cover both capability and limits. People need to know what the system can automate, when to verify it, how to correct it, and how to report a bad outcome. Trust grows from predictable behavior and visible recovery paths, not from promises of intelligence.
The durable advantage is judgment
Models will continue to become more capable, but the essential engineering question will remain: what should this system be allowed to do, under what evidence, and with what accountability?
The teams that get lasting value from AI will not be those with the longest prompts. They will be the ones that pair model capability with clear contracts, trusted data, measured evaluation, and thoughtful human control. That is how a compelling conversation becomes a system that can actually act.