Integracija umjetne inteligencije: prelazak s inženjeringa upita na agentske radne tijekove
Prompt engineering is useful, but it is not the finish line. A well-written prompt can produce a strong answer, a draft, or a code suggestion. It does not, by itself, create a dependable system that understands a goal, gathers the right context, takes bounded actions, checks its work, and knows when to stop.
That distinction matters as AI moves from a chat window into everyday software and business processes. The real opportunity is not simply asking a model better questions. It is designing agentic workflows: systems in which models reason over a task, use carefully chosen tools, maintain state, and operate inside clear technical and organizational guardrails.
From a single response to a controlled workflow
A prompt is an input to a model. An agentic workflow is a process around the model. The process may include retrieving information, calling an API, writing a draft, validating structured output, requesting approval, and recording the result for later review.
Consider a support-triage workflow. A prompt-only approach might ask a model to summarize an incoming ticket. That can be helpful, but it leaves the surrounding work untouched. An agentic version can classify urgency, retrieve relevant product documentation, identify missing diagnostic information, draft a response, and route high-risk cases to a human reviewer.
The model remains important, but it is no longer the entire application. Reliability comes from the workflow design: what the system is allowed to access, how it represents state, how it validates outputs, and where people retain control.
Think in capabilities, not magic
Agentic systems are often described as autonomous. In practice, the most useful systems are usually constrained. They have a small set of explicit capabilities and a clear definition of success.
A deployment assistant, for example, should not receive broad permission to alter any environment. It might be allowed to inspect a specific service’s health checks, compare a release version with an approved change record, and create a rollback recommendation. A separate, approved deployment mechanism can execute the change.
This principle makes systems easier to test and safer to operate. It also prevents a common failure mode: handing a model a vague objective and a powerful set of tools, then being surprised when it chooses an undesirable path.
Define the job before selecting the model
Start with the workflow, not the model. Describe the initial event, required inputs, allowed actions, expected output, and exit conditions. If those elements are unclear to a human operator, an AI agent will not make them clearer.
- Trigger: What starts the workflow: a ticket, pull request, form submission, alert, or scheduled review?
- Context: Which documents, records, and current-state signals are relevant?
- Actions: Which tools can the system call, and with what permissions?
- Decision points: What requires a rule, a confidence threshold, or human approval?
- Completion: What does success look like, and how is it recorded?
These questions produce a workflow that can be implemented with or without AI. That is a useful test. AI should improve judgment, language handling, prioritization, or adaptation where conventional code is brittle or expensive to maintain. It should not replace straightforward deterministic logic.
Build with deterministic boundaries
The strongest pattern is often a model inside deterministic boundaries. Let normal software handle authentication, authorization, retries, data persistence, schemas, and irreversible actions. Use the model where ambiguity genuinely exists: interpreting language, extracting intent, comparing alternatives, or composing a response.
Structured outputs are especially valuable. Instead of accepting a free-form answer and hoping downstream code can interpret it, require a defined shape that your application validates before acting on it.
{
"priority": "low | medium | high",
"category": "billing | access | defect | other",
"needs_human_review": true,
"summary": "short plain-language summary",
"next_action": "request_logs | draft_reply | escalate"
}
The application should reject malformed or incomplete output, provide corrective feedback when appropriate, and avoid taking an action merely because a field exists. A valid schema is not proof that a conclusion is correct; it is simply a safer interface between probabilistic reasoning and deterministic code.
Use tools as contracts
Tool calls should be treated like public APIs. Give each tool a narrow name, a clear description, typed inputs, explicit error responses, and the least privilege necessary. Avoid tools that accept vague commands such as “update the customer record” when a more precise operation such as “add an internal note” is sufficient.
Also plan for failure. A tool may time out, return stale data, reject an authorization check, or succeed without the model recognizing that it succeeded. The workflow should distinguish between a retriable technical failure, a business-rule rejection, and an uncertain outcome requiring review.
Retries need boundaries. Retrying a read-only lookup may be reasonable. Retrying a payment, deletion, or external message without an idempotency strategy can create duplicate or harmful actions. Agentic systems do not remove classic distributed-systems concerns; they make disciplined handling of them more important.
Context is a product decision
Models cannot make sound decisions from context they do not have, but more context is not automatically better. Large, unfiltered context can obscure the relevant facts, expose unnecessary sensitive information, and raise operating costs.
Retrieve information deliberately. For each workflow, identify authoritative sources and rank them. A support assistant may need the current account status and approved help content, but not unrestricted access to every internal conversation. A coding assistant may need the relevant repository files and build output, but not credentials or unrelated production data.
Context should carry provenance where possible. If an agent presents a recommendation, reviewers should be able to see which records or documents informed it. This is not only useful for trust; it makes debugging much faster when a result is wrong.
Human review is a design feature
Human-in-the-loop does not mean forcing a person to approve every low-risk action. It means placing review where judgment, accountability, or consequences demand it.
A useful approach is to divide actions by risk. Let the system draft, summarize, classify, and prepare routine work. Require review before it sends sensitive communication, changes permissions, modifies production configuration, commits financial actions, or makes decisions affecting people in meaningful ways.
Review screens should show the proposed action, supporting context, and a simple way to edit, approve, or reject it. A reviewer who must reconstruct the agent’s reasoning from logs will not trust the system for long.
Measure the workflow, not the demo
A compelling demonstration is easy to produce. A dependable workflow requires evaluation. Build a representative set of real task shapes, including awkward inputs, missing information, conflicting instructions, tool failures, and policy-sensitive cases.
Measure outcomes that matter to the workflow: correct routing, useful drafts, successful completion, unnecessary escalations, reviewer edits, and harmful actions prevented. Review failures by category. Is the issue missing context, ambiguous instructions, an unreliable tool, a weak validation rule, or a task that should never have been automated?
The goal is not to prove that an agent is intelligent. The goal is to establish that a specific system improves work while remaining observable, correctable, and safe enough for its scope.
The durable shift
Prompt engineering will remain a practical skill, much like writing a good query or API request. But the enduring work is systems design. It is deciding where intelligence belongs, what it may do, how it is checked, and how people recover when it fails.
The teams that get the most value from AI will not be the ones with the cleverest prompts. They will be the ones that turn uncertain model behavior into well-bounded workflows, pair autonomy with accountability, and treat every agent as a software component that must earn trust in production.