AI (Вештачка Интелигенција)

Beyond the AI Draft: Architecting Agents That Ship Real Software

Надвор од нацртот со ВИ: Архитектирање агенти што испорачуваат вистински софтвер

An AI draft is easy to admire. It appears in seconds, follows a prompt, and often looks close enough to production that the remaining work feels trivial. But software is not a document. It has interfaces, dependencies, state, users, failure modes, security boundaries, deployment constraints, and a long future of maintenance.

The useful question is not whether an agent can generate code. It is whether you can design a system in which that agent helps move a real change safely from intent to production. That requires architecture around the model: clear boundaries, reliable tools, verification, and human accountability at the points where judgment matters most.

Start with a bounded unit of work

Agents are most effective when they receive a task that has an observable outcome and limited blast radius. “Improve checkout” is an initiative. “Add server-side validation that rejects malformed postal codes, with tests for accepted and rejected inputs” is an implementable unit of work.

A good task definition gives the agent enough context to act without asking it to infer product strategy. Include the target behavior, relevant constraints, acceptance criteria, and the commands or tools available for verification. The goal is not exhaustive documentation; it is reducing ambiguity at the decision points that can produce the wrong implementation.

  • State what must change and what must remain unchanged.
  • Specify the component, service, or repository area in scope.
  • Define success in terms that a test, review, or user workflow can observe.
  • Call out sensitive areas such as authorization, payments, data migration, and public APIs.

This framing also makes delegation safer. If an agent fails, retries, or proposes an unexpected approach, the team can evaluate the result against a concrete contract rather than a vague sense of whether the output “looks right.”

Give the agent tools, not unlimited authority

A model without tools can explain how a change might work. An agent with tools can inspect code, edit files, run tests, query issue trackers, and prepare deployment artifacts. Those capabilities are what transform text generation into useful software automation, but each capability should have an appropriate permission boundary.

Read access is a natural starting point. Let an agent inspect relevant code, configuration, tests, and documentation before it proposes a change. Write access should be constrained to an isolated branch or workspace. Actions with irreversible or external effects deserve another boundary: creating a release, changing production data, sending customer communication, or altering access controls should require explicit approval.

Think of permissions as part of the product design. A broadly capable agent may save steps in a happy path, but a narrow toolset often produces a more dependable system because its possible mistakes are limited and easier to audit.

Design tools around stable contracts

Tool descriptions should be concrete about inputs, outputs, and side effects. An agent should not need to guess whether a deployment command is safe, whether a database operation is idempotent, or whether a failed request can be retried.

For example, a tool that creates a pull request can return its URL, changed files, and validation status. A tool that applies a migration should distinguish between generating a migration, checking it, and executing it. Small, well-defined actions are easier for both models and humans to reason about than one opaque command that does everything.

npm test -- --runInBand
npm run lint
npm run build

Even familiar commands need context. An agent should know whether these checks are local-only, whether they require credentials, and which failures are expected to block the change. Treat command output as evidence, not a ceremonial step.

Make verification a first-class stage

Code generation is a hypothesis. Verification is the process that turns it into a candidate change. The most reliable agent workflows separate planning, implementation, and validation rather than treating a successful file edit as completion.

Start with targeted checks: unit tests for the behavior changed, type checking, linting, and a build. Then add broader checks when the risk justifies them, such as integration tests, contract tests, browser flows, performance checks, or security scanning. The right suite depends on the system, but the principle is consistent: tests should examine the behavior the agent was asked to create.

Agents can also contribute by writing tests before implementation or by identifying untested edge cases. That does not eliminate review. Generated tests can reproduce the same mistaken assumption as generated code. A useful review asks whether the tests would fail if the intended behavior were absent or reversed.

Handle failures deliberately

Autonomy should not mean blind persistence. If a test fails, an agent needs a policy for deciding whether to investigate, retry, change its implementation, or stop for human input. Retrying a transient network operation may be sensible. Repeatedly changing code to silence an unfamiliar authorization failure is not.

Build workflows around explicit stop conditions. Examples include a failing security check, a database migration that cannot be proven safe, conflicting requirements, missing credentials, or a proposed change outside the approved scope. A clear escalation is often the most valuable output an agent can produce.

Keep humans responsible for decisions, not keystrokes

The best use of agents is not to remove developers from software delivery. It is to reduce the repetitive work that keeps developers from applying judgment. An agent can trace call sites, assemble a focused change, update routine tests, summarize a diff, and run standard checks. A developer can decide whether the behavior is desirable, whether the design fits the system, and whether the risk is acceptable.

Code review remains especially important for changes involving business rules, privacy, security, concurrency, data lifecycle, and user experience. Reviewers should inspect the diff, but also the agent’s assumptions: what files it read, what tools it invoked, what tests passed, and what uncertainty it reported.

A reliable agent system makes its work inspectable: intent, actions, evidence, and unresolved questions should all be visible.

Measure useful outcomes

Do not judge an agent program by the number of prompts submitted or lines of code produced. Look for outcomes that matter to engineering: shorter time from a well-defined task to a reviewed change, fewer repetitive interruptions, better test coverage for routine changes, clearer incident follow-up, and lower rework caused by misunderstanding.

Start with a workflow that is frequent, bounded, and reversible. Documentation updates linked to code changes, dependency upgrade preparation, test generation for existing behavior, or pull-request summaries are often good candidates. Learn where the agent is reliable, where it needs stronger context, and where human approval adds the most value.

Beyond the AI draft lies the real opportunity: software delivery systems that combine machine speed with engineering discipline. The agent is not the architecture. The architecture is the set of constraints, tools, checks, and decisions that allows useful work to move quickly without making trust optional.

Портрет на автор на блогот

Mihajlo

Јас сум Михајло - развивач поттикнат од љубопитност, дисциплина и постојаната желба да создадам нешто значајно. Споделувам увиди, упатства и бесплатни услуги за да им помогнам на другите да ја поедностават својата работа и да растат во постојано развивачкиот свет на софтверот и вештачката интелигенција.