ИТ развој

Refactor for Resilence: Architecting Systems Beyond the AI Hype

Рефакторирање за отпорност: Архитектирање системи надвор од возбудата околу ВИ

AI can generate a controller, suggest an index, and turn a vague ticket into a working-looking pull request in minutes. That is useful. It is not architecture.

The uncomfortable part of backend engineering has always been deciding what must remain true when dependencies are slow, data is incomplete, deployments go wrong, and requirements change. Those questions do not disappear when code arrives faster. In fact, they become more important because rapid generation can make a fragile design look finished.

Refactoring for resilience is the practice of making a system easier to understand, change, and recover. It is less about chasing an ideal design than deliberately reducing the cost of failure.

Start with the failure, not the framework

A system is resilient when it behaves predictably under stress. That includes obvious failures such as a database outage, but also quieter ones: a duplicate webhook, a delayed queue message, a partially completed batch job, or an API client retrying after a timeout.

Before extracting another service or introducing another abstraction, identify the important failure modes. For a payment-related endpoint, useful questions include:

  • What happens if the client submits the same request twice?
  • What happens if the payment provider succeeds but the response times out?
  • Can an administrator safely retry a failed operation?
  • Which state transitions are valid, and which must be rejected?
  • How will the team discover that processing is stuck?

These questions lead to concrete design choices: idempotency keys, database constraints, durable event records, explicit state machines, and useful logging. They are more valuable than a fashionable folder structure because they protect real business outcomes.

Make boundaries explicit

Many difficult bugs live at boundaries: between HTTP and application logic, between an application and its database, or between synchronous requests and asynchronous workers. Resilient refactoring makes those boundaries visible.

In PHP, a controller should generally translate an HTTP request into an application command and translate the result back into a response. It should not contain the complete lifecycle of an order, raw SQL, remote API calls, and retry policy in one method. The goal is not ceremony. It is giving each concern a home where it can be tested and changed safely.

final class CreateOrderController
{
    public function __invoke(CreateOrderRequest $request): JsonResponse
    {
        $order = $this->createOrder->handle(
            new CreateOrderCommand(
                customerId: $request->customerId(),
                items: $request->items(),
                idempotencyKey: $request->idempotencyKey()
            )
        );

        return new JsonResponse([
            'id' => $order->id()->toString(),
            'status' => $order->status()->value,
        ], 201);
    }
}

The command handler can own the transaction, validation against current state, and persistence. A separate integration component can own communication with an external service. This separation does not guarantee correctness, but it limits the blast radius when requirements evolve.

Use the database as a partner

Application code is not the only place where rules belong. If the business rule says one external payment reference can map to only one payment record, enforce that rule with a unique constraint. If a foreign key relationship must always exist, model it in the schema.

Database constraints remain effective when code paths multiply: web requests, console commands, queue workers, data repairs, and future integrations. They are especially valuable in systems where retries and concurrent processing are normal.

Consider an endpoint that creates invoices. A check-then-insert sequence can fail under concurrency: two requests can both observe that no invoice exists before either writes one. A unique index, combined with handling the resulting conflict intentionally, provides a stronger guarantee than an application-level check alone.

That does not mean every decision belongs in SQL. Complex workflow rules may be clearer in application code. The practical rule is simple: protect invariants at the lowest reliable layer available.

Design retries deliberately

Retries are not an error-handling afterthought. A network timeout does not reveal whether the remote system received a request. Retrying may be correct, or it may create two shipments, two emails, or two charges.

Safe retries require an operation to have a stable identity. For incoming requests, that may be an idempotency key stored alongside the resulting record. For outbound calls, it may be a provider-supported idempotency mechanism or an internal operation identifier included in the request.

When work can happen asynchronously, persist the intent before publishing the event. A transactional outbox is a common pattern: write the business change and an unpublished event in the same database transaction, then let a worker publish that event later. If publishing fails, the worker can retry without losing the business event.

The exact implementation varies, but the principle is stable: avoid designs that require a database commit and a network call to succeed as one imaginary transaction.

Keep Docker and deployment boring

Resilience is weakened when development, testing, and production behave like unrelated systems. Containers help when they make runtime assumptions explicit: PHP version, extensions, process startup, configuration, and health checks.

A production image should be small enough to reason about and deterministic enough to rebuild. Dependencies should be installed from a lock file, configuration should come from the environment or a managed secret store, and startup should fail clearly when required configuration is missing.

Deployment safety also deserves application-level support. Prefer backward-compatible database migrations where possible. Add a nullable column before requiring it. Deploy code that can read both old and new representations before removing the old one. Treat schema changes as staged transitions, not instant replacements.

Measure the system you actually have

Performance work is another place where confident-looking assumptions cause damage. An ORM relation loaded inside a loop may create an N+1 query problem; adding a cache may hide it while introducing invalidation complexity. Start by observing request timing, query counts, slow queries, queue latency, memory use, and error rates.

Then make the smallest change that addresses the evidence. Add an index after confirming the query pattern. Eager-load a relation when the response truly needs it. Paginate large collections. Move expensive, non-interactive work to a queue only after defining how failures and visibility will work.

Fast code that cannot be diagnosed is often slower to operate than slightly less efficient code with clear metrics and logs.

Refactor toward clearer ownership

The AI hype cycle can encourage teams to treat implementation as the scarce skill that has been solved. It has not. More code can now be produced with less effort, which makes judgment, review, and system design more scarce than before.

A resilient system has clear ownership of data, side effects, failures, and recovery. Its code can be changed without needing everyone to remember hidden assumptions. Its operational behavior is visible enough that a problem becomes a task, not a mystery.

The most durable refactoring is rarely the most dramatic. It is the steady removal of ambiguity: one invariant enforced, one retry made safe, one boundary clarified, one deployment made reversible. That is how software becomes capable of surviving both the next incident and the next wave of excitement.

Портрет на автор на блогот

Mihajlo

Јас сум Михајло - развивач поттикнат од љубопитност, дисциплина и постојаната желба да создадам нешто значајно. Споделувам увиди, упатства и бесплатни услуги за да им помогнам на другите да ја поедностават својата работа и да растат во постојано развивачкиот свет на софтверот и вештачката интелигенција.