ИТ развој

Stop Guessing API Performance: Measure, Don't Assume

Престанете да нагаѓате за перформансите на API: Мерете, не претпоставувајте

The Feeling of Fast Is Not a Metric

An API can feel fast in a browser, pass a quick manual test, and still disappoint under real traffic. That gap is where expensive assumptions live.

“It only does one database query” is not a performance conclusion. Neither is “the response payload is small” or “Docker is not the bottleneck.” These may be useful hypotheses, but they are not evidence. Performance work begins when a request becomes measurable from the caller’s perspective and traceable through the system.

The practical goal is not to make every endpoint as fast as possible. It is to understand which work is slow, when it becomes slow, and whether changing it improves the experience or reliability that actually matters.

Start With a Question You Can Measure

Vague goals produce vague optimizations. Replace “make the API faster” with a specific question: does GET /orders remain responsive when the caller requests a large page? Does creating an invoice become slower as an account accumulates records? Does latency increase only after a container restart, cache eviction, or database connection surge?

Define the request shape before measuring it. Include the authenticated role, request parameters, representative data volume, and expected response. A query that is fast against ten rows may behave very differently against a production-sized table with uneven data distribution.

For an endpoint, collect at least:

  • request rate and concurrency;
  • response time, including percentile values such as p50, p95, and p99;
  • error and timeout counts;
  • database query count and query duration;
  • external service calls and their duration;
  • CPU, memory, connection, and queue pressure where relevant.

An average response time can hide the problem. If most requests complete quickly while a smaller group stalls on a lock, a slow dependency, or a cold cache, the average can appear acceptable. Percentiles show how the slower part of the experience behaves.

Measure the Whole Request, Then Narrow the Scope

Begin at the edge: measure the elapsed time observed by the client or load-testing process. That number includes routing, middleware, application execution, serialization, network overhead, and downstream work. It is the right starting point because it represents the service being delivered.

Then break the request into meaningful spans. In a PHP application, that might include authentication, controller work, database access, a cache lookup, an HTTP call, and response encoding. Avoid timing every tiny function at first. Excessive instrumentation creates noise and can make the signal harder to interpret.

$startedAt = hrtime(true);

$result = $repository->findOrdersForCustomer($customerId);

$elapsedMs = (hrtime(true) - $startedAt) / 1_000_000;

$logger->info('order lookup completed', [
    'customer_id' => $customerId,
    'duration_ms' => $elapsedMs,
    'result_count' => count($result),
]);

This is useful for targeted investigation, not as a permanent substitute for structured observability. Log enough context to compare like with like, but do not record secrets, tokens, full personal data, or entire request bodies simply because they are convenient.

When a span is slow, ask what it is waiting for. CPU work, a database round trip, lock contention, disk I/O, DNS resolution, connection setup, and an upstream timeout can all look like “a slow controller” if the only measurement is at the controller boundary.

Databases Reward Evidence

Database queries are a common target because they are visible, but visibility does not prove they are the dominant cost. Count queries per request and measure their total duration before redesigning repositories or adding indexes.

Once a particular query is implicated, inspect its execution plan against representative data. The key question is not whether the SQL looks elegant. It is whether the database can find, join, sort, and return the needed rows efficiently for the actual filter and ordering pattern.

For example, a paginated endpoint may appear healthy until users reach later pages. Offset-based pagination can require the database to walk past rows that will not be returned. A cursor-based approach can be a better fit when the endpoint has a stable ordering and clients can continue from a known position. But it changes the API contract, so it should solve a measured problem rather than satisfy a fashion for “modern pagination.”

Beware the N+1 pattern as well. Loading a list and then fetching related data once per item may be harmless for a short list but costly for a larger one. Measure query count and total database time, then choose between eager loading, a join, a batch lookup, or a deliberately separate endpoint based on the response contract.

Load Tests Need Realistic Boundaries

A single benchmark command can tell you whether an endpoint responds, but it does not automatically describe production behavior. A useful test states what it is exercising: one application container or several, a local database or a managed one, warm cache or cold cache, isolated endpoint or competing traffic.

Increase load gradually. Watch latency, errors, resource use, database connections, and queue depth together. The first limit is often not the component developers expected. More PHP workers may shift pressure to the database. A larger connection pool may increase contention. More replicas may not help an endpoint dominated by a single write path.

Run the same scenario more than once. Transient network variation, cache state, background jobs, and container startup can all affect a result. A benchmark is more useful when it compares a controlled before-and-after change than when it attempts to produce one impressive number.

Make Failure Part of the Measurement

Fast success responses are only half the story. Set explicit timeouts for outbound dependencies, distinguish retryable failures from permanent ones, and measure what happens when an upstream service slows down.

Retries deserve particular caution. Retrying a safe, idempotent read after a transient failure can be reasonable. Retrying a non-idempotent write without an idempotency strategy can create duplicate work. Even safe retries add load precisely when a dependency may already be struggling, so their limits and backoff behavior should be deliberate and observable.

Change One Thing, Then Prove It Helped

Performance improvements are easy to over-credit. If a deployment also changes container settings, cache state, query code, and application logging, a better result does not identify the cause. Make a focused change, rerun the same workload, and compare the metrics that motivated the work.

Keep the outcome in maintainable form. A cache without invalidation rules is deferred correctness work. A hand-written query that nobody can safely modify is operational risk. A configuration change that relies on undocumented container behavior is a future incident waiting for a busy day.

The best performance work leaves behind more than a lower latency graph. It leaves a clearer request path, a measurable service-level expectation, and a team that knows where to look when behavior changes.

Replace Confidence With Feedback

Good engineering judgment still matters. It helps form the hypothesis, choose the measurement, and recognize a misleading result. But judgment should not be asked to impersonate a profiler, a trace, a query plan, or a load test.

Measure the real request. Find the expensive work. Change the smallest thing that addresses it. Measure again. That loop is less dramatic than guessing, but it is how an API becomes reliably fast instead of merely confidently described as fast.

Портрет на автор на блогот

Mihajlo

Јас сум Михајло - развивач поттикнат од љубопитност, дисциплина и постојаната желба да создадам нешто значајно. Споделувам увиди, упатства и бесплатни услуги за да им помогнам на другите да ја поедностават својата работа и да растат во постојано развивачкиот свет на софтверот и вештачката интелигенција.