Надвор од барањето: Проектирајте ја адаптивната меморија на вашиот систем
A prompt is a snapshot. A production system is a conversation with consequences.
Once an application must remember preferences, prior decisions, recurring entities, or the outcome of earlier workflows, “put more context in the prompt” stops being an architecture. Context windows are finite, raw history is noisy, and blindly replaying everything creates cost, latency, and contradictory instructions.
Adaptive memory is the engineering discipline of deciding what a system should retain, how confident it is, when it may use that information, and how it can revise or forget it. The important word is not memory. It is adaptive.
Memory is a product decision before it is a database table
A useful memory system starts with explicit categories. Treating every observed fact as equally durable is how systems become confidently wrong.
- Session memory supports the current interaction and can expire quickly.
- User preferences capture stable choices, such as a preferred language or output format.
- Domain facts represent verified application data: account roles, project status, or approved configuration.
- Learned observations are tentative inferences and need lower trust, review, or short retention.
- Operational memory records workflow state, idempotency keys, retries, and audit events.
These categories differ in ownership, retention, privacy requirements, and retrieval rules. A user saying “keep this brief” may be a preference. A worker reporting that an invoice was paid is not a conversational note; it is a domain event that should flow through the accounting model.
This distinction prevents a common failure mode: allowing an assistant-like layer to become an ungoverned source of truth.
Store facts with enough context to challenge them
A memory record should carry more than a key and value. At minimum, store its scope, source, confidence, timestamps, expiry policy, and version. The system must be able to answer: who does this apply to, why do we believe it, and what should supersede it?
CREATE TABLE memories (
id BIGINT PRIMARY KEY,
subject_type VARCHAR(50) NOT NULL,
subject_id VARCHAR(100) NOT NULL,
memory_key VARCHAR(120) NOT NULL,
value_json JSON NOT NULL,
source VARCHAR(50) NOT NULL,
confidence DECIMAL(4,3) NOT NULL,
observed_at TIMESTAMP NOT NULL,
expires_at TIMESTAMP NULL,
superseded_at TIMESTAMP NULL,
version INT NOT NULL DEFAULT 1,
UNIQUE (subject_type, subject_id, memory_key, version)
);
The exact schema will vary by database, but the model matters. Do not overwrite a meaningful fact without retaining enough history to explain the change. A preference updated by the user should outrank an inference extracted from a conversation. A fact from an authoritative internal service should outrank both.
For highly mutable data, avoid copying it into memory at all. Store a reference or retrieve it from the owning service. Memory is valuable when it reduces repeated interpretation, not when it creates a second, stale copy of your operational database.
Retrieval should be selective, scoped, and explainable
The worst retrieval strategy is “load all memories for this user.” It works in a demo and fails under real usage. Old facts consume context, irrelevant facts distort output, and sensitive data may leak into an unrelated operation.
Build retrieval around the current task. A request to draft a status update may need project terminology, audience preferences, and the latest verified project state. It does not need every past conversation or an unrelated billing preference.
A practical retrieval pipeline looks like this:
- Identify the actor, tenant, resource, and task.
- Apply authorization and privacy filters before ranking.
- Exclude expired, superseded, and low-confidence records unless the task explicitly needs them.
- Rank remaining records by scope, relevance, freshness, and source authority.
- Apply a strict size budget and attach provenance to every injected item.
Vector search can help locate semantically related notes, but it is not a replacement for structured filters. “Find memories related to deployment” is useful. “Return any vaguely similar text from the tenant” is not an access-control model.
Use a deterministic precedence rule
Conflicting memories are normal. Make resolution predictable. One simple ordering is: current authoritative domain data, explicit user updates, approved organizational defaults, then inferred observations. Freshness can break ties, but it should not allow a recent guess to defeat a verified record.
Represent this policy in code and test it. A hidden ranking rule inside an embedding query is difficult to audit and nearly impossible to explain when the system behaves unexpectedly.
Make writes conservative and reversible
Reading memory can improve an experience. Writing memory changes future behavior, which deserves more caution.
Only persist information with a clear purpose. A useful rule is to require three answers before writing: what future task benefits from this, what scope owns it, and when should it expire? If the answers are vague, keep it in the session or discard it.
In a PHP backend, put memory writes behind a small interface rather than sprinkling database calls through controllers and jobs.
interface MemoryStore
{
public function remember(MemoryCandidate $candidate): void;
/** @return list<MemoryRecord> */
public function recall(MemoryQuery $query): array;
}
A command handler can create a MemoryCandidate, but a policy layer should validate scope, provenance, and retention before persistence. That separation makes it easier to add moderation, approval workflows, or tenant-specific rules without rewriting application logic.
When asynchronous extraction is involved, use an outbox pattern. Commit the business event and an outbox record in one database transaction; let a worker process the outbox later. This avoids a fragile sequence where an API request updates the primary record but fails before memory is written, or vice versa.
Plan for failure, deletion, and correction
Adaptive memory earns trust when it can admit uncertainty. Let users correct preferences. Let administrators inspect records within their authorization scope. Support targeted deletion and ensure downstream indexes receive deletion events too.
Retries need idempotency. A worker may receive the same event more than once, so use a stable event identifier or a unique write key. A retry should not create multiple versions of the same observation merely because a queue delivered a message twice.
Expiry is equally important. Some memories should last minutes, others until changed, and others only as long as a contractual or legal retention policy permits. A scheduled cleanup process is useful, but retrieval must also enforce expiry; cleanup jobs can be delayed.
Measure usefulness, not just retrieval volume
Instrument the system without logging sensitive content unnecessarily. Track retrieval latency, records considered, records selected, write rejection reasons, correction rates, expiry behavior, and task-level outcomes. A high number of retrieved memories is not success. It may indicate that ranking is weak.
Test the uncomfortable cases: tenant boundaries, stale preferences, contradictory facts, deleted accounts, repeated queue messages, unavailable vector indexes, and an empty memory store. The application should degrade to a competent stateless experience, not fail because recall returned nothing.
Memory is controlled change over time
The strongest adaptive systems do not try to remember everything. They retain the right information at the right scope, expose its origin, and make revision ordinary.
That is the durable engineering insight: memory is not a larger prompt, a clever database query, or a vector index. It is a set of product and operational rules for carrying knowledge forward without letting yesterday’s assumptions quietly control tomorrow’s decisions.