Your agent remembers the wrong balance

How a banking agent repeated a stale payment status as fact, and the memory design we built so it can't happen again.

Josiah Yahaya

Cloove

10 min read

A user asked our banking agent a simple question: "Is Tunde still pending?"

Tunde was a member of a savings group the user managed through Flow, our WhatsApp banking agent. Twenty minutes earlier, the agent had called a tool, read the group's contribution status, and correctly reported that Tunde hadn't paid his ₦200.

So the agent answered: "Yes, his ₦200 is still outstanding."

Tunde had paid in the meantime. Everyone had.

The agent only got it right after the user pushed back with "I don't think you are using the correct data." Then it called the tool, saw that everyone had paid, and corrected itself.

The data was right the whole time. The tool was right. The model wasn't hallucinating in the usual sense. It was quoting itself, and nothing told it that what it had said twenty minutes ago had an expiry date.

(Names changed.)

If you are building an agent that touches money, you will hit this bug. This is how we fixed it.


Conversation history launders stale data

Most agent frameworks store conversation history the same way: a list of { role, content } messages replayed to the model on every turn.

Here is what that list loses. When the agent answered "Tunde hasn't paid" the first time, that sentence came from a live tool read at 15:35. Once it was written into history, it was just prose. No timestamp. No source. Nothing to separate "I read this from the ledger a moment ago" from "the user told me their name is Ada".

On the next turn, the model sees its own earlier sentence and treats it as a standing fact. Why would it re-check? It already said so.

We call this history laundering: a volatile tool result goes in, and a timeless assertion comes out.

Fig. 1. Without provenance, the stored row keeps the content but not its source or time.

A bigger context window doesn't fix this. It makes it worse, because more stale claims survive longer. A better model doesn't reliably fix it either. The model is behaving sensibly with the information it has. The information is the problem.


Fix 1: Every reply carries provenance

The first change was small, and it closed the bug on its own before we built anything else.

Every tool the agent can call is classified by volatility, meaning how long its result stays true:

  • LIVE: true only at the moment it was read. A balance, a payment status, a position in a savings rotation. Never safe to repeat without re-checking.
  • EVENT: something that happened, and having happened is permanent. A transfer completed. A payout released.
  • STABLE: changes only when the user changes it. A name, a preference, a referral link.

When the agent replies, we store which tools that turn used, the worst volatility among them, and when the read happened:

TypeScript
enum MemoryVolatility {
  LIVE = 'live',
  EVENT = 'event',
  STABLE = 'stable',
}
 
interface ReplyProvenance {
  tools: string[]
  volatility: MemoryVolatility
  observedAt: string
}

When history is replayed to the model, any LIVE reply gets a visible stamp:

 
[Checked live at 15:35. Re-check before repeating any status, balance or position from this.]

The model reading its own earlier turn now sees a snapshot with a shelf life, not a standing fact.

Fig. 3. With provenance, the same question is answered from a fresh read.

Three details turned out to matter more than they look.

Classification is exhaustive and enforced by the type checker. The map from tool to volatility is typed as Record<AgentTool, MemoryVolatility>. If someone adds a tool without classifying it, the build fails. The alternative is a default, and every default is wrong in the dangerous direction: an unclassified read is far more likely to be live state than a stable fact.

Volatility describes what a result means, not how the tool is shaped. Our "get group invite link" tool sits right beside the live group-status tools, but the link never changes, so it's STABLE. The test we use is: "Would repeating this later, out of context, ever be wrong?" Whether the tool name starts with get_ doesn't matter.

The stamp is added at replay time, not write time. That way it always reflects the current clock, and every LIVE message already in the database got stamped the moment this shipped, with no backfill.


Fix 2: Summaries that are never shown a live value

Provenance fixed the incident. But we still had a flat "last 24 messages" window, and anything older simply disappeared. The standard answer is compaction: fold older messages into a rolling summary.

Compaction built on unannotated history would have made the original bug permanent. A stale claim that used to scroll out of the window after 24 messages would get compressed into a durable summary and replayed forever. That's why we shipped provenance first. The order mattered more than usual here.

Our compactor works like this:

  1. A background job collects messages that have aged out of the verbatim window.
  2. A fast model folds them into the existing summary under a word limit.
  3. A validator checks the result against a set of rules.
  4. If it fails, the model gets one retry with the broken rule named.
  5. Only a summary that passes is written.

The rules are specific to the domain. In a banking context, a summary must never contain a balance, a member's payment status, a position in a rotation, or a member count ("2 of 3 members have paid").

The first real run broke the design anyway

We ran it against a real development account. It folded 35 messages into 145 tokens and passed every rule. It was still wrong. The summary included:

  • "...₦699,990 locked until his NIN is verified"
  • "He needs to verify his NIN to unlock the full amount"

Neither is a balance or a payment status, so neither tripped a rule. But both are live state. Held funds get released without the user doing anything. An outstanding step stops being true the moment it's done.

We deleted the summary and added two rules: no money described as held, locked, pending or restricted, and no step still to do. A step goes into memory only once it's finished, as an event. "Has verified his BVN" passes. "Needs to verify his NIN" doesn't.

Then we stopped showing the compactor live answers at all

At first, the compactor saw LIVE replies labelled as such, with an instruction not to copy them. Models don't reliably follow that. In a dry run, older replies from before provenance existed carried "balance ₦999,990, ₦300,000 available" straight into a candidate summary.

The fix was structural, not another line in the prompt. Every LIVE assistant reply is replaced with a placeholder before the compactor sees it. The user's question stays in, so "the user asked about their balance" can still be recorded. The answer is never shown, so it can't be copied.

And because older replies have no provenance at all, unknown provenance counts as live whenever the reply quotes an amount of money.

A compactor can't leak a value it was never shown.

Fig. 4. One compaction job: rejected, retried once, written, cursor advanced.

The validator is strict on purpose

"Has paid" is rejected even when it describes a completed event, because the rule can't reliably tell "Tunde has paid" (event) from "Tunde has paid, so he's clear" (status). The costs aren't symmetric. A false rejection means a slightly stale summary for a few more exchanges. A false acceptance is the original incident, made permanent.

A failure mode we didn't see coming

A week later, a production account showed memory enabled, 36 messages waiting, and no summary. The cause: if both attempts on a batch were rejected, the job wrote nothing and didn't move forward. The next job retried the same batch, was rejected again, and so on, forever. One bad batch blocked that account's memory completely.

Now a rejected batch moves coverage forward and keeps the previous summary. Those messages drop out of memory, which is exactly what happened to them before the engine existed. A failed call (timeout, provider error) still doesn't move anything, so it gets retried. Rejected and failed are different outcomes and need different handling.


Fix 3: Facts must be grounded in the user's own words

Once you have a compactor, it's tempting to have it extract durable facts too: "the user's wife is Martha", "rent of ₦250k is due on the 25th". We do this in the same model call, with no extra cost. But a model extracting facts will happily paraphrase, generalise and invent.

So every proposed fact has to pass checks against the user's actual messages, not the model's summary of them:

  • Evidence must be a real quote from a user message (ignoring case and punctuation). No quote, no fact.
  • Every number in the fact must appear in the user's words. "5,000" matches "5000". A figure the assistant said doesn't count.
  • Confidence must be at least 0.7.
  • It must be lasting. A preference needs a standing instruction ("always", "call me", "from now on"). A commitment needs a recurrence. A goal needs saving language. Without this, "check all the transactions I did since account creation" was proposed as a preference, and "want to buy airtime" became three near-duplicate financial goals.
  • At most five per batch.
Fig. 5. Checks run in order and stop at the first failure.

Some kinds of fact are never inferred, however confident the model is: beneficiary nicknames, bill preferences, how transactions are categorised, and transfer purposes. Each one changes where money goes or how it's recorded. A wrong guess there is a financial mistake, not a memory quirk. The user has to state these explicitly.

Inferred facts never overwrite something the user said directly, and the agent is told which facts were stated and which were picked up, so it can say so when asked "What do you remember about me?"


Forgetting has to actually forget

"Forget my wife's name" has to mean the name is gone from everything the model sees. Deleting the saved fact isn't enough. The name may still be in the summary, and in older messages that will age out and get summarised later.

When a user forgets a fact:

  1. We read the fact first and build the list of words that must no longer appear: the key ("wife"), the note ("Wife is Martha") and any names in it ("Martha").
  2. We delete it. The audit log keeps keyed hashes of those words (HMAC with an app secret), never the words themselves. A plain hash of a name is easy to reverse with a list of common names.
  3. We queue a redaction that rewrites the summary without those words. If the rewrite can't be verified clean after one retry, the summary is cleared instead. Losing context is recoverable. Breaking "I've forgotten that" isn't.
  4. Older messages mentioning a forgotten word are dropped when they age out, so they never get folded back in.
  5. A fact is never re-learned from evidence older than the forget. If the user says it again later, that's new information.

"Forget everything" writes an empty summary as a tombstone instead of just deleting. Otherwise coverage resets to the start of the conversation and the next job summarises everything straight back.

Message and transaction history are never deleted. They're the banking record, not memory, and the agent says so if asked.


Rolling it out without betting the product on it

Every part of this sits behind one mode switch with three values:

  • disabled: the engine does nothing and returns the exact history it was given. Zero overhead.
  • shadow: the engine runs and logs what it would have sent, but the model still gets today's history. No user ever sees a different reply because of shadow mode, and a test verifies exactly that.
  • enabled: the engine's output is used, for an allowlist of accounts (or everyone).

And everything fails open. A config error, an unexpected message shape or a telemetry sink that throws all fall back to the untouched history. Annotating memory is best-effort. It must never be the reason a turn fails.

One more lesson from rollout: an unrecognised mode resolves to disabled. A typo in an environment variable must never switch on a feature that changes what users see.


What we learned

Compaction is about accuracy, not cost. Our first enabled log showed a 22-message conversation costing about 362 tokens, next to a roughly 19K-token cached prefix (system prompt plus tool schemas). Compressing ordinary chat history saves almost nothing. Its real value is carrying forward what was said beyond the window while never carrying forward a stale live value. Measure it on that.

Keep the summary out of the system prompt. We put it in per-turn context so the cached prefix stays byte-identical and prompt caching keeps working.

Enforce structurally, not with instructions. "Don't copy live values" in a prompt failed. Not showing the compactor live values worked. Wherever you can, remove the model's opportunity to make the mistake instead of asking it not to.

Real data breaks rules that pass tests. Every rule we added after launch (held funds, outstanding steps, member counts, durability) came from reading real output. The offending sentences are now regression tests.

The default for unknown state is "live". Unclassified tools, replies without provenance, anything we can't vouch for: treat it as perishable until proven otherwise.


The general principle

LLMs are good at deciding what to do and explaining it. They're bad at being the system of record for anything that changes. In a financial agent, the ledger is the truth, tools are how you read it, and memory's job is to remember that you asked, not what the answer was.


We're thinking about open-sourcing this

Everything above lives in a memory engine whose core is domain-agnostic: volatility and provenance, a token-budgeted history window, validated compaction, grounded fact screening and real forgetting. The banking rules plug in from outside.

We're considering releasing it, along with a few other pieces we built for the same reason (a provider-neutral engine for bounded decisions, a deterministic insights framework, and eval tooling for agent traces), as an open-source harness for developers building agents that touch money.

If you'd use it, or you've hit this bug in your own agent, I'd like to hear from you at chat@clooveai.com.

Follow along

New posts on how we build Cloove: agents that move money, the systems behind them, and what we learn in production.