It Sounds Like It Remembers: The Memory AI Models Still Don’t Have

Published:

A model can answer as if it knows you. It doesn’t. It is reading the chat that is open in front of it. Close that chat and the facts leave with it. Start a new one next week and the model has no way to tell a fact that is still true from one that stopped being true. It may repeat the old one with the same confidence, because nothing inside the model marked it as old.

That is the problem this piece is about. A long prompt is not a memory. A cache is not a memory. And no shipping model was found, as of 3 October 2026, that keeps the right facts across weeks and drops the ones that went stale.

Picture a reply as many cheap steps adding up to one thought, if that helps. Treat it as a metaphor. It is not a claim that models work like brains.

First, What People Mean by Memory

Three different things get called memory, and only one of them is the hard problem.

A long context window is a bigger sheet of paper for a single call. Gemini 3.8 Flash, latest update September 2026, accepts up to 1,048,576 input tokens and can write up to 65,536. Mixtral’s window, from Mistral’s note of 11 December 2023, is 32,000 tokens. DeepSeek-V3’s model card says 128,000. More room for one conversation. Nothing about next month.

A prompt cache is a discount, not a diary. If you send the same opening text again, exactly, the repeat can be cheap. Change the text, or add tokens after the break, and you pay the normal price. Anthropic’s caching page, opened on 3 October 2026, prices a 5-minute cache write at 1.25 times normal input, a 1-hour write at 2 times, and a read at 0.1 times. The read is 0.025 times on Claude Fable 5.1 and Claude Mythos 5.1, and 0.05 times on Claude Opus 5.5. The page does not say when it was published.

OpenAI’s caching page, opened the same day, says cached input is discounted by up to 95 percent. For GPT-5.6 and later, a write is 1.25 times the normal input price and a read is 0.1 times, or 0.05 times on GPT-6.1 Sol. Output tokens are not part of the deal. That page does not date itself either.

A store beside the model is the thing that actually keeps notes. It is still not the model remembering.

What You Can Already Build

A check on 3 October 2026, not redone for this piece, looked at the tools that sit next to the model. Two limits matter here.

Graphiti keeps facts on a timeline, so a fact can expire. You bring the database. That is the closest shipping version of “this used to be true, and now it isn’t.”

Mem0, as of April 2026, uses an add-only algorithm. Its hosted test scores belong to the hosted product. They are not the open-source build, and they should not be used to judge the code you run yourself.

The other tools in that check write notes, hooks, or files an agent can read next time. Useful. Still a notebook. A chat log is still a chat log. Codex memories being off by default, and an OpenHands memory file you opt into at about 6,000 characters, were named in that same check. Those repositories were not opened again.

What Would Have to Be True

Three gaps are still open. None of them showed up as a product you can call.

The model would have to update, and forget. Google’s Titans paper, arXiv 2501.00663, is the research version of that idea. It updates a neural memory while it runs, and it has a forget gate. The experiments are research-scale models, not an API, and the paper should not be read as a deployed memory that replaces a chat log.

The practical version is narrower than the paper. Write a fact, stamp when it was true, and stop using it when a newer fact replaces it. Graphiti already does the expiring part in a graph you operate. Doing it inside the model, with no side store, was not found.

The hidden steps would have to stop inflating the bill. On 12 September 2024, OpenAI said o1 gets better when it spends more time thinking. The post does not say that time is cheap. The current reasoning guide says those thinking tokens are hidden from you, still take up context, and are billed as output tokens. A lower effort setting uses fewer of them. A higher one thinks more. Pro mode bills the extra work at that model’s normal rates, and it raises both token use and cost. The guide does not give its own publication date.

The bill is a price times tokens. A lower price does not cap it. Spend more tokens and you can pay more. No primary page checked here measured a named customer’s bill rising after a specific price cut. No page said the provider absorbs the hidden tokens.

There are cheaper paths for some calls, and only for those calls. OpenAI’s Batch API is 50 percent off the normal price if the job can wait up to 24 hours. Microsoft’s Foundry model router, version 2025-11-18, can send an easier call to a cheaper model. Balanced is the default. The page’s 1 to 2 percent and 5 to 6 percent figures are quality gaps against the best model for that prompt, not a cost discount. It does not say the router itself is free.

A long internal chain would have to cost about as much as one reply. No 2025 or 2026 primary source opened for this piece says that. NVIDIA’s NIM documentation, opened on 3 October 2026, describes something smaller. A draft model guesses several tokens and the larger model checks them in one pass. When the guess is right, several tokens cost about one check. When it is wrong, decoding continues as usual. That is not a chain of memory updates. The page has no publication date.

OpenAI’s Predicted Outputs is a real feature on the GPT-4o and GPT-4.1 series named on its page. The page does not call it speculative decoding. Tokens it rejects are still billed, so the feature can cost more, not less.

Mixture-of-experts is a different kind of saving, and it is about cost per token, not about remembering. Mistral, on 11 December 2023, said Mixtral picks 2 of 8 experts for every token. Mistral’s own figures are 46.7 billion parameters in total and 12.9 billion used, at the same speed and cost as a 12.9 billion parameter model. DeepSeek-V3’s card says 671 billion parameters in total and 37 billion active per token. The card’s needle-in-a-haystack score stays on DeepSeek’s page. The arXiv submission day was not read, because the abstract page timed out.

What an Honest Version Would Look Like

If someone built the missing piece, it would not look like a model that “just knows.” It would look like this.

Facts would carry a time. A new fact would replace the old one, so nothing stayed true forever by accident. The store would sit apart from the chat, so a long prompt would not be mistaken for a memory. Hosted scores would be labeled as hosted. Text you send again could use the cache. A fresh fact would be written at the normal price, because a write is not a read. Extra thinking would stay on the bill until a provider prices it at zero. None of them has said they do.

The Bottom Line

You can already give an agent a notebook, a very long sheet of paper, and a discount when you repeat yourself. You cannot yet hand it a memory that updates and forgets. The research sketch exists. The product does not. And until the hidden steps are priced differently, remembering more will keep showing up on the invoice as more tokens, not as one cheap thought.

Sources

The OpenAI reasoning guide is one source, not the only one. Undated pages were opened on 3 October 2026.

  1. OpenAI, reasoning models. https://developers.openai.com/api/docs/guides/reasoning
  2. OpenAI, “Learning to reason with LLMs,” 12 September 2024. https://openai.com/index/learning-to-reason-with-llms/
  3. OpenAI, prompt caching. https://developers.openai.com/api/docs/guides/prompt-caching
  4. OpenAI, Batch API. https://developers.openai.com/api/docs/guides/batch
  5. OpenAI, Predicted Outputs. https://developers.openai.com/api/docs/guides/predicted-outputs
  6. Anthropic, prompt caching. https://platform.claude.com/docs/en/build-with-claude/prompt-caching
  7. Microsoft, Foundry model router, version 2025-11-18. https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-router
  8. Mistral, “Mixtral of experts,” 11 December 2023. https://mistral.ai/news/mixtral-of-experts
  9. DeepSeek-V3 model card, citing arXiv:2412.19437. https://huggingface.co/deepseek-ai/DeepSeek-V3
  10. NVIDIA NIM, speculative decoding. https://docs.nvidia.com/nim/large-language-models/latest/advanced-use-cases/speculative-decoding.html
  11. Gemini 3.8 Flash. https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash
  12. Titans, arXiv 2501.00663. https://arxiv.org/abs/2501.00663
  13. Memory products checked 3 October 2026. https://app.notion.com/p/3ee93fe0b24081a196d5c76a14d4522f
TSN
TSNhttps://tsnmedia.org/
Welcome to TSN. I'm a data analyst who spent two decades mastering traditional analytics—then went all-in on AI. Here you'll find practical implementation guides, career transition advice, and the news that actually matters for deploying AI in enterprise. No hype. Just what works.

Related articles

Recent articles