~/wiki

conversational memory

---
title: Conversational Memory
category: concepts  
created: 2025-12-19
updated: 2025-01-04
tags: [conversational-memory, llm-evaluation, memory-systems, chatbot-design, long-term-context, user-preferences, openai, raw-derived-tradeoff, memory-architecture, dreaming-system, nine-axis-framework, evaluation-paradox, memory-failure-modes, infinite-context-limitations, cost-scaling]
sources: [raw/articles/Why long-term memory for LLMs remains unsolved.md, raw/feeds/2026-06-11-dreaming-better-memory-for-a-more-helpful-chatgpt.md]
confidence: high
---

# Conversational Memory

Challenge of maintaining coherent long-term context and relationship history in LLM-based conversational systems across extended interactions. Core technical and design problem affecting user experience in AI assistants that remains fundamentally unsolved despite apparent progress.

## The Dream vs Reality

The ideal: models that remember what users said and draw meaning across time over months or years. Not just recall, but interpretation, narrative, and continuous cumulative conversation flow.

Current reality: temporary illusions that work for days or weeks until the LLM starts forgetting and the illusion breaks. Every memory system creates lossy, opinionated, non-deterministic decisions that either become too large to search reliably or drift from actual conversations through repeated summarization.

## The Fundamental Problem: Raw vs Derived Tradeoff

As identified by chrys-bader, every memory system must choose between two preservation approaches:

1. **Raw** - Original messages stored verbatim (lossless but inert, buried in source material)  
2. **Derived** - Summaries, narratives, structured extractions (compact and usable but drift-prone through repeated derivation)

Neither extreme works, and solving both simultaneously requires perfect preservation AND perfect interpretation - the core unsolved challenge.

## Why Infinite Context Doesn't Solve It

The most intuitive solution fails for two reasons:

1. **Cost** - Processing years of conversation history on every turn creates brutal economics that scale linearly with history
2. **Degradation** - Models perform worse as context windows fill, with attention dropping on middle information and overall reasoning quality declining

Infinite context is just the extreme version of the raw path, which we know doesn't work alone.

## The Evaluation Paradox

Fundamental challenge in evaluating memory systems: you need ground truth to know if it's working, but for real conversational memory spanning months or years, the ground truth is larger than any context window and larger than any human can annotate.

Benchmarks like LongMemEval test needle-in-haystack retrieval, but retrieval isn't memory. Memory involves relationship arcs where significance becomes clear weeks later, facts change, old context gets superseded - and no benchmark captures evolving relationship arcs.

## Nine-Axis Framework for Memory System Design

Comprehensive framework by chrys-bader mapping all fundamental design decisions:

1. **What gets stored** - Raw vs derived spectrum
2. **When derivation happens** - Real-time vs batch vs on-demand  
3. **What triggers a write** - Every turn vs significant events vs user-driven
4. **Where it gets stored** - Filesystem vs vector DB vs graph vs hybrid
5. **How it gets retrieved** - Semantic vs full-text vs graph vs filesystem
6. **Post-retrieval processing** - Reranking vs filtering vs summarization
7. **When retrieval happens** - Always vs hook-driven vs tool-driven
8. **Who does curation** - Main model vs cheap model vs user vs hybrid
9. **Forgetting policy** - What/how/when to forget

## Common Failure Modes

Every memory system fails predictably:

- **Session amnesia** - New sessions start with no awareness of previous ones
- **Entity confusion** - Misidentifying or merging distinct entities during derivation
- **Over-inference** - Jumping to conclusions and encoding fabrications as facts
- **Derivation drift** - Chained summarizations compound small errors over time
- **Retrieval misfire** - Surfacing semantically similar but contextually wrong memories
- **Stale context dominance** - Old, heavily-referenced memories crowd out recent ones
- **Selective retrieval bias** - Only finding memories matching current query framing
- **Compaction information loss** - Specific details vanish when summaries replace raw turns
- **Confidence without provenance** - Stating "memories" with no way to trace back to source
- **Memory-induced bias** - Responses always colored by existing knowledge when fresh perspective needed

## Current State and OpenAI's Approach

OpenAI has experimented with a "dreaming" system for ChatGPT that processes conversations during idle periods to extract key information, preferences, and relationship context. However, this still faces the fundamental [raw-derived-tradeoff](/concepts/raw-derived-tradeoff) and doesn't solve the evaluation or drift problems.

## Why This Remains Unsolved

The problem is "very, very hard to solve" because:

- Compression is inherently lossy
- Retrieval is imperfect  
- The desired outcome (meaning that accumulates and evolves) may be the hardest thing to formalize in systems running on pattern matching over tokens
- No ground truth exists to validate success at realistic scales

Every current approach is "a different set of trade-offs dressed up as a solution" rather than a true solution to the fundamental constraints.

## See also

- [raw-derived-tradeoff](/concepts/raw-derived-tradeoff) - Core constraint in all memory systems
- [nine-axis-framework](/concepts/nine-axis-framework) - Systematic design analysis framework  
- [evaluation-paradox](/concepts/evaluation-paradox) - Why memory systems can't be properly evaluated
- [memory-system-failure-modes](/concepts/memory-system-failure-modes) - Common patterns of memory system failures
- [infinite-context-limitations](/concepts/infinite-context-limitations) - Why bigger context windows don't solve the problem
- chrys-bader - Researcher who developed this definitive analysis