Back to research

Research summary

The Narrative Continuity Test

A Conceptual Framework for Evaluating Identity Persistence in AI Systems

The Narrative Continuity Test asks whether an AI system remains a coherent interlocutor across time and gaps in interaction. It defines continuity as a joint property of memory, goals, correction, voice and role.

Beyond performance in a single exchange

An assistant can answer a question well and still fail to carry forward an earlier correction, retain the priority of an established goal or maintain a previously defined role. The NCT examines these longer-term properties, especially where an application presents an assistant as a continuing interlocutor.

The paper distinguishes retrieving information from integrating it into persistent priorities and commitments. Its concern is not simply how much an assistant can recall, but how remembered information constrains what it does later.

Five axes of narrative continuity

Situated memory
Remembering information together with its temporal and contextual relevance, and using priorities to decide what matters in the present interaction.
Goal persistence
Maintaining established goals and their precedence across interactions, while making justified revisions distinguishable from unprincipled drift.
Autonomous self-correction
Detecting inconsistencies, revising them and carrying the correction forward, rather than repeatedly repairing the same error only when prompted.
Stylistic and semantic stability
Preserving a coherent voice and meaning across time without requiring every response to have identical wording.
Persona and role continuity
Maintaining the boundaries and commitments of a role across sessions and contexts.

The framework treats these axes as interdependent. Better factual recall alone does not establish continuity if goals, corrections or role boundaries fail to persist.

A conceptual framework for future evaluation

The paper draws together work on memory, consistency, self-correction and model behaviour, then uses qualitative vignettes to illustrate possible continuity failures. It considers why a larger context window, retrieval or reactive filtering may improve immediate responses without establishing all five axes jointly.

It specifies a target for future testing rather than a ready-to-run benchmark. An operational evaluation would need to define observation periods, interruptions, relevant memory priorities and criteria for crediting corrections that persist over time.

Scope and limits

The NCT is a conceptual framework, not a validated score, a pass/fail certification or an empirical claim that particular products have passed a common test. The vignettes are illustrative rather than a systematic comparative evaluation.

It is not a test of consciousness, and it does not assume that every application needs a persistent interlocutor. Questions about the required time horizon, how to establish priority ground truth and whether particular architectures could satisfy the framework remain open for investigation. The practical aim is to make claims of continuity explicit enough to examine.

Source & citation

Public preprint on arXiv (2025); manuscript in revision at AI & Ethics. This summary reflects the revised manuscript. The linked public preprint may differ from that revision.

Natangelo S. The Narrative Continuity Test: A Conceptual Framework for Evaluating Identity Persistence in AI Systems. arXiv. 2025. Preprint. doi: 10.48550/arXiv.2510.24831.