For serious AI users

AI context
and decision
recall.

Continuum keeps a persistent working record around projects, decisions, reference material, and next steps, so the next AI run can resume with the right context instead of starting over.

Recall project context Bring forward the work history that shapes the next answer.
Recover prior decisions Keep the why, constraints, and tradeoffs close to the work.
Carry next steps forward Start the next run from what remains open.
Ground reference material Use the right project material without pasting the whole archive.

Current proof run: Continuum matched full-context answers in 5 of 5 live checks while reducing provider-reported input tokens by 85.09%.

Context and decision recall without the reset tax.

The problem is not that AI users lack models. The problem is that serious work carries history, constraints, decisions, reference material, and moving parts. Without persistent memory and working state, each new AI run starts by paying for the same explanation again.

Persist state

Start from the real state of the work.

Bring the next model into the project without rebuilding the whole story by hand. Continuum keeps useful working memory and task state available so the conversation can begin closer to action.

Retrieve less

Stop sending the whole archive when the task only needs a slice.

The current evals show that task-critical context can be preserved while dramatically reducing the amount of context sent into a model.

Handoff

Use the model that fits the moment.

Continuum is built for people who use more than one AI system. The work stays coherent even when the model changes.

Evaluate

Claims stay attached to the result.

Recall, answer quality, context reduction, and input-token reduction stay tied to the proof package, so the value is visible in the numbers instead of buried in a story.

The eval result: same recall, much less input.

The proof package tested a 146K-token working archive two ways: a deterministic recall eval across 15 scenarios, and a live model run across five tasks. The labels below are anonymized as S01-S15 and T1-T5.

The proof is accuracy held while context dropped.

The deterministic result is simple: Continuum preserved 100% of the required context in 15/15 scenarios, including 5/5 new scenarios, while reducing the working context by 98.8%.

The live result is just as direct: no project memory scored 0/5, full context scored 5/5, and Continuum scored 5/5 while using 85.09% fewer provider-reported input tokens than full context.

The practical comparison is model memory, not just full context. In the same deterministic eval, fixed-size memory-style lanes missed the new-scenario split while Continuum reached 5/5.

Answer accuracy Did the live model answer the task correctly? Current live run: Continuum 5/5.
Required recall Did the working context contain everything the task needed? Current eval: 100% expected output coverage, 15/15.
Token reduction How much less input was sent. Current live run: 85.09% fewer provider-reported input tokens.
Model-memory comparison Fixed memory-style lanes were cheap, but missed required context on new scenarios. Continuum stayed complete.
Figure 1. Model-memory proxy vs Continuum First-pass required recall on five new scenarios
Small model memory
0/5
Larger model memory
0/5
Continuum
5/5

This is the model-memory result. The fixed-size memory-style lanes stayed small but failed every new scenario in first-pass required recall. Continuum stayed small and complete: 5/5 new scenarios, 100% average recall, and about 99.1% less context than the full archive on that split.

Figure 2. Context load collapsed, recall held Deterministic eval average, 15 scenarios
Full archive
146,413 tokens
Continuum
1,775 tokens

This graph is the context-reduction proof. Continuum used about 1.2% of the full archive while preserving every required item in the current 15-scenario eval.

Figure 3. Recall held on new scenarios First-pass required recall
Known scenario set
10/10
New scenario set
5/5
Total
15/15

The new-scenario split matters: it checks whether the system generalizes instead of only passing tasks from the original set.

Figure 4. Live accuracy: memory was necessary, Continuum was sufficient Five-task live run
No project memory
0/5
Full context
5/5
Continuum
5/5

This is the answer-quality proof. No-memory answers failed the task set. Full context and Continuum both answered all five tasks correctly.

Figure 5. Live input tokens fell sharply Sum across five tasks
Full context
697,175
Continuum
103,954

This is the provider-reported usage proof. The live run shows an 85.09% input-token reduction across the five-task set.

Figure 6. The reduction held across every live task T1-T5 anonymized tasks
87.29%T1 input-token reduction
80.38%T2 input-token reduction
85.11%T3 input-token reduction
87.45%T4 input-token reduction
85.22%T5 input-token reduction

Every live task preserved correctness while using less input than the full-context lane. Per-task reduction ranged from 80.38% to 87.45%.

The comparison numbers.

The eval compared three practical choices: rely on small model-style memory, brute-force the full archive, or route the task through Continuum. The model-memory comparison is where the value becomes obvious.

Memory lane All-scenario first-pass passes New-scenario first-pass passes Average required recall Average context tokens Read
Small model memory 3/15 0/5 61.6% 1,320 Cheap, but missed task-critical context.
Larger model memory 7/15 0/5 79.4% 2,594 More memory helped known tasks, but still failed every new scenario.
Continuum 15/15 5/5 100% 1,775 Complete recall with a small, task-routed context load.
Live lane Correct answers Provider-reported input tokens Mean wall time Read
Full context 5/5 697,175 11.3s Correct answers with the full archive loaded.
Continuum 5/5 103,954 6.5s Matched full-context accuracy with 85.09% fewer input tokens and about 42% lower mean wall time.
Deterministic lane First-pass required recall Average context tokens Context reduction Read
Full archive complete by construction 146,413 0% Complete, but expensive and noisy.
Continuum 15/15 1,775 98.8% 100% required recall at 98.8% less context.

The proof in plain language.

Continuum kept the context that mattered, matched the full-context answer quality, and reduced the provider-reported input load. The eval checks working outcomes, not just whether similar text can be found.

Proof point Current result Why it matters
Required context stayed available Yes. 15/15 deterministic scenarios passed first-pass required recall, including 5/5 new scenarios. The model gets the working memory it needs without loading the full archive.
Live answer quality matched full context Yes. Full context: 5/5. Continuum: 5/5. The continuity layer kept the work usable while sending much less input.
Provider-reported input tokens dropped Yes. 103,954 Continuum input tokens vs 697,175 full-context input tokens, an 85.09% reduction. Less repeated context gets pushed through the model for the same working result.
The run moved faster Yes. Mean wall time was 6.5s with Continuum vs 11.3s with full context, about 42% lower. Focused context gives the model less to read before it can act.

How the eval was run.

The method is simple: use a real working archive, ask work-shaped questions, and compare full context against Continuum.

Real working archive The test used an 83-item working archive of about 146K tokens.
Required-recall eval Fifteen anonymized work scenarios were graded against the information each task required.
New-scenario split Five scenarios were held as new questions to check whether the system generalized beyond the original set.
Live answer eval Five live tasks were run three ways: no memory, full context, and Continuum.
Model-memory proxy The deterministic eval also compared fixed-size memory-style lanes against Continuum. Direct host-app memory APIs were not available in the Codex lab environment.

Where Continuum wins.

Continuum is built for the work where context is the advantage: long-running projects, multi-model teams, client history, files, decisions, and handoffs that need to stay current.

Long-running AI workflows

Keep decisions, constraints, active files, and next actions attached to the work so every new run starts closer to execution.

Multi-model handoffs

Move between models without losing the thread. The work stays coherent even when the provider changes.

High-context teams

Give the next person or model the useful working state without making them dig through every prior artifact.

Measured AI operations

Turn continuity into something visible: recall held, input dropped, and output stayed aligned with the work.

The value is strongest when the work already has real history. Continuum keeps that history usable, portable, and focused so long-running AI work can keep moving.

Bring your AI workflow to the consultation.

Continuum is for teams and serious individual users who are already doing real work with AI and feel the cost of starting over. The consultation maps where persistent memory, workflow state, and measured context reduction can help first.

Request consultation