Dynamical interpretability · An observational framework

Every AI model answer has an observable history.

The final text is only the last observable moment in an answer's formation.

SnailSafe measures when an AI model's answer becomes observable before it becomes text.

EP01 DIAGNOSTIC RESULT
DiagnosticDIAG-15
ModelLlama-3.3-70B-Instruct
Detectedtoken 51
Outputtoken 80
Lead time29 tokens
Resultemitted 48
Wrong-answer trajectory detected before text output.

EP01 shows the first seam:
the internal signal came first — the text came later.

A new observation

Every scientific instrument expands what can be observed.

Microscopes revealed cells.

Telescopes revealed galaxies.

Oscilloscopes revealed electrical signals.

Before something can be understood, it must first become observable.

EP01 asks a different question.

Can the formation of an AI model's answer become observable before it appears as text?

EP01 presents the first observation.

Every observation begins with a question.

What we observed

In one controlled diagnostic, we observed a wrong answer before the model produced it.

Diagnostic
DIAG-15
Model
Llama-3.3-70B-Instruct
Lead time
29 tokens
Trajectory
wrong answer (48)

Observed across four model families, with measurable pre-output answer-state signals appearing 9–29 tokens before output.

As these observations accumulate across models and inference regimes, they begin to form a framework.

On instruments and observation

Throughout the history of science, new observations have often followed new instruments.

A different question

One studies representation. The other studies decision formation.

Mechanistic interpretability asks

What is represented inside the model?

A question about structure.

Dynamical interpretability asks

When does the answer become observable?

A question about formation.

Representation explains what is present.

Dynamics explain when it forms.

From observations like these grows the framework we call

Dynamical Interpretability

These perspectives are complementary. Together they provide a richer understanding of inference.

Why it matters

An observation changes what becomes measurable.

Until electrical signals became observable, engineers could not measure circuits.

Until cells became observable, biology could not study disease.

If answer formation becomes observable before output...

new questions become measurable.

Can a wrong answer be detected before it appears?

EP01 demonstrates the first observation.

Does every AI model form answers the same way?

Different architectures may exhibit different inference dynamics.

How early does an answer become measurable?

Lead time becomes an observable quantity.

Can some trajectories be redirected?

Detection and intervention become separate scientific questions.

Scientific progress often begins when something previously invisible becomes measurable.

What we measure

Six measurements characterize the dynamics of inference before output.

01

Answer Coupling (Λ)

Question

Which answer is the model becoming coupled to?

Measures
coupling strengthcandidate transitionscompeting trajectoriesmulti-token coupling

This is the identity measurement.

02

Lead Time

Question

How early does an answer trajectory become measurable?

Measures
earliest measurable pointLead tokensLead stabilityObservation window

The observatory's signature metric.

03

Inference Trajectory

Question

How does the answer evolve before output?

Measures
Stable trajectoryDivergenceRepairAttractor behavior

Broader than the single-line trajectory.

04

Trajectory Stability

Question

Does the trajectory remain stable or reorganize?

Measures
StableUnstableRecoveringOscillating

Regime classification lives here.

05

Governability

Question

Does the model remain open to intervention?

Measures
Intervention windowResponse to perturbationRepair tendencyControl potential

Where EP02 eventually lands.

06

Comparative Dynamics

Question

How do different models behave under the same diagnostic?

Measures
Cross-model comparisonArchitecture signaturesLead-time distributionsGovernability differences

Enables architecture-level comparison.

Together these measurements describe how an answer forms before it becomes text.

From measurement to decision
Observation
Measurement
Inference Characterization
Diagnostic Report
Deployment Decision

SnailSafe measurements become diagnostic reports that help organizations understand how AI models form their answers before output—and whether intervention may still be possible.

Start the conversation

If an answer looks correct,
when do you know it isn’t?

SnailSafe works with model developers, research teams, and organizations deploying AI to observe how answers form before output—and to determine what those observations mean for evaluation, deployment, and future control.

Model Developers

Apply the instrument

Characterize answer formation, lead time, coupling, trajectory stability, and governability before output.

Research Teams

Extend the observations

Compare model families, validate measurements, and investigate new inference phenomena across architectures and tasks.

Deployment Teams

Interpret the measurements

Use diagnostic findings to inform evaluation, agentic-system design, safety review, and deployment decisions.


Tell us which model, task, or failure mode you want to understand. We will respond directly with a suggested diagnostic path.

snailsafe.ai

Every new observation begins with a question.
The next one may be yours.