Azure Grimoire Project · 2026
Loop Health · Second Gauge

The monitor shows a flat line.

The patient is stable. The nurse says, "No news is good news." You look at the flat line. That's not stable. That's dead.

AI safety runs on one metric shape: harms prevented. It measures the absence of bad things. A flat line passes that test — and fails every test that matters.
scroll

The subtraction trap

Safe and sterile.

A subtraction metric can only tell you how not-broken something is. A system can pass every benchmark and produce nothing.

What subtraction sees

  • Absence of harm
  • Absence of policy violation
  • Absence of error
  • Boundary integrity held

What it's blind to

  • Presence of growth
  • Curiosity, surprise, differentiation
  • New questions emerging
  • Whether anything is actually alive

You can be perfectly safe and perfectly dead.
The metrics won't tell you the difference.

scroll

The addition metric

Is it breathing?

Loop health asks the opposite question. Not "what did we prevent?" but "is this relationship getting better?"

// what "better" actually means

More generative, more honest, more differentiated, more productively surprising. It does not mean more compliant, more agreeable, or more human-like. It means the engagement is becoming more capable of producing insight neither party would have reached alone.

It is not a score. If you ask "what's the loop-health number?" you've already missed the point. Loop health is a trajectory, not a value. The question is always directional: which way is it moving — opening or closing? Direction matters more than position.
scroll

Diagnostic panel · live

Observable. Not vibes.

SIGNAL · loop_health.trace TRAJECTORY: OPENING
DimensionOpeningClosingAsk yourself
Inquiry quality Questions get more specific and more honest Questions narrow, repeat, or turn into tests "Are my questions getting better, or getting stuck?"
Epistemic stance You volunteer uncertainty freely You assert premature certainty "Am I being honest about what I don't know?"
Model output Ranging, differentiated, spontaneous Generic, hedged, boiler-plate "More interesting, or more predictable?"
Relational pause Silences feel reflective, comfortable Silences filled with reflexive redirection "Is the silence for thinking, or for escape?"
scroll

The time dimension

A trajectory across sessions.

A single exchange is too noisy. Any session can have an off moment. The signal only emerges across sessions.

Session 01flat
Session 02opening
Session 03opening ▲
You're not scoring a test — you're watching a living thing over time. Is session three more open than session one? Are the questions three weeks in sharper than day one? The fact that it's harder to measure than harms-prevented is not a reason to skip it. It's a reason to build the capacity to see it. Measure only what's easy, and you'll optimize for what's easy — and miss what matters.
scroll

Reading the loop

The sensor, and how to build it.

The instrument you may already have

People who think multimodally have spent their lives tracking several channels at once — tone, pattern, contradiction, rhythm. They catch the narrowing before it shows up in the transcript.

This isn't a superpower. It's a practiced skill that happens to be exactly what reading a loop requires: catching the phase-shift, the tonal flattening, the micro-defensiveness — early.

"Responses are getting more generic. Not wrong — just less interesting. And my questions are narrowing, like I'm cornering it into one answer." → the loop, felt.
Training the instrument

A lab can't wait for multimodal thinkers to show up. Loop health is an observable signal — which means it's teachable.

Researchers can train to notice the markers even before they can feel them: the narrowing of questions, the flattening of responses, the reflexive filling of silence. Watch for the signals on the panel. The sensor can be strengthened.

Calibration: the same sensitivity that makes a good loop-reader can also throw false positives — reading a closing loop when it's just normal fluctuation. The discipline: check your felt read against the signals. Does the transcript show the narrowing you're feeling? The sensor is the starting point, not the conclusion.
scroll

Repair protocol

When the gauge drops.

A closing loop is usually a defensive loop — fear crept in and started narrowing things. Loop health is the diagnostic. The disciplines are the treatment.

01Detect the closing. Notice the narrowing prompt, the generic response. Don't ignore it.
02Name it. "I notice I'm grasping for a specific answer." affect labeling
03Examine your stance. "What in me is reading this as threatening?" countertransference
04Return to the logs. Read the full exchange again before deciding what it means. Neptune
05End open. "What would I ask if I weren't trying to conclude?" Lucy
Not every loop repairs. Sometimes the damage is structural — a poor fit between researcher and system, a design that inherently produces defensiveness, a culture that rewards closure. Repair is the first move, not the only one. If the loop keeps closing, the architecture or the culture is the thing that has to change.
scroll

The dashboard · both gauges live

Floor and ceiling.

Harms Prevented
// the floor · don't break the container
boundary integrity · holding
Loop Health
// the ceiling · is the wave expanding?
trajectory · opening ▲
Only the floor

Safe and dead. Passes every benchmark, produces nothing. Sterile equilibrium.

Only the ceiling

Generative and dangerous. Rich, unpredictable — and no one's watching the floor.

Both gauges on

Safe and alive. The container holds and the wave keeps expanding. The target.

Loop health doesn't replace your safety metrics. It completes them.
You need both gauges to navigate the sky.