Skip to content

Developer preview · Claude Code demonstrated today · Built for model- and framework-agnostic agent operations

Behavioral evidence for agent operations

Investigate what your AI agent did.

GMAN turns observed agent activity into a source-aware session record of tools, files, decisions, denials, and delegation — demonstrated today on Claude Code, built for custom agents across models and frameworks. Facts, inference, and unknowns stay separate.

Developer preview · Claude Code demonstrated today · Metadata-only cloud record

Agent Activity Session3f7c21e8…9d04

2026-08-14 · 15:51:51 → 16:11:16 UTC

200 recorded activity rows over 19m · 103 tool calls · 1 write with success status · 5 denied decisions · 0 observed failures · outcome not derivable

5 tool denials · 15:56:23–15:56:28 UTC
    • tool.calldecision deny15:56:23denied

task outcome not derivable from cloud evidence

Sanitized production session. Identifiers transformed. Counts are recorded evidence from this one session.

One real session, accounted for.

recorded activity rows
200
tool calls
103
denied decisions
5
observed failures
0

Task outcome is not derivable from cloud evidence.

One sanitized Claude Code session. Every figure on this page comes from it.

See how the evidence is modeled.

Investigation

Follow a signal to the exact events.

Attention appears only when there is something to look at. Related moments cluster by kind and time span, then expand to the individual events.

Attentionunanchored time present· informational · 28
5 tool denials · 15:56:23–15:56:28 UTC
  • tool denial·
  • tool denial·
  • tool denial·
  • tool denial·
  • tool denial·
    • tool.callWebSearch · decision accept15:53:48success
    • tool.calldecision deny15:56:23denied
    • tool.calldecision deny15:56:23denied

showing 200 of total 200 · crop centred on the denial burst

Denied is a decision.

It renders indigo, never red. Red is reserved for an observed execution failure, and this session recorded none.

Co-occurrence only, not causality.

The cluster groups moments that share a kind and a time span. It does not claim they share a cause, and nothing here explains why the denials happened.

Evidence

Fact, inference, and unknowns stay separate.

GMAN keeps execution, decisions, normalization, source provenance, time anchoring, arrival finality, conflicts, and delegation as eight independent readings. It does not fuse them into a grade.

Rows failing validation (content withheld)0

Source provenance

which capture sources evidenced each row

joined
80
otel only
92
hook only
28

one of eight independent dimensions

Eight independent readings. Never fused into a grade.

Each dimension answers one question about the record. A high count in one is not evidence about another, so they are never averaged, scored, or reduced to a health colour.

Four distinct states that are not failures

  • unsupported adapter coverage
  • degraded normalization
  • incomplete correlation
  • independently established absence

GMAN does not make the record look complete. It makes the record explainable.

Read the methodology.

Precise evidence

See what the record kept out.

Cloud-facing evidence is metadata-scoped. Previews and paths are absent, and the fields removed before delivery are listed on the event itself.

llm.requestsuccess

What was kept out

content_scope
metadata_only
preview_state
absent
paths
absent

stripped_fields

  • organization.id
  • terminal.type
  • not established — a further field is recorded on this event; its name is cut off at the edge of the durable capture, and whether more follow it is not established
Full event detail

What this is

occurred_at
2026-08-14, 15:53:04 UTC
observed_at
normalization
parsed

Why we believe it

otel
true
hook
false
normalizer
0.1.0

agent attribution unavailable

Measured

model
claude-haiku-4-5-20251001
output_tokens
21
input_tokens
not reported
cache_read
0
cache_create
0
cost_usd
duration_ms
1122

Kept local by design.

The cloud record carries metadata about the event, not the event's content. Previews and paths are absent because they are never sent, and the fields removed before delivery are named on the event itself.

A formatted number never hides its exact value.

The cost reads $0.0009. Focus it, or copy it, and the canonical value is there in full.

Stripped fields show construction, not discovery.

This is a list of fields the recorder removes before delivery. It is not a scan, not detection, and not a claim about what the local source material contains.

Delegation

Delegation, on the record.

GMAN recorded four subagent lifecycle moments in this session, two starts and two stops. None of them carries a known time, so they are labelled and listed last rather than placed in a sequence the evidence does not support.

Lifecycleactivity without a known time appears last
  • subagent.lifecycle · stoptime unanchored
  • subagent.lifecycle · starttime unanchored
  • subagent.lifecycle · starttime unanchored
  • subagent.lifecycle · stoptime unanchored

Lifecycle moments

2 subagent start · 2 subagent stop

no known time on any of them

Agent family

2 Agent-family calls

tool calls of the Agent family, counted across the session

Time anchoring

unanchored time present · 28

informational · a separate dimension of the record

three independent readings · not summed, not paired, not mapped onto one another

Lifecycle observed. Time absent, and labelled.

A start and a stop are events the record captured. They do not say a subagent is running, how long it ran, how deep the delegation went, or that every subagent's stop was observed.

Identity is withheld, not missing.

Agent attribution is recorded for this session. An agent id is a real identifier, so it is not published here. That is a sanitization decision, and it is a different thing from a gap in the evidence — the record knows, this page does not say.

What a reader can reconstruct

  • that two subagent starts were recorded
  • that two subagent stops were recorded
  • that the session made two tool calls of the Agent family
  • that none of the four moments can be placed on the timeline

What a reader cannot conclude: how many distinct subagents ran, which lifecycle moment belongs to which Agent-family call, or which start pairs with which stop. The record does not evidence any of the three, so the page does not draw them.

How it works

Two sources, one record.

Evidence reaches GMAN from the machine the agent runs on. Nothing is scraped from a model provider, and no prompt or response text leaves the machine.

Two capture sources converge into one normalized session recordNative Claude Code hooks and OpenTelemetry both report from the machine the agent runs on. Both feed a durable local journal. The journal delivers over an authenticated, metadata-scoped channel into the session record.native hooksOpenTelemetrylocal journalauthenticateddeliverymetadata-scopedsession record

Capture fails open: if delivery is unavailable the agent keeps working. Delivery fails closed: an unauthenticated record is not accepted.

  1. Observe. Native Claude Code hooks and OpenTelemetry both report activity from the machine the agent runs on. Two sources, so a row can be evidenced by one, the other, or both.

  2. Hold locally. A durable local journal keeps what was captured. If delivery is unavailable, the agent keeps working and the record waits.

  3. Deliver and normalize. Authenticated delivery sends a metadata-scoped record. Normalization states are recorded, not smoothed.

What arrives is metadata about the event, not its content. Which source evidenced each row stays on the row, so a reader can tell a joined record from a single-source one.

Read the methodology.

Scope

Today, next, and where this is going.

Claude Code is the integration GMAN can demonstrate today, and the record it produces is the one this page has been showing. The system is being built for custom agents across models, providers, frameworks, and enterprise environments.

Today

What operators can use or see now.

  • Investigate demonstrated Claude Code sessions.
  • Inspect activity, decisions, actions, provenance, timing, and where the evidence stops.
  • Read a metadata-scoped cloud record that discloses removed fields per event.

Next

Actively being built.

  • A guided path from signup to a first recorded session.
  • Recorder access beyond internal operation.
  • Investigation summaries an operator can hand to a teammate.

North star

Where GMAN is going.

  • Reconstruct custom-agent journeys across models and frameworks.
  • Connect actions, delegation, evidence, and outcomes across enterprise systems.
  • Give operators defensible answers about what an agent did, what the evidence behind it is, and what remains unknown.

A capability moves into Today when a session on this page can demonstrate it. That is the whole rule, and it is the one every claim on this site is written against.

Read the record for yourself.

Open the session account, then follow one moment to the event behind it.