Behavioral evidence for agent operations
Investigate what your AI agent did.
GMAN turns observed agent activity into a source-aware session record of tools, files, decisions, denials, and delegation — demonstrated today on Claude Code, built for custom agents across models and frameworks. Facts, inference, and unknowns stay separate.
Developer preview · Claude Code demonstrated today · Metadata-only cloud record
2026-08-14 · 15:51:51 → 16:11:16 UTC
200 recorded activity rows over 19m · 103 tool calls · 1 write with success status · 5 denied decisions · 0 observed failures · outcome not derivable
Observed decisions
- accepted
- 75
- denied
- 5
co-occurrence only, not causality
Observed actions
- tool calls
- 103
- writes with success status
- 1
WebFetch ×66 · WebSearch ×26 · +5 more
- tool.calldecision deny15:56:23denied
- tool.calldecision deny15:56:23denied
- tool.calldecision deny15:56:23denied
task outcome not derivable from cloud evidence
One real session, accounted for.
- recorded activity rows
- 200
- tool calls
- 103
- denied decisions
- 5
- observed failures
- 0
Task outcome is not derivable from cloud evidence.
One sanitized Claude Code session. Every figure on this page comes from it.
See how the evidence is modeled.Investigation
Follow a signal to the exact events.
Attention appears only when there is something to look at. Related moments cluster by kind and time span, then expand to the individual events.
5 tool denials · 15:56:23–15:56:28 UTC
- tool denial·
- tool denial·
- tool denial·
- tool denial·
- tool denial·
- tool.callWebSearch · decision accept15:53:48success
- tool.calldecision deny15:56:23denied
- tool.calldecision deny15:56:23denied
- tool.calldecision deny15:56:23denied
- tool.calldecision deny15:56:27denied
- tool.calldecision deny15:56:28denied
showing 200 of total 200 · crop centred on the denial burst
Denied is a decision.
It renders indigo, never red. Red is reserved for an observed execution failure, and this session recorded none.
Co-occurrence only, not causality.
The cluster groups moments that share a kind and a time span. It does not claim they share a cause, and nothing here explains why the denials happened.
Evidence
Fact, inference, and unknowns stay separate.
GMAN keeps execution, decisions, normalization, source provenance, time anchoring, arrival finality, conflicts, and delegation as eight independent readings. It does not fuse them into a grade.
Source provenance
which capture sources evidenced each row
- joined
- 80
- otel only
- 92
- hook only
- 28
one of eight independent dimensionstwo of eight independent dimensionsthree of eight independent dimensions
Eight independent readings. Never fused into a grade.
Each dimension answers one question about the record. A high count in one is not evidence about another, so they are never averaged, scored, or reduced to a health colour.
Four distinct states that are not failures
- unsupported adapter coverage
- degraded normalization
- incomplete correlation
- independently established absence
Read the methodology.GMAN does not make the record look complete. It makes the record explainable.
Precise evidence
See what the record kept out.
Cloud-facing evidence is metadata-scoped. Previews and paths are absent, and the fields removed before delivery are listed on the event itself.
What was kept out
- content_scope
- metadata_only
- preview_state
- absent
- paths
- absent
stripped_fields
- organization.id
- terminal.type
- not established — a further field is recorded on this event; its name is cut off at the edge of the durable capture, and whether more follow it is not established
Full event detail
What this is
- occurred_at
- 2026-08-14, 15:53:04 UTC
- observed_at
- —
- normalization
- parsed
Why we believe it
- otel
- true
- hook
- false
- normalizer
- 0.1.0
agent attribution unavailable
Measured
- model
- claude-haiku-4-5-20251001
- output_tokens
- 21
- input_tokens
- not reported
- cache_read
- 0
- cache_create
- 0
- cost_usd
- duration_ms
- 1122
Kept local by design.
The cloud record carries metadata about the event, not the event's content. Previews and paths are absent because they are never sent, and the fields removed before delivery are named on the event itself.
A formatted number never hides its exact value.
The cost reads $0.0009. Focus it, or copy it, and the canonical value is there in full.
Stripped fields show construction, not discovery.
This is a list of fields the recorder removes before delivery. It is not a scan, not detection, and not a claim about what the local source material contains.
Delegation
Delegation, on the record.
GMAN recorded four subagent lifecycle moments in this session, two starts and two stops. None of them carries a known time, so they are labelled and listed last rather than placed in a sequence the evidence does not support.
- subagent.lifecycle · stoptime unanchored
- subagent.lifecycle · starttime unanchored
- subagent.lifecycle · starttime unanchored
- subagent.lifecycle · stoptime unanchored
Lifecycle moments
2 subagent start · 2 subagent stop
no known time on any of them
Agent family
2 Agent-family calls
tool calls of the Agent family, counted across the session
Time anchoring
unanchored time present · 28
informational · a separate dimension of the record
three independent readings · not summed, not paired, not mapped onto one another
Lifecycle observed. Time absent, and labelled.
A start and a stop are events the record captured. They do not say a subagent is running, how long it ran, how deep the delegation went, or that every subagent's stop was observed.
Identity is withheld, not missing.
Agent attribution is recorded for this session. An agent id is a real identifier, so it is not published here. That is a sanitization decision, and it is a different thing from a gap in the evidence — the record knows, this page does not say.
What a reader can reconstruct
- that two subagent starts were recorded
- that two subagent stops were recorded
- that the session made two tool calls of the Agent family
- that none of the four moments can be placed on the timeline
What a reader cannot conclude: how many distinct subagents ran, which lifecycle moment belongs to which Agent-family call, or which start pairs with which stop. The record does not evidence any of the three, so the page does not draw them.
How it works
Two sources, one record.
Evidence reaches GMAN from the machine the agent runs on. Nothing is scraped from a model provider, and no prompt or response text leaves the machine.
Capture fails open: if delivery is unavailable the agent keeps working. Delivery fails closed: an unauthenticated record is not accepted.
Observe. Native Claude Code hooks and OpenTelemetry both report activity from the machine the agent runs on. Two sources, so a row can be evidenced by one, the other, or both.
Hold locally. A durable local journal keeps what was captured. If delivery is unavailable, the agent keeps working and the record waits.
Deliver and normalize. Authenticated delivery sends a metadata-scoped record. Normalization states are recorded, not smoothed.
What arrives is metadata about the event, not its content. Which source evidenced each row stays on the row, so a reader can tell a joined record from a single-source one.
Read the methodology.Scope
Today, next, and where this is going.
Claude Code is the integration GMAN can demonstrate today, and the record it produces is the one this page has been showing. The system is being built for custom agents across models, providers, frameworks, and enterprise environments.
Today
What operators can use or see now.
- Investigate demonstrated Claude Code sessions.
- Inspect activity, decisions, actions, provenance, timing, and where the evidence stops.
- Read a metadata-scoped cloud record that discloses removed fields per event.
Next
Actively being built.
- A guided path from signup to a first recorded session.
- Recorder access beyond internal operation.
- Investigation summaries an operator can hand to a teammate.
North star
Where GMAN is going.
- Reconstruct custom-agent journeys across models and frameworks.
- Connect actions, delegation, evidence, and outcomes across enterprise systems.
- Give operators defensible answers about what an agent did, what the evidence behind it is, and what remains unknown.
A capability moves into Today when a session on this page can demonstrate it. That is the whole rule, and it is the one every claim on this site is written against.
Read the record for yourself.
Open the session account, then follow one moment to the event behind it.