Search Results

Inspector

The Dojo includes an inspector that exposes structured snapshots of the running sandbox via a simple file-based protocol. It was designed so that AI agents, scripts, or any external tool can verify what the application is doing — without needing a GUI debugger or an undocumented binary protocol.

How It Works

The inspector lives inside the Dojo process. It watches a request file, and when that file changes it reads the sandbox state and writes a response file:

<sandbox-dir>/.clay/inspect/
├── request.json      ← you (or your tool) write this
├── response.json     ← the inspector writes this (atomically)
├── state.json        ← lifecycle: phase, pid, protocolVersion, generation, reloadCount, openAnnotations
├── events.jsonl      ← append-only event stream (5 MB one-level rotation)
├── log.jsonl         ← every console/Qt message: ts, level, category, text
├── dojo.json         ← dojo supervisor state (generation, crash info)
├── crash.json        ← after crash loops: exit info + child output tail
├── autoflag_*.json   ← auto-captured evidence bundle on runtime errors
├── screenshot.png    ← default capture target; a request can name its own path
└── trace.jsonl       ← written during trace recording

The .clay/ directory is created automatically in the directory where the sandbox QML file lives.

Correlation: put a unique "id" into every request — the response echoes it as "requestId" so stale responses (from earlier roundtrips or a previous process generation) can be rejected. state.json carries protocolVersion (currently 3) so tools can check capabilities before relying on them, and runId — unique per loader process. On startup a loader removes a previous run’s state.json/response.json; drivers relaunching an instance should still either clean the instance dir first or wait for runId to change, since a just-spawned process needs a moment before its first write.

One request at a time: wait for your reply before writing the next request. Blocking actions (waitForRoot, a settling capture, a batch) spin the event loop, so a request written while one is in flight would be answered into the same response.json and lost; the loader drops it instead and reports reentrantDropped on the response it does send. Use batch when you want several things done without waiting in between.

Multiple instances: networked games run several instances of the same sandbox. Start each loader with --instance <name> and it serves its protocol under .clay/inspect/i/<name>/ instead (with instanceId stamped into its state.json); the flat layout remains the single-instance default. dojo.json stays at <sandbox-dir>/.clay/inspect/dojo.json in every case — there is one supervisor per sandbox dir, whatever the instance layout.

The status envelope

Every response carries a status object, whatever the action was (protocol v3):

"status": {
  "alive": true, "rootLoaded": true, "generation": 7, "phase": "ready",
  "reloadCount": 9, "runId": "84b55d33", "supervised": true, "restarts": 3,
  "sandbox": "/path/Sandbox.qml",
  "renderedAt": "2026-08-01T10:21:07.412",
  "lastError": "child exited 0 (normal)", "lastErrorAt": "2026-08-01T10:20:58.101"
}
Field Meaning
alive The inspector reached the point of writing this response. Not a measurement — a dead inspector writes nothing, but a stale response.json still says true, so pair it with requestId.
rootLoaded A root object exists.
generation Successful loads. Does not advance on a reload that failed, which is exactly what makes “am I measuring the scene I edited?” answerable. reloadCount counts attempts.
phase, runId, sandbox, instanceId The same facts as state.json, so one response answers them.
renderedAt Present only when this response carries a capture, and set to the moment the image was grabbed. A snapshot with "screenshot": true that returns no renderedAt produced no image (see screenshotError); any PNG on disk is then from an earlier request.
supervised, restarts Read from the dojo’s dojo.json. restarts counts respawns, i.e. every child after the first.
lastError, lastErrorAt The newer of the most recent QML error and how the supervisor’s last child died — so a crash loop is visible even from a loader that logged nothing.
supervisorGaveUp Present and true when the dojo stopped respawning (see dojo.json’s gave_up phase and its reason).

Because status is reserved for the envelope, actions that acknowledge themselves use their own key: reload answers reloadStatus, trace answers traceStatus, annotate answers annotateStatus. id is reserved the same way — it is the caller’s request-correlation id on every action, which is why annotate selects its target with annotationId.

On the supervisor side, claydojo forwards the child’s stdout as well as its stderr (a rejected command line prints its usage to stdout), rejects unknown positional arguments and missing --sbx files at the front door instead of passing them to the child, and stops respawning once the crash threshold is reached without any child ever having run stably — writing phase: "gave_up" and a reason into dojo.json and printing both. A silent infinite restart loop is the worst failure mode for headless use.

Actions

snapshot — Point-in-time state

{
  "action": "snapshot",
  "screenshot": true,
  "eval": ["player.health", "world.room.children.length"]
}

Returns: rootProperties (auto-captured primitive properties on the sandbox root), flagInfo (if the root defines a flagInfo() function), viewState (if the root implements it, see below), eval results, logTail (last 50 log entries), warnings, errors (plain strings — the errors action carries file and line), and optionally a screenshot path. A grab that failed returns screenshotError instead, and leaves status.renderedAt unset.

Capture — framing, settling, comparison

"screenshot": true writes .clay/inspect/screenshot.png, unchanged. Passing an object instead tells the renderer where the picture goes and how it is framed — the renderer already knows the framing, so nothing is left to copy, downscale or crop afterwards:

{
  "action": "snapshot",
  "settle": true,
  "screenshot": {"path": "shots/hud.png",
                 "crop": {"objectName": "hud"},
                 "scale": 0.5},
  "diff": "shots/hud-baseline.png"
}
Key Meaning
screenshot.path Where to write. A relative path resolves under .clay/inspect/, never the loader’s working directory; parent directories are created. Omit it to keep the default .clay/inspect/screenshot.png. Absolute paths go wherever you say — but not into the sandbox directory: the dojo watches that whole tree (it skips only .clay/), so a capture written there triggers a reload and your next step measures a different scene.
screenshot.crop [x, y, width, height] in captured-image pixels, or {"objectName": "..."} to frame exactly that item. Applied before scaling. A crop that lies outside the viewport is a screenshotError, not a silent clamp — a picture of the wrong thing is worse than no picture.
screenshot.scale, screenshot.width Downscale by a factor, or to a target width in pixels (width wins over scale).
settle true, or {timeoutMs, stableFrames, intervalMs, tolerance}. Waits for the picture to stop changing before grabbing it.
diff A baseline PNG path, or {"baseline": "...", "tolerance": N}. Compares the capture just taken against that file.

The response says what was actually produced: screenshot (the path written), screenshotSize ({width, height} after crop and scale) and status.renderedAt (the moment those pixels were grabbed).

settle answers {settled, waitedMs, framesCompared, lastDelta}. It compares successive frames instead of asking the animation system, so it covers animation, physics and shader-driven motion alike. settled: false means the timeout hit while the scene was still moving — a bounded, reported fact rather than an error, because a scene in continuous motion never settles. Read it before trusting the picture: it is the difference between “quiet” and “gave up while it was still moving”. settle works without a capture too, as a plain “wait until it is done”.

diff answers {baseline, tolerance, delta, changedPixels, changedBounds}, where delta is the fraction of pixels beyond tolerance (0..1) and changedBounds the rectangle they occupy (absent when nothing changed). tolerance defaults to 2 per channel, which absorbs the jitter between two GPU renders of the same scene. A baseline that cannot be read is a diffError — never a silently skipped comparison, which is how a regression check passes forever without comparing anything. diff without screenshot compares without writing a file.

settle waits for the picture; the time action controls the clock (pause, single-step, timescale). They are separate knobs, and a deterministic frame usually wants both.

errors — QML errors and warnings since load

{"action": "errors"}
{"action": "errors", "sinceGeneration": 8}

Returns errors and warnings as objects — {generation, ts, text, file, line}, with file/line filled in when the message carried a location — plus errorCount, warningCount and truncated (each buffer holds 200 entries). Unhandled QML/JS exceptions (TypeError, ReferenceError, …) arrive as warnings, so read both lists.

The buffers are cleared before every reload, so a bare errors request means “what has gone wrong with the current scene”. sinceGeneration: N drops everything older than generation N. Diagnostics raised while a load is in flight are tagged with the generation being attempted, so a reload that failed leaves its errors at status.generation + 1 — findable even though the generation itself never advanced.

annotations / annotate — the user’s remarks about the running scene

{"action": "annotations"}
{"action": "annotations", "status": "open", "sinceGeneration": 8}
{"action": "annotate", "annotationId": "a7", "note": "raised the button to 44px and re-centred the label"}

The user frames regions over the running sandbox and writes a note on each (Ctrl+F opens the surface). Those notes are the spatial channel from them to you — the thing a screenshot with an arrow drawn on it used to be, except each one carries the region, an image of it, and where resolvable the object it is about.

annotations lists them. Filters: status (open, addressed, any), sinceGeneration (drops everything from an older load), limit. Each entry carries what the surface stored — id, created, generation, scope (region or scene), rect (null for a scene-wide note), note, view — plus status, addressedNote, addressedAt, anchor and crop. The response adds cropPath (absolute, ready to read) when the crop file is there and cropMissing: true when it is not, and reports total, openCount and addressedCount alongside the filtered count. A missing store is an empty list, not an error — it only exists once something has been annotated.

"reproject": true adds a now object per anchored entry: where that anchor is on screen at this moment, as {x, y, rect, via, insideViewport}. via is objectName when the object was found again and asked where it is, world when it is gone and the stored world point was projected through the live camera, stored when there was nothing to project against.

annotate marks one addressed. It takes annotationId (not id — that key is the request’s own correlation id on every action) and a note saying what you actually did; a mark without one is refused, because “addressed” with no explanation leaves the user re-deriving it from a diff. It answers {annotated, annotateStatus: "addressed", addressedNote, addressedAt}. It is the only write an agent gets: nothing here deletes. Re-opening, clearing addressed and wiping the store are the user’s, in their own overlay.

What an anchor is

The rect is pixels, and pixels never fail — but they stop meaning anything the moment the camera moves. The anchor is the other half: what those pixels were about, resolved once when the annotation was made.

"anchor": {"resolved": true, "kind": "2d", "objectName": "player",
           "type": "RectBoxBody", "source": "Sandbox.qml",
           "space": "world", "world": [5, 10], "at": [452, 372],
           "under": "Rectangle"}
  • kind — 2d (an item, resolved through the item tree) or 3d (a node, picked through the View3D).
  • objectName / type / source — enough to go straight to a source line instead of hunting. objectName is absent when the item has none.
  • space / world — read world according to space: world is canvas world units, world3d is Qt Quick 3D scene coordinates, scene is scene pixels (a plain GUI app has no world, and calling pixels “world” would be a lie you could not detect).
  • at — the viewport point the resolution used, i.e. the rect’s centre.
  • under — the innermost thing actually at that point, when the reported anchor is something else. Resolution walks up from an anonymous internal item to the nearest thing a note can be acted on: something declared in a QML file on disk, preferring one with an objectName. under keeps that walk visible rather than silent.

An anchor can be unresolved, and that is a real answer:

"anchor": {"resolved": false, "at": [40, 40],
           "reason": "nothing specific enough under 40,40 (innermost: Item)"}

Empty space, a shader effect, an item that just fills the whole viewport, and instanced geometry (a LineBatch3D or a VoxelMap — Qt Quick 3D cannot pick those; use inspect for them) all resolve to nothing. A wrong anchor is worse than none: it sends you to the wrong file with confidence. When an anchor is unresolved, the crop and the rect are still the whole record, and they are enough.

The crop, and where all this lives

The store is <sandboxDir>/.clay/crew/annotations/index.json, with one PNG per annotation beside it. The crop is taken once, when the annotation is made, and never refreshed — it freezes the evidence, so what you look at is what the user framed rather than a scene that has moved on since.

A rect that hangs partly off-screen yields the part that was visible, flagged with cropClipped: true. A rect entirely outside the viewport yields no crop at all: crop is null and the failure is reported. That is deliberate — a clamped crop of somewhere else would be a picture of the wrong thing — and it costs nothing, because the note and the rect are the record; the crop is evidence about it.

Two writers share the store (the overlay creates entries, the inspector marks them addressed), so every write re-reads the file, patches only its own fields and commits atomically. Editing index.json by hand while a session runs is safe.

eval — Expression evaluation

{
  "action": "eval",
  "eval": ["player.health", "JSON.stringify(canvas.find({type: 'Enemy*'}))"]
}

Evaluates JavaScript/QML expressions in the sandbox root context. This is the bridge to canvas.find() and any other QML function.

tree — Structural dump

{"action": "tree", "maxDepth": 4}
{"action": "tree", "maxDepth": 6, "detail": "full"}

Returns a JSON tree of the QML item hierarchy. Overview mode (default) includes type, objectName, source file, custom properties, complex property names, and visible/enabled state. Full mode adds vector properties, z-order, opacity, clip, state, and childrenRect.

When a child list exceeds 20 items, the tree truncates to the first 5 items plus a summary of all children — type counts and mini-dumps of rare or named items. This surfaces interesting entities (player, enemies) among hundreds of walls.

{"action": "tree", "select": "Rectangle", "limit": 5}
{"action": "tree", "objectName": "hud", "maxDepth": 2}

A full tree costs a few hundred milliseconds and dumps the whole scene, which is too expensive inside a verification loop. With select (type) or objectName you get an items array of just those nodes, at depth 0 unless you ask for more. A selector that matches nothing returns an empty list — never the whole tree.

inspect — Ask the renderer what it actually got

{"action": "inspect", "select": "lines"}
{"action": "inspect", "select": "LabelBatch3D", "objectName": "roadNames"}

Most verification questions are numeric — is that arrowhead full size? did the second line get drawn at all? where is the player, in world units? — and a screenshot answers none of them. inspect walks the scene and returns the data as the renderer received it, after all bindings and JS have run.

Two kinds of answer, and every entry says which it gave via "via":

  • "via": "hook" — the type implements clayInspect() and answers in its own terms.
  • "via": "properties" — no hook, so the answer is read generically: geometry, app-level properties, source file. This only happens when you asked for something specific (select or objectName); with no selector, only hooked objects answer, which keeps the default cheap.

select matches the type name, the class name, any base type (so select: "PhysicsItem" finds every body, and select: "ClayWorld2d" finds a sandbox root that derives from it), or the type a hook reports for itself; lines is shorthand for LineBatch3D. Types shipping a hook today: LineBatch3D (per line: resolved world points, width, colour, styleId, length; plus batch bounds and instance count), LabelBatch3D (text, size, position, colour per label, including curved path labels), VoxelMap (grid, solid count, palette, storage), ClayCanvas (pixelPerUnit, the visible world rect, canvas size), ClayWorld2d (world bounds, entity count, registered components, gravity), ClayWorld2dCamera (mode, target, centre, look-ahead offset) and PhysicsItem (world-unit position and size, body type, velocity, sensor/categories).

The inspector knows none of these types by name — a type opts in by implementing Q_INVOKABLE QVariantMap clayInspect() const in C++ or a plain function clayInspect() in QML. The contract for anyone adding one: pull-only. Read state the renderer already keeps, compute nothing on your own schedule, never maintain bookkeeping for the inspector’s benefit. That is what keeps this free in a shipped app, where nothing ever calls it. A hook belongs to the object that declares it: children of the same file do not answer with it.

tree and inspect overlap but are not the same view: tree gives you the item hierarchy (2D items only — a LineBatch3D is a Model, not an Item), while inspect gives a flat list of matches across the whole scene, 2D and 3D alike. For “what is in my scene”, reach for inspect.

project / pick — Screen space and what’s under a pixel

{"action": "project", "world": [0, 0, 0]}
{"action": "pick", "x": 640, "y": 400}

project maps a world point through the live camera and viewport, answering x, y, depth, behindCamera and insideViewport — a point behind the camera is never reported as inside the viewport, since those are different failures. It removes the need to hand-roll mapFrom3DScene arithmetic in a shell script.

pick answers what is under that pixel: the object hit, distance, world position and normal, plus the colour actually rendered there. The colour works in a 2D-only scene too; the hit needs a View3D. Instanced geometry (a LineBatch3D) is not pickable in Qt Quick 3D — use inspect for those.

Both take an optional "view" (id or objectName) when the scene has several View3Ds.

trace — Temporal observation

Start recording:

{
  "action": "trace",
  "start": true,
  "watch": ["player.xWu", "boss.health", "boss.state"],
  "interval": 200,
  "stopWhen": "boss.health <= 0",
  "timeout": 30000
}

Stop manually:

{"action": "trace", "stop": true}

While running, the inspector evaluates the watched expressions at the given interval and writes samples to .clay/inspect/trace.jsonl. The first line is a meta record carrying the wall-clock start (epochMs, also present in the start response); each sample’s t is milliseconds relative to it, so the absolute time of a sample is epochMs + t — this is what makes traces from multiple instances correlatable:

{"meta":"trace_start","epochMs":1789450123456,"interval":200,"watch":["player.xWu","boss.health","boss.state"]}
{"t":0,"player.xWu":44.8,"boss.health":500,"boss.state":"idle"}
{"t":200,"player.xWu":45.1,"boss.health":500,"boss.state":"aggro"}
{"t":400,"player.xWu":45.5,"boss.health":480,"boss.state":"attacking"}

The trace stops when:

  • The stopWhen condition evaluates to true
  • The timeout is exceeded
  • A manual stop request is sent

The response includes a summary — often sufficient without reading the full trace:

{
  "traceStatus": "stopped",
  "stoppedBy": "condition",
  "samples": 42,
  "duration": 8400,
  "file": ".clay/inspect/trace.jsonl",
  "summary": {
    "boss.health": {"first": 500, "last": 0, "min": 0, "max": 500, "changes": 15},
    "boss.state": {"values": ["idle", "aggro", "attacking"], "changes": 8}
  }
}

reload — reload the sandbox, optionally into a scenario

{"action": "reload"}
{"action": "reload", "scenario": "boss-fight", "rearm": true}

Runs the same path as a file-watch reload (full engine recreation — no scene state survives). The response acknowledges with reloadStatus: "requested"; the reload itself has not happened yet, so follow with waitForRoot and confirm status.generation advanced. With scenario, the named checkpoint is applied via the root’s applyScenario() once the new root is ready. With rearm: true, the scenario is re-applied after every subsequent reload (including file-watch reloads while you edit) until a reload request passes rearm: false. The active rearm is visible as rearmedScenario in state.json; each apply is recorded as a scenario_applied event in events.jsonl.

waitForRoot — block until a load resolves

{"action": "waitForRoot", "timeoutMs": 5000}

Blocks (default 3000 ms) until the pending load succeeds or fails; returns immediately when the phase is already terminal. The response carries phase, ready, waited, and diagnostics. Use this instead of sleep/poll loops after any reload.

time — pause, single-step, timescale

{"action": "time", "paused": true}
{"action": "time", "step": 30}
{"action": "time", "scale": 0.1}

Drives Clayground.paused / Clayground.timeScale. ClayWorld2d binds them automatically: pause halts the physics world (restoring the user’s running value when lifted), scale slows or speeds the simulation, and step advances a paused world by exactly N fixed 1/60 s physics steps (deterministic — step implies pause). Worlds acknowledge steps; the response’s stepped count is 0 with a clean error when no world consumed the request (plain QML app). Non-world apps can opt in by binding their own timers/animations to the two Clayground properties.

input — synthesized player input

{"action": "input", "gamepad": {"axisX": 1.0, "buttonA": true, "durationMs": 600}}
{"action": "input", "key": {"key": "Right"}}
{"action": "input", "click": {"xWu": 12, "yWu": 10}}
{"action": "input", "click": {"objectName": "startButton"}}
{"action": "input", "click": {"x": 240, "y": 130, "button": "right"}}

Three channels, combinable in one request:

  • gamepad — feeds a synthetic source in every GameController (imperative writes, exactly like keyboard input, so human and agent input coexist). durationMs auto-resets to neutral — the “hold right for 600 ms” primitive.
  • key — synthesizes a key press/release on the window (press/release booleans, default both).
  • click — addressed by window pixels (x/y, works in any QML app), world units (xWu/yWu, canvas apps — resolved via canvas.worldToScene(), clean error otherwise), or objectName (resolves the item’s center; nicest for Controls-style UIs).

batch — several steps in one round trip

{"action": "batch", "steps": [
  {"action": "input", "key": {"key": "V"}},
  {"action": "snapshot", "screenshot": true},
  {"action": "input", "key": {"key": "V"}},
  {"action": "snapshot", "screenshot": true}
]}

A step is a request — same keys, same handler, nothing new to learn. Anything that works standalone works as a step, with one difference: a step must spell out its action. There is no snapshot default inside a batch, so a step written as {"input": {...}} fails loudly instead of quietly taking a picture.

Returns steps — one result per step, in the same order, each carrying the action it ran and the generation it ran at — plus stepsRun and stepsTotal. The status envelope is attached once, for the whole batch; the per-step generation is what makes a reload landing mid-batch visible instead of silently changing what the later steps measured.

A batch stops at the first failing step. The response then carries failedStep (the index) and an error reading batch: step 1 failed: …, and steps ends there — nothing after it ran. Continuing past a failure would be worse than not batching at all.

Batches do not nest, and 32 steps is the ceiling: a batch runs synchronously inside one file-watch callback. A trace step may start or stop a trace, but a batch never waits for one — a trace answers later, from its own timer, as its own response. Trace sampling is suspended for the duration of a batch so that completion cannot land in the middle and overwrite the batch’s own reply.

Scenario Checkpoints

The sandbox root may define named checkpoints so reloads land directly in the situation under test instead of at the title screen:

function scenarios() { return ["start", "boss-fight"] }

function applyScenario(name) {
    if (name === "boss-fight") {
        player.health = 100;
        placeEntities(bossRoomLayout);
    }
}

snapshot responses list scenarios when defined; reload applies and optionally rearms them (see above). Assign entity positions imperatively in applyScenario — initial QML property values do not fire change handlers, so PhysicsItem coordinates only sync on post-creation writes.

View State Preservation

The sandbox root may define viewState() (returns a JSON-serializable object — the user’s place: camera, tuning params, sim time) and applyViewState(state) (assigns it back). Both are optional siblings of flagInfo() / scenarios():

function viewState() { return { camX: cam.x, camY: cam.y, zoom: cam.zoom } }

function applyViewState(s) { cam.x = s.camX; cam.y = s.camY; cam.zoom = s.zoom }

When present, the loader captures viewState() from the outgoing root right before every reload — file-watch and agent-requested alike — and re-applies it to the new root once it is ready, after any rearmed scenario (so a sandbox that encodes scenario + sim time in its view state gets the last word). A null or failed capture keeps the previous one, so a fix applied after a load error still restores the user’s place. Each step is recorded in events.jsonl as a view_state_captured and a view_state_restored ({"ok": bool}) event.

snapshot responses carry a viewState key when the root implements it, and flag bundles (Ctrl+F flags and auto-flags) include the same key. The effect: the user keeps their camera, zoom, and place while an agent edits and reloads files underneath them.

The 2D canvas provides a find() function for spatial and conditional entity search. Call it via eval:

{
  "action": "eval",
  "eval": ["JSON.stringify(canvas.find({type: 'Enemy*', near: {objectName: 'player', radius: 10}, props: ['health', 'state']}))"]
}

Filters (all optional, combined with AND):

Filter Description
type Class name pattern (* wildcard)
objectName ObjectName pattern
near Spatial filter: {objectName: "player", radius: 10} or {x: 30, y: 40, radius: 15}
where JS expression evaluated per candidate, e.g. "health < 50"
props Array of property names to include in results
limit Max results (default 50)

Distance is always in world units — the canvas owns the coordinate system.

The flagInfo() Convention

The sandbox root item may optionally define a flagInfo() function that returns domain-specific state:

function flagInfo() {
    return {
        player: { x: player.xWu, y: player.yWu, hp: player.health },
        enemyCount: enemyRepeater.count,
        currentRoom: roomManager.activeRoom
    }
}

If present, the inspector calls it at snapshot and flag time. If absent, snapshots are still complete — just without the custom context.

Keyboard Shortcuts

Shortcut Action
Ctrl+F Flag a moment — captures screenshot, lets you type an annotation, saves to .clay/crew/
Ctrl+T Toggle trace recording — starts or stops the currently configured trace

Ctrl+F — Flag a Moment

Press Ctrl+F to capture the current moment:

  1. A screenshot is taken and displayed as a frozen overlay
  2. Type an annotation describing what you see (Shift+Return for newlines)
  3. Return confirms, Escape cancels

The flag JSON contains the annotation, screenshot path, root properties, flagInfo, an overview tree dump (depth 4), log tail, warnings, and errors. Max 5 flags are retained.

Ctrl+T — Toggle Trace

Starts or stops the currently configured trace. The agent configures what to watch via the file protocol; the human controls when to record.

Offscreen Mode

When no display is available (Docker, CI), the Dojo runs headlessly:

QT_QPA_PLATFORM=offscreen ./build/bin/claydojo --sbx Sandbox.qml

The inspector works identically in offscreen mode. Screenshots still capture the rendered scene via Qt’s offscreen framebuffer.

Next Steps