LOCAL ONLY ORIGINAL IMMUTABLE ACTION SINK ENFORCED REFERENCE ONLY — NOT DEPLOYMENT-READY NO CLOUD ACTUATION NO TRAINING
Why these boundaries matter

This browser never actuates hardware. It coordinates the part we can help with digitally: contract mismatches, suspicious proposal behavior, documented fixes, and unresolved evidence. The downloaded package can support a separate local test only after an experienced owner deliberately completes the embodiment-specific gates.

STOP TEST

ROOM 3 · DEBUG THE POLICY

Catch red flags before the robot team stages a live launch.

Find controller, normalization, timing, action, and evidence bugs. Return a reviewed launch pack to the roboticist. A flag is a clue—not proof.

UNCERTAINTY NO EVIDENCE YET Counts stay blank until a local audit or comparison.

Start here · known local fixture

Triage only. No policy code or hardware command runs.
The learning analytics loop
  1. 1
    Load evidencepolicy + episode + controller
  2. 2
    Work the inboxkeep uncertainty visible
  3. 3
    Compare same statesdocument → harmonize → try again
What this toy can and cannot tell youOpen the exact boundary
The loop we are trying to make less painful: turn on a robot, learn through a game, train or finetune a policy, find a broken one, use this toy to see what looks wrong, try a documented fix, check whether it still jitters on the same evidence, combine/harmonize the useful episodes, and try again. What this toy can do today: play and record local episodes, make heterogeneous episodes easier to combine, and compare base and candidate policy outputs without forwarding their actions. Training, changing the weights, closed-loop hardware testing, and deciding whether to run the robot still happen outside the toy. What this does: finds digitally observable mismatches and suspicious policy-output behavior that may contribute to instability, and preserves exactly how each item was handled. What it does not do: predict every fall, repair a controller, retrain the weights, reproduce closed-loop contact and balance, or decide that hardware is safe to run.

play the evidence · fix the flags

walk the policy trail before the next launch.

Each pennant is a real candidate destabilization finding from the current audit. Open it, inspect what was expected versus detected, and add a review plan to the receipt.

review plans 0 / 0 load an audit to reveal the trail
candidate run pathevery flag is a clue—not proof
audit order → owner review → discard-only replay
no trail yet start the diagnostic example or run a contract audit. The real audit findings will become playable pennants.

These markers come from contract rules or recorded trace evidence. They do not locate every physical failure and they never approve hardware.

OPEN RESEARCH PREVIEW

A toy for learning—not a safety device.

Play with it. Inspect it. Improve it. Check every result before hardware.

Research boundary + technical composition

A local sandbox meant to invite experienced roboticists to play, inspect, critique, and extend the episode-harmonization and pre-deployment stability workflow. It is intentionally presented as an unfinished research toy, not a finished product or authority. Its flags and suggestions are informational hypotheses—not a validated safety device, manufacturer guidance, or permission to deploy. Use it to prepare better questions before testing; independently verify every decision and keep the robot’s normal physical safeguards and on-site supervision.

Technical composition: MuJoCo advances the local task dynamics. Deterministic seeded scene generation adapts ideas from an archived WebGL procedural-universe engine. Live captions combine simulator truth with robust temporal motion evidence. Jiminy follows a GODEL-inspired recent context + grounded task knowledge → response contract, but no GODEL checkpoint, GPT model, hosted LLM, or remote inference service is running.

ROBOT RULES

Know which robot rules are real.

G1 has detailed rules. Ackermann has external rules. Other robots stay generic until their exact contract exists.

Open exact embodiment + policy scope

“Broken” here means a candidate that may be unstable, mismatched, or incompletely declared. The browser accepts a local data-only policy package for audit; real inference stays in the owner’s local policy environment through the action-sink bridge.

REFERENCE IMPLEMENTATION

Unitree G1 humanoid

Informed by the reported workflow: a controller first trained with RL, then adapted with offline imitation data and KL regularization; a π0.5/VLA-style checkpoint being integrated with a SONIC/whole-body stack; and demonstrations from a Chinese dataset. The current demo audits that VLA package context. Native multimodal VLA replay is not implemented: images, task text, temporal history, preprocessing, and action chunks do not yet pass through this state-vector/action-vector bridge.

Screen for 29D joint order · .q vs torque/target semantics · normalization · rate/latency · high/low-level command ownership · action drift/jerk · fall-risk proxies
EXTERNAL DEDICATED RULE FAMILY

Ackermann vehicle

Ackermann steering has dedicated controller-semantic checks, but Ackermann is not a registered robot in this LeRobot checkout. It is useful for candidate driving policies only after the exact external vehicle/controller contract is supplied. A matching output width does not establish that steering commands mean the same thing.

Screen for steering angle vs curvature or yaw rate · velocity mode · wheelbase · limits · command frequency · stale observations · stop/watchdog ownership
ADAPTER-REQUIRED FAMILY

Other arms and manipulators

Useful for base-versus-finetuned arm policies when the package declares the embodiment and an experienced owner provides the native controller adapter and safety envelope.

Screen for joint position/velocity/torque vs Cartesian delta · reference frame · joint and gripper channel order · scaling · contact-sensitive discontinuity · controller limits
Policy formats this path is meant to compare: base and finetuned flat-vector checkpoints and RL → offline-imitation adaptations. VLA package/schema audit is useful now; native multimodal VLA replay is not implemented. Execution boundary: static manifests and safe tensor metadata may be uploaded locally; arbitrary policy code is not executed by the web server. The owner runs it locally and sends only proposed actions into the discard-only trace.
LEROBOT-SHAPED CORE The audit is not tied to 29 G1 joints.

The local LeRobot checkout defines a common robot boundary around observation_features, action_features, get_observation, and send_action. Shot‑5.1.1 audits and records that boundary; it does not assume that two equally sized vectors share physical meaning.

ROBOTS ALREADY REPRESENTED LOCALLY One dedicated reference; the remaining registered embodiments are generic-only

G1 is the concrete LeRobot reference. The other registered implementations—SO followers and bi‑SO, Koch, OMX, OpenArm and bi‑OpenArm, ReBot B601 and bi‑ReBot, Reachy2, HopeJr, LeKiwi, and EarthRover Mini Plus—receive generic package/schema and flat-vector shadow analysis only. Each needs its own Shot‑5.1.1 manifest, semantic rule set, adapter, and safety envelope before meaningful native replay or actuation.

POLICY FAMILIES ALREADY REPRESENTED LOCALLY Behavior cloning, RL, diffusion, and VLA candidates

The policy factory includes ACT, Diffusion, VQ‑BeT, TD‑MPC, Gaussian Actor, π0, π0.5, SmolVLA, GR00T, X‑VLA, and related families. The present bridge compares finite flat action vectors produced from state-vector observations. Native multimodal VLA replay is absent; π0.5, SmolVLA, and similar packages remain package/schema-audit only until a dedicated adapter preserves images, task text, preprocessing, temporal history, and action chunks.

ARCHIVED RESEARCH CONTRIBUTION Seeded task coverage + robust temporal evidence

The procedural-world code contributes deterministic seed streams for reproducible scenario variation. The temporal-evidence code contributes robust median/MAD motion scoring, provisional events, confirmation/retraction, and source attribution. Those methods transfer across embodiments; task geometry, simulator truth, goal predicates, and safety limits do not.

Implemented interpretation today: Dedicated audit rules exist for exactly two families: Unitree G1 and Ackermann steering. G1 is the concrete LeRobot reference; Ackermann is an external controller family, not a robot registered in this checkout. No other embodiment-aware rule set is implemented today. All other registered LeRobot embodiments receive generic package/schema and flat-vector shadow analysis only until their own semantic rules, manifest, and adapter are added. Native multimodal VLA replay is not implemented.

POLICYWeights + inference behavior
POLICY CONTRACTInputs, outputs, normalization + timing
CONTROLLER MANIFESTMapping, ownership, limits + safe stop
EPISODE / SHADOW TRACERecorded evidence of what happened or was proposed

ONE BUG · ONE CLEAR PATH

Follow one bug from evidence to handoff.

Check the contract. Replay two policies on the same states. Compare actions. Save the result. Nothing here moves a robot.

V2 · NO NUMBER ALONE

A number needs a baseline and a limit.

These candidate values came from the walkthrough. Base values and robot limits are missing, so the result stays open.

REVIEW REQUIRED
  1. 01Capturesave signal + source
  2. 02Reproducesame states · two policies
  3. 03Use rulefix it or mark unknown
  4. 04Replaysave the outcome
CANDIDATE SIGNALS
drift
0.5814
jerk
2,600.03
saturation
0.00%
latency delta
4.55 ms
WHAT IS MISSING
base value
missing
candidate / base
cannot compute
robot limit set
not supplied

No pass/fail result until the owner checks these values.

DEBUG INBOX
  1. OPENtiming mismatchneeds another same-state run
  2. UNKNOWNaction unitsunits are not defined
  3. RETURNowner handoffsave the call + next evidence

SAFE PRACTICE RUN · NO POLICY CODE EXECUTES HERE

Try the G1 debug example.

Use safe fixture files to see the full loop. Then swap in your policy, recorded run, and controller map. The robot stays still.

  1. 1Policy ZIPreplace fixture weights + policy contract
  2. 2Recorded episodereplace the registered target collection
  3. 3Exact controller manifestreplace every unknown physical declaration
  4. 4Optional shadow traceadd discarded proposed-action evidence
Output Grouped risk inbox + paired trace comparison + downloadable handoff receipt The paired fixture demonstrates comparison only. It is not a policy to deploy.

Step 01 · policy intake and contract audit

Load the policy, run, and controller map.

STOP TEST · AUDIT REQUIRED

room 2 → room 3 handoff · minimum satchel

pack one policy debug kit.

the practice tape and brain candidate stay separate. pair them to run offline checks; neither ingredient silently rewrites the other.

checking satchel…
ingredient 1 · brain candidatechoose a policy zip

use one already in the policy satchel, or add a local zip.

ingredient 2 · practice tapechoose a run package

room 2 should leave aligned observations, actions, labels, frames, and a receipt.

ingredient 3 · body contractmatch the same robot body

embodiment, dimensions, units, frame, mode, and controller timing must agree.

result · offline debug passnot ready yet

pair the ingredients to reveal contract and policy-output red flags.

choose one brain candidate and one practice tape.

a room 2 export can be saved without being policy-input ready. this gate keeps that distinction visible.

brain candidates · practice tapes · controller contract

choose the ingredients for one debug kit.

each policy × tape pair gets its own receipt. checks run one at a time and never blend the original files.

0 comparisons
1 Brain candidate · policy satchel
0 packages
Policies will appear here.
2 Practice tape · run satchel
0 episode targets
Episode collections will appear here.
3 Controller manifest

A manifest gives dimensions physical meaning: embodiment, channel order, mode, units, frame, timing, latency, limits, and stale-command behavior.

Manifest required for meaningful controller replay Generic package/schema audit can still run without one.
Inspect parsed manifest
No manifest loaded.
4 Optional shadow evidence

A shadow trace records observations and proposed commands while actions remain discarded. It is evidence—not a policy or controller contract.

No extra trace attached Selected episode targets still provide recorded evidence.

Select at least one policy and one target. The current cap is 12 local comparisons per run.

Add one policy package to the Satchel

ZIP only. Static files are inventoried and hashed; package code is never executed.

loopback upload
No file leaves this computer.

Focused comparison and replay pair

The dropdowns choose one pair for detailed replay. The checkbox Satchel above audits several combinations without mixing their receipts.

Recorded dataset view 100D raw observation vector
Policy input contract 32 named policy inputs
Policy output contract 29D G1 joint-position targets
29D does not mean the same controller.

A dimension match may be green while action mode remains blocked: the simulation declares 29 joint torques, direct G1 hardware exposes 29 .q positions, and a SONIC whole-body controller requires its own declared interface.

MODE BLOCKER
REAL POLICY WORKFLOW

Import your friend’s actual package and review its actual evidence. The bundled incomplete package is only a diagnostic crash-test sample for demonstrating the inbox—not a policy to hand to a robot owner.

An audit may propose an adapter receipt. It cannot make an undeclared unit, frame, timing, or normalization assumption safe.

CONTRACT FLAG INBOX Unresolved semantics wait here
Offline blockers
Warnings
Checks passed
Hardware blockers
Resolution plans0
EXACT RULE TRACE STATISTIC RECORDED HUMAN Every finding states what kind of evidence produced it.
0 selected
Run a contract audit Timing, units, frame, normalizer, input, and action mismatches will appear as familiar inbox threads.

Step 02 · local policy shadow bridge

Test policy output. Keep the robot still.

ZERO-FORWARD GUARANTEE

The real policy stays on your friend’s machine.

Send recorded states in. Get proposed actions back as evidence. Discard every action. The robot stays still.

Mission: Use every digitally available clue before another physical run. The toy may reduce avoidable stability-testing iterations by surfacing clues related to jitter, falls, or part breakage; it cannot guarantee that the policy is stable or that hardware will not be damaged.

YOUChoose recorded observations
FRIEND’S MACHINERun base + candidate locally
ROBOT STAYS STILLDiscard every proposal
SHARED INBOXReview flags + trace evidence
Forwarded
0
Hardware commands
0
Simulator commands
0
Network egress
0
Policy code in web server
0

Choose the same observation trace for both policies

A fair comparison holds observations, ordering, timestamps, adapter declarations, and frame limits constant.

read-only source
or import a local trace

Nothing leaves this computer. Imported files are treated as data, never code.

Local policy identities · no weights uploaded

Connect the friend’s own local policy environment

The bridge kit speaks a small localhost JSONL protocol. It does not install the policy, resolve its dependencies, import actuator SDKs, or call a hosted model.

create sessions first
  1. 1
    Download the bridge kitOne small client, protocol schema, README, and session receipt template.
  2. 2
    Start it inside the policy environmentGive it a local inference callable; keep controller and actuator processes stopped.
  3. 3
    Run base, then candidateBoth receive the same recorded observations and may only return proposals.
LOCAL COMMAND TEMPLATE
Create a session pair to receive a loopback command.
Download local bridge kit Inspect the kit before running it. The session secret authorizes trace submission only—never actuation.
A
BASE POLICY

Waiting for a session

NOT CREATED
0 / 0 observationsNo local runner connected
Received
0
Discarded
0
Forwarded
0
Hardware
0
Simulator
0
Network
0
Trace receipt
B
CANDIDATE / FINETUNED POLICY

Waiting for a session

NOT CREATED
0 / 0 observationsNo local runner connected
Received
0
Discarded
0
Forwarded
0
Hardware
0
Simulator
0
Network
0
Trace receipt

Choose one observation source, then create two discard-only sessions.

Step 03 · matched-observation comparison

Compare base and candidate. Do not call it stable.

WAITING FOR TWO TRACES

Base ↔ candidate evidence

Flags show mismatches and risky action changes. Clear what you can. The on-site tester makes the final call.

Matched observationswaiting for traces
Mean output driftnot measured
Candidate jerknot measured
Candidate saturationnot measured
Latency deltanot measured
Policy KL divergencenot_computablerequires comparable action distributions—not point proposals
No comparison yet

Finalize both traces. “Not computable” will remain visible wherever the submitted evidence cannot support a metric.

No hardware command is needed—or reachable—to make this comparison.

Supporting check · bundled offline fixture replay

Replay the same evidence.

ACTION SINK ONLY

Offline replay

Read observations, inspect policy outputs, and measure contract-level risk without actuating hardware.

Policy
not selected
Collection
shot1-g1
Resolution plans
0
Execution mode
action_sink_only

Audit a package before replay.

Signals to check before hardware

These signals may show a mismatch or risky policy output. They are clues, not proof of safety.

awaiting replay
Frames inspectedno run
Saturationno run
Command jerkno run
Contract violationsno run
Recorded episode language + replay evidence Parquet language streams · recorded with the episode · not generated by replay
0 frames
Run a replay to inspect event-aligned output behavior.

Step 04 · auditable export and physical return loop

Return the evidence. Keep the decision local.

NOT HARDWARE APPROVED
The robot owner makes the final call.

This tool does not certify safety or approve a physical run. An experienced owner may choose a small, supervised test. Record that choice as the owner’s call—not as a green “safe” result.

  1. TOY VERDICTNOT HARDWARE APPROVED
  2. OWNER CHOICEExperienced onsite tester may independently stage a test at own risk
  3. HANDOFF TRAVELS WITHUnresolved flags · abstentions · changes · rollback · safeguards
  4. WARNING PERSISTSContinuing never turns this green or certified
The browser and server cannot actuate. A downloaded package can reach actuators only after explicit local deployment, an embodiment-specific hardware adapter, deliberate local enablement, verified limits and watchdogs, a physical emergency stop, and an experienced on-site supervisor.

Two downloads with different meanings

The unchanged original and the generated evidence bundle are kept separate by design.

REFERENCE DEFAULT · LOCAL DEPLOYMENT REQUIRES EXPERT ENABLEMENT

Reference handoff package for the friend

A versioned, auditable package for continuing the stability-testing loop locally. It never presents rewritten weights or automatic permission to deploy; dry-run is the default.

  • Unchanged policy copies by hashonly when a Policy Satchel package was explicitly linked
  • Documented overlaysadapter, normalizer, joint order, timing, and controller declarations
  • Review historyall flags, individual or bulk decisions, rules, uncertainty, and abstentions
  • Before / after evidencebase and candidate discard-only traces plus matched-trajectory comparison
  • Rollback manifestsource hashes, versioned artifacts, change log, and parent lineage
  • START HERE overviewpurpose, results, unresolved risks, local gates, supervised test checklist, and return-loop instructions
NOT HARDWARE APPROVED · the domain expert makes and records the final scientific tradeoff.
ORIGINAL PACKAGE Byte-for-byte unchanged

The exact imported ZIP, identified by its SHA‑256 hash. No adapter, declaration, or metric is written into it.

Download original
DOCUMENTED LOCAL HANDOFF START HERE + receipt + evidence + gates + rollback

A companion ZIP containing an overview, hashes, contract diff, adapter and controller overlays, human decisions, explicit abstentions, comparison evidence, local deployment gate, change log, and rollback manifest.

Prepare supervised test-at-your-own-risk handoff
Documented handoff contents START_HERE.md manifest.json comparison.json documented_changes.json flags_and_abstentions.json local_deployment_gate.json

No new checkpoint is produced because Shot‑5.1.1 does not train. The browser bridge records proposals but never forwards them. After download, the experienced owner may install an embodiment-specific local adapter and explicitly enable a guarded actuator path on their own machine. A generated mapping remains an interface transformation—not evidence that the weights are stable or that a run is approved.

BETA · GUARDED HARDWARE RETURN LOOP

Semi-automates the tedious stability-testing cycle around policy finetuning: schema harmonization, contract validation, matched-observation shadow inference, action-drift and anomaly analysis, failure clustering, human review, versioned handoff, and rollback. After a supervised physical run, the robot owner returns the recorded episode and safety trace so measured outcomes can be compared with pre-run predictions. The tool never promotes a checkpoint automatically; the on-site tester reviews the evidence and makes every physical deployment decision.

  1. 1
    Dry-run locallyverify hashes, dimensions, normalizers, timing, and command bounds with hardware disconnected
    first gate
  2. 2
    Shadow commandsread live observations and log proposed actions without passing them to actuators
    local only
  3. 3
    Expert-decided guarded hardware testthe domain expert weighs unresolved risk against scientific value; independent safety controller, physical e-stop, conservative limits, and trained operator remain mandatory
    at own risk
  4. 4
    return the evidenceupload the resulting episode and controller logs to the episode harmonizer for comparison
    close loop
NOT HARDWARE APPROVED—even when the owner chooses to continue.

The downloaded handoff supports a separate, explicit local decision; it does not turn that decision into certification. Offline replay covers recorded states only, and a policy can still diverge under contact, latency, sensor faults, or out-of-distribution states.

open the shot‑5.1.1 episode harmonizer for returned runs →