science Experiments

The lab notebook.

Question, hypothesis, method, result. Status is honest: "Proposed" means queued, not started — nothing here claims further progress than actually happened.

Proposed

Procedural Graph Runtime

Can a validation-gated, self-evolving execution graph improve long-horizon agent workflows without increasing unsafe tool use?

hub Inspectable Agent Systems Read the Experiment arrow_forward
Proposed

Independent Test Agent

Does separating an independent Test agent from a Repair agent reduce false-confidence patches without hurting resolution rate?

hub Inspectable Agent Systems Read the Experiment arrow_forward
Proposed

Revocable Memory Graph

Can a graph-based retriever enforce memory revocation structurally while still preserving useful personalized context?

hub Inspectable Agent Systems Read the Experiment arrow_forward
Proposed

Agentic Simulation Foundry

Does a generate → retrieve → mutate → execute → score pipeline transfer from autonomous driving to logistics, GIS, or facility-operations scenarios?

hub Spatial & Simulation Systems Read the Experiment arrow_forward
Proposed

Safety Judge Stress Lab

Which content-invariant wrapper transformations produce the highest disagreement across automated safety judges, and can simple ensembles resist them?

hub Inspectable Agent Systems Read the Experiment arrow_forward
Proposed

Context Graph vs. Context Database

Does explicit graph structure improve retrieval traceability and multi-hop reasoning over a context-database approach, at the same token cost?

hub Inspectable Agent Systems Read the Experiment arrow_forward