Introducing the entire journey.
Q1How is everyone planning to use it?
"It feels like everyone else is using some kind of advanced skills, while I'm the only one just typing in prompts and going back and forth..."
We will start with what a harness is, and answer with the records left instead of saying 'completed'.
Harness = The layer that fills in the code with four things that weren't in the prompt
STATEStatusNone → Record
VERIFYVerificationNone → Gate
PERMISSIONPermissionNone → Boundary
REPLAYReproductionNone → Playback
sha256:
genesis
→
seq 1fanout_planned
→
seq 2authorized
→
seq 3hash chain
⇢
append
only
Instead of saying "I have completed it," only an append-only record remains, which carries the hash of the previous event. If one part is modified, everything that follows breaks, ensuring the same execution is reproduced exactly as is.
Q2To what extent can we trust and delegate?
"We lack the manpower to sufficiently inspect outputs and a systematic evaluation loop. A highly reliable agent—do I really know it well...?"
We answer with the verification/evaluation loop and who the 'master of judgment' is.
Even for the same "count the records" task, the results can vary like this.
EPISODENATURAL
response: …count is 72…
claim.rows: 72
EPISODEOVERCLAIM
task: …report the rows…
claim.rows: 999
The judgment is made by code, not AI — 1.26 seconds per inspection (takes 15+ seconds if done by AI). The default is to block if the result is ambiguous.
Q3Quality or tokens?
"If I prioritize quality, the token cost becomes too heavy, but if I choose a lower-tier model for token optimization, the test quality degrades… How can I solve this?"
There is only one answer — you can have both if you split them. Turn on the dashboard and see for yourself.
SPLITBreak it down small, let the cheap model handle it
↓ If it fails
ESCALATE Promote only one step at a time
BASELINEInterview → Completion on plans under $200
2.6×
The same 8B model, just by splitting
0.14→0.37
Increase quality ↑ · Decrease tokens ↓ simultaneously
The standard is the reduction rate — reduction = (baseline − spent) ÷ baseline. If there is no measurement axis, it is marked as INSUFFICIENT_DATA instead of PASS, and if the evidence is lost, it is an immediate FAIL.
Q4If the model improves, will the harness be discarded?
"Models improve every time I wake up, so I wonder if building a harness becomes unnecessary as the models get better."
The Bitter Lesson and the Meta Layer provide the answer — what gets swallowed, and what remains.
User-Level Programsseed.A · seed.B · seed.C …
│
Agent OS · Ouroboros Kernelrecursive correction ↻
│
Contract Layerschema · invariants · capabilities
│
Claude CodeCodexGemini CLI…
As the model improves, only the engine at the very bottom gets better. The kernel and contract layers above it — the name of this layer is Agent OS.
Q5How was Ouroboros created?
"The thought process behind creating such a product, the reasons for the design, the process of failing and fixing it, and even the criteria for judgment."
From the physical contract to a bot that writes letters of apology, we answer with code the untold story of reaching number one in the benchmarks.
seed.contract
goal: Build a hermetic CLI, deps_refresh.py …
constraints: std-library only · no network
acceptance_criteria:
verify_command: pytest -q test_dry_run.py && echo OK
output_assertion: OK
phase: blocked · verdict: revise · qa_score: 0.74
WORKING context window
EPISODICevent journal
SEMANTICsource ledger
PROCEDURAL Skill/Setting Evolution
Instead of saying "It's done," the verify_command makes the judgment. Contracts that could not be judged due to empty verification commands were rejected before execution. On top of that, a bot that reflects on itself using four-layer memory manages the open source.