Blog · 13 July 2026· updated 19 July 2026

Part of: Ground · Pono· Status: Active development· For: Teams running coding agents

We charged our own tool rent. The bill was not what we expected.

We promised to publish the verdict either way. It was neither the clean win nor the clean loss we braced for — and the honest answer is more useful than both.

Last week we told you, in writing and in advance, how our own tool could fail, and promised to report the result either way. Here it is — and the honest version is more interesting, and more encouraging, than our first read of it.

Pono (Māori: truth) is a deterministic, queryable projection of a codebase’s history. We registered it in our real working sessions with one rule: no nudges. Nobody is told to use it. The failure line was pinned before the first session: near-zero use after three working days is a negative signal — redesign, don’t explain it away.

Three days in, the first number was near zero. And we almost stopped there and published exactly that: our tool didn’t earn its place. We’re glad we didn’t, because it would have been the wrong verdict — and the reasons it’s wrong are the actual result.

We had measured one narrow slice

Near-zero use turned out not to mean near-zero value. In that first setup the tool was never really being looked at — an agent deep in exactly the work the tool exists for was reconstructing the answer from what was already under its hand, and never turning to consider a specialist one it hadn’t reached for before. That is a fact about how a tool gets *surfaced*, not about whether the ground beneath it is worth standing on.

So we did the boring, honest thing: we changed how the tool was presented instead of explaining the number away, and we ran it again against a different agent. This time it wasn’t ignored and it didn’t need coaxing — it was reached for straight away, unprompted, and put to work. Availability isn’t adoption; a tool has to be met where the work is happening. That part is a solved problem.

And then it carried real weight

Here is the finding that mattered — the one our first draft got wrong. Once an agent was actually using Pono, it didn’t climb back down to git for the hard questions. It traced the why of a dozen different files through the ground, and built its entire assessment on that: which boundaries were real, which were legacy scaffolding, which decisions were already settled and shouldn’t be reopened — each one tied back to the actual commit that set it.

It got the deep architecture right, from the recorded history rather than from a guess, and it explicitly refused to reopen decisions the codebase had already made. That is the whole thesis in one session:

Grounded truth earns its keep when it out-answers the tools the agent already trusts — and, given a fair chance to be used, it did.

Our first read made the tool look worse than it is. The honest position is the more hopeful one: where the ground is both surfaced and leaned on, an agent will stand on it instead of re-deriving the past by hand — and it will keep settled questions settled, which is most of what “don’t hallucinate about this codebase” actually means.

Kano was the first world. The codebase is the second. The thesis still says any causally-recorded system can be next — and the work now is to make sure that wherever an agent stands on this ground, it never needs to climb back off it. If you’re building agents that need somewhere real to stand, we’d like to hear from you.

Update — July 19: we ran the controlled version, and we’re publishing the score

Everything above measured whether agents reach for ground. The harder question — the one this bet actually rides on — is whether ground prevents wrong statements. So we pre-registered a protocol before collecting any data: frozen questions per repository, written before anyone saw the incoming changes; identical twin sessions on the same commit — same model, same tools, same prompt — where one twin receives a small derived summary of what changed since the machine last looked, and the other doesn’t; and a failure condition we committed to publishing if it went against us.

Three qualifying pairs ran across three of our own repositories. The score:

The ungrounded sessions made two confident false claims about our own codebases. One declared a crate “load-bearing” that nothing imports. One declared an open architectural decision “resolved and binding,” attributing it to the wrong document entirely. Both were fluent, specific, and wrong — exactly the kind of claim you act on.

The grounded sessions made one false claim — a provenance detail pinned to a neighbouring commit. We’re reporting it at the same volume. An instrument that only finds what you hoped for isn’t an instrument.

Grounded sessions did far less rework where it mattered. Facing a 38-commit backlog of unseen changes, the ungrounded twin needed more than three times the tool calls to reach the same factual answers. Facing a 5-commit backlog, the twins were at parity. The value concentrates exactly where drift risk lives: large, unobserved change.

The most telling split wasn’t factual but behavioural. On questions whose answers live outside the repository — intent, decisions, “who chose this” — the ungrounded sessions force-fit the nearest plausible artifact and answered anyway. The grounded sessions checked, or said “confirm before acting.” Orientation didn’t just supply facts; it changed whether the agent verified before asserting.

Three pairs, one machine, one model. We’re not claiming a law. We’re claiming our own pre-registered bar was met: when real change had landed, grounded sessions made fewer false statements about the world than their ungrounded twins — and the same net caught our tool’s own miss. The protocol, runsheets, and raw scores live in the research record for anyone we work with to audit.

— The Taniwha team

Building agents that need somewhere real to stand?

See what ships today, or tell us what you're building.