Three task types, one run per cell. The median was 77.916%.
Tool scope dominated the measured difference. Removing rules mostly did not.
A 12-run local Hermes factorial separated tool-schema scoping from rules/context removal across three locked task types.
Cells passed locked automated task checks. These were internal checks, not external review.
Median reduction from skipping rules alone; one task used more tokens.
What we asked
In bounded local agent tasks, how much of the observed physical-token difference is associated with tool-schema scoping versus skipping rules and context, while model, provider, task inputs, and independent quality checks remain fixed within each task?
Terminology note: In the locked question, “independent quality checks” meant automated checks separated from task execution. No external party performed the checks.
What we changed
Four harness cells crossed a broad versus scoped tool surface with preserved versus skipped rules/context. Model, provider, prompt, workspace, locked inputs, and task-specific validators stayed fixed within each task.
What happened
Tool scoping reduced physical tokens on all three tasks. The reductions were 77.916%, 78.740%, and 20.787%. All 12 outputs passed locked automated task checks. Skipping rules/context alone was approximately neutral at the median and was not uniformly beneficial.
Why the inconvenient cell matters
The asset-manifest task showed a much smaller 20.787% reduction. Keeping that cell visible is part of the result. It suggests the effect depends on task behavior and call shape, and it is one reason the next study needs more tasks and repeated cells.
Registration timing
The underlying factorial design was fixed before the 12 runs. The HNI protocol wrapper and public research-note structure were created afterward, so this work is not described as publicly preregistered.
AI, funding, and conflicts
Self-funded. Inference was available through an existing subscription. No external research grant supported this study.
Hypernovelty Institute develops and uses the Hermes research harness and benefits reputationally from useful results. No Writer, Nous Research, or OpenAI funding supported the study.
AI roles included design assistance, Python implementation, task execution, receipt extraction, task-quality-check implementation, analysis, and drafting. The quality checks were implemented with AI assistance and locked before runs. They are not independent of the AI execution environment and do not constitute external review. Jordan Finneseth approved the research direction and public communication.
Terminology correction
The locked post hoc HNI protocol used the phrases “independently validate” and “independent validator.” A dated amendment clarifies that these were internal locked automated checks. The protocol bytes and numerical results were not changed.
Public discussion
Read the comment on the original paper’s alphaXiv discussion page.
32b237b6e05c6a597fd2358430684131ee8ca01e0f66a627daa0e477005f3e43