Experiment Note 001 · HNI-2026-001

Tool scope dominated the measured difference. Removing rules mostly did not.

A 12-run local Hermes factorial separated tool-schema scoping from rules/context removal across three locked task types.

ExploratoryEvidence checks passedNot peer reviewedOne run per cell
Per-task reductions
77.9 / 78.7 / 20.8%

Three task types, one run per cell. The median was 77.916%.

Quality checks
12 / 12

Cells passed locked automated task checks. These were internal checks, not external review.

Rules removal
0.014%

Median reduction from skipping rules alone; one task used more tokens.

What we asked

In bounded local agent tasks, how much of the observed physical-token difference is associated with tool-schema scoping versus skipping rules and context, while model, provider, task inputs, and independent quality checks remain fixed within each task?

Terminology note: In the locked question, “independent quality checks” meant automated checks separated from task execution. No external party performed the checks.

What we changed

Four harness cells crossed a broad versus scoped tool surface with preserved versus skipped rules/context. Model, provider, prompt, workspace, locked inputs, and task-specific validators stayed fixed within each task.

What happened

Tool scoping reduced physical tokens on all three tasks. The reductions were 77.916%, 78.740%, and 20.787%. All 12 outputs passed locked automated task checks. Skipping rules/context alone was approximately neutral at the median and was not uniformly beneficial.

Claim boundary: This is a descriptive three-task mechanism-attribution pilot with one run per cell. Within-cell variance is unobservable, so these data cannot support a confidence interval. This is not a universal 77.9% estimate, not a reproduction of the cited paper’s 38% result, and not evidence that every task should receive the same tool allowlist.

Why the inconvenient cell matters

The asset-manifest task showed a much smaller 20.787% reduction. Keeping that cell visible is part of the result. It suggests the effect depends on task behavior and call shape, and it is one reason the next study needs more tasks and repeated cells.

Registration timing

The underlying factorial design was fixed before the 12 runs. The HNI protocol wrapper and public research-note structure were created afterward, so this work is not described as publicly preregistered.

AI, funding, and conflicts

Self-funded. Inference was available through an existing subscription. No external research grant supported this study.

Hypernovelty Institute develops and uses the Hermes research harness and benefits reputationally from useful results. No Writer, Nous Research, or OpenAI funding supported the study.

AI roles included design assistance, Python implementation, task execution, receipt extraction, task-quality-check implementation, analysis, and drafting. The quality checks were implemented with AI assistance and locked before runs. They are not independent of the AI execution environment and do not constitute external review. Jordan Finneseth approved the research direction and public communication.

Terminology correction

The locked post hoc HNI protocol used the phrases “independently validate” and “independent validator.” A dated amendment clarifies that these were internal locked automated checks. The protocol bytes and numerical results were not changed.

Public discussion

Read the comment on the original paper’s alphaXiv discussion page.

Protocol identity

32b237b6e05c6a597fd2358430684131ee8ca01e0f66a627daa0e477005f3e43