The follow-up program addresses alternatives left unresolved by the completed studies. The materials on this page are prepared protocols and software, not new model results. Request manifests, exact graph answers, and software fixtures can be checked offline. Execution will require a separately recorded budget and current service conditions.
Renaming
Question: Does name sensitivity persist when the literal questions, candidate identifiers, and marginal identifier-role frequencies remain fixed?
The prepared design uses 384 fresh forests across three specified topology families. It compares original names, anchor-fixed unrestricted internal renaming, and anchor-fixed renaming within equal role-frequency groups. Query and candidate identifiers remain fixed. Both renaming treatments change the same eligible internal nodes. Shuffled and grouped layouts, root and ancestor questions, exact repeats, and separate direct-parent controls are included.
The primary endpoints measure constrained-renaming disagreement in excess of matched exact-request disagreement for shuffled root and ancestor questions. All copies and pairings are aggregated within a forest. The protocol specifies missingness, multiplicity, effect-size interpretation, and precision assumptions before model outcomes are collected.
Read the renaming protocol and obtain its executable materials from the study preparation download.
Layout conflict
Question: With factual truth fixed, do decisions move toward an independently manipulated incorrect layout cue?
The prepared design uses 256 fresh forests. It contrasts layouts associated with the current root, the old root, a third incorrect root, and a neutral arrangement. The factual record set stays unchanged. In the three cued conditions, non-root records remain at the same positions while truthful root declarations change position. The neutral condition changes root-to-target distance; that difference is measured and retained as a limitation.
The main comparisons measure old-root and third-root selection relative to neutral layout. Direct-parent questions, repeated requests, exact graph answers, and simple positional heuristics provide controls. The current-root layout is explicitly classified as assisted because the grouping exposes membership.
Read the layout-conflict protocol and obtain its executable materials from the study preparation download.
Application evaluation
Question: Does a reproducible evidence-construction method improve useful semantic decisions on independent documents?
The application package separates exact source resolution from semantic policy interpretation. It prepares full-context, canonical-order, deterministic-retrieval, and code-resolved conditions, with source correctness, semantic correctness, joint correctness, abstention, cost, and timing recorded separately.
Development fixtures support software checks and protocol refinement. They are not a completed held-out application benchmark. Independent document families, reviewed semantic annotations, and comparator-model measurements remain to be collected. Rewordings from one template do not become independent semantic cases.
Read the application protocol and obtain the schema, annotation guide, fixtures, and adapter materials from the study preparation download.
Analysis and execution readiness
Each study package records its actual generated request count, tests, audit results, and remaining limitations. The package manifest hashes those files. Request counts can differ from the original planning estimate when separate controls are added; the generated manifest is authoritative for execution accounting.
Before live execution, preserve the reviewed protocol and complete manifest, record the selected model and provider settings, estimate cost from the actual request inventory, and set the run's explicit budget. Keep failures, untouched suffixes, and any retries distinguishable. An outcome-dependent change starts a new protocol version rather than silently replacing the original test.
The preparation process makes no paid model calls. Analyzer fixtures are synthetic and remain labeled as such. A fixture that produces a numerical estimate demonstrates software behavior only.
How results will change the publication
| Future observation | Appropriate interpretation |
|---|---|
| Constrained naming sensitivity remains substantial | Query-name and marginal-frequency changes are insufficient explanations under the new controls. |
| The constrained effect becomes much smaller | The initial interpretation should narrow; the removed factors explain part of the earlier instability. |
| Incorrect layout cues attract answers | The service has an observable dependence on the manipulated layout; this alone does not identify an internal algorithm. |
| Preprocessing improves held-out semantic decisions | Report a bounded recommendation for the tested workload, costs, and error policy. |
| Preprocessing only reduces context cost | Report the cost benefit without claiming an accuracy improvement. |
| Conventional models show similar effects | Present a shared evaluation problem with comparative magnitudes, rather than a Jev-specific mechanism. |