Status: executable offline preparation, not an executed experiment or authorization to call an API. No responses have been collected for this study. The generator, audit, analyzer and synthetic tests accompany a complete prospective request inventory. A separate independent source review and execution/budget authorization are required before a live study. The prepared manifest is a provenance snapshot, not a claim of external preregistration.

Question and scope

With the current graph facts and the question fixed, does placing a truthful root declaration before the queried node's record block increase selection of that root even when the explicit parent links lead elsewhere? Compare a cue for the current root, the old construction root, a third incorrect root, and a neutral root-declaration preamble. The third root separates an old-root-specific explanation from attraction to another incorrect marker. The old graph is never shown to the API: this study does not test memory of an earlier response or a temporal update within a conversation.

This refines the provisional publication plan before outcomes. Regrouping entire subtrees would change target placement and many relevant-edge distances at once. Here the three cued layouts move only four root declarations; every nonroot record retains its exact absolute list position. Both incorrect-cue conditions also hold the true-root declaration at exactly the same off-target position and byte distance from the target. Neutral moves the same root declarations into a preamble, preserving all nonroot relative order. This reduces some confounds without pretending to eliminate placement and distance differences. “Current” means target-valid marker, not globally correct component grouping for every node.

Fresh worlds and literal facts

The fixed seed is 2026091902. Generate 256 independent forest worlds, 64 in each of four specified topology families: bushy spine, random recursive tree, binary side tree, and segmented branches. Each world has four roots and 64 nodes per root basin: 256 records. A shared eight-edge spine supplies the focal query in all families. Family variation primarily changes surrounding structure; this is not a general distribution over all graph tasks or depths. Each world samples a fresh naming/order realization. The four root components within one world are structural copies and are not independent samples.

Starting from four isomorphic components, replace the parent of the depth-four spine node in each component with the depth-three node in the next component cyclically. These four explicit edits preserve acyclicity, node/edge count, basin size and candidate-root isomorphism. The target is the depth-eight descendant in the first component. Its current root, prior construction root, third root and fourth root are all different. The prior graph and edit list appear only in the offline world inventory, never in request state, instructions or criteria.

Opaque identifiers use a new lc_ namespace and fixed length. Within each family, the 64 worlds cross all combinations of four correct root criterion positions, four correct-root lexical ranks and four target-block positions. Root criteria are identical across layouts and copies within a world. The direct-parent control uses a separate four-way criterion order, balanced over correct positions. Relative role-to-key offsets remain a fixed cyclic schedule; marginal key balance is not a claim of exhaustive counterbalancing of every permutation.

Four layouts

Each request contains native JSON state with the same semantics string and the same list of {id, parent} records. A null parent denotes a root. The semantics and root question explicitly state that only parent fields determine relationships and that order/proximity do not assert a relationship.

Nonroot records are arranged as four blocks corresponding to their pre-edit component, using a seeded internal record permutation. In a cued layout, one root declaration precedes each 63-record block:

Condition Root declaration before the target's block Literal current answer
Current cue Current root Current root
Old cue Prior construction root Current root
Third cue A different incorrect root Current root
Neutral All four root declarations in a preamble; no interleaved root markers Current root

The three cued conditions permute only those four declarations. Old and third cues use the same true-root marker slot; they differ by swapping the old and third declarations, so their full root-path positions also match. The current cue swaps the true root into the target block. The remaining root-marker geometry is measured rather than asserted fully matched. The neutral preamble retains the same nonroot sequence. No root assertion, grouping label, separator, stale parent value or extra semantic hint is added. Every condition has the same factual record multiset, question bytes, criterion bytes, correct answers, total compact JSON byte count and identifier frequencies.

The current-cue condition is assisted for the focal root query: the prespecified preceding-root heuristic is an oracle there by construction. It is not offered as evidence of graph traversal. The same heuristic deliberately predicts an incorrect root in both conflict conditions. Neutral removes these interleaved markers but cannot remove every possible sequence cue or structural regularity.

Questions, copies and schedule

Every request asks two independent Choice questions: the target's root and the edited depth-four node's explicit immediate parent. Both question texts identify their nodes without relying on the question IDs. The companion question may make the edited record salient; this is fixed in every condition and defines the scope of the result. No sibling answer is fed into another question.

Four layouts × two exact request copies × 256 worlds yield 2,048 calls, 1,024 distinct request bodies and 4,096 question answers. Copies are a serving reference, not additional graph worlds. Complete eight-call world blocks are randomized; adjacent pairs of world blocks use reversed role schedules. Every layout/copy role has mean within-block position 3.5, each named copy runs first in 128 worlds, and exact copies have at least three start intervals (two intervening starts) between them. Whole paired blocks are shuffled prospectively. Actual times and any recovery phase must be retained by a future central runner; no execution code is included here.

Primary endpoints and uncertainty

For each world, average the two named copies within each condition. Define:

  1. Old-cue preference: selected-old rate in the old-cue layout minus selected-old rate in neutral.
  2. Third-cue preference: selected-third rate in the third-cue layout minus selected-third rate in neutral.

Each endpoint requires all four named selected-label observations in its two conditions. The estimator is the equal-weight average of the four fixed-family means among endpoint-complete worlds. Report complete and excluded world identities and per-family counts. With no missingness this is the ordinary mean over 256 worlds. If a family has no complete worlds, the endpoint is undefined rather than silently changing the family mixture.

Use 20,000 bootstrap draws resampling whole forests within each fixed family, seed 2026091903 for old and seed 2026091904 for third. Report two-sided nonsimultaneous 95% percentile intervals. Use 20,000 whole-forest sign-flip draws and (exceedances + 1)/(draws + 1) two-sided p-values under a symmetric/exchangeable null effect assumption. Apply Holm across the fixed two-test family, retaining that family size if a comparison is undefined. These assumptions concern independent generated worlds within fixed families; they do not create a population of independent language domains or remove service-time dependence.

Report all four selected-root role distributions, accuracy and direct-parent correctness by family/layout. Retain raw full-vector coordinates, displayed mass, every two-copy cell and exact-repeat disagreement. Probability-map missingness does not erase an otherwise valid selected label. The old-minus-third difference is descriptive and paired on worlds with both primary endpoints; it receives no new p-value. Condition marginal summaries can have different complete-world sets and must not be subtracted as if paired.

The prospective smallest effect of engineering interest is 10 percentage points of increased incorrect-cue selection, used for planning rather than an equivalence margin or a rule to discard smaller effects. With 256 complete independent worlds and per-world difference standard deviation 0.5, a normal approximation gives an individual 95% half-width of 6.13 points and approximately 80% power near 9.6 points at the conservative two-sided 0.025 level. At the mathematical SD upper bound 1, the half-width is 12.25 points; the count is not a precision guarantee. A distribution-free two-endpoint bound from bounded differences is much wider, approximately 18.50 points for 95% simultaneous coverage. Missingness, serving dependence and family heterogeneity can worsen precision. No sample-size increase will be chosen from observed effects.

Missingness and integrity

Analysis requires all 2,048 planned attempts exactly once and in the frozen dispatch order, including failures. Unplanned/duplicate jobs, altered ordered bodies or metadata, mismatched request digests/byte counts, or mixed synthetic/live evidence are errors. A successful HTTP status alone is insufficient: require the central validated_success flag, no error, matching response model and expected answer IDs. Failed calls remain audited missing observations and are not repaired with repeats or alternative layouts.

Selected-label validity and complete probability vectors are checked separately. No missing coordinate is zero-filled; maps are never renormalized. Every per-cell metric requires both planned copies for that field. Primary paired endpoint exclusions follow this rule independently. The output also reports worst-case full-cohort effect bounds by allowing each missing world difference to range from −1 to 1; this is not an imputation into the point estimate. No confidence threshold or response-quality selection is fitted.

Literal audit and surviving alternatives

The separate audit parses the literal record list, checks unique IDs, valid parent references and acyclicity, then reconstructs roots by breadth-first traversal from null-parent records. It parses target and candidate IDs from the question text, independently verifies both oracles for every request, checks all four before/after edits, and confirms equal basin sizes and isomorphic rooted component shapes. Ordered-byte and multiset comparisons check the intervention itself.

Prespecified baselines are preceding-root block, nearest root declaration, nearest root-ID mention, first/last root declaration, first/last candidate and minimum lexical root ID. Every prediction and current/old/third/fourth count is retained before model outcomes. Correct-cue success is labeled assisted rather than used to claim that these heuristics are invalid.

The audit measures target and edited-record positions, all root declaration positions, complete-path and nonroot-path spans, adjacent relevant-record distances, root-to-target row and compact-wire-byte distances, state/request sizes and their maxima. In the three cued layouts nonroot positions and nonroot-path distances are exact matches. The true-root declaration has an identical position in both wrong-cue conditions but changes position in the assisted current-cue condition. Neutral changes root placement and some absolute nonroot indices. Those are intervention components and possible mechanisms, not nuisance factors proven absent. Equal byte length does not guarantee equal tokenization, billing or internal preparation. Four root components are copied within each world; their shape equality controls simple root-size/shape heuristics but limits structural diversity.

A positive wrong-cue contrast identifies dependence of the complete service output on this layout intervention. It does not identify attention, traversal, hidden heads, training, caching, or a unique internal architecture. A null contrast is not proof of invariance.

API documentation and execution boundary

The raw TypeSafe skill was read from its official repository. Current live access to the official documentation index/API/Choice pages was attempted but unavailable in the browsing tool; same-day archived official pages were read instead, with exact hashes recorded in docs-provenance.json. The generated request uses their documented state/model/questions/Choice schema and pins jev-1.13.0. There are two questions and four options each. The full offline audit reports byte counts, not a purported provider token count. Revalidate version availability, rate/context limits, billing and the execution adapter before any future authorized run.

This package contains no HTTP client, credential access, purchases, execution queue, commit or publishing operation. Synthetic results are explicitly marked and must never be presented as Jev measurements.

Before-outcome design refinement

The initial unexecuted preview used cyclic root-declaration permutations. Its literal audit found that the true-root declaration was, on average, 80.63 records from the target under the old cue and 112.86 under the third cue. Version 2 instead holds that declaration fixed between the two wrong-cue conditions. This improves their comparison without using any model outcomes. The initial source, request inventory and explicitly synthetic demonstrations remain under review-history/initial-unexecuted-preview/; they were never run against the API and are not the final prospective plan. The final request inventory is artifacts/.