Historical evidence

The focused paper contains three completed studies: naming confirmation, counterfactual rewiring, and locality confirmation. Their 496 graph worlds produced 6,560 attempted requests and 6,559 usable responses. These counts are not pooled into a single statistical test.

Download the complete evidence package.

The package contains:

  • Frozen request plans, local freeze manifests, and protocol text for all three main studies and four supporting cohorts.
  • Exported literal request and response objects, including the retained failed attempt.
  • Expected analysis outputs and the historical Python dependencies required to reproduce them.
  • Five request/response pairs for the illustrative edited-graph example.
  • A file manifest, source hashes, and an offline verification command.
  • Independent literal, statistical, supporting-count, and estimator-identity audits, labeled as post hoc review.

The archive size and SHA-256 appear in the download summary; the complete manifest specifies every file and export transformation. Follow the reproduction instructions to recompute the results.

Publication figures and tables

The following figures are generated from the saved analysis outputs. SVG versions preserve text and are suitable for publication; PNG versions provide a simple raster alternative.

Figure Vector Raster
Historical edit and regrouping example SVG PNG
Naming effects beyond serving repeats SVG PNG
Per-forest regrouping effects SVG PNG
Path position and compactness SVG PNG

The publication claim ledger maps the principal numerical claims to exact source values and files. It is a compact navigation aid, not a substitute for the full analyses. The preparation verification record reports whether the exported analyses reproduce the historical result sections.

Supporting campaign reports

The paper cites several supporting findings outside its three principal cohorts. The following static supplements preserve the relevant methods and qualifications:

The four supporting cohorts contribute 576, 1,824, 144, and 384 attempts for header position, confidence, the initial application, and policy, respectively. They are not included in the 6,560-request main-study count. Their literal records, plans, and analyses are now in the download, for 9,488 packaged attempts and two retained failures. The confidence threshold analysis excludes 96 serving-repeat calls and uses 1,728 planned root cases. These datasets retain separate sampling units and estimators; the download does not claim to reproduce every experiment in the larger program.

Prepared follow-up materials

The study preparation page links the new protocols, generators, audits, software-test fixtures, and analysis procedures. Those materials are prospective. Synthetic fixtures exercise software and are explicitly excluded from empirical evidence.

The completed historical evidence and the prepared future studies have separate manifests. Preparing a request file does not imply that it was submitted, charged, or answered by the model.

The prepared-study archive preserves the original frozen files byte-for-byte. For the separately served renaming HTML overview, static styles are extracted to a local stylesheet to meet the publication site's HTML rules. The study manifest records both original archive hashes and served-file hashes. This presentation change does not alter protocols, requests, answers, or analysis code.

Evidence boundaries

This evidence accompanies the published report by OpenProse Research. The publication record states authorship, internal review, rights, and correspondence details. No external peer review is implied. The evidence supports the stated behavioral effects under the tested conditions; it does not reveal model weights, internal attention, parameter count, or training procedure.

The selected historical evidence is pinned to commit 149932ddd4e09fcdfcb7833c6dea7955fa5c5985. This identifier locates the research snapshot; it is not an assertion that the private working repository itself is a public download.