Primary research
The references below identify established work relevant to representation, context use, and probability interpretation. They are cited to position the case study accurately, not to attribute their findings to Jev.
-
Bahare Fatemi, Jonathan Halcrow, and Bryan Perozzi. 2024. Talk like a Graph: Encoding Graphs for Large Language Models. Proceedings of ICLR 2024. Original preprint: arXiv:2310.04560 (2023). Studies graph encoding, task, and structure as factors in language-model performance.
-
Yuyao Ge et al. 2025. Can Graph Descriptive Order Affect Solving Graph Problems with LLMs?. Proceedings of ACL 2025, pages 6404–6420. Examines graph description order across tasks and models.
-
Hamed Firooz, Maziar Sanjabi, Wenlong Jiang, and Xiaoling Zhai. 2024; revised January 2025. Lost-in-Distance: Impact of Contextual Proximity on LLM Performance in Graph Tasks. arXiv:2410.01985v2; accepted poster at the ICBINB 2025 workshop. This is not an assertion of ICLR main-conference acceptance. Examines distance between relevant graph facts and distinguishes it from absolute contextual position.
-
Nelson F. Liu et al. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, volume 12, pages 157–173. Studies how relevant-information position affects use of long contexts.
-
Zike Yuan et al. August 2026. GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL. arXiv:2608.27142v1, preprint; no accepted venue verified. Its GRIT evaluation crosses four naming schemes with two task formulations. It also investigates deterministic graph extraction and invariance-oriented training. The methods comparison for this publication checked the full paper.
-
Charles J. Lovering et al. 2025. Language Model Probabilities are Not Calibrated in Numeric Contexts. Proceedings of ACL 2025, pages 29218–29257. Examines probability outputs under known numerical event distributions. It provides close prior context for the separate numerical investigations, which are not the main contribution of this paper.
-
Pouya Pezeshkpour and Estevam Hruschka. 2024. Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions. Findings of NAACL 2024. Provides prior context for option-order effects and their relationship to uncertainty among alternatives.
-
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander Smola. 2012. A Kernel Two-Sample Test. Journal of Machine Learning Research, 13(25):723–773. Lemma 6, equation 3 defines the unbiased squared maximum mean discrepancy estimator. With the categorical equality kernel and two samples per condition, one half of that estimator equals the naming statistic used here.
Service documentation
The historical experiments record the explicit model identifier jev-1.13.0. Service documentation can change and should be refreshed before any live replication.
- TypeSafe documentation index.
- HTTP API reference.
- Question primitives.
- Confidence guide.
- Jev 1.13 limitations, reviewed 18 September 2026. The page identifies relevant limitations and recommends computing exact relations in code.
- TypeSafe agent skill.
Vendor descriptions of speed, calibration, training, and architecture are not empirical findings of this report. The article and paper distinguish structured output guarantees from task correctness.
Research record
The article's principal claims are derived from the completed naming, rewiring, and locality analyses. Their source hashes, original study identifiers, and exact numerical values are available through the evidence package. Prospective studies and synthetic software fixtures are labeled separately on the study preparation page.
The primary-source and full-method checks were completed on 18 September 2026. The related-work search was targeted rather than systematic or exhaustive. These references support a narrowly framed service case study, not a first demonstration of naming or order sensitivity.