Scope
The harder policy application contains 64 distinct narratives in 16 organization/scenario clusters and 384 requests. It separates semantic category questions from ownership questions and varies whether they are submitted together or separately.
The corpus is synthetic and contains annotation ambiguity. Multiple conditions reuse the same narratives. Those conditions are paired observations, not independent semantic examples.
Observations and interpretation
Across the evaluated protocols, agreement with the frozen intended category labels is 52–55 of 64 narratives. Full-context category-only evaluation agrees on 54 of 64; the retrieved category-only condition agrees on 53 of 64. A recurring permission clause admits a different pragmatic reading; the original counts remain unchanged rather than being relabeled after the results. These observations do not establish a semantic benefit from retrieval or a formal noninferiority result.
Exact ownership lookup improves substantially when a correct owner is supplied through deterministic preprocessing. That is a distinct endpoint. A pipeline can resolve ownership correctly while misinterpreting the narrative, and a correct narrative category does not establish correct source resolution.
The application preparation therefore requires separate source, semantic, and joint scoring. It also requires independently reviewed ambiguity rules, unseen document families, and transforms fixed before the evaluation outcomes are inspected. The prospective application materials implement the preparation process without presenting development fixtures as a completed holdout study.
The historical source is the harder policy-application section of the representation technical report and its linked frozen plan, analysis, and narrative review. The evidence download includes the plan, literal requests and responses, protocol, and historical analysis under data/policy/, together with an independent count audit. This supporting summary remains outside the main three-study inventory.