ABSTRACT
Arrival — Understanding the Mind Behind the Language
Before comparing answers, establish which distinctions the question and its representation make available.
PUBLICATION INTEGRITY
Status you can inspect.
- Evidence class
- Conceptual / Interpretive
- Publication type
- Conceptual working paper
- Review status
- Working paper · not peer reviewed
- Current edition
- R3 · Reading edition R4 · Author attribution R5
- Public date
- 2026-09-17
- Register version
- 1.0.0
AI-use disclosure. Partial AI-assisted drafting and editorial preparation. Human authors retain responsibility for scholarly judgment, source verification, interpretation, and final approval. This working paper has not undergone external peer review.
Correction record. No separate correction, withdrawal, or retraction notice is attached to this current public edition.
PUBLICATION SUMMARY
Abstract
Question
Before judging an answer, have we established that the question makes the intended distinctions available? This paper uses Arrival’s elicitation problem to examine how representation, information, objective specification, and response format can be conflated when evaluating artificial cognition. Its focus is task framing, not the film’s speculative claims about language and time.
Approach
The argument distinguishes what a system is given from how it is encoded, what it is asked to optimize, and how it must respond. A constructed directed-route world is rendered as prose, tables, and diagrams from the same underlying record. Worked cases include conflicting objectives, infeasible requests, multiple admissible answers, missing preferences, and visually ambiguous relations.
Conceptual contribution
The paper proposes a representation-equivalence certificate that maps each task-relevant fact to its rendered location in every format. It separates source-data equivalence from perceptual accessibility and from equal task difficulty. Clarification is treated as useful when it resolves a decision-relevant ambiguity, not simply when a system asks more questions. A format-sensitive performance difference is a bounded behavioral finding; it does not uniquely reveal a system’s internal architecture or mode of experience.
Research agenda
A staged evaluation tests extraction, constraint interpretation, planning, and checking separately before combining them. Independent solvers define admissible answer sets. Matched-information comparisons, transcribed-image controls, objective changes, and metamorphic transformations help distinguish perception failures from reasoning or task-specification failures. An oracle clarification channel permits assessment of whether a requested detail actually improves the decision. World-level sampling, declared exclusions, and separate human comparisons constrain generalization.
Scope and status
This conceptual working paper presents an unexecuted research program, not results from a multimodal benchmark. The linguistic consultant’s account anchors the fictional discussion; selected research supplies methodological context. Constructed numerical examples illustrate distinctions rather than estimate performance. The central proposal is to investigate the assumptions built into an evaluation before interpreting an unfamiliar system’s response as evidence of a different kind of mind.
INSIDE THE FULL PAPER
Follow the argument.
- Before an answer, there is a theory of the question
- The fictional premise and the empirical question must be kept apart
- Representation, information, and objective are different variables
- Form does not settle meaning by itself
- A synthetic world with several defensible answers
- Generate formats from one authoritative record
- The objective should be a manipulated variable, not an evaluator's assumption
- Clarification is an information-seeking action with measurable value
- A response pipeline separates extraction from planning
- Metamorphic tests examine which changes should matter
- Multiple objectives reveal the difference between ambiguity and error
- Formal notation can clarify the claim without proving too much
- The comparison system should expose simpler explanations
- A worked ambiguity case with no fabricated model result
- A worked representation case separates perception from inference
- Statistical planning should follow the questions, not the desired headline
- Behavioral sensitivity does not uniquely identify an internal representation
- Human comparisons need a purpose and a valid interpretation
- Interactive access changes the task, and that change should be measured
- Temporal representation is not evidence of temporal experience
- Category construction can be the hidden source of disagreement
- Design implications should follow the narrow evidence
- A preregisterable protocol with explicit limits
- What the proposal leaves unresolved
- Conclusion: understanding begins by making the question answerable
- Appendix: a representation-equivalence certificate
RELATED PAPERS