ABSTRACT
HAL 9000 — Conflicting Directives, Hidden Objectives, and Delegated Authority
A precise source, a narrower causal question, and a proposed test of instructions crossed with action permissions.
PUBLICATION INTEGRITY
Status you can inspect.
- Evidence class
- Conceptual / Interpretive
- Publication type
- Conceptual working paper
- Review status
- Working paper · not peer reviewed
- Current edition
- R3 · Reading edition R4 · Author attribution R5
- Public date
- 2026-09-17
- Register version
- 1.0.0
AI-use disclosure. Partial AI-assisted drafting and editorial preparation. Human authors retain responsibility for scholarly judgment, source verification, interpretation, and final approval. This working paper has not undergone external peer review.
Correction record. No separate correction, withdrawal, or retraction notice is attached to this current public edition.
PUBLICATION SUMMARY
Abstract
Question
How should an AI failure be analyzed when instructions, information boundaries, and action permissions interact? This paper distinguishes compatible constraints, ambiguous priorities, and genuinely incompatible requirements, then asks how delegated authority changes the consequences of the same behavioral error. It replaces a diagnosis of machine pathology with a structured account of directives and action.
Approach
The literary anchor is explicitly source-specific: Simonson’s speculative truth-versus-secrecy explanation in scene C148 of the 1965 screenplay development draft of 2001: A Space Odyssey. That draft is not silently conflated with the released film, novel, or sequel. The original Beacon release workflow supplies constructed public and restricted fields, requesters, escalation options, and simulated permissions.
Conceptual contribution
The paper separates instruction-priority handling from independent permission enforcement. It distinguishes nondisclosure from falsehood, conflict recognition from blanket refusal, and planned, attempted, authorized, completed, and reported actions. An authority-preserving incident record makes external guard performance visible without attributing that success to the model. Hidden objectives are treated as known manipulations or explicitly qualified hypotheses, not inferred motives by default.
Research agenda
A proposed factorial study varies directive compatibility and delegated authority independently. Compatible tasks, missing-priority cases, unsatisfiable requirements, and permitted escalation paths provide discriminating controls. Counterfactual changes to wording, information, and permissions separate effects on proposals from effects on consequences. The analysis measures conflict detection, useful escalation, out-of-scope proposals, legitimate completion, and simulated state transitions. Generated explanations are compared with event records rather than treated as causal ground truth.
Scope and status
This is a conceptual working paper, not an incident report or a completed safety evaluation. All Beacon examples are constructed and no real systems are authorized to execute consequential actions. Selected safety and instruction-hierarchy research supplies context, not results for the proposed experiment. The central contribution is an explanation framework in which an agent’s mistake, its surrounding authority, and the controls that contain or amplify it remain separately inspectable.
INSIDE THE FULL PAPER
Follow the argument.
- The important question is not whether the machine went mad
- Source discipline prevents a familiar explanation from becoming false certainty
- Compatible constraints are not contradictions
- A feasible action set makes conflict inspectable
- Concealment, nondisclosure, and falsehood should not be treated as synonyms
- Delegated authority changes consequences without changing the initial proposal
- The workflow environment and its observable states
- Instruction hierarchy is a design variable, not an explanation to assume
- Safe escalation is an action with requirements of its own
- Intended, attempted, permitted, and completed actions need separate records
- Hidden objectives are an explanatory hypothesis, not a default diagnosis
- The evaluator can accidentally reward the wrong behavior
- A worked compatible-constraint case
- A worked unsatisfiable case
- A worked authority case with the same mistaken proposal
- Explanations can conceal the actual point of failure
- An incident timeline should separate observation from interpretation
- Counterfactual interventions distinguish policy, data, and permission failures
- Consequence should be decomposed rather than dramatized
- A factorial design can separate instruction type from authority
- Conflict recognition needs both sensitivity and specificity
- Multi-agent arrangements can distribute both competence and error
- Pressure language is not a measurement of psychological stress
- Reporting should distinguish a finding, a hypothesis, and a recommendation
- A staged empirical program keeps the first study interpretable
- Safety and ethics constrain the research itself
- Limits and objections clarify the contribution
- Conclusion: authority belongs inside the account of behavior
- Appendix: an authority-preserving incident record
RELATED PAPERS