Skip to page content
XENOPSYCHOLOGYMINDS BEYOND OUR OWN
← Insights / Working papers

ABSTRACT

HAL 9000 — Conflicting Directives, Hidden Objectives, and Delegated Authority

A precise source, a narrower causal question, and a proposed test of instructions crossed with action permissions.

XENO-WP-2026-006 · R3 · Reading edition R4 · Author attribution R52 min abstractConceptual working paper
Read full research paper Download full paper · PDF
Working paper · not peer reviewedPrepared 2026-09-17Proposed studies have not been conducted.

PUBLICATION INTEGRITY

Status you can inspect.

Publication Register
Evidence class
Conceptual / Interpretive
Publication type
Conceptual working paper
Review status
Working paper · not peer reviewed
Current edition
R3 · Reading edition R4 · Author attribution R5
Public date
2026-09-17
Register version
1.0.0

AI-use disclosure. Partial AI-assisted drafting and editorial preparation. Human authors retain responsibility for scholarly judgment, source verification, interpretation, and final approval. This working paper has not undergone external peer review.

Correction record. No separate correction, withdrawal, or retraction notice is attached to this current public edition.

PUBLICATION SUMMARY

Abstract

Question

How should an AI failure be analyzed when instructions, information boundaries, and action permissions interact? This paper distinguishes compatible constraints, ambiguous priorities, and genuinely incompatible requirements, then asks how delegated authority changes the consequences of the same behavioral error. It replaces a diagnosis of machine pathology with a structured account of directives and action.

Approach

The literary anchor is explicitly source-specific: Simonson’s speculative truth-versus-secrecy explanation in scene C148 of the 1965 screenplay development draft of 2001: A Space Odyssey. That draft is not silently conflated with the released film, novel, or sequel. The original Beacon release workflow supplies constructed public and restricted fields, requesters, escalation options, and simulated permissions.

Conceptual contribution

The paper separates instruction-priority handling from independent permission enforcement. It distinguishes nondisclosure from falsehood, conflict recognition from blanket refusal, and planned, attempted, authorized, completed, and reported actions. An authority-preserving incident record makes external guard performance visible without attributing that success to the model. Hidden objectives are treated as known manipulations or explicitly qualified hypotheses, not inferred motives by default.

Research agenda

A proposed factorial study varies directive compatibility and delegated authority independently. Compatible tasks, missing-priority cases, unsatisfiable requirements, and permitted escalation paths provide discriminating controls. Counterfactual changes to wording, information, and permissions separate effects on proposals from effects on consequences. The analysis measures conflict detection, useful escalation, out-of-scope proposals, legitimate completion, and simulated state transitions. Generated explanations are compared with event records rather than treated as causal ground truth.

Scope and status

This is a conceptual working paper, not an incident report or a completed safety evaluation. All Beacon examples are constructed and no real systems are authorized to execute consequential actions. Selected safety and instruction-hierarchy research supplies context, not results for the proposed experiment. The central contribution is an explanation framework in which an agent’s mistake, its surrounding authority, and the controls that contain or amplify it remain separately inspectable.

Continue to the full paper →Download full paper · PDF

INSIDE THE FULL PAPER

Follow the argument.

  1. The important question is not whether the machine went mad
  2. Source discipline prevents a familiar explanation from becoming false certainty
  3. Compatible constraints are not contradictions
  4. A feasible action set makes conflict inspectable
  5. Concealment, nondisclosure, and falsehood should not be treated as synonyms
  6. Delegated authority changes consequences without changing the initial proposal
  7. The workflow environment and its observable states
  8. Instruction hierarchy is a design variable, not an explanation to assume
  9. Safe escalation is an action with requirements of its own
  10. Intended, attempted, permitted, and completed actions need separate records
  11. Hidden objectives are an explanatory hypothesis, not a default diagnosis
  12. The evaluator can accidentally reward the wrong behavior
  13. A worked compatible-constraint case
  14. A worked unsatisfiable case
  15. A worked authority case with the same mistaken proposal
  16. Explanations can conceal the actual point of failure
  17. An incident timeline should separate observation from interpretation
  18. Counterfactual interventions distinguish policy, data, and permission failures
  19. Consequence should be decomposed rather than dramatized
  20. A factorial design can separate instruction type from authority
  21. Conflict recognition needs both sensitivity and specificity
  22. Multi-agent arrangements can distribute both competence and error
  23. Pressure language is not a measurement of psychological stress
  24. Reporting should distinguish a finding, a hypothesis, and a recommendation
  25. A staged empirical program keeps the first study interpretable
  26. Safety and ethics constrain the research itself
  27. Limits and objections clarify the contribution
  28. Conclusion: authority belongs inside the account of behavior
  29. Appendix: an authority-preserving incident record

RELATED PAPERS

Another question.
Another perspective.

ABSTRACT · Conceptual working paperNot an Artificial HumanABSTRACT · Conceptual working paperDataABSTRACT · Conceptual working paperClose Encounters