Abstract
Models deployed outside controlled settings encounter inputs unlike their training data and may generate outputs unsupported by what they learned. A reliable system should recognize these cases. Out-of-distribution detection and hallucination detection are usually studied separately, but both depend on the same condition: the model's representation must retain the information that distinguishes an anomalous case from an ordinary one. If training never captures that information, or later suppresses it, no downstream detector can recover it. We study this condition from two complementary perspectives. Mutual information measures how much information about the relevant distinction remains in a representation, while feature-space geometry shows whether the directions needed to detect a shift retain sufficient variation. The first study proves that unlabeled detection must fail when the self-supervised or unsupervised objective is independent of the label-relevant features. It also introduces the Adjacent OOD benchmark to evaluate this high-overlap regime. The second study identifies domain-sensitivity collapse, in which single-domain training suppresses shift-sensitive directions, and introduces Teacher-Guided Training to restore them without adding inference cost. The third study develops a hallucination detector from cross-layer activations. Its contrastive probe reaches parity with the strongest engineered probe and outperforms it in the mean, consistent with the study's feasibility argument. The unifying representation-structure framework is the central contribution. The dissertation establishes a conditional limit on unlabeled OOD detection and demonstrates interventions when relevant structure can be restored or read out; it does not claim these methods are optimal.
Publication Date
8-2026
Document Type
Dissertation
Student Type
Graduate
Degree Name
Computing and Information Sciences (Ph.D.)
Department, Program, or Center
Computing and Information Sciences Ph.D, Department of
College
Golisano College of Computing and Information Sciences
Advisor
Travis Desell
Advisor/Committee Member
Qi Yu
Advisor/Committee Member
Alexander Ororbia
Recommended Citation
Yang, Hong, "What Representations Preserve: Information-Theoretic and Geometric Conditions for Out-of-Distribution and Hallucination Detection" (2026). Thesis. Rochester Institute of Technology. Accessed from
https://repository.rit.edu/theses/12805
Campus
RIT – Main Campus
