Abstract

In the United States political sphere, growing polarization between Democrats and Republicans, particularly their more ideologically extreme wings, has given rise to distinct English-adjacent languages, shaped by social tendencies and referred to as “sociolects”. These sociolects offer a new avenue to analyze and further quantify growing polarization, and as Generative Artificial Intelligence (GenAI) and Large Language Models (LLMs) become increasingly embedded in political and social analysis, understanding their depth of political reasoning, biases, and classification behavior has become imperative. In this study, we sourced comment data from the CNN, Fox, and MSNBC YouTube channels, spanning from 2020 through 2024, trained FastText embeddings, and identified misaligned linguistic pairs: words that carry different meanings across sociolects despite similarity in lexicon (e.g. undocumented vs. illegal), on a year-by-year and network-by-network basis. We then prompted three independently trained LLMs to classify and rationalize each misalignment. We found that lexical misalignment within and between these news networks generally increased across the study timeframe, consistent with continued political polarization in U.S. news media. Furthermore, we found that LLMs are not interchangeable, reliable interpreters of this divergence. One of the three models evaluated was both less reliable at recognizing trivial baseline data and substantially more likely to attribute politically charged categories to lexical divergence than the other two, indicating that model selection is a meaningful factor in using LLMs for political interpretation.

Publication Date

8-12-2026

Document Type

Thesis

Student Type

Graduate

Degree Name

Software Engineering (MS)

Department, Program, or Center

Software Engineering, Department of

College

Golisano College of Computing and Information Sciences

Advisor

Ashique KhudaBukhsh

Advisor/Committee Member

Larry Kiser

Campus

RIT – Main Campus

Share

COinS