Abstract
Cyber Threat Intelligence (CTI) helps security analysts respond to evolving threats, but the evidence it draws on is often scattered across noisy reports, advisories, and threat descriptions. Large language models (LLMs) provide a flexible interface for security analysis, yet reliable use requires grounding their outputs in source evidence, domain schemas, and CTI standards. We address this problem through structured threat representation, task-specific evaluation, and verifier-guided post-training. First, we study how unstructured CTI reports can be converted into structured threat knowledge. We develop an ontology-guided extraction framework that normalizes attack patterns, extracts typed entities and relations, and constructs a threat intelligence knowledge graph for querying and link prediction. We also examine an LLM-based extension that uses prompting, constrained generation, and fine-tuning to produce typed triples within the same structured representation. Second, we investigate LLM capabilities on CTI tasks through systematic benchmarking. The benchmarks cover closed-form knowledge, structured problem solving, extraction and mapping, and open-ended reasoning across cybersecurity advisory and CTI analysis tasks. Results show that reliable CTI performance cannot be inferred from general language ability alone; it must be measured through concrete analyst-facing tasks and task-specific metrics. Finally, motivated by the capability gaps revealed by benchmarking, we develop a CTI post-training framework based on reinforcement learning with verifiable rewards. The framework uses structured CTI targets to guide model adaptation, while self-training supplements this process when verifier feedback is sparse. Experiments show that structured domain knowledge can guide the adaptation of CTI-specialized LLMs across tasks and model backbones. Together, these contributions show that reliable LLM use in CTI depends on structured representations, task-grounded evaluation, and training signals grounded in verifiable domain knowledge.
Publication Date
7-2026
Document Type
Dissertation
Student Type
Graduate
Degree Name
Computing and Information Sciences (Ph.D.)
Department, Program, or Center
Computing and Information Sciences Ph.D, Department of
College
Golisano College of Computing and Information Sciences
Advisor
Nidhi Rastogi
Advisor/Committee Member
Matthew Wright
Advisor/Committee Member
Dongfang Liu
Recommended Citation
Alam, Md Tanvirul, "Towards Reliable Large Language Models for Cyber Threat Intelligence" (2026). Thesis. Rochester Institute of Technology. Accessed from
https://repository.rit.edu/theses/12773
Campus
RIT – Main Campus
