Evaluating the Reliability of LLMs in OSINT Investigations: A Friend or Foe?

Authors

  • Errol Baloyi Council for Scientific and Industrial Research (CSIR)
  • Nokuthaba Siphambili Council for Scientific and Industrial Research (CSIR)
  • Ntomfuthi Ntshangase Council for Scientific and Industrial Research (CSIR)
  • Mpho Letshwenyo Council for Scientific and Industrial Research (CSIR)
  • Fhatuwani Makharamedzha Council for Scientific and Industrial Research (CSIR)
  • Rendani Mmbodi Council for Scientific and Industrial Research (CSIR)
  • Ndabezinhle Hlongwane Council for Scientific and Industrial Research (CSIR)

DOI:

https://doi.org/10.34190/eccws.25.1.4818

Keywords:

Large Language Models (LLMs), Open-Source Intelligence (OSINT), ChatGPT, Investigation

Abstract

This study investigates the use of Large Language Models (LLMs), including GPT, Claude 3, Gemini, Meta, DeepSeek, Qwen 2.5, Mistral Large, and Grok, in Open-Source Intelligence (OSINT) investigations, focusing on their capabilities, limitations, and practical implications. Using a controlled fictional organization, the CtrlZ Society, to simulate a plausible online footprint, the LLMs were provided with a synthetic dataset purportedly linked to the organization. The dataset included social media posts, a Reddit thread, a Pastebin document, a GitHub repository, a blog post, a Telegram broadcast, a WHOIS record, and a news article. Based on this dataset, the models were evaluated on accuracy, timeline reconstruction, account attribution, evidence traceability, susceptibility to hallucination, and handling of ambiguity and incomplete information. Results revealed substantial variation among models: Claude 3, GPT, and Qwen 2.5 demonstrated strong analytical performance and reliable synthesis of investigative outputs, while Gemini and DeepSeek exhibited weaker capabilities. Some models, including Meta were also prone to forced narrative construction when prompted adversarially, highlighting risks of misinterpretation or overreach. Despite these limitations, all LLMs provided valuable support for structuring and summarising complex data, demonstrating their potential as efficiency multipliers in OSINT workflows. Based on these findings, the study provides recommendations for practitioners, including rigorous human oversight, multi-model validation, adherence to verification protocols, and careful evaluation of outputs to mitigate risks and maximise the reliability of LLM-assisted investigations.

Author Biographies

Errol Baloyi, Council for Scientific and Industrial Research (CSIR)

Errol Baloyi is a cybersecurity professional with a multidisciplinary background spanning the military, academia, and the research sector. He is currently a cybersecurity researcher at the Council for Scientific and Industrial Research (CSIR) and a cybersecurity instructor. He is a Certified Ethical Hacker and an Associate Member of the Institute of Commercial Forensic Practitioners (South Africa). His research interests and areas of expertise include open-source intelligence, threat intelligence, penetration testing, and digital forensics.

Nokuthaba Siphambili, Council for Scientific and Industrial Research (CSIR)

Nokuthaba Siphambili is a cybersecurity researcher. Her current research interests lie at the intersection of cybersecurity, governance, privacy, and trust; particularly focused on exploring innovative approaches to enhance cybersecurity measures, address governance challenges in information systems, promote privacy and trust in digital environments.

Ntomfuthi Ntshangase, Council for Scientific and Industrial Research (CSIR)

Ntomfuthi Lungile Ntshangase is a Cybersecurity Researcher at the Council for Scientific and Industrial Research (CSIR) in South Africa, working within the Defense and Security unit. Her research focuses on Governance, Risk, and Compliance (GRC), bridging the gap between regulatory requirements and technical security controls to help organizations achieve measurable, defensible compliance.

 

Mpho Letshwenyo, Council for Scientific and Industrial Research (CSIR)

Mpho Letshwenyo is a cybersecurity specialist specializing in penetration testing, vulnerability assessments, and cybersecurity governance. Her work focuses on identifying and mitigating security threats to ensure robust and secure systems. In addition to her technical expertise, she conducts cybersecurity research to stay ahead of emerging threats and industry trends. She also has a background in software development, having previously worked as a front-end developer.

Fhatuwani Makharamedzha, Council for Scientific and Industrial Research (CSIR)

Fhatuwani Makharamedzha is an emerging researcher and technology enthusiast, with interests in artificial intelligence, software development, cybersecurity and data-driven innovation. committed to advancing innovative research that contributes to sustainable development and technological advancement through the effective use of emerging technologist.

Rendani Mmbodi , Council for Scientific and Industrial Research (CSIR)

Rendani Mmbodi is an emerging researcher specializing in cybersecurity and software engineering. His research interests include digital learning technologies, cyber-awareness systems, and secure software development practices.

Ndabezinhle Hlongwane, Council for Scientific and Industrial Research (CSIR)

Ndabe Hlongwane is a software developer and researcher whose work focuses on cybersecurity education, digital awareness systems, and secure application development. He has contributed to projects involving hackathons, mobile applications, and IT training programmes.

Downloads

Published

2026-06-15