Evaluating the Reliability of LLMs in OSINT Investigations: A Friend or Foe?
DOI:
https://doi.org/10.34190/eccws.25.1.4818Keywords:
Large Language Models (LLMs), Open-Source Intelligence (OSINT), ChatGPT, InvestigationAbstract
This study investigates the use of Large Language Models (LLMs), including GPT, Claude 3, Gemini, Meta, DeepSeek, Qwen 2.5, Mistral Large, and Grok, in Open-Source Intelligence (OSINT) investigations, focusing on their capabilities, limitations, and practical implications. Using a controlled fictional organization, the CtrlZ Society, to simulate a plausible online footprint, the LLMs were provided with a synthetic dataset purportedly linked to the organization. The dataset included social media posts, a Reddit thread, a Pastebin document, a GitHub repository, a blog post, a Telegram broadcast, a WHOIS record, and a news article. Based on this dataset, the models were evaluated on accuracy, timeline reconstruction, account attribution, evidence traceability, susceptibility to hallucination, and handling of ambiguity and incomplete information. Results revealed substantial variation among models: Claude 3, GPT, and Qwen 2.5 demonstrated strong analytical performance and reliable synthesis of investigative outputs, while Gemini and DeepSeek exhibited weaker capabilities. Some models, including Meta were also prone to forced narrative construction when prompted adversarially, highlighting risks of misinterpretation or overreach. Despite these limitations, all LLMs provided valuable support for structuring and summarising complex data, demonstrating their potential as efficiency multipliers in OSINT workflows. Based on these findings, the study provides recommendations for practitioners, including rigorous human oversight, multi-model validation, adherence to verification protocols, and careful evaluation of outputs to mitigate risks and maximise the reliability of LLM-assisted investigations.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 European Conference on Cyber Warfare and Security

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.