GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
Massive Data Ingestion and Aggregation Module for Threat Intelligence
Authors
Ranjana Jadhav, Shireen Kulkarni, Samruddhi Mache, Kaivalya Mane, Himanshu Mantri, Hana Khan
Abstract
The cybersecurity world’s gotten a whole lot messier lately. Companies have to handle an overwhelming flood of threat data from all over: open intelligence feeds, vendor reports, the dark web, and non-stop streams of vulnerability updates. The real headache isn’t just gathering this chaotic jumble of information it’s trying to turn it into something organized, useful, and ready for analysis. Right now, most tools split data collection and analysis into separate steps. That might sound sensible, but in practice, it just slows everything down and risks losing important context or even data itself. This paper takes a different route. It lays out a Massive Data Ingestion and Aggregation Module that pulls everything together into one fast, smooth workflow. There’s automated scheduling, collection from multiple sources, schema normalization to standardize weird formats, machine learning to enrich the data, and full-text indexing to make searches easier and also all in a single pipeline. This system grabs data from places like NVD, CIRCL, OpenCTI, Exploit-DB, PhishTank, GitHub Security Advisories, and CISA KEV, ending up with a huge 342,168 vulnerability records ready for real analysis.
Pages:
5658 - 5665