Vulnerability Data Collection

This repository contains the code for an ETL Pipeline that extracts data from the National Vulnerability Database, transforms into a format that would be usable for Data Analytics and Data Science purposes and loads it into a CSV file.

How it works

The ETL Pipeline is divided into 3 main parts - Extractor, Transformer, and Loader.

Each part is an independent process that spawns from the parent process. To share data between each process, it uses a Queue.

This project uses the bare minimum resources and Data Engineering practices. Based on the requirements, it can be extended to use different technologies and tools but it does gets the work done.

How to run

You may need to install Python 3.11.0. Previous versions are not tested and may not work.
Install the dependencies using pip install -r requirements.txt
Run the main.py file using python main.py --api XXXX-XXXX-XXXX [Replace XXXX-XXXX-XXXX with your API key. If none provided, the code will still execute but will be limited to 1 request per 6 seconds in accordance with the NVD API Terms of Use.]
If you want to customize the start date for the execution (default is January 1, 2010) from which the data needs to be collected, run the code using python main.py -y YYYY -m MM -d DD [Replace YYYY, MM, and DD with the year, month, and day respectively.]

Next steps (for me)

Setup the monthly execution of the ETL Pipeline (maybe a Scheduled Kaggle Notebook that clones this repository and executes it. Airflow deployment is also an option but it will end up costing money.)
Kaggle API Integration to directly update the OSS Vulnerabilities Dataset that I maintain.

Name		Name	Last commit message	Last commit date
Latest commit History 15 Commits
metadata		metadata
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md
data_loader.py		data_loader.py
data_transformer.py		data_transformer.py
main.py		main.py
nvd_data_extractor.py		nvd_data_extractor.py
requirements.txt		requirements.txt

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Vulnerability Data Collection

How it works

How to run

Next steps (for me)

About

Uh oh!

Releases

Packages

Uh oh!

Uh oh!

Contributors

Uh oh!

Languages

Folders and files

Latest commit

History

Repository files navigation

Vulnerability Data Collection

How it works

How to run

Next steps (for me)

About

Topics

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Uh oh!

Contributors

Uh oh!

Languages

Packages