Skip to content

Latest commit

Β 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Cyber Security Attacks Classifier

πŸš€ End-to-end multiclass classification of network attacks - a complete, reproducible ML pipeline from Kaggle download to interactive dashboard, with an honest reading of the results

A machine learning project that classifies network security events into three attack types β€” DDoS, Malware and Intrusion β€” using the Cyber Security Attacks dataset from Kaggle. The primary model is a Random Forest; Gradient Boosting and k-NN are trained alongside it for comparison. A Streamlit dashboard walks through exploratory analysis, preprocessing, model design, comparison and evaluation.

The project deliberately reports what it actually found rather than what would look good. After identifiers and free-text fields are removed, the remaining tabular features carry little usable signal for Attack Type in this dataset β€” all three models land near the random baseline. The pipeline and metrics are correct; the finding is the result. See Notes.

Python scikit-learn pandas Streamlit Kaggle License


🎯 Key Features

  • 🎯 Balanced three-class target β€” Attack Type with DDoS, Malware and Intrusion at roughly one third each, over 40,000 instances and 25 attributes.
  • πŸ“₯ Automatic data acquisition β€” download_data.py fetches the official CSV via Kaggle Hub into data/cybersecurity_attacks.csv, caching under .kaggle_cache/.
  • πŸ”¬ Full EDA suite β€” class distribution, missing-value profile, numeric and categorical distributions, a correlation heatmap and a mutual-information ranking against the target.
  • βš™οΈ Documented preprocessing β€” drops non-generalizing columns (timestamps, IPs, payload, geo/proxy data), converts sparse fields to binary presence flags, label-encodes categoricals, standard-scales numerics, then makes an 80/20 stratified split.
  • 🌲 Three models compared β€” Random Forest (n_estimators=200, class_weight="balanced"), Gradient Boosting (n_estimators=100) and k-NN (k=7), each with 5-fold cross-validation.
  • πŸ“ˆ Complete evaluation β€” accuracy, macro F1, precision, recall, one-vs-rest ROC AUC, confusion matrices (counts and normalized), ROC curves and Random Forest feature importances.
  • πŸ–₯️ Eight-tab Streamlit dashboard β€” every stage of the study explorable, including an interactive feature explorer.
  • ♻️ One-command reproduction β€” ./start.sh (or start.bat) resolves Python, builds the venv, downloads data, runs the pipeline and opens the dashboard.

πŸ“Š Results & Visualizations

Streamlit dashboard preview showing the Cyber Security Attacks classifier interface with exploratory analysis and model results

Running python pipeline.py regenerates every figure and metric into results/ (git-ignored, see Notes):

Artifact Content
class_distribution.png Attack Type distribution with counts and percentages
missing_values.png Missing values per column
numeric_distributions.png Histograms for Source Port, Destination Port, Packet Length, Anomaly Scores
categorical_distributions.png Eight categorical feature distributions
correlation_heatmap.png Correlation across the numeric features
mutual_information.png Mutual information with Attack Type, against a 0.01 threshold
model_comparison.png Random Forest vs Gradient Boosting vs k-NN across five metrics, with the random baseline (0.333) marked
confusion_matrix.png Random Forest confusion matrix, counts and normalized
roc_curves.png One-vs-rest ROC curves per class
feature_importance.png Random Forest feature importances

Machine-readable counterparts are written beside them: metrics.json, model_comparison.json, cv_scores.json, eda_summary.json, preprocessing_info.json, roc_data.json, feature_importance.json and confusion_matrix.npy.

Concrete numbers are intentionally not quoted in this README β€” read them from your own results/metrics.json after running the pipeline, and see Notes for how to interpret them.


πŸ—οΈ Pipeline

Pipeline diagram: Kaggle download, exploratory analysis, preprocessing, an 80/20 stratified split, training of Random Forest with Gradient Boosting and k-NN comparison models, evaluation, and artifact export to the Streamlit dashboard

download_data.py  β†’  pipeline.py  β†’  app.py
   Kaggle CSV        EDA, preprocessing,      Streamlit dashboard
   into data/        training, evaluation     reading results/ and models/
                     into results/, models/

Stages in pipeline.py

  1. Load β€” ensure_dataset() downloads the CSV if absent, then reads it with pandas.
  2. EDA β€” six plot groups plus a mutual-information analysis, summarized to eda_summary.json.
  3. Preprocess β€” drop ten non-predictive columns; convert Malware Indicators and Alerts/Warnings to binary presence flags; label-encode the target and remaining categoricals; StandardScaler on the four numeric columns; train_test_split(test_size=0.2, random_state=42, stratify=y).
  4. Train β€” Random Forest (primary) with 5-fold cross_val_score, plus Gradient Boosting and k-NN under the same protocol.
  5. Evaluate β€” detailed Random Forest metrics, classification report, confusion matrices, OvR ROC curves and feature importances.

🧩 Modules

Path Purpose
download_data.py Download dataset from Kaggle into data/
pipeline.py Full ML pipeline and evaluation
app.py Streamlit dashboard
description.md Detailed project description (Polish)
data/ Dataset CSV (generated; see .gitignore)
results/ Metrics, JSON, PNG plots, NumPy confusion matrix (generated)
models/ Saved random_forest.joblib (generated)
start.sh / start.bat One-command setup + pipeline + Streamlit

Dashboard tabs (app.py)

Tab Content
πŸ“‹ Project Overview Dataset description and the attribute table
πŸ“Š Exploratory Data Analysis Distributions, missing values, correlations, mutual information
βš™οΈ Preprocessing Dropped columns, encodings, scaling, split sizes
🌲 Model & Training Random Forest design and hyperparameters
βš–οΈ Model Comparison Random Forest vs Gradient Boosting vs k-NN
πŸ“ˆ Results & Evaluation Metrics, confusion matrix, ROC curves, feature importance
πŸ” Interactive Explorer Scatter and distribution exploration over the raw features
ℹ️ Informacje Supplementary notes

πŸ› οΈ Technology Stack

Machine Learning

  • scikit-learn (>=1.3) β€” RandomForestClassifier, GradientBoostingClassifier, KNeighborsClassifier, LabelEncoder, StandardScaler, train_test_split, cross_val_score, mutual_info_classif, the full metrics suite
  • joblib (>=1.3) β€” model persistence to models/random_forest.joblib
  • NumPy (>=1.24) β€” numeric arrays and confusion-matrix export

Data

  • pandas (>=2.0) β€” loading, cleaning and aggregation
  • kagglehub (>=0.2) β€” dataset download from Kaggle

Visualization & UI

  • Matplotlib (>=3.7) β€” all pipeline figures (Agg backend, headless-safe)
  • seaborn (>=0.13) β€” heatmaps and confusion matrices
  • Plotly (>=5.18) β€” interactive dashboard charts
  • Streamlit (>=1.30) β€” the eight-tab dashboard

πŸš€ Getting Started

Prerequisites

  • Docker β€” for the containerised path below (recommended)

  • Python 3.10–3.13 (tested with 3.12; 3.14+ is not supported yet for this stack).

  • Internet access on first run to download the dataset (~5 MB) unless data/cybersecurity_attacks.csv is already present.

Run with Docker (recommended)

The container installs the scientific stack, runs the experiment and serves the dashboard in one step β€” no local Python setup and no virtual environment:

docker compose -f .tools/docker/docker-compose.yml up --build

The dashboard is then available at http://localhost:8501.

First run downloads the dataset. The Kaggle CSV is not committed, so the container needs outbound network access the first time it boots. Later boots reuse the copy on the mounted data/ volume.

Generated artefacts (results/, data/) are bind-mounted back to the host, so charts and metrics written inside the container survive it being removed. Stop the stack with:

docker compose -f .tools/docker/docker-compose.yml down

1. Clone the Repository

git clone https://github.com/dawidolko/CyberAttack-Classifier-Python.git
cd CyberAttack-Classifier-Python

2. Install Dependencies

python3 -m venv venv
source venv/bin/activate          # Windows: venv\Scripts\activate
pip install -r requirements.txt

3. Run

One command (recommended)

Linux / macOS

chmod +x start.sh
./start.sh

Windows β€” double-click start.bat or run in cmd / PowerShell:

start.bat

The script will:

  1. Resolve Python 3.10–3.13.
  2. Create venv/ if needed and pip install -r requirements.txt.
  3. Run python pipeline.py (downloads data if missing, trains models, writes results/).
  4. Start the Streamlit app at http://localhost:8501 (streamlit run app.py).

Stop the server with Ctrl+C.

Manual steps

python download_data.py           # optional; pipeline also downloads if needed
python pipeline.py
streamlit run app.py

πŸ“ Project Structure

CyberAttack-Classifier-Python/
β”œβ”€β”€ πŸ“₯ download_data.py        # Kaggle download via kagglehub
β”œβ”€β”€ πŸ”¬ pipeline.py             # Full ML pipeline: EDA, preprocessing, training, evaluation
β”œβ”€β”€ πŸ–₯️ app.py                  # Streamlit dashboard (8 tabs)
β”œβ”€β”€ πŸ“Š data/                   # Dataset CSV (generated, git-ignored)
β”‚   └── README.md              # How the CSV is obtained
β”œβ”€β”€ πŸ“ˆ results/                # Metrics JSON, PNG plots, confusion matrix (generated)
β”œβ”€β”€ πŸ€– models/                 # random_forest.joblib (generated)
β”œβ”€β”€ πŸ–ΌοΈ img/
β”‚   β”œβ”€β”€ cyberattack-preview.png # Dashboard preview
β”‚   └── logo.svg               # Sidebar logo
β”œβ”€β”€ πŸ“š docs/
β”‚   β”œβ”€β”€ diagrams/pipeline.svg  # Pipeline diagram
β”‚   └── dokumentacja_do125148.docx
β”œβ”€β”€ πŸ“ description.md          # Detailed project description (Polish)
β”œβ”€β”€ πŸš€ start.sh / start.bat    # One-command setup + pipeline + dashboard
β”œβ”€β”€ πŸ“¦ requirements.txt
└── πŸ“– README.md

πŸŽ“ Academic report (Polish course outline)

For the Sztuczna inteligencja report, map sections as follows:

  1. Student data β€” name, program, year, academic year (fill in manually).
  2. Course β€” Artificial Intelligence (or your exact course title).
  3. Project topic β€” Multiclass classification of cyber security attacks (Random Forest on Kaggle dataset).
  4. Problem characterization β€” Supervised multiclass classification; balanced three-class target; network and security features with missing values in several columns.
  5. Number of instances β€” 40,000.
  6. Attributes β€” 25; use the table in the Streamlit Project Overview tab and the dataset documentation on Kaggle.
  7. Preprocessing β€” Summarize steps from the Preprocessing tab / preprocessing_info.json (dropped columns, binary flags, encodings, scaling, split).
  8. Model design β€” Random Forest (primary), plus Gradient Boosting and k-NN for comparison; hyperparameters as in the Model & Training tab and pipeline.py.
  9. Results β€” Accuracy, macro F1, precision, recall, ROC AUC, confusion matrix, per-class metrics (metrics.json / dashboard).
  10. Conclusions β€” Strengths of RF on this task, role of important features, limitations (e.g. label encoding of IPs removed; text fields dropped).

πŸ“Œ Notes

  • Empirical performance: On this Kaggle release, test accuracy is often near the random baseline (β‰ˆ1/3) for balanced three-class prediction, and ROC AUC is near 0.5, with all three models behaving similarly. That is a valid finding for your report: after removing identifiers and free text, the remaining tabular features may carry little usable signal for Attack Type in this synthetic split. The pipeline and metrics are still correct; interpret results honestly in section Wnioski / Conclusions.
  • Git: venv/, .kaggle_cache/, data/cybersecurity_attacks.csv, and generated results/ / models/ artifacts are listed in .gitignore. Clone the repo and run ./start.sh to regenerate everything.
  • Kaggle authentication: Public dataset download via kagglehub typically works without extra setup; if you hit auth errors, follow Kaggle API credentials and set KAGGLE_USERNAME / KAGGLE_KEY or place kaggle.json in ~/.kaggle/.

πŸ“„ License

Dataset usage is subject to the Kaggle dataset license. This repository code is provided for educational use β€” see the LICENSE file.


πŸ‘¨β€πŸ’» Author

Created by Dawid Olko

About

Multiclass classification of network attacks (DDoS / Malware / Intrusion) based on the Cyber Security Attacks dataset from Kaggle. Built with Python using Random Forest and Gradient Boosting models. University project for Artificial Intelligence course.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages