{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"datasetVersion","sourceId":15876589,"datasetId":10179197,"databundleVersionId":16829770},{"sourceType":"kernelVersion","sourceId":313611246},{"sourceType":"kernelVersion","sourceId":313616188},{"sourceType":"kernelVersion","sourceId":313616207},{"sourceType":"kernelVersion","sourceId":313619655},{"sourceType":"kernelVersion","sourceId":313873640}],"dockerImageVersionId":31328,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Malware Classification with Stacking Ensemble\n\n## Step 1: Construction of the Meta-Training Dataset\n\nThis notebook implements the second layer of the Stacking Ensemble for the Microsoft Malware Classification Challenge.\n\nThe first layer consists of four diverse base learners:\n\n- Random Forest.\n- LightGBM.\n- XGBoost.\n- Multi-Layer Perceptron.\n\nEach base learner produces a probability distribution across the nine malware classes. Instead of using the original 413 malware features, the second-layer model learns from these probability outputs.\n\n### Out-of-Fold Training Probabilities\n\nThe following files are loaded from the four base-model notebooks:\n\n- `rf_OOF_train_for_stacking.csv`\n- `lgbm_OOF_train_for_stacking.csv`\n- `xgb_OOF_train_for_stacking.csv`\n- `mlp_OOF_train_for_stacking.csv`\n\nEach file contains:\n\n- `ID`: the original malware sample identifier.\n- Nine class-probability columns generated by the corresponding base learner.\n- `True_Class`: the ground-truth malware label.\n\nThe training probabilities are Out-of-Fold predictions. Each sample was predicted while it belonged to a validation fold and was excluded from the corresponding fold-level training subset. This reduces direct data leakage between the base learners and the second-layer model.\n\n### Meta-Feature Construction\n\nThe four probability tables are merged using the `ID` column.\n\nEach base learner contributes nine probability features:\n\n- 9 Random Forest probabilities.\n- 9 LightGBM probabilities.\n- 9 XGBoost probabilities.\n- 9 MLP probabilities.\n\nThe resulting meta-training matrix therefore contains:\n\n`4 models × 9 classes = 36 meta-features`\n\nThe `ID` column is used only for synchronization and is removed before model training. The target labels are obtained from the `True_Class` column of the Random Forest OOF table.\n\nThe resulting objects are:\n\n- `X_train_final`: the 36-dimensional meta-feature matrix.\n- `y_train_meta`: the corresponding ground-truth class labels.\n\nThis meta-training dataset is used to optimize and train the Logistic Regression meta-learner.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import classification_report, f1_score\nimport joblib\n\nprint(\"--- BƯỚC 1: HỢP NHẤT DỮ LIỆU OOF TRAIN (DÀNH CHO HUẤN LUYỆN TẦNG 2) ---\")\n\n# Nạp các file OOF Train đã sinh ra từ Tầng 1\nrf_train = pd.read_csv('/kaggle/input/notebooks/nguynlthanhhuy/random-forest-malware-classification/rf_OOF_train_for_stacking.csv')\nlgbm_train = pd.read_csv('/kaggle/input/notebooks/nguynlthanhhuy/lightgbm-malware-classification/lgbm_OOF_train_for_stacking.csv')\nxgb_train = pd.read_csv('/kaggle/input/notebooks/nguynlthanhhuy/xgboost-malware-classification/xgb_OOF_train_for_stacking.csv')\nmlp_train = pd.read_csv('/kaggle/input/notebooks/nguynlthanhhuy/multi-layer-perceptron-malware-classification/mlp_OOF_train_for_stacking.csv')\n\n# Hợp nhất dựa trên ID\nX_train_meta = pd.merge(rf_train.drop(columns=['True_Class']), lgbm_train.drop(columns=['True_Class']), on='ID')\nX_train_meta = pd.merge(X_train_meta, xgb_train.drop(columns=['True_Class']), on='ID')\nX_train_meta = pd.merge(X_train_meta, mlp_train.drop(columns=['True_Class']), on='ID')\n\n# Lấy nhãn thật (Labels)\ny_train_meta = rf_train['True_Class']\n\n# Loại bỏ cột ID để đưa vào huấn luyện\nX_train_final = X_train_meta.drop(columns=['ID'])\n\nprint(f\"✅ Đã chuẩn bị xong Meta-features cho tập Train: {X_train_final.shape}\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-07-23T07:19:00.270032Z","iopub.execute_input":"2026-07-23T07:19:00.270296Z","iopub.status.idle":"2026-07-23T07:19:03.017705Z","shell.execute_reply.started":"2026-07-23T07:19:00.270206Z","shell.execute_reply":"2026-07-23T07:19:03.016950Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 2: Construction of the Meta-Test Dataset\n\nThis step constructs an independent meta-test dataset for evaluating the complete Stacking Ensemble.\n\nThe following probability files are loaded from the four base-model notebooks:\n\n- `rf_OOF_test_for_stacking.csv`\n- `lgbm_OOF_test_for_stacking.csv`\n- `xgb_OOF_test_for_stacking.csv`\n- `mlp_OOF_test_for_stacking.csv`\n\nDespite the shared OOF naming convention, these files do not contain Out-of-Fold predictions. They contain probability predictions for the isolated test set, averaged across the fold-level models used by each base learner.\n\n### Test Meta-Feature Construction\n\nThe four test probability tables are merged using the `ID` column.\n\nThe merge order is kept consistent with the meta-training dataset:\n\n1. Random Forest.\n2. LightGBM.\n3. XGBoost.\n4. MLP.\n\nEach model contributes nine probabilities, producing another 36-dimensional feature matrix.\n\nThe `ID` column is removed before evaluation because it is an identifier rather than a predictive feature. The true test labels are obtained from the `True_Class` column of the Random Forest test table.\n\nThe resulting objects are:\n\n- `X_test_final`: the 36-dimensional meta-test matrix.\n- `y_test_meta`: the corresponding ground-truth malware labels.\n\nThis isolated meta-test dataset is not used to train the Logistic Regression model. It is reserved for evaluating the complete Stacking pipeline after hyperparameter optimization.","metadata":{}},{"cell_type":"code","source":"print(\"--- BƯỚC 2: HỢP NHẤT DỮ LIỆU OOF TEST (DÀNH CHO DỰ ĐOÁN CUỐI CÙNG) ---\")\n\n# Nạp các file OOF Test\nrf_test = pd.read_csv('/kaggle/input/notebooks/nguynlthanhhuy/random-forest-malware-classification/rf_OOF_test_for_stacking.csv')\nlgbm_test = pd.read_csv('/kaggle/input/notebooks/nguynlthanhhuy/lightgbm-malware-classification/lgbm_OOF_test_for_stacking.csv')\nxgb_test = pd.read_csv('/kaggle/input/notebooks/nguynlthanhhuy/xgboost-malware-classification/xgb_OOF_test_for_stacking.csv')\nmlp_test = pd.read_csv('/kaggle/input/notebooks/nguynlthanhhuy/multi-layer-perceptron-malware-classification/mlp_OOF_test_for_stacking.csv')\n\n# Hợp nhất dựa trên ID\nX_test_meta = pd.merge(rf_test.drop(columns=['True_Class']), lgbm_test.drop(columns=['True_Class']), on='ID')\nX_test_meta = pd.merge(X_test_meta, xgb_test.drop(columns=['True_Class']), on='ID')\nX_test_meta = pd.merge(X_test_meta, mlp_test.drop(columns=['True_Class']), on='ID')\n\n# Lấy nhãn Test thật để đánh giá cuối cùng\ny_test_meta = rf_test['True_Class']\nX_test_final = X_test_meta.drop(columns=['ID'])\n\nprint(f\"✅ Đã chuẩn bị xong Meta-features cho tập Test: {X_test_final.shape}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-23T07:19:03.019528Z","iopub.execute_input":"2026-07-23T07:19:03.019698Z","iopub.status.idle":"2026-07-23T07:19:03.087440Z","shell.execute_reply.started":"2026-07-23T07:19:03.019680Z","shell.execute_reply":"2026-07-23T07:19:03.086859Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 3: Logistic Regression Optimization and Stacking Evaluation\n\nAfter constructing the meta-training and meta-test datasets, a multinomial Logistic Regression classifier is used as the second-layer meta-learner.\n\nLogistic Regression is selected because the meta-feature space is relatively small and highly structured. The model learns how much confidence should be assigned to the probability outputs of each base learner for each malware class.\n\nIts linear structure also reduces the risk of overfitting compared with using a highly complex model at the second layer.\n\n### Meta-Learner Configuration\n\nThe Logistic Regression model is configured with:\n\n- `solver='lbfgs'` for multinomial optimization.\n- `max_iter=1000` to provide sufficient iterations for convergence.\n- `random_state=42` for reproducibility.\n- Multinomial classification for the nine malware classes.\n\n### Regularization Search\n\nThe regularization parameter `C` controls the inverse strength of L2 regularization:\n\n- A smaller `C` applies stronger regularization and restricts the model coefficients more heavily.\n- A larger `C` applies weaker regularization and allows the model to fit the meta-training data more closely.\n\nThe following candidate values are evaluated:\n\n`0.1, 0.5, 1.0, 5.0, 10.0`\n\n`GridSearchCV` evaluates every candidate using 5-fold cross-validation on the meta-training dataset.\n\nThe optimization metric is:\n\n`scoring='f1_macro'`\n\nMacro F1-Score gives equal importance to all nine malware classes. This is particularly important because the original dataset is highly imbalanced and overall Accuracy may hide weak performance on minority classes.\n\nAfter the search, the best fitted estimator is stored in `meta_model`.\n\n### Final Stacking Evaluation\n\nThe optimized meta-model generates:\n\n- Final class predictions through `predict`.\n- Final class probabilities through `predict_proba`.\n\nThe complete Stacking Ensemble is evaluated on the isolated meta-test dataset using:\n\n- Classification report.\n- Accuracy.\n- Macro F1-Score.\n- Multi-class Log Loss.\n\nAccuracy measures the overall proportion of correct predictions. Macro F1 evaluates the classes equally, while Log Loss evaluates the quality and confidence of the predicted probability distributions.\n\n### Confusion Matrix\n\nA confusion matrix is generated to show how predictions are distributed across the nine true and predicted malware classes.\n\nThe matrix is visualized as a heatmap and saved as:\n\n`confusion_matrix_stacking.png`\n\nThis visualization helps identify which malware families are classified correctly and which class pairs are occasionally confused.\n\n### Meta-Model Persistence\n\nThe optimized Logistic Regression model is saved as:\n\n`meta_model_stacking_final_v2.joblib`\n\nThe saved model is later loaded during inference on the real unlabeled competition test set.","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.metrics import classification_report, f1_score, log_loss, accuracy_score, confusion_matrix\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport joblib\n\nprint(\"--- BƯỚC 3: TỐI ƯU HÓA & HUẤN LUYỆN META-MODEL ---\")\n\n# 1. Khởi tạo Meta-model và không gian tìm kiếm C (Regularization)\n# C nhỏ = phạt nặng (chống overfit), C lớn = tin vào dữ liệu hơn\nlr = LogisticRegression(multi_class='multinomial', solver='lbfgs', max_iter=1000, random_state=42)\nparam_grid = {'C': [0.1, 0.5, 1.0, 5.0, 10.0]}\n\n# Dò tìm C tối ưu bằng 5-Fold trên tập Meta-Train\ngrid_search = GridSearchCV(lr, param_grid, cv=5, scoring='f1_macro', n_jobs=-1)\ngrid_search.fit(X_train_final, y_train_meta)\n\nmeta_model = grid_search.best_estimator_\nprint(f\"✅ Đã tìm thấy tham số tối ưu: {grid_search.best_params_}\")\n\n# 2. DỰ ĐOÁN CUỐI CÙNG TRÊN TẬP TEST\ny_final_pred = meta_model.predict(X_test_final)\ny_final_proba = meta_model.predict_proba(X_test_final) # Bắt buộc để tính Log Loss\n\n# 3. IN BÁO CÁO TỔNG KẾT CHO ĐỒ ÁN\nprint(\"\\n\" + \"=\"*60)\nprint(\"BÁO CÁO PHÂN LOẠI STACKING ENSEMBLE (KẾT QUẢ CUỐI CÙNG)\")\nprint(\"=\"*60)\ntarget_names = [f'Class {i}' for i in range(1, 10)]\nprint(classification_report(y_test_meta, y_final_pred, target_names=target_names))\n\n# Tính toán các chỉ số quan trọng\nfinal_f1 = f1_score(y_test_meta, y_final_pred, average='macro')\nfinal_logloss = log_loss(y_test_meta, y_final_proba)\n\n\n# 4. LƯU MÔ HÌNH CHỐT SỔ\n# Tính toán các chỉ số quan trọng (Bổ sung thêm accuracy_score)\nfinal_accuracy = accuracy_score(y_test_meta, y_final_pred)\nfinal_f1 = f1_score(y_test_meta, y_final_pred, average='macro')\nfinal_logloss = log_loss(y_test_meta, y_final_proba)\n\nprint(f\"==> ACCURACY:       {final_accuracy:.4f}\")\nprint(f\"==> MACRO F1-SCORE: {final_f1:.4f}\")\nprint(f\"==> LOG LOSS:       {final_logloss:.4f}\")\n\n#VẼ VÀ LƯU CONFUSION MATRIX\nprint(\"\\n\" + \"=\"*60)\nprint(\"TRỰC QUAN HÓA MA TRẬN NHẦM LẪN (CONFUSION MATRIX)\")\nprint(\"=\"*60)\n\n# Tạo ma trận\ncm = confusion_matrix(y_test_meta, y_final_pred)\n\n# Thiết lập kích thước ảnh và vẽ Heatmap\nplt.figure(figsize=(10, 8))\nsns.heatmap(cm, annot=True, fmt='d', cmap='Blues', \n            xticklabels=target_names, yticklabels=target_names)\n\n# Thêm tiêu đề và nhãn\nplt.title('Confusion Matrix - Stacking Ensemble (Logistic Regression)')\nplt.ylabel('True Class (Nhãn thực tế)')\nplt.xlabel('Predicted Class (Nhãn dự đoán)')\n\n# Lưu ảnh chất lượng cao để chèn vào báo cáo\nplt.savefig('confusion_matrix_stacking.png', dpi=300, bbox_inches='tight')\nplt.show()\n\nprint(\"\\n✅ Đã lưu ảnh Confusion Matrix thành 'confusion_matrix_stacking.png'\")\n\n# 5. LƯU MÔ HÌNH CHỐT SỔ (Giữ nguyên)\njoblib.dump(meta_model, 'meta_model_stacking_final_v2.joblib')\nprint(\"\\n✅ Đã lưu Meta-model thành công!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-23T07:19:03.088120Z","iopub.execute_input":"2026-07-23T07:19:03.088346Z","iopub.status.idle":"2026-07-23T07:19:07.370362Z","shell.execute_reply.started":"2026-07-23T07:19:03.088322Z","shell.execute_reply":"2026-07-23T07:19:07.369626Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 4: Real Competition Inference and Submission Generation\n\nAfter training and evaluating the complete Stacking Ensemble, the saved models are applied to the real unlabeled test set from the Microsoft Malware Classification Challenge.\n\n### Real Test Feature Input\n\nThe test feature table is read directly from the saved output of the feature-extraction notebook:\n\n`hnlanh/test-data-feature-malware-classification`\n\nThe input file is located at:\n\n`/kaggle/input/notebooks/hnlanh/test-data-feature-malware-classification/REAL_TEST_FEATURES_KAGGLE.csv`\n\nThis file contains:\n\n- `ID`: the identifier of each unlabeled malware sample.\n- The same extracted features used to train the four base learners.\n\nThe `ID` column is preserved for submission generation and removed from the model input matrix.\n\n### Loading the Trained Models\n\nThe four fitted base learners are loaded from their corresponding notebook outputs:\n\n- `best_rf_optimized_oof.joblib`\n- `best_lgbm_optimized.joblib`\n- `best_xgb.joblib`\n- `best_mlp_optimized_oof.joblib`\n\nThe fitted Logistic Regression meta-learner is also loaded from:\n\n`meta_model_stacking_final_v2.joblib`\n\n### First-Layer Inference\n\nEach base learner predicts a nine-class probability distribution for every real test sample:\n\n- `prob_rf`: Random Forest probabilities.\n- `prob_lgbm`: LightGBM probabilities.\n- `prob_xgb`: XGBoost probabilities.\n- `prob_mlp`: MLP probabilities.\n\nThese probability matrices are concatenated in the same order used during meta-model training:\n\n`Random Forest → LightGBM → XGBoost → MLP`\n\nThe resulting `X_real_meta` matrix contains 36 meta-features for every unlabeled malware sample.\n\nMaintaining the same model and column order is essential because the Logistic Regression coefficients were learned according to this exact arrangement.\n\n### Second-Layer Inference\n\nThe 36-dimensional meta-feature matrix is passed to the saved Logistic Regression model.\n\nThe meta-learner produces the final probability distribution across the nine malware classes:\n\n`final_probs`\n\n### Inference-Time Measurement\n\nExecution time is measured from the beginning of first-layer prediction until the second-layer probabilities have been produced.\n\nTwo timing measurements are reported:\n\n- Total inference time for the complete real test set.\n- Average inference time per malware sample.\n\nThis measurement reflects the prediction stage of the Stacking Ensemble and does not include the earlier feature-extraction process.\n\n### Submission Generation\n\nThe final probability matrix is converted into the Kaggle submission format:\n\n- `Id`\n- `Prediction1`\n- `Prediction2`\n- ...\n- `Prediction9`\n\nThe original malware identifiers are inserted as the first column, and the final file is saved as:\n\n`submission.csv`\n\nThis file contains the complete Stacking Ensemble predictions for the unlabeled competition test set.","metadata":{}},{"cell_type":"code","source":"print(\"--- BƯỚC 4: QUÉT MÃ ĐỘC VÀ ĐO THỜI GIAN THỰC TẾ ---\")\nimport pandas as pd\nimport numpy as np\nimport joblib\nimport time\n\n# 1. Đọc dữ liệu test thực tế\nreal_test_df = pd.read_csv(\n    \"/kaggle/input/notebooks/hnlanh/test-data-feature-malware-classification/REAL_TEST_FEATURES_KAGGLE.csv\"\n)\nX_real_test = real_test_df.drop(columns=['ID'])\n\n# 2. Load lại TOÀN BỘ các mô hình đã lưu\nbest_rf = joblib.load('/kaggle/input/notebooks/nguynlthanhhuy/random-forest-malware-classification/best_rf_optimized_oof.joblib')\nbest_lgbm = joblib.load('/kaggle/input/notebooks/nguynlthanhhuy/lightgbm-malware-classification/best_lgbm_optimized.joblib')\nbest_mlp = joblib.load('/kaggle/input/notebooks/nguynlthanhhuy/multi-layer-perceptron-malware-classification/best_mlp_optimized_oof.joblib')\nbest_xgb = joblib.load('/kaggle/input/notebooks/nguynlthanhhuy/xgboost-malware-classification/best_xgb.joblib') \nmeta_model = joblib.load('meta_model_stacking_final_v2.joblib')\n\nprint(\"⏳ Bắt đầu quét mã độc...\")\nstart_inference = time.time() # <--- BẮT ĐẦU BẤM GIỜ\n\n# 3. Đi qua Tầng 1\nprob_rf = best_rf.predict_proba(X_real_test)\nprob_lgbm = best_lgbm.predict_proba(X_real_test)\nprob_mlp = best_mlp.predict_proba(X_real_test)\nprob_xgb = best_xgb.predict_proba(X_real_test)\n\n# 4. Hợp nhất Meta-features\nX_real_meta = np.hstack((prob_rf, prob_lgbm, prob_xgb, prob_mlp))\n\n# 5. Đi qua Tầng 2\nfinal_probs = meta_model.predict_proba(X_real_meta)\n\nend_inference = time.time() # <--- KẾT THÚC BẤM GIỜ\n\n# 6. In kết quả thời gian quét\ntotal_time = end_inference - start_inference\ntime_per_sample = total_time / len(X_real_test) \n\nprint(f\"⏱️ Tổng thời gian quét {len(X_real_test)} files: {total_time:.4f} giây\")\nprint(f\"⚡ Tốc độ trung bình (Inference time): {time_per_sample:.6f} giây/file\")\n\n# 7. Xuất file submission để nộp\nsubmission = pd.DataFrame(final_probs, columns=[f'Prediction{i}' for i in range(1, 10)])\nsubmission.insert(0, 'Id', real_test_df['ID'])\nsubmission.to_csv('submission.csv', index=False)\nprint(\"✅ Đã tạo file submission.csv thành công!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-23T07:19:07.371275Z","iopub.execute_input":"2026-07-23T07:19:07.371520Z","iopub.status.idle":"2026-07-23T07:19:12.516539Z","shell.execute_reply.started":"2026-07-23T07:19:07.371491Z","shell.execute_reply":"2026-07-23T07:19:12.515547Z"}},"outputs":[],"execution_count":null}]}