{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"colab":{"provenance":[]},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":4117,"databundleVersionId":46665,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"1-55WKas28Lr","cell_type":"markdown","source":"# Microsoft Malware Classification Challenge\n\n---\n\n## Problem Statement  \nThe increasing sophistication and frequency of malware attacks present a significant threat to cybersecurity. Traditional signature-based detection methods are often ineffective against rapidly evolving malware families.  \n\nThis project aims to build a robust classification model to automatically detect and categorize malware files into **9 distinct families** based on both raw binary content and disassembled metadata logs.  \n\nThe model leverages machine learning techniques to classify malware effectively, enabling faster and more accurate threat detection and response.\n\n---\n\n## Dataset Overview\n\n- **Train Set:** Hexadecimal and disassembly log files for malware samples\n- **Labels:** Provided in `trainLabels.csv` (class integers 1-9)\n- **Test Set:** Similar structure, without labels\n- **Sample Submission:** Format provided in `sampleSubmission.csv`\n- **Preview:** `dataSample.csv` allows for a sneak peek before full download\n\n---\n\n## Malware Families:\n1. Ramnit  \n2. Lollipop  \n3. Kelihos_ver3  \n4. Vundo  \n5. Simda  \n6. Tracur  \n7. Kelihos_ver1  \n8. Obfuscator.ACY  \n9. Gatak  \n\n---\n\n## Model Performance Summary\n\n- **Best-fit XGBoost** → **0.0368 Log-loss on public leaderboard**\n\n---\n\n","metadata":{"id":"1-55WKas28Lr"}},{"id":"fe75cf05-4a32-471b-a527-68b9a1bbaf20","cell_type":"markdown","source":"### **GitHub Repository:** [Malware Classification – Hex Unigram Approach](https://github.com/Yashvj22/microsoft-malware-classification-using-hex-unigrams)\n\n---","metadata":{}},{"id":"B29uw6ql6G_4","cell_type":"markdown","source":"## This Notebook is on only .bytes Files\n\n---","metadata":{"id":"B29uw6ql6G_4"}},{"id":"1_XTV8aB4ejd","cell_type":"markdown","source":"## Importing Libraries","metadata":{"id":"1_XTV8aB4ejd"}},{"id":"5d02d5bf-4a1b-4702-9d03-7300718953bd","cell_type":"code","source":"import warnings\nwarnings.filterwarnings(\"ignore\")\n\nimport pandas as pd\nimport os\nimport sklearn\nimport numpy as np\nimport shutil\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.manifold import TSNE\nfrom sklearn.preprocessing import StandardScaler","metadata":{"id":"5d02d5bf-4a1b-4702-9d03-7300718953bd"},"outputs":[],"execution_count":null},{"id":"Eqsox6ULimzV","cell_type":"code","source":"from google.colab import drive\ndrive.mount('/content/drive')","metadata":{"id":"Eqsox6ULimzV","outputId":"c550c2ff-11b7-40f5-b382-b2eed47a45ee"},"outputs":[],"execution_count":null},{"id":"46v9UUOKiu7F","cell_type":"code","source":"df = pd.read_csv(\"/content/drive/MyDrive/bytes_df.csv\")","metadata":{"id":"46v9UUOKiu7F"},"outputs":[],"execution_count":null},{"id":"pjLHPViCi1wk","cell_type":"code","source":"df","metadata":{"id":"pjLHPViCi1wk","outputId":"478e9539-dc48-4283-b8fd-1f8b51375202"},"outputs":[],"execution_count":null},{"id":"f4f2e506-6e9e-4316-8f53-a645d5ee8d78","cell_type":"markdown","source":"## Preparing the Data\n\n---","metadata":{"id":"f4f2e506-6e9e-4316-8f53-a645d5ee8d78"}},{"id":"e16a2ea5-56d2-43d8-bcd7-84081e35c9fd","cell_type":"code","source":"os.makedirs(\"byteFiles\")\nos.makedirs(\"asmFiles\")\n\ndata_path = r\"E:\\Malware_Classification_ML_Project\\train\"\n\nfor file in os.listdir(data_path):\n    file_path = os.path.join(data_path, file)\n\n    if file.endswith(\".bytes\"):\n        shutil.move(file_path, os.path.join(\"byteFiles\", file))\n\n    elif file.endswith(\".asm\"):\n        shutil.move(file_path, os.path.join(\"asmFiles\", file))","metadata":{"id":"e16a2ea5-56d2-43d8-bcd7-84081e35c9fd"},"outputs":[],"execution_count":null},{"id":"e8f273fb-6028-4aea-b809-1c8fd4e8c7c2","cell_type":"markdown","source":"## In this section working with only Bytes Files","metadata":{"id":"e8f273fb-6028-4aea-b809-1c8fd4e8c7c2"}},{"id":"27f5632a-a9cc-4c97-9d1f-0d02b53205f1","cell_type":"code","source":"## Getting class_lables\n\nclass_label = pd.read_csv(r\"E:\\Malware_Classification_ML_Project\\trainLabels.csv\")\n\nclass_label","metadata":{"id":"27f5632a-a9cc-4c97-9d1f-0d02b53205f1","outputId":"e649a3a9-5ca1-4286-ece8-983d03c7a401"},"outputs":[],"execution_count":null},{"id":"3295010b-004b-44ff-b4ad-eb2157782ad3","cell_type":"markdown","source":"### B) Making File Size as a Feature\n\n+ because Malware behavior affects file size\n+ File families often have consistent size patterns : For example, Ramnit files might be around 800 KB, while Kelihos_ver1 might typically be larger or smaller","metadata":{"id":"3295010b-004b-44ff-b4ad-eb2157782ad3"}},{"id":"15cd45d6-0be5-42be-8895-18c229ff2019","cell_type":"code","source":"# Load ID and Class info from your label DataFrame\nid_to_class = dict(zip(class_label['Id'], class_label['Class']))\n\nfile_ids = []\nfile_sizes = []\nfile_classes = []\n\nfor filename in os.listdir('byteFiles'):\n    if filename.endswith('.bytes'):\n        file_id = filename.split('.')[0]\n\n        if file_id in id_to_class:\n\n            # Get file size in MB for consistent data\n            size_mb = os.path.getsize(os.path.join('byteFiles', filename)) / (1024.0 * 1024.0)\n\n            file_ids.append(file_id)\n            file_sizes.append(size_mb)\n            file_classes.append(id_to_class[file_id])\n\ndata_size_byte = pd.DataFrame({\n    'ID': file_ids,\n    'size': file_sizes,\n    'Class': file_classes\n})","metadata":{"id":"15cd45d6-0be5-42be-8895-18c229ff2019"},"outputs":[],"execution_count":null},{"id":"636db0ff-6dfa-41d7-b97e-360ee9b097b6","cell_type":"code","source":"data_size_byte","metadata":{"id":"636db0ff-6dfa-41d7-b97e-360ee9b097b6","outputId":"d4bb796e-b5d0-450b-a04e-8e3ba403697d"},"outputs":[],"execution_count":null},{"id":"5ff9f9df-b383-4b63-bebe-77c30e69ceca","cell_type":"code","source":"data_size_byte.to_csv(\"bytes_sizes_df.csv\", index=False)","metadata":{"id":"5ff9f9df-b383-4b63-bebe-77c30e69ceca"},"outputs":[],"execution_count":null},{"id":"9da3d3e1-493d-40f2-98b0-c8b07267c824","cell_type":"markdown","source":"## Simple EDA\n\n---","metadata":{"id":"9da3d3e1-493d-40f2-98b0-c8b07267c824"}},{"id":"6ff080f4-0f00-4480-985d-508c9449fc36","cell_type":"code","source":"data = pd.read_csv(\"bytes_sizes_df.csv\")\ndata.head()","metadata":{"id":"6ff080f4-0f00-4480-985d-508c9449fc36","outputId":"4b96c7de-1d81-4afc-90e5-4a29f2d95f85"},"outputs":[],"execution_count":null},{"id":"d386a285-d77a-420d-a60c-5e4eed7fc4f9","cell_type":"code","source":"data['Class'].value_counts()","metadata":{"id":"d386a285-d77a-420d-a60c-5e4eed7fc4f9","outputId":"4211188c-ea8d-4588-cb05-c7e8226f799a"},"outputs":[],"execution_count":null},{"id":"7a2c8cd8-44dc-4254-9911-57a1b5c220ca","cell_type":"markdown","source":"### Class Distribution","metadata":{"id":"7a2c8cd8-44dc-4254-9911-57a1b5c220ca"}},{"id":"2e11120c-eabd-43ca-941e-7bc57a6ac306","cell_type":"code","source":"plt.figure(figsize=(10,6))\nax = sns.countplot(data=data, x='Class')\nplt.title(\"Malware Class Distribution\")\n\n# Calculate total samples for percentage\ntotal = len(data)\n\nfor p in ax.patches:\n    count = p.get_height()\n    percentage = f'{100 * count / total:.2f}%'\n    x = p.get_x() + p.get_width() / 2\n    y = p.get_height()\n    ax.annotate(percentage, (x, y), ha='center', va='bottom', fontsize=10, color='black')\n\nplt.show()","metadata":{"id":"2e11120c-eabd-43ca-941e-7bc57a6ac306","outputId":"81bc73f8-6787-46fd-a430-8a37e451bf9c"},"outputs":[],"execution_count":null},{"id":"4d12b369-93bb-4b4e-b955-121dd34d7391","cell_type":"markdown","source":"`Insight:`\n\n+ We can see class 5 malware is very rare in our dataset\n+ Class 3 and 2 malware are having most occurances in our data","metadata":{"id":"4d12b369-93bb-4b4e-b955-121dd34d7391"}},{"id":"ef6d710a-d981-49e0-a3de-220e4b66eea4","cell_type":"markdown","source":"### Distribution of File Sizes across Class","metadata":{"id":"ef6d710a-d981-49e0-a3de-220e4b66eea4"}},{"id":"f9a4ab5e-30f4-4d6e-b429-2d9ecab68a8c","cell_type":"code","source":"plt.figure(figsize=(10, 6))\nsns.boxplot(data=data, x='Class', y='size', palette='Set2')\nplt.title(\"Distribution of File Sizes Across Malware Classes\")\nplt.xlabel(\"Malware Class\")\nplt.ylabel(\"File Size (MB)\")\nplt.tight_layout()\nplt.show()","metadata":{"id":"f9a4ab5e-30f4-4d6e-b429-2d9ecab68a8c","outputId":"65895dd2-2d1f-42c8-8565-0d01df6c424b"},"outputs":[],"execution_count":null},{"id":"505e39a2-f9fb-48a2-ada7-279cb89a5ed1","cell_type":"markdown","source":"`Insight:`\n\n+ We can conclude that Malware 1, 4 and 6 i.e. `Ramnit`, `Vundo` and `Tracur` having similar kind of spread","metadata":{"id":"505e39a2-f9fb-48a2-ada7-279cb89a5ed1"}},{"id":"15acd5e4-f909-4d5c-bae0-4a8eff019055","cell_type":"markdown","source":"### Average Size per Class","metadata":{"id":"15acd5e4-f909-4d5c-bae0-4a8eff019055"}},{"id":"4ee530fc-bb72-4559-a1b6-359d1572cf3a","cell_type":"code","source":"# Compute average size per class\navg_size_per_class = data.groupby('Class')['size'].mean().reset_index()\n\nplt.figure(figsize=(10,6))\nsns.barplot(data=avg_size_per_class, x='Class', y='size', palette='crest')\nplt.title(\"Average File Size per Malware Class\")\nplt.ylabel(\"Avg Size (MB)\")\n\nfor index, row in avg_size_per_class.iterrows():\n    plt.text(index, row['size'] + 0.1, f\"{row['size']:.2f} MB\", ha='center', fontsize=9)\n\nplt.tight_layout()\nplt.show()","metadata":{"id":"4ee530fc-bb72-4559-a1b6-359d1572cf3a","outputId":"983109fc-2144-44c8-86ad-8535a25467a2"},"outputs":[],"execution_count":null},{"id":"a0bb7d8d-e3db-42db-944f-9b3f44619a1a","cell_type":"markdown","source":"`Insight:`\n\n+ We can conclude that Malware 3 i.e. `Kelihos_ver3` having highest average size\n+ Malware 1 and 6 i.e. `Ramnit` and `Tracur` having similar average sizes which we already concluded by previous box plot\n+ Malware 8 i.e. `Obfuscator.ACY` having smallest avergae sizes","metadata":{"id":"a0bb7d8d-e3db-42db-944f-9b3f44619a1a"}},{"id":"6b6fcfc0-ee17-444e-ad75-e3a6bb039b50","cell_type":"markdown","source":"---\n## Data Pre-processing\n\n+ I have only use only uni-gram hexadecimal counts\n+ Attempted bigram extraction, but failed due to `insufficient RAM`\n+ So finally only using uni-gram","metadata":{"id":"6b6fcfc0-ee17-444e-ad75-e3a6bb039b50"}},{"id":"40509aea-7389-45d0-a5c9-c4b77a8230cb","cell_type":"code","source":"## This Code I use to Extract bytes files and get the occurance of each uni-gram hexa-decimal values in dataframe\n\nimport os\nfrom collections import Counter\nimport pandas as pd\nfrom concurrent.futures import ThreadPoolExecutor\nfrom tqdm import tqdm\n\nBYTE_FOLDER = r\"E:\\\\Malware_Classification_ML_Project\\\\Code\\byteFiles\"\nBATCH_SIZE = 500\nOUTPUT_FOLDER = \"hex_batches\"\nos.makedirs(OUTPUT_FOLDER, exist_ok=True)\n\n# Step 1: Get list of .bytes files and remove already processed\ndef get_remaining_files():\n    processed_ids = set()\n    for f in os.listdir(OUTPUT_FOLDER):\n        if f.endswith('.csv'):\n            df = pd.read_csv(os.path.join(OUTPUT_FOLDER, f), usecols=['ID'])\n            processed_ids.update(df['ID'].tolist())\n\n    all_files = [os.path.join(BYTE_FOLDER, f) for f in os.listdir(BYTE_FOLDER) if f.endswith('.bytes')]\n    return [f for f in all_files if os.path.basename(f).split('.')[0] not in processed_ids]\n\n# Step 2: Process a single file\ndef process_file(file_path):\n    file_id = os.path.basename(file_path).split('.')[0]\n    freq = Counter()\n    try:\n        with open(file_path, 'r', errors='ignore') as file:\n            for line in file:\n                parts = line.strip().split()\n                hex_values = parts[1:]  # Skip address part\n                for val in hex_values:\n                    freq[val.upper() if val != '??' else '??'] += 1\n    except Exception as e:\n        print(f\"Error in {file_id}: {e}\")\n        return None\n\n    hex_keys = [f'{i:02X}' for i in range(256)] + ['??']\n    return [file_id] + [freq.get(k, 0) for k in hex_keys]\n\n# Step 3: Process in batches and save\ndef process_in_batches(files, batch_size):\n    hex_keys = [f'{i:02X}' for i in range(256)] + ['??']\n    columns = ['ID'] + hex_keys\n\n    total_batches = (len(files) + batch_size - 1) // batch_size\n    for i in range(total_batches):\n        batch_files = files[i * batch_size: (i + 1) * batch_size]\n        print(f\"\\nProcessing Batch {i + 1}/{total_batches}... ({len(batch_files)} files)\")\n\n        results = []\n        with ThreadPoolExecutor(max_workers=3) as executor:  ## Multi threading\n            for result in tqdm(executor.map(process_file, batch_files), total=len(batch_files)):\n                if result:\n                    results.append(result)\n\n\n        batch_df = pd.DataFrame(results, columns=columns)\n        batch_df.to_csv(os.path.join(OUTPUT_FOLDER, f'batch_{i+1}.csv'), index=False)\n        print(f\"Saved batch_{i+1}.csv\")\n\n# Step 4: Combine all batch CSVs into final DataFrame\ndef combine_batches():\n    all_batches = [os.path.join(OUTPUT_FOLDER, f) for f in os.listdir(OUTPUT_FOLDER) if f.endswith('.csv')]\n    final_df = pd.concat([pd.read_csv(f) for f in all_batches], ignore_index=True)\n    final_df.to_csv(\"hex_frequency_all_files.csv\", index=False)\n    print(\"All batches combined into 'hex_frequency_all_files.csv'\")\n\nif __name__ == \"__main__\":\n    remaining_files = get_remaining_files()\n    print(f\"Remaining files to process: {len(remaining_files)}\")\n    process_in_batches(remaining_files, BATCH_SIZE)\n    combine_batches()","metadata":{"id":"40509aea-7389-45d0-a5c9-c4b77a8230cb"},"outputs":[],"execution_count":null},{"id":"46afa00a-146a-49ff-be02-55e7b84248c2","cell_type":"markdown","source":"### Getting all data into single dataframe","metadata":{"id":"46afa00a-146a-49ff-be02-55e7b84248c2"}},{"id":"774882fc-d2a2-474f-9bf5-bfb2c38d6e53","cell_type":"code","source":"# Folder containing all batch CSV files\ncsv_folder = \"hex_batches\"\n\ncsv_files = [os.path.join(csv_folder, f) for f in os.listdir(csv_folder) if f.endswith('.csv')]\n\nfinal_df = pd.concat([pd.read_csv(file) for file in csv_files], ignore_index=True)","metadata":{"id":"774882fc-d2a2-474f-9bf5-bfb2c38d6e53"},"outputs":[],"execution_count":null},{"id":"07ddf4e2-3b3e-40b5-b45e-c5a5461a9e99","cell_type":"code","source":"final_df.head()","metadata":{"id":"07ddf4e2-3b3e-40b5-b45e-c5a5461a9e99","outputId":"afb75705-b953-4cff-a8be-b4aeb8f94385"},"outputs":[],"execution_count":null},{"id":"95f2d65a-8243-4728-ac29-4c172edd8b2c","cell_type":"code","source":"df2 = pd.read_csv(\"bytes_sizes_df.csv\")\n\nbytes_df = final_df.merge(df2, on='ID', how='left')","metadata":{"id":"95f2d65a-8243-4728-ac29-4c172edd8b2c"},"outputs":[],"execution_count":null},{"id":"d41d9e0a-7f3c-486a-866c-849ed9f8dc98","cell_type":"code","source":"## Final Dataframe\n\nbytes_df","metadata":{"id":"d41d9e0a-7f3c-486a-866c-849ed9f8dc98","outputId":"ce3f9d38-1c3d-4034-cded-d99e48f4322a"},"outputs":[],"execution_count":null},{"id":"cc0b1918-3880-44b0-8501-3270abb55895","cell_type":"code","source":"bytes_df.to_csv(\"bytes_df.csv\",index=False)","metadata":{"id":"cc0b1918-3880-44b0-8501-3270abb55895"},"outputs":[],"execution_count":null},{"id":"9f0c79e6-c22f-48e2-a85d-437087ea73eb","cell_type":"markdown","source":"---\n\n## Exploratory Data Analysis","metadata":{"id":"9f0c79e6-c22f-48e2-a85d-437087ea73eb"}},{"id":"tkMzOMeHA4PC","cell_type":"code","source":"df = pd.read_csv(\"/content/drive/MyDrive/bytes_df.csv\")","metadata":{"id":"tkMzOMeHA4PC"},"outputs":[],"execution_count":null},{"id":"426fe36a-2dcf-498b-b13a-263bcccc1efd","cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\nX = df.drop(columns=['ID', 'Class'])\ny = df['Class']\n\nX_scaled = StandardScaler().fit_transform(X)\n\ntsne = TSNE(n_components=2, perplexity=30, n_iter=1000, random_state=42, n_jobs=-1)\nX_tsne = tsne.fit_transform(X_scaled)\n\ntsne_df = pd.DataFrame(X_tsne, columns=['Dim1', 'Dim2'])\ntsne_df['Class'] = y\n\nplt.figure(figsize=(10, 8))\nsns.scatterplot(data=tsne_df, x='Dim1', y='Dim2', hue='Class', palette='tab10', s=60, alpha=0.8)\nplt.title('t-SNE Visualization of Byte Frequency Features')\nplt.legend(title='Class')\nplt.show()","metadata":{"id":"426fe36a-2dcf-498b-b13a-263bcccc1efd","outputId":"0dd7ca26-0f5d-4364-8a4b-677aa21cbf47"},"outputs":[],"execution_count":null},{"id":"6bb657a8-e8f9-41bf-898d-f32102269ccc","cell_type":"markdown","source":"`Insight:`\n\n+ We can conclude that, class 3 which is `Kelihos_ver3` is some how `separable` from other classes","metadata":{"id":"6bb657a8-e8f9-41bf-898d-f32102269ccc"}},{"id":"PxXtosqndYN9","cell_type":"markdown","source":"### Compare class-wise mean frequency for each hex value","metadata":{"id":"PxXtosqndYN9"}},{"id":"kRVIe2vwdWqg","cell_type":"code","source":"# Filter out hex columns\nhex_columns = [col for col in df.columns if col not in ['ID', 'Class', 'size', '??', '00']]\n\nclasswise_sum = df.groupby('Class')[hex_columns].sum()\n\n# Extract top 3 hex values per class\nrecords = []\nfor cls in classwise_sum.index:\n    top_hexes = classwise_sum.loc[cls].sort_values(ascending=False).head(3)\n    for hex_val, freq in top_hexes.items():\n        records.append({'Class': cls, 'Hex': hex_val, 'Frequency': freq})\n\ntop3_df = pd.DataFrame(records)\n\nplt.figure(figsize=(10, 6))\nsns.barplot(data=top3_df, x='Class', y='Frequency', hue='Hex')\nplt.title(\"Top 3 Most Frequent Hex Values per Class\")\nplt.ylabel(\"Total Frequency\")\nplt.xticks(rotation=45)\nplt.tight_layout()\nplt.legend(title=\"Hex Byte\")\nplt.show()","metadata":{"id":"kRVIe2vwdWqg","outputId":"1da3a362-712d-4bc1-90ea-bb1fd0dc76df"},"outputs":[],"execution_count":null},{"id":"sZ-HUbXejRc5","cell_type":"markdown","source":"`Insight:`\n\n+ For `Kelihos_ver3` class (class 3) `01` hex is occuring which is not at all occuring in any other classes\n+ For `Lollipop` class (class2) `CC` and `02` hex are mostly occuring\n+ For `Ramnit` class (class1) `8B` hex is occuring more than other classes","metadata":{"id":"sZ-HUbXejRc5"}},{"id":"QZnLWlw0knP6","cell_type":"markdown","source":"### Excluding Majority classes which are Ramnit, Lollipop, Kelihos_ver3","metadata":{"id":"QZnLWlw0knP6"}},{"id":"RGtrAlM8k5A6","cell_type":"code","source":"hex_columns = [col for col in df.columns if col not in ['ID', 'Class', 'size', '??', '00']]\n\ndf_filtered = df[~df['Class'].isin([1, 2, 3])]\n\n# Compute class-wise sum for the filtered dataframe\nclasswise_sum = df_filtered.groupby('Class')[hex_columns].sum()\n\n# Extract top 3 hex values per class\nrecords = []\nfor cls in classwise_sum.index:\n    top_hexes = classwise_sum.loc[cls].sort_values(ascending=False).head(3)\n    for hex_val, freq in top_hexes.items():\n        records.append({'Class': cls, 'Hex': hex_val, 'Frequency': freq})\n\ntop3_df = pd.DataFrame(records)\n\nplt.figure(figsize=(10, 6))\nsns.barplot(data=top3_df, x='Class', y='Frequency', hue='Hex')\nplt.title(\"Top 3 Most Frequent Hex Values per Class (Excluding Classes 1, 2, and 3)\")\nplt.ylabel(\"Total Frequency\")\nplt.xticks(rotation=45)\nplt.tight_layout()\nplt.legend(title=\"Hex Byte\")\nplt.show()","metadata":{"id":"RGtrAlM8k5A6","outputId":"d16e2cd6-51a9-4335-edb6-f8191110231b"},"outputs":[],"execution_count":null},{"id":"o3sZk7FOllzF","cell_type":"markdown","source":"`Insight:`\n\n+ For `Tracur` class (class 6) `20` and `65` hex are occuring which is not at all occuring in any other classes for majority time\n+ For `Kelihos_ver1` class (class 7) `80` and `10` hex are mostly occuring which also not occuring in other classes for majority time\n+ `8B` hex occurs in only `Obfuscator.ACY` (class 8), `Gatak` (class 9) and `Ramint` (class 1) for most frequent","metadata":{"id":"o3sZk7FOllzF"}},{"id":"whx5kO_fo6Ot","cell_type":"markdown","source":"### For minority classes which are class 4 and 5","metadata":{"id":"whx5kO_fo6Ot"}},{"id":"KeV7KjPTopxT","cell_type":"code","source":"bhex_columns = [col for col in df.columns if col not in ['ID', 'Class', 'size', '??', '00']]\n\ndf_filtered = df[~df['Class'].isin([1, 2, 3, 6, 7, 8, 9])]\n\n# Compute class-wise sum for the filtered dataframe\nclasswise_sum = df_filtered.groupby('Class')[hex_columns].sum()\n\n# Extract top 3 hex values per class\nrecords = []\nfor cls in classwise_sum.index:\n    top_hexes = classwise_sum.loc[cls].sort_values(ascending=False).head(3)\n    for hex_val, freq in top_hexes.items():\n        records.append({'Class': cls, 'Hex': hex_val, 'Frequency': freq})\n\ntop3_df = pd.DataFrame(records)\n\nplt.figure(figsize=(10, 6))\nsns.barplot(data=top3_df, x='Class', y='Frequency', hue='Hex')\nplt.title(\"Top 3 Most Frequent Hex Values per Class (Including Classes 4 and 5)\")\nplt.ylabel(\"Total Frequency\")\nplt.xticks(rotation=45)\nplt.tight_layout()\nplt.legend(title=\"Hex Byte\")\nplt.show()","metadata":{"id":"KeV7KjPTopxT","outputId":"00088fc2-7551-4700-ccb2-453c9aa1b9f4"},"outputs":[],"execution_count":null},{"id":"TgekZD16pVz9","cell_type":"markdown","source":"\n+ `01` hex is only occuring in `Simda` (class 5) and `Kelihos_ver3` (class 3) for maximun number of times , which we saw in previous plots\n+ `EB` and `68` hex are only hex which occurs mostly in only `Vundo` (class 4)","metadata":{"id":"TgekZD16pVz9"}},{"id":"5f9ef128-87ce-4d5e-9851-5c1b1d9a0d59","cell_type":"markdown","source":"---\n\n## Data Pre-processing","metadata":{"id":"5f9ef128-87ce-4d5e-9851-5c1b1d9a0d59"}},{"id":"SPUl3cpyyt5v","cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import MinMaxScaler, LabelEncoder\nfrom xgboost import XGBClassifier\nfrom sklearn.utils import compute_class_weight\nfrom sklearn.metrics import log_loss, classification_report, confusion_matrix\nfrom sklearn.calibration import CalibratedClassifierCV\nfrom lightgbm import LGBMClassifier\nimport lightgbm as lgb","metadata":{"id":"SPUl3cpyyt5v"},"outputs":[],"execution_count":null},{"id":"oDv4bJGBdSlk","cell_type":"markdown","source":"#### Mapping 1 to 9 class lables to start them from 0 to 8\n\n+ For future use while doing `inverse_transform` with LabelEncoder()\n+ Some tools or metrics expect labels to start from 0, especially:\n  + XGBoost with multi:softprob or multi:softmax.\n  + Scikit-learn classification metrics","metadata":{"id":"oDv4bJGBdSlk"}},{"id":"sdehY2EDtaBX","cell_type":"code","source":"label_map = {\n    1: 'Ramnit',\n    2: 'Lollipop',\n    3: 'Kelihos_ver3',\n    4: 'Vundo',\n    5: 'Simda',\n    6: 'Tracur',\n    7: 'Kelihos_ver1',\n    8: 'Obfuscator.ACY',\n    9: 'Gatak'\n}\n\ndf['Class'] = df['Class'].map(label_map)","metadata":{"id":"sdehY2EDtaBX"},"outputs":[],"execution_count":null},{"id":"CgB4oOYgtbFF","cell_type":"code","source":"df['Class'].head()","metadata":{"id":"CgB4oOYgtbFF","outputId":"cbabc50b-af17-43df-801e-7e6e6e53dd71"},"outputs":[],"execution_count":null},{"id":"xBRCvrzByuDN","cell_type":"code","source":"X = df.drop(['Class', 'ID'], axis=1)  # Drop ID and target\ny = df['Class']\n\nx_train, x_test, y_train, y_test = train_test_split(\n    X, y,\n    test_size=0.2,       # 80% train, 20% test\n    stratify=y,\n    random_state=42\n)","metadata":{"id":"xBRCvrzByuDN"},"outputs":[],"execution_count":null},{"id":"mx-aRC-OuQrK","cell_type":"code","source":"le = LabelEncoder()\n\ny_train = le.fit_transform(y_train)\ny_test = le.transform(y_test)","metadata":{"id":"mx-aRC-OuQrK"},"outputs":[],"execution_count":null},{"id":"gUKnV7Uq_vyc","cell_type":"code","source":"y_train_series = pd.Series(y_train)\n\n# Compute class weights\nclasses = np.unique(y_train_series)\nweights = compute_class_weight(class_weight='balanced', classes=classes, y=y_train_series)\nclass_weights = dict(zip(classes, weights))\n\n# Map sample weights\nsample_weights = y_train_series.map(class_weights)","metadata":{"id":"gUKnV7Uq_vyc"},"outputs":[],"execution_count":null},{"id":"cr_yeN9c4etH","cell_type":"markdown","source":"---\n\n## Random XGBoost Model","metadata":{"id":"cr_yeN9c4etH"}},{"id":"H93ikYBFunRE","cell_type":"code","source":"xgb_model = XGBClassifier(\n    objective='multi:softmax',\n    num_class= len(le.classes_),\n    eval_metric='mlogloss',\n    use_label_encoder=False,\n    learning_rate=0.05,\n    n_estimators=1000,\n    max_depth=6,                    # Controls model complexity\n    subsample=0.8,                  # To prevent overfitting (80% of data used per tree)\n    colsample_bytree=0.8,           # To prevent overfitting (80% of features used per tree)\n    random_state=42,\n    verbosity=1\n)","metadata":{"id":"H93ikYBFunRE"},"outputs":[],"execution_count":null},{"id":"7dDk-eqY44ko","cell_type":"code","source":"xgb_model.fit(x_train, y_train, sample_weight=sample_weights)","metadata":{"id":"7dDk-eqY44ko","outputId":"eaad9181-90ff-4e57-a3bb-640c725ec3ae"},"outputs":[],"execution_count":null},{"id":"yJ2QL2zE5Vba","cell_type":"code","source":"y_proba = xgb_model.predict_proba(x_test)\n\nloss = log_loss(y_test, y_proba)","metadata":{"id":"yJ2QL2zE5Vba"},"outputs":[],"execution_count":null},{"id":"GH1XntXQ6Kh5","cell_type":"code","source":"loss","metadata":{"id":"GH1XntXQ6Kh5","outputId":"a5050330-65c6-4efa-8554-9c9c5eed40ac"},"outputs":[],"execution_count":null},{"id":"5LDFHYMI7-ov","cell_type":"markdown","source":"## log-loss with Random XGBoost : 0.0561","metadata":{"id":"5LDFHYMI7-ov"}},{"id":"TDmtGoDH8egM","cell_type":"code","source":"print(\"Classification Report:\\n\")\nprint(classification_report(y_test, y_pred, target_names=le.classes_))\n\ncm = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(8,6))\nsns.heatmap(cm, annot=True, fmt='d', xticklabels=le.classes_, yticklabels=le.classes_, cmap='Blues')\nplt.title(\"Confusion Matrix\")\nplt.xlabel(\"Predicted\")\nplt.ylabel(\"Actual\")\nplt.show()","metadata":{"id":"TDmtGoDH8egM","outputId":"79589329-e16c-4af8-9ff6-f89f301345ea"},"outputs":[],"execution_count":null},{"id":"UfauyAvA9lRE","cell_type":"markdown","source":"`Insight:`\n\n+ Most classes have very strong performance, particularly the large ones like `Kelihos_ver3` and `Lollipop`\n+ Class imbalance issue visible: `Simda` is clearly underrepresented or underlearned","metadata":{"id":"UfauyAvA9lRE"}},{"id":"y_gd_rmNF6YZ","cell_type":"markdown","source":"## XGBoost with CalibratedClassifierCV","metadata":{"id":"y_gd_rmNF6YZ"}},{"id":"aa2Pbe5TAWlm","cell_type":"code","source":"calibrated_model = CalibratedClassifierCV( estimator=xgb_model, method='isotonic', cv=3)\ncalibrated_model.fit(x_train, y_train)","metadata":{"id":"aa2Pbe5TAWlm","outputId":"cf561c23-a6dc-484b-ad90-4cd2b9cd8aec"},"outputs":[],"execution_count":null},{"id":"x6PgYpfqDSGB","cell_type":"code","source":"# Predict probabilities\ny_prob_calib = calibrated_model.predict_proba(x_test)","metadata":{"id":"x6PgYpfqDSGB"},"outputs":[],"execution_count":null},{"id":"Sr0r1mdRECPs","cell_type":"code","source":"log_loss(y_test, y_prob_calib)","metadata":{"id":"Sr0r1mdRECPs","outputId":"fb740f7e-8c08-4a22-a332-efdc28a26536"},"outputs":[],"execution_count":null},{"id":"KwbjNHVoFps6","cell_type":"markdown","source":"## log-loss with XGBoost with CalibratedClassifierCV: 0.0559","metadata":{"id":"KwbjNHVoFps6"}},{"id":"UCLL7Vw7GJnY","cell_type":"markdown","source":"---\n\n## XGBoost with Optuna","metadata":{"id":"UCLL7Vw7GJnY"}},{"id":"jdS8PqAix0GG","cell_type":"code","source":"from sklearn.model_selection import cross_validate\nimport optuna\nfrom sklearn.model_selection import StratifiedKFold","metadata":{"id":"jdS8PqAix0GG"},"outputs":[],"execution_count":null},{"id":"e-yYScnJ0lSc","cell_type":"code","source":"def objective(trial):\n    params = {\n        'max_depth'        : trial.suggest_int('max_depth', 3, 10),\n        'learning_rate'    : trial.suggest_float('learning_rate', 1e-3, 5e-2, log=True),\n        'n_estimators'     : trial.suggest_int('n_estimators', 200, 1000),\n        'subsample'        : trial.suggest_float('subsample', 0.6, 1.0),\n        'colsample_bytree' : trial.suggest_float('colsample_bytree', 0.6, 1.0),\n        'gamma'            : trial.suggest_float('gamma', 1e-5, 1e-1, log=True),\n        'reg_alpha'        : trial.suggest_float('reg_alpha', 1e-5, 1.0, log=True),\n        'reg_lambda'       : trial.suggest_float('reg_lambda', 1e-5, 1.0, log=True),\n        'min_child_weight' : trial.suggest_int('min_child_weight', 1, 10),\n        'tree_method'      : 'gpu_hist',\n        'predictor'        : 'gpu_predictor',\n        'use_label_encoder': False,\n        'objective'        : 'multi:softprob',\n        'num_class'        : len(np.unique(y_train)),\n        'random_state'     : 42,\n        'eval_metric'      : 'mlogloss',\n    }\n\n    sample_weights = compute_sample_weight('balanced', y=y_train)\n    skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\n    losses = []\n\n    for tr_idx, val_idx in skf.split(x_train, y_train):\n        X_tr, X_val = x_train.iloc[tr_idx], x_train.iloc[val_idx]\n        y_tr, y_val = y_train.iloc[tr_idx], y_train.iloc[val_idx]\n        sw_tr = sample_weights[tr_idx]\n\n        scaler = MinMaxScaler()\n        X_tr = scaler.fit_transform(X_tr)\n        X_val = scaler.transform(X_val)\n\n        model = XGBClassifier(**params)\n\n        model.fit(\n            X_tr,\n            y_tr,\n            sample_weight=sw_tr,\n            eval_set=[(X_val, y_val)],\n            verbose=False\n        )\n\n        y_pred = model.predict_proba(X_val)\n        losses.append(log_loss(y_val, y_pred))\n\n    trial.set_user_attr('user_attrs_train_log_loss', losses)\n    return np.mean(losses)\n","metadata":{"id":"e-yYScnJ0lSc"},"outputs":[],"execution_count":null},{"id":"g-R_WTJN2gJj","cell_type":"code","source":"study = optuna.create_study(direction=\"minimize\", sampler= optuna.samplers.TPESampler())","metadata":{"id":"g-R_WTJN2gJj","outputId":"1ce5d378-1737-48e2-9bb5-d33acb1d6b12"},"outputs":[],"execution_count":null},{"id":"DDjJbO5L2gcW","cell_type":"code","source":"study.optimize(objective,n_trials=40)","metadata":{"id":"DDjJbO5L2gcW","outputId":"5724dec2-08b8-4a24-a872-80f556be7301"},"outputs":[],"execution_count":null},{"id":"8ZGp7SQJ25tN","cell_type":"code","source":"df = study.trials_dataframe()\ndf","metadata":{"id":"8ZGp7SQJ25tN","outputId":"90c160ea-79db-40a9-ab41-41e01528e3f0"},"outputs":[],"execution_count":null},{"id":"mwXGJSPUEWXW","cell_type":"code","source":"plt.figure(figsize=(14,7))\nplt.plot(df[\"number\"],df[\"user_attrs_train_log_loss\"],label = \"Training loss\")\nplt.plot(df[\"number\"],df[\"value\"],label = \"cv loss\")\nplt.grid()\nplt.xticks(df[\"number\"],rotation=45)\nplt.legend()\nplt.show()","metadata":{"id":"mwXGJSPUEWXW","outputId":"3fdfd8fe-cb6a-4ee7-a8e7-5d8f85e443fb"},"outputs":[],"execution_count":null},{"id":"VAkiA6YtE8fs","cell_type":"code","source":"## Best-fit\n\ndf.iloc[17]","metadata":{"id":"VAkiA6YtE8fs","outputId":"738c08cf-4db9-41f0-bd1b-07b4832d7ccc"},"outputs":[],"execution_count":null},{"id":"kX56WTXzFrXV","cell_type":"code","source":"xgb_model_best_fit = XGBClassifier(\n        objective='multi:softprob',\n        num_class=len(np.unique(y_train)),\n        learning_rate=0.146953,\n        max_depth=8,\n        min_child_weight=8,\n        subsample=0.837142,\n        colsample_bytree=0.962213,\n        gamma=0.66026,\n        reg_alpha=0.70606,\n        reg_lambda=0.674609,\n        n_estimators=423,\n        use_label_encoder=False,\n        eval_metric='mlogloss',\n        random_state=42\n    )","metadata":{"id":"kX56WTXzFrXV"},"outputs":[],"execution_count":null},{"id":"zFPQIOyuGKlW","cell_type":"code","source":"xgb_model_best_fit.fit(x_train, y_train, sample_weight=sample_weights)","metadata":{"id":"zFPQIOyuGKlW","outputId":"b8e95374-4b15-46b7-b041-2539f051ef69"},"outputs":[],"execution_count":null},{"id":"WTjShFeGV401","cell_type":"code","source":"y_proba = xgb_model_best_fit.predict_proba(x_test)\n\nloss = log_loss(y_test, y_proba)\n\nloss","metadata":{"id":"WTjShFeGV401","outputId":"f9759d52-e4d2-4af5-cfa2-290ef1b34178"},"outputs":[],"execution_count":null},{"id":"aY-gSQ_kWNG7","cell_type":"markdown","source":"## Best fit XGBoost model loss : 0.0725","metadata":{"id":"aY-gSQ_kWNG7"}},{"id":"bwKUqNBOW5yt","cell_type":"markdown","source":"---\n\n## LightGBM Model\n","metadata":{"id":"bwKUqNBOW5yt"}},{"id":"Ylif0kxl6fav","cell_type":"code","source":"import lightgbm as lgb\nfrom sklearn.model_selection import StratifiedKFold\nimport optuna","metadata":{"id":"Ylif0kxl6fav"},"outputs":[],"execution_count":null},{"id":"3Bx426QS_Mtc","cell_type":"code","source":"def objective(trial):\n    param = {\n        'objective': 'multiclass',\n        'num_class': len(np.unique(y_train)),\n        'metric': 'multi_logloss',\n        'boosting_type': 'gbdt',\n        'learning_rate': trial.suggest_float(\"learning_rate\", 0.01, 0.3),\n        'max_depth': trial.suggest_int(\"max_depth\", 3, 10),\n        'num_leaves': trial.suggest_int(\"num_leaves\", 15, 200),\n        'min_child_weight': trial.suggest_int(\"min_child_weight\", 1, 10),\n        'subsample': trial.suggest_float(\"subsample\", 0.6, 1.0),\n        'colsample_bytree': trial.suggest_float(\"colsample_bytree\", 0.6, 1.0),\n        'reg_alpha': trial.suggest_float(\"reg_alpha\", 0.0, 1.0),\n        'reg_lambda': trial.suggest_float(\"reg_lambda\", 0.0, 1.0),\n        'n_estimators': trial.suggest_int(\"n_estimators\", 100, 500),\n        'verbosity': -1,\n        'random_state': 42\n    }\n\n    skf = StratifiedKFold(n_splits=3, shuffle=True, random_state=42)\n    train_losses, val_losses = [], []\n\n    for train_idx, val_idx in skf.split(x_train, y_train):\n        X_tr, X_val = x_train.iloc[train_idx], x_train.iloc[val_idx]\n        y_tr, y_val = y_train[train_idx], y_train[val_idx]\n        w_tr, w_val = sample_weights[train_idx], sample_weights[val_idx]\n\n        lgb_model = lgb.LGBMClassifier(**param)\n\n        lgb_model.fit(\n            X_tr, y_tr,\n            sample_weight=w_tr,\n            eval_set=[(X_val, y_val)],\n            eval_sample_weight=[w_val],\n            eval_metric='multi_logloss'\n        )\n\n        y_tr_pred = lgb_model.predict_proba(X_tr)\n        y_val_pred = lgb_model.predict_proba(X_val)\n\n        train_losses.append(log_loss(y_tr, y_tr_pred, sample_weight=w_tr))\n        val_losses.append(log_loss(y_val, y_val_pred, sample_weight=w_val))\n\n    trial.set_user_attr(\"train_log_loss\", np.mean(train_losses))\n    return np.mean(val_losses)\n","metadata":{"id":"3Bx426QS_Mtc"},"outputs":[],"execution_count":null},{"id":"lC8LmK2YBO0K","cell_type":"code","source":"study = optuna.create_study(direction=\"minimize\", sampler= optuna.samplers.TPESampler())","metadata":{"id":"lC8LmK2YBO0K","outputId":"c063ea49-442b-4814-d5af-4a5a0087a924"},"outputs":[],"execution_count":null},{"id":"YFz4ep2DBlm9","cell_type":"code","source":"study.optimize(objective,n_trials=100)","metadata":{"id":"YFz4ep2DBlm9","outputId":"766e7b7f-79fd-43fd-b940-c13f643bd2ee"},"outputs":[],"execution_count":null},{"id":"BNRWWj_6CxYM","cell_type":"code","source":"study.best_value","metadata":{"id":"BNRWWj_6CxYM","outputId":"ec60204b-8a9c-47a3-ebbe-964ef23e3d83"},"outputs":[],"execution_count":null},{"id":"gX8QdAHvnbwA","cell_type":"code","source":"df = study.trials_dataframe()\ndf","metadata":{"id":"gX8QdAHvnbwA","outputId":"3636a202-8e06-4f13-cf8b-89cb85c0db72"},"outputs":[],"execution_count":null},{"id":"W-XPOIzAneba","cell_type":"code","source":"plt.figure(figsize=(14,7))\nplt.plot(df[\"number\"],df[\"user_attrs_train_log_loss\"],label = \"Training loss\")\nplt.plot(df[\"number\"],df[\"value\"],label = \"cv loss\")\nplt.grid()\nplt.xticks(df[\"number\"],rotation=45)\nplt.legend()\nplt.show()","metadata":{"id":"W-XPOIzAneba","outputId":"cc23f9d9-ca9f-4b53-8eed-135664eee57f"},"outputs":[],"execution_count":null},{"id":"HSmE4CIcnplo","cell_type":"code","source":"## best fit\n\ndf.iloc[54]","metadata":{"id":"HSmE4CIcnplo","outputId":"692682bd-8c9f-491b-8f6d-dcf9ed569ba4"},"outputs":[],"execution_count":null},{"id":"XgpReiAypFB3","cell_type":"code","source":"Lgb_best_fit = LGBMClassifier(\n    colsample_bytree=0.707874,\n    learning_rate=0.233835,\n    max_depth=6,\n    min_child_weight=3,\n    n_estimators=287,\n    num_leaves=186,\n    reg_alpha=0.361441,\n    reg_lambda=0.774936,\n    subsample=0.97339,\n    random_state=42\n)","metadata":{"id":"XgpReiAypFB3"},"outputs":[],"execution_count":null},{"id":"7_ymv6R9p5vd","cell_type":"code","source":"Lgb_best_fit.fit(x_train, y_train)","metadata":{"collapsed":true,"id":"7_ymv6R9p5vd","outputId":"52cf35d3-3952-42ba-bbbd-9a74e3b2222e","jupyter":{"outputs_hidden":true}},"outputs":[],"execution_count":null},{"id":"CcL_SRRXpjei","cell_type":"code","source":"y_cap_proba = Lgb_best_fit.predict_proba(x_test)","metadata":{"id":"CcL_SRRXpjei"},"outputs":[],"execution_count":null},{"id":"3q2nJPCApxAC","cell_type":"code","source":"log_loss(y_test, y_cap_proba)","metadata":{"id":"3q2nJPCApxAC","outputId":"7a094e73-71f1-4f23-ddae-611ce967406b"},"outputs":[],"execution_count":null},{"id":"d8Zm1RNjqJRj","cell_type":"markdown","source":"### Best-Fit LGBMClassifier log-loss : 0.0590\n\n---","metadata":{"id":"d8Zm1RNjqJRj"}},{"id":"vKksxnpZsGKQ","cell_type":"code","source":"y_cap = Lgb_best_fit.predict(x_test)","metadata":{"id":"vKksxnpZsGKQ"},"outputs":[],"execution_count":null},{"id":"Au3VUAavqzHV","cell_type":"code","source":"print(\"Classification Report:\\n\")\nprint(classification_report(y_test, y_cap, target_names=le.classes_))\n\ncm = confusion_matrix(y_test, y_cap)\nplt.figure(figsize=(8,6))\nsns.heatmap(cm, annot=True, fmt='d', xticklabels=le.classes_, yticklabels=le.classes_, cmap='Blues')\nplt.title(\"Confusion Matrix\")\nplt.xlabel(\"Predicted\")\nplt.ylabel(\"Actual\")\nplt.show()","metadata":{"id":"Au3VUAavqzHV","outputId":"af288f66-6835-47b4-ad06-53654dc3ed56"},"outputs":[],"execution_count":null},{"id":"iwqnu9-pswpu","cell_type":"markdown","source":"---\n\n## Making outliers cap in File_size column to see performance change","metadata":{"id":"iwqnu9-pswpu"}},{"id":"3v5gFlQHxsOU","cell_type":"code","source":"x_train['size'] = np.log1p(x_train['size'])\nx_test['size'] = np.log1p(x_test['size'])","metadata":{"id":"3v5gFlQHxsOU"},"outputs":[],"execution_count":null},{"id":"WBIIWHckxsfj","cell_type":"code","source":"bplt.figure(figsize=(10, 6))\nsns.boxplot(data=x_train, x=y_train, y='size', palette='Set2')\nplt.title(\"Distribution of File Sizes Across Malware Classes\")\nplt.xlabel(\"Malware Class\")\nplt.ylabel(\"File Size (MB)\")\nplt.tight_layout()\nplt.show()","metadata":{"id":"WBIIWHckxsfj","outputId":"fe52e659-5dd7-4330-a794-802bf7880817"},"outputs":[],"execution_count":null},{"id":"4n_EPsbSsxJJ","cell_type":"code","source":"def objective(trial):\n    param = {\n        'objective': 'multiclass',\n        'num_class': len(np.unique(y_train)),\n        'metric': 'multi_logloss',\n        'boosting_type': 'gbdt',\n        'learning_rate': trial.suggest_float(\"learning_rate\", 0.005, 0.1),\n        'max_depth': trial.suggest_int(\"max_depth\", 3, 7),\n        'num_leaves': trial.suggest_int(\"num_leaves\", 15, 63),\n        'min_child_weight': trial.suggest_int(\"min_child_weight\", 5, 20),\n        'subsample': trial.suggest_float(\"subsample\", 0.6, 0.9),\n        'colsample_bytree': trial.suggest_float(\"colsample_bytree\", 0.6, 0.9),\n        'reg_alpha': trial.suggest_float(\"reg_alpha\", 0.1, 5.0),\n        'reg_lambda': trial.suggest_float(\"reg_lambda\", 0.1, 5.0),\n        'n_estimators': trial.suggest_int(\"n_estimators\", 300, 1000),\n        'verbosity': -1,\n        'random_state': 42\n    }\n\n    skf = StratifiedKFold(n_splits=3, shuffle=True, random_state=42)\n    train_losses, val_losses = [], []\n\n    for train_idx, val_idx in skf.split(x_train, y_train):\n        X_tr, X_val = x_train.iloc[train_idx], x_train.iloc[val_idx]\n        y_tr, y_val = y_train[train_idx], y_train[val_idx]\n        w_tr, w_val = sample_weights.iloc[train_idx], sample_weights.iloc[val_idx]\n\n        model = lgb.LGBMClassifier(**param)\n\n        model.fit(\n            X_tr, y_tr,\n            sample_weight=w_tr,\n            eval_set=[(X_val, y_val)],\n            eval_sample_weight=[w_val],\n            eval_metric='multi_logloss',\n            callbacks=[lgb.log_evaluation(0)]\n        )\n\n        y_tr_pred = model.predict_proba(X_tr)\n        y_val_pred = model.predict_proba(X_val)\n\n        train_losses.append(log_loss(y_tr, y_tr_pred, sample_weight=w_tr))\n        val_losses.append(log_loss(y_val, y_val_pred, sample_weight=w_val))\n\n    trial.set_user_attr(\"train_log_loss\", np.mean(train_losses))\n    return np.mean(val_losses)","metadata":{"id":"4n_EPsbSsxJJ"},"outputs":[],"execution_count":null},{"id":"gT8cjZEiz3g6","cell_type":"code","source":"study = optuna.create_study(direction=\"minimize\", sampler= optuna.samplers.TPESampler())","metadata":{"id":"gT8cjZEiz3g6","outputId":"0b3f7697-044d-4e3a-97b6-bd5be24de701"},"outputs":[],"execution_count":null},{"id":"hvQTWrE0z64l","cell_type":"code","source":"study.optimize(objective,n_trials=80)","metadata":{"id":"hvQTWrE0z64l","outputId":"84a5135d-6b64-40f5-888c-22a56a6281ce"},"outputs":[],"execution_count":null},{"id":"I8rQaYgn0wiz","cell_type":"code","source":"study.best_value","metadata":{"id":"I8rQaYgn0wiz","outputId":"9969da3d-aff1-437e-ef29-afbcdeb4de7b"},"outputs":[],"execution_count":null},{"id":"2-AK4XltPR9s","cell_type":"code","source":"df = study.trials_dataframe()\ndf","metadata":{"id":"2-AK4XltPR9s","outputId":"1acd6802-9681-4c91-d606-7660c1678520"},"outputs":[],"execution_count":null},{"id":"-UAVVKZFPYkK","cell_type":"code","source":"plt.figure(figsize=(14,7))\nplt.plot(df[\"number\"],df[\"user_attrs_train_log_loss\"],label = \"Training loss\")\nplt.plot(df[\"number\"],df[\"value\"],label = \"cv loss\")\nplt.grid()\nplt.xticks(df[\"number\"],rotation=45)\nplt.legend()\nplt.show()","metadata":{"id":"-UAVVKZFPYkK","outputId":"29cee597-cb48-481e-b9f8-2bca00ccf7a6"},"outputs":[],"execution_count":null},{"id":"8AOx49mCPkvU","cell_type":"code","source":"np.argmin(df[\"value\"] - df[\"user_attrs_train_log_loss\"])","metadata":{"id":"8AOx49mCPkvU","outputId":"fc8dbd14-f06f-4bae-fad4-a161ad1607f7"},"outputs":[],"execution_count":null},{"id":"UpBeQF1YQRKD","cell_type":"code","source":"df.iloc[3]","metadata":{"id":"UpBeQF1YQRKD","outputId":"93cd3bb0-2018-44d5-e9c7-3531727e7192"},"outputs":[],"execution_count":null},{"id":"FAWvcIVkQgtk","cell_type":"code","source":"Lgb_best_fit = LGBMClassifier(\n    colsample_bytree=0.825103,\n    learning_rate=0.068884,\n    max_depth=6,\n    min_child_weight=20,\n    n_estimators=329,\n    num_leaves=53,\n    reg_alpha=3.858204,\n    reg_lambda=0.155637,\n    subsample=0.719234,\n    random_state=42\n)","metadata":{"id":"FAWvcIVkQgtk"},"outputs":[],"execution_count":null},{"id":"5D_wzJv1Q7TF","cell_type":"code","source":"Lgb_best_fit.fit(x_train, y_train)","metadata":{"id":"5D_wzJv1Q7TF","outputId":"5aecdbcd-caf7-42b2-92ce-80595bfafca3"},"outputs":[],"execution_count":null},{"id":"CTuj1oaJRA53","cell_type":"code","source":"y_cap = Lgb_best_fit.predict_proba(x_test)\n\nlog_loss(y_test, y_cap)","metadata":{"id":"CTuj1oaJRA53","outputId":"198d91ef-a8ab-48f4-c51e-8a5f785ebd75"},"outputs":[],"execution_count":null},{"id":"1a3omM3sRNS3","cell_type":"markdown","source":"### So by adding np.log1p on class label, our loss gets worse than previous, so we will drop this concept","metadata":{"id":"1a3omM3sRNS3"}},{"id":"YaROOt541iLB","cell_type":"markdown","source":"---\n\n## FInal Model","metadata":{"id":"YaROOt541iLB"}},{"id":"tICQybrcRKSU","cell_type":"markdown","source":"---\n\n## Using XGBoost Best Fit model on Test Data\n\n+ because I used LGBMClassifier on test data, but final log_loss was not good","metadata":{"id":"tICQybrcRKSU"}},{"id":"nd6LNJ8PUEQ5","cell_type":"code","source":"test_df = pd.read_csv(\"/content/final_test_df.csv\")","metadata":{"id":"nd6LNJ8PUEQ5"},"outputs":[],"execution_count":null},{"id":"vdX5D4hdUi-Q","cell_type":"code","source":"test_df.drop(\"ID\", axis=1, inplace=True)","metadata":{"id":"vdX5D4hdUi-Q"},"outputs":[],"execution_count":null},{"id":"vdLkJTuwUnkK","cell_type":"code","source":"test_df","metadata":{"id":"vdLkJTuwUnkK","outputId":"0ab99dbe-4117-4346-9e05-a0a6403b0884"},"outputs":[],"execution_count":null},{"id":"9B3OteA6T6UI","cell_type":"code","source":"y_cap = xgb_model_best_fit.predict_proba(test_df)","metadata":{"id":"9B3OteA6T6UI"},"outputs":[],"execution_count":null},{"id":"8Kii8CWYUxLA","cell_type":"code","source":"pd.DataFrame(y_cap)","metadata":{"id":"8Kii8CWYUxLA","outputId":"a8c73123-1688-47df-eb03-8584d53827d0"},"outputs":[],"execution_count":null},{"id":"hvkdPcq9UyH5","cell_type":"code","source":"final_submission_df = pd.concat([pd.DataFrame(test_df[\"ID\"]), pd.DataFrame(y_cap)], axis=1)","metadata":{"id":"hvkdPcq9UyH5"},"outputs":[],"execution_count":null},{"id":"WQRXAxmMVBvk","cell_type":"code","source":"final_submission_df.columns = [\"Id\",\"Prediction1\",\"Prediction2\",\"Prediction3\",\"Prediction4\",\"Prediction5\",\"Prediction6\",\"Prediction7\",\"Prediction8\",\"Prediction9\"]","metadata":{"id":"WQRXAxmMVBvk"},"outputs":[],"execution_count":null},{"id":"oBh1DofBV7nB","cell_type":"code","source":"final_submission_df","metadata":{"id":"oBh1DofBV7nB","outputId":"18a880f3-7809-480d-a688-0bb43c940174"},"outputs":[],"execution_count":null},{"id":"6106520d-9b6c-43e5-bfcf-ecfd1f97a6f3","cell_type":"markdown","source":"## Coming Soon: `.asm` File-Based Analysis\n\nThis project focused exclusively on hexadecimal unigram-based classification due to hardware constraints.  \nIn the next phase, I plan to extract and analyze features from **assembly (.asm) files**, such as:\n\n- API call patterns  \n- Function frequencies  \n- Control flow structures  \n- Opcode sequences\n\nStay tuned for **deeper static analysis** using `.asm` metadata to enhance malware family prediction. 🚀","metadata":{}},{"id":"3c094c8a-6fd6-4ed0-9b15-0e8679f5fcbb","cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}