{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":4117,"databundleVersionId":46665}],"dockerImageVersionId":31328,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Data Preparation & Preprocessing\n\n## Step 1: Data Loading, Cleaning & Exploratory Data Analysis (EDA)\nBefore performing any data splitting, we must load the metadata (`trainLabels.csv`) to clean it and understand the target variable's distribution.\n\n**Data Cleaning Steps:**\n1. **Remove Duplicates:** Ensure no overlapping IDs exist.\n2. **Handle Missing Values:** Drop records without an assigned class.\n3. **Validate Target Variable:** Ensure all labels fall strictly within the 1-9 class range.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# 1. Load the labels\ndf_labels = pd.read_csv('/kaggle/input/competitions/malware-classification/trainLabels.csv')\n\n# 2. Data Cleaning\ndf_labels = df_labels.drop_duplicates(subset=['Id'])\ndf_labels = df_labels.dropna(subset=['Class'])\nvalid_classes = range(1, 10)\ndf_labels = df_labels[df_labels['Class'].isin(valid_classes)]\n\nprint(f\"Total clean samples: {len(df_labels)}\")\n\n# 3. Exploratory Data Analysis (EDA) - Checking Class Distribution\nclass_counts = df_labels['Class'].value_counts().sort_index()\nprint(\"\\nClass Distribution:\")\nprint(class_counts)\n\n# Plotting the distribution \nplt.figure(figsize=(10, 5))\nsns.barplot(x=class_counts.index, y=class_counts.values, hue=class_counts.index, palette='viridis', legend=False)\nplt.ylabel('Number of Samples')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-10T16:24:13.790713Z","iopub.execute_input":"2026-04-10T16:24:13.791437Z","iopub.status.idle":"2026-04-10T16:24:14.068040Z","shell.execute_reply.started":"2026-04-10T16:24:13.791397Z","shell.execute_reply":"2026-04-10T16:24:14.066296Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 2: Data Splitting \n**70% Train, 15% Validation, 15% Test**\n\n**Handling Imbalanced Data:**\nAs clearly seen in the EDA step above, our dataset is highly imbalanced (e.g., Class 5 has only 42 samples compared to thousands in other classes). \n\nTo prevent the minority classes from being completely wiped out in the Validation or Test sets, we MUST use a **Stratified Split**. This ensures the proportion of each class is strictly maintained across all data subsets. We also store the IDs in Hash Sets for O(1) lookup time during the extraction phase.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\n# 1. First split: 70% Train, 30% Temporary (Val + Test)\ntrain_df, temp_df = train_test_split(\n    df_labels, \n    test_size=0.3, \n    stratify=df_labels['Class'], # Justified by EDA\n    random_state=42\n)\n\n# 2. Second split: Divide the 30% Temporary into 15% Validation and 15% Test\nval_df, test_df = train_test_split(\n    temp_df, \n    test_size=0.5, \n    stratify=temp_df['Class'],   # Justified by EDA\n    random_state=42\n)\n\n# 3. Create Hash Sets for blazing-fast ID lookups later\ntrain_ids = set(train_df['Id'])\nval_ids = set(val_df['Id'])\ntest_ids = set(test_df['Id'])\n\n# Print the results to verify the split\nprint(f\"Train set: {len(train_ids)} samples\")\nprint(f\"Validation set: {len(val_ids)} samples\")\nprint(f\"Test set: {len(test_ids)} samples\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-10T16:24:46.146629Z","iopub.execute_input":"2026-04-10T16:24:46.146963Z","iopub.status.idle":"2026-04-10T16:24:46.485405Z","shell.execute_reply.started":"2026-04-10T16:24:46.146932Z","shell.execute_reply":"2026-04-10T16:24:46.484300Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 3.1: Define Core Extraction Functions\nWe isolate the core extraction logic into helper functions. The dummy placeholders have been replaced with the actual implementation:\n* **Deep Learning Pipeline:** Parses Hex strings, converts them to numerical pixels, reshapes them into a square matrix, and resizes them to 224x224 using PIL for ResNet consumption.\n* **Machine Learning Pipeline:** Extracts 18 tabular features (3 from `.bytes` focusing on file size and obfuscation, and 15 from `.asm` tracking Opcode and Section frequencies).","metadata":{}},{"cell_type":"code","source":"import numpy as np\nfrom PIL import Image\n\n# ==========================================\n# HELPER FUNCTIONS (FULLY IMPLEMENTED)\n# ==========================================\n\ndef convert_hex_to_image(raw_content):\n    \"\"\"\n    Converts raw Hex content from a .bytes file into a 224x224 Numpy image array.\n    Filters out memory addresses and handles unreadable '??' bytes.\n    \"\"\"\n    lines = raw_content.split('\\n')\n    byte_list = []\n    \n    for line in lines:\n        if len(line) < 10:\n            continue\n        # Skip the first 8 characters (memory address) and extract hex values\n        hex_str = line[9:]\n        parts = hex_str.split()\n        \n        for p in parts:\n            if p == '??':\n                byte_list.append(0) # Unreadable bytes are rendered as black pixels\n            else:\n                try:\n                    byte_list.append(int(p, 16))\n                except ValueError:\n                    pass\n\n    # Failsafe for empty or fully corrupted extraction\n    if len(byte_list) == 0:\n        return np.zeros((224, 224), dtype=np.uint8)\n\n    arr = np.array(byte_list, dtype=np.uint8)\n\n    # Calculate the dimensions for a square image\n    length = len(arr)\n    width = int(np.ceil(length ** 0.5))\n    padded_length = width * width\n\n    # Zero-pad the array to form a perfect square\n    padded_arr = np.pad(arr, (0, padded_length - length), mode='constant')\n    img_matrix = padded_arr.reshape((width, width))\n\n    # Resize to the standard ResNet input size (224x224) using Nearest Neighbor\n    img = Image.fromarray(img_matrix, mode='L')\n    img_resized = img.resize((224, 224), Image.NEAREST) \n\n    return np.array(img_resized)\n\ndef extract_bytes_features(raw_content):\n    \"\"\"\n    Extracts basic numerical features from a .bytes file.\n    Returns a list of 3 features.\n    \"\"\"\n    features = []\n    features.append(len(raw_content))        # Feature 1: Total file length\n    features.append(raw_content.count('??')) # Feature 2: Frequency of obfuscated bytes\n    features.append(raw_content.count('00')) # Feature 3: Frequency of null bytes\n    return features\n\ndef extract_asm_features(raw_content):\n    \"\"\"\n    Extracts Opcode and Section frequencies from a .asm file.\n    Returns a list of 15 features.\n    \"\"\"\n    features = []\n    content_lower = raw_content.lower()\n    \n    # 1. Frequency of common operational codes (Opcodes)\n    opcodes = ['jmp', 'mov', 'retf', 'push', 'pop', 'call', 'sub', 'add', 'dec', 'inc']\n    for op in opcodes:\n        features.append(content_lower.count(op))\n        \n    # 2. Frequency of executable sections\n    sections = ['.text', '.data', '.rdata', '.bss', '.idata']\n    for sec in sections:\n        features.append(content_lower.count(sec))\n        \n    return features","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-10T16:27:23.652550Z","iopub.execute_input":"2026-04-10T16:27:23.652996Z","iopub.status.idle":"2026-04-10T16:27:23.667764Z","shell.execute_reply.started":"2026-04-10T16:27:23.652962Z","shell.execute_reply":"2026-04-10T16:27:23.666272Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 3.2: The Multi-Task Processing Pipeline\nTo avoid Out-of-Memory (OOM) and disk space errors when handling the massive archive, we use a **Temp-Disk Buffer Strategy**. \n\n**Late Filtering Integration:** During the extraction process, we check for `os.path.getsize(path) > 0`. Any corrupted zero-byte files generated during the Kaggle archival process are automatically identified and safely dropped (zero-padded) to prevent pipeline failures.","metadata":{}},{"cell_type":"code","source":"import os\nimport shutil \nimport gc\nimport csv\nimport subprocess\n\nARCHIVE_PATH = '/kaggle/input/competitions/malware-classification/train.7z'\nBASE_DATA_DIR = '/kaggle/working/data'\nTEMP_DIR = '/kaggle/working/temp_buffer'\nCHUNK_SIZE = 200 # Process in batches to manage Disk Space\n\n# --- Map Setup ---\nlabel_mapping = dict(zip(df_labels['Id'], df_labels['Class']))\nid_list = list(df_labels['Id'])\n\n# --- CSV Initialization ---\ncsv_paths = {\n    'train': '/kaggle/working/rf_xgb_train.csv',\n    'val': '/kaggle/working/rf_xgb_val.csv',\n    'test': '/kaggle/working/rf_xgb_test.csv'\n}\n\nheader_names = [\n    'ID', 'Size_Bytes', 'Byte_??', 'Byte_00',\n    'op_jmp', 'op_mov', 'op_retf', 'op_push', 'op_pop', 'op_call', 'op_sub', 'op_add', 'op_dec', 'op_inc',\n    'sec_text', 'sec_data', 'sec_rdata', 'sec_bss', 'sec_idata', 'Class'\n]\n\nfor path in csv_paths.values():\n    with open(path, 'w', newline='') as f:\n        writer = csv.writer(f)\n        writer.writerow(header_names)\n\nprint(f\"Starting High-Speed Pipeline for {len(id_list)} samples...\")\n\nfor i in range(0, len(id_list), CHUNK_SIZE):\n    chunk_ids = id_list[i : i + CHUNK_SIZE]\n    os.makedirs(TEMP_DIR, exist_ok=True)\n    \n    # 1. Create a list of files to extract\n    extract_list_path = '/kaggle/working/extract_list.txt'\n    with open(extract_list_path, 'w') as f:\n        for file_id in chunk_ids:\n            f.write(f\"train/{file_id}.bytes\\n\")\n            f.write(f\"train/{file_id}.asm\\n\")\n            \n    # 2. Extract using Native OS 7-Zip\n    cmd = [\"7z\", \"e\", ARCHIVE_PATH, f\"-o{TEMP_DIR}\", f\"@{extract_list_path}\", \"-y\", \"-bsp0\", \"-bso0\"]\n    subprocess.run(cmd, check=True) \n    \n    # 3. Process the extracted chunk\n    with open(csv_paths['train'], 'a', newline='') as f_train, \\\n         open(csv_paths['val'], 'a', newline='') as f_val, \\\n         open(csv_paths['test'], 'a', newline='') as f_test:\n        \n        writers = {\n            'train': csv.writer(f_train),\n            'val': csv.writer(f_val),\n            'test': csv.writer(f_test)\n        }\n        \n        for file_id in chunk_ids:\n            class_label = label_mapping[file_id]\n            \n            # Determine split\n            if file_id in train_ids:   split_target = 'train'\n            elif file_id in val_ids:   split_target = 'val'\n            elif file_id in test_ids:  split_target = 'test'\n            else: continue\n                \n            bytes_path = os.path.join(TEMP_DIR, f\"{file_id}.bytes\")\n            asm_path = os.path.join(TEMP_DIR, f\"{file_id}.asm\")\n            \n            # Initialize default feature vectors reflecting the new sizes (3 and 15)\n            bytes_features = [0] * 3\n            asm_features = [0] * 15\n            \n            # --- Branch 1: Deep Learning (Images) & ML Features (.bytes) ---\n            # Late Filtering: Ensure file exists and is not a corrupted 0-byte file\n            if os.path.exists(bytes_path) and os.path.getsize(bytes_path) > 0:\n                with open(bytes_path, 'r', encoding='utf-8', errors='ignore') as f_bytes:\n                    b_content = f_bytes.read()\n                    \n                # Save Image for ResNet\n                img_array = convert_hex_to_image(b_content)\n                img = Image.fromarray(img_array, mode='L')\n                save_path = os.path.join(BASE_DATA_DIR, split_target, str(class_label), f\"{file_id}.png\")\n                os.makedirs(os.path.dirname(save_path), exist_ok=True)\n                img.save(save_path)\n                \n                # Extract tabular features\n                bytes_features = extract_bytes_features(b_content)\n            \n            # --- Branch 2: ML Features (.asm) ---\n            if os.path.exists(asm_path) and os.path.getsize(asm_path) > 0:\n                with open(asm_path, 'r', encoding='utf-8', errors='ignore') as f_asm:\n                    a_content = f_asm.read()\n                asm_features = extract_asm_features(a_content)\n            \n            # --- Stream Record to CSV ---\n            # Row structure: [ID, f1..f3, f4..f18, Class_Label]\n            csv_row = [file_id] + bytes_features + asm_features + [class_label]\n            writers[split_target].writerow(csv_row)\n            \n    # 4. Clean up disk and RAM for the next chunk\n    shutil.rmtree(TEMP_DIR)\n    os.remove(extract_list_path)\n    gc.collect()\n    \n    print(f\"Processed chunk {min(i + CHUNK_SIZE, len(id_list))} / {len(id_list)} files.\")\n\nprint(\"Pipeline Fully Completed!\")","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}