{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":4117,"databundleVersionId":46665}],"dockerImageVersionId":31328,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Data Preparation & Preprocessing\n\n## Step 1: Data Loading, Cleaning & Exploratory Data Analysis (EDA)\nBefore performing any data splitting, we must load the metadata (`trainLabels.csv`) to clean it and understand the target variable's distribution.\n\n**Data Cleaning Steps:**\n1. **Remove Duplicates:** Ensure no overlapping IDs exist.\n2. **Handle Missing Values:** Drop records without an assigned class.\n3. **Validate Target Variable:** Ensure all labels fall strictly within the 1-9 class range.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# 1. Load the labels\ndf_labels = pd.read_csv('/kaggle/input/competitions/malware-classification/trainLabels.csv')\n\n# 2. Data Cleaning\ndf_labels = df_labels.drop_duplicates(subset=['Id'])\ndf_labels = df_labels.dropna(subset=['Class'])\nvalid_classes = range(1, 10)\ndf_labels = df_labels[df_labels['Class'].isin(valid_classes)]\n\nprint(f\"Total clean samples: {len(df_labels)}\")\n\n# 3. Exploratory Data Analysis (EDA) - Checking Class Distribution\nclass_counts = df_labels['Class'].value_counts().sort_index()\nprint(\"\\nClass Distribution:\")\nprint(class_counts)\n\n# Plotting the distribution \nplt.figure(figsize=(10, 5))\nsns.barplot(x=class_counts.index, y=class_counts.values, hue=class_counts.index, palette='viridis', legend=False)\nplt.ylabel('Number of Samples')\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 2: Data Splitting \n**70% Train, 15% Validation, 15% Test**\n\n**Handling Imbalanced Data:**\nAs clearly seen in the EDA step above, our dataset is highly imbalanced (e.g., Class 5 has only 42 samples compared to thousands in other classes). \n\nTo prevent the minority classes from being completely wiped out in the Validation or Test sets, we MUST use a **Stratified Split**. This ensures the proportion of each class is strictly maintained across all data subsets. We also store the IDs in Hash Sets for O(1) lookup time during the extraction phase.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\n# 1. First split: 70% Train, 30% Temporary (Val + Test)\ntrain_df, temp_df = train_test_split(\n    df_labels, \n    test_size=0.3, \n    stratify=df_labels['Class'], # Justified by EDA\n    random_state=42\n)\n\n# 2. Second split: Divide the 30% Temporary into 15% Validation and 15% Test\nval_df, test_df = train_test_split(\n    temp_df, \n    test_size=0.5, \n    stratify=temp_df['Class'],   # Justified by EDA\n    random_state=42\n)\n\n# 3. Create Hash Sets for blazing-fast ID lookups later\ntrain_ids = set(train_df['Id'])\nval_ids = set(val_df['Id'])\ntest_ids = set(test_df['Id'])\n\n# Print the results to verify the split\nprint(f\"Train set: {len(train_ids)} samples\")\nprint(f\"Validation set: {len(val_ids)} samples\")\nprint(f\"Test set: {len(test_ids)} samples\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 3.1: Define Core Extraction Functions\nParses Hex strings, converts them to numerical pixels, reshapes them into a square matrix, and resizes them to 224x224 using PIL (Bilinear Interpolation) for ResNet consumption.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nfrom PIL import Image\n\n# ==========================================\n# HELPER FUNCTIONS (FULLY IMPLEMENTED)\n# ==========================================\n\ndef convert_hex_to_image(raw_content):\n    \"\"\"\n    Converts raw Hex content from a .bytes file into a 224x224 Numpy image array.\n    Filters out memory addresses and handles unreadable '??' bytes.\n    \"\"\"\n    lines = raw_content.split('\\n')\n    byte_list = []\n    \n    for line in lines:\n        if len(line) < 10:\n            continue\n        # Skip the first 8 characters (memory address) and extract hex values\n        hex_str = line[9:]\n        parts = hex_str.split()\n        \n        for p in parts:\n            if p == '??':\n                byte_list.append(0) # Unreadable bytes are rendered as black pixels\n            else:\n                try:\n                    byte_list.append(int(p, 16))\n                except ValueError:\n                    pass\n\n    # Failsafe for empty or fully corrupted extraction\n    if len(byte_list) == 0:\n        return np.zeros((224, 224), dtype=np.uint8)\n\n    arr = np.array(byte_list, dtype=np.uint8)\n\n    # Calculate the dimensions for a square image\n    length = len(arr)\n    width = int(np.ceil(length ** 0.5))\n    padded_length = width * width\n\n    # Zero-pad the array to form a perfect square\n    padded_arr = np.pad(arr, (0, padded_length - length), mode='constant')\n    img_matrix = padded_arr.reshape((width, width))\n\n    # Resize to the standard ResNet input size (224x224) using Nearest Neighbor\n    img = Image.fromarray(img_matrix)\n    img_resized = img.resize((224, 224), Image.BILINEAR) \n\n    return np.array(img_resized)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 3.2: The Image Extraction Pipeline\nTo avoid Out-of-Memory (OOM) and disk space errors when handling the massive archive, we use a **Temp-Disk Buffer Strategy**. Since we are strictly preparing data for a Deep Learning model (ResNet50), we optimize the process by **only extracting `.bytes` files** and completely ignoring the `.asm` files. This drastically reduces extraction time and halves the temporary storage requirements.\n\n**Late Filtering & Directory Structuring:**\n* **Corrupt File Handling:** During the extraction process, we check for `os.path.getsize(path) > 0`. Any corrupted zero-byte files generated during the Kaggle archival process are automatically identified and safely skipped to prevent pipeline failures.\n* **Ready-to-Train Format:** The processed hex-to-image arrays are directly saved into a structured hierarchy (`image_data/split/class_label/id.png`). This specific folder layout is natively supported by modern Deep Learning frameworks (such as PyTorch's `ImageFolder` or TensorFlow's `image_dataset_from_directory`), making the transition to the training phase seamless.","metadata":{}},{"cell_type":"code","source":"import os\nimport shutil \nimport gc\nimport subprocess\n\nARCHIVE_PATH = '/kaggle/input/competitions/malware-classification/train.7z'\nBASE_DATA_DIR = '/kaggle/working/image_data' # Updated directory name for clarity\nTEMP_DIR = '/kaggle/working/temp_buffer'\nCHUNK_SIZE = 200 # Process in batches to manage Disk Space\n\n# --- Map Setup ---\nlabel_mapping = dict(zip(df_labels['Id'], df_labels['Class']))\nid_list = list(df_labels['Id'])\n\nprint(f\"Starting Image Extraction Pipeline for {len(id_list)} samples...\")\n\nfor i in range(0, len(id_list), CHUNK_SIZE):\n    chunk_ids = id_list[i : i + CHUNK_SIZE]\n    os.makedirs(TEMP_DIR, exist_ok=True)\n    \n    # 1. Create a list of files to extract (ONLY .bytes files to save time and space)\n    extract_list_path = '/kaggle/working/extract_list.txt'\n    with open(extract_list_path, 'w') as f:\n        for file_id in chunk_ids:\n            f.write(f\"train/{file_id}.bytes\\n\")\n            \n    # 2. Extract using Native OS 7-Zip\n    cmd = [\"7z\", \"e\", ARCHIVE_PATH, f\"-o{TEMP_DIR}\", f\"@{extract_list_path}\", \"-y\", \"-bsp0\", \"-bso0\"]\n    subprocess.run(cmd, check=True) \n    \n    # 3. Process the extracted chunk (Image Generation ONLY)\n    for file_id in chunk_ids:\n        class_label = label_mapping[file_id]\n        \n        # Determine split target directory based on the hash sets from Step 2\n        if file_id in train_ids:   split_target = 'train'\n        elif file_id in val_ids:   split_target = 'val'\n        elif file_id in test_ids:  split_target = 'test'\n        else: continue\n            \n        bytes_path = os.path.join(TEMP_DIR, f\"{file_id}.bytes\")\n        \n        # Late Filtering: Ensure file exists and is not a corrupted 0-byte file\n        if os.path.exists(bytes_path) and os.path.getsize(bytes_path) > 0:\n            with open(bytes_path, 'r', encoding='utf-8', errors='ignore') as f_bytes:\n                b_content = f_bytes.read()\n                \n            # Convert Hex to Image array\n            img_array = convert_hex_to_image(b_content)\n            img = Image.fromarray(img_array).convert('RGB')   \n            \n            # Define save path structure: /kaggle/working/image_data/train/1/id.png\n            save_path = os.path.join(BASE_DATA_DIR, split_target, str(class_label), f\"{file_id}.png\")\n            \n            # Create directories if they don't exist and save the image\n            os.makedirs(os.path.dirname(save_path), exist_ok=True)\n            img.save(save_path)\n            \n    # 4. Clean up disk and RAM for the next chunk\n    shutil.rmtree(TEMP_DIR)\n    os.remove(extract_list_path)\n    gc.collect()\n    \n    print(f\"Processed chunk {min(i + CHUNK_SIZE, len(id_list))} / {len(id_list)} files.\")\n\nprint(\"Image Generation Pipeline Fully Completed!\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import shutil\n\nprint(\"Zipping the final image dataset to prevent Kaggle output file limits...\")\nshutil.make_archive('/kaggle/working/malware_images', 'zip', '/kaggle/working/image_data')\n\nshutil.rmtree('/kaggle/working/image_data')\n\nprint(\"Dataset is zipped!\")","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}