{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":6799,"databundleVersionId":4225553,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This notebook is my project in Deep Learning Course, ..","metadata":{}},{"cell_type":"markdown","source":"<div style=\"font-size:28px; font-weight:bold; margin:20px 0;\">\nProject \"Visualizing and Understanding Convolutional Networks\"\n</div>","metadata":{}},{"cell_type":"markdown","source":"The project is associated with the paper **“Visualizing and Understanding Convolutional Networks” (2014) by Matthew D. Zeiler and Rob Fergus**. \n\n**The task** is to write a summary of the paper’s contributions and a code reproducing the paper’s empirical results, including comments associating each part of the code to the relevant part of the paper.\n\n**The summary** is in my Medium Article: [https://medium.com/p/839667fd80e7/edit](https://medium.com/p/839667fd80e7/edit)\n\n**The code** The code is devided to 3 parts:\n\n* Part 1 - overview on the image dataset for this project - ILSVRC-2012 (ImageNet-1K)\n\n --> you are here -- https://www.kaggle.com/code/anako2020/cnn-deconvnet-part-1\n\n* Part 2 - overview on AlexNet, it's workflow and prediction\n\n  Part2 - https://www.kaggle.com/code/anako2020/cnn-deconvnet-part-2\n\n \n* Part 3 - decovnet\n\n  Part3 - https://www.kaggle.com/code/anako2020/cnn-deconvnet-part-3\n","metadata":{}},{"cell_type":"markdown","source":"# Introduction","metadata":{}},{"cell_type":"markdown","source":"Back in 2014, convolutional neural networks (CNNs) had just started making waves in computer vision. A couple of years earlier, Krizhevsky et al. (2012) had shocked the research community by winning the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) with AlexNet, cutting the top-5 classification error from 26.1% to 16.4%—a massive leap. (AlexNet's success was thanks to deeper architectures, ReLU activations, dropout, and—critically—training on GPUs.)\nBut even though CNNs were performing well, they were like a ‘black box’, no one really understood how they worked. What was each layer learning? What kind of features were being extracted? Were the models just memorizing data? \n\nIn their paper, Matthew D. Zeiler and Rob Fergus (2014), introduced a deconvolutional network (“deconvnet”)—a way to project internal activations back into pixel space so we can see what parts of the image triggered certain neurons. This visualization technique revealed the hierarchy of features learned at different layers—from edges and textures to object parts and full objects—and helped diagnose architectural problems in CNNs. \n\n","metadata":{}},{"cell_type":"markdown","source":"# Project overview\n\nSo, i aiming here to reproduce the visualization of learned pattern - as the authors did.  \n\nThe authors of the paper, used AlexNet for the visualization.\n\nAfter the visualization, they understood how to finetune, changed a few parameters and created ZFNet, that improved the AlexNet.\n\n## Training the model\n\nI am not going to train a model. By the paper, their training took around 12days on a single GTX580 GPU. Even using uptodated (jul 2025) tools like Colab Pro - it will take hours. Additionally, the ImageSet train data is about 135GB, which make it harder to operate on remoted.\n\nLuckily there is a pytorch modul that load pretrained AlexNet, here the code:\n\n```\nfrom torchvision.models import alexnet\nmodel = alexnet(weights=\"DEFAULT\").to(device)  ## ...\n```","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport os\nimport random\n\nfrom PIL import Image\nimport matplotlib.pyplot as plt\n\nimport torch\nfrom torchvision import transforms\nfrom torchvision.models import alexnet, AlexNet_Weights","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:11.190725Z","iopub.execute_input":"2025-09-19T09:00:11.1912Z","iopub.status.idle":"2025-09-19T09:00:11.196923Z","shell.execute_reply.started":"2025-09-19T09:00:11.191175Z","shell.execute_reply":"2025-09-19T09:00:11.19593Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Loading the test data\n\n\n\nIn the Article: Zeiler & Fergus, ECCV 2014 — Visualizing and Understanding Convolutional Networks, in Section 3.1 (\"Image Classification Performance\"), the authors explicitly state:\n\n“We trained our models on the ILSVRC 2012 classification dataset, which contains 1.3 million training images, 50,000 validation images and 100,000 test images.”\n\nSo, they used the full ILSVRC 2012 dataset, which includes:\n\n*  ~1.2M images for training (with labels)\n*  50K for validation (with labels)\n*  100K for test \n\n\n\nHere in this work, i will use pre-trained model to show the Decovnet process and to visualize the filters.\n\nI use here the validation set 50K that contains 1000 classes with 50 images per class. The dataset I use - thanks to the competition. It can be found on ImageNet site.","metadata":{}},{"cell_type":"markdown","source":"**a synset = a “synonym set.”**\n\nLong version (but still chill):\n\n* From **WordNet**, a synset is a group of words/phrases that mean the same concept (e.g., `{dog, domestic dog, Canis familiaris}`).\n* **ImageNet classes = WordNet synsets.** Each class is literally one synset.\n* The **ID looks like** `n01440764`:\n\n  * `n` = noun (ImageNet uses noun synsets)\n  * `01440764` = the WordNet offset (unique ID)\n* A synset has:\n\n  * **ID** (e.g., `n01440764`)\n  * **Gloss/label(s)** (human-readable names, sometimes multiple synonyms)\n  * **Hierarchy** (what it’s a kind of; useful for taxonomy)\n\nWhy you care in your workflow:\n\n* When you “pick a class,” you actually pick its **synset ID**.\n* **Train split** already has folders by synset (easy mode).\n* **Val split** needs a mapping file to link images → synset.\n* Any per-class analysis you run (activations, Grad-CAM, filter stats) keys off that **synset**.\n\nTL;DR: Think of a synset as the **canonical class handle** in ImageNet—a stable ID + a bundle of synonymous labels for one concept.\n","metadata":{}},{"cell_type":"markdown","source":"### Dataset overview\n##### Showing first 20 images from the dataset","metadata":{}},{"cell_type":"code","source":"# 1. Paths\nval_path = \"/kaggle/input/imagenet-object-localization-challenge/ILSVRC/Data/CLS-LOC/val\"\nlabels_path = \"/kaggle/input/imagenet-object-localization-challenge/LOC_val_solution.csv\"\nmapping_path = \"/kaggle/input/imagenet-object-localization-challenge/LOC_synset_mapping.txt\"\n\n# 2. Load synset-to-name mapping\nsynset_to_name = {}\nwith open(mapping_path, 'r') as f:\n    for line in f:\n        line = line.strip()\n        if not line:\n            continue\n        parts = line.split(' ', 1)\n        if len(parts) == 2:\n            synset, name = parts\n            name = name.split(',')[0]  # Take only first label\n            synset_to_name[synset] = name\n\n# 3. Load image labels (synsets)\ndf = pd.read_csv(labels_path)\ndf.columns = ['ImageId', 'Label']\ndf['Label'] = df['Label'].str.split().str[0]\n\n# 4. Dataset info printout\nnum_images = len(df)\nnum_classes = df['Label'].nunique()\nexample_synsets = df['Label'].unique()[:5]\nexample_names = [synset_to_name.get(s, s) for s in example_synsets]\n\nprint(f\"- Using the ImageNet validation set\")\nprint(f\"- Total validation images: {num_images}\")\nprint(f\"- Number of unique classes: {num_classes}\")\nprint(f\"- Example classes: {', '.join(example_names)}\")\n\n# 5. Show sample images\nsample_imgs = sorted(os.listdir(val_path))[:20]\ncols = 5\nrows = (len(sample_imgs) + cols - 1) // cols\n\nplt.figure(figsize=(15, 3 * rows))\nfor i, filename in enumerate(sample_imgs):\n    img_id = filename.split('.')[0]\n    label_row = df[df['ImageId'] == img_id]\n\n    if label_row.empty:\n        label = \"Unknown\"\n    else:\n        synset = label_row.iloc[0]['Label']\n        label = synset_to_name.get(synset, synset)\n\n    img_path = os.path.join(val_path, filename)\n    img = Image.open(img_path)\n    width, height = img.size\n\n    plt.subplot(rows, cols, i + 1)\n    plt.imshow(img)\n    plt.axis(\"off\")\n    plt.title(f\"{label}\\n{width}×{height}\", fontsize=12, fontweight='bold')\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T08:54:24.199411Z","iopub.execute_input":"2025-09-19T08:54:24.199907Z","iopub.status.idle":"2025-09-19T08:54:28.51193Z","shell.execute_reply.started":"2025-09-19T08:54:24.199871Z","shell.execute_reply":"2025-09-19T08:54:28.510569Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Showing 100 out of 1000 labels","metadata":{}},{"cell_type":"code","source":"# Step 1: Load synset → human-readable name using space split\nsynset_to_name = {}\nwith open(\"/kaggle/input/imagenet-object-localization-challenge/LOC_synset_mapping.txt\") as f:\n    for line in f:\n        parts = line.strip().split(\" \", 1)  # split at first space only\n        if len(parts) < 2:\n            continue\n        synset = parts[0]\n        name = parts[1]\n        synset_to_name[synset] = name\n\n# Step 2: Load validation labels\nimport pandas as pd\ncsv_path = \"/kaggle/input/imagenet-object-localization-challenge/LOC_val_solution.csv\"\ndf = pd.read_csv(csv_path)\ndf.columns = ['ImageId', 'Label']\ndf['Label'] = df['Label'].apply(lambda x: x.split()[0])  # just synset ID\n\n# Step 3: Get first 100 unique labels and map them\nunique_labels = df['Label'].unique()[:100]\nlabel_names = [(wnid, synset_to_name.get(wnid, \"UNKNOWN\")) for wnid in unique_labels]\n\nprint('The list of first 100 labels:')\n# Step 4: Print results\nfor i, (wnid, name) in enumerate(label_names, 1):\n    print(f\"{i:3}. {wnid} → {name}\")\nprint('The full list maybe also found on the ImageNet page https://www.image-net.org/challenges/LSVRC/2012/browse-synsets')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T08:54:28.51311Z","iopub.execute_input":"2025-09-19T08:54:28.513388Z","iopub.status.idle":"2025-09-19T08:54:28.652087Z","shell.execute_reply.started":"2025-09-19T08:54:28.513367Z","shell.execute_reply":"2025-09-19T08:54:28.650944Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Showing all 50 images of the same class","metadata":{}},{"cell_type":"code","source":"csv_path = \"/kaggle/input/imagenet-object-localization-challenge/LOC_val_solution.csv\"\ndf = pd.read_csv(csv_path)\ndf.columns = ['ImageId', 'Label']\ndf['Label'] = df['Label'].apply(lambda x: x.split()[0])  # just the synset\n\n# Extract synset for ILSVRC2012_val_00000017\ntarget_id = \"ILSVRC2012_val_00000017\"\nsynset = df[df['ImageId'] == target_id]['Label'].values[0]\nprint(\"Synset:\", synset)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T08:54:28.653355Z","iopub.execute_input":"2025-09-19T08:54:28.653859Z","iopub.status.idle":"2025-09-19T08:54:28.777178Z","shell.execute_reply.started":"2025-09-19T08:54:28.653824Z","shell.execute_reply":"2025-09-19T08:54:28.775906Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Synset for \"sea star\"\ntarget_synset = \"n01914609\"\n\n# Load and filter the CSV\ncsv_path = \"/kaggle/input/imagenet-object-localization-challenge/LOC_val_solution.csv\"\ndf = pd.read_csv(csv_path)\ndf.columns = ['ImageId', 'Label']\ndf['Label'] = df['Label'].apply(lambda x: x.split()[0])\nstar_imgs = df[df['Label'] == target_synset]['ImageId'].values\nprint(f\"Found {len(star_imgs)} sea star images\")\n\n# Number to show\nnum_to_show = 50\ncols = 5\nrows = (num_to_show + cols - 1) // cols\n\n# Path to validation images\nval_path = \"/kaggle/input/imagenet-object-localization-challenge/ILSVRC/Data/CLS-LOC/val\"\n\n# Plot in batches to save memory\nfig, axes = plt.subplots(rows, cols, figsize=(15, 3 * rows))\naxes = axes.flatten()\n\nfor ax, img_id in zip(axes, star_imgs[:num_to_show]):\n    img_path = os.path.join(val_path, img_id + \".JPEG\")\n    try:\n        with Image.open(img_path) as img:\n            img = img.resize((96, 96))  # smaller size = less memory\n            ax.imshow(img)\n    except:\n        ax.text(0.5, 0.5, \"Error loading\", ha='center', va='center')\n    ax.set_title(img_id, fontsize=8)\n    ax.axis(\"off\")\n\n# Hide unused axes\nfor ax in axes[len(star_imgs[:num_to_show]):]:\n    ax.axis(\"off\")\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T08:54:28.779493Z","iopub.execute_input":"2025-09-19T08:54:28.779789Z","iopub.status.idle":"2025-09-19T08:54:36.1924Z","shell.execute_reply.started":"2025-09-19T08:54:28.779766Z","shell.execute_reply":"2025-09-19T08:54:36.190752Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<hr>","metadata":{}},{"cell_type":"markdown","source":"# Loading pretrained AlexNet ","metadata":{}},{"cell_type":"code","source":"# 1. Paths\nval_path = \"/kaggle/input/imagenet-object-localization-challenge/ILSVRC/Data/CLS-LOC/val\"\nlabels_path = \"/kaggle/input/imagenet-object-localization-challenge/LOC_val_solution.csv\"\nmapping_path = \"/kaggle/input/imagenet-object-localization-challenge/LOC_synset_mapping.txt\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T08:54:36.194149Z","iopub.execute_input":"2025-09-19T08:54:36.194615Z","iopub.status.idle":"2025-09-19T08:54:36.203384Z","shell.execute_reply.started":"2025-09-19T08:54:36.194568Z","shell.execute_reply":"2025-09-19T08:54:36.201496Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import torch\nfrom torchvision.models import alexnet\n\n# Define device\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n# Load pretrained model\nmodel = alexnet(weights=\"DEFAULT\").to(device)  ## ...\nmodel.eval()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T08:54:36.205116Z","iopub.execute_input":"2025-09-19T08:54:36.205723Z","iopub.status.idle":"2025-09-19T08:54:38.829977Z","shell.execute_reply.started":"2025-09-19T08:54:36.205695Z","shell.execute_reply":"2025-09-19T08:54:38.828557Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"AlexNet has **5 convolutional layers**, each with a different number of filters.\n\nHere’s a breakdown of the **feature extractor part** (`model.features`) of AlexNet:\n\n| Conv Layer | Model Layer | Type      | Filters | Kernel Size | Stride | Padding |\n| -----------|------------ | --------- | ------- | ----------- | ------ |-------\n| 1          | 0           | Conv2d    | 64      | 11×11       | 4      | 2       |\n| 1          | 1           | ReLU      | —       | —           | —      | —       |\n| 1          | 2           | MaxPool2d | —       | 3×3         | 2      | —       |\n| 2          | 3           | Conv2d    | 192     | 5×5         | 1      | 2       |\n| 2          | 4           | ReLU      | —       | —           | —      | —       |\n| 2          | 5           | MaxPool2d | —       | 3×3         | 2      | —       |\n| 3          | 6           | Conv2d    | 384     | 3×3         | 1      | 1       |\n| 3          | 7           | ReLU      | —       | —           | —      | —       |\n| 4          | 8           | Conv2d    | 256     | 3×3         | 1      | 1       |\n| 4          | 9           | ReLU      | —       | —           | —      | —       |\n| 5          | 10          | Conv2d    | 256     | 3×3         | 1      | 1       |\n| 5          | 11          | ReLU      | —       | —           | —      | —       |\n| 5          | 12          | MaxPool2d | —       | 3×3         | 2      | —       |\n\n### So the convolutional **filter layers** are at:\n\n* Layer 0 (`Conv2d(3 → 64)`)\n* Layer 3 (`Conv2d(64 → 192)`)\n* Layer 6 (`Conv2d(192 → 384)`)\n* Layer 8 (`Conv2d(384 → 256)`)\n* Layer 10 (`Conv2d(256 → 256)`)\n\nThere are **5 convolutional layers total**, and each has its own set of learned filters.\n\nWe need to extract and visualize those filters ...","metadata":{}},{"cell_type":"markdown","source":"## Using the loaded Alex-Net to predict 20 random images","metadata":{}},{"cell_type":"code","source":"\n# 1. Load pretrained AlexNet\n# device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n# model = alexnet(weights=\"DEFAULT\").to(device)\n# model.eval()\n\n# 2. Load val image paths\nval_path = \"/kaggle/input/imagenet-object-localization-challenge/ILSVRC/Data/CLS-LOC/val\"\nsample_imgs = random.sample(sorted(os.listdir(val_path)), 20)\n\n# 3. Load synset-to-name mapping\nlabel_map_path = \"/kaggle/input/imagenet-object-localization-challenge/LOC_synset_mapping.txt\"\nsynset_to_label = {}\nwith open(label_map_path, \"r\") as f:\n    for line in f:\n        parts = line.strip().split(\" \", 1)\n        if len(parts) == 2:\n            synset, label = parts\n            synset_to_label[synset] = label.split(\",\")[0]\n\n# 4. Load ground truth labels\ndf_labels = pd.read_csv(\"/kaggle/input/imagenet-object-localization-challenge/LOC_val_solution.csv\")\ndf_labels.columns = ['ImageId', 'Label']\ndf_labels['Label'] = df_labels['Label'].str.split().str[0]\n\n# 5. ImageNet class index → label mapping\nidx_to_label = AlexNet_Weights.IMAGENET1K_V1.meta[\"categories\"]\n\n# 6. Define transform\ntransform = transforms.Compose([\n    transforms.Resize(256),\n    transforms.CenterCrop(224),\n    transforms.ToTensor(),\n    transforms.Normalize(mean=[0.485, 0.456, 0.406],\n                         std=[0.229, 0.224, 0.225])\n])\n\n# 7. Show 20 random images with predictions\nrows, cols = 4, 5\nplt.figure(figsize=(20, 15))\n\nfor i, fname in enumerate(sample_imgs):\n    img_id = fname.split(\".\")[0]\n    img_path = os.path.join(val_path, fname)\n    img = Image.open(img_path).convert('RGB')\n    input_tensor = transform(img).unsqueeze(0).to(device)\n\n    with torch.no_grad():\n        output = model(input_tensor)\n        pred_class = output.argmax(dim=1).item()\n\n    # Predicted label\n    predicted = idx_to_label[pred_class]\n\n    # True label\n    true_synset = df_labels[df_labels[\"ImageId\"] == img_id][\"Label\"].values[0]\n    true_label = synset_to_label.get(true_synset, true_synset)\n\n    # Plot\n    plt.subplot(rows, cols, i + 1)\n    plt.imshow(img)\n    plt.axis(\"off\")\n    plt.title(f\"Pred: {predicted}\\nTrue: {true_label}\", fontsize=12)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:27.349549Z","iopub.execute_input":"2025-09-19T09:00:27.349911Z","iopub.status.idle":"2025-09-19T09:00:32.570181Z","shell.execute_reply.started":"2025-09-19T09:00:27.349887Z","shell.execute_reply":"2025-09-19T09:00:32.568371Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Looking on the filters  \n","metadata":{}},{"cell_type":"code","source":"for name, p in model.named_parameters():\n    if name.endswith(\".weight\") and \"features\" in name:\n        print(name, tuple(p.shape))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.572383Z","iopub.execute_input":"2025-09-19T09:00:32.572845Z","iopub.status.idle":"2025-09-19T09:00:32.58149Z","shell.execute_reply.started":"2025-09-19T09:00:32.572784Z","shell.execute_reply":"2025-09-19T09:00:32.58032Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.features[0]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.582607Z","iopub.execute_input":"2025-09-19T09:00:32.583041Z","iopub.status.idle":"2025-09-19T09:00:32.612124Z","shell.execute_reply.started":"2025-09-19T09:00:32.583006Z","shell.execute_reply":"2025-09-19T09:00:32.611046Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.features[0].weight.size()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.614231Z","iopub.execute_input":"2025-09-19T09:00:32.614551Z","iopub.status.idle":"2025-09-19T09:00:32.629018Z","shell.execute_reply.started":"2025-09-19T09:00:32.614526Z","shell.execute_reply":"2025-09-19T09:00:32.627914Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.features[0].weight.shape   # (64, 3, 11, 11)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.630169Z","iopub.execute_input":"2025-09-19T09:00:32.630459Z","iopub.status.idle":"2025-09-19T09:00:32.64883Z","shell.execute_reply.started":"2025-09-19T09:00:32.630426Z","shell.execute_reply":"2025-09-19T09:00:32.647404Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.features[0].out_channels   # 64","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.650522Z","iopub.execute_input":"2025-09-19T09:00:32.650869Z","iopub.status.idle":"2025-09-19T09:00:32.66648Z","shell.execute_reply.started":"2025-09-19T09:00:32.650841Z","shell.execute_reply":"2025-09-19T09:00:32.665291Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.features[0].bias.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.667474Z","iopub.execute_input":"2025-09-19T09:00:32.667779Z","iopub.status.idle":"2025-09-19T09:00:32.684086Z","shell.execute_reply.started":"2025-09-19T09:00:32.667756Z","shell.execute_reply":"2025-09-19T09:00:32.682953Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.features[1]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.685229Z","iopub.execute_input":"2025-09-19T09:00:32.68556Z","iopub.status.idle":"2025-09-19T09:00:32.711027Z","shell.execute_reply.started":"2025-09-19T09:00:32.685531Z","shell.execute_reply":"2025-09-19T09:00:32.709668Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.features[3].weight.size()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.712233Z","iopub.execute_input":"2025-09-19T09:00:32.712472Z","iopub.status.idle":"2025-09-19T09:00:32.73514Z","shell.execute_reply.started":"2025-09-19T09:00:32.712453Z","shell.execute_reply":"2025-09-19T09:00:32.734044Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.features[6].weight.size()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.738524Z","iopub.execute_input":"2025-09-19T09:00:32.738853Z","iopub.status.idle":"2025-09-19T09:00:32.760977Z","shell.execute_reply.started":"2025-09-19T09:00:32.73883Z","shell.execute_reply":"2025-09-19T09:00:32.760026Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.features[8].weight.size()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.761923Z","iopub.execute_input":"2025-09-19T09:00:32.762191Z","iopub.status.idle":"2025-09-19T09:00:32.783976Z","shell.execute_reply.started":"2025-09-19T09:00:32.762171Z","shell.execute_reply":"2025-09-19T09:00:32.782674Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.features[10].weight.size()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.785171Z","iopub.execute_input":"2025-09-19T09:00:32.785524Z","iopub.status.idle":"2025-09-19T09:00:32.806621Z","shell.execute_reply.started":"2025-09-19T09:00:32.785493Z","shell.execute_reply":"2025-09-19T09:00:32.805604Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.features[0].kernel_size         # -> (11, 11)\nmodel.features[3].kernel_size         # -> (5, 5)\nmodel.features[6].kernel_size         # -> (3, 3)\nmodel.features[8].kernel_size         # -> (3, 3)\nmodel.features[10].kernel_size        # -> (3, 3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.807534Z","iopub.execute_input":"2025-09-19T09:00:32.80789Z","iopub.status.idle":"2025-09-19T09:00:32.828129Z","shell.execute_reply.started":"2025-09-19T09:00:32.80781Z","shell.execute_reply":"2025-09-19T09:00:32.827155Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"params_list = list(model.parameters())\nlen(params_list) ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.829202Z","iopub.execute_input":"2025-09-19T09:00:32.829473Z","iopub.status.idle":"2025-09-19T09:00:32.849095Z","shell.execute_reply.started":"2025-09-19T09:00:32.829453Z","shell.execute_reply":"2025-09-19T09:00:32.848156Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"total = sum(p.numel() for p in model.parameters())\ntrainable = sum(p.numel() for p in model.parameters() if p.requires_grad)\nprint(f\"Total: {total:,}  |  Trainable: {trainable:,}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.850104Z","iopub.execute_input":"2025-09-19T09:00:32.850362Z","iopub.status.idle":"2025-09-19T09:00:32.871956Z","shell.execute_reply.started":"2025-09-19T09:00:32.850341Z","shell.execute_reply":"2025-09-19T09:00:32.870755Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Load pretrained AlexNet\nweights = AlexNet_Weights.DEFAULT\nmodel = alexnet(weights=weights).eval()\n\n# Grab the first conv layer\nconv1 = model.features[0]\n\n# Extract weights: shape [64, 3, 11, 11]\nkernels = conv1.weight.data.clone()\n\n# Normalize each kernel for visualization\ndef normalize_kernel(k):\n    k = k - k.min()\n    k = k / k.max()\n    return k\n\n\n# Plot all 64 kernels\nfig, axes = plt.subplots(8, 8, figsize=(12, 12))\nfor i, ax in enumerate(axes.flat):\n    kernel = normalize_kernel(kernels[i])\n    # Move from [3,11,11] -> [11,11,3] for RGB plotting\n    kernel = kernel.permute(1, 2, 0).numpy()\n    ax.imshow(kernel)\n    ax.axis(\"off\")\nplt.suptitle(\"All 64 Conv1 Kernels (AlexNet)\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:32.873166Z","iopub.execute_input":"2025-09-19T09:00:32.873516Z","iopub.status.idle":"2025-09-19T09:00:35.445971Z","shell.execute_reply.started":"2025-09-19T09:00:32.873484Z","shell.execute_reply":"2025-09-19T09:00:35.444881Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import torch, torch.nn.functional as F\nfrom torchvision.utils import make_grid\n\n\n# --- utils ---\ndef save_grid(chw_list, nrow, path, norm=True, scale_each=True):\n    tiles = []\n    for t in chw_list:\n        if t.ndim == 2: t = t[None, ...]      # (H,W) -> (1,H,W)\n        if t.shape[0] == 1: t = t.repeat(3,1,1)  # grayscale -> RGB for viewing\n        tiles.append(t.cpu())\n    grid = make_grid(torch.stack(tiles,0), nrow=nrow, normalize=norm, scale_each=scale_each, padding=2)\n    plt.figure(figsize=(12,12)); plt.axis('off'); plt.imshow(grid.permute(1,2,0))\n    plt.tight_layout(); plt.savefig(path, dpi=160, bbox_inches='tight'); plt.close()\n    print(\"saved:\", path)\n\ndef full_corr2d(a, b):  # “full” cross-correlation of 2D kernels (for composing kernels)\n    ph, pw = b.shape[-2]-1, b.shape[-1]-1\n    a = F.pad(a[None,None], (pw,pw,ph,ph), mode='constant', value=0)  # 1x1xH'×W'\n    b = b[None,None]\n    y = F.conv2d(a, b)  # 1x1x(Ha+Hb-1)x(Wa+Wb-1)\n    return y[0,0]\n\n# --- 1) Conv1 filters as RGB tiles ---\nW1 = model.features[0].weight.data.clone()          # (64, 3, 11, 11) in torchvision AlexNet\n# sort by L2 norm just to put chunky filters first\nidx1 = torch.topk(W1.view(W1.size(0), -1).norm(dim=1), k=W1.size(0)).indices\nsave_grid([W1[i] for i in idx1], nrow=8, path=\"conv1_rgb_filters.png\")\n\n# --- 2) Conv2 filters projected to pixel space via Conv1 (≈ 11×11 RGB) ---\nW2 = model.features[3].weight.data.clone()          # (192, 64, 5, 5)\nk_eff = W1.size(-1) + W2.size(-1) - 1               # 11 + 5 - 1 = 15 for AlexNet\nW2_px = torch.zeros(W2.size(0), 3, k_eff, k_eff)    # (192,3,15,15)\n\n# compose: for each Conv2 filter m and RGB channel c, sum over Conv1 ch k of (W1[k,c] (*) W2[m,k])\nwith torch.no_grad():\n    for m in range(W2.size(0)):\n        for c in range(3):\n            acc = torch.zeros(k_eff, k_eff)\n            for k in range(W1.size(0)):\n                acc += full_corr2d(W1[k, c], W2[m, k])\n            W2_px[m, c] = acc\n\nidx2 = torch.topk(W2.view(W2.size(0), -1).norm(dim=1), k=min(64, W2.size(0))).indices\nsave_grid([W2_px[i] for i in idx2], nrow=8, path=\"conv2_pixelspace_top64.png\")\n\n# --- 3) Channel-slice heatmaps for any deeper conv (e.g., Conv3) ---\ndef visualize_conv_channel_slices(conv_module, out_index=0, top_slices=64, fname=\"convX_filter_slices.png\"):\n    W = conv_module.weight.data.clone()             # (out_ch, in_ch, kH, kW)\n    m = int(torch.topk(W.view(W.size(0), -1).norm(dim=1), k=1).indices) if out_index is None else out_index\n    Wm = W[m]                                       # (in_ch, kH, kW)\n    # rank input-channel slices by energy\n    idx = torch.topk(Wm.view(Wm.size(0), -1).norm(dim=1), k=min(top_slices, Wm.size(0))).indices\n    save_grid([Wm[i] for i in idx], nrow=16, path=fname)\n\n# examples:\nvisualize_conv_channel_slices(model.features[6], out_index=None, top_slices=64, fname=\"conv3_filter_slices_top64.png\")\nvisualize_conv_channel_slices(model.features[8], out_index=None, top_slices=64, fname=\"conv4_filter_slices_top64.png\")\nvisualize_conv_channel_slices(model.features[10], out_index=None, top_slices=64, fname=\"conv5_filter_slices_top64.png\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:35.447114Z","iopub.execute_input":"2025-09-19T09:00:35.447368Z","iopub.status.idle":"2025-09-19T09:00:44.51233Z","shell.execute_reply.started":"2025-09-19T09:00:35.447348Z","shell.execute_reply":"2025-09-19T09:00:44.511365Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Filters Layer 1","metadata":{}},{"cell_type":"code","source":"from IPython.display import Image, display\ndisplay(Image(filename=\"/kaggle/working/conv1_rgb_filters.png\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:44.513237Z","iopub.execute_input":"2025-09-19T09:00:44.514144Z","iopub.status.idle":"2025-09-19T09:00:44.521905Z","shell.execute_reply.started":"2025-09-19T09:00:44.514121Z","shell.execute_reply":"2025-09-19T09:00:44.520991Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Filters Layer 2","metadata":{}},{"cell_type":"code","source":"display(Image(filename=\"/kaggle/working/conv2_pixelspace_top64.png\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:44.522887Z","iopub.execute_input":"2025-09-19T09:00:44.523223Z","iopub.status.idle":"2025-09-19T09:00:44.55175Z","shell.execute_reply.started":"2025-09-19T09:00:44.523194Z","shell.execute_reply":"2025-09-19T09:00:44.550732Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Filters Layer 3","metadata":{}},{"cell_type":"code","source":"display(Image(filename=\"/kaggle/working/conv3_filter_slices_top64.png\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:44.55298Z","iopub.execute_input":"2025-09-19T09:00:44.553247Z","iopub.status.idle":"2025-09-19T09:00:44.574284Z","shell.execute_reply.started":"2025-09-19T09:00:44.553222Z","shell.execute_reply":"2025-09-19T09:00:44.573116Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from PIL import Image, ImageDraw, ImageFont\nfrom IPython.display import Image as IPyImage, display\n\npaths = [\n    \"/kaggle/working/conv3_filter_slices_top64.png\",\n    \"/kaggle/working/conv4_filter_slices_top64.png\",\n    \"/kaggle/working/conv5_filter_slices_top64.png\",\n]\nlabels = [\"Conv3 slices\", \"Conv4 slices\", \"Conv5 slices\"]\n\n# ----- sizing & layout -----\ntarget_w = 900   # <- make wider/narrower here\ngutter   = 20    # vertical space between panels\nlabel_h  = 40    # label bar height\nside_pad = 20    # left/right padding\n\n# Pillow 10+ resampling compatibility\nResampling = getattr(Image, \"Resampling\", Image)\n\n# load & resize (keep aspect)\nimgs = [Image.open(p).convert(\"RGB\") for p in paths]\nresized = []\nfor im in imgs:\n    w, h = im.size\n    scale = target_w / max(1, w)\n    nh = int(round(h * scale))\n    resized.append(im.resize((target_w, nh), Resampling.BILINEAR))\n\n# canvas size\ntotal_w = target_w + side_pad * 2\ntotal_h = sum(im.size[1] for im in resized) + (label_h * len(resized)) + gutter * (len(resized) - 1)\ncanvas = Image.new(\"RGB\", (total_w, total_h), (255, 255, 255))\ndraw = ImageDraw.Draw(canvas)\n\n# font\ntry:\n    font = ImageFont.truetype(\"/usr/share/fonts/truetype/dejavu/DejaVuSans.ttf\", 20)\nexcept Exception:\n    font = ImageFont.load_default()\n\ndef text_wh(draw_obj, text, font_obj):\n    if hasattr(draw_obj, \"textbbox\"):  # Pillow ≥10\n        l, t, r, b = draw_obj.textbbox((0, 0), text, font=font_obj)\n        return r - l, b - t\n    if hasattr(font_obj, \"getsize\"):   # fallback\n        return font_obj.getsize(text)\n    w = draw_obj.textlength(text, font=font_obj)\n    ascent, descent = font_obj.getmetrics()\n    return int(w), int(ascent + descent)\n\n# paste vertically with labels\ny = 0\nfor im, lab in zip(resized, labels):\n    # image\n    canvas.paste(im, (side_pad, y))\n    y += im.size[1]\n\n    # label centered under panel\n    tw, th = text_wh(draw, lab, font)\n    cx = side_pad + im.size[0] // 2\n    draw.text((cx - tw // 2, y + (label_h - th) // 2), lab, fill=(30, 30, 30), font=font)\n\n    y += label_h + gutter  # move down for next panel\n\nout_path = \"/kaggle/working/conv3to5_slices_stack.png\"\ncanvas.save(out_path, quality=95)\ndisplay(IPyImage(filename=out_path))\nprint(\"Saved:\", out_path)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-19T09:00:44.575686Z","iopub.execute_input":"2025-09-19T09:00:44.576114Z","iopub.status.idle":"2025-09-19T09:00:44.726519Z","shell.execute_reply.started":"2025-09-19T09:00:44.576088Z","shell.execute_reply":"2025-09-19T09:00:44.72547Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"[Next part - CNN DeconcNet part 2](https://www.kaggle.com/code/anako2020/cnn-deconvnet-part-2)\n","metadata":{}}]}