{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h2 style=\"text-align:center;font-size:200%;;\">UltraMNIST Classification Challenge</h2>","metadata":{}},{"cell_type":"markdown","source":"# Table of Contents<a id='top'></a>\n>1. [Overview](#1.-Overview)  \n>    * [Project Detail](#Project-Detail)\n>    * [Goal of this notebook](#Goal-of-this-notebook)\n>1. [Import libraries](#2.-Import-libraries)\n>1. [Load the dataset](#3.-Load-the-dataset)\n>1. [EDA](#4.-EDA)\n>1. [Preprocessing](#5.-Preprocessing)\n>1. [Modeling](#6.-Modeling)\n>    * [6.1 Randomized Random Model](#6.1-Randomized-Random-Model)\n>    * [6.2 CNN Model](#6.2-CNN-Model)","metadata":{}},{"cell_type":"markdown","source":"<a href=\"#top\" class=\"btn btn-sm active\" role=\"button\" aria-pressed=\"true\"> 🔝 Table of Contents</a>","metadata":{}},{"cell_type":"markdown","source":"# 1. Overview\n## Project Detail\n> In this project, we use [UltraMNIST dataset](https://www.kaggle.com/c/ultra-mnist/data)\n\n> Dataset comprises very large-scale images, each of 4000x4000 pixels with 3-5 digits per image. Each of these digits has been extracted from the original MNIST dataset. Your task is to predict the sum of the digits per image, and this number can be anything from 0 to 27.\n\n\n## Goal of this notebook\n>* Perform data pre-processing techniques\n>* Perform EDA techniques\n>* Perform Modeling","metadata":{}},{"cell_type":"markdown","source":"<a href=\"#top\" class=\"btn btn-sm active\" role=\"button\" aria-pressed=\"true\"> 🔝 Table of Contents</a>","metadata":{}},{"cell_type":"markdown","source":"# 2. Import libraries","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom matplotlib.image import imread\nfrom warnings import filterwarnings\n\nfilterwarnings(\"ignore\")\nplt.style.use('fivethirtyeight')\nplt.rcParams[\"figure.figsize\"] = (20, 10)","metadata":{"execution":{"iopub.status.busy":"2022-03-10T09:56:48.638877Z","iopub.execute_input":"2022-03-10T09:56:48.639838Z","iopub.status.idle":"2022-03-10T09:56:48.648336Z","shell.execute_reply.started":"2022-03-10T09:56:48.639780Z","shell.execute_reply":"2022-03-10T09:56:48.647329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a href=\"#top\" class=\"btn btn-sm active\" role=\"button\" aria-pressed=\"true\"> 🔝 Table of Contents</a>","metadata":{}},{"cell_type":"markdown","source":"# 3. Load the dataset\n> Dataset contains two directories `train` and `test` with images for the classification\n\n> Respective directories metadata is avaibale in the `train.csv` and `sample_submission.csv` files.\n\n> `train.csv` file contains `id` (image id resides in the `train` directory) and `digit_sum` (sum of 3-5 digits in the image)\n\n> `sample_submission.csv` file contains `id` (image id resides in the `test` directory) and `digit_sum` (needs to be predicted)","metadata":{}},{"cell_type":"code","source":"ROOT_PATH = \"/kaggle/input/ultra-mnist/\"\nTRAIN_PATH = ROOT_PATH + \"train/\"\nTEST_PATH = ROOT_PATH + \"test/\"\n\ndf = pd.read_csv(ROOT_PATH + \"train.csv\")\ndf_test = pd.read_csv(ROOT_PATH + \"sample_submission.csv\")\n\nprint(f'Images in the Train dataset: {df.shape}, Images in the Test dataset: {df_test.shape}')","metadata":{"execution":{"iopub.status.busy":"2022-03-10T09:56:48.650879Z","iopub.execute_input":"2022-03-10T09:56:48.651304Z","iopub.status.idle":"2022-03-10T09:56:48.732939Z","shell.execute_reply.started":"2022-03-10T09:56:48.651256Z","shell.execute_reply":"2022-03-10T09:56:48.732285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a href=\"#top\" class=\"btn btn-sm active\" role=\"button\" aria-pressed=\"true\"> 🔝 Table of Contents</a>","metadata":{}},{"cell_type":"markdown","source":"# 4. EDA","metadata":{}},{"cell_type":"code","source":"df[\"digit_sum\"].value_counts().plot(kind=\"barh\", colormap=\"gist_rainbow\")\nplt.grid(False)\nplt.title(\"Target Distribution\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-10T09:56:48.734338Z","iopub.execute_input":"2022-03-10T09:56:48.734588Z","iopub.status.idle":"2022-03-10T09:56:49.374756Z","shell.execute_reply.started":"2022-03-10T09:56:48.734559Z","shell.execute_reply":"2022-03-10T09:56:49.373691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_df = df[\"id\"].reset_index().sample(12)\n\nfor i, (_, row) in enumerate(sample_df.iterrows()):\n    plt.subplot(3, 4, i+1)\n    plt.imshow(imread(TRAIN_PATH + row[1] + \".jpeg\"))\n    plt.title(f'Sum of Digits: {df.iloc[row[0]][\"digit_sum\"]}')\n    plt.xticks([])\n    plt.yticks([])\nplt.suptitle(\"Sample Images in Train Dataset\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-10T09:56:49.379095Z","iopub.execute_input":"2022-03-10T09:56:49.379414Z","iopub.status.idle":"2022-03-10T09:57:00.150270Z","shell.execute_reply.started":"2022-03-10T09:56:49.379375Z","shell.execute_reply":"2022-03-10T09:57:00.149283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a href=\"#top\" class=\"btn btn-sm active\" role=\"button\" aria-pressed=\"true\"> 🔝 Table of Contents</a>","metadata":{}},{"cell_type":"markdown","source":"# 5. Preprocessing","metadata":{}},{"cell_type":"markdown","source":"### Clean the images.\n\nTaken the code from: https://www.kaggle.com/lukaszborecki/digit-cleaner-concept. Thanks to Lukasz. Please upvote his work.","metadata":{}},{"cell_type":"code","source":"def sector_cleaner(img):\n    img=img/255\n    for i in range(4):\n        for j in range(4):\n            border = img[1000*i:1000*(i+1),min(1000*j,3999)].sum() +img[1000*i:1000*(i+1),1000*(j+1)-1].sum() +img[min(1000*i,3999),1000*j:1000*(j+1)].sum()+img[1000*(i+1)-1,1000*j:1000*(j+1)].sum()\n            if border>=2000:\n                img[1000*i:1000*(i+1),1000*j:1000*(j+1)]=np.abs(img[1000*i:1000*(i+1),1000*j:1000*(j+1)]-1)\n    return img","metadata":{"execution":{"iopub.status.busy":"2022-03-10T09:57:00.151713Z","iopub.execute_input":"2022-03-10T09:57:00.151988Z","iopub.status.idle":"2022-03-10T09:57:00.162998Z","shell.execute_reply.started":"2022-03-10T09:57:00.151953Z","shell.execute_reply":"2022-03-10T09:57:00.161883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i, (_, row) in enumerate(sample_df.iterrows()):\n    plt.subplot(3, 4, i+1)\n    plt.imshow(sector_cleaner(imread(TRAIN_PATH + row[1] + \".jpeg\")))\n    plt.title(f'Sum of Digits: {df.iloc[row[0]][\"digit_sum\"]}')\n    plt.xticks([])\n    plt.yticks([])\nplt.suptitle(\"Sample Images in Train Dataset\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-10T09:57:00.164427Z","iopub.execute_input":"2022-03-10T09:57:00.164676Z","iopub.status.idle":"2022-03-10T09:57:13.547226Z","shell.execute_reply.started":"2022-03-10T09:57:00.164647Z","shell.execute_reply":"2022-03-10T09:57:13.546292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a href=\"#top\" class=\"btn btn-sm active\" role=\"button\" aria-pressed=\"true\"> 🔝 Table of Contents</a>","metadata":{}},{"cell_type":"markdown","source":"# 6. Modeling","metadata":{}},{"cell_type":"markdown","source":"## 6.1 Randomized Random Model","metadata":{}},{"cell_type":"code","source":"df_test['digit_sum'] = [np.random.choice(np.random.randint(0, 27, 2), 1)[0] for idx in range(28000)]\ndf_test.to_csv('submission_randomized.csv', index=False)\ndf_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-10T09:57:13.548598Z","iopub.execute_input":"2022-03-10T09:57:13.548849Z","iopub.status.idle":"2022-03-10T09:57:14.817223Z","shell.execute_reply.started":"2022-03-10T09:57:13.548819Z","shell.execute_reply":"2022-03-10T09:57:14.816133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6.2 CNN Model","metadata":{}},{"cell_type":"code","source":"import gc\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torchvision import transforms as T\nfrom torchvision.models import resnext101_32x8d\nfrom torch.utils.data import DataLoader, Dataset\n\nfrom typing import Dict, Tuple, Any\n\nfrom torchmetrics import Accuracy\nfrom pytorch_lightning.callbacks import EarlyStopping\nfrom pytorch_lightning import seed_everything, LightningModule, Trainer\n\ngc.collect();","metadata":{"execution":{"iopub.status.busy":"2022-03-10T10:00:47.774456Z","iopub.execute_input":"2022-03-10T10:00:47.774800Z","iopub.status.idle":"2022-03-10T10:00:48.004366Z","shell.execute_reply.started":"2022-03-10T10:00:47.774765Z","shell.execute_reply":"2022-03-10T10:00:48.003101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class SectorCleaner:\n    def __call__(self, img):\n        img = imread(img) / 255\n        for i in range(4):\n            for j in range(4):\n                border = img[1000*i:1000*(i+1),min(1000*j,3999)].sum() +img[1000*i:1000*(i+1),1000*(j+1)-1].sum() +img[min(1000*i,3999),1000*j:1000*(j+1)].sum()+img[1000*(i+1)-1,1000*j:1000*(j+1)].sum()\n                if border>=2000:\n                    img[1000*i:1000*(i+1),1000*j:1000*(j+1)]=np.abs(img[1000*i:1000*(i+1),1000*j:1000*(j+1)]-1)\n        return img.astype(\"uint8\")","metadata":{"execution":{"iopub.status.busy":"2022-03-10T09:57:20.104011Z","iopub.execute_input":"2022-03-10T09:57:20.104361Z","iopub.status.idle":"2022-03-10T09:57:20.116353Z","shell.execute_reply.started":"2022-03-10T09:57:20.104320Z","shell.execute_reply":"2022-03-10T09:57:20.115120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class UltraMNISTDataset(Dataset):\n    def __init__(self, df, training=True) -> None:\n        super().__init__()\n        self.df = df\n        self.image_path = TRAIN_PATH if training else TEST_PATH\n        self.transform = T.Compose([\n            SectorCleaner(),\n            T.ToPILImage(),\n            T.Resize((1000, 1000)),\n            T.ToTensor(),\n            T.Normalize((0.1307,), (0.3081,)),\n        ])\n    \n    def __len__(self) -> int:\n        return len(self.df)\n\n    def __getitem__(self, idx) -> Tuple[torch.Tensor]:\n        row = self.df.iloc[idx]\n        images = self.transform(self.image_path + row[0] + \".jpeg\")\n        labels = torch.LongTensor([row[1]])\n        return images, labels","metadata":{"execution":{"iopub.status.busy":"2022-03-10T09:57:20.119511Z","iopub.execute_input":"2022-03-10T09:57:20.119842Z","iopub.status.idle":"2022-03-10T09:57:20.141137Z","shell.execute_reply.started":"2022-03-10T09:57:20.119803Z","shell.execute_reply":"2022-03-10T09:57:20.140127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class UltraMNISTModel(LightningModule):\n    def __init__(self, *args: Any, **kwargs: Any) -> None:\n        super().__init__(*args, **kwargs)\n        # Layers\n        self.input = nn.Conv2d(1, 3, kernel_size=(7, 7), stride=(2, 2), padding=(3, 3), bias=False)\n        self.model_ft = resnext101_32x8d(pretrained=True, progress=False)\n        self.model_ft.fc = nn.Linear(self.model_ft.fc.in_features, 28)\n        self.accuracy = Accuracy()\n        # Hyperparameters\n        self.lr = 0.001\n        self.batch_size = 4\n\n    def forward(self, x: torch.Tensor) -> torch.Tensor:\n        x = self.input(x)\n        return self.model_ft(x)\n    \n    def setup(self, stage=None):\n        self.train_df = df.sample(frac=0.8, random_state=42)\n        self.valid_df = df.drop(self.train_df.index)\n\n    def configure_optimizers(self) -> torch.optim:\n        return torch.optim.Adam(self.parameters(), lr=self.lr)\n\n    def train_dataloader(self) -> DataLoader:\n        return DataLoader(\n            UltraMNISTDataset(self.train_df),\n            batch_size=self.batch_size,\n            num_workers=4,\n            shuffle=True\n            )\n\n    def training_step(self, batch, batch_nb):\n        x, y = batch\n        logits = self(x)\n        y = y.squeeze(1)\n        loss = nn.CrossEntropyLoss()(logits, y)\n        acc = self.accuracy(logits, y)\n        return {\n            'loss': loss,\n            'acc': acc,\n            'log': {\n                'train_loss': loss,\n                'train_acc': acc,\n            }\n        }\n\n    def val_dataloader(self) -> DataLoader:\n        return DataLoader(\n            UltraMNISTDataset(self.valid_df),\n            batch_size=self.batch_size,\n            num_workers=4,\n            shuffle=False\n            )\n\n    def validation_step(self, batch, batch_nb):\n        x, y = batch\n        logits = self(x)\n        y = y.squeeze(1)\n        loss = nn.CrossEntropyLoss()(logits, y)\n        acc = self.accuracy(logits, y)\n        return {\n            'val_loss': loss,\n            'val_acc': acc,\n            'log': {\n                'val_loss': loss,\n                'val_acc': acc,\n            }\n        }\n\n    def validation_epoch_end(self, outputs) -> Dict:\n        val_loss_mean = sum([o['val_loss'] for o in outputs]) / len(outputs)\n        val_acc_mean = sum([o['val_acc'] for o in outputs]) / len(outputs)\n        return {\n            'progress_bar': {\n                'val_loss': val_loss_mean,\n                'val_acc': val_acc_mean,\n            },\n            'val_loss': val_loss_mean,\n            'val_acc': val_acc_mean,\n        }","metadata":{"execution":{"iopub.status.busy":"2022-03-10T10:07:44.240793Z","iopub.execute_input":"2022-03-10T10:07:44.241487Z","iopub.status.idle":"2022-03-10T10:07:44.265122Z","shell.execute_reply.started":"2022-03-10T10:07:44.241438Z","shell.execute_reply":"2022-03-10T10:07:44.263912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"seed_everything(42)\ndevice = 'cpu'\n\nearly_stopping = EarlyStopping(monitor='val_loss', min_delta=0.00, patience=5, verbose=True)\n\nmodel = UltraMNISTModel().to(device)\n\ntrainer = Trainer(\n    max_epochs=5,\n    min_epochs=1,\n    auto_lr_find=False,\n    auto_scale_batch_size=False,\n    callbacks=[early_stopping]\n)","metadata":{"execution":{"iopub.status.busy":"2022-03-10T10:07:45.057422Z","iopub.execute_input":"2022-03-10T10:07:45.058411Z","iopub.status.idle":"2022-03-10T10:07:46.702369Z","shell.execute_reply.started":"2022-03-10T10:07:45.058364Z","shell.execute_reply":"2022-03-10T10:07:46.701395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# trainer.tune(model)","metadata":{"execution":{"iopub.status.busy":"2022-03-10T10:07:46.704378Z","iopub.execute_input":"2022-03-10T10:07:46.704643Z","iopub.status.idle":"2022-03-10T10:07:46.708805Z","shell.execute_reply.started":"2022-03-10T10:07:46.704610Z","shell.execute_reply":"2022-03-10T10:07:46.707908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trainer.fit(model)","metadata":{"execution":{"iopub.status.busy":"2022-03-10T10:07:46.710403Z","iopub.execute_input":"2022-03-10T10:07:46.710696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}