{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# UTMIST Workshop #3: Histopathologic Cancer Detection\nHello UTMIST community!\n\nIn this workshop, we will be building a model that will compete at a Histopathologic Cancer Detection Competition on Kaggle, which can be found here https://www.kaggle.com/competitions/histopathologic-cancer-detection.\nAs stated on their website, our algorithm will identify metastatic cancer in small image patches taken from larger digital pathology scans. The data for this competition is a slightly modified version of the PatchCamelyon (PCam) benchmark dataset. In this dataset, we are provided with a large number of small pathology images to classify. Files are named with an image id. The train_labels.csv file provides the ground truth for the images in the train folder. You are predicting the labels for the images in the test folder. A positive label indicates that the center 32x32px region of a patch contains at least one pixel of tumor tissue. Tumor tissue in the outer region of the patch does not influence the label. This outer region is provided to enable fully-convolutional models that do not use zero-padding, to ensure consistent behavior when applied to a whole-slide image.\n\n\nAuthors: Berke Altiparmak (altiparmak.berke@gmail.com), Lindy Zhai (lindy.zhai@mail.utoronto.ca)\n\n\nReferences:\n* https://www.kaggle.com/code/bonhart/pytorch-cnn-from-scratch for the model.\n* https://www.kaggle.com/code/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93 for data visualization.","metadata":{}},{"cell_type":"markdown","source":"# Step 1: Importing libraries\n\nWe will import the libraries that we need to build our CNN model. Keep in mind that those libraries are not specific for this project; we are keep using the same libraries, so it is useful for you to get familiar with them.","metadata":{}},{"cell_type":"code","source":"import os\nfrom glob import glob \nimport time\nimport numpy as np  # for math operations, one of the most commonly used libraries\nimport pandas as pd  # for handling data and data frames, another essential library\nimport cv2  # OpenCV library, which we will use to \"read\" images and transform them\nimport matplotlib.pyplot as plt  # to visualize data\nfrom tqdm import tqdm_notebook,trange  # to see the progress\nimport gc #garbage collection to save RAM\n\n# the output of plotting commands is displayed inline within Jupyter notebook:\n%matplotlib inline  \n\nfrom sklearn.model_selection import train_test_split  # to split train and validation data\n\n# PyTorch libraries to build a Machine Learning Model:\nimport torch \nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torchvision\nimport torchvision.transforms as transforms\nfrom torch.utils.data import TensorDataset, DataLoader, Dataset","metadata":{"_uuid":"b0b625dddcf44cd8db7690bf380c6cb25f555cbd","execution":{"iopub.status.busy":"2022-11-19T18:26:11.236712Z","iopub.execute_input":"2022-11-19T18:26:11.237001Z","iopub.status.idle":"2022-11-19T18:26:13.356480Z","shell.execute_reply.started":"2022-11-19T18:26:11.236944Z","shell.execute_reply":"2022-11-19T18:26:13.355656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 2: Understanding and Visualizing The Data\n\nIn this section, we will first get our training, validation, and test data, and then we will visualize some of the data we have. Moving on, we'll take a deeper look into our data by analyzing the positive and negative data distribution, and then we will compare the properties of positive samples and negative samples to see if it is actually possible to tell a positive and negative sample apart just from an image.","metadata":{}},{"cell_type":"code","source":"#set paths to training and test data\npath = \"../input/\"\ntrain_path = path + \"train/\"\ntest_path = path + \"test/\"\nlabels = pd.read_csv(path + \"train_labels.csv\")\nsubmission = pd.read_csv(path + \"sample_submission.csv\")\n\n\n#Splitting data into train and val\ntrain, val = train_test_split(labels, stratify=labels.label, test_size = 0.1)\n\ndf = pd.DataFrame({'path': glob(os.path.join(train_path,'*.tif'))}) # load the filenames\ndf['id'] = df.path.map(lambda x: x.split('/')[3].split(\".\")[0]) # keep only the file names in 'id'\ndf = df.merge(labels, on = \"id\") # merge labels and filepaths\n\ndf.head(10) # print the first ten entries","metadata":{"execution":{"iopub.status.busy":"2022-11-19T18:34:08.749111Z","iopub.execute_input":"2022-11-19T18:34:08.749414Z","iopub.status.idle":"2022-11-19T18:34:10.258811Z","shell.execute_reply.started":"2022-11-19T18:34:08.749365Z","shell.execute_reply":"2022-11-19T18:34:10.257938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_data(N,df):\n    \"\"\" This functions loads N images using the data df\n    \"\"\"\n    # allocate a numpy array for the images (N, 96x96px, 3 channels, values 0 - 255)\n    X = np.zeros([N,96,96,3],dtype=np.uint8) \n    #convert the labels to a numpy array too\n    y = np.squeeze(df.as_matrix(columns=['label']))[0:N]\n    #read images one by one, tdqm notebook displays a progress bar\n    for i, row in tqdm_notebook(df.iterrows(), total=N):\n        if i == N:\n            break\n        X[i] = cv2.imread(row['path'])\n          \n    return X,y","metadata":{"execution":{"iopub.status.busy":"2022-11-19T18:39:07.601486Z","iopub.execute_input":"2022-11-19T18:39:07.601775Z","iopub.status.idle":"2022-11-19T18:39:07.607075Z","shell.execute_reply.started":"2022-11-19T18:39:07.601723Z","shell.execute_reply":"2022-11-19T18:39:07.606313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#load data\nN=1000\nX, y = load_data(N, df)","metadata":{"execution":{"iopub.status.busy":"2022-11-19T18:39:45.216787Z","iopub.execute_input":"2022-11-19T18:39:45.217142Z","iopub.status.idle":"2022-11-19T18:39:53.606665Z","shell.execute_reply.started":"2022-11-19T18:39:45.217080Z","shell.execute_reply":"2022-11-19T18:39:53.605826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(8, 8), dpi=150)\nnp.random.seed(100) #we can use the seed to get a different set of random images\nimage_count = 16  # number of images we want to have\n\n# iterate over a list of random numbers, then plot images with labels\nfor plot_idx, image_idx in enumerate(np.random.randint(0, N, image_count)):\n    ax = fig.add_subplot(4, 4, plot_idx+1)\n    plt.imshow(X[image_idx])\n    ax.set_title(\"Label: \" + str(y[image_idx]))","metadata":{"execution":{"iopub.status.busy":"2022-11-19T18:45:48.549190Z","iopub.execute_input":"2022-11-19T18:45:48.549477Z","iopub.status.idle":"2022-11-19T18:45:51.424461Z","shell.execute_reply.started":"2022-11-19T18:45:48.549422Z","shell.execute_reply":"2022-11-19T18:45:51.423813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let's analyze the positive and negative sample distribution in our data to see if it has any bias.","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(4, 2),dpi=100)\n#plot bars of numbers of positive and negative samples\nnumber_of_negatives = (y==0).sum()\nnumber_of_positives = (y==1).sum()\nnegative_samples = X[y==0]\npositive_samples = X[y==1]\nplt.bar([1, 0], [number_of_negatives, number_of_positives])\nplt.xticks([1, 0], [\"Negative N = \" + str(number_of_negatives), \n                    \"Positive N = \" + str(number_of_positives)])\nplt.ylabel(\"# of samples\")\n...","metadata":{"execution":{"iopub.status.busy":"2022-11-19T18:51:02.027395Z","iopub.execute_input":"2022-11-19T18:51:02.027689Z","iopub.status.idle":"2022-11-19T18:51:02.257095Z","shell.execute_reply.started":"2022-11-19T18:51:02.027638Z","shell.execute_reply":"2022-11-19T18:51:02.255991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Now, we wonder if it's possible to tell a positive and a negative sample apart from each other just from an image. To check that, we will look at whether negative samples generally have different color or brightness compare to the positive samples.\n* First, compare each color channel (red, green, blue) of positive and negatives.\n* Then, make a cumulative comparison.\n* Finally, compare overall brightness.\n* Scroll back up and compare your conclusions with the 16 image sample you have. Is it accurate? Specifically, check if your observations for red and brightness is generally true so far.","metadata":{}},{"cell_type":"code","source":"nr_of_bins = 256 #each possible pixel value will get a bin in the following histograms\nfig,axs = plt.subplots(4,2,sharey=True,figsize=(8,8),dpi=100)  # 4 by 2 plot\nrgb_list = [\"Red\", \"Green\", \"Blue\", \"RGB\"]\n\n#append a histogram to each of the 4*2=8 positions of the axs\nfor row_idx in range(0, 4):\n    for col_idx in range(0, 2):\n        if row_idx < 3:\n            axs[row_idx, 0].set_ylabel(\"Relative Frequency\")\n            axs[row_idx, 1].set_ylabel(rgb_list[row_idx], rotation=\"horizontal\",\n                                       labelpad=35, fontsize=12)\n            # show the color channels\n            if col_idx == 0: #show positive samples\n                axs[row_idx, 0].hist(positive_samples[:, :, :, row_idx].flatten(),\n                                     bins=nr_of_bins, density = True)\n            elif col_idx == 1: #show negative sampels\n                axs[row_idx, 1].hist(negative_samples[:, :, :, row_idx].flatten(),\n                                     bins=nr_of_bins, density = True)\n                \n        else:\n            # show the rgb as cumulative\n            if col_idx == 0: #show positive samples\n                axs[row_idx, 0].hist(positive_samples.flatten(),\n                                     bins=nr_of_bins, density = True)\n            elif col_idx == 1: #show negative sampels\n                axs[row_idx, 1].hist(negative_samples.flatten(),\n                                     bins=nr_of_bins, density = True)\n                \naxs[0, 0].set_title(\"Positive Samples\")\naxs[0, 1].set_title(\"Negative Samples\")\n\n...","metadata":{"execution":{"iopub.status.busy":"2022-11-19T19:02:13.437851Z","iopub.execute_input":"2022-11-19T19:02:13.438204Z","iopub.status.idle":"2022-11-19T19:02:22.178855Z","shell.execute_reply.started":"2022-11-19T19:02:13.438148Z","shell.execute_reply":"2022-11-19T19:02:22.177979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 3: Preprocessing the Data\nNow that we have our data, we want to turn it into something that our CNN model can take as an input. In order to improve the accuracy of our model in the real world, we will introduce something called transformations. Transformations allow us to play with our data so that they are imperfect and unpredictable, which intuitively sounds weird, but it actually helps our model to be have a greater test accuracy because it prevents overfitting. Transformations also allow us to have more variety, which again imrpoves the real world application of it.","metadata":{}},{"cell_type":"code","source":"#transform training data\ntrans_train = transforms.Compose([transforms.ToPILImage(),\n                                 transforms.Pad(64, padding_mode=\"reflect\"),\n                                 transforms.RandomHorizontalFlip(),\n                                 transforms.RandomVerticalFlip(),\n                                 transforms.ToTensor(), # (h, w, c) -> (c, h, w)\n                                 transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]) \n])\n\n#transform validation data\ntrans_valid = transforms.Compose([transforms.ToPILImage(),\n                                 transforms.Pad(64, padding_mode=\"reflect\"),\n                                 transforms.ToTensor(), # (h, w, c) -> (c, h, w)\n                                 transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]) \n])","metadata":{"execution":{"iopub.status.busy":"2022-11-19T19:31:13.700239Z","iopub.execute_input":"2022-11-19T19:31:13.700545Z","iopub.status.idle":"2022-11-19T19:31:13.707402Z","shell.execute_reply.started":"2022-11-19T19:31:13.700490Z","shell.execute_reply":"2022-11-19T19:31:13.706639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class MyDataset(Dataset):\n    def __init__(self, df_data, data_dir = './', transform=None):\n        super().__init__()\n        self.df = df_data.values\n        self.data_dir = data_dir\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.df)\n    \n    def __getitem__(self, index):\n        img_name,label = self.df[index]\n        img_path = os.path.join(self.data_dir, img_name+'.tif')\n        image = cv2.imread(img_path)\n        if self.transform is not None:\n            image = self.transform(image)\n        return image, label","metadata":{"_uuid":"a5f0097544dc230bb794e75a163196f08de0c13c","execution":{"iopub.status.busy":"2022-11-19T19:31:15.493796Z","iopub.execute_input":"2022-11-19T19:31:15.494273Z","iopub.status.idle":"2022-11-19T19:31:15.500613Z","shell.execute_reply.started":"2022-11-19T19:31:15.494070Z","shell.execute_reply":"2022-11-19T19:31:15.499443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#get train and validation datasets\ndataset_train = MyDataset(train, train_path, trans_train)\ndataset_valid = MyDataset(val, train_path, trans_valid)\n\n#load train and validation datasets\nloader_train = DataLoader(dataset=dataset_train, batch_size=batch_size, shuffle=True, num_workers=0)\nloader_valid = DataLoader(dataset=dataset_valid, batch_size=batch_size//2, shuffle=False, num_workers=0)","metadata":{"_uuid":"cbacdb7bcd20d414d115cde200f564934b69f308","execution":{"iopub.status.busy":"2022-11-19T20:04:58.886668Z","iopub.execute_input":"2022-11-19T20:04:58.887012Z","iopub.status.idle":"2022-11-19T20:04:58.909198Z","shell.execute_reply.started":"2022-11-19T20:04:58.886953Z","shell.execute_reply":"2022-11-19T20:04:58.908261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#plot a transformed image to see how different it is!\nimport random\nrandom.seed(30)\nrand_idx = int(random.random() * len(dataset_train))\nimage, label = dataset_train[rand_idx]\nplt.imshow(image.permute(1, 2, 0)) # (c, h, w) -> (h, w, c)\nplt.title(label)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2022-11-19T19:31:24.223787Z","iopub.execute_input":"2022-11-19T19:31:24.224114Z","iopub.status.idle":"2022-11-19T19:31:24.511528Z","shell.execute_reply.started":"2022-11-19T19:31:24.224049Z","shell.execute_reply":"2022-11-19T19:31:24.510456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 4: Building the CNN Model!\nWe've made it! We analyzed and processed our data, and now we are ready to build our model that will train on it to later predict images that it has never seen before.","metadata":{"_uuid":"ac596a6914e7455f1326c01eb532d7b36b972e7f"}},{"cell_type":"code","source":"#build a simple CNN model of 6 convolutional layers, 6 batch normalization layers, \n#additional pooling layers, and a fully connected layer at the end to connect nodes to an output.\n#then, create a feed forward function that uses Leaky ReLU.\nclass SimpleCNN(nn.Module):\n    def __init__(self):\n        super(SimpleCNN, self).__init__()\n        self.conv1 = nn.Conv2d(in_channels=3, out_channels=32, kernel_size=3, padding=2)\n        self.conv2 = nn.Conv2d(in_channels=32, out_channels=64, kernel_size=3, padding=2)\n        self.conv3 = nn.Conv2d(in_channels=64, out_channels=128, kernel_size=3, padding=2)\n        self.conv4 = nn.Conv2d(in_channels=128, out_channels=256, kernel_size=3, padding=2)\n        self.conv5 = nn.Conv2d(in_channels=256, out_channels=512, kernel_size=3, padding=2)\n        self.conv6 = nn.Conv2d(in_channels=512, out_channels=1024, kernel_size=3, padding=2)\n        self.bn1 = nn.BatchNorm2d(32)\n        self.bn2 = nn.BatchNorm2d(64)\n        self.bn3 = nn.BatchNorm2d(128)\n        self.bn4 = nn.BatchNorm2d(256)\n        self.bn5 = nn.BatchNorm2d(512)\n        self.bn6 = nn.BatchNorm2d(1024)\n        self.pool = nn.MaxPool2d(kernel_size=2, stride=2)\n        self.avg = nn.AvgPool2d(8)\n        self.fc = nn.Linear(1024, 2)\n    def forward(self, x):\n        x = self.pool(F.leaky_relu(self.bn1(self.conv1(x))))\n        x = self.pool(F.leaky_relu(self.bn2(self.conv2(x))))\n        x = self.pool(F.leaky_relu(self.bn3(self.conv3(x))))\n        x = self.pool(F.leaky_relu(self.bn4(self.conv4(x))))\n        x = self.pool(F.leaky_relu(self.bn5(self.conv5(x))))\n        x = self.pool(F.leaky_relu(self.bn6(self.conv6(x))))\n        x = self.avg(x)\n        x = x.view(x.size(0), -1)\n        x = self.fc(x)\n        return x\n        ","metadata":{"_uuid":"f32c84715ab1c585a56b41fc71c05c42d55bfd4b","execution":{"iopub.status.busy":"2022-11-19T20:00:11.104376Z","iopub.execute_input":"2022-11-19T20:00:11.104682Z","iopub.status.idle":"2022-11-19T20:00:11.115097Z","shell.execute_reply.started":"2022-11-19T20:00:11.104620Z","shell.execute_reply":"2022-11-19T20:00:11.114116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Set up your parameters and create your model\nMake sure you understand what each parameter and object is used for.","metadata":{}},{"cell_type":"code","source":"# Hyper parameters\nnum_epochs = 8\nnum_classes = 2\nbatch_size = 32\nlearning_rate = 0.002\n\n# Device configuration\ndevice = torch.device('cuda:0' if torch.cuda.is_available() else 'cpu')\n\ngc.collect()\ntorch.cuda.empty_cache()","metadata":{"_uuid":"67522fc88f034ddb684f9fba0e154c341760d1d1","execution":{"iopub.status.busy":"2022-11-19T20:00:17.980567Z","iopub.execute_input":"2022-11-19T20:00:17.980856Z","iopub.status.idle":"2022-11-19T20:00:18.094522Z","shell.execute_reply.started":"2022-11-19T20:00:17.980802Z","shell.execute_reply":"2022-11-19T20:00:18.093453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = SimpleCNN().to(device)","metadata":{"execution":{"iopub.status.busy":"2022-11-19T20:00:20.837456Z","iopub.execute_input":"2022-11-19T20:00:20.837755Z","iopub.status.idle":"2022-11-19T20:00:24.588220Z","shell.execute_reply.started":"2022-11-19T20:00:20.837701Z","shell.execute_reply":"2022-11-19T20:00:24.586936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Backpropagation: Choosing your Cost Function and Optimizer\nWe are using CrossEntropyLoss as our cost function, and Adamax as our optimizer, a Stochastic Optimization.","metadata":{}},{"cell_type":"code","source":"# Loss and optimizer\ncriterion = nn.CrossEntropyLoss()\noptimizer = torch.optim.Adamax(model.parameters(), lr=learning_rate)","metadata":{"_uuid":"a56a159ba70e958c82e518d718be99bd41c17299","execution":{"iopub.status.busy":"2022-11-19T20:09:55.745833Z","iopub.execute_input":"2022-11-19T20:09:55.746147Z","iopub.status.idle":"2022-11-19T20:09:55.752714Z","shell.execute_reply.started":"2022-11-19T20:09:55.746085Z","shell.execute_reply":"2022-11-19T20:09:55.752085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create your Training Loop\nThis is where your model actually trains. The training loop should keep the model with the lowest loss, which we will use to predict labels for the images in our test set","metadata":{}},{"cell_type":"code","source":"#create a training loop that iterates over the training set for epoch times,\n#and improves the model in each iteration.\n\n#the actual data has 198022 images. If you want to train on all of them, remove\n#the lines below. We will only train on 1000 images to show you the result\n#within the duration of workshop.\ntotal_num = 1000\ntotal_step = int(total_num/batch_size)\nif total_num%batch_size != 0:\n    total_step += 1 \nprint('total is ', total_step)\n\n#write your training loop here\nfor epoch in range(num_epochs):\n    step, data_count = 0, 0\n    start = time.time()\n    \n    for images, labels in iter(loader_train):\n        data_count += len(images)\n        if torch.cuda.is_available():\n            images = images.cuda()\n            labels = labels.cuda()\n        \n        outputs = model(images)\n        loss = criterion(outputs, labels)\n        \n        optimizer.zero_grad()\n        loss.backward()\n        optimizer.step()\n        \n        if (step+1) % 10 == 0:\n            end = time.time()\n            print(\"============================== step runtime from start of current epoch {} ====================================\".format(end - start))\n            print ('Epoch [{}/{}], Step [{}/{}], Loss: {:.4f}' \n                   .format(epoch+1, num_epochs, step+1, total_step, loss.item()))\n            \n        step += 1\n        if (data_count > total_num):\n            end = time.time()\n            print(\"============================== EPOCH RUN COMPLETED ====================================\",end - start)\n            print ('Epoch [{}/{}], Step [{}/{}], Loss: {:.4f}' \n                   .format(epoch+1, num_epochs, step+1, total_step, loss.item()))\n            break","metadata":{"_uuid":"dca29061dd5d6a60ba95f16dbaea70e6dc822150","execution":{"iopub.status.busy":"2022-11-19T20:10:11.955715Z","iopub.execute_input":"2022-11-19T20:10:11.956036Z","iopub.status.idle":"2022-11-19T20:11:23.121956Z","shell.execute_reply.started":"2022-11-19T20:10:11.955977Z","shell.execute_reply":"2022-11-19T20:11:23.120121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Validate the Accuracy of your Model\nWe are using the validation data set to compare the model's predictions with the actual results. Notice that we are no longer training our model here.","metadata":{}},{"cell_type":"code","source":"# Test the model\n\ntesting_num = 2000\n\nmodel.eval()  # eval mode (batchnorm uses moving mean/variance instead of mini-batch mean/variance)\nwith torch.no_grad():  # prevents the model from running gradient descent, thus saving time since we are not training\n    correct = 0\n    total = 0\n    for images, labels in loader_valid:\n        images = images.to(device)\n        labels = labels.to(device)\n        outputs = model(images)\n        _, predicted = torch.max(outputs.data, 1)\n        total += labels.size(0)\n        correct += (predicted == labels).sum().item()\n        \n        if(total>testing_num):\n            break\n    \n    #calculate the accuracy here in terms of \"correct\" and \"total\"\n    accuracy = correct / total * 100\n    \n    print('Test Accuracy of the model on the 22003 test images: {} %'.format(accuracy))\n\n# Save the model checkpoint\ntorch.save(model.state_dict(), 'model.ckpt')","metadata":{"_uuid":"998c3c05a3c3fb6fc7f2523dddff50496490d051","execution":{"iopub.status.busy":"2022-11-19T20:14:37.414119Z","iopub.execute_input":"2022-11-19T20:14:37.414444Z","iopub.status.idle":"2022-11-19T20:14:54.505516Z","shell.execute_reply.started":"2022-11-19T20:14:37.414391Z","shell.execute_reply":"2022-11-19T20:14:54.504712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submit your Model to the Competition!\nWe will not do this in our workshop because it takes time to evaluate our model, but feel free to do it on your own! The code is provided below.","metadata":{"_uuid":"c1c0709b4a1e7edea538eda5168819a7b949eedf"}},{"cell_type":"code","source":"dataset_valid = MyDataset(df_data=submission, data_dir=test_path, transform=trans_valid)\nloader_test = DataLoader(dataset = dataset_valid, batch_size=32, shuffle=False, num_workers=0)","metadata":{"_uuid":"4d03acf3711d8036703041276749bbaf2ef64e71","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nmodel.eval()\n\npreds = []\nfor batch_i, (data, target) in enumerate(loader_test):\n    data, target = data.cuda(), target.cuda()\n    output = model(data)\n\n    pr = output[:,1].detach().cpu().numpy()\n    for i in pr:\n        preds.append(i)\nsubmission.shape, len(preds)\nsubmission['label'] = preds\nsubmission.to_csv('s.csv', index=False)","metadata":{"_uuid":"b85e1661e660b45486ab1b46a4d925e18f541b75","trusted":true},"execution_count":null,"outputs":[]}]}