{"cells":[{"metadata":{"_uuid":"2d6da9673d8622bb6edbadb79b98b77263fc1d70"},"cell_type":"markdown","source":"# Welcome to the Human Protein Atlas Competition!\n\nI started this kernel notebook to explore the data and build a simple baseline model to play with. I have never mind that the results are useful for many kagglers that like to start with the competition but need some starter code or some inspiration. Thank you for pushing this kernel that far! :-)\n\nIf you have just found this kernel, here is a **short summary of what you can find**:\n\n1. Encoding of binary target labels out of the given multilabel list per image,\n2. Visual analysis of target protein distribution in the train set,\n3. A simple image generator that yields images of a target-protein-wishlist. Each sample that has at least one match with this list is returned.\n4. Some ideas on validation.\n5. A baseline model build with keras that is supported by:\n    * A modelparameter class that holds all parameters that are necessary to build the model, to load the data and to preprocess the images.\n    * A data generator that can be used with CPU/GPU computing to perform training and validation.\n    * An image preprocessor that rescales, reshapes and normalizes the images for feeding into the model.\n6. Ideas on how to improve the baseline model by tracking loss with a keras callback. \n7. Some ideas on how to proceed. (coming soon)"},{"metadata":{"_uuid":"cbec118af20f078fd71b23c8ee0b7aa0cb89daa5"},"cell_type":"markdown","source":"## Latest important updates\n\nBeside building competition code I'm still updating this kernel with ideas and code. To make it easier for you to checkout new content, here is an **update summary** of the latest changes:\n\n* Ideas how to solve jiggle-wiggle losses and gradients: dropout\n* DataGenerator: The **multithreaded version to perform predictions with DataGenerator was dropped**. I've tried out various different methods the last days (using ImageGenerator, flow_form_dataframe etc.) but every time there were some annoying problems. As training is the most expensive part, I decided to keep multithreading for fitting but to drop if for making predictions. \n* **Bug-fix** in modelparameter class. **Shuffle was always set to True**. Now default is False and you can change it interactively. \n* **Modelparameter class** has now an **attribute basepath** that points to your desired image directory. If you perform a prediction with same modelparameters, the path is overwritten during prediction with predict_path. Feels dirty, this will be changed soon. \n* The **image preprocessor is now responsible for loading the images**. With default model parameter n_channels=1 it only returns the green image. If you like to change that increase n_channels as you like or do it by hand but make sure than modelparameters has a suitable n_channels. \n* The **f1 score** of the improved metric is extended by **min, max, std** besides the mean to gain more insights of the f1 score distribution between different target classes. \n* Some ideas what causes the gradient noise and how to set up solutions. "},{"metadata":{"_uuid":"16c5e9f6b0cc1e2be83d54adff6ec5ed16dbe0d3"},"cell_type":"markdown","source":"### Our goal\n\n* Predict various protein structures in cellular images\n* there are 28 different target proteins\n* multiple proteins can be present in one image (multilabel classification)\n* 27 different cell types of highly different morphology"},{"metadata":{"_uuid":"bbba5a9f15329ac61510f60a84337c1fe2a6e98e"},"cell_type":"markdown","source":"## Kernel settings\n\nThis notebook contains model fitting that may take some time. **If you don't like to wait for compution, you can set fit_baseline and/or fit_improved_baseline of the KernelSettings class to False**:"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"550027fb14c0774f8a58b5220548429126549131"},"cell_type":"code","source":"class KernelSettings:\n    \n    def __init__(self, fit_baseline=False, fit_improved_baseline=False):\n        self.fit_baseline = fit_baseline\n        self.fit_improved_baseline = fit_improved_baseline","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"efd75058e48c0b7e962f4592ad14ec7d705196cb"},"cell_type":"code","source":"kernelsettings = KernelSettings(fit_baseline=False, fit_improved_baseline=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b2735faa22f122ed33728f0782bf911e402774ca"},"cell_type":"markdown","source":"This way the prediction probabilities of the corresponding model are loaded as csv from added data source. "},{"metadata":{"_uuid":"51824d453581725c0c20a38ae056b9d80a893bdf"},"cell_type":"markdown","source":"## Improved model settings\n\nWithin this kernel you can play with different improvement steps of the baseline model in a plug and play style. To make it easier to understand which improvements you currently use, let's set them here in front of the analysis:"},{"metadata":{"trusted":true,"_uuid":"32888a471d6e55cbcc2e1f88d1a6c974ae0e991f"},"cell_type":"code","source":"use_dropout=False","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cce6944f02a874a8773859db189ab119393822ac"},"cell_type":"markdown","source":"## Loading packages and data"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\nfrom PIL import Image\nfrom scipy.misc import imread\n\nimport tensorflow as tf\nsns.set()\n\nimport os\nprint(os.listdir(\"../input\"))\n\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=DeprecationWarning)\nwarnings.filterwarnings(\"ignore\", category=UserWarning)\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train_labels = pd.read_csv(\"../input/human-protein-atlas-image-classification/train.csv\")\ntrain_labels.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"69a6b9c1584ebcf68b5aab128199f524daccd365"},"cell_type":"markdown","source":"How many samples do we have?"},{"metadata":{"trusted":true,"_uuid":"abaa2a7225e485102ceae9bdda12f4cf835c492f","_kg_hide-input":true},"cell_type":"code","source":"train_labels.shape[0]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"882bba3d7f68f209051fea3899f046bd96dc2915"},"cell_type":"markdown","source":"## Helper code"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"beaac70d0392fab7c5b1d49649a8c994310320e4"},"cell_type":"code","source":"label_names = {\n    0:  \"Nucleoplasm\",  \n    1:  \"Nuclear membrane\",   \n    2:  \"Nucleoli\",   \n    3:  \"Nucleoli fibrillar center\",   \n    4:  \"Nuclear speckles\",\n    5:  \"Nuclear bodies\",   \n    6:  \"Endoplasmic reticulum\",   \n    7:  \"Golgi apparatus\",   \n    8:  \"Peroxisomes\",   \n    9:  \"Endosomes\",   \n    10:  \"Lysosomes\",   \n    11:  \"Intermediate filaments\",   \n    12:  \"Actin filaments\",   \n    13:  \"Focal adhesion sites\",   \n    14:  \"Microtubules\",   \n    15:  \"Microtubule ends\",   \n    16:  \"Cytokinetic bridge\",   \n    17:  \"Mitotic spindle\",   \n    18:  \"Microtubule organizing center\",   \n    19:  \"Centrosome\",   \n    20:  \"Lipid droplets\",   \n    21:  \"Plasma membrane\",   \n    22:  \"Cell junctions\",   \n    23:  \"Mitochondria\",   \n    24:  \"Aggresome\",   \n    25:  \"Cytosol\",   \n    26:  \"Cytoplasmic bodies\",   \n    27:  \"Rods & rings\"\n}\n\nreverse_train_labels = dict((v,k) for k,v in label_names.items())\n\ndef fill_targets(row):\n    row.Target = np.array(row.Target.split(\" \")).astype(np.int)\n    for num in row.Target:\n        name = label_names[int(num)]\n        row.loc[name] = 1\n    return row","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"707280dec8e4b10dc743b5473ba8d1c4492141ff"},"cell_type":"markdown","source":"## Which proteins occur most often in images?"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"13a1f4fd6b594cd4a225f6d79d966096d1c800ae"},"cell_type":"code","source":"for key in label_names.keys():\n    train_labels[label_names[key]] = 0","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"3662893c91d0f74fe666300c214f67cfb03060c7"},"cell_type":"code","source":"train_labels = train_labels.apply(fill_targets, axis=1)\ntrain_labels.head()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"f516b3a1a297a167264b0f22803de231096c9719"},"cell_type":"code","source":"target_counts = train_labels.drop([\"Id\", \"Target\"],axis=1).sum(axis=0).sort_values(ascending=False)\nplt.figure(figsize=(15,15))\nsns.barplot(y=target_counts.index.values, x=target_counts.values, order=target_counts.index)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"af190deeab9f05afcba339f31c7162d841e58b33"},"cell_type":"markdown","source":"### Take-Away\n\n* We can see that most common protein structures belong to coarse grained cellular components like the plasma membrane, the cytosol and the nucleus. \n* In contrast small components like the lipid droplets, peroxisomes, endosomes, lysosomes, microtubule ends, rods and rings are very seldom in our train data. For these classes the prediction will be very difficult as we have only a few examples that may not cover all variabilities and as our model probably will be confused during ins learning process by the major classes. Due to this confusion we will make less accurate predictions on the minor classes.\n* Consequently accuracy is not the right score here to measure your performance and validation strategy should be very fine. "},{"metadata":{"_uuid":"40bff26b2d170e9936c5e6ce4fbef12e7cb4f2ba"},"cell_type":"markdown","source":"## How many targets are most common?"},{"metadata":{"trusted":true,"_uuid":"313f47ac92fa1ff241148af6282c200451edbea7","_kg_hide-input":true},"cell_type":"code","source":"train_labels[\"number_of_targets\"] = train_labels.drop([\"Id\", \"Target\"],axis=1).sum(axis=1)\ncount_perc = np.round(100 * train_labels[\"number_of_targets\"].value_counts() / train_labels.shape[0], 2)\nplt.figure(figsize=(20,5))\nsns.barplot(x=count_perc.index.values, y=count_perc.values, palette=\"Reds\")\nplt.xlabel(\"Number of targets per image\")\nplt.ylabel(\"% of data\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d6e2a72c6f544f4d03a7f0e273a5819982b58108"},"cell_type":"markdown","source":"### Take-away\n\n* Most train images only have 1 or two target labels.\n* More than 3 targets are very seldom!"},{"metadata":{"_uuid":"cb957f819453857d62a96a236febabcfb9ce5a22"},"cell_type":"markdown","source":"## Which targets are correlated?\n\nLet's see if we find some correlations between our targets. This way we may already see that some proteins often come together."},{"metadata":{"trusted":true,"_uuid":"b3c0f5a77e34296416136ca2800b8623f94885d8","_kg_hide-input":true},"cell_type":"code","source":"plt.figure(figsize=(15,15))\nsns.heatmap(train_labels[train_labels.number_of_targets>1].drop(\n    [\"Id\", \"Target\", \"number_of_targets\"],axis=1\n).corr(), cmap=\"RdYlBu\", vmin=-1, vmax=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c8083f38bf943a2d38c86544cf322ffe721d8a19"},"cell_type":"markdown","source":"### Take-away\n\n* We can see that many targets only have very slight correlations. \n* In contrast, endosomes and lysosomes often occur together and sometimes seem to be located at the endoplasmatic reticulum. \n* In addition we find that the mitotic spindle often comes together with the cytokinetic bridge. This makes sense as both are participants for cellular division. And in this process microtubules and thier ends are active and participate as well. Consequently we find a positive correlation between these targets.\n\nIf you like to dive deeper into target correlations you may like to take a look at my kernel notebook about [target clustering.](https://www.kaggle.com/allunia/in-depth-protein-correlations) :-)"},{"metadata":{"_uuid":"3b7bc009eb7ec4b6d24e4e0951d5d799ac42d289"},"cell_type":"markdown","source":"## How are special and seldom targets grouped?"},{"metadata":{"_uuid":"6f0bf1de2bf7c2550cd7bb320a16ce0d79b25087"},"cell_type":"markdown","source":"### Lysosomes and endosomes\n\nLet's start with these high correlated features!"},{"metadata":{"trusted":true,"_uuid":"86c06ed4c0dbf02c605e447a5ecc40373b35771a","_kg_hide-input":true},"cell_type":"code","source":"def find_counts(special_target, labels):\n    counts = labels[labels[special_target] == 1].drop(\n        [\"Id\", \"Target\", \"number_of_targets\"],axis=1\n    ).sum(axis=0)\n    counts = counts[counts > 0]\n    counts = counts.sort_values()\n    return counts","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e2af724b721a15ebd294771629fb270c9ad9f0ae","_kg_hide-input":true},"cell_type":"code","source":"lyso_endo_counts = find_counts(\"Lysosomes\", train_labels)\n\nplt.figure(figsize=(10,3))\nsns.barplot(x=lyso_endo_counts.index.values, y=lyso_endo_counts.values, palette=\"Blues\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dc6dbcea143b60d0ef360e75d14b53b09d36e7be"},"cell_type":"markdown","source":"### Rods and rings"},{"metadata":{"trusted":true,"_uuid":"5f13e4113cfb00ec8b2de6d01fb548e10b794279","_kg_hide-input":true},"cell_type":"code","source":"rod_rings_counts = find_counts(\"Rods & rings\", train_labels)\nplt.figure(figsize=(15,3))\nsns.barplot(x=rod_rings_counts.index.values, y=rod_rings_counts.values, palette=\"Greens\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fea4a1a472bf60fdc1a5317a2e616c98b781f7ab"},"cell_type":"markdown","source":"### Peroxisomes"},{"metadata":{"trusted":true,"_uuid":"11fef619f19eafc3657d693d05f7a50c786e0621","_kg_hide-input":true},"cell_type":"code","source":"peroxi_counts = find_counts(\"Peroxisomes\", train_labels)\n\nplt.figure(figsize=(15,3))\nsns.barplot(x=peroxi_counts.index.values, y=peroxi_counts.values, palette=\"Reds\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"555447f455cac28ec0c84b09b89f32b657371470"},"cell_type":"markdown","source":"### Microtubule ends"},{"metadata":{"trusted":true,"_uuid":"a543faa3994e6bd7b57fda135ea4b5bea40efed3","_kg_hide-input":true},"cell_type":"code","source":"tubeends_counts = find_counts(\"Microtubule ends\", train_labels)\n\nplt.figure(figsize=(15,3))\nsns.barplot(x=tubeends_counts.index.values, y=tubeends_counts.values, palette=\"Purples\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0a04ae9a48f60f02021c0dba920b2ac3d6d65b38"},"cell_type":"markdown","source":"### Nuclear speckles"},{"metadata":{"trusted":true,"_uuid":"85d838037fc3310b9eef31dc4e6c64e1bc35efec","_kg_hide-input":true},"cell_type":"code","source":"nuclear_speckles_counts = find_counts(\"Nuclear speckles\", train_labels)\n\nplt.figure(figsize=(15,3))\nsns.barplot(x=nuclear_speckles_counts.index.values, y=nuclear_speckles_counts.values, palette=\"Oranges\")\nplt.xticks(rotation=\"70\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ba2b973b09d2436f65e8775b214d377d46b29188"},"cell_type":"markdown","source":"### Take-away\n\n* We can see that even with very seldom targets we find some kind of grouping with other targets that reveal where the protein structure seems to be located. \n* For example, we can see that rods and rings have something to do with the nucleus whereas peroxisomes may be located in the nucleus as well as in the cytosol.\n* Perhaps this patterns might help to build a more robust model!  "},{"metadata":{"_uuid":"0714e168ea347df2111124d0f341de3172bac00d"},"cell_type":"markdown","source":"## How do the images look like?\n\n"},{"metadata":{"_uuid":"4af4b12200fd62e7726ff8a481d7566e564947ad"},"cell_type":"markdown","source":"### Peek into the directory\n\nBefore we start loading images, let's have a look into the train directory to get an impression of what we can find there:"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"1d11fa62aaa798cfcb8671b7aac0e4cd4fa18776"},"cell_type":"code","source":"from os import listdir\n\nfiles = listdir(\"../input/human-protein-atlas-image-classification/train\")\nfor n in range(10):\n    print(files[n])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f8804380faaaeb428227dec5ee280f27c57baea"},"cell_type":"markdown","source":"Ah, ok, great! It seems that for one image id, there are different color channels present. Looking into the data description of this competition we can find that:\n\n* Each image is actually splitted into 4 different image files. \n* These 4 files correspond to 4 different filter:\n    * a **green** filter for the **target protein structure** of interest\n    * **blue** landmark filter for the **nucleus**\n    * **red** landmark filter for **microtubules**\n    * **yellow** landmark filter for the **endoplasmatic reticulum**\n* Each image is of size 512 x 512"},{"metadata":{"_uuid":"e29074f7ddba22fbdd2275fdc8782dc2dd5885c6"},"cell_type":"markdown","source":"Let's check if the number of files divided by 4 yields the number of target samples:"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"8545dd9afec3e0cdc02873375ba81ff4f4658aab"},"cell_type":"code","source":"len(files) / 4 == train_labels.shape[0]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b89a5ae1b6e358d7100c79e5a2cdc5d77ff88cc2"},"cell_type":"markdown","source":"## How do images of specific targets look like?\n\nWhile looking at examples, we can build an batch loader:"},{"metadata":{"trusted":true,"_uuid":"ce954da724c8a400c12080f8c3f986a537d4a4c4"},"cell_type":"code","source":"train_path = \"../input/human-protein-atlas-image-classification/train/\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c7e4ffb20ebaee046d888620d0cf016f0cebf2c4","_kg_hide-input":true},"cell_type":"code","source":"def load_image(basepath, image_id):\n    images = np.zeros(shape=(4,512,512))\n    images[0,:,:] = imread(basepath + image_id + \"_green\" + \".png\")\n    images[1,:,:] = imread(basepath + image_id + \"_red\" + \".png\")\n    images[2,:,:] = imread(basepath + image_id + \"_blue\" + \".png\")\n    images[3,:,:] = imread(basepath + image_id + \"_yellow\" + \".png\")\n    return images\n\ndef make_image_row(image, subax, title):\n    subax[0].imshow(image[0], cmap=\"Greens\")\n    subax[1].imshow(image[1], cmap=\"Reds\")\n    subax[1].set_title(\"stained microtubules\")\n    subax[2].imshow(image[2], cmap=\"Blues\")\n    subax[2].set_title(\"stained nucleus\")\n    subax[3].imshow(image[3], cmap=\"Oranges\")\n    subax[3].set_title(\"stained endoplasmatic reticulum\")\n    subax[0].set_title(title)\n    return subax\n\ndef make_title(file_id):\n    file_targets = train_labels.loc[train_labels.Id==file_id, \"Target\"].values[0]\n    title = \" - \"\n    for n in file_targets:\n        title += label_names[n] + \" - \"\n    return title","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"090d1b258d2641e747e58a49376ab3d32b3893b8","_kg_hide-input":true},"cell_type":"code","source":"class TargetGroupIterator:\n    \n    def __init__(self, target_names, batch_size, basepath):\n        self.target_names = target_names\n        self.target_list = [reverse_train_labels[key] for key in target_names]\n        self.batch_shape = (batch_size, 4, 512, 512)\n        self.basepath = basepath\n    \n    def find_matching_data_entries(self):\n        train_labels[\"check_col\"] = train_labels.Target.apply(\n            lambda l: self.check_subset(l)\n        )\n        self.images_identifier = train_labels[train_labels.check_col==1].Id.values\n        train_labels.drop(\"check_col\", axis=1, inplace=True)\n    \n    def check_subset(self, targets):\n        return np.where(set(targets).issubset(set(self.target_list)), 1, 0)\n    \n    def get_loader(self):\n        filenames = []\n        idx = 0\n        images = np.zeros(self.batch_shape)\n        for image_id in self.images_identifier:\n            images[idx,:,:,:] = load_image(self.basepath, image_id)\n            filenames.append(image_id)\n            idx += 1\n            if idx == self.batch_shape[0]:\n                yield filenames, images\n                filenames = []\n                images = np.zeros(self.batch_shape)\n                idx = 0\n        if idx > 0:\n            yield filenames, images\n            ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"61ba9ebe5b6bfbea57d24e43bc8f4a1c2638d5eb"},"cell_type":"markdown","source":"Let's try to visualize specific target groups. **In this example we will see images that contain the protein structures lysosomes or endosomes**. Set target values of your choice and the target group iterator will collect all images that are subset of your choice:"},{"metadata":{"trusted":true,"_uuid":"445f454dc36c1f4056bc870e8810f219f1c560ea"},"cell_type":"code","source":"your_choice = [\"Lysosomes\", \"Endosomes\"]\nyour_batch_size = 3","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6dcd0c8f57651ba21b82420edd37f70f93953215","_kg_hide-input":true},"cell_type":"code","source":"imageloader = TargetGroupIterator(your_choice, your_batch_size, train_path)\nimageloader.find_matching_data_entries()\niterator = imageloader.get_loader()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7d1f440b6ed70150f46a76970b544752f0d194a5"},"cell_type":"markdown","source":"To keep the kernel dense, the target group iterator has a batch size which stands for the number of examples you like to look at once. In this example you can see a maximum amount of 3 images at one iteration.  **To observe the next 3 examples of your target group, just run the cell below again.** This way you can run the cell until you have seen all images of your group without polluting the kernel:"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"5f7c98a6e15c081fcbe54c4e678ea2ea5ae4fbd7"},"cell_type":"code","source":"file_ids, images = next(iterator)\n\nfig, ax = plt.subplots(len(file_ids),4,figsize=(20,5*len(file_ids)))\nif ax.shape == (4,):\n    ax = ax.reshape(1,-1)\nfor n in range(len(file_ids)):\n    make_image_row(images[n], ax[n], make_title(file_ids[n]))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a66576fbdab4678c42d183017335d1ef27eed655"},"cell_type":"markdown","source":"### Take-Away\n\n* Looking at this few examples we can already obtain some insights:\n    * The staining of target proteins in the green channel was not equally successful. The **images differ in their intensities and the target proteins are not always located the same way**. The first image you can get by the loader shows endosomes that are spread all over the cells and in the second and third you can find endosomes and lysosomes more concetrated around the nucleus. \n    * Especially **in the red channel we can see morphological differences**. It looks like if the cells are of different types. This is just an assumption but perhaps one could use the red channel information to reveal cell types. "},{"metadata":{"_uuid":"ad5c5a22a659540bcb3da81453a6ef2cfaa958cc"},"cell_type":"markdown","source":"## Building a baseline model"},{"metadata":{"_uuid":"8dd65575d8a244f142b060aeb5cca30db9f27576"},"cell_type":"markdown","source":"### K-Fold Cross-Validation"},{"metadata":{"_uuid":"63437e877ed24bd86d34d5c7e8cd9f9e98d29d59"},"cell_type":"markdown","source":"Let's see how many test and train samples we have in this competition:"},{"metadata":{"trusted":true,"_uuid":"b2d74fa662ec90ebc2912105848daf7301845ef8","_kg_hide-input":true},"cell_type":"code","source":"train_files = listdir(\"../input/human-protein-atlas-image-classification/train\")\ntest_files = listdir(\"../input/human-protein-atlas-image-classification/test\")\npercentage = np.round(len(test_files) / len(train_files) * 100)\n\nprint(\"The test set size turns out to be {} % compared to the train set.\".format(percentage))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2d9613edd4f1a05b54d79dbe142b78e8cbec7625"},"cell_type":"markdown","source":"To understand the performance of our model we will use **k-fold cross validation**. The train data is splitted into k chunks and each chunk is used once for testing the prediction performance whereas the others are used for training. As our targets show relationships seemed to be grouped somehow the performance per test chunk probably highly depends on the target distribution per test chunk. For example there could be chunks with very seldom targets that may obtain a bad score and some chunks with very common targets and a very good score. To reduce this effect, we will **repeat the K-Fold several times** and look at scoing distributions in the end."},{"metadata":{"_uuid":"d94c71d94925daa6fe0a71a07c1da985925f7d91"},"cell_type":"markdown","source":"As our test data is 38 % of size compared to the train set it makes sense to use 3-Fold cross validation where the test set is 33 % of size compared to the train set. As we are working with neural networks that can be demanding in computational resources, let's only use 2 repetitions. "},{"metadata":{"trusted":true,"_uuid":"29a024cbd71145018c36342a5e8a483b6e010f5b"},"cell_type":"code","source":"from sklearn.model_selection import RepeatedKFold\n\nsplitter = RepeatedKFold(n_splits=3, n_repeats=1, random_state=0)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"37fc0f4495c93fc45379aab2eb156fdbd4d3f7c3"},"cell_type":"markdown","source":"This splitter is now a generator. Hence if you call splitters split method it will yield one Fold of the repeated K-Folds. Consequently if we choose n_repeats=2 we will end up with 6 Folds in total: 3 Folds for the first cross validation and again 3 Folds for the repeated cross validation. We will perform the splitting on the image ids. This way we can easily load images and targets given the chunk ids. **Due to performance reasons I will only use one cv-fold to explore results and one repeat!**"},{"metadata":{"trusted":true,"_uuid":"539a35dfd7d735d7068f32277146e60e15729e55"},"cell_type":"code","source":"partitions = []\n\nfor train_idx, test_idx in splitter.split(train_labels.index.values):\n    partition = {}\n    partition[\"train\"] = train_labels.Id.values[train_idx]\n    partition[\"validation\"] = train_labels.Id.values[test_idx]\n    partitions.append(partition)\n    print(\"TRAIN:\", train_idx, \"TEST:\", test_idx)\n    print(\"TRAIN:\", len(train_idx), \"TEST:\", len(test_idx))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"65f4c326d8f32f13872d0dfaaad2bbbfdda257a7"},"cell_type":"code","source":"partitions[0][\"train\"][0:5]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e0c336986b46a35c224dd447190fbb8c553bb64a"},"cell_type":"markdown","source":"### Collecting ideas\n\nNext we need to setup a simple baseline model. This need not be very complex or very good. Its our first attempt to play with and to figure out how to improve. For this purpose let's use the deep learning library [keras](https://keras.io/). This tools makes it easy for us to build and train neural networks. First of all, we should collect some ideas:\n\n* To **stay simple let's use only the green channel image of our images per id**. The competition says that it shows the stained target proteins and consequently it's hopefully the most informative one. The other images are like references showing microtubules, nucleus and endoplasmatic reticulum. We don't acutally now how informative they are and in our current state they would blow up our neural network with a huge amount of network weigths that we might not need.\n* Let's use **generators to only load data images of our batch and not all in once**. Using keras fit_generator, evaluate_generator and predict_generator we can directly connect them to keras without worrying much about how keras does its job. For this purpose I highly follow a descprition of a post in the www for which you will find the link below.\n* It could be advantegous to write a **small class that does simple preprocessing per image.** This way we can easily change something of this phase without producing chaos in the model itself or during data loading.   \n* I'm going to use a **small class that hold parameters that are used or shared between the data loader, the image preprocessor and the baseline model**. Passing an instance of this class to them reduced the risk of setting different parameters and obtaining mismatch errors for example during build & compile of the network layers. \n"},{"metadata":{"_uuid":"5cd298c9360c772a9c8b7946425449db8be0ab8f"},"cell_type":"markdown","source":"### Shared Parameter class"},{"metadata":{"trusted":true,"_uuid":"96df4f1aa1085b95b806b48071c0737d45a5b708","_kg_hide-input":true},"cell_type":"code","source":"class ModelParameter:\n    \n    def __init__(self, basepath,\n                 num_classes=28,\n                 image_rows=512,\n                 image_cols=512,\n                 batch_size=200,\n                 n_channels=1,\n                 row_scale_factor=4,\n                 col_scale_factor=4,\n                 shuffle=False,\n                 n_epochs=1):\n        self.basepath = basepath\n        self.num_classes = num_classes\n        self.image_rows = image_rows\n        self.image_cols = image_cols\n        self.batch_size = batch_size\n        self.n_channels = n_channels\n        self.shuffle = shuffle\n        self.row_scale_factor = row_scale_factor\n        self.col_scale_factor = col_scale_factor\n        self.scaled_row_dim = np.int(self.image_rows / self.row_scale_factor)\n        self.scaled_col_dim = np.int(self.image_cols / self.col_scale_factor)\n        self.n_epochs = n_epochs","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bc20cc4fdebf53b66bda3899974bd8758159d85f"},"cell_type":"markdown","source":"Ok, now we will create an instance of this class and pass it to the DataGenerator, the BaseLineModel and the ImagePreprocessor."},{"metadata":{"trusted":true,"_uuid":"96b80b33c817ad07deac345e36c46c062255c60d"},"cell_type":"code","source":"parameter = ModelParameter(train_path)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4497282fcdc0d68554e3878bb82bae8942eda681"},"cell_type":"markdown","source":"### Image Preprocessor\n\nLet's write a simple image preprocessor that handles for example the rescaling of the images. Perhaps we can expand its functionality during improvement of the baseline model. "},{"metadata":{"trusted":true,"_uuid":"2674c1331715aeab66a0194b3ad31afa7cfc3444","_kg_hide-input":true},"cell_type":"code","source":"from skimage.transform import resize\n\nclass ImagePreprocessor:\n    \n    def __init__(self, modelparameter):\n        self.parameter = modelparameter\n        self.basepath = self.parameter.basepath\n        self.scaled_row_dim = self.parameter.scaled_row_dim\n        self.scaled_col_dim = self.parameter.scaled_col_dim\n        self.n_channels = self.parameter.n_channels\n    \n    def preprocess(self, image):\n        image = self.resize(image)\n        image = self.reshape(image)\n        image = self.normalize(image)\n        return image\n    \n    def resize(self, image):\n        image = resize(image, (self.scaled_row_dim, self.scaled_col_dim))\n        return image\n    \n    def reshape(self, image):\n        image = np.reshape(image, (image.shape[0], image.shape[1], self.n_channels))\n        return image\n    \n    def normalize(self, image):\n        image /= 255 \n        return image\n    \n    def load_image(self, image_id):\n        image = np.zeros(shape=(512,512,4))\n        image[:,:,0] = imread(self.basepath + image_id + \"_green\" + \".png\")\n        image[:,:,1] = imread(self.basepath + image_id + \"_blue\" + \".png\")\n        image[:,:,2] = imread(self.basepath + image_id + \"_red\" + \".png\")\n        image[:,:,3] = imread(self.basepath + image_id + \"_yellow\" + \".png\")\n        return image[:,:,0:self.parameter.n_channels]\n        ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d93eeb59fa1f35d3e46f24832a40bee0086abb5b"},"cell_type":"markdown","source":"Let's create an instance of this preprocessor and pass it to the data generator."},{"metadata":{"trusted":true,"_uuid":"589121752a0ca1693c7e794655047f947a476b8e"},"cell_type":"code","source":"preprocessor = ImagePreprocessor(parameter)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c83a1888064cf71767aee2d462171227836600f5"},"cell_type":"markdown","source":"#### Looking at a preprocessed example image"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"cc9f0ff3eecc870337c475533b4c1cffed704403"},"cell_type":"code","source":"example = images[0,0]\npreprocessed = preprocessor.preprocess(example)\nprint(example.shape)\nprint(preprocessed.shape)\n\nfig, ax = plt.subplots(1,2,figsize=(20,10))\nax[0].imshow(example, cmap=\"Greens\")\nax[1].imshow(preprocessed.reshape(parameter.scaled_row_dim,parameter.scaled_col_dim), cmap=\"Greens\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"615d734ed90fb1def88151c5365443d61b411536"},"cell_type":"markdown","source":"You can see that we have lost a lot of information by downscaling the image!"},{"metadata":{"_uuid":"2b3cecd1dde678563c35405a7ddde5c35314677b"},"cell_type":"markdown","source":"### Data Generator\n\nI highly build upon the [nice data generator presented by Shervine Amidi.](https://stanford.edu/~shervine/blog/keras-how-to-generate-data-on-the-fly) Thank you! :-) \n"},{"metadata":{"trusted":true,"_uuid":"6e77e06f77164446a8689b55fd56734f943368b1","_kg_hide-input":true},"cell_type":"code","source":"import keras\n\nclass DataGenerator(keras.utils.Sequence):\n    \n    def __init__(self, list_IDs, labels, modelparameter, imagepreprocessor):\n        self.params = modelparameter\n        self.labels = labels\n        self.list_IDs = list_IDs\n        self.dim = (self.params.scaled_row_dim, self.params.scaled_col_dim)\n        self.batch_size = self.params.batch_size\n        self.n_channels = self.params.n_channels\n        self.num_classes = self.params.num_classes\n        self.shuffle = self.params.shuffle\n        self.preprocessor = imagepreprocessor\n        self.on_epoch_end()\n    \n    def on_epoch_end(self):\n        self.indexes = np.arange(len(self.list_IDs))\n        if self.shuffle == True:\n            np.random.shuffle(self.indexes)\n    \n    def get_targets_per_image(self, identifier):\n        return self.labels.loc[self.labels.Id==identifier].drop(\n                [\"Id\", \"Target\", \"number_of_targets\"], axis=1).values\n            \n    def __data_generation(self, list_IDs_temp):\n        'Generates data containing batch_size samples' # X : (n_samples, *dim, n_channels)\n        # Initialization\n        X = np.empty((self.batch_size, *self.dim, self.n_channels))\n        y = np.empty((self.batch_size, self.num_classes), dtype=int)\n        # Generate data\n        for i, identifier in enumerate(list_IDs_temp):\n            # Store sample\n            image = self.preprocessor.load_image(identifier)\n            image = self.preprocessor.preprocess(image)\n            X[i] = image\n            # Store class\n            y[i] = self.get_targets_per_image(identifier)\n        return X, y\n    \n    def __len__(self):\n        'Denotes the number of batches per epoch'\n        return int(np.floor(len(self.list_IDs) / self.batch_size))\n    \n    def __getitem__(self, index):\n        'Generate one batch of data'\n        # Generate indexes of the batch\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        # Find list of IDs\n        list_IDs_temp = [self.list_IDs[k] for k in indexes]\n        # Generate data\n        X, y = self.__data_generation(list_IDs_temp)\n        return X, y","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7867efb076524384d57556b13283aeb45b0ea1dd"},"cell_type":"code","source":"class PredictGenerator:\n    \n    def __init__(self, predict_Ids, imagepreprocessor, predict_path):\n        self.preprocessor = imagepreprocessor\n        self.preprocessor.basepath = predict_path\n        self.identifiers = predict_Ids\n    \n    def predict(self, model):\n        y = np.empty(shape=(len(self.identifiers), self.preprocessor.parameter.num_classes))\n        for n in range(len(self.identifiers)):\n            image = self.preprocessor.load_image(self.identifiers[n])\n            image = self.preprocessor.preprocess(image)\n            image = image.reshape((1, *image.shape))\n            y[n] = model.predict(image)\n        return y","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"48e03c679baabe6f4a57cb9da335ccb6d6f9710b"},"cell_type":"markdown","source":"### CNN Baseline model using keras"},{"metadata":{"trusted":true,"_uuid":"76a48f7a606ef5293bfc465871c707ce9b3900a8","_kg_hide-input":true},"cell_type":"code","source":"from keras.models import Sequential\nfrom keras.layers import Dense, Dropout, Flatten\nfrom keras.layers import Conv2D, MaxPooling2D\nfrom keras.losses import binary_crossentropy\nfrom keras.optimizers import Adadelta\nfrom keras.models import load_model\n\n\nclass BaseLineModel:\n    \n    def __init__(self, modelparameter):\n        self.params = modelparameter\n        self.num_classes = self.params.num_classes\n        self.img_rows = self.params.scaled_row_dim\n        self.img_cols = self.params.scaled_col_dim\n        self.n_channels = self.params.n_channels\n        self.input_shape = (self.img_rows, self.img_cols, self.n_channels)\n        self.my_metrics = ['accuracy']\n    \n    def build_model(self):\n        self.model = Sequential()\n        self.model.add(Conv2D(16, kernel_size=(3, 3), activation='relu', input_shape=self.input_shape))\n        self.model.add(Conv2D(32, (3, 3), activation='relu'))\n        self.model.add(MaxPooling2D(pool_size=(2, 2)))\n        self.model.add(Dropout(0.25))\n        self.model.add(Flatten())\n        self.model.add(Dense(64, activation='relu'))\n        self.model.add(Dropout(0.5))\n        self.model.add(Dense(self.num_classes, activation='sigmoid'))\n    \n    def compile_model(self):\n        self.model.compile(loss=keras.losses.binary_crossentropy,\n              optimizer=keras.optimizers.Adadelta(),\n              metrics=self.my_metrics)\n    \n    def set_generators(self, train_generator, validation_generator):\n        self.training_generator = train_generator\n        self.validation_generator = validation_generator\n    \n    def learn(self):\n        return self.model.fit_generator(generator=self.training_generator,\n                    validation_data=self.validation_generator,\n                    epochs=self.params.n_epochs, \n                    use_multiprocessing=True,\n                    workers=8)\n    \n    def score(self):\n        return self.model.evaluate_generator(generator=self.validation_generator,\n                                      use_multiprocessing=True, \n                                      workers=8)\n    \n    def predict(self, predict_generator):\n        y = predict_generator.predict(self.model)\n        return y\n    \n    def save(self, modeloutputpath):\n        self.model.save(modeloutputpath)\n    \n    def load(self, modelinputpath):\n        self.model = load_model(modelinputpath)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"79ae4ddf40420cb60ed564e515dc6a8fe3fd2d82"},"cell_type":"markdown","source":"### Training the baseline on the first cv-fold"},{"metadata":{"trusted":true,"_uuid":"819cb636a11da079e2f2a387a8cad6c3a7bd3105"},"cell_type":"code","source":"# Datasets\npartition = partitions[0]\nlabels = train_labels\n\nprint(\"Number of samples in train: {}\".format(len(partition[\"train\"])))\nprint(\"Number of samples in validation: {}\".format(len(partition[\"validation\"])))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"453d8dc45e28db3d49ed281dd698f03fe1aa1cb8"},"cell_type":"code","source":"training_generator = DataGenerator(partition['train'], labels, parameter, preprocessor)\nvalidation_generator = DataGenerator(partition['validation'], labels, parameter, preprocessor)\npredict_generator = PredictGenerator(partition['validation'], preprocessor, train_path)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6bdc3802ae488499de2afacb32e2e93bee5dccf4"},"cell_type":"code","source":"# Run computation and store results as csv\nif kernelsettings.fit_baseline == True:\n    model = BaseLineModel(parameter)\n    model.build_model()\n    model.compile_model()\n    model.set_generators(training_generator, validation_generator)\n    history = model.learn()\n    #model.save(\"baseline_model.h5\")\n    proba_predictions = model.predict(predict_generator)\n    baseline_proba_predictions = pd.DataFrame(proba_predictions, columns=train_labels.drop(\n        [\"Target\", \"number_of_targets\", \"Id\"], axis=1).columns)\n    baseline_proba_predictions.to_csv(\"baseline_predictions.csv\")\n# If you already have done a baseline fit once, \n# you can load predictions as csv and further fitting is not neccessary:\nelse:\n    baseline_proba_predictions = pd.read_csv(\"../input/protein-atlas-eab-predictions/baseline_predictions.csv\", index_col=0)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"98c5206497d690c27fca5fb0cc2866ab31b3f673"},"cell_type":"markdown","source":"Even though this accuracy looks nice it's an illusion! We are far away from a good model. Let's try to understand why..."},{"metadata":{"_uuid":"5e093fa4cd74f6ceda7f395c2017a39a335f2fe9"},"cell_type":"markdown","source":"## What do the results tell us?\n\nLet's have a look at predicted probabilites per target class:"},{"metadata":{"trusted":true,"_uuid":"8ed1e7784c82006e8aa08eb1e18a06cd4af8ecd8"},"cell_type":"code","source":"validation_labels = train_labels.loc[train_labels.Id.isin(partition[\"validation\"])]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"94b4239cfb2f1c62b66bb685c503ba668fb4a3e2"},"cell_type":"markdown","source":"With our new strategy to perform predictions we should obtain the same shape[0] of the true targets and the predicted ones:"},{"metadata":{"trusted":true,"_uuid":"22fb79c1ff6d84f7403ad26418b66669e73fab6a"},"cell_type":"code","source":"print(validation_labels.shape)\nprint(baseline_proba_predictions.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"175383cd82e42c699b163e0f533a58655b4c3b85"},"cell_type":"markdown","source":"Hmmmm this still looks not good :-( Have to search even further. "},{"metadata":{"trusted":true,"_uuid":"127f03493b264398faaa55e2d5a81251619059ea"},"cell_type":"code","source":"baseline_proba_predictions.tail()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1bdefea146c0c50bc9d78fb2ddf41c2f9e4b1f14"},"cell_type":"code","source":"proba_predictions = baseline_proba_predictions.values","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"851060fb988a7942a3325349e9999f57f2a6225f"},"cell_type":"code","source":"hot_values = validation_labels.drop([\"Id\", \"Target\", \"number_of_targets\"], axis=1).values.flatten()\none_hot = (hot_values.sum()) / hot_values.shape[0] * 100\nzero_hot = (hot_values.shape[0] - hot_values.sum()) / hot_values.shape[0] * 100\n\nfig, ax = plt.subplots(1,2, figsize=(20,5))\nsns.distplot(proba_predictions.flatten() * 100, color=\"DodgerBlue\", ax=ax[0])\nax[0].set_xlabel(\"Probability in %\")\nax[0].set_ylabel(\"Density\")\nax[0].set_title(\"Predicted probabilities\")\nsns.barplot(x=[\"label = 0\", \"label = 1\"], y=[zero_hot, one_hot], ax=ax[1])\nax[1].set_ylim([0,100])\nax[1].set_title(\"True target label count\")\nax[1].set_ylabel(\"Percentage\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4ee24c3b3799e06b86d54c5a24323b6b5902d0ed"},"cell_type":"markdown","source":"### Take-Away\n\n* We can see that our model was always very uncertain to predict the presence of a target protein. All probabilities are close to zero and there are only a few with targets where our model predicted a protein structure with higher than 10 %.\n* If we take a look at the true target label count we can see that most of our targets are filled with zero. This corresponds to an absence of corresponding target proteins. This makes sense: For each image we have a high probability to contain either 1 or 2 target protein structures. Their label values are one whereas all others are zero. \n* Consequently our high accuracy belongs to the high correct prediction of the absence of target proteins. In contrast we weren't able to predict the presence of a target protein which is the most relevant part! \n* Now a bell should ring :-) Have you ever heard about imbalanced classes and model confusion? "},{"metadata":{"_uuid":"6475127409a86d7c4da3544478456d0d17c148eb"},"cell_type":"markdown","source":"### To which targets do the high and small predicted probabilities belong to?"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"7174d05d4e824c290df5149e6990827a076755a7"},"cell_type":"code","source":"mean_predictions = np.mean(proba_predictions, axis=0)\nstd_predictions = np.std(proba_predictions, axis=0)\nmean_targets = validation_labels.drop([\"Id\", \"Target\", \"number_of_targets\"], axis=1).mean()\n\nlabels = validation_labels.drop([\"Id\", \"Target\", \"number_of_targets\"], axis=1).columns.values\n\nfig, ax = plt.subplots(1,2,figsize=(20,5))\nsns.barplot(x=labels,\n            y=mean_predictions,\n            ax=ax[0])\nax[0].set_xticklabels(labels=labels,\n                      rotation=90)\nax[0].set_ylabel(\"Mean predicted probability\")\nax[0].set_title(\"Mean predicted probability per class over all samples\")\nsns.barplot(x=labels,\n           y=std_predictions,\n           ax=ax[1])\nax[1].set_xticklabels(labels=labels,\n                      rotation=90)\nax[1].set_ylabel(\"Standard deviation\")\nax[1].set_title(\"Standard deviation of predicted probability per class over all samples\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b058c1f00898dad04957433672e726822426e478"},"cell_type":"code","source":"fig, ax = plt.subplots(1,1,figsize=(20,5))\nsns.barplot(x=labels, y=mean_targets.values, ax=ax)\nax.set_xticklabels(labels=labels,\n                      rotation=90)\nax.set_ylabel(\"Percentage of hot (1)\")\nax.set_title(\"Percentage of hot counts (ones) per target class\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e3717835b45904306c0b00f66949ef8f8f7f86a2"},"cell_type":"markdown","source":"### Take-Away\n\n* Our baseline model seemed to learn something even if this something does not look very nice. \n* Taking a look at the standard deviation we can see that all samples have nearly the same predicted values. There is no deviation, no difference between them. This is of course very bad! :-(\n\nLet's go one step deeper and take a look at the Cytosol (choose another feature if you like ;-)). Here we can see a higher standard deviation than for all other samples and perhaps its corresponding distribution starts to diverge, trying to get bimodal. This would be great at it indicates that the model starts solving the problem of binary classification for this target:"},{"metadata":{"trusted":true,"_uuid":"7a42ae9d8ae2881ab558c3b140588e961ac17545"},"cell_type":"code","source":"feature = \"Cytosol\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d96f5b4f4b62fb463bbe2266998d09e759289c37"},"cell_type":"code","source":"plt.figure(figsize=(20,5))\nsns.distplot(baseline_proba_predictions[feature].values[0:-10], color=\"Purple\")\nplt.xlabel(\"Predicted probabilites of {}\".format(feature))\nplt.ylabel(\"Density\")\nplt.xlim([0,1])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ee00e75c559f74b7f1b50c8f55bf1e7c22107591"},"cell_type":"markdown","source":"No, it does not seem that our model starts to separate well. The mode is close to the fraction of one-hot-counts over all samples. At least the flat tail gives hope that learning could be in progress.  But even though our next goal should be to find out what to tune in such a way that our model really starts learning! "},{"metadata":{"_uuid":"2ae9aa27baa88d6a0eaf913e9215339881ad6444"},"cell_type":"markdown","source":"## One-Step-Improvement\n\n\nOk, again let's go one step back and choose the most common target proteins that are present in our data: nucleoplasm, cytosol and plasma membrane. If we are not able to predict them we can go home and stay in bed ;-) ."},{"metadata":{"trusted":true,"_uuid":"bd5d3dc8105a731c940e1996241fa1b0e32fcec1"},"cell_type":"code","source":"wishlist = [\"Nucleoplasm\", \"Cytosol\", \"Plasma membrane\"]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e9c23453f175fc3768f4b94b659c96f8fb72c547"},"cell_type":"markdown","source":"### The target wish list\n\nTo introduce a target wishlist that we can change whenever we want we need to improve the data generator. For this purpose we're going to extend the class we have already written. Taking a closer look at the base generator you can see that there is just one line code in def data_generation(self, list_IDs_temp) we have to change, namely the part with y[i] = ... inside the for loop over temp list ids (image identifiers of the batch). To make things easier, I added a small method to the DataGenerator we already had:\n\n```\ndef get_targets_per_image(self, identifier):\n        return self.labels.loc[self.labels.Id==identifier].drop(\n                [\"Id\", \"Target\", \"number_of_targets\"], axis=1).values\n```\n\nThis method just avoids the direct pass to y[i], the targets per image in a batch. Now, we can overwrite this method in our ImprovedDataGenerator without loosing functionality:"},{"metadata":{"trusted":true,"_uuid":"24b1f61b38bc5ec85150f0ffbeb1000d2b4c4ee2"},"cell_type":"code","source":"class ImprovedDataGenerator(DataGenerator):\n    \n    # in contrast to the base DataGenerator we add a target wishlist to init\n    def __init__(self, list_IDs, labels, modelparameter, imagepreprocessor, target_wishlist):\n        super().__init__(list_IDs, labels, modelparameter, imagepreprocessor)\n        self.target_wishlist = target_wishlist\n    \n    def get_targets_per_image(self, identifier):\n        return self.labels.loc[self.labels.Id==identifier][self.target_wishlist].values","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b3577681a7dc0efaffc9773706b2ecad4a88499b"},"cell_type":"markdown","source":"### Plug and Play\n\n:-) \n\nThis part makes fun! Actually I don't know if our model learns something meaningful. But we can try to find out by improving our model adding **new features** and playing with **different parameter settings**. The latter I like to do in a **plug-and-play style**: Setting the flag *improve=True* adds a change to our model whereas *improve=False* uses the old concept we already used in the baseline. \n\n"},{"metadata":{"_uuid":"66ad84e730aea51fed90d03a40e6615726ea88fb"},"cell_type":"markdown","source":"### Add scoring metrics\n\nWe have already seen that the accuracy score is an illusion and does not help to figure out how good our predictions are. Let's take a closer look to the competition scoring and alternatives:\n\n* **F1 macro score**: Check out this [nice implementation of Guglielmo Camporese](https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras). Thank you very much! We can easily add it to our model.\n* But even with that score we should be careful! We have 28 different classes that are **very different in their frequency of being present**.  In addition we have to deal with **highly imbalanced classes per single target**. Even for the most common target nucleoplasm there are only 40 % of samples that show it and 60 % not. This imbalance becomes even more dramatic for seldom targets like rods and rings. We should **attach more importance to true positives**. \n* Well, there is **one problem with the competition score** that we can use to measure the performance of our model: **The f1 mean**. The mean is not robust towards outliers and consequently not very informative to understand the distribution of f1 scores for each target class. Perhaps we will make nice predictions for Cytosol but bad ones for Nucleoplasmn... who knows? Consequently it could be nice to introduce some further statistical quantities like **min** as well as **max** and the **standard deviation** . This way we can see the worst and gain some insights how the scores are spread over the classes. "},{"metadata":{"trusted":true,"_uuid":"b1d64d6e4dbc439fb135cd77d0d0ef2d65706c96","_kg_hide-input":true},"cell_type":"code","source":"import keras.backend as K\n\ndef base_f1(y_true, y_pred):\n    y_pred = K.round(y_pred)\n    tp = K.sum(K.cast(y_true*y_pred, 'float'), axis=0)\n    tn = K.sum(K.cast((1-y_true)*(1-y_pred), 'float'), axis=0)\n    fp = K.sum(K.cast((1-y_true)*y_pred, 'float'), axis=0)\n    fn = K.sum(K.cast(y_true*(1-y_pred), 'float'), axis=0)\n\n    p = tp / (tp + fp + K.epsilon())\n    r = tp / (tp + fn + K.epsilon())\n\n    f1 = 2*p*r / (p+r+K.epsilon())\n    f1 = tf.where(tf.is_nan(f1), tf.zeros_like(f1), f1)\n    return f1\n\ndef f1_min(y_true, y_pred):\n    f1 = base_f1(y_true, y_pred)\n    return K.min(f1)\n\ndef f1_max(y_true, y_pred):\n    f1 = base_f1(y_true, y_pred)\n    return K.max(f1)\n\ndef f1_mean(y_true, y_pred):\n    f1 = base_f1(y_true, y_pred)\n    return K.mean(f1)\n\ndef f1_std(y_true, y_pred):\n    f1 = base_f1(y_true, y_pred)\n    return K.std(f1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"43cc22388e3799c8387fe35d35f865928924a1c4"},"cell_type":"markdown","source":"### Track losses and scores\n\nOk, after adding the metrics we like to observe we should try to obtain more insights into the learning process of our model. One question on my mind draws circles: **What happened to the loss after each batch during one epoch?** Does it converge? Has our model started to learn or does nothing happen? Currently we obtain a history after calling fit_generator, but this history only contains the loss of train and validation data after one epoch. It does not contain losses that are obtained after computing each batch. But wait a minute... **How does our model update gradients? After each batch? After one epoch?** \n\nIf we take a look at simple feedforward networks that are close related to CNNs add gradient descent, we can see that learning means to compute the derivatives of the loss with respect to the weights over all samples:\n\n$$ w_{i,j}^{new} = w_{i,j}^{old} - \\eta \\cdot \\partial_{w_{i,j}} E$$\n\nWith a set of independent observation samples, we can obtain the gradients this way:\n\n$$ \\partial_{w_{i,j}} E = \\sum_{n=1}^{N} \\partial_{w_{i,j}} E_{n}$$\n"},{"metadata":{"_uuid":"6f0edac71f3ad51f238324ff2c573a6b6af792ad"},"cell_type":"markdown","source":"We are computing the change of the loss with respect to a change in the weights for each sample one after another. Consequently in original gradient descent we need to pass the whole dataset once for just one single update step of gradient descent. As our initial weights are not sufficent to solve the classification task we need many such update steps. **But what if the dataset it too large to wait a long time just for one of these steps?**\n\nPerhaps it would be sufficient and already good to use only some $M < N$ of the $N$ samples to compute the gradients $\\partial_{w_{i,j}} E$:\n\n$$ \\partial_{w_{i,j}} E = \\sum_{m=1}^{M} \\partial_{w_{i,j}} E_{m} $$\n\nThis way **we could use only a batch of samples, compute the gradients and perform the update of weights**. After that we continue by computing the gradients with the next batch. This could be done in sequence of samples or random with replacement. This means that we already start learning without passing the whole dataset. In my case I chose a batch_size of 200. With a total number of samples of around 20000 in my training set this means that I have already done 200 weight updates. **After doing so I have used each sample once and this means the whole dataset was passed through gradient descent**. As far as I know this means, **one epoch**. In addition we can now say that we want to shuffle the samples before doing weight updates after each batch again in the next epoch. This way our model sees a different set, computes some other gradients and hence the loss minimziation could be more robust.   "},{"metadata":{"_uuid":"d404bced8febfc3a8b5e3900a539e0e630c74d65"},"cell_type":"markdown","source":"Improve! :-) We can already see that we might need **more than 1 epoch to learn the pattern in the data**. Computing the gradient with batches is a stochastic process: Depending on the batch samples the gradients may sometimes lead to some good update directions in weight space that points to some minimum of loss.... and sometimes they could lead to the opposite. Well this really depends on the samples within one batch. This becomes more robust with **more samples per batch.** But then we will **need even more epochs** as we make less weight updata aka learning steps! But even if this stochastic gradient descent looks wiggly it has an advantage too: It can escape from local minima of the loss function. That's nice. \n\nBefore we improve the model using more epochs and smaller batches, let's **implement a new feature: a loss callback**. This way we can see if the loss decreases during weight update steps! And we can see if we need more epochs or if the loss has already converged and settled down. "},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"0a007ca6bde3b124e400c0b1b89e6075b3d8c5b0"},"cell_type":"code","source":"class TrackHistory(keras.callbacks.Callback):\n    \n    def on_train_begin(self, logs={}):\n        self.losses = []\n\n    def on_batch_end(self, batch, logs={}):\n        self.losses.append(logs.get('loss'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0b763fe315bbc0960599bb4149bce9f600ef45f3","_kg_hide-input":true},"cell_type":"code","source":"class ImprovedModel(BaseLineModel):\n    \n    def __init__(self, modelparameter,\n                 use_dropout,\n                 my_metrics=[f1_mean, f1_std, f1_min, f1_max]):\n        \n        super().__init__(modelparameter)\n        self.my_metrics = my_metrics\n        self.use_dropout = use_dropout\n        \n    def learn(self):\n        self.history = TrackHistory()\n        return self.model.fit_generator(generator=self.training_generator,\n                    validation_data=self.validation_generator,\n                    epochs=self.params.n_epochs, \n                    use_multiprocessing=True,\n                    workers=8,\n                    callbacks = [self.history])\n    \n    def build_model(self):\n        self.model = Sequential()\n        self.model.add(Conv2D(16, kernel_size=(3, 3), activation='relu', input_shape=self.input_shape))\n        self.model.add(Conv2D(32, (3, 3), activation='relu'))\n        self.model.add(MaxPooling2D(pool_size=(2, 2)))\n        if self.use_dropout:\n            self.model.add(Dropout(0.25))\n        self.model.add(Flatten())\n        self.model.add(Dense(64, activation='relu'))\n        if self.use_dropout:\n            self.model.add(Dropout(0.5))\n        self.model.add(Dense(self.num_classes, activation='sigmoid'))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c1ef7652070b897e274ee3adbb0a8e1afae2bc6a"},"cell_type":"markdown","source":"Ok, now let's increase the number of epochs and decrease the batch_size. This way we use more weight update steps and hopefully makes our model learn more than before: "},{"metadata":{"trusted":true,"_uuid":"42c9c106d61b9abebfa072a2ff3b731788ac81a1"},"cell_type":"code","source":"parameter = ModelParameter(train_path, num_classes=len(wishlist), n_epochs=5, batch_size=64)\npreprocessor = ImagePreprocessor(parameter)\nlabels = train_labels","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"50c9977084282e37842555172e3484aa769f95fc"},"cell_type":"code","source":"training_generator = ImprovedDataGenerator(partition['train'], labels,\n                                           parameter, preprocessor, wishlist)\nvalidation_generator = ImprovedDataGenerator(partition['validation'], labels,\n                                             parameter, preprocessor, wishlist)\npredict_generator = PredictGenerator(partition['validation'], preprocessor, train_path)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"85e7b6b70322d7d0b7739ebe75455e0275379b32","_kg_hide-input":true},"cell_type":"code","source":"# Run computation and store results as csv\nif kernelsettings.fit_improved_baseline == True:\n    model = ImprovedModel(parameter, use_dropout=use_dropout)\n    model.build_model()\n    model.compile_model()\n    model.set_generators(training_generator, validation_generator)\n    epoch_history = model.learn()\n    proba_predictions = model.predict(predict_generator)\n    #model.save(\"improved_model.h5\")\n    improved_proba_predictions = pd.DataFrame(proba_predictions, columns=wishlist)\n    improved_proba_predictions.to_csv(\"improved_predictions.csv\")\n# If you already have done a baseline fit once, \n# you can load predictions as csv and further fitting is not neccessary:\nelse:\n    improved_proba_predictions = pd.read_csv(\"../input/protein-atlas-eab-predictions/improved_predictions.csv\", index_col=0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"09ecd6d62d6a725d39bf9f52bbc4ef882cfbd52a"},"cell_type":"code","source":"if kernelsettings.fit_improved_baseline == True:\n    print(epoch_history.history.keys())\n    fig, ax = plt.subplots(2,1,figsize=(20,10))\n    ax[0].plot(model.history.losses, color=\"Red\")\n    ax[0].set_xlabel(\"Update step after one batch\")\n    ax[0].set_ylabel(\"Train loss\")\n    ax[0].set_title(\"Loss evolution batch per batch\")\n    ax[1].plot(epoch_history.history[\"val_loss\"], color=\"Green\")\n    ax[1].plot(epoch_history.history[\"loss\"])\n    ax[1].set_title(\"Loss evolution per epoch\")\n    ax[1].set_xlabel(\"Update step after one epoch\")\n    ax[1].set_ylabel(\"Loss\")\n    fig.savefig(\"losses_improved_model\", format=\"eps\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a9a9f5d6b095c92334d2f878a0dd16fc79157a3c"},"cell_type":"markdown","source":"### What does the loss tell us?\n![Losses](https://www.kaggleusercontent.com/kf/7185016/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..3D5GOiotTaFacOc_0ztMkQ.xELnVFi-FwGr8RjlVpZt65KXnyCK5GtztyHgErbUR1iSjxSROPp2jai_q_MHvZMSZ56KGVPQPtCx5j4AegtwsDtZRa8QAT3urcmvDW80qKcuApN7DJWnV7xJxGvfVIInJS_ml5Rma43kvo5NWA3OgXE9qfTJ_zG4C1p95FMUJN8.1bNcwgsup-EREVBNPymfFA/__results___files/__results___118_1.png)\n\nThe loss is very noisy! While decreasing the batch size we increased the number of learning steps. Hence our model learns faster. But... with smaller batch size there **are fewer samples to learn from, to compute gradients from**! The gradients we obtain may be very specific to the images and class labels that are covered by the batch of the current learning step. **There was a tradeoff we made**. We gained more learning speed but payed with a reduced gradient quality. Before increasing the batch size again and waiting too long for predictions we might improve by choosing another way:\n\n1. Weight regularization\n2. Gradient clipping\n\nThese two will be the next improvement steps. Nonetheless, one question remains: Has our model started learning? Can we see a separating force that tries to split zero and one predictions?"},{"metadata":{"_uuid":"a4b4695cf526a3cd4209e75e665eb71e2eacdb87"},"cell_type":"markdown","source":"### Does our model try to classify?\n\nIf this is the case and our model starts learning we should see more bimodal distributions of the predicted probability per target label:"},{"metadata":{"trusted":true,"_uuid":"47c0b8d3e90ea26ceed1298255db8688e2aab942","_kg_hide-input":true},"cell_type":"code","source":"fig, ax = plt.subplots(3,1,figsize=(25,15))\nsns.distplot(improved_proba_predictions.values[:,0], color=\"Orange\", ax=ax[0])\nax[0].set_xlabel(\"Predicted probabilites of {}\".format(improved_proba_predictions.columns.values[0]))\nax[0].set_xlim([0,1])\nsns.distplot(improved_proba_predictions.values[:,1], color=\"Purple\", ax=ax[1])\nax[1].set_xlabel(\"Predicted probabilites of {}\".format(improved_proba_predictions.columns.values[1]))\nax[1].set_xlim([0,1])\nsns.distplot(improved_proba_predictions.values[:,2], color=\"Limegreen\", ax=ax[2])\nax[2].set_xlabel(\"Predicted probabilites of {}\".format(improved_proba_predictions.columns.values[2]))\nax[2].set_xlim([0,1])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"504a4336ccac32dee731574f11c705e93292ea75"},"cell_type":"markdown","source":"Jeahy! This looks far better than the distributions we obtained with the baseline model! :-) **Thus increasing epochs while decreasing the batch size helped our model as we made more learning steps.** But... as we have seen by the noise of our losses as well we make large jumps in weight space with each update step. That's not nice! Sometimes these jumps could even lead to exploding losses as well. Then the jump was so big that we escaped from nice regions that lead to a local minimum of loss."},{"metadata":{"_uuid":"39049822930c3d8219f1552fdeefe17ab98bbe90"},"cell_type":"markdown","source":"## How can we tackle gradient jiggles?\n\nLet's try to dive deeper into the problem. We compute the gradients with respect to the weights after processing each batch this way:\n\n$$ \\partial_{w_{i,j}} E = \\sum_{m=1}^{M} \\partial_{w_{i,j}} E_{m} $$"},{"metadata":{"_uuid":"4fb02efae72d20df876cc9bb7ac45082d097f59b"},"cell_type":"markdown","source":"Looking at the sum we can see one disadvantage... it's mainly driven by high contributions. **An image in the batch that causes very high positive or negative gradients for the weight $w_{i,j}$ have more impact on the overall gradient than images with low absolute values**. This can be bad especially in the case of outlier images that are not representative to explain the pattern in the data. Consequently our model may try to learn from exotics. In addition we have to be very **careful with small batches as its target distribution might not reflect the overall pattern**. Imagine we would try to distinguish dogs from cats. With a batch size of 10 we are likely to fill up these places with imbalanced targets. For example it could be occupied with cats only. This would yield gradients that try to improve the detection of cats thereby changing the weights we might need to identify dogs. Hence beside image outliers the target distribution itself influences the learning as well. This can cause jiggles as well. One step we try to improve nucleoplasmn and the next perhaps cytosol but with a downgrade of the nucleoplasmn predictions and the next steps it could be the other way round. \n\nHow to solve this jiggle-wiggle problem?\n\nWell, first of all we might choose a higher batch_size again :-) Decreasing the batch size we made a tradeoff: We increased the learning speed but increased the risk of low quality gradients. Hence before playing with further strategies, we should make a step backwards again. With a batch_size of 128 and number of epochs 10, we obtain these losses:"},{"metadata":{"_uuid":"8162b0af9e6e4f75aad399868fb452ca51944082"},"cell_type":"markdown","source":"### Dropout Layer\n\nI often read that dropout helps to avoid overfitting but for me it seems that there is one more useful advantage: Imagine the cat-problem above - Given a batch full of cats we only compute gradients for making the predictions of cats better. And this can be done to the detriment of the dogs as both classes share weights. Hence changing the weights with batches of imbalanced classes can lead to jiggles. \n\nNow let's consider a dropout-layer: It randomly selects a given percentage of input neurons and drops them during the current training step. Taking a look at [this paper](http://www.jmlr.org/papers/volume15/srivastava14a/srivastava14a.pdf) (linked at keras documentation) you can see that the forward and backward propagation is done only with this reduced, thinned kind of network. In our case this means that each learning step some weights will be untouched and not used to compute gradients. This is great as this would mean that some weight that could be good for predicting dogs will not change after we process a batch with cats only. ;-)\n\nConsequently if you have a **problem with overfitting** (learning too much out of your training data and loosing generalization performance) **or you have very small, imbalanced batches you should consider dropout** as a strategy. But before turning happy and starting to use dropout frequently, we should think about its downside: Dropping neruons during training and learning with a thinned kind of networks means that **we freeze learning each step a bit**. And the randomness of frosty the dropout snowman can turn to a problem difficult to graps: Which neurons should be dropped, are there some for which dropping is good and some for which it is bad? After which layer in our network does it make sense to use it?Perhaps it would have been better to use some neurons that were dropped a random dropout session during one batch learning step... we don't know. Perhaps we have prevented the success of a learning step given one batch and improved learning given another batch. The information flow through the network is somehow a blackbox for us and this randomness of thinned network learning makes it more difficult to understand what's going on. \n\nYou can see that in my baseline model I'm already using dropout. This choice was somehow arbitrarily as I used an example network that can be found to classify mnist digits as a starting point. Hence let's improve again and **turn dropout to a plug-and-play feature**:\n\n```\nclass ImprovedModel(BaseLineModel):\n    \n    def __init__(self, modelparameter,\n                 my_metrics=[f1_mean, f1_std, f1_min, f1_max],\n                 use_dropout=True):\n        \n        super().__init__(modelparameter)\n        self.my_metrics = my_metrics\n        self.use_dropout = use_dropout\n\n    def build_model(self):\n        self.model = Sequential()\n        self.model.add(Conv2D(16, kernel_size=(3, 3), activation='relu', input_shape=self.input_shape))\n        self.model.add(Conv2D(32, (3, 3), activation='relu'))\n        self.model.add(MaxPooling2D(pool_size=(2, 2)))\n        if self.use_dropout:\n            self.model.add(Dropout(0.25))\n        self.model.add(Flatten())\n        self.model.add(Dense(64, activation='relu'))\n        if self.use_dropout:\n        self.model.add(Dropout(0.5))\n        self.model.add(Dense(self.num_classes, activation='sigmoid'))\n```"},{"metadata":{"_uuid":"a0fb6a83c1a1cc6730f9b852ae656cda111d09d7"},"cell_type":"markdown","source":"Our last losses were obtained by using dropout with high percentage (25 % and 50 %) of dropped neurons. **What will happen if we do not use our dropout layers?**\n\n![LossWithoutDropout](https://www.kaggleusercontent.com/kf/7212722/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..ZSQ_3svPXTPSUJWGcnFFjQ.IfPN45Kb77RLHEkt96pfewTm-O_B93fu-vw-weY_q8_pXlM8HdkyiFr1bhbJ9BmzzSckS4KvDKa2HNRb4aCJwa9btO4vwA4lUToAsCpRqZeHE5w6u4K1-AG_K6FWdDACoYvy7tP52dJku4zJsEySvTuoUt1qcb4xAsDoIlUMedg.z5L-MIPi65oUJmTJIRqXig/__results___files/__results___120_1.png)"},{"metadata":{"_uuid":"9a0fbfd1d2ab88ea2be51687311a5da8b6be7bc1"},"cell_type":"markdown","source":"## Where to go next?\n\n* optimization objective alias loss choice\n* hyperparameter search"},{"metadata":{"trusted":true,"_uuid":"22e06e9e8f271085b893854ecd25c07cd303bbd1"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c2ec53fb49b9595ab6887b402e2e51960746c468"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"268033bf99ade750313ee55cb3efe3cb575917ba"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"90e60b44ab05430575b18f5f297e93067e386144"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"577baa9e44ce1582afcb6902edfeb68419187362"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1e4a6bcc70f11c31c473e305ccd193acb7da3831"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3406284394ea3bb6b30909f230a44781194dd7ae"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4f753da3e464cafec57307339e6f4b5febcfa060"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"31470fc6e9680a5c3c0f132550a3996b180ec2f4"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"21bc52fcb39c21be3b370ede272ec1b307fa7f33"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dc1a988df0aa1067e64967cbadbb1eaeda37eb09"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cdf112650357806494a38b6a5998b60c265e8fd9"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2186f01c34ba03d6c3067004a17bcda8e2d9170f"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"76f72f002c3073ee1bc3c9760ff4bfe6d1d81aed"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"52f28ddf28d174bb319d7532636960893543f288"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"49057a2c9f15f6a1b6680e24217c703f668fc667"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7acac9a57ec69f8074445841fd8da73560c187fa"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"98bbc3f48f216bc6fa5f1d909fe0e14fec8369fc"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2515e907888777029d07dc3188cf262c6c2cfedf"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c20e5e5de0bc76fa4a599ee9e59094e4d898059e"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e760cad56cf540da639e8005ab5099d1bf4215e3"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bce86e094b052e2def9d97034d9fb93d78a07f4e"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f97080a9b177c29dec03f468fb46a7ab63d1dbec"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0f19d0f8f454b08f8b3aec272a150f5133329182"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9366c3e160b8273e362e2577f0a212cd53aacd8d"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2e9b6e6a4ec748947d8e758f520caa0a7b928654"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4ac21edf75e3ad5e7f33c0a6986ec84771f84330"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"854383d7bf2447728b6bf7dc5f0424709c57677d"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5a2e70f0a4298e3eefc891554adfa0ed03c472c4"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"30b722580b21a649e39f093a557074dbd46c2ea9"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"47b13e92f80c8f3e57707f344b3bb29e48b93fe0"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dd23b421be4e8901ab06e74ce40e41782edd2d97"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"927e97e177e908dab548174f5153206820f81545"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9414a97a45d05eacad6485479ec330f72a4628c3"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ede05d3e3e3dd187a6fe007706743499017e054c"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"91b144f0fc971b561b459cf1d59056ce26f1f9c2"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"74d0cffd2bc04c166313a3f066fef7bb20b1292e"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2eb45eab49d96d6d112df5802bb85263a311d9e5"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e75871684953707c2beff8a2dbae1fca51123c5d"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6ddd5b5a91d12c2fea634640a343e0a05b795b41"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}