{"cells":[{"metadata":{"_uuid":"7448c68b26712d0e22d961bb0434e24cb30f736b"},"cell_type":"markdown","source":"Throught this kernel I will show different images of some combination of labels on the training dataset of the Protein Atlas Image Classification problem in order to better understand what is what we are predicting.\n\nAll image samples are represented by four filters (stored as individual files), the protein of interest (green) plus three cellular landmarks: nucleus (blue), microtubules (red), endoplasmic reticulum (yellow). The green filter should hence be used to predict the label, and the other filters are used as references. \n\nI will print the images as RGB, using the yellow channel as the green one and I will use the green channel as a mask to determine what part of the cell have the desired protein."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"%matplotlib inline \nimport matplotlib.pyplot as plt\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport time\nimport copy\n\n\nfrom PIL import Image\n\nDATASET_SIZE = 3500\nBATCH_SIZE = 500\nW = H = 256\n\ntrain_path = '../input/train/'\ntest_path = '../input/test/'\n\nLABEL_MAP = {\n0: \"Nucleoplasm\" ,\n1: \"Nuclear membrane\"   ,\n2: \"Nucleoli\"   ,\n3: \"Nucleoli fibrillar center\",   \n4: \"Nuclear speckles\"   ,\n5: \"Nuclear bodies\"   ,\n6: \"Endoplasmic reticulum\"   ,\n7: \"Golgi apparatus\"  ,\n8: \"Peroxisomes\"   ,\n9:  \"Endosomes\"   ,\n10: \"Lysosomes\"   ,\n11: \"Intermediate filaments\"  , \n12: \"Actin filaments\"   ,\n13: \"Focal adhesion sites\"  ,\n14: \"Microtubules\"   ,\n15: \"Microtubule ends\"   ,\n16: \"Cytokinetic bridge\"   ,\n17: \"Mitotic spindle\"  ,\n18: \"Microtubule organizing center\",  \n19: \"Centrosome\",\n20: \"Lipid droplets\"   ,\n21: \"Plasma membrane\"  ,\n22: \"Cell junctions\"   ,\n23: \"Mitochondria\"   ,\n24: \"Aggresome\"   ,\n25: \"Cytosol\" ,\n26: \"Cytoplasmic bodies\",\n27: \"Rods & rings\"}\n\nLABELS = []\n\nfor label in LABEL_MAP.values():\n    LABELS.append(label)\n    \ntrain_csv_path = '../input/train.csv'","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"import torch\nimport random\nimport numpy as np\nfrom skimage import filters\nimport pandas as pd\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms, utils\nimport warnings\nwarnings.filterwarnings(\"ignore\")\nfrom skimage import io, transform\nfrom skimage.util import img_as_ubyte\nfrom skimage.transform import resize\nfrom sklearn.preprocessing import MultiLabelBinarizer\nfrom skimage.morphology import erosion, dilation, opening, closing, white_tophat\nfrom skimage.morphology import black_tophat, skeletonize, convex_hull_image\nfrom skimage.morphology import disk\nclasses = np.arange(0,28)\nmlb = MultiLabelBinarizer(classes)\nmlb.fit(classes)\n\nclass HumanProteinDataset(Dataset):\n\n    def __init__(self, csv_file,transform=None, test=False):\n        self.test = test\n        self.df = pd.read_csv(csv_file)\n        self.transform = transform\n        \n        if not test:\n            self.path = train_path\n            self.df['Targets'] = self.df['Target'].map(lambda x: list(map(int, x.strip().split())))\n\n        else:\n            self.path = test_path\n            \n    def load_image(path, image_id):\n        images = np.zeros(shape=(256,256,4))\n        r = resize(img_as_ubyte(io.imread(os.path.join(path+image_id+\"_red.png\"),as_gray=True)),(W,H))\n        g = resize(img_as_ubyte(io.imread(os.path.join(path+image_id+\"_green.png\"),as_gray=True)),(W,H))\n        b = resize(img_as_ubyte(io.imread(os.path.join(path+image_id+\"_blue.png\"),as_gray=True)),(W,H))\n        y = resize(img_as_ubyte(io.imread(os.path.join(path+image_id+\"_yellow.png\"),as_gray=True)),(W,H))\n\n        images[:,:,0] = np.asarray(r)\n        images[:,:,1] = np.asarray(g)\n        images[:,:,2] = np.asarray(b)\n        images[:,:,3] = np.asarray(y)\n\n        return images\n            \n    def __getitem__(self, idx):\n        image = HumanProteinDataset.load_image(self.path, self.df['Id'].iloc[idx])\n        sample = {'image': image}\n\n        if not self.test:\n            target = np.array(self.df['Targets'].iloc[idx])\n            target = mlb.transform([target])\n            sample['Target'] = target\n        else:\n            sample['Id'] = self.df['Id'].iloc[idx]\n\n        if self.transform:\n            sample = self.transform(sample)\n        return sample\n    \n    def IndexOfOne(self, targets):\n        fdf = self.df[self.df['Target'] == targets]\n        idx = random.randint(0,fdf.shape[0])\n        return fdf.index[idx]\n        \n    \n    def __len__(self):\n        return self.df.shape[0]\n    \n    def shape(self):\n        return self.df.shape\n    \n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d6a5f967b14e9d8201f56600b0495dcab3192b47"},"cell_type":"code","source":"dataset = HumanProteinDataset(train_csv_path)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1989e766fca717a2670f5a9bbdd399e45710760b"},"cell_type":"code","source":"def Show(sample):\n    f, (ax1,ax2,ax3,ax4) = plt.subplots(1, 4, figsize=(25,15), sharey=True)\n\n    title = ''\n    \n    labels = sample['Target'][0]\n                \n    for i, label in enumerate(LABELS):\n        if labels[i] == 1:\n            if title == '':\n                title += label\n            else:\n                title += \" & \" + label\n                \n    rgb = np.zeros([W,H,3])\n    \n    rgb[:,:,0] = sample['image'][:,:,0]\n    rgb[:,:,1] = sample['image'][:,:,3] # green channel will be the yellow one\n    rgb[:,:,2] = sample['image'][:,:,2]\n    \n    protein = sample['image'][:,:,1]\n    \n    ax2.imshow(rgb)\n    ax2.set_title('Reference')\n    ax1.imshow(protein)\n    ax1.set_title('Protein')\n    \n    protein = filters.gaussian(protein, sigma= 0.8)\n    protein = (protein - protein.min())/(protein.max() - protein.min())\n    #protein = protein > 0.95 * protein.mean()\n    \n    rgb[:,:,0] *= protein\n    rgb[:,:,1] *= protein\n    rgb[:,:,2] *= protein\n    \n    ax3.imshow(rgb)\n    ax3.set_title('Reference (filterd)')\n    ax4.imshow(protein)\n    ax4.set_title('Protein mask')\n    f.suptitle(title, fontsize=15, y=0.68)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"236451830c5b655266e720ea683b57f9f4b3d86c"},"cell_type":"markdown","source":"# just Nucleoplasm (0)\n\n![nucleoplasm](https://upload.wikimedia.org/wikipedia/commons/thumb/3/38/Diagram_human_cell_nucleus.svg/462px-Diagram_human_cell_nucleus.svg.png)\n![nucleoplasm images](https://i.imgur.com/MrX7Nvn.png)\n\nThe nucleoplasm is one of the types of protoplasm, and it is enveloped by the nuclear membrane (also known as the nuclear envelope). The nucleoplasm includes the chromosomes and nucleolus. \n\nIf we visualice a fews samples labeled as Nucleoplasm we can see that they are quite different... what they have in common is that, in all of them, the protein is surrounded by something (nuclear membrane) and the protein appear arround all the nucleous of different cells. In some images that nucleous is circular in others is rather oval... but, in all of them the protein not appear in the whole cell but only in a limited zone of it (nucleous). With out protein mask we can see it better.\n\n¿What need our model to learn in order to classify that?"},{"metadata":{"trusted":true,"_uuid":"55aece09385a9ed88364208bd5abbbbe56a25057"},"cell_type":"code","source":"for i in range(8):\n    idx = dataset.IndexOfOne('0')\n    Show(dataset[idx])\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7a8346c93f5a5c771bcb267301e9db672e29fac2"},"cell_type":"markdown","source":"# just Cytosol (25)\n\n![cytosol](https://upload.wikimedia.org/wikipedia/commons/thumb/7/71/Simple_diagram_of_yeast_cell_%28es%29.svg/800px-Simple_diagram_of_yeast_cell_%28es%29.svg.png)\n![cytosol](https://i.imgur.com/2wKXDqn.png)\n\nThe cytosol, also known as intracellular fluid (ICF) or cytoplasmic matrix, is the liquid found inside cells.[2] It is separated into compartments by membranes. For example, the mitochondrial matrix separates the mitochondrion into many compartments. \n\nIn this case the protein is in the cytosol which is the liquid found inside cells that contains all the organelles... That samples are quite similat to nucleoplasm ones because, as nucleoplasm is what are inside nucleous, cytosol is what is inside the cell membrane (all the cell) and they look similar. The main difference is that, inside cytosol, there is the nucleous... so, as our protein isn't present at nucleous, in that samlples the nucleous is black on the filtered reference."},{"metadata":{"trusted":true,"_uuid":"43f6e9cab62976e32a882d6d4632ef74bf3df5bb"},"cell_type":"code","source":"for i in range(8):\n    idx = dataset.IndexOfOne('25')\n    Show(dataset[idx])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c2e041414316446bf5930107391737fb21828a26"},"cell_type":"markdown","source":"# Both, Cytosol and Nucleoplasm (25 0)\n\n![both cytosol and nucleoplasm](https://i.imgur.com/iXxpySu.png)\n\nThis is a sample where the protein appears in two different parts of the cell, both cytosol and nucleoplasm. In this case we can see how our protein mask keept the whole cell and we can distinguish the blue part (nucleoud) and the yellow one (cytosol) in most cases."},{"metadata":{"trusted":true,"_uuid":"3ec5899880a195b4f8e69e467554e8c1c03442a0"},"cell_type":"code","source":"for i in range(8):\n    idx = dataset.IndexOfOne('25 0')\n    Show(dataset[idx])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5390aed554d4cd4669748e41b34b1606eeeac70a"},"cell_type":"markdown","source":"# Mitochondria (23)\n![mytochondria](https://upload.wikimedia.org/wikipedia/commons/0/0c/Mitochondria%2C_mammalian_lung_-_TEM.jpg)\n\nMitochondrias are a double-membrane-bound organelle commonly between 0.75 and 3 μm in diameter[5] but vary considerably in size and structure. Unless specifically stained, they are not visible. In addition to supplying cellular energy, mitochondria are involved in other tasks, such as signaling, cellular differentiation, and cell death, as well as maintaining control of the cell cycle and cell growth. They are in the cytoplasm embedded by cytosol.  In this kind of samples we can have a picture of a entire cell and small scattered dots that represent mitochondrias:\n\n![mitochondria](https://i.imgur.com/LVvOjWJ.png)\n\nor "},{"metadata":{"trusted":true,"_uuid":"12e5c6bf49d355ad2503f1b3a8b9303b93226d27"},"cell_type":"code","source":"for i in range(8):\n    idx = dataset.IndexOfOne('23')\n    Show(dataset[idx])","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}