{"cells":[{"metadata":{},"cell_type":"markdown","source":"This notebook covers the following topic\n\n- [Input Data Structure](#Input-Data-Structure)  \n- [Exploration](#Exploration)  \n    - [Q1. What is the occurrence of each label?](#Q1.)\n    - [Q2. How many labels do most images have?](#Q2.)\n    - [Q3. Are there any correlations between the pairs of labels?](#Q3.)\n\n**Credits** \n- https://www.kaggle.com/allunia/protein-atlas-exploration-and-baseline"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\nfrom PIL import Image\nimport os\nimport torch\nfrom torch.utils.data import Dataset, random_split, DataLoader\nimport torchvision.transforms as tt\nfrom torchvision.utils import make_grid","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Input Data Structure\n\nLet's take a look at the structure of the input folder"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"! ls /kaggle/input/human-protein-atlas-image-classification","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"DATA_DIR = '/kaggle/input/human-protein-atlas-image-classification'\n\nTRAIN_CSV = DATA_DIR + '/train.csv'\nTRAIN_DIR = DATA_DIR + '/train'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df = pd.read_csv(TRAIN_CSV)\ndisplay(df.head())\nprint(f\"df.shape: {df.shape}\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len(df['Id'].value_counts())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In the train.csv, we have 31072 unique image ids and their labels.\n\nIn the training image folder,\n```\n/kaggle/input/human-protein-atlas-image-classification/train\n```\neach Image Id is associated with the following four PNGs:  \n\n- the protein of interest (green) \n- nucleus (blue)  \n- microtubules (red)  \n- endoplasmic reticulum (yellow)\n\nThe labels indicate the localization of the protein of interest, which can be within multiple organelles, meaning this is a multi-label classification problem.\n\nNow let's add the integer to text mapping of the target."},{"metadata":{"trusted":true},"cell_type":"code","source":"text_labels = {\n0:  \"Nucleoplasm\", \n1:  \"Nuclear membrane\",   \n2:  \"Nucleoli\",   \n3:  \"Nucleoli fibrillar center\" ,  \n4:  \"Nuclear speckles\",\n5:  \"Nuclear bodies\",\n6:  \"Endoplasmic reticulum\",   \n7:  \"Golgi apparatus\",\n8:  \"Peroxisomes\",\n9:  \"Endosomes\",\n10:  \"Lysosomes\",\n11:  \"Intermediate filaments\",   \n12:  \"Actin filaments\",\n13:  \"Focal adhesion sites\",   \n14:  \"Microtubules\",\n15:  \"Microtubule ends\",   \n16:  \"Cytokinetic bridge\",   \n17:  \"Mitotic spindle\",\n18:  \"Microtubule organizing center\",  \n19:  \"Centrosome\",\n20:  \"Lipid droplets\",   \n21:  \"Plasma membrane\",   \n22:  \"Cell junctions\", \n23:  \"Mitochondria\",\n24:  \"Aggresome\",\n25:  \"Cytosol\",\n26:  \"Cytoplasmic bodies\",   \n27:  \"Rods & rings\" \n}\n\nNUM_LABELS = len(text_labels)\nprint(f\"There are {NUM_LABELS} labels\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We are ready to plot out the four channels of a random image id"},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"FILTERS = ['red', 'green', 'blue','yellow']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"def get_image(ddir, filename):\n    r = Image.open(f'{ddir}/{filename}_red.png')\n    g = Image.open(f'{ddir}/{filename}_green.png')\n    b = Image.open(f'{ddir}/{filename}_blue.png')\n    y = Image.open(f'{ddir}/{filename}_yellow.png')\n    return r, g, b, y\n\n\ndef display_image(image, ax):\n    [a.axis('off') for a in ax]\n    r, g, b, y = image\n    ax[0].imshow(r,cmap='Reds')\n    ax[0].set_title('Microtubules')\n    ax[1].imshow(g,cmap='Greens')\n    ax[1].set_title('Protein of Interest')\n    ax[2].imshow(b,cmap='Blues')\n    ax[2].set_title('Nucleus')\n    ax[3].imshow(y,cmap='Oranges') \n    ax[3].set_title('Endoplasmic Reticulum')\n    return ax","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"filename = df.Id.sample(1, random_state=9473).values[0]\nimgs = get_image(TRAIN_DIR, filename)\n\nfig, ax = plt.subplots(figsize=(15,5),nrows=1, ncols=4)\ndisplay_image(imgs, ax);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"For the image id, we are provided with staining of the protein of interest, microtubules, nucleus, and endoplasmic reticulum."},{"metadata":{},"cell_type":"markdown","source":"# Exploration\n\n## Q1.\n\n**What is the occurrence of each label?**"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"for key in text_labels.keys():\n    df[key] = df['Target'].apply(lambda x: int(str(key) in x.split()))\n\ntargets_df = df.drop(labels=['Id', 'Target'], axis=1)\n\n# targets_df.head()\n\ntarget_counts = pd.DataFrame({'Localization': [v + ' ' + str(k) for k, v in text_labels.items()],\n                              'Count': targets_df[text_labels.keys()].sum().values})\ntarget_counts.sort_values('Count', inplace=True)\nax = target_counts.plot.barh(x='Localization', y='Count',figsize=(15,10), legend=False)\n\nfor i, v in enumerate(target_counts['Count']):\n    ax.text(v + 3, i - 0.25, str(v) + ', ' + str(round(v / len(df) * 100, 2)) + '%')\nax.set_xlabel('Count');\nax.set_ylabel('');\n# plt.axis('off')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## A1. \n\nIt makes sense that nucleoplasm, cytosol, plasma membrane are the top occuring organelle with localized proteins, given that they are much larger in size comparing to other organelles. The plot above also agrees with Fig.2 (Protein distribution in the human cell) on [this page](https://www.proteinatlas.org/humanproteome/cell/organelle)\n\nFor details about each organelle, structure and substrucutre, check out this [awesome interactive graph](https://www.proteinatlas.org/humanproteome/cell)\n\nThe class imbalance means that we need to,\n- do train-validation split wisely\n- maybe resample the dataset\n- look into precision-recall curve\n- use loss function that giving FN more weight than TN, for example Focal Loss"},{"metadata":{},"cell_type":"markdown","source":"## Q2. \n\n**We know this is a multi-label classification problem. How many labels does the majority of images have?**"},{"metadata":{"trusted":true,"collapsed":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"targets_df['num_labels'] = targets_df.sum(axis=1)\ntargets_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"label_counts = pd.DataFrame({'image_count': targets_df['num_labels'].value_counts(),\n                             'pct_of_dataset': targets_df['num_labels'].value_counts() / len(df) * 100})\nlabel_counts.columns = ['image_count','pct_of_dataset']\nlabel_counts","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## A2 \n\nIt is most common for protein to localize in exactly 1 organelle (at 48.7%). The other 51.3% of the dataset are [multilocalizing proteins](https://www.proteinatlas.org/humanproteome/cell/multilocalizing) with 2 to up to 5 labels."},{"metadata":{},"cell_type":"markdown","source":"# Q3.\n\n**Are there any correlations between labels?**"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt.figure(figsize=(10, 10))\nsns.heatmap(targets_df[targets_df['num_labels']>1].drop(['num_labels'], axis=1).rename(\n    columns={k: f\"{v} ({str(k)})\" for k, v in text_labels.items()}\n).corr(), cmap='YlGnBu');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## A3:\n\nFrom the heatmap we can see that,\n\n- Endosomes (9) and Lysosomes (10) are higly correlated, and the two of them occassionaly show up together with Endoplasmatic Reticulum (6)\n\n- Mitotik spindle (17), Cytokinetic bridges (16) and Microtubules (14) are correlated."},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"class LocalizationDataLoader():\n    def __init__(self, labels, batch_size, ddir):\n        self.labels = labels\n        self.batch_size = batch_size\n        self.ddir = ddir\n        self.get_image_ids()\n    \n    \n    def are_labels_subset_of_targets(self, s):\n        targets = [int(i) for i in s.split()]\n        return np.where(set(self.labels).issubset(targets), 1, 0)\n    \n    \n    def get_image_ids(self):\n        df['check_condition'] = df.Target.apply(lambda s: self.are_labels_subset_of_targets(s))\n        self.identified_image_ids = df[df['check_condition'] == 1].Id.values\n        df.drop('check_condition', axis=1, inplace=True)\n        \n    \n    def get_loader(self):\n        idx = 0\n        batch_images = []\n        batch_image_ids = []\n        for image_id in self.identified_image_ids:\n            idx += 1\n            batch_images.append(get_image(self.ddir, image_id))\n            batch_image_ids.append(image_id)\n            if idx == self.batch_size:\n                yield batch_images, batch_image_ids\n                idx = 0\n                batch_images = []\n                batch_image_ids = []\n        if batch_images != []:\n            yield batch_images, batch_image_ids","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"def get_image_targets(image_id):\n    targets = df[df.Id==image_id].Target.values[0]\n    targets = ', '.join([f\"{text_labels[int(t)]} {t}\" for t in targets.split()])\n    return f\"{targets}\"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now let's take a look at a batch of images where the protein of interest is present in both the Endosomes and the Lysosomes"},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"batch_size = 5\nendo_lyso = LocalizationDataLoader([9, 10], batch_size, TRAIN_DIR)\nloader = endo_lyso.get_loader()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"imgs, img_ids = next(loader)\n\nfig, ax = plt.subplots(nrows=len(imgs), ncols=4, figsize=(15, 5 * len(imgs)))\nfor i, img in enumerate(imgs):\n    display_image(img, ax[i])\n#     ax[i][1].set_title(get_image_targets(img_ids[i]), y=-0.1)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}