{"cells":[{"metadata":{"_uuid":"730d52725d43d80f8658b684e3ef5fad5030c73b"},"cell_type":"markdown","source":"This Kernel contains basic analysis of data that includes the following:\n\n1) Converting *train.csv*'s \"Target\" column from string to list [required for one-hot-encoding]\n\n2) One Hot Encoding of \"Target\" column.\n\n3) Frequency of protein occurance\n\n4) Correlations among protein types.\n\nNote: This kernal is a work in progress. I will update above description periodically.\n\nAcknowledgements:\n1) \"Protein Atlas - Exploration and Baseline\" Kernel by Allunia\n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train_labels = pd.read_csv(\"../input/train.csv\")\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b0bc9ecb407d333b4a4f9c6691bc7a77c0059c3c"},"cell_type":"code","source":"label_names = {\n    0:  \"Nucleoplasm\",  \n    1:  \"Nuclear membrane\",   \n    2:  \"Nucleoli\",   \n    3:  \"Nucleoli fibrillar center\",   \n    4:  \"Nuclear speckles\",\n    5:  \"Nuclear bodies\",   \n    6:  \"Endoplasmic reticulum\",   \n    7:  \"Golgi apparatus\",   \n    8:  \"Peroxisomes\",   \n    9:  \"Endosomes\",   \n    10:  \"Lysosomes\",   \n    11:  \"Intermediate filaments\",   \n    12:  \"Actin filaments\",   \n    13:  \"Focal adhesion sites\",   \n    14:  \"Microtubules\",   \n    15:  \"Microtubule ends\",   \n    16:  \"Cytokinetic bridge\",   \n    17:  \"Mitotic spindle\",   \n    18:  \"Microtubule organizing center\",   \n    19:  \"Centrosome\",   \n    20:  \"Lipid droplets\",   \n    21:  \"Plasma membrane\",   \n    22:  \"Cell junctions\",   \n    23:  \"Mitochondria\",   \n    24:  \"Aggresome\",   \n    25:  \"Cytosol\",   \n    26:  \"Cytoplasmic bodies\",   \n    27:  \"Rods & rings\"\n}","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5a365ca323236efc166f63d48f2ffba711442231"},"cell_type":"markdown","source":"**Converting Target: string to list **"},{"metadata":{"trusted":true,"_uuid":"7fa7d795bff64a8fd77f6fe6026f8e0489b93420"},"cell_type":"code","source":"\ntrain_labels_modified = train_labels.copy()\nfor i in range(train_labels_modified.shape[0]):\n    train_labels_modified.Target[i] = train_labels_modified.Target[i].split()\n    for k,j in enumerate(train_labels_modified.Target[i]):\n        train_labels_modified.Target[i][k] = label_names[int(j)]\n    \ntrain_labels_modified.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5d0d3dff65b88e8e0a80cf43d4d8448018086695"},"cell_type":"markdown","source":"**One Hot Encoding Target Column**\n\n(I am using MultiLabelBinarizer here. For one hot encoding using first principles, please refer Allunia's kernel mentioned in Acknowledgement.)"},{"metadata":{"trusted":true,"_uuid":"00dbb0ad8d3631e64a7c9ceda723fd7858c7341e"},"cell_type":"code","source":"from sklearn.preprocessing import MultiLabelBinarizer\nmlb = MultiLabelBinarizer()\ns = pd.DataFrame(mlb.fit_transform(train_labels_modified.Target),columns=mlb.classes_, index=train_labels_modified.index)\ntrain_labels_one_hot = train_labels_modified.join(s)\ntrain_labels_one_hot","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7d360ee714b34e3bb8706b312cc4b6f936deac6e"},"cell_type":"markdown","source":"**Protein Histogram: Frequency of occurence of proteins** "},{"metadata":{"trusted":true,"_uuid":"3c467f963a6645f0b8aa032b9614394ed2e94202"},"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\ntarget_counts = train_labels_one_hot.drop([\"Id\", \"Target\"],axis =1).sum(axis=0).sort_values(ascending=False)\nplt.figure(figsize=(10,10))\nsns.barplot(x = target_counts.values, y=target_counts.index, order=target_counts.index)\nplt.xlabel(\"Number of Occurences in training set\")\n#plt.ylabel(\"Protein Type\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"856cbd1a96967a86c2e494b685e34524751ff22e"},"cell_type":"markdown","source":"Above Plot tells us that \"Nucleoplasm\" is the most frequent protein followed by cytosol and so on.\n\nIt is worthwhile to see: In what percentage of training images/cells, a particular protein is identified"},{"metadata":{"trusted":true,"_uuid":"525cffdd4c06cafb8e73af1a397c7c0a90e41d2b"},"cell_type":"code","source":"plt.figure(figsize=(10,10))\n# Using the same code as above, but divided by number of training images to get percentage\nsns.barplot(x = (target_counts.values)/train_labels_modified.shape[0], y=target_counts.index, order=target_counts.index)\n\nplt.xlabel(\"Fraction of Occurences in training set\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a097e6789570fa060d9f241e75bdcd8a51fbb789"},"cell_type":"markdown","source":"Now it is clear that \"Nucleoplasm\" is present is about 40% of training images,\"cytosol\" in about 30% and so on.\n\nSo, even without any modeling, we know that predicting \"Nucleoplasm\" in a cell is an educated guess that will get us right ~40% of time!!"},{"metadata":{"trusted":true,"_uuid":"f0ec1f4daa6be1edf71825535e09f9bd9d385e85"},"cell_type":"code","source":"print(\"Number of training images with no protein identified:\", train_labels_modified.Target.isnull().sum())\n# every image has atleast one protein identified, so we don't have to worry about  missing data.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"58b5e44860680560b222eafc72aea1d195ae59a7"},"cell_type":"code","source":"# Distribution of number of proteins per image\n\noccurances  = [len(train_labels_modified.Target[i]) for i in range(train_labels_modified.shape[0])]\nplt.hist(occurances, align = \"left\",range = [0,5])\n\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d25bb8bd706690cc3dd75cfb30cce4065d1c8403"},"cell_type":"markdown","source":"**Observations**: Almost half of the training images have one protein and most of other half have 2 proteins.  "},{"metadata":{"_uuid":"1c98869cbf77cdcc560d0cdd88b27cf0d7861236"},"cell_type":"markdown","source":"**Are there any proteins that (almost always) occur in pairs ?**"},{"metadata":{"trusted":true,"_uuid":"cfef57f0c3b457078b08b980643bc78434630a4d"},"cell_type":"code","source":"plt.figure(figsize=(10,10))\nsns.heatmap(train_labels_one_hot.drop([\"Id\", \"Target\"],axis=1).corr(),cmap=\"RdYlBu\", vmin=-1, vmax=1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7923b28c2198dca4382eb494b8297e36c6eaae8a"},"cell_type":"markdown","source":"1) *endosomes* and *lysosomes* occur in pairs(most of the time).\n\n2)* mitotic spindle* and *cytokinetic bridge* occur in pairs(sometimes)"},{"metadata":{"trusted":true,"_uuid":"d50e94e67c0f3b36e2af4fa4bc7d08ba5834d764"},"cell_type":"code","source":"#print(os.listdir(\"../input/train\"))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d26a178f58c3e2c0e9e060775829371252b80082"},"cell_type":"markdown","source":"**Working on data loading and display**(Incomplete)"},{"metadata":{"trusted":true,"_uuid":"450e395683804e2d23f7fbd45363b63e326da5ad"},"cell_type":"code","source":"from skimage.io import imread\ndef load_image(image_id, path=\"../input/train/\"):\n    images = np.zeros((4,512,512))\n    images[0,:,:] = imread(path + image_id + \"_green\" + \".png\")\n    images[1,:,:] = imread(path + image_id + \"_red\" + \".png\")\n    images[2,:,:] = imread(path + image_id + \"_blue\" + \".png\")\n    images[3,:,:] = imread(path + image_id + \"_yellow\" + \".png\")\n    return images\n#load_image(image_id = \"000a6c98-bb9b-11e8-b2b9-ac1f6b6435d0\")\ndef display_image_row(image, axis, title):\n    axis[0].imshow(image[0,:,:], cmap = \"Greens\")\n    axis[1].imshow(image[1,:,:], cmap = \"Reds\")\n    axis[2].imshow(image[2,:,:], cmap = \"Blues\")    \n    axis[3].imshow(image[3,:,:], cmap = \"Oranges\")\n    axis[1].set_title(\"microtubules\")\n    axis[2].set_title(\"nucleus\")\n    axis[3].set_title(\"endoplasmatic reticulum\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5a4eed0e76d0855dca86116f0fa187cd36681b8f"},"cell_type":"markdown","source":""}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}