{"cells":[{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true},"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nfrom tqdm import tqdm\nimport glob\nimport os\nimport matplotlib.pyplot as plt\n\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Welcome to Recursion Cellular Image Classification competition**.\n<br>This kernel is a basic idea about the data and we will do a basic EDA."},{"metadata":{},"cell_type":"markdown","source":"## Loading the metadata with Description of metadata"},{"metadata":{"trusted":true},"cell_type":"code","source":"BASE_DIR = '../input'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"os.listdir(BASE_DIR)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So, we basically have train.csv and test.csv files which contains the major details about our image data, other than these, we also have train_controls.csv, test_controls.csv, pixel_stats.csv, sample_submission.csv files and train and test folder which contains our image data."},{"metadata":{},"cell_type":"markdown","source":"### Train metadata"},{"metadata":{},"cell_type":"markdown","source":"Let's start by reading train.csv:"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df_train = pd.read_csv(os.path.join(BASE_DIR, 'train.csv'))\nprint(df_train.sample(5))\nprint(\"*\"*40)\nprint(df_train.info())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So, we can see that in our training data, there are five columns in it, among which \"sirna\" will be our label which we will need to predict. So we will talk about \"sirna\" later. For now, we can go to the info section of our data, according to it, our training metadata contains 36515 different rows and there are no missing values, due to which we can consider our data to be already cleaned."},{"metadata":{},"cell_type":"markdown","source":"Also, from the dataset info available at [RxRx.ai](https://www.rxrx.ai/#the-data),\n* Total of 308 wells on which the experiments are performed(277 wells in our training data)\n* Each having 4 plates(namely 1, 2, 3 and 4)\n* Total of 51 experiments(33 experiments in our training data)\n* Each sample having a unique id_code. \n* Also, each well has two different sites both marked with same sirna, details about which can be found in the pixel data.\n\nAll above points can be seen easily from the following code cell. "},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df_train.nunique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now, our target value, i.e. \"sirna\" has 1108 different values, that basically means that we will need to classify our images into any of the 1108 classes. More details about our target value can be found in the above mentioned link."},{"metadata":{},"cell_type":"markdown","source":"### Test metadata"},{"metadata":{},"cell_type":"markdown","source":"Let's now come to our test data:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df_test = pd.read_csv(os.path.join(BASE_DIR, 'test.csv'))\nprint(df_test.sample(5))\nprint(\"*\"*40)\nprint(df_test.info())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Our test dataset is defined in the similar fashion excepting the \"sirna\" column as expected. Also this dataset too do not contain any missing values."},{"metadata":{},"cell_type":"markdown","source":"### Train/Test controls metadata"},{"metadata":{},"cell_type":"markdown","source":"Let's talk a bit about train_controls and test_controls dataset:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df_train_controls = pd.read_csv(os.path.join(BASE_DIR, 'train_controls.csv'))\nprint(df_train_controls.sample(5))\nprint(\"*\"*40)\nprint(df_train_controls.info())\nprint(\"*\"*80)\ndf_test_controls = pd.read_csv(os.path.join(BASE_DIR, 'test_controls.csv'))\nprint(df_test_controls.sample(5))\nprint(\"*\"*40)\nprint(df_test_controls.info())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So, according to the data description given with the problem, in each experiment, the same 30 siRNAs appear on every plate as positive controls. In addition, there is at least one well per plate with untreated cells as a negative control. It has the same schema as [train/test].csv, plus a well_type field denoting the type of control. In our [train/test]_controls.csv file, we can see that there are two unique well_types given as can be seen from below code cell:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(df_train_controls[\"well_type\"].unique())\nprint(\"*\"*40)\nprint(df_test_controls[\"well_type\"].unique())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Pixel metadata"},{"metadata":{},"cell_type":"markdown","source":"We are provided with statistics on all the images in the pixel_stats.csv file. This information will allow us to normalize the data, for example by reducing the mean and dividing by the standard deviation."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df_pixel_stats = pd.read_csv(os.path.join(BASE_DIR, 'pixel_stats.csv'))\ndf_pixel_stats_updated = df_pixel_stats.set_index(['id_code','site', 'channel'])\nprint(df_pixel_stats.sample(5))\nprint(\"*\"*40)\nprint(df_pixel_stats_updated.sample(5))\nprint(\"*\"*40)\nprint(df_pixel_stats.info())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Basic EDA"},{"metadata":{},"cell_type":"markdown","source":"#### Distribution of \"plate\" attribute among the train dataset"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df_train[\"plate\"].value_counts().plot(kind = \"bar\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* So, this shows us that each experiment has a equal number of different plates and dataset is stable."},{"metadata":{},"cell_type":"markdown","source":"#### Distribution of \"experiment\" attribute among train dataset"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.rcParams['figure.figsize'] = [15, 5]\ndf_train[\"experiment\"].value_counts().plot(kind = \"bar\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Similarly, the number of samples from each experiment is also very stable."},{"metadata":{},"cell_type":"markdown","source":"#### Distribution of \"well_type\" among train/test control dataset"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.rcParams['figure.figsize'] = [10, 5]\nprint(df_train_controls[\"well_type\"].value_counts())\ndf_train_controls[\"well_type\"].value_counts().plot(kind = \"bar\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* So as already told in the metadata description, the same 30 siRNAs appear on every plate as positive controls. In addition, there is at least one well per plate with untreated cells as a negative control. That's why there is a high number of positive_controls in comparison to negative_controls in our train_controls.csv metadata.\n* Similar fashion can be seen in test_controls.csv metadata as follows:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.rcParams['figure.figsize'] = [10, 5]\nprint(df_test_controls[\"well_type\"].value_counts())\ndf_test_controls[\"well_type\"].value_counts().plot(kind = \"bar\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Distribution of \"site\" among pixel stats dataset"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.rcParams['figure.figsize'] = [10, 5]\nprint(df_pixel_stats[\"site\"].value_counts())\ndf_pixel_stats[\"site\"].value_counts().plot( kind = \"bar\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* So in above figure, we can see that an experiment is performed on two different sites(namely 1 and 2) in a same well."},{"metadata":{},"cell_type":"markdown","source":"#### Distribution of \"channel\" among pixel stats dataset"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.rcParams['figure.figsize'] = [10, 8]\nprint(df_pixel_stats[\"channel\"].value_counts())\ndf_pixel_stats[\"channel\"].value_counts().plot(kind = \"pie\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Above figure here shows that every image is having 6 different channels, actually all the images are taken with adjusting the wavelength of the apparatus and then can be combined if one wants to visualise them as a single image."},{"metadata":{},"cell_type":"markdown","source":"### Visualising image data"},{"metadata":{},"cell_type":"markdown","source":"* At the very last of this notebook, we would like to conclude this very basic EDA with a visualisation of the image data provided to us."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"path = []\nfor index in range(6):\n    path.append(os.path.join(\"../input/train/HUVEC-01/Plate1\", os.listdir(\"../input/train/HUVEC-01/Plate1\")[index]))","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":false,"trusted":true},"cell_type":"code","source":"fig, axes = plt.subplots(2, 3, figsize=(24, 16))\nfor index, ax in enumerate(axes.flatten()):\n    img = plt.imread(path[index])\n    ax.axis('off')\n    ax.set_title(os.listdir(\"../input/train/HUVEC-01/Plate1\")[index])\n    _ = ax.imshow(img, cmap='gray')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Thanks!"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.5.3"}},"nbformat":4,"nbformat_minor":1}