{"cells":[{"metadata":{},"cell_type":"markdown","source":"# RCIC - Basic EDA v1"},{"metadata":{},"cell_type":"markdown","source":"Basic EDA of the train and test csv files for the Recursion Cellular Image Classification (RCIC) challenge on kaggle."},{"metadata":{},"cell_type":"markdown","source":"## Imports"},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T14:11:11.318862Z","start_time":"2019-06-28T14:11:11.091158Z"},"trusted":false},"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{"heading_collapsed":true},"cell_type":"markdown","source":"## Unzip data and fix chmod"},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T13:40:08.378386Z","start_time":"2019-06-28T13:40:08.260221Z"},"hidden":true,"trusted":false},"cell_type":"code","source":"# list files in directory\n!ls -lh \"../input\"","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T13:39:13.561480Z","start_time":"2019-06-28T13:39:13.444195Z"},"hidden":true,"trusted":false},"cell_type":"code","source":"# unzip train.csv, is not needed in the kaggle kernel\n#!unzip train.csv.zip","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T13:40:06.522228Z","start_time":"2019-06-28T13:40:06.411282Z"},"hidden":true,"trusted":false},"cell_type":"code","source":"# set chmod to read train.csv, is not needed in the kaggle kernel\n#!chmod +r train.csv","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Load data"},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:08:53.901308Z","start_time":"2019-06-28T15:08:53.857637Z"},"trusted":false},"cell_type":"code","source":"# read csv data to pandas data frame\ndf_train = pd.read_csv('../input/train.csv')\ndf_test = pd.read_csv('../input/test.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## EDA for train and test dataset"},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:09:27.135164Z","start_time":"2019-06-28T15:09:27.121884Z"},"trusted":false},"cell_type":"code","source":"# check for missing values\nassert ~df_train.isnull().values.any()\nassert ~df_test.isnull().values.any()\n\n# check for NaN values\nassert ~df_train.isna().values.any()\nassert ~df_test.isna().values.any()","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:11:58.186089Z","start_time":"2019-06-28T15:11:58.170938Z"},"trusted":false},"cell_type":"code","source":"# check column name, non-null values and dtypes\nprint(df_train.info(),'\\n')\nprint(df_test.info())","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T14:44:29.933613Z","start_time":"2019-06-28T14:44:29.926586Z"},"trusted":false},"cell_type":"code","source":"# have a look at the first rows\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:09:48.259173Z","start_time":"2019-06-28T15:09:48.251553Z"},"trusted":false},"cell_type":"code","source":"df_test.head()","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:12:20.710304Z","start_time":"2019-06-28T15:12:20.654019Z"},"trusted":false},"cell_type":"code","source":"# cast every column to object and get the unique elements\nprint(df_train.astype('object').describe(include='all').loc['unique', :],'\\n')\nprint(df_test.astype('object').describe(include='all').loc['unique', :])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The experiment count in the train and test set is different. We will look into this in more detail."},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:26:25.454305Z","start_time":"2019-06-28T15:26:25.441581Z"},"trusted":false},"cell_type":"code","source":"# get unique values of every column\ncol_values = [col for col in df_train]\nunique_col_values_train = [df_train[col].unique() for col in df_train]\nunique_col_values_test = [df_test[col].unique() for col in df_test]","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:49:18.640538Z","start_time":"2019-06-28T15:49:18.634528Z"},"trusted":false},"cell_type":"code","source":"# check if there are difference in the columns and if print them\nfor i, (c, a, b) in enumerate(zip(col_values, unique_col_values_train, unique_col_values_test)):\n    \n    if i == 0: continue # skip id_code\n        \n    a = set(a)\n    b = set(b)\n    \n    print('\\n'+c+':', a == b)\n    \n    # if the column elements are not equal, check if they are disjoint\n    if not(a == b):\n        print('disjoint:', a.isdisjoint(b))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Based on this result, we can derive that the train and test dataset are based on different experiments."},{"metadata":{},"cell_type":"markdown","source":"## Visual EDA train data"},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:00:20.761374Z","start_time":"2019-06-28T15:00:20.516192Z"},"trusted":false},"cell_type":"code","source":"sns.catplot(x='experiment', kind='count', data=df_train, height=3, aspect=10);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The dataset count for all the experiments is comparable."},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:00:02.003534Z","start_time":"2019-06-28T15:00:01.441578Z"},"trusted":false},"cell_type":"code","source":"sns.catplot(x='plate', hue='experiment', kind='count', data=df_train, height=4, aspect=5);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In every experiment we have 4 plates and the count is evenly distributed."},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T14:39:23.251771Z","start_time":"2019-06-28T14:39:21.100232Z"},"trusted":false},"cell_type":"code","source":"sns.catplot(y='well', kind='count', data=df_train, height=40, aspect=0.25);","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:00:58.101961Z","start_time":"2019-06-28T15:00:33.566033Z"},"trusted":false},"cell_type":"code","source":"sns.catplot(y='well', hue='experiment', kind='count', data=df_train, height=150, aspect=0.1);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In general, the wells are evenly distributed. However, some wells are not present in every experiment, i.e., G3, G4, G11, G12, M08. (Note: If a experiment is not present in the first or the last position it can be tricky to see it in the visualisation due to the white background.)"},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:03:53.235212Z","start_time":"2019-06-28T15:02:02.125030Z"},"trusted":false},"cell_type":"code","source":"sns.catplot(y='sirna', hue='experiment', kind='count', data=df_train, height=400, aspect=0.05);","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:06:50.455267Z","start_time":"2019-06-28T15:06:50.446397Z"}},"cell_type":"markdown","source":"The siRNA data looks also very balanced. Nevertheless, some siRNAs were not used in every experiment."},{"metadata":{},"cell_type":"markdown","source":"## Visual EDA test dataset"},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:14:50.679373Z","start_time":"2019-06-28T15:14:50.501378Z"},"trusted":false},"cell_type":"code","source":"sns.catplot(x='experiment', kind='count', data=df_test, height=3, aspect=10);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The data set count for all the experiments is comparable."},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:14:55.694147Z","start_time":"2019-06-28T15:14:55.346806Z"},"trusted":false},"cell_type":"code","source":"sns.catplot(x='plate', hue='experiment', kind='count', data=df_test, height=4, aspect=5);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In every experiment we have 4 plates. The count is evenly distributed."},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:15:07.203070Z","start_time":"2019-06-28T15:15:01.706031Z"},"trusted":false},"cell_type":"code","source":"sns.catplot(y='well', kind='count', data=df_test, height=40, aspect=0.25);","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2019-06-28T15:15:18.915643Z","start_time":"2019-06-28T15:15:07.204142Z"},"trusted":false},"cell_type":"code","source":"sns.catplot(y='well', hue='experiment', kind='count', data=df_test, height=150, aspect=0.1);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In general, the wells are evenly distributed. However, some wells are not present in certain experiments."},{"metadata":{},"cell_type":"markdown","source":"# Conclusion"},{"metadata":{},"cell_type":"markdown","source":"The train and test dataset contain different experiments. All in all, the train and test dataset looks very balanced."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.7.3"},"toc":{"base_numbering":1,"nav_menu":{},"number_sections":true,"sideBar":true,"skip_h1_title":false,"title_cell":"Table of Contents","title_sidebar":"Contents","toc_cell":false,"toc_position":{},"toc_section_display":true,"toc_window_display":true}},"nbformat":4,"nbformat_minor":1}