{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Recursion Cellular Image Classification</span>\n### Disentangling biological signal from experimental noise in cellular images</span>\n![](https://assets.website-files.com/5cb63fe47eb5472014c3dae6/5d040176f0a2fd66df939c51_figure1%400.75x.png)"},{"metadata":{},"cell_type":"markdown","source":"Finding hard to follow this dataset? \n\nThis notebook should help you visualise and understand what the hell is going on! I have just started exploring the dataset and I am aiming to refresh this kernel soon with more always more visualisations.\n\nMore technical details are provided here: https://www.rxrx.ai/"},{"metadata":{},"cell_type":"markdown","source":"# Getting setup"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nprint(os.listdir(\"../input\"))\nimport sys\nimport matplotlib.pyplot as plt\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load the metadata"},{"metadata":{"trusted":true},"cell_type":"code","source":"test_controls = pd.read_csv('../input/test_controls.csv')\ntest = pd.read_csv('../input/test.csv')\ntest['sirna'] = 0\ntest['well_type'] = 'treatment'\ntest['dataset'] = 'test'\ntest_controls['dataset'] = 'test'\n\n\ntrain_controls = pd.read_csv('../input/train_controls.csv')\ntrain = pd.read_csv('../input/train.csv')\ntrain['sirna'] = 0\ntrain['well_type'] = 'treatment'\ntrain['dataset'] = 'train'\ntrain_controls['dataset'] = 'train'\n\nmd = pd.concat([train, train_controls, test, test_controls]).reset_index(drop=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Understand the dataset"},{"metadata":{"trusted":true},"cell_type":"code","source":"unique_experiments = md.experiment.unique()\nlen(unique_experiments)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The dataset is composed of 51 experiments."},{"metadata":{"trusted":true},"cell_type":"code","source":"names, counts = np.unique([experiment.split('-')[0] for experiment in unique_experiments], return_counts=True)\nfor experiment_name, experiment_count in zip(names, counts):\n    print('{} experiments focused on the cell type {}'.format(experiment_count, experiment_name))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The same setup experiments have been tested on four cell types: HUVEC, RPE, HepG2, U2OS!"},{"metadata":{},"cell_type":"markdown","source":"# What is an experiment?\nLet's select an experiment and explore."},{"metadata":{"trusted":true},"cell_type":"code","source":"one_experiment_md = md[(md.experiment=='RPE-03')]\none_experiment_md.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"An experiment is made with:\n - 4 plates\n     - in which 308 wells have been photographed."},{"metadata":{"trusted":true},"cell_type":"code","source":"one_experiment_md.groupby(['plate']).count()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Visualise the plats!\nI believe that a drawing is always better than a long paragraph, let's transform those well code into actual locations."},{"metadata":{"trusted":true},"cell_type":"code","source":"import geopandas as gpd\nfrom shapely.geometry import Point\n\ndef letter_to_int(letter):\n    alphabet = list('abcdefghijklmnopqrstuvwxyz'.upper())\n    return alphabet.index(letter) + 1\n\ndef well_to_point(well):\n    letter = letter_to_int(well[0])\n    number = int(well[1:])\n    return Point(letter, number)\nmd['geometry'] = md.well.apply(lambda well: well_to_point(well))\nmd = gpd.GeoDataFrame(md)\nmd.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_well_type_positions_for_experiment(experiment_name):\n    fig, axes = plt.subplots(1, 4, figsize=(20, 5))\n    for plate, ax in zip([1,2,3,4], axes):\n        one_plate_md = md[md.experiment==experiment_name]\n        if plate==1:\n            legend=True\n        else:\n            legend=False\n        one_plate_md.plot(column='well_type',legend=legend, ax=ax);\n        if plate==1:\n            leg = ax.get_legend()\n            leg.set_bbox_to_anchor((-0.3, 0., 0.2, 0.2))\n    _ = fig.suptitle('Plate 1 to 4 - Experiment {}'.format(experiment_name))\n    \nplot_well_type_positions_for_experiment(md.experiment.unique()[0])\nplot_well_type_positions_for_experiment(md.experiment.unique()[-1])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Each plate holds the same 30 control siRNA conditions (positive controlled), 277 different non-control siRNA (treatment), and one untreated well (negative control). The well types follow the same architecture from one plat to another and experiment to another."},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_sirna_positions_for_experiment(experiment_name):\n    fig, axes = plt.subplots(1, 4, figsize=(20, 5))\n    for plate, ax in zip([1,2,3,4], axes):\n        one_plate_md = md[md.experiment==experiment_name]\n        one_plate_md = one_plate_md[one_plate_md.well_type=='positive_control'] \n        one_plate_md.plot(column='sirna', ax=ax, categorical=True, cmap='tab20c');\n    _ = fig.suptitle('Plate 1 to 4 - Experiment {}'.format(experiment_name))\n    \nplot_sirna_positions_for_experiment(md.experiment.unique()[0])\nplot_sirna_positions_for_experiment(md.experiment.unique()[-1])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The location of each of the 1,108 non-control siRNA conditions is randomized in each experiment to prevent confounding effects of the location of a particular well."},{"metadata":{},"cell_type":"markdown","source":"# To Be Continued..."},{"metadata":{},"cell_type":"markdown","source":"* I hope you enjoyed this kernel, I will commit more soon! Please let a thumbs up or a comment if you appreciated the kernel! :)"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}