{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n    <h1 align=\"center\" style=\"color:darkgreen;\">Multimodal Single-Cell Integration</h1>  \n</div>\n\n<div>\n    <h2 align=\"center\" style=\"color:darkgray;\">Predict how DNA, RNA & protein measurements co-vary in single cells</h2>\n</div>\n\n<img src=\"https://openproblems.bio/media/learning/central-dogma-large.png\">\n\n### [Open Problems - About multimodal single-cell data](https://openproblems.bio/neurips_docs/data/about_multimodal/) <span style=\"color:darkred;\">(Resources)</span>\n\nBecause we know that the difference between types of cells has to do with different levels of RNA and proteins, it’s very useful to be able to measure the abundance of these molecules at the level of individual cells. Not only does this give us a fine-resolution view into the different kinds of cells in the body, it also provides insight into how the same set of DNA instructions can be interpreted so differently throughout the body.\n\nThe promise of single-cell measurements of genetics information is that by better understanding how this information flow within our cells and tissues, we might better understand what goes wrong in the context of disease.\n\n<img src=\"https://raw.githubusercontent.com/MehranKazeminia/Multimodal-Single-Cell/main/differing_expression_20.png\">\n\nRegulation of gene expression affects the amount of RNA and protein in the cells.\n\nIn this example, gene A is more upregulated than gene B resulting more RNA and more protein.","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"import warnings # suppress warnings\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os, gc\nimport random\n\nimport numpy as np \nimport pandas as pd\nimport seaborn as sns\n\nfrom tqdm import tqdm\nfrom scipy import stats\nfrom pathlib import Path\n\nimport matplotlib.pyplot as plt\nimport plotly.figure_factory as ff\nimport plotly.express as px\n%matplotlib inline\n!ls ../input/*","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install tables","metadata":{"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">metadata.csv</h1>\n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\nThe dataset for this competition comprises single-cell multiomics data collected from mobilized peripheral CD34+ hematopoietic stem and progenitor cells (HSPCs) isolated from four healthy human donors. ","metadata":{}},{"cell_type":"code","source":"metadata = pd.read_csv('../input/open-problems-multimodal/metadata.csv')\nmetadata","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:darkred;\">Descriptions:</span>\n\n- **cell_id** - A unique identifier for each observed cell.\n\n- **donor** - An identifier for the four cell donors.\n\n- **day** - The day of the experiment the observation was made.\n\n- **technology** - Either citeseq or multiome.\n\n- **cell_type** - One of the above cell types or else hidden.","metadata":{}},{"cell_type":"code","source":"print('Duplicates in Cell_Id:', metadata.index.duplicated().sum())\nprint('Missing_Values:') \nmetadata.isna().sum()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Number of unique cells: <span style=\"color:darkred;\">281528</span>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"sns.set()\nmetadata['cell_type'].value_counts(normalize=True).plot(kind='barh', figsize=(15,3))\nmetadata['cell_type'].value_counts(normalize=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Number of cell types: <span style=\"color:darkred;\">8 (MasP, MkP, NeuP, MoP, EryP, HSC, BP, hidden)</span>\n\n- **MasP** = Mast Cell Progenitor\n- **MkP** = Megakaryocyte Progenitor\n- **NeuP** = Neutrophil Progenitor\n- **MoP** = Monocyte Progenitor\n- **EryP** = Erythrocyte Progenitor\n- **HSC** = Hematoploetic Stem Cell\n- **BP** = B-Cell Progenitor","metadata":{}},{"cell_type":"code","source":"metadata.set_index('cell_id').pivot(columns=['technology', 'donor'], values= 'cell_type')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata.set_index('cell_id').pivot(columns=['technology', 'day'], values= 'cell_type')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata.groupby('day')['cell_type'].value_counts(normalize=True).plot(kind='barh', figsize=(15,10))\n#metadata.groupby('day')['cell_type'].value_counts(normalize=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The number of days the cells have been donated: <span style=\"color:darkred;\">5 (d2, d3, d4, d7, d10)</span>","metadata":{}},{"cell_type":"code","source":"metadata.groupby('donor')['cell_type'].value_counts(normalize=True).plot(kind='barh', figsize=(15,10))\n#metadata.groupby('donor')['cell_type'].value_counts(normalize=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Number of donors: <span style=\"color:darkred;\">4 (27678, 32606, 13176, 31800)</span>","metadata":{}},{"cell_type":"code","source":"metadata.groupby('technology')['cell_type'].value_counts(normalize=True).plot(kind='barh', figsize=(15,5))\n#metadata.groupby('technology')['cell_type'].value_counts(normalize=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Number of technologies: <span style=\"color:darkred;\">2 (citeseq, multiome)</span>\n\n### The two technologies do not share cells.","metadata":{}},{"cell_type":"code","source":"metadata['technology'].value_counts(normalize=True).plot(kind='barh', figsize=(15,1))\nmetadata['technology'].value_counts(normalize=True), metadata['technology'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### metadata.csv: <span style=\"color:darkred;\">281528 rows x 5 columns</span>\n\n> #### Cells that used **CITEseq** technology: <span style=\"color:darkred;\">119651 (42%)</span>\n\n> #### Cells that used **Multiome** technology: <span style=\"color:darkred;\">161877 (58%)</span>","metadata":{}},{"cell_type":"code","source":"del metadata\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">train_cite_inputs.h5, train_cite_targets.h5, test_cite_inputs.h5</h1>\n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">CITEseq technology:</span>\n\n### <span style=\"color:darkblue;\">given gene expression, predict protein levels.</span>\n\n<img src=\"https://openproblems.bio/media/learning/CITE-seq.svg\">\n\n<div>\n    <h2 align=\"center\" style=\"color:black;\">Multimodal scRNA and protein abundance from individual cells</h2>  \n    <h4 align=\"center\" style=\"color:black;\">CITE-seq provides a measure of gene expression at the level of RNA and measure of protein abundance for 10-200 cell surface proteins.</h4>\n</div>","metadata":{"_kg_hide-input":false}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">train_cite_inputs.h5</span>","metadata":{}},{"cell_type":"code","source":"train_cite_10 = pd.read_hdf('../input/open-problems-multimodal/train_cite_inputs.h5', stop=10)\ntrain_cite_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_cite_10 = pd.read_hdf('../input/open-problems-multimodal/train_cite_inputs.h5').index\n#The number of rows: 70988\n#train_cite_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <span style=\"color:darkred;\">Feature Analysis:</span>\n\nDue to the large volume of the data set, only seven thousand rows of the **\"train_cite_inputs.h5\"** file were randomly selected and then a histogram was made of all the available columns (features). And of course this has been repeated several times.\n\nBut in this information, the number of zeros is very high (about eighty percent). There are even columns where all the rows are zero (the zero value indicates that there are cells that do not express these genes). Therefore, all zero values should be removed when plotting the histogram. If you don't, your histograms won't show the rest of the values.","metadata":{}},{"cell_type":"code","source":"NUMBER = 5 # The number of iteration\nL = 7000 # The number of rows selected for each iteration\ndata_len = 70988 # The number of rows in the train_cite_inputs.h5\n\nsns.set()\nfor i in range(NUMBER):\n    split = random.randrange(data_len-L)\n    \n    train_cite_L = pd.read_hdf('../input/open-problems-multimodal/train_cite_inputs.h5', start=split, stop=split+L)\n    train_cite_L = train_cite_L.values.ravel()\n    train_cite_L = train_cite_L[train_cite_L != 0]\n\n    plt.figure(figsize=(15, 5))\n    plt.gca().set_facecolor('lightblue')\n    plt.hist(train_cite_L, bins=500, color='red')\n    \n    plt.xlabel('Gene expression levels')\n    plt.title(f'Histogram for row number {split}  until row number {split+L}', fontsize=16)\n    plt.show() \n    print()\n    \ndel train_cite_10\ndel train_cite_L\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <span style=\"color:darkred;\">Histogram for individual columns:</span>\n\nFor better visualization, five columns are randomly selected and the histogram of the values in each of the columns (features) is drawn separately.","metadata":{}},{"cell_type":"code","source":"NUMBER = 5 # The number of iteration\ndata_F = 22050 # The number of columns in the train_cite_inputs.h5\n\ntrain_cite = pd.read_hdf('../input/open-problems-multimodal/train_cite_inputs.h5')\ncols = train_cite.columns\n\nsns.set()\nfor i in range(NUMBER):\n    feature = random.randrange(data_F)   \n    train_cite_F = train_cite.iloc[:,feature].values\n    train_cite_F = train_cite_F[train_cite_F != 0]\n\n    plt.figure(figsize=(15, 5))\n    plt.gca().set_facecolor('lightblue')\n    plt.hist(train_cite_F, bins=500, color='red')\n      \n    plt.xlabel('Gene expression levels')\n    plt.title(f'Histogram for Column number {feature} <<{cols[feature]}>>', fontsize=16)\n    plt.show() \n    print()\n\ndel train_cite_F   \ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <span style=\"color:darkred;\">Shapiro–Wilk test:</span>","metadata":{}},{"cell_type":"code","source":"cols = list(train_cite.columns)\nlen(cols)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_select = []\nalpha = 0.05\n\nfor col in cols:\n    _, p_value = stats.shapiro(train_cite[col])\n    \n    if (p_value <= alpha): \n        cols_select.append(col)     \n         \nlen(cols_select)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, the minimum number (22050-21601=449) of columns (features) will not have a positive effect for training or will have a very small effect.","metadata":{}},{"cell_type":"code","source":"del train_cite  \ngc.collect()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">test_cite_inputs.h5</span>","metadata":{}},{"cell_type":"code","source":"test_cite_10 = pd.read_hdf('../input/open-problems-multimodal/test_cite_inputs.h5', stop=10)\ntest_cite_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#test_cite_10 = pd.read_hdf('../input/open-problems-multimodal/test_cite_inputs.h5').index\n#The number of rows: 48663\n#test_cite_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <span style=\"color:darkred;\">Feature Analysis:</span>\n\nThe above tasks are repeated for the **\"test_cite_inputs.h5\"** file. Histograms are somewhat different.","metadata":{}},{"cell_type":"code","source":"NUMBER = 5 # The number of iteration\nL = 7000 # The number of rows selected for each iteration\ndata_len = 48663 # The number of rows in the test_cite_inputs.h5\n\nsns.set()\nfor i in range(NUMBER):\n    split = random.randrange(data_len-L)\n    \n    test_cite_L = pd.read_hdf('../input/open-problems-multimodal/test_multi_inputs.h5', start=split, stop=split+L)\n    test_cite_L = test_cite_L.values.ravel()\n    test_cite_L = test_cite_L[test_cite_L != 0]\n\n    plt.figure(figsize=(15, 5))\n    plt.gca().set_facecolor('lightblue')\n    plt.hist(test_cite_L, bins=500, color='orange')\n    \n    plt.xlabel('Gene expression levels')\n    plt.title(f'Histogram for row number {split}  until row number {split+L}', fontsize=16)\n    plt.show() \n    print()\n    \ndel test_cite_10\ndel test_cite_L\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <span style=\"color:darkred;\">Histogram for individual columns:</span>\n\nFor better visualization, five columns are randomly selected and the histogram of the values in each of the columns (features) is drawn separately.","metadata":{}},{"cell_type":"code","source":"NUMBER = 5 # The number of iteration\ndata_F = 22050 # The number of columns in the test_cite_inputs.h5\n\ntest_cite = pd.read_hdf('../input/open-problems-multimodal/test_cite_inputs.h5')\ncols = test_cite.columns\n\nsns.set()\nfor i in range(NUMBER):\n    feature = random.randrange(data_F)   \n    test_cite_F = test_cite.iloc[:,feature].values\n    test_cite_F = test_cite_F[test_cite_F != 0]\n\n    plt.figure(figsize=(15, 5))\n    plt.gca().set_facecolor('lightblue')\n    plt.hist(test_cite_F, bins=500, color='orange')\n      \n    plt.xlabel('Gene expression levels')\n    plt.title(f'Histogram for Column number {feature} <<{cols[feature]}>>', fontsize=16)\n    plt.show() \n    print()\n\ndel test_cite_F   \ndel test_cite\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">train_cite_targets.h5</span>\n\nSurface protein levels for the same cells that have been dsb normalized.","metadata":{}},{"cell_type":"code","source":"target_cite_10 = pd.read_hdf('../input/open-problems-multimodal/train_cite_targets.h5', stop=10)\ntarget_cite_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#target_cite_id = pd.read_hdf('../input/open-problems-multimodal/train_cite_targets.h5').index\n#The number of rows: 70988\n#target_cite_id","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <span style=\"color:darkred;\">Target Analysis:</span>\n\nThe target file has 140 columns. That is, the challenge of this part is regression with 140 outputs. Five columns are randomly selected and the histogram of the values in each of the columns is drawn separately.","metadata":{}},{"cell_type":"code","source":"NUMBER = 5 # The number of iteration\ndata_F = 140 # The number of columns in the train_cite_targets.h5\n\ntarget_cite = pd.read_hdf('../input/open-problems-multimodal/train_cite_targets.h5')\ncols = target_cite.columns\n\nsns.set()\nfor i in range(NUMBER):\n    protein = random.randrange(data_F)   \n    target_cite_F = target_cite.iloc[:,protein].values\n    target_cite_F = target_cite_F[target_cite_F != 0]\n\n    plt.figure(figsize=(15, 5))\n    plt.gca().set_facecolor('lightgray')\n    plt.hist(target_cite_F, bins=500, color='green')\n      \n    plt.xlabel('Surface protein levels')\n    plt.title(f'Histogram for Column number {protein} < {cols[protein]} >', fontsize=16)\n    plt.show() \n    print()\n\ndel target_cite_10\ndel target_cite_F   \ndel target_cite\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### CITEseq technology:\n\n> #### train_cite_inputs.h5: <span style=\"color:darkred;\">70988 rows x 22050 columns</span>\n\n> #### train_cite_targets.h5: <span style=\"color:darkred;\">70988  rows x 140 columns</span>\n\n> #### test_cite_inputs.h5: <span style=\"color:darkred;\">48663  rows x 22050 columns</span>","metadata":{}},{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">train_multi_inputs.h5, train_multi_targets.h5, test_multi_inputs.h5</h1>\n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">Multiome technology:</span>\n\n### <span style=\"color:darkblue;\">given chromatin accessibility, predict gene expression.</span>\n\n<img src=\"https://openproblems.bio/media/learning/ATAC-seq.svg\">\n\n<div>\n    <h2 align=\"center\" style=\"color:black;\">Multimodal scRNA and scATAC from cell nuclei</h2>\n    <h4 align=\"center\" style=\"color:black;\">With the 10X Genomics Single-Cell Multiome ATAC + Gene Expression kit, it is possible to measure chromatin accessibility and RNA expression in tens of thousands of cells. These methods only measure RNA within the nucleus of the cell.</h4>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">train_multi_inputs.h5</span>","metadata":{}},{"cell_type":"code","source":"train_multi_10 = pd.read_hdf('../input/open-problems-multimodal/train_multi_inputs.h5', stop=10)\ntrain_multi_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_multi_10 = pd.read_hdf('../input/open-problems-multimodal/train_multi_inputs.h5').index\n#The number of rows: 105942\n#train_multi_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <span style=\"color:darkred;\">Feature Analysis:</span>\n\nThis dataset is very big. Of course, more than ninety percent of the numbers are zero. Also, the number of columns (features) is more than the number of rows. These topics should not be forgotten during training. However, only three thousand rows of the \"train_multi_inputs.h5\" file were randomly selected and then a histogram was made from all the available columns (features). And of course this has been repeated several times. Also, as mentioned earlier, when drawing the histogram, all zero values must be removed.\n\nIn order to prevent the explosion of this notebook :), the histogram of the rest of the files is not drawn. Of course, more topics will be discussed in future notebooks and during training. It may even be necessary to use programs that work with \"data flow\" such as \"mrjob\".","metadata":{}},{"cell_type":"code","source":"NUMBER = 5 # The number of iteration\nL = 3000 # The number of rows selected for each iteration\ndata_len = 105942 # The number of rows in the train_multi_inputs.h5\n\nsns.set()\nfor i in range(NUMBER):\n    split = random.randrange(data_len-L)\n    \n    train_multi_L = pd.read_hdf('../input/open-problems-multimodal/train_multi_inputs.h5', start=split, stop=split+L)\n    train_multi_L = train_multi_L.values.ravel()\n    train_multi_L = train_multi_L[train_multi_L != 0]\n\n    plt.figure(figsize=(15, 5))\n    plt.gca().set_facecolor('lightgray')\n    plt.hist(train_multi_L, bins=500, color='violet')\n    \n    plt.xlabel('ATAC-seq peak counts')\n    plt.title(f'Histogram for row number {split}  until row number {split+L}', fontsize=16)\n    plt.show() \n    print()\n    \ndel train_multi_10\ndel train_multi_L\ngc.collect()","metadata":{"_kg_hide-output":false},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">test_multi_inputs.h5</span>","metadata":{}},{"cell_type":"code","source":"test_multi_10 = pd.read_hdf('../input/open-problems-multimodal/test_multi_inputs.h5', stop=10)\ntest_multi_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#test_multi_10 = pd.read_hdf('../input/open-problems-multimodal/test_multi_inputs.h5').index\n#The number of rows: 55935\n#test_multi_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">train_multi_targets.h5</span>","metadata":{}},{"cell_type":"code","source":"targets_multi_10 = pd.read_hdf('../input/open-problems-multimodal/train_multi_targets.h5', stop=10)\ntargets_multi_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#targets_multi_id = pd.read_hdf('../input/open-problems-multimodal/train_multi_targets.h5').index\n#The number of rows: 105942\n#targets_multi_id","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del targets_multi_10\ndel test_multi_10\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Multiome technology:\n\n> #### train_multi_inputs.h5: <span style=\"color:darkred;\">105942 rows x 228942 columns</span>\n\n> #### train_multi_targets.h5: <span style=\"color:darkred;\">105942 rows x 23418 columns</span>\n\n> #### test_multi_inputs.h5: <span style=\"color:darkred;\">55935 rows x 228942 columns</span>","metadata":{}},{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">evaluation_ids.csv & sample_submission.csv</h1>\n</div>\n\n<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"evaluation_10 = pd.read_csv('../input/open-problems-multimodal/evaluation_ids.csv', nrows=10)\nevaluation_10.set_index('row_id', inplace=True)\nevaluation_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#evaluation_10 = pd.read_csv('../input/open-problems-multimodal/evaluation_ids.csv').index\n#The number of rows: 65744180\n#evaluation_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_10 = pd.read_csv('../input/open-problems-multimodal/sample_submission.csv', nrows=10)\nsubmission_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#submission_10 = pd.read_csv('../input/open-problems-multimodal/sample_submission.csv').index\n#The number of rows: 65744180\n#submission_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del evaluation_10\ndel submission_10\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Evaluation & Submission:\n\n> #### evaluation_ids.csv: <span style=\"color:darkred;\">65744180 rows x 3 columns</span>\n\n> #### sample_submission.csv: <span style=\"color:darkred;\">65744180 rows x 2 columns</span>","metadata":{}},{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">Summary</h1>\n</div>\n<div class=\"alert alert-success\">  \n</div>\n\n### metadata.csv: <span style=\"color:darkblue;\">281528 rows x 5 columns</span>\n\n> #### Cells that used **CITEseq** technology: <span style=\"color:darkred;\">119651 (42%)</span>\n>> #### For train:<span style=\"color:darkred;\"> 70988 Cells</span>\n>> #### For test:<span style=\"color:darkred;\"> 48663 Cells</span>\n\n> #### Cells that used **Multiome** technology: <span style=\"color:darkred;\">161877 (58%)</span>\n>> #### For train:<span style=\"color:darkred;\"> 105942 Cells</span>\n>> #### For test:<span style=\"color:darkred;\"> 55935 Cells</span>\n-----\n\n### CITEseq technology: <span style=\"color:darkblue;\">(given gene expression, predict protein levels)</span>\n\n> #### train_cite_inputs.h5: <span style=\"color:darkred;\">70988 rows x 22050 columns</span>\n>> #### <span style=\"color:darkred;\"> [ 70988 ] Cells</span>\n>> #### <span style=\"color:darkred;\"> [ 22050 ] RNA library-size normalized and log1p transformed counts (gene expression levels) </span>\n\n> #### train_cite_targets.h5: <span style=\"color:darkred;\">70988 rows x 140 columns</span>\n>> #### <span style=\"color:darkred;\"> [ 70988 ] Cells</span>\n>> #### <span style=\"color:darkred;\"> [ 140 ] Surface protein levels for the same cells that have been dsb normalized </span>\n\n> #### test_cite_inputs.h5: <span style=\"color:darkred;\">48663 rows x 22050 columns</span>\n>> #### <span style=\"color:darkred;\"> [ 48663 ] Cells</span>\n>> #### <span style=\"color:darkred;\"> [ 22050 ] RNA library-size normalized and log1p transformed counts (gene expression levels) </span>\n\n**train/test_cite_inputs.h5** - RNA library-size normalized and log1p transformed counts (gene expression levels), with rows corresponding to cells and columns corresponding to genes given by {gene_name}_{gene_ensemble-ids}.\n\n**train_cite_targets.h5** - Surface protein levels for the same cells that have been dsb normalized.\n\n-----\n\n### Multiome technology: <span style=\"color:darkblue;\">(given chromatin accessibility, predict gene expression)</span>\n\n> #### train_multi_inputs.h5: <span style=\"color:darkred;\">105942 rows x 228942 columns</span>\n>> #### <span style=\"color:darkred;\"> [ 105942 ] Cells</span>\n>> #### <span style=\"color:darkred;\"> [ 228942 ] ATAC-seq peak counts</span>\n\n> #### train_multi_targets.h5: <span style=\"color:darkred;\">105942 rows x 23418 columns</span>\n>> #### <span style=\"color:darkred;\"> [ 105942 ] Cells</span>\n>> #### <span style=\"color:darkred;\"> [ 23418 ] RNA gene expression levels for the same cells</span>\n\n> #### test_multi_inputs.h5: <span style=\"color:darkred;\">55935 rows x 228942 columns</span>\n>> #### <span style=\"color:darkred;\"> [ 55935 ] Cells</span>\n>> #### <span style=\"color:darkred;\"> [ 228942 ] ATAC-seq peak counts</span>\n\n**train/test_multi_inputs.h5** - ATAC-seq peak counts transformed with TF-IDF using the default log(TF) * log(IDF) output (chromatin accessibility), with rows corresponding to cells and columns corresponding to the location of the genome whose level of accessibility is measured, here identified by the genomic coordinates on reference genome GRCh38 provided in the 10x References - 2020-A (July 7, 2020).\n\n**train_multi_targets.h5** - RNA gene expression levels as library-size normalized and log1p transformed counts for the same cells.\n\n-----\n\n### Evaluation & Submission:\n\n> #### evaluation_ids.csv: <span style=\"color:darkred;\">65744180 rows x 3 columns</span>\n\n> #### sample_submission.csv: <span style=\"color:darkred;\">65744180 rows x 2 columns</span>\n\n**evaluation_ids.csv** - Identifies the labels from the test set to be evaluated. It provides a join key from the cell_id / gene_id identifiers of the label matrix to the row_id needed for the submission file.\n\n**sample_submission.csv** - A sample submission file in the correct format. See the Evaluation page for more information.\n\n-----","metadata":{}},{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">Splits</h1>\n</div>\n<div class=\"alert alert-success\">  \n</div>\n\n### <span style=\"color:darkred;\">The data splits are arranged as follows:</span>\n\n- **The training set** - Comprises samples only from donors 13176, 31800, and 32606.\n\n- **The public test set** - Comprises samples only from donor 27678. \n\n- **The private test set** - Comprises samples from all four donors.\n\n### <span style=\"color:darkred;\">For the Multiome samples:</span>\n\n- **The training set** - Comprises samples only from days 2, 3, 4, and 7. \n\n- **The public test set** - Comprises samples only from days 2, 3, and 7. \n\n- **The private test set** - Comprises data only from day 10.\n\n### <span style=\"color:darkred;\">For the CITEseq samples</span>\n\n- **The training set** - Comprises samples only from days 2, 3, and 4. \n\n- **The public test set** - Comprises samples only from days 2, 3, and 4. \n\n- **The private test set** - Comprises samples only from day 7. There are no day 10 CITEseq samples in any split.","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4308072%2F23e8c1f6faea1453998544cdc116a20e%2FNeurIPS%202022%20-%20Frame%204.jpg?generation=1660755395301873&alt=media\">","metadata":{}},{"cell_type":"markdown","source":"### <span style=\"color:darkred;\">Our task:</span>\n\n- Our task is to predict the labels corresponding to the inputs in the test set. To facilitate submission scoring, we only require predictions on a subset of the Multiome data. This subset was created by sampling 30% of the Multiome rows, and for each row, 15% of the columns. The sample of columns varies from row-to-row. All of the CITEseq labels are scored.","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">Competition Evaluation:</span>\n\n#### [Taken from the main page of the competition.](https://www.kaggle.com/competitions/open-problems-multimodal/overview/evaluation)\n\nWe use the Pearson correlation coefficient to rank submissions. For each observation in the Multiome data set, we compute the correlation between the ground-truth gene expressions and the predicted gene expressions. For each observation in the CITEseq data set, we compute the correlation between ground-truth surface protein levels and predicted surface protein levels. The overall score is the average of each sample's correlation score. If a sample's predictions are all the same, the correlation for that sample is scored as -1.0.","metadata":{}},{"cell_type":"code","source":"def corr_score(y_true, y_pred):\n    corr = 0\n    \n    for i in range(len(y_true)):\n        corr += np.corrcoef(y_true[i], y_pred[i])[1, 0]\n    return corr / len(y_true)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n    <h3 align=\"center\" style=\"color:darkgreen;\">Good Luck</h3>  \n</div>","metadata":{}},{"cell_type":"markdown","source":"\n \n","metadata":{}}]}