{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n    <h1 align=\"center\" style=\"color:darkgreen;\">Multimodal Single-Cell Integration</h1>  \n</div>\n\n<div>\n    <h1 align=\"center\" style=\"color:darkgray;\">Splitting up - 1</h1>\n</div>\n\n<img src=\"https://openproblems.bio/media/learning/central-dogma-large.png\">\n\n<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"> In the [first notebook](https://www.kaggle.com/code/mehrankazeminia/1-5-msci22-eda), it became clear that in this challenge, in addition to train_data and test_data, there is interesting and important information in the \"metadata\" file. This information includes four items: \"**day**\", \"**donor**\", \"**cell_type**\" and \"**technology**\".\n>  \n> We have already seen that the two **technologies**  used in this challenge (**CITEseq** and **Multiome**) are different in terms of features and targets, and they must be modeled separately.\n> \n> Also, it seems that the importance of \"**cell_type**\" is much higher than \"**day**\" and \"**donor**\". Of course, this issue can be checked by clustering (or using the NearestNeighbors library) on the train_cite_targets and train_multi_targets files and then comparing the results.\n> \n> However, we can use \"**cell_type**\" in two ways. The first method is to add \"**cell_type**\" information (or even \"**donor**\" information) as a column (feature) to train_data and test_data. The second method is that we can **split** train_data and test_data based on each different \"**cell_type**\" so that completely separate calculations are possible (similar to what we usually do in cross-validation), of course, here the models will be independent. Also, in the second case, we will have smaller files that are easier to work with.\n> \n> In this notebook; I am trying to **split** all files related to **CITEseq** technology. The files associated with **Multiome** technology are larger and difficult to work with on these types of notebooks. Methods that use \"**data flow**\" are more suitable. For example, using \"**mrjob**\" which should be done outside the notebook. Of course, if there is a chance; I just publish the \"**mrjob**\" codes separately in a notebook.\n>\n> You can also find the results of this notebook at [this address](https://www.kaggle.com/datasets/mehrankazeminia/msci22-citeseqsplit).","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"import warnings # suppress warnings\nwarnings.filterwarnings('ignore')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os, gc\nimport numpy as np \nimport pandas as pd\nimport seaborn as sns\nfrom scipy import stats\n\n!ls ../input/*","metadata":{"execution":{"iopub.status.busy":"2022-09-13T23:14:51.6046Z","iopub.execute_input":"2022-09-13T23:14:51.605177Z","iopub.status.idle":"2022-09-13T23:14:53.312764Z","shell.execute_reply.started":"2022-09-13T23:14:51.605066Z","shell.execute_reply":"2022-09-13T23:14:53.311022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install tables","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-09-13T23:14:53.315698Z","iopub.execute_input":"2022-09-13T23:14:53.316091Z","iopub.status.idle":"2022-09-13T23:15:04.557301Z","shell.execute_reply.started":"2022-09-13T23:14:53.316053Z","shell.execute_reply":"2022-09-13T23:15:04.55567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.preprocessing import MinMaxScaler, StandardScaler","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">Meta Data</h1>\n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"metadata = pd.read_csv('../input/open-problems-multimodal/metadata.csv')\nmetadata","metadata":{"execution":{"iopub.status.busy":"2022-09-13T23:15:04.560391Z","iopub.execute_input":"2022-09-13T23:15:04.561083Z","iopub.status.idle":"2022-09-13T23:15:04.892135Z","shell.execute_reply.started":"2022-09-13T23:15:04.560971Z","shell.execute_reply":"2022-09-13T23:15:04.890665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:darkred;\">Descriptions:</span>\n\n- **cell_id** - A unique identifier for each observed cell.\n\n- **donor** - An identifier for the four cell donors.\n\n- **day** - The day of the experiment the observation was made.\n\n- **technology** - Either citeseq or multiome.\n\n- **cell_type** - One of the above cell types or else hidden.\n\n-----\n\n### <span style=\"color:darkred;\">Cell types:</span>\n\n- **MasP** = Mast Cell Progenitor\n\n- **MkP** = Megakaryocyte Progenitor\n\n- **NeuP** = Neutrophil Progenitor\n\n- **MoP** = Monocyte Progenitor\n\n- **EryP** = Erythrocyte Progenitor\n\n- **HSC** = Hematoploetic Stem Cell\n\n- **BP** = B-Cell Progenitor\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"sns.set()\nmetadata['technology'].value_counts(normalize=True).plot(kind='barh', figsize=(15,1))\nmetadata['technology'].value_counts(normalize=True), metadata['technology'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-09-13T23:15:04.895088Z","iopub.execute_input":"2022-09-13T23:15:04.895531Z","iopub.status.idle":"2022-09-13T23:15:05.214269Z","shell.execute_reply.started":"2022-09-13T23:15:04.895492Z","shell.execute_reply":"2022-09-13T23:15:05.212848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata['cell_type'].value_counts(normalize=True).plot(kind='barh', figsize=(15,3))\nmetadata['cell_type'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-09-13T23:15:05.216044Z","iopub.execute_input":"2022-09-13T23:15:05.216432Z","iopub.status.idle":"2022-09-13T23:15:05.560317Z","shell.execute_reply.started":"2022-09-13T23:15:05.216398Z","shell.execute_reply":"2022-09-13T23:15:05.559152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata.groupby('technology')['cell_type'].value_counts(normalize=True).plot(kind='barh', figsize=(15,5))\nmetadata.groupby('technology')['cell_type'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-09-13T23:15:05.562235Z","iopub.execute_input":"2022-09-13T23:15:05.562652Z","iopub.status.idle":"2022-09-13T23:15:06.135705Z","shell.execute_reply.started":"2022-09-13T23:15:05.562614Z","shell.execute_reply":"2022-09-13T23:15:06.134511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata.set_index('cell_id').pivot(columns=['technology', 'cell_type'], values= 'cell_type')","metadata":{"execution":{"iopub.status.busy":"2022-09-13T23:15:06.137294Z","iopub.execute_input":"2022-09-13T23:15:06.138239Z","iopub.status.idle":"2022-09-13T23:15:07.394693Z","shell.execute_reply.started":"2022-09-13T23:15:06.138199Z","shell.execute_reply":"2022-09-13T23:15:07.393409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-13T23:15:07.39683Z","iopub.execute_input":"2022-09-13T23:15:07.397297Z","iopub.status.idle":"2022-09-13T23:15:07.564665Z","shell.execute_reply.started":"2022-09-13T23:15:07.397257Z","shell.execute_reply":"2022-09-13T23:15:07.563315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">metadata_cite</span>","metadata":{}},{"cell_type":"code","source":"metadata_cite = metadata [metadata.technology == 'citeseq']\nmetadata_cite","metadata":{"execution":{"iopub.status.busy":"2022-09-13T23:15:07.566707Z","iopub.execute_input":"2022-09-13T23:15:07.567377Z","iopub.status.idle":"2022-09-13T23:15:07.615083Z","shell.execute_reply.started":"2022-09-13T23:15:07.567335Z","shell.execute_reply":"2022-09-13T23:15:07.613878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cite1_masp = metadata_cite.cell_id [metadata_cite.cell_type == 'MasP']\ncite2_mkp  = metadata_cite.cell_id [metadata_cite.cell_type == 'MkP' ]\ncite3_neup = metadata_cite.cell_id [metadata_cite.cell_type == 'NeuP']\ncite4_mop  = metadata_cite.cell_id [metadata_cite.cell_type == 'MoP' ]\ncite5_eryp = metadata_cite.cell_id [metadata_cite.cell_type == 'EryP']\ncite6_hsc  = metadata_cite.cell_id [metadata_cite.cell_type == 'HSC' ]\ncite7_bp   = metadata_cite.cell_id [metadata_cite.cell_type == 'BP'  ]\n\ncite_list  = [cite1_masp, cite2_mkp, cite3_neup, cite4_mop, cite5_eryp, \n              cite6_hsc, cite7_bp]","metadata":{"execution":{"iopub.status.busy":"2022-09-13T23:15:07.618608Z","iopub.execute_input":"2022-09-13T23:15:07.619293Z","iopub.status.idle":"2022-09-13T23:15:07.704223Z","shell.execute_reply.started":"2022-09-13T23:15:07.619254Z","shell.execute_reply":"2022-09-13T23:15:07.703106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lencite = 0\nfor n, cite in enumerate(cite_list): \n    \n    print(len(cite))\n    lencite += len(cite)\n    cite.to_csv(f'cite_id_cell_{n+1}.csv', index=False)\n    gc.collect()\n    \nprint('Total number of rows (to check):', lencite)\nprint('---------------------------------------\\n')\n!ls","metadata":{"execution":{"iopub.status.busy":"2022-09-13T23:15:07.70577Z","iopub.execute_input":"2022-09-13T23:15:07.706644Z","iopub.status.idle":"2022-09-13T23:15:09.763909Z","shell.execute_reply.started":"2022-09-13T23:15:07.706608Z","shell.execute_reply":"2022-09-13T23:15:09.762217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata_cite_cell_id = metadata_cite.cell_id","metadata":{"execution":{"iopub.status.busy":"2022-09-13T23:15:09.766374Z","iopub.execute_input":"2022-09-13T23:15:09.767023Z","iopub.status.idle":"2022-09-13T23:15:09.772713Z","shell.execute_reply.started":"2022-09-13T23:15:09.766979Z","shell.execute_reply":"2022-09-13T23:15:09.771485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del metadata_cite\ndel cite1_masp, cite2_mkp, cite3_neup, cite4_mop, cite5_eryp, cite6_hsc, cite7_bp\ndel lencite \ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">metadata_multi</span>\n#### <span style=\"color:darkblue;\">(For use in the next notebooks)</span>","metadata":{}},{"cell_type":"code","source":"metadata_multi = metadata [metadata.technology == 'multiome']\nmetadata_multi","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"multi1_masp = metadata_multi.cell_id [metadata_multi.cell_type == 'MasP']\nmulti2_mkp  = metadata_multi.cell_id [metadata_multi.cell_type == 'MkP' ]\nmulti3_neup = metadata_multi.cell_id [metadata_multi.cell_type == 'NeuP']\nmulti4_mop  = metadata_multi.cell_id [metadata_multi.cell_type == 'MoP' ]\nmulti5_eryp = metadata_multi.cell_id [metadata_multi.cell_type == 'EryP']\nmulti6_hsc  = metadata_multi.cell_id [metadata_multi.cell_type == 'HSC' ]\nmulti7_bp   = metadata_multi.cell_id [metadata_multi.cell_type == 'BP'  ]\nmulti8_hide = metadata_multi.cell_id [metadata_multi.cell_type == 'hidden']\n\nmulti_list  = [multi1_masp, multi2_mkp, multi3_neup, multi4_mop, multi5_eryp, \n               multi6_hsc, multi7_bp, multi8_hide]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lenmulti = 0\nfor n, multi in enumerate(multi_list): \n      \n    print(len(multi))\n    lenmulti += len(multi)\n    multi.to_csv(f'multi_id_cell_{n+1}.csv', index=False)\n    gc.collect()\n    \n                \nprint('Total number of rows (to check):', lenmulti)\nprint('---------------------------------------\\n')\n!ls","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del metadata\ndel metadata_multi\ndel multi1_masp, multi2_mkp, multi3_neup, multi4_mop, multi5_eryp, multi6_hsc, multi7_bp, multi8_hide \ndel multi_list\ndel lenmulti\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">CITEseq Technology</h1>\n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">train_cite_inputs</span>","metadata":{}},{"cell_type":"code","source":"train_cite = pd.read_hdf('../input/open-problems-multimodal/train_cite_inputs.h5')\ntrain_cite","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:darkred;\">Shapiro–Wilk test:</span>","metadata":{}},{"cell_type":"code","source":"cols = list(train_cite.columns)\nlen(cols)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_select = []\nalpha = 0.05\n\nfor col in cols:\n    _, p_value = stats.shapiro(train_cite[col])\n    \n    if (p_value <= alpha): \n        cols_select.append(col) \n        \n    gc.collect()        \nlen(cols_select)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_cite = train_cite[cols_select]\ntrain_cite.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:darkred;\">Columns with a fixed value of zero:</span>","metadata":{}},{"cell_type":"code","source":"print('The number of columns with a fixed value of zero:', (train_cite == 0).all(axis=0).sum())\n\n# cols_z = (train_cite == 0).all(axis=0)\n# cols_z = cols_z[cols_z == False] \n\n# train_cite = train_cite[cols_z.index]\n# train_cite.shape  ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### So the columns that have only zero value have already been removed with the Shapiro test.","metadata":{}},{"cell_type":"code","source":"# train_cite.to_csv('cite_train.csv', index=False)\n# train_cite.to_hdf('cite_train.h5', key='train_cite')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del cols \ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkred;\">Scaling:</span>","metadata":{}},{"cell_type":"code","source":"# index = train_cite.index\n# scaler = MinMaxScaler()\n# train_cite = pd.DataFrame(scaler.fit_transform(train_cite), index=index)\n# del index\n# gc.collect()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:darkred;\">Common elements comparison (to check):</span>","metadata":{}},{"cell_type":"code","source":"common = list(set(list(train_cite.index)).intersection(list(metadata_cite_cell_id)))\nlen(common)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">To Separate:</span>","metadata":{}},{"cell_type":"code","source":"lencite_train = 0\nfor n, cite in enumerate(cite_list):    \n    cite_train = train_cite.loc[train_cite.index.isin(list(cite))]\n    \n    print(len(cite_train))\n    lencite_train += len(cite_train)\n    cite_train.to_csv(f'cite_train_{n+1}.csv') \n    # cite_train.to_hdf(f'cite_train_{n+1}.h5', key='cite_train') \n    gc.collect()\n    \nprint('Total number of rows (to check):', lencite_train)\nprint('---------------------------------------\\n')\n!ls","metadata":{"_kg_hide-output":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del common\ndel train_cite\ndel lencite_train\ngc.collect()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">target_cite_inputs</span>","metadata":{}},{"cell_type":"code","source":"target_cite = pd.read_hdf('../input/open-problems-multimodal/train_cite_targets.h5')\ntarget_cite","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:darkred;\">Common elements comparison (to check):</span>","metadata":{}},{"cell_type":"code","source":"common = list(set(list(target_cite.index)).intersection(list(metadata_cite_cell_id)))\nlen(common)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">To Separate:</span>","metadata":{}},{"cell_type":"code","source":"lencite_target = 0\nfor n, cite in enumerate(cite_list):     \n    cite_target = target_cite.loc[target_cite.index.isin(list(cite))]\n    \n    print(len(cite_target))\n    lencite_target += len(cite_target)\n    cite_target.to_csv(f'cite_target_{n+1}.csv') \n    # cite_target.to_hdf(f'cite_target_{n+1}.h5', key='cite_target')\n    gc.collect()\n    \nprint('Total number of rows (to check):', lencite_target)\nprint('---------------------------------------\\n')\n!ls","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del common\ndel target_cite\ndel lencite_target\ngc.collect()","metadata":{"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">test_cite_inputs</span>","metadata":{}},{"cell_type":"code","source":"test_cite = pd.read_hdf('../input/open-problems-multimodal/test_cite_inputs.h5')\ntest_cite","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_cite = test_cite[cols_select]\ntest_cite.shape","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test_cite.to_csv('cite_test.csv', index=False)\n# test_cite.to_hdf('cite_test.h5', key='test_cite')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkred;\">Scaling:</span>","metadata":{}},{"cell_type":"code","source":"# index = test_cite.index\n# test_cite = pd.DataFrame(scaler.transform(test_cite), index=index)\n# del index, scaler\n# gc.collect()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:darkred;\">Common elements comparison (to check):</span>","metadata":{}},{"cell_type":"code","source":"common = list(set(list(test_cite.index)).intersection(list(metadata_cite_cell_id)))\nlen(common)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del common\ndel cols_select\ndel metadata_cite_cell_id\ngc.collect()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">To Separate:</span>","metadata":{}},{"cell_type":"code","source":"lencite_test = 0\nfor n, cite in enumerate(cite_list[:2]):  \n    '''\n    Due to the memory limitation of this notebook;\n    You should replace the rest of the cite_list in another version \n    so that the results are complete.\n    '''  \n    cite_test = test_cite.loc[test_cite.index.isin(list(cite))]\n \n    print(len(cite_test))\n    lencite_test += len(cite_test)\n    cite_test.to_csv(f'cite_test_{n+1}.csv')        \n    # cite_test.to_hdf(f'cite_test_{n+1}.h5', key='cite_test')\n    gc.collect()\n    \nprint('Total number of rows (to check):', lencite_test)\nprint('---------------------------------------\\n')\n!ls","metadata":{"execution":{"iopub.status.busy":"2022-09-14T21:53:54.983214Z","iopub.execute_input":"2022-09-14T21:53:54.983716Z","iopub.status.idle":"2022-09-14T21:53:54.994175Z","shell.execute_reply.started":"2022-09-14T21:53:54.983674Z","shell.execute_reply":"2022-09-14T21:53:54.992853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">  \n    <h3 align=\"center\" style=\"color:darkgreen;\">Good Luck</h3>  \n</div>","metadata":{}}]}