{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n    <h1 align=\"center\" style=\"color:darkgreen;\">Multimodal Single-Cell Integration</h1>  \n</div>\n\n<div>\n    <h1 align=\"center\" style=\"color:darkgray;\">CITEseq - \nKNeighborsRegressor</h1>\n</div>\n\n<img src=\"https://openproblems.bio/media/learning/central-dogma-large.png\">\n\n<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"> In the [third notebook](https://www.kaggle.com/code/mehrankazeminia/3-5-msci22-baseline-separate-models-citeseq), there was a separate model for each \"cell_type\" and they were taught separately. But in this notebook, all the samples participate in the training and in the fifth notebook, new results are combined with previous results.\n>\n> I have used KNeighborsRegressor in this notebook, but you can use better methods. ","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"import warnings # suppress warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-10-14T14:16:52.282183Z","iopub.execute_input":"2022-10-14T14:16:52.28304Z","iopub.status.idle":"2022-10-14T14:16:52.309285Z","shell.execute_reply.started":"2022-10-14T14:16:52.282912Z","shell.execute_reply":"2022-10-14T14:16:52.308531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os, gc\nimport numpy as np \nimport pandas as pd\nfrom tqdm import tqdm\n\n!ls ../input/*","metadata":{"execution":{"iopub.status.busy":"2022-10-14T14:16:58.784738Z","iopub.execute_input":"2022-10-14T14:16:58.785084Z","iopub.status.idle":"2022-10-14T14:16:59.833069Z","shell.execute_reply.started":"2022-10-14T14:16:58.785057Z","shell.execute_reply":"2022-10-14T14:16:59.830949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.neighbors import KNeighborsRegressor","metadata":{"execution":{"iopub.status.busy":"2022-10-14T14:17:03.102992Z","iopub.execute_input":"2022-10-14T14:17:03.103485Z","iopub.status.idle":"2022-10-14T14:17:03.743929Z","shell.execute_reply.started":"2022-10-14T14:17:03.103442Z","shell.execute_reply":"2022-10-14T14:17:03.742974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">CITEseq Technology</h1>\n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\n<div>\n    <h2 align=\"center\" style=\"color:darkblue;\">(Given gene expression, predict protein levels)</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"#### \"cell_type\" is sorted according to the following order:\n\n1 - **MasP** = Mast Cell Progenitor (Number of **18090**)\n\n2 - **MkP** = Megakaryocyte Progenitor (Number of **10800**)\n\n3 - **NeuP** = Neutrophil Progenitor (Number of **21418**)\n\n4 - **MoP** = Monocyte Progenitor (Number of **1822**)\n\n5 - **EryP** = Erythrocyte Progenitor (Number of **24344**)\n\n6 - **HSC** = Hematoploetic Stem Cell (Number of **42874**)\n\n7 - **BP** = B-Cell Progenitor (Number of **303**)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">KNeighborsRegressor</span>","metadata":{}},{"cell_type":"code","source":"X = pd.read_hdf('../input/open-problems-multimodal/train_cite_inputs.h5')\ny = pd.read_hdf('../input/open-problems-multimodal/train_cite_targets.h5')\n\nneigh = KNeighborsRegressor(n_neighbors=9)\nneigh.fit(X, y) \ngc.collect()\ndel X, y \n       \nXX = pd.read_hdf('../input/open-problems-multimodal/test_cite_inputs.h5')\ndf_pred = pd.DataFrame(neigh.predict(XX))\ndel XX, neigh\ngc.collect()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.read_csv('../input/all-in-one-citeseq-multiome-with-keras/submission_lolo_total_ensembling.csv',\n                         index_col='row_id', squeeze=True)\n\nsubmission.iloc[:len(df_pred.values.ravel())] = df_pred.values.ravel()\nsubmission.to_csv('submission.csv')\n!ls","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">  \n    <h3 align=\"center\" style=\"color:darkgreen;\">Good Luck</h3>  \n</div>","metadata":{}}]}