{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59094,"databundleVersionId":7010844,"sourceType":"competition"}],"dockerImageVersionId":30587,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Prediction only using linear algebra as unbiased and reproducible approach, a good ensemble target\n\nMy final submission is an ensemble of Transformer NN, Pyboost, AE based NN, and this Linear Algebra approach.  It turns out that the Linear Algebra model gets the best private leaderboard score, though it could have been better in the public score. \n\n## Least Square Transformation - conceptually simple\n1. Solve linear system, transformer, from a base cell line (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) to a target cell line ('B cells', 'Myeloid cells') of the DE using the chemicals tested in the two cell lines (15 plus 2 positive control).\n2. Apply the transformer to the DE of the base cell line/target chemicals and get the DE of the target cell line/chemicals.\n\n\n## SVD Projection and LS  - a bit more memory friendly, but the same result\n1. Project the DE data by SVD with the whole dimension. \n2. Solve linear system from a base cell line (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) to a target cell line ('B cells', 'Myeloid cells') of the projected DE using the chemicals tested in the two cell lines (15 plus 2 positive control).  \n3. Apply the transformer to the projected DE of the base cell line/target chemicals and get the projected DE of the target cell line/chemicals. \n4. Inverse the projected DE of the target cell line/chemicals to the predicted DE.\n\n\n### Both method produces essentially the same result. Instead of SVD, QR decomposition can also produce the same, sure.  In practice, there is still negligible difference due to the numerial stability though.\n\n\n\n\n# Key Findings\n\n1. Differences in DE among certain cell lines can be captured linearly very well and the linear relation is well transferable across different chemical response, 0.768 in private and 0.582 in public \n2. DE of NK and T4 cells are quite predictable for that of B Cell and Myeloid cells\n3. T4 is better predictor for B cell, and NK is a better predictor for Myeloid cells","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-12-01T19:47:04.635271Z","iopub.execute_input":"2023-12-01T19:47:04.635962Z","iopub.status.idle":"2023-12-01T19:47:05.078992Z","shell.execute_reply.started":"2023-12-01T19:47:04.635927Z","shell.execute_reply":"2023-12-01T19:47:05.077889Z"}}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2023-12-07T17:20:20.900682Z","iopub.execute_input":"2023-12-07T17:20:20.902089Z","iopub.status.idle":"2023-12-07T17:20:21.443628Z","shell.execute_reply.started":"2023-12-07T17:20:20.902036Z","shell.execute_reply":"2023-12-07T17:20:21.442221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\nde_train = pd.read_parquet(fn)# , index_col = 0)\nprint(de_train.shape)\n#display(de_train)\n\nfn = '/kaggle/input/open-problems-single-cell-perturbations/id_map.csv'\nid_map = pd.read_csv(fn)\nprint(id_map.shape)\n#display(id_map)\n\nfn = '/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv'\ndf_smp = pd.read_csv(fn, index_col = 0)\nprint(df_smp.shape)\n#display(df_smp)\n\ngenes = de_train.columns[5:].tolist() # 18211 genes","metadata":{"execution":{"iopub.status.busy":"2023-12-07T17:20:21.447931Z","iopub.execute_input":"2023-12-07T17:20:21.449181Z","iopub.status.idle":"2023-12-07T17:20:29.774650Z","shell.execute_reply.started":"2023-12-07T17:20:21.449119Z","shell.execute_reply":"2023-12-07T17:20:29.769983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train and Target Cell Types, Core Chemicals shared by both Cell types","metadata":{}},{"cell_type":"code","source":"temp = de_train.groupby(['cell_type']).size() > 20\ntrn_cell_types = temp[temp].index.tolist() # 4 cell types\n\ntemp2 = de_train.groupby(['cell_type']).size() < 20\ntargt_cell_types = temp2[temp2].index.tolist() # 2 cell types\n\nall_sm_name = de_train.loc[(de_train['cell_type'] == 'NK cells') & (de_train['control']==False)].sm_name.values\ncore_sm_names = de_train.query(\"cell_type == 'B cells'\").sm_name.values # 17 compounds including the two control compounds","metadata":{"execution":{"iopub.status.busy":"2023-12-07T17:20:29.776206Z","iopub.execute_input":"2023-12-07T17:20:29.776813Z","iopub.status.idle":"2023-12-07T17:20:31.276412Z","shell.execute_reply.started":"2023-12-07T17:20:29.776762Z","shell.execute_reply":"2023-12-07T17:20:31.275131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(trn_cell_types)\nprint(targt_cell_types)\nprint(core_sm_names)","metadata":{"execution":{"iopub.status.busy":"2023-12-07T17:20:31.279408Z","iopub.execute_input":"2023-12-07T17:20:31.280866Z","iopub.status.idle":"2023-12-07T17:20:31.288437Z","shell.execute_reply.started":"2023-12-07T17:20:31.280797Z","shell.execute_reply":"2023-12-07T17:20:31.287008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cell_ab = {'NK cells':'NK', 'T cells CD4+':'T4', 'T cells CD8+':'T8', 'T regulatory cells':'RG', 'avg':'avg'}","metadata":{"execution":{"iopub.status.busy":"2023-12-07T17:20:31.290250Z","iopub.execute_input":"2023-12-07T17:20:31.290602Z","iopub.status.idle":"2023-12-07T17:20:31.299946Z","shell.execute_reply.started":"2023-12-07T17:20:31.290573Z","shell.execute_reply":"2023-12-07T17:20:31.298289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create DE for Cell lines","metadata":{}},{"cell_type":"code","source":"_de_list = ['cell_type','sm_lincs_id','SMILES','control']\n\ndic_X_de={}\n\nfor tct in trn_cell_types:\n    t_msk = de_train['cell_type']==tct\n    dic_X_de[tct]=de_train[t_msk].drop(columns=_de_list).set_index('sm_name')\n    print(tct, dic_X_de[tct].shape)\n\nfor bmt in targt_cell_types:\n    t_msk = de_train['cell_type']==bmt\n    dic_X_de[bmt]=de_train[t_msk].drop(columns=_de_list).set_index('sm_name')\n    print(bmt, dic_X_de[bmt].shape)\n    ","metadata":{"execution":{"iopub.status.busy":"2023-12-07T17:20:32.910456Z","iopub.execute_input":"2023-12-07T17:20:32.910957Z","iopub.status.idle":"2023-12-07T17:20:33.109386Z","shell.execute_reply.started":"2023-12-07T17:20:32.910924Z","shell.execute_reply":"2023-12-07T17:20:33.107855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Fill the missing chemicals in T8 by T4 data","metadata":{}},{"cell_type":"code","source":"missing_chem = list(set(dic_X_de['T cells CD4+'].index.to_list()) - set(dic_X_de['T cells CD8+'].index.to_list()))\nT8_miss_de = dic_X_de['T cells CD4+'].loc[missing_chem]\ndic_X_de['T cells CD8+'] = pd.concat((dic_X_de['T cells CD8+'], T8_miss_de), axis=0)\n\nfor cl in dic_X_de.keys():\n    print(cl, dic_X_de[cl].shape)","metadata":{"execution":{"iopub.status.busy":"2023-12-07T17:20:36.578465Z","iopub.execute_input":"2023-12-07T17:20:36.579083Z","iopub.status.idle":"2023-12-07T17:20:36.611637Z","shell.execute_reply.started":"2023-12-07T17:20:36.579037Z","shell.execute_reply":"2023-12-07T17:20:36.610201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# SVD model","metadata":{}},{"cell_type":"code","source":"U_i, S_i, Vh_i = np.linalg.svd(de_train.iloc[:, 5:], full_matrices=False)","metadata":{"execution":{"iopub.status.busy":"2023-12-07T17:20:41.333988Z","iopub.execute_input":"2023-12-07T17:20:41.334562Z","iopub.status.idle":"2023-12-07T17:20:50.335753Z","shell.execute_reply.started":"2023-12-07T17:20:41.334516Z","shell.execute_reply":"2023-12-07T17:20:50.334490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# DEA: Observation of Similarity among cell lines","metadata":{}},{"cell_type":"code","source":"prj_NK = np.matmul(dic_X_de['NK cells'].loc[core_sm_names], Vh_i.T)\nprj_T4 = np.matmul(dic_X_de['T cells CD4+'].loc[core_sm_names], Vh_i.T)\nprj_T8 = np.matmul(dic_X_de['T cells CD8+'].loc[core_sm_names], Vh_i.T)\nprj_RG = np.matmul(dic_X_de['T regulatory cells'].loc[core_sm_names], Vh_i.T)\nprj_BC = np.matmul(dic_X_de['B cells'].loc[core_sm_names], Vh_i.T)\nprj_MC = np.matmul(dic_X_de['Myeloid cells'].loc[core_sm_names], Vh_i.T)\n\nprj_all = pd.concat([prj_NK, prj_T4, prj_T8, prj_RG, prj_BC, prj_MC], axis=0).reset_index()\nprj_all.columns = prj_all.columns.map(str)\nprj_all.insert(loc=0, column='cell_type', value=np.repeat(['NK cells','T cells CD4+','T cells CD8+',\n                                                        'T regulatory cells','B cells','Myeloid cells'],17))\nprj_all","metadata":{"execution":{"iopub.status.busy":"2023-12-07T17:21:57.124662Z","iopub.execute_input":"2023-12-07T17:21:57.125133Z","iopub.status.idle":"2023-12-07T17:21:57.410767Z","shell.execute_reply.started":"2023-12-07T17:21:57.125100Z","shell.execute_reply":"2023-12-07T17:21:57.409243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2023-12-07T17:22:10.456042Z","iopub.execute_input":"2023-12-07T17:22:10.456468Z","iopub.status.idle":"2023-12-07T17:22:11.341244Z","shell.execute_reply.started":"2023-12-07T17:22:10.456437Z","shell.execute_reply":"2023-12-07T17:22:11.339788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Well, it is a bit difficult to recognize a clear pattern.\n\nsns.scatterplot(prj_all, x=\"0\", y = \"1\", hue=\"cell_type\")","metadata":{"execution":{"iopub.status.busy":"2023-12-07T17:23:16.496268Z","iopub.execute_input":"2023-12-07T17:23:16.497223Z","iopub.status.idle":"2023-12-07T17:23:17.099462Z","shell.execute_reply.started":"2023-12-07T17:23:16.497182Z","shell.execute_reply":"2023-12-07T17:23:17.097897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here is a grid plot of the previous one by chemicals. In most of the chemicals, there is not much difference among 6 cell cline. However, NK cell seems to be similar to the B Cell and Myoloid cell on the chemical perturbations that create the larger difference among 6 cell lines such as Belinostat (one of the positive controls), MNL 2238, and Oprozomib. In the end, accuracy in predicting such chemicals may be important. ","metadata":{}},{"cell_type":"code","source":"sg_plot = sns.relplot(prj_all, x=\"0\", y = \"1\", hue=\"cell_type\", col=\"sm_name\", col_wrap=2)","metadata":{"execution":{"iopub.status.busy":"2023-12-07T17:25:04.612706Z","iopub.execute_input":"2023-12-07T17:25:04.613142Z","iopub.status.idle":"2023-12-07T17:25:14.084131Z","shell.execute_reply.started":"2023-12-07T17:25:04.613108Z","shell.execute_reply":"2023-12-07T17:25:14.082732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train Transformer","metadata":{}},{"cell_type":"markdown","source":"## Create chem set to be used for creating transformer\nTo identify chemicals that are available for both base and target cell lines","metadata":{}},{"cell_type":"code","source":"# Core chemicals\ndic_trn_names = {}\n\nfor tct in trn_cell_types:\n    dic_trn_names[tct]={}\n    for bmt in targt_cell_types:\n        dic_trn_names[tct][bmt]=[]\n        \nfor tct in trn_cell_types:\n    for bmt in targt_cell_types:\n        for ccn in core_sm_names:\n            if de_train[de_train['sm_name']==ccn]['cell_type'].isin([tct, bmt]).sum()==2:\n                dic_trn_names[tct][bmt].append(ccn)\n            \n                \nfor tct in trn_cell_types:\n    for bmt in targt_cell_types:\n        print(tct, bmt, len(dic_trn_names[tct][bmt]))      ","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:35:16.208935Z","iopub.execute_input":"2023-12-02T15:35:16.209304Z","iopub.status.idle":"2023-12-02T15:35:16.583768Z","shell.execute_reply.started":"2023-12-02T15:35:16.209274Z","shell.execute_reply":"2023-12-02T15:35:16.582411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## DE to create transformer","metadata":{}},{"cell_type":"code","source":"dic_trns_in={}\ndic_trns_true={}\n\nfor tct in trn_cell_types:\n    dic_trns_in[tct]={}\n    dic_trns_true[tct]={}\n    for bmt in targt_cell_types:\n         dic_trns_in[tct][bmt]  = dic_X_de[tct].loc[dic_trn_names[tct][bmt]]\n         dic_trns_true[tct][bmt]= dic_X_de[bmt].loc[dic_trn_names[tct][bmt]]","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:35:16.585338Z","iopub.execute_input":"2023-12-02T15:35:16.585790Z","iopub.status.idle":"2023-12-02T15:35:16.643736Z","shell.execute_reply.started":"2023-12-02T15:35:16.585759Z","shell.execute_reply":"2023-12-02T15:35:16.642721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Transformer Model","metadata":{}},{"cell_type":"code","source":"# Creating transformers by SVD projection and direct LS, just for fun\n\ndic_transfer = {}\ndic_transfer_ls = {}\n\nfor tct in trn_cell_types: \n    \n    dic_transfer[tct] = {}\n    dic_transfer_ls[tct] = {}\n\n    for bmt in targt_cell_types:  \n\n        # Using SVD\n        # Step 1: Project the DE data by SVD with the whole dimension. \n        prjct_in =  np.matmul(dic_trns_in[tct][bmt],   Vh_i.T)\n        prjct_tg =  np.matmul(dic_trns_true[tct][bmt], Vh_i.T)\n\n        # Step 2: Solve the linear transformation from a base cell line\n        \n        ## get pseudo invers of the projection \n        prjct_in_inv = np.linalg.pinv(prjct_in)\n\n        ## get transformer by the target and the pseudo invers of the projection \n        trans  = np.matmul(prjct_in_inv, prjct_tg)\n        \n        dic_transfer[tct][bmt] = trans\n        \n        #----------------------------------------------------------\n        \n        # Using Leqst Square, simpler but using more memory\n        # Using direct LS solution\n        in_inv_s = np.linalg.pinv(dic_trns_in[tct][bmt])\n        trans_s = np.matmul(in_inv_s , dic_trns_true[tct][bmt])\n        \n        # lstsq uses too much memory for notebook\n        #trans_s = np.linalg.lstsq(dic_trns_in['NK cells']['B cells'], dic_trns_true['NK cells']['B cells'], rcond=None)[0]\n        \n        dic_transfer_ls[tct][bmt] = trans_s\n        ","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:36:14.559706Z","iopub.execute_input":"2023-12-02T15:36:14.560096Z","iopub.status.idle":"2023-12-02T15:36:30.525567Z","shell.execute_reply.started":"2023-12-02T15:36:14.560065Z","shell.execute_reply":"2023-12-02T15:36:30.523207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction","metadata":{}},{"cell_type":"code","source":"def prediction_svd(pred_cell_b, pred_cell_m):\n    \n    id_map_b = id_map[id_map['cell_type']=='B cells']\n    id_map_m = id_map[id_map['cell_type']=='Myeloid cells']\n\n    chms_b = id_map_b['sm_name']\n    chms_m = id_map_m['sm_name']\n    \n    de_x_b = dic_X_de[pred_cell_b].loc[chms_b].values\n    de_x_m = dic_X_de[pred_cell_m].loc[chms_m].values\n    \n    \n    # Step 3: for B cells\n    # Apply the transformer to the projected DE of the base cell line/target chemicals \n    # and get the projected DE of the target cell line/chemicals\n    prjct_b =  np.matmul(de_x_b, Vh_i.T)\n    transfer_b = dic_transfer[pred_cell_b]['B cells']\n    transformed_b = np.matmul(prjct_b, transfer_b)\n    \n    # Step 4: for B cells, Inverse the projected DE of the target cell line/chemicals to the predicted DE.\n    pred_b = np.matmul(transformed_b.values, Vh_i)\n    \n    # Step 3: for Myeloid cells\n    prjct_m =  np.matmul(de_x_m, Vh_i.T)\n    transfer_m = dic_transfer[pred_cell_m]['Myeloid cells']\n    transformed_m = np.matmul(prjct_m, transfer_m)\n    \n    # Step 4: for Myeloid cells, Inverse the projected DE of the target cell line/chemicals to the predicted DE.\n    pred_m = np.matmul(transformed_m.values, Vh_i)\n    \n    pred_b = pd.DataFrame(pred_b, columns=genes, index=chms_b)\n    pred_m = pd.DataFrame(pred_m, columns=genes, index=chms_m)\n    \n    pred_b_s = pred_b.loc[chms_b]\n    pred_m_s = pred_m.loc[chms_m]\n\n    pred_submit = pd.concat([pred_b_s, pred_m_s], ignore_index=True)\n    pred_submit.columns=genes\n    pred_submit.index.name=\"id\"\n    \n    return pred_submit\n    ","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:36:54.172240Z","iopub.execute_input":"2023-12-02T15:36:54.172680Z","iopub.status.idle":"2023-12-02T15:36:54.185878Z","shell.execute_reply.started":"2023-12-02T15:36:54.172650Z","shell.execute_reply":"2023-12-02T15:36:54.184822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def prediction_ls(pred_cell_b, pred_cell_m):\n    \n    id_map_b = id_map[id_map['cell_type']=='B cells']\n    id_map_m = id_map[id_map['cell_type']=='Myeloid cells']\n\n    chms_b = id_map_b['sm_name']\n    chms_m = id_map_m['sm_name']\n    \n    de_x_b = dic_X_de[pred_cell_b].loc[chms_b].values\n    de_x_m = dic_X_de[pred_cell_m].loc[chms_m].values\n    \n    # Step 2: for B cells\n    transfer_b = dic_transfer_ls[pred_cell_b]['B cells']\n    pred_b = np.matmul(de_x_b, transfer_b.values)\n    \n    # Step 2: for Myeloid cells\n    transfer_m = dic_transfer_ls[pred_cell_m]['Myeloid cells']\n    pred_m = np.matmul(de_x_m, transfer_m.values)\n    \n    pred_b = pd.DataFrame(pred_b, columns=genes, index=chms_b)\n    pred_m = pd.DataFrame(pred_m, columns=genes, index=chms_m)\n    \n    pred_b_s = pred_b.loc[chms_b]\n    pred_m_s = pred_m.loc[chms_m]\n\n    pred_submit = pd.concat([pred_b_s, pred_m_s], ignore_index=True)\n    pred_submit.columns=genes\n    pred_submit.index.name=\"id\"\n    \n    return pred_submit","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:37:00.312841Z","iopub.execute_input":"2023-12-02T15:37:00.313331Z","iopub.status.idle":"2023-12-02T15:37:00.325734Z","shell.execute_reply.started":"2023-12-02T15:37:00.313297Z","shell.execute_reply":"2023-12-02T15:37:00.323384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sol_svd = {}\nsol_ls = {}\n\n\nfor pred_cell in trn_cell_types:\n\n   # out_path = os.path.join(out_file_name)\n\n    pred_submit = prediction_svd(pred_cell, pred_cell)\n    pred_submit_s = prediction_ls(pred_cell, pred_cell)\n    \n    sol_svd[pred_cell] = pred_submit\n    sol_ls[pred_cell] = pred_submit_s\n    ","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:37:12.643725Z","iopub.execute_input":"2023-12-02T15:37:12.644751Z","iopub.status.idle":"2023-12-02T15:37:32.415156Z","shell.execute_reply.started":"2023-12-02T15:37:12.644693Z","shell.execute_reply":"2023-12-02T15:37:32.413740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## There is slightl difference between SVD and Direct LS solution, due to numerical stability","metadata":{}},{"cell_type":"code","source":"sol_ls['NK cells'].loc[0][0]","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:37:32.417945Z","iopub.execute_input":"2023-12-02T15:37:32.419059Z","iopub.status.idle":"2023-12-02T15:37:32.431381Z","shell.execute_reply.started":"2023-12-02T15:37:32.419008Z","shell.execute_reply":"2023-12-02T15:37:32.429987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sol_svd['NK cells'].loc[0][0]","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:37:32.432849Z","iopub.execute_input":"2023-12-02T15:37:32.433790Z","iopub.status.idle":"2023-12-02T15:37:32.444139Z","shell.execute_reply.started":"2023-12-02T15:37:32.433754Z","shell.execute_reply":"2023-12-02T15:37:32.443212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### but you may see they are basically the same","metadata":{}},{"cell_type":"code","source":"sol_ls['NK cells']","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:37:36.918366Z","iopub.execute_input":"2023-12-02T15:37:36.918785Z","iopub.status.idle":"2023-12-02T15:37:36.966282Z","shell.execute_reply.started":"2023-12-02T15:37:36.918754Z","shell.execute_reply":"2023-12-02T15:37:36.964980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sol_svd['NK cells']","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:37:41.807320Z","iopub.execute_input":"2023-12-02T15:37:41.807743Z","iopub.status.idle":"2023-12-02T15:37:41.850004Z","shell.execute_reply.started":"2023-12-02T15:37:41.807711Z","shell.execute_reply":"2023-12-02T15:37:41.848718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for pred_cell in trn_cell_types:\n    out_file_name = 'SVD_'+ cell_ab[pred_cell]+'.csv'\n    sol_svd[pred_cell].to_csv(out_file_name)\n\n    out_file_name = 'LS_'+ cell_ab[pred_cell]+'.csv'\n    sol_ls[pred_cell].to_csv(out_file_name)","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:37:44.899864Z","iopub.execute_input":"2023-12-02T15:37:44.900402Z","iopub.status.idle":"2023-12-02T15:39:29.639467Z","shell.execute_reply.started":"2023-12-02T15:37:44.900358Z","shell.execute_reply":"2023-12-02T15:39:29.637968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# by NK private:0.784 , public:0.596\nens_v21_nk =  pd.read_csv('/kaggle/working/SVD_NK.csv', index_col=0)\n\n# by T4 private:0.775 , public:607\nens_v21_t4 =  pd.read_csv('/kaggle/working/SVD_T4.csv', index_col=0)\n\n# by T8  private:0.959, public:0.706, worse than sample submission\nens_v21_t8 =  pd.read_csv('/kaggle/working/SVD_T8.csv', index_col=0)\n\n# by RG private:0.834, public:0.680, worse than sample submission\nens_v21_rg =  pd.read_csv('/kaggle/working/SVD_RG.csv', index_col=0)\n","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:39:51.433774Z","iopub.execute_input":"2023-12-02T15:39:51.434264Z","iopub.status.idle":"2023-12-02T15:40:12.890101Z","shell.execute_reply.started":"2023-12-02T15:39:51.434224Z","shell.execute_reply":"2023-12-02T15:40:12.889144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Stacking all, private: 0.771, public: 0.583","metadata":{}},{"cell_type":"code","source":"# private: 0.771, public: 0.583\nens_all = ens_v21_nk * 0.4 + ens_v21_t4 * 0.4 + ens_v21_t8 * 0.1 + ens_v21_rg * 0.1\nens_all.to_csv(\"SVD_nk_t4_t8_rg.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:40:12.893204Z","iopub.execute_input":"2023-12-02T15:40:12.893537Z","iopub.status.idle":"2023-12-02T15:40:26.092789Z","shell.execute_reply.started":"2023-12-02T15:40:12.893506Z","shell.execute_reply":"2023-12-02T15:40:26.091399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Only using NK and T4, private: 0.768, public: 0.582\n\nSimple stacking of NK and T4 based prediction is better than stacking all in both public and private.  This result makes sense as T8 and RG doesn't perform well, even worse than the sample submission. \n","metadata":{}},{"cell_type":"code","source":"# private: 0.768, public: 0.582\nens_all = ens_v21_nk * 0.5 + ens_v21_t4 * 0.5\nens_all.to_csv(\"SVD_nk_t4.csv\")\nens_all.to_csv(\"submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:40:26.094713Z","iopub.execute_input":"2023-12-02T15:40:26.095191Z","iopub.status.idle":"2023-12-02T15:40:52.197972Z","shell.execute_reply.started":"2023-12-02T15:40:26.095151Z","shell.execute_reply":"2023-12-02T15:40:52.196851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Chekc to Predict B Call and Myeloid cells by Different Base cell","metadata":{}},{"cell_type":"markdown","source":"T4 is better predictor for B cell, and NK is a better predictor for Myeloid cells","metadata":{}},{"cell_type":"code","source":"# by NK private:0.784 , public:0.596\n# by T4 private:0.775 , public:607\n\n# Predict B Cell by NK, and Myeloid cells by T4, private: 0.786, public: 0.616 \npred_submit_nk_t4 = prediction_svd('NK cells', 'T cells CD4+')\npred_submit_nk_t4.to_csv(\"SVD_b_nk_m_t4.csv\")\n\n# Predict B Cell by T4, and Myeloid cells by NK, private: 0.773, public: 0.587\npred_submit_t4_nk = prediction_svd('T cells CD4+', 'NK cells')\npred_submit_t4_nk.to_csv(\"SVD_b_t4_m_nk.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-12-02T15:41:18.851778Z","iopub.execute_input":"2023-12-02T15:41:18.852672Z","iopub.status.idle":"2023-12-02T15:41:45.539748Z","shell.execute_reply.started":"2023-12-02T15:41:18.852611Z","shell.execute_reply":"2023-12-02T15:41:45.538138Z"},"trusted":true},"execution_count":null,"outputs":[]}]}