{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nFor selected protein e.g. CD36 we compare several simple Ridge based models for predicting it:\n\nUse only one feature - own RNA for protein e.g. for CD36 protein  we use CD36 rna - and predict by Ridge only based on that model\n\nUse cell type (one-hot encoded for prediction) \n\nUse cell type + own RNA\n\nUse top 100  PCA feautures as predictors \n\nAnd also (in version 2 of the notebook): All RNA, ALL RNA + Proteins, ALL Proteins + RNA (except own), All RNA (except own RNA). It takes longer time - about 6 minutes for each. \n\n\nConclusions/Results (Pearson correlation is mentioned): \n\n    0) CD36: Pearson correlation RNA vs PROTEIN = 0.73, (Spearman just 0.5) \n    1) Cell type only is much worse than own RNA - 0.55 cs 0.73\n    2) All RNA except own - 0.72 - it is a bit worse than correlation with own rna, but just about 0.01 (!!!) \n    3) All proteins a bit better than own RNA: 0.73 vs 0.74\n    4) All RNA even more better 0.758\n    5) PCA100 is almost similar 0.761\n    6) LightGBM on just one feature - own RNA - 0.765\n    7) Proteins + RNA (except own) a very bit better - 0.765 \n    8) allRNA + Proteins - quite better results - 0.797 \n\nThe same in the table: \n\n        CD36\n    Corr Ridge cell types\t0.553421 \n    Corr Ridge all RNA except own\t0.720771    \n    Corr Ridge ownRNA\t0.738055\n    Corr ownRNA\t0.738335\n    Corr Ridge CT + ownRNA\t0.741463\n    Corr Ridge allProteins\t0.742843\n    Corr Ridge allRNA\t0.757663\n    Corr Ridge pca100\t0.761470\n    Corr LightGBM on just one feature - own RNA - 0.765\n    Corr Ridge Proteins + RNA except own\t0.765259\n    Corr Ridge allRNA+Proteins\t0.797522\n\n\n\n#### Versions\n\n##### 3,4 - added LightGBM on one feature - own RNA - uplift to 0.765\n\n##### 2 - added Ridge models on ALL rna features - take about 6 minutes each \n\n    Overall 22 minutes \n\n##### 1 - most simple and fast Ridge models\n    \n    Including PCA100\n\n    The whole notebooks runs in about 68 seconds. (Models on ALL RNA are not yet included in that version - they run quite longer time - about 5 minutes each). \n\n\n","metadata":{}},{"cell_type":"markdown","source":"# Key Params","metadata":{"execution":{"iopub.status.busy":"2022-12-18T21:18:08.728359Z","iopub.execute_input":"2022-12-18T21:18:08.729709Z","iopub.status.idle":"2022-12-18T21:18:08.735176Z","shell.execute_reply.started":"2022-12-18T21:18:08.729662Z","shell.execute_reply":"2022-12-18T21:18:08.733772Z"}}},{"cell_type":"code","source":"target_name = 'CD36'\nn_splits_for_cross_valdition = 2 # if bigger, then for models on all rna (22K features) it may cause RAM crash and would be slower - \n\nrandom_state_cross_validation = 0 # random to be used in generating the folds - changing it we can study how stable are our results\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-12-26T14:41:33.738749Z","iopub.execute_input":"2022-12-26T14:41:33.739613Z","iopub.status.idle":"2022-12-26T14:41:33.763894Z","shell.execute_reply.started":"2022-12-26T14:41:33.739506Z","shell.execute_reply":"2022-12-26T14:41:33.762936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preparations","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\n# Main results will be stored here: \ndf_scores = pd.DataFrame()","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:41:33.766474Z","iopub.execute_input":"2022-12-26T14:41:33.767370Z","iopub.status.idle":"2022-12-26T14:41:33.776272Z","shell.execute_reply.started":"2022-12-26T14:41:33.767320Z","shell.execute_reply":"2022-12-26T14:41:33.774408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.preprocessing import OneHotEncoder\n\nimport time\nt0start = time.time()\n\nimport os","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:41:33.778072Z","iopub.execute_input":"2022-12-26T14:41:33.778608Z","iopub.status.idle":"2022-12-26T14:41:34.936783Z","shell.execute_reply.started":"2022-12-26T14:41:33.778560Z","shell.execute_reply":"2022-12-26T14:41:34.935326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data for target","metadata":{}},{"cell_type":"code","source":"%%time\ndf_y = pd.read_hdf('/kaggle/input/open-problems-multimodal/train_cite_targets.h5')\ndf_y","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:41:46.082594Z","iopub.execute_input":"2022-12-26T14:41:46.082995Z","iopub.status.idle":"2022-12-26T14:41:46.944485Z","shell.execute_reply.started":"2022-12-26T14:41:46.082963Z","shell.execute_reply":"2022-12-26T14:41:46.943279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = df_y[target_name]\ny","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:41:46.946311Z","iopub.execute_input":"2022-12-26T14:41:46.946677Z","iopub.status.idle":"2022-12-26T14:41:46.974638Z","shell.execute_reply.started":"2022-12-26T14:41:46.946642Z","shell.execute_reply":"2022-12-26T14:41:46.973253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_scores.loc['Mean', target_name ]  = np.mean(y)\ndf_scores.loc['Std', target_name ]  = np.std(y)\ndf_scores\n","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:41:47.261735Z","iopub.execute_input":"2022-12-26T14:41:47.262163Z","iopub.status.idle":"2022-12-26T14:41:47.281558Z","shell.execute_reply.started":"2022-12-26T14:41:47.262129Z","shell.execute_reply":"2022-12-26T14:41:47.280433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#  Predicting based on -  own RNA \n\ne.g. for CD36 protein we take CD36 as the only predictor ","metadata":{}},{"cell_type":"code","source":"%%time\ndf_rna = pd.read_hdf('/kaggle/input/open-problems-multimodal/train_cite_inputs.h5')\ndisplay(df_rna) \n\nfor t in df_rna.columns:\n    if target_name in t: \n        print(t)\n        rna_name = t # 'ENSG00000135218_CD36'\n","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:41:53.657872Z","iopub.execute_input":"2022-12-26T14:41:53.658312Z","iopub.status.idle":"2022-12-26T14:43:10.981660Z","shell.execute_reply.started":"2022-12-26T14:41:53.658277Z","shell.execute_reply":"2022-12-26T14:43:10.980792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = df_rna[[rna_name]]\nX","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:43:10.983460Z","iopub.execute_input":"2022-12-26T14:43:10.984081Z","iopub.status.idle":"2022-12-26T14:43:10.997385Z","shell.execute_reply.started":"2022-12-26T14:43:10.984042Z","shell.execute_reply":"2022-12-26T14:43:10.996038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"s = np.corrcoef(y,X.iloc[:,0])[0,1]\ndf_scores.loc['Corr ownRNA', target_name ]  = s\ndf_scores\n\n","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:43:10.998796Z","iopub.execute_input":"2022-12-26T14:43:10.999254Z","iopub.status.idle":"2022-12-26T14:43:11.019118Z","shell.execute_reply.started":"2022-12-26T14:43:10.999199Z","shell.execute_reply":"2022-12-26T14:43:11.017819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from scipy import stats\ns = stats.spearmanr(y,X.iloc[:,0])[0] # Take only correlation , skip p-value\ndf_scores.loc['Corr Spearman ownRNA', target_name ]  = s\ndf_scores\n","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:43:11.020988Z","iopub.execute_input":"2022-12-26T14:43:11.022077Z","iopub.status.idle":"2022-12-26T14:43:11.053654Z","shell.execute_reply.started":"2022-12-26T14:43:11.022034Z","shell.execute_reply":"2022-12-26T14:43:11.052392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport lightgbm as lgbm\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold\nimport time\n\nmodel = lgbm.LGBMRegressor(random_state = 0) # Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr LGB ownRNA', target_name ]  = s\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:49:58.359619Z","iopub.execute_input":"2022-12-26T14:49:58.360231Z","iopub.status.idle":"2022-12-26T14:49:58.860482Z","shell.execute_reply.started":"2022-12-26T14:49:58.360190Z","shell.execute_reply":"2022-12-26T14:49:58.859419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nprint(X.shape, y.shape)\nmodel = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\nalpha_selected = model.alpha_\nprint('Optimal alpha found by cross-validation: ', alpha_selected )\n\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge ownRNA', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge ownRNA', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge ownRNA', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:49:58.862973Z","iopub.execute_input":"2022-12-26T14:49:58.864112Z","iopub.status.idle":"2022-12-26T14:49:59.049094Z","shell.execute_reply.started":"2022-12-26T14:49:58.864062Z","shell.execute_reply":"2022-12-26T14:49:59.047800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#  Predicting based on - cell type information ","metadata":{}},{"cell_type":"code","source":"%%time\ndf_meta = pd.read_csv('/kaggle/input/feature-shop-for-multimodal-singlecell-competition/_citeseq_meta_all_text_also.csv', index_col = 0 )\ndf_meta = df_meta[df_meta['Train0OrTest1']==0]\ndf_meta","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:49:59.051048Z","iopub.execute_input":"2022-12-26T14:49:59.051510Z","iopub.status.idle":"2022-12-26T14:49:59.480196Z","shell.execute_reply.started":"2022-12-26T14:49:59.051465Z","shell.execute_reply":"2022-12-26T14:49:59.479002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = df_meta[['HSC cell_type', 'EryP cell_type', 'NeuP cell_type','MasP cell_type', 'MkP cell_type', 'BP cell_type', 'MoP cell_type' ]]\nX ","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:49:59.484055Z","iopub.execute_input":"2022-12-26T14:49:59.485262Z","iopub.status.idle":"2022-12-26T14:49:59.504221Z","shell.execute_reply.started":"2022-12-26T14:49:59.485211Z","shell.execute_reply":"2022-12-26T14:49:59.503026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nprint(X.shape, y.shape)\nmodel = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\nalpha_selected = model.alpha_\nprint('Optimal alpha found by cross-validation: ', alpha_selected )\n\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge cell types', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge cell types', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge cell types', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:49:59.505662Z","iopub.execute_input":"2022-12-26T14:49:59.506044Z","iopub.status.idle":"2022-12-26T14:49:59.731080Z","shell.execute_reply.started":"2022-12-26T14:49:59.506011Z","shell.execute_reply":"2022-12-26T14:49:59.729435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predicting based on - cell type +  own RNA \n\n(own RNA -   RNA corresponding to the protein, e.g. for CD36 we take CD36  )","metadata":{}},{"cell_type":"code","source":"df_tmp = df_meta[['HSC cell_type', 'EryP cell_type', 'NeuP cell_type','MasP cell_type', 'MkP cell_type', 'BP cell_type', 'MoP cell_type' ]]\nX = pd.concat([ df_rna[[rna_name]],  df_tmp ], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:49:59.734006Z","iopub.execute_input":"2022-12-26T14:49:59.735217Z","iopub.status.idle":"2022-12-26T14:49:59.794839Z","shell.execute_reply.started":"2022-12-26T14:49:59.735124Z","shell.execute_reply":"2022-12-26T14:49:59.793066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.metrics import mean_squared_error\n\nprint(X.shape, y.shape)\nmodel = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\nalpha_selected = model.alpha_\nprint('Optimal alpha found by cross-validation: ', alpha_selected )\n\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge CT + ownRNA', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge CT + ownRNA', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge CT + ownRNA', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:49:59.797745Z","iopub.execute_input":"2022-12-26T14:49:59.798999Z","iopub.status.idle":"2022-12-26T14:50:00.055043Z","shell.execute_reply.started":"2022-12-26T14:49:59.798928Z","shell.execute_reply":"2022-12-26T14:50:00.053820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predicting based on PCA top 100 features","metadata":{}},{"cell_type":"code","source":"%%time\ndf_pca = pd.read_csv('/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA200.csv', index_col = 0 )\ndf_pca","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:50:00.056575Z","iopub.execute_input":"2022-12-26T14:50:00.057746Z","iopub.status.idle":"2022-12-26T14:50:10.514512Z","shell.execute_reply.started":"2022-12-26T14:50:00.057698Z","shell.execute_reply":"2022-12-26T14:50:10.513302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = df_pca.iloc[:len(y), :100]\nX","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:50:10.516199Z","iopub.execute_input":"2022-12-26T14:50:10.517130Z","iopub.status.idle":"2022-12-26T14:50:10.573155Z","shell.execute_reply.started":"2022-12-26T14:50:10.517082Z","shell.execute_reply":"2022-12-26T14:50:10.571993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.metrics import mean_squared_error\n\nprint(X.shape, y.shape)\nmodel = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\nalpha_selected = model.alpha_\nprint('Optimal alpha found by cross-validation: ', alpha_selected )\n\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= 0 )\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge pca100', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge pca100', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge pca100', target_name ]  = s\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:50:10.576940Z","iopub.execute_input":"2022-12-26T14:50:10.577330Z","iopub.status.idle":"2022-12-26T14:50:12.288090Z","shell.execute_reply.started":"2022-12-26T14:50:10.577295Z","shell.execute_reply":"2022-12-26T14:50:12.286523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:50:12.295480Z","iopub.execute_input":"2022-12-26T14:50:12.300495Z","iopub.status.idle":"2022-12-26T14:50:12.486154Z","shell.execute_reply.started":"2022-12-26T14:50:12.300413Z","shell.execute_reply":"2022-12-26T14:50:12.484811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction based on proteins - all except the target one ","metadata":{}},{"cell_type":"code","source":"X = df_y.drop(target_name , axis = 1)\nX.shape","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:50:12.488068Z","iopub.execute_input":"2022-12-26T14:50:12.488586Z","iopub.status.idle":"2022-12-26T14:50:12.513705Z","shell.execute_reply.started":"2022-12-26T14:50:12.488535Z","shell.execute_reply":"2022-12-26T14:50:12.512514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nalpha_selected = 1e4\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\ny_pred = cross_val_predict(model, X, y, cv=kf)\n\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge allProteins', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge allProteins', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge allProteins', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:50:12.515483Z","iopub.execute_input":"2022-12-26T14:50:12.515956Z","iopub.status.idle":"2022-12-26T14:50:12.834995Z","shell.execute_reply.started":"2022-12-26T14:50:12.515894Z","shell.execute_reply":"2022-12-26T14:50:12.833275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:50:12.842796Z","iopub.execute_input":"2022-12-26T14:50:12.847940Z","iopub.status.idle":"2022-12-26T14:50:12.890099Z","shell.execute_reply.started":"2022-12-26T14:50:12.847834Z","shell.execute_reply":"2022-12-26T14:50:12.888081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:50:12.899065Z","iopub.execute_input":"2022-12-26T14:50:12.905382Z","iopub.status.idle":"2022-12-26T14:50:13.055768Z","shell.execute_reply.started":"2022-12-26T14:50:12.905298Z","shell.execute_reply":"2022-12-26T14:50:13.054659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ridge on ALL RNA features ( 11 minutes of calculations ) ","metadata":{}},{"cell_type":"code","source":"X = df_rna.values\nX","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:50:13.057379Z","iopub.execute_input":"2022-12-26T14:50:13.057704Z","iopub.status.idle":"2022-12-26T14:50:13.072212Z","shell.execute_reply.started":"2022-12-26T14:50:13.057675Z","shell.execute_reply":"2022-12-26T14:50:13.070898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nalpha_selected = 1e4\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\ny_pred = cross_val_predict(model, X, y, cv=kf)\n\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge allRNA', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge allRNA', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge allRNA', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:50:13.074082Z","iopub.execute_input":"2022-12-26T14:50:13.074555Z","iopub.status.idle":"2022-12-26T14:58:17.017493Z","shell.execute_reply.started":"2022-12-26T14:50:13.074509Z","shell.execute_reply":"2022-12-26T14:58:17.015729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:58:17.019761Z","iopub.execute_input":"2022-12-26T14:58:17.021147Z","iopub.status.idle":"2022-12-26T14:58:17.047694Z","shell.execute_reply.started":"2022-12-26T14:58:17.021075Z","shell.execute_reply":"2022-12-26T14:58:17.046492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction based on both rna and proteins (except target one ) \n","metadata":{"execution":{"iopub.status.busy":"2022-12-23T18:04:12.858384Z","iopub.execute_input":"2022-12-23T18:04:12.859084Z","iopub.status.idle":"2022-12-23T18:04:12.865728Z","shell.execute_reply.started":"2022-12-23T18:04:12.859047Z","shell.execute_reply":"2022-12-23T18:04:12.864980Z"}}},{"cell_type":"code","source":"X  = pd.concat( [ df_rna, df_y], axis = 1 )# .values\nX = X.drop(target_name , axis = 1)\nX.shape","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:58:17.049383Z","iopub.execute_input":"2022-12-26T14:58:17.050615Z","iopub.status.idle":"2022-12-26T14:58:29.750944Z","shell.execute_reply.started":"2022-12-26T14:58:17.050565Z","shell.execute_reply":"2022-12-26T14:58:29.749749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del df_rna\nimport gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:58:29.752405Z","iopub.execute_input":"2022-12-26T14:58:29.752754Z","iopub.status.idle":"2022-12-26T14:58:29.893278Z","shell.execute_reply.started":"2022-12-26T14:58:29.752722Z","shell.execute_reply":"2022-12-26T14:58:29.892013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nalpha_selected = 1e4\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\ny_pred = cross_val_predict(model, X, y, cv=kf)\n\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge allRNA+Proteins', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge allRNA+Proteins', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge allRNA+Proteins', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:58:29.895224Z","iopub.execute_input":"2022-12-26T14:58:29.895644Z","iopub.status.idle":"2022-12-26T15:04:49.347745Z","shell.execute_reply.started":"2022-12-26T14:58:29.895608Z","shell.execute_reply":"2022-12-26T15:04:49.345795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:04:49.350292Z","iopub.execute_input":"2022-12-26T15:04:49.351745Z","iopub.status.idle":"2022-12-26T15:04:49.378749Z","shell.execute_reply.started":"2022-12-26T15:04:49.351674Z","shell.execute_reply":"2022-12-26T15:04:49.376831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_y.corr()[target_name].sort_values(ascending= False)","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:04:49.381688Z","iopub.execute_input":"2022-12-26T15:04:49.382670Z","iopub.status.idle":"2022-12-26T15:04:53.386882Z","shell.execute_reply.started":"2022-12-26T15:04:49.382614Z","shell.execute_reply":"2022-12-26T15:04:53.385603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction based on RNA (except own RNA) and Proteins (except own proteins) ","metadata":{}},{"cell_type":"code","source":"print(X.shape)\nX.drop(rna_name , axis = 1, inplace = True)\nX.shape","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:04:53.389084Z","iopub.execute_input":"2022-12-26T15:04:53.389995Z","iopub.status.idle":"2022-12-26T15:04:55.954120Z","shell.execute_reply.started":"2022-12-26T15:04:53.389942Z","shell.execute_reply":"2022-12-26T15:04:55.952523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:04:55.955961Z","iopub.execute_input":"2022-12-26T15:04:55.956309Z","iopub.status.idle":"2022-12-26T15:04:56.144118Z","shell.execute_reply.started":"2022-12-26T15:04:55.956277Z","shell.execute_reply":"2022-12-26T15:04:56.142499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nalpha_selected = 1e4\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\ny_pred = cross_val_predict(model, X, y, cv=kf)\n\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge Proteins + RNA except own', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge Proteins + RNA except own', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge Proteins + RNA except own', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:04:56.146275Z","iopub.execute_input":"2022-12-26T15:04:56.147187Z","iopub.status.idle":"2022-12-26T15:11:10.978844Z","shell.execute_reply.started":"2022-12-26T15:04:56.147137Z","shell.execute_reply":"2022-12-26T15:11:10.977719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:11:10.984202Z","iopub.execute_input":"2022-12-26T15:11:10.986991Z","iopub.status.idle":"2022-12-26T15:11:11.008393Z","shell.execute_reply.started":"2022-12-26T15:11:10.986936Z","shell.execute_reply":"2022-12-26T15:11:11.006978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction based on RNA (all except own RNA) ","metadata":{}},{"cell_type":"code","source":"l = [t for t in X.columns if 'ENSG' in t ]\nprint( len(l), X.shape )\nX = X[l]\nX.shape","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:11:11.020642Z","iopub.execute_input":"2022-12-26T15:11:11.023847Z","iopub.status.idle":"2022-12-26T15:11:13.579504Z","shell.execute_reply.started":"2022-12-26T15:11:11.023777Z","shell.execute_reply":"2022-12-26T15:11:13.578436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:11:13.581155Z","iopub.execute_input":"2022-12-26T15:11:13.581585Z","iopub.status.idle":"2022-12-26T15:11:13.728186Z","shell.execute_reply.started":"2022-12-26T15:11:13.581543Z","shell.execute_reply":"2022-12-26T15:11:13.726968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nalpha_selected = 1e4\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\ny_pred = cross_val_predict(model, X, y, cv=kf)\n\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge all RNA except own', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge all RNA except own', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge all RNA except own', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:11:13.729584Z","iopub.execute_input":"2022-12-26T15:11:13.729929Z","iopub.status.idle":"2022-12-26T15:17:29.319374Z","shell.execute_reply.started":"2022-12-26T15:11:13.729880Z","shell.execute_reply":"2022-12-26T15:17:29.317478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:17:29.323252Z","iopub.execute_input":"2022-12-26T15:17:29.325300Z","iopub.status.idle":"2022-12-26T15:17:29.348750Z","shell.execute_reply.started":"2022-12-26T15:17:29.325220Z","shell.execute_reply":"2022-12-26T15:17:29.347213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Show slices by corr, r2, mse, ","metadata":{}},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n\nl = []\nfor IX in df_scores.index:\n    if ('r2' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n\nl = []\nfor IX in df_scores.index:\n    if ('RMSE' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:17:29.351277Z","iopub.execute_input":"2022-12-26T15:17:29.352460Z","iopub.status.idle":"2022-12-26T15:17:29.392616Z","shell.execute_reply.started":"2022-12-26T15:17:29.352389Z","shell.execute_reply":"2022-12-26T15:17:29.391669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save resuls","metadata":{"execution":{"iopub.status.busy":"2022-12-23T17:06:55.028528Z","iopub.execute_input":"2022-12-23T17:06:55.028970Z","iopub.status.idle":"2022-12-23T17:06:55.034005Z","shell.execute_reply.started":"2022-12-23T17:06:55.028933Z","shell.execute_reply":"2022-12-23T17:06:55.032823Z"}}},{"cell_type":"code","source":"df_scores.to_csv('Scores_'+target_name+'.csv')","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:17:29.394015Z","iopub.execute_input":"2022-12-26T15:17:29.394360Z","iopub.status.idle":"2022-12-26T15:17:29.403544Z","shell.execute_reply.started":"2022-12-26T15:17:29.394327Z","shell.execute_reply":"2022-12-26T15:17:29.402154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"t = time.time()-t0start\nprint('%.1f hours  = %.1f minutes = %.1f seconds passed total '%( t/3600,t/60, t ) )","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:17:29.404849Z","iopub.execute_input":"2022-12-26T15:17:29.406063Z","iopub.status.idle":"2022-12-26T15:17:29.411866Z","shell.execute_reply.started":"2022-12-26T15:17:29.406015Z","shell.execute_reply":"2022-12-26T15:17:29.410998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}