{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nFor selected protein e.g. CD36 we compare several simple Ridge based for predicting it:\n\nUse only one feature - own RNA for protein e.g. for CD36 protein  we use CD36 rna - and predict by Ridge only based on that model\n\nUse cell type (one-hot encoded for prediction) \n\nUse cell type + own RNA\n\nUse top 100  PCA feautures as predictors \n\nAnd also (in version 2 of the notebook): All RNA, ALL RNA + Proteins, ALL Proteins + RNA (except own), All RNA (except own RNA). It takes longer time - about 6 minutes for each. \n\n\nConclusions/Results (Pearson correlation is mentioned): \n\n    1) Cell type only is much worse than own RNA - 0.55 cs 0.73\n    2) All RNA except own - 0.72 \n    3) All proteins a bit better than own RNA: 0.73 vs 0.74\n    4) All RNA even more better 0.758\n    5) PCA100 is almost similar 0.761\n    6) Proteins + RNA (except own) a very bit better - 0.765 \n    7) allRNA + Proteins - quite better results - 0.797 \n\nThe same in the table: \n\n        CD36\n    Corr Ridge cell types\t0.553421 \n    Corr Ridge all RNA except own\t0.720771    \n    Corr Ridge ownRNA\t0.738055\n    Corr ownRNA\t0.738335\n    Corr Ridge CT + ownRNA\t0.741463\n    Corr Ridge allProteins\t0.742843\n    Corr Ridge allRNA\t0.757663\n    Corr Ridge pca100\t0.761470\n    Corr Ridge Proteins + RNA except own\t0.765259\n    Corr Ridge allRNA+Proteins\t0.797522\n\n\n\n#### Versions\n\n##### 2 - added Ridge models on ALL rna features - take about 6 minutes each \n\n    Overall 22 minutes \n\n##### 1 - most simple and fast Ridge models\n    \n    Including PCA100\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# Key Params","metadata":{"execution":{"iopub.status.busy":"2022-12-18T21:18:08.728359Z","iopub.execute_input":"2022-12-18T21:18:08.729709Z","iopub.status.idle":"2022-12-18T21:18:08.735176Z","shell.execute_reply.started":"2022-12-18T21:18:08.729662Z","shell.execute_reply":"2022-12-18T21:18:08.733772Z"}}},{"cell_type":"code","source":"target_name = 'CD36'\nn_splits_for_cross_valdition = 2 # if bigger, then for models on all rna (22K features) it may cause RAM crash and would be slower - \n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-12-23T19:39:53.531124Z","iopub.execute_input":"2022-12-23T19:39:53.53167Z","iopub.status.idle":"2022-12-23T19:39:53.557441Z","shell.execute_reply.started":"2022-12-23T19:39:53.531558Z","shell.execute_reply":"2022-12-23T19:39:53.556062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preparations","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\n# Main results will be stored here: \ndf_scores = pd.DataFrame()\nprint(type(df_scores))","metadata":{"execution":{"iopub.status.busy":"2022-12-24T18:04:59.874760Z","iopub.execute_input":"2022-12-24T18:04:59.875239Z","iopub.status.idle":"2022-12-24T18:04:59.907346Z","shell.execute_reply.started":"2022-12-24T18:04:59.875139Z","shell.execute_reply":"2022-12-24T18:04:59.906498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.preprocessing import OneHotEncoder\n\nimport time\nt0start = time.time()\n\nimport os","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:39:53.575392Z","iopub.execute_input":"2022-12-23T19:39:53.576379Z","iopub.status.idle":"2022-12-23T19:39:54.269615Z","shell.execute_reply.started":"2022-12-23T19:39:53.576258Z","shell.execute_reply":"2022-12-23T19:39:54.268312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data for target","metadata":{}},{"cell_type":"code","source":"%%time\ndf_y = pd.read_hdf('/kaggle/input/open-problems-multimodal/train_cite_targets.h5')\ndf_y","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:39:54.2717Z","iopub.execute_input":"2022-12-23T19:39:54.272848Z","iopub.status.idle":"2022-12-23T19:39:55.174174Z","shell.execute_reply.started":"2022-12-23T19:39:54.272801Z","shell.execute_reply":"2022-12-23T19:39:55.172966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = df_y[target_name]\ny","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:39:55.175429Z","iopub.execute_input":"2022-12-23T19:39:55.175754Z","iopub.status.idle":"2022-12-23T19:39:55.203563Z","shell.execute_reply.started":"2022-12-23T19:39:55.175723Z","shell.execute_reply":"2022-12-23T19:39:55.201968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_scores.loc['Mean', target_name ]  = np.mean(y)\ndf_scores.loc['Std', target_name ]  = np.std(y)\ndf_scores\n","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:39:55.206596Z","iopub.execute_input":"2022-12-23T19:39:55.20697Z","iopub.status.idle":"2022-12-23T19:39:55.227448Z","shell.execute_reply.started":"2022-12-23T19:39:55.206937Z","shell.execute_reply":"2022-12-23T19:39:55.226046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#  Predicting based on -  own RNA \n\ne.g. for CD36 protein we take CD36 as the only predictor ","metadata":{}},{"cell_type":"code","source":"%%time\ndf_rna = pd.read_hdf('/kaggle/input/open-problems-multimodal/train_cite_inputs.h5')\ndisplay(df_rna) \n\nfor t in df_rna.columns:\n    if target_name in t: \n        print(t)\n        rna_name = t # 'ENSG00000135218_CD36'\n","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:39:55.228962Z","iopub.execute_input":"2022-12-23T19:39:55.229462Z","iopub.status.idle":"2022-12-23T19:40:50.49947Z","shell.execute_reply.started":"2022-12-23T19:39:55.229427Z","shell.execute_reply":"2022-12-23T19:40:50.49807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = df_rna[[rna_name]]\nX","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:40:50.500898Z","iopub.execute_input":"2022-12-23T19:40:50.501276Z","iopub.status.idle":"2022-12-23T19:40:50.517016Z","shell.execute_reply.started":"2022-12-23T19:40:50.501243Z","shell.execute_reply":"2022-12-23T19:40:50.515442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"s = np.corrcoef(y,X.iloc[:,0])[0,1]\ndf_scores.loc['Corr ownRNA', target_name ]  = s\ndf_scores\n\n","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:40:50.518749Z","iopub.execute_input":"2022-12-23T19:40:50.519172Z","iopub.status.idle":"2022-12-23T19:40:50.53931Z","shell.execute_reply.started":"2022-12-23T19:40:50.519138Z","shell.execute_reply":"2022-12-23T19:40:50.537862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from scipy import stats\ns = stats.spearmanr(y,X.iloc[:,0])[0] # Take only correlation , skip p-value\ndf_scores.loc['Corr Spearman ownRNA', target_name ]  = s\ndf_scores\n","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:40:50.540638Z","iopub.execute_input":"2022-12-23T19:40:50.54099Z","iopub.status.idle":"2022-12-23T19:40:50.571117Z","shell.execute_reply.started":"2022-12-23T19:40:50.540946Z","shell.execute_reply":"2022-12-23T19:40:50.570131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nprint(X.shape, y.shape)\nmodel = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\nalpha_selected = model.alpha_\nprint('Optimal alpha found by cross-validation: ', alpha_selected )\n\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= 0 )\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge ownRNA', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge ownRNA', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge ownRNA', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:40:50.572405Z","iopub.execute_input":"2022-12-23T19:40:50.573453Z","iopub.status.idle":"2022-12-23T19:40:50.719048Z","shell.execute_reply.started":"2022-12-23T19:40:50.573403Z","shell.execute_reply":"2022-12-23T19:40:50.717641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#  Predicting based on - cell type information ","metadata":{}},{"cell_type":"code","source":"%%time\ndf_meta = pd.read_csv('/kaggle/input/feature-shop-for-multimodal-singlecell-competition/_citeseq_meta_all_text_also.csv', index_col = 0 )\ndf_meta = df_meta[df_meta['Train0OrTest1']==0]\ndf_meta","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:40:50.72504Z","iopub.execute_input":"2022-12-23T19:40:50.725876Z","iopub.status.idle":"2022-12-23T19:40:51.02799Z","shell.execute_reply.started":"2022-12-23T19:40:50.725826Z","shell.execute_reply":"2022-12-23T19:40:51.027072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = df_meta[['HSC cell_type', 'EryP cell_type', 'NeuP cell_type','MasP cell_type', 'MkP cell_type', 'BP cell_type', 'MoP cell_type' ]]\nX ","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:40:51.02932Z","iopub.execute_input":"2022-12-23T19:40:51.029858Z","iopub.status.idle":"2022-12-23T19:40:51.046636Z","shell.execute_reply.started":"2022-12-23T19:40:51.029825Z","shell.execute_reply":"2022-12-23T19:40:51.045307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nprint(X.shape, y.shape)\nmodel = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\nalpha_selected = model.alpha_\nprint('Optimal alpha found by cross-validation: ', alpha_selected )\n\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= 0 )\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge cell types', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge cell types', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge cell types', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:40:51.049116Z","iopub.execute_input":"2022-12-23T19:40:51.049524Z","iopub.status.idle":"2022-12-23T19:40:51.192052Z","shell.execute_reply.started":"2022-12-23T19:40:51.049476Z","shell.execute_reply":"2022-12-23T19:40:51.190604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predicting based on - cell type +  own RNA \n\n(own RNA -   RNA corresponding to the protein, e.g. for CD36 we take CD36  )","metadata":{}},{"cell_type":"code","source":"df_tmp = df_meta[['HSC cell_type', 'EryP cell_type', 'NeuP cell_type','MasP cell_type', 'MkP cell_type', 'BP cell_type', 'MoP cell_type' ]]\nX = pd.concat([ df_rna[[rna_name]],  df_tmp ], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:40:51.193877Z","iopub.execute_input":"2022-12-23T19:40:51.194651Z","iopub.status.idle":"2022-12-23T19:40:51.240628Z","shell.execute_reply.started":"2022-12-23T19:40:51.194582Z","shell.execute_reply":"2022-12-23T19:40:51.23922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.metrics import mean_squared_error\n\nprint(X.shape, y.shape)\nmodel = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\nalpha_selected = model.alpha_\nprint('Optimal alpha found by cross-validation: ', alpha_selected )\n\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= 0 )\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge CT + ownRNA', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge CT + ownRNA', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge CT + ownRNA', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:40:51.242365Z","iopub.execute_input":"2022-12-23T19:40:51.243531Z","iopub.status.idle":"2022-12-23T19:40:51.399137Z","shell.execute_reply.started":"2022-12-23T19:40:51.243478Z","shell.execute_reply":"2022-12-23T19:40:51.397804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predicting based on PCA top 100 features","metadata":{}},{"cell_type":"code","source":"%%time\ndf_pca = pd.read_csv('/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA200.csv', index_col = 0 )\ndf_pca","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:40:51.401168Z","iopub.execute_input":"2022-12-23T19:40:51.402001Z","iopub.status.idle":"2022-12-23T19:40:59.175434Z","shell.execute_reply.started":"2022-12-23T19:40:51.401938Z","shell.execute_reply":"2022-12-23T19:40:59.174181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = df_pca.iloc[:len(y), :100]\nX","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:40:59.176782Z","iopub.execute_input":"2022-12-23T19:40:59.177157Z","iopub.status.idle":"2022-12-23T19:40:59.235185Z","shell.execute_reply.started":"2022-12-23T19:40:59.177125Z","shell.execute_reply":"2022-12-23T19:40:59.233669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.metrics import mean_squared_error\n\nprint(X.shape, y.shape)\nmodel = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\nalpha_selected = model.alpha_\nprint('Optimal alpha found by cross-validation: ', alpha_selected )\n\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= 0 )\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge pca100', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge pca100', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge pca100', target_name ]  = s\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:40:59.23817Z","iopub.execute_input":"2022-12-23T19:40:59.238543Z","iopub.status.idle":"2022-12-23T19:41:00.715179Z","shell.execute_reply.started":"2022-12-23T19:40:59.238512Z","shell.execute_reply":"2022-12-23T19:41:00.71386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:41:00.721237Z","iopub.execute_input":"2022-12-23T19:41:00.722033Z","iopub.status.idle":"2022-12-23T19:41:00.879298Z","shell.execute_reply.started":"2022-12-23T19:41:00.721963Z","shell.execute_reply":"2022-12-23T19:41:00.878067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction based on proteins - all except the target one ","metadata":{}},{"cell_type":"code","source":"X = df_y.drop(target_name , axis = 1)\nX.shape","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:41:00.880886Z","iopub.execute_input":"2022-12-23T19:41:00.881313Z","iopub.status.idle":"2022-12-23T19:41:00.906552Z","shell.execute_reply.started":"2022-12-23T19:41:00.881276Z","shell.execute_reply":"2022-12-23T19:41:00.905149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nalpha_selected = 1e4\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= 0 )\ny_pred = cross_val_predict(model, X, y, cv=kf)\n\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge allProteins', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge allProteins', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge allProteins', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:41:00.907995Z","iopub.execute_input":"2022-12-23T19:41:00.908407Z","iopub.status.idle":"2022-12-23T19:41:01.209367Z","shell.execute_reply.started":"2022-12-23T19:41:00.908367Z","shell.execute_reply":"2022-12-23T19:41:01.206372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:41:01.211274Z","iopub.execute_input":"2022-12-23T19:41:01.214581Z","iopub.status.idle":"2022-12-23T19:41:01.239449Z","shell.execute_reply.started":"2022-12-23T19:41:01.214525Z","shell.execute_reply":"2022-12-23T19:41:01.237774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:41:01.241945Z","iopub.execute_input":"2022-12-23T19:41:01.242543Z","iopub.status.idle":"2022-12-23T19:41:01.389449Z","shell.execute_reply.started":"2022-12-23T19:41:01.242494Z","shell.execute_reply":"2022-12-23T19:41:01.388267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ridge on ALL RNA features ( 11 minutes of calculations ) ","metadata":{}},{"cell_type":"code","source":"X = df_rna.values\nX","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:41:01.390662Z","iopub.execute_input":"2022-12-23T19:41:01.391835Z","iopub.status.idle":"2022-12-23T19:41:01.404229Z","shell.execute_reply.started":"2022-12-23T19:41:01.391795Z","shell.execute_reply":"2022-12-23T19:41:01.40288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nalpha_selected = 1e4\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= 0 )\ny_pred = cross_val_predict(model, X, y, cv=kf)\n\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge allRNA', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge allRNA', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge allRNA', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:41:01.406254Z","iopub.execute_input":"2022-12-23T19:41:01.407298Z","iopub.status.idle":"2022-12-23T19:46:26.903193Z","shell.execute_reply.started":"2022-12-23T19:41:01.40722Z","shell.execute_reply":"2022-12-23T19:46:26.901845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:46:26.905284Z","iopub.execute_input":"2022-12-23T19:46:26.906777Z","iopub.status.idle":"2022-12-23T19:46:26.926786Z","shell.execute_reply.started":"2022-12-23T19:46:26.906723Z","shell.execute_reply":"2022-12-23T19:46:26.925211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction based on both rna and proteins (except target one ) \n","metadata":{"execution":{"iopub.status.busy":"2022-12-23T18:04:12.858384Z","iopub.execute_input":"2022-12-23T18:04:12.859084Z","iopub.status.idle":"2022-12-23T18:04:12.865728Z","shell.execute_reply.started":"2022-12-23T18:04:12.859047Z","shell.execute_reply":"2022-12-23T18:04:12.86498Z"}}},{"cell_type":"code","source":"X  = pd.concat( [ df_rna, df_y], axis = 1 )# .values\nX = X.drop(target_name , axis = 1)\nX.shape","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:46:26.929282Z","iopub.execute_input":"2022-12-23T19:46:26.930655Z","iopub.status.idle":"2022-12-23T19:46:37.862726Z","shell.execute_reply.started":"2022-12-23T19:46:26.930598Z","shell.execute_reply":"2022-12-23T19:46:37.861522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del df_rna\nimport gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:46:37.864159Z","iopub.execute_input":"2022-12-23T19:46:37.865648Z","iopub.status.idle":"2022-12-23T19:46:37.988594Z","shell.execute_reply.started":"2022-12-23T19:46:37.865603Z","shell.execute_reply":"2022-12-23T19:46:37.987305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nalpha_selected = 1e4\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= 0 )\ny_pred = cross_val_predict(model, X, y, cv=kf)\n\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge allRNA+Proteins', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge allRNA+Proteins', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge allRNA+Proteins', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:46:37.994687Z","iopub.execute_input":"2022-12-23T19:46:37.995111Z","iopub.status.idle":"2022-12-23T19:51:58.187385Z","shell.execute_reply.started":"2022-12-23T19:46:37.995076Z","shell.execute_reply":"2022-12-23T19:51:58.185924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:51:58.189792Z","iopub.execute_input":"2022-12-23T19:51:58.190852Z","iopub.status.idle":"2022-12-23T19:51:58.209509Z","shell.execute_reply.started":"2022-12-23T19:51:58.190794Z","shell.execute_reply":"2022-12-23T19:51:58.208059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_y.corr()[target_name].sort_values(ascending= False)","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:51:58.211442Z","iopub.execute_input":"2022-12-23T19:51:58.212072Z","iopub.status.idle":"2022-12-23T19:52:02.072688Z","shell.execute_reply.started":"2022-12-23T19:51:58.21202Z","shell.execute_reply":"2022-12-23T19:52:02.071429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction based on RNA (except own RNA) and Proteins (except own proteins) ","metadata":{}},{"cell_type":"code","source":"print(X.shape)\nX.drop(rna_name , axis = 1, inplace = True)\nX.shape","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:52:02.074481Z","iopub.execute_input":"2022-12-23T19:52:02.074842Z","iopub.status.idle":"2022-12-23T19:52:03.997598Z","shell.execute_reply.started":"2022-12-23T19:52:02.07481Z","shell.execute_reply":"2022-12-23T19:52:03.996432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:52:03.999062Z","iopub.execute_input":"2022-12-23T19:52:03.999929Z","iopub.status.idle":"2022-12-23T19:52:04.12042Z","shell.execute_reply.started":"2022-12-23T19:52:03.999891Z","shell.execute_reply":"2022-12-23T19:52:04.119041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nalpha_selected = 1e4\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= 0 )\ny_pred = cross_val_predict(model, X, y, cv=kf)\n\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge Proteins + RNA except own', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge Proteins + RNA except own', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge Proteins + RNA except own', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:52:04.121924Z","iopub.execute_input":"2022-12-23T19:52:04.122582Z","iopub.status.idle":"2022-12-23T19:57:24.4271Z","shell.execute_reply.started":"2022-12-23T19:52:04.122542Z","shell.execute_reply":"2022-12-23T19:57:24.425046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:57:24.430131Z","iopub.execute_input":"2022-12-23T19:57:24.432233Z","iopub.status.idle":"2022-12-23T19:57:24.454516Z","shell.execute_reply.started":"2022-12-23T19:57:24.432172Z","shell.execute_reply":"2022-12-23T19:57:24.452932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction based on RNA (all except own RNA) ","metadata":{}},{"cell_type":"code","source":"l = [t for t in X.columns if 'ENSG' in t ]\nprint( len(l), X.shape )\nX = X[l]\nX.shape","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:57:24.456362Z","iopub.execute_input":"2022-12-23T19:57:24.457261Z","iopub.status.idle":"2022-12-23T19:57:26.416339Z","shell.execute_reply.started":"2022-12-23T19:57:24.457209Z","shell.execute_reply":"2022-12-23T19:57:26.414861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:57:26.418156Z","iopub.execute_input":"2022-12-23T19:57:26.418642Z","iopub.status.idle":"2022-12-23T19:57:26.591915Z","shell.execute_reply.started":"2022-12-23T19:57:26.418594Z","shell.execute_reply":"2022-12-23T19:57:26.59023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\nalpha_selected = 1e4\n\nmodel = Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= 0 )\ny_pred = cross_val_predict(model, X, y, cv=kf)\n\ns = r2_score(y,y_pred)\ndf_scores.loc['r2 Ridge all RNA except own', target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr Ridge all RNA except own', target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc['RMSE Ridge all RNA except own', target_name ]  = s\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-23T19:57:26.594419Z","iopub.execute_input":"2022-12-23T19:57:26.594999Z","iopub.status.idle":"2022-12-23T20:02:41.773446Z","shell.execute_reply.started":"2022-12-23T19:57:26.594928Z","shell.execute_reply":"2022-12-23T20:02:41.771709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )","metadata":{"execution":{"iopub.status.busy":"2022-12-23T20:02:41.775644Z","iopub.execute_input":"2022-12-23T20:02:41.776525Z","iopub.status.idle":"2022-12-23T20:02:41.795637Z","shell.execute_reply.started":"2022-12-23T20:02:41.776472Z","shell.execute_reply":"2022-12-23T20:02:41.79428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Show slices by corr, r2, mse, ","metadata":{}},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n\nl = []\nfor IX in df_scores.index:\n    if ('r2' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n\nl = []\nfor IX in df_scores.index:\n    if ('RMSE' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n","metadata":{"execution":{"iopub.status.busy":"2022-12-23T20:02:41.797474Z","iopub.execute_input":"2022-12-23T20:02:41.798381Z","iopub.status.idle":"2022-12-23T20:02:41.841217Z","shell.execute_reply.started":"2022-12-23T20:02:41.79833Z","shell.execute_reply":"2022-12-23T20:02:41.839799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save resuls","metadata":{"execution":{"iopub.status.busy":"2022-12-23T17:06:55.028528Z","iopub.execute_input":"2022-12-23T17:06:55.02897Z","iopub.status.idle":"2022-12-23T17:06:55.034005Z","shell.execute_reply.started":"2022-12-23T17:06:55.028933Z","shell.execute_reply":"2022-12-23T17:06:55.032823Z"}}},{"cell_type":"code","source":"df_scores.to_csv('Scores_'+target_name+'.csv')","metadata":{"execution":{"iopub.status.busy":"2022-12-23T20:02:41.842765Z","iopub.execute_input":"2022-12-23T20:02:41.843404Z","iopub.status.idle":"2022-12-23T20:02:41.853024Z","shell.execute_reply.started":"2022-12-23T20:02:41.843364Z","shell.execute_reply":"2022-12-23T20:02:41.851739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"t = time.time()-t0start\nprint('%.1f hours  = %.1f minutes = %.1f seconds passed total '%( t/3600,t/60, t ) )","metadata":{"execution":{"iopub.status.busy":"2022-12-23T20:02:41.854788Z","iopub.execute_input":"2022-12-23T20:02:41.856081Z","iopub.status.idle":"2022-12-23T20:02:41.863063Z","shell.execute_reply.started":"2022-12-23T20:02:41.85603Z","shell.execute_reply":"2022-12-23T20:02:41.861687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}