{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nFor selected protein e.g. CD36 we compare several models.\nIn particular LightGBM , Ridge, etc. Compare on several sets of features - pca100, all rna, etc...\n\n\nConclusion: We see quite a singificant uplift for LGB models, comparing to Ridge.\nNote: here we use just random split validation. While when we used \nvalidation  on new days - the situation was different.\nSuprisingly - Ridge was better than boosting models - see script:\nhttps://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold\n\nRemark: More simple models considered in the previous notebook:\nhttps://www.kaggle.com/code/alexandervc/simple-models-for-cds\n\n\n#### Versions\n\n##### 5,6 - added LGB on all RNA + proteins \n\n    LGB Default Corr RNA + Proteins: 0.886272\n\n##### 1,2,3,4 - LGB default, RIDGE on PCA100 and all orginal for CD36\n\n    Correlation with own RNA 0.738\n    Correlation with LGB on just own RNA - 0.765 \n    LGB on pca100 gives quite better result than Ridge: 0.795 vs 0.761 \n    Using out favorite params LGB we get even more: 0.802704 (a bit better, not crucial)\n    On all features Ridge uplifts to 0.777, LGB to 0.84+ \n    \n    V5: on RNA+Proteins LGB: 0.886 ( Ridge: 0.797 - previous Notebook)\n\n    So: gain of LGB grows with number of the features involved - strarting from 0.03 for 1 feature and pca100, to 0.07 - all rna, 0.08 - all rna+proteins. \n\n","metadata":{}},{"cell_type":"markdown","source":"## From the previous notebook:\n\nhttps://www.kaggle.com/code/alexandervc/simple-models-for-cds\n\nFor CD36. \nConclusions/Results (Pearson correlation is mentioned): \n\n    1) Cell type only is much worse than own RNA - 0.55 cs 0.73\n    2) All RNA except own - 0.72 \n    3) All proteins a bit better than own RNA: 0.73 vs 0.74\n    4) All RNA even more better 0.758\n    5) PCA100 is almost similar 0.761\n    6) LightGBM on just one feature - own RNA - 0.765\n    7) Proteins + RNA (except own) a very bit better - 0.765 \n    8) allRNA + Proteins - quite better results - 0.797 \n\nThe same in the table: \n\n        CD36\n    Corr Ridge cell types\t0.553421 \n    Corr Ridge all RNA except own\t0.720771    \n    Corr Ridge ownRNA\t0.738055\n    Corr ownRNA\t0.738335\n    Corr Ridge CT + ownRNA\t0.741463\n    Corr Ridge allProteins\t0.742843\n    Corr Ridge allRNA\t0.757663\n    Corr Ridge pca100\t0.761470\n    Corr LightGBM on just one feature - own RNA - 0.765\n    Corr Ridge Proteins + RNA except own\t0.765259\n    Corr Ridge allRNA+Proteins\t0.797522\n","metadata":{}},{"cell_type":"markdown","source":"# Key Params","metadata":{"execution":{"iopub.status.busy":"2022-12-18T21:18:08.728359Z","iopub.execute_input":"2022-12-18T21:18:08.729709Z","iopub.status.idle":"2022-12-18T21:18:08.735176Z","shell.execute_reply.started":"2022-12-18T21:18:08.729662Z","shell.execute_reply":"2022-12-18T21:18:08.733772Z"}}},{"cell_type":"code","source":"target_name = 'CD36'\nn_splits_for_cross_valdition = 2 # if bigger, then for models on all rna (22K features) it may cause RAM crash and would be slower - \n\nrandom_state_cross_validation = 0 # random to be used in generating the folds - changing it we can study how stable are our results\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-12-26T14:57:56.634398Z","iopub.execute_input":"2022-12-26T14:57:56.635006Z","iopub.status.idle":"2022-12-26T14:57:56.665584Z","shell.execute_reply.started":"2022-12-26T14:57:56.634883Z","shell.execute_reply":"2022-12-26T14:57:56.664204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preparations","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\n# Main results will be stored here: \ndf_scores = pd.DataFrame()\ndf_predictions = pd.DataFrame()","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:57:56.667646Z","iopub.execute_input":"2022-12-26T14:57:56.668337Z","iopub.status.idle":"2022-12-26T14:57:56.684453Z","shell.execute_reply.started":"2022-12-26T14:57:56.668297Z","shell.execute_reply":"2022-12-26T14:57:56.683248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.preprocessing import OneHotEncoder\n\nimport time\nt0start = time.time()\n\nimport os","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:57:56.686051Z","iopub.execute_input":"2022-12-26T14:57:56.686696Z","iopub.status.idle":"2022-12-26T14:57:57.749864Z","shell.execute_reply.started":"2022-12-26T14:57:56.686651Z","shell.execute_reply":"2022-12-26T14:57:57.748354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data for target","metadata":{}},{"cell_type":"code","source":"%%time\ndf_y = pd.read_hdf('/kaggle/input/open-problems-multimodal/train_cite_targets.h5')\ndf_y","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:57:57.752603Z","iopub.execute_input":"2022-12-26T14:57:57.753016Z","iopub.status.idle":"2022-12-26T14:57:58.871459Z","shell.execute_reply.started":"2022-12-26T14:57:57.752981Z","shell.execute_reply":"2022-12-26T14:57:58.870234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = df_y[target_name]\ny","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:57:58.873151Z","iopub.execute_input":"2022-12-26T14:57:58.873536Z","iopub.status.idle":"2022-12-26T14:57:58.898587Z","shell.execute_reply.started":"2022-12-26T14:57:58.873500Z","shell.execute_reply":"2022-12-26T14:57:58.897038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_scores.loc['Mean', target_name ]  = np.mean(y)\ndf_scores.loc['Std', target_name ]  = np.std(y)\ndf_scores\n","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:57:58.900211Z","iopub.execute_input":"2022-12-26T14:57:58.900618Z","iopub.status.idle":"2022-12-26T14:57:58.924834Z","shell.execute_reply.started":"2022-12-26T14:57:58.900582Z","shell.execute_reply":"2022-12-26T14:57:58.923289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load features - rna expressions","metadata":{}},{"cell_type":"code","source":"%%time\ndf_rna = pd.read_hdf('/kaggle/input/open-problems-multimodal/train_cite_inputs.h5')\ndisplay(df_rna) ","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:57:58.926408Z","iopub.execute_input":"2022-12-26T14:57:58.926835Z","iopub.status.idle":"2022-12-26T14:59:00.135419Z","shell.execute_reply.started":"2022-12-26T14:57:58.926782Z","shell.execute_reply":"2022-12-26T14:59:00.134198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nfor t in df_rna.columns:\n    if target_name in t: \n        print(t)\n        rna_name = t # 'ENSG00000135218_CD36'\n        \ns = np.corrcoef(y,df_rna[rna_name])[0,1]\ndf_scores.loc['Corr ownRNA', target_name ]  = s\ndf_scores        ","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:59:00.136785Z","iopub.execute_input":"2022-12-26T14:59:00.137198Z","iopub.status.idle":"2022-12-26T14:59:00.162210Z","shell.execute_reply.started":"2022-12-26T14:59:00.137163Z","shell.execute_reply":"2022-12-26T14:59:00.160858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport lightgbm as lgbm\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold\nimport time\n\nmodel = lgbm.LGBMRegressor(random_state = 0) # Ridge(alpha = alpha_selected )\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\ny_pred = cross_val_predict(model, df_rna[ [rna_name] ], y, cv=kf)\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc['Corr LGB ownRNA', target_name ]  = s\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:59:00.166106Z","iopub.execute_input":"2022-12-26T14:59:00.166508Z","iopub.status.idle":"2022-12-26T14:59:01.702244Z","shell.execute_reply.started":"2022-12-26T14:59:00.166475Z","shell.execute_reply":"2022-12-26T14:59:01.700732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction based on PCA top 100 features","metadata":{}},{"cell_type":"code","source":"%%time\ndf_pca = pd.read_csv('/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA200.csv', index_col = 0 )\ndf_pca","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:59:01.704118Z","iopub.execute_input":"2022-12-26T14:59:01.704551Z","iopub.status.idle":"2022-12-26T14:59:11.898248Z","shell.execute_reply.started":"2022-12-26T14:59:01.704507Z","shell.execute_reply":"2022-12-26T14:59:11.893817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = df_pca.iloc[:len(y), :100]\nstr_data_inf = 'pca100'\nprint(str_data_inf, X.shape, y.shape)\nX","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:59:11.902558Z","iopub.execute_input":"2022-12-26T14:59:11.903639Z","iopub.status.idle":"2022-12-26T14:59:12.053070Z","shell.execute_reply.started":"2022-12-26T14:59:11.903595Z","shell.execute_reply":"2022-12-26T14:59:12.052065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Ridge","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.metrics import mean_squared_error\nimport lightgbm as lgbm\nimport time \n\nt0 = time.time()\n\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\n\nprint(X.shape, y.shape, str_data_inf)\nmodel = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\nalpha_selected = model.alpha_\nprint('Optimal alpha found by cross-validation: ', alpha_selected )\nmodel = Ridge(alpha = alpha_selected )\nif alpha_selected == int(alpha_selected): alpha_selected = int( alpha_selected ) # For better outout in str( ... )\nstr_model_inf = 'Ridge'+str(alpha_selected)\nprint('Time for RidgeCV:',  np.round( time.time() - t0, 2)  )\nt0 = time.time()\n\n# model = lgbm.LGBMRegressor(random_state = 0)\n# str_model_inf = 'LGB Default'\n# print('Model: ', str_model_inf )\n\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc[str_model_inf + ' r2 '+str_data_inf, target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc[str_model_inf + ' Corr '+str_data_inf, target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc[str_model_inf+ ' RMSE '+str_data_inf, target_name ]  = s\n\ndf_scores.loc[str_model_inf+ ' Time '+str_data_inf, target_name ]  = np.round( time.time() - t0, 2) \n\ndf_predictions[str_model_inf + ' '+str_data_inf ] = y_pred\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:59:12.054872Z","iopub.execute_input":"2022-12-26T14:59:12.056084Z","iopub.status.idle":"2022-12-26T14:59:14.847350Z","shell.execute_reply.started":"2022-12-26T14:59:12.056029Z","shell.execute_reply":"2022-12-26T14:59:14.846020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## LGB default params","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.metrics import mean_squared_error\nimport lightgbm as lgbm\nimport time \n\nt0 = time.time()\n\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\n\nprint(X.shape, y.shape, str_data_inf)\n#model = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\n#alpha_selected = model.alpha_\n#print('Optimal alpha found by cross-validation: ', alpha_selected )\n# model = Ridge(alpha = alpha_selected )\n\nmodel = lgbm.LGBMRegressor(random_state = 0)\nstr_model_inf = 'LGB Default'\nprint('Model: ', str_model_inf )\n\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc[str_model_inf + ' r2 '+str_data_inf, target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc[str_model_inf + ' Corr '+str_data_inf, target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc[str_model_inf+ ' RMSE '+str_data_inf, target_name ]  = s\n\ndf_scores.loc[str_model_inf+ ' Time '+str_data_inf, target_name ]  = np.round( time.time() - t0, 2) \n\ndf_predictions[str_model_inf + ' '+str_data_inf ] = y_pred\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:59:14.849396Z","iopub.execute_input":"2022-12-26T14:59:14.850867Z","iopub.status.idle":"2022-12-26T14:59:20.644407Z","shell.execute_reply.started":"2022-12-26T14:59:14.850795Z","shell.execute_reply":"2022-12-26T14:59:20.643048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## LGB with pre-selected params1","metadata":{}},{"cell_type":"code","source":"%%time\n\n# Use LGB params dicsussed here: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/367254\n\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.metrics import mean_squared_error\nimport lightgbm as lgbm\nimport time \n\nt0 = time.time()\n\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\n\nprint(X.shape, y.shape, str_data_inf)\n#model = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\n#alpha_selected = model.alpha_\n#print('Optimal alpha found by cross-validation: ', alpha_selected )\n# model = Ridge(alpha = alpha_selected )\n\nparams = {'random_state': 0, \n 'n_estimators': 500, 'reg_alpha': 6.734991732483901, 'reg_lambda': 3.0151837674804667, 'colsample_bytree': 0.9, \n 'subsample': 1.0, 'max_depth': 6, 'learning_rate': 0.03814860248977263, 'num_leaves': 859, 'min_child_samples': 13}\n\nmodel = lgbm.LGBMRegressor(**params)\nstr_model_inf = 'LGB Selected1'\nprint('Model: ', str_model_inf )\n\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc[str_model_inf + ' r2 '+str_data_inf, target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc[str_model_inf + ' Corr '+str_data_inf, target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc[str_model_inf+ ' RMSE '+str_data_inf, target_name ]  = s\n\ndf_scores.loc[str_model_inf+ ' Time '+str_data_inf, target_name ]  = np.round( time.time() - t0, 2) \n\ndf_predictions[str_model_inf + ' '+str_data_inf ] = y_pred\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:59:20.646079Z","iopub.execute_input":"2022-12-26T14:59:20.646572Z","iopub.status.idle":"2022-12-26T14:59:41.211281Z","shell.execute_reply.started":"2022-12-26T14:59:20.646529Z","shell.execute_reply":"2022-12-26T14:59:41.209834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:59:41.212842Z","iopub.execute_input":"2022-12-26T14:59:41.213321Z","iopub.status.idle":"2022-12-26T14:59:41.354298Z","shell.execute_reply.started":"2022-12-26T14:59:41.213277Z","shell.execute_reply":"2022-12-26T14:59:41.353152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\npd.set_option('display.max_rows', 500)\npd.set_option('display.max_columns', 500)\npd.set_option('display.width', 1000)","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:59:41.355892Z","iopub.execute_input":"2022-12-26T14:59:41.356308Z","iopub.status.idle":"2022-12-26T14:59:41.363999Z","shell.execute_reply.started":"2022-12-26T14:59:41.356273Z","shell.execute_reply":"2022-12-26T14:59:41.362775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predictions based on all rna ","metadata":{}},{"cell_type":"code","source":"X = df_rna.values\nstr_data_inf = 'RNAall'\nprint(str_data_inf, X.shape, y.shape)\n\nX","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:59:41.365429Z","iopub.execute_input":"2022-12-26T14:59:41.365906Z","iopub.status.idle":"2022-12-26T14:59:41.382964Z","shell.execute_reply.started":"2022-12-26T14:59:41.365872Z","shell.execute_reply":"2022-12-26T14:59:41.381711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Ridge model ","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.metrics import mean_squared_error\nimport lightgbm as lgbm\nimport time \n\nt0 = time.time()\n\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\n\nprint(X.shape, y.shape, str_data_inf)\n# model = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\n# alpha_selected = model.alpha_\n# print('Optimal alpha found by cross-validation: ', alpha_selected )\n# model = Ridge(alpha = alpha_selected )\n# if alpha_selected == int(alpha_selected): alpha_selected = int( alpha_selected ) # For better outout in str( ... )\n# str_model_inf = 'Ridge'+str(alpha_selected)\n# print('Time for RidgeCV:',  np.round( time.time() - t0, 2)  )\n# t0 = time.time()\n\nfor alpha_selected in  [10_000, 100_000, 1_000_000]:\n    model = Ridge(alpha = alpha_selected )\n    if alpha_selected == int(alpha_selected): alpha_selected = int( alpha_selected ) # For better outout in str( ... )\n    str_model_inf = 'Ridge'+str(alpha_selected)\n    t0 = time.time()\n\n    # model = lgbm.LGBMRegressor(random_state = 0)\n    # str_model_inf = 'LGB Default'\n    # print('Model: ', str_model_inf )\n\n    y_pred = cross_val_predict(model, X, y, cv=kf)\n    s = r2_score(y,y_pred)\n    df_scores.loc[str_model_inf + ' r2 '+str_data_inf, target_name ]  = s\n    s = np.corrcoef(y,y_pred)[0,1]\n    df_scores.loc[str_model_inf + ' Corr '+str_data_inf, target_name ]  = s\n    s = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \n    df_scores.loc[str_model_inf+ ' RMSE '+str_data_inf, target_name ]  = s\n\n    df_scores.loc[str_model_inf+ ' Time '+str_data_inf, target_name ]  = np.round( time.time() - t0, 2) \n    \n    df_predictions[str_model_inf + ' '+str_data_inf ] = y_pred\n\ndisplay( df_scores )","metadata":{"execution":{"iopub.status.busy":"2022-12-26T14:59:41.384942Z","iopub.execute_input":"2022-12-26T14:59:41.385311Z","iopub.status.idle":"2022-12-26T15:16:30.873383Z","shell.execute_reply.started":"2022-12-26T14:59:41.385281Z","shell.execute_reply":"2022-12-26T15:16:30.870941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## LGB Default","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.metrics import mean_squared_error\nimport lightgbm as lgbm\nimport time \n\nt0 = time.time()\n\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\n\nprint(X.shape, y.shape, str_data_inf)\n#model = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\n#alpha_selected = model.alpha_\n#print('Optimal alpha found by cross-validation: ', alpha_selected )\n# model = Ridge(alpha = alpha_selected )\n\nmodel = lgbm.LGBMRegressor(random_state = 0)\nstr_model_inf = 'LGB Default'\nprint('Model: ', str_model_inf )\n\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc[str_model_inf + ' r2 '+str_data_inf, target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc[str_model_inf + ' Corr '+str_data_inf, target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc[str_model_inf+ ' RMSE '+str_data_inf, target_name ]  = s\n\ndf_scores.loc[str_model_inf+ ' Time '+str_data_inf, target_name ]  = np.round( time.time() - t0, 2) \n\ndf_predictions[str_model_inf + ' '+str_data_inf ] = y_pred\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:16:30.876465Z","iopub.execute_input":"2022-12-26T15:16:30.879011Z","iopub.status.idle":"2022-12-26T15:28:50.491553Z","shell.execute_reply.started":"2022-12-26T15:16:30.878942Z","shell.execute_reply":"2022-12-26T15:28:50.490180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## LGB with pre-selected params1","metadata":{}},{"cell_type":"code","source":"%%time\n\n# Use LGB params dicsussed here: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/367254\n\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.metrics import mean_squared_error\nimport lightgbm as lgbm\nimport time \n\nt0 = time.time()\n\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\n\nprint(X.shape, y.shape, str_data_inf)\n#model = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\n#alpha_selected = model.alpha_\n#print('Optimal alpha found by cross-validation: ', alpha_selected )\n# model = Ridge(alpha = alpha_selected )\n\nparams = {'random_state': 0, \n 'n_estimators': 500, 'reg_alpha': 6.734991732483901, 'reg_lambda': 3.0151837674804667, 'colsample_bytree': 0.9, \n 'subsample': 1.0, 'max_depth': 6, 'learning_rate': 0.03814860248977263, 'num_leaves': 859, 'min_child_samples': 13}\n\nmodel = lgbm.LGBMRegressor(**params)\nstr_model_inf = 'LGB Selected1'\nprint('Model: ', str_model_inf )\n\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc[str_model_inf + ' r2 '+str_data_inf, target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc[str_model_inf + ' Corr '+str_data_inf, target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc[str_model_inf+ ' RMSE '+str_data_inf, target_name ]  = s\n\ndf_scores.loc[str_model_inf+ ' Time '+str_data_inf, target_name ]  = np.round( time.time() - t0, 2) \n\ndf_predictions[str_model_inf + ' '+str_data_inf ] = y_pred\n\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T15:28:50.493702Z","iopub.execute_input":"2022-12-26T15:28:50.494098Z","iopub.status.idle":"2022-12-26T16:11:26.185036Z","shell.execute_reply.started":"2022-12-26T15:28:50.494065Z","shell.execute_reply":"2022-12-26T16:11:26.183862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predictions based on all rna + Proteins ","metadata":{}},{"cell_type":"code","source":"X  = pd.concat( [ df_rna, df_y], axis = 1 )# .values\nX = X.drop(target_name , axis = 1)\nstr_data_inf = 'RNA + Proteins'\nprint(str_data_inf, X.shape, y.shape)\n\nX","metadata":{"execution":{"iopub.status.busy":"2022-12-26T16:14:01.144442Z","iopub.execute_input":"2022-12-26T16:14:01.145054Z","iopub.status.idle":"2022-12-26T16:14:28.676743Z","shell.execute_reply.started":"2022-12-26T16:14:01.145011Z","shell.execute_reply":"2022-12-26T16:14:28.675288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.metrics import mean_squared_error\nimport lightgbm as lgbm\nimport time \n\nt0 = time.time()\n\nkf = KFold(n_splits=n_splits_for_cross_valdition,  shuffle=True, random_state= random_state_cross_validation )\n\nprint(X.shape, y.shape, str_data_inf)\n#model = RidgeCV(alphas=[1e-1, 1, 1e1,1e2,1e3,1e4,1e5]).fit(X, y)\n#alpha_selected = model.alpha_\n#print('Optimal alpha found by cross-validation: ', alpha_selected )\n# model = Ridge(alpha = alpha_selected )\n\nmodel = lgbm.LGBMRegressor(random_state = 0)\nstr_model_inf = 'LGB Default'\nprint('Model: ', str_model_inf )\n\ny_pred = cross_val_predict(model, X, y, cv=kf)\ns = r2_score(y,y_pred)\ndf_scores.loc[str_model_inf + ' r2 '+str_data_inf, target_name ]  = s\ns = np.corrcoef(y,y_pred)[0,1]\ndf_scores.loc[str_model_inf + ' Corr '+str_data_inf, target_name ]  = s\ns = mean_squared_error(y,y_pred, squared=False) # squared=False -> RMSE, not MSE \ndf_scores.loc[str_model_inf+ ' RMSE '+str_data_inf, target_name ]  = s\n\ndf_scores.loc[str_model_inf+ ' Time '+str_data_inf, target_name ]  = np.round( time.time() - t0, 2) \n\ndf_predictions[str_model_inf + ' '+str_data_inf ] = y_pred\n\n\ndf_scores","metadata":{"execution":{"iopub.status.busy":"2022-12-26T16:28:18.087245Z","iopub.execute_input":"2022-12-26T16:28:18.087818Z","iopub.status.idle":"2022-12-26T16:41:31.335637Z","shell.execute_reply.started":"2022-12-26T16:28:18.087779Z","shell.execute_reply":"2022-12-26T16:41:31.333860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Show results by each metric ","metadata":{}},{"cell_type":"code","source":"l = []\nfor IX in df_scores.index:\n    if ('Corr' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n\nl = []\nfor IX in df_scores.index:\n    if ('r2' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )\n\nl = []\nfor IX in df_scores.index:\n    if ('RMSE' in IX) and ('Spearman' not in IX): l.append(IX)\ndisplay( df_scores.loc[l,:].sort_values(target_name) )","metadata":{"execution":{"iopub.status.busy":"2022-12-26T16:11:26.186763Z","iopub.execute_input":"2022-12-26T16:11:26.187470Z","iopub.status.idle":"2022-12-26T16:11:26.220621Z","shell.execute_reply.started":"2022-12-26T16:11:26.187434Z","shell.execute_reply":"2022-12-26T16:11:26.219152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save resuls","metadata":{"execution":{"iopub.status.busy":"2022-12-23T17:06:55.028528Z","iopub.execute_input":"2022-12-23T17:06:55.028970Z","iopub.status.idle":"2022-12-23T17:06:55.034005Z","shell.execute_reply.started":"2022-12-23T17:06:55.028933Z","shell.execute_reply":"2022-12-23T17:06:55.032823Z"}}},{"cell_type":"code","source":"df_scores.to_csv('Scores_WithBoostings_'+target_name+'.csv')","metadata":{"execution":{"iopub.status.busy":"2022-12-26T16:11:26.222461Z","iopub.execute_input":"2022-12-26T16:11:26.223043Z","iopub.status.idle":"2022-12-26T16:11:26.240488Z","shell.execute_reply.started":"2022-12-26T16:11:26.222999Z","shell.execute_reply":"2022-12-26T16:11:26.239124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_predictions.to_csv('Predictions_WithBoostings_'+target_name+'.csv')\ndf_predictions.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-12-26T16:11:26.242292Z","iopub.execute_input":"2022-12-26T16:11:26.243111Z","iopub.status.idle":"2022-12-26T16:11:27.050059Z","shell.execute_reply.started":"2022-12-26T16:11:26.243060Z","shell.execute_reply":"2022-12-26T16:11:27.048467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"t = time.time()-t0start\nprint('%.1f hours  = %.1f minutes = %.1f seconds passed total '%( t/3600,t/60, t ) )","metadata":{"execution":{"iopub.status.busy":"2022-12-26T16:11:27.058184Z","iopub.execute_input":"2022-12-26T16:11:27.058664Z","iopub.status.idle":"2022-12-26T16:11:27.066666Z","shell.execute_reply.started":"2022-12-26T16:11:27.058624Z","shell.execute_reply":"2022-12-26T16:11:27.064916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}