{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\n### Briefly - downsample and make quick experiments\n\nHere we will downsample the data. I.e. select 10% of data as a kind of \"Playground\" - for quick experiments. \nAnd will train/tune different models on that playground first.\n\n\n### About CV scheme\n\nPublic and private test sets are quite different. (Private - contains new DAY (and donor), while public ONLY new donor).\nSo one should be careful with the validation schemes.\nSome proposals are described here: \n\nNotebook: https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\nTopic: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\n\nWe will be based on them. \n\n### Start with selection of \"Playground\" - 10% part of data - to quickly test models, ideas\n\nData is quite big, and so training models, tuning params might take long time. That is not always affordable.\nIn the present script we first choose some 10% part of data - to make quick experiments.\nThat part is chosen to have SAME proporitions of key characteristics: donors, days, cell types as the initial data\n\nSo we can create CV scheme for that \"playground\" part.\nBut we can also simplify even further - for start - use  not 6-fold scheme but just splite by days. And have only 1 train subset for quick experiments. \n\nTo simplify even further we can start play with just one target only. \n\n\n\n### Versions\n\n#### 59 and later - testing Optuna parameters for 1 target with model trained on full targets\n\n#### 18,19 - Added: Mode to run only Ridge, Lgb, Catboost\n\n    19 - minor error corrected\n    \n    Run on 100 features, full train: Catboost  - 20 seconds, Lgb - 2 seconds (default params both) ,  Ridge - 0.05seconds\n    \n    Mode to run only Ridge, Lgb, Catboost\n    mode_what_slow_models_to_allow = 'LightGBM_Ridge_CatBoost'\n\n\n#### 17 - returned to small sizes of data: 50 features, 10% of samples \n\n#### 14,15,16 - run on 500 features and \"Full train\"\n\n    Here \"full train\"  means we use our valiation scheme - by day and donor\n    16 - successful - but about 8-10 hours, Just one kernel Ridge - 1 hours, tuning of Kernel Ridge - 5 hours, SVR was OFF\n    14 , 15 - crashed by time limit 12 hours. Some models are VERY slow on such data: KernelRidge, SVR, RandomForest.\n    \n#### 13 - switch to big size of playground - FULL train - it takes long time - 1.5 hour\n\n    Support for playground_fraction_of_train = 1 - everything (the whole train) is playground \n    \n    Here \"full train\"  means we use our valiation scheme - by day and donor\n    \n\n#### 12 - switch to big size of playground - half of full train - it takes long time - 1.5 hour\n\n    LightGBM - runs 0.6 seconds.\n    It still worse than Ridge, but difference is not big. \n    RF tuned - runs for 10 mintes  (500 interators)\n    Kernel Ridge - is very slow - 5 minutes - and again terrbible quality \n    XGBOost is 10 times slower with default params than Lgb, but quality is quite lower\n    Catboost default is second after the Ridge, but it takes 8 seconds, comparing to 0.6 lgb \n    \n#### 11 - cosmetic change\n    small error corrected - LGB Tuned2 params were not correct\n    Resulting table with statistcs - has been saved.  \n    \n#### 10 added parameter to control size of the play-ground\n\n    playground_fraction_of_train = 10# Should be Integer N - we create play-ground of the size Full_Train / N. \n\n#### 9 Blend and BIG SURPRISE(!) - how to explain ? ideas - welcome ! \n\n    added blending part for all solutions - and got suprise:\n    and the WORST/TERRIBLE (r2 negative, mse solution 10 times worse than others -   KernelRidge\n    enters the blends and improve it ! \n    \n    The only thing - that it is very uncorrelated with the other solutions. \n\nSee discussion: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/363230\n\n\n#### 8: Added - statistics on all methods - collected in one table (dataframe)\n    \n    Still Ridge, SVR are the best. Cat\n\n#### 7: added XGBoost+Optuna, CatBoost, MLP\n    \n    CatBoost is quite good with default params\n\n#### 3,4,5,6 - search for optimal LightGBM params; added: models and param tuning for other models,\n    Surprise - Ridge is better than LightGBM ! Even tuning of params of LightGBM does not help much !\n    \n    Version 5 - changed number of features to 100 - Ridge - quite improved, boosting - less. \n    Version 6 - seems around 40 PCA features is better for LightGBM  - current_best_r2 = 0.46792863038644295 # \n    \n#### 1,2 Test several models - Ridge,LGB, SVR, RF etc...\n\n    Target chosen - only CD31\n    No params tuning - just fist look \n","metadata":{}},{"cell_type":"markdown","source":"# Key params","metadata":{}},{"cell_type":"code","source":"rescale_Y_to_mean0_std1 = True\n\nflag_prepare_submission = True # \n\nn_features = 3790 # use first N features - to make quick experiments\n\n#selected_targets = 'ALL'# 'CD31'","metadata":{"execution":{"iopub.status.busy":"2022-11-10T20:36:04.077362Z","iopub.execute_input":"2022-11-10T20:36:04.077943Z","iopub.status.idle":"2022-11-10T20:36:04.121808Z","shell.execute_reply.started":"2022-11-10T20:36:04.077804Z","shell.execute_reply":"2022-11-10T20:36:04.120515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Install/import modules, load technical data\n","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-10T20:36:21.733405Z","iopub.execute_input":"2022-11-10T20:36:21.733971Z","iopub.status.idle":"2022-11-10T20:36:21.780629Z","shell.execute_reply.started":"2022-11-10T20:36:21.733897Z","shell.execute_reply":"2022-11-10T20:36:21.778730Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-10T20:36:26.122518Z","iopub.execute_input":"2022-11-10T20:36:26.123138Z","iopub.status.idle":"2022-11-10T20:36:56.029885Z","shell.execute_reply.started":"2022-11-10T20:36:26.123091Z","shell.execute_reply":"2022-11-10T20:36:56.026907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-10T20:46:12.908858Z","iopub.execute_input":"2022-11-10T20:46:12.909376Z","iopub.status.idle":"2022-11-10T20:46:36.889669Z","shell.execute_reply.started":"2022-11-10T20:46:12.909339Z","shell.execute_reply":"2022-11-10T20:46:36.888325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load prepared Features for CITE-seq part of task","metadata":{}},{"cell_type":"code","source":"%%time\n\nprint('Load prepared features for CITE-seq')\n# These files contain both train and test parts .\n# For CITEseq part - first 70988 elements - train, and later 48663 - test. Overall 119651 samples.\nfn = '../input/features-multimodal-singlecell-integration/features_2719_CITEseq_source_df.csv'\ndf_cite = pd.read_csv(fn, index_col= 0)\n#df_cite.drop('Unnamed: 0', axis=1, inplace=True)\ndf_cite","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:21:16.605164Z","iopub.execute_input":"2022-11-10T21:21:16.605640Z","iopub.status.idle":"2022-11-10T21:23:03.544705Z","shell.execute_reply.started":"2022-11-10T21:21:16.605601Z","shell.execute_reply":"2022-11-10T21:23:03.542954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fn2 = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA500.csv'\ndf2 = pd.read_csv(fn2)\ndf2","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:23:28.051629Z","iopub.execute_input":"2022-11-10T21:23:28.052147Z","iopub.status.idle":"2022-11-10T21:23:39.625286Z","shell.execute_reply.started":"2022-11-10T21:23:28.052108Z","shell.execute_reply":"2022-11-10T21:23:39.623926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite['cell_id'] = df2['cell_id']\ndf_cite = df_cite.set_index('cell_id')\ndf_cite","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:21:04.241426Z","iopub.execute_input":"2022-11-10T21:21:04.242883Z","iopub.status.idle":"2022-11-10T21:21:10.981846Z","shell.execute_reply.started":"2022-11-10T21:21:04.242819Z","shell.execute_reply":"2022-11-10T21:21:10.980961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_cite.mean(axis = 0).head(3))\nprint('Pay attention - PCA features are ordered by magnitude:')\nprint(df_cite.std(axis=0).head(5))","metadata":{"execution":{"iopub.status.busy":"2022-11-10T20:47:44.751980Z","iopub.execute_input":"2022-11-10T20:47:44.752440Z","iopub.status.idle":"2022-11-10T20:47:51.582856Z","shell.execute_reply.started":"2022-11-10T20:47:44.752402Z","shell.execute_reply":"2022-11-10T20:47:51.581597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n#if 1:\nprint('Load CITE-seq targets and ')\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndisplay(df_cite_train_y)","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:20:55.037493Z","iopub.execute_input":"2022-11-10T21:20:55.038083Z","iopub.status.idle":"2022-11-10T21:21:00.008717Z","shell.execute_reply.started":"2022-11-10T21:20:55.038046Z","shell.execute_reply":"2022-11-10T21:21:00.006755Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite_train_y.columns[]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfn2 = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/_citeseq_meta_all_text_also.csv'\ndf_meta_full = pd.read_csv(fn2,index_col = 0)\ndisplay(df_meta_full)\n#if 1:\ndf_meta = pd.DataFrame(index = df_cite_train_y.index) \ndf_meta = df_meta.join(df_cell.set_index('cell_id') )\ndisplay(df_meta)","metadata":{"execution":{"iopub.status.busy":"2022-11-10T20:54:59.391776Z","iopub.execute_input":"2022-11-10T20:54:59.392280Z","iopub.status.idle":"2022-11-10T20:55:01.765222Z","shell.execute_reply.started":"2022-11-10T20:54:59.392244Z","shell.execute_reply":"2022-11-10T20:55:01.763861Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Since metric is - correlation coefficient - any rescaling aY+b will not be change it , so one can do like that:\n# it is not clear is it optimal or not (see https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360253 )\n\nlist_selected_targets = []\nY = df_cite_train_y[list_selected_targets].values\n\nif  rescale_Y_to_mean0_std1 and (Y.shape[1]>1 ) :\n    Y -= Y.mean(axis=1).reshape(-1, 1)\n    Y /= Y.std(axis=1).reshape(-1, 1)\n    print('Rescaling to mean 0 and std 1 has been done')\n    \n    \nn_features_loc = n_features\n\nX = df_cite.iloc[:,:n_features_loc].values\nprint('X.shape:', X.shape,'Y.shape:', Y.shape) ","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:23:55.052073Z","iopub.execute_input":"2022-11-10T21:23:55.053623Z","iopub.status.idle":"2022-11-10T21:23:55.848076Z","shell.execute_reply.started":"2022-11-10T21:23:55.053559Z","shell.execute_reply":"2022-11-10T21:23:55.846749Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta['Index'] = range(len(df_meta))\nscol = 'donor&day&CT'\ndf_meta[scol] =df_meta['donor'].apply(lambda x:str(x)+'_') + df_meta['day'].apply(lambda x:str(x)+'_') + df_meta['cell_type']\ndf_meta","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:24:02.914229Z","iopub.execute_input":"2022-11-10T21:24:02.914695Z","iopub.status.idle":"2022-11-10T21:24:03.368013Z","shell.execute_reply.started":"2022-11-10T21:24:02.914659Z","shell.execute_reply":"2022-11-10T21:24:03.366876Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ntrain_size_loc = 1 # from 0 to 1 # To create diversity we have an option to train on subsets. \nrandom_seed4folds = 1 # If train_size_loc < 1 - that controls the random seed to create subset of train part where we will actually train the model\nverbose = 1000\n\nfrom sklearn.model_selection import train_test_split\n\nlist_folds_indices_loc = []\n\n\n# df_folds = pd.DataFrame(index = df_meta.index)\n# df_folds['Index'] = range(df_folds.shape[0])\n\nscol = 'donor&day&CT'\nif train_size_loc < 1:\n    IX_main, IX_complementary = train_test_split(df_meta['Index'].values, \n                         train_size = train_size_loc,  random_state= random_seed4folds,  \n                         stratify = df_meta[scol] ,  shuffle=True )\nelse:\n    IX_main = np.arange(df_meta.shape[0] ) # Everything will be included \n    IX_complementary = np.array([])\n\nmask_train_main = (df_meta['Index'].isin( IX_main) )\nc = 0\nfor day2exclude in [2,3,4]: # Day to be excluded from the TRAIN , it defines the test-private-like  \n    for donor2exclude in [32606,  31800]: # We will need to predict always MALE (not female) - like on LB. (# donor 13176 - female)\n        train_index =\\\n            np.where(mask_train_main &  (df_meta['day']  != day2exclude) & ( df_meta['donor']  != donor2exclude  ) )  [0]\n        test_index1_like_private_lb =\\\n            np.where(   (df_meta['day']  == day2exclude)  ) [0]\n        test_index2_like_public_lb =\\\n            np.where( (df_meta['day']  != day2exclude) &  (df_meta['donor']  == donor2exclude ) ) [0]\n\n        list_folds_indices_loc.append( (train_index,  test_index1_like_private_lb , test_index2_like_public_lb ) )\n        \n        \n        if verbose >= 100:\n            str_fold_inf = 'Fold ' +str(c) + ': Train: excludes Day '+str(day2exclude) + ' and Donor ' + str( donor2exclude )\n            print(str_fold_inf, 'Sizes: train:',len(train_index), 'Test Like Priv'  ,len(test_index1_like_private_lb), \n                  'Test Like Publ',   len(test_index2_like_public_lb),  \n                 ); \n        c+=1        ","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:24:12.263655Z","iopub.execute_input":"2022-11-10T21:24:12.264120Z","iopub.status.idle":"2022-11-10T21:24:12.362837Z","shell.execute_reply.started":"2022-11-10T21:24:12.264084Z","shell.execute_reply":"2022-11-10T21:24:12.361012Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list_folds_indices = list_folds_indices_loc.copy()\n","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:24:19.503584Z","iopub.execute_input":"2022-11-10T21:24:19.504218Z","iopub.status.idle":"2022-11-10T21:24:19.593721Z","shell.execute_reply.started":"2022-11-10T21:24:19.504181Z","shell.execute_reply":"2022-11-10T21:24:19.591237Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import lightgbm as lgbm\nfrom sklearn.multioutput import MultiOutputRegressor\nLGBM_PARAMETERS = {\n    \"random_state\": 0, \n    \"learning_rate\":  0.03814860248977263 ,  # hypersensitive - change to 0.0671 - worsens about 2% from 0.43590181064083766\n    \"num_leaves\": 859, # hypersensitive - change to 11, worsens about 2%\n    \"max_depth\": 6, # hypersensitive - change to 10 - worsens about 1%\n    \"n_estimators\": 500, # senstitive - change by 10 - worsents about 0.5%\n    \"min_child_samples\": 13,# hypersensitive - change to 21 - worsens about 3% \n    \"reg_lambda\": 3.0151837674804667, # seems only makes worse\n    \"reg_alpha\": 6.734991732483901, # seems only makes worse\n    \"colsample_bytree\": 0.9, # only worsens\n    \"subsample\": 1, # Does not seem to influence at all\n    }\nmodel = MultiOutputRegressor(lgbm.LGBMRegressor(**LGBM_PARAMETERS)) \nstr_model_inf = str(model)\nprint(str_model_inf)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport time\n\nflag_train_on_full_data4submit = True\nflag_save_predictions_on_each_sub_fold = True\nflag_calculate_and_save_predictions_for_test_public_like = True\n\ndict_model_results = {}\n\nY_pred_oof_private_like = np.zeros_like(Y)\nif flag_calculate_and_save_predictions_for_test_public_like:\n    Y_pred_oof_public_like = np.zeros_like(Y) # FEMALE is always in train - thus we have no predictions for female donor for that test part\n\nY_pred_submission_Kaggle_way = np.zeros( (119651-70988 , Y.shape[1]) )  # Way 2 (Kaggle way) - average predictions of models on each fold\nif flag_train_on_full_data4submit: \n    Y_pred_submission_classical_way = np.zeros( (119651-70988 , Y.shape[1]) )  # Way 1 (Classical way) - retrain model on the whole data\n\nt0 = time.time()\nfor fold, indices_tuple  in enumerate( list_folds_indices ):\n    t00 = time.time()\n    train_index = indices_tuple[0]\n    model.fit(X[train_index], Y[train_index])\n    if verbose >= 100:\n        print('fold:',fold, 'Train time:%.2f'%(time.time()-t00 ), 'X_train.shape:',X[train_index].shape )\n        \n    # Calculate predictions for test-private-like\n    i_loc = 1\n    indices_loc = indices_tuple[i_loc] # Private-like test\n    y_pred = model.predict( X[indices_loc ])\n    Y_pred_oof_private_like[indices_loc ] += y_pred * 0.5   # 0.5 - comes from the fact that our \"quasi-folds\" have intersection OOF parts - which correspond to \"DAYs\"\n    if flag_save_predictions_on_each_sub_fold:\n        dict_model_results['Pred_fold'+str(fold)+'_part'+str(i_loc)] = y_pred\n\n\n    if (flag_save_predictions_on_each_sub_fold) or ( flag_calculate_and_save_predictions_for_test_public_like):\n        # Calculate predictions for test-public-like\n        i_loc = 2\n        indices_loc = indices_tuple[i_loc] # Public-like test  # FEMALE is always in train - thus we have no predictions for female donor for that test part\n        y_pred = model.predict( X[indices_loc ])\n        if flag_calculate_and_save_predictions_for_test_public_like:\n            Y_pred_oof_public_like[indices_loc ] += y_pred * 0.5 # 0.5 - comes from the fact that our \"quasi-folds\" have intersection OOF parts\n                                                # Since we THREE DAYs we should put  1/(3-1) = 1/2\n        if flag_save_predictions_on_each_sub_fold:\n            dict_model_results['Pred_fold'+str(fold)+'_part'+str(i_loc)] = y_pred\n\n    if flag_save_predictions_on_each_sub_fold:\n        # Save train predictions in case some deep analysis will be done \n        i_loc = 0  \n        indices_loc = indices_tuple[i_loc] # Public-like test  # FEMALE is always in train - thus we have no predictions for female donor for that test part\n        y_pred = model.predict( X[indices_loc ])\n        dict_model_results['Pred_fold'+str(fold)+'_part'+str(i_loc)] = y_pred\n        \n    # Prepare submission data in way 2 - \"Kaggle way\"\n    Y_pred_submission_Kaggle_way += model.predict( X[70988:,: ])   # Way 2 (Kaggle way) - average predictions of models on each fold   \n\n    if verbose >= 100:\n        print('Finished fold:',fold, 'Train&Inference time:%.2f'%(time.time()-t00 ), 'X_train.shape:',X[train_index].shape )\n    \n    \nY_pred_submission_Kaggle_way /= (1+fold)  # Do not forget to average  predictions from all folds \n\nif flag_train_on_full_data4submit: \n    model.fit(X[:70988,:], Y) # Way 1 (Classical way) - retrain model on the whole data\n    Y_pred_submission_classical_way = model.predict( X[70988:,: ])   \n    \ndict_model_results['Time'] = time.time() - t0\ndict_model_results['list_folds_indices'] = list_folds_indices[:fold]\ndict_model_results['Y_pred_oof_private_like'] = Y_pred_oof_private_like\ndict_model_results['Y_pred_oof_public_like'] = Y_pred_oof_public_like\nif flag_train_on_full_data4submit: \n    dict_model_results['Y_pred_submission_classical_way'] = Y_pred_submission_classical_way\ndict_model_results['Y_pred_submission_Kaggle_way'] = Y_pred_submission_Kaggle_way\n\n    \nif verbose >= 10:\n    print('X.shape, Y.shape:',X.shape, Y.shape)\n    print('Finished. Folds:',fold+1, 'Time:%.2f'%(time.time()-t0), 'Y_pred_oof_private_like.shape:',Y_pred_oof_private_like.shape,\n         'Y_pred_submission_Kaggle_way.shape:',Y_pred_submission_Kaggle_way.shape )    \n","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:25:41.784367Z","iopub.execute_input":"2022-11-10T21:25:41.786947Z","iopub.status.idle":"2022-11-10T21:25:41.796175Z","shell.execute_reply.started":"2022-11-10T21:25:41.786823Z","shell.execute_reply":"2022-11-10T21:25:41.794229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n\nlist_targets_names = list_selected_targets # list( df_cite_train_y.columns )\n\nfilename_prefix = 'LGBM'\nfilename_postfix = ''\nfor key_loc in ['Y_pred_oof_private_like', 'Y_pred_oof_public_like', 'Y_pred_submission_classical_way', 'Y_pred_submission_Kaggle_way'  ]:\n    if key_loc in dict_model_results.keys():\n        fn_loc = filename_prefix + key_loc + filename_postfix + '.csv'\n        print('Key:', key_loc,'Filename:',  fn_loc)\n        d = pd.DataFrame(dict_model_results[key_loc],  columns = list_targets_names )\n        d.to_csv( fn_loc)\n        print(d.shape)\n        display(d.head(5))\n        display(d.tail(5))\n        print(); print();\n        \n","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:25:43.344777Z","iopub.execute_input":"2022-11-10T21:25:43.345512Z","iopub.status.idle":"2022-11-10T21:25:43.361124Z","shell.execute_reply.started":"2022-11-10T21:25:43.345402Z","shell.execute_reply":"2022-11-10T21:25:43.359360Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nimport gc\ndef correlation_score(y_true, y_pred):\n    \"\"\"Scores the predictions according to the competition rules. \n    \n    It is assumed that the predictions are not constant.\n    \n    Returns the average of each sample's Pearson correlation coefficient\"\"\"\n    \n    # Input should be matrices - does not make sense for vectors \n    if  len(y_pred.shape)< 2: return -10 # Some result to inform for incorrect input\n    if  y_pred.shape[1] < 2: return -10 # Some result to inform for incorrect input\n\n    y2 = y_pred.copy()\n    y2 -= y2.mean(axis=1).reshape(-1, 1);    y2 /= y2.std(axis=1).reshape(-1, 1)    \n    if rescale_Y_to_mean0_std1:\n        y1 = y_true # Already rescaled \n    else:\n        y1 = y_true.copy(); \n        y1 -= y1.mean(axis=1).reshape(-1, 1);    y1 /= y1.std(axis=1).reshape(-1, 1) \n        \n    c = (y1*y2).mean().mean()# Correlation for rescaled matrices is just matrix product and average \n    \n    # Memory control:\n    if not rescale_Y_to_mean0_std1:\n        del y1\n    del y2\n    gc.collect()\n    \n    return c\n\n    # Slow way:\n    \n    if type(y_true) == pd.DataFrame: y_true = y_true.values\n    if type(y_pred) == pd.DataFrame: y_pred = y_pred.values\n    corrsum = 0\n    for i in range(len(y_true)):\n        corrsum += np.corrcoef(y_true[i], y_pred[i])[1, 0]\n    return corrsum / len(y_true)\n  ","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:25:45.435224Z","iopub.execute_input":"2022-11-10T21:25:45.435718Z","iopub.status.idle":"2022-11-10T21:25:45.450264Z","shell.execute_reply.started":"2022-11-10T21:25:45.435682Z","shell.execute_reply":"2022-11-10T21:25:45.449020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_error\nfrom sklearn.metrics import r2_score","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# 'Y_pred_oof_private_like', 'Y_pred_oof_public_like',\nY_pred = dict_model_results['Y_pred_oof_private_like']\nY_true = Y\n\ns_cor = correlation_score( Y_true , Y_pred  ) # It takes about 5 seconds ! \n# If only one target it will automatically return -10\n\n\ns_r2 = r2_score( Y_true , Y_pred  )\ns_mse = mean_squared_error( Y_true , Y_pred  )\n\nprint(str(model)[:80], 'n_feat:', X.shape[1],)\nprint('OOF scores: ', s_cor, s_r2, s_mse, 'n_feat:', X.shape[1], 'Modeling time:%.3f secs.'%dict_model_results['Time'],  \n      'Oof size:', Y_true.shape[0], 'Y_pred.shape:', Y_pred.shape )\n      ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds1 = pd.DataFrame()\nIX_loc = 0\nds1.loc[IX_loc, 'Model'] = str_model_inf\nds1.loc[IX_loc, 'Corr OOF'] = s_cor\nds1.loc[IX_loc, 'r2 OOF'] = s_r2\nds1.loc[IX_loc, 'mse OOF'] = s_mse\nds1.loc[IX_loc, 'n_feat'] = X.shape[1]\nds1.loc[IX_loc, 'n_targets'] = Y_pred.shape[1]\nds1.loc[IX_loc, 'Time'] = np.round(dict_model_results['Time'],3)\nfn_loc = filename_prefix + 'ModelStat1Main_' + filename_postfix + '.csv'\nds1.to_csv(fn_loc)\nprint(fn_loc)\nds1\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dict_model_results.keys()","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:25:58.111197Z","iopub.execute_input":"2022-11-10T21:25:58.111777Z","iopub.status.idle":"2022-11-10T21:27:35.746295Z","shell.execute_reply.started":"2022-11-10T21:25:58.111734Z","shell.execute_reply":"2022-11-10T21:27:35.744846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n%%time\nverbose = 1000\nlist_folds_indices =  dict_model_results['list_folds_indices']\ndf_fold_score_stat = pd.DataFrame();df_fold_score_stat.index.name = 'Fold'\nt0 = time.time()\nfor fold, indices_tuple  in enumerate( list_folds_indices ):\n    # Calculate metrics on test and train folds \n    list_scores = []; list_scores_r2 = []\n    for i_loc in [1,2,0]: #, indices_loc in enumerate( indices_tuple) :\n        if i_loc == 1:\n            name_loc = 'Test Priv Like'\n        if i_loc == 2:\n            name_loc = 'Test Publ Like'\n        if i_loc == 0:\n            name_loc = 'Train'\n            \n        indices_loc = indices_tuple[i_loc]\n        key_loc = 'Pred_fold'+str(fold)+'_part'+str(i_loc)\n        if key_loc in dict_model_results.keys():\n            y_pred = dict_model_results[key_loc]\n            s = correlation_score(Y[indices_loc],  y_pred    )\n            df_fold_score_stat.loc[fold,'Corr '+name_loc] = s\n            s = r2_score( Y[indices_loc] , y_pred  )\n            df_fold_score_stat.loc[fold,'r2 '+name_loc] = s\n            s = mean_squared_error( Y[indices_loc] , y_pred  )\n            df_fold_score_stat.loc[fold,'mse '+name_loc] = s\n    if verbose >= 1000:\n        print('Fold:', fold, 'Shapes of train:', train_index.shape, 'Tests: ',[t.shape for t in indices_tuple[1:] ] )        \ndisplay( df_fold_score_stat )\ndisplay(df_fold_score_stat.describe(percentiles=[] ).iloc[1:,:])\nds2 = pd.concat([df_fold_score_stat.describe(percentiles=[] ).iloc[1:,:],  df_fold_score_stat])\nds2\n","metadata":{"execution":{"iopub.status.busy":"2022-11-10T21:00:33.478623Z","iopub.execute_input":"2022-11-10T21:00:33.479126Z","iopub.status.idle":"2022-11-10T21:00:33.498845Z","shell.execute_reply.started":"2022-11-10T21:00:33.479088Z","shell.execute_reply":"2022-11-10T21:00:33.497281Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fn_loc = filename_prefix + 'ModelStat2Foldwise_' + filename_postfix + '.csv'\nds2.to_csv(fn_loc)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nt = df_fold_score_stat.describe(percentiles=[] ).iloc[1,: ].T\nt= t.to_frame().T.reset_index()\n\nds1b = pd.concat([ds1, t], axis = 1 )\nfn_loc = filename_prefix + 'ModelStat1Main_' + filename_postfix + '.csv'\nds1b.to_csv(fn_loc)\nds1b\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndisplay(ds1)\ndisplay(ds1b)\ndisplay( ds2 )\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )","metadata":{},"execution_count":null,"outputs":[]}]}