{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59094,"databundleVersionId":6541963,"sourceType":"competition"}],"dockerImageVersionId":30558,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\n#### Brief\n\nWe consider simple \"gentle\" param tuner, and apply it to several models and category encoders.  It tunes params \"gently\" one by one using the predescribed regions. It allows to tune not only ONE score, but several scores simulataneosly.  It might be useful in the following circumstances which are present for the Kaggle \"Open Problems – Single-Cell Perturbations\":\n\n    CV-LB correspondence is NOT good, so\n    1) Small gentle param changes are preferable to check each step with LB. \n    2) Uplift of several scores (different folds/CV) might be better indication of the better params, rather than just one score. \n    \n    And also:\n    3) For multi-target tasks one may prefer to optimize params for all targets separately - thus: \n        number of params is multiplied by number of targets and such huge number may not be easily approached with standard optimizers like Optuna.\n        And simple approach here might be the better choice. \n    4) Such optimization process is easy to visualize / interpret.\n    5) We provide ready to use cases for many models/encoders \n\n#### Application -  benchmarking 8 category encoders with several models  \n\n    We consider 8  category Encoders: 'QuantileEncoder','TargetEncoder', 'CatBoostEncoder', 'LeaveOneOutEncoder' , 'JamesSteinEncoder', 'OneHotEncoder', 'HelmertEncoder', 'BackwardDifferenceEncoder'\n        from the Python package Category Encoders \n    As well as several models: Ridge, KernelRidge (with different kernels), etc\n    \n    Params tuned for both - model + encoder (at least in cases when encoder have tunable params:  'QuantileEncoder','TargetEncoder', 'CatBoostEncoder' ). \n   \n    Params tuned, separately for three CV-schemes: AmbrosM, MT, Random (5 folds).\n\n    Conclusions and notes:\n    1) KRR-lin with LOO encoder is top in all CV, but the alpha parameter is different, so not clear what should be right choice for LB - we check all - see below. \n    2) LOO encoder  is consitetly top for the Ridge model (see e.g. the previous notebook),  but for KRR models it is not always the case - for KRR-rbf it is the last one, but for KRR-linear is top1\n    3) In general KRR-linear behaves more similar to Ridge, than KRR-rbf. In particular optimal alpha are quite similar while for KRR-rbf - 1000-100000  times smaller. (The same for SVR model)\n    4) Result for Random and AmbrosM CV are quite correlated , while not that much with MT\n    5) AmbrosM and Random schemes requires  stronger regularization - alpha - often - 10 times greater, but that might not well correspond to LB as noted by MT originally\n\n    6) The LB scores for the top1 model (tsvd30_KRRlin_LeaveOneOutEncoder blended with \"priors\")  with alpha chosen by different CV-schemes:\n        LB 0.605 - AmbrosM (alpha = 5e5) , LB 0.613 - MT-scheme (alpha = 1e4) , LB 0.605 - Random CV-scheme (alpha = 2e5)\n        Thus in that case AmbrosM and Random schemes better correspond to LB.\n        That is different from the original setup of AmrbosM and MT (Ridge+OHE-compound), where MT scheme  better corresponded to LB\n        \n\n    The results are the following (Ordered by Mean rank for AmbrosM and MT scores, see output file: \"report_multi_model.csv\" )\n    \n| Model Name                                     | AmbrosM score | MT score   | Random score | AmbrosM Rank | MT Rank | Mean Rank AmbrosM MT | Random Rank | AmbrosM alpha | MT alpha | Random alpha |\n| ---------------------------------------------- | ------------- | ---------- | ------------ | ------------ | ------- | -------------------- | ----------- | ------------ | -------- | ------------ |\n| tsvd30_KRRlin_LeaveOneOutEncoderCompound      | 0.9598   | 2.5352 | 1.2059  | 1            | 1       | 1                    | 1           | 500000       | 10000    | 200000       |\n| tsvd30_Ridge_LeaveOneOutEncoderCompound       | 0.9721   | 2.5353 | 1.2134  | 7            | 2       | 4.5                  | 2           | 200000       | 10000    | 100000       |\n| tsvd30_KRRrbf_QuantileEncoderCompound         | 0.9662   | 2.65636 | 1.219  | 3            | 9       | 6                    | 5           | 5            | 0.5      | 5            |\n| tsvd30_KRRlin_JamesSteinEncoderCompound       | 0.9823    | 2.5942 | 1.2571  | 12           | 3       | 7.5                  | 16          | 10000000     | 2000     | 10000000     |\n| tsvd30_KRRrbf_TargetEncoderCompound           | 0.9672    | 2.66 | 1.2213  | 5            | 11      | 8                    | 8           | 5            | 0.5      | 5            |\n| tsvd30_Ridge_QuantileEncoderCompound          | 0.9821   | 2.6546 | 1.217  | 11           | 6       | 8.5                  | 3           | 500000       | 500000   | 200000       |\n| tsvd30_KRRrbf_JamesSteinEncoderCompound       | 0.9662   | 2.6648 | 1.2211  | 2            | 17      | 9.5                  | 6           | 5            | 0.5      | 5            |\n| tsvd30_KRRlin_OneHotEncoderCompound           | 0.9665   | 2.6647   | 1.2183  | 4            | 16      | 10                   | 4           | 100          | 10       | 100          |\n| tsvd30_KRRrbf_HelmertEncoderCompound           | 0.9677    | 2.6644 | 1.2213  | 6            | 15      | 10.5                | 7           | 5            | 0.5      | 5            |\n| tsvd30_Ridge_OneHotEncoderCompound             | 0.9934   | 2.6557 | 1.2331  | 15           | 7       | 11                   | 10          | 50           | 20       | 50           |\n| tsvd30_Ridge_TargetEncoderCompound             | 0.9941   | 2.6320 | 1.237  | 17           | 5       | 11                   | 13          | 100000       | 20000    | 50000        |\n| tsvd30_KRRrbf_OneHotEncoderCompound             | 0.9940   | 2.6562 | 1.234  | 16           | 8       | 12                   | 12          | 0.02         | 0.01     | 0.02         |\n| tsvd30_KRRlin_QuantileEncoderCompound          | 0.9736   | 2.6764 | 1.2247  | 9            | 19      | 14                   | 9           | 10000000     | 0.0001   | 200000       |\n| tsvd30_Ridge_JamesSteinEncoderCompound         | 1.0441   | 2.5951 | 1.2974  | 24           | 4       | 14                   | 24          | 2000000      | 2000     | 5000000      |\n| tsvd30_KRRlin_CatBoostEncoderCompound          | 0.9722   | 2.8784 | 1.2331  | 8            | 22      | 15                   | 11          | 5000000      | 2000000  | 5000000      |\n| tsvd30_KRRlin_TargetEncoderCompound             | 0.9814   | 2.6814 | 1.2408  | 10           | 20      | 15                   | 14          | 20000000     | 3000     | 10000000     |\n| tsvd30_Ridge_HelmertEncoderCompound             | 1.0076   | 2.6574 | 1.2598  | 22           | 10      | 16                   | 17          | 100000       | 20000    | 100000       |\n| tsvd30_KRRrbf_BackwardDifferenceEncoderCompound  | 0.9966   | 2.6637 | 1.2664  | 18           | 14      | 16                   | 18          | 1000         | 0.002    | 1000         |\n| tsvd30_KRRlin_BackwardDifferenceEncoderCompound  | 1.002   | 2.6631 | 1.2682  | 19           | 13      | 16                   | 19          | 100000       | 5        | 10           |\n| tsvd30_Ridge_CatBoostEncoderCompound            | 0.9882   | 2.8773 | 1.2443  | 14           | 21      | 17.5                | 15          | 500000       | 2000000  | 1000000      |\n| tsvd30_Ridge_BackwardDifferenceEncoderCompound  | 1.0192   | 2.6631 | 1.2682  | 23           | 12      | 17.5                | 20          | 5            | 5        |  10  |\n| tsvd30_Ridge_CatBoostEncoderCompound            |0.9877 |2.9104|1.2691   |13|23|18|21|1|0.5|200    |\n| tsvd30_KRRlin_HelmertEncoderCompound            | 1.003791      |2.6743 |1.273|20|18|19|22|10000000|0.0001|5000000 |\n| tsvd30_KRRrbf_LeaveOneOutEncoderCompound        | 1.005138      |3.2283|1.2805|21|24|22.5|23|1|0.0001  | 30 |\n    \n\n    \nSummaries with LB/CV scores here: https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing and discussed here: https://docs.google.com/presentation/d/1wiz0Wmt4D54pqMMsIOyJHuQYMZ3hTBZQQnjbLzwoGYY/edit?usp=sharing. Top models can achieve around LB 0.60*. \n    \n\n#### Organization and Navigating the Notebook    \n\n    The simple example of the tuner is in the section \"Gentle param tuner. Example 1\".\n    (Look there for detailed explanations). \n    \n    In later sections with wrap the code of the tuner and auxiliaries into functions.\n    \n    Main simulation in the section \"Process many models and encoders\"\n    Which loops accross 1) models 2) encoders 3) CV-schemes.\n    For each triplet model+encoder+CV-scheme the params for model and encoder are tuned. \n    Submission files are prepared and saved to csv for each optimal params set.  \n    Scores are scored. Optimization process is visualized - see plot in that section.\n    \n    The obtained benchmarking scoring results are presented in the section \"Main Report\" \n\n#### \"One config to rule them all\" - key coding principle\n\n    The main organizing idea is to pack all the params for model, encoder, reducer into ONE config, such that all functions can work with it. \n    That is the dictionary \"main_config_model_feature_etc_initial\" \n    Tuner changes only it during the optimization process. \n    Auxilliary wrapper functions for modeling/encoding takes that main_config, transforms it to format specific to model/encoder/reducer to use them.\n    Thus we can achieve that same code can work with different models/encoders/whatever. \n\n    PS\n    In basic versions use tsvd 30 components reduction of targets and encoding only the compound based on the previous experience.\n    That can be changed in the first section \"Key params\" - pca/ica can also be considered and encoding for the both - cell and type and compound.\n    Further results can be found in the notebooks mention below and the google docs mentioned above. \n    \n#### The model(s) strcuture\n\n    In short: we make dimensional reduction  for targets (tsvd by default), encode categorical features by all obtained tsvd components, use model (sklearn or boosting or whatever) to predict them, make tsvd.inverse_transform to get original targets. \n    The models considered: Ridge, Kernel Ridge (rbf,linear), SVR, \n    \n\nThe notebook is build upon previous more simple notebooks having similar structure:\n\nCustom CV-schemes: https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes, \n\nCategory encoders, chembert, Morgan fingerprints, molecular descriptors: https://www.kaggle.com/code/alexandervc/op2-category-encoders-chembert-fingerpints-moldes, \n\nTarget category encoders: https://www.kaggle.com/alexandervc/op2-target-encoders, \n\nAnd more advanced modeling combining all that in: https://www.kaggle.com/code/alexandervc/op2-advanced-modeling-tuning-featengineering-etc.\n\n\n\n    \n\n","metadata":{}},{"cell_type":"markdown","source":"# Key params","metadata":{}},{"cell_type":"code","source":"reducer_main = 'tsvd' # 'pca' # 'ica'\nn_components_main = 50 # Reduction of targets is done by tsvd to n_components, after predicting components we take tsvd.inverse_transform\n\nlist_features_to_encode = ['cell_type', 'sm_name'] #  ['sm_name']#  [ 'sm_name'] #  ['cell_type'],  \n# For simple models like Ridge and that kind of features seems encoding only the compounds just forgetting about cell type at all gives sometimes better results. ","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:04:25.351108Z","iopub.execute_input":"2023-11-01T10:04:25.351463Z","iopub.status.idle":"2023-11-01T10:04:25.384847Z","shell.execute_reply.started":"2023-11-01T10:04:25.351439Z","shell.execute_reply":"2023-11-01T10:04:25.383701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create info string on params:\nstr_inf_cfg = ''# encoding #+ ' '\nif 'sm_name' in list_features_to_encode: str_inf_cfg += 'Compound'\nif 'cell_type' in list_features_to_encode: str_inf_cfg += 'CellType'\nstr_inf_cfg += ' '+reducer_main + str(n_components_main)\nprint(str_inf_cfg)\n\nimport category_encoders as ce","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:04:25.42545Z","iopub.execute_input":"2023-11-01T10:04:25.42601Z","iopub.status.idle":"2023-11-01T10:04:26.571711Z","shell.execute_reply.started":"2023-11-01T10:04:25.425985Z","shell.execute_reply":"2023-11-01T10:04:26.570751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport time\nt0start = time.time() \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport pandas as pd\nimport tensorflow as tf\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:04:26.573485Z","iopub.execute_input":"2023-11-01T10:04:26.574221Z","iopub.status.idle":"2023-11-01T10:04:36.554349Z","shell.execute_reply.started":"2023-11-01T10:04:26.574189Z","shell.execute_reply":"2023-11-01T10:04:36.553471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\ndf_de_train = pd.read_parquet(fn)# , index_col = 0)\nprint(df_de_train.shape)\ndisplay(df_de_train )\n\nplt.figure(figsize = (20,4) )\nv = df_de_train.iloc[:,5:].max(axis = 0 ).sort_values(ascending = False, key = abs )\nplt.plot(v.values,'*-')\nplt.title('Max-abs DE for genes',fontsize = 20 )\nplt.grid()\nplt.show()\ndisplay(v.head(15))\n\n# %%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/id_map.csv'\ndf_id_map = pd.read_csv(fn,index_col = 0)\nprint(df_id_map.shape)\ndisplay(df_id_map)\nfn = '/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv'\ndf_sample_submit = pd.read_csv(fn, index_col = 0)\nprint(df_sample_submit.shape)\ndisplay( df_sample_submit )\n\n\nfeatures_columns = ['cell_type','sm_name']\nX_full_categorical = df_de_train[ list_features_to_encode ].values\nX_submit_categorical = df_id_map[list_features_to_encode].values # pd.DataFrame(df_id_map, columns= features_columns )\nY_full = df_de_train.iloc[:,5:].values\n\nprint();\nprint('X_full_categorical.shape, X_submit_categorical.shape ,  Y_full.shape', X_full_categorical.shape, X_submit_categorical.shape ,  Y_full.shape)  ; print();\nprint(X_full_categorical[:3,:5]); print();\nprint(X_submit_categorical[:3,:5]); print();\nprint(Y_full[:3,:5]); print();\n\n","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:04:36.555646Z","iopub.execute_input":"2023-11-01T10:04:36.556417Z","iopub.status.idle":"2023-11-01T10:04:41.928869Z","shell.execute_reply.started":"2023-11-01T10:04:36.556382Z","shell.execute_reply":"2023-11-01T10:04:41.927206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.decomposition import TruncatedSVD\nreducer = TruncatedSVD(n_components=50, n_iter=7, random_state=42)\nY_full_red = reducer.fit_transform( Y_full )\nprint( Y_full_red.shape )\n\ndf_stat_loc = pd.DataFrame(Y_full_red[:,:50] ).describe()    \npd.set_option('display.max_columns', 100)\n# pd.set_option('display.max_rows', None)    \ndisplay(df_stat_loc.round() )\nfig = plt.figure(figsize = (20,3))\nplt.plot(df_stat_loc.loc['std',:].values,'*-' )\nplt.grid()\nplt.title('std for targets reduced', fontsize = 20 )\nplt.xlabel('tsvd i_component', fontsize = 20)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:04:41.932422Z","iopub.execute_input":"2023-11-01T10:04:41.932882Z","iopub.status.idle":"2023-11-01T10:04:44.172584Z","shell.execute_reply.started":"2023-11-01T10:04:41.932839Z","shell.execute_reply":"2023-11-01T10:04:44.170949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Custom CV-schemes","metadata":{}},{"cell_type":"code","source":"%%time\nimport random \nfrom sklearn.model_selection import KFold\n\n# Prepare indexes for AmbrosM scheme:\n# For AmbrosM scheme: \nfolds_index_data_AmbrosM = [ ]\ntrain_sm_names = ['Idelalisib', 'Crizotinib', 'Linagliptin', 'Palbociclib', 'Dabrafenib', 'Alvocidib', 'LDN 193189', 'R428', 'Porcn Inhibitor III', \n  'Belinostat', 'Foretinib', 'MLN 2238', 'Penfluridol', 'Dactolisib', 'O-Demethylated Adapalene', 'Oprozomib (ONX 0912)', 'CHIR-99021']\nlist_fold_ids =  ['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells'] \nfor fold_id in list_fold_ids:\n    mask_va = (df_de_train.cell_type == fold_id) & ~df_de_train.sm_name.isin(train_sm_names)\n    mask_tr = ~mask_va # 485 or 487 training rows\n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_AmbrosM.append( [IX_train, IX_test ])\n\n# For MT CV scheme:\nfolds_index_data_MT = []\nfold_to_compounds = {0: ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189',  'Linagliptin', 'O-Demethylated Adapalene'],\n 1: ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III'],\n 2: ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol',  'R428']}\nfor fold_id in [0,1,2]:\n    mask_va = df_de_train['cell_type'].isin(['Myeloid cells', 'B cells']) & df_de_train['sm_name'].isin(fold_to_compounds[fold_id])\n    mask_tr = ~mask_va    \n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_MT.append( [IX_train, IX_test ])    \n    \n    \ndict_folds_Antonina = {0: [4, 15, 18, 41, 47, 53, 56, 57, 59, 61, 62, 66, 78, 83, 86, 93, 100, 110, 115, 116, 117, 120, 123, 131, 148, 152, 161, 185, 189, 193, 197, 198, 205, 207, 209, 210, 213, 218, 219, 225, 228, 230, 250, 252, 257, 263, 286, 292, 294, 304, 306, 310, 325, 327, 336, 338, 339, 350, 351, 356, 357, 360, 362, 369, 379, 385, 395, 419, 424, 430, 432, 434, 438, 441, 443, 445, 448, 452, 456, 465, 467, 471, 472, 474, 498, 508, 523, 534, 542, 544, 574, 581, 584, 589, 601, 607], 1: [1, 3, 5, 14, 24, 38, 40, 49, 50, 54, 55, 60, 79, 82, 87, 102, 112, 114, 118, 119, 127, 132, 146, 149, 166, 179, 183, 184, 186, 190, 192, 201, 203, 204, 216, 227, 242, 253, 261, 264, 271, 289, 295, 296, 299, 300, 309, 323, 324, 329, 333, 337, 354, 387, 390, 393, 406, 408, 415, 416, 420, 421, 422, 436, 442, 449, 453, 458, 462, 463, 469, 484, 487, 489, 490, 492, 497, 501, 505, 506, 509, 513, 515, 516, 546, 548, 566, 569, 585, 587, 588, 592, 593, 595, 600, 605], 2: [0, 16, 19, 44, 48, 51, 52, 58, 63, 65, 70, 81, 85, 88, 89, 92, 111, 113, 122, 125, 135, 137, 139, 150, 164, 175, 191, 202, 206, 211, 217, 220, 221, 229, 239, 243, 247, 248, 251, 258, 259, 260, 267, 270, 274, 281, 288, 290, 291, 303, 322, 326, 328, 332, 347, 359, 364, 370, 383, 384, 389, 392, 401, 423, 428, 454, 455, 461, 464, 494, 496, 504, 510, 511, 524, 525, 526, 530, 540, 541, 551, 565, 568, 570, 572, 573, 575, 577, 579, 580, 597, 598, 603, 606, 610, 611, 613], 3: [21, 22, 26, 36, 42, 45, 46, 64, 67, 69, 71, 80, 84, 90, 101, 103, 134, 136, 140, 147, 151, 160, 163, 165, 167, 174, 176, 177, 180, 181, 182, 187, 188, 195, 196, 199, 200, 212, 215, 226, 231, 240, 241, 249, 255, 256, 262, 265, 269, 285, 287, 297, 298, 307, 312, 319, 321, 330, 340, 342, 344, 345, 348, 358, 366, 368, 382, 391, 400, 403, 404, 405, 425, 427, 431, 433, 435, 444, 450, 451, 457, 459, 460, 473, 485, 499, 500, 502, 503, 507, 512, 514, 527, 528, 529, 531, 543, 547, 553, 554, 564, 571, 576, 583, 586, 591, 596, 602, 604, 608, 609], 4: [2, 6, 7, 17, 20, 23, 25, 27, 28, 29, 37, 39, 43, 68, 91, 121, 124, 126, 128, 129, 130, 133, 138, 153, 162, 178, 194, 208, 214, 222, 223, 224, 232, 244, 245, 246, 254, 266, 268, 272, 273, 282, 283, 284, 293, 301, 302, 305, 308, 311, 320, 331, 334, 335, 341, 343, 346, 349, 352, 353, 355, 361, 363, 365, 367, 377, 378, 380, 381, 386, 388, 394, 396, 397, 398, 399, 402, 407, 417, 418, 426, 429, 437, 439, 440, 446, 447, 466, 468, 470, 481, 482, 483, 486, 488, 491, 493, 495, 532, 533, 545, 549, 550, 552, 555, 562, 563, 567, 578, 582, 590, 594, 599, 612], 5: [8, 9, 10, 11, 12, 13, 30, 31, 32, 33, 34, 35, 72, 73, 74, 75, 76, 77, 94, 95, 96, 97, 98, 99, 104, 105, 106, 107, 108, 109, 141, 142, 143, 144, 145, 154, 155, 156, 157, 158, 159, 168, 169, 170, 171, 172, 173, 233, 234, 235, 236, 237, 238, 275, 276, 277, 278, 279, 280, 313, 314, 315, 316, 317, 318, 371, 372, 373, 374, 375, 376, 409, 410, 411, 412, 413, 414, 475, 476, 477, 478, 479, 480, 517, 518, 519, 520, 521, 522, 535, 536, 537, 538, 539, 556, 557, 558, 559, 560, 561]}\n\nfolds_index_data_Antonina = []    \nfor i in range(5):  \n    l_valid = np.array( dict_folds_Antonina[i] )\n    l_train = np.array( [k for k in range(614) if k not in l_valid ] )\n    folds_index_data_Antonina.append( [l_train, l_valid ])    \nprint( len(folds_index_data_Antonina ) )\n\nclass KFold_custom:\n    '''\n    Class with similar to sklearn \"KFold\" class, interface.\n    Support of CV schemes proposed by AmbrosM, MT, Kishan and standard sklearn KFold\n    Supports options to return IX_train,IX_valid,IX_test - triplet - if valid_size > 0\n    Examples:\n    kf = KFold_custom('AmbrosM'):     \n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n        pass\n    kf = KFold_custom('AmbrosM', valid_size = 0.1 )     \n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split()) :\n        pass\n        \n    kf = KFold_custom(CV_scheme = 'Random') # wrapper for sklearn Kfolds with shuffle = True\n    kf = KFold_custom(CV_scheme = CV_scheme, random_state = 42 ) # fixing random_state ensures reprodicibility of folds \n    '''\n    def __init__(self, CV_scheme,  valid_size = 0, n_splits=5, random_state = 42, verbose = 0 ):\n        self.CV_scheme = CV_scheme\n        self.CV_scheme_inf =  CV_scheme\n        self.valid_size = np.clip(valid_size,0,1)\n        self.random_state = random_state\n\n        self.folds_index_data = [ [np.arange(0), np.arange(0), np.arange(0)] ] # Example data\n        self.n_splits = 1 # Example data\n        if CV_scheme == 'Kishan1':\n            # One fold scheme which contains train, valid, test parts \n            IX_train_Kishan1 = np.arange(429)\n            IX_val_Kishan1 = np.arange(429,521)\n            IX_test_Kishan1 = np.arange(521,614)\n            self.n_splits = 1\n            self.folds_index_data =  [   [IX_train_Kishan1, IX_val_Kishan1, IX_test_Kishan1]  ]\n        elif CV_scheme == 'Tonya':\n            self.n_splits = 5\n            self.folds_index_data = folds_index_data_Antonina\n        elif CV_scheme == 'AmbrosM':\n            self.n_splits = 4\n            self.folds_index_data = folds_index_data_AmbrosM\n        elif CV_scheme == 'MT':\n            self.n_splits = 3\n            self.folds_index_data = folds_index_data_MT\n        elif 'Random'.lower() in CV_scheme.lower():\n            self.n_splits = n_splits\n            self.CV_scheme_inf = 'Random_'+str(n_splits) +'_'+str(random_state )\n            kf = KFold(n_splits=n_splits, random_state = random_state, shuffle=True )#,\n            self.folds_index_data = list( kf.split( np.arange(614) ) )\n        elif 'Full'.lower() in CV_scheme.lower():\n            # Just return the full set as both train and test - it useful to re-train the model on the entire data\n            IX_train_full = np.arange(614)\n            IX_test_full = np.arange(614)\n            self.n_splits = 1\n            self.folds_index_data = [   [IX_train_full, IX_test_full]  ]\n        else:\n            s = 'Uncrecognized CV_scheme ' + str(CV_scheme) \n            raise ValueError(s)\n\n        self.verbose = verbose \n        if verbose >= 10:\n            print(self.CV_scheme, self.n_splits, len(self.folds_index_data) )\n            #print(self.folds_index_data)\n            \n    def split(self, X=None):\n        '''\n        X - NOT used, just for compatibility with sklearn \n        '''\n        for item in self.folds_index_data:\n            if self.CV_scheme in ['Kishan1']:\n                if self.valid_size > 0:\n                    yield item[0],item[1],item[2]\n                else :\n                    yield np.array( list(set(item[0])|set(item[1]))  ) , item[2] # Return train, test only \n            elif self.CV_scheme in ['Tonya', 'AmbrosM','MT', 'Random', 'Full']:\n                if self.valid_size == 0:\n                    yield item[0],item[1]\n                else:\n                    # Split \"full-train\"->( real-train, valid ) \n                    index_for_real_train_part = int(  (1-self.valid_size) * len(item[0]) )\n                    if self.random_state is None:\n                        IX_train =  item[0][:index_for_real_train_part]\n                        IX_valid =  item[0][index_for_real_train_part:]\n                    elif self.random_state == -1:\n                        p = np.random.permutation(len(item[0])) \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    else:\n                        np.random.seed(self.random_state) # Set temporary random seed\n                        p = np.random.permutation(len(item[0])) \n                        np.random.seed(random.randint(0,30000)) # Randomize seed again - use Python random, not numpy       \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    yield IX_train, IX_valid, item[1]\n                    \n        \n    def get_list_main_CV_schemes(self):\n        return ['AmbrosM','MT','Kishan1','Random']\n    def get_n_splits(self, X=np.array([])):\n        return self.n_splits","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:04:44.174594Z","iopub.execute_input":"2023-11-01T10:04:44.174947Z","iopub.status.idle":"2023-11-01T10:04:44.220454Z","shell.execute_reply.started":"2023-11-01T10:04:44.174917Z","shell.execute_reply":"2023-11-01T10:04:44.219043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Simple Modeling\n\n\nResults change NOT much after smoothing 100: 100-1000 difference in 3-th digit, 1000-10_000 - 4-th digit, 10_000 - inf - 5-th digit of score.\n\n\n","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\nimport category_encoders as ce\n\nverbose = 1\n\nn_components = 30\nreducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n\nalpha = 20000.0\nmodel = Ridge(alpha=alpha)\n\nlist_smoothing = [0,1,5,8,9,10,11,12,15, 20,50,100,200,500, 1000,10_000, np.inf]\n# enc = ce.TargetEncoder(smoothing = smoothing )\n# enc = ce.LeaveOneOutEncoder(sigma = 0.05 )\n\nCV_scheme = 'MT'\nprint('CV_scheme:', CV_scheme); print()\nvalid_size = 0 \nkf = KFold_custom(CV_scheme, valid_size = valid_size)   \n\nY_oof_pred_blend = np.zeros( (Y_full.shape) ); i_self_blend = 0; Y_oof_pred = np.zeros( (Y_full.shape) );\nlist_result_scores = []\nfor smoothing in list_smoothing:\n    print('smoothing:', smoothing, 'i:', i_self_blend)\n    enc = ce.TargetEncoder(smoothing = smoothing )\n    \n    list_mrrmse = [];list_r2 = []\n    kf = KFold_custom(CV_scheme, valid_size = valid_size)   \n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n        # \n        # Get train, test:\n        X_train_categorical,X_test_categorical = X_full_categorical[IX_train], X_full_categorical[IX_test]\n\n        # Reduce the targets \n        Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n\n        if verbose >= 10:\n            print('fold:', i_fold, 'X_train.shape, X_test.shape, Y_train_red.shape', X_train.shape, X_test.shape, Y_train_red.shape)\n\n        ### -------------- Target Encoding -------------------------------------------------------------------\n        for i_target in range(Y_train_red.shape[1]):\n            if i_target == 0:\n                X_train_encoded = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n                X_test_encoded = enc.transform(X_test_categorical)\n            else:\n                X_encoded_tmp = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n                X_train_encoded = np.concatenate( [X_train_encoded, X_encoded_tmp], axis = 1)\n                X_encoded_tmp = enc.transform(X_test_categorical)\n                X_test_encoded = np.concatenate( [X_test_encoded, X_encoded_tmp], axis = 1)\n\n\n        model.fit(X_train_encoded, Y_train_red)\n        Y_test_pred_red = model.predict( X_test_encoded )\n        Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n        Y_oof_pred[ IX_test ] =  Y_test_pred\n\n        Y_test = Y_full[IX_test]\n        mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n        print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n        list_mrrmse.append(mrrmse); list_r2.append(r2)\n        \n        \n    Y_oof_pred_blend =  ( Y_oof_pred_blend * i_self_blend + Y_oof_pred ) / (i_self_blend + 1)\n    i_self_blend += 1\n    print('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ),  'Average r2:', np.round(np.mean(list_r2 ),4 ))\n    print()\n    list_result_scores.append( np.mean(list_mrrmse ) )\n    \nprint('i_self_blend', i_self_blend)    \nY_oof_pred_blend =  Y_oof_pred_blend\n\nprint('Blend scoring:')\nkf = KFold_custom(CV_scheme, valid_size = valid_size)   \nlist_mrrmse = [];list_r2 = []\nfor i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n    Y_test = Y_full[IX_test]\n    mrrmse = np.sqrt(np.square(Y_test - Y_oof_pred_blend[IX_test]).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_oof_pred_blend[IX_test] ) \n    print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n    list_mrrmse.append(mrrmse); list_r2.append(r2)\nprint('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ),  'Average r2:', np.round(np.mean(list_r2 ),4 ))\nprint()\n\nres = pd.Series(index = list_smoothing, data = list_result_scores)\nres.index.name = 'smoothing'; res.name = 'mrrmse'\ndisplay( res.sort_values().to_frame() )\nres.sort_values().to_csv('scores_from_smoothing_first_example.csv')","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:04:44.222573Z","iopub.execute_input":"2023-11-01T10:04:44.222984Z","iopub.status.idle":"2023-11-01T10:06:11.262581Z","shell.execute_reply.started":"2023-11-01T10:04:44.222951Z","shell.execute_reply":"2023-11-01T10:06:11.261791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nprint('Ridge alpha was optimized for smoothing = 10, so not surprising that we get it as optimal smoothing')\nfig = plt.figure(figsize = (15,3) )\nplt.plot(res.values , '*-')\nplt.title('Scores dependence on smoothing. Ridge. TargetEncoder',fontsize = 20 )\nplt.xticks(range(len(list_smoothing)), list_smoothing)  # Set x-axis tick labels\nplt.grid()\nplt.xlabel('Smoothing',fontsize = 20 )\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:06:11.264107Z","iopub.execute_input":"2023-11-01T10:06:11.264732Z","iopub.status.idle":"2023-11-01T10:06:11.529358Z","shell.execute_reply.started":"2023-11-01T10:06:11.2647Z","shell.execute_reply":"2023-11-01T10:06:11.52788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Get model, encoder, reducer - wrapper functions","metadata":{"execution":{"iopub.status.busy":"2023-10-30T09:27:54.507085Z","iopub.execute_input":"2023-10-30T09:27:54.507494Z","iopub.status.idle":"2023-10-30T09:27:54.513283Z","shell.execute_reply.started":"2023-10-30T09:27:54.50746Z","shell.execute_reply":"2023-10-30T09:27:54.51205Z"}}},{"cell_type":"code","source":"%%time\nfrom sklearn.metrics import r2_score\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.decomposition import FastICA\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.linear_model import Ridge\nfrom sklearn.svm import LinearSVR\nfrom sklearn.svm import SVR\nfrom sklearn.kernel_ridge import KernelRidge\n\nimport lightgbm as lgb\nimport catboost\nfrom catboost import CatBoostRegressor, Pool\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.ensemble import ExtraTreesRegressor\n\nfrom sklearn.multioutput import MultiOutputRegressor\n\n\nimport category_encoders as ce\n\nimport warnings\n\n\ndef get_brief_string_info_on_config(main_config_model_feature_etc):\n\n    str_inf_cfg = ''\n    if 'reducer' in main_config_model_feature_etc.keys():\n        str_inf_cfg += main_config_model_feature_etc['reducer']\n        n_components = main_config_model_feature_etc.get('n_components', 25)\n        str_inf_cfg += str(n_components)\n    for key_loc in ['model', 'encoder'  ]:\n        if key_loc in main_config_model_feature_etc.keys():\n            str_inf_cfg += '_'+main_config_model_feature_etc[key_loc]\n            \n    if 'sm_name' in list_features_to_encode: str_inf_cfg += 'Compound'\n    if 'cell_type' in list_features_to_encode: str_inf_cfg += 'CellType'\n    \n    return str_inf_cfg\n\ncfg1 = {'model':'Ridge', 'reducer': 'tsvd',  'n_components': 2, 'encoder':'TargetEncoder'   }\nstr_inf_cfg = get_brief_string_info_on_config(cfg1)\nprint(str_inf_cfg)\n\n\n\ndef get_encoder( main_config_model_feature_etc ):\n    '''\n    Returns category encoder, intialized with appropriate params from the input config\n    \n    enc, str_encoder_id   =  get_encoder( main_config_model_feature_etc )\n    '''\n    \n    features_mode = main_config_model_feature_etc['encoder']\n    str_encoder_id =  main_config_model_feature_etc['encoder']\n    if features_mode == 'TargetEncoder':\n        # Target encoding for categorical features.\n        # For the case of continuous target: \n        # features are replaced with a blend of the expected value of the target given particular categorical value \n        # and the expected value of the target over all the training data.\n        # The blend weight is controlled by \"smoothing\"\n        smoothing = main_config_model_feature_etc.get('smoothing', 10)\n        min_samples_leaf = main_config_model_feature_etc.get('min_samples_leaf', 20)\n        enc = ce.TargetEncoder(smoothing = smoothing, min_samples_leaf = min_samples_leaf  )\n    elif features_mode == 'QuantileEncoder':\n        # Quantile Encoding for categorical features.\n        # This a statistically modified version of target MEstimate encoder \n        # where selected features are replaced by the statistical quantile instead of the mean. \n        # Replacing with the median is a particular case where self.quantile = 0.5. \n        # In comparison to MEstimateEncoder it has two tunable parameter m and quantile\n\n        quantile = main_config_model_feature_etc.get('quantile', .5)\n        m = main_config_model_feature_etc.get('smoothing', 1) \n        m = main_config_model_feature_etc.get('m', 1) \n        # this is the “m” in the m-probability estimate. Higher value of m results into stronger shrinking. M is non-negative. 0 for no smoothing.\n        enc = ce.QuantileEncoder(quantile = quantile, m = m  )\n    elif features_mode == 'CatBoostEncoder':\n        # CatBoost Encoding for categorical features. \n        \n        a = main_config_model_feature_etc.get('smoothing', 1) \n        a = main_config_model_feature_etc.get('a', 1) \n        # additive smoothing (it is the same variable as “m” in m-probability estimate). By default set to 1.\n        random_state = main_config_model_feature_etc.get('random_state', None) \n        sigma = main_config_model_feature_etc.get('sigma', None)\n        # adds normal (Gaussian) distribution noise into training data in order to decrease overfitting (testing data are untouched). \n        # sigma gives the standard deviation (spread or “width”) of the normal distribution.\n        # See example here: https://www.kaggle.com/code/alexandervc/op2-target-encoders/notebook#Example:-add-Gaussian-noise-into-training-data-(%22sigma%22-parameter-of-encoder)\n        enc = ce.CatBoostEncoder(a = a, sigma = sigma  , random_state = random_state)\n    elif features_mode == 'JamesSteinEncoder':\n        #  James-Stein estimator - it is a kind of target encoder but uses theoretical  estimate for smoothing parameter.\n        sigma = main_config_model_feature_etc.get('sigma', None)\n        if (sigma is not None) and (sigma > 0):             randomized = True\n        else: randomized = False \n        # adds normal (Gaussian) distribution noise into training data in order to decrease overfitting (testing data are untouched). \n        # sigma gives the standard deviation (spread or “width”) of the normal distribution.\n        # It works only if randomized = True \n        # See example here: https://www.kaggle.com/code/alexandervc/op2-target-encoders/notebook#Example:-add-Gaussian-noise-into-training-data-(%22sigma%22-parameter-of-encoder)\n        random_state = main_config_model_feature_etc.get('random_state', None) \n        enc = ce.JamesSteinEncoder( sigma = sigma  , randomized = randomized, random_state = random_state)\n        \n    elif features_mode == 'LeaveOneOutEncoder':\n        # This is very similar to target encoding but excludes the current row’s target when calculating the mean target for a level to reduce the effect of outliers.\n        \n        sigma = main_config_model_feature_etc.get('sigma', None)\n        # adds normal (Gaussian) distribution noise into training data in order to decrease overfitting (testing data are untouched). \n        # sigma gives the standard deviation (spread or “width”) of the normal distribution.\n        # It works only if randomized = True \n        # See example here: https://www.kaggle.com/code/alexandervc/op2-target-encoders/notebook#Example:-add-Gaussian-noise-into-training-data-(%22sigma%22-parameter-of-encoder)\n        random_state = main_config_model_feature_etc.get('random_state', None) \n        enc = ce.LeaveOneOutEncoder( sigma = sigma  , random_state = random_state)\n        \n    elif features_mode in  ['OneHotEncoder', 'HelmertEncoder', 'BackwardDifferenceEncoder', 'CountEncoder', 'OrdinalEncoder'  ]:\n        enc = getattr(ce, features_mode )()\n        \n    return enc, str_encoder_id\n\ndef get_reducer( main_config_model_feature_etc , verbose = 0): \n    '''\n    Returns the dimensional reduction method initialized with appropriate params from the config.\n    \n    reducer, str_reducer_id  = get_reducer( main_config_model_feature_etc , verbose = 0)\n    '''\n    \n    str_reducer_id =  str( main_config_model_feature_etc['reducer'] )\n    if main_config_model_feature_etc['reducer']=='tsvd':\n        n_components = main_config_model_feature_etc.get('n_components',30)\n        reducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n        str_reducer_id = 'tsvd'+str( n_components )\n    elif main_config_model_feature_etc['reducer']=='ica':\n        n_components = main_config_model_feature_etc.get('n_components',30)\n        reducer = FastICA(n_components=n_components, random_state=0, whiten='unit-variance')\n        str_reducer_id = 'ica'+str( n_components )\n    elif main_config_model_feature_etc['reducer']=='pca':\n        n_components = main_config_model_feature_etc.get('n_components',30)\n        reducer =  PCA(n_components=n_components)\n        str_reducer_id = 'pca'+str( n_components )\n    \n    if verbose >= 100:\n        print( str_reducer_id )\n        print( reducer )\n        print( main_config_model_feature_etc )\n\n    return reducer, str_reducer_id    \n\n\ndef get_model(main_config_model_feature_etc , verbose = 0):\n    '''\n    Returns model specified by the config\n    \n    model, str_model_id = get_model(main_config_model_feature_etc , verbose = 0)\n    '''\n    \n    str_model_id =  str( main_config_model_feature_etc['model'] )\n    if main_config_model_feature_etc['model'] == 'Ridge':\n        alpha = main_config_model_feature_etc.get('alpha',1)\n        fit_intercept = main_config_model_feature_etc.get('fit_intercept',True)\n        positive=main_config_model_feature_etc.get('positive',False)\n        model = Ridge( alpha ,  fit_intercept = fit_intercept,  positive = positive,)\n        str_model_id  = 'Ridge'+str(alpha)\n    elif main_config_model_feature_etc['model'] == 'KRRrbf':\n        # main_config_model_feature_etc = {'model': 'KRR', 'alpha':1, 'kernel': 'linear', 'gamma': None,  'degree': 3,  'coef0': 1, 'reducer': 'tsvd', 'n_components': n_components, 'features_mode': 'target_enc_i_th_target', 'list_features_in': ['cell_type', 'sm_name']}\n        alpha = main_config_model_feature_etc.get('alpha',1)\n        kernel = 'rbf'\n        gamma = main_config_model_feature_etc.get('gamma',None ) # Gamma parameter for the RBF, laplacian, polynomial, exponential chi2 and sigmoid kernels. Interpretation of the default value is left to the kernel; see the documentation for sklearn.metrics.pairwise. Ignored by other kernels.\n        coef0 = main_config_model_feature_etc.get('coef0',1.0 )  #  Independent term in kernel function. It is only significant in ‘poly’ and ‘sigmoid’.\n        degree = main_config_model_feature_etc.get('degree',3 ) ##Degree of the polynomial kernel function (‘poly’). Must be non-negative. Ignored by all other kernels.\n        model = KernelRidge(alpha=alpha, kernel = kernel,  gamma = gamma, coef0=coef0, degree = degree)\n    elif main_config_model_feature_etc['model'] in [ 'KRR', 'KRRrbf', 'KRRlin' ] :\n        # main_config_model_feature_etc = {'model': 'KRR', 'alpha':1, 'kernel': 'linear', 'gamma': None,  'degree': 3,  'coef0': 1, 'reducer': 'tsvd', 'n_components': n_components, 'features_mode': 'target_enc_i_th_target', 'list_features_in': ['cell_type', 'sm_name']}\n        alpha = main_config_model_feature_etc.get('alpha',1)\n        kernel = main_config_model_feature_etc.get('kernel','linear' ) # {‘linear’, ‘poly’, ‘rbf’, ‘sigmoid’, ‘precomputed’} or callable, default=’rbf’\n        if main_config_model_feature_etc['model'] == 'KRRrbf': kernel = 'rbf'\n        elif main_config_model_feature_etc['model'] == 'KRRlin': kernel = 'linear'\n        gamma = main_config_model_feature_etc.get('gamma',None ) # Gamma parameter for the RBF, laplacian, polynomial, exponential chi2 and sigmoid kernels. Interpretation of the default value is left to the kernel; see the documentation for sklearn.metrics.pairwise. Ignored by other kernels.\n        coef0 = main_config_model_feature_etc.get('coef0',1.0 )  #  Independent term in kernel function. It is only significant in ‘poly’ and ‘sigmoid’.\n        degree = main_config_model_feature_etc.get('degree',3 ) ##Degree of the polynomial kernel function (‘poly’). Must be non-negative. Ignored by all other kernels.\n        model = KernelRidge(alpha=alpha, kernel = kernel,  gamma = gamma, coef0=coef0, degree = degree)\n        \n    elif main_config_model_feature_etc['model'] == 'LSVR':\n        C = main_config_model_feature_etc.get('C',1)#  default=1.0 Regularization parameter. The strength of the regularization is inversely proportional to C. Must be strictly positive.\n        max_iter = main_config_model_feature_etc.get('max_iter',1000) # sklearn default default=1000 \n        epsilon = main_config_model_feature_etc.get('epsilon',0.1) # Epsilon parameter in the epsilon-insensitive loss function. Note that the value of this parameter depends on the scale of the target variable y. If unsure, set epsilon=0. \n        # 0.1 is used: https://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain?scriptVersionId=145592242&cellId=88\n        model = MultiOutputRegressor( LinearSVR(max_iter= max_iter, epsilon= epsilon, C=C) )\n        \n    elif main_config_model_feature_etc['model'] in ['SVR','SVRrbf','SVRlin' ]:\n        # main_config_model_feature_etc = {'model':'SVR', 'kernel':'rbf','C':1 }\n        kernel = main_config_model_feature_etc.get('kernel','rbf' ) # {‘linear’, ‘poly’, ‘rbf’, ‘sigmoid’, ‘precomputed’} or callable, default=’rbf’\n        if 'rbf' in main_config_model_feature_etc['model']: kernel = 'rbf'\n        elif 'lin' in main_config_model_feature_etc['model']: kernel = 'linear'\n        C = main_config_model_feature_etc.get('C',1)#  default=1.0 Regularization parameter. The strength of the regularization is inversely proportional to C. Must be strictly positive.\n        epsilon = main_config_model_feature_etc.get('epsilon',0.1) # Epsilon parameter in the epsilon-insensitive loss function. Note that the value of this parameter depends on the scale of the target variable y. If unsure, set epsilon=0. \n        gamma = main_config_model_feature_etc.get('gamma','scale' ) # Kernel coefficient for ‘rbf’, ‘poly’ and ‘sigmoid’. \n        coef0 = main_config_model_feature_etc.get('coef0',0.0 )  #  Independent term in kernel function. It is only significant in ‘poly’ and ‘sigmoid’.\n        tol = main_config_model_feature_etc.get('tol',0.001 ) #\n        degree = main_config_model_feature_etc.get('degree',3 ) ##Degree of the polynomial kernel function (‘poly’). Must be non-negative. Ignored by all other kernels.\n        max_iter = main_config_model_feature_etc.get('max_iter',-1) # sklearn default default=-1 - it may cause too long\n        shrinking = main_config_model_feature_etc.get('shrinking', True) # Whether to use the shrinking heuristic. See the User Guide.        \n        model = MultiOutputRegressor( SVR(kernel = kernel, max_iter= max_iter, epsilon= epsilon, C=C, gamma = gamma, coef0=coef0, degree = degree, tol = tol, shrinking=shrinking) )\n        \n\n    elif main_config_model_feature_etc['model'] == 'CATB':\n        params_loc = main_config_model_feature_etc.get('CATB_params',{} )\n        params_loc['loss_function'] = main_config_model_feature_etc.get('loss_function', 'RMSE' )\n        categorical_features = main_config_model_feature_etc.get('categorical_features',[])\n        for prm_name in ['iterations', 'depth','learning_rate','subsample', 'colsample_bylevel', 'min_data_in_leaf', 'random_strength','random_seed','random_state', 'l2_leaf_reg' ]: #  \n            if prm_name in main_config_model_feature_etc.keys(): params_loc[prm_name] = main_config_model_feature_etc[prm_name]\n        model = MultiOutputRegressor( CatBoostRegressor(cat_features=categorical_features, verbose = 0, **params_loc ) )  \n        \n        # https://forecastegy.com/posts/catboost-hyperparameter-tuning-guide-with-optuna/#subsample-subsample\n        # def objective(trial):\n        #     params = {\n        #         \"iterations\": 1000,\n        #         \"learning_rate\": trial.suggest_float(\"learning_rate\", 1e-3, 0.1, log=True),\n        #         \"depth\": trial.suggest_int(\"depth\", 1, 10),\n        #         \"subsample\": trial.suggest_float(\"subsample\", 0.05, 1.0),\n        #         \"colsample_bylevel\": trial.suggest_float(\"colsample_bylevel\", 0.05, 1.0),\n        #         \"min_data_in_leaf\": trial.suggest_int(\"min_data_in_leaf\", 1, 100),\n        #     }\n\n        #     model = cb.CatBoostRegressor(**params, silent=True)\n        #     model.fit(X_train, y_train)\n        #     predictions = model.predict(X_val)\n        #     rmse = mean_squared_error(y_val, predictions, squared=False)\n        #     return rmse\n\n            \n    elif main_config_model_feature_etc['model'] == 'RFR':\n        #n_estimators=100,*, criterion='squared_error', max_depth=None, min_samples_split=2, min_samples_leaf=1, min_weight_fraction_leaf=0.0, max_features=1.0, max_leaf_nodes=None, min_impurity_decrease=0.0, bootstrap=True, oob_score=False, n_jobs=None, random_state=None, verbose=0, warm_start=False, ccp_alpha=0.0, max_samples=None\n        params_loc = main_config_model_feature_etc.get('RFR_params',{} )\n        for prm_name in ['criterion', 'n_estimators', 'max_depth', 'min_samples_split', 'min_samples_split', 'min_samples_leaf', 'min_weight_fraction_leaf',  'max_features', 'max_leaf_nodes','min_impurity_decrease','random_state','ccp_alpha','max_samples' ]:\n            if prm_name in main_config_model_feature_etc.keys(): params_loc[prm_name] = main_config_model_feature_etc[prm_name]\n        model = MultiOutputRegressor( RandomForestRegressor( **params_loc )  )\n        \n    elif main_config_model_feature_etc['model'] == 'ETR':\n        # n_estimators=100, *, criterion='squared_error', max_depth=None, min_samples_split=2, min_samples_leaf=1, min_weight_fraction_leaf=0.0, max_features=1.0, max_leaf_nodes=None, min_impurity_decrease=0.0, bootstrap=False, oob_score=False, n_jobs=None, random_state=None, verbose=0, warm_start=False, ccp_alpha=0.0, max_samples=None\n        params_loc = main_config_model_feature_etc.get('ETR_params',{} )\n        for prm_name in ['criterion', 'n_estimators', 'max_depth', 'min_samples_split', 'min_samples_split', 'min_samples_leaf', 'min_weight_fraction_leaf',  'max_features', 'max_leaf_nodes','min_impurity_decrease','random_state','ccp_alpha','max_samples' ]:\n            if prm_name in main_config_model_feature_etc.keys(): params_loc[prm_name] = main_config_model_feature_etc[prm_name]\n        model = MultiOutputRegressor(  ExtraTreesRegressor( **params_loc )  )\n        \n\n    elif main_config_model_feature_etc['model'] == 'LGB':\n        params_loc = main_config_model_feature_etc.get('LGB_params',{} )\n        for prm_name in ['n_estimators', 'max_depth','learning_rate', 'colsample_bytree', 'subsample', 'random_state', 'reg_alpha',  'reg_lambda', 'num_leaves','min_child_samples' ]:\n            if prm_name in main_config_model_feature_etc.keys(): params_loc[prm_name] = main_config_model_feature_etc[prm_name]\n        model = MultiOutputRegressor(  lgb.LGBMRegressor( **params_loc ) )\n    elif main_config_model_feature_etc['model'] == 'LGBcv1_036':\n        # Params for LGB just on two categorical features as it is  found by optimization in https://www.kaggle.com/code/alexandervc/op2-models-cv-tuning#Optuna+LightGBM\n        # But LB score is terrible - \n        params_best1_cv1_036 = {'random_state': 0, 'n_estimators': 20, 'reg_alpha': 6.764079452929363, 'reg_lambda': 0.41900776876588564, \n        'colsample_bytree': 0.3, 'subsample': 0.7, 'max_depth': 1, 'learning_rate': 0.08456104070184789,  'num_leaves': 682, 'min_child_samples': 102}\n        model = MultiOutputRegressor( lgb.LGBMRegressor( **params_best1_cv1_036 )  )\n        \n    if verbose >= 100:\n        print( str_model_id )\n        print( model )\n        print( main_config_model_feature_etc )\n        \n    return model, str_model_id\n\nmodel, str_model_id = get_model({'model':'Ridge'}, verbose = 100)\nprint(model, str_model_id  ); print()\nmodel, str_model_id = get_model({'model':'CATB'}, verbose = 100)\nprint(model, str_model_id  ); print()\n","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:06:11.531164Z","iopub.execute_input":"2023-11-01T10:06:11.531613Z","iopub.status.idle":"2023-11-01T10:06:13.666534Z","shell.execute_reply.started":"2023-11-01T10:06:11.531576Z","shell.execute_reply":"2023-11-01T10:06:13.664722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Core modeling function ","metadata":{}},{"cell_type":"code","source":"%%time \nimport warnings\n\ndef do_modeling( main_config_model_feature_etc, list_CV_scheme = ['AmbrosM'],  scoring_mode = 'return_all_folds_results',    verbose = 0): \n    '''\n    Makes modelling with model, encodings, dim-reductions specified in the input config\n    Computes CV-scores, submit predictions, oof predictions. \n    \n    scoring_mode: 'return_all_folds_results' - return list of scores for all folds and CV-schemes.\n                  'return_average_over_folds' - return list for each CV-scheme - average over folds \n    '''\n    \n    warnings.filterwarnings(\"ignore\") \n\n        \n    dict_modeling_res = {}\n    list_final_scores_output = []\n    dict_modeling_res['scores'] = list_final_scores_output\n    \n    for CV_scheme in list_CV_scheme:\n        dict_modeling_res[CV_scheme] = {} \n        \n        if verbose >= 1:\n            print('CV_scheme:', CV_scheme)\n        kf = KFold_custom(CV_scheme)# , valid_size = valid_size)  # valid_size = 0 \n    \n    \n        list_mrrmse = [];list_r2 = []; i_blend_submit  = 0\n        for i_fold,(IX_train,IX_test) in enumerate( kf.split()) :\n            # \n            # Get train, test:\n            X_train_categorical,X_test_categorical = X_full_categorical[IX_train], X_full_categorical[IX_test]\n\n            enc, str_encoder_id   =  get_encoder( main_config_model_feature_etc )\n            reducer, str_reducer_id  = get_reducer( main_config_model_feature_etc , verbose = 0)\n            model, str_model_id = get_model(main_config_model_feature_etc , verbose = 0)\n            \n            # Reduce the targets \n            Y_train_red = reducer.fit_transform(Y_full[IX_train] )\n            \n            ### -------------- Target Encoding -------------------------------------------------------------------\n            for i_target in range(Y_train_red.shape[1]):\n                if i_target == 0:\n                    X_train_encoded = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n                    X_test_encoded = enc.transform(X_test_categorical)\n                    X_submit_encoded = enc.transform(X_submit_categorical)\n                else:\n                    X_encoded_tmp = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n                    X_train_encoded = np.concatenate( [X_train_encoded, X_encoded_tmp], axis = 1)\n                    X_encoded_tmp = enc.transform(X_test_categorical)\n                    X_test_encoded = np.concatenate( [X_test_encoded, X_encoded_tmp], axis = 1)\n                    X_encoded_tmp = enc.transform(X_submit_categorical)\n                    X_submit_encoded = np.concatenate( [X_submit_encoded, X_encoded_tmp], axis = 1)\n            \n            \n            ##--------------- Train model ---------------------------------------------------\n            model.fit(X_train_encoded, Y_train_red)\n            Y_test_pred_red = model.predict( X_test_encoded )\n            ##--------------- Predictions and scoring ----------------------------------------\n            Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n            Y_test = Y_full[IX_test]\n            mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n            list_mrrmse.append(mrrmse); list_r2.append(r2)\n            if verbose >= 1:\n                print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n                \n            if i_blend_submit == 0:\n                Y_submit_pred = reducer.inverse_transform( model.predict( X_submit_encoded  ) )\n            else:\n                Y_submit_pred = ( Y_submit_pred * i_blend_submit + reducer.inverse_transform( model.predict( X_submit_encoded  ) ) ) / ( i_blend_submit + 1 )\n            i_blend_submit += 1\n            \n            \n        dict_modeling_res[CV_scheme]['mrrmse'] = list_mrrmse\n        dict_modeling_res[CV_scheme]['r2'] = list_r2\n        dict_modeling_res[CV_scheme]['Y_submit'] = Y_submit_pred\n        if scoring_mode == 'return_all_folds_results':\n            list_final_scores_output += list(list_mrrmse)\n        elif scoring_mode == 'return_average_over_folds':\n            list_final_scores_output.append(  np.mean( list_mrrmse ) )\n        else:\n            list_final_scores_output.append(  np.mean( list_mrrmse ) )\n \n        if verbose >= 1:            \n            print('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ),  'Average r2:', np.round(np.mean(list_r2 ),4 ))\n            print()   \n            \n            \n    dict_modeling_res['scores'] = list_final_scores_output\n    return dict_modeling_res\n\nmain_config_model_feature_etc_tmp = {'model':'Ridge', 'alpha':1e5, 'reducer':'tsvd','n_components':30, 'encoder':'TargetEncoder', 'smoothing':100 }\ndict_modeling_res = do_modeling( main_config_model_feature_etc_tmp, list_CV_scheme = ['MT'],  scoring_mode = 'return_all_folds_results',    verbose = 1000)\ndict_modeling_res","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:06:13.668032Z","iopub.execute_input":"2023-11-01T10:06:13.668394Z","iopub.status.idle":"2023-11-01T10:06:18.953711Z","shell.execute_reply.started":"2023-11-01T10:06:13.668369Z","shell.execute_reply":"2023-11-01T10:06:18.95181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Gentle param tuner. Example 1 \n\nThe idea of the tuner is quite simple. \nWe define possible sets of values for each parameter, tuner goes through each set one by one taking the best value. \nSuch process would rarely lead to a global extremum, but our purpose is different - change params gently, by the reasons explained in the introduction.\n\nSo technically:\nwe create dictionary \"dict_prms_lists\" which contains the sets of variation for all params, like that:\n\n    dict_prms_lists = {'alpha':[1e4,1e5,1e6],  'quantile':[0.5,0.6], 'm': [100,1000,10_000]}\n\n    Loop over param names e.g. keys of that dictionary\n\n    Loop over variation set for param chosen at the loop above. \n    \n    That is basically all.\n    \n    Even more technically:\n    There is outer loop which allows to repeat the entire process several times (it is rarely necessary - that is a bit surprising ).\n    The key computation is wrapped in the function \n    do_modeling( main_config_model_feature_etc, list_CV_scheme = list_CV_scheme,  scoring_mode = scoring_mode, verbose = 0 )\n    \n    The function return VECTOR(!) of scores (can be different folds or different CV-schemes) \n    the parameter is beeing accepted as new optimal if ALL components of the vector are better than the previous values.\n    (One can encode easily code other conditions).\n    \n    \"One config to rule them all\"\n    The main point is that it takes the config main_config_model_feature_etc which describes all: model+encoder+reducer params in the standartized way,\n    such that we do not need to change anything in the tuner except specifying the config to add new models/encoders/etc...\n    \n     \n    \n    \n\n\n\n","metadata":{}},{"cell_type":"code","source":"%%time\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")     # To suppress all warnings\n\n# ---------------------------------------- Key  Params --------------------------------------------------------------\n\nlist_CV_scheme = ['MT'] # ['AmbrosM','MT','Random']\nscoring_mode = 'return_all_folds_results' # 'return_average_over_folds'\n    # Controls do we want to find scores uplift over ALL folds , or only average over folds \n    \nn_rounds = 1\n\nverbose = 100 \n\n\n### --------------------------- Initial config (config describes ALL(!) - model, encoder, all params for model and encoder)\nmain_config_model_feature_etc = {}\nmain_config_model_feature_etc['model'] = 'KRRlin' # 'KRRrbf' # 'Ridge'\nmain_config_model_feature_etc['alpha'] = 1e4\nmain_config_model_feature_etc['reducer'] = reducer_main #  'tsvd'# 'pca' # 'ica'\nmain_config_model_feature_etc['n_components'] = n_components_main\nmain_config_model_feature_etc['encoder'] = 'LeaveOneOutEncoder'# 'QuantileEncoder'\n# 8 Encoders: 'QuantileEncoder','TargetEncoder', 'CatBoostEncoder', 'LeaveOneOutEncoder' , 'JamesSteinEncoder', 'OneHotEncoder', 'HelmertEncoder', 'BackwardDifferenceEncoder'\nmain_config_model_feature_etc['quantile'] = 0.5\nmain_config_model_feature_etc['m'] = 1\n\n# Get scores for initial params:\ndict_modeling_res = do_modeling( main_config_model_feature_etc, list_CV_scheme = list_CV_scheme,  scoring_mode = scoring_mode, verbose = 0 )\nlist_best_scores = dict_modeling_res['scores']\nprint('Initial list_best_scores:', np.round( list_best_scores, 4) , 'average:',  np.round(np.mean( list_best_scores) , 4), )\n\n### --------------------------- Dictionary with optimized params and their intervals \n\n#dict_prms_lists = {'alpha':[1e4,1e5,1e6],  'quantile':[0.5,0.6], 'm':list(range(20)) + [100,1000,10_000]}\ndict_prms_lists = {'alpha':[1e2,1e3,1e4, 2e4,5e4, 1e5,2e5, 5e5, 1e6]} # ,  'quantile':[0.5,0.6], 'm':list(range(20)) + [100,1000,10_000]}\n\n\n### --------------------------- Starting tuner \nt00 = time.time() ; i_total_count = 0\nlist_best_scores_afer_prm_optimization = []\ndf_report_accepted_params = pd.DataFrame(); \ndf_report_all_steps = pd.DataFrame(); \n\nfor i_round in range(n_rounds): # rounds to repeat the entire search \n    for i_prm in range( len( dict_prms_lists )) : # Choose what param will be  tuned\n        \n        prm_name = list(dict_prms_lists.keys())[i_prm]\n        list_prm_values = dict_prms_lists[prm_name] \n        flag_achieved_better_result = False \n        save_previous_prm_value = main_config_model_feature_etc[prm_name]\n        \n        for prm_value in list_prm_values: # Grid the list of the params \n            i_total_count += 1\n\n            main_config_model_feature_etc[prm_name]  = prm_value\n\n            dict_modeling_res = do_modeling( main_config_model_feature_etc, list_CV_scheme = list_CV_scheme,  scoring_mode = scoring_mode, verbose = 0 )\n             \n            flag_accept_prm =  np.all(np.asarray(dict_modeling_res['scores']) <  np.asarray(list_best_scores)   ) \n            \n            if flag_accept_prm:\n                list_best_scores = dict_modeling_res['scores']\n                best_prm = prm_value\n                flag_achieved_better_result = True\n                # Report prepare:\n                IX = len(df_report_accepted_params)+1\n                df_report_accepted_params.loc[IX,'prm name' ] = prm_name\n                df_report_accepted_params.loc[IX,'prm value' ] = prm_value\n                df_report_accepted_params.loc[IX,'trial' ] = i_total_count\n                for ii in range(len( list_best_scores )) :\n                    df_report_accepted_params.loc[IX,'score '+str(ii) ] = list_best_scores[ii]\n                df_report_accepted_params.loc[IX,'score average' ] = np.mean( list_best_scores ) \n                df_report_accepted_params.loc[IX,'Full Prm' ] = str( main_config_model_feature_etc ) \n                \n                \n            if verbose >= 100:\n                print('trial:',i_total_count,  'prm_name', prm_name, 'prm_value', prm_value,'scores:', np.round(dict_modeling_res['scores'], 4), \n                      'average:',  np.round(np.mean( dict_modeling_res['scores']) , 4),\n                     'best scores:', np.round(list_best_scores, 4), 'average:',  np.round(np.mean( list_best_scores) , 4),\n                      'flag_accept_prm', flag_accept_prm )\n             \n            # Report prepare:\n            IX = len(df_report_all_steps)+1\n            df_report_all_steps.loc[IX,'prm name' ] = prm_name\n            df_report_all_steps.loc[IX,'prm value' ] = prm_value\n            df_report_all_steps.loc[IX,'trial' ] = i_total_count\n            df_report_all_steps.loc[IX,'score average' ] = np.mean( dict_modeling_res['scores'] ) \n            df_report_all_steps.loc[IX,'score average best' ] = np.mean( list_best_scores ) \n            for ii in range(len( dict_modeling_res['scores'] )) :\n                df_report_all_steps.loc[IX,'score '+str(ii) ] = dict_modeling_res['scores'][ii]\n            for ii in range(len( dict_modeling_res['scores'] )) :\n                df_report_all_steps.loc[IX,'score best '+str(ii) ] = list_best_scores[ii]\n            df_report_all_steps.loc[IX,'Full Prm' ] = str( main_config_model_feature_etc ) \n                \n        if flag_achieved_better_result:\n            main_config_model_feature_etc[prm_name] = best_prm\n        else: \n            # If all prms from list gave worse results than previous - restore previous param \n            main_config_model_feature_etc[prm_name] = save_previous_prm_value\n\n\n            \nwarnings.filterwarnings(\"default\") # To restore all warnings\n            \nprint()\nprint('Final best scores:')\nprint( np.round(list_best_scores, 4), 'average:',  np.round(np.mean( list_best_scores) , 4) )\nprint('Final config:')            \nprint(main_config_model_feature_etc)            ","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:42:32.523214Z","iopub.execute_input":"2023-11-01T10:42:32.523578Z","iopub.status.idle":"2023-11-01T10:43:20.387873Z","shell.execute_reply.started":"2023-11-01T10:42:32.523549Z","shell.execute_reply":"2023-11-01T10:43:20.387163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize Tuning Process","metadata":{}},{"cell_type":"code","source":"%%time\nimport datetime\nstr_inf_cfg = get_brief_string_info_on_config(main_config_model_feature_etc)\nprint( str_inf_cfg )\n\ncurrent_datetime = datetime.datetime.now()\nstr_datetime =  str(current_datetime)[5:16].replace(' ','-').replace(':','-') \n\nfn = 'gentle_tuner_report_all_steps_'+str_inf_cfg+'_' + str_datetime + '.csv'\nprint(fn)\ndf_report_all_steps.to_csv(fn)\n\nfig = plt.figure(figsize = (20,5) )\nfor col in df_report_all_steps.columns:\n    if 'score average' not in col: continue\n    plt.plot( df_report_all_steps[col].values,  'o-', label = col )\n\nN = len(df_report_all_steps)\nx_tick_labels = [ (df_report_all_steps['prm name'].iat[i] + ' ' + str( df_report_all_steps['prm value'].iat[i]) ) for i in range(N) ]\nplt.xticks(range( N ), x_tick_labels, rotation=90)\nplt.grid()\nplt.title(str_inf_cfg + ' Gentle Tuner', fontsize = 20  )\nplt.legend()\nplt.show()\n\nfig = plt.figure(figsize = (20,8) )\nfor col in df_report_all_steps.columns:\n    if 'score' not in col: continue\n    plt.plot( df_report_all_steps[col].values, 'o-', label = col )\nplt.xticks(range( N ), x_tick_labels, rotation=90)\nplt.grid()\nplt.title(str_inf_cfg + ' Gentle Tuner' , fontsize = 20  )\nplt.legend()\nplt.show()\n\n\n","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:43:32.497633Z","iopub.execute_input":"2023-11-01T10:43:32.497984Z","iopub.status.idle":"2023-11-01T10:43:33.134095Z","shell.execute_reply.started":"2023-11-01T10:43:32.497954Z","shell.execute_reply":"2023-11-01T10:43:33.132833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport datetime\n\nif len(df_report_accepted_params) > 0:\n    df_report_loc = df_report_accepted_params\n    \n    str_inf_cfg = get_brief_string_info_on_config(main_config_model_feature_etc)\n    print( str_inf_cfg )\n\n    current_datetime = datetime.datetime.now()\n    str_datetime =  str(current_datetime)[5:16].replace(' ','-').replace(':','-') \n\n    fn = 'gentle_tuner_report_accepted_params_'+str_inf_cfg+'_' + str_datetime + '.csv'\n    print(fn)\n    df_report_loc.to_csv(fn)\n\n    fig = plt.figure(figsize = (20,5) )\n    for col in df_report_loc.columns:\n        if 'score average' not in col: continue\n        plt.plot( df_report_loc[col].values,  'o-', label = col )\n\n    N = len(df_report_loc)\n    x_tick_labels = [ (df_report_loc['prm name'].iat[i] + ' ' + str( df_report_loc['prm value'].iat[i]) ) for i in range(N) ]\n    plt.xticks(range( N ), x_tick_labels, rotation=90)\n    plt.grid()\n    plt.title(str_inf_cfg + ' Gentle Tuner', fontsize = 20  )\n    plt.legend()\n    plt.show()\n\n    fig = plt.figure(figsize = (20,8) )\n    for col in df_report_loc.columns:\n        if 'score' not in col: continue\n        plt.plot( df_report_loc[col].values, 'o-', label = col )\n    plt.xticks(range( N ), x_tick_labels, rotation=90)\n    plt.grid()\n    plt.title(str_inf_cfg + ' Gentle Tuner' , fontsize = 20  )\n    plt.legend()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:46:38.374244Z","iopub.execute_input":"2023-11-01T10:46:38.374638Z","iopub.status.idle":"2023-11-01T10:46:38.384827Z","shell.execute_reply.started":"2023-11-01T10:46:38.37461Z","shell.execute_reply":"2023-11-01T10:46:38.384003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare submits","metadata":{}},{"cell_type":"markdown","source":"## Prepare \"priors\"\n\nThe trick typically uplifts the score about 0.01\n\nTrick to add \"priors\" (aggregates over the cell type and compound - which typically improves scores by 0.01 ): ( df_submit = (1-w_compound - w_cell_type)(df_submit) + w_compound df_submit_aggr_compound + w_cell_type * df_submit_aggr_cell_type )\n\nLIUDA CHELDIEVA: https://www.kaggle.com/code/liudacheldieva/streamlined-baseline-approach\n\nZXMKCD : https://www.kaggle.com/code/zmcxjt/streamlined-baseline-approach","metadata":{}},{"cell_type":"code","source":"%%time\ndef get_submit_priors():\n    #group by drug and take the mean\n    df_tmp = df_de_train.iloc[:, [1] + list(range(5, df_de_train.shape[1]))] # Take only numeric columns and \"sm_name\"\n    df_aggr = df_tmp.groupby('sm_name').mean().reset_index()\n    print(df_aggr.shape)\n    df_submit_aggr_compound = pd.merge( df_id_map.reset_index(),  df_aggr, on='sm_name', how = 'left' ).sort_values('id').drop(columns = ['cell_type', 'sm_name']).set_index('id')\n    print(df_submit_aggr_compound.shape)\n    display(df_submit_aggr_compound.head(3))\n\n    print( )\n\n    df_tmp = df_de_train.iloc[:, [0] + list(range(5, df_de_train.shape[1]))] # Take only numeric columns and \"cell_type\"\n    df_aggr = df_tmp.groupby('cell_type').mean().reset_index()\n    print(df_aggr.shape)\n    df_submit_aggr_cell_type = pd.merge( df_id_map.reset_index(),  df_aggr, on='cell_type', how = 'left' ).sort_values('id').drop(columns = ['cell_type', 'sm_name']).set_index('id')\n    print(df_submit_aggr_cell_type.shape)\n    display(df_submit_aggr_cell_type.head(3))\n    \n    return df_submit_aggr_cell_type, df_submit_aggr_compound\n\ndf_submit_aggr_cell_type, df_submit_aggr_compound = get_submit_priors()\n","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:09:42.858527Z","iopub.execute_input":"2023-11-01T10:09:42.860022Z","iopub.status.idle":"2023-11-01T10:09:43.326268Z","shell.execute_reply.started":"2023-11-01T10:09:42.859989Z","shell.execute_reply":"2023-11-01T10:09:43.325301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Actual submits prepare and save ","metadata":{}},{"cell_type":"code","source":"%%time\nflag_save_submits_loc = False\n\nstr_inf_cfg = get_brief_string_info_on_config(main_config_model_feature_etc)\nprint( str_inf_cfg )\n\ndict_modeling_res = do_modeling( main_config_model_feature_etc, list_CV_scheme = ['Full'] )\nY_submit_pred = dict_modeling_res['Full']['Y_submit'] \ndf_submit = pd.DataFrame(Y_submit_pred, columns = df_de_train.columns[5:])\n\nfn_save = 'submission_' + str_inf_cfg +'_tuned.csv' \nprint( df_submit.shape )\ndisplay(df_submit.head(2))\nif flag_save_submits_loc:\n    df_submit.to_csv(fn_save )            \n\nw_cell_type = 0.1\nw_compound = 0.45\ndf_submit_2 = (1-w_compound - w_cell_type)*(df_submit) + w_compound * df_submit_aggr_compound +  w_cell_type * df_submit_aggr_cell_type\nfn_save2 = 'submission_' + str_inf_cfg +'_blendWithPriors_'+str(w_compound)+'_'+str( w_cell_type )+'_tuned.csv' \nprint(fn_save2)\nif flag_save_submits_loc:\n    df_submit_2.to_csv(fn_save2)\nprint(df_submit_2.shape)\ndisplay(df_submit_2.head(3) )","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:09:43.327778Z","iopub.execute_input":"2023-11-01T10:09:43.328639Z","iopub.status.idle":"2023-11-01T10:09:58.55993Z","shell.execute_reply.started":"2023-11-01T10:09:43.328588Z","shell.execute_reply":"2023-11-01T10:09:58.558704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Wrappers to functions - reporting and submit prepare","metadata":{}},{"cell_type":"markdown","source":"## Report for tuner ","metadata":{}},{"cell_type":"code","source":"%%time\nimport datetime\n\ndef report_for_tuner( df_report_all_steps , df_report_accepted_params, str_postfix = '' ):\n    '''\n    Create plots on score changing during the tuning.\n    Save reports to csv\n    '''\n    str_inf_cfg = get_brief_string_info_on_config(main_config_model_feature_etc)\n    print( str_inf_cfg )\n\n    current_datetime = datetime.datetime.now()\n    str_datetime =  str(current_datetime)[5:16].replace(' ','-').replace(':','-') \n\n    fn = 'gentle_tuner_report_all_steps_'+str_inf_cfg+'_' + str_datetime + str_postfix +'.csv'\n    print(fn)\n    df_report_all_steps.to_csv(fn)\n\n    fig = plt.figure(figsize = (20,5) )\n    for col in df_report_all_steps.columns:\n        if 'score average' not in col: continue\n        plt.plot( df_report_all_steps[col].values,  'o-', label = col )\n\n    N = len(df_report_all_steps)\n    x_tick_labels = [ (df_report_all_steps['prm name'].iat[i] + ' ' + str( df_report_all_steps['prm value'].iat[i]) ) for i in range(N) ]\n    plt.xticks(range( N ), x_tick_labels, rotation=90)\n    plt.grid()\n    plt.title(str_inf_cfg + ' Gentle Tuner'  + ' ' + str_postfix , fontsize = 20  )\n    plt.legend()\n    plt.show()\n\n    fig = plt.figure(figsize = (20,8) )\n    for col in df_report_all_steps.columns:\n        if 'score' not in col: continue\n        plt.plot( df_report_all_steps[col].values, 'o-', label = col )\n    plt.xticks(range( N ), x_tick_labels, rotation=90)\n    plt.grid()\n    plt.title(str_inf_cfg + ' Gentle Tuner'  + ' ' + str_postfix, fontsize = 20  )\n    plt.legend()\n    plt.show()\n\n    if len(df_report_accepted_params) > 0:\n        df_report_loc = df_report_accepted_params\n\n        str_inf_cfg = get_brief_string_info_on_config(main_config_model_feature_etc)\n        print( str_inf_cfg )\n\n        current_datetime = datetime.datetime.now()\n        str_datetime =  str(current_datetime)[5:16].replace(' ','-').replace(':','-') \n\n        fn = 'gentle_tuner_report_accepted_params_'+str_inf_cfg+'_' + str_datetime +  str_postfix + '.csv'\n        print(fn)\n        df_report_loc.to_csv(fn)\n\n        fig = plt.figure(figsize = (20,5) )\n        for col in df_report_loc.columns:\n            if 'score average' not in col: continue\n            plt.plot( df_report_loc[col].values,  'o-', label = col )\n\n        N = len(df_report_loc)\n        x_tick_labels = [ (df_report_loc['prm name'].iat[i] + ' ' + str( df_report_loc['prm value'].iat[i]) ) for i in range(N) ]\n        plt.xticks(range( N ), x_tick_labels, rotation=90)\n        plt.grid()\n        plt.title(str_inf_cfg + ' Gentle Tuner  Uplifts only' + ' ' + str_postfix, fontsize = 20  )\n        plt.legend()\n        plt.show()\n\n        fig = plt.figure(figsize = (20,8) )\n        for col in df_report_loc.columns:\n            if 'score' not in col: continue\n            plt.plot( df_report_loc[col].values, 'o-', label = col )\n        plt.xticks(range( N ), x_tick_labels, rotation=90)\n        plt.grid()\n        plt.title(str_inf_cfg + ' Gentle Tuner Uplifts only' + ' ' + str_postfix , fontsize = 20  )\n        plt.legend()\n        plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:09:58.56108Z","iopub.execute_input":"2023-11-01T10:09:58.561344Z","iopub.status.idle":"2023-11-01T10:09:58.577606Z","shell.execute_reply.started":"2023-11-01T10:09:58.561323Z","shell.execute_reply":"2023-11-01T10:09:58.576245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submit prepare","metadata":{}},{"cell_type":"code","source":"%%time\n\ndf_submit_aggr_cell_type, df_submit_aggr_compound = get_submit_priors()\n\ndef prepare_submits(main_config_model_feature_etc, str_postfix = '' ):\n    '''\n    Prepare and save submits to csv.\n    Direct and blended with \"priors\"\n    '''\n    \n    str_inf_cfg = get_brief_string_info_on_config(main_config_model_feature_etc)\n    print( str_inf_cfg )\n\n    dict_modeling_res = do_modeling( main_config_model_feature_etc, list_CV_scheme = ['Full'] )\n    Y_submit_pred = dict_modeling_res['Full']['Y_submit'] \n    df_submit = pd.DataFrame(Y_submit_pred, columns = df_de_train.columns[5:])\n    df_submit.index.name = 'id'\n    fn_save = 'submission_' + str_inf_cfg + str_postfix +  str_postfix + '_tuned.csv' \n    print( df_submit.shape )\n    display(df_submit.head(2))\n    df_submit.to_csv(fn_save )            \n\n    w_cell_type = 0.1\n    w_compound = 0.45\n    df_submit_2 = (1-w_compound - w_cell_type)*(df_submit) + w_compound * df_submit_aggr_compound +  w_cell_type * df_submit_aggr_cell_type\n    df_submit_2.index.name = 'id'\n    fn_save2 = 'submission_' + str_inf_cfg +'_blendWithPriors_'+str(w_compound)+'_'+str( w_cell_type )+ str_postfix +'_tuned.csv' \n    df_submit_2.to_csv(fn_save2)\n    print(fn_save2)\n    print(df_submit_2.shape)\n    display(df_submit_2.head(3) )","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:09:58.579025Z","iopub.execute_input":"2023-11-01T10:09:58.57943Z","iopub.status.idle":"2023-11-01T10:09:59.032003Z","shell.execute_reply.started":"2023-11-01T10:09:58.579396Z","shell.execute_reply":"2023-11-01T10:09:59.030448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Gentle Tuner Wrapper to Function\n\nHere we wrap into a function the code from tuner code (see section \"Gentle param tuner. Example 1\") \n\n\nThe idea of the tuner is quite simple. \nWe define possible sets of values for each parameter, tuner goes through each set one by one taking the best value. \nSuch process would rarely lead to a global extremum, but our purpose is different - change params gently, by the reasons explained in the introduction.\n\nSo technically:\nwe create dictionary \"dict_prms_lists\" which contains the sets of variation for all params, like that:\n\n    dict_prms_lists = {'alpha':[1e4,1e5,1e6],  'quantile':[0.5,0.6], 'm': [100,1000,10_000]}\n\n    Loop over param names e.g. keys of that dictionary\n\n    Loop over variation set for param chosen at the loop above. \n    \n    That is basically all.\n    \n    Even more technically:\n    There is outer loop which allows to repeat the entire process several times (it is rarely necessary - that is a bit surprising ).\n    The key computation is wrapped in the function \n    do_modeling( main_config_model_feature_etc, list_CV_scheme = list_CV_scheme,  scoring_mode = scoring_mode, verbose = 0 )\n    \n    The function return VECTOR(!) of scores (can be different folds or different CV-schemes) \n    the parameter is beeing accepted as new optimal if ALL components of the vector are better than the previous values.\n    (One can encode easily code other conditions).\n    \n    \"One config to rule them all\"\n    The main point is that it takes the config main_config_model_feature_etc which describes all: model+encoder+reducer params in the standartized way,\n    such that we do not need to change anything in the tuner except specifying the config to add new models/encoders/etc...","metadata":{}},{"cell_type":"code","source":"%%time\n\n\nimport warnings\n\ndef gentle_tuner(main_config_model_feature_etc_initial, dict_prms_lists, dict_tuner_params, str_postfix = '', flag_save_submits = True, verbose = 0 ):\n    warnings.filterwarnings(\"ignore\")     # To suppress all warnings\n    \n    dict_tuner_results  = {} # Output\n\n    list_CV_scheme = dict_tuner_params['list_CV_scheme']# = ['MT'] # ['AmbrosM','MT','Random']\n    scoring_mode = dict_tuner_params['scoring_mode']# = 'return_average_over_folds' # 'return_all_folds_results' # \n        # Controls do we want to find scores uplift over ALL folds , or only average over folds \n    n_rounds = dict_tuner_params['n_rounds']# = 1\n\n    main_config_model_feature_etc = main_config_model_feature_etc_initial\n    \n    # Get scores for initial params:\n    dict_modeling_res = do_modeling( main_config_model_feature_etc, list_CV_scheme = list_CV_scheme,  scoring_mode = scoring_mode, verbose = 0 )\n    list_best_scores = dict_modeling_res['scores']\n    if verbose >= 10:\n        print('Initial list_best_scores:', np.round( list_best_scores, 4) , 'average:',  np.round(np.mean( list_best_scores) , 4), )\n    \n\n    ### --------------------------- Starting tuner \n    t00 = time.time() ; i_total_count = 0\n    list_best_scores_afer_prm_optimization = []\n    df_report_accepted_params = pd.DataFrame(); \n    df_report_all_steps = pd.DataFrame(); \n\n    for i_round in range(n_rounds): # rounds to repeat the entire search \n        for i_prm in range( len( dict_prms_lists )) : # Choose what param will be  tuned\n\n            prm_name = list(dict_prms_lists.keys())[i_prm]\n            list_prm_values = dict_prms_lists[prm_name] \n            flag_achieved_better_result = False \n            save_previous_prm_value = main_config_model_feature_etc[prm_name]\n\n            for prm_value in list_prm_values: # Grid the list of the params \n                i_total_count += 1\n                t0 = time.time()\n\n                main_config_model_feature_etc[prm_name]  = prm_value\n\n                dict_modeling_res = do_modeling( main_config_model_feature_etc, list_CV_scheme = list_CV_scheme,  scoring_mode = scoring_mode, verbose = 0 )\n\n                flag_accept_prm =  np.all(np.asarray(dict_modeling_res['scores']) <  np.asarray(list_best_scores)   ) \n\n                if flag_accept_prm:\n                    list_best_scores = dict_modeling_res['scores']\n                    best_prm = prm_value\n                    flag_achieved_better_result = True\n                    # Report prepare:\n                    IX = len(df_report_accepted_params)+1\n                    df_report_accepted_params.loc[IX,'prm name' ] = prm_name\n                    df_report_accepted_params.loc[IX,'prm value' ] = prm_value\n                    df_report_accepted_params.loc[IX,'trial' ] = i_total_count\n                    for ii in range(len( list_best_scores )) :\n                        df_report_accepted_params.loc[IX,'score '+str(ii) ] = list_best_scores[ii]\n                    df_report_accepted_params.loc[IX,'score average' ] = np.mean( list_best_scores ) \n                    df_report_accepted_params.loc[IX,'Full Prm' ] = str( main_config_model_feature_etc ) \n\n\n                if verbose >= 100:\n                    print('trial:',i_total_count,  'prm_name', prm_name, 'prm_value', prm_value,'scores:', np.round(dict_modeling_res['scores'], 4), \n                          'average:',  np.round(np.mean( dict_modeling_res['scores']) , 4),\n                         'best scores:', np.round(list_best_scores, 4), 'average:',  np.round(np.mean( list_best_scores) , 4),\n                          'flag_accept_prm', flag_accept_prm ,'time:', np.round( time.time()-t0, 1), np.round( time.time()-t00, 1)  ) \n\n                # Report prepare:\n                IX = len(df_report_all_steps)+1\n                df_report_all_steps.loc[IX,'prm name' ] = prm_name\n                df_report_all_steps.loc[IX,'prm value' ] = prm_value\n                df_report_all_steps.loc[IX,'trial' ] = i_total_count\n                df_report_all_steps.loc[IX,'score average' ] = np.mean( dict_modeling_res['scores'] ) \n                df_report_all_steps.loc[IX,'score average best' ] = np.mean( list_best_scores ) \n                for ii in range(len( dict_modeling_res['scores'] )) :\n                    df_report_all_steps.loc[IX,'score '+str(ii) ] = dict_modeling_res['scores'][ii]\n                for ii in range(len( dict_modeling_res['scores'] )) :\n                    df_report_all_steps.loc[IX,'score best '+str(ii) ] = list_best_scores[ii]\n                df_report_all_steps.loc[IX,'Full Prm' ] = str( main_config_model_feature_etc ) \n                df_report_all_steps.loc[IX,'Time' ] = time.time()-t0\n                df_report_all_steps.loc[IX,'Time Full' ] = time.time()-t00\n\n\n            if flag_achieved_better_result:\n                main_config_model_feature_etc[prm_name] = best_prm\n            else: \n                # If all prms from list gave worse results than previous - restore previous param \n                main_config_model_feature_etc[prm_name] = save_previous_prm_value\n\n    warnings.filterwarnings(\"default\") # To restore all warnings\n\n    if verbose >= 1:\n        print()\n        print('Final best scores:')\n        print( np.round(list_best_scores, 4), 'average:',  np.round(np.mean( list_best_scores) , 4) )\n        print('Final config:')            \n        print(main_config_model_feature_etc)      \n\n    dict_tuner_results['best_params'] = main_config_model_feature_etc\n    dict_tuner_results['list_best_scores'] = list_best_scores\n    dict_tuner_results['df_report_all_steps'] = df_report_all_steps\n    dict_tuner_results['df_report_accepted_params'] = df_report_accepted_params\n    \n    report_for_tuner( df_report_all_steps , df_report_accepted_params, str_postfix = str_postfix ) \n    if flag_save_submits:\n        prepare_submits(main_config_model_feature_etc, str_postfix = str_postfix )\n        \n    return dict_tuner_results\n\n\n# --------------------- Simple Example  --------------------------\n\n     \n# ---------------------------------------- Key  Params --------------------------------------------------------------\ndict_tuner_params = {}\ndict_tuner_params['list_CV_scheme'] = ['MT'] # ['AmbrosM','MT','Random']\ndict_tuner_params['scoring_mode'] = 'return_average_over_folds' # 'return_all_folds_results' # \n    # Controls do we want to find scores uplift over ALL folds , or only average over folds \ndict_tuner_params['n_rounds'] = 1\n\n### --------------------------- Initial config (config describes ALL(!) - model, encoder, all params for model and encoder)\nmain_config_model_feature_etc = {}\nmain_config_model_feature_etc['model'] =  'KRR'# 'Ridge'\nmain_config_model_feature_etc['kernel'] =  'rbf' # 'linear'\nmain_config_model_feature_etc['alpha'] = 1e4\nmain_config_model_feature_etc['reducer'] = reducer_main #  'tsvd'# 'pca' # 'ica'\nmain_config_model_feature_etc['n_components'] = n_components_main\nmain_config_model_feature_etc['encoder'] = 'QuantileEncoder'\nmain_config_model_feature_etc['quantile'] = 0.5\nmain_config_model_feature_etc['m'] = 1\n### --------------------------- Dictionary with optimized params and their intervals \n# dict_prms_lists = {'alpha':[1e-3, 1e-2, 1e-1,1,1e1,1e2,1e3, 1e4,1e5,1e6],  'quantile':[0.5,0.95], 'm':[1,10, 100,1000,10_000]}\n# dict_prms_lists = {'alpha':[0.007, 0.08, 0.09,0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8],  'quantile':[0.3, 0.5,0.7], 'm':[10,20,30, 40,50,60, 70,80,90, 100,150,200,300,400,500]}\ndict_prms_lists = {'alpha':[.5,1],  'quantile':[0.5,], 'm':[10, 100]}\n    \n    \ndict_tuner_results = gentle_tuner(main_config_model_feature_etc, dict_prms_lists, dict_tuner_params, str_postfix = '_test', flag_save_submits = False, verbose = 100 )        ","metadata":{"execution":{"iopub.status.busy":"2023-11-01T10:09:59.033411Z","iopub.execute_input":"2023-11-01T10:09:59.033754Z","iopub.status.idle":"2023-11-01T10:11:00.399119Z","shell.execute_reply.started":"2023-11-01T10:09:59.033723Z","shell.execute_reply":"2023-11-01T10:11:00.397692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Process many models and encoders ","metadata":{}},{"cell_type":"code","source":"%%time \nimport warnings\nwarnings.filterwarnings(\"ignore\")     # To suppress all warnings\n# 8 Encoders: 'QuantileEncoder','TargetEncoder', 'CatBoostEncoder', 'LeaveOneOutEncoder' , 'JamesSteinEncoder', 'OneHotEncoder', 'HelmertEncoder', 'BackwardDifferenceEncoder'\n\nlist_alpha = [0.0001,0.0002,.0005, 0.001,0.002,.005, 0.01,0.02,.05, 0.1,0.2,.5,1,2,5,10,20,30,50,100,200,500,1000,2000,3000,5000, 1e4,2e4,5e4, 1e5,2e5,5e5, 1e6,2e6,5e6, 1e7,2e7,5e7,1e8]\nprint(len(list_alpha ))\nlist_C_for_SVR = [0.00001,0.00002,.00005,0.0001,0.0002,.0005, 0.001,0.002,.005, 0.01,0.02,.05, 0.1,0.2,.5,1,2,5,10,20,30,50,100,200,500,1000,2000,3000,5000, 1e4,2e4,5e4, 1e5]\nprint(len(list_C_for_SVR ))\n\ndf_grand_report = pd.DataFrame()\nfor model_id in ['Ridge']: # ['RFR']: #  ['CATB']:#  ['LGB']:# ['KRRlin']: #  ['SVRrbf','SVRlin', ]: #['Ridge' ,'KRRlin','KRRrbf']\n    # Some rough estimate of optimal alpha for different encoders and Ridge/KRR models (AmrbsoM CV)\n    # Conclusion: min alpha = 1;   max alpha = 1e7\n    # Remark: typically KRRrbf requires much smaller alpha, than KRR-linear and Ridge, among these two - the one: KRR-linear requires typically larger alpha \n    # HelmertEncoder: Ridge - 1e5 ; KRRlin - 1e7; KRRrbf - 10 \n    # OnehotEncoder:  Ridge - 100  score: 0.99; KRRlin - 100 score: 0.96; KRRrbf - 1000 score: 0.99, \n    # BackwardDifferenceEncoder + Ridge->10 score: 1.02, +KRRlin - same!? 10 score: 1.02, KRRrbf 1000 score:  0.9966; \n    # LeaveOneOutEncoder  +Ridge 100_000.0 scores: [0.9746] ; +KRRlin 1000_000 score: 0.963; +KRRrbf -  1  score: 1.0051 \n    # CatBoostEncoder  +Ridge 1000000.0 scores: [0.9897] ; +KRRlin  10_000_000.0 scores: [0.9744] average ; +KRRrbf  1 scores: [1.0048] \n    # TargetEncoder  +Ridge 100000.0 scores: [0.9942] ; +KRRlin  10000000.0 scores: [0.9978] ; +KRRrbf   10 scores: [0.9798]\n    # TE smoothing: Ridge 10 scores: [0.9942], KRRlin 1000 scores: [0.9853];  KRRrbf 50 scores: [0.9733] \n    # QuantileEncoder: Ridge 1000_000 scores: [0.9837] ,  KRRlin 10_000_000.0 scores: [0.9802]; KRRrbf 10 scores: [0.9736] \n    # KRRrbf: quantile - very small effect, \"m\" - around 1 is good\n    # Ridge - quantile - 0.5-0.6 - good,  m - around 1 is good \n    # Addtional analysis show that for MT scheme for three encoders ( Onehot, Backward,LOO) and KRRrbf we get optimal alpha = 0.01 - the boundary for initial interval.\n    # Manually checking these case we see that it is better to extent alpha interval more : \n    #\n    # SVR \n    # Note(!) - it works timing typically quite bigger than KernelRidge , especially for high C (low regularization) it becomes very slow (probably poor convergence)\n    # Rough estimate for optimal \"C\"\n    # Quantile Encoder SVRlin 0.0001 scores: [0.9807], SVRrbf 10 scores: [0.9719] (AmbrosM)\n    # OHE: SVRrbf C=10 scores: [0.9786] average:  201.8 - timing is quite bigger than for other encoders (may be because we have many features )  (AmbrosM)\n    # OHE: SVRlin: C prm_value 1 scores: [0.9799] average: 0.9799 best scores: [0.9799] average: 0.9799 flag_accept_prm True time: 155.8  (AmbrosM)\n    # MT-CV:  C prm_value 100 scores: [2.8194] average: 2.8194 best scores: [2.8194] average: 2.8194 flag_accept_prm True time: 145.6\n    # SVRrbf, Backward: C prm_value 20000.0 scores: [2.8109]  time: 65.2 \n    \n    for encoder_id in ['LeaveOneOutEncoder']: # , 'OneHotEncoder', 'HelmertEncoder', 'BackwardDifferenceEncoder']:#'TargetEncoder', 'CatBoostEncoder', 'LeaveOneOutEncoder' , 'JamesSteinEncoder', 'OneHotEncoder', 'HelmertEncoder', 'BackwardDifferenceEncoder' ]:\n        # 8 Encoders: 'QuantileEncoder','TargetEncoder', 'CatBoostEncoder', 'LeaveOneOutEncoder' , 'JamesSteinEncoder', 'OneHotEncoder', 'HelmertEncoder', 'BackwardDifferenceEncoder'\n        for CV_scheme in ['Tonya' , 'AmbrosM' ,'MT' , 'Random']:# , 'AmbrosM' ,'MT' , 'Random']:\n            dict_tuner_params['list_CV_scheme'] = [ CV_scheme  ]\n            \n            # Initial params:\n            main_config_model_feature_etc = {} # Initial params \n            main_config_model_feature_etc['reducer'] = reducer_main #  'tsvd'# 'pca' # 'ica'\n            main_config_model_feature_etc['n_components'] = n_components_main\n\n            main_config_model_feature_etc['model'] = model_id # 'KRR'#  'KRRlin','KRRrbf' 'Ridge'\n            main_config_model_feature_etc['encoder'] = encoder_id\n \n\n            dict_prms_lists = {}\n            dict_tuner_params['scoring_mode'] = 'return_average_over_folds' # 'return_all_folds_results' # \n                # Controls do we want to find scores uplift over ALL folds , or only average over folds \n            dict_tuner_params['n_rounds'] = 1\n            if ( 'Ridge' in model_id) or ('KRR' in model_id): \n                dict_prms_lists['alpha'] = list_alpha\n                main_config_model_feature_etc['alpha'] = 1e4\n            elif ('SVR' in model_id):\n                dict_prms_lists['C'] = list_C_for_SVR\n                main_config_model_feature_etc['C'] = 0.0001 # Default C\n                main_config_model_feature_etc['max_iter'] = 32000# Number of iterations for SVR, is some cases \"-1\" option (no limit) causes super large timing\n            elif ('CATB' in model_id):\n                main_config_model_feature_etc['iterations'] = 250\n                main_config_model_feature_etc['depth'] = 6\n                main_config_model_feature_etc['learning_rate'] = 0.03\n                main_config_model_feature_etc['subsample'] = 1\n                main_config_model_feature_etc['colsample_bylevel'] = 0.5\n                main_config_model_feature_etc['min_data_in_leaf'] = 5\n                main_config_model_feature_etc['loss_function'] ='RMSE'\n                main_config_model_feature_etc['random_strength'] = None\n                main_config_model_feature_etc['random_seed'] = 42# Alias - random_sate. Default - None.   main_config_model_feature_etc['random_state'] = 42\n                main_config_model_feature_etc['l2_leaf_reg'] = 1 # 3.0 # Any positive value is allowed. Default 3. https://catboost.ai/en/docs/references/training-parameters/common#l2_leaf_reg\n                #dict_prms_lists['l2_leaf_reg' ]  = [0.1,1,3,10,100]\n                #dict_prms_lists['loss_function'] = ['Huber:delta=0.5', \n                #                                    'Huber:delta=1', 'Huber:delta=2', 'Huber:delta=10', 'Huber:delta=100']\n                # ,['RMSE','MAE','MAPE']# ,'Huber'] # default is 6\n                #dict_prms_lists['iterations'] =[100,150,200,250,300,350,400,450,500,550,600] # default - 1000\n                #dict_prms_lists['depth'] =[2,3,4,5,6,7,8,9]# [2,3,4,5,6,7,8,9,10] # default is 6\n                #dict_prms_lists['learning_rate'] =[0.035, 0.03,0.025 ] # default 0.03\n                # dict_prms_lists['subsample'] =[0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8, 1] # default is 1 \n                #dict_prms_lists['colsample_bylevel'] =[0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8, 1] # default is 1 \n                #dict_prms_lists['random_strength'] = [0.1,0.3,0.5,0.8,1,2,5,10,100,1e3]\n                #dict_prms_lists['min_data_in_leaf'] = [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,20,30]\n                #dict_prms_lists['min_data_in_leaf'] =[5] # default is 1 - for bin class, 5 for regres/multi-targ \n            elif ('LGB' in model_id):\n                main_config_model_feature_etc['n_estimators'] = 100\n                main_config_model_feature_etc['depth'] = -1\n                main_config_model_feature_etc['learning_rate'] = 0.1\n                dict_prms_lists['n_estimators'] =[50,100,1000] # default - 100\n                dict_prms_lists['depth'] =[-1 , 2, 6] # default \"-1\" (unlimited)\n                dict_prms_lists['learning_rate'] =[0.1,0.001] # default 0.1\n                # Default for regressor:  can check it by: model = lgb.LGBMRegressor(); print( model.get_params() )\n                # {'boosting_type': 'gbdt',\n                # 'class_weight': None,\n                # 'colsample_bytree': 1.0,\n                # 'importance_type': 'split',\n                # 'learning_rate': 0.1,\n                # 'max_depth': -1,\n                # 'min_child_samples': 20,\n                # 'min_child_weight': 0.001,\n                # 'min_split_gain': 0.0,\n                # 'n_estimators': 100,\n                # 'n_jobs': -1,\n                # 'num_leaves': 31,\n                # 'objective': None,\n                # 'random_state': None,\n                # 'reg_alpha': 0.0,\n                # 'reg_lambda': 0.0,\n                # 'silent': 'warn',\n                # 'subsample': 1.0,\n                # 'subsample_for_bin': 200000,\n                # 'subsample_freq': 0 }\n\n            elif ('RFR' in model_id):\n                # RandomForestRegressor\n                main_config_model_feature_etc['n_estimators'] = 500\n                main_config_model_feature_etc['max_depth'] = None\n                main_config_model_feature_etc['criterion'] = 'squared_error' # 'absolute_error' #  “absolute_error”, “friedman_mse”, “poisson”\n                main_config_model_feature_etc['max_features'] = 1 # # {“sqrt”, “log2”, None}, int or float, default=1.0\n                #dict_prms_lists['n_estimators'] =[50,100, 200,1000] # default - 100\n                #dict_prms_lists['max_depth'] =[None , 2, 6,10 ] # default \"-1\" (unlimited)\n                #dict_prms_lists['max_features'] =['sqrt','log2',None, 0.1,0.2,0.3]#,0.4,0.5,0.6,0.7,0.8,0.9,1 ] # {“sqrt”, “log2”, None}, int or float, default=1.0\n                #n_estimators=100, *, criterion='squared_error', max_depth=None, min_samples_split=2, min_samples_leaf=1, min_weight_fraction_leaf=0.0, max_features=1.0, max_leaf_nodes=None, min_impurity_decrease=0.0, bootstrap=True, oob_score=False, n_jobs=None, random_state=None, verbose=0, warm_start=False, ccp_alpha=0.0, max_samples=None\n            elif ('ETR' in model_id):\n                # ExtraTreesRegressor\n                main_config_model_feature_etc['n_estimators'] = 100\n                main_config_model_feature_etc['max_depth'] = None\n                main_config_model_feature_etc['criterion'] = 'absolute_error' # “squared_error”, “absolute_error”, “friedman_mse”, “poisson”\n                dict_prms_lists['n_estimators'] =[50,100, 200,500,1000] # default - 100\n                dict_prms_lists['max_depth'] =[None , 2, 6,10,12, 15 ] # default \"-1\" (unlimited)\n                # n_estimators=100, *, criterion='squared_error', max_depth=None, min_samples_split=2, min_samples_leaf=1, min_weight_fraction_leaf=0.0, max_features=1.0, max_leaf_nodes=None, min_impurity_decrease=0.0, bootstrap=False, oob_score=False, n_jobs=None, random_state=None, verbose=0, warm_start=False, ccp_alpha=0.0, max_samples=None\n            \n            \n            if encoder_id == 'QuantileEncoder':\n                dict_prms_lists['quantile'] = [0.1, 0.2,0.3,0.4, 0.5, 0.6,0.7,0.8,0.9 ] # [0.75,0.8,0.85,0.9,0.95 ] #  [0.01,0.1,0.15,0.2,0.25,0.3]#  # 65 tsvd75 CatB Quantile[0.75,0.8,0.85,0.9,0.95,0.99 ] #  [0.2,0.3,0.4, 0.5, 0.6,0.7,0.8 ] 0.01-0.3\n                dict_prms_lists['m'] =  [0,0.5,1,2,5,10,20, 50, 100,1e3]\n                main_config_model_feature_etc['m'] = 1\n                main_config_model_feature_etc['quantile'] = 0.5\n            elif encoder_id == 'TargetEncoder':\n                dict_prms_lists['smoothing'] = [1,5, 10,15, 20,50,70,100,120,150,200, 1000]\n                main_config_model_feature_etc['smoothing'] = 10\n            elif encoder_id == 'CatBoostEncoder':\n                # 'a' - smoothing-like - but higher values - more shrinking to constant\n                dict_prms_lists['a'] = [0.1, 0.5,1,2,3, 5,10,20, 50, 100,1e3]\n                main_config_model_feature_etc['a'] = 1 \n\n        \n            str_inf_cfg = get_brief_string_info_on_config(main_config_model_feature_etc)\n            print()\n            print( model_id , encoder_id, str_inf_cfg,  CV_scheme )\n            #\n            #\n            ## ----------------------------------------------------------------- Launch Tuner -----------------------------------------------\n            #\n            dict_tuner_results = gentle_tuner(main_config_model_feature_etc, dict_prms_lists, dict_tuner_params, str_postfix = CV_scheme,  verbose = 100 )        \n            \n            \n            # Saving data for report \n            df_grand_report.loc[str_inf_cfg, CV_scheme + ' score'  ] = np.mean( dict_tuner_results[ 'list_best_scores' ] )\n            if 'alpha' in dict_tuner_results[ 'best_params' ].keys():\n                df_grand_report.loc[str_inf_cfg, CV_scheme + ' alpha'  ] =  dict_tuner_results[ 'best_params' ]['alpha'] \n            if 'C' in dict_tuner_results[ 'best_params' ].keys():\n                df_grand_report.loc[str_inf_cfg, CV_scheme + ' C'  ] =  dict_tuner_results[ 'best_params' ]['C'] \n            df_grand_report.loc[str_inf_cfg, CV_scheme + ' params'  ] = str( dict_tuner_results[ 'best_params' ])\n            \ndf_grand_report.to_csv('report_multi_model.csv')\nl1 = [t for t in df_grand_report.columns if 'score' in t]\nd_tmp = df_grand_report[l1].rank()\nl1b = [t.replace('score','Rank') for t in d_tmp.columns ]\nd_tmp['Mean Rank'] = d_tmp.mean(axis =1 ).round(1)\nl1b.append('Mean Rank')\nd_tmp.columns = l1b\ndf_grand_report = pd.concat([df_grand_report, d_tmp], axis = 1)\nl2 = [t for t in df_grand_report.columns if 'alpha'  in t]\nl2b = [t for t in df_grand_report.columns if ' C'  in t]\nl3 = [t for t in df_grand_report.columns if 'params'  in t]\ndf_grand_report = df_grand_report[ l1+l1b + l2 +l2b  + l3 ]\nprint(df_grand_report.columns)\ndisplay(df_grand_report[l1].corr().round(2) )\ndisplay(df_grand_report[l1].corr(method = 'spearman').round(2) )\ndisplay(df_grand_report.sort_values( 'Mean Rank' ) )\n\ndf_grand_report.sort_values( 'Mean Rank' ).to_csv('report_multi_model_sorted.csv')","metadata":{"execution":{"iopub.status.busy":"2023-11-01T11:52:03.334566Z","iopub.execute_input":"2023-11-01T11:52:03.335607Z","iopub.status.idle":"2023-11-01T12:05:45.385139Z","shell.execute_reply.started":"2023-11-01T11:52:03.335535Z","shell.execute_reply":"2023-11-01T12:05:45.384036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Main Report","metadata":{}},{"cell_type":"code","source":"%%time\npd.set_option('display.max_colwidth', 200)\n\n\ndisplay(df_grand_report[l1].corr().round(2) )\ndisplay(df_grand_report[l1].corr(method = 'spearman').round(2) )\n\ndisplay(df_grand_report.sort_values( 'Mean Rank' ) )","metadata":{"execution":{"iopub.status.busy":"2023-10-31T21:05:12.19107Z","iopub.execute_input":"2023-10-31T21:05:12.191392Z","iopub.status.idle":"2023-10-31T21:05:12.217857Z","shell.execute_reply.started":"2023-10-31T21:05:12.19137Z","shell.execute_reply":"2023-10-31T21:05:12.216859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Final timing","metadata":{"execution":{"iopub.status.busy":"2023-10-25T10:04:59.25663Z","iopub.execute_input":"2023-10-25T10:04:59.25698Z","iopub.status.idle":"2023-10-25T10:04:59.261681Z","shell.execute_reply.started":"2023-10-25T10:04:59.256955Z","shell.execute_reply":"2023-10-25T10:04:59.260221Z"}}},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )\nprint('%.1f minutes passed total '%( (time.time()-t0start)/60)  )\nprint('%.2f hours passed total '%( (time.time()-t0start)/3600)  )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}