{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59094,"databundleVersionId":6541963,"sourceType":"competition"}],"dockerImageVersionId":30558,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\n#### Brief\n\nWe explore target encoders modeling for \"Open Problems – Single-Cell Perturbations\". The ones: \"TargetEncoder\", \"QuantileEncoder\", \"JamesSteinEncoder\", \"CatBoostEncoder\", 'LeaveOneOutEncoder' from the Python package https://contrib.scikit-learn.org/category_encoders/ . The other 'contrast' encoders (not using the targets) are explored in the previous (more simple)  notebook: https://www.kaggle.com/code/alexandervc/op2-category-encoders-chembert-fingerpints-moldes (may be take a look on it first). \n\nIn short: we make tsvd for targets, encode categorical features by all obtained tsvd components, use Ridge model to predict them, make tsvd.inverse_transform to get original targets. We do it for: all encoders; wide range of the regularization alpha parameter for Ridge; and for several CV-schemes: Random, AmbrosM, MT. Present analysis and plots for all encoders/CV-scores dependence on alpha.  Choose optimal alpha for each encoder+CV-scheme, prepare submissions with the chosen optimal alpha for each encoder. \n\nIntroduction and examples for target encoders is also included - see section: \"Target encoders - demo examples\".\n\n#### Outcomes and conslusions: \n\n    LeaveOneOutEncoder consistently outperforms other encoders being top1 in 5 cases out of 6. \n    Quantile encoder is the second, Target - 3, CatBoost - 4, JamesStein - 5.  \n\n    Encoding compound only shows better results than both - compound + cell type for the present setup.\n    Encoding only cell type - quite poor results. \n    \n    \nNote: encoders considred with default params. Tuning encoder's params to be considered elsewhere (https://www.kaggle.com/alexandervc/op2-gentle-param-tuner )\n    \n#### Ranking encoders by CV(s) - encoding only compound \n\nSummaries with LB/CV scores here: https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing (sheet \"Encoders\") and discussed here: https://docs.google.com/presentation/d/1wiz0Wmt4D54pqMMsIOyJHuQYMZ3hTBZQQnjbLzwoGYY/edit?usp=sharing. Top models can achieve around LB 0.60*. \n\n\nResults from the file \"df_ranking_and_scores_report.csv\" (output tab)\n\n\n| Encoder | Rank Mean | Rank AmbrosM alpha opt for AmbrosM | Rank MT alpha opt for MT | Rank Random alpha opt for Random | mrrmse AmbrosM alpha opt for AmbrosM | mrrmse MT alpha opt for MT | mrrmse Random alpha opt for Random |\n|-----------|-----------|--------------------------------------|-----------------------------|----------------------------------|--------------------------------------|-----------------------------|----------------------------------|\n| LeaveOneOutEncoder Ridge Compound tsvd30 | 1.0 | 1.0 | 1.0 | 1.0 | 0.97213 | 2.53533 | 1.21349 |\n| QuantileEncoder Ridge Compound tsvd30 | 3.0 | 3.0 | 4.0 | 2.0 | 0.98972 | 2.69507 | 1.23648 |\n| TargetEncoder Ridge Compound tsvd30 | 3.3 | 4.0 | 3.0 | 3.0 | 0.99415 | 2.63205 | 1.23705 |\n| CatBoostEncoder Ridge Compound tsvd30 | 3.7 | 2.0 | 5.0 | 4.0 | 0.98823 | 2.87992 | 1.24851 |\n| JamesSteinEncoder Ridge Compound tsvd30 | 4.0 | 5.0 | 2.0 | 5.0 | 1.04416 | 2.59518 | 1.29749 |\n\n\n#### Ranking encoders by CV(s) - encoding both cell type and  compound \n\n\n| Encoder| Rank Mean | Rank AmbrosM alpha opt for AmbrosM | Rank MT alpha opt for MT | Rank Random alpha opt for Random | mrrmse AmbrosM alpha opt for AmbrosM | mrrmse MT alpha opt for MT | mrrmse Random alpha opt for Random |\n|-----------|-----------|-----------------------------------|--------------------------|----------------------------|-----------------------------------|--------------------------|----------------------------|\n| LeaveOneOutEncoder Ridge CompoundCellType tsvd30 | 1.3 | 2.0 | 1.0 | 1.0 | 1.02758 | 2.54844 | 1.21432 |\n| QuantileEncoder Ridge CompoundCellType tsvd30 | 2.0 | 1.0 | 3.0 | 2.0 | 0.98054 | 2.65518 | 1.23221 |\n| TargetEncoder Ridge CompoundCellType tsvd30 | 3.3 | 5.0 | 2.0 | 3.0 | 1.10447 | 2.61953 | 1.24089 |\n| CatBoostEncoder Ridge CompoundCellType tsvd30 | 4.0 | 3.0 | 5.0 | 4.0 | 1.03409 | 2.86031 | 1.24738 |\n| JamesSteinEncoder Ridge CompoundCellType tsvd30 | 4.3 | 4.0 | 4.0 | 5.0 | 1.08256 | 2.75086 | 1.29198 |\n\n\n#### Navigation through the notebook:\n\n    Versions of the notebook correspond changes in the encoding features: both cell-type and compound; only compound; only cell-type.  \n        Version 1: CellType and Compound Encoded\n        Version 2: CellType only Encoded\n        Version 3: Compound only Encoded\n                (LB0.608 - submitted LeaveOneOut+priors - file: submission_LeaveOneOutEncoder_Ridge_Compound_tsvd30_blendWithPriors_0.45_0.1.csv )\n            \n            Version 4: minor changes - supplementary example section added: non-zero \"sigma\" for encoder - add noise to the train part.\n            Verision 5: minor changes - supplementary example section added: change both Ridge alpha and target encoder smoothing.\n                Suprisingly we see better results when smoothing is practically turned off and alpha is increased.\n\n    The plots with dependence of scores on alpha-Ridge-regularization are in the section: \"Plots scores/alphas\" \n    \n    The ranking of the encoders with respect to several CV-schemes scores:  \"Ranking encoders report\"\n    \n    Introduction and examples for target encoders: \"Target encoders - demo examples\".\n    \n    For each encoder we find optimal alpha(s) by CV score(s) and prepare submission files for that encoder + RidgeOptimalAlpha - see output \"tab\" of the notebook \n    \n    Main modelling is placed in the section: \"Modeling several encoders with different alphas\" - it takes about 1 hour. The loops over encoders, CV-schemes and alphas is there. It saves scoring results into dataframes for each combination of the params and these dataframes processed below for visualizations and analysis. The core function to cross-validate is \"compute_cv( model, target_encoder,  list_CV_scheme , verbose = 0 )\" which is placed in the section \"Wrapper CV-computatation function\". \n\n#### Some notes, details: \n\n    Use finer grid for alpha like [1,2,5,10,20,50,100...] - not just [1,10,100] is important, otherwise for some validation schemes optimal alpha appears near 0 or infinity, missing the real optimum which can be seen by finer grid. That cause may cause big descrepancy between optimal values for different CV-schemes. (The effect observed in original notebook by MT - but we observe that this effect mitigates with finer alpha grid, although it is still present).  \n    \n\n\n#### PS\n\nThe related notebooks: custom CV-schemes: https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes, category encoders, chembert, Morgan fingerprints, molecular descriptors: https://www.kaggle.com/code/alexandervc/op2-category-encoders-chembert-fingerpints-moldes, target category encoders: https://www.kaggle.com/alexandervc/op2-target-encoders, \"Gentle\" param tuner: https://www.kaggle.com/alexandervc/op2-gentle-param-tuner, advanced modeling combining all that: https://www.kaggle.com/code/alexandervc/op2-advanced-modeling-tuning-featengineering-etc.\n\n","metadata":{}},{"cell_type":"markdown","source":"# Key params","metadata":{}},{"cell_type":"code","source":"n_components = 30 # Reduction of targets is done by tsvd to n_components, after predicting components we take tsvd.inverse_transform\n\nlist_features_to_encode = ['cell_type','sm_name' ]#  ['sm_name'] # ['cell_type','sm_name'] #  ['cell_type'],  \n# For simple models like Ridge and that kind of features seems encoding only the compounds just forgetting about cell type at all gives sometimes better results. \n\n# encoding = 'TargetEncoder' #  \"QuantileEncoder\", \"JamesSteinEncoder\", \"CatBoostEncoder\", 'LeaveOneOutEncoder'\n","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:02:45.311079Z","iopub.execute_input":"2023-11-24T10:02:45.311921Z","iopub.status.idle":"2023-11-24T10:02:45.319083Z","shell.execute_reply.started":"2023-11-24T10:02:45.311876Z","shell.execute_reply":"2023-11-24T10:02:45.317634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create info string on params:\nstr_inf_cfg = ''# encoding #+ ' '\nif 'sm_name' in list_features_to_encode: str_inf_cfg += 'Compound'\nif 'cell_type' in list_features_to_encode: str_inf_cfg += 'CellType'\nstr_inf_cfg += ' tsvd' + str(n_components)\nprint(str_inf_cfg)","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:02:45.322087Z","iopub.execute_input":"2023-11-24T10:02:45.322863Z","iopub.status.idle":"2023-11-24T10:02:45.335929Z","shell.execute_reply.started":"2023-11-24T10:02:45.322819Z","shell.execute_reply":"2023-11-24T10:02:45.334560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport time\nt0start = time.time() \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport pandas as pd\nimport tensorflow as tf\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:02:45.337269Z","iopub.execute_input":"2023-11-24T10:02:45.337618Z","iopub.status.idle":"2023-11-24T10:02:56.390325Z","shell.execute_reply.started":"2023-11-24T10:02:45.337589Z","shell.execute_reply":"2023-11-24T10:02:56.388999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\ndf_de_train = pd.read_parquet(fn)# , index_col = 0)\nprint(df_de_train.shape)\ndisplay(df_de_train )\n\nplt.figure(figsize = (20,4) )\nv = df_de_train.iloc[:,5:].max(axis = 0 ).sort_values(ascending = False, key = abs )\nplt.plot(v.values,'*-')\nplt.title('Max-abs DE for genes',fontsize = 20 )\nplt.grid()\nplt.show()\ndisplay(v.head(15))\n\n# %%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/id_map.csv'\ndf_id_map = pd.read_csv(fn,index_col = 0)\nprint(df_id_map.shape)\ndisplay(df_id_map)\nfn = '/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv'\ndf_sample_submit = pd.read_csv(fn, index_col = 0)\nprint(df_sample_submit.shape)\ndisplay( df_sample_submit )\n\n\nfeatures_columns = ['cell_type','sm_name']\nX_full_categorical = df_de_train[ list_features_to_encode ].values\nX_submit_categorical = df_id_map[list_features_to_encode].values # pd.DataFrame(df_id_map, columns= features_columns )\nY_full = df_de_train.iloc[:,5:].values\n\nprint();\nprint('X_full_categorical.shape, X_submit_categorical.shape ,  Y_full.shape', X_full_categorical.shape, X_submit_categorical.shape ,  Y_full.shape)  ; print();\nprint(X_full_categorical[:3,:5]); print();\nprint(X_submit_categorical[:3,:5]); print();\nprint(Y_full[:3,:5]); print();\n\n","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:02:56.392083Z","iopub.execute_input":"2023-11-24T10:02:56.392870Z","iopub.status.idle":"2023-11-24T10:03:05.243265Z","shell.execute_reply.started":"2023-11-24T10:02:56.392831Z","shell.execute_reply":"2023-11-24T10:03:05.241997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.decomposition import TruncatedSVD\nreducer = TruncatedSVD(n_components=50, n_iter=7, random_state=42)\nY_full_red = reducer.fit_transform( Y_full )\nprint( Y_full_red.shape )\n\ndf_stat_loc = pd.DataFrame(Y_full_red[:,:50] ).describe()    \npd.set_option('display.max_columns', 100)\n# pd.set_option('display.max_rows', None)    \ndisplay(df_stat_loc.round() )\nfig = plt.figure(figsize = (20,3))\nplt.plot(df_stat_loc.loc['std',:].values,'*-' )\nplt.grid()\nplt.title('std for targets reduced', fontsize = 20 )\nplt.xlabel('tsvd i_component', fontsize = 20)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:03:05.247205Z","iopub.execute_input":"2023-11-24T10:03:05.247596Z","iopub.status.idle":"2023-11-24T10:03:08.719008Z","shell.execute_reply.started":"2023-11-24T10:03:05.247562Z","shell.execute_reply":"2023-11-24T10:03:08.717377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Target encoders - demo examples\n\n\nHere are some demo examples for encoders: \"TargetEncoder\", \"QuantileEncoder\", \"JamesSteinEncoder\", \"CatBoostEncoder\" , 'LeaveOneOutEncoder'\n\nWe also show effect of these encoders params on the transformed data.\n\n\n#### Smoothing param\n\nMost encoderes have the \"smoothing\" param which a sort of regularization. It controls the balance between target mean over category and target mean overall. Thus extreme \"smoothing\" will convert feature to just the constant equal to overal target mean overall set. \n\nHowever notations, direction(!) and scale of smoothing is different for different encoders:\n\n    TargetEnocder smoothing name is \"smoothing\".  Default smoothing=10.\n        LOWER(!) value means stronger regularization, smoothing = 0 - shrinks all values to constant.  \n    \n    QuantileEncoder, smoothing name is \"m\". Default is 1.\n        HIGHER(!) value of m results into stronger shrinking. \"m\" is non-negative. 0 for NO smoothing.\n        \n        Additional param is \"quantile\" - which is quantile itself, default quantile = 0.5 - corresponds to median. \n        \n    CatBoostEncoder, smoothing name is \"a\". Default a=1.\n        HIGHER(!) value of \"a\" results into stronger shrinking. \"a\" is non-negative. 0 for NO smoothing.\n        \n     JamesSteinEncoder, NO smoothing parameter - it is somewhat under-the-carpet the James-Stein provided a clever way to provide closed estimate for it.  \n     \n     LeaveOneOutEncoder,  NO smoothing parameter \n     This is very similar to target encoding but excludes the current row’s target when calculating the mean target for a level to reduce the effect of outliers.\n    \n    \n#### Sigma (adding noise to train only) param\n\nSee example and more details in the section \"Example: add Gaussian noise into training data (\"sigma\" parameter of encoder)\" below.\n\n    All the encoders have the same param - \"sigma\" - which adds random gaussion noise with std \"sigma\" to the train part of the data.\n    I.e. is added ONLY when \"fit_transform\" used, but NOT when \"transform\". \n    NO noise is added by default.\n    It is a sort of the regularization method. \n    \n    One should pay attention that noise is random - thus results will be different from run to run (unless random seed is fixed). So possible uplift might be caused by randomness and does not work on LB - one should be careful with that. ","metadata":{}},{"cell_type":"markdown","source":"## TargetEncoder","metadata":{}},{"cell_type":"markdown","source":"### Example 1 - toy data","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\nprint('smoothing 0 - sends everything to constant , difference with 1 is very small')\nprint('smoothing 10_000 or np.inf almost the same - no \"prior\" effect ')\nv =  np.random.randint(0, 10, size=100)\nv_targets = np.random.randint(0, 10, size=100)\nv = [str(t) for t in v ]\nd = pd.DataFrame(); d.index.name = 'category'\nfor smoothing in [0,0.5,1,5, 1e4, np.inf]:\n    enc = ce.TargetEncoder(smoothing = smoothing)\n    d['smoothing'+str(smoothing)] = enc.fit_transform(v, v_targets) \ndisplay( d.head(10).round(3) )\n\nprint(enc)","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:03:08.720492Z","iopub.execute_input":"2023-11-24T10:03:08.720900Z","iopub.status.idle":"2023-11-24T10:03:09.289900Z","shell.execute_reply.started":"2023-11-24T10:03:08.720865Z","shell.execute_reply":"2023-11-24T10:03:09.288722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example 2 - challenge data ( tsvd reduced)","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nlist_smoothing = [1, 10, 100, 1_000, 10_000, np.inf]\n\nfig = plt.figure(figsize = (20,5)); cc = 0\ndf_stat_loc = pd.DataFrame()\nfor smoothing in list_smoothing:\n    enc = ce.TargetEncoder(smoothing = smoothing )\n    for i_target in range(20):\n        if i_target == 0:\n            X_encoded = enc.fit_transform(X_full_categorical, Y_full_red[:,i_target])\n        else:\n            X_encoded_tmp = enc.fit_transform(X_full_categorical, Y_full_red[:,i_target])\n            X_encoded = np.concatenate( [X_encoded, X_encoded_tmp], axis = 1)\n\n    print('Smoothing: '+str(smoothing) )        \n    print('X_encoded.shape', X_encoded.shape)        \n    print(X_encoded[:2,:5])\n    print( )\n    df_stat_loc_tmp = pd.Series(X_encoded.ravel() ).describe().to_frame()\n    df_stat_loc_tmp.columns = ['Smoothing ' + str(smoothing)]\n    df_stat_loc = pd.concat( [df_stat_loc, df_stat_loc_tmp ], axis = 1)\n    cc+=1; fig.add_subplot(1,len(list_smoothing),cc)\n    plt.hist(X_encoded.ravel(), bins = 100 )\n    plt.title('Smoothing: '+str(smoothing), fontsize = 17 )\n    plt.grid()\ndisplay( df_stat_loc )\nfig.suptitle('Distributions for all encoded components at once: X_encoded.ravel() ', fontsize = 17 )\nprint('Distributions for all encoded components at once: X_encoded.ravel() ')\nplt.show()\n\nenc = ce.TargetEncoder(  smoothing=10 ) #  smoothing=10 - default smoothing \nprint('Default smoothing 10 encoding:', enc)\nlist_i_component = list(range(6))\nfig = plt.figure(figsize = (20,3)); cc = 0\nfor i_component in list_i_component:\n    cc+=1; fig.add_subplot(1,len(list_i_component),cc)\n    plt.hist(X_encoded[:,i_component], bins = 100 )\n    plt.title('i_component: '+str(i_component), fontsize = 12 )\n    plt.grid()\nplt.show()\nlist_i_component = list(range(6,12))\nfig = plt.figure(figsize = (20,3)); cc = 0\nfor i_component in list_i_component:\n    cc+=1; fig.add_subplot(1,len(list_i_component),cc)\n    plt.hist(X_encoded[:,i_component], bins = 100 )\n    plt.title('i_component: '+str(i_component), fontsize = 12 )\n    plt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:03:09.294317Z","iopub.execute_input":"2023-11-24T10:03:09.294740Z","iopub.status.idle":"2023-11-24T10:03:17.911649Z","shell.execute_reply.started":"2023-11-24T10:03:09.294680Z","shell.execute_reply":"2023-11-24T10:03:17.910146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example 3 - challenge data (no tsvd)","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nlist_smoothing = [1, 10, 100, 1_000, 10_000, np.inf]\n\nfig = plt.figure(figsize = (20,3)); cc = 0\ndf_stat_loc = pd.DataFrame()\nfor smoothing in list_smoothing:\n    enc = ce.TargetEncoder(smoothing = smoothing )\n    for i_target in range(20):\n        if i_target == 0:\n            X_encoded = enc.fit_transform(X_full_categorical, Y_full[:,i_target])\n        else:\n            X_encoded_tmp = enc.fit_transform(X_full_categorical, Y_full[:,i_target])\n            X_encoded = np.concatenate( [X_encoded, X_encoded_tmp], axis = 1)\n\n    print('Smoothing: '+str(smoothing) )        \n    print('X_encoded.shape', X_encoded.shape)        \n    print(X_encoded[:2,:5])\n    print( )\n    df_stat_loc_tmp = pd.Series(X_encoded.ravel() ).describe().to_frame()\n    df_stat_loc_tmp.columns = ['Smoothing ' + str(smoothing)]\n    df_stat_loc = pd.concat( [df_stat_loc, df_stat_loc_tmp ], axis = 1)\n    cc+=1; fig.add_subplot(1,len(list_smoothing),cc)\n    plt.hist(X_encoded.ravel(), bins = 100 )\n    plt.title('Smoothing: '+str(smoothing), fontsize = 20 )\n    plt.grid()\ndisplay( df_stat_loc )    \nfig.suptitle('TargetEncoder. Distributions for all encoded components at once: X_encoded.ravel() ', fontsize = 17, y= 1.1 )\nplt.show()\n\nenc = ce.TargetEncoder(  smoothing=10 ) #  smoothing=10 - default smoothing \nprint('Default smoothing 10 encoding:', enc)\nlist_i_component = list(range(6))\nfig = plt.figure(figsize = (20,3)); cc = 0\nfig.suptitle('TargetEncoder Default', fontsize = 17, y= 1.05 )\nfor i_component in list_i_component:\n    cc+=1; fig.add_subplot(1,len(list_i_component),cc)\n    plt.hist(X_encoded[:,i_component], bins = 100 )\n    plt.title('i_component: '+str(i_component), fontsize = 12 )\n    plt.grid()\nplt.show()\nlist_i_component = list(range(6,12))\nfig = plt.figure(figsize = (20,3)); cc = 0\nfig.suptitle('TargetEncoder Default', fontsize = 17, y= 1.05 )\nfor i_component in list_i_component:\n    cc+=1; fig.add_subplot(1,len(list_i_component),cc)\n    plt.hist(X_encoded[:,i_component], bins = 100 )\n    plt.title('i_component: '+str(i_component), fontsize = 12 )\n    plt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:03:17.913626Z","iopub.execute_input":"2023-11-24T10:03:17.914672Z","iopub.status.idle":"2023-11-24T10:03:26.658360Z","shell.execute_reply.started":"2023-11-24T10:03:17.914632Z","shell.execute_reply":"2023-11-24T10:03:26.656975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## QuantileEncoder","metadata":{}},{"cell_type":"markdown","source":"### Example 1 - toy data ","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nprint(''' Info on QuantileEncoder Category Encoder. \n\nTuneable params: two tunable parameter m (default 1) and quantile (default -  0.5)\n\nHigher value of m results into stronger shrinking. M is non-negative. 0 for no smoothing.\n\nThis a statistically modified version of target MEstimate encoder where selected features are replaced by the statistical quantile instead of the mean. Replacing with the median is a particular case where self.quantile = 0.5. In comparison to MEstimateEncoder it has two tunable parameter m and quantile\n''')\nv =  np.random.randint(0, 10, size=100)\nv_targets = np.random.randint(0, 10, size=100)\nv = [str(t) for t in v ]\nd = pd.DataFrame(); d.index.name = 'category'\nfor quantile in [ 0.5, 0.1,0.9]:\n    for m in [0.0, 1,10, 100,1000,10000,]:\n        enc = ce.QuantileEncoder(quantile = quantile , m =m )# smoothing = smoothing)\n        d['QuantileEncoder quantile' +str(quantile) + ' m(smoothing)'+str(m)] = enc.fit_transform(v, v_targets) \ndisplay( d.head(10).round(3) )","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:03:26.660329Z","iopub.execute_input":"2023-11-24T10:03:26.660743Z","iopub.status.idle":"2023-11-24T10:03:27.004946Z","shell.execute_reply.started":"2023-11-24T10:03:26.660706Z","shell.execute_reply":"2023-11-24T10:03:27.003599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example 2 - challenge data (tsvd reduced)","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nlist_smoothing = [0.0, 1,10, 100,1000,10000,] #  [1, 10, 100, 1_000, 10_000, np.inf]\n\n#for smoothing in list_smoothing:\ndf_stat_loc = pd.DataFrame()\nfor quantile in [ 0.5, 0.1,0.9]:\n    fig = plt.figure(figsize = (20,4)); cc = 0\n    \n    for m in list_smoothing:\n        str_enc_info = 'Quantile ' + str( quantile ) + '\\n Smoothing ' + str(m)\n        enc = ce.QuantileEncoder(quantile = quantile , m =m )# smoothing = smoothing)\n        for i_target in range(20):\n            if i_target == 0:\n                X_encoded = enc.fit_transform(X_full_categorical, Y_full_red[:,i_target])\n            else:\n                X_encoded_tmp = enc.fit_transform(X_full_categorical, Y_full_red[:,i_target])\n                X_encoded = np.concatenate( [X_encoded, X_encoded_tmp], axis = 1)\n\n#         print(str_enc_info )        \n#         print('X_encoded.shape', X_encoded.shape)        \n#         print(X_encoded[:2,:5])\n#         print( )\n        df_stat_loc_tmp = pd.Series(X_encoded.ravel() ).describe().to_frame()\n        df_stat_loc_tmp.columns = [str_enc_info]\n#         if cc == 0:\n#             df_stat_loc = df_stat_loc_tmp\n#         else:\n        df_stat_loc = pd.concat( [df_stat_loc, df_stat_loc_tmp ], axis = 1)\n        cc+=1; fig.add_subplot(1,len(list_smoothing),cc)\n        plt.hist(X_encoded.ravel(), bins = 100 )\n        plt.title(str_enc_info, fontsize = 17 )\n        plt.grid()\n    fig.suptitle('QuantileEncoder. Distributions for all encoded components at once: X_encoded.ravel() ', fontsize = 17, y= 1.1 )\n    print('Distributions for all encoded components at once: X_encoded.ravel() ')\n    plt.show()\ndisplay( df_stat_loc )    \n\n\nenc = ce.QuantileEncoder( quantile=0.5, m=1.0) #  quantile=0.5, m=1.0 - default smoothing \nprint()\nprint('Default  quantile=0.5, m=1.0 encoding:', enc)\nlist_i_component = list(range(6))\nfig = plt.figure(figsize = (20,3)); cc = 0\nplt.suptitle(' Default QuantileEncoder quantile=0.5, m=1.0 encoding ', fontsize = 17, y = 1.05)\n#plt.suptitle('Default  quantile=0.5, m=1.0 encoding', fontsize = 17)\nfor i_component in list_i_component:\n    cc+=1; fig.add_subplot(1,len(list_i_component),cc)\n    plt.hist(X_encoded[:,i_component], bins = 100 )\n    plt.title('i_component: '+str(i_component), fontsize = 12 )\n    plt.grid()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:03:27.008157Z","iopub.execute_input":"2023-11-24T10:03:27.008678Z","iopub.status.idle":"2023-11-24T10:03:53.502842Z","shell.execute_reply.started":"2023-11-24T10:03:27.008628Z","shell.execute_reply":"2023-11-24T10:03:53.501555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## JamesSteinEncoder","metadata":{}},{"cell_type":"markdown","source":"### Example 1 - toy data","metadata":{}},{"cell_type":"code","source":"%%time\n\nimport category_encoders as ce\n\nprint('''\nJames-Stein estimator.\n\nSee also: Stein's example (paradox) https://en.wikipedia.org/wiki/Stein%27s_example, https://en.wikipedia.org/wiki/James%E2%80%93Stein_estimator,\nhttps://mathoverflow.net/questions/93745/james-stein-phenomenon-what-does-it-mean-that-a-james-stein-estimator-beats-lea\n\nTuneable params:\n\nrandomized: bool,\nadds normal (Gaussian) distribution noise into training data in order to decrease overfitting (testing data are untouched).\n\nsigma: float (only used if randomized = True )\nstandard deviation (spread or “width”) of the normal distribution.\n\nSupported targets: binomial and continuous. For polynomial target support, see PolynomialWrapper.\n\nFor feature value i, James-Stein estimator returns a weighted average of:\n\nThe mean target value for the observed feature value i.\n\nThe mean target value (regardless of the feature value).\n\nThis can be written as:\n\nJS_i = (1-B)*mean(y_i) + B*mean(y)\nThe question is, what should be the weight B? If we put too much weight on the conditional mean value, we will overfit. If we put too much weight on the global mean, we will underfit. The canonical solution in machine learning is to perform cross-validation. However, Charles Stein came with a closed-form solution to the problem. The intuition is: If the estimate of mean(y_i) is unreliable (y_i has high variance), we should put more weight on mean(y). Stein put it into an equation as:\n\nB = var(y_i) / (var(y_i)+var(y))\nThe only remaining issue is that we do not know var(y), let alone var(y_i). Hence, we have to estimate the variances. But how can we reliably estimate the variances, when we already struggle with the estimation of the mean values?! There are multiple solutions:\n\n1. If we have the same count of observations for each feature value i and all y_i are close to each other, we can pretend that all var(y_i) are identical. This is called a pooled model. 2. If the observation counts are not equal, it makes sense to replace the variances with squared standard errors, which penalize small observation counts:\n\nSE^2 = var(y)/count(y)\nThis is called an independent model.\n\nJames-Stein estimator has, however, one practical limitation - it was defined only for normal distributions. If you want to apply it for binary classification, which allows only values {0, 1}, it is better to first convert the mean target value from the bound interval <0,1> into an unbounded interval by replacing mean(y) with log-odds ratio:\n\nlog-odds_ratio_i = log(mean(y_i)/mean(y_not_i))\nThis is called binary model. The estimation of parameters of this model is, however, tricky and sometimes it fails fatally. In these situations, it is better to use beta model, which generally delivers slightly worse accuracy than binary model but does not suffer from fatal failures.\n''')\n\nv =  np.random.randint(0, 10, size=100)\nv_targets = np.random.randint(0, 10, size=100)\nv = [str(t) for t in v ]\nd = pd.DataFrame(); d.index.name = 'category'\nfor randomized in [False, True]:\n    for sigma in [0.01,0.05, 0.1]:\n        enc = ce.JamesSteinEncoder(randomized = randomized, sigma = sigma)# \n        d['JamesSteinEncoder sigma' +str(sigma) + ' randomized'+str(randomized)] = enc.fit_transform(v, v_targets) \n        \ndisplay( d.head(10).round(3) )","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:03:53.504924Z","iopub.execute_input":"2023-11-24T10:03:53.506252Z","iopub.status.idle":"2023-11-24T10:03:53.668539Z","shell.execute_reply.started":"2023-11-24T10:03:53.506201Z","shell.execute_reply":"2023-11-24T10:03:53.667067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example 2 - challenge data (tsvd reduced)","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nlist_smoothing = [0.0, 1,10, 100,1000,10000,] #  [1, 10, 100, 1_000, 10_000, np.inf]\n\n#for smoothing in list_smoothing:\ndf_stat_loc = pd.DataFrame()\n# for sigma in [0, 0.05, 0.5, 0.1,0.9]:\nif 1:\n    fig = plt.figure(figsize = (20,4)); cc = 0\n    fig.suptitle('JamesSteinEncoder',fontsize = 17, y = 1.02)\n    \n    for sigma in [0, 0.05, 0.5, 0.1,0.9,10]:\n    #for m in list_smoothing:\n        str_enc_info = 'sigma ' + str( sigma )# + '\\n Smoothing ' + str(m)\n        enc = ce.JamesSteinEncoder(sigma = sigma, randomized=True ) # \n        for i_target in range(20):\n            if i_target == 0:\n                X_encoded = enc.fit_transform(X_full_categorical, Y_full_red[:,i_target])\n            else:\n                X_encoded_tmp = enc.fit_transform(X_full_categorical, Y_full_red[:,i_target])\n                X_encoded = np.concatenate( [X_encoded, X_encoded_tmp], axis = 1)\n\n#         print(str_enc_info )        \n#         print('X_encoded.shape', X_encoded.shape)        \n#         print(X_encoded[:2,:5])\n#         print( )\n        df_stat_loc_tmp = pd.Series(X_encoded.ravel() ).describe().to_frame()\n        df_stat_loc_tmp.columns = [str_enc_info]\n#         if cc == 0:\n#             df_stat_loc = df_stat_loc_tmp\n#         else:\n        df_stat_loc = pd.concat( [df_stat_loc, df_stat_loc_tmp ], axis = 1)\n        cc+=1; fig.add_subplot(1,len(list_smoothing),cc)\n        plt.hist(X_encoded.ravel(), bins = 100 )\n        plt.title(str_enc_info, fontsize = 17 )\n        plt.grid()\n    #fig.suptitle('Distributions for all encoded components at once: X_encoded.ravel() ', fontsize = 17 )\n    print('Distributions for all encoded components at once: X_encoded.ravel() ')\n    plt.show()\ndisplay( df_stat_loc )    \n\n\nenc = ce.JamesSteinEncoder(randomized=False) # \nprint()\nprint('Default   encoding:', enc)\nlist_i_component = list(range(6))\nfig = plt.figure(figsize = (20,3)); cc = 0\nplt.suptitle('JamesSteinEncoder ', fontsize = 17, y = 1.05)\nfor i_component in list_i_component:\n    cc+=1; fig.add_subplot(1,len(list_i_component),cc)\n    plt.hist(X_encoded[:,i_component], bins = 100 )\n    plt.title('i_component: '+str(i_component), fontsize = 12 )\n    plt.grid()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:03:53.671080Z","iopub.execute_input":"2023-11-24T10:03:53.671633Z","iopub.status.idle":"2023-11-24T10:04:00.402019Z","shell.execute_reply.started":"2023-11-24T10:03:53.671589Z","shell.execute_reply":"2023-11-24T10:04:00.400583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Catboost Encoder","metadata":{}},{"cell_type":"markdown","source":"### Example 1 - toy data","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nprint(''' Info on Catboost Category Encoder. \n\nTuneable params:\n\nsigma: float (default = None)\nadds normal (Gaussian) distribution noise into training data in order to decrease overfitting (testing data are untouched). sigma gives the standard deviation (spread or “width”) of the normal distribution.\n\na: float = big values converts to constant  (opposite direction to \"smoothing in TE\")\nadditive smoothing (it is the same variable as “m” in m-probability estimate). By default set to 1\n''')\nv =  np.random.randint(0, 10, size=100)\nv_targets = np.random.randint(0, 10, size=100)\nv = [str(t) for t in v ]\nd = pd.DataFrame(); d.index.name = 'category'\nfor sigma in [None,0.1, 1,100]:\n    for a in [0.001, 1,10, 100]:\n        enc = ce.CatBoostEncoder(sigma = sigma , a =a )# smoothing = smoothing)\n        enc.fit(v, v_targets)\n        d['CatBoostEnc sigma' +str(sigma) + ' a(smoothing)'+str(a)] =  enc.transform(v)#, v_targets)\ndisplay( d.head(10).round(3) )\nenc.get_params()","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:04:00.404183Z","iopub.execute_input":"2023-11-24T10:04:00.407079Z","iopub.status.idle":"2023-11-24T10:04:00.620815Z","shell.execute_reply.started":"2023-11-24T10:04:00.407029Z","shell.execute_reply":"2023-11-24T10:04:00.619612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example 2 - challenge data (tsvd reduced)","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nlist_smoothing = [0.0, 1,10, 100,1000,10000,] #  [1, 10, 100, 1_000, 10_000, np.inf]\n\n#for smoothing in list_smoothing:\ndf_stat_loc = pd.DataFrame()\nfor sigma in [0, 1]:\n# if 1:\n    fig = plt.figure(figsize = (20,4)); cc = 0\n    fig.suptitle('CatBoostEncoder',fontsize = 17, y = 1.02)\n    \n    for a in [0.001, 1,10, 100]:\n    #for m in list_smoothing:\n        str_enc_info = 'a '+str(a)+ ' sigma ' + str( sigma )# + '\\n Smoothing ' + str(m)\n        enc = ce.CatBoostEncoder(sigma = sigma,a = a)# ,  randomized=True ) # \n        for i_target in range(20):\n            if i_target == 0:\n                X_encoded = enc.fit_transform(X_full_categorical, Y_full_red[:,i_target])\n            else:\n                X_encoded_tmp = enc.fit_transform(X_full_categorical, Y_full_red[:,i_target])\n                X_encoded = np.concatenate( [X_encoded, X_encoded_tmp], axis = 1)\n\n#         print(str_enc_info )        \n#         print('X_encoded.shape', X_encoded.shape)        \n#         print(X_encoded[:2,:5])\n#         print( )\n        df_stat_loc_tmp = pd.Series(X_encoded.ravel() ).describe().to_frame()\n        df_stat_loc_tmp.columns = [str_enc_info]\n#         if cc == 0:\n#             df_stat_loc = df_stat_loc_tmp\n#         else:\n        df_stat_loc = pd.concat( [df_stat_loc, df_stat_loc_tmp ], axis = 1)\n        cc+=1; fig.add_subplot(1,len(list_smoothing),cc)\n        plt.hist(X_encoded.ravel(), bins = 100 )\n        plt.title(str_enc_info, fontsize = 17 )\n        plt.grid()\n    #fig.suptitle('Distributions for all encoded components at once: X_encoded.ravel() ', fontsize = 17 )\n    print('Distributions for all encoded components at once: X_encoded.ravel() ')\n    plt.show()\ndisplay( df_stat_loc )    \n\n\nenc = ce.CatBoostEncoder()# randomized=False) # \nprint()\nprint('Default   encoding:', enc)\nlist_i_component = list(range(6))\nfig = plt.figure(figsize = (20,3)); cc = 0\nfig.suptitle('CatBoostEncoder Default',fontsize = 17, y = 1.05)\n#plt.suptitle('Default  quantile=0.5, m=1.0 encoding', fontsize = 17)\nfor i_component in list_i_component:\n    cc+=1; fig.add_subplot(1,len(list_i_component),cc)\n    plt.hist(X_encoded[:,i_component], bins = 100 )\n    plt.title('i_component: '+str(i_component), fontsize = 12 )\n    plt.grid()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:04:00.627534Z","iopub.execute_input":"2023-11-24T10:04:00.627972Z","iopub.status.idle":"2023-11-24T10:04:08.397118Z","shell.execute_reply.started":"2023-11-24T10:04:00.627935Z","shell.execute_reply":"2023-11-24T10:04:08.395729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## LeaveOneOutEncoder\n\n","metadata":{}},{"cell_type":"markdown","source":"### Example 1 - toy data example","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nprint(''' \nLeave one out coding for categorical features.\n\nThis is very similar to target encoding but excludes the current row’s target when calculating the mean target for a level to reduce the effect of outliers.\n\nTune: sigma: float\nadds normal (Gaussian) distribution noise into training data in order to decrease overfitting (testing data are untouched). Sigma gives the standard deviation (spread or “width”) of the normal distribution. The optimal value is commonly between 0.05 and 0.6. The default is to not add noise, but that leads to significantly suboptimal results.\n\n''')\n\nv =  np.random.randint(0, 10, size=100)\nv_targets = np.random.randint(0, 10, size=100)\nv = [str(t) for t in v ]\nd = pd.DataFrame(); d.index.name = 'category'\nfor sigma in [None,0.01,0.05,0.1,0.2,0.3,0.4, 0.5,0.6,0.7,0.8,0.9, 1,100]:\n    prm_enc_loc = {'sigma':sigma}\n    enc = ce.LeaveOneOutEncoder(**prm_enc_loc )\n    d['LeaveOneOutEncoder sigma' +str(sigma)] = enc.fit_transform(v, v_targets) \ndisplay( d.head(10).round(3) )\nenc.get_params()","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:04:08.398923Z","iopub.execute_input":"2023-11-24T10:04:08.399958Z","iopub.status.idle":"2023-11-24T10:04:08.606455Z","shell.execute_reply.started":"2023-11-24T10:04:08.399892Z","shell.execute_reply":"2023-11-24T10:04:08.604901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example 2 - challenge data (tsvd reduced)","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nlist_smoothing = [0.0, 1,10, 100,1000,10000,] #  [1, 10, 100, 1_000, 10_000, np.inf]\n\n#for smoothing in list_smoothing:\ndf_stat_loc = pd.DataFrame()\n# for sigma in [0, 0.05, 0.5, 0.1,0.9]:\nif 1:\n    fig = plt.figure(figsize = (20,4)); cc = 0\n    fig.suptitle('LeaveOneOutEncoder',fontsize = 17, y = 1.02)\n    \n    for sigma in [0, 0.05, 0.5, 0.1,0.9,10]:\n    #for m in list_smoothing:\n        str_enc_info = 'sigma ' + str( sigma )# + '\\n Smoothing ' + str(m)\n        enc = ce.LeaveOneOutEncoder(sigma = sigma) # \n        for i_target in range(20):\n            if i_target == 0:\n                X_encoded = enc.fit_transform(X_full_categorical, Y_full_red[:,i_target])\n            else:\n                X_encoded_tmp = enc.fit_transform(X_full_categorical, Y_full_red[:,i_target])\n                X_encoded = np.concatenate( [X_encoded, X_encoded_tmp], axis = 1)\n\n#         print(str_enc_info )        \n#         print('X_encoded.shape', X_encoded.shape)        \n#         print(X_encoded[:2,:5])\n#         print( )\n        df_stat_loc_tmp = pd.Series(X_encoded.ravel() ).describe().to_frame()\n        df_stat_loc_tmp.columns = [str_enc_info]\n#         if cc == 0:\n#             df_stat_loc = df_stat_loc_tmp\n#         else:\n        df_stat_loc = pd.concat( [df_stat_loc, df_stat_loc_tmp ], axis = 1)\n        cc+=1; fig.add_subplot(1,len(list_smoothing),cc)\n        plt.hist(X_encoded.ravel(), bins = 100 )\n        plt.title(str_enc_info, fontsize = 17 )\n        plt.grid()\n    #fig.suptitle('Distributions for all encoded components at once: X_encoded.ravel() ', fontsize = 17 )\n    print('Distributions for all encoded components at once: X_encoded.ravel() ')\n    plt.show()\ndisplay( df_stat_loc )    \n\n\nprint()\nenc = ce.LeaveOneOutEncoder()\nprint('Default   encoding:', enc)\nlist_i_component = list(range(6))\nfig = plt.figure(figsize = (20,3)); cc = 0\nfig.suptitle('LeaveOneOutEncoder Default',fontsize = 17, y = 1.05)\nfor i_component in list_i_component:\n    cc+=1; fig.add_subplot(1,len(list_i_component),cc)\n    plt.hist(X_encoded[:,i_component], bins = 100 )\n    plt.title('i_component: '+str(i_component), fontsize = 12 )\n    plt.grid()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:04:08.608006Z","iopub.execute_input":"2023-11-24T10:04:08.608392Z","iopub.status.idle":"2023-11-24T10:04:14.998508Z","shell.execute_reply.started":"2023-11-24T10:04:08.608358Z","shell.execute_reply":"2023-11-24T10:04:14.996946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## GLMMEncoder \n\nGenerates some errors","metadata":{}},{"cell_type":"code","source":"# %%time\n\n# import category_encoders as ce\n\n# print(''' Info on GLMMEncoder Category Encoder. \n\n# Tuneable params:\n\n# randomized: bool,\n# adds normal (Gaussian) distribution noise into training data in order to decrease overfitting (testing data are untouched).\n\n# sigma: float (only used if randomized = True )\n# standard deviation (spread or “width”) of the normal distribution.\n\n\n# This is a supervised encoder similar to TargetEncoder or MEstimateEncoder, but there are some advantages:\n\n# Solid statistical theory behind the technique. Mixed effects models are a mature branch of statistics.\n\n# 2. No hyper-parameters to tune. The amount of shrinkage is automatically determined through the estimation process. In short, the less observations a category has and/or the more the outcome varies for a category then the higher the regularization towards “the prior” or “grand mean”. 3. The technique is applicable for both continuous and binomial targets. If the target is continuous, the encoder returns regularized difference of the observation’s category from the global mean.\n\n# If the target is binomial, the encoder returns regularized log odds per category.\n\n# In comparison to JamesSteinEstimator, this encoder utilizes generalized linear mixed models from statsmodels library.\n\n# Note: This is an alpha implementation. The API of the method may change in the future.\n# ''')\n# v =  np.random.randint(0, 10, size=100)\n# v_targets = np.random.randint(0, 10, size=100)\n# v = [str(t) for t in v ]\n# d = pd.DataFrame(); d.index.name = 'category'\n# for randomized in [False, True]:\n#     for sigma in [0.01,0.05, 0.1]:\n#         enc = ce.GLMMEncoder(randomized = randomized, sigma = sigma)# \n#         d['GLMMEncoder sigma' +str(sigma) + ' randomized'+str(randomized)] = enc.fit_transform(v, v_targets) \n# display( d.head(10).round(3) )","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:04:15.000367Z","iopub.execute_input":"2023-11-24T10:04:15.000823Z","iopub.status.idle":"2023-11-24T10:04:15.007619Z","shell.execute_reply.started":"2023-11-24T10:04:15.000789Z","shell.execute_reply":"2023-11-24T10:04:15.006559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Custom CV-schemes","metadata":{}},{"cell_type":"code","source":"%%time\nimport random \nfrom sklearn.model_selection import KFold\n\n# Prepare indexes for AmbrosM scheme:\n# For AmbrosM scheme: \nfolds_index_data_AmbrosM = [ ]\ntrain_sm_names = ['Idelalisib', 'Crizotinib', 'Linagliptin', 'Palbociclib', 'Dabrafenib', 'Alvocidib', 'LDN 193189', 'R428', 'Porcn Inhibitor III', \n  'Belinostat', 'Foretinib', 'MLN 2238', 'Penfluridol', 'Dactolisib', 'O-Demethylated Adapalene', 'Oprozomib (ONX 0912)', 'CHIR-99021']\nlist_fold_ids =  ['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells'] \nfor fold_id in list_fold_ids:\n    mask_va = (df_de_train.cell_type == fold_id) & ~df_de_train.sm_name.isin(train_sm_names)\n    mask_tr = ~mask_va # 485 or 487 training rows\n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_AmbrosM.append( [IX_train, IX_test ])\n\n# For MT CV scheme:\nfolds_index_data_MT = []\nfold_to_compounds = {0: ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189',  'Linagliptin', 'O-Demethylated Adapalene'],\n 1: ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III'],\n 2: ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol',  'R428']}\nfor fold_id in [0,1,2]:\n    mask_va = df_de_train['cell_type'].isin(['Myeloid cells', 'B cells']) & df_de_train['sm_name'].isin(fold_to_compounds[fold_id])\n    mask_tr = ~mask_va    \n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_MT.append( [IX_train, IX_test ])    \n    \n    \ndict_folds_Antonina = {0: [4, 15, 18, 41, 47, 53, 56, 57, 59, 61, 62, 66, 78, 83, 86, 93, 100, 110, 115, 116, 117, 120, 123, 131, 148, 152, 161, 185, 189, 193, 197, 198, 205, 207, 209, 210, 213, 218, 219, 225, 228, 230, 250, 252, 257, 263, 286, 292, 294, 304, 306, 310, 325, 327, 336, 338, 339, 350, 351, 356, 357, 360, 362, 369, 379, 385, 395, 419, 424, 430, 432, 434, 438, 441, 443, 445, 448, 452, 456, 465, 467, 471, 472, 474, 498, 508, 523, 534, 542, 544, 574, 581, 584, 589, 601, 607], 1: [1, 3, 5, 14, 24, 38, 40, 49, 50, 54, 55, 60, 79, 82, 87, 102, 112, 114, 118, 119, 127, 132, 146, 149, 166, 179, 183, 184, 186, 190, 192, 201, 203, 204, 216, 227, 242, 253, 261, 264, 271, 289, 295, 296, 299, 300, 309, 323, 324, 329, 333, 337, 354, 387, 390, 393, 406, 408, 415, 416, 420, 421, 422, 436, 442, 449, 453, 458, 462, 463, 469, 484, 487, 489, 490, 492, 497, 501, 505, 506, 509, 513, 515, 516, 546, 548, 566, 569, 585, 587, 588, 592, 593, 595, 600, 605], 2: [0, 16, 19, 44, 48, 51, 52, 58, 63, 65, 70, 81, 85, 88, 89, 92, 111, 113, 122, 125, 135, 137, 139, 150, 164, 175, 191, 202, 206, 211, 217, 220, 221, 229, 239, 243, 247, 248, 251, 258, 259, 260, 267, 270, 274, 281, 288, 290, 291, 303, 322, 326, 328, 332, 347, 359, 364, 370, 383, 384, 389, 392, 401, 423, 428, 454, 455, 461, 464, 494, 496, 504, 510, 511, 524, 525, 526, 530, 540, 541, 551, 565, 568, 570, 572, 573, 575, 577, 579, 580, 597, 598, 603, 606, 610, 611, 613], 3: [21, 22, 26, 36, 42, 45, 46, 64, 67, 69, 71, 80, 84, 90, 101, 103, 134, 136, 140, 147, 151, 160, 163, 165, 167, 174, 176, 177, 180, 181, 182, 187, 188, 195, 196, 199, 200, 212, 215, 226, 231, 240, 241, 249, 255, 256, 262, 265, 269, 285, 287, 297, 298, 307, 312, 319, 321, 330, 340, 342, 344, 345, 348, 358, 366, 368, 382, 391, 400, 403, 404, 405, 425, 427, 431, 433, 435, 444, 450, 451, 457, 459, 460, 473, 485, 499, 500, 502, 503, 507, 512, 514, 527, 528, 529, 531, 543, 547, 553, 554, 564, 571, 576, 583, 586, 591, 596, 602, 604, 608, 609], 4: [2, 6, 7, 17, 20, 23, 25, 27, 28, 29, 37, 39, 43, 68, 91, 121, 124, 126, 128, 129, 130, 133, 138, 153, 162, 178, 194, 208, 214, 222, 223, 224, 232, 244, 245, 246, 254, 266, 268, 272, 273, 282, 283, 284, 293, 301, 302, 305, 308, 311, 320, 331, 334, 335, 341, 343, 346, 349, 352, 353, 355, 361, 363, 365, 367, 377, 378, 380, 381, 386, 388, 394, 396, 397, 398, 399, 402, 407, 417, 418, 426, 429, 437, 439, 440, 446, 447, 466, 468, 470, 481, 482, 483, 486, 488, 491, 493, 495, 532, 533, 545, 549, 550, 552, 555, 562, 563, 567, 578, 582, 590, 594, 599, 612], 5: [8, 9, 10, 11, 12, 13, 30, 31, 32, 33, 34, 35, 72, 73, 74, 75, 76, 77, 94, 95, 96, 97, 98, 99, 104, 105, 106, 107, 108, 109, 141, 142, 143, 144, 145, 154, 155, 156, 157, 158, 159, 168, 169, 170, 171, 172, 173, 233, 234, 235, 236, 237, 238, 275, 276, 277, 278, 279, 280, 313, 314, 315, 316, 317, 318, 371, 372, 373, 374, 375, 376, 409, 410, 411, 412, 413, 414, 475, 476, 477, 478, 479, 480, 517, 518, 519, 520, 521, 522, 535, 536, 537, 538, 539, 556, 557, 558, 559, 560, 561]}\n\nfolds_index_data_Antonina = []    \nfor i in range(5):  \n    l_valid = np.array( dict_folds_Antonina[i] )\n    l_train = np.array( [k for k in range(614) if k not in l_valid ] )\n    folds_index_data_Antonina.append( [l_train, l_valid ])    \nprint( len(folds_index_data_Antonina ) )\n\nclass KFold_custom:\n    '''\n    Class with similar to sklearn \"KFold\" class, interface.\n    Support of CV schemes proposed by AmbrosM, MT, Kishan and standard sklearn KFold\n    Supports options to return IX_train,IX_valid,IX_test - triplet - if valid_size > 0\n    Examples:\n    kf = KFold_custom('AmbrosM'):     \n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n        pass\n    kf = KFold_custom('AmbrosM', valid_size = 0.1 )     \n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split()) :\n        pass\n        \n    kf = KFold_custom(CV_scheme = 'Random') # wrapper for sklearn Kfolds with shuffle = True\n    kf = KFold_custom(CV_scheme = CV_scheme, random_state = 42 ) # fixing random_state ensures reprodicibility of folds \n    '''\n    def __init__(self, CV_scheme,  valid_size = 0, n_splits=5, random_state = 42, verbose = 0 ):\n        self.CV_scheme = CV_scheme\n        self.CV_scheme_inf =  CV_scheme\n        self.valid_size = np.clip(valid_size,0,1)\n        self.random_state = random_state\n\n        self.folds_index_data = [ [np.arange(0), np.arange(0), np.arange(0)] ] # Example data\n        self.n_splits = 1 # Example data\n        if CV_scheme == 'Kishan1':\n            # One fold scheme which contains train, valid, test parts \n            IX_train_Kishan1 = np.arange(429)\n            IX_val_Kishan1 = np.arange(429,521)\n            IX_test_Kishan1 = np.arange(521,614)\n            self.n_splits = 1\n            self.folds_index_data =  [   [IX_train_Kishan1, IX_val_Kishan1, IX_test_Kishan1]  ]\n        elif CV_scheme == 'Tonya':\n            self.n_splits = 5\n            self.folds_index_data = folds_index_data_Antonina\n        elif CV_scheme == 'AmbrosM':\n            self.n_splits = 4\n            self.folds_index_data = folds_index_data_AmbrosM\n        elif CV_scheme == 'MT':\n            self.n_splits = 3\n            self.folds_index_data = folds_index_data_MT\n        elif 'Random'.lower() in CV_scheme.lower():\n            self.n_splits = n_splits\n            self.CV_scheme_inf = 'Random_'+str(n_splits) +'_'+str(random_state )\n            kf = KFold(n_splits=n_splits, random_state = random_state, shuffle=True )#,\n            self.folds_index_data = list( kf.split( np.arange(614) ) )\n        elif 'Full'.lower() in CV_scheme.lower():\n            # Just return the full set as both train and test - it useful to re-train the model on the entire data\n            IX_train_full = np.arange(614)\n            IX_test_full = np.arange(614)\n            self.n_splits = 1\n            self.folds_index_data = [   [IX_train_full, IX_test_full]  ]\n        else:\n            s = 'Uncrecognized CV_scheme ' + str(CV_scheme) \n            raise ValueError(s)\n\n        self.verbose = verbose \n        if verbose >= 10:\n            print(self.CV_scheme, self.n_splits, len(self.folds_index_data) )\n            #print(self.folds_index_data)\n            \n    def split(self, X=None):\n        '''\n        X - NOT used, just for compatibility with sklearn \n        '''\n        for item in self.folds_index_data:\n            if self.CV_scheme in ['Kishan1']:\n                if self.valid_size > 0:\n                    yield item[0],item[1],item[2]\n                else :\n                    yield np.array( list(set(item[0])|set(item[1]))  ) , item[2] # Return train, test only \n            elif self.CV_scheme in ['Tonya', 'AmbrosM','MT', 'Random', 'Full']:\n                if self.valid_size == 0:\n                    yield item[0],item[1]\n                else:\n                    # Split \"full-train\"->( real-train, valid ) \n                    index_for_real_train_part = int(  (1-self.valid_size) * len(item[0]) )\n                    if self.random_state is None:\n                        IX_train =  item[0][:index_for_real_train_part]\n                        IX_valid =  item[0][index_for_real_train_part:]\n                    elif self.random_state == -1:\n                        p = np.random.permutation(len(item[0])) \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    else:\n                        np.random.seed(self.random_state) # Set temporary random seed\n                        p = np.random.permutation(len(item[0])) \n                        np.random.seed(random.randint(0,30000)) # Randomize seed again - use Python random, not numpy       \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    yield IX_train, IX_valid, item[1]\n                    \n        \n    def get_list_main_CV_schemes(self):\n        return ['AmbrosM','MT','Kishan1','Random']\n    def get_n_splits(self, X=np.array([])):\n        return self.n_splits\n    \nprint();print();    \nprint('--------------------------------Examples-----------------------------------------------')\nprint();print();    \n        \nCV_scheme = 'Kishan1'\nprint(CV_scheme,  )\nkf = KFold_custom('Kishan1', valid_size = 0.18)     \n#list( kf.split() )\nprint(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test)  \" )\nfor i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split(None)) :\n    print(i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test) )\nprint()\n\nfor CV_scheme in ['AmbrosM','MT','Random']:\n    valid_size = 0\n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size)     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train),type(IX_test)  \" )\n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train), len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train), type(IX_test) )\n    \n    print()\n    valid_size = 0.2 \n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size )     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test)  \" )\n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test) )\n    print()","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:04:15.009186Z","iopub.execute_input":"2023-11-24T10:04:15.009812Z","iopub.status.idle":"2023-11-24T10:04:15.121774Z","shell.execute_reply.started":"2023-11-24T10:04:15.009779Z","shell.execute_reply":"2023-11-24T10:04:15.120457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Simple Modeling","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\nimport category_encoders as ce\n\nverbose = 1\n\nn_components = 30\nreducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n\nalpha = 1\nmodel = Ridge(alpha=alpha)\n\nsmoothing = 10\nenc = ce.TargetEncoder(smoothing = smoothing )\n\nCV_scheme = 'MT'\nprint('CV_scheme:', CV_scheme)\nvalid_size = 0 \nkf = KFold_custom(CV_scheme, valid_size = valid_size)   \n\nlist_mrrmse = [];list_r2 = []\nfor i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n    # \n    # Get train, test:\n    X_train_categorical,X_test_categorical = X_full_categorical[IX_train], X_full_categorical[IX_test]\n    \n    # Reduce the targets \n    Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n    \n    if verbose >= 10:\n        print('fold:', i_fold, 'X_train.shape, X_test.shape, Y_train_red.shape', X_train.shape, X_test.shape, Y_train_red.shape)\n        \n    ### -------------- Target Encoding -------------------------------------------------------------------\n    for i_target in range(Y_train_red.shape[1]):\n        if i_target == 0:\n            X_train_encoded = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n            X_test_encoded = enc.transform(X_test_categorical)\n        else:\n            X_encoded_tmp = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n            X_train_encoded = np.concatenate( [X_train_encoded, X_encoded_tmp], axis = 1)\n            X_encoded_tmp = enc.transform(X_test_categorical)\n            X_test_encoded = np.concatenate( [X_test_encoded, X_encoded_tmp], axis = 1)\n    \n    \n    model.fit(X_train_encoded, Y_train_red)\n    Y_test_pred_red = model.predict( X_test_encoded )\n    Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n    \n    Y_test = Y_full[IX_test]\n    mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n    print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n    list_mrrmse.append(mrrmse); list_r2.append(r2)\nprint()\nprint('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ))\nprint('Average r2:', np.round(np.mean(list_r2 ),4 ))","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:04:15.123819Z","iopub.execute_input":"2023-11-24T10:04:15.124586Z","iopub.status.idle":"2023-11-24T10:04:23.977345Z","shell.execute_reply.started":"2023-11-24T10:04:15.124541Z","shell.execute_reply":"2023-11-24T10:04:23.972535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Wrapper CV-computatation function\n\nLet us wrap the code above into in a function","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\n\ndef compute_cv( model, target_encoder,  list_CV_scheme , verbose = 0 ):\n    '''\n    Compute cross validation(s) for model.\n    Returns dictionary with scores for each CV-scheme and fold\n    \n    dict_res = compute_cv( Ridge(alpha=1), ce.TargetEncoder(smoothing = 10), ['MT','AmbrosM'] , verbose = 0 )\n    '''\n\n    dict_res = {}\n    \n    enc = target_encoder\n    \n    reducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n    \n    #alpha = 1\n    #model = Ridge(alpha=alpha)\n    \n    for CV_scheme in list_CV_scheme:\n        dict_res[CV_scheme] = {}\n        #CV_scheme = 'MT'\n        if verbose >= 1:\n            print('CV_scheme:', CV_scheme)\n         \n        kf = KFold_custom(CV_scheme)# , valid_size = valid_size)  # valid_size = 0 \n        \n        list_mrrmse = [];list_r2 = []; i_blend_submit  = 0\n        for i_fold,(IX_train,IX_test) in enumerate( kf.split()) :\n            # \n            # Get train, test:\n            X_train_categorical,X_test_categorical = X_full_categorical[IX_train], X_full_categorical[IX_test]\n\n            # Reduce the targets \n            Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n\n            if verbose >= 10:\n                print('fold:', i_fold, 'X_train_categorical.shape, X_test_categorical.shape, Y_train_red.shape', X_train_categorical.shape, X_test_categorical.shape, Y_train_red.shape)\n\n            ### -------------- Target Encoding -------------------------------------------------------------------\n            for i_target in range(Y_train_red.shape[1]):\n                if i_target == 0:\n                    X_train_encoded = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n                    X_test_encoded = enc.transform(X_test_categorical)\n                    X_submit_encoded = enc.transform(X_submit_categorical)\n                else:\n                    X_encoded_tmp = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n                    X_train_encoded = np.concatenate( [X_train_encoded, X_encoded_tmp], axis = 1)\n                    X_encoded_tmp = enc.transform(X_test_categorical)\n                    X_test_encoded = np.concatenate( [X_test_encoded, X_encoded_tmp], axis = 1)\n                    X_encoded_tmp = enc.transform(X_submit_categorical)\n                    X_submit_encoded = np.concatenate( [X_submit_encoded, X_encoded_tmp], axis = 1)\n\n\n            model.fit(X_train_encoded, Y_train_red)\n            Y_test_pred_red = model.predict( X_test_encoded )\n            Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n            Y_test = Y_full[IX_test]\n            mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n            list_mrrmse.append(mrrmse); list_r2.append(r2)\n            if verbose >= 1:\n                print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n                \n            if i_blend_submit == 0:\n                Y_submit_pred = reducer.inverse_transform( model.predict( X_submit_encoded  ) )\n            else:\n                Y_submit_pred = ( Y_submit_pred * i_blend_submit + reducer.inverse_transform( model.predict( X_submit_encoded  ) ) ) / ( i_blend_submit + 1 )\n            i_blend_submit += 1\n            \n        dict_res[CV_scheme]['mrrmse'] = list_mrrmse\n        dict_res[CV_scheme]['r2'] = list_mrrmse\n        dict_res[CV_scheme]['Y_submit'] = Y_submit_pred\n        \n        if verbose >= 1:            \n            print('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ))\n            print('Average r2:', np.round(np.mean(list_r2 ),4 ))\n            print()\n            \n    return dict_res\n\ndict_res = compute_cv( Ridge(alpha=1), ce.TargetEncoder(smoothing = 10), ['MT','AmbrosM'] , verbose = 100 )\ndict_res","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:04:23.986032Z","iopub.execute_input":"2023-11-24T10:04:23.987063Z","iopub.status.idle":"2023-11-24T10:04:43.771574Z","shell.execute_reply.started":"2023-11-24T10:04:23.986995Z","shell.execute_reply":"2023-11-24T10:04:43.770032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling several encoders with different alphas","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\nimport warnings\nwarnings.filterwarnings(\"ignore\") \n\nverbose = 1\n\nlist_CV_scheme = ['Tonya', 'MT','AmbrosM', 'Random' ] \n\n# list_alpha = [0.0001,0.0002,0.0005,0.001, 0.002, 0.005, 0.01,0.02, 0.05,0.1,0.2,0.5,1,2,5,10,20,50 ,1e2,2e2,5e2,1e3,2e3,5e3,1e4,1e5,1e6]\nlist_alpha = [1e1, 1e2,2e2,5e2, 1e3,2e3,5e3,  1e4,2e4,5e4,8e4, 1e5,2e5,5e5, 1e6, 2e6,5e6, 1e7, 1e8] #   [1e1, 1e2, 1e3, 1e4, 1e5, 1e6,1e7, 1e8] # [1e1, 1e5,1e8]\n# list_alpha = [ 1e5,1e8]\n\ndict_df_stat = {}\nfor encoder_id in ['TargetEncoder', \"QuantileEncoder\", \"JamesSteinEncoder\", \"CatBoostEncoder\", 'LeaveOneOutEncoder' ]: # \"TargetEncoder\", \"QuantileEncoder\", \"JamesSteinEncoder\", \"CatBoostEncoder\", 'LeaveOneOutEncoder'\n    enc = getattr(ce,encoder_id)()\n    print(enc)\n    \n    df_stat = pd.DataFrame(); IX_stat = -1\n    for alpha in list_alpha:\n        model = Ridge(alpha = alpha)\n        str_model_id = 'Ridge'\n        dict_res = compute_cv( model, enc, list_CV_scheme , verbose =  0 )\n        \n        \n        # Pack obtained scores into dataframe: \n        IX_stat += 1\n        df_stat.loc[IX_stat,'alpha'] = alpha\n        for CV_scheme in  list_CV_scheme:\n            score_name = 'mrrmse'\n            col_average_name = score_name + ' ' + CV_scheme   + ' ' +  'average'\n            df_stat.loc[IX_stat, col_average_name ] =  np.mean( dict_res[CV_scheme]['mrrmse'] ) # Average over folds\n            for i_fold in range(len(dict_res[CV_scheme][score_name])):\n                s = dict_res[CV_scheme][score_name][i_fold]\n                df_stat.loc[IX_stat,score_name+' fold '+str(i_fold) +  ' ' + CV_scheme] = s\n\n                \n    display(df_stat)\n    fn = 'df_stat_'+ encoder_id  + '.csv'\n    print(fn)\n    df_stat.to_csv(fn)\n    \n    dict_df_stat[encoder_id] = df_stat\n    \nwarnings.filterwarnings(\"default\") \n# for k in dict_df_stat:\n#     df_stat = dict_df_stat[k]\n#     #display(df_stat.sort_values('mrrmse MT average',ascending = False).head(5))\n#     display(df_stat.sort_values('mrrmse MT average',ascending = False).head(5))\n#     fn = 'df_stat_'+ k  + '.csv'\n#     print(fn)\n#     df_stat.to_csv(fn)\n","metadata":{"execution":{"iopub.status.busy":"2023-11-24T10:04:43.773939Z","iopub.execute_input":"2023-11-24T10:04:43.774832Z","iopub.status.idle":"2023-11-24T11:25:07.331392Z","shell.execute_reply.started":"2023-11-24T10:04:43.774771Z","shell.execute_reply":"2023-11-24T11:25:07.329741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plots scores/alphas","metadata":{}},{"cell_type":"code","source":"%%time\nstr_model_inf = str_model_id + ' '+ str_inf_cfg\n\nfor encoder_id in dict_df_stat: # ['TargetEncoder' ]:    \n    df_stat = dict_df_stat[encoder_id]\n    for score_name in ['mrrmse']:# ,'r2']:\n        print();print(); print('Score:', score_name ); print()\n        for CV_scheme in list_CV_scheme:\n            list_selected_cols =[ col for col in df_stat.columns if (score_name in col) and (CV_scheme in col) ]\n            \n            plt.figure(figsize = (20,4))\n\n            str_inf = 'Best alphas: '\n            for col in list_selected_cols:\n                v = df_stat[col]\n                plt.plot(np.log10(list_alpha), df_stat[col] ,'*-' , label = col)\n                #print('Best alpha:', df_stat.sort_values(col)['alpha'].iat[0], 'for ', col )\n                if score_name in ['mrrmse']:\n                    str_score = str( np.round(df_stat.sort_values(col)[col].iat[0],4) )\n                    str_inf += str(df_stat.sort_values(col)['alpha'].iat[0]) + ' for ' + col + ' ' + str_score +' ; '\n                else:\n                    str_score = str( np.round(df_stat.sort_values(col,ascending = False)[col].iat[0],4) )\n                    str_inf += str(df_stat.sort_values(col,ascending = False)['alpha'].iat[0]) + ' for ' + col  + ' ' + str_score + ' ; '\n\n            print(str_inf)\n            plt.xlabel('LOG10 alpha',fontsize = 20 )\n            plt.legend(fontsize = 12 )\n            plt.title(encoder_id + '. ' + 'scores: '  + score_name + ', CV: ' + CV_scheme + ', Model: ' + str_model_inf, fontsize = 20 )\n            plt.grid()\n            plt.show()\n\n            ","metadata":{"execution":{"iopub.status.busy":"2023-11-24T11:25:07.333622Z","iopub.execute_input":"2023-11-24T11:25:07.334058Z","iopub.status.idle":"2023-11-24T11:25:16.358786Z","shell.execute_reply.started":"2023-11-24T11:25:07.334023Z","shell.execute_reply":"2023-11-24T11:25:16.357518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Brief final report(s)","metadata":{}},{"cell_type":"code","source":"%%time\n\ndf_brief_report = pd.DataFrame()\nfor encoder_id in dict_df_stat: # ['TargetEncoder' ]:    \n    enc = getattr(ce,encoder_id)()\n    print(encoder_id)\n\n    df_stat = dict_df_stat[encoder_id]\n    list_alpha = []\n    for col in ['mrrmse AmbrosM average', 'mrrmse MT average', 'mrrmse Random average',  ]:\n        alpha_tmp = df_stat.sort_values(col)['alpha'].iloc[0]\n        score_tmp = df_stat.sort_values(col)[col].iat[0]\n        list_alpha.append(alpha_tmp)\n        CV_scheme_loc = col.split(' ')[1]\n        str_score_inf = col.split(' ')[0]\n        df_brief_report.loc[encoder_id + ' ' + str_model_id + ' '+str_inf_cfg   , 'alpha opt '+ CV_scheme_loc    ] = alpha_tmp\n\n        dict_res = compute_cv( Ridge(alpha=alpha_tmp), enc, ['MT','AmbrosM','Random'] , verbose = 0 )\n        for CV2 in ['MT','AmbrosM','Random']:\n            df_brief_report.loc[encoder_id + ' ' +  str_model_id + ' '+str_inf_cfg   ,  str_score_inf +' '  +  CV2 + ' alpha opt for '+ CV_scheme_loc   ] = np.mean( dict_res[CV2][ str_score_inf ] )\n\n    str_score_inf = 'mrrmse'          \n    alpha_average =    np.round( np.mean(list_alpha[:2]), 4)\n    df_brief_report.loc[encoder_id + ' ' +  str_model_id + ' ' +str_inf_cfg   , 'alpha average MT,AmbrosM '   ] = alpha_average\n    dict_res = compute_cv( Ridge(alpha=alpha_average), enc,  ['MT','AmbrosM','Random'] , verbose = 0 )\n    for CV2 in ['MT','AmbrosM','Random']:\n        df_brief_report.loc[encoder_id + ' ' +  str_model_id + ' '+str_inf_cfg   ,  str_score_inf+' '  +  CV2 + ' alpha average MT,AmbrosM '    ] = np.mean( dict_res[CV2][ str_score_inf ] )\n\n    alpha_selected = 1 #   \n    df_brief_report.loc[ encoder_id + ' ' + str_model_id + ' '+str_inf_cfg   , 'alpha selected '   ] = alpha_selected\n    dict_res = compute_cv( Ridge(alpha=alpha_selected),enc,  ['MT','AmbrosM','Random'] , verbose = 0 )\n    for CV2 in ['MT','AmbrosM','Random']:\n        df_brief_report.loc[encoder_id + ' ' +  str_model_id + ' '+str_inf_cfg   ,  str_score_inf +' ' +  CV2 + ' alpha(selected): ' + str(alpha_selected)    ] = np.mean( dict_res[CV2][ str_score_inf ] )\n\n    print( list_alpha     )\n\ndf_brief_report.round(5).to_csv('df_brief_report.csv')\ndf_brief_report.round(4)","metadata":{"execution":{"iopub.status.busy":"2023-11-24T11:25:16.360509Z","iopub.execute_input":"2023-11-24T11:25:16.361919Z","iopub.status.idle":"2023-11-24T11:40:18.094074Z","shell.execute_reply.started":"2023-11-24T11:25:16.361876Z","shell.execute_reply":"2023-11-24T11:40:18.092400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Report for alpha averaged optimal for AmbrosM and MT","metadata":{}},{"cell_type":"code","source":"%%time\n\ndf_very_brief_report = pd.DataFrame()\n\nfor encoder_id in dict_df_stat: # ['TargetEncoder' ]:    \n\n    enc = getattr(ce,encoder_id)()\n\n    df_stat = dict_df_stat[encoder_id]\n\n    list_alpha = []\n    for col in ['mrrmse AmbrosM average', 'mrrmse MT average']:\n        alpha_tmp = df_stat.sort_values(col)['alpha'].iloc[0]\n        list_alpha.append(alpha_tmp)\n\n    alpha_average  = np.round( np.mean(list_alpha), 4)\n    print(encoder_id,'Alpha half sum of optimal for MT and AmbrosM: ' , alpha_average )\n    print('The two alpha: ',  list_alpha )\n\n    dict_res = compute_cv( Ridge(alpha=alpha_average), enc, ['MT','AmbrosM','Random'] , verbose = 0 )\n    dict_tmp_foldwise = {}; dict_tmp = {}\n    for CV_scheme in ['MT','AmbrosM','Random']:\n        s = dict_res[CV_scheme]['mrrmse']\n        dict_tmp_foldwise[CV_scheme] = np.round(s,5)\n        s = np.mean(dict_res[CV_scheme]['mrrmse'])\n        dict_tmp[CV_scheme] = np.round(s,5)\n    print(dict_tmp )\n    print(dict_tmp_foldwise )\n    print()\n\n    str_IX = encoder_id + ' ' +  str_model_id + ' ' +str_inf_cfg\n    df_very_brief_report.loc[ str_IX  , 'alpha optimal average '   ] = alpha_average    \n    for CV_scheme in ['MT','AmbrosM','Random']:\n        s = dict_res[CV_scheme]['mrrmse']\n        df_very_brief_report.loc[ str_IX  , 'mrrmse ' + CV_scheme   ] = np.round(np.mean(s),5)\n    for CV_scheme in ['MT','AmbrosM','Random']:\n        vec_scores = dict_res[CV_scheme]['mrrmse']\n        for i_fold in range(len(vec_scores)):\n            s = vec_scores[i_fold]\n            df_very_brief_report.loc[ str_IX  , 'mrrmse ' + CV_scheme + ' ' + str(i_fold)    ] = np.round(s,5)\n    \ndf_very_brief_report.to_csv( 'df_very_brief_report.csv')    \ndisplay( df_very_brief_report.round(5) )","metadata":{"execution":{"iopub.status.busy":"2023-11-24T11:40:18.099820Z","iopub.execute_input":"2023-11-24T11:40:18.104973Z","iopub.status.idle":"2023-11-24T11:43:18.959531Z","shell.execute_reply.started":"2023-11-24T11:40:18.104880Z","shell.execute_reply":"2023-11-24T11:43:18.957938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ranking encoders report","metadata":{}},{"cell_type":"code","source":"list_selected_cols = ['mrrmse AmbrosM alpha opt for AmbrosM', 'mrrmse MT alpha opt for MT', 'mrrmse Random alpha opt for Random']\ndf_ranking_report = df_brief_report[list_selected_cols].rank()\ndf_ranking_report.columns = [t.replace('mrrmse','Rank') for t in df_ranking_report.columns  ]\ndf_ranking_report.index.name = 'Encoder'\ndf_ranking_report['Rank Mean'] = df_ranking_report.mean(axis = 1).round(1)\ndf_ranking_report =  df_ranking_report[ ['Rank Mean'] + [t for t in df_ranking_report.columns if t != 'Rank Mean'] ] # Put on the first position\ndf_ranking_report = df_ranking_report.sort_values( 'Rank Mean') \ndf_ranking_report.to_csv('df_ranking_report.csv')\ndisplay( df_ranking_report )   ","metadata":{"execution":{"iopub.status.busy":"2023-11-24T11:43:18.962392Z","iopub.execute_input":"2023-11-24T11:43:18.963535Z","iopub.status.idle":"2023-11-24T11:43:19.002939Z","shell.execute_reply.started":"2023-11-24T11:43:18.963474Z","shell.execute_reply":"2023-11-24T11:43:19.001838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_ranking_and_scores_report = pd.concat( [df_ranking_report ,  df_brief_report[list_selected_cols].round(5) ] , axis = 1)\ndf_ranking_and_scores_report.to_csv('df_ranking_and_scores_report.csv')\ndisplay( df_ranking_and_scores_report )   ","metadata":{"execution":{"iopub.status.busy":"2023-11-24T11:43:19.004924Z","iopub.execute_input":"2023-11-24T11:43:19.005716Z","iopub.status.idle":"2023-11-24T11:43:19.031079Z","shell.execute_reply.started":"2023-11-24T11:43:19.005655Z","shell.execute_reply":"2023-11-24T11:43:19.029685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare submit(s) ","metadata":{}},{"cell_type":"markdown","source":"## Prepare  blend with \"priors\"\n\nThe trick typically uplifts the score about 0.01\n\nTrick to add \"priors\" (aggregates over the cell type and compound - which typically improves scores by 0.01 ): ( df_submit = (1-w_compound - w_cell_type)(df_submit) + w_compound df_submit_aggr_compound + w_cell_type * df_submit_aggr_cell_type )\n\nLIUDA CHELDIEVA: https://www.kaggle.com/code/liudacheldieva/streamlined-baseline-approach\n\nZXMKCD : https://www.kaggle.com/code/zmcxjt/streamlined-baseline-approach","metadata":{}},{"cell_type":"code","source":"%%time\n#group by drug and take the mean\ndf_tmp = df_de_train.iloc[:, [1] + list(range(5, df_de_train.shape[1]))] # Take only numeric columns and \"sm_name\"\ndf_aggr = df_tmp.groupby('sm_name').mean().reset_index()\nprint(df_aggr.shape)\ndf_submit_aggr_compound = pd.merge( df_id_map.reset_index(),  df_aggr, on='sm_name', how = 'left' ).sort_values('id').drop(columns = ['cell_type', 'sm_name']).set_index('id')\nprint(df_submit_aggr_compound.shape)\ndisplay(df_submit_aggr_compound.head(3))\n\nprint( )\n\ndf_tmp = df_de_train.iloc[:, [0] + list(range(5, df_de_train.shape[1]))] # Take only numeric columns and \"cell_type\"\ndf_aggr = df_tmp.groupby('cell_type').mean().reset_index()\nprint(df_aggr.shape)\ndf_submit_aggr_cell_type = pd.merge( df_id_map.reset_index(),  df_aggr, on='cell_type', how = 'left' ).sort_values('id').drop(columns = ['cell_type', 'sm_name']).set_index('id')\nprint(df_submit_aggr_cell_type.shape)\ndisplay(df_submit_aggr_cell_type.head(3))\n\n","metadata":{"execution":{"iopub.status.busy":"2023-11-24T11:43:19.033505Z","iopub.execute_input":"2023-11-24T11:43:19.033937Z","iopub.status.idle":"2023-11-24T11:43:19.913643Z","shell.execute_reply.started":"2023-11-24T11:43:19.033901Z","shell.execute_reply":"2023-11-24T11:43:19.912118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Save main submits and their blends","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\n\nverbose = 1\n\ni_blend = 0\nfor encoder_id in dict_df_stat: # ['TargetEncoder' ]:    \n    print(); print()\n    print(encoder_id)\n    enc = getattr(ce,encoder_id)()\n\n    df_stat = dict_df_stat[encoder_id]\n\n    list_alpha = []\n    for col in ['mrrmse AmbrosM average', 'mrrmse MT average']:\n        alpha_tmp = df_stat.sort_values(col)['alpha'].iloc[0]\n        list_alpha.append(alpha_tmp)\n\n    alpha_average  = np.round( np.mean(list_alpha), 4)\n    print(encoder_id,'Alpha half sum of optimal for MT and AmbrosM: ' , alpha_average )\n    print('The two alpha: ',  list_alpha )\n    \n    dict_res = compute_cv( Ridge(alpha=alpha_average), enc, ['Full'] , verbose = 0 )\n\n    Y_submit_pred = dict_res['Full']['Y_submit'] \n\n    df_submit = pd.DataFrame(Y_submit_pred, columns = df_de_train.columns[5:])\n    df_submit.index.name = 'id'\n    print( df_submit.shape )\n    display(df_submit)\n    str_IX = (encoder_id + ' ' +  str_model_id + ' ' +str_inf_cfg).replace(' ','_')\n    fn_save = 'submission_' + str_IX +'.csv' \n    print(fn_save)\n    df_submit.to_csv(fn_save)\n    \n    if i_blend == 0:\n        df_submit_blend = df_submit.copy()\n    else:\n        df_submit_blend = ( df_submit_blend*i_blend + df_submit  )/i_blend\n    i_blend += 1\n    \n    \n    w_cell_type = 0.1\n    w_compound = 0.45\n    df_submit_2 = (1-w_compound - w_cell_type)*(df_submit) + w_compound * df_submit_aggr_compound +  w_cell_type * df_submit_aggr_cell_type\n    fn_save2 = 'submission_' + str_IX +'_blendWithPriors_'+str(w_compound)+'_'+str( w_cell_type )+'.csv' \n    df_submit_2.to_csv(fn_save2)\n    print(fn_save2)\n    print(df_submit_2.shape)\n    display(df_submit_2.head(3) )\n\n\nfn_save = 'submission_blend_' + str( i_blend )  +'.csv' \nprint(fn_save)\ndf_submit = df_submit_blend\ndf_submit.to_csv(fn_save)\ndisplay(df_submit )\n\ndf_submit_2 = (1-w_compound - w_cell_type)*(df_submit) + w_compound * df_submit_aggr_compound +  w_cell_type * df_submit_aggr_cell_type\n#fn_save2 = 'submission_' + str_IX +'_blendWithPriors_'+str(w_compound)+'_'+str( w_cell_type )+'.csv' \nfn_save2 = 'submission_blend_' + str( i_blend )   +'_blendWithPriors_'+str(w_compound)+'_'+str( w_cell_type )+'.csv' \ndf_submit_2.to_csv(fn_save2)\nprint(fn_save2)\nprint(df_submit_2.shape)\ndisplay(df_submit_2.head(3) )","metadata":{"execution":{"iopub.status.busy":"2023-11-24T11:43:19.915234Z","iopub.execute_input":"2023-11-24T11:43:19.915572Z","iopub.status.idle":"2023-11-24T11:46:21.330543Z","shell.execute_reply.started":"2023-11-24T11:43:19.915541Z","shell.execute_reply":"2023-11-24T11:46:21.329537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Supplementary 1. Example: add Gaussian noise into training data (\"sigma\" parameter of encoder)\n\n\nAs discussed above there is an option to add noise to the training data. \nIn category encoders it is controlled by the \"sigma\" parameter. (Sigma gives the standard deviation (spread or “width”) of the normal distribution).\nBy default NO noise is added. Here we give an example to add noise.\nPay attention on the following:\n\n    Since it random - the results may differ from run to run (or fix the random seed otherwise).\n    Tuning the \"sigma\" one may uplift the score - but should be careful that uplift might be cause only by randomness above.\n    \nIn example below we take best obtained encoder case obtained so far - LeaveOneOut and show how the results varies when \"sigma\" is turned on.   \n\nSo we repeat the same modeling 10 times with only difference - each time added random noise is different. \n\nNo noise result -  mrrmse: 2.5353. Result with sigma = 0.1 noise varies from  2.5069 to 2.58 \n\n    Surprisingly: \n    Blending predictions from different runs - does not bring essential uplift - results are still around 2.5353  (no noise case). \n    ","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\nimport category_encoders as ce\n\nverbose = 1\n\nn_components = 30\nreducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n\nalpha = 10000.0\nmodel = Ridge(alpha=alpha)\n\n# smoothing = 10\n# enc = ce.TargetEncoder(smoothing = smoothing )\nenc = ce.LeaveOneOutEncoder(sigma = 0.1 )\n\nCV_scheme = 'MT'\nprint('CV_scheme:', CV_scheme); print()\nvalid_size = 0 \nkf = KFold_custom(CV_scheme, valid_size = valid_size)   \n\nY_oof_pred_blend = np.zeros( (Y_full.shape) ); i_self_blend = 0; Y_oof_pred = np.zeros( (Y_full.shape) );\nfor ii_self_blend in range(20):\n    if ii_self_blend == 0:\n        print('i_self_blend', i_self_blend, 'First run - without noise.')\n        enc = ce.LeaveOneOutEncoder( )\n    else:\n        print('i_self_blend', i_self_blend)\n        enc = ce.LeaveOneOutEncoder(sigma = 0.1)\n    \n    list_mrrmse = [];list_r2 = []\n    kf = KFold_custom(CV_scheme, valid_size = valid_size)   \n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n        # \n        # Get train, test:\n        X_train_categorical,X_test_categorical = X_full_categorical[IX_train], X_full_categorical[IX_test]\n\n        # Reduce the targets \n        Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n\n        if verbose >= 10:\n            print('fold:', i_fold, 'X_train.shape, X_test.shape, Y_train_red.shape', X_train.shape, X_test.shape, Y_train_red.shape)\n\n        ### -------------- Target Encoding -------------------------------------------------------------------\n        for i_target in range(Y_train_red.shape[1]):\n            if i_target == 0:\n                X_train_encoded = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n                X_test_encoded = enc.transform(X_test_categorical)\n            else:\n                X_encoded_tmp = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n                X_train_encoded = np.concatenate( [X_train_encoded, X_encoded_tmp], axis = 1)\n                X_encoded_tmp = enc.transform(X_test_categorical)\n                X_test_encoded = np.concatenate( [X_test_encoded, X_encoded_tmp], axis = 1)\n\n\n        model.fit(X_train_encoded, Y_train_red)\n        Y_test_pred_red = model.predict( X_test_encoded )\n        Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n        Y_oof_pred[ IX_test ] =  Y_test_pred\n\n        Y_test = Y_full[IX_test]\n        mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n        print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n        list_mrrmse.append(mrrmse); list_r2.append(r2)\n        \n        \n    Y_oof_pred_blend =  ( Y_oof_pred_blend * i_self_blend + Y_oof_pred ) / (i_self_blend + 1)\n    i_self_blend += 1\n    print('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ),  'Average r2:', np.round(np.mean(list_r2 ),4 ))\n    print()\n    \nprint('i_self_blend', i_self_blend)    \nY_oof_pred_blend =  Y_oof_pred_blend\n\nprint('Blend scoring:')\nkf = KFold_custom(CV_scheme, valid_size = valid_size)   \nlist_mrrmse = [];list_r2 = []\nfor i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n    Y_test = Y_full[IX_test]\n    mrrmse = np.sqrt(np.square(Y_test - Y_oof_pred_blend[IX_test]).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_oof_pred_blend[IX_test] ) \n    print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n    list_mrrmse.append(mrrmse); list_r2.append(r2)\nprint('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ),  'Average r2:', np.round(np.mean(list_r2 ),4 ))\nprint()\n    ","metadata":{"execution":{"iopub.status.busy":"2023-11-24T11:46:21.332194Z","iopub.execute_input":"2023-11-24T11:46:21.333239Z","iopub.status.idle":"2023-11-24T11:49:13.198106Z","shell.execute_reply.started":"2023-11-24T11:46:21.333204Z","shell.execute_reply":"2023-11-24T11:49:13.196545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Supplementary 2. Change both smoothing and alpha\n\n\nHere we change both Ridge alpha and target encoder smoothing.\n\nSuprisingly we see better results when smoothing is practically turned off and alpha is increased.\n\nAlpha = 300_000 ","metadata":{}},{"cell_type":"code","source":"%%time\n\nverbose = 1\n\nlist_smoothing = [10, 20, 30,40,50,60,70,80,90,100,1000, 1e4, np.inf ]\ndf_stat_scores = pd.DataFrame(index = list_smoothing )\n\nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\nimport category_encoders as ce\n\n\nn_components = 30\nreducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\nCV_scheme = 'MT'\nprint('CV_scheme:', CV_scheme); print()\nvalid_size = 0 \nkf = KFold_custom(CV_scheme, valid_size = valid_size)   \n\n\nfor alpha in [ 10_000, 50_000, 100_000, 150_000, 200_000,250_000,300_000, 400_000, 500_000,600_000,700_000, 800_000, 1_000_000]:\n    # alpha = 10_000.0\n    if verbose >= 10:\n        print()\n        print('alpha:',alpha); \n    model = Ridge(alpha=alpha)\n\n    # enc = ce.TargetEncoder(smoothing = smoothing )\n    # enc = ce.LeaveOneOutEncoder(sigma = 0.05 )\n\n\n    Y_oof_pred_blend = np.zeros( (Y_full.shape) ); i_self_blend = 0; Y_oof_pred = np.zeros( (Y_full.shape) );\n    list_result_scores = []\n    for smoothing in list_smoothing:\n        if verbose >= 10:\n            print('smoothing:', smoothing, 'i:', i_self_blend)\n        enc = ce.TargetEncoder(smoothing = smoothing )\n\n        list_mrrmse = [];list_r2 = []\n        kf = KFold_custom(CV_scheme, valid_size = valid_size)   \n        for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n            # \n            # Get train, test:\n            X_train_categorical,X_test_categorical = X_full_categorical[IX_train], X_full_categorical[IX_test]\n\n            # Reduce the targets \n            Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n\n            if verbose >= 10:\n                print('fold:', i_fold, 'X_train.shape, X_test.shape, Y_train_red.shape', X_train_categorical.shape, X_test_categorical.shape, Y_train_red.shape)\n\n            ### -------------- Target Encoding -------------------------------------------------------------------\n            for i_target in range(Y_train_red.shape[1]):\n                if i_target == 0:\n                    X_train_encoded = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n                    X_test_encoded = enc.transform(X_test_categorical)\n                else:\n                    X_encoded_tmp = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n                    X_train_encoded = np.concatenate( [X_train_encoded, X_encoded_tmp], axis = 1)\n                    X_encoded_tmp = enc.transform(X_test_categorical)\n                    X_test_encoded = np.concatenate( [X_test_encoded, X_encoded_tmp], axis = 1)\n\n\n            model.fit(X_train_encoded, Y_train_red)\n            Y_test_pred_red = model.predict( X_test_encoded )\n            Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n            Y_oof_pred[ IX_test ] =  Y_test_pred\n\n            Y_test = Y_full[IX_test]\n            mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n            if verbose >= 100:\n                print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n            list_mrrmse.append(mrrmse); list_r2.append(r2)\n\n\n        Y_oof_pred_blend =  ( Y_oof_pred_blend * i_self_blend + Y_oof_pred ) / (i_self_blend + 1)\n        i_self_blend += 1\n        if verbose >= 10:\n            print('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ),  'Average r2:', np.round(np.mean(list_r2 ),4 ))\n            print()\n        list_result_scores.append( np.mean(list_mrrmse ) )\n\n    if verbose >= 1:\n        print('i_self_blend', i_self_blend)    \n    Y_oof_pred_blend =  Y_oof_pred_blend\n\n    if verbose >= 1:\n        print('Blend scoring:')\n    kf = KFold_custom(CV_scheme, valid_size = valid_size)   \n    list_mrrmse = [];list_r2 = []\n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n        Y_test = Y_full[IX_test]\n        mrrmse = np.sqrt(np.square(Y_test - Y_oof_pred_blend[IX_test]).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_oof_pred_blend[IX_test] ) \n        if verbose >= 1:\n            print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n        list_mrrmse.append(mrrmse); list_r2.append(r2)\n    if verbose >= 1:\n        print('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ),  'Average r2:', np.round(np.mean(list_r2 ),4 ))\n        print()\n\n#     res = pd.Series(index = list_smoothing, data = list_result_scores)\n#     res.index.name = 'smoothing'; res.name = 'mrrmse'\n#     display( res.sort_values().to_frame() )\n#     res.sort_values().to_csv('scores_from_smoothing_first_example.csv')\n\n    df_stat_scores['mrrmse Alpha '+str(alpha)] = list_result_scores\n    \ndf_stat_scores.to_csv('scores_change_alpha_and_smoothing_target_encoder_ridge.csv')    \ndisplay( df_stat_scores    )\ndisplay( df_stat_scores.describe() )\n\nfig = plt.figure(figsize = (15,3) )\nfor col in df_stat_scores.columns:\n    plt.plot(df_stat_scores[col].values , '*-', label = col)\nplt.legend()    \nplt.title('Scores dependence on smoothing. Ridge. TargetEncoder',fontsize = 20 )\nplt.xticks(range(len(list_smoothing)), list_smoothing)  # Set x-axis tick labels\nplt.grid()\nplt.xlabel('Smoothing',fontsize = 20 )\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-24T11:49:13.200239Z","iopub.execute_input":"2023-11-24T11:49:13.200761Z","iopub.status.idle":"2023-11-24T12:12:48.602167Z","shell.execute_reply.started":"2023-11-24T11:49:13.200713Z","shell.execute_reply":"2023-11-24T12:12:48.600769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Final timing","metadata":{"execution":{"iopub.status.busy":"2023-10-25T10:04:59.25663Z","iopub.execute_input":"2023-10-25T10:04:59.25698Z","iopub.status.idle":"2023-10-25T10:04:59.261681Z","shell.execute_reply.started":"2023-10-25T10:04:59.256955Z","shell.execute_reply":"2023-10-25T10:04:59.260221Z"}}},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )\nprint('%.1f minutes passed total '%( (time.time()-t0start)/60)  )\nprint('%.2f hours passed total '%( (time.time()-t0start)/3600)  )","metadata":{"execution":{"iopub.status.busy":"2023-11-24T12:12:48.604447Z","iopub.execute_input":"2023-11-24T12:12:48.605536Z","iopub.status.idle":"2023-11-24T12:12:48.616308Z","shell.execute_reply.started":"2023-11-24T12:12:48.605489Z","shell.execute_reply":"2023-11-24T12:12:48.614766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}