{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nConsider neural network from Kishan's notebook: https://www.kaggle.com/code/liudacheldieva/neural-network-regression/notebook#Predicting-on-test-data (please upvote).\n\n\nAnd (self)-blend it over random seeds many times e.g. 100 and more - see different versions of the notebook. \n\nTypically blend in general and in particular of NN with different random seeds improves the score, unfortunately that not seems to be the case here:\n\n    LB 0.621 - 100 blended random seeds - version 2 \n    LB 0.620 - 2000 blended random seeds - version 6\n\nWhile original was 0.604 , and some follow-up ahieved 0.599: https://www.kaggle.com/code/liudacheldieva/neural-network-regression\nSo seems to be the selected random seed 42 and other randomness in original notebook were lucky.    \n\n#### Updates: CV-schemes scores (AmbrsoM, MT).\n\nIn later versions (from v9) we add analysis of the NN by CV-schemes proposed by AmbrosM and MT. \nWe see that in average over random seeds - one the scores is a bit better than for Ridge baseline, another score a bit worse than for Ridge baseline.\nSee here:  https://www.kaggle.com/alexandervc/op2-class-for-custom-cv-schemes\nA collection of CV (LB) scores for differnt models can be found here: https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing\n\n    Averages over 5 random seeds:  \n    CV-AmbrosM: 0.981833\n    CV-MT: 2.810055\n    CV-Random5: 1.219919\n    \n    For Ridge baseline:\n    CV-AmbrosM:  0.997847\n    CV-MT: 2.660317\n    CV-Random: 1.240134\n\n\nPS \n\nFor  some analysis we compute average correlation between columns of the predictions - blended and individual. \nIt shows how different individual predictions are depending on the random seeds. The average correlation is 0.872029, with some outliers below 0.75. \nThe typical straregy for blend is to try to select most uncorrelated predictions, but we also need to know that their scores are good, that is problematic - we do not have so many submits, while local validation may not fully correspond to LB.\n\n\n####  Notes:\n\n    Kishan's NN is predicting all 18211 targets directly  NOT using pca/tsvd reduction of targets as most other approaches.\n    It is trained with \"mae\" loss, not typically mmse, may be it helps woth outliers or in general ill-distributedness of targets \n    It contains 6 layers with shrinking layers: 256->128->64->32->16->final=18211 , with Batchnorms, droupouts (0.2,0.1), relu activations. Trained with Adam\n    \n","metadata":{"execution":{"iopub.status.busy":"2023-10-22T20:04:46.417569Z","iopub.execute_input":"2023-10-22T20:04:46.417989Z","iopub.status.idle":"2023-10-22T20:04:46.449994Z","shell.execute_reply.started":"2023-10-22T20:04:46.417956Z","shell.execute_reply":"2023-10-22T20:04:46.448951Z"}}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport time\nt0start = time.time() \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport pandas as pd\nimport tensorflow as tf\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-10-24T19:44:03.148286Z","iopub.execute_input":"2023-10-24T19:44:03.148748Z","iopub.status.idle":"2023-10-24T19:44:14.355057Z","shell.execute_reply.started":"2023-10-24T19:44:03.148714Z","shell.execute_reply":"2023-10-24T19:44:14.353942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data","metadata":{}},{"cell_type":"code","source":"%%time\nde_train =   pd.read_parquet(\"/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet\")\nid_map = pd.read_csv(\"/kaggle/input/open-problems-single-cell-perturbations/id_map.csv\")\ndf_de_train = de_train; df_id_map = id_map # for compatibility between parts borrowed from different noteboks\nsample_submission = pd.read_csv(\"/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv\")\nfeatures_columns = [\"cell_type\", \"sm_name\"]\nlabels_columns=[\"cell_type\",\"sm_name\",\"sm_lincs_id\",\"SMILES\",\"control\"]\nlabels = de_train.drop(columns=labels_columns)\nfeatures = pd.DataFrame(de_train, columns=features_columns)\nfeatures","metadata":{"execution":{"iopub.status.busy":"2023-10-24T19:45:10.726406Z","iopub.execute_input":"2023-10-24T19:45:10.726932Z","iopub.status.idle":"2023-10-24T19:45:17.203953Z","shell.execute_reply.started":"2023-10-24T19:45:10.726891Z","shell.execute_reply":"2023-10-24T19:45:17.202441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Onehot encoding","metadata":{}},{"cell_type":"code","source":"%%time\n\ntest_data = pd.DataFrame(id_map, columns=features_columns)\n\nfrom sklearn.preprocessing import OneHotEncoder\n\n# Create an instance of the encoder\nencoder = OneHotEncoder()\n\n# Fit the encoder on features\nencoder.fit(features)\n\n# Transform the features into one-hot encoded format\none_hot_encode_features = encoder.transform(features)\n\n# Transform the test data(id_map)\none_hot_test = encoder.transform(test_data)\n\none_hot_encode_features.toarray().shape, one_hot_test.toarray().shape\nX_full = one_hot_encode_features.toarray()\nY_full = labels.values\nX_submit = one_hot_test.toarray()\nprint(X_full.shape, X_submit.shape ,  Y_full.shape)","metadata":{"execution":{"iopub.status.busy":"2023-10-24T19:44:22.259639Z","iopub.execute_input":"2023-10-24T19:44:22.260396Z","iopub.status.idle":"2023-10-24T19:44:22.280257Z","shell.execute_reply.started":"2023-10-24T19:44:22.260358Z","shell.execute_reply":"2023-10-24T19:44:22.279373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# NN train and blend","metadata":{}},{"cell_type":"code","source":"%%time\nfrom tensorflow.keras.layers import Dense, Dropout, BatchNormalization, Activation\nfrom tensorflow.keras.models import Sequential\n\nfrom sklearn.preprocessing import StandardScaler\nscaler = StandardScaler()\n\nn_blends = 10\n\n#dict_Y_submits = {}\ni_blend = 0\nlist_average_corr_coefs_blend_to_individual_prediction = [] # for some analysis\nfor rand_seed in range(42,42+n_blends): # [42,43,44,45]:\n    tf.random.set_seed(rand_seed)\n    # clone model 2\n    model_4 = Sequential([ \n        Dense(256),\n        BatchNormalization(),\n        Activation(\"relu\"),\n        Dropout(0.2),\n        Dense(128, activation=\"relu\"),\n        Dropout(0.2),\n        Dense(64, activation=\"relu\"),\n        BatchNormalization(),\n        Dropout(0.2),\n        Dense(32, activation=\"relu\"),\n        Dropout(0.2),\n        Dense(16, activation=\"relu\"),\n        Dropout(0.1),\n        Dense(18211,activation= \"linear\")\n    ])\n\n    model_4.compile(loss=\"mae\", \n                    optimizer=tf.keras.optimizers.Adam(),\n                    metrics=[\"mae\"])\n    model_4.fit(X_full, Y_full, epochs=50, verbose  = None )\n    Y_submit = model_4.predict(X_submit, batch_size=1)\n    if i_blend == 0:\n        Y_submit_blend = Y_submit\n    else:\n        Y_submit_blend = (Y_submit_blend *i_blend +  Y_submit)/(i_blend + 1)\n    i_blend += 1\n        \n    # dict_Y_submits[rand_seed] = Y_submit\n    \n    # Analysis\n    c = np.mean( np.sum( scaler.fit_transform(Y_submit_blend)*scaler.fit_transform(Y_submit), axis = 0) )/len(Y_submit)# average correlation coefficient, fast way\n    list_average_corr_coefs_blend_to_individual_prediction.append(c)\n    if (i_blend%10 == 0) or (i_blend <10):\n        print(i_blend, rand_seed, np.round(c,3))","metadata":{"execution":{"iopub.status.busy":"2023-10-23T06:57:36.518602Z","iopub.execute_input":"2023-10-23T06:57:36.519033Z","iopub.status.idle":"2023-10-23T07:04:08.521328Z","shell.execute_reply.started":"2023-10-23T06:57:36.518995Z","shell.execute_reply":"2023-10-23T07:04:08.519901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Y_submit_blend.nbytes/1e6","metadata":{"execution":{"iopub.status.busy":"2023-10-23T07:04:08.523282Z","iopub.execute_input":"2023-10-23T07:04:08.524194Z","iopub.status.idle":"2023-10-23T07:04:08.532181Z","shell.execute_reply.started":"2023-10-23T07:04:08.524148Z","shell.execute_reply":"2023-10-23T07:04:08.530784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submit prepare","metadata":{}},{"cell_type":"code","source":"%%time\nsample_columns = sample_submission.columns\nsample_columns= sample_columns[1:]\n\nsubmission_df = pd.DataFrame(Y_submit_blend, columns=sample_columns)\nsubmission_df.index.name = 'id'\nsubmission_df.to_csv('submission_NN_blends_'+str(n_blends)+'.csv')\ndisplay(submission_df)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analysis","metadata":{}},{"cell_type":"code","source":"plt.hist(list_average_corr_coefs_blend_to_individual_prediction, bins = 100)\nplt.title('Correlation of  N-th prediction with blend of all previous', fontsize = 20 )\nplt.show()\nsr = pd.Series(list_average_corr_coefs_blend_to_individual_prediction)\nprint( sr.describe() )\nsr.to_csv('list_average_corr_coefs_blend_to_individual_prediction.csv')\n\nplt.plot( list_average_corr_coefs_blend_to_individual_prediction )\nplt.title('Correlation of  N-th prediction with blend of all previous', fontsize = 20 )\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-24T19:46:01.539083Z","iopub.execute_input":"2023-10-24T19:46:01.539539Z","iopub.status.idle":"2023-10-24T19:46:01.597671Z","shell.execute_reply.started":"2023-10-24T19:46:01.539506Z","shell.execute_reply":"2023-10-24T19:46:01.595632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CV schemes class\n\nSee https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes#Custom-CV-(AmbrosM,-MT,-etc-)\n","metadata":{}},{"cell_type":"code","source":"import random \nfrom sklearn.model_selection import KFold\n\n# Prepare indexes for AmbrosM scheme:\n# For AmbrosM scheme: \nfolds_index_data_AmbrosM = [ ]\ntrain_sm_names = ['Idelalisib', 'Crizotinib', 'Linagliptin', 'Palbociclib', 'Dabrafenib', 'Alvocidib', 'LDN 193189', 'R428', 'Porcn Inhibitor III', \n  'Belinostat', 'Foretinib', 'MLN 2238', 'Penfluridol', 'Dactolisib', 'O-Demethylated Adapalene', 'Oprozomib (ONX 0912)', 'CHIR-99021']\nlist_fold_ids =  ['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells'] \nfor fold_id in list_fold_ids:\n    mask_va = (df_de_train.cell_type == fold_id) & ~df_de_train.sm_name.isin(train_sm_names)\n    mask_tr = ~mask_va # 485 or 487 training rows\n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_AmbrosM.append( [IX_train, IX_test ])\n\n# For MT CV scheme:\nfolds_index_data_MT = []\nfold_to_compounds = {0: ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189',  'Linagliptin', 'O-Demethylated Adapalene'],\n 1: ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III'],\n 2: ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol',  'R428']}\nfor fold_id in [0,1,2]:\n    mask_va = df_de_train['cell_type'].isin(['Myeloid cells', 'B cells']) & df_de_train['sm_name'].isin(fold_to_compounds[fold_id])\n    mask_tr = ~mask_va    \n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_MT.append( [IX_train, IX_test ])    \n    \nclass KFold_custom:\n    '''\n    Class with similar to sklearn \"KFold\" class, interface.\n    Support of CV schemes proposed by AmbrosM, MT, Kishan and standard sklearn KFold\n    Supports options to return IX_train,IX_valid,IX_test - triplet - if valid_size > 0\n    Examples:\n    kf = KFold_custom('AmbrosM'):     \n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n        pass\n    kf = KFold_custom('AmbrosM', valid_size = 0.1 )     \n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split()) :\n        pass\n    '''\n    def __init__(self, CV_scheme,  valid_size = 0, n_splits=5, random_state = 42, verbose = 0 ):\n        self.CV_scheme = CV_scheme\n        self.CV_scheme_inf =  CV_scheme\n        self.valid_size = np.clip(valid_size,0,1)\n        self.random_state = random_state\n\n        self.folds_index_data = [ [np.arange(0), np.arange(0), np.arange(0)] ] # Example data\n        self.n_splits = 1 # Example data\n        if CV_scheme == 'Kishan1':\n            # One fold scheme which contains train, valid, test parts \n            IX_train_Kishan1 = np.arange(429)\n            IX_val_Kishan1 = np.arange(429,521)\n            IX_test_Kishan1 = np.arange(521,614)\n            self.n_splits = 1\n            self.folds_index_data =  [   [IX_train_Kishan1, IX_val_Kishan1, IX_test_Kishan1]  ]\n        elif CV_scheme == 'AmbrosM':\n            self.n_splits = 4\n            self.folds_index_data = folds_index_data_AmbrosM\n        elif CV_scheme == 'MT':\n            self.n_splits = 3\n            self.folds_index_data = folds_index_data_MT\n        elif 'Random'.lower() in CV_scheme.lower():\n            self.n_splits = n_splits\n            self.CV_scheme_inf = 'Random_'+str(n_splits) +'_'+str(random_state )\n            kf = KFold(n_splits=n_splits, random_state = random_state, shuffle=True )#,\n            self.folds_index_data = list( kf.split( np.arange(614) ) )\n            \n        else:\n            s = 'Uncrecognized CV_scheme ' + str(CV_scheme) \n            raise ValueError(s)\n\n        self.verbose = verbose \n        if verbose >= 10:\n            print(self.CV_scheme, self.n_splits, len(self.folds_index_data) )\n            #print(self.folds_index_data)\n            \n    def split(self, X=None):\n        '''\n        X - NOT used, just for compatibility with sklearn \n        '''\n        for item in self.folds_index_data:\n            if self.CV_scheme in ['Kishan1']:\n                if self.valid_size > 0:\n                    yield item[0],item[1],item[2]\n                else :\n                    yield np.array( list(set(item[0])|set(item[1]))  ) , item[2] # Return train, test only \n            elif self.CV_scheme in ['AmbrosM','MT', 'Random']:\n                if self.valid_size == 0:\n                    yield item[0],item[1]\n                else:\n                    # Split \"full-train\"->( real-train, valid ) \n                    index_for_real_train_part = int(  (1-self.valid_size) * len(item[0]) )\n                    if self.random_state is None:\n                        IX_train =  item[0][:index_for_real_train_part]\n                        IX_valid =  item[0][index_for_real_train_part:]\n                    elif self.random_state == -1:\n                        p = np.random.permutation(len(item[0])) \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    else:\n                        np.random.seed(self.random_state) # Set temporary random seed\n                        p = np.random.permutation(len(item[0])) \n                        np.random.seed(random.randint(0,1000)) # Randomize seed again - use Python random, not numpy       \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    yield IX_train, IX_valid, item[1]\n                    \n        \n    def get_list_main_CV_schemes(self):\n        return ['AmbrosM','MT','Kishan1','Random']\n    def get_n_splits(self, X=np.array([])):\n        return self.n_splits\n        \n    \nprint();print();    \nprint('--------------------------------Examples-----------------------------------------------')\nprint();print();    \n        \nCV_scheme = 'Kishan1'\nprint(CV_scheme,  )\nkf = KFold_custom('Kishan1', valid_size = 0.18)     \n#list( kf.split() )\nprint(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test)  \" )\nfor i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split(None)) :\n    print(i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test) )\nprint()\n\nfor CV_scheme in ['AmbrosM','MT','Random']:\n    valid_size = 0\n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size)     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train),type(IX_test)  \" )\n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train), len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train), type(IX_test) )\n    \n    print()\n    valid_size = 0.2 \n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size )     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test)  \" )\n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test) )\n    print()","metadata":{"execution":{"iopub.status.busy":"2023-10-24T19:45:36.885565Z","iopub.execute_input":"2023-10-24T19:45:36.886083Z","iopub.status.idle":"2023-10-24T19:45:36.951291Z","shell.execute_reply.started":"2023-10-24T19:45:36.886046Z","shell.execute_reply":"2023-10-24T19:45:36.949770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Encoding (2)","metadata":{}},{"cell_type":"code","source":"%%time\n\nfrom sklearn.preprocessing import OneHotEncoder\n\nfeatures_columns = ['cell_type','sm_name']\nselected_features_columns = [ 'cell_type', 'sm_name']\nX_submit_categorical = df_id_map[selected_features_columns] # pd.DataFrame(df_id_map, columns= features_columns )\nX_full_categorical = df_de_train[ selected_features_columns ]\n\n\n# Create an instance of the encoder\nencoder = OneHotEncoder()\n\n\nX_full = encoder.fit_transform(X_full_categorical).toarray()\nX_submit = encoder.transform(X_submit_categorical).toarray()\n\nY_full = df_de_train.iloc[:,5:].values\n\nprint('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-24T19:50:37.520051Z","iopub.execute_input":"2023-10-24T19:50:37.520467Z","iopub.status.idle":"2023-10-24T19:50:37.592791Z","shell.execute_reply.started":"2023-10-24T19:50:37.520436Z","shell.execute_reply":"2023-10-24T19:50:37.591596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Preliminaries for modeling","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.layers import Dense, Dropout, BatchNormalization, Activation\nfrom tensorflow.keras.models import Sequential\n\ndef get_model( rand_seed) :\n    tf.random.set_seed(rand_seed)\n    # clone model 2\n    model_4 = Sequential([ \n        Dense(256),\n        BatchNormalization(),\n        Activation(\"relu\"),\n        Dropout(0.2),\n        Dense(128, activation=\"relu\"),\n        Dropout(0.2),\n        Dense(64, activation=\"relu\"),\n        BatchNormalization(),\n        Dropout(0.2),\n        Dense(32, activation=\"relu\"),\n        Dropout(0.2),\n        Dense(16, activation=\"relu\"),\n        Dropout(0.1),\n        Dense(18211,activation= \"linear\")\n    ])\n\n    model_4.compile(loss=\"mae\", \n                    optimizer=tf.keras.optimizers.Adam(),\n                    metrics=[\"mae\"])\n    \n    return model_4","metadata":{"execution":{"iopub.status.busy":"2023-10-24T20:00:47.030985Z","iopub.execute_input":"2023-10-24T20:00:47.031605Z","iopub.status.idle":"2023-10-24T20:00:47.046944Z","shell.execute_reply.started":"2023-10-24T20:00:47.031554Z","shell.execute_reply":"2023-10-24T20:00:47.045253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CV Modeling","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.metrics import r2_score\n\nverbose = 100\n# CV_scheme = 'AmbrosM'\nlist_CV_scheme =  ['AmbrosM','MT','Random']\nvalid_size = 0\nprint(CV_scheme, 'valid_size', valid_size )\nkf = KFold_custom(CV_scheme, valid_size = valid_size)     \nlist_prms = list(range(42, 72))# [ 42,43,44,45,46 ] # rand_seed = 42\n\ndf_stat = pd.DataFrame();\n\nfor CV_scheme in  list_CV_scheme:\n    kf = KFold_custom(CV_scheme, valid_size = valid_size) \n    print();print('CV_scheme', CV_scheme)\n    list_mrrmse = []; list_r2 = []\n    IX_stat = 0;\n    for rand_seed in list_prms:\n        print(rand_seed)\n        list_mrrmse_folds = []; list_r2_folds = []    \n        for i_fold,(IX_train,IX_test)   in enumerate( kf.split(None)) :\n            if verbose>= 1000:\n                print(i_fold, len(IX_train), len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train), type(IX_test) )\n            X_train,X_test = X_full[IX_train], X_full[IX_test]\n            Y_train,Y_test = Y_full[IX_train], Y_full[IX_test]\n\n            model_4 = get_model(rand_seed)\n            model_4.fit(X_train, Y_train, epochs=50, verbose  = None )\n            Y_test_pred = model_4.predict(X_test, batch_size=1)    \n            mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n            print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  CV_scheme )\n            list_mrrmse_folds.append( mrrmse ); list_r2_folds.append(r2)\n\n            df_stat.loc[IX_stat,'seed' ] = rand_seed\n            df_stat.loc[IX_stat,'mrrmse fold '+str(i_fold) +  ' ' + CV_scheme] = mrrmse\n            df_stat.loc[IX_stat,'r2 fold '+str(i_fold) + ' ' + CV_scheme] = r2\n            \n        IX_stat += 1\n        list_mrrmse.append(np.mean( list_mrrmse_folds )); list_r2.append(np.mean( list_r2_folds ))\n        print('Average:', 'mrrmse:', np.round(np.mean(list_mrrmse_folds),4), 'r2:',np.round(np.mean(list_r2_folds),4)  )\n    ","metadata":{"execution":{"iopub.status.busy":"2023-10-24T20:10:22.823779Z","iopub.execute_input":"2023-10-24T20:10:22.824773Z","iopub.status.idle":"2023-10-24T20:23:33.995786Z","shell.execute_reply.started":"2023-10-24T20:10:22.824734Z","shell.execute_reply":"2023-10-24T20:23:33.994324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Example one launch gave that (AmbrosM) scheme:\n#print('Average:', 'mrrmse:', np.round(np.mean(list_mrrmse_folds),4), 'r2:',np.round(np.mean(list_r2_folds),4)  )\n#Average: mrrmse: 0.9773 r2: -0.0191\n ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_stat_save  = df_stat.copy()\ndf_stat.to_csv('df_stat.csv')\ndisplay( df_stat.describe() )\ndisplay( df_stat )","metadata":{"execution":{"iopub.status.busy":"2023-10-24T20:23:55.489234Z","iopub.execute_input":"2023-10-24T20:23:55.489680Z","iopub.status.idle":"2023-10-24T20:23:55.611474Z","shell.execute_reply.started":"2023-10-24T20:23:55.489646Z","shell.execute_reply":"2023-10-24T20:23:55.610160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nstr_prm_name = 'seed'\n# list_prms\ndf_stat = df_stat_save.copy()\nstr_model_inf = 'NN (Kishans); one-hot'\nfor score_name in ['mrrmse','r2']:\n    print();print(); print('Score:', score_name ); print()\n    for CV_scheme in list_CV_scheme:\n        list_selected_cols =[ col for col in df_stat.columns if (score_name in col) and (CV_scheme in col) ]\n        col_average_name = score_name + ' ' + CV_scheme   + ' ' +  'average'\n        df_stat[col_average_name] = df_stat[ list_selected_cols ].mean(axis = 1)\n        list_selected_cols = [ col_average_name ] + list_selected_cols\n        plt.figure(figsize = (20,4))\n        \n        str_inf = 'Best params: '\n        for col in list_selected_cols:\n            v = df_stat[col]\n            #plt.plot(np.log10(list_alpha), df_stat[col] ,'*-' , label = col)\n            plt.plot(list_prms, df_stat[col] ,'*-' , label = col)\n            #print('Best alpha:', df_stat.sort_values(col)['alpha'].iat[0], 'for ', col )\n            if score_name in ['mrrmse']:\n                str_inf += str(df_stat.sort_values(col)[str_prm_name].iat[0]) + ' for ' + col  + ' ; '\n            else:\n                str_inf += str(df_stat.sort_values(col,ascending = False)[str_prm_name].iat[0]) + ' for ' + col  + ' ; '\n                \n        print(str_inf)\n        #plt.xlabel('LOG10 alpha',fontsize = 20 )\n        plt.xlabel( str_prm_name ,fontsize = 20 )\n        plt.legend(fontsize = 12 )\n        plt.title('scores: '  + score_name + ', CV: ' + CV_scheme + ', Model: ' + str_model_inf, fontsize = 20 )\n        plt.grid()\n        plt.show()\n\n\n# df_stat","metadata":{"execution":{"iopub.status.busy":"2023-10-24T20:24:12.122957Z","iopub.execute_input":"2023-10-24T20:24:12.123450Z","iopub.status.idle":"2023-10-24T20:24:14.583413Z","shell.execute_reply.started":"2023-10-24T20:24:12.123416Z","shell.execute_reply":"2023-10-24T20:24:14.582171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\npd.set_option('display.max_columns', 500)\npd.set_option('display.max_rows', 500)\n\nprint( df_stat.shape )\nlist_cols1 = [t for t in df_stat.columns if ('average' in t) or (str_prm_name == t) ]\nlist_cols2 = [t for t in df_stat.columns if 'average' not in t ]\ndf_stat = df_stat[list_cols1+list_cols2]\ndisplay( df_stat[list_cols1+list_cols2] )\ndisplay( df_stat[list_cols1+list_cols2].describe() )\n\ndf_stat.to_csv('df_stat.csv')","metadata":{"execution":{"iopub.status.busy":"2023-10-24T20:24:55.670008Z","iopub.execute_input":"2023-10-24T20:24:55.670445Z","iopub.status.idle":"2023-10-24T20:24:55.835565Z","shell.execute_reply.started":"2023-10-24T20:24:55.670407Z","shell.execute_reply":"2023-10-24T20:24:55.834233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Final timing","metadata":{}},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )\nprint('%.1f minutes passed total '%( (time.time()-t0start)/60)  )\nprint('%.2f hours passed total '%( (time.time()-t0start)/3600)  )","metadata":{"execution":{"iopub.status.busy":"2023-10-23T07:04:19.986008Z","iopub.execute_input":"2023-10-23T07:04:19.986631Z","iopub.status.idle":"2023-10-23T07:04:19.996385Z","shell.execute_reply.started":"2023-10-23T07:04:19.986578Z","shell.execute_reply":"2023-10-23T07:04:19.994560Z"},"trusted":true},"execution_count":null,"outputs":[]}]}