{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":59094,"databundleVersionId":7010844,"sourceType":"competition"},{"sourceId":5835808,"sourceType":"datasetVersion","datasetId":3354626}],"dockerImageVersionId":30559,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Neural Network based on NLP-like SMILES embedding\n\nHere is a componet of U900 team (#13) solution for the \"Open Problems – Single-Cell Perturbations\" challenge.\n\nSee main writeup here:  https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/460858 \n\nThe present notebook implements the part denoted \"Neural Network based on NLP-like SMILES embedding\" in the write-up.\n\nIt achieves 0.574 public score (0.766 private) - version 11 of the present notebook. (To test stability we rerun it 3 times, averaged and made another submit - again achieving 0.574 and even bettter 0.762 private [(version 1 here)](https://www.kaggle.com/code/alexandervc/op2-u900-team-blend?scriptVersionId=152310281)). \nThe solution originates from the [public one](https://www.kaggle.com/code/kishanvavdara/nlp-regression) by Kishan Vavdara though substantially reworked from architectural and training points of view uplifting the score from 0.607 (original) to 0.574. The model has been mainly developed by Dmitry Rudenko  and  further improved by Alexander Chervov. \n\n\n\n### Highlights:\n\n- Lion - new powerful optimizer - outperformed Adam\n- Magic (simple) train duplicating trick improved score 0.582 -> 0.574\n- SMILES encoding by the embedding layer\n\n### Modeling organization :\n\n- Feature Encoding: SMILES - by Embedding layer, Cell Types - One-hot; both concatenated\n- Architecture: 5-Layer (1558,512,256, 128, 256, 18211) Perceptron with carefully chosen Batchnorm and Dropout layers positions, activation: “elu”\n- Preprocessing: Standard Scaler for targets, Add Gaussian Noise for features\n- Training: Lion optimizer;  loss: competition loss - MRRMSE (custom); 5 almost random folds; best (by validation score) epoch (out of 300) is restored for each fold - that appears to be quite important\n- Prediction scheme: 18211 targets directly, (TSVD - not used at all)\n- The trick with duplicating the train for each fold yields 0.582->0.574 uplift. Similar to our other NN models.\n- Tuning: params were optimized by CV\n\n### What did not worked well for that version of the NN:\n- SMILES augmentation package\n- LSTM/CNN architectures, other optimizers, pseudolabeling, dropping out noisy samples.\n- The trick to retrain model on the entire train also was not successful for that NN (in contrast to other models) because the epoch number determined by early stopping was different from fold to fold and fixing it to some particular number - degraded the CV-scores and so we did not want to risk employing the models not having good CV scores. Spent quite a lot efforts to resolve it, but unsuccessful.\n\n### Model \n\nThe model is speciefed in the section [\"The model\"](https://www.kaggle.com/code/alexandervc/nlp-regression-custom-kfold-update1#The-model) below, the [figure](https://www.kaggle.com/code/alexandervc/nlp-regression-custom-kfold-update1?scriptVersionId=154711440&cellId=64) shows the architecture. \n\nSo the model organized as follows: SMILES are encoded via the embedding layer and Cell Type via one-hot; both encodings concatenated; that followed by 5-dense-layers perceptron (1558,512,256, 128, 256, 18211) carefully interchanged with batchnorm and dropout layers; activation is “elu”.\nWe checked the stability of the model as follows. Rerun it with several times with similar params compare CV scores and submit. We stably observed similar CV scores and moreover LB scores at the range 0.581-0.583 - before adding train duplication trick and 0.574 after. That is quite in contrast to the original model - which has larger score variance: 0.600 - 0.617 at least (see experiments here ).\n\n### Notebook orgnization\n\nThe notebook contains an option to tune model params. One can run several models sequentially and compare the scores. Launching the model(s) managed in the seciton: \"Main modeling\".\n\n- First section devoted to preparations (install,import, data load, etc)\n- Section \"The model\" specifies and visualizes the model\n- Section \"Core modeling function - k_fold_predict\" contains the key modeling function - loop via folds, statistic accumulation, etc \n- Section \"Main modeling\" - key launcher section, here model(s) params specified and modelling launched (optinally several modeling sequentially to tune the params)\n- Section \"Show scores, stat, etc\" shows final scoring for all models lauched \n\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"!pip install -U tensorflow==2.14.0","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:03:28.089977Z","iopub.execute_input":"2023-12-12T14:03:28.090477Z","iopub.status.idle":"2023-12-12T14:04:30.937302Z","shell.execute_reply.started":"2023-12-12T14:03:28.090440Z","shell.execute_reply":"2023-12-12T14:04:30.936371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\n\nimport time\nt0start = time.time() \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport pandas as pd\nimport tensorflow as tf\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\ni = 0\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        i += 1\n        if i <= 15:\n            print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:30.939315Z","iopub.execute_input":"2023-12-12T14:04:30.939622Z","iopub.status.idle":"2023-12-12T14:04:37.108791Z","shell.execute_reply.started":"2023-12-12T14:04:30.939595Z","shell.execute_reply":"2023-12-12T14:04:37.107654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.__version__","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:37.110308Z","iopub.execute_input":"2023-12-12T14:04:37.110969Z","iopub.status.idle":"2023-12-12T14:04:37.118385Z","shell.execute_reply.started":"2023-12-12T14:04:37.110933Z","shell.execute_reply":"2023-12-12T14:04:37.117365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read data","metadata":{}},{"cell_type":"code","source":"data = pd.read_parquet(\"/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet\")\n\n# Same:\nfn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\ndf_de_train = pd.read_parquet(fn)# , index_col = 0)\nprint(df_de_train.shape)\ndisplay(df_de_train )\n\n\nid_map = pd.read_csv(\"/kaggle/input/open-problems-single-cell-perturbations/id_map.csv\")\nsample_submission = pd.read_csv(\"/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv\")\nsample_columns = sample_submission.columns\nsample_columns = sample_columns[1:]\nRNDST1=271828\nRNDST2=314159","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:37.121233Z","iopub.execute_input":"2023-12-12T14:04:37.121792Z","iopub.status.idle":"2023-12-12T14:04:49.062930Z","shell.execute_reply.started":"2023-12-12T14:04:37.121758Z","shell.execute_reply":"2023-12-12T14:04:49.062142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" **🔑Tip:** Stare at the data for a while to get some insights and ideas to run experiments.","metadata":{}},{"cell_type":"code","source":"data.head(20)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:49.063995Z","iopub.execute_input":"2023-12-12T14:04:49.064252Z","iopub.status.idle":"2023-12-12T14:04:49.103592Z","shell.execute_reply.started":"2023-12-12T14:04:49.064228Z","shell.execute_reply":"2023-12-12T14:04:49.102721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Aux functions ","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.callbacks import ModelCheckpoint\n\ndef create_model_checkpoint(filepath, monitor='val_mae', save_best_only=True,\n                            save_weights_only=True, mode='auto', verbose=0):\n    \"\"\"\n    Create a ModelCheckpoint callback for saving the best model weights during training.\n\n    Args:\n        filepath (str): Filepath to save the best weights.\n        monitor (str): Metric to monitor (e.g., 'val_loss' or 'val_mae').\n        save_best_only (bool): Save only the best weights.\n        save_weights_only (bool): Save only the model's weights, not the entire model.\n        mode (str): One of {'auto', 'min', 'max'}. In 'min' mode, it saves when the monitored metric decreases.\n        verbose (int): Verbosity mode. 0 = silent, 1 = progress bar, 2 = one line per epoch.\n\n    Returns:\n        keras.callbacks.ModelCheckpoint: ModelCheckpoint callback.\n    \"\"\"\n    checkpoint = ModelCheckpoint(\n        filepath=filepath,\n        monitor=monitor,\n        save_best_only=save_best_only,\n        save_weights_only=save_weights_only,\n        mode=mode,\n        verbose=verbose\n    )\n    return checkpoint\n","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:49.104620Z","iopub.execute_input":"2023-12-12T14:04:49.104869Z","iopub.status.idle":"2023-12-12T14:04:49.131172Z","shell.execute_reply.started":"2023-12-12T14:04:49.104847Z","shell.execute_reply":"2023-12-12T14:04:49.130209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_training_history(history, metrics):\n    \"\"\"\n    Plot training history curves for loss and evaluation metrics on the same line.\n\n    Args:\n        history (keras.callbacks.History): Training history object.\n        metrics (list): List of metric names to plot.\n\n    Returns:\n        None\n    \"\"\"\n    loss = history.history['loss']\n    val_loss = history.history['val_loss']\n\n    epochs = range(len(loss))\n\n    plt.figure(figsize=(12, 6))\n\n    # Plot loss\n    plt.subplot(1, 2, 1)\n    plt.plot(epochs, loss, label='Training Loss', color=\"blue\")\n    plt.plot(epochs, val_loss, label='Validation Loss', color=\"red\")\n    plt.title('Loss')\n    plt.xlabel('Epochs')\n    plt.legend()\n\n    # Plot specified evaluation metrics on the same line\n    for metric in metrics:\n        train_metric_name = f'Training {metric.capitalize()}'\n        val_metric_name = f'Validation {metric.capitalize()}'\n        train_metric = history.history[metric]\n        val_metric = history.history['val_' + metric]\n\n        plt.subplot(1, 2, 2)\n        plt.plot(epochs, train_metric, label=train_metric_name, color=\"green\")\n        plt.plot(epochs, val_metric, label=val_metric_name, color=\"orange\")\n\n    plt.title('Metrics')\n    plt.xlabel('Epochs')\n    plt.legend(loc='upper right')\n\n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:49.132297Z","iopub.execute_input":"2023-12-12T14:04:49.132576Z","iopub.status.idle":"2023-12-12T14:04:49.141468Z","shell.execute_reply.started":"2023-12-12T14:04:49.132553Z","shell.execute_reply":"2023-12-12T14:04:49.140519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def add_columns(data, id_map):\n    sm_name_to_smiles = data.set_index('sm_name')['SMILES'].to_dict()\n    sm_lincs_id = data.set_index('sm_name')[\"sm_lincs_id\"].to_dict()\n\n    id_map['SMILES'] = id_map['sm_name'].map(sm_name_to_smiles)\n    id_map['sm_lincs_id'] = id_map['sm_name'].map(sm_lincs_id)\n\n    return id_map","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:49.142382Z","iopub.execute_input":"2023-12-12T14:04:49.142624Z","iopub.status.idle":"2023-12-12T14:04:49.156056Z","shell.execute_reply.started":"2023-12-12T14:04:49.142602Z","shell.execute_reply":"2023-12-12T14:04:49.155302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\ndf_de_train = pd.read_parquet(fn)# , index_col = 0)\nprint(df_de_train.shape)\ndisplay(df_de_train )","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:49.157535Z","iopub.execute_input":"2023-12-12T14:04:49.157849Z","iopub.status.idle":"2023-12-12T14:04:50.357972Z","shell.execute_reply.started":"2023-12-12T14:04:49.157819Z","shell.execute_reply":"2023-12-12T14:04:50.357004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Prepare indexes for AmbrosM scheme:\n# For AmbrosM scheme: \nfolds_index_data_AmbrosM = [ ]\ntrain_sm_names = ['Idelalisib', 'Crizotinib', 'Linagliptin', 'Palbociclib', 'Dabrafenib', 'Alvocidib', 'LDN 193189', 'R428', 'Porcn Inhibitor III', \n  'Belinostat', 'Foretinib', 'MLN 2238', 'Penfluridol', 'Dactolisib', 'O-Demethylated Adapalene', 'Oprozomib (ONX 0912)', 'CHIR-99021']\nlist_fold_ids =  ['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells'] \nfor fold_id in list_fold_ids:\n    mask_va = (df_de_train.cell_type == fold_id) & ~df_de_train.sm_name.isin(train_sm_names)\n    mask_tr = ~mask_va # 485 or 487 training rows\n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_AmbrosM.append( [IX_train, IX_test ])\n\n# For MT CV scheme:\nfolds_index_data_MT = []\nfold_to_compounds = {0: ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189',  'Linagliptin', 'O-Demethylated Adapalene'],\n 1: ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III'],\n 2: ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol',  'R428']}\nfor fold_id in [0,1,2]:\n    mask_va = df_de_train['cell_type'].isin(['Myeloid cells', 'B cells']) & df_de_train['sm_name'].isin(fold_to_compounds[fold_id])\n    mask_tr = ~mask_va    \n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_MT.append( [IX_train, IX_test ])    \n    \n    \ndict_folds_Antonina = {0: [4, 15, 18, 41, 47, 53, 56, 57, 59, 61, 62, 66, 78, 83, 86, 93, 100, 110, 115, 116, 117, 120, 123, 131, 148, 152, 161, 185, 189, 193, 197, 198, 205, 207, 209, 210, 213, 218, 219, 225, 228, 230, 250, 252, 257, 263, 286, 292, 294, 304, 306, 310, 325, 327, 336, 338, 339, 350, 351, 356, 357, 360, 362, 369, 379, 385, 395, 419, 424, 430, 432, 434, 438, 441, 443, 445, 448, 452, 456, 465, 467, 471, 472, 474, 498, 508, 523, 534, 542, 544, 574, 581, 584, 589, 601, 607], 1: [1, 3, 5, 14, 24, 38, 40, 49, 50, 54, 55, 60, 79, 82, 87, 102, 112, 114, 118, 119, 127, 132, 146, 149, 166, 179, 183, 184, 186, 190, 192, 201, 203, 204, 216, 227, 242, 253, 261, 264, 271, 289, 295, 296, 299, 300, 309, 323, 324, 329, 333, 337, 354, 387, 390, 393, 406, 408, 415, 416, 420, 421, 422, 436, 442, 449, 453, 458, 462, 463, 469, 484, 487, 489, 490, 492, 497, 501, 505, 506, 509, 513, 515, 516, 546, 548, 566, 569, 585, 587, 588, 592, 593, 595, 600, 605], 2: [0, 16, 19, 44, 48, 51, 52, 58, 63, 65, 70, 81, 85, 88, 89, 92, 111, 113, 122, 125, 135, 137, 139, 150, 164, 175, 191, 202, 206, 211, 217, 220, 221, 229, 239, 243, 247, 248, 251, 258, 259, 260, 267, 270, 274, 281, 288, 290, 291, 303, 322, 326, 328, 332, 347, 359, 364, 370, 383, 384, 389, 392, 401, 423, 428, 454, 455, 461, 464, 494, 496, 504, 510, 511, 524, 525, 526, 530, 540, 541, 551, 565, 568, 570, 572, 573, 575, 577, 579, 580, 597, 598, 603, 606, 610, 611, 613], 3: [21, 22, 26, 36, 42, 45, 46, 64, 67, 69, 71, 80, 84, 90, 101, 103, 134, 136, 140, 147, 151, 160, 163, 165, 167, 174, 176, 177, 180, 181, 182, 187, 188, 195, 196, 199, 200, 212, 215, 226, 231, 240, 241, 249, 255, 256, 262, 265, 269, 285, 287, 297, 298, 307, 312, 319, 321, 330, 340, 342, 344, 345, 348, 358, 366, 368, 382, 391, 400, 403, 404, 405, 425, 427, 431, 433, 435, 444, 450, 451, 457, 459, 460, 473, 485, 499, 500, 502, 503, 507, 512, 514, 527, 528, 529, 531, 543, 547, 553, 554, 564, 571, 576, 583, 586, 591, 596, 602, 604, 608, 609], 4: [2, 6, 7, 17, 20, 23, 25, 27, 28, 29, 37, 39, 43, 68, 91, 121, 124, 126, 128, 129, 130, 133, 138, 153, 162, 178, 194, 208, 214, 222, 223, 224, 232, 244, 245, 246, 254, 266, 268, 272, 273, 282, 283, 284, 293, 301, 302, 305, 308, 311, 320, 331, 334, 335, 341, 343, 346, 349, 352, 353, 355, 361, 363, 365, 367, 377, 378, 380, 381, 386, 388, 394, 396, 397, 398, 399, 402, 407, 417, 418, 426, 429, 437, 439, 440, 446, 447, 466, 468, 470, 481, 482, 483, 486, 488, 491, 493, 495, 532, 533, 545, 549, 550, 552, 555, 562, 563, 567, 578, 582, 590, 594, 599, 612], 5: [8, 9, 10, 11, 12, 13, 30, 31, 32, 33, 34, 35, 72, 73, 74, 75, 76, 77, 94, 95, 96, 97, 98, 99, 104, 105, 106, 107, 108, 109, 141, 142, 143, 144, 145, 154, 155, 156, 157, 158, 159, 168, 169, 170, 171, 172, 173, 233, 234, 235, 236, 237, 238, 275, 276, 277, 278, 279, 280, 313, 314, 315, 316, 317, 318, 371, 372, 373, 374, 375, 376, 409, 410, 411, 412, 413, 414, 475, 476, 477, 478, 479, 480, 517, 518, 519, 520, 521, 522, 535, 536, 537, 538, 539, 556, 557, 558, 559, 560, 561]}\n\nfolds_index_data_Antonina = []    \nfor i in range(5):  \n    l_valid = np.array( dict_folds_Antonina[i] )\n    l_train = np.array( [k for k in range(614) if k not in l_valid ] )\n    folds_index_data_Antonina.append( [l_train, l_valid ])    \nprint( len(folds_index_data_Antonina ) )","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.362004Z","iopub.execute_input":"2023-12-12T14:04:50.362276Z","iopub.status.idle":"2023-12-12T14:04:50.420000Z","shell.execute_reply.started":"2023-12-12T14:04:50.362253Z","shell.execute_reply":"2023-12-12T14:04:50.419100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class KFold_custom:\n    '''\n    Class with similar to sklearn \"KFold\" class, interface.\n    Support of CV schemes proposed by AmbrosM, MT, Kishan and standard sklearn KFold\n    Supports options to return IX_train,IX_valid,IX_test - triplet - if valid_size > 0\n    Examples:\n    kf = KFold_custom('AmbrosM'):     \n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n        pass\n    kf = KFold_custom('AmbrosM', valid_size = 0.1 )     \n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split()) :\n        pass\n        \n    kf = KFold_custom(CV_scheme = 'Random') # wrapper for sklearn Kfolds with shuffle = True\n    kf = KFold_custom(CV_scheme = CV_scheme, random_state = 42 ) # fixing random_state ensures reprodicibility of folds \n    '''\n    def __init__(self, CV_scheme,  valid_size = 0, n_splits=5, random_state = 42, verbose = 0 ):\n        self.CV_scheme = CV_scheme\n        self.CV_scheme_inf =  CV_scheme\n        self.valid_size = np.clip(valid_size,0,1)\n        self.random_state = random_state\n\n        self.folds_index_data = [ [np.arange(0), np.arange(0), np.arange(0)] ] # Example data\n        self.n_splits = 1 # Example data\n        if CV_scheme == 'Kishan1':\n            # One fold scheme which contains train, valid, test parts \n            IX_train_Kishan1 = np.arange(429)\n            IX_val_Kishan1 = np.arange(429,521)\n            IX_test_Kishan1 = np.arange(521,614)\n            self.n_splits = 1\n            self.folds_index_data =  [   [IX_train_Kishan1, IX_val_Kishan1, IX_test_Kishan1]  ]\n        elif CV_scheme == 'Tonya':\n            self.n_splits = 5\n            self.folds_index_data = folds_index_data_Antonina\n        elif CV_scheme == 'AmbrosM':\n            self.n_splits = 4\n            self.folds_index_data = folds_index_data_AmbrosM\n        elif CV_scheme == 'MT':\n            self.n_splits = 3\n            self.folds_index_data = folds_index_data_MT\n        elif 'Random'.lower() in CV_scheme.lower():\n            self.n_splits = n_splits\n            self.CV_scheme_inf = 'Random_'+str(n_splits) +'_'+str(random_state )\n            kf = KFold(n_splits=n_splits, random_state = random_state, shuffle=True )#,\n            self.folds_index_data = list( kf.split( np.arange(614) ) )\n        elif 'Full'.lower() in CV_scheme.lower():\n            # Just return the full set as both train and test - it useful to re-train the model on the entire data\n            IX_train_full = np.arange(614)\n            IX_test_full = np.arange(614)\n            self.n_splits = 1\n            self.folds_index_data = [   [IX_train_full, IX_test_full]  ]\n        else:\n            s = 'Uncrecognized CV_scheme ' + str(CV_scheme) \n            raise ValueError(s)\n\n        self.verbose = verbose \n        if verbose >= 10:\n            print(self.CV_scheme, self.n_splits, len(self.folds_index_data) )\n            #print(self.folds_index_data)\n            \n    def split(self, X=None):\n        '''\n        X - NOT used, just for compatibility with sklearn \n        '''\n        for item in self.folds_index_data:\n            if self.CV_scheme in ['Kishan1']:\n                if self.valid_size > 0:\n                    yield item[0],item[1],item[2]\n                else :\n                    yield np.array( list(set(item[0])|set(item[1]))  ) , item[2] # Return train, test only \n            elif self.CV_scheme in ['Tonya', 'AmbrosM','MT', 'Random', 'Full']:\n                if self.valid_size == 0:\n                    yield item[0],item[1]\n                else:\n                    # Split \"full-train\"->( real-train, valid ) \n                    index_for_real_train_part = int(  (1-self.valid_size) * len(item[0]) )\n                    if self.random_state is None:\n                        IX_train =  item[0][:index_for_real_train_part]\n                        IX_valid =  item[0][index_for_real_train_part:]\n                    elif self.random_state == -1:\n                        p = np.random.permutation(len(item[0])) \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    else:\n                        np.random.seed(self.random_state) # Set temporary random seed\n                        p = np.random.permutation(len(item[0])) \n                        np.random.seed(random.randint(0,30000)) # Randomize seed again - use Python random, not numpy       \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    yield IX_train, IX_valid, item[1]\n                    \n        \n    def get_list_main_CV_schemes(self):\n        return ['AmbrosM','MT','Kishan1','Random']\n    def get_n_splits(self, X=np.array([])):\n        return self.n_splits","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.421425Z","iopub.execute_input":"2023-12-12T14:04:50.421936Z","iopub.status.idle":"2023-12-12T14:04:50.442529Z","shell.execute_reply.started":"2023-12-12T14:04:50.421900Z","shell.execute_reply":"2023-12-12T14:04:50.441648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import mean_absolute_error\n\ndef calculate_mae_and_mrrmse(model, data, y_true, scaler=None):\n    \"\"\"\n    Calculate Mean Absolute Error (MAE) and Mean Rowwise Root Mean Squared Error (MRRMSE).\n\n    Parameters:\n    - model: The trained  model.\n    - data: The input data for prediction.\n    - y_true: The true target values.\n    - scaler: The scaler used for data normalization.\n\n    Returns:\n    - None\n    \"\"\"\n    # Predict using the model\n    y_pred_original = model.predict(data, batch_size=1)\n\n    if scaler:\n        \n        y_pred = scaler.inverse_transform(y_pred_original)\n        y_true = scaler.inverse_transform(y_true)\n        # Calculate Mean Absolute Error (MAE)\n        mae = mean_absolute_error(y_true , y_pred)\n\n        # Calculate Mean Rowwise Root Mean Squared Error (MRRMSE)\n        rowwise_rmse = np.sqrt(np.mean(np.square(y_true - y_pred), axis=1))\n        mrrmse_score = np.mean(rowwise_rmse)\n    else:\n       # Calculate Mean Absolute Error (MAE)\n        mae = mean_absolute_error(y_true , y_pred_original)\n\n        # Calculate Mean Rowwise Root Mean Squared Error (MRRMSE)\n        rowwise_rmse = np.sqrt(np.mean(np.square(y_true - y_pred_original), axis=1))\n        mrrmse_score = np.mean(rowwise_rmse)\n    # Print the results\n    print(f\"Mean Absolute Error (MAE): {mae}\")\n    print(f\"Mean Rowwise Root Mean Squared Error (MRRMSE): {mrrmse_score}\")","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.443722Z","iopub.execute_input":"2023-12-12T14:04:50.445846Z","iopub.status.idle":"2023-12-12T14:04:50.590942Z","shell.execute_reply.started":"2023-12-12T14:04:50.445812Z","shell.execute_reply":"2023-12-12T14:04:50.589857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nfrom sklearn.model_selection import KFold, StratifiedKFold\nimport matplotlib.pyplot as plt\nfrom sklearn.metrics import mean_absolute_error\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.layers import Dense, Flatten, concatenate, GaussianNoise, Concatenate\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.callbacks import EarlyStopping\nfrom scipy.stats import t\n\nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\n\ndef train_and_predict(model, X_train_list, y_train, X_val_list, y_val, X_test_list, scaler, epochs, early_stopping=10, restore_best_weights = True ):\n    # Define Early Stopping callback\n    early_stopping_callback = EarlyStopping(monitor='val_loss', patience=early_stopping, restore_best_weights=restore_best_weights ) # True)\n\n    # Train the model with early stopping\n    history = model.fit(X_train_list, y_train, epochs=epochs, verbose=0, \n                        validation_data=(X_val_list, y_val), callbacks=[early_stopping_callback])\n\n    preds_ = model.predict(X_val_list, batch_size=1)\n    # Make predictions on the test set\n    preds = model.predict(X_test_list, batch_size=1)\n    preds = scaler.inverse_transform(preds)\n\n    return history, preds, preds_\n\n\ndef plot_training_history_kfold(histories_and_folds):\n    fig, axes = plt.subplots(nrows=2, ncols=1, figsize=(15, 10))\n\n    for history, fold in histories_and_folds:\n        color = plt.cm.jet(fold / len(histories_and_folds))\n\n        # Plot Training Loss\n        train_loss = history.history['loss']\n        min_train_loss_epoch = np.argmin(train_loss)\n        axes[0].plot(train_loss, label=f'Fold {fold + 1}', color=color)\n        axes[0].axvline(min_train_loss_epoch, linestyle='--', color=color, label=f'Min Train Loss (Fold {fold + 1}, Epoch {min_train_loss_epoch + 1})')\n\n        # Plot Validation Loss\n        val_loss = history.history['val_loss']\n        min_val_loss_epoch = np.argmin(val_loss)\n        axes[1].plot(val_loss, label=f'Fold {fold + 1}', color=color)\n        axes[1].axvline(min_val_loss_epoch, linestyle='--', color=color, label=f'Min Val Loss (Fold {fold + 1}, Epoch {min_val_loss_epoch + 1})')\n\n    axes[0].set_xlabel('Epochs')\n    axes[0].set_ylabel('Training Loss')\n    axes[0].set_title('Training Loss Across Folds')\n    axes[0].legend()\n\n    axes[1].set_xlabel('Epochs')\n    axes[1].set_ylabel('Validation Loss')\n    axes[1].set_title('Validation Loss Across Folds')\n    axes[1].legend()\n\n    plt.tight_layout()\n    plt.show()\n\nimport time    \n","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.592620Z","iopub.execute_input":"2023-12-12T14:04:50.593135Z","iopub.status.idle":"2023-12-12T14:04:50.621128Z","shell.execute_reply.started":"2023-12-12T14:04:50.593099Z","shell.execute_reply":"2023-12-12T14:04:50.620111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preprocessing data ","metadata":{"execution":{"iopub.status.busy":"2023-10-10T09:17:28.606984Z","iopub.execute_input":"2023-10-10T09:17:28.607326Z","iopub.status.idle":"2023-10-10T09:17:28.611868Z","shell.execute_reply.started":"2023-10-10T09:17:28.607301Z","shell.execute_reply":"2023-10-10T09:17:28.610695Z"}}},{"cell_type":"code","source":"# Shuffle data because we use full features for final training \ndata = data.sample(frac=1.0, random_state=RNDST2)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.622591Z","iopub.execute_input":"2023-12-12T14:04:50.622950Z","iopub.status.idle":"2023-12-12T14:04:50.671029Z","shell.execute_reply.started":"2023-12-12T14:04:50.622919Z","shell.execute_reply":"2023-12-12T14:04:50.670030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"add_columns(data, id_map)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.672293Z","iopub.execute_input":"2023-12-12T14:04:50.672600Z","iopub.status.idle":"2023-12-12T14:04:50.763638Z","shell.execute_reply.started":"2023-12-12T14:04:50.672576Z","shell.execute_reply":"2023-12-12T14:04:50.762684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kf = KFold_custom('Kishan1')\nlen([elem for elem in kf.split()])","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.765100Z","iopub.execute_input":"2023-12-12T14:04:50.765442Z","iopub.status.idle":"2023-12-12T14:04:50.772980Z","shell.execute_reply.started":"2023-12-12T14:04:50.765418Z","shell.execute_reply":"2023-12-12T14:04:50.771998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cell_type_feature = pd.DataFrame(data[\"cell_type\"], columns=[\"cell_type\"])\nsmiles_feature = pd.DataFrame(data[\"SMILES\"], columns=[\"SMILES\"])\nlabels = data.drop([\"cell_type\",\"sm_name\",\"sm_lincs_id\",\"SMILES\",\"control\"], axis=1)\n\n# for test\ntest_feature_smiles = pd.DataFrame(id_map[\"SMILES\"], columns=[\"SMILES\"])\ntest_feature_cell_type = pd.DataFrame(id_map[\"cell_type\"], columns=[\"cell_type\"])","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.774039Z","iopub.execute_input":"2023-12-12T14:04:50.774319Z","iopub.status.idle":"2023-12-12T14:04:50.819400Z","shell.execute_reply.started":"2023-12-12T14:04:50.774296Z","shell.execute_reply":"2023-12-12T14:04:50.818605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cell_type_feature.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.820502Z","iopub.execute_input":"2023-12-12T14:04:50.820804Z","iopub.status.idle":"2023-12-12T14:04:50.828591Z","shell.execute_reply.started":"2023-12-12T14:04:50.820779Z","shell.execute_reply":"2023-12-12T14:04:50.827752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"smiles_feature.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.829738Z","iopub.execute_input":"2023-12-12T14:04:50.830000Z","iopub.status.idle":"2023-12-12T14:04:50.842200Z","shell.execute_reply.started":"2023-12-12T14:04:50.829978Z","shell.execute_reply":"2023-12-12T14:04:50.841495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_feature_cell_type.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.843475Z","iopub.execute_input":"2023-12-12T14:04:50.844075Z","iopub.status.idle":"2023-12-12T14:04:50.854412Z","shell.execute_reply.started":"2023-12-12T14:04:50.844043Z","shell.execute_reply":"2023-12-12T14:04:50.853595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_feature_smiles.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.855407Z","iopub.execute_input":"2023-12-12T14:04:50.855674Z","iopub.status.idle":"2023-12-12T14:04:50.870680Z","shell.execute_reply.started":"2023-12-12T14:04:50.855652Z","shell.execute_reply":"2023-12-12T14:04:50.869711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.871683Z","iopub.execute_input":"2023-12-12T14:04:50.871956Z","iopub.status.idle":"2023-12-12T14:04:50.902184Z","shell.execute_reply.started":"2023-12-12T14:04:50.871935Z","shell.execute_reply":"2023-12-12T14:04:50.901519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Normalize labels\nfrom sklearn.preprocessing import StandardScaler\n\nscaler = StandardScaler()\nnorm_label = scaler.fit_transform(labels)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:50.903201Z","iopub.execute_input":"2023-12-12T14:04:50.903478Z","iopub.status.idle":"2023-12-12T14:04:51.435854Z","shell.execute_reply.started":"2023-12-12T14:04:50.903456Z","shell.execute_reply":"2023-12-12T14:04:51.435047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Onehotencode\nfrom sklearn.preprocessing import OneHotEncoder\n\nencoder = OneHotEncoder()\none_hot_celltype = encoder.fit_transform(cell_type_feature)\none_hot_test_cell_type = encoder.transform(test_feature_cell_type)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.437113Z","iopub.execute_input":"2023-12-12T14:04:51.437545Z","iopub.status.idle":"2023-12-12T14:04:51.446270Z","shell.execute_reply.started":"2023-12-12T14:04:51.437513Z","shell.execute_reply":"2023-12-12T14:04:51.445338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\ntrain_cell_type, temp_cell_type, train_labels, temp_labels = train_test_split(one_hot_celltype.toarray(),\n                                                                            norm_label,\n                                                                            test_size=0.3, random_state=RNDST2)\n\nval_cell_type, test_cell_type, val_labels , test_labels = train_test_split(temp_cell_type, temp_cell_type, test_size=0.6, random_state=RNDST2)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.447253Z","iopub.execute_input":"2023-12-12T14:04:51.447534Z","iopub.status.idle":"2023-12-12T14:04:51.575659Z","shell.execute_reply.started":"2023-12-12T14:04:51.447511Z","shell.execute_reply":"2023-12-12T14:04:51.574880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_smiles, temp_smiles, train_labels, temp_labels = train_test_split(smiles_feature[\"SMILES\"].to_numpy(),\n                                                                            norm_label,\n                                                                            test_size=0.3, random_state=RNDST2)\n\nval_smiles, test_smiles, val_labels , test_labels = train_test_split(temp_smiles, temp_labels, test_size=0.6, random_state=RNDST2)\n","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.576791Z","iopub.execute_input":"2023-12-12T14:04:51.577116Z","iopub.status.idle":"2023-12-12T14:04:51.705201Z","shell.execute_reply.started":"2023-12-12T14:04:51.577085Z","shell.execute_reply":"2023-12-12T14:04:51.704409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"full_smiles = smiles_feature[\"SMILES\"].values \nfull_cell_type = one_hot_celltype.toarray()\nfull_labels = norm_label","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.706494Z","iopub.execute_input":"2023-12-12T14:04:51.706844Z","iopub.status.idle":"2023-12-12T14:04:51.712122Z","shell.execute_reply.started":"2023-12-12T14:04:51.706812Z","shell.execute_reply":"2023-12-12T14:04:51.711194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_test_smiles = test_feature_smiles[\"SMILES\"].values\nfinal_cell_type = one_hot_test_cell_type.toarray()","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.719492Z","iopub.execute_input":"2023-12-12T14:04:51.719754Z","iopub.status.idle":"2023-12-12T14:04:51.725695Z","shell.execute_reply.started":"2023-12-12T14:04:51.719731Z","shell.execute_reply":"2023-12-12T14:04:51.724937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"full_smiles[:3]","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.726651Z","iopub.execute_input":"2023-12-12T14:04:51.726921Z","iopub.status.idle":"2023-12-12T14:04:51.738536Z","shell.execute_reply.started":"2023-12-12T14:04:51.726897Z","shell.execute_reply":"2023-12-12T14:04:51.737733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the lengths\nlen(train_smiles), len(train_smiles), len(val_smiles), len(val_smiles), len(test_smiles), len(test_smiles)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.739541Z","iopub.execute_input":"2023-12-12T14:04:51.739799Z","iopub.status.idle":"2023-12-12T14:04:51.751341Z","shell.execute_reply.started":"2023-12-12T14:04:51.739777Z","shell.execute_reply":"2023-12-12T14:04:51.750517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Split chars \nWe will split the `SMILES` into chars and then vectorize.","metadata":{}},{"cell_type":"code","source":"# Function to split sentences into characters\ndef split_chars(text):\n    return \" \".join(list(text))","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.752469Z","iopub.execute_input":"2023-12-12T14:04:51.752761Z","iopub.status.idle":"2023-12-12T14:04:51.761109Z","shell.execute_reply.started":"2023-12-12T14:04:51.752730Z","shell.execute_reply":"2023-12-12T14:04:51.760233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample = \"CC(C)c1cc(C(=O)N2Cc3ccc(CN4CCN(C)CC4)cc3C2)c(O)cc1O\"\nsplit_chars(sample)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.762038Z","iopub.execute_input":"2023-12-12T14:04:51.762291Z","iopub.status.idle":"2023-12-12T14:04:51.773364Z","shell.execute_reply.started":"2023-12-12T14:04:51.762269Z","shell.execute_reply":"2023-12-12T14:04:51.772617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Split the smiles","metadata":{}},{"cell_type":"code","source":"train_char_smiles = [split_chars(feature) for feature in train_smiles]\nval_char_smiles = [split_chars(feature) for feature in val_smiles]\ntest_char_smiles = [split_chars(feature) for feature in test_smiles]","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.774333Z","iopub.execute_input":"2023-12-12T14:04:51.774615Z","iopub.status.idle":"2023-12-12T14:04:51.785700Z","shell.execute_reply.started":"2023-12-12T14:04:51.774593Z","shell.execute_reply":"2023-12-12T14:04:51.784750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For final training\nfull_train_smiles = [split_chars(feature) for feature in full_smiles]","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.786718Z","iopub.execute_input":"2023-12-12T14:04:51.787020Z","iopub.status.idle":"2023-12-12T14:04:51.798764Z","shell.execute_reply.started":"2023-12-12T14:04:51.786998Z","shell.execute_reply":"2023-12-12T14:04:51.797914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For prediction\nfinal_test_smiles = [split_chars(feature) for feature in final_test_smiles]","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.799777Z","iopub.execute_input":"2023-12-12T14:04:51.800007Z","iopub.status.idle":"2023-12-12T14:04:51.809675Z","shell.execute_reply.started":"2023-12-12T14:04:51.799987Z","shell.execute_reply":"2023-12-12T14:04:51.808797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Char vectorizer","metadata":{}},{"cell_type":"code","source":"# What's the average character length?\nchar_len_smiles = [len(feature) for feature in full_smiles]\nmean_char_smiles = np.mean(char_len_smiles)\nmean_char_smiles","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.810879Z","iopub.execute_input":"2023-12-12T14:04:51.811507Z","iopub.status.idle":"2023-12-12T14:04:51.822992Z","shell.execute_reply.started":"2023-12-12T14:04:51.811474Z","shell.execute_reply":"2023-12-12T14:04:51.822179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#  Find what character length covers 95% of sequences\noutput_seq_char_len = int(np.percentile(char_len_smiles, 95))\noutput_seq_char_len","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.823994Z","iopub.execute_input":"2023-12-12T14:04:51.824312Z","iopub.status.idle":"2023-12-12T14:04:51.835288Z","shell.execute_reply.started":"2023-12-12T14:04:51.824281Z","shell.execute_reply":"2023-12-12T14:04:51.834419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the distribution of our sequences at character-level\nplt.hist(char_len_smiles, bins=7);","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:51.836301Z","iopub.execute_input":"2023-12-12T14:04:51.836629Z","iopub.status.idle":"2023-12-12T14:04:52.294278Z","shell.execute_reply.started":"2023-12-12T14:04:51.836597Z","shell.execute_reply":"2023-12-12T14:04:52.293418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Get all the unique characters in smiles","metadata":{"execution":{"iopub.status.busy":"2023-10-10T09:36:47.578256Z","iopub.execute_input":"2023-10-10T09:36:47.578652Z","iopub.status.idle":"2023-10-10T09:36:47.585849Z","shell.execute_reply.started":"2023-10-10T09:36:47.578623Z","shell.execute_reply":"2023-10-10T09:36:47.5846Z"}}},{"cell_type":"code","source":"unique_characters = []\n\n# Iterate over the \"features\" and extract unique characters\nfor feature in full_train_smiles:\n    unique_characters.extend(set(feature))\n\n# Remove duplicates by converting the list to a set and then back to a list\nunique_characters = list(set(unique_characters))\n\n# Print the list of unique characters\nprint(unique_characters)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:52.295393Z","iopub.execute_input":"2023-12-12T14:04:52.295664Z","iopub.status.idle":"2023-12-12T14:04:52.303061Z","shell.execute_reply.started":"2023-12-12T14:04:52.295640Z","shell.execute_reply":"2023-12-12T14:04:52.302124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(unique_characters)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:52.304035Z","iopub.execute_input":"2023-12-12T14:04:52.304280Z","iopub.status.idle":"2023-12-12T14:04:52.314738Z","shell.execute_reply.started":"2023-12-12T14:04:52.304258Z","shell.execute_reply":"2023-12-12T14:04:52.313891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras.layers.experimental.preprocessing import TextVectorization\n\n# Create char-level token vectorizer instance\nNUM_CHAR_TOKENS = len(unique_characters) + 2 \nchar_vectorizer_smiles = TextVectorization(max_tokens=NUM_CHAR_TOKENS,\n                                    output_sequence_length=output_seq_char_len,\n                                    standardize=None,\n                                    split=\"character\",\n                                    name=\"char_vectorizer_SMILES\")\n\n# Adapt character vectorizer to training characters\nchar_vectorizer_smiles.adapt(train_char_smiles)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:52.315663Z","iopub.execute_input":"2023-12-12T14:04:52.315940Z","iopub.status.idle":"2023-12-12T14:04:54.320176Z","shell.execute_reply.started":"2023-12-12T14:04:52.315916Z","shell.execute_reply":"2023-12-12T14:04:54.319403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the config of our char vectorizer\nchar_vectorizer_smiles.get_config()","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:54.321265Z","iopub.execute_input":"2023-12-12T14:04:54.321545Z","iopub.status.idle":"2023-12-12T14:04:54.330193Z","shell.execute_reply.started":"2023-12-12T14:04:54.321521Z","shell.execute_reply":"2023-12-12T14:04:54.329398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import random\n\n# Test out character vectorizer\nrandom_train_feature = random.choice(train_char_smiles)\nprint(f\"Charified text:\\n{random_train_feature}\")\nprint(f\"\\nLength of feature: {len(random_train_feature.split())}\")\nvectorized_feature = char_vectorizer_smiles([random_train_feature])\nprint(f\"\\nVectorized feature:\\n{vectorized_feature}\")\nprint(f\"\\nLength of vectorized feature: {len(vectorized_feature[0])}\")","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:54.331347Z","iopub.execute_input":"2023-12-12T14:04:54.331641Z","iopub.status.idle":"2023-12-12T14:04:55.309534Z","shell.execute_reply.started":"2023-12-12T14:04:54.331618Z","shell.execute_reply":"2023-12-12T14:04:55.308585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You'll notice sequences with a length shorter than 193 (output_seq_char_length) get padded with zeros on the end, this ensures all sequences passed to our model are the same length.","metadata":{}},{"cell_type":"code","source":"# Check character vocabulary characteristics\nchar_vocab = char_vectorizer_smiles.get_vocabulary()\nprint(f\"Number of different characters in character vocab: {len(char_vocab)}\")\nprint(f\"5 most common characters: {char_vocab[:5]}\")\nprint(f\"5 least common characters: {char_vocab[-5:]}\")","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.310757Z","iopub.execute_input":"2023-12-12T14:04:55.311028Z","iopub.status.idle":"2023-12-12T14:04:55.318397Z","shell.execute_reply.started":"2023-12-12T14:04:55.311004Z","shell.execute_reply":"2023-12-12T14:04:55.317544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Char embedding","metadata":{}},{"cell_type":"code","source":"# Create char embedding layer\nchar_embedding = tf.keras.layers.Embedding(input_dim=NUM_CHAR_TOKENS, # number of different characters\n                              output_dim=16, # embedding dimension of each character \n                              name=\"char_embedding_SMILES\")","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.319283Z","iopub.execute_input":"2023-12-12T14:04:55.319586Z","iopub.status.idle":"2023-12-12T14:04:55.331427Z","shell.execute_reply.started":"2023-12-12T14:04:55.319554Z","shell.execute_reply":"2023-12-12T14:04:55.330601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Test out character embedding layer\nprint(f\"Charified text (before vectorization and embedding):\\n{random_train_feature}\\n\")\nchar_embed_example = char_embedding(char_vectorizer_smiles([random_train_feature]))\nprint(f\"Embedded chars (after vectorization and embedding):\\n{char_embed_example}\\n\")\nprint(f\"Character embedding shape: {char_embed_example.shape}\")","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.332551Z","iopub.execute_input":"2023-12-12T14:04:55.333589Z","iopub.status.idle":"2023-12-12T14:04:55.364708Z","shell.execute_reply.started":"2023-12-12T14:04:55.333564Z","shell.execute_reply":"2023-12-12T14:04:55.363815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert list to array\ntrain_chars = np.array(train_char_smiles)\nval_chars = np.array(val_char_smiles)\ntest_chars = np.array(test_char_smiles)\nfull_train_chars = np.array(full_train_smiles)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.365854Z","iopub.execute_input":"2023-12-12T14:04:55.366126Z","iopub.status.idle":"2023-12-12T14:04:55.372573Z","shell.execute_reply.started":"2023-12-12T14:04:55.366102Z","shell.execute_reply":"2023-12-12T14:04:55.371597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Building models","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.utils import plot_model\nfrom tensorflow.keras import Model\nfrom tensorflow.keras.layers import LSTM, Dense, concatenate, Conv1D, GlobalMaxPooling1D, Input, Flatten, GaussianNoise, GlobalAveragePooling1D","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.373762Z","iopub.execute_input":"2023-12-12T14:04:55.374042Z","iopub.status.idle":"2023-12-12T14:04:55.382841Z","shell.execute_reply.started":"2023-12-12T14:04:55.374019Z","shell.execute_reply":"2023-12-12T14:04:55.382112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras.layers import Input, Conv1D, Dense, GlobalMaxPooling1D, concatenate\nfrom tensorflow.keras.models import Model\n\ndef build_smiles_input(input_shape=(1,)):\n    # Define the SMILES input\n    smiles_input = Input(shape=input_shape, dtype=\"string\", name=\"smiles_input\")\n    char_vectors = char_vectorizer_smiles(smiles_input)\n    char_embeddings = char_embedding(char_vectors)\n    conv_layer = Conv1D(64, kernel_size=5, padding=\"same\", activation=\"relu\")(char_embeddings)\n    smiles_output = GlobalMaxPooling1D()(conv_layer)\n    return smiles_input, smiles_output\n\ndef build_cell_type_input(cell_type_shape=(6,)):\n    # Define the cell type input\n    cell_type_input = Input(shape=cell_type_shape, name=\"cell_type_input\")\n    cell_type_output = Dense(32, activation=\"relu\")(cell_type_input)\n    return cell_type_input, cell_type_output\n\ndef build_neural_network(concatenated):\n    # Define the neural network architecture\n    x = Dense(128, activation=\"relu\")(concatenated)\n    output = Dense(18211, activation=\"linear\")(x)\n    return output","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.383855Z","iopub.execute_input":"2023-12-12T14:04:55.384104Z","iopub.status.idle":"2023-12-12T14:04:55.393458Z","shell.execute_reply.started":"2023-12-12T14:04:55.384082Z","shell.execute_reply":"2023-12-12T14:04:55.392748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# The model","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras.layers import Input, Flatten, concatenate, GaussianNoise, Dense, Dropout\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.callbacks import EarlyStopping\nfrom tensorflow.keras.layers import BatchNormalization\n\ndef mean_rowwise_rmse(y_true, y_pred):\n    # Calculate RMSE for each row\n    rowwise_rmse = tf.sqrt(tf.reduce_mean(tf.square(y_true - y_pred), axis=1))\n    \n    # Return the mean of rowwise RMSE\n    return tf.reduce_mean(rowwise_rmse)\n\ndef build_model_3(input_shape=(1,), cell_type_shape=(6,),\n                  embedding_dim=128, dropout_rate=0.9, optimizer=None,\n                  prm_GaussianNoise = 0.09   ):\n    \n    # SMILES input\n    smiles_input = Input(shape=(1,), dtype=\"string\", name=\"smiles_input\")\n    char_vectors = char_vectorizer_smiles(smiles_input)\n    char_embeddings = char_embedding(char_vectors)\n    embed_flatten = Flatten()(char_embeddings)\n\n    # Cell type input\n    cell_type_input = Input(shape=cell_type_shape, dtype=tf.float32, name=\"cell_type_input\")\n\n    # Concatenate SMILES embedding and cell type data\n    concatenated_data = Concatenate(name=\"concatenate\")([embed_flatten, cell_type_input])\n\n    # Apply Gaussian Noise\n    x = GaussianNoise(prm_GaussianNoise)(concatenated_data) # prm_GaussianNoise = 0.09\n\n    # Neural network layers\n    x = Dense(512, activation='elu')(x)\n    x = BatchNormalization()(x)\n#     x = Dropout(dropout_rate)(x)\n\n    x = Dense(256, activation='elu')(x)\n#     x = BatchNormalization()(x)\n#     x = Dropout(dropout_rate)(x)\n\n    x = Dense(128, activation='elu')(x)\n    x = BatchNormalization()(x)\n    x = Dropout(dropout_rate)(x)\n\n    x = Dense(256, activation='elu')(x)\n#     x = BatchNormalization()(x)\n#     x = Dropout(dropout_rate)(x)\n\n    x = Dense(512, activation='elu')(x)\n    x = BatchNormalization()(x)\n    x = Dropout(dropout_rate)(x)\n\n    # Output layer\n    output = Dense(18211, activation='linear')(x)\n\n    # Model creation\n    model = Model(inputs=[smiles_input, cell_type_input], outputs=output, name=\"custom_model\")\n\n    # Compile the model\n    if optimizer:\n        model.compile(loss=mean_rowwise_rmse, optimizer=optimizer, metrics=[\"mae\"])\n    else:\n        model.compile(loss=mean_rowwise_rmse, optimizer=tf.keras.optimizers.Lion(\n    ), metrics=[\"mae\"])\n        \n\n    return model\n","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.394998Z","iopub.execute_input":"2023-12-12T14:04:55.395275Z","iopub.status.idle":"2023-12-12T14:04:55.407836Z","shell.execute_reply.started":"2023-12-12T14:04:55.395253Z","shell.execute_reply":"2023-12-12T14:04:55.406853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model(build_model_3(), show_shapes=True, show_layer_activations=True)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.408848Z","iopub.execute_input":"2023-12-12T14:04:55.409100Z","iopub.status.idle":"2023-12-12T14:04:55.940746Z","shell.execute_reply.started":"2023-12-12T14:04:55.409078Z","shell.execute_reply":"2023-12-12T14:04:55.939811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif 0:\n    model_2 = build_model_3()\n    model_2_history = model_2.fit(\n        x=[train_chars, train_cell_type],\n        y=train_labels,\n        epochs=40,\n        verbose=0,\n        validation_data=([val_chars, val_cell_type], val_labels),\n        callbacks=[create_model_checkpoint(\"model_2\")])","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.941943Z","iopub.execute_input":"2023-12-12T14:04:55.942231Z","iopub.status.idle":"2023-12-12T14:04:55.948541Z","shell.execute_reply.started":"2023-12-12T14:04:55.942200Z","shell.execute_reply":"2023-12-12T14:04:55.947701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif 0:\n    model_2.load_weights(\"model_2\")\n    print(\"Scores on test:\\n\")\n    calculate_mae_and_mrrmse(model=model_2, data=[test_chars, test_cell_type], y_true=test_labels, scaler=scaler)\n    print(\"\\nScores on full data:\\n\")\n    calculate_mae_and_mrrmse(model=model_2, data=[full_train_chars, full_cell_type], y_true=full_labels, scaler=scaler)\n    print(\"\\nPlot training and validation curves:\\n\")\n    plot_training_history(model_2_history, metrics=[\"mae\"])","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.949892Z","iopub.execute_input":"2023-12-12T14:04:55.950258Z","iopub.status.idle":"2023-12-12T14:04:55.963376Z","shell.execute_reply.started":"2023-12-12T14:04:55.950219Z","shell.execute_reply":"2023-12-12T14:04:55.962411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.__version__","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.964407Z","iopub.execute_input":"2023-12-12T14:04:55.964660Z","iopub.status.idle":"2023-12-12T14:04:55.976750Z","shell.execute_reply.started":"2023-12-12T14:04:55.964638Z","shell.execute_reply":"2023-12-12T14:04:55.975720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def build_sub(res, idx=0):\n    df = pd.DataFrame(np.mean(res, axis = 0), columns=sample_columns)\n    df.insert(0, 'id', range(255))\n    df.to_csv(f\"submission_{idx}.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.977678Z","iopub.execute_input":"2023-12-12T14:04:55.977929Z","iopub.status.idle":"2023-12-12T14:04:55.987348Z","shell.execute_reply.started":"2023-12-12T14:04:55.977907Z","shell.execute_reply":"2023-12-12T14:04:55.986523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Core modeling function - k_fold_predict","metadata":{}},{"cell_type":"code","source":"%%time\n\ndef k_fold_predict(create_model_func, features_list=[], full_labels=None,\n                   n_selfblend = 1,n_multiplex_train = 1,  test_data=None, num_folds=7, \n                   random_state=None, scaler=None, optimizer=None, epochs=100, prm_GaussianNoise = 0.09, \n                   restore_best_weights = True,\n                   mode_change_train_set = '',\n                   flag_save_oof = False,\n                   early_stopping=10, dropout=0.7):\n    # Initialize lists to store the predictions\n    all_preds = []\n    all_histories = []\n    val_losses = []\n\n\n\n    Y_pred_oof_blend = np.zeros( (614, 18211) )\n    for i_selfblend in range(n_selfblend ):\n        \n        mrrmses = []\n        r2s = []\n        mae_1_folds = []\n        list_corr_row = []\n        list_corr_col = []\n        \n        \n        # Initialize the KFold object\n    #     kf = StratifiedKFold(n_splits=num_folds, shuffle=True, random_state=random_state)\n        kf = KFold_custom('Tonya')\n        # Loop through the K folds\n    #     for fold, (train_index, val_index) in enumerate(kf.split(full_labels, features_list[1].argmax(axis=1))):\n\n        display(df_stat.head(1))\n        t0_start_folds_loop = time.time()\n        for fold, (train_index, val_index) in enumerate(kf.split()):\n            # Convert indices to integers and split the data\n            train_index = train_index.astype(int)\n            val_index = val_index.astype(int)\n\n            IX_train = train_index\n            if mode_change_train_set == 'Exclude T cells CD8+':            \n                IX_train = np.array( [i for i in IX_train if df_de_train['cell_type'].iat[i] != 'T cells CD8+'  ] )\n            \n            # Multiplex train several times :\n            train_index_basic = list(train_index)\n            train_index = list(train_index)\n            for k in range(1,n_multiplex_train): train_index +=  train_index_basic\n            train_index = np.array(train_index)\n\n\n            X_train_list = [features[train_index] for features in features_list] # list\n            #print('features_list', type(features_list[0]))\n            #print('features_list', (features_list[0].shape)) # shape (614,), (614,6) \n            #print('features_list', (features_list[1].shape)) # \n            #print('features_list', type(features_list), len( features_list  ) )\n            \n            \n            X_val_list = [features[val_index] for features in features_list]\n            y_train = full_labels[train_index]\n            y_val = full_labels[val_index]\n\n            # Create and compile the model\n            model = create_model_func(input_shape=[X_train.shape[1:] for X_train in X_train_list], optimizer=optimizer, dropout_rate=dropout, prm_GaussianNoise = prm_GaussianNoise)\n\n            # Train the model and get predictions on the test set\n            history, preds, preds_single_model_ = train_and_predict(model, X_train_list, y_train, X_val_list, y_val, test_data, scaler, \n                            epochs=epochs, early_stopping=early_stopping, restore_best_weights = restore_best_weights )\n            \n            #print('y_val.shape', y_val.shape,'preds_single_model_.shape', preds_single_model_.shape  )\n            Y_pred_oof_blend[val_index, :] =  ( Y_pred_oof_blend[val_index, :]  * i_selfblend   + preds_single_model_  )    / (i_selfblend + 1)\n            preds_ = Y_pred_oof_blend[val_index, :]  \n            \n            mrrmse = np.sqrt(np.square(y_val - preds_).mean(axis=1)).mean()\n            mrrmses.append(mrrmse)\n            r2 = r2_score( y_val , preds_ )\n            r2s.append(r2)\n            mae_1_fold = mean_absolute_error( y_val , preds_  )\n            mae_1_folds.append(mae_1_fold)\n            list_tmp =[ np.corrcoef( y_val[k,:] , preds_[k,:])[0,1] for k in range(y_val.shape[0]) ]\n            list_corr_row.append( np.mean(list_tmp ))\n            list_tmp_col =[ np.corrcoef( y_val[:,k] , preds_[:,k])[0,1] for k in range(y_val.shape[1]) ]\n            list_corr_col.append( np.mean(list_tmp_col ))\n\n            print(f\"Fold {fold + 1} Mrrmse {mrrmse}\")\n            print(f\"Fold {fold + 1} r2 {r2}\")\n            print(f\"Fold {fold + 1} MAE: {mae_1_fold}\")\n            print(f\"Fold {fold + 1} corrcoef row: {np.mean(list_tmp )}\")\n            print(f\"Fold {fold + 1} corrcoef col: {np.mean(list_tmp_col )}\")\n\n            # Print validation metric for every batch\n            print(f'Fold {fold + 1} - Mean Validation Loss for each batch:')\n            print(np.mean(history.history['val_loss']))\n\n            # Store validation loss for this fold\n            val_losses.append(np.min(history.history['val_loss']))\n\n            # Store predictions for this fold\n            all_preds.append(preds)\n            all_histories.append((history, fold))\n\n        IX = len(df_stat)+1\n        df_stat.loc[IX,'model'] = 'NLPr_SB'+str(i_selfblend+1)+ '_MUL'+str(n_multiplex_train) + '_Epo'+str(epochs) + '_Noi'+str(prm_GaussianNoise) +\\\n            '_RBW' +str(int(restore_best_weights))\n        df_stat.loc[IX,'mrrmse'] = np.mean(mrrmse)\n        df_stat.loc[IX,'mae'] = np.mean(mae_1_folds)\n        df_stat.loc[IX,'corr row'] = np.mean(list_corr_row)\n        df_stat.loc[IX,'corr col'] = np.mean(list_corr_col)\n        df_stat.loc[IX,'corr row med'] = np.median(list_corr_row)\n        df_stat.loc[IX,'corr col med'] = np.median(list_corr_col)\n        df_stat.loc[IX,'r2s'] = np.mean(r2s)\n        df_stat.loc[IX,'time'] = np.round(  time.time()- t0_start_folds_loop  ,1 )\n        for kk in range( len( mrrmses ) ):\n            df_stat.loc[IX,'mrrmse ' + str(kk)] = mrrmses[kk]\n        for kk in range( len( mae_1_folds ) ):\n            df_stat.loc[IX,'mae ' + str(kk)] = mae_1_folds[kk]\n        for kk in range( len( list_corr_row ) ):\n            df_stat.loc[IX,'corr row ' + str(kk)] = list_corr_row[kk]\n        for kk in range( len( list_corr_col ) ):\n            df_stat.loc[IX,'corr col ' + str(kk)] = list_corr_col[kk]\n        \n        display(df_stat)\n        df_stat.round(6).to_csv('df_stat.csv')\n        \n        # Save oof\n        if flag_save_oof:\n            fn4save = 'Y_pred_oof_blend'\n            np.save( fn4save+'.npy', Y_pred_oof_blend )\n            df_oof_for_save =  pd.DataFrame( Y_pred_oof_blend )\n            df_oof_for_save.index = range(1,615)\n            df_oof_for_save.index.name = 'id'\n            df_oof_for_save.to_csv(  fn4save + '.csv')\n            print( df_oof_for_save.shape )\n        \n\n        # Calculate mean and confidence interval for best val_losses\n        mean_val_loss = np.mean(val_losses)\n        confidence_interval = t.interval(0.95, len(val_losses) - 1, loc=mean_val_loss, scale=np.std(val_losses) / np.sqrt(len(val_losses)))\n        print(f\"\\nMean of Best Validation Losses: {mean_val_loss}\")\n        print(f\"Confidence Interval (95%): {confidence_interval}\")\n\n        mean_val_loss = np.mean(mrrmse)\n        confidence_interval = t.interval(0.95, len(mrrmses) - 1, loc=mrrmses, scale=np.std(mrrmses) / np.sqrt(len(mrrmses)))\n        print(f\"\\nMean of Best Validation mrrmses: {mean_val_loss}\")\n        print(f\"Confidence Interval (95%): {confidence_interval}\")\n\n        mean_val_loss = np.mean(r2s)\n        confidence_interval = t.interval(0.95, len(r2s) - 1, loc=r2s, scale=np.std(r2s) / np.sqrt(len(r2s)))\n        print(f\"\\nMean of Best Validation r2: {mean_val_loss}\")\n        print(f\"Confidence Interval (95%): {confidence_interval}\")\n\n        mean_val_loss = np.mean(mae_1_folds)\n        confidence_interval = t.interval(0.95, len(mae_1_folds) - 1, loc=mae_1_folds, scale=np.std(mae_1_folds) / np.sqrt(len(mae_1_folds)))\n        print(f\"\\nMean of Best Validation mae: {mean_val_loss}\")\n        print(f\"Confidence Interval (95%): {confidence_interval}\")\n\n\n    plot_training_history_kfold(all_histories)\n    return all_preds\n\n# Example usage:\n# final_test_data = [np.array(final_test_smiles), final_cell_type]\n# all_preds = k_fold_predict(... ) \n","metadata":{"execution":{"iopub.status.busy":"2023-12-12T14:04:55.988691Z","iopub.execute_input":"2023-12-12T14:04:55.988970Z","iopub.status.idle":"2023-12-12T14:04:56.023464Z","shell.execute_reply.started":"2023-12-12T14:04:55.988947Z","shell.execute_reply":"2023-12-12T14:04:56.022506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Main modeling","metadata":{}},{"cell_type":"code","source":"%%time\npd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None)\ndf_stat = pd.DataFrame()\n\n# Versions of the notebooks:\n\nprint('------------------------')\nprint('Start main modeling part')\nprint()\n\n#dropout 0.7\nfor i_trial in range(1,2):\n    print()\n#     prm_GaussianNoise = 0.09 # 0.12 - 0.03*i_trial\n#     if i_trial == 1: prm_GaussianNoise = 0.01\n#     if i_trial == 2: prm_GaussianNoise = 0.001\n#     if i_trial == 3: prm_GaussianNoise = 0.0001\n        \n    print('Trial', i_trial)# , 'prm_GaussianNoise', prm_GaussianNoise)\n    res = k_fold_predict(build_model_3, \n                         features_list=[full_train_chars, np.array(full_cell_type)], \n                         full_labels=full_labels, \n                         n_selfblend = 1, n_multiplex_train = 25,\n                         num_folds=10, \n                         random_state= i_trial, # RNDST1, \n                         epochs= 300,\n                         prm_GaussianNoise = 0.09, # prm_GaussianNoise, #  1e-3,\n                         restore_best_weights = True, # False, \n                         mode_change_train_set = 'Exclude T cells CD8+', #  '',\n                         flag_save_oof = True,\n                         scaler=scaler,\n                         test_data=[np.array(final_test_smiles), final_cell_type], \n                         early_stopping=200,\n                         dropout=0.425)\n    build_sub(res, idx = i_trial)\n    \n# 125 ExcCD8 25,300 \n\n# 122 - repeat best params 0.574 - 25,300\n    \n# ------------------- restore_best_weights = False, \n\n# 119 long run noise: 0.01, 0.001,0.0001\n\n# 117 long run gpu - n_selfblend = 1, n_multiplex_train = 25,epochs= 500,  prm_GaussianNoise = 0.12 - 0.03*i_trial\n\n# 120 Noi09 sb1m20e5 RBW0 repeat\n# 1\tNLPr_SB1_MUL20_Epo5_Noi0.09_RBW1\t0.607524\t0.436299\t0.315357\t0.491471\t0.314746\t0.507401\t0.144458\t1351.1\n# 113 Noi09 sb1m20e5 RBW0\n# 1\tNLPr_SB1_MUL20_Epo5_Noi0.09_RBW0\t0.604943\t0.44154\t0.315096\t0.481275\t0.321174\t0.480752\t0.134006\t1217.\n\n# Compere 10epochs \n# NLPr_SB1_MUL20_Epo10_Noi0.09_RBW0\t\t1\tС88\t0.572\t0.425\t0.361\t0.529\t0.363\t0.528\t0.152\t172.300\n\n# 112 Noi09 sb1m20e7 RBW0 repeat\n# NLPr_SB1_MUL20_Epo7_Noi0.09_RBW0\t0.588433\t0.427588\t0.334955\t0.530374\t0.334339\t0.509172\t0.165577\t1614.0\n# 111 Noi09 sb1m20e7 RBW0\n# NLPr_SB1_MUL20_Epo7_Noi0.09_RBW0\t0.580138\t0.428156\t0.33704\t0.494878\t0.344791\t0.489412\t0.156175\t2106.4\t0.60156\n    \n# 110 Noi1e-3 sb5m20e7 RBW0 Repeat \n# 109 Noi1e-3 sb5m20e7 RBW0 \n\n# 108 Noi09 sb5m20e7 RBW0 Repeat \n# 107 Noi09 sb5m20e7 RBW0\n\n# 106 Noi09 sb5m20e5 RBW0 Repeat \n# NLPr_SB5_MUL20_Epo5_Noi0.09_RBW0\t0.577995\t0.419039\t0.354628\t0.514035\t0.352540\t0.508756\t0.208101\t1034.\n# 105 Noi09 sb5m20e5 RBW0\n\n# 104 Noi09 sb5m20e10 RBW0  preprare launch\n# NLPr_SB5_MUL20_Epo10_Noi0.09_RBW0\t0.566163\t0.410955\t0.398505\t0.549047\t0.398243\t0.527772\t0.212369\t2023.8\n\n# 100 Noi09 sb1m20e10 RBW0  preprare launch\n\n# 100 Noi09 sb1m20e10 RBW0  preprare launch\n\n# 99 Noi1e-5 sb5m20e40 RBW0 CPU\n# 98 Noi1e-5 sb5m20e30 RBW0 CPU\n\n# 97 Noi1e-5 100e5sb3m RBW0 CPU\n# 96 Noi1e-6 100e5sb3m RBW0 CPU\n\n# 95 Noi0.001 30e5sb20m RBW0 CPU\n# 94 Noi0.001 20e5sb20m RBW0 CPU\n# 93 Noi0.001 10e5sb20m RBW0 CPU\n\n# 92 50e5sb20m RBW0 CPU\n# 91 40e5sb20m RBW0 CPU\n# 90 30e5sb20m RBW0 CPU\n# 89 20e5sb20m RBW0 CPU\n\n# 85 100ep5sb3m0.001 RBW=0 repeat\n# 84 100ep5sb3m0.01 RBW=0 repeat\n# 83 100ep5sb3m0.01 RBW=0\n# 82 100ep5sb3m0.001 RBW=0\n\n# 78 50epo5sb3m009noi RBW=0 repeat\n# 77 50epo5sb3m009noi RBW=0 repeat CPU\n# 76 100epo5sb3m009noi RBW=0 repaeat\n# 75 50epo5sb3m009noi RBW=0\n\n# 72 100e5sb3m  \n\n# Exception:  RBW=1 # 71 500epo1sb3m009noi restore_best_weights = True\n\n\n# 70 500epo5sb3m009noi\n# 69 500epo1sb3m009noi\n\n# 68 CPU 400epo5sb3m00001noi \n# 67 CPU 300epo5sb3m00001noi \n# 66 CPU 300epo5sb3m001noi \n\n# 65 CPU 500epo5sb3m009noi \n# 64 CPU 400epo5sb3m009noi \n# 63 CPU 400epo5sb3m009noi \n# 62 CPU 300epo5sb3m009noi \n# 61 CPU 200epo5sb3m009noi \n# 60 CPU 100epo5sb3m009noi \n# 59 CPU 50epo5sb3m009noi \n# ---------------- restore_best_weights = False, \n\n# 51 CPU noi0.09, 5sb,2m,\n# 50 CPU noi0.09, 5sb,1m, (Basic version - repeat)\n# 49 CPU noi0.07, 5sb,1m,\n# 48 CPU noi0.05, 5sb,1m,\n# V47 CPU noi0.03, 5sb,1m,\n    ","metadata":{"execution":{"iopub.status.busy":"2023-11-22T20:43:08.524859Z","iopub.execute_input":"2023-11-22T20:43:08.525076Z","iopub.status.idle":"2023-11-22T20:43:48.712747Z","shell.execute_reply.started":"2023-11-22T20:43:08.525057Z","shell.execute_reply":"2023-11-22T20:43:48.711963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"prediction_variance = np.var(res, axis=0)\n\n# Calculate a metric representing the overall difficulty of prediction\ndifficulty_metric = np.mean(prediction_variance, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-11-22T20:43:48.713774Z","iopub.execute_input":"2023-11-22T20:43:48.714210Z","iopub.status.idle":"2023-11-22T20:43:48.766353Z","shell.execute_reply.started":"2023-11-22T20:43:48.714181Z","shell.execute_reply":"2023-11-22T20:43:48.765668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"difficulty_metric.shape","metadata":{"execution":{"iopub.status.busy":"2023-11-22T20:43:48.767275Z","iopub.execute_input":"2023-11-22T20:43:48.767644Z","iopub.status.idle":"2023-11-22T20:43:48.790241Z","shell.execute_reply.started":"2023-11-22T20:43:48.767620Z","shell.execute_reply":"2023-11-22T20:43:48.789396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_ = plt.hist(difficulty_metric, bins=100)","metadata":{"execution":{"iopub.status.busy":"2023-11-22T20:43:48.791495Z","iopub.execute_input":"2023-11-22T20:43:48.791795Z","iopub.status.idle":"2023-11-22T20:43:49.072602Z","shell.execute_reply.started":"2023-11-22T20:43:48.791766Z","shell.execute_reply":"2023-11-22T20:43:49.071731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Show scores, stat, etc","metadata":{}},{"cell_type":"code","source":"pd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None)\ndisplay(df_stat.sort_values(df_stat.columns[0]  ) )","metadata":{"execution":{"iopub.status.busy":"2023-11-22T20:43:49.073685Z","iopub.execute_input":"2023-11-22T20:43:49.073997Z","iopub.status.idle":"2023-11-22T20:43:49.095476Z","shell.execute_reply.started":"2023-11-22T20:43:49.073960Z","shell.execute_reply":"2023-11-22T20:43:49.094669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Final timing","metadata":{}},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )\nprint('%.1f minutes passed total '%( (time.time()-t0start)/60)  )\nprint('%.2f hours passed total '%( (time.time()-t0start)/3600)  )","metadata":{"execution":{"iopub.status.busy":"2023-11-22T20:43:49.096460Z","iopub.execute_input":"2023-11-22T20:43:49.097240Z","iopub.status.idle":"2023-11-22T20:43:49.102320Z","shell.execute_reply.started":"2023-11-22T20:43:49.097212Z","shell.execute_reply":"2023-11-22T20:43:49.101535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}