{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Resources\n```\nhttps://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\n\nhttps://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109850852&cellId=28 \n\n\nhttps://github.com/salvaba94/G2Net\nThe model implemented for the competition (see the image below) has been created following an end-to-end philosophy, meaning that even the time-series pre-processing logic is included as part of the model and might be made trainable. To know more details about the building blocks of the model, refer to any of the Colab Guides provided by the project.\nPreprocess module:\nAugmentation.py: Implements several augmentations in the form of Keras layers, including Gaussian noise, spectral masking (TPU-compatible and TPU-incompatible versions) and channel permutation.\nPreprocessing.py: Implements several preprocessing layers in the form of trainable Keras layers, including time windows (TPU-incompatible Tukey window and generic TPU-compatible window), bandpass filtering and spectral whitening.\nSpectrogram.py: Includes a TensorFlow version of CQT1992v2 implemented in nnAudio with PyTorch. Being in the form of a Keras layer, it also adds functionality to adapt the output range to that recommended as per stability by 2D convolutional models.\nTrain module:\nAcceleration.py: Includes the logic to automatically configure the TPU if any.\nLosses.py: Implements a differentiable loss whose minimisation directly maximises the AUC score.\nSchedulers.py: Implements a wrapper to make CosineDecayRestarts learning rate scheduler compatible with ReduceLROnPlateau.\nUtilities module:\nGeneralUtilities.py: General utilities used all along the project mainly to perform automatic Tensor broadcast and determine mean and standard deviation from a dataset with multiprocessing capabilities.\nPlottingUtilities.py: Includes all the logic behind the plots.\n\n```","metadata":{"_uuid":"e85ad2dd-4548-45c2-bd22-91d2a8e4c38a","_cell_guid":"6f7edf4e-a866-408c-9f90-083b4e2ba675","trusted":true}},{"cell_type":"markdown","source":"","metadata":{"_uuid":"51600441-6c40-4f20-a165-f4380cbb5626","_cell_guid":"cc51e09c-c472-4c65-b21d-230e27930807","trusted":true}},{"cell_type":"code","source":"## Resources\n\n\n\n# https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\n# https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\n\n# https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109850852&cellId=28 \n\n\n# https://github.com/salvaba94/G2Net\n# The model implemented for the competition (see the image below) has been created following an end-to-end philosophy, meaning that even the time-series pre-processing logic is included as part of the model and might be made trainable. To know more details about the building blocks of the model, refer to any of the Colab Guides provided by the project.\n# Preprocess module:\n# Augmentation.py: Implements several augmentations in the form of Keras layers, including Gaussian noise, spectral masking (TPU-compatible and TPU-incompatible versions) and channel permutation.\n# Preprocessing.py: Implements several preprocessing layers in the form of trainable Keras layers, including time windows (TPU-incompatible Tukey window and generic TPU-compatible window), bandpass filtering and spectral whitening.\n# Spectrogram.py: Includes a TensorFlow version of CQT1992v2 implemented in nnAudio with PyTorch. Being in the form of a Keras layer, it also adds functionality to adapt the output range to that recommended as per stability by 2D convolutional models.\n# Train module:\n# Acceleration.py: Includes the logic to automatically configure the TPU if any.\n# Losses.py: Implements a differentiable loss whose minimisation directly maximises the AUC score.\n# Schedulers.py: Implements a wrapper to make CosineDecayRestarts learning rate scheduler compatible with ReduceLROnPlateau.\n# Utilities module:\n# GeneralUtilities.py: General utilities used all along the project mainly to perform automatic Tensor broadcast and determine mean and standard deviation from a dataset with multiprocessing capabilities.\n# PlottingUtilities.py: Includes all the logic behind the plots.\n\n","metadata":{"execution":{"iopub.status.busy":"2022-11-08T07:08:57.935806Z","iopub.execute_input":"2022-11-08T07:08:57.936221Z","iopub.status.idle":"2022-11-08T07:08:57.942783Z","shell.execute_reply.started":"2022-11-08T07:08:57.936186Z","shell.execute_reply":"2022-11-08T07:08:57.941675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"playground_size = 0.1# from 0 to 1 \nholdout_size = 0.1# from 0 to 1 \nrandom_seed4folds = 42 \nn_features = 100 # use first N features - to make quick experiments\n\nrescale_Y_to_mean0_std1 = True\n\nflag_prepare_submission = True # \n\n#selected_targets = 'ALL'# 'CD31'","metadata":{"_uuid":"6b246588-79af-4af7-aeaa-28c40ef039e5","_cell_guid":"90284018-0e3f-4e49-941a-d77ce2d84e9d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:35:54.574467Z","iopub.execute_input":"2022-11-07T19:35:54.575077Z","iopub.status.idle":"2022-11-07T19:35:54.590355Z","shell.execute_reply.started":"2022-11-07T19:35:54.574939Z","shell.execute_reply":"2022-11-07T19:35:54.588911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{"_uuid":"a42ce9d6-ee5d-4ec0-9c2a-31710ef3b8b5","_cell_guid":"f6b5ae5b-a81f-47da-8027-c514cc207aac","trusted":true}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"104f0242-958b-4ef1-af40-8bec2dc9bea0","_cell_guid":"98e6bc60-36d3-4d45-855a-11748f831a0d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:35:54.592596Z","iopub.execute_input":"2022-11-07T19:35:54.593113Z","iopub.status.idle":"2022-11-07T19:35:54.62244Z","shell.execute_reply.started":"2022-11-07T19:35:54.593063Z","shell.execute_reply":"2022-11-07T19:35:54.621098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndisplay(df_cell)","metadata":{"_uuid":"01798d74-2f72-4000-b474-cdb4d74e5f3e","_cell_guid":"d3410584-a62c-4543-8fac-a4fe8fe946f6","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:35:54.624379Z","iopub.execute_input":"2022-11-07T19:35:54.624755Z","iopub.status.idle":"2022-11-07T19:36:19.957418Z","shell.execute_reply.started":"2022-11-07T19:35:54.624724Z","shell.execute_reply":"2022-11-07T19:36:19.955748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"_uuid":"0bdd40f9-eb9c-4eaa-9e08-5e91c0ed728b","_cell_guid":"bfdcb1cc-f7ee-4688-a601-d8661a95ec53","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:36:19.96272Z","iopub.execute_input":"2022-11-07T19:36:19.963346Z","iopub.status.idle":"2022-11-07T19:36:44.343023Z","shell.execute_reply.started":"2022-11-07T19:36:19.96329Z","shell.execute_reply":"2022-11-07T19:36:44.341286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load prepared Features for CITE-seq part of task","metadata":{"_uuid":"32413e3a-7b7f-4a2e-b028-4410b034ea14","_cell_guid":"d47e3bef-4916-43c5-921c-512ca1df8b9d","trusted":true}},{"cell_type":"code","source":"%%time\n\nprint('Load prepared features for CITE-seq')\n# These files contain both train and test parts .\n# For CITEseq part - first 70988 elements - train, and later 48663 - test. Overall 119651 samples.\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_TruncatedSVD200_niter7_rs42.csv'\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA500.csv'\ndf_cite = pd.read_csv(fn,index_col = 0)\ndisplay(df_cite)","metadata":{"_uuid":"14d13305-d62b-4822-b783-09f607eb2cb6","_cell_guid":"0495e7d2-ea36-469e-819a-26ab27a41c58","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:36:44.345866Z","iopub.execute_input":"2022-11-07T19:36:44.346932Z","iopub.status.idle":"2022-11-07T19:37:00.438419Z","shell.execute_reply.started":"2022-11-07T19:36:44.346863Z","shell.execute_reply":"2022-11-07T19:37:00.437053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_cite.mean(axis = 0).head(3))\nprint('Pay attention - PCA features are ordered by magnitude:')\nprint(df_cite.std(axis=0).head(5))","metadata":{"_uuid":"afe9baa3-48d2-4d56-b027-159880e77522","_cell_guid":"bc595bf0-e292-4bfa-bd19-821051bf8579","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:00.440605Z","iopub.execute_input":"2022-11-07T19:37:00.441596Z","iopub.status.idle":"2022-11-07T19:37:01.67191Z","shell.execute_reply.started":"2022-11-07T19:37:00.441534Z","shell.execute_reply":"2022-11-07T19:37:01.670383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Targets for CITE-seq","metadata":{"_uuid":"11f0481b-2dbc-4835-b7c2-98755ca1f63b","_cell_guid":"2e5213c8-69c3-4832-a5ce-f2471ccabb52","trusted":true}},{"cell_type":"code","source":"%%time\n#if 1:\nprint('Load CITE-seq targets and ')\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndisplay(df_cite_train_y)","metadata":{"_uuid":"93f0dd7a-8a6a-4224-bc07-ff982ab834f7","_cell_guid":"915efa3d-6965-4b7f-a759-d1e3848d165d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:01.673602Z","iopub.execute_input":"2022-11-07T19:37:01.674019Z","iopub.status.idle":"2022-11-07T19:37:02.761919Z","shell.execute_reply.started":"2022-11-07T19:37:01.673983Z","shell.execute_reply":"2022-11-07T19:37:02.760313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load and start prepare some metadata (cut only CITE-seq train part)","metadata":{"_uuid":"60de91b6-5173-4461-9dd6-7094e3173f1a","_cell_guid":"3e9c27ac-e90e-44a4-ab94-dda2b95239f9","trusted":true}},{"cell_type":"code","source":"%%time\nfn2 = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/_citeseq_meta_all_text_also.csv'\ndf_meta_full = pd.read_csv(fn2,index_col = 0)\ndisplay(df_meta_full)\n#if 1:\ndf_meta = pd.DataFrame(index = df_cite_train_y.index) \ndf_meta = df_meta.join(df_cell.set_index('cell_id') )\ndisplay(df_meta)","metadata":{"_uuid":"a3de63b6-a27d-4924-9764-6ef3fcf03c1d","_cell_guid":"f5fa99ae-4511-45ec-907c-b872e8744acf","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:02.763921Z","iopub.execute_input":"2022-11-07T19:37:02.764801Z","iopub.status.idle":"2022-11-07T19:37:03.103852Z","shell.execute_reply.started":"2022-11-07T19:37:02.764758Z","shell.execute_reply":"2022-11-07T19:37:03.102535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{"_uuid":"7a5aa539-257d-4e8b-9bcd-61843a2ebc5a","_cell_guid":"499ad22c-c00c-441c-94a7-f069d2dcecb3","trusted":true}},{"cell_type":"code","source":"%%time\n# Create columns in df_meta which label the \"Playground\" and \"Holdout\"\nfrom sklearn.model_selection import train_test_split\n\ndf_meta['Index'] = range(len(df_meta))\nscol = 'donor&day&CT'\ndf_meta[scol] =df_meta['donor'].apply(lambda x:str(x)+'_') + df_meta['day'].apply(lambda x:str(x)+'_') + df_meta['cell_type']\n\n\nflagged_column_name = 'Playground'\nif playground_size > 0 :\n    d_tmp_train, d_tmp_test = train_test_split(df_meta, test_size = playground_size, random_state= random_seed4folds, \n                                               stratify = df_meta[scol] ,  shuffle=True )\n    df_meta[flagged_column_name] = df_meta['Index'].isin(d_tmp_test['Index'].values)\nelse:\n    d_tmp_train = df_meta\n    df_meta[flagged_column_name] = False\nprint('flagged_column_name:',df_meta[flagged_column_name].sum(), df_meta[flagged_column_name].sum()/len(df_meta) )\n\nflagged_column_name = 'Holdout'\nif holdout_size > 0 :\n    d_tmp_train, d_tmp_test = train_test_split(d_tmp_train, test_size = holdout_size, random_state= random_seed4folds, \n                                               stratify = d_tmp_train[scol] ,  shuffle=True )\n    df_meta[flagged_column_name] = df_meta['Index'].isin(d_tmp_test['Index'].values)\nelse:\n    df_meta[flagged_column_name] = False\nprint('flagged_column_name:', df_meta[flagged_column_name].sum(), df_meta[flagged_column_name].sum()/len(df_meta) )\n\ndisplay(df_meta)","metadata":{"_uuid":"9b4dc5b2-2978-466d-b4c4-1ac1e73baa13","_cell_guid":"d2b7ecb9-51df-4676-9b96-20f8dac5702b","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:03.105409Z","iopub.execute_input":"2022-11-07T19:37:03.105779Z","iopub.status.idle":"2022-11-07T19:37:03.488236Z","shell.execute_reply.started":"2022-11-07T19:37:03.105745Z","shell.execute_reply":"2022-11-07T19:37:03.486861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create validation/folds schemes data","metadata":{"_uuid":"5109c030-65a8-4e67-9cae-0aa01392cec1","_cell_guid":"c3d12c63-4e9a-45dd-9c6b-ea503d434ddf","trusted":true}},{"cell_type":"markdown","source":"## Create Playground folds Structure","metadata":{"_uuid":"eb88a2cb-94a1-4c36-96db-091aa57696ba","_cell_guid":"cda7479e-f9e6-4543-beab-498b95470527","execution":{"iopub.status.busy":"2022-11-05T08:03:06.263638Z","iopub.execute_input":"2022-11-05T08:03:06.264028Z","iopub.status.idle":"2022-11-05T08:03:06.273381Z","shell.execute_reply.started":"2022-11-05T08:03:06.263996Z","shell.execute_reply":"2022-11-05T08:03:06.272164Z"},"trusted":true}},{"cell_type":"code","source":"%%time\n# Validation scheme for playground\n\ndict_playground_folds_info = {}\ndict_playground_folds_info['name'] = 'Playground Folds'\ndict_playground_folds_info['Info'] = ['Train','Test1 Priv Like','Test2 Public Like', 'Not Playground' ]\ndict_playground_folds_info['OOF Multiplicity'] = [0,1/2,0,1/2 ]\ndict_playground_folds_info['OOF Indices Main'] = df_meta['Index'][df_meta['Playground'] == 1].values\ndict_playground_folds_info['OOF Indices Additional'] = df_meta['Index'][df_meta['Playground'] == 0].values\n\n#dict_playground_folds_info['List Indices'] = list_folds_main    \n\n\nlist_folds_playground = []\nif playground_size > 0: \n    mask_playground = (df_meta['Playground'] == 1)\n    c = 0\n    for day2exclude in [2,3,4]:\n        for donor2exclude in [32606,  31800]: # We will need to predict always MALE (not female) - like on LB. (# donor 13176 - female)\n            train_index =\\\n                np.where(mask_playground &  (df_meta['day']  != day2exclude) & ( df_meta['donor']  != donor2exclude  ) )  [0]\n            test_index1_like_private_lb =\\\n                np.where( mask_playground &  (df_meta['day']  == day2exclude)  ) [0]\n            test_index2_like_public_lb =\\\n                np.where(mask_playground &  (df_meta['day']  != day2exclude) &  (df_meta['donor']  == donor2exclude ) ) [0]\n\n            test_outside_playground = np.where( (mask_playground == 0) &  (df_meta['day']  == day2exclude)   )[0]\n\n            list_folds_playground.append( (train_index,  test_index1_like_private_lb , test_index2_like_public_lb,\n                                          test_outside_playground ) )\n\n            str_fold_inf = 'Fold ' +str(c) + ': Train: excludes Day '+str(day2exclude) + ' and Donor ' + str( donor2exclude )\n            print(str_fold_inf, 'Sizes: train:',len(train_index), 'Test Like Priv'  ,len(test_index1_like_private_lb), \n                  'Test Like Publ',   len(test_index2_like_public_lb),  \n                  'test_outside_playground',   len(test_outside_playground),  \n                 ); c+=1\n\ndict_playground_folds_info['List Indices'] = list_folds_playground                \nprint(len(list_folds_playground))","metadata":{"_uuid":"a0109ea6-317a-4d84-a02b-75ce1b107e14","_cell_guid":"61867b0c-cbd7-4a90-b9bc-481e793a0fbe","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:03.490402Z","iopub.execute_input":"2022-11-07T19:37:03.491049Z","iopub.status.idle":"2022-11-07T19:37:03.528309Z","shell.execute_reply.started":"2022-11-07T19:37:03.490973Z","shell.execute_reply":"2022-11-07T19:37:03.526973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create Main Folds Structure","metadata":{"_uuid":"8c6c6d8f-3c62-4f5e-b93f-3e219ce6ea03","_cell_guid":"2276ca0f-7b17-4f18-9a22-b5e793542771","execution":{"iopub.status.busy":"2022-11-05T08:32:00.891523Z","iopub.execute_input":"2022-11-05T08:32:00.892175Z","iopub.status.idle":"2022-11-05T08:32:00.897837Z","shell.execute_reply.started":"2022-11-05T08:32:00.892142Z","shell.execute_reply":"2022-11-05T08:32:00.896742Z"},"trusted":true}},{"cell_type":"code","source":"%%time\ndict_main_folds_info = {}\ndict_main_folds_info['name'] = 'Main Folds'\ndict_main_folds_info['Info'] = ['Train','Test1 Priv Like','Test2 Public Like', 'Test3 Priv Like HO', 'Test4 Publ Like HO', 'Playground' ]\ndict_main_folds_info['OOF Multiplicity'] = [0,1/2,0,1/2,0,1/2 ]\nm = (df_meta['Holdout'] == 0) & ( df_meta['Playground'] == 0 )\ndict_main_folds_info['OOF Indices Main'] = df_meta['Index'][m].values\ndict_main_folds_info['OOF Indices Additional'] = df_meta['Index'][df_meta['Holdout'] == 1].values\n\n#dict_main_folds_info['List Indices'] = list_folds_main    \n\nlist_folds_main = []\nmask_not_playground = (df_meta['Playground'] == 0) \nmask_holdout = (df_meta['Holdout']==1)\nc = 0\nfor day2exclude in [2,3,4]: # The day excluded from the TRAIN \n    for donor2exclude in [32606,  31800]: # We will need to predict always MALE (not female) - like on LB. (# donor 13176 - female)\n        # donor2exclude - the donor excluded from the train \n        \n        train_index = np.where(\\\n           mask_not_playground & (df_meta['day']  != day2exclude) & ( df_meta['donor']  != donor2exclude  ) )  [0]\n        test_index1_like_private_lb = np.where(\\\n           mask_not_playground & (df_meta['day']  == day2exclude) & (~mask_holdout) ) [0]\n        test_index1_like_private_lb_holdout = np.where(\\\n           mask_not_playground & (df_meta['day']  == day2exclude) & (mask_holdout) ) [0]\n        test_index2_like_public_lb = np.where(\\\n           mask_not_playground & (df_meta['day']  != day2exclude) &  (df_meta['donor']  == donor2exclude ) & (~mask_holdout) ) [0]\n        test_index2_like_public_lb_holdout = np.where(\\\n           mask_not_playground & (df_meta['day']  != day2exclude) &  (df_meta['donor']  == donor2exclude ) & mask_holdout ) [0]\n        \n        test_playground = np.where( (mask_not_playground == 0) &  (df_meta['day']  == day2exclude) )[0]\n        \n        list_folds_main.append( (train_index,  test_index1_like_private_lb , test_index2_like_public_lb, \n                              test_index1_like_private_lb_holdout, test_index2_like_public_lb_holdout, test_playground ) )\n    \n        str_fold_inf = 'Fold ' +str(c) + ': Train: excludes Day '+str(day2exclude) + ' and Donor ' + str( donor2exclude )\n        print(str_fold_inf, 'Sizes: train:',len(train_index), 'Test Like Priv'  ,len(test_index1_like_private_lb), \n              'Test Like Publ',   len(test_index2_like_public_lb),  \n              'Test Like Priv HoldOut',   len(test_index1_like_private_lb_holdout),  \n              'Test Like Publ HoldOut',   len(test_index2_like_public_lb_holdout),  \n              'Test Like Publ HoldOut',   len(test_index2_like_public_lb_holdout),  \n             ); c+=1\n    \ndict_main_folds_info['List Indices'] = list_folds_main    \nprint(len(list_folds_main))","metadata":{"_uuid":"3589e878-57ab-4b98-aa39-51424a654bb3","_cell_guid":"bddaea9c-6375-482e-a4d7-978a6ae1faf9","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:03.530167Z","iopub.execute_input":"2022-11-07T19:37:03.530641Z","iopub.status.idle":"2022-11-07T19:37:03.59099Z","shell.execute_reply.started":"2022-11-07T19:37:03.530582Z","shell.execute_reply":"2022-11-07T19:37:03.589573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Optional Rescaling of Targets","metadata":{"_uuid":"9b83bdd2-29d9-4358-b1bb-982a45d07779","_cell_guid":"7f7f64db-34dc-4f60-83dc-7e82329481c2","trusted":true}},{"cell_type":"code","source":"# Since metric is - correlation coefficient - any rescaling aY+b will not be change it , so one can do like that:\n# it is not clear is it optimal or not (see https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360253 )\n\nY = df_cite_train_y.values\n\nif rescale_Y_to_mean0_std1:\n    Y -= Y.mean(axis=1).reshape(-1, 1)\n    Y /= Y.std(axis=1).reshape(-1, 1)\n    print('Rescaling to mean 0 and std 1 has been done')","metadata":{"_uuid":"97b44d5e-03a8-4b2e-9c2c-92223b6318bd","_cell_guid":"117ab723-71af-45cb-affb-4eb0c8ce487b","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:03.592631Z","iopub.execute_input":"2022-11-07T19:37:03.593094Z","iopub.status.idle":"2022-11-07T19:37:03.637612Z","shell.execute_reply.started":"2022-11-07T19:37:03.593053Z","shell.execute_reply":"2022-11-07T19:37:03.636121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling Preparations","metadata":{"_uuid":"36ecbbd7-f9a3-4b97-843d-1b787d6bd3c9","_cell_guid":"a5757f02-9083-4caf-a0b8-fff7d956fb21","trusted":true}},{"cell_type":"code","source":"import gc\ndef correlation_score(y_true, y_pred):\n    \"\"\"Scores the predictions according to the competition rules. \n    \n    It is assumed that the predictions are not constant.\n    \n    Returns the average of each sample's Pearson correlation coefficient\"\"\"\n\n    y2 = y_pred.copy()\n    y2 -= y2.mean(axis=1).reshape(-1, 1);    y2 /= y2.std(axis=1).reshape(-1, 1)    \n    if rescale_Y_to_mean0_std1:\n        y1 = y_true # Already rescaled \n    else:\n        y1 = y_true.copy(); \n        y1 -= y1.mean(axis=1).reshape(-1, 1);    y1 /= y1.std(axis=1).reshape(-1, 1) \n        \n    c = (y1*y2).mean().mean()# Correlation for rescaled matrices is just matrix product and average \n    \n    # Memory control:\n    if not rescale_Y_to_mean0_std1:\n        del y1\n    del y2\n    gc.collect()\n    \n    return c\n\n    # Slow way:\n    \n    if type(y_true) == pd.DataFrame: y_true = y_true.values\n    if type(y_pred) == pd.DataFrame: y_pred = y_pred.values\n    corrsum = 0\n    for i in range(len(y_true)):\n        corrsum += np.corrcoef(y_true[i], y_pred[i])[1, 0]\n    return corrsum / len(y_true)","metadata":{"_uuid":"0e59893b-3efb-4c89-bb2a-5dee2cbb4239","_cell_guid":"eda4e4b5-494e-4f86-aac2-ef97b9884596","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:03.642598Z","iopub.execute_input":"2022-11-07T19:37:03.642996Z","iopub.status.idle":"2022-11-07T19:37:03.654181Z","shell.execute_reply.started":"2022-11-07T19:37:03.642964Z","shell.execute_reply":"2022-11-07T19:37:03.652862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_error\nfrom sklearn.metrics import r2_score","metadata":{"_uuid":"f89b4512-7d1c-406d-8f37-67c07d1bedfc","_cell_guid":"c19ef489-3fd3-4e6e-bd2a-e642789651de","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:03.655779Z","iopub.execute_input":"2022-11-07T19:37:03.656126Z","iopub.status.idle":"2022-11-07T19:37:03.67356Z","shell.execute_reply.started":"2022-11-07T19:37:03.656097Z","shell.execute_reply":"2022-11-07T19:37:03.671743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Define X \n","metadata":{"_uuid":"cf22a937-d737-44fe-af69-7f294b0875d4","_cell_guid":"7693d572-9fc4-47fe-9a02-aad5a1ceaefc","trusted":true}},{"cell_type":"code","source":"n_features_loc = n_features\n\nX = df_cite.iloc[:,:n_features_loc].values\nprint('X.shape:', X.shape,'Y.shape:', Y.shape)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T19:37:03.675267Z","iopub.execute_input":"2022-11-07T19:37:03.675657Z","iopub.status.idle":"2022-11-07T19:37:03.684547Z","shell.execute_reply.started":"2022-11-07T19:37:03.675625Z","shell.execute_reply":"2022-11-07T19:37:03.683705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Setting the model used below","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import Ridge\nalpha4Ridge  = 1e4\nmodel = Ridge(alpha=alpha4Ridge)\n\n\nimport lightgbm as lgbm\nfrom sklearn.multioutput import MultiOutputClassifier, MultiOutputRegressor\n\nif 0:\n    # https://www.kaggle.com/code/yaroshenko/mmscel-crossvalidation-schemes?scriptVersionId=109641833&cellId=50\n\n    model = MultiOutputRegressor(estimator=lgbm.LGBMRegressor(colsample_bytree=1,\n                                                 learning_rate=0.09, max_depth=15,\n                                                 min_child_samples=50,\n                                                 min_child_weight=0.1,\n                                                 min_split_gain=3.119890406724053,\n                                                 n_estimators=150, num_leaves=12,\n                                                 other_rate=1, random_state=0,\n                                                 reg_lambda=0.8575143843171859,\n                                                 subsample=1,\n                                                 subsample_for_bin=10000,\n                                                 subsample_freq=10))\n    \n    # From : https://www.kaggle.com/code/vuonglam/tune-lgbm-only-final-cite-task?scriptVersionId=104599626&cellId=13\n    params = {\n         'learning_rate': 0.1, \n         'metric': 'mae', \n         \"seed\": 42,\n        'reg_alpha': 0.0014, \n        'reg_lambda': 0.2, \n        'colsample_bytree': 0.8, \n        'subsample': 0.5, \n        'max_depth': 10, \n        'num_leaves': 722, \n        'min_child_samples': 83, \n        }\n\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params, n_estimators=400))   \n\n\n    #  https://www.kaggle.com/code/swimmy/lgbm-baseline-msci-citeseq\n    params = {\n         'learning_rate': 0.1, \n         'metric': 'mae', \n         \"seed\": 42,\n        'reg_alpha': 0.0014, \n        'reg_lambda': 0.2, \n        'colsample_bytree': 0.8, \n        'subsample': 0.5, \n        'max_depth': 10, \n        'num_leaves': 722, \n        'min_child_samples': 83, \n        }\n\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params, n_estimators=1000))\n\n\n    # https://www.kaggle.com/code/xiafire/lb0-830-lgbm-optuna-msci-citeseq \n    # [LightGBM] [Warning] Unknown parameter: min_data_per_groups\n    params = {'metric': 'mae', 'random_state': 42, 'n_estimators': 2000, 'reg_alpha': 0.03645857751758206, \n              'reg_lambda': 0.0025972855120393492, 'colsample_bytree': 1.0, 'subsample': 0.6, \n              'learning_rate': 0.013262872399411381, 'max_depth': 10, 'num_leaves': 186, \n              'min_child_samples': 263, 'min_data_per_groups': 46}\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n\n    # https://www.kaggle.com/code/user327934/mmscel-cv-modeling-with-playground-lgbm-opt?scriptVersionId=109668468&cellId=26\n    params =  {'n_estimators': 950, 'learning_rate': 0.02861279493150065, 'max_depth': 16, 'num_leaves': 5, \n               'min_child_samples': 17, 'min_child_weight': 0.1340996547520917, 'colsample_bytree': 0.4, \n               'subsample': 0.8, 'subsample_freq': 47, 'cat_smooth': 27}\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n\n    # Same from above - for 'CD31\n    params =  {'random_state': 42,\n      'min_split_gain': 12.2,\n      'subsample_for_bin': 10000,\n      'other_rate': 1,\n      'reg_alpha': 0.0,\n      'reg_lambda': 0.0,\n      'n_estimators': 585,\n      'learning_rate': 0.015583982144506184,\n      'max_depth': 18,\n      'num_leaves': 14,\n      'min_child_samples': 48,\n      'min_child_weight': 0.16452279543326437,\n      'colsample_bytree': 1.0,\n      'subsample': 0.6,\n      'subsample_freq': 5,\n      'cat_smooth': 20}\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n\n\n    # Same from above - for 'CD44\n    {'random_state': 42,\n      'min_split_gain': 12.2,\n      'subsample_for_bin': 10000,\n      'other_rate': 1,\n      'reg_alpha': 0.0,\n      'reg_lambda': 0.0,\n      'n_estimators': 950,\n      'learning_rate': 0.02861279493150065,\n      'max_depth': 16,\n      'num_leaves': 5,\n      'min_child_samples': 17,\n      'min_child_weight': 0.1340996547520917,\n      'colsample_bytree': 0.4,\n      'subsample': 0.8,\n      'subsample_freq': 47,\n      'cat_smooth': 27}\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n\n    # Optuna 100 for CD31 (CD31: Best r2: 0.491408 Mse: 20.766) https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109847986&cellId=28\n    params = {'n_estimators': 500, 'reg_alpha': 4.319240701302627, 'reg_lambda': 1.9095488166970294, 'colsample_bytree': 1.0, 'subsample': 0.7, 'max_depth': 10, 'learning_rate': 0.02890880408831122, 'num_leaves': 54, 'min_child_samples': 22}\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n\n    # Same as above, but results from more trials: https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109850852&cellId=28\n    #Optimization starts.  n_trials =  1600 # Best r2: 0.487042 Mse: 20.944 # Wall time: 3h 40min 52s\n    #Best params:\n    params = {'n_estimators': 500, 'reg_alpha': 0.12904988002901469, 'reg_lambda': 2.83434368782451, 'colsample_bytree': 0.5, 'subsample': 0.6, 'max_depth': 6, 'learning_rate': 0.036631995226817246, 'num_leaves': 19, 'min_child_samples': 10}\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n\n    # Same as above, but results from most trials:\n    # Optimization starts.  n_trials =  4000\n    # Best r2: 0.494498 Mse: 20.639\n    # CPU times: user 1d 6h 7min 22s, sys: 37min 27s, total: 1d 6h 44min 50s\n    # Wall time: 8h 7min 9s\n    #  Best params:\n    params = {'n_estimators': 500, 'reg_alpha': 7.950134341890093, 'reg_lambda': 6.950583421079477, 'colsample_bytree': 0.3, 'subsample': 0.6, 'max_depth': 10, 'learning_rate': 0.08004284256392488, 'num_leaves': 4, 'min_child_samples': 6}\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n\n    # Same as above, but results from other trials:\n    # Optimization starts.  n_trials =  3200\n    # Best r2: 0.488103 Mse: 20.900\n    # CPU times: user 21h 35min 47s, sys: 21min 52s, total: 21h 57min 39s\n    # Wall time: 5h 44min 30s\n    # Best params:\n    params = {'n_estimators': 500, 'reg_alpha': 7.461080505689749, 'reg_lambda': 4.442312699482935, 'colsample_bytree': 0.5, 'subsample': 0.6, 'max_depth': 6, 'learning_rate': 0.03168107096628742, 'num_leaves': 34, 'min_child_samples': 11}\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n\n\n    # Optimization starts.  n_trials =  800\n    # Best r2: 0.492957 Mse: 20.702\n    # CPU times: user 7h 4min 20s, sys: 2min 15s, total: 7h 6min 35s\n    # Wall time: 1h 49min 54s\n    # Best params:\n    params = {'n_estimators': 500, 'reg_alpha': 7.609216976336635, 'reg_lambda': 4.412148983414323, 'colsample_bytree': 0.8, 'subsample': 0.7, 'max_depth': 6, 'learning_rate': 0.030037500939340295, 'num_leaves': 19, 'min_child_samples': 6}\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n\n    # Optimization starts.  n_trials =  400\n    # Best r2: 0.484921 Mse: 21.030\n    # CPU times: user 4h 41min 43s, sys: 1min 22s, total: 4h 43min 6s\n    # Wall time: 1h 12min 41s\n    # Best params:\n    params = {'n_estimators': 500, 'reg_alpha': 6.676110060308106, 'reg_lambda': 9.264320368469557, 'colsample_bytree': 0.5, 'subsample': 0.7, 'max_depth': 6, 'learning_rate': 0.033090306623644414, 'num_leaves': 53, 'min_child_samples': 8}\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n\n    # Optimization starts.  n_trials =  200\n    # Best r2: 0.489367 Mse: 20.849\n    # CPU times: user 2h 8min 36s, sys: 41 s, total: 2h 9min 17s\n    # Wall time: 33min 3s\n    # Best params:\n    params = {'n_estimators': 500, 'reg_alpha': 6.734991732483901, 'reg_lambda': 3.0151837674804667, 'colsample_bytree': 0.9, 'subsample': 1.0, 'max_depth': 6, 'learning_rate': 0.03814860248977263, 'num_leaves': 859, 'min_child_samples': 13}\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n\n\n    # Default boosting: \n    params = {} #  {'n_estimators': 500, 'reg_alpha': 6.734991732483901, 'reg_lambda': 3.0151837674804667, 'colsample_bytree': 0.9, 'subsample': 1.0, 'max_depth': 6, 'learning_rate': 0.03814860248977263, 'num_leaves': 859, 'min_child_samples': 13}\n    model = MultiOutputRegressor(lgbm.LGBMRegressor())\n\n    # Manually tuned LGBM  (for CD31) # https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109895207&cellId=35\n    params = {\"random_state\":  0, \n        #extra_trees = True, # For 0.46792863038644295 it worsens - probably need to tune other params \n        \"learning_rate\": 0.07 ,  # hypersensitive - change to 0.0671 - worsens about 2% from 0.43590181064083766\n        \"num_leaves\":  12, # hypersensitive - change to 11, worsens about 2%\n        \"max_depth\":  9, # hypersensitive - change to 10 - worsens about 1%\n        \"n_estimators\": 150, # senstitive - change by 10 - worsents about 0.5%\n        \"min_child_samples\":  20,# hypersensitive - change to 21 - worsens about 3% \n        \"min_split_gain\":  12.2,# sometimes hypersenstive - change by 0.1 worsens by 0.9%, But for 100 features changes 11.8-12.2 - not change AT ALLL !!!  # Uplifts to  0.43590.. !!! \n        \"reg_lambda\":  0.0, # seems only makes worse\n        \"reg_alpha\":  0.0, # seems only makes worse\n        \"subsample_for_bin\":  10000, # seems  less 10000 - worsens, but after 10 000 does not influence  \n        \"colsample_bytree\":  1, # only worsens\n        \"subsample\":  1, # Does not seem to influence at all\n        \"other_rate\":  1,# no influence ?  \n        \"min_child_weight\":  0.1, # no influence ?\n        \"subsample_freq\":  10,# no influence ? # Integer; alias: bagging_freq; k means perform bagging at every k iteration\n             } # \n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n\n    # From https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109895207&cellId=36\n    params = {\n        \"random_state\": 0, \n        \"learning_rate\": 0.09 ,  # hypersensitive - change to 0.0671 - worsens about 2% from 0.43590181064083766\n        \"num_leaves\": 12, # hypersensitive - change to 11, worsens about 2%\n        \"max_depth\": 15, # hypersensitive - change to 10 - worsens about 1%\n        \"n_estimators\": 150, # senstitive - change by 10 - worsents about 0.5%\n        \"min_child_samples\": 50,# hypersensitive - change to 21 - worsens about 3% \n        \"min_split_gain\": 3.119890406724053,# sometimes hypersenstive - change by 0.1 worsens by 0.9%, But for 100 features changes 11.8-12.2 - not change AT ALLL !!!  # Uplifts to  0.43590.. !!! \n        \"reg_lambda\": 0.8575143843171859, # seems only makes worse\n        \"reg_alpha\": 0.0, # seems only makes worse\n        \"subsample_for_bin\": 10000, # seems  less 10000 - worsens, but after 10 000 does not influence  \n        \"colsample_bytree\": 1, # only worsens\n        \"subsample\": 1, # Does not seem to influence at all\n        \"other_rate\": 1,# no influence ?  \n        \"min_child_weight\":  0.1, # no influence ?\n        \"subsample_freq\": 10,# no influence ? # Integer; alias: bagging_freq; k means perform bagging at every k iteration\n            }\n    model = MultiOutputRegressor(lgbm.LGBMRegressor(**params))\n","metadata":{"_uuid":"6c75eb10-c597-4be8-ae69-4b9b7758ba2f","_cell_guid":"138fb1ab-f680-4436-a025-7f352d111f08","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:03.686193Z","iopub.execute_input":"2022-11-07T19:37:03.686559Z","iopub.status.idle":"2022-11-07T19:37:04.299545Z","shell.execute_reply.started":"2022-11-07T19:37:03.686527Z","shell.execute_reply":"2022-11-07T19:37:04.298153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"7187ae1d-25c6-4a59-805d-25c71534afbd","_cell_guid":"61712ef6-6727-4333-a876-f98e9388b9be","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Core model train/predict\n","metadata":{"_uuid":"d805f54e-a486-41d0-a8e7-ca007a9be5b3","_cell_guid":"2b79a8e7-be1c-49e4-97f9-33af416ca554","trusted":true}},{"cell_type":"code","source":"%%time\ndict_folds_info = dict_main_folds_info\n\nflag_train_on_full_data4submit = True\nverbose = 1000\n\ndict_model_results = {}\n\nlist_folds_indices = dict_folds_info['List Indices'] \nt0 = time.time()\nY_oof_pred = np.zeros_like(Y)\nY_submission_way1_pred = np.zeros( (119651-70988 , Y.shape[1]) )  # Way 1 (Classical way) - retrain model on the whole data\nY_submission_way2_pred = np.zeros( (119651-70988 , Y.shape[1]) )  # Way 2 (Kaggle way) - average predictions of models on each fold\nfor fold, indices_tuple  in enumerate( list_folds_indices ):\n    \n    t00 = time.time()\n    train_index = indices_tuple[0]\n    model.fit(X[train_index], Y[train_index])\n    if verbose >= 100:\n        print('fold:',fold, 'Train time:%.2f'%(time.time()-t00 ), 'X_train.shape:',X[train_index].shape )\n        \n    # Calculate predictions on test and train folds \n    for i_loc in range(0,len( indices_tuple)  ):\n        indices_loc = indices_tuple[i_loc]\n        y_pred = model.predict( X[indices_loc ])\n        dict_model_results['Pred_fold'+str(fold)+'_part'+str(i_loc)] = y_pred\n        if i_loc > 0: # Skip train \n            Y_oof_pred[indices_loc] += y_pred * dict_main_folds_info['OOF Multiplicity'][i_loc] \n    # Prepare submission data in way 2 - \"Kaggle way\"\n    Y_submission_way2_pred += model.predict( X[70988:,: ])  # Way 2 (Kaggle way) - average predictions of models on each fold   \n\n# Main loop over folds - finished just do some final calculations:  \nY_submission_way2_pred /= (fold+1)      \nif flag_train_on_full_data4submit: \n    model.fit(X[:70988,:], Y) # Way 1 (Classical way) - retrain model on the whole data\n    Y_submission_way1_pred = model.predict( X[70988:,: ])   \n\nif verbose >= 10:\n    print('X.shape, Y.shape:',X.shape, Y.shape)\n    print('Finished. Folds:',fold, 'Time:%.2f'%(time.time()-t0), 'Y_oof_pred.shape:',Y_oof_pred.shape,\n         'Y_submission_way2_pred.shape:',Y_submission_way2_pred.shape )\n\n\ndict_model_results['Time'] = time.time() - t0\ndict_model_results['Y_pred'] = Y_oof_pred\ndict_model_results['OOF Indices Main'] = dict_folds_info['OOF Indices Main'] \ndict_model_results['OOF Indices Additional'] = dict_folds_info['OOF Indices Additional'] \ndict_model_results['Submit1'] = Y_submission_way1_pred\ndict_model_results['Submit2'] = Y_submission_way2_pred\ndict_model_results['Dict Folds Info' ]= dict_folds_info.copy() # \ndict_model_results['df_meta'] = df_meta # Just save, just in case","metadata":{"_uuid":"bc1f04c7-6ff6-448c-a440-59ebc8c2364a","_cell_guid":"92dc558c-a17b-4e29-a9f0-137bfcdc4b6e","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:04.301796Z","iopub.execute_input":"2022-11-07T19:37:04.302326Z","iopub.status.idle":"2022-11-07T19:37:08.254157Z","shell.execute_reply.started":"2022-11-07T19:37:04.302281Z","shell.execute_reply":"2022-11-07T19:37:08.242083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Main SCORING - caculate and save MAIN scores","metadata":{"_uuid":"78f3fc8c-93e8-4624-9f44-68f05d41a87c","_cell_guid":"641df5a1-559e-4329-a442-426ad974c7ed","trusted":true}},{"cell_type":"code","source":"%%time\nY_pred = dict_model_results['Y_pred']\nindices_oof = dict_model_results['OOF Indices Main']\nY_pred = Y_pred[  indices_oof , : ]\nY_true = Y[  indices_oof , : ]\n\ns_cor = correlation_score( Y_true , Y_pred  ) # It takes about 5 seconds ! \n\ns_r2 = r2_score( Y_true , Y_pred  )\ns_mse = mean_squared_error( Y_true , Y_pred  )\n\nprint(str(model)[:40], 'n_feat:', X.shape[1],)\nprint('OOF scores: ', s_cor, s_r2, s_mse, 'n_feat:', X.shape[1], 'Oof size:', Y_true.shape[0] )","metadata":{"_uuid":"11fcbb93-bb7d-4775-808c-a5da8763a6fe","_cell_guid":"75243105-be40-40f7-8f04-626031e7fc97","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:08.256544Z","iopub.execute_input":"2022-11-07T19:37:08.257125Z","iopub.status.idle":"2022-11-07T19:37:08.693724Z","shell.execute_reply.started":"2022-11-07T19:37:08.257058Z","shell.execute_reply":"2022-11-07T19:37:08.692095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"    100 features - better than 40 for LGBM \n    MultiOutputRegressor LightGBM: (Calculation time - 19 minutes)\n    OOF scores:  0.8843042 0.14324448260463332 0.21749122 n_feat: 100 Oof size: 57500\n\n    On 200 Features LGBM is sligly worse than on 100:\n    OOF scores:  0.88414896 0.14308675934626502 0.21776518 n_feat: 200 Oof size: 57500 (36 minutes - version 10 ) \n\n    MultiOutputRegressor(estimator=LGBMRegre n_feat: 40\n    OOF scores:  0.88402647 0.14195155980790192 0.21797003 n_feat: 40 Oof size: 57500\n\n    MultiOutputRegressor LightGBM: (Calculation time - 19 minutes)\n    OOF scores:  0.8843042 0.14324448260463332 0.21749122 n_feat: 100 Oof size: 57500\n\n    Ridge(alpha=10000.0) n_feat: 100\n    OOF scores:  0.8849721 0.1520688839180771 0.21653545 n_feat: 100 Oof size: 57500\n\n    Ridge 50: \n    OOF scores:  0.88314193 0.1486980690002863 0.21981749\n    MeanByFolds Test1: 0.882145 Test2: 0.883643\n    \n    Ridge 100:\n    OOF scores:  0.8849221 0.15140698792122334 0.21665995\n    MeanByFolds Test1: 0.883907\t0.136736\t0.218520\t0.885304\t0.131960\t0.215981\t0.885119\t0.136830\t0.216512\t0.886379\t0.135263\t0.214223\t0.887717\t0.158759\t0.211723\n    \n    Ridge 150\n    OOF scores:  0.88490844 0.15052967052133392 0.21670045\n    MeanByFolds Test1:  0.883843\t0.135412\t0.218653\t0.885232\t0.130625\t0.216116\t0.885073\t0.135405\t0.216610\t0.886348\t0.134102\t0.214291\t0.887650\t0.157636\t0.211866\n    \n    Ridge 200\n    OOF scores:  0.88479453 0.1494652531252453 0.21691151\n    MeanByFolds Test1: 0.883674\t0.133911\t0.218972\t0.885063\t0.129137\t0.216421\t0.884907\t0.133874\t0.216923\t0.886183\t0.132540\t0.214590\t0.887495\t0.156314\t0.212149","metadata":{"_uuid":"e513a1e3-2034-472b-8065-932dc87fa23f","_cell_guid":"e3ea991f-5a6b-4e71-9c71-665a7b4590b7","trusted":true}},{"cell_type":"code","source":"","metadata":{"_uuid":"e118e4fb-e518-4bb5-bf27-d4b172e50e4a","_cell_guid":"e6336559-e353-4fe6-8746-baaa1df0e6b5","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Additional info on scores - foldwise","metadata":{"_uuid":"4e08dde7-b0ff-49d5-b32a-5cfdebdd5320","_cell_guid":"4083aaf8-b9b3-466a-9758-5816a714b159","trusted":true}},{"cell_type":"code","source":"%%time\nverbose = 1000\nlist_folds_indices =  dict_model_results['Dict Folds Info']['List Indices']\nlist_folds_indices_info = dict_model_results['Dict Folds Info']['Info'] \ndf_fold_score_stat = pd.DataFrame();df_fold_score_stat.index.name = 'Fold'\nt0 = time.time()\nfor fold, indices_tuple  in enumerate( list_folds_indices ):\n    # Calculate metrics on test and train folds \n    list_scores = []; list_scores_r2 = []\n    for i_loc, indices_loc in enumerate( indices_tuple) :\n        y_pred = dict_model_results['Pred_fold'+str(fold)+'_part'+str(i_loc)]\n        name_loc = list_folds_indices_info[i_loc]\n        s = correlation_score(Y[indices_loc],  y_pred    )\n        df_fold_score_stat.loc[fold,'Corr '+name_loc] = s\n        s = r2_score( Y[indices_loc] , y_pred  )\n        df_fold_score_stat.loc[fold,'r2 '+name_loc] = s\n        s = mean_squared_error( Y[indices_loc] , y_pred  )\n        df_fold_score_stat.loc[fold,'mse '+name_loc] = s\n    if verbose >= 1000:\n        print('Fold:', fold, 'Shapes of train:', train_index.shape, 'Tests: ',[t.shape for t in indices_tuple[1:] ] )        \ndisplay( df_fold_score_stat )\ndisplay(df_fold_score_stat.describe(percentiles=[] ).iloc[1:,:])","metadata":{"_uuid":"909dcb65-0c59-475f-bc3a-b6cdcd4aaab4","_cell_guid":"bfd1aaae-bdd8-4d05-a031-184493215821","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:08.696052Z","iopub.execute_input":"2022-11-07T19:37:08.696585Z","iopub.status.idle":"2022-11-07T19:37:17.227329Z","shell.execute_reply.started":"2022-11-07T19:37:08.696522Z","shell.execute_reply":"2022-11-07T19:37:17.225983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization of Scores TargetWISELY - Optional look on scores\n\nAs it is expected high MSE is highly correlated to just high Mean.\n\nProbably those targets are contribute most to the correlation metric.","metadata":{"_uuid":"efed7dc6-ff06-4901-bdad-7890d97164d7","_cell_guid":"86405e0b-95fd-4dfc-85a2-64edcede60ab","execution":{"iopub.status.busy":"2022-11-05T09:07:34.033666Z","iopub.execute_input":"2022-11-05T09:07:34.034094Z","iopub.status.idle":"2022-11-05T09:07:34.044978Z","shell.execute_reply.started":"2022-11-05T09:07:34.034057Z","shell.execute_reply":"2022-11-05T09:07:34.043899Z"},"trusted":true}},{"cell_type":"code","source":"pd.set_option('display.max_columns', None)  # or 1000\npd.set_option('display.max_rows', None)  # or 1000\npd.set_option('display.max_colwidth', None)  # or 199","metadata":{"_uuid":"20e4ff28-683c-4a9b-b03e-6a9b815442ce","_cell_guid":"fe2e95aa-7eb1-49ec-86f3-3ae68bffac6d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:17.229211Z","iopub.execute_input":"2022-11-07T19:37:17.229994Z","iopub.status.idle":"2022-11-07T19:37:17.236019Z","shell.execute_reply.started":"2022-11-07T19:37:17.229948Z","shell.execute_reply":"2022-11-07T19:37:17.234903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ndf_targetwise = pd.DataFrame()\n\nY_pred = dict_model_results['Y_pred']\nindices_oof = dict_model_results['OOF Indices Main']\nY_pred = Y_pred[  indices_oof , : ]\nY_true = Y[  indices_oof , : ]\nfor i,col in enumerate(df_cite_train_y.columns):\n    IX = col#.columns[i_loc]\n    df_targetwise.loc[IX, 'r2'] = r2_score( Y_true[:,i] , Y_pred[:,i]  )\n    df_targetwise.loc[IX, 'mse'] = mean_squared_error( Y_true[:,i] , Y_pred[:,i]  )    \n    df_targetwise.loc[IX, 'corr'] = np.corrcoef( Y_true[:,i] , Y_pred[:,i]  )[0,1]    \n    df_targetwise.loc[IX, 'std'] = df_cite_train_y[col].std()\n    df_targetwise.loc[IX, 'absMean'] = df_cite_train_y[col].abs().mean()\n    df_targetwise.loc[IX, 'mean'] = df_cite_train_y[col].mean()\n    df_targetwise.loc[IX, 'max'] = df_cite_train_y[col].max()\n    \nprint('Pearson correlation between scores (and other characteristics):')\ndisplay(df_targetwise.corr()   )\nprint('Spearman correlation between scores (and other characteristics):')\ndisplay(df_targetwise.corr(method = 'spearman')   )\n\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['r2'].plot(); plt.legend(); plt.grid();\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['mse'].plot(); plt.legend(); plt.grid();\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['corr'].plot(); plt.legend(); plt.grid();\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['absMean'].plot(); plt.legend(); plt.grid();\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['std'].plot(); plt.legend(); plt.grid();\n\nd1 = df_targetwise.sort_values('r2'); d1.index.name = 'r2 sorted';  d1 = d1.reset_index()\nd2 = df_targetwise.sort_values('mse',ascending = False); d2.index.name = 'mse sorted';  d2 = d2.reset_index()\nd3 = df_targetwise.sort_values('corr'); d3.index.name = 'corr sorted';  d3 = d3.reset_index()\nd = pd.concat([d1,d2,d3],axis = 1)\nprint('\\n Worst:')\ndisplay(d.head(10))\nprint('\\n Best:')\ndisplay(d.tail(10))","metadata":{"_uuid":"21700999-98c8-42b6-8448-aa5037e51d15","_cell_guid":"21362122-af7a-4872-a2bf-0ae98248af08","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:17.237856Z","iopub.execute_input":"2022-11-07T19:37:17.238295Z","iopub.status.idle":"2022-11-07T19:37:20.61046Z","shell.execute_reply.started":"2022-11-07T19:37:17.238257Z","shell.execute_reply":"2022-11-07T19:37:20.608896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization of the main scores - correlations","metadata":{"_uuid":"31bed355-fa2d-4159-ae2d-5690b3da64c7","_cell_guid":"cb95733a-ccce-43fb-b4e6-f98efcd98c66","trusted":true}},{"cell_type":"code","source":"%%time\ncol = 'OOF Indices Main'\nprint('Predictions on:', col)\nY_pred = dict_model_results['Y_pred']\nindices_oof = dict_model_results[col]\nY_pred = Y_pred[  indices_oof , : ]\nY_true = Y[  indices_oof , : ]\nprint('Y_pred.shape: ', Y_pred.shape, 'Y_true.shape: ', Y_true.shape,  )\nlist_corr = []\nfor i in range(Y_pred.shape[0]):\n    list_corr.append(np.corrcoef(Y_true[i,:], Y_pred[i,:] )[0,1] )\n    \n\nfig = plt.figure( figsize = (10,5))\nplt.title('Main score - correlation \\n'+str(col))\nplt.hist(list_corr, bins = 100)\nplt.grid()\nplt.show()\ndisplay(pd.Series(list_corr).describe() )","metadata":{"_uuid":"1b23a96e-e7aa-414f-b652-264dcc8ad132","_cell_guid":"cafa68bd-7f64-4cb7-ba78-6f1c7bf0a41a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:20.612396Z","iopub.execute_input":"2022-11-07T19:37:20.61283Z","iopub.status.idle":"2022-11-07T19:37:27.450432Z","shell.execute_reply.started":"2022-11-07T19:37:20.612773Z","shell.execute_reply":"2022-11-07T19:37:27.448898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ncol = 'OOF Indices Additional'\nprint('Predictions on:', col)\nY_pred = dict_model_results['Y_pred']\nindices_oof = dict_model_results[col]\nY_pred = Y_pred[  indices_oof , : ]\nY_true = Y[  indices_oof , : ]\nprint('Y_pred.shape: ', Y_pred.shape, 'Y_true.shape: ', Y_true.shape,  )\nlist_corr = []\nfor i in range(Y_pred.shape[0]):\n    list_corr.append(np.corrcoef(Y_true[i,:], Y_pred[i,:] )[0,1] )\n    \n\nfig = plt.figure( figsize = (10,5))\nplt.title('Main score - correlation \\n'+str(col))\nplt.hist(list_corr, bins = 100)\nplt.grid()\nplt.show()\ndisplay(pd.Series(list_corr).describe() )","metadata":{"_uuid":"028fb6ab-aea0-4e64-95b9-36af75927beb","_cell_guid":"7f8e634a-cdf3-470d-b783-15cb45325b3b","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:27.452555Z","iopub.execute_input":"2022-11-07T19:37:27.453091Z","iopub.status.idle":"2022-11-07T19:37:28.622584Z","shell.execute_reply.started":"2022-11-07T19:37:27.453038Z","shell.execute_reply":"2022-11-07T19:37:28.621043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preparation of the submission file","metadata":{"_uuid":"752f9889-c538-47ad-bb16-853de39e16bc","_cell_guid":"50a47445-c6e4-4686-bff1-53a874ded929","jupyter":{"outputs_hidden":false},"trusted":true}},{"cell_type":"code","source":"%%time\n# Normally that part works Wall time: 1min 11s - however sometimes Kaggle system does not work correctly and it never stops\nif flag_prepare_submission:\n    mode_subm = 'USE: sskknts MSCI CITEseq Keras Quickstart + Dropout' # 'put_zeros_to_multiome_part'\n\n    if mode_subm == 'USE: sskknts MSCI CITEseq Keras Quickstart + Dropout' :\n        fn_loc = '/kaggle/input/msci-citeseq-keras-quickstart-dropout/submission.csv'\n        fn_loc = '/kaggle/input/solutions-for-multimodal-singlecell-integration/msci-citeseq-keras-quickstart-dropout_submission.csv'\n        df_submission_full = pd.read_csv(fn_loc,\n                                 index_col='row_id', squeeze=True)\n        df_submission_full = df_submission_full.to_frame()\n    elif mode_subm == 'put_zeros_to_multiome_part':\n        df_submission_full = pd.DataFrame(index = range(65_744_180), columns = ['target'], data = np.zeros(65_744_180) )\n        df_submission_full.index.name = 'row_id'\n\n    display(df_submission_full.info() )\n    print()\n    display(df_submission_full.head(10))","metadata":{"_uuid":"9861b2e6-dc33-4c62-9965-81dc2eaf3d93","_cell_guid":"e3d08b1e-d469-439b-ba7a-5299bc95ddab","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:37:28.624231Z","iopub.execute_input":"2022-11-07T19:37:28.624745Z","iopub.status.idle":"2022-11-07T19:38:37.957894Z","shell.execute_reply.started":"2022-11-07T19:37:28.624708Z","shell.execute_reply":"2022-11-07T19:38:37.956359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"2+2","metadata":{"execution":{"iopub.status.busy":"2022-11-07T19:38:37.959352Z","iopub.execute_input":"2022-11-07T19:38:37.959824Z","iopub.status.idle":"2022-11-07T19:38:37.967864Z","shell.execute_reply.started":"2022-11-07T19:38:37.959774Z","shell.execute_reply":"2022-11-07T19:38:37.966564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif flag_prepare_submission:\n    #Y_pred4submit = dict_model_results['Submit1']\n    Y_pred4submit = dict_model_results['Submit2']\n    print(Y_pred4submit.shape)","metadata":{"_uuid":"1230587a-49e6-45f3-9883-253b244ef753","_cell_guid":"797b5aed-7533-43c4-a594-cd4099676bd9","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:38:37.96952Z","iopub.execute_input":"2022-11-07T19:38:37.97009Z","iopub.status.idle":"2022-11-07T19:38:37.98416Z","shell.execute_reply.started":"2022-11-07T19:38:37.970049Z","shell.execute_reply":"2022-11-07T19:38:37.982611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif flag_prepare_submission:\n    # We can prepare CITE-seq part as: Y_pred4submit.values.ravel() -- see notebook: https://www.kaggle.com/code/alexandervc/mmscel-analysis-of-evaluation-and-submission\n    df_submission_full['target'].iloc[:6_812_820] = Y_pred4submit.ravel()\n    display(df_submission_full.head(20))\n    display(df_submission_full.tail(20))\n    ","metadata":{"_uuid":"f175b5c0-9177-492b-ab43-d160e325f4fb","_cell_guid":"98b5d7cf-5741-4f0e-ab4f-aa2377679b55","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:38:37.985952Z","iopub.execute_input":"2022-11-07T19:38:37.986344Z","iopub.status.idle":"2022-11-07T19:38:38.028884Z","shell.execute_reply.started":"2022-11-07T19:38:37.986313Z","shell.execute_reply":"2022-11-07T19:38:38.027275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Wall time: 1min 46s\nif flag_prepare_submission:\n    submit_filename_postfix = 'CVadvV5_Ridge100'\n    df_submission_full.to_csv('submission_cite_seq_'+submit_filename_postfix +'.csv')","metadata":{"_uuid":"72332661-0de3-4fc2-a769-716579e47c35","_cell_guid":"d833b02f-6cb9-49fa-ac3a-7ab7eb61f91d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:38:38.030631Z","iopub.execute_input":"2022-11-07T19:38:38.031272Z","iopub.status.idle":"2022-11-07T19:40:59.819584Z","shell.execute_reply.started":"2022-11-07T19:38:38.031219Z","shell.execute_reply":"2022-11-07T19:40:59.818354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"b372da4f-a87d-4b71-a071-414e63564d37","_cell_guid":"de9b5839-be50-4f31-8fae-21d9a4b8cd7e","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"5d8d0a31-a09c-490d-8387-464a59f1dd23","_cell_guid":"5ca85d26-228a-4f42-910c-a380f4827022","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )","metadata":{"_uuid":"f55e2443-fd24-4fc6-9238-fec592d87d02","_cell_guid":"d3eca9d9-7cdc-4323-beec-8e7d1c64fc84","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T19:40:59.821619Z","iopub.execute_input":"2022-11-07T19:40:59.822734Z","iopub.status.idle":"2022-11-07T19:40:59.830447Z","shell.execute_reply.started":"2022-11-07T19:40:59.822675Z","shell.execute_reply":"2022-11-07T19:40:59.829026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"c1bbfdc5-d83f-44e3-8a64-6d4b2f4ff50b","_cell_guid":"a27b9c08-7e28-47e3-a2de-5a4eed9df4a2","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]}]}