{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nSome advanced modeling planned to be here... \nSeveral models, blending, tuning, etc...\n\nIdeas planned to be implemented:\n\n0) Select some part of dataset for \"playground\", it can be used for preliminary feature selection and/or preparing stacking and/or quick experiments and/or varying that part we will create more diversity for models, which can be blended later. \n\n1) Various models, hyper-param tuning\n\n2) Blending possibly targetwisely, possibly with solutions defined only for some targets\n\n3) Optimization of hyper-params for blend, not for individual performance\n\n4) ...\n\nBased on previous scripts:\nhttps://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold\n\n\n### About CV scheme\n\nPublic and private test sets are quite different. (Private - contains new DAY (and donor), while public ONLY new donor).\nSo one should be careful with the validation schemes.\nSome proposals are described here: \n\nNotebook: https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\nTopic: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\n\nWe will be based on them. \n\n### Option to create \"Playground\" \n\nSome part of the data can be optionally distinguished. And used for preliminary feature selection and/or preparing stacking and/or quick experiments and/or varying that part we will create more diversity for models, which can be blended later. \n\n\n### Versions:\n\n#### 8, 9, 10 , 11 , 12\n\n    Lightgbm model from \n    https://www.kaggle.com/code/yaroshenko/mmscel-crossvalidation-schemes?scriptVersionId=109641833&cellId=50\n    \n    On 100 features it is slightly worse that Ridge 100, 150, 200 ..  (Version 8 - 20 minutes) \n    OOF scores:  0.8843042 0.14324448260463332 0.21749122 n_feat: 100 Oof size: 57500\n    On 40 Features LGBM is sligly worse than on 100: (Version 9 - 11 minutes to run)\n    OOF scores:  0.88402647 0.14195155980790192 0.21797003 n_feat: 40 Oof size: 57500\n    On 200 Features LGBM is sligly worse than on 100:\n    OOF scores:  0.88414896 0.14308675934626502 0.21776518 n_feat: 200 Oof size: 57500 (36 minutes - version 10 ) \n    \n#### 5,6, 7\n    Added: \n    Submission preparation - multiome part - taken from public (Normally that part takes about 3 minutes, however currently Kaggle does not work correctly and that part never finishes)\n    Param: flag_prepare_submission = True/False - to switch off (for speed) that part\n\n    Problem: Kaggle system works not fine and some problems occur at loading the file with multiome part. \n\n    More updates to speed-up correlation score calculation\n    \n#### 2 , 3 , 4\n\n    Added:\n    Ridge model with 100 (on CV it seems better than 50 and 150)\n    Foldwise analysis of scores.\n    Speed-up for correlation metric. \n    \n    Ridge 100 features - Wall time: 3.58 s - on all training and predictions (very fast). \n    The longer part - calculation of correlation metric. \n\n#### 1 - draft\n\n    Creation of Playground and \"Mainground\" (with respect to statification by the day donor cell-type).\n    Modeling:\n    Trivial model - Ridge over 1 feature - so it is basically the same as MEAN prediction. \n    Calculation of the model prediction according to CV-scheme described above.\n    In particular  \"OOF\" (out-of-fold) predictions - that is not straigthford since CV schem is not standard one. \n    Scoring:\n    Main scores (correlation) calculation.\n    Analysis: \n    Histograms of scores  \n    Analysis \"targetwisely\" i.e. for each particular target - look on r2, mse, corr scores.","metadata":{"_uuid":"e85ad2dd-4548-45c2-bd22-91d2a8e4c38a","_cell_guid":"6f7edf4e-a866-408c-9f90-083b4e2ba675","trusted":true}},{"cell_type":"markdown","source":"# Key params","metadata":{"_uuid":"51600441-6c40-4f20-a165-f4380cbb5626","_cell_guid":"cc51e09c-c472-4c65-b21d-230e27930807","trusted":true}},{"cell_type":"code","source":"playground_size = 0.75 # from 0 to 1 \nholdout_size = 0.8# from 0 to 1 \nrandom_seed4folds = 42 \nn_features = 2 # use first N features - to make quick experiments\n\nrescale_Y_to_mean0_std1 = True\n\nflag_prepare_submission = False # \n\n#selected_targets = 'ALL'# 'CD31'","metadata":{"_uuid":"6b246588-79af-4af7-aeaa-28c40ef039e5","_cell_guid":"90284018-0e3f-4e49-941a-d77ce2d84e9d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:13:41.217234Z","iopub.execute_input":"2022-11-07T08:13:41.218347Z","iopub.status.idle":"2022-11-07T08:13:41.248630Z","shell.execute_reply.started":"2022-11-07T08:13:41.218236Z","shell.execute_reply":"2022-11-07T08:13:41.247731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Install/import modules, load technical data","metadata":{"_uuid":"a42ce9d6-ee5d-4ec0-9c2a-31710ef3b8b5","_cell_guid":"f6b5ae5b-a81f-47da-8027-c514cc207aac","trusted":true}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"104f0242-958b-4ef1-af40-8bec2dc9bea0","_cell_guid":"98e6bc60-36d3-4d45-855a-11748f831a0d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:13:41.326642Z","iopub.execute_input":"2022-11-07T08:13:41.326925Z","iopub.status.idle":"2022-11-07T08:13:41.380778Z","shell.execute_reply.started":"2022-11-07T08:13:41.326901Z","shell.execute_reply":"2022-11-07T08:13:41.379740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndisplay(df_cell)","metadata":{"_uuid":"01798d74-2f72-4000-b474-cdb4d74e5f3e","_cell_guid":"d3410584-a62c-4543-8fac-a4fe8fe946f6","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:13:41.383245Z","iopub.execute_input":"2022-11-07T08:13:41.384056Z","iopub.status.idle":"2022-11-07T08:14:07.166085Z","shell.execute_reply.started":"2022-11-07T08:13:41.384018Z","shell.execute_reply":"2022-11-07T08:14:07.164945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"_uuid":"0bdd40f9-eb9c-4eaa-9e08-5e91c0ed728b","_cell_guid":"bfdcb1cc-f7ee-4688-a601-d8661a95ec53","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:14:07.168602Z","iopub.execute_input":"2022-11-07T08:14:07.169009Z","iopub.status.idle":"2022-11-07T08:14:27.394959Z","shell.execute_reply.started":"2022-11-07T08:14:07.168968Z","shell.execute_reply":"2022-11-07T08:14:27.393656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load prepared Features for CITE-seq part of task","metadata":{"_uuid":"32413e3a-7b7f-4a2e-b028-4410b034ea14","_cell_guid":"d47e3bef-4916-43c5-921c-512ca1df8b9d","trusted":true}},{"cell_type":"code","source":"%%time\n\nprint('Load prepared features for CITE-seq')\n# These files contain both train and test parts .\n# For CITEseq part - first 70988 elements - train, and later 48663 - test. Overall 119651 samples.\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_TruncatedSVD200_niter7_rs42.csv'\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA500.csv'\ndf_cite = pd.read_csv(fn,index_col = 0)\ndisplay(df_cite)","metadata":{"_uuid":"14d13305-d62b-4822-b783-09f607eb2cb6","_cell_guid":"0495e7d2-ea36-469e-819a-26ab27a41c58","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:14:27.396574Z","iopub.execute_input":"2022-11-07T08:14:27.397498Z","iopub.status.idle":"2022-11-07T08:14:45.062378Z","shell.execute_reply.started":"2022-11-07T08:14:27.397455Z","shell.execute_reply":"2022-11-07T08:14:45.058152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_cite.mean(axis = 0).head(3))\nprint('Pay attention - PCA features are ordered by magnitude:')\nprint(df_cite.std(axis=0).head(5))","metadata":{"_uuid":"afe9baa3-48d2-4d56-b027-159880e77522","_cell_guid":"bc595bf0-e292-4bfa-bd19-821051bf8579","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:14:45.065712Z","iopub.execute_input":"2022-11-07T08:14:45.066134Z","iopub.status.idle":"2022-11-07T08:14:45.833386Z","shell.execute_reply.started":"2022-11-07T08:14:45.066092Z","shell.execute_reply":"2022-11-07T08:14:45.832053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Targets for CITE-seq","metadata":{"_uuid":"11f0481b-2dbc-4835-b7c2-98755ca1f63b","_cell_guid":"2e5213c8-69c3-4832-a5ce-f2471ccabb52","trusted":true}},{"cell_type":"code","source":"%%time\n#if 1:\nprint('Load CITE-seq targets and ')\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndisplay(df_cite_train_y)","metadata":{"_uuid":"93f0dd7a-8a6a-4224-bc07-ff982ab834f7","_cell_guid":"915efa3d-6965-4b7f-a759-d1e3848d165d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:14:45.835484Z","iopub.execute_input":"2022-11-07T08:14:45.835932Z","iopub.status.idle":"2022-11-07T08:14:46.586147Z","shell.execute_reply.started":"2022-11-07T08:14:45.835875Z","shell.execute_reply":"2022-11-07T08:14:46.584062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load and start prepare some metadata (cut only CITE-seq train part)","metadata":{"_uuid":"60de91b6-5173-4461-9dd6-7094e3173f1a","_cell_guid":"3e9c27ac-e90e-44a4-ab94-dda2b95239f9","trusted":true}},{"cell_type":"code","source":"%%time\nfn2 = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/_citeseq_meta_all_text_also.csv'\ndf_meta_full = pd.read_csv(fn2,index_col = 0)\ndisplay(df_meta_full)\n#if 1:\ndf_meta = pd.DataFrame(index = df_cite_train_y.index) \ndf_meta = df_meta.join(df_cell.set_index('cell_id') )\ndisplay(df_meta)","metadata":{"_uuid":"a3de63b6-a27d-4924-9764-6ef3fcf03c1d","_cell_guid":"f5fa99ae-4511-45ec-907c-b872e8744acf","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:14:46.588617Z","iopub.execute_input":"2022-11-07T08:14:46.589501Z","iopub.status.idle":"2022-11-07T08:14:46.906410Z","shell.execute_reply.started":"2022-11-07T08:14:46.589447Z","shell.execute_reply":"2022-11-07T08:14:46.905263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create \"Playground\" and   (i.e. downsample)","metadata":{"_uuid":"7a5aa539-257d-4e8b-9bcd-61843a2ebc5a","_cell_guid":"499ad22c-c00c-441c-94a7-f069d2dcecb3","trusted":true}},{"cell_type":"code","source":"%%time\n# Create columns in df_meta which label the \"Playground\" and \"Holdout\"\nfrom sklearn.model_selection import train_test_split\n\ndf_meta['Index'] = range(len(df_meta))\nscol = 'donor&day&CT'\ndf_meta[scol] =df_meta['donor'].apply(lambda x:str(x)+'_') + df_meta['day'].apply(lambda x:str(x)+'_') + df_meta['cell_type']\n\n\nflagged_column_name = 'Playground'\n# if playground_size > 0 :\n#     d_tmp_train, d_tmp_test = train_test_split(df_meta, test_size = playground_size, random_state= random_seed4folds, stratify = df_meta[scol] ,  shuffle=True )\n#     df_meta[flagged_column_name] = df_meta['Index'].isin(d_tmp_test['Index'].values)\n# else:\nd_tmp_train = df_meta\ndf_meta[flagged_column_name] = False\nprint('flagged_column_name:',df_meta[flagged_column_name].sum(), df_meta[flagged_column_name].sum()/len(df_meta) )\n\nflagged_column_name = 'Holdout'\n# if holdout_size > 0 :\n#     d_tmp_train, d_tmp_test = train_test_split(d_tmp_train, test_size = holdout_size, random_state= random_seed4folds, stratify = d_tmp_train[scol] ,  shuffle=True )\n#     df_meta[flagged_column_name] = df_meta['Index'].isin(d_tmp_test['Index'].values)\n# else:\ndf_meta[flagged_column_name] = False\nprint('flagged_column_name:', df_meta[flagged_column_name].sum(), df_meta[flagged_column_name].sum()/len(df_meta) )\n\ndisplay(df_meta)","metadata":{"_uuid":"9b4dc5b2-2978-466d-b4c4-1ac1e73baa13","_cell_guid":"d2b7ecb9-51df-4676-9b96-20f8dac5702b","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:28:16.359616Z","iopub.execute_input":"2022-11-07T08:28:16.360040Z","iopub.status.idle":"2022-11-07T08:28:16.460030Z","shell.execute_reply.started":"2022-11-07T08:28:16.360005Z","shell.execute_reply":"2022-11-07T08:28:16.458892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create validation/folds schemes data","metadata":{"_uuid":"5109c030-65a8-4e67-9cae-0aa01392cec1","_cell_guid":"c3d12c63-4e9a-45dd-9c6b-ea503d434ddf","trusted":true}},{"cell_type":"markdown","source":"## Create Playground folds Structure","metadata":{"_uuid":"eb88a2cb-94a1-4c36-96db-091aa57696ba","_cell_guid":"cda7479e-f9e6-4543-beab-498b95470527","execution":{"iopub.status.busy":"2022-11-05T08:03:06.263638Z","iopub.execute_input":"2022-11-05T08:03:06.264028Z","iopub.status.idle":"2022-11-05T08:03:06.273381Z","shell.execute_reply.started":"2022-11-05T08:03:06.263996Z","shell.execute_reply":"2022-11-05T08:03:06.272164Z"},"trusted":true}},{"cell_type":"code","source":"%%time\n# Validation scheme for playground\n\ndict_playground_folds_info = {}\ndict_playground_folds_info['name'] = 'Playground Folds'\ndict_playground_folds_info['Info'] = ['Train','Test1 Priv Like','Test2 Public Like', 'Not Playground' ]\ndict_playground_folds_info['OOF Multiplicity'] = [0,1/2,0,1/2 ]\ndict_playground_folds_info['OOF Indices Main'] = df_meta['Index'][df_meta['Playground'] == 1].values\ndict_playground_folds_info['OOF Indices Additional'] = df_meta['Index'][df_meta['Playground'] == 0].values\n\n#dict_playground_folds_info['List Indices'] = list_folds_main    \n\n\nlist_folds_playground = []\nif playground_size > 0: \n    mask_playground = (df_meta['Playground'] == 1)\n    c = 0\n    for day2exclude in [2,3,4]:\n        for donor2exclude in [32606,  31800]: # We will need to predict always MALE (not female) - like on LB. (# donor 13176 - female)\n            train_index =\\\n                np.where(mask_playground &  (df_meta['day']  != day2exclude) & ( df_meta['donor']  != donor2exclude  ) )  [0]\n            test_index1_like_private_lb =\\\n                np.where( mask_playground &  (df_meta['day']  == day2exclude)  ) [0]\n            test_index2_like_public_lb =\\\n                np.where(mask_playground &  (df_meta['day']  != day2exclude) &  (df_meta['donor']  == donor2exclude ) ) [0]\n\n            test_outside_playground = np.where( (mask_playground == 0) &  (df_meta['day']  == day2exclude)   )[0]\n\n            list_folds_playground.append( (train_index,  test_index1_like_private_lb , test_index2_like_public_lb,\n                                          test_outside_playground ) )\n\n            str_fold_inf = 'Fold ' +str(c) + ': Train: excludes Day '+str(day2exclude) + ' and Donor ' + str( donor2exclude )\n            print(str_fold_inf, 'Sizes: train:',len(train_index), 'Test Like Priv'  ,len(test_index1_like_private_lb), \n                  'Test Like Publ',   len(test_index2_like_public_lb),  \n                  'test_outside_playground',   len(test_outside_playground),  \n                 ); c+=1\n\ndict_playground_folds_info['List Indices'] = list_folds_playground                \nprint(len(list_folds_playground))","metadata":{"_uuid":"a0109ea6-317a-4d84-a02b-75ce1b107e14","_cell_guid":"61867b0c-cbd7-4a90-b9bc-481e793a0fbe","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:28:18.783192Z","iopub.execute_input":"2022-11-07T08:28:18.785003Z","iopub.status.idle":"2022-11-07T08:28:18.812804Z","shell.execute_reply.started":"2022-11-07T08:28:18.784953Z","shell.execute_reply":"2022-11-07T08:28:18.811361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create Main Folds Structure","metadata":{"_uuid":"8c6c6d8f-3c62-4f5e-b93f-3e219ce6ea03","_cell_guid":"2276ca0f-7b17-4f18-9a22-b5e793542771","execution":{"iopub.status.busy":"2022-11-05T08:32:00.891523Z","iopub.execute_input":"2022-11-05T08:32:00.892175Z","iopub.status.idle":"2022-11-05T08:32:00.897837Z","shell.execute_reply.started":"2022-11-05T08:32:00.892142Z","shell.execute_reply":"2022-11-05T08:32:00.896742Z"},"trusted":true}},{"cell_type":"code","source":"%%time\ndict_main_folds_info = {}\ndict_main_folds_info['name'] = 'Main Folds'\ndict_main_folds_info['Info'] = ['Train','Test1 Priv Like','Test2 Public Like', 'Test3 Priv Like HO', 'Test4 Publ Like HO', 'Playground' ]\ndict_main_folds_info['OOF Multiplicity'] = [0,1/2,0,1/2,0,1/6 ]\nm = (df_meta['Holdout'] == 0) & ( df_meta['Playground'] == 0 )\ndict_main_folds_info['OOF Indices Main'] = df_meta['Index'][m].values\ndict_main_folds_info['OOF Indices Additional'] = df_meta['Index'][df_meta['Holdout'] == 1].values\n\n#dict_main_folds_info['List Indices'] = list_folds_main    \n\nlist_folds_main = []\nmask_not_playground = (df_meta['Playground'] == 0) \nmask_holdout = (df_meta['Holdout']==1)\nc = 0\nfor day2exclude in [2,3,4]:\n    for donor2exclude in [32606,  31800]: # We will need to predict always MALE (not female) - like on LB. (# donor 13176 - female)\n        train_index = np.where(\\\n           mask_not_playground & (df_meta['day']  != day2exclude) & ( df_meta['donor']  != donor2exclude  ) )  [0]\n        test_index1_like_private_lb = np.where(\\\n           mask_not_playground & (df_meta['day']  == day2exclude) & (~mask_holdout) ) [0]\n        test_index1_like_private_lb_holdout = np.where(\\\n           mask_not_playground & (df_meta['day']  == day2exclude) & (mask_holdout) ) [0]\n        test_index2_like_public_lb = np.where(\\\n           mask_not_playground & (df_meta['day']  != day2exclude) &  (df_meta['donor']  == donor2exclude ) & (~mask_holdout) ) [0]\n        test_index2_like_public_lb_holdout = np.where(\\\n           mask_not_playground & (df_meta['day']  != day2exclude) &  (df_meta['donor']  == donor2exclude ) & mask_holdout ) [0]\n        \n        test_playground = np.where(mask_not_playground == 0)[0]\n        \n        list_folds_main.append( (train_index,  test_index1_like_private_lb , test_index2_like_public_lb, \n                              test_index1_like_private_lb_holdout, test_index2_like_public_lb_holdout, test_playground ) )\n    \n        str_fold_inf = 'Fold ' +str(c) + ': Train: excludes Day '+str(day2exclude) + ' and Donor ' + str( donor2exclude )\n        print(str_fold_inf, 'Sizes: train:',len(train_index), 'Test Like Priv'  ,len(test_index1_like_private_lb), \n              'Test Like Publ',   len(test_index2_like_public_lb),  \n              'Test Like Priv HoldOut',   len(test_index1_like_private_lb_holdout),  \n              'Test Like Publ HoldOut',   len(test_index2_like_public_lb_holdout),  \n              'Test Like Publ HoldOut',   len(test_index2_like_public_lb_holdout),  \n             ); c+=1\n    \ndict_main_folds_info['List Indices'] = list_folds_main    \nprint(len(list_folds_main))","metadata":{"_uuid":"3589e878-57ab-4b98-aa39-51424a654bb3","_cell_guid":"bddaea9c-6375-482e-a4d7-978a6ae1faf9","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:28:20.874653Z","iopub.execute_input":"2022-11-07T08:28:20.875373Z","iopub.status.idle":"2022-11-07T08:28:20.912257Z","shell.execute_reply.started":"2022-11-07T08:28:20.875335Z","shell.execute_reply":"2022-11-07T08:28:20.911168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Optional Rescaling of Targets","metadata":{"_uuid":"9b83bdd2-29d9-4358-b1bb-982a45d07779","_cell_guid":"7f7f64db-34dc-4f60-83dc-7e82329481c2","trusted":true}},{"cell_type":"code","source":"# Since metric is - correlation coefficient - any rescaling aY+b will not be change it , so one can do like that:\n# it is not clear is it optimal or not (see https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360253 )\n\nY = df_cite_train_y.values\n\nif rescale_Y_to_mean0_std1:\n    Y -= Y.mean(axis=1).reshape(-1, 1)\n    Y /= Y.std(axis=1).reshape(-1, 1)\n    print('Rescaling to mean 0 and std 1 has been done')","metadata":{"_uuid":"97b44d5e-03a8-4b2e-9c2c-92223b6318bd","_cell_guid":"117ab723-71af-45cb-affb-4eb0c8ce487b","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:28:24.010202Z","iopub.execute_input":"2022-11-07T08:28:24.010574Z","iopub.status.idle":"2022-11-07T08:28:24.048205Z","shell.execute_reply.started":"2022-11-07T08:28:24.010540Z","shell.execute_reply":"2022-11-07T08:28:24.046482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling Preparations","metadata":{"_uuid":"36ecbbd7-f9a3-4b97-843d-1b787d6bd3c9","_cell_guid":"a5757f02-9083-4caf-a0b8-fff7d956fb21","trusted":true}},{"cell_type":"code","source":"import gc\ndef correlation_score(y_true, y_pred):\n    \"\"\"Scores the predictions according to the competition rules. \n    \n    It is assumed that the predictions are not constant.\n    \n    Returns the average of each sample's Pearson correlation coefficient\"\"\"\n\n    y2 = y_pred.copy()\n    y2 -= y2.mean(axis=1).reshape(-1, 1);    y2 /= y2.std(axis=1).reshape(-1, 1)    \n    if rescale_Y_to_mean0_std1:\n        y1 = y_true # Already rescaled \n    else:\n        y1 = y_true.copy(); \n        y1 -= y1.mean(axis=1).reshape(-1, 1);    y1 /= y1.std(axis=1).reshape(-1, 1) \n        \n    c = (y1*y2).mean().mean()# Correlation for rescaled matrices is just matrix product and average \n    \n    # Memory control:\n    if not rescale_Y_to_mean0_std1:\n        del y1\n    del y2\n    gc.collect()\n    \n    return c\n\n    # Slow way:\n    \n    if type(y_true) == pd.DataFrame: y_true = y_true.values\n    if type(y_pred) == pd.DataFrame: y_pred = y_pred.values\n    corrsum = 0\n    for i in range(len(y_true)):\n        corrsum += np.corrcoef(y_true[i], y_pred[i])[1, 0]\n    return corrsum / len(y_true)","metadata":{"_uuid":"0e59893b-3efb-4c89-bb2a-5dee2cbb4239","_cell_guid":"eda4e4b5-494e-4f86-aac2-ef97b9884596","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:28:26.080825Z","iopub.execute_input":"2022-11-07T08:28:26.081181Z","iopub.status.idle":"2022-11-07T08:28:26.092442Z","shell.execute_reply.started":"2022-11-07T08:28:26.081150Z","shell.execute_reply":"2022-11-07T08:28:26.091397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_error\nfrom sklearn.metrics import r2_score","metadata":{"_uuid":"f89b4512-7d1c-406d-8f37-67c07d1bedfc","_cell_guid":"c19ef489-3fd3-4e6e-bd2a-e642789651de","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:28:28.273107Z","iopub.execute_input":"2022-11-07T08:28:28.273467Z","iopub.status.idle":"2022-11-07T08:28:28.278116Z","shell.execute_reply.started":"2022-11-07T08:28:28.273432Z","shell.execute_reply":"2022-11-07T08:28:28.277123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Setting the model used below","metadata":{"_uuid":"cf22a937-d737-44fe-af69-7f294b0875d4","_cell_guid":"7693d572-9fc4-47fe-9a02-aad5a1ceaefc","trusted":true}},{"cell_type":"code","source":"n_features_loc = n_features\n\nX = df_cite.iloc[:,:n_features_loc].values\nprint('X.shape:', X.shape,'Y.shape:', Y.shape)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T08:28:28.928623Z","iopub.execute_input":"2022-11-07T08:28:28.929372Z","iopub.status.idle":"2022-11-07T08:28:28.936516Z","shell.execute_reply.started":"2022-11-07T08:28:28.929336Z","shell.execute_reply":"2022-11-07T08:28:28.935356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"xgb_params = {'n_estimators': 500, \n              'max_depth': 6, \n              'learning_rate': 0.01881258066472956, \n              'gamma': 3.3414783586188253, \n              'base_score': 0.4001126687131699, \n              'tree_method': 'gpu_hist', \n              'predictor': 'gpu_predictor', \n              'max_leaves': 735}","metadata":{"execution":{"iopub.status.busy":"2022-11-07T08:28:30.839168Z","iopub.execute_input":"2022-11-07T08:28:30.839579Z","iopub.status.idle":"2022-11-07T08:28:30.844819Z","shell.execute_reply.started":"2022-11-07T08:28:30.839545Z","shell.execute_reply":"2022-11-07T08:28:30.843856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.linear_model import Ridge\nalpha4Ridge  = 1e4\nmodel = Ridge(alpha=alpha4Ridge)\n\n\n# import lightgbm as lgbm\nimport xgboost as xgb\nfrom sklearn.multioutput import MultiOutputClassifier, MultiOutputRegressor\n\nif 1:\n    model = MultiOutputRegressor(estimator=xgb.XGBRegressor(**xgb_params))","metadata":{"_uuid":"6c75eb10-c597-4be8-ae69-4b9b7758ba2f","_cell_guid":"138fb1ab-f680-4436-a025-7f352d111f08","execution":{"iopub.status.busy":"2022-11-07T08:28:34.102660Z","iopub.execute_input":"2022-11-07T08:28:34.103375Z","iopub.status.idle":"2022-11-07T08:28:34.301352Z","shell.execute_reply.started":"2022-11-07T08:28:34.103339Z","shell.execute_reply":"2022-11-07T08:28:34.300348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Calculate and Store various predictions","metadata":{"_uuid":"d805f54e-a486-41d0-a8e7-ca007a9be5b3","_cell_guid":"2b79a8e7-be1c-49e4-97f9-33af416ca554","trusted":true}},{"cell_type":"code","source":"%%time\ndict_folds_info = dict_main_folds_info\n\nflag_train_on_full_data4submit = True\nverbose = 1000\n\ndict_model_results = {}\n\nlist_folds_indices = dict_folds_info['List Indices'] \nt0 = time.time()\nY_oof_pred = np.zeros_like(Y)\nY_submission_way1_pred = np.zeros( (119651-70988 , Y.shape[1]) )  # Way 1 (Classical way) - retrain model on the whole data\nY_submission_way2_pred = np.zeros( (119651-70988 , Y.shape[1]) )  # Way 2 (Kaggle way) - average predictions of models on each fold\nfor fold, indices_tuple  in enumerate( list_folds_indices ):\n    \n    t00 = time.time()\n    train_index = indices_tuple[0]\n    model.fit(X[train_index], Y[train_index])\n    if verbose >= 100:\n        print('fold:',fold, 'Train time:%.2f'%(time.time()-t00 ), 'X_train.shape:',X[train_index].shape )\n        \n    # Calculate predictions on test and train folds \n    for i_loc in range(0,len( indices_tuple)  ):\n        indices_loc = indices_tuple[i_loc]\n        y_pred = model.predict( X[indices_loc ])\n        dict_model_results['Pred_fold'+str(fold)+'_part'+str(i_loc)] = y_pred\n        if i_loc > 0: # Skip train \n            Y_oof_pred[indices_loc] += y_pred * dict_main_folds_info['OOF Multiplicity'][i_loc] \n    # Prepare submission data in way 2 - \"Kaggle way\"\n    Y_submission_way2_pred += model.predict( X[70988:,: ])  # Way 2 (Kaggle way) - average predictions of models on each fold   \n\n# Main loop over folds - finished just do some final calculations:  \nY_submission_way2_pred /= (fold+1)      \nif flag_train_on_full_data4submit: \n    model.fit(X[:70988,:], Y) # Way 1 (Classical way) - retrain model on the whole data\n    Y_submission_way1_pred = model.predict( X[70988:,: ])   \n\nif verbose >= 10:\n    print('X.shape, Y.shape:',X.shape, Y.shape)\n    print('Finished. Folds:',fold, 'Time:%.2f'%(time.time()-t0), 'Y_oof_pred.shape:',Y_oof_pred.shape,\n         'Y_submission_way2_pred.shape:',Y_submission_way2_pred.shape )\n\n\ndict_model_results['Time'] = time.time() - t0\ndict_model_results['Y_pred'] = Y_oof_pred\ndict_model_results['OOF Indices Main'] = dict_folds_info['OOF Indices Main'] \ndict_model_results['OOF Indices Additional'] = dict_folds_info['OOF Indices Additional'] \ndict_model_results['Submit1'] = Y_submission_way1_pred\ndict_model_results['Submit2'] = Y_submission_way2_pred\ndict_model_results['Dict Folds Info' ]= dict_folds_info.copy() # \ndict_model_results['df_meta'] = df_meta # Just save, just in case","metadata":{"_uuid":"bc1f04c7-6ff6-448c-a440-59ebc8c2364a","_cell_guid":"92dc558c-a17b-4e29-a9f0-137bfcdc4b6e","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-07T08:28:42.996128Z","iopub.execute_input":"2022-11-07T08:28:42.996484Z","iopub.status.idle":"2022-11-07T08:29:26.805802Z","shell.execute_reply.started":"2022-11-07T08:28:42.996451Z","shell.execute_reply":"2022-11-07T08:29:26.804623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Main SCORING - caculate and save MAIN scores","metadata":{"_uuid":"78f3fc8c-93e8-4624-9f44-68f05d41a87c","_cell_guid":"641df5a1-559e-4329-a442-426ad974c7ed","trusted":true}},{"cell_type":"code","source":"%%time\nY_pred = dict_model_results['Y_pred']\nindices_oof = dict_model_results['OOF Indices Main']\nY_pred = Y_pred[  indices_oof , : ]\nY_true = Y[  indices_oof , : ]\n\ns_cor = correlation_score( Y_true , Y_pred  ) # It takes about 5 seconds ! \n\ns_r2 = r2_score( Y_true , Y_pred  )\ns_mse = mean_squared_error( Y_true , Y_pred  )\n\nprint(str(model)[:40], 'n_feat:', X.shape[1],)\nprint('OOF scores: ', s_cor, s_r2, s_mse, 'n_feat:', X.shape[1], 'Oof size:', Y_true.shape[0] )","metadata":{"_uuid":"11fcbb93-bb7d-4775-808c-a5da8763a6fe","_cell_guid":"75243105-be40-40f7-8f04-626031e7fc97","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-06T14:07:31.10326Z","iopub.execute_input":"2022-11-06T14:07:31.103675Z","iopub.status.idle":"2022-11-06T14:07:31.429428Z","shell.execute_reply.started":"2022-11-06T14:07:31.10364Z","shell.execute_reply":"2022-11-06T14:07:31.427876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"    100 features - better than 40 for LGBM \n    MultiOutputRegressor LightGBM: (Calculation time - 19 minutes)\n    OOF scores:  0.8843042 0.14324448260463332 0.21749122 n_feat: 100 Oof size: 57500\n\n    MultiOutputRegressor(estimator=LGBMRegre n_feat: 40\n    OOF scores:  0.88402647 0.14195155980790192 0.21797003 n_feat: 40 Oof size: 57500\n\n    MultiOutputRegressor LightGBM: (Calculation time - 19 minutes)\n    OOF scores:  0.8843042 0.14324448260463332 0.21749122 n_feat: 100 Oof size: 57500\n\n    Ridge(alpha=10000.0) n_feat: 100\n    OOF scores:  0.8849721 0.1520688839180771 0.21653545 n_feat: 100 Oof size: 57500\n\n    Ridge 50: \n    OOF scores:  0.88314193 0.1486980690002863 0.21981749\n    MeanByFolds Test1: 0.882145 Test2: 0.883643\n    \n    Ridge 100:\n    OOF scores:  0.8849221 0.15140698792122334 0.21665995\n    MeanByFolds Test1: 0.883907\t0.136736\t0.218520\t0.885304\t0.131960\t0.215981\t0.885119\t0.136830\t0.216512\t0.886379\t0.135263\t0.214223\t0.887717\t0.158759\t0.211723\n    \n    Ridge 150\n    OOF scores:  0.88490844 0.15052967052133392 0.21670045\n    MeanByFolds Test1:  0.883843\t0.135412\t0.218653\t0.885232\t0.130625\t0.216116\t0.885073\t0.135405\t0.216610\t0.886348\t0.134102\t0.214291\t0.887650\t0.157636\t0.211866\n    \n    Ridge 200\n    OOF scores:  0.88479453 0.1494652531252453 0.21691151\n    MeanByFolds Test1: 0.883674\t0.133911\t0.218972\t0.885063\t0.129137\t0.216421\t0.884907\t0.133874\t0.216923\t0.886183\t0.132540\t0.214590\t0.887495\t0.156314\t0.212149","metadata":{"_uuid":"e513a1e3-2034-472b-8065-932dc87fa23f","_cell_guid":"e3ea991f-5a6b-4e71-9c71-665a7b4590b7","trusted":true}},{"cell_type":"code","source":"","metadata":{"_uuid":"e118e4fb-e518-4bb5-bf27-d4b172e50e4a","_cell_guid":"e6336559-e353-4fe6-8746-baaa1df0e6b5","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Additional info on scores - foldwise","metadata":{"_uuid":"4e08dde7-b0ff-49d5-b32a-5cfdebdd5320","_cell_guid":"4083aaf8-b9b3-466a-9758-5816a714b159","trusted":true}},{"cell_type":"code","source":"%%time\nverbose = 1000\nlist_folds_indices =  dict_model_results['Dict Folds Info']['List Indices']\nlist_folds_indices_info = dict_model_results['Dict Folds Info']['Info'] \ndf_fold_score_stat = pd.DataFrame();df_fold_score_stat.index.name = 'Fold'\nt0 = time.time()\nfor fold, indices_tuple  in enumerate( list_folds_indices ):\n    # Calculate metrics on test and train folds \n    list_scores = []; list_scores_r2 = []\n    for i_loc, indices_loc in enumerate( indices_tuple) :\n        y_pred = dict_model_results['Pred_fold'+str(fold)+'_part'+str(i_loc)]\n        name_loc = list_folds_indices_info[i_loc]\n        s = correlation_score(Y[indices_loc],  y_pred    )\n        df_fold_score_stat.loc[fold,'Corr '+name_loc] = s\n        s = r2_score( Y[indices_loc] , y_pred  )\n        df_fold_score_stat.loc[fold,'r2 '+name_loc] = s\n        s = mean_squared_error( Y[indices_loc] , y_pred  )\n        df_fold_score_stat.loc[fold,'mse '+name_loc] = s\n    if verbose >= 1000:\n        print('Fold:', fold, 'Shapes of train:', train_index.shape, 'Tests: ',[t.shape for t in indices_tuple[1:] ] )        \ndisplay( df_fold_score_stat )\ndisplay(df_fold_score_stat.describe(percentiles=[] ).iloc[1:,:])","metadata":{"_uuid":"909dcb65-0c59-475f-bc3a-b6cdcd4aaab4","_cell_guid":"bfd1aaae-bdd8-4d05-a031-184493215821","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-06T13:56:54.335922Z","iopub.execute_input":"2022-11-06T13:56:54.336257Z","iopub.status.idle":"2022-11-06T13:57:01.101666Z","shell.execute_reply.started":"2022-11-06T13:56:54.336229Z","shell.execute_reply":"2022-11-06T13:57:01.100363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization of Scores TargetWISELY - Optional look on scores\n\nAs it is expected high MSE is highly correlated to just high Mean.\n\nProbably those targets are contribute most to the correlation metric.","metadata":{"_uuid":"efed7dc6-ff06-4901-bdad-7890d97164d7","_cell_guid":"86405e0b-95fd-4dfc-85a2-64edcede60ab","execution":{"iopub.status.busy":"2022-11-05T09:07:34.033666Z","iopub.execute_input":"2022-11-05T09:07:34.034094Z","iopub.status.idle":"2022-11-05T09:07:34.044978Z","shell.execute_reply.started":"2022-11-05T09:07:34.034057Z","shell.execute_reply":"2022-11-05T09:07:34.043899Z"},"trusted":true}},{"cell_type":"code","source":"pd.set_option('display.max_columns', None)  # or 1000\npd.set_option('display.max_rows', None)  # or 1000\npd.set_option('display.max_colwidth', None)  # or 199","metadata":{"_uuid":"20e4ff28-683c-4a9b-b03e-6a9b815442ce","_cell_guid":"fe2e95aa-7eb1-49ec-86f3-3ae68bffac6d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-06T13:57:01.103568Z","iopub.execute_input":"2022-11-06T13:57:01.104003Z","iopub.status.idle":"2022-11-06T13:57:01.109665Z","shell.execute_reply.started":"2022-11-06T13:57:01.103966Z","shell.execute_reply":"2022-11-06T13:57:01.10849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ndf_targetwise = pd.DataFrame()\n\nY_pred = dict_model_results['Y_pred']\nindices_oof = dict_model_results['OOF Indices Main']\nY_pred = Y_pred[  indices_oof , : ]\nY_true = Y[  indices_oof , : ]\nfor i,col in enumerate(df_cite_train_y.columns):\n    IX = col#.columns[i_loc]\n    df_targetwise.loc[IX, 'r2'] = r2_score( Y_true[:,i] , Y_pred[:,i]  )\n    df_targetwise.loc[IX, 'mse'] = mean_squared_error( Y_true[:,i] , Y_pred[:,i]  )    \n    df_targetwise.loc[IX, 'corr'] = np.corrcoef( Y_true[:,i] , Y_pred[:,i]  )[0,1]    \n    df_targetwise.loc[IX, 'std'] = df_cite_train_y[col].std()\n    df_targetwise.loc[IX, 'absMean'] = df_cite_train_y[col].abs().mean()\n    df_targetwise.loc[IX, 'mean'] = df_cite_train_y[col].mean()\n    df_targetwise.loc[IX, 'max'] = df_cite_train_y[col].max()\n    \nprint('Pearson correlation between scores (and other characteristics):')\ndisplay(df_targetwise.corr()   )\nprint('Spearman correlation between scores (and other characteristics):')\ndisplay(df_targetwise.corr(method = 'spearman')   )\n\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['r2'].plot(); plt.legend(); plt.grid();\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['mse'].plot(); plt.legend(); plt.grid();\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['corr'].plot(); plt.legend(); plt.grid();\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['absMean'].plot(); plt.legend(); plt.grid();\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['std'].plot(); plt.legend(); plt.grid();\n\nd1 = df_targetwise.sort_values('r2'); d1.index.name = 'r2 sorted';  d1 = d1.reset_index()\nd2 = df_targetwise.sort_values('mse',ascending = False); d2.index.name = 'mse sorted';  d2 = d2.reset_index()\nd3 = df_targetwise.sort_values('corr'); d3.index.name = 'corr sorted';  d3 = d3.reset_index()\nd = pd.concat([d1,d2,d3],axis = 1)\nprint('\\n Worst:')\ndisplay(d.head(10))\nprint('\\n Best:')\ndisplay(d.tail(10))","metadata":{"_uuid":"21700999-98c8-42b6-8448-aa5037e51d15","_cell_guid":"21362122-af7a-4872-a2bf-0ae98248af08","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-06T13:57:01.111345Z","iopub.execute_input":"2022-11-06T13:57:01.112085Z","iopub.status.idle":"2022-11-06T13:57:04.23293Z","shell.execute_reply.started":"2022-11-06T13:57:01.112047Z","shell.execute_reply":"2022-11-06T13:57:04.231697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization of the main scores - correlations","metadata":{"_uuid":"31bed355-fa2d-4159-ae2d-5690b3da64c7","_cell_guid":"cb95733a-ccce-43fb-b4e6-f98efcd98c66","trusted":true}},{"cell_type":"code","source":"%%time\ncol = 'OOF Indices Main'\nprint('Predictions on:', col)\nY_pred = dict_model_results['Y_pred']\nindices_oof = dict_model_results[col]\nY_pred = Y_pred[  indices_oof , : ]\nY_true = Y[  indices_oof , : ]\nprint('Y_pred.shape: ', Y_pred.shape, 'Y_true.shape: ', Y_true.shape,  )\nlist_corr = []\nfor i in range(Y_pred.shape[0]):\n    list_corr.append(np.corrcoef(Y_true[i,:], Y_pred[i,:] )[0,1] )\n    \n\nfig = plt.figure( figsize = (10,5))\nplt.title('Main score - correlation \\n'+str(col))\nplt.hist(list_corr, bins = 100)\nplt.grid()\nplt.show()\ndisplay(pd.Series(list_corr).describe() )","metadata":{"_uuid":"1b23a96e-e7aa-414f-b652-264dcc8ad132","_cell_guid":"cafa68bd-7f64-4cb7-ba78-6f1c7bf0a41a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-06T13:57:04.234468Z","iopub.execute_input":"2022-11-06T13:57:04.234826Z","iopub.status.idle":"2022-11-06T13:57:10.724893Z","shell.execute_reply.started":"2022-11-06T13:57:04.234793Z","shell.execute_reply":"2022-11-06T13:57:10.723533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ncol = 'OOF Indices Additional'\nprint('Predictions on:', col)\nY_pred = dict_model_results['Y_pred']\nindices_oof = dict_model_results[col]\nY_pred = Y_pred[  indices_oof , : ]\nY_true = Y[  indices_oof , : ]\nprint('Y_pred.shape: ', Y_pred.shape, 'Y_true.shape: ', Y_true.shape,  )\nlist_corr = []\nfor i in range(Y_pred.shape[0]):\n    list_corr.append(np.corrcoef(Y_true[i,:], Y_pred[i,:] )[0,1] )\n    \n\nfig = plt.figure( figsize = (10,5))\nplt.title('Main score - correlation \\n'+str(col))\nplt.hist(list_corr, bins = 100)\nplt.grid()\nplt.show()\ndisplay(pd.Series(list_corr).describe() )","metadata":{"_uuid":"028fb6ab-aea0-4e64-95b9-36af75927beb","_cell_guid":"7f8e634a-cdf3-470d-b783-15cb45325b3b","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-06T13:57:10.726919Z","iopub.execute_input":"2022-11-06T13:57:10.727435Z","iopub.status.idle":"2022-11-06T13:57:11.786639Z","shell.execute_reply.started":"2022-11-06T13:57:10.727382Z","shell.execute_reply":"2022-11-06T13:57:11.785226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preparation of the submission file","metadata":{"_uuid":"752f9889-c538-47ad-bb16-853de39e16bc","_cell_guid":"50a47445-c6e4-4686-bff1-53a874ded929","jupyter":{"outputs_hidden":false},"trusted":true}},{"cell_type":"code","source":"%%time\n\nif flag_prepare_submission:\n    mode_subm = 'USE: sskknts MSCI CITEseq Keras Quickstart + Dropout' # 'put_zeros_to_multiome_part'\n\n    if mode_subm == 'USE: sskknts MSCI CITEseq Keras Quickstart + Dropout' :\n        df_submission_full = pd.read_csv('/kaggle/input/msci-citeseq-keras-quickstart-dropout/submission.csv',\n                                 index_col='row_id', squeeze=True)\n        df_submission_full = df_submission_full.to_frame()\n    elif mode_subm == 'put_zeros_to_multiome_part':\n        df_submission_full = pd.DataFrame(index = range(65_744_180), columns = ['target'], data = np.zeros(65_744_180) )\n        df_submission_full.index.name = 'row_id'\n\n    display(df_submission_full.info() )\n    print()\n    display(df_submission_full)","metadata":{"_uuid":"9861b2e6-dc33-4c62-9965-81dc2eaf3d93","_cell_guid":"e3d08b1e-d469-439b-ba7a-5299bc95ddab","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-06T13:57:11.788411Z","iopub.execute_input":"2022-11-06T13:57:11.789307Z","iopub.status.idle":"2022-11-06T13:57:11.798447Z","shell.execute_reply.started":"2022-11-06T13:57:11.789257Z","shell.execute_reply":"2022-11-06T13:57:11.797077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif flag_prepare_submission:\n    Y_pred4submit = dict_model_results['Submit1']\n    print(Y_pred4submit.shape)","metadata":{"_uuid":"1230587a-49e6-45f3-9883-253b244ef753","_cell_guid":"797b5aed-7533-43c4-a594-cd4099676bd9","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-06T13:57:11.800438Z","iopub.execute_input":"2022-11-06T13:57:11.800942Z","iopub.status.idle":"2022-11-06T13:57:11.811323Z","shell.execute_reply.started":"2022-11-06T13:57:11.800891Z","shell.execute_reply":"2022-11-06T13:57:11.809919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif flag_prepare_submission:\n    # We can prepare CITE-seq part as: Y_pred4submit.values.ravel() -- see notebook: https://www.kaggle.com/code/alexandervc/mmscel-analysis-of-evaluation-and-submission\n    df_submission_full['target'].iloc[:6_812_820] = Y_pred4submit.ravel()\n    display(df_submission_full)","metadata":{"_uuid":"f175b5c0-9177-492b-ab43-d160e325f4fb","_cell_guid":"98b5d7cf-5741-4f0e-ab4f-aa2377679b55","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-06T13:57:11.813445Z","iopub.execute_input":"2022-11-06T13:57:11.813954Z","iopub.status.idle":"2022-11-06T13:57:11.826732Z","shell.execute_reply.started":"2022-11-06T13:57:11.813906Z","shell.execute_reply":"2022-11-06T13:57:11.825888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Wall time: 1min 46s\nif flag_prepare_submission:\n    submit_filename_postfix = 'CVadvV5_Ridge100'\n    df_submission_full.to_csv('submission_cite_seq_'+submit_filename_postfix +'.csv')","metadata":{"_uuid":"72332661-0de3-4fc2-a769-716579e47c35","_cell_guid":"d833b02f-6cb9-49fa-ac3a-7ab7eb61f91d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-06T13:57:11.828936Z","iopub.execute_input":"2022-11-06T13:57:11.829403Z","iopub.status.idle":"2022-11-06T13:57:11.841971Z","shell.execute_reply.started":"2022-11-06T13:57:11.829358Z","shell.execute_reply":"2022-11-06T13:57:11.840917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"b372da4f-a87d-4b71-a071-414e63564d37","_cell_guid":"de9b5839-be50-4f31-8fae-21d9a4b8cd7e","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"5d8d0a31-a09c-490d-8387-464a59f1dd23","_cell_guid":"5ca85d26-228a-4f42-910c-a380f4827022","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )","metadata":{"_uuid":"f55e2443-fd24-4fc6-9238-fec592d87d02","_cell_guid":"d3eca9d9-7cdc-4323-beec-8e7d1c64fc84","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-06T13:57:11.84348Z","iopub.execute_input":"2022-11-06T13:57:11.843812Z","iopub.status.idle":"2022-11-06T13:57:11.854901Z","shell.execute_reply.started":"2022-11-06T13:57:11.843782Z","shell.execute_reply":"2022-11-06T13:57:11.85351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"c1bbfdc5-d83f-44e3-8a64-6d4b2f4ff50b","_cell_guid":"a27b9c08-7e28-47e3-a2de-5a4eed9df4a2","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]}]}