{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nSome advanced modeling planned to be here... \nSeveral models, blending, tuning, etc...\n\nIdeas planned to be implemented:\n\n0) Select some part of dataset for \"playground\", it can be used for preliminary feature selection and/or preparing stacking and/or quick experiments and/or varying that part we will create more diversity for models, which can be blended later. \n\n1) Various models, hyper-param tuning\n\n2) Blending possibly targetwisely, possibly with solutions defined only for some targets\n\n3) Optimization of hyper-params for blend, not for individual performance\n\n4) ...\n\nBased on previous scripts:\nhttps://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold\n\n\n### About CV scheme\n\nPublic and private test sets are quite different. (Private - contains new DAY (and donor), while public ONLY new donor).\nSo one should be careful with the validation schemes.\nSome proposals are described here: \n\nNotebook: https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\nTopic: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\n\nWe will be based on them. \n\n\n    Depricated from Version 45: ### Option to create \"Playground\" \n    Some part of the data can be optionally distinguished. And used for preliminary feature selection and/or preparing stacking and/or quick experiments and/or varying that part we will create more diversity for models, which can be blended later. \n\n\n### Versions:\n\n#### 50 - 5*    Some modeling \n\n    66 Quick save\n    \n    65 New data: fn = '/kaggle/input/dataframe-with-removed-features/df_r_filter.csv' # features 1987 \n    65 Ridge 0\tRidge(alpha=10000.0)\t0.887469\t0.157243\t0.212128\t1987.0\t140.0\t26.259\n\n    64 Quick save \n    \n    63 KNN 100 features, trainsize = 0.3\n    62 KNN 100 features \n    \n    61 Quick save\n    \n    60 same KNN, but turn off all the optinal inference \n    59 KNN regressor on just two features - 1h 30 min  and r2 is a bit below zero - poor quality \n\n    58 - Quick Save\n    \n    57 MLP2 on 651 features is almost the same as top-Ridge 2719, but still a bit worse\n    start MLP2 with correct features: /kaggle/input/features-651-multimodal-singlecell-integration/features_651_CITEseq_better4MLP.csv\n    MLPRegressor(activation='logistic', early_stop...\t0.890939\t0.161648\t0.205977\t651.0\t140.0\t1040.779\tmean\t0.887327\t0.137725\t0.212882\t0.888985\t0.134247\t0.209749\t0.910277\t0.211343\t0.170746\n\n    55-56 - used incorrect feature set - one needs 651 features: /kaggle/input/features-651-multimodal-singlecell-integration/features_651_CITEseq_better4MLP.csv\n    \n    56 MLP2 - better than MLP1, but worse the best Ridge \n    56 MLP 2 - Updated From https://www.kaggle.com/code/tttzof351/mmscel-crossvalidation-schemes-193f49  \n    0.889691\t0.158286\t0.208065\t2274.0\t140.0\t3409.6\n    0.889691\t0.158286\t0.208065\t2274.0\t140.0\t3409.6\tmean\t0.885469\t0.133817\t0.216063\t0.886978\t0.130619\t0.213255\t0.908189\t0.206095\t0.174642\n    \n    55 MLP1 - slightly worse than Ridge - as it was before. \n    55 MLP on 2273 features around 0.808 LB: https://www.kaggle.com/code/visualcomments/mmscel-crossvalidation-schemes\n    MLPRegressor(activation='logistic', alpha=1e-0...\t\n    0.889154\t0.157723\t0.209004\t2274.0\t140.0\t3567.768\tmean\t0.884736\t0.132232\t0.217273\t0.885922\t0.127582\t0.215027\t0.907195\t0.202555\t0.176431\n\n    \n    51 Ridge on 2719 features Corr OOF: 0.890993\t0.162407\t0.205793\t2719.0\t140.0\t76.279\tmean\t0.889617 0.145192\t0.208354\n    2719 features: https://www.kaggle.com/datasets/alexandervc/features-multimodal-singlecell-integration\n    \n    50 Ridge on 3790 features Corr OOF: 0.89025\t0.154837\t0.2072\t3789.0\t140.0\t94.987\tmean\t0.888543\t0.134474\t0.2104\t\n    44 Ridge on 100 features Corr OOF: 0.8849721 0.1520688839180771 0.21653545 n_feat: 100 Oof size: 57500\n\n    53 Ridge change Alfa = 1e5,  on 3790 features 0.889385\t0.160949\t0.208602\t3789.0\t140.0\t129.022\tmean\t0.88843\n    54 Ridge change Alfa = 1e5, on 2719 features 0.889274\t0.163049\t0.20876\t2719.0\t140.0\t53.764\tmean\t0.888446\t0.149614\t0.210256\t0.889628\t0.144439\t0.208202\t0.900737\t0.201367\t0.188176\n\n#### 46-49 \n    \n    Updating, debugging, polishing the notebook. \n    \n\n### 45 - Start of major update of the pipeline \n\n    Quite simplify and clarify the scheme.\n    Introduce OOF predictions and mainly focus on them.  Prepare saving results for future blend\n    \n    Create option to train models on part of the train data in order to gain diversity from averaging several such models.  \n    \n\n#### 40 - 4*  plan to remake the pipeline - thus return to Ridge for fast experiments\n    \n    0.805Ridge100, Alpha = 1e4, \"\"\n    \n\n### 39:  Conclusions from boosting performance with different params:\n\n    100 features seems to be the best (among PCA500) for vaious boostings params as well as for Ridge.\n    \n    Boosting params which were good on ONLY one target CD31 - seems to perform well here also. \n    Seems to be better than params from vuonglam \n    \n    Default params - seems to perform quite well - better than many other params sets, but still 0.001 worse than best CD31 params. Better than Ridge ! \n\n    \n#### 14 - 38 testing more boosting params testing\n\n\n\n    # From n_trials =  200 trials for CD31 https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109850852&cellId=28 \n    Wall time: 3h 12min 31s\n    OOF scores:  0.8867753 0.15452355703071607 0.21307571 n_feat: 100 Oof size: 57500\n\n    From 100 trials on CD31 (Version 24 here) :  https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109847986&cellId=28\n    # 4h 14min 8s\n    OOF scores:  0.88677084 0.15495630150980477 0.21306391 n_feat: 100 Oof size: 57500\n\n    From 1600 trials - also not bad (Version 25 here ): \n    OOF scores:  0.8863187 0.15582536340176648 0.21396601 n_feat: 100 Oof size: 57500\n    \n    Default LGBM - beats the Ridge !  Version 32 \n    OOF scores:  0.88551927 0.15129287544868444 0.2153129 n_feat: 100 Oof size: 57500\n\n#### 8, 9, 10 , 11 , 12, 13\n\n    Lightgbm model from \n    https://www.kaggle.com/code/yaroshenko/mmscel-crossvalidation-schemes?scriptVersionId=109641833&cellId=50\n    \n    On 100 features it is slightly worse that Ridge 100, 150, 200 ..  (Version 8 - 20 minutes) \n    OOF scores:  0.8843042 0.14324448260463332 0.21749122 n_feat: 100 Oof size: 57500\n    On 40 Features LGBM is sligly worse than on 100: (Version 9 - 11 minutes to run)\n    OOF scores:  0.88402647 0.14195155980790192 0.21797003 n_feat: 40 Oof size: 57500\n    On 200 Features LGBM is sligly worse than on 100:\n    OOF scores:  0.88414896 0.14308675934626502 0.21776518 n_feat: 200 Oof size: 57500 (36 minutes - version 10 )\n    On 150 Features LGBM is sligly worse than on 100, better than on 200 :\n    OOF scores:  0.88421464 0.14314857962629174 0.21765445 n_feat: 150 Oof size: 57500 (25 minutes - version 12 )\n    On 500 Features LGBM is sligly worse  100, 200, and even 40\n    OOF scores:  0.88396496 0.1428889367211743 0.21808921 n_feat: 500 Oof size: 57500 (90 minute - version 11 )\n    \n#### 5,6, 7\n    Added: \n    Submission preparation - multiome part - taken from public (Normally that part takes about 3 minutes, however currently Kaggle does not work correctly and that part never finishes)\n    Param: flag_prepare_submission = True/False - to switch off (for speed) that part\n\n    Problem: Kaggle system works not fine and some problems occur at loading the file with multiome part. \n\n    More updates to speed-up correlation score calculation\n   \n    # Ridge 100 features with Alpha 10 000 seems to be the best Ridge \n    # See: \n    Ridge(alpha=10000.0) n_feat: 100\n    OOF scores:  0.8849721 0.1520688839180771 0.21653545 n_feat: 100 Oof size: 57500\n\n    Ridge 50: \n    OOF scores:  0.88314193 0.1486980690002863 0.21981749\n    MeanByFolds Test1: 0.882145 Test2: 0.883643\n\n    Ridge 100:\n    OOF scores:  0.8849221 0.15140698792122334 0.21665995\n    MeanByFolds Test1: 0.883907 0.136736    0.218520    0.885304    0.131960    0.215981    0.885119    0.136830    0.216512    0.886379    0.135263    0.214223    0.887717    0.158759    0.211723\n\n    Ridge 150\n    OOF scores:  0.88490844 0.15052967052133392 0.21670045\n    MeanByFolds Test1:  0.883843    0.135412    0.218653    0.885232    0.130625    0.216116    0.885073    0.135405    0.216610    0.886348    0.134102    0.214291    0.887650    0.157636    0.211866\n\n    Ridge 200\n    OOF scores:  0.88479453 0.1494652531252453 0.21691151\n    MeanByFolds Test1: 0.883674 0.133911    0.218972    0.885063    0.129137    0.216421    0.884907    0.133874    0.216923    0.886183    0.132540    0.214590    0.887495    0.156314    0.212149   \n   \n#### 2 , 3 , 4\n\n    Added:\n    Ridge model with 100 (on CV it seems better than 50 and 150)\n    Foldwise analysis of scores.\n    Speed-up for correlation metric. \n    \n    Ridge 100 features - Wall time: 3.58 s - on all training and predictions (very fast). \n    The longer part - calculation of correlation metric. \n\n#### 1 - draft\n\n    Creation of Playground and \"Mainground\" (with respect to statification by the day donor cell-type).\n    Modeling:\n    Trivial model - Ridge over 1 feature - so it is basically the same as MEAN prediction. \n    Calculation of the model prediction according to CV-scheme described above.\n    In particular  \"OOF\" (out-of-fold) predictions - that is not straigthford since CV schem is not standard one. \n    Scoring:\n    Main scores (correlation) calculation.\n    Analysis: \n    Histograms of scores  \n    Analysis \"targetwisely\" i.e. for each particular target - look on r2, mse, corr scores.","metadata":{"_uuid":"e85ad2dd-4548-45c2-bd22-91d2a8e4c38a","_cell_guid":"6f7edf4e-a866-408c-9f90-083b4e2ba675","trusted":true}},{"cell_type":"markdown","source":"# Key params","metadata":{"_uuid":"51600441-6c40-4f20-a165-f4380cbb5626","_cell_guid":"cc51e09c-c472-4c65-b21d-230e27930807","trusted":true}},{"cell_type":"code","source":"rescale_Y_to_mean0_std1 = True\n\nflag_prepare_submission = False  # \n\nflag_save_foldwise_predictions = False # In basic scenario we will not use them, but just in case we can save \n\nn_features = 2200 # use first N features - to make quick experiments\n\n\nconstant_number_of_original_targets = 140 # For CITE-seq part \n#selected_targets = 'ALL'# 'CD31'","metadata":{"_uuid":"6b246588-79af-4af7-aeaa-28c40ef039e5","_cell_guid":"90284018-0e3f-4e49-941a-d77ce2d84e9d","execution":{"iopub.status.busy":"2022-11-15T10:55:29.529207Z","iopub.execute_input":"2022-11-15T10:55:29.529702Z","iopub.status.idle":"2022-11-15T10:55:29.535687Z","shell.execute_reply.started":"2022-11-15T10:55:29.529664Z","shell.execute_reply":"2022-11-15T10:55:29.534513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Install/import modules, load technical data","metadata":{"_uuid":"a42ce9d6-ee5d-4ec0-9c2a-31710ef3b8b5","_cell_guid":"f6b5ae5b-a81f-47da-8027-c514cc207aac","trusted":true}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"104f0242-958b-4ef1-af40-8bec2dc9bea0","_cell_guid":"98e6bc60-36d3-4d45-855a-11748f831a0d","execution":{"iopub.status.busy":"2022-11-15T10:55:35.644229Z","iopub.execute_input":"2022-11-15T10:55:35.644719Z","iopub.status.idle":"2022-11-15T10:55:35.716582Z","shell.execute_reply.started":"2022-11-15T10:55:35.644683Z","shell.execute_reply":"2022-11-15T10:55:35.715544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndisplay(df_cell)","metadata":{"_uuid":"01798d74-2f72-4000-b474-cdb4d74e5f3e","_cell_guid":"d3410584-a62c-4543-8fac-a4fe8fe946f6","execution":{"iopub.status.busy":"2022-11-15T10:55:37.502198Z","iopub.execute_input":"2022-11-15T10:55:37.503372Z","iopub.status.idle":"2022-11-15T10:56:00.254986Z","shell.execute_reply.started":"2022-11-15T10:55:37.503328Z","shell.execute_reply":"2022-11-15T10:56:00.253698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"_uuid":"0bdd40f9-eb9c-4eaa-9e08-5e91c0ed728b","_cell_guid":"bfdcb1cc-f7ee-4688-a601-d8661a95ec53","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-15T10:56:36.972604Z","iopub.execute_input":"2022-11-15T10:56:36.973239Z","iopub.status.idle":"2022-11-15T10:56:59.396349Z","shell.execute_reply.started":"2022-11-15T10:56:36.973181Z","shell.execute_reply":"2022-11-15T10:56:59.395096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load prepared Features for CITE-seq part of task","metadata":{"_uuid":"32413e3a-7b7f-4a2e-b028-4410b034ea14","_cell_guid":"d47e3bef-4916-43c5-921c-512ca1df8b9d","trusted":true}},{"cell_type":"code","source":"%%time\n\nprint('Load prepared features for CITE-seq')\n# These files contain both train and test parts .\n# For CITEseq part - first 70988 elements - train, and later 48663 - test. Overall 119651 samples.\n#fn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_TruncatedSVD200_niter7_rs42.csv'\n#fn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA500.csv'\n#fn = '/kaggle/input/3790-features-in-df/3790_features_in_df.csv'\n## #fn = '/kaggle/input/mmscel-crossvalidation-schemes-features-select/source_df.csv'\n#fn = '/kaggle/input/features-multimodal-singlecell-integration/features_2719_CITEseq_source_df.csv'\nfn = '/kaggle/input/features-2274-multimodal-singlecell-integration/features_2274_CITEseq_good4MLP.csv'\n#fn = '/kaggle/input/features-651-multimodal-singlecell-integration/features_651_CITEseq_better4MLP.csv'\n# fn = '/kaggle/input/dataframe-with-removed-features/df_r_filter.csv' # features 1987 \ndf_cite = pd.read_csv(fn,index_col = 0)\nif fn == '/kaggle/input/3790-features-in-df/3790_features_in_df.csv':\n    df_cite = df_cite.set_index('cell_id')\ndisplay(df_cite.head(5))\ndisplay(df_cite.tail(5))\n","metadata":{"_uuid":"14d13305-d62b-4822-b783-09f607eb2cb6","_cell_guid":"0495e7d2-ea36-469e-819a-26ab27a41c58","execution":{"iopub.status.busy":"2022-11-15T10:43:28.000735Z","iopub.execute_input":"2022-11-15T10:43:28.001256Z","iopub.status.idle":"2022-11-15T10:45:05.428215Z","shell.execute_reply.started":"2022-11-15T10:43:28.001213Z","shell.execute_reply":"2022-11-15T10:45:05.427045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_cite.mean(axis = 0).head(3))\nprint('Pay attention - PCA features are ordered by magnitude:')\nprint(df_cite.std(axis=0).head(5))","metadata":{"_uuid":"afe9baa3-48d2-4d56-b027-159880e77522","_cell_guid":"bc595bf0-e292-4bfa-bd19-821051bf8579","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-15T10:46:50.468440Z","iopub.execute_input":"2022-11-15T10:46:50.468958Z","iopub.status.idle":"2022-11-15T10:46:55.276425Z","shell.execute_reply.started":"2022-11-15T10:46:50.468921Z","shell.execute_reply":"2022-11-15T10:46:55.275053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Targets for CITE-seq","metadata":{"_uuid":"11f0481b-2dbc-4835-b7c2-98755ca1f63b","_cell_guid":"2e5213c8-69c3-4832-a5ce-f2471ccabb52","trusted":true}},{"cell_type":"code","source":"%%time\n#if 1:\nprint('Load CITE-seq targets and ')\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndisplay(df_cite_train_y)","metadata":{"_uuid":"93f0dd7a-8a6a-4224-bc07-ff982ab834f7","_cell_guid":"915efa3d-6965-4b7f-a759-d1e3848d165d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-15T10:46:56.587271Z","iopub.execute_input":"2022-11-15T10:46:56.587784Z","iopub.status.idle":"2022-11-15T10:46:57.474648Z","shell.execute_reply.started":"2022-11-15T10:46:56.587745Z","shell.execute_reply":"2022-11-15T10:46:57.473236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load and start prepare some metadata (cut only CITE-seq train part)","metadata":{"_uuid":"60de91b6-5173-4461-9dd6-7094e3173f1a","_cell_guid":"3e9c27ac-e90e-44a4-ab94-dda2b95239f9","trusted":true}},{"cell_type":"code","source":"%%time\n# fn2 = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/_citeseq_meta_all_text_also.csv'\n# df_meta_full = pd.read_csv(fn2,index_col = 0)\n# display(df_meta_full)\n#if 1:\ndf_meta = pd.DataFrame(index = df_cite_train_y.index) \ndf_meta = df_meta.join(df_cell.set_index('cell_id') )\ndisplay(df_meta)","metadata":{"_uuid":"a3de63b6-a27d-4924-9764-6ef3fcf03c1d","_cell_guid":"f5fa99ae-4511-45ec-907c-b872e8744acf","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-15T10:47:07.093817Z","iopub.execute_input":"2022-11-15T10:47:07.095228Z","iopub.status.idle":"2022-11-15T10:47:07.273601Z","shell.execute_reply.started":"2022-11-15T10:47:07.095172Z","shell.execute_reply":"2022-11-15T10:47:07.272181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list_all_targets = ['CD86', 'CD274', 'CD270', 'CD155', 'CD112', 'CD47', 'CD48', 'CD40',\n       'CD154', 'CD52', 'CD3', 'CD8', 'CD56', 'CD19', 'CD33', 'CD11c',\n       'HLA-A-B-C', 'CD45RA', 'CD123', 'CD7', 'CD105', 'CD49f', 'CD194', 'CD4',\n       'CD44', 'CD14', 'CD16', 'CD25', 'CD45RO', 'CD279', 'TIGIT',\n       'Mouse-IgG1', 'Mouse-IgG2a', 'Mouse-IgG2b', 'Rat-IgG2b', 'CD20',\n       'CD335', 'CD31', 'Podoplanin', 'CD146', 'IgM', 'CD5', 'CD195', 'CD32',\n       'CD196', 'CD185', 'CD103', 'CD69', 'CD62L', 'CD161', 'CD152', 'CD223',\n       'KLRG1', 'CD27', 'CD107a', 'CD95', 'CD134', 'HLA-DR', 'CD1c', 'CD11b',\n       'CD64', 'CD141', 'CD1d', 'CD314', 'CD35', 'CD57', 'CD272', 'CD278',\n       'CD58', 'CD39', 'CX3CR1', 'CD24', 'CD21', 'CD11a', 'CD79b', 'CD244',\n       'CD169', 'integrinB7', 'CD268', 'CD42b', 'CD54', 'CD62P', 'CD119',\n       'TCR', 'Rat-IgG1', 'Rat-IgG2a', 'CD192', 'CD122', 'FceRIa', 'CD41',\n       'CD137', 'CD163', 'CD83', 'CD124', 'CD13', 'CD2', 'CD226', 'CD29',\n       'CD303', 'CD49b'] + ['CD81', 'IgD', 'CD18', 'CD28', 'CD38', 'CD127', 'CD45', 'CD22', 'CD71',\n       'CD26', 'CD115', 'CD63', 'CD304', 'CD36', 'CD172a', 'CD72', 'CD158',\n       'CD93', 'CD49a', 'CD49d', 'CD73', 'CD9', 'TCRVa7.2', 'TCRVd2', 'LOX-1',\n       'CD158b', 'CD158e1', 'CD142', 'CD319', 'CD352', 'CD94', 'CD162',\n       'CD85j', 'CD23', 'CD328', 'HLA-E', 'CD82', 'CD101', 'CD88', 'CD224']\nprint(len(list_all_targets))","metadata":{"execution":{"iopub.status.busy":"2022-11-15T10:47:10.392872Z","iopub.execute_input":"2022-11-15T10:47:10.393796Z","iopub.status.idle":"2022-11-15T10:47:10.405983Z","shell.execute_reply.started":"2022-11-15T10:47:10.393746Z","shell.execute_reply":"2022-11-15T10:47:10.404566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Define X,Y, optinal rescaling of Y, cutting feature number in X ","metadata":{}},{"cell_type":"code","source":"# Since metric is - correlation coefficient - any rescaling aY+b will not be change it , so one can do like that:\n# it is not clear is it optimal or not (see https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360253 )\n\nlist_selected_targets = list_all_targets# [:2] #  ['CD31'] # list_all_targets\n\nY = df_cite_train_y[list_selected_targets].values\n\nif  rescale_Y_to_mean0_std1 and (Y.shape[1]>1 ) :\n    Y -= Y.mean(axis=1).reshape(-1, 1)\n    Y /= Y.std(axis=1).reshape(-1, 1)\n    print('Rescaling to mean 0 and std 1 has been done')\n    \n    \nn_features_loc = n_features\n\nl = list(range(1,100)) + list(np.arange(100, 2200)[np.random.permutation(2100)][:600])\n\nX = df_cite.iloc[:,l].values\nprint('X.shape:', X.shape,'Y.shape:', Y.shape)    ","metadata":{"execution":{"iopub.status.busy":"2022-11-15T10:57:12.749458Z","iopub.execute_input":"2022-11-15T10:57:12.753704Z","iopub.status.idle":"2022-11-15T10:57:13.306412Z","shell.execute_reply.started":"2022-11-15T10:57:12.753573Z","shell.execute_reply":"2022-11-15T10:57:13.305183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## start preparation of the specific cross-validation like scheme  ","metadata":{}},{"cell_type":"code","source":"df_meta['Index'] = range(len(df_meta))\nscol = 'donor&day&CT'\ndf_meta[scol] =df_meta['donor'].apply(lambda x:str(x)+'_') + df_meta['day'].apply(lambda x:str(x)+'_') + df_meta['cell_type']\ndf_meta","metadata":{"execution":{"iopub.status.busy":"2022-11-14T12:47:10.675333Z","iopub.execute_input":"2022-11-14T12:47:10.675699Z","iopub.status.idle":"2022-11-14T12:47:10.749224Z","shell.execute_reply.started":"2022-11-14T12:47:10.675673Z","shell.execute_reply":"2022-11-14T12:47:10.747582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# KEY Preparation of the specific cross-validation like scheme  ","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.model_selection import train_test_split\n\ntrain_size_loc = 0.9 # 0.3 # from 0 to 1 # To create diversity we have an option to train on subsets. \nn_blends = 7 # If train_size_loc < 1 - that controls the random seed to create subset of train part where we will actually train the model\nverbose = 1000\n\n\ndef get_list_folds_indices(train_size_loc, n_blends):\n    list_folds_indices_loc = []\n\n\n    # df_folds = pd.DataFrame(index = df_meta.index)\n    # df_folds['Index'] = range(df_folds.shape[0])\n\n    scol = 'donor&day&CT'\n    if train_size_loc < 1:\n        IX_main, IX_complementary = train_test_split(df_meta['Index'].values, \n                             train_size = train_size_loc,  random_state= n_blends,  \n                             stratify = df_meta[scol] ,  shuffle=True )\n    else:\n        IX_main = np.arange(df_meta.shape[0] ) # Everything will be included \n        IX_complementary = np.array([])\n\n    mask_train_main = (df_meta['Index'].isin( IX_main) )\n    c = 0\n    for day2exclude in [2,3,4]: # Day to be excluded from the TRAIN , it defines the test-private-like  \n        for donor2exclude in [32606,  31800]: # We will need to predict always MALE (not female) - like on LB. (# donor 13176 - female)\n            train_index =\\\n                np.where(mask_train_main &  (df_meta['day']  != day2exclude) & ( df_meta['donor']  != donor2exclude  ) )  [0]\n            test_index1_like_private_lb =\\\n                np.where(   (df_meta['day']  == day2exclude)  ) [0]\n            test_index2_like_public_lb =\\\n                np.where( (df_meta['day']  != day2exclude) &  (df_meta['donor']  == donor2exclude ) ) [0]\n\n            list_folds_indices_loc.append( (train_index,  test_index1_like_private_lb , test_index2_like_public_lb ) )\n\n\n            if verbose >= 100:\n                str_fold_inf = 'Fold ' +str(c) + ': Train: excludes Day '+str(day2exclude) + ' and Donor ' + str( donor2exclude )\n                print(str_fold_inf, 'Sizes: train:',len(train_index), 'Test Like Priv'  ,len(test_index1_like_private_lb), \n                      'Test Like Publ',   len(test_index2_like_public_lb),  \n                     ); \n            c+=1        \n\n    list_folds_indices = list_folds_indices_loc.copy()\n    return list_folds_indices    ","metadata":{"execution":{"iopub.status.busy":"2022-11-14T12:47:31.165413Z","iopub.execute_input":"2022-11-14T12:47:31.165767Z","iopub.status.idle":"2022-11-14T12:47:31.177143Z","shell.execute_reply.started":"2022-11-14T12:47:31.165735Z","shell.execute_reply":"2022-11-14T12:47:31.175577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Set the model","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import Ridge\nfrom sklearn.neural_network import MLPRegressor\nfrom sklearn.multioutput import MultiOutputRegressor\nimport xgboost as xgb  \n\nmodel_type_loc = 3\nif model_type_loc == 1:\n    alpha4Ridge  = 1e4\n    model = Ridge(alpha=alpha4Ridge)\n    str_model_inf = str(model)\n    model_inf_str_short = 'RidgeNfeat'+str(X.shape[1])+'_Alpha1e4_'\n    print(str_model_inf, model_inf_str_short)\nif model_type_loc == 2:\n    # From: https://www.kaggle.com/code/visualcomments/mmscel-crossvalidation-schemes/notebook#Concatenating-additional-features\n    model = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-5, random_state=42, \n                         hidden_layer_sizes=(300, 200))\n    str_model_inf = str(model)\n    model_inf_str_short = 'MLP1_NFeat'+str(X.shape[1])\n    print(str_model_inf ) \n    print( model_inf_str_short)\nif model_type_loc == 3:\n    # From: https://www.kaggle.com/code/tttzof351/mmscel-crossvalidation-schemes-193f49\n    model = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-4, random_state=42, \n                         hidden_layer_sizes=(300, 300))\n    str_model_inf = str(model)\n    model_inf_str_short = 'MLP2_NFeat'+str(X.shape[1])\n    print(str_model_inf ) \n    print( model_inf_str_short)\nif model_type_loc == 4:\n    from sklearn.neighbors import KNeighborsRegressor\n    \n    n_neighbors_knn = 15\n    model = MultiOutputRegressor( KNeighborsRegressor(n_neighbors=n_neighbors_knn) )\n    str_model_inf = str(model)\n    model_inf_str_short = 'KNN'+str(n_neighbors_knn)+'_NFeat'+str(X.shape[1])\n    print(str_model_inf ) \n    print( model_inf_str_short)\n    \nif model_type_loc == 5:\n    XGB_PARAMETERS = {         'n_estimators': 225,           'max_depth': 6,           'learning_rate': 0.07190990240552594,\n      'gamma': 2.0460152365496342,           'base_score': 0.5794691454636103,           'tree_method': 'gpu_hist',           \n      'predictor': 'gpu_predictor',           'max_leaves': 13,           'random_state': 123     }     \n    model = MultiOutputRegressor(xgb.XGBRegressor(**XGB_PARAMETERS))     \n    str_model_inf = str(model)     \n    model_inf_str_short = 'XGB_NFeat'+str(X.shape[1])     \n    print(str_model_inf )      \n    print( model_inf_str_short)    \n    \n\n","metadata":{"execution":{"iopub.status.busy":"2022-11-14T12:47:34.849738Z","iopub.execute_input":"2022-11-14T12:47:34.850992Z","iopub.status.idle":"2022-11-14T12:47:34.864565Z","shell.execute_reply.started":"2022-11-14T12:47:34.850939Z","shell.execute_reply":"2022-11-14T12:47:34.863280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Core model train/predict with respect to specific CV-like scheme ","metadata":{}},{"cell_type":"code","source":"model1 = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-4, random_state=42, \n                         hidden_layer_sizes=(300, 300))\nmodel2 = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-4, random_state=42, \n                         hidden_layer_sizes=(200, 200))\nmodel3 = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-4, random_state=42, \n                         hidden_layer_sizes=(300, 200))\nmodel4 = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-3, random_state=0, \n                         hidden_layer_sizes=(300, 300))\nmodel5 = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-3, random_state=0, \n                         hidden_layer_sizes=(200, 200))\nmodel6 = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-3, random_state=0, \n                         hidden_layer_sizes=(300, 200))\nmodel7 = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-5, random_state=17, \n                         hidden_layer_sizes=(300, 300))\nmodel8 = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-5, random_state=17, \n                         hidden_layer_sizes=(200, 200))\nmodel9 = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-5, random_state=17, \n                         hidden_layer_sizes=(300, 200))","metadata":{"execution":{"iopub.status.busy":"2022-11-14T12:48:48.140470Z","iopub.execute_input":"2022-11-14T12:48:48.140836Z","iopub.status.idle":"2022-11-14T12:48:48.153562Z","shell.execute_reply.started":"2022-11-14T12:48:48.140811Z","shell.execute_reply":"2022-11-14T12:48:48.152178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"models= [model1, model2, model3, model4, model5, model6, model7, model8, model9]","metadata":{"execution":{"iopub.status.busy":"2022-11-14T12:49:31.539642Z","iopub.execute_input":"2022-11-14T12:49:31.540039Z","iopub.status.idle":"2022-11-14T12:49:31.546808Z","shell.execute_reply.started":"2022-11-14T12:49:31.540008Z","shell.execute_reply":"2022-11-14T12:49:31.545592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport time\n\nflag_train_on_full_data4submit = False\nflag_save_predictions_on_each_sub_fold = False\nflag_calculate_and_save_predictions_for_test_public_like = False \n\ndict_model_results = {}\n\nY_pred_oof_private_like = np.zeros_like(Y)\nif flag_calculate_and_save_predictions_for_test_public_like:\n    Y_pred_oof_public_like = np.zeros_like(Y) # FEMALE is always in train - thus we have no predictions for female donor for that test part\n\nY_pred_submission_Kaggle_way = np.zeros( (119651-70988 , Y.shape[1]) )  # Way 2 (Kaggle way) - average predictions of models on each fold\nif flag_train_on_full_data4submit: \n    Y_pred_submission_classical_way = np.zeros( (119651-70988 , Y.shape[1]) )  # Way 1 (Classical way) - retrain model on the whole data\n\nt0 = time.time()\n\nfor i in models:\n\n    for n_seed in range(n_blends):    \n        \n        list_folds_indices = get_list_folds_indices(train_size_loc, n_seed ) \n\n        l = list(range(1,100)) + list(np.arange(100, 2200)[np.random.permutation(2100)][:600])\n        X = df_cite.iloc[:,l].values\n        \n        \n        for fold, indices_tuple  in enumerate( list_folds_indices ):\n            t00 = time.time()\n            random_state_MLP = (fold+1)*(n_seed)\n            model = i\n            train_index = indices_tuple[0]\n            model.fit(X[train_index], Y[train_index])\n            if verbose >= 100:\n                print('n_seed: ', n_seed, 'fold:',fold, 'Train time:%.2f'%(time.time()-t00 ), 'X_train.shape:',X[train_index].shape )\n\n            # Calculate predictions for test-private-like\n            i_loc = 1\n            indices_loc = indices_tuple[i_loc] # Private-like test\n            y_pred = model.predict( X[indices_loc ])\n            Y_pred_oof_private_like[indices_loc ] += y_pred * 0.5  / n_blends  # 0.5 - comes from the fact that our \"quasi-folds\" have intersection OOF parts - which correspond to \"DAYs\"\n            if flag_save_predictions_on_each_sub_fold:\n                dict_model_results['n_seed: ', n_seed,'Pred_fold'+str(fold)+'_part'+str(i_loc)] = y_pred\n\n\n            if (flag_save_predictions_on_each_sub_fold) or ( flag_calculate_and_save_predictions_for_test_public_like):\n                # Calculate predictions for test-public-like\n                i_loc = 2\n                indices_loc = indices_tuple[i_loc] # Public-like test  # FEMALE is always in train - thus we have no predictions for female donor for that test part\n                y_pred = model.predict( X[indices_loc ])\n                if flag_calculate_and_save_predictions_for_test_public_like:\n                    Y_pred_oof_public_like[indices_loc ] += y_pred * 0.5 / n_blends # 0.5 - comes from the fact that our \"quasi-folds\" have intersection OOF parts\n                                                    # Since we THREE DAYs we should put  1/(3-1) = 1/2\n                if flag_save_predictions_on_each_sub_fold:\n                    dict_model_results['n_seed: ', n_seed,'Pred_fold'+str(fold)+'_part'+str(i_loc)] = y_pred\n\n            if flag_save_predictions_on_each_sub_fold:\n                # Save train predictions in case some deep analysis will be done \n                i_loc = 0  \n                indices_loc = indices_tuple[i_loc] # Public-like test  # FEMALE is always in train - thus we have no predictions for female donor for that test part\n                y_pred = model.predict( X[indices_loc ])\n                dict_model_results['n_seed: ', n_seed,'Pred_fold'+str(fold)+'_part'+str(i_loc)] = y_pred    \n             \n            # Prepare submission data in way 2 - \"Kaggle way\"\n            Y_pred_submission_Kaggle_way += model.predict( X[70988:,: ])   # Way 2 (Kaggle way) - average predictions of models on each fold   \n\n            if verbose >= 100:\n                print('n_seed: ', n_seed,'Finished fold:',fold, 'Train&Inference time:%.2f'%(time.time()-t00 ), 'X_train.shape:',X[train_index].shape )\n        \n        if flag_train_on_full_data4submit: \n            model.fit(X[:70988,:], Y) # Way 1 (Classical way) - retrain model on the whole data\n            Y_pred_submission_classical_way += model.predict( X[70988:,: ]) \n\n    \n    Y_pred_submission_Kaggle_way /= (n_blends*(1+fold))  # Do not forget to average  predictions from all folds \n\n    if flag_train_on_full_data4submit:\n        Y_pred_submission_classical_way /= n_blends\n  \n\ndict_model_results['Time'] = time.time() - t0\n#dict_model_results['train_size '] = train_size(train_size_loc, n_seed)[:(fold+1)]\ndict_model_results['list_folds_indices'] = list_folds_indices[:(fold+1)]\ndict_model_results['Y_pred_oof_private_like'] = Y_pred_oof_private_like\nif flag_calculate_and_save_predictions_for_test_public_like:\n    dict_model_results['Y_pred_oof_public_like'] = Y_pred_oof_public_like\nif flag_train_on_full_data4submit: \n    dict_model_results['Y_pred_submission_classical_way'] = Y_pred_submission_classical_way\ndict_model_results['Y_pred_submission_Kaggle_way'] = Y_pred_submission_Kaggle_way\n\n\nif verbose >= 10:\n    print('n_seed: ', n_seed,'X.shape, Y.shape:',X.shape, Y.shape)\n    print('n_seed: ', n_seed,'Finished. Folds:',fold+1, 'Time:%.2f'%(time.time()-t0), 'Y_pred_oof_private_like.shape:',Y_pred_oof_private_like.shape,\n         'Y_pred_submission_Kaggle_way.shape:',Y_pred_submission_Kaggle_way.shape )    \n","metadata":{"execution":{"iopub.status.busy":"2022-11-14T12:49:34.868021Z","iopub.execute_input":"2022-11-14T12:49:34.868339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save predictions to CSV","metadata":{}},{"cell_type":"markdown","source":"## Save key OOF and submit predictions","metadata":{}},{"cell_type":"code","source":"%%time \n\nlist_targets_names = list_selected_targets # list( df_cite_train_y.columns )\n\nfilename_prefix = model_inf_str_short # 'Ridge100_1e4_'\nfilename_postfix = ''\nfor key_loc in ['Y_pred_oof_private_like', 'Y_pred_oof_public_like', 'Y_pred_submission_classical_way', 'Y_pred_submission_Kaggle_way'  ]:\n    if key_loc in dict_model_results.keys():\n        fn_loc = filename_prefix + key_loc + filename_postfix + '.csv'\n        print('Key:', key_loc,'Filename:',  fn_loc)\n        d = pd.DataFrame(dict_model_results[key_loc],  columns = list_targets_names )\n        d.to_csv( fn_loc)\n        print(d.shape)\n        display(d.head(2))\n        display(d.tail(2))\n        print(); print();\n        \n","metadata":{"execution":{"iopub.status.busy":"2022-11-13T09:55:53.999511Z","iopub.execute_input":"2022-11-13T09:55:53.999951Z","iopub.status.idle":"2022-11-13T09:56:07.706833Z","shell.execute_reply.started":"2022-11-13T09:55:53.999916Z","shell.execute_reply":"2022-11-13T09:56:07.705353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Save (optionaly) foldwise predictions ","metadata":{}},{"cell_type":"code","source":"%%time\nif flag_save_foldwise_predictions:\n    verbose = 1000\n    \n    list_folds_indices =  dict_model_results['list_folds_indices']\n    df_fold_score_stat = pd.DataFrame();df_fold_score_stat.index.name = 'Fold'\n    t0 = time.time()\n    for fold, indices_tuple  in enumerate( list_folds_indices ):\n        for i_loc in [1,2,0]: #, indices_loc in enumerate( indices_tuple) :\n            if i_loc == 1:\n                name_loc = 'Test Priv Like'\n            if i_loc == 2:\n                name_loc = 'Test Publ Like'\n            if i_loc == 0:\n                name_loc = 'Train'\n\n            indices_loc = indices_tuple[i_loc]\n            key_loc = 'Pred_fold'+str(fold)+'_part'+str(i_loc)\n            if key_loc in dict_model_results.keys():\n                y_pred = dict_model_results[key_loc]\n                if verbose >= 100:\n                    fn_loc = 'unimportant_foldwise_'+filename_prefix + key_loc + '_'+name_loc + filename_postfix + '.csv'\n                    df_tmp = pd.DataFrame(index = indices_loc, data = y_pred, columns =  list_selected_targets )\n                    df_tmp.to_csv(fn_loc)\n                    if verbose >= 1000:\n                        print(key_loc)\n                        display(df_tmp.head(3))\n                        display(df_tmp.tail(2))\n                    print(key_loc,'File saved:', fn_loc, 'shape:', df_tmp.shape  ); print(); print()\n                \n                \n        if verbose >= 100:\n            print('Fold:', fold, 'Shapes of train:', train_index.shape, 'Tests: ',[t.shape for t in indices_tuple[1:] ] )        \n","metadata":{"execution":{"iopub.status.busy":"2022-11-13T09:57:10.742030Z","iopub.execute_input":"2022-11-13T09:57:10.742451Z","iopub.status.idle":"2022-11-13T09:57:10.754045Z","shell.execute_reply.started":"2022-11-13T09:57:10.742413Z","shell.execute_reply":"2022-11-13T09:57:10.752255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Correlation Scoring Function\n","metadata":{"_uuid":"36ecbbd7-f9a3-4b97-843d-1b787d6bd3c9","_cell_guid":"a5757f02-9083-4caf-a0b8-fff7d956fb21","trusted":true}},{"cell_type":"code","source":"import gc\ndef correlation_score(y_true, y_pred):\n    \"\"\"Scores the predictions according to the competition rules. \n    \n    It is assumed that the predictions are not constant.\n    \n    Returns the average of each sample's Pearson correlation coefficient\"\"\"\n    \n    # Input should be matrices - does not make sense for vectors \n    if  len(y_pred.shape)< 2: return -10 # Some result to inform for incorrect input\n    if  y_pred.shape[1] < 2: return -10 # Some result to inform for incorrect input\n\n    y2 = y_pred.copy()\n    y2 -= y2.mean(axis=1).reshape(-1, 1);    y2 /= y2.std(axis=1).reshape(-1, 1)    \n    if rescale_Y_to_mean0_std1:\n        y1 = y_true # Already rescaled \n    else:\n        y1 = y_true.copy(); \n        y1 -= y1.mean(axis=1).reshape(-1, 1);    y1 /= y1.std(axis=1).reshape(-1, 1) \n        \n    c = (y1*y2).mean().mean()# Correlation for rescaled matrices is just matrix product and average \n    \n    # Memory control:\n    if not rescale_Y_to_mean0_std1:\n        del y1\n    del y2\n    gc.collect()\n    \n    return c\n\n    # Slow way:\n    \n    if type(y_true) == pd.DataFrame: y_true = y_true.values\n    if type(y_pred) == pd.DataFrame: y_pred = y_pred.values\n    corrsum = 0\n    for i in range(len(y_true)):\n        corrsum += np.corrcoef(y_true[i], y_pred[i])[1, 0]\n    return corrsum / len(y_true)","metadata":{"_uuid":"0e59893b-3efb-4c89-bb2a-5dee2cbb4239","_cell_guid":"eda4e4b5-494e-4f86-aac2-ef97b9884596","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-13T09:57:13.088305Z","iopub.execute_input":"2022-11-13T09:57:13.088717Z","iopub.status.idle":"2022-11-13T09:57:13.101559Z","shell.execute_reply.started":"2022-11-13T09:57:13.088684Z","shell.execute_reply":"2022-11-13T09:57:13.099864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_error\nfrom sklearn.metrics import r2_score","metadata":{"_uuid":"f89b4512-7d1c-406d-8f37-67c07d1bedfc","_cell_guid":"c19ef489-3fd3-4e6e-bd2a-e642789651de","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-13T08:40:36.320947Z","iopub.execute_input":"2022-11-13T08:40:36.321354Z","iopub.status.idle":"2022-11-13T08:40:36.328028Z","shell.execute_reply.started":"2022-11-13T08:40:36.321319Z","shell.execute_reply":"2022-11-13T08:40:36.326189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Main scoring - via correlation metric ","metadata":{}},{"cell_type":"code","source":"%%time\n# 'Y_pred_oof_private_like', 'Y_pred_oof_public_like',\nY_pred = dict_model_results['Y_pred_oof_private_like']\nY_true = Y\n\ns_cor = correlation_score( Y_true , Y_pred  ) # It takes about 5 seconds ! \n# If only one target it will automatically return -10\n\n\ns_r2 = r2_score( Y_true , Y_pred  )\ns_mse = mean_squared_error( Y_true , Y_pred  )\n\nprint(str(model)[:80], 'n_feat:', X.shape[1],)\nprint('OOF scores: ', s_cor, s_r2, s_mse, 'n_feat:', X.shape[1], 'Modeling time:%.3f secs.'%dict_model_results['Time'],  \n      'Oof size:', Y_true.shape[0], 'Y_pred.shape:', Y_pred.shape )","metadata":{"execution":{"iopub.status.busy":"2022-11-13T09:57:15.835002Z","iopub.execute_input":"2022-11-13T09:57:15.835720Z","iopub.status.idle":"2022-11-13T09:57:16.327160Z","shell.execute_reply.started":"2022-11-13T09:57:15.835685Z","shell.execute_reply":"2022-11-13T09:57:16.325950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds1 = pd.DataFrame()\nIX_loc = 0\nds1.loc[IX_loc, 'Model'] = str_model_inf\nds1.loc[IX_loc, 'Corr OOF'] = s_cor\nds1.loc[IX_loc, 'r2 OOF'] = s_r2\nds1.loc[IX_loc, 'mse OOF'] = s_mse\nds1.loc[IX_loc, 'n_feat'] = X.shape[1]\nds1.loc[IX_loc, 'n_targets'] = Y_pred.shape[1]\nds1.loc[IX_loc, 'Time'] = np.round(dict_model_results['Time'],3)\nfn_loc = filename_prefix + 'ModelStat1Main_' + filename_postfix + '.csv'\nds1.to_csv(fn_loc)\nprint(fn_loc)\nds1\n","metadata":{"execution":{"iopub.status.busy":"2022-11-13T09:57:22.140013Z","iopub.execute_input":"2022-11-13T09:57:22.140547Z","iopub.status.idle":"2022-11-13T09:57:22.164652Z","shell.execute_reply.started":"2022-11-13T09:57:22.140517Z","shell.execute_reply":"2022-11-13T09:57:22.163365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Additional (optional) info on scores - foldwise","metadata":{}},{"cell_type":"code","source":"%%time\nverbose = 1000\nlist_folds_indices =  dict_model_results['list_folds_indices']\ndf_fold_score_stat = pd.DataFrame();df_fold_score_stat.index.name = 'Fold'\nt0 = time.time()\nfor fold, indices_tuple  in enumerate( list_folds_indices ):\n    # Calculate metrics on test and train folds \n    list_scores = []; list_scores_r2 = []\n    for i_loc in [1,2,0]: #, indices_loc in enumerate( indices_tuple) :\n        if i_loc == 1:\n            name_loc = 'Test Priv Like'\n        if i_loc == 2:\n            name_loc = 'Test Publ Like'\n        if i_loc == 0:\n            name_loc = 'Train'\n            \n        indices_loc = indices_tuple[i_loc]\n        key_loc = 'Pred_fold'+str(fold)+'_part'+str(i_loc)\n        if key_loc in dict_model_results.keys():\n            y_pred = dict_model_results[key_loc]\n            s = correlation_score(Y[indices_loc],  y_pred    )\n            df_fold_score_stat.loc[fold,'Corr '+name_loc] = s\n            s = r2_score( Y[indices_loc] , y_pred  )\n            df_fold_score_stat.loc[fold,'r2 '+name_loc] = s\n            s = mean_squared_error( Y[indices_loc] , y_pred  )\n            df_fold_score_stat.loc[fold,'mse '+name_loc] = s\n    if verbose >= 1000:\n        print('Fold:', fold, 'Shapes of train:', train_index.shape, 'Tests: ',[t.shape for t in indices_tuple[1:] ] )        \ndisplay( df_fold_score_stat )\nif len(df_fold_score_stat)> 0:\n    display(df_fold_score_stat.describe(percentiles=[] ).iloc[1:,:])\n    ds2 = pd.concat([df_fold_score_stat.describe(percentiles=[] ).iloc[1:,:],  df_fold_score_stat])\nelse:\n    ds2 = df_fold_score_stat\n    \nds2\n","metadata":{"execution":{"iopub.status.busy":"2022-11-13T08:47:36.310944Z","iopub.execute_input":"2022-11-13T08:47:36.311315Z","iopub.status.idle":"2022-11-13T08:47:36.343577Z","shell.execute_reply.started":"2022-11-13T08:47:36.311281Z","shell.execute_reply":"2022-11-13T08:47:36.341610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfn_loc = filename_prefix + 'ModelStat2Foldwise_' + filename_postfix + '.csv'\nds2.to_csv(fn_loc)","metadata":{"execution":{"iopub.status.busy":"2022-11-12T09:47:00.735297Z","iopub.status.idle":"2022-11-12T09:47:00.735791Z","shell.execute_reply.started":"2022-11-12T09:47:00.735524Z","shell.execute_reply":"2022-11-12T09:47:00.735546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif len(df_fold_score_stat) > 0:\n    t = df_fold_score_stat.describe(percentiles=[] ).iloc[1,: ].T\n    t= t.to_frame().T.reset_index()\n\n    ds1b = pd.concat([ds1, t], axis = 1 )\n    fn_loc = filename_prefix + 'ModelStat1Main_' + filename_postfix + '.csv'\n    ds1b.to_csv(fn_loc)\n    ds1b","metadata":{"execution":{"iopub.status.busy":"2022-11-12T09:47:00.737573Z","iopub.status.idle":"2022-11-12T09:47:00.738142Z","shell.execute_reply.started":"2022-11-12T09:47:00.737836Z","shell.execute_reply":"2022-11-12T09:47:00.737859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Show main and additonal (foldwise) scores","metadata":{}},{"cell_type":"code","source":"%%time\ntry:\n    display(ds1)\nexcept:\n    pass\ntry:\n    display(ds1b)\nexcept:\n    pass\ntry:\n    display( ds2 )\nexcept:\n    pass\n","metadata":{"execution":{"iopub.status.busy":"2022-11-12T09:47:00.740335Z","iopub.status.idle":"2022-11-12T09:47:00.740808Z","shell.execute_reply.started":"2022-11-12T09:47:00.740563Z","shell.execute_reply":"2022-11-12T09:47:00.740586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analysis predictions TargetWISELY (Optionals) ","metadata":{}},{"cell_type":"code","source":"%%time\n\ndf_targetwise = pd.DataFrame()\n\nY_pred = dict_model_results['Y_pred_oof_private_like']\nY_true = Y\nfor i,col in enumerate(list_selected_targets):\n    IX = col#.columns[i_loc]\n    df_targetwise.loc[IX, 'absMean'] = df_cite_train_y[col].abs().mean()\n    df_targetwise.loc[IX, 'std'] = df_cite_train_y[col].std()\n    df_targetwise.loc[IX, 'mse'] = mean_squared_error( Y_true[:,i] , Y_pred[:,i]  )    \n    df_targetwise.loc[IX, 'r2'] = r2_score( Y_true[:,i] , Y_pred[:,i]  )\n    df_targetwise.loc[IX, 'corr'] = np.corrcoef( Y_true[:,i] , Y_pred[:,i]  )[0,1]    \n    df_targetwise.loc[IX, 'mean'] = df_cite_train_y[col].mean()\n    df_targetwise.loc[IX, 'max'] = df_cite_train_y[col].max()\n    \nprint('Pearson correlation between scores (and other characteristics):')\ncm = df_targetwise.corr()  \ndisplay(cm )\nsns.heatmap(cm)\nplt.show()\nprint('Spearman correlation between scores (and other characteristics):')\ncm = df_targetwise.corr(method = 'spearman') \ndisplay( cm )\nsns.heatmap(cm)\nplt.show()\n\n\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['std'].plot(); plt.legend(); plt.grid();\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['mse'].plot(); plt.legend(); plt.grid();\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['r2'].plot(); plt.legend(); plt.grid();\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['corr'].plot(); plt.legend(); plt.grid();\nfig = plt.figure(figsize = (20,4) )\ndf_targetwise['absMean'].plot(); plt.legend(); plt.grid();\n\nd1 = df_targetwise.sort_values('r2'); d1.index.name = 'r2 sorted';  d1 = d1.reset_index()\nd2 = df_targetwise.sort_values('mse',ascending = False); d2.index.name = 'mse sorted';  d2 = d2.reset_index()\nd3 = df_targetwise.sort_values('corr'); d3.index.name = 'corr sorted';  d3 = d3.reset_index()\nd_concat_several_sorting = pd.concat([d1,d2,d3],axis = 1)\nprint('\\n Worst:')\ndisplay(d_concat_several_sorting.head(20))\nprint('\\n Best:')\ndisplay(d_concat_several_sorting.tail(20))","metadata":{"_uuid":"b372da4f-a87d-4b71-a071-414e63564d37","_cell_guid":"de9b5839-be50-4f31-8fae-21d9a4b8cd7e","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-12T09:47:00.743069Z","iopub.status.idle":"2022-11-12T09:47:00.743616Z","shell.execute_reply.started":"2022-11-12T09:47:00.743309Z","shell.execute_reply":"2022-11-12T09:47:00.743332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfor nm_loc in ['Y_pred_oof_private_like', 'Y_pred_oof_public_like']:\n    if nm_loc not in dict_model_results.keys(): continue \n        \n    print('Predictions on:', nm_loc)\n    Y_pred = dict_model_results[nm_loc]\n    Y_pred = Y_pred#[  indices_oof , : ]\n    Y_true = Y#[  indices_oof , : ]\n    print('Y_pred.shape: ', Y_pred.shape, 'Y_true.shape: ', Y_true.shape,  )\n    list_corr = []\n    for i in range(Y_pred.shape[0]):\n        list_corr.append(np.corrcoef(Y_true[i,:], Y_pred[i,:] )[0,1] )\n\n\n    fig = plt.figure( figsize = (10,5))\n    plt.title('Main score - correlation \\n'+str(nm_loc))\n    plt.hist(list_corr, bins = 100)\n    plt.grid()\n    plt.show()\n    display(pd.Series(list_corr).describe() )","metadata":{"execution":{"iopub.status.busy":"2022-11-12T09:47:00.745691Z","iopub.status.idle":"2022-11-12T09:47:00.746801Z","shell.execute_reply.started":"2022-11-12T09:47:00.746518Z","shell.execute_reply":"2022-11-12T09:47:00.746544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfor nm_loc in ['Y_pred_oof_private_like']:# , 'Y_pred_oof_public_like']:\n    print('Predictions on:', nm_loc)\n    Y_pred = dict_model_results[nm_loc]\n    Y_pred = Y_pred#[  indices_oof , : ]\n    Y_true = Y#[  indices_oof , : ]\n    t = (Y_pred - Y_true)\n    print('Y_pred.shape: ', Y_pred.shape, 'Y_true.shape: ', Y_true.shape,  )\n\n    fig = plt.figure( figsize = (10,5))\n    plt.title('Error \\n std= %.3f MAE = %.3f ' %( np.std(t.ravel() ), np.mean(np.abs(t.ravel())))   )\n    plt.hist(t.ravel(), bins = 100)\n    plt.grid()\n    plt.show()\n    display(pd.Series(t.ravel()).describe() )","metadata":{"execution":{"iopub.status.busy":"2022-11-12T09:47:00.748012Z","iopub.status.idle":"2022-11-12T09:47:00.748881Z","shell.execute_reply.started":"2022-11-12T09:47:00.748617Z","shell.execute_reply":"2022-11-12T09:47:00.74864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare submission","metadata":{}},{"cell_type":"markdown","source":"# Load public multiome part or fill it by zeros ","metadata":{}},{"cell_type":"code","source":"%%time\n# Normally that part works Wall time: 1min 11s - however sometimes Kaggle system does not work correctly and it never stops\nif (flag_prepare_submission) and (Y.shape[1] == constant_number_of_original_targets):\n    mode_subm = 'USE: sskknts MSCI CITEseq Keras Quickstart + Dropout' # 'put_zeros_to_multiome_part'\n\n    if mode_subm == 'USE: sskknts MSCI CITEseq Keras Quickstart + Dropout' :\n        fn_loc = '/kaggle/input/msci-citeseq-keras-quickstart-dropout/submission.csv'\n        fn_loc = '/kaggle/input/solutions-for-multimodal-singlecell-integration/msci-citeseq-keras-quickstart-dropout_submission.csv'\n        df_submission_full = pd.read_csv(fn_loc,\n                                 index_col='row_id', squeeze=True)\n        df_submission_full = df_submission_full.to_frame()\n    elif mode_subm == 'put_zeros_to_multiome_part':\n        df_submission_full = pd.DataFrame(index = range(65_744_180), columns = ['target'], data = np.zeros(65_744_180) )\n        df_submission_full.index.name = 'row_id'\n\n    display(df_submission_full.info() )\n    print()\n    display(df_submission_full.head(10))","metadata":{"_uuid":"5d8d0a31-a09c-490d-8387-464a59f1dd23","_cell_guid":"5ca85d26-228a-4f42-910c-a380f4827022","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-12T09:47:00.750291Z","iopub.status.idle":"2022-11-12T09:47:00.751122Z","shell.execute_reply.started":"2022-11-12T09:47:00.750856Z","shell.execute_reply":"2022-11-12T09:47:00.75088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif (flag_prepare_submission) and (Y.shape[1] == constant_number_of_original_targets):\n    Y_pred4submit = dict_model_results['Y_pred_submission_Kaggle_way']\n    print(Y_pred4submit.shape)","metadata":{"execution":{"iopub.status.busy":"2022-11-12T09:47:00.752484Z","iopub.status.idle":"2022-11-12T09:47:00.753474Z","shell.execute_reply.started":"2022-11-12T09:47:00.753207Z","shell.execute_reply":"2022-11-12T09:47:00.753241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif (flag_prepare_submission) and (Y.shape[1] == constant_number_of_original_targets):\n    # We can prepare CITE-seq part as: Y_pred4submit.values.ravel() -- see notebook: https://www.kaggle.com/code/alexandervc/mmscel-analysis-of-evaluation-and-submission\n    df_submission_full['target'].iloc[:6_812_820] = Y_pred4submit.ravel()\n    display(df_submission_full.head(20))\n    display(df_submission_full.tail(20))\n    ","metadata":{"execution":{"iopub.status.busy":"2022-11-12T09:47:00.755151Z","iopub.status.idle":"2022-11-12T09:47:00.755628Z","shell.execute_reply.started":"2022-11-12T09:47:00.755386Z","shell.execute_reply":"2022-11-12T09:47:00.755408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Wall time: 1min 46s\nif (flag_prepare_submission) and (Y.shape[1] == constant_number_of_original_targets):\n    submit_filename_postfix = model_inf_str_short\n    df_submission_full.to_csv('submission_'+submit_filename_postfix +'.csv')","metadata":{"execution":{"iopub.status.busy":"2022-11-12T09:47:00.75773Z","iopub.status.idle":"2022-11-12T09:47:00.758471Z","shell.execute_reply.started":"2022-11-12T09:47:00.75822Z","shell.execute_reply":"2022-11-12T09:47:00.758244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )","metadata":{"_uuid":"f55e2443-fd24-4fc6-9238-fec592d87d02","_cell_guid":"d3eca9d9-7cdc-4323-beec-8e7d1c64fc84","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-11-12T09:47:00.760023Z","iopub.status.idle":"2022-11-12T09:47:00.761058Z","shell.execute_reply.started":"2022-11-12T09:47:00.760777Z","shell.execute_reply":"2022-11-12T09:47:00.760802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}