{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59094,"databundleVersionId":7010844,"sourceType":"competition"},{"sourceId":6891925,"sourceType":"datasetVersion","datasetId":3841755},{"sourceId":7068220,"sourceType":"datasetVersion","datasetId":3955392},{"sourceId":7122895,"sourceType":"datasetVersion","datasetId":3948965}],"dockerImageVersionId":30558,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nHere we demonstrate quite poor relation between CV scores and LB scores for the challenge. We analyse what metrics and what folds/cell types are mostly correlated with the public LB scores. \n\nWe generated about 47 submissions by very different models, features, encoding schemes etc, and saved their out-of-fold predictions (to calculate various CV-metrics). In the present notebook we go through these saved out-of-fold predictions, calculate several CV-metrics , compare them to LB score , analyse and present conclusions.\n\nFurther comments can be found in the post: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458834\n\n\n### Conclusions:\n\n    0 NO metric/fold shows good correlation with LB (at most 0.5)\n    1 Suprisingly: NOT the competition metric ( mrrmse ) is the best correlated with the LB score, but the row-wise correlation\n    2 Moreover mrrmse  CV metric for AmbrosM and MT CV-schemes - shows practically ZERO correlation with LB  ! \n    2b  For random folds it is a bit better : 0.2\n    3 NK-cells shows the best relation to LB with respect to MRRMSE  0.22 , but still it is very low.\n    4 T-regs cells show the best correlation with LB with respect to the row-wise correlation metric\n    5 The next are CD8 cells - which is quite surprising - since they are quite different from test - their exclusion uplifts scores of some models\n    Later added:\n    6a for MRRMSE-corrected i.e. taking into account only those Y_true which might be biologically meaningful (greater than threshold like 1 or 2 or 3)\n        (proposed in  Antoine Passemiers and Jalil Nourisa writeup: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461159 )\n        we see better correlation with LB than for local straightforward-original MRRMSE - for AmbrosM CV\n        That can be explained - p-values above say 0.01 correspond to noise, and metric should not take them into account.\n        So excluding them from metric - gives metric more sense. \n    6b If we look by folds, then suprisingly T-cells CD4+ cells arise as best correlated to LB fold, NK-cells are the next (we have already seen NK cells above).\n        (see tables below, and section \"MRRMSE-corrected (Antoine Passemiers and Jalil Nourisa )\")\n    6c For Random CV correlations are quite very low even for MRRMS-correct, but again corrected is better correlated to LB,\n        while for MT scheme it is different(!) , suprisignly MRRMSE-corrected is NOT better correlated to LB. \n\n        \n    \n    \n(See more details on CD8+ T-cells here: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458842 )\n\nThe dataset with submits/oof-predictions: https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc/, the information on the models https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing  (both made public **during** the challenge for the sake of the community befefit - see links in the post: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444825 ). \n\nSome general look how diverse/correlated are the predictions can be seen from the clustermap picture here:\nhttps://www.kaggle.com/code/alexandervc/op2-submits-correlations-and-analysis?scriptVersionId=150274308&cellId=11\n\n\n### Navigation and main tables\n\nThe main tables with analysis can be found in the notebook section: \nhttps://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis/notebook#Main---LB-vs-CV-correlation-for-different-metrics\n\nWe explore different CV-schemes (AmbrosM (all drugs/only public drugs) , MT scheme, Random folds) in different versions of the present notebook:\n\nResults on standard AmbrosM scheme (all drugs):\n(Notebook version 20: https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis?scriptVersionId=153470048&cellId=19 )\n\n|         Metric      | Pearson correlation LB to CV: AmbrosM All drugs | Spearman correlation LB to CV: AmbrosM all drugs | Fold Info                 |\n|---------------------|---------------------------------------|----------------------------------------|---------------------------|\n| LB                  | 1.00                                  | 1.00                                   | mean over folds           |\n| corr row3           | -0.43                                 | -0.46                                  | T regulatory cells       |\n| corr row mean       | -0.41                                 | -0.51                                  | mean over folds           |\n| corr col1           | -0.36                                 | -0.42                                  | T cells CD4+              |\n| corr row0           | -0.35                                 | -0.47                                  | NK cells                  |\n| corr col0           | -0.28                                 | -0.49                                  | NK cells                  |\n| corr row2           | -0.27                                 | -0.37                                  | T cells CD8+              |\n| corr col mean       | -0.26                                 | -0.39                                  | mean over folds           |\n| mrrmse0             | 0.25                                  | 0.36                                   | NK cells                  |\n| corr row1           | -0.21                                 | -0.21                                  | T cells CD4+              |\n| corr col2           | -0.20                                 | -0.32                                  | T cells CD8+              |\n| mrrmse1             | 0.17                                  | 0.15                                   | T cells CD4+              |\n| mrrmse2             | -0.13                                 | -0.13                                  | T cells CD8+              |\n| mrrmse mean         | 0.11                                  | 0.04                                   | mean over folds           |\n| mrrmse3             | 0.09                                  | -0.00                                  | T regulatory cells       |\n| corr col3           | 0.07                                  | -0.01                                  | T regulatory cells       |\n| CV                  | 0.02                                  | 0.06                                   | mean over folds           |\n\n\nResults on modified AmbrosM scheme (use drugs from public LB only):\n(Notebook version 19: https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis?scriptVersionId=153469064&cellId=20 )\n\n\n|        Metric       | Pearson correlation LB to CV: AmbrosM only public drugs | Spearman correlation LB to CV: AmbrosM only public drugs | Fold Info                 |\n|---------------------|---------------------------------------|----------------------------------------|---------------------------|\n| LB                  | 1.00                                  | 1.00                                   | mean over folds           |\n| corr row3           | -0.50                                 | -0.47                                  | T regulatory cells       |\n| corr row mean       | -0.42                                 | -0.43                                  | mean over folds           |\n| corr col2           | -0.37                                 | -0.32                                  | T cells CD8+              |\n| corr row2           | -0.29                                 | -0.20                                  | T cells CD8+              |\n| corr col1           | -0.28                                 | -0.42                                  | T cells CD4+              |\n| corr row0           | -0.27                                 | -0.32                                  | NK cells                  |\n| corr col3           | 0.25                                  | 0.16                                   | T regulatory cells       |\n| corr row1           | -0.23                                 | -0.21                                  | T cells CD4+              |\n| mrrmse0             | 0.22                                  | 0.26                                   | NK cells                  |\n| corr col mean       | -0.19                                 | -0.23                                  | mean over folds           |\n| mrrmse2             | -0.14                                 | -0.16                                  | T cells CD8+              |\n| mrrmse1             | 0.09                                  | -0.05                                  | T cells CD4+              |\n| CV                  | 0.02                                  | 0.06                                   | mean over folds           |\n| mrrmse mean         | 0.02                                  | -0.07                                  | mean over folds           |\n| corr col0           | 0.02                                  | -0.12                                  | NK cells                  |\n| mrrmse3             | -0.01                                 | -0.09                                  | T regulatory cells       |\n\n\nResults on MT-scheme: \n(Notebook version 21: https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis?scriptVersionId=153471382&cellId=19 )\n\n|     Metric          | Pearson correlation LB to CV: MT | Spearman correlation LB to CV: MT | Fold Info                                                        |\n|---------------------|----------------------------------|-----------------------------------|------------------------------------------------------------------|\n| LB                  | 1.00                             | 1.00                              | mean over folds                                                  |\n| corr col mean       | -0.33                            | -0.30                             | mean over folds                                                  |\n| corr col1           | -0.31                            | -0.07                             | ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III']                                   |\n| corr row0           | -0.30                            | -0.47                             | ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189', 'Linagliptin', 'O-Demethylated Adapalene']                              |\n| mrrmse0             | 0.28                             | 0.38                              | ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189', 'Linagliptin', 'O-Demethylated Adapalene']                              |\n| corr row mean       | -0.24                            | -0.30                             | mean over folds                                                  |\n| corr row2           | -0.19                            | -0.15                             | ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol', 'R428']                                                    |\n| corr row1           | -0.17                            | -0.16                             | ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III']                                   |\n| mrrmse mean         | 0.16                             | 0.05                              | mean over folds                                                  |\n| corr col0           | -0.14                            | -0.19                             | ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189', 'Linagliptin', 'O-Demethylated Adapalene']                              |\n| corr col2           | -0.13                            | 0.06                              | ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol', 'R428']                                                    |\n| mrrmse1             | 0.11                             | -0.00                             | ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III']                                   |\n| CV                  | 0.04                             | 0.01                              | mean over folds                                                  |\n| mrrmse2             | 0.01                             | -0.04                             | ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol', 'R428']                                                    |\n\n\nResults on random CV-scheme: \n(Notebook version 22: https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis?scriptVersionId=153476379&cellId=19 )\n\n\n\n|         Metric      | Pearson correlation LB to CV: Random | Spearman correlation LB to CV: Random |\n|---------------------|------------------------------------|----------------------------------------|\n| LB                  | 1.00                               | 1.00                                   |\n| corr row2           | -0.38                              | -0.49                                  |\n| corr row1           | -0.37                              | -0.45                                  |\n| corr row3           | -0.36                              | -0.59                                  |\n| corr row mean       | -0.36                              | -0.50                                  |\n| corr row4           | -0.31                              | -0.43                                  |\n| corr col4           | -0.28                              | -0.38                                  |\n| corr row0           | -0.27                              | -0.29                                  |\n| corr col0           | -0.25                              | -0.30                                  |\n| mrrmse4             | 0.22                               | 0.40                                   |\n| CV                  | 0.22                               | 0.24                                   |\n| corr col mean       | -0.09                              | -0.27                                  |\n| corr col1           | 0.06                               | -0.02                                  |\n| mrrmse3             | 0.05                               | 0.27                                   |\n| mrrmse1             | -0.05                              | -0.02                                  |\n| corr col2           | 0.05                               | 0.02                                   |\n| mrrmse mean         | 0.04                               | 0.15                                   |\n| mrrmse2             | 0.02                               | -0.01                                  |\n| corr col3           | 0.01                               | -0.35                                  |\n| mrrmse0             | -0.00                              | 0.07                                   |\n\n\n## MRRMSE-corrected ( Antoine Passemiers and Jalil Nourisa )\n\nSee their writeup: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461159 \nAnd section \"MRRMSE-corrected (Antoine Passemiers and Jalil Nourisa )\" here below.\n\nTables for AmbrosM CV\nhttps://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team?scriptVersionId=156200706&cellId=19\n\n\n| Metric               | Pearson correlation LB to CV: AmbrosM | Spearman correlation LB to CV: AmbrosM | Fold Info         |\n|----------------------|--------------------------------------|----------------------------------------|-------------------|\n| mrrmse corrected threshold = 1    | 0.23                                 | 0.39                                   | mean over folds   |\n| mrrmse corrected threshold = 2    | 0.22                                 | 0.40                                   | mean over folds   |\n| mrrmse corrected threshold = 3    | 0.16                                 | 0.29                                   | mean over folds   |\n| mrrmse standard           | 0.11                               | 0.04                                   | mean over folds   |\n\n\nIf we look by folds, than suprisingly T-cells CD4+ cells arise as best correlated to LB fold, NK-cells are the next (we have already seen NK cells above). \n\n\n| Metric                | Pearson correlation LB to CV: AmbrosM | Spearman correlation LB to CV: AmbrosM | Fold Info     |\n|-----------------------|--------------------------------------|----------------------------------------|---------------|\n| mrrmse corrected threshold = 1 | 0.32                                 | 0.49                                   | T cells CD4+  |\n| mrrmse corrected threshold = 1 | 0.28                                 | 0.39                                   | NK cells      |\n| mrrmse corrected threshold = 2 | 0.25                                 | 0.34                                   | T cells CD4+  |\n| mrrmse standard               | 0.25                                 | 0.36                                   | NK cells      |\n\n\nFor Random folds correlations are quite low, but still MRRMSE corrected is better:\nhttps://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team?scriptVersionId=156201219&cellId=19\n\n| Metric               | Pearson correlation LB to CV: Random | Spearman correlation LB to CV: Random |\n|----------------------|-------------------------------------|---------------------------------------|\n| mrrmse corrected threshold = 1    | 0.11                                | 0.29                                  |\n| mrrmse corrected threshold = 2    | 0.09                                | 0.27                                  |\n| mrrmse corrected threshold = 3    | 0.09                                | 0.19                                  |\n| mrrmse mean           | 0.04                                | 0.15                                  |\n\n\nFor MT-CV situation is different(!) - standard MRRMSE is better - surprisingly ! \nFor average over folds - all the  correlations are near zero. But if look by folds then fold-0 with drugs: \"['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189', 'Linagliptin', 'O-Demethylated Adapalene']\" gives better correlation - for all MRRMSE and MRRMSE-corrected for all thresholds\n\nhttps://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team?scriptVersionId=156202049&cellId=21\n\n\n| Metric                        | Pearson correlation LB to CV: MT | Spearman correlation LB to CV: MT | Fold Info                                                                                   |\n|-------------------------------|---------------------------------|------------------------------------|-----------------------------------------------------------------------------------------------|\n| mrrmse standard                       | 0.28                            | 0.38                               | ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189', 'Linagliptin', 'O-Demethylated Adapalene'] |\n| mrrmse corrected threshold = 1         | 0.23                            | 0.30                               | ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189', 'Linagliptin', 'O-Demethylated Adapalene'] |\n| mrrmse corrected threshold = 2         | 0.22                            | 0.28                               | ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189', 'Linagliptin', 'O-Demethylated Adapalene'] |\n| mrrmse corrected threshold = 3         | 0.21                            | 0.29                               | ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189', 'Linagliptin', 'O-Demethylated Adapalene'] |\n| mrrmse standard                   | 0.16                            | 0.05                               | mean over folds                                                                             |\n| mrrmse corrected threshold = 1     | 0.12                            | 0.03                               | mean over folds                                                                             |\n\n\n\nThe presentation with comments is here: \nhttps://docs.google.com/presentation/d/1wiz0Wmt4D54pqMMsIOyJHuQYMZ3hTBZQQnjbLzwoGYY/edit?usp=sharing\n\n\nSome comments on the collection are here: \nhttps://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing\n    ","metadata":{}},{"cell_type":"markdown","source":"# Preparations\n\n## Imports\n## Train data load \n## CV schemes class","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport time\nt0start = time.time() \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport pandas as pd\nimport tensorflow as tf\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nc = 0\nprint('Fist 30 files:')\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        c+=1\n        if c<=30:\n            #print(os.path.join(dirname, filename))\n            print(filename) # os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-12-23T16:25:43.388011Z","iopub.execute_input":"2023-12-23T16:25:43.388454Z","iopub.status.idle":"2023-12-23T16:25:43.480141Z","shell.execute_reply.started":"2023-12-23T16:25:43.388423Z","shell.execute_reply":"2023-12-23T16:25:43.478918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Data","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\ndf_de_train = pd.read_parquet(fn)# , index_col = 0)\nprint(df_de_train.shape)\ndisplay(df_de_train )\n\nplt.figure(figsize = (20,4) )\nv = df_de_train.iloc[:,5:].max(axis = 0 ).sort_values(ascending = False, key = abs )\nplt.plot(v.values,'*-')\nplt.title('Max-abs DE for genes',fontsize = 20 )\nplt.grid()\nplt.show()\ndisplay(v.head(15))\n\n# %%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/id_map.csv'\ndf_id_map = pd.read_csv(fn,index_col = 0)\nprint(df_id_map.shape)\ndisplay(df_id_map)\nfn = '/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv'\ndf_sample_submit = pd.read_csv(fn, index_col = 0)\nprint(df_sample_submit.shape)\ndisplay( df_sample_submit )","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-12-23T16:25:43.481870Z","iopub.execute_input":"2023-12-23T16:25:43.483019Z","iopub.status.idle":"2023-12-23T16:25:48.998506Z","shell.execute_reply.started":"2023-12-23T16:25:43.482983Z","shell.execute_reply":"2023-12-23T16:25:48.997209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##  Public LB drugs","metadata":{}},{"cell_type":"code","source":"list_ids_public = [\n    'LSM-43216', 'LSM-1050', 'LSM-45849', 'LSM-42800', 'LSM-1131', 'LSM-6335', 'LSM-1211',\n    'LSM-45239', 'LSM-1130', 'LSM-45786', 'LSM-5199', 'LSM-45281',\n    'LSM-6324', # 'ACY-1215' -> 'Ricolinostat'\n    'LSM-3309', 'LSM-1056', 'LSM-45591', 'LSM-46203', 'LSM-5662',\n    'LSM-47134',  # 'SB-2342' -> '5-(9-Isopropyl-8-methyl-2-morpholino-9H-purin-6-yl)pyrimidin-2-amine  '\n    'LSM-45637', 'LSM-1127', 'LSM-46971', 'LSM-1172', 'LSM-46042', 'LSM-1101', 'LSM-45758',\n    'LSM-5218', 'LSM-2287', 'LSM-1014',\n    'LSM-1040', #  'fostamatinib' -> 'Tamatinib'\n    'LSM-1476;LSM-5290',\n    'LSM-45680',  # 'basimglurant' -> 'RG7090'\n    'LSM-4349',  # '5-iodotubercidin' -> 'IN1451'\n    'LSM-3425', 'LSM-45806',\n    'LSM-45616',  # 'SB-683698' -> 'TR-14035'\n    'LSM-1055',\n    'LSM-43281',  # 'C-646' -> 'STK219801'\n    'LSM-5690', 'LSM-1155', 'LSM-2499',\n    'LSM-2382',  # 'JTC-801' -> 'UNII-BXU45ZH6LI'\n    'LSM-45220', 'LSM-1037', 'LSM-1005', 'LSM-1180', 'LSM-36812',\n    'LSM-45924',  # 'filgotinib' -> 'GLPG0634'\n    'LSM-2013',  # 'TL-HRAS-61' -> TL_HRAS26'\n    'LSM-4738'\n]\nprint(len(list_ids_public))\n\nIX_drugs_public_lb = [i for i in range(len(df_de_train)) if df_de_train['sm_lincs_id'].iat[i] in list_ids_public]\n# print( len( IX_drugs_public_lb ) )\n# print( IX_drugs_public_lb )\n\nm = df_de_train['sm_lincs_id'].isin(list_ids_public)\n# print(m.sum())\ndf_de_train[m]['sm_lincs_id'].nunique()\ndisplay( df_de_train[m]['sm_name'].unique() )\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-12-23T16:25:49.000023Z","iopub.execute_input":"2023-12-23T16:25:49.000405Z","iopub.status.idle":"2023-12-23T16:25:49.045067Z","shell.execute_reply.started":"2023-12-23T16:25:49.000374Z","shell.execute_reply":"2023-12-23T16:25:49.044144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## CV schemes\n\nHere is a class to work with several CV schemes used in OP2 challenge - AmbrosM, MT, random ","metadata":{}},{"cell_type":"code","source":"# Prepare indexes for AmbrosM scheme:\n# For AmbrosM scheme: \nfolds_index_data_AmbrosM = [ ]\ntrain_sm_names = ['Idelalisib', 'Crizotinib', 'Linagliptin', 'Palbociclib', 'Dabrafenib', 'Alvocidib', 'LDN 193189', 'R428', 'Porcn Inhibitor III', \n  'Belinostat', 'Foretinib', 'MLN 2238', 'Penfluridol', 'Dactolisib', 'O-Demethylated Adapalene', 'Oprozomib (ONX 0912)', 'CHIR-99021']\nlist_fold_ids =  ['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells'] \nfor fold_id in list_fold_ids:\n    mask_va = (df_de_train.cell_type == fold_id) & ~df_de_train.sm_name.isin(train_sm_names)\n    mask_tr = ~mask_va # 485 or 487 training rows\n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_AmbrosM.append( [IX_train, IX_test ])\n\n# For MT CV scheme:\nfolds_index_data_MT = []\nfold_to_compounds = {0: ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189',  'Linagliptin', 'O-Demethylated Adapalene'],\n 1: ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III'],\n 2: ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol',  'R428']}\nfor fold_id in [0,1,2]:\n    mask_va = df_de_train['cell_type'].isin(['Myeloid cells', 'B cells']) & df_de_train['sm_name'].isin(fold_to_compounds[fold_id])\n    mask_tr = ~mask_va    \n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_MT.append( [IX_train, IX_test ])    \n    \n    \ndict_folds_Antonina = {0: [4, 15, 18, 41, 47, 53, 56, 57, 59, 61, 62, 66, 78, 83, 86, 93, 100, 110, 115, 116, 117, 120, 123, 131, 148, 152, 161, 185, 189, 193, 197, 198, 205, 207, 209, 210, 213, 218, 219, 225, 228, 230, 250, 252, 257, 263, 286, 292, 294, 304, 306, 310, 325, 327, 336, 338, 339, 350, 351, 356, 357, 360, 362, 369, 379, 385, 395, 419, 424, 430, 432, 434, 438, 441, 443, 445, 448, 452, 456, 465, 467, 471, 472, 474, 498, 508, 523, 534, 542, 544, 574, 581, 584, 589, 601, 607], 1: [1, 3, 5, 14, 24, 38, 40, 49, 50, 54, 55, 60, 79, 82, 87, 102, 112, 114, 118, 119, 127, 132, 146, 149, 166, 179, 183, 184, 186, 190, 192, 201, 203, 204, 216, 227, 242, 253, 261, 264, 271, 289, 295, 296, 299, 300, 309, 323, 324, 329, 333, 337, 354, 387, 390, 393, 406, 408, 415, 416, 420, 421, 422, 436, 442, 449, 453, 458, 462, 463, 469, 484, 487, 489, 490, 492, 497, 501, 505, 506, 509, 513, 515, 516, 546, 548, 566, 569, 585, 587, 588, 592, 593, 595, 600, 605], 2: [0, 16, 19, 44, 48, 51, 52, 58, 63, 65, 70, 81, 85, 88, 89, 92, 111, 113, 122, 125, 135, 137, 139, 150, 164, 175, 191, 202, 206, 211, 217, 220, 221, 229, 239, 243, 247, 248, 251, 258, 259, 260, 267, 270, 274, 281, 288, 290, 291, 303, 322, 326, 328, 332, 347, 359, 364, 370, 383, 384, 389, 392, 401, 423, 428, 454, 455, 461, 464, 494, 496, 504, 510, 511, 524, 525, 526, 530, 540, 541, 551, 565, 568, 570, 572, 573, 575, 577, 579, 580, 597, 598, 603, 606, 610, 611, 613], 3: [21, 22, 26, 36, 42, 45, 46, 64, 67, 69, 71, 80, 84, 90, 101, 103, 134, 136, 140, 147, 151, 160, 163, 165, 167, 174, 176, 177, 180, 181, 182, 187, 188, 195, 196, 199, 200, 212, 215, 226, 231, 240, 241, 249, 255, 256, 262, 265, 269, 285, 287, 297, 298, 307, 312, 319, 321, 330, 340, 342, 344, 345, 348, 358, 366, 368, 382, 391, 400, 403, 404, 405, 425, 427, 431, 433, 435, 444, 450, 451, 457, 459, 460, 473, 485, 499, 500, 502, 503, 507, 512, 514, 527, 528, 529, 531, 543, 547, 553, 554, 564, 571, 576, 583, 586, 591, 596, 602, 604, 608, 609], 4: [2, 6, 7, 17, 20, 23, 25, 27, 28, 29, 37, 39, 43, 68, 91, 121, 124, 126, 128, 129, 130, 133, 138, 153, 162, 178, 194, 208, 214, 222, 223, 224, 232, 244, 245, 246, 254, 266, 268, 272, 273, 282, 283, 284, 293, 301, 302, 305, 308, 311, 320, 331, 334, 335, 341, 343, 346, 349, 352, 353, 355, 361, 363, 365, 367, 377, 378, 380, 381, 386, 388, 394, 396, 397, 398, 399, 402, 407, 417, 418, 426, 429, 437, 439, 440, 446, 447, 466, 468, 470, 481, 482, 483, 486, 488, 491, 493, 495, 532, 533, 545, 549, 550, 552, 555, 562, 563, 567, 578, 582, 590, 594, 599, 612], 5: [8, 9, 10, 11, 12, 13, 30, 31, 32, 33, 34, 35, 72, 73, 74, 75, 76, 77, 94, 95, 96, 97, 98, 99, 104, 105, 106, 107, 108, 109, 141, 142, 143, 144, 145, 154, 155, 156, 157, 158, 159, 168, 169, 170, 171, 172, 173, 233, 234, 235, 236, 237, 238, 275, 276, 277, 278, 279, 280, 313, 314, 315, 316, 317, 318, 371, 372, 373, 374, 375, 376, 409, 410, 411, 412, 413, 414, 475, 476, 477, 478, 479, 480, 517, 518, 519, 520, 521, 522, 535, 536, 537, 538, 539, 556, 557, 558, 559, 560, 561]}\n\nfolds_index_data_Antonina = []    \nfor i in range(5):  \n    l_valid = np.array( dict_folds_Antonina[i] )\n    l_train = np.array( [k for k in range(614) if k not in l_valid ] )\n    folds_index_data_Antonina.append( [l_train, l_valid ])    \nprint( len(folds_index_data_Antonina ) )\n    \n    \nimport random \nfrom sklearn.model_selection import KFold\n\nclass KFold_custom:\n    '''\n    Class with similar to sklearn \"KFold\" class, interface.\n    Support of CV schemes proposed by AmbrosM, MT, Kishan and standard sklearn KFold\n    Supports options to return IX_train,IX_valid,IX_test - triplet - if valid_size > 0\n    Examples:\n    kf = KFold_custom('AmbrosM'):     \n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n        pass\n    kf = KFold_custom('AmbrosM', valid_size = 0.1 )     \n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split()) :\n        pass\n        \n    kf = KFold_custom(CV_scheme = 'Random') # wrapper for sklearn Kfolds with shuffle = True\n    kf = KFold_custom(CV_scheme = CV_scheme, random_state = 42 ) # fixing random_state ensures reprodicibility of folds \n    '''\n    def __init__(self, CV_scheme,  valid_size = 0, n_splits=5, random_state = 42, verbose = 0 ):\n        self.CV_scheme = CV_scheme\n        self.CV_scheme_inf =  CV_scheme\n        self.valid_size = np.clip(valid_size,0,1)\n        self.random_state = random_state\n\n        self.folds_index_data = [ [np.arange(0), np.arange(0), np.arange(0)] ] # Example data\n        self.n_splits = 1 # Example data\n        if CV_scheme == 'Kishan1':\n            # One fold scheme which contains train, valid, test parts \n            IX_train_Kishan1 = np.arange(429)\n            IX_val_Kishan1 = np.arange(429,521)\n            IX_test_Kishan1 = np.arange(521,614)\n            self.n_splits = 1\n            self.folds_index_data =  [   [IX_train_Kishan1, IX_val_Kishan1, IX_test_Kishan1]  ]\n        elif CV_scheme == 'Tonya':\n            self.n_splits = 5\n            self.folds_index_data = folds_index_data_Antonina\n        elif CV_scheme == 'AmbrosM':\n            self.n_splits = 4\n            self.folds_index_data = folds_index_data_AmbrosM\n        elif CV_scheme == 'MT':\n            self.n_splits = 3\n            self.folds_index_data = folds_index_data_MT\n        elif 'Random'.lower() in CV_scheme.lower():\n            self.n_splits = n_splits\n            self.CV_scheme_inf = 'Random_'+str(n_splits) +'_'+str(random_state )\n            kf = KFold(n_splits=n_splits, random_state = random_state, shuffle=True )#,\n            self.folds_index_data = list( kf.split( np.arange(614) ) )\n        elif 'Full'.lower() in CV_scheme.lower():\n            # Just return the full set as both train and test - it useful to re-train the model on the entire data\n            IX_train_full = np.arange(614)\n            IX_test_full = np.arange(614)\n            self.n_splits = 1\n            self.folds_index_data = [   [IX_train_full, IX_test_full]  ]\n        else:\n            s = 'Uncrecognized CV_scheme ' + str(CV_scheme) \n            raise ValueError(s)\n\n        self.verbose = verbose \n        if verbose >= 10:\n            print(self.CV_scheme, self.n_splits, len(self.folds_index_data) )\n            #print(self.folds_index_data)\n            \n    def split(self, X=None):\n        '''\n        X - NOT used, just for compatibility with sklearn \n        '''\n        for item in self.folds_index_data:\n            if self.CV_scheme in ['Kishan1']:\n                if self.valid_size > 0:\n                    yield item[0],item[1],item[2]\n                else :\n                    yield np.array( list(set(item[0])|set(item[1]))  ) , item[2] # Return train, test only \n            elif self.CV_scheme in ['Tonya', 'AmbrosM','MT', 'Random', 'Full']:\n                if self.valid_size == 0:\n                    yield item[0],item[1]\n                else:\n                    # Split \"full-train\"->( real-train, valid ) \n                    index_for_real_train_part = int(  (1-self.valid_size) * len(item[0]) )\n                    if self.random_state is None:\n                        IX_train =  item[0][:index_for_real_train_part]\n                        IX_valid =  item[0][index_for_real_train_part:]\n                    elif self.random_state == -1:\n                        p = np.random.permutation(len(item[0])) \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    else:\n                        np.random.seed(self.random_state) # Set temporary random seed\n                        p = np.random.permutation(len(item[0])) \n                        np.random.seed(random.randint(0,30000)) # Randomize seed again - use Python random, not numpy       \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    yield IX_train, IX_valid, item[1]\n                    \n        \n    def get_list_main_CV_schemes(self):\n        return ['AmbrosM','MT','Kishan1','Random']\n    def get_n_splits(self, X=np.array([])):\n        return self.n_splits\n        \n    \nprint();print();    \nprint('--------------------------------Examples-----------------------------------------------')\nprint();print();    \n        \nCV_scheme = 'Kishan1'\nprint(CV_scheme,  )\nkf = KFold_custom('Kishan1', valid_size = 0.18)     \n#list( kf.split() )\nprint(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test)  \" )\nfor i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split(None)) :\n    print(i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test) )\nprint()\n\nfor CV_scheme in ['Tonya','AmbrosM','MT','Random']:\n    valid_size = 0\n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size)     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train),type(IX_test)  \" )\n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train), len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train), type(IX_test) )\n    \n    print()\n    valid_size = 0.2 \n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size )     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test)  \" )\n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test) )\n    print()    \n    ","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-12-23T16:25:49.047397Z","iopub.execute_input":"2023-12-23T16:25:49.048070Z","iopub.status.idle":"2023-12-23T16:25:49.127520Z","shell.execute_reply.started":"2023-12-23T16:25:49.048036Z","shell.execute_reply":"2023-12-23T16:25:49.126511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Test example load OOF  and score it","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations-submitsetc/LB0608_0607_tsvd50_SVR_target_enc_i_th_target_BothEncoded_10-17-01-33_V202/oof_AmbrosM_CV09661.npy'\noof = np.load(fn)\noof\nprint(oof.shape)\n\n# for CV_scheme in ['Tonya','AmbrosM','MT','Random']:\nif 1:    \n    CV_scheme = 'AmbrosM'\n    valid_size = 0\n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size)     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train),type(IX_test)  \" )\n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train), len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train), type(IX_test) )\n        Y_pred = oof[IX_test]\n        Y_true = df_de_train.iloc[:,5:].values[IX_test]\n        mrrmse = np.sqrt(np.square(Y_true - Y_pred).mean(axis=1)).mean();  \n        print(mrrmse)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-12-23T16:25:49.128665Z","iopub.execute_input":"2023-12-23T16:25:49.130094Z","iopub.status.idle":"2023-12-23T16:25:49.501849Z","shell.execute_reply.started":"2023-12-23T16:25:49.130054Z","shell.execute_reply":"2023-12-23T16:25:49.499505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Core load/compute - load many saved OOF predictions and compute CV scores\n\nLoad stored predictions and compute CV/OOF scores. \n\nIn the dataset:\nhttps://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc/\nwe stored more than 30 submission (with their public LB scores) for OP2 challenge - jointly with their OOF-predictions for several CV schemes (AmbrosM, MT, random).\nSo using these OOF we can calculate the CV scores and compare with saved LB scores.\n","metadata":{}},{"cell_type":"code","source":"%%time\nmode_use_public_or_private_or_all_drugs = 'ALL' #  'not_public_drugs' #  'public_drugs'\n\nCV_scheme = 'AmbrosM' # 'MT'# 'MT' #      'Random' #   \ndf_stat = pd.DataFrame(); \ncc = 0\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        if 'open-problems-single-cell-perturbations-submitsetc' in dirname:\n            if 'oof_'+CV_scheme in filename:\n            #if 1:# ('LB0608' in dirname):\n            #if  ('Random' in filename) and ('blend' not in filename) and ('.csv' in filename):\n            # if  ('simple' in filename) and ('blend' not in filename) and ('.csv' in filename):\n            #   ('Random' in filename) and ('blend' not in filename) and ('.csv' in filename):\n            #if  ('Random' in filename) and ('priors' in filename) and ('.csv' in filename):\n                cc += 1 \n                f = os.path.join(dirname, filename)\n                #list_fn.append( f )\n                #list_ids.append(f.split('/')[-2][:30]) \n                IX = len(df_stat)+1\n                df_stat.loc[IX,'LB'] = float( f.split('/')[-2][:30].split('_')[0][3:] )/1000\n                if CV_scheme != 'Random':\n                    df_stat.loc[IX,'CV'] = float( filename.split('.')[0].split('_')[2][2:] ) / 10000\n                else:\n                    df_stat.loc[IX,'CV'] = float( filename.split('.')[0].split('_')[-1][2:] ) / 10000\n                    \n                oof = np.load(f)\n                #print(oof.shape)\n                print(cc, f )\n                kf = KFold_custom(CV_scheme)#, valid_size = valid_size )  \n\n                #print(\" i_fold, len(IX_train),len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train),type(IX_test)  \" )\n                list_mrrmse = []; list_mrrmse_corrected0 = []; list_mrrmse_corrected1 = []; list_mrrmse_corrected2 = []; \n                for i_fold,(IX_train,IX_test)   in enumerate( kf.split(None)) :\n                    #print(i_fold, len(IX_train), len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train), type(IX_test) )\n\n                    \n                    if mode_use_public_or_private_or_all_drugs == 'ALL':\n                        pass\n                    elif mode_use_public_or_private_or_all_drugs == 'not_public_drugs':\n                        IX_test = np.array( list( set(IX_test) - set(IX_drugs_public_lb) ) )\n                    elif mode_use_public_or_private_or_all_drugs == 'public_drugs':\n                        IX_test = np.array( list( set(IX_test) & set(IX_drugs_public_lb) ) )\n                    \n                    Y_pred = oof[IX_test]\n                    Y_true = df_de_train.iloc[:,5:].values[IX_test]\n                    mrrmse = np.sqrt(np.square(Y_true - Y_pred).mean(axis=1)).mean();  list_mrrmse.append( mrrmse)\n                    list_corr_rows_fold = [np.corrcoef(Y_true[i,:], Y_pred[i,:])[0,1] for i in range(Y_pred.shape[0])]\n                    list_corr_columns_fold = [np.corrcoef(Y_true[:,i], Y_pred[:,i])[0,1] for i in range(Y_pred.shape[1])]\n                    #print(i_fold,mrrmse)\n                    df_stat.loc[IX,'mrrmse' + str(i_fold)] = mrrmse\n                    thres_Y0 = 1\n                    tmp = (Y_true - Y_pred) * (np.abs(Y_true)>thres_Y0).astype( float )\n                    mrrmse = np.sqrt(np.square(tmp).mean(axis=1)).mean();  list_mrrmse_corrected0.append(mrrmse)\n                    df_stat.loc[IX,'mrrmse corrected'+str(thres_Y0) +' ' + str(i_fold)] = mrrmse\n                    thres_Y1 = 2\n                    tmp = (Y_true - Y_pred) * (np.abs(Y_true)>thres_Y1).astype( float )\n                    mrrmse = np.sqrt(np.square(tmp).mean(axis=1)).mean();  list_mrrmse_corrected1.append(mrrmse)\n                    df_stat.loc[IX,'mrrmse corrected'+str(thres_Y1) +' ' + str(i_fold)] = mrrmse\n                    thres_Y2 = 3\n                    tmp = (Y_true - Y_pred) * (np.abs(Y_true)>thres_Y2).astype( float )\n                    mrrmse = np.sqrt(np.square(tmp).mean(axis=1)).mean();  list_mrrmse_corrected2.append(mrrmse)\n                    df_stat.loc[IX,'mrrmse corrected'+str(thres_Y2) +' ' + str(i_fold)] = mrrmse\n\n                    df_stat.loc[IX,'corr row' + str(i_fold)] = np.mean(list_corr_rows_fold)\n                    df_stat.loc[IX,'corr col' + str(i_fold)] = np.mean(list_corr_columns_fold)\n                    \n                df_stat.loc[IX,'mrrmse mean'] = np.mean(list_mrrmse)     \n                df_stat.loc[IX,'mrrmse corrected'+str(thres_Y0)] = np.mean(list_mrrmse_corrected0)    \n                df_stat.loc[IX,'mrrmse corrected'+str(thres_Y1)] = np.mean(list_mrrmse_corrected1)    \n                df_stat.loc[IX,'mrrmse corrected'+str(thres_Y2)] = np.mean(list_mrrmse_corrected2)    \n                df_stat.loc[IX,'corr row mean'] = np.mean( [df_stat.loc[IX,:].iat[i] for i in range(df_stat.shape[1]) if  ('corr row' in df_stat.columns[i]) and ('corr row mean' not in df_stat.columns[i])   ] )     \n                df_stat.loc[IX,'corr col mean'] = np.mean( [df_stat.loc[IX,:].iat[i] for i in range(df_stat.shape[1]) if  ('corr col' in df_stat.columns[i]) and ('corr col mean' not in df_stat.columns[i])   ] )     \n                \n                df_stat.loc[IX,'Id'] = f.split('/')[-2][:50]\n                df_stat.loc[IX,'fn'] = filename.split('.')[0]\n                print(cc,'mrrmse:', np.round( np.mean(list_mrrmse)    ,4), 'mrrmse by folds:',   str(np.round(list_mrrmse,4)) )\n                if cc<5:\n                    display(df_stat.tail(1))    \n                \n                \ndf_stat.to_csv('df_stat.csv')\ndf_stat \n","metadata":{"execution":{"iopub.status.busy":"2023-12-23T17:22:19.275071Z","iopub.execute_input":"2023-12-23T17:22:19.275525Z","iopub.status.idle":"2023-12-23T17:28:01.215998Z","shell.execute_reply.started":"2023-12-23T17:22:19.275493Z","shell.execute_reply":"2023-12-23T17:28:01.214927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Main analysis - LB vs CV correlation for different metrics","metadata":{}},{"cell_type":"code","source":"# N = 3+ len(list_mrrmse)                \n#print('Pearson:')        \ncm1 = df_stat.iloc[:,:-2].corr().round(2)\ncm2 = df_stat.iloc[:,:-2].corr(method = 'spearman').round(2)\n# print('Spearman:')                \n# display( cm2)                \n\nd = pd.DataFrame()\nd['Pearson corellation LB to CV: ' + CV_scheme ] = cm1['LB']\nd['Spearman corellation LB to CV: ' + CV_scheme] = cm2['LB']\nif CV_scheme == 'AmbrosM':\n    list_fold_ids =  ['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells'] \n    for k in d.index:\n        if k[-1].isdigit():\n            if not( (k[-2] != ' ') and ('mrrmse corrected' in k ) ) :  \n                # print(k)\n                d.loc[k, 'Fold Info' ]  = list_fold_ids[ int(k[-1] )]\n    # l = [ list_do]\n    d['Fold Info'] = d['Fold Info'].fillna('mean over folds')\nelif CV_scheme == 'MT':\n    fold_to_compounds = {0: ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189',  'Linagliptin', 'O-Demethylated Adapalene'],\n     1: ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III'],\n     2: ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol',  'R428']}\n    for k in d.index:\n        if k[-1].isdigit(): # \n            if not( (k[-2] != ' ') and ('mrrmse corrected' in k ) ) :  \n        #         print(k)\n                d.loc[k, 'Fold Info' ]  = str(fold_to_compounds[ int(k[-1] )])\n    # l = [ list_do]\n    d['Fold Info'] = d['Fold Info'].fillna('mean over folds')\n\n\n\nd2 = d.sort_values(d.columns[0], key = abs, ascending = False) \npd.set_option('display.max_colwidth',100)\npd.set_option('display.max_rows', None)\nd2.to_csv('df_LB_to_CV_corrs_for_different_metrics.csv')\nprint('Sorted by Pearson correlation:')\ndisplay( d2 )                \nprint('Sorted by Spearman correlation:')\nd2b = d.sort_values(d.columns[1], key = abs, ascending = False) \ndisplay( d2b )                \n\ndisplay( d2.iloc[:,:2].corr() )                \n","metadata":{"execution":{"iopub.status.busy":"2023-12-23T17:29:36.601196Z","iopub.execute_input":"2023-12-23T17:29:36.601637Z","iopub.status.idle":"2023-12-23T17:29:36.661317Z","shell.execute_reply.started":"2023-12-23T17:29:36.601608Z","shell.execute_reply":"2023-12-23T17:29:36.659957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"v1 = d2.iloc[:,0]\nv2 = d2.iloc[:,1]\nsns.scatterplot(x=v1,y=v2)\nplt.title('Correlations to LB for different metrics',fontsize = 20 )\nplt.xlabel('Pearson correlation to LB', fontsize = 20)\nplt.ylabel('Spearman correlation to LB', fontsize = 20)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-12-23T17:29:37.696565Z","iopub.execute_input":"2023-12-23T17:29:37.697261Z","iopub.status.idle":"2023-12-23T17:29:37.959572Z","shell.execute_reply.started":"2023-12-23T17:29:37.697225Z","shell.execute_reply":"2023-12-23T17:29:37.958376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# MRRMSE-corrected (Antoine Passemiers and Jalil Nourisa )\n\nI.e. only those Y_true which have value greater than threshold are takeing into account, for others error set to zero. \n\n\nSee  MRRMSE corrected in  (Antoine Passemiers and Jalil Nourisa writeup): https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461159\n","metadata":{}},{"cell_type":"markdown","source":"## Full train (not splitted by folds)","metadata":{}},{"cell_type":"code","source":"l = ['mrrmse corrected1','mrrmse corrected2','mrrmse corrected3', 'mrrmse mean']\nm = d2.index.isin(l)\nd2[m]","metadata":{"execution":{"iopub.status.busy":"2023-12-23T17:29:42.098763Z","iopub.execute_input":"2023-12-23T17:29:42.099176Z","iopub.status.idle":"2023-12-23T17:29:42.112129Z","shell.execute_reply.started":"2023-12-23T17:29:42.099142Z","shell.execute_reply":"2023-12-23T17:29:42.110927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Including each fold metrics ","metadata":{}},{"cell_type":"code","source":"m = ['mrrmse' in t for t in d2.index] \n# m = d2.index.isin(l)\nprint('Sorted by Pearson correlation to LB')\ndisplay(d2[m])\nprint('Sorted by Spearman correlation to LB')\ndisplay(d2[m].sort_values(d2.columns[1], key = abs, ascending = False))\n\nv1 = d2[m].iloc[:,0]\nv2 = d2[m].iloc[:,1]\nsns.scatterplot(x=v1,y=v2)\nplt.title('Correlation to LB. Only MRRMSE corrected metrics',fontsize = 20 )\nplt.xlabel('Pearson correlation to LB', fontsize = 20)\nplt.ylabel('Spearman correlation to LB', fontsize = 20)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-12-23T17:15:07.033018Z","iopub.execute_input":"2023-12-23T17:15:07.033462Z","iopub.status.idle":"2023-12-23T17:15:07.314499Z","shell.execute_reply.started":"2023-12-23T17:15:07.033427Z","shell.execute_reply":"2023-12-23T17:15:07.313236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Scatteplots LB vs CV","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (20,8) )\nplt.suptitle('CV scheme: '+CV_scheme, fontsize = 20)\nplt.subplot(1,2,1)\nsns.scatterplot( x= df_stat['LB'], y = df_stat['corr row mean'] )\nplt.subplot(1,2,2)\nsns.scatterplot( x= df_stat['LB'], y = df_stat['corr col mean'] )\nplt.show()\n\nplt.figure(figsize = (20,8) )\nplt.suptitle('CV scheme: '+CV_scheme, fontsize = 20)\nplt.subplot(1,2,1)\nsns.scatterplot( x= df_stat['LB'], y = df_stat['mrrmse mean'] )\nplt.subplot(1,2,2)\nsns.scatterplot( x= df_stat['LB'], y = df_stat['mrrmse0'] )\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-12-23T16:31:59.247833Z","iopub.execute_input":"2023-12-23T16:31:59.248168Z","iopub.status.idle":"2023-12-23T16:32:00.301573Z","shell.execute_reply.started":"2023-12-23T16:31:59.248141Z","shell.execute_reply":"2023-12-23T16:32:00.300299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Final timing","metadata":{}},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )\nprint('%.1f minutes passed total '%( (time.time()-t0start)/60)  )\nprint('%.2f hours passed total '%( (time.time()-t0start)/3600)  )","metadata":{"execution":{"iopub.status.busy":"2023-12-23T16:32:00.303115Z","iopub.execute_input":"2023-12-23T16:32:00.303427Z","iopub.status.idle":"2023-12-23T16:32:00.309801Z","shell.execute_reply.started":"2023-12-23T16:32:00.303399Z","shell.execute_reply":"2023-12-23T16:32:00.308689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}