{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59094,"databundleVersionId":7010844,"sourceType":"competition"},{"sourceId":7210153,"sourceType":"datasetVersion","datasetId":3765688}],"dockerImageVersionId":30582,"isInternetEnabled":true,"language":"r","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Tricks and metrics for the [Open Problems – Single-Cell Perturbations](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations) competition \n\nIn this notebook, I compare my validation scheme (on test drugs) with several others in terms of how well they reflect the changes in LB score. For this, using [the MLP NN](https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-20th-place-solution) I ran several experiments with filtration of train data, adding the noise to features, and data augmentation, - tricks which improved or did not affect LB scores.","metadata":{}},{"cell_type":"code","source":"library(data.table)\nlibrary(ggplot2)\nlibrary(patchwork)\nlibrary(DT)\nlibrary(ggrepel)\n\nshape <- function(dt) {\n    cat(\"\\nShape: \",\n      nrow(dt), \" rows x \",\n      ncol(dt), \" columns\",\n      sep = \"\")\n    return(head(dt))  \n}\nfig <- function(width, heigth){\n  options(repr.plot.width = width, repr.plot.height = heigth)\n}","metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","execution":{"iopub.status.busy":"2023-12-16T05:36:21.889677Z","iopub.execute_input":"2023-12-16T05:36:21.891539Z","iopub.status.idle":"2023-12-16T05:36:22.482243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"border: 3px solid #3B2F2F; border-radius: 10px; padding: 15px; background-color: #FFCC99; text-align: center;\">Load data</div>","metadata":{}},{"cell_type":"markdown","source":"**The metrics of the experiments with noise, filtration and augmentation**","metadata":{}},{"cell_type":"code","source":"all_data <- fread(\"/kaggle/input/op2-supplementary-calcs-for-ml/MLP_experiments/MLP_metrics_aug_noise_filt.csv\")\n\n# all_data[, unique(features)]\nall_data[, augm := ifelse(grepl(\"augm 1\", features), 1, 50)]\nall_data[, noise := ifelse(grepl(\"s0.1/0.1\", features), \"0.1/0.1\", \"1/0.1\")]\n\nall_data[, epochs := ifelse(use_valid == FALSE,\n                            mean_num_epochs,\n                            \"early_stop\")]\nall_data <- all_data[epochs %in% c(\"early_stop\", \"15\")]\n\nall_data[, `:=` (validation = paste(cv, \", \", epochs),\n                 cv = NULL, epochs = NULL, mean_num_epochs = NULL)]\nrem_cols <- c('n_groups', \"train_prop\", 'i', 'fold', 'n_PC', 'features',\n              'd_hid1', 'wd', 'bs', 'max_lr', 'pt', 'time, s', 'use_valid')\nall_data <- all_data[, !rem_cols, with = FALSE]\n\nshape(all_data)","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:32.947966Z","iopub.execute_input":"2023-12-16T05:36:32.950039Z","iopub.status.idle":"2023-12-16T05:36:33.022527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Validation folds**","metadata":{}},{"cell_type":"code","source":"all_data[, unique(gsub(\"\\\\, .*\", \"\", validation))]","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:33.026677Z","iopub.execute_input":"2023-12-16T05:36:33.028330Z","iopub.status.idle":"2023-12-16T05:36:33.054038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- on test drugs: 5-fold cross validation on test drugs. The training part of each fold contains all available data from the test cells (myeloid and B), while the test part containes only the 129 drugs for which we have to make predictions.\n- on test drugs, 1ct: 4-fold cross validation on test drugs and one cell type. The training part of each fold contains all available data from the test cells (myeloid and B), while the test part of each fold containes only one cell type (NK cells, T cells CD4+, T cells CD8+ or T regulatory cells) and the 129 drugs for which we have to make predictions.\n- rand: 5-fold cross validation with randomly splitted data (train/test = ~80/20)\n- on NK cells: 1 fold with validation on 129 sampless with test drugs on NK cells\n- 1st MT fold: 1 fold with validation on 12 samples with some test drugs ('Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189',  'Linagliptin', 'O-Demethylated Adapalene') and Myeloid and B cells\n- Unnamed fold: 1 fold with validation on 22 samples with some test drugs ('Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III', 'CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol',  'R428') and Myeloid and B cells\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"**LB scores**","metadata":{}},{"cell_type":"markdown","source":"- MLP version 12 (augm 1, s0.1/0.1, 614 samples): public LB 0.606\n- MLP version 32 (augm 50, s0.1/0.1, 614 samples): public LB 0.582\n- MLP version 39 (augm 50, s0.1/0.1. 603 samples): public LB 0.581\n- MLP versions 49, 121, 164, 154 (augm 50, s1/0.1, 603 samples): public LB 0.579-0.583\n\nSo in LB we can wee the notable score improvements for only train augmentation, the effects of noise and cell filtration are unclear","metadata":{}},{"cell_type":"markdown","source":"**Some functions**","metadata":{}},{"cell_type":"code","source":"metrics_cols <- c(\"mrRMSE\", \"RMSE\", \"MAE\", \"r2\",\n          \"mean_corr_rows\", \"mean_corr_cols\",\n          \"median_corr_rows\", \"median_corr_cols\")\n\nall_data <- all_data[order(augm, -n_train)]\n\ndelta <- function (x) {    \n    delt <- round(max(x) - min(x), 3)\n    #paste0(delt, \" (n = \", length(x), \")\")\n    return(delt)\n    }","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:33.057533Z","iopub.execute_input":"2023-12-16T05:36:33.059166Z","iopub.status.idle":"2023-12-16T05:36:33.081935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"exp_cols <- c(\"validation\", \"augm\",\"noise\",\"n_train\")","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:33.084769Z","iopub.execute_input":"2023-12-16T05:36:33.086230Z","iopub.status.idle":"2023-12-16T05:36:33.099091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"border: 3px solid #3B2F2F; border-radius: 10px; padding: 15px; background-color: #FFCC99; text-align: center;\">Metrics variability by different model parameters</div>","metadata":{}},{"cell_type":"markdown","source":"Variability (max-min) of the metrics obtained after 3-4 runs of the same model with different parameters (with or without augmentation, different noise to features, with or without data filtering)","metadata":{}},{"cell_type":"code","source":"dat <- all_data[, n_obs := .N, by = exp_cols][n_obs > 2][, n_obs := NULL]\ndat[, `:=` (validation = paste(validation, \"(n_Yoof = \", n_Y_oof, \")\"),\n            n_Y_oof = NULL)]\nfuns = c('delta')\nres <- dat[, lapply(.SD, function(col) {\n    \n    sapply(funs, function(f) do.call(f, list(col)))\n           \n     }), .SDcols = metrics_cols, by = exp_cols][\n        order(augm, noise, n_train)\n    ]\nres[order(validation, augm, noise)]","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:33.102039Z","iopub.execute_input":"2023-12-16T05:36:33.104790Z","iopub.status.idle":"2023-12-16T05:36:33.195235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"res <- melt(res, measure.vars = metrics_cols,\n           variable.name = \"metric\")\nres[, augm := factor(augm)]","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:33.197723Z","iopub.execute_input":"2023-12-16T05:36:33.199003Z","iopub.status.idle":"2023-12-16T05:36:33.218723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(30, 30)\n    \nggplot(data = res, aes(x = validation, y = value, fill = augm)) +\n    geom_bar(stat = \"identity\") +\n    scale_fill_manual(values=c('#C9E5C4', '#FFCC99')) +\n    theme_bw(base_size = 18) +\n    xlab(\"\") +\n    facet_wrap(~ metric, ncol = 2, scales = \"free\") +\n    theme(axis.text.x = element_text(angle = 45, hjust=1),\n         legend.position = \"bottom\")","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:33.224147Z","iopub.execute_input":"2023-12-16T05:36:33.225789Z","iopub.status.idle":"2023-12-16T05:36:35.927787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"line-height:24px; font-size:16px;border-left: 5px solid silver; padding-left: 26px;\"> \n    💡 Note: \n    <ul style=\"list-style:circle\">\n<li>We can see the all metrics for MT fold are highly variable, thus unreliable for the detection of the improvement of model performance\n<li>When the data are available for both experiments (green and orange), we see that the metrics are less variable and more robust for those with train augmentation (orange)\n    </ul>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"border: 3px solid #3B2F2F; border-radius: 10px; padding: 15px; background-color: #FFCC99; text-align: center;\">Metrics delta for each CV scheme</div>","metadata":{}},{"cell_type":"markdown","source":"Now we average mertics for each experiment.","metadata":{}},{"cell_type":"code","source":"dat <- all_data[, !\"n_obs\"]\n\nsuppressWarnings(dat <- unique(\n    dat[, (metrics_cols) := lapply(.SD, mean, na.rm = TRUE),\n        .SDcols = metrics_cols,\n        by = exp_cols]\n))\ndat[ , (metrics_cols) := lapply(.SD, round, 3), .SDcols = metrics_cols]\n\nhead(dat[order(validation)])","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:35.930764Z","iopub.execute_input":"2023-12-16T05:36:35.932477Z","iopub.status.idle":"2023-12-16T05:36:35.974014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filt <- NULL\naugm <- NULL\nnoise <- NULL\n\nfor(v in unique(all_data$validation)) {\n    \n    res <- dat[validation == v, lapply(.SD, diff),\n           by = c(\"validation\", \"augm\", \"noise\"),\n           .SD = metrics_cols]\n    filt <- rbind(filt, res)\n\n    res <- dat[validation == v, lapply(.SD, diff),\n               .SD = metrics_cols, by = c(\"validation\", \"n_train\", \"noise\")]\n    augm <- rbind(augm, res)\n\n    res <- dat[validation == v, lapply(.SD, diff),\n               .SD = metrics_cols, by = c(\"validation\", \"n_train\", \"augm\")]\n    noise <- rbind(noise, res)\n    \n}","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:35.977302Z","iopub.execute_input":"2023-12-16T05:36:35.978544Z","iopub.status.idle":"2023-12-16T05:36:36.106050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"border: 3px solid #3B2F2F; border-radius: 10px; padding: 15px; background-color: #FFCC99; text-align: center;\">Summary</div>","metadata":{}},{"cell_type":"code","source":"filt <- filt[ , (metrics_cols) := lapply(.SD, round, 3), .SDcols = metrics_cols][order(validation)]\naugm <- augm[ , (metrics_cols) := lapply(.SD, round, 3), .SDcols = metrics_cols][order(validation)]\nnoise <- noise[ , (metrics_cols) := lapply(.SD, round, 3), .SDcols = metrics_cols][order(validation)]","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:36.111146Z","iopub.execute_input":"2023-12-16T05:36:36.122268Z","iopub.status.idle":"2023-12-16T05:36:36.150601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data augmentation by 50 (LB uplift ~ 0.025)","metadata":{}},{"cell_type":"code","source":"datatable(augm[, !\"noise\"]) %>% formatStyle(metrics_cols, backgroundColor = styleInterval(0, c('#C9E5C4', '#FFCC99')))","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:36.154985Z","iopub.execute_input":"2023-12-16T05:36:36.156683Z","iopub.status.idle":"2023-12-16T05:36:36.525576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"line-height:24px; font-size:16px;border-left: 5px solid silver; padding-left: 26px;\"> \n    💡 Note: \n    Only validation on test drugs and randomly split train corresponds well with the scores improvements in the LB\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Filtration of data from > 2 cells, no augmentation  (LB uplift - unknown)","metadata":{}},{"cell_type":"code","source":"datatable(filt[augm == 1, !c(\"augm\", \"noise\")]) %>% formatStyle(metrics_cols, backgroundColor = styleInterval(0, c('#C9E5C4', '#FFCC99')))","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:36.543269Z","iopub.execute_input":"2023-12-16T05:36:36.545743Z","iopub.status.idle":"2023-12-16T05:36:36.712270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Filtration of data from > 2 cells, with augmentation (LB uplift - none)","metadata":{}},{"cell_type":"code","source":"datatable(filt[augm == 50, !c(\"augm\", \"noise\")]) %>% formatStyle(metrics_cols, backgroundColor = styleInterval(0, c('#C9E5C4', '#FFCC99')))","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:36.714627Z","iopub.execute_input":"2023-12-16T05:36:36.715938Z","iopub.status.idle":"2023-12-16T05:36:36.792862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"line-height:24px; font-size:16px;border-left: 5px solid silver; padding-left: 26px;\"> \n    💡 Note: \n   I did not submit predictions with de_train filtration and without data augmentation, so we cannot assess the performance of the metrics. In the experiments with data augmentation, metrics from the validation on test drugs and randomly split train are changed only slightly, which corresponds with the absence of the scores improvements in the LB. Other folds, like MT and NK cells, show misleading metric changes.\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Noise, with augmentation (LB uplift - none)","metadata":{}},{"cell_type":"code","source":"datatable(noise[, !\"augm\"]) %>% formatStyle(metrics_cols, backgroundColor = styleInterval(0, c('#C9E5C4', '#FFCC99')))","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:36.795480Z","iopub.execute_input":"2023-12-16T05:36:36.797208Z","iopub.status.idle":"2023-12-16T05:36:36.880846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"line-height:24px; font-size:16px;border-left: 5px solid silver; padding-left: 26px;\"> \n    💡 Note: \n    Metrics from the validation on test drugs and randomly split train are changed only slightly, which corresponds with the absence of the scores improvements in the LB. Other folds, like MT and NK cells, show misleading metric changes.\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"border: 3px solid #3B2F2F; border-radius: 10px; padding: 15px; background-color: #FFCC99; text-align: center;\">Public versus private LB scores</div>","metadata":{}},{"cell_type":"markdown","source":"Here I load the information about submissions I saved during the competition in the Google Sheets with added private scores. Sometimes I was lazy so this is not exhaustive. Also, I filtered only models validated on the same CV scheme. Here is the head of the data:","metadata":{}},{"cell_type":"code","source":"cv_lb <- fread(\"/kaggle/input/op2-supplementary-calcs-for-ml/MLP_experiments/CV_LB_metrics_for_MLP.csv\")\nhead(cv_lb)","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:36.884411Z","iopub.execute_input":"2023-12-16T05:36:36.886347Z","iopub.status.idle":"2023-12-16T05:36:36.936957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat(\"Number of model variations in the dataset:\", cv_lb[, .N])","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:36.939638Z","iopub.execute_input":"2023-12-16T05:36:36.940982Z","iopub.status.idle":"2023-12-16T05:36:36.958269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(20, 10)\ncorr <- round(cor(cv_lb$`public LB`, cv_lb$`privat LB`,\n                  method = c(\"pearson\")), 2)\np1 <- ggplot(cv_lb,\n       aes(x = `public LB`,\n           y = `privat LB`)) + \n    geom_point(shape = 21, fill = \"steelblue\", colour = \"black\") + \n    theme_bw(base_size = 22) +\n    geom_text(x = 0.6,\n              y = 1.3,\n              label = paste0('R = ', corr),\n             size = 6) + \n    ggtitle(\"All submissions\")\n\ncorr <- round(cor(cv_lb[`public LB` < 0.6, `public LB`],\n                  cv_lb[`public LB` < 0.6, `privat LB`],\n                  method = c(\"pearson\")), 2)\np2 <- ggplot(cv_lb[`public LB` < 0.6],\n       aes(x = `public LB`,\n           y = `privat LB`)) + \n    geom_point(shape = 21, fill = \"steelblue\", colour = \"black\") + \n    geom_label_repel(aes(label = ifelse(`public LB` < 0.585 &\n                                        `privat LB` > 0.78  |\n                                        `public LB` > 0.594, version, \"\"),\n                            size = 3)) + \n    theme_bw(base_size = 22) +\n    geom_text(x = 0.58,\n              y = 0.808,\n              label = paste0('R = ', corr),\n             size = 6) +\n    theme(legend.position = \"none\") + \n    ggtitle(\"Submissions with public score < 0.6\")\n\np1 + p2","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:36.960566Z","iopub.execute_input":"2023-12-16T05:36:36.961748Z","iopub.status.idle":"2023-12-16T05:36:38.214073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"line-height:24px; font-size:16px;border-left: 5px solid silver; padding-left: 26px;\"> \n    💡 Note: \nIn a closer look, if filter only submissions with the public score < 0.6, we can see that the very good public-private correspondence is misleading.\n</div>","metadata":{}},{"cell_type":"markdown","source":"**Versions with similar public and bad private score**","metadata":{}},{"cell_type":"code","source":"cv_lb[`public LB` < 0.585 & `privat LB` > 0.78]","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:38.217237Z","iopub.execute_input":"2023-12-16T05:36:38.220034Z","iopub.status.idle":"2023-12-16T05:36:38.261927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"line-height:24px; font-size:16px;border-left: 5px solid silver; padding-left: 26px;\"> \n    💡 Note: \nThe filtration of samples was misleading on the public LB: all these versions involve de_train filtration (most of them - before features generation), which did not affect the public score, but it appears it significantly worsened the private score.</div>","metadata":{}},{"cell_type":"markdown","source":"**Versions with bad public and good private score**","metadata":{}},{"cell_type":"code","source":"cv_lb[`public LB` > 0.594 & `privat LB` < 0.78]","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:38.264192Z","iopub.execute_input":"2023-12-16T05:36:38.265343Z","iopub.status.idle":"2023-12-16T05:36:38.293558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"line-height:24px; font-size:16px;border-left: 5px solid silver; padding-left: 26px;\"> \n    💡 Note: \n    The higher number of neurons and lower learning rate (MLPv38) improve privat and local CV scores but worsen the public score.\n</div>","metadata":{}},{"cell_type":"markdown","source":"**Versions with the best private score**","metadata":{}},{"cell_type":"code","source":"cv_lb[`privat LB` < 0.767][order(`privat LB`)]","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:38.295940Z","iopub.execute_input":"2023-12-16T05:36:38.297477Z","iopub.status.idle":"2023-12-16T05:36:38.334428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"line-height:24px; font-size:16px;border-left: 5px solid silver; padding-left: 26px;\"> \n    💡 Note: \n    The best private scores were achieved without filtration of with filtration after the features were generated with all samples. Also, the noise 1.5 maybe better than 1 and 0.5.\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"border: 3px solid #3B2F2F; border-radius: 10px; padding: 15px; background-color: #FFCC99; text-align: center;\">CV versus LB scores</div>","metadata":{}},{"cell_type":"code","source":"m_cols <- c(\"RMSE\",\n            \"mrRMSE\",\n            \"R2\",\n            \"overfit_loss\",\n            \"corr_rows\",\n            \"corr_cols\")\n\nplt <- melt(cv_lb, measure.vars = m_cols,\n           variable.name = \"metric\")\n\ncorrs <- plt[, .(corr_vs_public = \n                 round(cor(value, `public LB`, use = \"complete.obs\"), 2),\n                corr_vs_private = \n                 round(cor(value, `privat LB`, use = \"complete.obs\"), 2)\n                ),\n                 by = \"metric\"]\ncorrs","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:38.417535Z","iopub.execute_input":"2023-12-16T05:36:38.420508Z","iopub.status.idle":"2023-12-16T05:36:38.467085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"line-height:24px; font-size:16px;border-left: 5px solid silver; padding-left: 26px;\"> \n    💡 Note: \n    The local CV metrics correlate much worse with private scores compared to public\n</div>","metadata":{}},{"cell_type":"code","source":"fig(20, 20)\ncorrs <- plt[, .(corr = cor(value, `public LB`, use = \"complete.obs\"),\n                 min_val = min(value, na.rm = TRUE)),\n                 by = \"metric\"]\n\nggplot(plt, aes(x = value, y = `public LB`)) + \n    geom_point(shape = 21, fill = \"steelblue\", colour = \"black\") + \n    geom_text(data = corrs, \n              aes(x = min_val + 0.05, y = 0.9, label = paste0('R = ', round(corr, 3))),\n              size = 6, inherit.aes = FALSE) +\n    facet_wrap(~ metric, scales = \"free\", ncol = 2) +\n    theme_bw(base_size = 22)","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:38.469682Z","iopub.execute_input":"2023-12-16T05:36:38.470997Z","iopub.status.idle":"2023-12-16T05:36:39.750055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check the two outliers","metadata":{}},{"cell_type":"code","source":"cv_lb[`public LB` > 0.7]","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:39.752956Z","iopub.execute_input":"2023-12-16T05:36:39.754416Z","iopub.status.idle":"2023-12-16T05:36:39.799316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"line-height:24px; font-size:16px;border-left: 5px solid silver; padding-left: 26px;\"> \n    💡 Note: \n    The outliers were models trained on the features, that could require tuning of model parameters, so this may be an explanation.\n</div>","metadata":{}},{"cell_type":"code","source":"plt <- plt[`public LB` < 0.7 ]\ncorrs <- plt[, .(corr_vs_public = \n                 round(cor(value, `public LB`, use = \"complete.obs\"), 2),\n                corr_vs_private = \n                 round(cor(value, `privat LB`, use = \"complete.obs\"), 2)\n                ),\n                 by = \"metric\"]\ncorrs","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:39.805413Z","iopub.execute_input":"2023-12-16T05:36:39.808439Z","iopub.status.idle":"2023-12-16T05:36:39.853492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(20, 20)\n\ncorrs <- plt[, .(corr = cor(value, `public LB`, use = \"complete.obs\"),\n                 min_val = min(value, na.rm = TRUE)),\n                 by = \"metric\"]\nggplot(plt, aes(x = value, y = `public LB`)) + \n    geom_point(shape = 21, fill = \"steelblue\", colour = \"black\") + \n    geom_text(data = corrs, \n              aes(x = min_val + 0.05, y = 0.7, label = paste0('R = ', round(corr, 3))),\n              size = 6, inherit.aes = FALSE) +\n    facet_wrap(~ metric, scales = \"free\", ncol = 2) +\n    theme_bw(base_size = 22)","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:39.856038Z","iopub.execute_input":"2023-12-16T05:36:39.858863Z","iopub.status.idle":"2023-12-16T05:36:41.053571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"line-height:24px; font-size:16px;border-left: 5px solid silver; padding-left: 26px;\"> \n    💡 Note: \n    After the outliers removal, the CV-LB correspondence seems much better. The correlation between Yoof predictions and true target values shows the best correspondence to the public LB (correlation 0.786).\n</div>","metadata":{}},{"cell_type":"code","source":"#","metadata":{"execution":{"iopub.status.busy":"2023-12-16T05:36:41.056317Z","iopub.execute_input":"2023-12-16T05:36:41.057816Z","iopub.status.idle":"2023-12-16T05:36:41.074036Z"},"trusted":true},"execution_count":null,"outputs":[]}]}