{
  "id": 460251,
  "title": "CV-LB correspondence analysis: row-wise correlation (CV) - better fits LB, local mrrmse is near ZERO correlated with LB",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/460251",
  "author_name": "",
  "post_date": "2023-12-08T12:12:29.466110Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>CV - LB correspondence is quite problematic in the current challenge. </p>\n<p>Its better understanding would be important for research community future works.<br>\nEven aftermath writeups analysis seems to reveal that good solution for CV-LB correspondence is not found yet. </p>\n<p>Here is some analysis by U900 team:</p>\n<h2>Highlights</h2>\n<ol>\n<li>Local (CV) row-wise correlation score is better  related (0.5)  to LB (mrrmse score) than other metrics </li>\n<li>Local (CV) mrrmse score is near zero correlated with the LB (mrrmse score)  for all CV schemes considered</li>\n<li>NK cells local mrrmse is better  correlated to LB ( 0.2+), while for T-cells CD8+ it is negative (-0.1+) </li>\n<li>NK-cells local mrrmse is well related with LB for Pyboost models, but not for other e.g. NN models</li>\n<li>Random folds are NOT worse than more logical and sophisticated CV-schemes; and seems preferable for NN models</li>\n<li>Public and private LB scores -  highly correlated:  0.98, despite poor CV-LB correspondence</li>\n</ol>\n<p>So, there seems to be many surprises: despite the LB metric is mrrmse - the best locally related to it - is the OTHER metric - row-wise correlation;  while local mrrmse performs near zero. Another surprise - CV-LB correspondence is poor - while public-private LB is very good - 0.98 correlation. And also it is surprising that random folds performs not worse than more logical schemes.</p>\n<h3>Further notes/suggestions:</h3>\n<ol>\n<li>Models of the same nature/features  - the CV-LB correspondence  somehow  working (not so good but still) for all CV schemes. So strategy can be - tune each particular model by CV, verifying by LB - that what we used. </li>\n<li>The main problem to compare different models - even close models Pyboost and Catboost,  with same public LB scores e.g. 0.584 may show quite different CV score like 0.92 vs 0.89, and even worse for boosting vs NN. So for final blend we decided to rely more on LB score, rather than on CV. </li>\n</ol>\n<p><strong>Setup. Potential bias.</strong> The analysis is based on more than 50 quite diverse models, still it can be biased by models choice. We see very clearly that local metrics better corresponding to LB are quite dependent on the model/features/etc (see examples below).</p>\n<h3>Public sharing</h3>\n<p>For community benefits we shared part of code and analysis during the challenge: <a href=\"https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes\" target=\"_blank\">notebook</a> wrapping CV-schemes by AmbrosM, MT,   etc into the class with same interface as sklearn Kfold. <a href=\"https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc\" target=\"_blank\">Kaggle dataset</a> with submits and out-of-fold CV predictions,  <a href=\"https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing\" target=\"_blank\">google doc sheet</a>  with comparison  of CV-LB scores for these CV-schemes for variety of models, <a href=\"https://docs.google.com/presentation/d/1wiz0Wmt4D54pqMMsIOyJHuQYMZ3hTBZQQnjbLzwoGYY/edit?usp=sharing\" target=\"_blank\">slides</a> documenting some subtle moments for CV-LB correspondence.  See also posts: <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458939\" target=\"_blank\">public vs private</a>,  <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458834\" target=\"_blank\">what CV you used</a>, <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444494\" target=\"_blank\">CV - LB thread</a></p>\n<h2>Details</h2>\n<p>The data below is borrowed from our notebooks: <a href=\"https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team\" target=\"_blank\">OP2 CV vs LB analysis U900 team</a>, which analyses the stored submits and local out-of-fold predictions from our <a href=\"https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc\" target=\"_blank\">Kaggle dataset</a> , also notebook:  <a href=\"https://www.kaggle.com/code/alexandervc/op2-public-vs-private-scores\" target=\"_blank\">OP2 Public vs Private scores</a>, and also analysis of many public notebooks.  For the first two tables - we take collection (<a href=\"https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc\" target=\"_blank\">Kaggle dataset</a> ) of more that 50 models and compare a) LB scores  b) CV-scores by AmrbosM scheme c) CV-scores by MT scheme d) CV-scores for Random 5-fold scheme.  The comparison is done for several local metrics.  </p>\n<h4>Table 1. MRRMSE performs worse than row-wise correlation - all CV schemes:</h4>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Pearson correlation LB to CV</th>\n<th>Spearman correlation LB to CV</th>\n<th>CV Scheme</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>corr row</td>\n<td>-0.41</td>\n<td>-0.51</td>\n<td>AmbrosM CV scheme</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>0.11</td>\n<td>0.04</td>\n<td>AmbrosM CV scheme</td>\n</tr>\n<tr>\n<td>corr row</td>\n<td>-0.24</td>\n<td>-0.30</td>\n<td>MT CV scheme</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>0.16</td>\n<td>0.05</td>\n<td>MT CV scheme</td>\n</tr>\n<tr>\n<td>corr row</td>\n<td>-0.36</td>\n<td>-0.50</td>\n<td>Random  CV scheme</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>0.04</td>\n<td>0.15</td>\n<td>Random  CV scheme</td>\n</tr>\n</tbody>\n</table>\n<p>See notebook  <a href=\"https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team\" target=\"_blank\">OP2 CV vs LB analysis U900 team</a> for further details. </p>\n<h4>Table 2. NK-cells perform better for mrrmse, while  for T cells CD8+ it is even negatively correlated to LB</h4>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Pearson correlation LB to CV</th>\n<th>Spearman correlation LB to CV</th>\n<th>Fold Info</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>corr row</td>\n<td>-0.43</td>\n<td>-0.46</td>\n<td>T regulatory cells</td>\n</tr>\n<tr>\n<td>corr row</td>\n<td>-0.35</td>\n<td>-0.47</td>\n<td>NK cells</td>\n</tr>\n<tr>\n<td>corr row</td>\n<td>-0.27</td>\n<td>-0.37</td>\n<td>T cells CD8+</td>\n</tr>\n<tr>\n<td>corr row</td>\n<td>-0.21</td>\n<td>-0.21</td>\n<td>T cells CD4+</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>0.25</td>\n<td>0.36</td>\n<td>NK cells</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>0.17</td>\n<td>0.15</td>\n<td>T cells CD4+</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>0.09</td>\n<td>-0.00</td>\n<td>T regulatory cells</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>-0.13</td>\n<td>-0.13</td>\n<td>T cells CD8+</td>\n</tr>\n</tbody>\n</table>\n<p>See notebook  <a href=\"https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team\" target=\"_blank\">OP2 CV vs LB analysis U900 team</a> for further details. </p>\n<h4>Table 3.  NK-cells local mrrmse well related with LB for Pyboost models,  but not for some other models</h4>\n<p>Compare several modifications of the pyboost model - the LB change is mostly reflected by mrrmse on NK-fold (green color highlight) and opposite to LB change direction for other folds (purple color highlight) </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fda5bb856dc92e2cda3b6ed1a48fc9318%2FScreenshot%202023-12-08%20124908.png?generation=1702036221482020&amp;alt=media\" alt=\"\"></p>\n<p>See <a href=\"https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing\" target=\"_blank\">doc</a></p>\n<h4>Table 4.  NK-cells local mrrmse NOT well related with LB for several AmbrosM models</h4>\n<p>Compare several models by AmbrosM - the LB change is mostly reflected by CD4+ T-cells fold  (green color highlight) and opposite to LB change direction for other folds (purple color highlight) </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Ff3b49e760e4808c7c214b361feb164a4%2FScreenshot%202023-12-08%20125216.png?generation=1702036375142547&amp;alt=media\" alt=\"\"></p>\n<p>See <a href=\"https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing\" target=\"_blank\">doc</a></p>\n<h4>Figures 1,2.  Public and private LB scores - very highly correlated:  0.98, despite poor CV-LB correspondence</h4>\n<p>Figure 1. Scatter plot of public and private scores for the current challenge  Open Problems Single Cell Perturbations - correlation is very high - 0.98<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F8e5b3ab084e4b2d2f483b19d2f214b28%2F__results___10_0%20(3).png?generation=1702035733406836&amp;alt=media\" alt=\"\"></p>\n<p>Figure 2. Scatter plot of public and private scores for the LAST YEAR (2022) challenge  Open Problems Multi modal integration - correlation is also very high - 0.99, but Spearman correlation is much worse - 0.86. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fe708cf356e5437afefb2ac6b851fb477%2F__results___10_0%20(4).png?generation=1702035898423940&amp;alt=media\" alt=\"\"></p>\n<p>See notebook:  <a href=\"https://www.kaggle.com/code/alexandervc/op2-public-vs-private-scores\" target=\"_blank\">OP2 Public vs Private scores</a> for further details.</p>",
  "messages": [
    {
      "id": "2553637",
      "postDate": "12/08/2023 12:12:29",
      "content": "<p>CV - LB correspondence is quite problematic in the current challenge. </p>\n<p>Its better understanding would be important for research community future works.<br>\nEven aftermath writeups analysis seems to reveal that good solution for CV-LB correspondence is not found yet. </p>\n<p>Here is some analysis by U900 team:</p>\n<h2>Highlights</h2>\n<ol>\n<li>Local (CV) row-wise correlation score is better  related (0.5)  to LB (mrrmse score) than other metrics </li>\n<li>Local (CV) mrrmse score is near zero correlated with the LB (mrrmse score)  for all CV schemes considered</li>\n<li>NK cells local mrrmse is better  correlated to LB ( 0.2+), while for T-cells CD8+ it is negative (-0.1+) </li>\n<li>NK-cells local mrrmse is well related with LB for Pyboost models, but not for other e.g. NN models</li>\n<li>Random folds are NOT worse than more logical and sophisticated CV-schemes; and seems preferable for NN models</li>\n<li>Public and private LB scores -  highly correlated:  0.98, despite poor CV-LB correspondence</li>\n</ol>\n<p>So, there seems to be many surprises: despite the LB metric is mrrmse - the best locally related to it - is the OTHER metric - row-wise correlation;  while local mrrmse performs near zero. Another surprise - CV-LB correspondence is poor - while public-private LB is very good - 0.98 correlation. And also it is surprising that random folds performs not worse than more logical schemes.</p>\n<h3>Further notes/suggestions:</h3>\n<ol>\n<li>Models of the same nature/features  - the CV-LB correspondence  somehow  working (not so good but still) for all CV schemes. So strategy can be - tune each particular model by CV, verifying by LB - that what we used. </li>\n<li>The main problem to compare different models - even close models Pyboost and Catboost,  with same public LB scores e.g. 0.584 may show quite different CV score like 0.92 vs 0.89, and even worse for boosting vs NN. So for final blend we decided to rely more on LB score, rather than on CV. </li>\n</ol>\n<p><strong>Setup. Potential bias.</strong> The analysis is based on more than 50 quite diverse models, still it can be biased by models choice. We see very clearly that local metrics better corresponding to LB are quite dependent on the model/features/etc (see examples below).</p>\n<h3>Public sharing</h3>\n<p>For community benefits we shared part of code and analysis during the challenge: <a href=\"https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes\" target=\"_blank\">notebook</a> wrapping CV-schemes by AmbrosM, MT,   etc into the class with same interface as sklearn Kfold. <a href=\"https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc\" target=\"_blank\">Kaggle dataset</a> with submits and out-of-fold CV predictions,  <a href=\"https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing\" target=\"_blank\">google doc sheet</a>  with comparison  of CV-LB scores for these CV-schemes for variety of models, <a href=\"https://docs.google.com/presentation/d/1wiz0Wmt4D54pqMMsIOyJHuQYMZ3hTBZQQnjbLzwoGYY/edit?usp=sharing\" target=\"_blank\">slides</a> documenting some subtle moments for CV-LB correspondence.  See also posts: <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458939\" target=\"_blank\">public vs private</a>,  <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458834\" target=\"_blank\">what CV you used</a>, <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444494\" target=\"_blank\">CV - LB thread</a></p>\n<h2>Details</h2>\n<p>The data below is borrowed from our notebooks: <a href=\"https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team\" target=\"_blank\">OP2 CV vs LB analysis U900 team</a>, which analyses the stored submits and local out-of-fold predictions from our <a href=\"https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc\" target=\"_blank\">Kaggle dataset</a> , also notebook:  <a href=\"https://www.kaggle.com/code/alexandervc/op2-public-vs-private-scores\" target=\"_blank\">OP2 Public vs Private scores</a>, and also analysis of many public notebooks.  For the first two tables - we take collection (<a href=\"https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc\" target=\"_blank\">Kaggle dataset</a> ) of more that 50 models and compare a) LB scores  b) CV-scores by AmrbosM scheme c) CV-scores by MT scheme d) CV-scores for Random 5-fold scheme.  The comparison is done for several local metrics.  </p>\n<h4>Table 1. MRRMSE performs worse than row-wise correlation - all CV schemes:</h4>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Pearson correlation LB to CV</th>\n<th>Spearman correlation LB to CV</th>\n<th>CV Scheme</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>corr row</td>\n<td>-0.41</td>\n<td>-0.51</td>\n<td>AmbrosM CV scheme</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>0.11</td>\n<td>0.04</td>\n<td>AmbrosM CV scheme</td>\n</tr>\n<tr>\n<td>corr row</td>\n<td>-0.24</td>\n<td>-0.30</td>\n<td>MT CV scheme</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>0.16</td>\n<td>0.05</td>\n<td>MT CV scheme</td>\n</tr>\n<tr>\n<td>corr row</td>\n<td>-0.36</td>\n<td>-0.50</td>\n<td>Random  CV scheme</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>0.04</td>\n<td>0.15</td>\n<td>Random  CV scheme</td>\n</tr>\n</tbody>\n</table>\n<p>See notebook  <a href=\"https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team\" target=\"_blank\">OP2 CV vs LB analysis U900 team</a> for further details. </p>\n<h4>Table 2. NK-cells perform better for mrrmse, while  for T cells CD8+ it is even negatively correlated to LB</h4>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Pearson correlation LB to CV</th>\n<th>Spearman correlation LB to CV</th>\n<th>Fold Info</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>corr row</td>\n<td>-0.43</td>\n<td>-0.46</td>\n<td>T regulatory cells</td>\n</tr>\n<tr>\n<td>corr row</td>\n<td>-0.35</td>\n<td>-0.47</td>\n<td>NK cells</td>\n</tr>\n<tr>\n<td>corr row</td>\n<td>-0.27</td>\n<td>-0.37</td>\n<td>T cells CD8+</td>\n</tr>\n<tr>\n<td>corr row</td>\n<td>-0.21</td>\n<td>-0.21</td>\n<td>T cells CD4+</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>0.25</td>\n<td>0.36</td>\n<td>NK cells</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>0.17</td>\n<td>0.15</td>\n<td>T cells CD4+</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>0.09</td>\n<td>-0.00</td>\n<td>T regulatory cells</td>\n</tr>\n<tr>\n<td>mrrmse</td>\n<td>-0.13</td>\n<td>-0.13</td>\n<td>T cells CD8+</td>\n</tr>\n</tbody>\n</table>\n<p>See notebook  <a href=\"https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team\" target=\"_blank\">OP2 CV vs LB analysis U900 team</a> for further details. </p>\n<h4>Table 3.  NK-cells local mrrmse well related with LB for Pyboost models,  but not for some other models</h4>\n<p>Compare several modifications of the pyboost model - the LB change is mostly reflected by mrrmse on NK-fold (green color highlight) and opposite to LB change direction for other folds (purple color highlight) </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fda5bb856dc92e2cda3b6ed1a48fc9318%2FScreenshot%202023-12-08%20124908.png?generation=1702036221482020&amp;alt=media\" alt=\"\"></p>\n<p>See <a href=\"https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing\" target=\"_blank\">doc</a></p>\n<h4>Table 4.  NK-cells local mrrmse NOT well related with LB for several AmbrosM models</h4>\n<p>Compare several models by AmbrosM - the LB change is mostly reflected by CD4+ T-cells fold  (green color highlight) and opposite to LB change direction for other folds (purple color highlight) </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Ff3b49e760e4808c7c214b361feb164a4%2FScreenshot%202023-12-08%20125216.png?generation=1702036375142547&amp;alt=media\" alt=\"\"></p>\n<p>See <a href=\"https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing\" target=\"_blank\">doc</a></p>\n<h4>Figures 1,2.  Public and private LB scores - very highly correlated:  0.98, despite poor CV-LB correspondence</h4>\n<p>Figure 1. Scatter plot of public and private scores for the current challenge  Open Problems Single Cell Perturbations - correlation is very high - 0.98<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F8e5b3ab084e4b2d2f483b19d2f214b28%2F__results___10_0%20(3).png?generation=1702035733406836&amp;alt=media\" alt=\"\"></p>\n<p>Figure 2. Scatter plot of public and private scores for the LAST YEAR (2022) challenge  Open Problems Multi modal integration - correlation is also very high - 0.99, but Spearman correlation is much worse - 0.86. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fe708cf356e5437afefb2ac6b851fb477%2F__results___10_0%20(4).png?generation=1702035898423940&amp;alt=media\" alt=\"\"></p>\n<p>See notebook:  <a href=\"https://www.kaggle.com/code/alexandervc/op2-public-vs-private-scores\" target=\"_blank\">OP2 Public vs Private scores</a> for further details.</p>",
      "rawMarkdown": "CV - LB correspondence is quite problematic in the current challenge. \n\nIts better understanding would be important for research community future works.\nEven aftermath writeups analysis seems to reveal that good solution for CV-LB correspondence is not found yet. \n\nHere is some analysis by U900 team:\n\n## Highlights\n1.  Local (CV) row-wise correlation score is better  related (0.5)  to LB (mrrmse score) than other metrics \n2. Local (CV) mrrmse score is near zero correlated with the LB (mrrmse score)  for all CV schemes considered\n3. NK cells local mrrmse is better  correlated to LB ( 0.2+), while for T-cells CD8+ it is negative (-0.1+) \n4. NK-cells local mrrmse is well related with LB for Pyboost models, but not for other e.g. NN models\n5. Random folds are NOT worse than more logical and sophisticated CV-schemes; and seems preferable for NN models\n6. Public and private LB scores -  highly correlated:  0.98, despite poor CV-LB correspondence\n\nSo, there seems to be many surprises: despite the LB metric is mrrmse - the best locally related to it - is the OTHER metric - row-wise correlation;  while local mrrmse performs near zero. Another surprise - CV-LB correspondence is poor - while public-private LB is very good - 0.98 correlation. And also it is surprising that random folds performs not worse than more logical schemes.\n\n\n###  Further notes/suggestions:\n1.  Models of the same nature/features  - the CV-LB correspondence  somehow  working (not so good but still) for all CV schemes. So strategy can be - tune each particular model by CV, verifying by LB - that what we used. \n2. The main problem to compare different models - even close models Pyboost and Catboost,  with same public LB scores e.g. 0.584 may show quite different CV score like 0.92 vs 0.89, and even worse for boosting vs NN. So for final blend we decided to rely more on LB score, rather than on CV. \n\n\n**Setup. Potential bias.** The analysis is based on more than 50 quite diverse models, still it can be biased by models choice. We see very clearly that local metrics better corresponding to LB are quite dependent on the model/features/etc (see examples below).\n\n### Public sharing \n\nFor community benefits we shared part of code and analysis during the challenge: [notebook](https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes) wrapping CV-schemes by AmbrosM, MT,   etc into the class with same interface as sklearn Kfold. [Kaggle dataset](https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc) with submits and out-of-fold CV predictions,  [google doc sheet](https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing)  with comparison  of CV-LB scores for these CV-schemes for variety of models, [slides](https://docs.google.com/presentation/d/1wiz0Wmt4D54pqMMsIOyJHuQYMZ3hTBZQQnjbLzwoGYY/edit?usp=sharing) documenting some subtle moments for CV-LB correspondence.  See also posts: \n[public vs private](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458939),  [what CV you used](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458834), [CV - LB thread](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444494)\n\n## Details\n\n The data below is borrowed from our notebooks: [OP2 CV vs LB analysis U900 team](https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team), which analyses the stored submits and local out-of-fold predictions from our [Kaggle dataset](https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc) , also notebook:  [OP2 Public vs Private scores](https://www.kaggle.com/code/alexandervc/op2-public-vs-private-scores), and also analysis of many public notebooks.  For the first two tables - we take collection ([Kaggle dataset](https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc) ) of more that 50 models and compare a) LB scores  b) CV-scores by AmrbosM scheme c) CV-scores by MT scheme d) CV-scores for Random 5-fold scheme.  The comparison is done for several local metrics.  \n\n#### Table 1. MRRMSE performs worse than row-wise correlation - all CV schemes:\n\n|         Metric      | Pearson correlation LB to CV | Spearman correlation LB to CV | CV Scheme                 |\n|---------------------|---------------------------------------|----------------------------------------|---------------------------|\n| corr row       | -0.41                                 | -0.51                                  | AmbrosM CV scheme           |\n| mrrmse        | 0.11                                  | 0.04                                   | AmbrosM CV scheme           |\n| corr row       | -0.24                            | -0.30                             | MT CV scheme |\n| mrrmse        | 0.16                             | 0.05                              | MT CV scheme   |\n| corr row       | -0.36                              | -0.50                  | Random  CV scheme   |\n| mrrmse        | 0.04                               | 0.15          | Random  CV scheme   |\n\nSee notebook  [OP2 CV vs LB analysis U900 team](https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team) for further details. \n\n#### Table 2. NK-cells perform better for mrrmse, while  for T cells CD8+ it is even negatively correlated to LB \n\n|         Metric      | Pearson correlation LB to CV | Spearman correlation LB to CV | Fold Info                 |\n|---------------------|---------------------------------------|----------------------------------------|---------------------------|\n| corr row           | -0.43                                 | -0.46                                  | T regulatory cells       |\n| corr row           | -0.35                                 | -0.47                                  | NK cells                  |\n| corr row           | -0.27                                 | -0.37                                  | T cells CD8+              |\n| corr row           | -0.21                                 | -0.21                                  | T cells CD4+              |\n| mrrmse             | 0.25                                  | 0.36                                   | NK cells                  |\n| mrrmse             | 0.17                                  | 0.15                                   | T cells CD4+              |\n| mrrmse             | 0.09                                  | -0.00                                  | T regulatory cells       |\n| mrrmse             | -0.13                                 | -0.13                                  | T cells CD8+              |\n\nSee notebook  [OP2 CV vs LB analysis U900 team](https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team) for further details. \n \n#### Table 3.  NK-cells local mrrmse well related with LB for Pyboost models,  but not for some other models \n\nCompare several modifications of the pyboost model - the LB change is mostly reflected by mrrmse on NK-fold (green color highlight) and opposite to LB change direction for other folds (purple color highlight) \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fda5bb856dc92e2cda3b6ed1a48fc9318%2FScreenshot%202023-12-08%20124908.png?generation=1702036221482020&alt=media)\n\nSee [doc](https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing)\n\n#### Table 4.  NK-cells local mrrmse NOT well related with LB for several AmbrosM models\n\nCompare several models by AmbrosM - the LB change is mostly reflected by CD4+ T-cells fold  (green color highlight) and opposite to LB change direction for other folds (purple color highlight) \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Ff3b49e760e4808c7c214b361feb164a4%2FScreenshot%202023-12-08%20125216.png?generation=1702036375142547&alt=media)\n\nSee [doc](https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing)\n\n\n#### Figures 1,2.  Public and private LB scores - very highly correlated:  0.98, despite poor CV-LB correspondence\n\nFigure 1. Scatter plot of public and private scores for the current challenge  Open Problems Single Cell Perturbations - correlation is very high - 0.98\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F8e5b3ab084e4b2d2f483b19d2f214b28%2F__results___10_0%20(3).png?generation=1702035733406836&alt=media)\n\nFigure 2. Scatter plot of public and private scores for the LAST YEAR (2022) challenge  Open Problems Multi modal integration - correlation is also very high - 0.99, but Spearman correlation is much worse - 0.86. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fe708cf356e5437afefb2ac6b851fb477%2F__results___10_0%20(4).png?generation=1702035898423940&alt=media)\n\nSee notebook:  [OP2 Public vs Private scores](https://www.kaggle.com/code/alexandervc/op2-public-vs-private-scores) for further details.",
      "votes": null
    },
    {
      "id": "2553943",
      "postDate": "12/08/2023 16:30:48",
      "content": "<p>Thank you for sharing! You indeed raised some interesting questions. I have to add, that some points mentioned do not apply to the MLP NN (shown in <a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-13th-place-solution\" target=\"_blank\">this notebook</a>), which had almost a half of the weight in the final solution.</p>\n<p><a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-tricks-and-metrics#Public-versus-private-LB-scores\" target=\"_blank\">Here</a> I did a separate analysis of LB-CV and private-public scores correspondence for 51 different variations of this NN.</p>\n<p><strong>Public versus private LB scores</strong><br>\nIf we take a closer look at submissions with scores below 0.6, we see that the very good public-private correspondence is misleading (correlation 0.98 -&gt; 0.18). </p>\n<p>Our and the organizer's interest lies in having good, quality predictions that are equally good for different test data. But this isn't the case. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F03319158fc5f503ca56769b6bb32529e%2FScreenshot%202023-12-08%20184927.png?generation=1702050873083119&amp;alt=media\" alt=\"\"></p>\n<p>It's noticeable that in your picture this isn't the case too.  If you zoom in on the area with public scores below 0.6, you will see points with nearly the same public score and very different private scores.</p>\n<p>More interestingly, I was able to identify a few model parameters that led to the improvement of the private but not public score (e.g. higher number of neurons and lower learning rate), as well as some misleading model changes that improved the public but worsened private score (e.g. filtering out samples from 1-2 cells based on adata analysis ).</p>\n<p><strong>CV versus LB scores</strong></p>\n<p>Also, with this NN, when I used validation on test drugs, the CV-LB correspondence was quite good if, as you correctly pointed out, we carefully selected a subset of more similar models to compare (e.g. trained on the same or similar features):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2Fe1efdfa5949a1d0f4e5a66bbe6e5eb22%2FScreenshot%202023-12-08%20190440.png?generation=1702051560092653&amp;alt=media\" alt=\"\">.</p>\n<p>We also can see that the local CV metrics correlate much worse with private scores compared to public. And yes, quite an interesting observation that the row-wise correlation score for this NN is much better agrees with the LB scores than other metrics do, as with Pyboost.</p>\n<p>Finally, to illustrate that different schemes are better suited for different models, I have done a fairly detailed analysis with more than 100 experiments with repeats <a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-tricks-and-metrics#Summary\" target=\"_blank\">here</a>. For this MLP NN, validation on test drugs corresponds better with the scores improvements in the LB compared to other schemes, including NK cells. Actually, other schemes, like MT and NK cells, show metric changes opposite to LB changes.</p>",
      "rawMarkdown": "Thank you for sharing! You indeed raised some interesting questions. I have to add, that some points mentioned do not apply to the MLP NN (shown in [this notebook](https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-13th-place-solution)), which had almost a half of the weight in the final solution.\n\n[Here](https://www.kaggle.com/code/antoninadolgorukova/op2-tricks-and-metrics#Public-versus-private-LB-scores) I did a separate analysis of LB-CV and private-public scores correspondence for 51 different variations of this NN.\n\n**Public versus private LB scores**\nIf we take a closer look at submissions with scores below 0.6, we see that the very good public-private correspondence is misleading (correlation 0.98 -> 0.18). \n\nOur and the organizer's interest lies in having good, quality predictions that are equally good for different test data. But this isn't the case. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F03319158fc5f503ca56769b6bb32529e%2FScreenshot%202023-12-08%20184927.png?generation=1702050873083119&alt=media)\n\nIt's noticeable that in your picture this isn't the case too.  If you zoom in on the area with public scores below 0.6, you will see points with nearly the same public score and very different private scores.\n\nMore interestingly, I was able to identify a few model parameters that led to the improvement of the private but not public score (e.g. higher number of neurons and lower learning rate), as well as some misleading model changes that improved the public but worsened private score (e.g. filtering out samples from 1-2 cells based on adata analysis ).\n\n**CV versus LB scores**\n\nAlso, with this NN, when I used validation on test drugs, the CV-LB correspondence was quite good if, as you correctly pointed out, we carefully selected a subset of more similar models to compare (e.g. trained on the same or similar features):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2Fe1efdfa5949a1d0f4e5a66bbe6e5eb22%2FScreenshot%202023-12-08%20190440.png?generation=1702051560092653&alt=media).\n\nWe also can see that the local CV metrics correlate much worse with private scores compared to public. And yes, quite an interesting observation that the row-wise correlation score for this NN is much better agrees with the LB scores than other metrics do, as with Pyboost.\n\nFinally, to illustrate that different schemes are better suited for different models, I have done a fairly detailed analysis with more than 100 experiments with repeats [here](https://www.kaggle.com/code/antoninadolgorukova/op2-tricks-and-metrics#Summary). For this MLP NN, validation on test drugs corresponds better with the scores improvements in the LB compared to other schemes, including NK cells. Actually, other schemes, like MT and NK cells, show metric changes opposite to LB changes.",
      "votes": null
    },
    {
      "id": "2554117",
      "postDate": "12/08/2023 19:40:47",
      "content": "<p>Aha, you are a statistician!  I like your deep dive.</p>",
      "rawMarkdown": "Aha, you are a statistician!  I like your deep dive.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2553943,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "12/08/2023 16:30:48",
      "content": "<p>Thank you for sharing! You indeed raised some interesting questions. I have to add, that some points mentioned do not apply to the MLP NN (shown in <a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-13th-place-solution\" target=\"_blank\">this notebook</a>), which had almost a half of the weight in the final solution.</p>\n<p><a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-tricks-and-metrics#Public-versus-private-LB-scores\" target=\"_blank\">Here</a> I did a separate analysis of LB-CV and private-public scores correspondence for 51 different variations of this NN.</p>\n<p><strong>Public versus private LB scores</strong><br>\nIf we take a closer look at submissions with scores below 0.6, we see that the very good public-private correspondence is misleading (correlation 0.98 -&gt; 0.18). </p>\n<p>Our and the organizer's interest lies in having good, quality predictions that are equally good for different test data. But this isn't the case. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F03319158fc5f503ca56769b6bb32529e%2FScreenshot%202023-12-08%20184927.png?generation=1702050873083119&amp;alt=media\" alt=\"\"></p>\n<p>It's noticeable that in your picture this isn't the case too.  If you zoom in on the area with public scores below 0.6, you will see points with nearly the same public score and very different private scores.</p>\n<p>More interestingly, I was able to identify a few model parameters that led to the improvement of the private but not public score (e.g. higher number of neurons and lower learning rate), as well as some misleading model changes that improved the public but worsened private score (e.g. filtering out samples from 1-2 cells based on adata analysis ).</p>\n<p><strong>CV versus LB scores</strong></p>\n<p>Also, with this NN, when I used validation on test drugs, the CV-LB correspondence was quite good if, as you correctly pointed out, we carefully selected a subset of more similar models to compare (e.g. trained on the same or similar features):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2Fe1efdfa5949a1d0f4e5a66bbe6e5eb22%2FScreenshot%202023-12-08%20190440.png?generation=1702051560092653&amp;alt=media\" alt=\"\">.</p>\n<p>We also can see that the local CV metrics correlate much worse with private scores compared to public. And yes, quite an interesting observation that the row-wise correlation score for this NN is much better agrees with the LB scores than other metrics do, as with Pyboost.</p>\n<p>Finally, to illustrate that different schemes are better suited for different models, I have done a fairly detailed analysis with more than 100 experiments with repeats <a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-tricks-and-metrics#Summary\" target=\"_blank\">here</a>. For this MLP NN, validation on test drugs corresponds better with the scores improvements in the LB compared to other schemes, including NK cells. Actually, other schemes, like MT and NK cells, show metric changes opposite to LB changes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2554117,
          "author_name": "makio323",
          "author_url": "",
          "post_date": "12/08/2023 19:40:47",
          "content": "<p>Aha, you are a statistician!  I like your deep dive.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2553637": "CV - LB correspondence is quite problematic in the current challenge. \n\nIts better understanding would be important for research community future works.\nEven aftermath writeups analysis seems to reveal that good solution for CV-LB correspondence is not found yet. \n\nHere is some analysis by U900 team:\n\n## Highlights\n1.  Local (CV) row-wise correlation score is better  related (0.5)  to LB (mrrmse score) than other metrics \n2. Local (CV) mrrmse score is near zero correlated with the LB (mrrmse score)  for all CV schemes considered\n3. NK cells local mrrmse is better  correlated to LB ( 0.2+), while for T-cells CD8+ it is negative (-0.1+) \n4. NK-cells local mrrmse is well related with LB for Pyboost models, but not for other e.g. NN models\n5. Random folds are NOT worse than more logical and sophisticated CV-schemes; and seems preferable for NN models\n6. Public and private LB scores -  highly correlated:  0.98, despite poor CV-LB correspondence\n\nSo, there seems to be many surprises: despite the LB metric is mrrmse - the best locally related to it - is the OTHER metric - row-wise correlation;  while local mrrmse performs near zero. Another surprise - CV-LB correspondence is poor - while public-private LB is very good - 0.98 correlation. And also it is surprising that random folds performs not worse than more logical schemes.\n\n\n###  Further notes/suggestions:\n1.  Models of the same nature/features  - the CV-LB correspondence  somehow  working (not so good but still) for all CV schemes. So strategy can be - tune each particular model by CV, verifying by LB - that what we used. \n2. The main problem to compare different models - even close models Pyboost and Catboost,  with same public LB scores e.g. 0.584 may show quite different CV score like 0.92 vs 0.89, and even worse for boosting vs NN. So for final blend we decided to rely more on LB score, rather than on CV. \n\n\n**Setup. Potential bias.** The analysis is based on more than 50 quite diverse models, still it can be biased by models choice. We see very clearly that local metrics better corresponding to LB are quite dependent on the model/features/etc (see examples below).\n\n### Public sharing \n\nFor community benefits we shared part of code and analysis during the challenge: [notebook](https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes) wrapping CV-schemes by AmbrosM, MT,   etc into the class with same interface as sklearn Kfold. [Kaggle dataset](https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc) with submits and out-of-fold CV predictions,  [google doc sheet](https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing)  with comparison  of CV-LB scores for these CV-schemes for variety of models, [slides](https://docs.google.com/presentation/d/1wiz0Wmt4D54pqMMsIOyJHuQYMZ3hTBZQQnjbLzwoGYY/edit?usp=sharing) documenting some subtle moments for CV-LB correspondence.  See also posts: \n[public vs private](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458939),  [what CV you used](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458834), [CV - LB thread](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444494)\n\n## Details\n\n The data below is borrowed from our notebooks: [OP2 CV vs LB analysis U900 team](https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team), which analyses the stored submits and local out-of-fold predictions from our [Kaggle dataset](https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc) , also notebook:  [OP2 Public vs Private scores](https://www.kaggle.com/code/alexandervc/op2-public-vs-private-scores), and also analysis of many public notebooks.  For the first two tables - we take collection ([Kaggle dataset](https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc) ) of more that 50 models and compare a) LB scores  b) CV-scores by AmrbosM scheme c) CV-scores by MT scheme d) CV-scores for Random 5-fold scheme.  The comparison is done for several local metrics.  \n\n#### Table 1. MRRMSE performs worse than row-wise correlation - all CV schemes:\n\n|         Metric      | Pearson correlation LB to CV | Spearman correlation LB to CV | CV Scheme                 |\n|---------------------|---------------------------------------|----------------------------------------|---------------------------|\n| corr row       | -0.41                                 | -0.51                                  | AmbrosM CV scheme           |\n| mrrmse        | 0.11                                  | 0.04                                   | AmbrosM CV scheme           |\n| corr row       | -0.24                            | -0.30                             | MT CV scheme |\n| mrrmse        | 0.16                             | 0.05                              | MT CV scheme   |\n| corr row       | -0.36                              | -0.50                  | Random  CV scheme   |\n| mrrmse        | 0.04                               | 0.15          | Random  CV scheme   |\n\nSee notebook  [OP2 CV vs LB analysis U900 team](https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team) for further details. \n\n#### Table 2. NK-cells perform better for mrrmse, while  for T cells CD8+ it is even negatively correlated to LB \n\n|         Metric      | Pearson correlation LB to CV | Spearman correlation LB to CV | Fold Info                 |\n|---------------------|---------------------------------------|----------------------------------------|---------------------------|\n| corr row           | -0.43                                 | -0.46                                  | T regulatory cells       |\n| corr row           | -0.35                                 | -0.47                                  | NK cells                  |\n| corr row           | -0.27                                 | -0.37                                  | T cells CD8+              |\n| corr row           | -0.21                                 | -0.21                                  | T cells CD4+              |\n| mrrmse             | 0.25                                  | 0.36                                   | NK cells                  |\n| mrrmse             | 0.17                                  | 0.15                                   | T cells CD4+              |\n| mrrmse             | 0.09                                  | -0.00                                  | T regulatory cells       |\n| mrrmse             | -0.13                                 | -0.13                                  | T cells CD8+              |\n\nSee notebook  [OP2 CV vs LB analysis U900 team](https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis-u900-team) for further details. \n \n#### Table 3.  NK-cells local mrrmse well related with LB for Pyboost models,  but not for some other models \n\nCompare several modifications of the pyboost model - the LB change is mostly reflected by mrrmse on NK-fold (green color highlight) and opposite to LB change direction for other folds (purple color highlight) \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fda5bb856dc92e2cda3b6ed1a48fc9318%2FScreenshot%202023-12-08%20124908.png?generation=1702036221482020&alt=media)\n\nSee [doc](https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing)\n\n#### Table 4.  NK-cells local mrrmse NOT well related with LB for several AmbrosM models\n\nCompare several models by AmbrosM - the LB change is mostly reflected by CD4+ T-cells fold  (green color highlight) and opposite to LB change direction for other folds (purple color highlight) \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Ff3b49e760e4808c7c214b361feb164a4%2FScreenshot%202023-12-08%20125216.png?generation=1702036375142547&alt=media)\n\nSee [doc](https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing)\n\n\n#### Figures 1,2.  Public and private LB scores - very highly correlated:  0.98, despite poor CV-LB correspondence\n\nFigure 1. Scatter plot of public and private scores for the current challenge  Open Problems Single Cell Perturbations - correlation is very high - 0.98\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F8e5b3ab084e4b2d2f483b19d2f214b28%2F__results___10_0%20(3).png?generation=1702035733406836&alt=media)\n\nFigure 2. Scatter plot of public and private scores for the LAST YEAR (2022) challenge  Open Problems Multi modal integration - correlation is also very high - 0.99, but Spearman correlation is much worse - 0.86. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fe708cf356e5437afefb2ac6b851fb477%2F__results___10_0%20(4).png?generation=1702035898423940&alt=media)\n\nSee notebook:  [OP2 Public vs Private scores](https://www.kaggle.com/code/alexandervc/op2-public-vs-private-scores) for further details.",
    "2553943": "Thank you for sharing! You indeed raised some interesting questions. I have to add, that some points mentioned do not apply to the MLP NN (shown in [this notebook](https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-13th-place-solution)), which had almost a half of the weight in the final solution.\n\n[Here](https://www.kaggle.com/code/antoninadolgorukova/op2-tricks-and-metrics#Public-versus-private-LB-scores) I did a separate analysis of LB-CV and private-public scores correspondence for 51 different variations of this NN.\n\n**Public versus private LB scores**\nIf we take a closer look at submissions with scores below 0.6, we see that the very good public-private correspondence is misleading (correlation 0.98 -> 0.18). \n\nOur and the organizer's interest lies in having good, quality predictions that are equally good for different test data. But this isn't the case. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F03319158fc5f503ca56769b6bb32529e%2FScreenshot%202023-12-08%20184927.png?generation=1702050873083119&alt=media)\n\nIt's noticeable that in your picture this isn't the case too.  If you zoom in on the area with public scores below 0.6, you will see points with nearly the same public score and very different private scores.\n\nMore interestingly, I was able to identify a few model parameters that led to the improvement of the private but not public score (e.g. higher number of neurons and lower learning rate), as well as some misleading model changes that improved the public but worsened private score (e.g. filtering out samples from 1-2 cells based on adata analysis ).\n\n**CV versus LB scores**\n\nAlso, with this NN, when I used validation on test drugs, the CV-LB correspondence was quite good if, as you correctly pointed out, we carefully selected a subset of more similar models to compare (e.g. trained on the same or similar features):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2Fe1efdfa5949a1d0f4e5a66bbe6e5eb22%2FScreenshot%202023-12-08%20190440.png?generation=1702051560092653&alt=media).\n\nWe also can see that the local CV metrics correlate much worse with private scores compared to public. And yes, quite an interesting observation that the row-wise correlation score for this NN is much better agrees with the LB scores than other metrics do, as with Pyboost.\n\nFinally, to illustrate that different schemes are better suited for different models, I have done a fairly detailed analysis with more than 100 experiments with repeats [here](https://www.kaggle.com/code/antoninadolgorukova/op2-tricks-and-metrics#Summary). For this MLP NN, validation on test drugs corresponds better with the scores improvements in the LB compared to other schemes, including NK cells. Actually, other schemes, like MT and NK cells, show metric changes opposite to LB changes.",
    "2554117": "Aha, you are a statistician!  I like your deep dive."
  },
  "source": "meta"
}