{
  "id": 458834,
  "title": "What CV you used  ? What tricks/lessons to overcome poor CV - LB correspondence ?  NK-cells - as validation - good for you ? ",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/458834",
  "author_name": "",
  "post_date": "2023-12-01T18:01:17.401454Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>What CV you used  ?  How bad was CV-LB correspondence ?<br>\nWhat tricks/lessons to overcome poor CV - LB correspondence ? <br>\nHave you used the NK-cells as local validation fold - was it good for you ? If yes - how did you come to it ?  </p>\n<p>Our side: </p>\n<ol>\n<li><p><strong>Pains:</strong> Started with AmbrosM and MT schemes - but observed a poor LB CV correspondence for both.  Especially comparing different types of models say boosting and NN seemed hopeless.  Lacked ideas and so worked with them in the following ways a) still sometimes local improvement was good on LB b) searched for cases which improve BOTH these CV (even that does not always work) and such search was not easy </p></li>\n<li><p><strong>Good news 1:</strong> At some late point observed that for many cases of PYBOOST model - the first fold on AmbrosM scheme - i.e. NK-cells have pretty good correspondence with LB. Started to think what was going on.</p>\n<p>2.1 Clever people observed  that NK cells close to LB  immediately looking on EDA : <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/457793#2541075\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/457793#2541075</a><br>\nbut we come to that point - only after much pain , is there any trick/lesson to reduce that pain in future ? </p>\n<p>2.2 Actually we have also seen hint for that by the EDA with clustermap:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=147818286&amp;cellId=21\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=147818286&amp;cellId=21</a><br>\nbut did not believe that should pay attention on that… </p></li>\n<li><p><strong>Trick/lesson 1:</strong>  Started to look on accumulated statistics of submits - just calculated correlation between LB scores and scores on each folds - observed: the NK-fold (AmbrosM-1) and first fold in MT scheme are much better - about 0.5 correlated with LB.  While the other folds correlated badly.</p></li>\n<li><p><strong>Trick/lesson 2:</strong>  tried several other metrics to correlate with the LB for our submits - observed the average row-wise correlation - is better correlated with LB, than competition metric mrrmse itself  ! Strange enough.</p></li>\n<li><p><strong>Lesson:</strong> save your OOF predicts - then you will be able analyze with many metrics and folds - how to fit LB better   ( after  many submits are accumulated). May be even train a small model to predict LB from your local metrics (not actually tried that now). </p></li>\n<li><p><strong>do not overthink:</strong> Turned back to random folds. Observed that a) results on random CV quite similar to results on AmbrosM CV b) in some cases modified Random folds (with only test drugs) shows better corresponde to LB - for some NN correlation was about 0.9 ! </p></li>\n<li><p><strong>Check on big uplifts</strong> Begin to compare  some cases when we have model-1 and model-2 and param change gives strong uplift - checked metrics on those cases. Observed that for some NN the NK-fold is not perfect - uplift on LB - does not correspond to uplfit on CV, but works find with modified random folds.<br>\nStarted to work mainly with random folds. </p></li>\n</ol>",
  "messages": [
    {
      "id": "2545763",
      "postDate": "12/01/2023 18:01:17",
      "content": "<p>What CV you used  ?  How bad was CV-LB correspondence ?<br>\nWhat tricks/lessons to overcome poor CV - LB correspondence ? <br>\nHave you used the NK-cells as local validation fold - was it good for you ? If yes - how did you come to it ?  </p>\n<p>Our side: </p>\n<ol>\n<li><p><strong>Pains:</strong> Started with AmbrosM and MT schemes - but observed a poor LB CV correspondence for both.  Especially comparing different types of models say boosting and NN seemed hopeless.  Lacked ideas and so worked with them in the following ways a) still sometimes local improvement was good on LB b) searched for cases which improve BOTH these CV (even that does not always work) and such search was not easy </p></li>\n<li><p><strong>Good news 1:</strong> At some late point observed that for many cases of PYBOOST model - the first fold on AmbrosM scheme - i.e. NK-cells have pretty good correspondence with LB. Started to think what was going on.</p>\n<p>2.1 Clever people observed  that NK cells close to LB  immediately looking on EDA : <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/457793#2541075\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/457793#2541075</a><br>\nbut we come to that point - only after much pain , is there any trick/lesson to reduce that pain in future ? </p>\n<p>2.2 Actually we have also seen hint for that by the EDA with clustermap:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=147818286&amp;cellId=21\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=147818286&amp;cellId=21</a><br>\nbut did not believe that should pay attention on that… </p></li>\n<li><p><strong>Trick/lesson 1:</strong>  Started to look on accumulated statistics of submits - just calculated correlation between LB scores and scores on each folds - observed: the NK-fold (AmbrosM-1) and first fold in MT scheme are much better - about 0.5 correlated with LB.  While the other folds correlated badly.</p></li>\n<li><p><strong>Trick/lesson 2:</strong>  tried several other metrics to correlate with the LB for our submits - observed the average row-wise correlation - is better correlated with LB, than competition metric mrrmse itself  ! Strange enough.</p></li>\n<li><p><strong>Lesson:</strong> save your OOF predicts - then you will be able analyze with many metrics and folds - how to fit LB better   ( after  many submits are accumulated). May be even train a small model to predict LB from your local metrics (not actually tried that now). </p></li>\n<li><p><strong>do not overthink:</strong> Turned back to random folds. Observed that a) results on random CV quite similar to results on AmbrosM CV b) in some cases modified Random folds (with only test drugs) shows better corresponde to LB - for some NN correlation was about 0.9 ! </p></li>\n<li><p><strong>Check on big uplifts</strong> Begin to compare  some cases when we have model-1 and model-2 and param change gives strong uplift - checked metrics on those cases. Observed that for some NN the NK-fold is not perfect - uplift on LB - does not correspond to uplfit on CV, but works find with modified random folds.<br>\nStarted to work mainly with random folds. </p></li>\n</ol>",
      "rawMarkdown": "What CV you used  ?  How bad was CV-LB correspondence ?\nWhat tricks/lessons to overcome poor CV - LB correspondence ? \nHave you used the NK-cells as local validation fold - was it good for you ? If yes - how did you come to it ?  \n\nOur side: \n\n1.  **Pains:** Started with AmbrosM and MT schemes - but observed a poor LB CV correspondence for both.  Especially comparing different types of models say boosting and NN seemed hopeless.  Lacked ideas and so worked with them in the following ways a) still sometimes local improvement was good on LB b) searched for cases which improve BOTH these CV (even that does not always work) and such search was not easy \n\n2. **Good news 1:** At some late point observed that for many cases of PYBOOST model - the first fold on AmbrosM scheme - i.e. NK-cells have pretty good correspondence with LB. Started to think what was going on.\n\n    2.1 Clever people observed  that NK cells close to LB  immediately looking on EDA : https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/457793#2541075\nbut we come to that point - only after much pain , is there any trick/lesson to reduce that pain in future ? \n\n    2.2 Actually we have also seen hint for that by the EDA with clustermap:\nhttps://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=147818286&cellId=21\nbut did not believe that should pay attention on that... \n\n3. **Trick/lesson 1:**  Started to look on accumulated statistics of submits - just calculated correlation between LB scores and scores on each folds - observed: the NK-fold (AmbrosM-1) and first fold in MT scheme are much better - about 0.5 correlated with LB.  While the other folds correlated badly.\n\n4. **Trick/lesson 2:**  tried several other metrics to correlate with the LB for our submits - observed the average row-wise correlation - is better correlated with LB, than competition metric mrrmse itself  ! Strange enough.\n\n5. **Lesson:** save your OOF predicts - then you will be able analyze with many metrics and folds - how to fit LB better   ( after  many submits are accumulated). May be even train a small model to predict LB from your local metrics (not actually tried that now). \n\n6. **do not overthink:** Turned back to random folds. Observed that a) results on random CV quite similar to results on AmbrosM CV b) in some cases modified Random folds (with only test drugs) shows better corresponde to LB - for some NN correlation was about 0.9 ! \n\n7. **Check on big uplifts** Begin to compare  some cases when we have model-1 and model-2 and param change gives strong uplift - checked metrics on those cases. Observed that for some NN the NK-fold is not perfect - uplift on LB - does not correspond to uplfit on CV, but works find with modified random folds.\nStarted to work mainly with random folds.",
      "votes": null
    },
    {
      "id": "2545824",
      "postDate": "12/01/2023 19:22:55",
      "content": "<p>Thank you for sharing the details of your cross-validation journey! Finding a suitable local validation set was arguably the most challenging aspect of this competition and the lack of a solid validation set had me worried about overfitting to the public leaderboard all through the competition. I tried to be clever about it in the beginning but gave up pretty quickly and started relying mostly on random splits for my neural nets, whereas AmbrosM's CV scheme seemed to work better for random and boosted forests.</p>",
      "rawMarkdown": "Thank you for sharing the details of your cross-validation journey! Finding a suitable local validation set was arguably the most challenging aspect of this competition and the lack of a solid validation set had me worried about overfitting to the public leaderboard all through the competition. I tried to be clever about it in the beginning but gave up pretty quickly and started relying mostly on random splits for my neural nets, whereas AmbrosM's CV scheme seemed to work better for random and boosted forests.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2545824,
      "author_name": "frenio",
      "author_url": "",
      "post_date": "12/01/2023 19:22:55",
      "content": "<p>Thank you for sharing the details of your cross-validation journey! Finding a suitable local validation set was arguably the most challenging aspect of this competition and the lack of a solid validation set had me worried about overfitting to the public leaderboard all through the competition. I tried to be clever about it in the beginning but gave up pretty quickly and started relying mostly on random splits for my neural nets, whereas AmbrosM's CV scheme seemed to work better for random and boosted forests.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2545763": "What CV you used  ?  How bad was CV-LB correspondence ?\nWhat tricks/lessons to overcome poor CV - LB correspondence ? \nHave you used the NK-cells as local validation fold - was it good for you ? If yes - how did you come to it ?  \n\nOur side: \n\n1.  **Pains:** Started with AmbrosM and MT schemes - but observed a poor LB CV correspondence for both.  Especially comparing different types of models say boosting and NN seemed hopeless.  Lacked ideas and so worked with them in the following ways a) still sometimes local improvement was good on LB b) searched for cases which improve BOTH these CV (even that does not always work) and such search was not easy \n\n2. **Good news 1:** At some late point observed that for many cases of PYBOOST model - the first fold on AmbrosM scheme - i.e. NK-cells have pretty good correspondence with LB. Started to think what was going on.\n\n    2.1 Clever people observed  that NK cells close to LB  immediately looking on EDA : https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/457793#2541075\nbut we come to that point - only after much pain , is there any trick/lesson to reduce that pain in future ? \n\n    2.2 Actually we have also seen hint for that by the EDA with clustermap:\nhttps://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=147818286&cellId=21\nbut did not believe that should pay attention on that... \n\n3. **Trick/lesson 1:**  Started to look on accumulated statistics of submits - just calculated correlation between LB scores and scores on each folds - observed: the NK-fold (AmbrosM-1) and first fold in MT scheme are much better - about 0.5 correlated with LB.  While the other folds correlated badly.\n\n4. **Trick/lesson 2:**  tried several other metrics to correlate with the LB for our submits - observed the average row-wise correlation - is better correlated with LB, than competition metric mrrmse itself  ! Strange enough.\n\n5. **Lesson:** save your OOF predicts - then you will be able analyze with many metrics and folds - how to fit LB better   ( after  many submits are accumulated). May be even train a small model to predict LB from your local metrics (not actually tried that now). \n\n6. **do not overthink:** Turned back to random folds. Observed that a) results on random CV quite similar to results on AmbrosM CV b) in some cases modified Random folds (with only test drugs) shows better corresponde to LB - for some NN correlation was about 0.9 ! \n\n7. **Check on big uplifts** Begin to compare  some cases when we have model-1 and model-2 and param change gives strong uplift - checked metrics on those cases. Observed that for some NN the NK-fold is not perfect - uplift on LB - does not correspond to uplfit on CV, but works find with modified random folds.\nStarted to work mainly with random folds.",
    "2545824": "Thank you for sharing the details of your cross-validation journey! Finding a suitable local validation set was arguably the most challenging aspect of this competition and the lack of a solid validation set had me worried about overfitting to the public leaderboard all through the competition. I tried to be clever about it in the beginning but gave up pretty quickly and started relying mostly on random splits for my neural nets, whereas AmbrosM's CV scheme seemed to work better for random and boosted forests."
  },
  "source": "meta"
}