{
  "id": 365016,
  "title": "Questions about the relation between local CV and publi LB",
  "url": "/competitions/open-problems-multimodal/discussion/365016",
  "author_name": "",
  "post_date": "2022-11-09T12:59:17.905218200Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all, I just have a very confusing question. I read some discussion notes and I think if we find score improvement on public LB but not in local CV, it means that we overfit the public LB. But if we meet higher CV but lower local LB, does it mean we should choose different CV method?</p>",
  "messages": [
    {
      "id": "2022982",
      "postDate": "11/09/2022 12:59:17",
      "content": "<p>Hi all, I just have a very confusing question. I read some discussion notes and I think if we find score improvement on public LB but not in local CV, it means that we overfit the public LB. But if we meet higher CV but lower local LB, does it mean we should choose different CV method?</p>",
      "rawMarkdown": "Hi all, I just have a very confusing question. I read some discussion notes and I think if we find score improvement on public LB but not in local CV, it means that we overfit the public LB. But if we meet higher CV but lower local LB, does it mean we should choose different CV method?",
      "votes": null
    },
    {
      "id": "2023519",
      "postDate": "11/09/2022 20:00:40",
      "content": "<p>The question has no simple answer. The facts are:</p>\n<ol>\n<li>You want to tune your model for the maximum possible score on the private leaderboard.</li>\n<li>Tuning decisions must be based on an estimate of the private leaderboard score. The estimate should measure the model's quality in a setting which resembles real life (i.e., the private leaderboard) as much as possible.</li>\n<li>The private leaderboard is calculated on predictions for an unseen day which comes three days after the last training day (day 7 for CITEseq and day 10 for Multiome).</li>\n<li>We know that there are batch effects depending on donor and day.</li>\n</ol>\n<p>The question now turns into: What score is the best estimator for the private leaderboard score? The answer depends not only on facts, but also on your judgement of the situation. Some options are:</p>\n<ol>\n<li>A simple KFold cross-validation doesn't take into account that we train on some batches and the leaderboard is based on other batches.</li>\n<li>The public leaderboard is based on an unknown donor. It measures the model's performance in presence of unknown batch effects, but it ignores the time series aspect.</li>\n<li>You can simulate the unknown day by implementing a GroupKFold on days.</li>\n<li>You can simulate the last day in the sequence by implementing a TimeSeriesSplit.</li>\n</ol>\n<p>You have the choice…</p>",
      "rawMarkdown": "The question has no simple answer. The facts are:\n1. You want to tune your model for the maximum possible score on the private leaderboard.\n2. Tuning decisions must be based on an estimate of the private leaderboard score. The estimate should measure the model's quality in a setting which resembles real life (i.e., the private leaderboard) as much as possible.\n3. The private leaderboard is calculated on predictions for an unseen day which comes three days after the last training day (day 7 for CITEseq and day 10 for Multiome).\n4. We know that there are batch effects depending on donor and day.\n\nThe question now turns into: What score is the best estimator for the private leaderboard score? The answer depends not only on facts, but also on your judgement of the situation. Some options are:\n1. A simple KFold cross-validation doesn't take into account that we train on some batches and the leaderboard is based on other batches.\n2. The public leaderboard is based on an unknown donor. It measures the model's performance in presence of unknown batch effects, but it ignores the time series aspect.\n3. You can simulate the unknown day by implementing a GroupKFold on days.\n4. You can simulate the last day in the sequence by implementing a TimeSeriesSplit.\n\nYou have the choice...",
      "votes": null
    },
    {
      "id": "2023872",
      "postDate": "11/10/2022 05:28:05",
      "content": "<p>This competition is unusual in that the public LB probably has little relationship to the private LB as AmbroseM clarified so well. But if you felt relatively confident about your public LB predictions, it could maybe help with pseudo labels for training/models for private test since all donors are included there just a later day for them.  (Feeling more confused than confident though so not sure if it will help or not!)</p>\n<p>Higher CV but lower public LB could be the split strategy, since it is an unknown donor but if using just a KFold not GroupKFold there could be leakage about the donor and or day to other folds.  Also some public notebooks use a random split subset like 70K of train for Multiome which may or may not work well for CV or public LB depending on what is in that subset.  One other aspect of this data is that although Multiome Train has day 4 which tends to score well in CV if you split out for days, Public test does not have day 4 for Multiome. And from the Data page,  only a subset of Multiome is scored - </p>\n<p>\" To facilitate submission scoring, we only require predictions on a subset of the Multiome data. This subset was created by sampling 30% of the Multiome rows, and for each row, 15% of the columns. The sample of columns varies from row-to-row. All of the CITEseq labels are scored.\"</p>\n<p>Generally you hope to see increase in CV and increase in public LB.  </p>\n<p>But as AmbroseM posted, it is the private LB that will matter in the end.  So thinking about which submissions to select for final score, public LB may not really matter.  Good Luck!</p>",
      "rawMarkdown": "This competition is unusual in that the public LB probably has little relationship to the private LB as AmbroseM clarified so well. But if you felt relatively confident about your public LB predictions, it could maybe help with pseudo labels for training/models for private test since all donors are included there just a later day for them.  (Feeling more confused than confident though so not sure if it will help or not!)\n\nHigher CV but lower public LB could be the split strategy, since it is an unknown donor but if using just a KFold not GroupKFold there could be leakage about the donor and or day to other folds.  Also some public notebooks use a random split subset like 70K of train for Multiome which may or may not work well for CV or public LB depending on what is in that subset.  One other aspect of this data is that although Multiome Train has day 4 which tends to score well in CV if you split out for days, Public test does not have day 4 for Multiome. And from the Data page,  only a subset of Multiome is scored - \n\n\" To facilitate submission scoring, we only require predictions on a subset of the Multiome data. This subset was created by sampling 30% of the Multiome rows, and for each row, 15% of the columns. The sample of columns varies from row-to-row. All of the CITEseq labels are scored.\"\n\nGenerally you hope to see increase in CV and increase in public LB.  \n\nBut as AmbroseM posted, it is the private LB that will matter in the end.  So thinking about which submissions to select for final score, public LB may not really matter.  Good Luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2023519,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "11/09/2022 20:00:40",
      "content": "<p>The question has no simple answer. The facts are:</p>\n<ol>\n<li>You want to tune your model for the maximum possible score on the private leaderboard.</li>\n<li>Tuning decisions must be based on an estimate of the private leaderboard score. The estimate should measure the model's quality in a setting which resembles real life (i.e., the private leaderboard) as much as possible.</li>\n<li>The private leaderboard is calculated on predictions for an unseen day which comes three days after the last training day (day 7 for CITEseq and day 10 for Multiome).</li>\n<li>We know that there are batch effects depending on donor and day.</li>\n</ol>\n<p>The question now turns into: What score is the best estimator for the private leaderboard score? The answer depends not only on facts, but also on your judgement of the situation. Some options are:</p>\n<ol>\n<li>A simple KFold cross-validation doesn't take into account that we train on some batches and the leaderboard is based on other batches.</li>\n<li>The public leaderboard is based on an unknown donor. It measures the model's performance in presence of unknown batch effects, but it ignores the time series aspect.</li>\n<li>You can simulate the unknown day by implementing a GroupKFold on days.</li>\n<li>You can simulate the last day in the sequence by implementing a TimeSeriesSplit.</li>\n</ol>\n<p>You have the choice…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2023872,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "11/10/2022 05:28:05",
      "content": "<p>This competition is unusual in that the public LB probably has little relationship to the private LB as AmbroseM clarified so well. But if you felt relatively confident about your public LB predictions, it could maybe help with pseudo labels for training/models for private test since all donors are included there just a later day for them.  (Feeling more confused than confident though so not sure if it will help or not!)</p>\n<p>Higher CV but lower public LB could be the split strategy, since it is an unknown donor but if using just a KFold not GroupKFold there could be leakage about the donor and or day to other folds.  Also some public notebooks use a random split subset like 70K of train for Multiome which may or may not work well for CV or public LB depending on what is in that subset.  One other aspect of this data is that although Multiome Train has day 4 which tends to score well in CV if you split out for days, Public test does not have day 4 for Multiome. And from the Data page,  only a subset of Multiome is scored - </p>\n<p>\" To facilitate submission scoring, we only require predictions on a subset of the Multiome data. This subset was created by sampling 30% of the Multiome rows, and for each row, 15% of the columns. The sample of columns varies from row-to-row. All of the CITEseq labels are scored.\"</p>\n<p>Generally you hope to see increase in CV and increase in public LB.  </p>\n<p>But as AmbroseM posted, it is the private LB that will matter in the end.  So thinking about which submissions to select for final score, public LB may not really matter.  Good Luck!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2022982": "Hi all, I just have a very confusing question. I read some discussion notes and I think if we find score improvement on public LB but not in local CV, it means that we overfit the public LB. But if we meet higher CV but lower local LB, does it mean we should choose different CV method?",
    "2023519": "The question has no simple answer. The facts are:\n1. You want to tune your model for the maximum possible score on the private leaderboard.\n2. Tuning decisions must be based on an estimate of the private leaderboard score. The estimate should measure the model's quality in a setting which resembles real life (i.e., the private leaderboard) as much as possible.\n3. The private leaderboard is calculated on predictions for an unseen day which comes three days after the last training day (day 7 for CITEseq and day 10 for Multiome).\n4. We know that there are batch effects depending on donor and day.\n\nThe question now turns into: What score is the best estimator for the private leaderboard score? The answer depends not only on facts, but also on your judgement of the situation. Some options are:\n1. A simple KFold cross-validation doesn't take into account that we train on some batches and the leaderboard is based on other batches.\n2. The public leaderboard is based on an unknown donor. It measures the model's performance in presence of unknown batch effects, but it ignores the time series aspect.\n3. You can simulate the unknown day by implementing a GroupKFold on days.\n4. You can simulate the last day in the sequence by implementing a TimeSeriesSplit.\n\nYou have the choice...",
    "2023872": "This competition is unusual in that the public LB probably has little relationship to the private LB as AmbroseM clarified so well. But if you felt relatively confident about your public LB predictions, it could maybe help with pseudo labels for training/models for private test since all donors are included there just a later day for them.  (Feeling more confused than confident though so not sure if it will help or not!)\n\nHigher CV but lower public LB could be the split strategy, since it is an unknown donor but if using just a KFold not GroupKFold there could be leakage about the donor and or day to other folds.  Also some public notebooks use a random split subset like 70K of train for Multiome which may or may not work well for CV or public LB depending on what is in that subset.  One other aspect of this data is that although Multiome Train has day 4 which tends to score well in CV if you split out for days, Public test does not have day 4 for Multiome. And from the Data page,  only a subset of Multiome is scored - \n\n\" To facilitate submission scoring, we only require predictions on a subset of the Multiome data. This subset was created by sampling 30% of the Multiome rows, and for each row, 15% of the columns. The sample of columns varies from row-to-row. All of the CITEseq labels are scored.\"\n\nGenerally you hope to see increase in CV and increase in public LB.  \n\nBut as AmbroseM posted, it is the private LB that will matter in the end.  So thinking about which submissions to select for final score, public LB may not really matter.  Good Luck!"
  },
  "source": "meta"
}