{
  "id": 443395,
  "title": "How are others handling cross-validation?",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/443395",
  "author_name": "",
  "post_date": "2023-09-27T03:58:03.014782100Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I heard from kaggle that it's important to trust your local CV score and not overfit to the public leaderboard. For this competition, I noticed that if you just do a normal k-fold CV, you'll get nulls in the validation split. To fix this, I learned I could do a stratified k-fold, stratifying on <code>sm_name</code>. However, this limits to just 3 folds max before getting the minimum categories per split error. Does anyone have a good take on this?</p>",
  "messages": [
    {
      "id": "2457571",
      "postDate": "09/27/2023 03:58:03",
      "content": "<p>I heard from kaggle that it's important to trust your local CV score and not overfit to the public leaderboard. For this competition, I noticed that if you just do a normal k-fold CV, you'll get nulls in the validation split. To fix this, I learned I could do a stratified k-fold, stratifying on <code>sm_name</code>. However, this limits to just 3 folds max before getting the minimum categories per split error. Does anyone have a good take on this?</p>",
      "rawMarkdown": "I heard from kaggle that it's important to trust your local CV score and not overfit to the public leaderboard. For this competition, I noticed that if you just do a normal k-fold CV, you'll get nulls in the validation split. To fix this, I learned I could do a stratified k-fold, stratifying on `sm_name`. However, this limits to just 3 folds max before getting the minimum categories per split error. Does anyone have a good take on this?",
      "votes": null
    },
    {
      "id": "2457831",
      "postDate": "09/27/2023 07:23:29",
      "content": "<p>The competition data is split into train and test so that the test data contains 90 % of the myeloid and B cells. This setting should be simulated by the cross-validation strategy. </p>\n<p>Because there are six cell types of which two are used for the test data, I suggest to make four cross-validation folds, each one based on 90 % of one of the other four cell-types. The diagram shows training data (green), test data (red) and one of the four validation folds (blue). The few missing cell_type–sm_name combinations (black) can be ignored:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F11f9f88d4fa3d032ce05734bc464ce66%2Fcv.png?generation=1695798735354050&amp;alt=media\" alt=\"cross-validation\"></p>\n<p>The complete source code is in the <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart\" target=\"_blank\">SCP Quickstart</a> notebook.</p>",
      "rawMarkdown": "The competition data is split into train and test so that the test data contains 90 % of the myeloid and B cells. This setting should be simulated by the cross-validation strategy. \n\nBecause there are six cell types of which two are used for the test data, I suggest to make four cross-validation folds, each one based on 90 % of one of the other four cell-types. The diagram shows training data (green), test data (red) and one of the four validation folds (blue). The few missing cell_type–sm_name combinations (black) can be ignored:\n\n![cross-validation](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F11f9f88d4fa3d032ce05734bc464ce66%2Fcv.png?generation=1695798735354050&alt=media)\n\nThe complete source code is in the [SCP Quickstart](https://www.kaggle.com/code/ambrosm/scp-quickstart) notebook.",
      "votes": null
    },
    {
      "id": "2458385",
      "postDate": "09/27/2023 14:53:40",
      "content": "<p>Awesome, thanks for the info and code! I'll take a look at this strategy, sounds practical.</p>",
      "rawMarkdown": "Awesome, thanks for the info and code! I'll take a look at this strategy, sounds practical.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2457831,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "09/27/2023 07:23:29",
      "content": "<p>The competition data is split into train and test so that the test data contains 90 % of the myeloid and B cells. This setting should be simulated by the cross-validation strategy. </p>\n<p>Because there are six cell types of which two are used for the test data, I suggest to make four cross-validation folds, each one based on 90 % of one of the other four cell-types. The diagram shows training data (green), test data (red) and one of the four validation folds (blue). The few missing cell_type–sm_name combinations (black) can be ignored:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F11f9f88d4fa3d032ce05734bc464ce66%2Fcv.png?generation=1695798735354050&amp;alt=media\" alt=\"cross-validation\"></p>\n<p>The complete source code is in the <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart\" target=\"_blank\">SCP Quickstart</a> notebook.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2458385,
          "author_name": "yuqizheng",
          "author_url": "",
          "post_date": "09/27/2023 14:53:40",
          "content": "<p>Awesome, thanks for the info and code! I'll take a look at this strategy, sounds practical.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2457571": "I heard from kaggle that it's important to trust your local CV score and not overfit to the public leaderboard. For this competition, I noticed that if you just do a normal k-fold CV, you'll get nulls in the validation split. To fix this, I learned I could do a stratified k-fold, stratifying on `sm_name`. However, this limits to just 3 folds max before getting the minimum categories per split error. Does anyone have a good take on this?",
    "2457831": "The competition data is split into train and test so that the test data contains 90 % of the myeloid and B cells. This setting should be simulated by the cross-validation strategy. \n\nBecause there are six cell types of which two are used for the test data, I suggest to make four cross-validation folds, each one based on 90 % of one of the other four cell-types. The diagram shows training data (green), test data (red) and one of the four validation folds (blue). The few missing cell_type–sm_name combinations (black) can be ignored:\n\n![cross-validation](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F11f9f88d4fa3d032ce05734bc464ce66%2Fcv.png?generation=1695798735354050&alt=media)\n\nThe complete source code is in the [SCP Quickstart](https://www.kaggle.com/code/ambrosm/scp-quickstart) notebook.",
    "2458385": "Awesome, thanks for the info and code! I'll take a look at this strategy, sounds practical."
  },
  "source": "meta"
}