{
  "id": 358209,
  "title": "About the cross-validation strategy",
  "url": "/competitions/open-problems-multimodal/discussion/358209",
  "author_name": "shunto nakamura",
  "post_date": "2022-10-07T02:28:21.201000",
  "votes": 9,
  "comment_count": 6,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6658947%2F5f191ab34a8c773e32bc1730f0103232%2F2022-10-07%2011.07.17.png?generation=1665108495487036&amp;alt=media\" alt=\"\"></p>\n<p>Hi, I am new to KAGGLE.<br>\nAs discussed <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202\" target=\"_blank\">here</a>, I think that the correct cross-validation that reflects private is necessary to win the prize, because of the different way of dividing public and private data.</p>\n<p>The private consists of 3 known donors and 1 unknown donor at a future date. So I think the validation dataset for cross-validation should consist of 2 known donors and 1 unknown donor at a future date.</p>\n<p>Is this concept correct?<br>\nAnd please share your cross-validation strategies.</p>",
  "messages": [
    {
      "id": 1975774,
      "postDate": "2022-10-07T02:28:21.200Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6658947%2F5f191ab34a8c773e32bc1730f0103232%2F2022-10-07%2011.07.17.png?generation=1665108495487036&amp;alt=media\" alt=\"\"></p>\n<p>Hi, I am new to KAGGLE.<br>\nAs discussed <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202\" target=\"_blank\">here</a>, I think that the correct cross-validation that reflects private is necessary to win the prize, because of the different way of dividing public and private data.</p>\n<p>The private consists of 3 known donors and 1 unknown donor at a future date. So I think the validation dataset for cross-validation should consist of 2 known donors and 1 unknown donor at a future date.</p>\n<p>Is this concept correct?<br>\nAnd please share your cross-validation strategies.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6658947%2F5f191ab34a8c773e32bc1730f0103232%2F2022-10-07%2011.07.17.png?generation=1665108495487036&alt=media)\n\n\nHi, I am new to KAGGLE.\nAs discussed [here](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202), I think that the correct cross-validation that reflects private is necessary to win the prize, because of the different way of dividing public and private data.\n\nThe private consists of 3 known donors and 1 unknown donor at a future date. So I think the validation dataset for cross-validation should consist of 2 known donors and 1 unknown donor at a future date.\n\nIs this concept correct?\nAnd please share your cross-validation strategies.",
      "votes": 8
    },
    {
      "id": 1984971,
      "postDate": "2022-10-13T02:36:28.803Z",
      "content": "<p>I now am using GroupbyKFold with a combined donor and date feature.<br>\nI am currently thinking that as long as Train and Test do not include the same donor on the same day, it should be fine.</p>",
      "rawMarkdown": "I now am using GroupbyKFold with a combined donor and date feature.\nI am currently thinking that as long as Train and Test do not include the same donor on the same day, it should be fine.",
      "votes": 1
    },
    {
      "id": 1979935,
      "postDate": "2022-10-09T20:26:48.453Z",
      "content": "<p>My proposal on CV is the following:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860</a></p>",
      "rawMarkdown": "My proposal on CV is the following:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860",
      "votes": 1
    },
    {
      "id": 1981953,
      "postDate": "2022-10-11T06:48:15.990Z",
      "content": "<blockquote>\n  <p>The private consists of 3 known donors and 1 unknown donor at a future date. So I think the validation dataset for cross-validation should consist of 2 known donors and 1 unknown donor at a future date.</p>\n</blockquote>\n<p>I agree with this for predicting the unseen donor. For the others, I think you should have either all three \"known\" donors, or just the donor being predicted in the training set, but still maintain the time split idea. </p>",
      "rawMarkdown": "> The private consists of 3 known donors and 1 unknown donor at a future date. So I think the validation dataset for cross-validation should consist of 2 known donors and 1 unknown donor at a future date.\n\n\n\nI agree with this for predicting the unseen donor. For the others, I think you should have either all three \"known\" donors, or just the donor being predicted in the training set, but still maintain the time split idea. ",
      "replies": [
        {
          "id": 1981977,
          "postDate": "2022-10-11T06:59:38.613Z",
          "content": "<p>Take a look on a  proposal on CV :<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860</a></p>",
          "rawMarkdown": "Take a look on a  proposal on CV :\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860"
        }
      ]
    },
    {
      "id": 2025695,
      "postDate": "2022-11-11T12:03:15.980Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2014707,
      "postDate": "2022-11-02T19:12:24.307Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1984971,
      "author_name": "Ryota",
      "author_url": "",
      "post_date": "2022-10-13T02:36:28.803000",
      "content": "<p>I now am using GroupbyKFold with a combined donor and date feature.<br>\nI am currently thinking that as long as Train and Test do not include the same donor on the same day, it should be fine.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1979935,
      "author_name": "Alexander Chervov",
      "author_url": "",
      "post_date": "2022-10-09T20:26:48.453000",
      "content": "<p>My proposal on CV is the following:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1981953,
      "author_name": "Chris Miles",
      "author_url": "",
      "post_date": "2022-10-11T06:48:15.990000",
      "content": "<blockquote>\n  <p>The private consists of 3 known donors and 1 unknown donor at a future date. So I think the validation dataset for cross-validation should consist of 2 known donors and 1 unknown donor at a future date.</p>\n</blockquote>\n<p>I agree with this for predicting the unseen donor. For the others, I think you should have either all three \"known\" donors, or just the donor being predicted in the training set, but still maintain the time split idea. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1981977,
          "author_name": "Alexander Chervov",
          "author_url": "",
          "post_date": "2022-10-11T06:59:38.613000",
          "content": "<p>Take a look on a  proposal on CV :<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2025695,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-11T12:03:15.980000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2014707,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-02T19:12:24.307000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1975774": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6658947%2F5f191ab34a8c773e32bc1730f0103232%2F2022-10-07%2011.07.17.png?generation=1665108495487036&alt=media)\n\n\nHi, I am new to KAGGLE.\nAs discussed [here](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202), I think that the correct cross-validation that reflects private is necessary to win the prize, because of the different way of dividing public and private data.\n\nThe private consists of 3 known donors and 1 unknown donor at a future date. So I think the validation dataset for cross-validation should consist of 2 known donors and 1 unknown donor at a future date.\n\nIs this concept correct?\nAnd please share your cross-validation strategies.",
    "1984971": "I now am using GroupbyKFold with a combined donor and date feature.\nI am currently thinking that as long as Train and Test do not include the same donor on the same day, it should be fine.",
    "1979935": "My proposal on CV is the following:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860",
    "1981953": "> The private consists of 3 known donors and 1 unknown donor at a future date. So I think the validation dataset for cross-validation should consist of 2 known donors and 1 unknown donor at a future date.\n\n\n\nI agree with this for predicting the unseen donor. For the others, I think you should have either all three \"known\" donors, or just the donor being predicted in the training set, but still maintain the time split idea. ",
    "2025695": "",
    "2014707": ""
  }
}