{
  "id": 349830,
  "title": "Cite Private Data Day 7",
  "url": "/competitions/open-problems-multimodal/discussion/349830",
  "author_name": "",
  "post_date": "2022-09-02T23:41:44.318861200Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>If you check the train data for cite, you will see that we have days 2, 3, 4. In the data description they say the following: \"The private test set comprises samples only from day 7\"</p>\n<p>Does this means that for the best kfold strategy would be to do GroupKFold and use day as the group (3 folds)?</p>\n<p>Any thoughts?</p>",
  "messages": [
    {
      "id": "1924310",
      "postDate": "09/02/2022 23:41:44",
      "content": "<p>If you check the train data for cite, you will see that we have days 2, 3, 4. In the data description they say the following: \"The private test set comprises samples only from day 7\"</p>\n<p>Does this means that for the best kfold strategy would be to do GroupKFold and use day as the group (3 folds)?</p>\n<p>Any thoughts?</p>",
      "rawMarkdown": "If you check the train data for cite, you will see that we have days 2, 3, 4. In the data description they say the following: \"The private test set comprises samples only from day 7\"\n\nDoes this means that for the best kfold strategy would be to do GroupKFold and use day as the group (3 folds)?\n\nAny thoughts?",
      "votes": null
    },
    {
      "id": "1924549",
      "postDate": "09/03/2022 06:46:32",
      "content": "<p>Of course, one must take that into account ! <br>\nBut there are many further issues to be resolved with the validation scheme. For example cells at day 4 of course are better predictors for day 7, cell types should be taken into account, for Multiome you predict not the day 7, but the day 10 and so on.<br>\nSome one can hopefully  spent some time on  details and make publicly available reasonable folds splits.<br>\nAs Chris Deotte did for MoA for example. <br>\nIt would be a favor for community.<br>\nPS<br>\nAlso take into account that public LB is quite different from private LB and shake-up is inevitable.<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202</a></p>",
      "rawMarkdown": "Of course, one must take that into account ! \nBut there are many further issues to be resolved with the validation scheme. For example cells at day 4 of course are better predictors for day 7, cell types should be taken into account, for Multiome you predict not the day 7, but the day 10 and so on.\nSome one can hopefully  spent some time on  details and make publicly available reasonable folds splits.\nAs Chris Deotte did for MoA for example. \nIt would be a favor for community.\nPS\nAlso take into account that public LB is quite different from private LB and shake-up is inevitable.\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1924549,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/03/2022 06:46:32",
      "content": "<p>Of course, one must take that into account ! <br>\nBut there are many further issues to be resolved with the validation scheme. For example cells at day 4 of course are better predictors for day 7, cell types should be taken into account, for Multiome you predict not the day 7, but the day 10 and so on.<br>\nSome one can hopefully  spent some time on  details and make publicly available reasonable folds splits.<br>\nAs Chris Deotte did for MoA for example. <br>\nIt would be a favor for community.<br>\nPS<br>\nAlso take into account that public LB is quite different from private LB and shake-up is inevitable.<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1924310": "If you check the train data for cite, you will see that we have days 2, 3, 4. In the data description they say the following: \"The private test set comprises samples only from day 7\"\n\nDoes this means that for the best kfold strategy would be to do GroupKFold and use day as the group (3 folds)?\n\nAny thoughts?",
    "1924549": "Of course, one must take that into account ! \nBut there are many further issues to be resolved with the validation scheme. For example cells at day 4 of course are better predictors for day 7, cell types should be taken into account, for Multiome you predict not the day 7, but the day 10 and so on.\nSome one can hopefully  spent some time on  details and make publicly available reasonable folds splits.\nAs Chris Deotte did for MoA for example. \nIt would be a favor for community.\nPS\nAlso take into account that public LB is quite different from private LB and shake-up is inevitable.\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202"
  },
  "source": "meta"
}