{
  "id": 356786,
  "title": "GroupKFold is the right scheme.... or not?",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/356786",
  "author_name": "",
  "post_date": "2022-10-01T22:41:18.734760300Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>After thinking about this for some time, I think the best CV scheme is GroupKFold using each file as a group (or game_id); however, keep in mind that this is computationally expensive.</p>\n<p>My reasons are:</p>\n<ol>\n<li>Training and validating on the same game will inflate our scores, records from the same game cannot be considered independent.</li>\n<li>The description doesn't mention that each virtual player corresponds to the same person in real life. Model must understand the game situation regardless of who is behind the wheel.</li>\n<li>All chunks share more o less the same distribution of events. See my <a href=\"https://www.kaggle.com/code/jcaliz/tps-oct22-quickstart-eda-keras-baseline?scriptVersionId=106990515#Events-per-game\" target=\"_blank\">notebook</a> for more information.</li>\n<li>Lag features will not leak data.</li>\n</ol>\n<p>What do you think?</p>",
  "messages": [
    {
      "id": "1966356",
      "postDate": "10/01/2022 22:41:18",
      "content": "<p>After thinking about this for some time, I think the best CV scheme is GroupKFold using each file as a group (or game_id); however, keep in mind that this is computationally expensive.</p>\n<p>My reasons are:</p>\n<ol>\n<li>Training and validating on the same game will inflate our scores, records from the same game cannot be considered independent.</li>\n<li>The description doesn't mention that each virtual player corresponds to the same person in real life. Model must understand the game situation regardless of who is behind the wheel.</li>\n<li>All chunks share more o less the same distribution of events. See my <a href=\"https://www.kaggle.com/code/jcaliz/tps-oct22-quickstart-eda-keras-baseline?scriptVersionId=106990515#Events-per-game\" target=\"_blank\">notebook</a> for more information.</li>\n<li>Lag features will not leak data.</li>\n</ol>\n<p>What do you think?</p>",
      "rawMarkdown": "After thinking about this for some time, I think the best CV scheme is GroupKFold using each file as a group (or game_id); however, keep in mind that this is computationally expensive.\n\nMy reasons are:\n\n1. Training and validating on the same game will inflate our scores, records from the same game cannot be considered independent.\n2. The description doesn't mention that each virtual player corresponds to the same person in real life. Model must understand the game situation regardless of who is behind the wheel.\n3. All chunks share more o less the same distribution of events. See my [notebook](https://www.kaggle.com/code/jcaliz/tps-oct22-quickstart-eda-keras-baseline?scriptVersionId=106990515#Events-per-game) for more information.\n4. Lag features will not leak data.\n\nWhat do you think?",
      "votes": null
    },
    {
      "id": "1987673",
      "postDate": "10/14/2022 19:28:53",
      "content": "<p>Oh dear… This should explain why there is a whole bunch of kernels getting ~0.18 loss on validation sets but then only get a ~0.19-0.21 loss on the LB. </p>",
      "rawMarkdown": "Oh dear... This should explain why there is a whole bunch of kernels getting ~0.18 loss on validation sets but then only get a ~0.19-0.21 loss on the LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1987673,
      "author_name": "fabianbong",
      "author_url": "",
      "post_date": "10/14/2022 19:28:53",
      "content": "<p>Oh dear… This should explain why there is a whole bunch of kernels getting ~0.18 loss on validation sets but then only get a ~0.19-0.21 loss on the LB. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1966356": "After thinking about this for some time, I think the best CV scheme is GroupKFold using each file as a group (or game_id); however, keep in mind that this is computationally expensive.\n\nMy reasons are:\n\n1. Training and validating on the same game will inflate our scores, records from the same game cannot be considered independent.\n2. The description doesn't mention that each virtual player corresponds to the same person in real life. Model must understand the game situation regardless of who is behind the wheel.\n3. All chunks share more o less the same distribution of events. See my [notebook](https://www.kaggle.com/code/jcaliz/tps-oct22-quickstart-eda-keras-baseline?scriptVersionId=106990515#Events-per-game) for more information.\n4. Lag features will not leak data.\n\nWhat do you think?",
    "1987673": "Oh dear... This should explain why there is a whole bunch of kernels getting ~0.18 loss on validation sets but then only get a ~0.19-0.21 loss on the LB."
  },
  "source": "meta"
}