{
  "id": 393400,
  "title": "Fold split with less player duplication",
  "url": "/competitions/nfl-player-contact-detection/discussion/393400",
  "author_name": "",
  "post_date": "2023-03-09T08:33:28.862502700Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I wrote a <a href=\"https://www.kaggle.com/code/tatamikenn/nfl-pcd-fold-split-with-less-player-duplication\" target=\"_blank\">notebook</a> to split folds with less duplication of players.<br>\nIf the train/test sets are sampled from completely different ones, your CV might only shows the overfitted performance to the train dataset. For example, if the players/teams in the train and test dataset differ, the CV shows your model's overfitted performance to the specific players/teams.</p>\n<p>I suggested a naive method to choose split set of games so that each folds contains less duplicated players if fold differs. The result shows the suggested method reduced average similarity of folds about ~4.1% comparing to the usual StratifiedGroupKFold.</p>\n<p>Although the method is too naive, I think the idea of choosing difficult fold using the specific domain knowledge is applicable to other competitions that have a risk of overfit.</p>\n<p>Enjoy!</p>",
  "messages": [
    {
      "id": "2174572",
      "postDate": "03/09/2023 08:33:28",
      "content": "<p>I wrote a <a href=\"https://www.kaggle.com/code/tatamikenn/nfl-pcd-fold-split-with-less-player-duplication\" target=\"_blank\">notebook</a> to split folds with less duplication of players.<br>\nIf the train/test sets are sampled from completely different ones, your CV might only shows the overfitted performance to the train dataset. For example, if the players/teams in the train and test dataset differ, the CV shows your model's overfitted performance to the specific players/teams.</p>\n<p>I suggested a naive method to choose split set of games so that each folds contains less duplicated players if fold differs. The result shows the suggested method reduced average similarity of folds about ~4.1% comparing to the usual StratifiedGroupKFold.</p>\n<p>Although the method is too naive, I think the idea of choosing difficult fold using the specific domain knowledge is applicable to other competitions that have a risk of overfit.</p>\n<p>Enjoy!</p>",
      "rawMarkdown": "I wrote a [notebook] to split folds with less duplication of players.\nIf the train/test sets are sampled from completely different ones, your CV might only shows the overfitted performance to the train dataset. For example, if the players/teams in the train and test dataset differ, the CV shows your model's overfitted performance to the specific players/teams.\n\nI suggested a naive method to choose split set of games so that each folds contains less duplicated players if fold differs. The result shows the suggested method reduced average similarity of folds about ~4.1% comparing to the usual StratifiedGroupKFold.\n\nAlthough the method is too naive, I think the idea of choosing difficult fold using the specific domain knowledge is applicable to other competitions that have a risk of overfit.\n\nEnjoy!\n\n[notebook]: https://www.kaggle.com/code/tatamikenn/nfl-pcd-fold-split-with-less-player-duplication",
      "votes": null
    },
    {
      "id": "2175718",
      "postDate": "03/10/2023 05:34:14",
      "content": "<p>I also shared my experiment result using this fold split.<br>\nYou can check <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/393570\" target=\"_blank\">this topic</a></p>",
      "rawMarkdown": "I also shared my experiment result using this fold split.\nYou can check [this topic](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/393570)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2175718,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "03/10/2023 05:34:14",
      "content": "<p>I also shared my experiment result using this fold split.<br>\nYou can check <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/393570\" target=\"_blank\">this topic</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2174572": "I wrote a [notebook] to split folds with less duplication of players.\nIf the train/test sets are sampled from completely different ones, your CV might only shows the overfitted performance to the train dataset. For example, if the players/teams in the train and test dataset differ, the CV shows your model's overfitted performance to the specific players/teams.\n\nI suggested a naive method to choose split set of games so that each folds contains less duplicated players if fold differs. The result shows the suggested method reduced average similarity of folds about ~4.1% comparing to the usual StratifiedGroupKFold.\n\nAlthough the method is too naive, I think the idea of choosing difficult fold using the specific domain knowledge is applicable to other competitions that have a risk of overfit.\n\nEnjoy!\n\n[notebook]: https://www.kaggle.com/code/tatamikenn/nfl-pcd-fold-split-with-less-player-duplication",
    "2175718": "I also shared my experiment result using this fold split.\nYou can check [this topic](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/393570)"
  },
  "source": "meta"
}