{
  "id": 393570,
  "title": "Quality of CV matters: fold split with less player duplication made CV/LB correlated",
  "url": "/competitions/nfl-player-contact-detection/discussion/393570",
  "author_name": "Bilzard",
  "post_date": "2023-03-09T21:05:06.033000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi. I found interesting experiment result about the train/test split of this competition, so I share it.</p>\n<p>As I wrote in my <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391607\" target=\"_blank\">solution writeup</a>, my solution was seriouly suffered from uncorrelated CV/LB. One of my hypothesis of this is that <strong>train and test datasets consist of data with different players or teams</strong>. To verify this hypothesis, I tried to split games in train dataset into folds with less duplication of players. (Detail of the method is explained in <a href=\"https://www.kaggle.com/code/tatamikenn/nfl-pcd-fold-split-with-less-player-duplication\" target=\"_blank\">this notebook</a>).</p>\n<p>Using this split and submitted my 3-staged pipeline solution, the CVs and LBs becomes correlated. This experimental result supports the hypothesis that train and test sets consists of different team/players.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F218f17f407db05930caa6433652fafc3%2FScreen%20Shot%202023-03-10%20at%2013.47.13.png?generation=1678423650605611&amp;alt=media\" alt=\"\"></p>\n<h2>Dataset</h2>\n<p>For anyone want to verify their solution on the late sub, I shared <a href=\"https://www.kaggle.com/datasets/tatamikenn/nfl-pcd-fold-split-of-less-player-duplication\" target=\"_blank\">my fold split dataset</a>.<br>\nYou can test if this split make your solution CV/LB correlated.<br>\nIf you test using this split, please inform me the result. If this split makes CV/LB correlated on other solutions, it supports my hypothesis more robustly.</p>",
  "messages": [
    {
      "id": 2175444,
      "postDate": "2023-03-09T21:05:06.033Z",
      "content": "<p>Hi. I found interesting experiment result about the train/test split of this competition, so I share it.</p>\n<p>As I wrote in my <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391607\" target=\"_blank\">solution writeup</a>, my solution was seriouly suffered from uncorrelated CV/LB. One of my hypothesis of this is that <strong>train and test datasets consist of data with different players or teams</strong>. To verify this hypothesis, I tried to split games in train dataset into folds with less duplication of players. (Detail of the method is explained in <a href=\"https://www.kaggle.com/code/tatamikenn/nfl-pcd-fold-split-with-less-player-duplication\" target=\"_blank\">this notebook</a>).</p>\n<p>Using this split and submitted my 3-staged pipeline solution, the CVs and LBs becomes correlated. This experimental result supports the hypothesis that train and test sets consists of different team/players.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F218f17f407db05930caa6433652fafc3%2FScreen%20Shot%202023-03-10%20at%2013.47.13.png?generation=1678423650605611&amp;alt=media\" alt=\"\"></p>\n<h2>Dataset</h2>\n<p>For anyone want to verify their solution on the late sub, I shared <a href=\"https://www.kaggle.com/datasets/tatamikenn/nfl-pcd-fold-split-of-less-player-duplication\" target=\"_blank\">my fold split dataset</a>.<br>\nYou can test if this split make your solution CV/LB correlated.<br>\nIf you test using this split, please inform me the result. If this split makes CV/LB correlated on other solutions, it supports my hypothesis more robustly.</p>",
      "rawMarkdown": "Hi. I found interesting experiment result about the train/test split of this competition, so I share it.\n\nAs I wrote in my [solution writeup](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391607), my solution was seriouly suffered from uncorrelated CV/LB. One of my hypothesis of this is that **train and test datasets consist of data with different players or teams**. To verify this hypothesis, I tried to split games in train dataset into folds with less duplication of players. (Detail of the method is explained in [this notebook](https://www.kaggle.com/code/tatamikenn/nfl-pcd-fold-split-with-less-player-duplication)).\n\nUsing this split and submitted my 3-staged pipeline solution, the CVs and LBs becomes correlated. This experimental result supports the hypothesis that train and test sets consists of different team/players.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F218f17f407db05930caa6433652fafc3%2FScreen%20Shot%202023-03-10%20at%2013.47.13.png?generation=1678423650605611&alt=media)\n\n## Dataset\n\nFor anyone want to verify their solution on the late sub, I shared [my fold split dataset](https://www.kaggle.com/datasets/tatamikenn/nfl-pcd-fold-split-of-less-player-duplication).\nYou can test if this split make your solution CV/LB correlated.\nIf you test using this split, please inform me the result. If this split makes CV/LB correlated on other solutions, it supports my hypothesis more robustly.",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2175444": "Hi. I found interesting experiment result about the train/test split of this competition, so I share it.\n\nAs I wrote in my [solution writeup](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391607), my solution was seriouly suffered from uncorrelated CV/LB. One of my hypothesis of this is that **train and test datasets consist of data with different players or teams**. To verify this hypothesis, I tried to split games in train dataset into folds with less duplication of players. (Detail of the method is explained in [this notebook](https://www.kaggle.com/code/tatamikenn/nfl-pcd-fold-split-with-less-player-duplication)).\n\nUsing this split and submitted my 3-staged pipeline solution, the CVs and LBs becomes correlated. This experimental result supports the hypothesis that train and test sets consists of different team/players.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F218f17f407db05930caa6433652fafc3%2FScreen%20Shot%202023-03-10%20at%2013.47.13.png?generation=1678423650605611&alt=media)\n\n## Dataset\n\nFor anyone want to verify their solution on the late sub, I shared [my fold split dataset](https://www.kaggle.com/datasets/tatamikenn/nfl-pcd-fold-split-of-less-player-duplication).\nYou can test if this split make your solution CV/LB correlated.\nIf you test using this split, please inform me the result. If this split makes CV/LB correlated on other solutions, it supports my hypothesis more robustly."
  }
}