{
  "id": 174938,
  "title": "Leak-free KFold CV",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/174938",
  "author_name": "",
  "post_date": "2020-08-16T10:24:38.375965800Z",
  "votes": 23,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>Leak-free KFold CV</h1>\n<p>As in the SIIM-ISIC competition, there are a lot more images than unique patients. More importantly, there are much more tabular training samples than unique patients which suggest that several patients had their FVC tested multiple times.</p>\n<h2>No overlapping group of patients across folds</h2>\n<p>Hence, one of the trap we want to avoid is to introduce a data leak by having one patient present in multiple folds.</p>\n<p>As shown, by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\" target=\"_blank\">SIIM-ISIC competition</a>, I try to replicate <a href=\"https://www.kaggle.com/rftexas/osic-eda-leak-free-k-fold-cv-lgb-baseline\" target=\"_blank\">here</a> his Triple Stratified Leak-Free CV partition.</p>\n<h2>Balance Patient count distribution across folds</h2>\n<p>I've made sure to balance patient count distribution across folds.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1302247%2F3e452073a7f12e31cee511c68c3d6ad3%2FScreen%20Shot%202020-08-16%20at%2012.23.01%20PM.png?generation=1597573405015832&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "972196",
      "postDate": "08/16/2020 10:24:38",
      "content": "<h1>Leak-free KFold CV</h1>\n<p>As in the SIIM-ISIC competition, there are a lot more images than unique patients. More importantly, there are much more tabular training samples than unique patients which suggest that several patients had their FVC tested multiple times.</p>\n<h2>No overlapping group of patients across folds</h2>\n<p>Hence, one of the trap we want to avoid is to introduce a data leak by having one patient present in multiple folds.</p>\n<p>As shown, by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\" target=\"_blank\">SIIM-ISIC competition</a>, I try to replicate <a href=\"https://www.kaggle.com/rftexas/osic-eda-leak-free-k-fold-cv-lgb-baseline\" target=\"_blank\">here</a> his Triple Stratified Leak-Free CV partition.</p>\n<h2>Balance Patient count distribution across folds</h2>\n<p>I've made sure to balance patient count distribution across folds.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1302247%2F3e452073a7f12e31cee511c68c3d6ad3%2FScreen%20Shot%202020-08-16%20at%2012.23.01%20PM.png?generation=1597573405015832&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "# Leak-free KFold CV\n\nAs in the SIIM-ISIC competition, there are a lot more images than unique patients. More importantly, there are much more tabular training samples than unique patients which suggest that several patients had their FVC tested multiple times.\n\n## No overlapping group of patients across folds\n\nHence, one of the trap we want to avoid is to introduce a data leak by having one patient present in multiple folds.\n\nAs shown, by @cdeotte in [SIIM-ISIC competition](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526), I try to replicate [here](https://www.kaggle.com/rftexas/osic-eda-leak-free-k-fold-cv-lgb-baseline) his Triple Stratified Leak-Free CV partition.\n\n## Balance Patient count distribution across folds\n\nI've made sure to balance patient count distribution across folds.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1302247%2F3e452073a7f12e31cee511c68c3d6ad3%2FScreen%20Shot%202020-08-16%20at%2012.23.01%20PM.png?generation=1597573405015832&alt=media)",
      "votes": null
    },
    {
      "id": "972253",
      "postDate": "08/16/2020 11:45:21",
      "content": "<p>Agree. I do the same. Upvote.</p>",
      "rawMarkdown": "Agree. I do the same. Upvote.",
      "votes": null
    },
    {
      "id": "973458",
      "postDate": "08/17/2020 10:04:08",
      "content": "<p>great sir <br>\nThanks for sharing 😊😊 </p>",
      "rawMarkdown": "great sir \nThanks for sharing 😊😊",
      "votes": null
    },
    {
      "id": "973980",
      "postDate": "08/17/2020 16:38:34",
      "content": "<p>Good point! thanks for sharing.</p>",
      "rawMarkdown": "Good point! thanks for sharing.",
      "votes": null
    },
    {
      "id": "974018",
      "postDate": "08/17/2020 17:03:08",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rftexas\" target=\"_blank\">@rftexas</a>, thanks for sharing. To be clear, what you did was GroupKFold, but you did not stratify by other variables, such as Age, Sex, etc. - is that correct? Given that the goal is to avoid data leakage (and I agree with the idea of making sure Patients don't overlap between groups), I'm confused by why you focused on the count distribution to support this form of validation.</p>",
      "rawMarkdown": "Hi @rftexas, thanks for sharing. To be clear, what you did was GroupKFold, but you did not stratify by other variables, such as Age, Sex, etc. - is that correct? Given that the goal is to avoid data leakage (and I agree with the idea of making sure Patients don't overlap between groups), I'm confused by why you focused on the count distribution to support this form of validation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 972253,
      "author_name": "eliasgreen",
      "author_url": "",
      "post_date": "08/16/2020 11:45:21",
      "content": "<p>Agree. I do the same. Upvote.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 973458,
      "author_name": "palaksood97",
      "author_url": "",
      "post_date": "08/17/2020 10:04:08",
      "content": "<p>great sir <br>\nThanks for sharing 😊😊 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 973980,
      "author_name": "amolayari",
      "author_url": "",
      "post_date": "08/17/2020 16:38:34",
      "content": "<p>Good point! thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974018,
      "author_name": "jjinho",
      "author_url": "",
      "post_date": "08/17/2020 17:03:08",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rftexas\" target=\"_blank\">@rftexas</a>, thanks for sharing. To be clear, what you did was GroupKFold, but you did not stratify by other variables, such as Age, Sex, etc. - is that correct? Given that the goal is to avoid data leakage (and I agree with the idea of making sure Patients don't overlap between groups), I'm confused by why you focused on the count distribution to support this form of validation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "972196": "# Leak-free KFold CV\n\nAs in the SIIM-ISIC competition, there are a lot more images than unique patients. More importantly, there are much more tabular training samples than unique patients which suggest that several patients had their FVC tested multiple times.\n\n## No overlapping group of patients across folds\n\nHence, one of the trap we want to avoid is to introduce a data leak by having one patient present in multiple folds.\n\nAs shown, by @cdeotte in [SIIM-ISIC competition](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526), I try to replicate [here](https://www.kaggle.com/rftexas/osic-eda-leak-free-k-fold-cv-lgb-baseline) his Triple Stratified Leak-Free CV partition.\n\n## Balance Patient count distribution across folds\n\nI've made sure to balance patient count distribution across folds.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1302247%2F3e452073a7f12e31cee511c68c3d6ad3%2FScreen%20Shot%202020-08-16%20at%2012.23.01%20PM.png?generation=1597573405015832&alt=media)",
    "972253": "Agree. I do the same. Upvote.",
    "973458": "great sir \nThanks for sharing 😊😊",
    "973980": "Good point! thanks for sharing.",
    "974018": "Hi @rftexas, thanks for sharing. To be clear, what you did was GroupKFold, but you did not stratify by other variables, such as Age, Sex, etc. - is that correct? Given that the goal is to avoid data leakage (and I agree with the idea of making sure Patients don't overlap between groups), I'm confused by why you focused on the count distribution to support this form of validation."
  },
  "source": "meta"
}