{
  "id": 221008,
  "title": "Best strategy to create folds ?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/221008",
  "author_name": "",
  "post_date": "2021-02-20T15:12:44.896116Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am still a newbie in computer vision. I just wanted to know some of the best strategies to split the folds for this data. The dataset is highly skewed and I observed there are around 30k images in train data but only 3255 patients. So, a patient will have multiple images, some have in the range of 150+ and some of them only one.<br>\nA good CV strategy will definitely help in this problem.</p>",
  "messages": [
    {
      "id": "1211798",
      "postDate": "02/20/2021 15:12:44",
      "content": "<p>I am still a newbie in computer vision. I just wanted to know some of the best strategies to split the folds for this data. The dataset is highly skewed and I observed there are around 30k images in train data but only 3255 patients. So, a patient will have multiple images, some have in the range of 150+ and some of them only one.<br>\nA good CV strategy will definitely help in this problem.</p>",
      "rawMarkdown": "I am still a newbie in computer vision. I just wanted to know some of the best strategies to split the folds for this data. The dataset is highly skewed and I observed there are around 30k images in train data but only 3255 patients. So, a patient will have multiple images, some have in the range of 150+ and some of them only one.\nA good CV strategy will definitely help in this problem.",
      "votes": null
    },
    {
      "id": "1212789",
      "postDate": "02/21/2021 15:45:01",
      "content": "<p>As the data is highly imbalanced, its good if you could stratify your splits. you could refer to my previous discussion <a href=\"discussion\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215327</a>.<br>\nIn my opinion, split made by <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> in this notebook <a href=\"notebook\" target=\"_blank\">https://www.kaggle.com/underwearfitting/how-to-properly-split-folds</a> will work well in this competition.<a href=\"url\" target=\"_blank\"></a></p>",
      "rawMarkdown": "As the data is highly imbalanced, its good if you could stratify your splits. you could refer to my previous discussion [https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215327](discussion).\nIn my opinion, split made by @underwearfitting in this notebook [https://www.kaggle.com/underwearfitting/how-to-properly-split-folds](notebook) will work well in this competition.[](url)",
      "votes": null
    },
    {
      "id": "1212897",
      "postDate": "02/21/2021 17:46:08",
      "content": "<p>Thank you for sharing. I just went through the notebook and it is actually a good way to prevent data leakage in multiple folds.</p>",
      "rawMarkdown": "Thank you for sharing. I just went through the notebook and it is actually a good way to prevent data leakage in multiple folds.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1212789,
      "author_name": "nischaydnk",
      "author_url": "",
      "post_date": "02/21/2021 15:45:01",
      "content": "<p>As the data is highly imbalanced, its good if you could stratify your splits. you could refer to my previous discussion <a href=\"discussion\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215327</a>.<br>\nIn my opinion, split made by <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> in this notebook <a href=\"notebook\" target=\"_blank\">https://www.kaggle.com/underwearfitting/how-to-properly-split-folds</a> will work well in this competition.<a href=\"url\" target=\"_blank\"></a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1212897,
          "author_name": "atharvaingle",
          "author_url": "",
          "post_date": "02/21/2021 17:46:08",
          "content": "<p>Thank you for sharing. I just went through the notebook and it is actually a good way to prevent data leakage in multiple folds.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1211798": "I am still a newbie in computer vision. I just wanted to know some of the best strategies to split the folds for this data. The dataset is highly skewed and I observed there are around 30k images in train data but only 3255 patients. So, a patient will have multiple images, some have in the range of 150+ and some of them only one.\nA good CV strategy will definitely help in this problem.",
    "1212789": "As the data is highly imbalanced, its good if you could stratify your splits. you could refer to my previous discussion [https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215327](discussion).\nIn my opinion, split made by @underwearfitting in this notebook [https://www.kaggle.com/underwearfitting/how-to-properly-split-folds](notebook) will work well in this competition.[](url)",
    "1212897": "Thank you for sharing. I just went through the notebook and it is actually a good way to prevent data leakage in multiple folds."
  },
  "source": "meta"
}