{
  "id": 221415,
  "title": "Speedup your training by creating your own TFRecord folds",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/221415",
  "author_name": "",
  "post_date": "2021-02-22T16:16:23.937776800Z",
  "votes": 13,
  "comment_count": 1,
  "views": 0,
  "content": "<h3>Hello!</h3>\n<p>As mentioned numerous times on the competition forum, the most proper way to organize folds is as follows:</p>\n<ul>\n<li>all folds must share nearly the same number of samples</li>\n<li>label-wise distributions must be kept close to those in the entire dataset, as there are some extremely rare cases (e.g. <code>ETT - Abnormal</code>)</li>\n<li>no <code>PatientID</code> must appear in different folds to prevent data leaks</li>\n</ul>\n<p>I've found two solutions so far: one by <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> <strong><a href=\"https://www.kaggle.com/underwearfitting/how-to-properly-split-folds\" target=\"_blank\">here</a></strong> and another by <a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a> <strong><a href=\"https://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3\" target=\"_blank\">here</a></strong>.</p>\n<p>If you are a <strong>TensorFlow</strong> user, having this splits the most straightforward way to start your efficient data workflow is <code>tf.data.Dataset.from_tensor_slices</code> which makes it just as easy as feeding a <code>DataFrame</code> into the network but results in longer (really longer) runtime. On the other hand, TFRecords serialization from 30K images takes just about 10-15 minutes and saves up to few hours on TPU when training large models.</p>\n<p>So I've decided to make this <strong><a href=\"https://www.kaggle.com/nickuzmenkov/ranzcr-clip-groupkfold-with-tfrecords?scriptVersionId=54973556\" target=\"_blank\">short starter notebook</a></strong> to either create GroupKFold TFRecords with any resolution on your taste on the fly or save them as a dataset for later use. Outputs for 600x600 image resolution are also published to <strong><a href=\"https://www.kaggle.com/nickuzmenkov/ranzcr-clip-kfold-tfrecords\" target=\"_blank\">this dataset</a></strong>.</p>\n<p>Hope this helps someone. Happy coding!</p>",
  "messages": [
    {
      "id": "1214168",
      "postDate": "02/22/2021 16:16:23",
      "content": "<h3>Hello!</h3>\n<p>As mentioned numerous times on the competition forum, the most proper way to organize folds is as follows:</p>\n<ul>\n<li>all folds must share nearly the same number of samples</li>\n<li>label-wise distributions must be kept close to those in the entire dataset, as there are some extremely rare cases (e.g. <code>ETT - Abnormal</code>)</li>\n<li>no <code>PatientID</code> must appear in different folds to prevent data leaks</li>\n</ul>\n<p>I've found two solutions so far: one by <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> <strong><a href=\"https://www.kaggle.com/underwearfitting/how-to-properly-split-folds\" target=\"_blank\">here</a></strong> and another by <a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a> <strong><a href=\"https://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3\" target=\"_blank\">here</a></strong>.</p>\n<p>If you are a <strong>TensorFlow</strong> user, having this splits the most straightforward way to start your efficient data workflow is <code>tf.data.Dataset.from_tensor_slices</code> which makes it just as easy as feeding a <code>DataFrame</code> into the network but results in longer (really longer) runtime. On the other hand, TFRecords serialization from 30K images takes just about 10-15 minutes and saves up to few hours on TPU when training large models.</p>\n<p>So I've decided to make this <strong><a href=\"https://www.kaggle.com/nickuzmenkov/ranzcr-clip-groupkfold-with-tfrecords?scriptVersionId=54973556\" target=\"_blank\">short starter notebook</a></strong> to either create GroupKFold TFRecords with any resolution on your taste on the fly or save them as a dataset for later use. Outputs for 600x600 image resolution are also published to <strong><a href=\"https://www.kaggle.com/nickuzmenkov/ranzcr-clip-kfold-tfrecords\" target=\"_blank\">this dataset</a></strong>.</p>\n<p>Hope this helps someone. Happy coding!</p>",
      "rawMarkdown": "### Hello!\n\nAs mentioned numerous times on the competition forum, the most proper way to organize folds is as follows:\n* all folds must share nearly the same number of samples\n* label-wise distributions must be kept close to those in the entire dataset, as there are some extremely rare cases (e.g. `ETT - Abnormal`)\n* no `PatientID` must appear in different folds to prevent data leaks\n\nI've found two solutions so far: one by @underwearfitting **[here](https://www.kaggle.com/underwearfitting/how-to-properly-split-folds)** and another by @virilo **[here](https://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3)**.\n\nIf you are a **TensorFlow** user, having this splits the most straightforward way to start your efficient data workflow is `tf.data.Dataset.from_tensor_slices` which makes it just as easy as feeding a `DataFrame` into the network but results in longer (really longer) runtime. On the other hand, TFRecords serialization from 30K images takes just about 10-15 minutes and saves up to few hours on TPU when training large models.\n\nSo I've decided to make this **[short starter notebook](https://www.kaggle.com/nickuzmenkov/ranzcr-clip-groupkfold-with-tfrecords?scriptVersionId=54973556)** to either create GroupKFold TFRecords with any resolution on your taste on the fly or save them as a dataset for later use. Outputs for 600x600 image resolution are also published to **[this dataset](https://www.kaggle.com/nickuzmenkov/ranzcr-clip-kfold-tfrecords)**.\n\nHope this helps someone. Happy coding!",
      "votes": null
    },
    {
      "id": "1215056",
      "postDate": "02/23/2021 09:54:19",
      "content": "<p>Nice work! Thanks for a list of curated notebooks and dataset. I needed them for starting with this competition!</p>",
      "rawMarkdown": "Nice work! Thanks for a list of curated notebooks and dataset. I needed them for starting with this competition!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1215056,
      "author_name": "jiny333",
      "author_url": "",
      "post_date": "02/23/2021 09:54:19",
      "content": "<p>Nice work! Thanks for a list of curated notebooks and dataset. I needed them for starting with this competition!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1214168": "### Hello!\n\nAs mentioned numerous times on the competition forum, the most proper way to organize folds is as follows:\n* all folds must share nearly the same number of samples\n* label-wise distributions must be kept close to those in the entire dataset, as there are some extremely rare cases (e.g. `ETT - Abnormal`)\n* no `PatientID` must appear in different folds to prevent data leaks\n\nI've found two solutions so far: one by @underwearfitting **[here](https://www.kaggle.com/underwearfitting/how-to-properly-split-folds)** and another by @virilo **[here](https://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3)**.\n\nIf you are a **TensorFlow** user, having this splits the most straightforward way to start your efficient data workflow is `tf.data.Dataset.from_tensor_slices` which makes it just as easy as feeding a `DataFrame` into the network but results in longer (really longer) runtime. On the other hand, TFRecords serialization from 30K images takes just about 10-15 minutes and saves up to few hours on TPU when training large models.\n\nSo I've decided to make this **[short starter notebook](https://www.kaggle.com/nickuzmenkov/ranzcr-clip-groupkfold-with-tfrecords?scriptVersionId=54973556)** to either create GroupKFold TFRecords with any resolution on your taste on the fly or save them as a dataset for later use. Outputs for 600x600 image resolution are also published to **[this dataset](https://www.kaggle.com/nickuzmenkov/ranzcr-clip-kfold-tfrecords)**.\n\nHope this helps someone. Happy coding!",
    "1215056": "Nice work! Thanks for a list of curated notebooks and dataset. I needed them for starting with this competition!"
  },
  "source": "meta"
}