{
  "id": 268562,
  "title": "TFRecord for faster training and using TPU ",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/268562",
  "author_name": "Kaveh Shahhosseini",
  "post_date": "2021-08-27T20:13:45.681000",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As you may know, TFRecord stores the data in a binary file format and can have a significant impact on the performance of training time of your model. <strong>Binary data takes up less space on disk, takes less time to copy and can be read much more efficiently from disk</strong>.[<a href=\"https://medium.com/mostly-ai/tensorflow-records-what-they-are-and-how-to-use-them-c46bc4bbb564\" target=\"_blank\">1</a>]</p>\n<p>Another <strong>advantage of using TFRecord is that you can leverage TPU</strong> for training your model much faster, otherwise reading inputs will be a bottleneck for TPU, since TPUs are very fast.  </p>\n<p>In order to convert DICOMs to TFRecord I have created a notebook which you can find it <a href=\"https://www.kaggle.com/kavehshahhosseini/rsna-brain-tumor-convert-dicom-to-tfrecord\" target=\"_blank\">here</a>. But don't worry, I have created a dataset from output files, which you can find it <a href=\"https://www.kaggle.com/kavehshahhosseini/rsna-brain-tumor-classification-tfrecords\" target=\"_blank\">here</a>, and import it to your kernel and use it for faster training. </p>\n<p>There is also a notebook, which I have provided an example of how use it, <a href=\"https://www.kaggle.com/kavehshahhosseini/how-to-load-and-use-data\" target=\"_blank\">here</a>. </p>\n<p>Note that, the dataset has 2 files: 1 for training with 465 samples and 1 for validation with 117 samples. All samples stored in the <code>(128,128,32,4)</code> shape. Last dimension (channel) is for modalities. So it includes FLAIR, T1W and… . </p>\n<p>Also note that you can merge all samples (train and validation samples) in 1 files by something like <code>dataset=train_set.concatenate(valid_set)</code>.</p>",
  "messages": [
    {
      "id": 1493370,
      "postDate": "2021-08-27T20:13:45.683Z",
      "content": "<p>As you may know, TFRecord stores the data in a binary file format and can have a significant impact on the performance of training time of your model. <strong>Binary data takes up less space on disk, takes less time to copy and can be read much more efficiently from disk</strong>.[<a href=\"https://medium.com/mostly-ai/tensorflow-records-what-they-are-and-how-to-use-them-c46bc4bbb564\" target=\"_blank\">1</a>]</p>\n<p>Another <strong>advantage of using TFRecord is that you can leverage TPU</strong> for training your model much faster, otherwise reading inputs will be a bottleneck for TPU, since TPUs are very fast.  </p>\n<p>In order to convert DICOMs to TFRecord I have created a notebook which you can find it <a href=\"https://www.kaggle.com/kavehshahhosseini/rsna-brain-tumor-convert-dicom-to-tfrecord\" target=\"_blank\">here</a>. But don't worry, I have created a dataset from output files, which you can find it <a href=\"https://www.kaggle.com/kavehshahhosseini/rsna-brain-tumor-classification-tfrecords\" target=\"_blank\">here</a>, and import it to your kernel and use it for faster training. </p>\n<p>There is also a notebook, which I have provided an example of how use it, <a href=\"https://www.kaggle.com/kavehshahhosseini/how-to-load-and-use-data\" target=\"_blank\">here</a>. </p>\n<p>Note that, the dataset has 2 files: 1 for training with 465 samples and 1 for validation with 117 samples. All samples stored in the <code>(128,128,32,4)</code> shape. Last dimension (channel) is for modalities. So it includes FLAIR, T1W and… . </p>\n<p>Also note that you can merge all samples (train and validation samples) in 1 files by something like <code>dataset=train_set.concatenate(valid_set)</code>.</p>",
      "rawMarkdown": "As you may know, TFRecord stores the data in a binary file format and can have a significant impact on the performance of training time of your model. **Binary data takes up less space on disk, takes less time to copy and can be read much more efficiently from disk**.[[1](https://medium.com/mostly-ai/tensorflow-records-what-they-are-and-how-to-use-them-c46bc4bbb564)]\n\nAnother **advantage of using TFRecord is that you can leverage TPU** for training your model much faster, otherwise reading inputs will be a bottleneck for TPU, since TPUs are very fast.  \n\nIn order to convert DICOMs to TFRecord I have created a notebook which you can find it [here](https://www.kaggle.com/kavehshahhosseini/rsna-brain-tumor-convert-dicom-to-tfrecord). But don't worry, I have created a dataset from output files, which you can find it [here](https://www.kaggle.com/kavehshahhosseini/rsna-brain-tumor-classification-tfrecords), and import it to your kernel and use it for faster training. \n\nThere is also a notebook, which I have provided an example of how use it, [here](https://www.kaggle.com/kavehshahhosseini/how-to-load-and-use-data). \n\nNote that, the dataset has 2 files: 1 for training with 465 samples and 1 for validation with 117 samples. All samples stored in the `(128,128,32,4)` shape. Last dimension (channel) is for modalities. So it includes FLAIR, T1W and... . \n\nAlso note that you can merge all samples (train and validation samples) in 1 files by something like `dataset=train_set.concatenate(valid_set)`.",
      "votes": 5
    },
    {
      "id": 1495189,
      "postDate": "2021-08-29T11:31:16.077Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1495295,
          "postDate": "2021-08-29T12:54:24.157Z",
          "content": "<p>No. Just wanted to get everything ready. </p>",
          "rawMarkdown": "No. Just wanted to get everything ready. "
        },
        {
          "id": 1495411,
          "postDate": "2021-08-29T14:41:13.927Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1495189,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-29T11:31:16.077000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1495295,
          "author_name": "Kaveh Shahhosseini",
          "author_url": "",
          "post_date": "2021-08-29T12:54:24.157000",
          "content": "<p>No. Just wanted to get everything ready. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1495411,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-08-29T14:41:13.927000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1493370": "As you may know, TFRecord stores the data in a binary file format and can have a significant impact on the performance of training time of your model. **Binary data takes up less space on disk, takes less time to copy and can be read much more efficiently from disk**.[[1](https://medium.com/mostly-ai/tensorflow-records-what-they-are-and-how-to-use-them-c46bc4bbb564)]\n\nAnother **advantage of using TFRecord is that you can leverage TPU** for training your model much faster, otherwise reading inputs will be a bottleneck for TPU, since TPUs are very fast.  \n\nIn order to convert DICOMs to TFRecord I have created a notebook which you can find it [here](https://www.kaggle.com/kavehshahhosseini/rsna-brain-tumor-convert-dicom-to-tfrecord). But don't worry, I have created a dataset from output files, which you can find it [here](https://www.kaggle.com/kavehshahhosseini/rsna-brain-tumor-classification-tfrecords), and import it to your kernel and use it for faster training. \n\nThere is also a notebook, which I have provided an example of how use it, [here](https://www.kaggle.com/kavehshahhosseini/how-to-load-and-use-data). \n\nNote that, the dataset has 2 files: 1 for training with 465 samples and 1 for validation with 117 samples. All samples stored in the `(128,128,32,4)` shape. Last dimension (channel) is for modalities. So it includes FLAIR, T1W and... . \n\nAlso note that you can merge all samples (train and validation samples) in 1 files by something like `dataset=train_set.concatenate(valid_set)`.",
    "1495189": ""
  }
}