{
  "id": 219344,
  "title": "Add this to your TF pipeline to run 6x faster",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/219344",
  "author_name": "",
  "post_date": "2021-02-14T14:15:36.638848900Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<h3>Hello!</h3>\n<p>As the competition end is approaching, one should think about using the last 30h of available TPU quota wisely. As you know, using TFRecords can really boost up training, but sometimes those provided in the original competition dataset are not enough, e.g.:</p>\n<ul>\n<li>you need to add more data, but external data is only available in .jpeg format</li>\n<li>you need to prepare augmented data instead of creating it on the fly</li>\n<li>you need to perform knowledge distillation on soft labels, etc.</li>\n</ul>\n<p>A one-stop solution is <code>from_tensor_slices</code> method, which I came across in many notebooks. This makes it just as easy as feeding a DataFrame into the network but results in slower (really slower) runtime. E.g. training EfficientNetB4 on dataset of 25K 512x512 images takes nearly 6x more time:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>from tensor slices</th>\n<th>from tfrecord files</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Max time per epoch</td>\n<td>996s</td>\n<td><strong>132s</strong></td>\n</tr>\n<tr>\n<td>Average time per epoch</td>\n<td>502s</td>\n<td><strong>81s</strong></td>\n</tr>\n</tbody>\n</table>\n<p>On the other hand, serializing this dataset to TFRecords can be done in just 30 lines of code. Taking only 10 minutes on GPU, this step would save you up to a few hours on TPU when training an ensemble or a large model.</p>\n<p>So I've decided to create <strong><a href=\"https://www.kaggle.com/nickuzmenkov/tfrecord-lab\" target=\"_blank\">this short notebook</a></strong> as an example, where I take <a href=\"https://www.kaggle.com/nickuzmenkov/cassava-leaf-disease-soft-targets-09-model\" target=\"_blank\">soft targets</a> for this competition and create new TFRecords. Hope this helps someone. Happy coding!</p>",
  "messages": [
    {
      "id": "1200250",
      "postDate": "02/14/2021 14:15:36",
      "content": "<h3>Hello!</h3>\n<p>As the competition end is approaching, one should think about using the last 30h of available TPU quota wisely. As you know, using TFRecords can really boost up training, but sometimes those provided in the original competition dataset are not enough, e.g.:</p>\n<ul>\n<li>you need to add more data, but external data is only available in .jpeg format</li>\n<li>you need to prepare augmented data instead of creating it on the fly</li>\n<li>you need to perform knowledge distillation on soft labels, etc.</li>\n</ul>\n<p>A one-stop solution is <code>from_tensor_slices</code> method, which I came across in many notebooks. This makes it just as easy as feeding a DataFrame into the network but results in slower (really slower) runtime. E.g. training EfficientNetB4 on dataset of 25K 512x512 images takes nearly 6x more time:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>from tensor slices</th>\n<th>from tfrecord files</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Max time per epoch</td>\n<td>996s</td>\n<td><strong>132s</strong></td>\n</tr>\n<tr>\n<td>Average time per epoch</td>\n<td>502s</td>\n<td><strong>81s</strong></td>\n</tr>\n</tbody>\n</table>\n<p>On the other hand, serializing this dataset to TFRecords can be done in just 30 lines of code. Taking only 10 minutes on GPU, this step would save you up to a few hours on TPU when training an ensemble or a large model.</p>\n<p>So I've decided to create <strong><a href=\"https://www.kaggle.com/nickuzmenkov/tfrecord-lab\" target=\"_blank\">this short notebook</a></strong> as an example, where I take <a href=\"https://www.kaggle.com/nickuzmenkov/cassava-leaf-disease-soft-targets-09-model\" target=\"_blank\">soft targets</a> for this competition and create new TFRecords. Hope this helps someone. Happy coding!</p>",
      "rawMarkdown": "### Hello! \n\nAs the competition end is approaching, one should think about using the last 30h of available TPU quota wisely. As you know, using TFRecords can really boost up training, but sometimes those provided in the original competition dataset are not enough, e.g.:\n* you need to add more data, but external data is only available in .jpeg format\n* you need to prepare augmented data instead of creating it on the fly\n* you need to perform knowledge distillation on soft labels, etc.\n\nA one-stop solution is `from_tensor_slices` method, which I came across in many notebooks. This makes it just as easy as feeding a DataFrame into the network but results in slower (really slower) runtime. E.g. training EfficientNetB4 on dataset of 25K 512x512 images takes nearly 6x more time:\n\n| | from tensor slices | from tfrecord files |\n| --- | --- | --- |\n| Max time per epoch | 996s | **132s** |\n| Average time per epoch | 502s | **81s** | \n\nOn the other hand, serializing this dataset to TFRecords can be done in just 30 lines of code. Taking only 10 minutes on GPU, this step would save you up to a few hours on TPU when training an ensemble or a large model.\n\nSo I've decided to create **[this short notebook](https://www.kaggle.com/nickuzmenkov/tfrecord-lab)** as an example, where I take [soft targets](https://www.kaggle.com/nickuzmenkov/cassava-leaf-disease-soft-targets-09-model) for this competition and create new TFRecords. Hope this helps someone. Happy coding!",
      "votes": null
    },
    {
      "id": "1200361",
      "postDate": "02/14/2021 15:44:13",
      "content": "<p>Thank you really much, I will finally give these soft targets a try!</p>",
      "rawMarkdown": "Thank you really much, I will finally give these soft targets a try!",
      "votes": null
    },
    {
      "id": "1204158",
      "postDate": "02/15/2021 22:16:02",
      "content": "<p>Thank you for sharing  this tips 👌</p>",
      "rawMarkdown": "Thank you for sharing  this tips 👌",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1200361,
      "author_name": "mviola",
      "author_url": "",
      "post_date": "02/14/2021 15:44:13",
      "content": "<p>Thank you really much, I will finally give these soft targets a try!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1204158,
      "author_name": "milobele",
      "author_url": "",
      "post_date": "02/15/2021 22:16:02",
      "content": "<p>Thank you for sharing  this tips 👌</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1200250": "### Hello! \n\nAs the competition end is approaching, one should think about using the last 30h of available TPU quota wisely. As you know, using TFRecords can really boost up training, but sometimes those provided in the original competition dataset are not enough, e.g.:\n* you need to add more data, but external data is only available in .jpeg format\n* you need to prepare augmented data instead of creating it on the fly\n* you need to perform knowledge distillation on soft labels, etc.\n\nA one-stop solution is `from_tensor_slices` method, which I came across in many notebooks. This makes it just as easy as feeding a DataFrame into the network but results in slower (really slower) runtime. E.g. training EfficientNetB4 on dataset of 25K 512x512 images takes nearly 6x more time:\n\n| | from tensor slices | from tfrecord files |\n| --- | --- | --- |\n| Max time per epoch | 996s | **132s** |\n| Average time per epoch | 502s | **81s** | \n\nOn the other hand, serializing this dataset to TFRecords can be done in just 30 lines of code. Taking only 10 minutes on GPU, this step would save you up to a few hours on TPU when training an ensemble or a large model.\n\nSo I've decided to create **[this short notebook](https://www.kaggle.com/nickuzmenkov/tfrecord-lab)** as an example, where I take [soft targets](https://www.kaggle.com/nickuzmenkov/cassava-leaf-disease-soft-targets-09-model) for this competition and create new TFRecords. Hope this helps someone. Happy coding!",
    "1200361": "Thank you really much, I will finally give these soft targets a try!",
    "1204158": "Thank you for sharing  this tips 👌"
  },
  "source": "meta"
}