{
  "id": 226737,
  "title": "Images packed in TFRecord files for TPUs",
  "url": "/competitions/herbarium-2021-fgvc8/discussion/226737",
  "author_name": "Luigi Saetta",
  "post_date": "2021-03-17T15:24:34.472000",
  "votes": 12,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>in order to make effective use of TPU one approach is to pre-process and pack images and labels in <strong>TFRecord</strong> files.<br>\nI'm going to explore this approach in this competition, and I'll publish a set of datasets, with different resolutions.<br>\nI have started publishing an initial dataset: <a href=\"https://www.kaggle.com/luigisaetta/herb2021-256\" target=\"_blank\">HERB2021-256</a>. Now the dataset contains only <strong>500k</strong> images of the total train set, but I hope soon to publish a new, complete version.<br>\nI'll publish, as soon as it works, a Notebook explaining how to read TFRecord, with a classification model based on EfficientNet.</p>\n<p>If anyone decides to try with this dataset, I'm greatly interested to hear comments and results.</p>",
  "messages": [
    {
      "id": 1242384,
      "postDate": "2021-03-17T15:24:34.473Z",
      "content": "<p>Hi,</p>\n<p>in order to make effective use of TPU one approach is to pre-process and pack images and labels in <strong>TFRecord</strong> files.<br>\nI'm going to explore this approach in this competition, and I'll publish a set of datasets, with different resolutions.<br>\nI have started publishing an initial dataset: <a href=\"https://www.kaggle.com/luigisaetta/herb2021-256\" target=\"_blank\">HERB2021-256</a>. Now the dataset contains only <strong>500k</strong> images of the total train set, but I hope soon to publish a new, complete version.<br>\nI'll publish, as soon as it works, a Notebook explaining how to read TFRecord, with a classification model based on EfficientNet.</p>\n<p>If anyone decides to try with this dataset, I'm greatly interested to hear comments and results.</p>",
      "rawMarkdown": "Hi,\n\nin order to make effective use of TPU one approach is to pre-process and pack images and labels in **TFRecord** files.\nI'm going to explore this approach in this competition, and I'll publish a set of datasets, with different resolutions.\nI have started publishing an initial dataset: [HERB2021-256](https://www.kaggle.com/luigisaetta/herb2021-256). Now the dataset contains only **500k** images of the total train set, but I hope soon to publish a new, complete version.\nI'll publish, as soon as it works, a Notebook explaining how to read TFRecord, with a classification model based on EfficientNet.\n\nIf anyone decides to try with this dataset, I'm greatly interested to hear comments and results.\n",
      "votes": 11
    },
    {
      "id": 1243732,
      "postDate": "2021-03-18T13:04:24.703Z",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/luigisaetta\" target=\"_blank\">@luigisaetta</a>, looking forward to the complete dataset. Ran out of GPU for this week 😅. Having TFRecords would be helpful.</p>",
      "rawMarkdown": "Thanks, @luigisaetta, looking forward to the complete dataset. Ran out of GPU for this week 😅. Having TFRecords would be helpful.",
      "votes": 1,
      "replies": [
        {
          "id": 1245338,
          "postDate": "2021-03-19T17:37:39.417Z",
          "content": "<p>why don't you try with TPU? It is much faster</p>",
          "rawMarkdown": "why don't you try with TPU? It is much faster"
        },
        {
          "id": 1245795,
          "postDate": "2021-03-20T07:00:46.950Z",
          "content": "<p>AFAIK, Keras just plays well with TPU.  I haven't worked with TFRecords much. I tried making them on my own just couldn't. Hopefully can do some stuff once there's a public TFRecords dataset available.</p>\n<p>Also, was saving TPU for Plant Pathology 😅</p>",
          "rawMarkdown": "AFAIK, Keras just plays well with TPU.  I haven't worked with TFRecords much. I tried making them on my own just couldn't. Hopefully can do some stuff once there's a public TFRecords dataset available.\n\nAlso, was saving TPU for Plant Pathology 😅",
          "votes": 1
        }
      ]
    },
    {
      "id": 1243323,
      "postDate": "2021-03-18T06:30:51.600Z",
      "content": "<p><a href=\"https://www.kaggle.com/luigisaetta\" target=\"_blank\">@luigisaetta</a> Thanks for the TF Records dataset Luigi</p>",
      "rawMarkdown": "@luigisaetta Thanks for the TF Records dataset Luigi",
      "votes": 1
    },
    {
      "id": 1242799,
      "postDate": "2021-03-17T20:13:58.047Z",
      "content": "<p>Training on the dataset (500K images) achieve 50% accuracy. Good start.</p>",
      "rawMarkdown": "Training on the dataset (500K images) achieve 50% accuracy. Good start.",
      "votes": 1
    },
    {
      "id": 1242796,
      "postDate": "2021-03-17T20:11:56.720Z",
      "content": "<p>I have also added a second dataset with <strong>all test images</strong>: <a href=\"https://www.kaggle.com/luigisaetta/herb2021-test-256\" target=\"_blank\">Test dataset</a></p>",
      "rawMarkdown": "I have also added a second dataset with **all test images**: [Test dataset](https://www.kaggle.com/luigisaetta/herb2021-test-256)",
      "votes": 1
    },
    {
      "id": 1548209,
      "postDate": "2021-10-18T04:44:04.120Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/luigisaetta\" target=\"_blank\">@luigisaetta</a>, thanks so much for your tfrecords. I finally was able to train using them. It does ok in terms of performance 0.5 test F1 score. I also found your 384 tfrecords and I am training on those. I wonder if you have available your code for tfrecords creation? Will let you know how the 384 goes. Thanks so much! </p>",
      "rawMarkdown": "Hey @luigisaetta, thanks so much for your tfrecords. I finally was able to train using them. It does ok in terms of performance 0.5 test F1 score. I also found your 384 tfrecords and I am training on those. I wonder if you have available your code for tfrecords creation? Will let you know how the 384 goes. Thanks so much! "
    },
    {
      "id": 1278220,
      "postDate": "2021-04-19T16:56:32.643Z",
      "content": "<p>I have updated the dataset. Now it contains all the train images.</p>",
      "rawMarkdown": "I have updated the dataset. Now it contains all the train images."
    },
    {
      "id": 1255699,
      "postDate": "2021-03-29T04:51:12.113Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/luigisaetta\" target=\"_blank\">@luigisaetta</a>, I was wondering how did you manage to create a <code>feature_map</code> for parsing the TFRecords ?</p>",
      "rawMarkdown": "Hey @luigisaetta, I was wondering how did you manage to create a `feature_map` for parsing the TFRecords ?"
    },
    {
      "id": 1261355,
      "postDate": "2021-04-03T01:53:22.163Z",
      "content": "<p>Thanks for the initial dataset.</p>",
      "rawMarkdown": "Thanks for the initial dataset."
    }
  ],
  "comments": [
    {
      "id": 1243732,
      "author_name": "Saurav Maheshkar ☕️",
      "author_url": "",
      "post_date": "2021-03-18T13:04:24.703000",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/luigisaetta\" target=\"_blank\">@luigisaetta</a>, looking forward to the complete dataset. Ran out of GPU for this week 😅. Having TFRecords would be helpful.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1245338,
          "author_name": "Luigi Saetta",
          "author_url": "",
          "post_date": "2021-03-19T17:37:39.417000",
          "content": "<p>why don't you try with TPU? It is much faster</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1245795,
          "author_name": "Saurav Maheshkar ☕️",
          "author_url": "",
          "post_date": "2021-03-20T07:00:46.950000",
          "content": "<p>AFAIK, Keras just plays well with TPU.  I haven't worked with TFRecords much. I tried making them on my own just couldn't. Hopefully can do some stuff once there's a public TFRecords dataset available.</p>\n<p>Also, was saving TPU for Plant Pathology 😅</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1243323,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-03-18T06:30:51.600000",
      "content": "<p><a href=\"https://www.kaggle.com/luigisaetta\" target=\"_blank\">@luigisaetta</a> Thanks for the TF Records dataset Luigi</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1242799,
      "author_name": "Luigi Saetta",
      "author_url": "",
      "post_date": "2021-03-17T20:13:58.047000",
      "content": "<p>Training on the dataset (500K images) achieve 50% accuracy. Good start.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1242796,
      "author_name": "Luigi Saetta",
      "author_url": "",
      "post_date": "2021-03-17T20:11:56.720000",
      "content": "<p>I have also added a second dataset with <strong>all test images</strong>: <a href=\"https://www.kaggle.com/luigisaetta/herb2021-test-256\" target=\"_blank\">Test dataset</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1548209,
      "author_name": "Ivan Felipe Rodriguez",
      "author_url": "",
      "post_date": "2021-10-18T04:44:04.120000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/luigisaetta\" target=\"_blank\">@luigisaetta</a>, thanks so much for your tfrecords. I finally was able to train using them. It does ok in terms of performance 0.5 test F1 score. I also found your 384 tfrecords and I am training on those. I wonder if you have available your code for tfrecords creation? Will let you know how the 384 goes. Thanks so much! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1278220,
      "author_name": "Luigi Saetta",
      "author_url": "",
      "post_date": "2021-04-19T16:56:32.643000",
      "content": "<p>I have updated the dataset. Now it contains all the train images.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1255699,
      "author_name": "Saurav Maheshkar ☕️",
      "author_url": "",
      "post_date": "2021-03-29T04:51:12.113000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/luigisaetta\" target=\"_blank\">@luigisaetta</a>, I was wondering how did you manage to create a <code>feature_map</code> for parsing the TFRecords ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1261355,
      "author_name": "ToPu",
      "author_url": "",
      "post_date": "2021-04-03T01:53:22.163000",
      "content": "<p>Thanks for the initial dataset.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1242384": "Hi,\n\nin order to make effective use of TPU one approach is to pre-process and pack images and labels in **TFRecord** files.\nI'm going to explore this approach in this competition, and I'll publish a set of datasets, with different resolutions.\nI have started publishing an initial dataset: [HERB2021-256](https://www.kaggle.com/luigisaetta/herb2021-256). Now the dataset contains only **500k** images of the total train set, but I hope soon to publish a new, complete version.\nI'll publish, as soon as it works, a Notebook explaining how to read TFRecord, with a classification model based on EfficientNet.\n\nIf anyone decides to try with this dataset, I'm greatly interested to hear comments and results.\n",
    "1243732": "Thanks, @luigisaetta, looking forward to the complete dataset. Ran out of GPU for this week 😅. Having TFRecords would be helpful.",
    "1243323": "@luigisaetta Thanks for the TF Records dataset Luigi",
    "1242799": "Training on the dataset (500K images) achieve 50% accuracy. Good start.",
    "1242796": "I have also added a second dataset with **all test images**: [Test dataset](https://www.kaggle.com/luigisaetta/herb2021-test-256)",
    "1548209": "Hey @luigisaetta, thanks so much for your tfrecords. I finally was able to train using them. It does ok in terms of performance 0.5 test F1 score. I also found your 384 tfrecords and I am training on those. I wonder if you have available your code for tfrecords creation? Will let you know how the 384 goes. Thanks so much! ",
    "1278220": "I have updated the dataset. Now it contains all the train images.",
    "1255699": "Hey @luigisaetta, I was wondering how did you manage to create a `feature_map` for parsing the TFRecords ?",
    "1261355": "Thanks for the initial dataset."
  }
}