{
  "id": 154475,
  "title": "Organizers, thank you <3",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/154475",
  "author_name": "",
  "post_date": "2020-05-28T13:45:22.296514Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>As I'm doing my EDA, I discovered with great joy that TFRecords file are provided. I'd like to thank the organizers for that since it will help us to build fast pipeline with TF and TPU. </p>\n\n<p>That's the first time that I see it on computer vision competitions. I appreciate that!</p>",
  "messages": [
    {
      "id": "865231",
      "postDate": "05/28/2020 13:45:22",
      "content": "<p>As I'm doing my EDA, I discovered with great joy that TFRecords file are provided. I'd like to thank the organizers for that since it will help us to build fast pipeline with TF and TPU. </p>\n\n<p>That's the first time that I see it on computer vision competitions. I appreciate that!</p>",
      "rawMarkdown": "As I'm doing my EDA, I discovered with great joy that TFRecords file are provided. I'd like to thank the organizers for that since it will help us to build fast pipeline with TF and TPU. \n\nThat's the first time that I see it on computer vision competitions. I appreciate that!",
      "votes": null
    },
    {
      "id": "865565",
      "postDate": "05/28/2020 17:45:04",
      "content": "<p>Sounds nice, what do these TFRecords specifically do?</p>",
      "rawMarkdown": "Sounds nice, what do these TFRecords specifically do?",
      "votes": null
    },
    {
      "id": "866390",
      "postDate": "05/29/2020 11:03:49",
      "content": "<p>They are protobuf objects that can store multiple images and labels at the same time so they can be loaded more efficently (instead of loading and decoding one image at a time). To maximize performance they should weight around 100MB. You can load this type of files with tf.data.TFRecordDataset (in tensorflow, obviously). It is usually a pain to create these objects, so it is nice that the organizers already provide them. Plus, it make it easier to run on TPUs. </p>",
      "rawMarkdown": "They are protobuf objects that can store multiple images and labels at the same time so they can be loaded more efficently (instead of loading and decoding one image at a time). To maximize performance they should weight around 100MB. You can load this type of files with tf.data.TFRecordDataset (in tensorflow, obviously). It is usually a pain to create these objects, so it is nice that the organizers already provide them. Plus, it make it easier to run on TPUs.",
      "votes": null
    },
    {
      "id": "868117",
      "postDate": "05/31/2020 00:09:12",
      "content": "<p>The dataset is extremely imbalanced.  Unfortunately the TFRecords containers won't let you stratify it. </p>",
      "rawMarkdown": "The dataset is extremely imbalanced.  Unfortunately the TFRecords containers won't let you stratify it.",
      "votes": null
    },
    {
      "id": "869078",
      "postDate": "05/31/2020 18:14:20",
      "content": "<p>I suggest class weighting in this regard or programming a custom generator.</p>",
      "rawMarkdown": "I suggest class weighting in this regard or programming a custom generator.",
      "votes": null
    },
    {
      "id": "869268",
      "postDate": "05/31/2020 22:03:54",
      "content": "<p>Yes that's what I've discovered. Quite frustrating...</p>",
      "rawMarkdown": "Yes that's what I've discovered. Quite frustrating...",
      "votes": null
    },
    {
      "id": "869269",
      "postDate": "05/31/2020 22:04:50",
      "content": "<p>Unfortunately, Keras custom generator does not work well with TPU. You need to pass by tf.data.Dataset API.</p>",
      "rawMarkdown": "Unfortunately, Keras custom generator does not work well with TPU. You need to pass by tf.data.Dataset API.",
      "votes": null
    },
    {
      "id": "869279",
      "postDate": "05/31/2020 22:21:38",
      "content": "<p>You can still use stratified splitted jpeg files for TPU.   That's what I am doing now.  </p>\n\n<p>The first epoch is relatively slow but it's fast for the rest</p>",
      "rawMarkdown": "You can still use stratified splitted jpeg files for TPU.   That's what I am doing now.  \n\nThe first epoch is relatively slow but it's fast for the rest",
      "votes": null
    },
    {
      "id": "869291",
      "postDate": "05/31/2020 22:46:44",
      "content": "<p><a href=\"/serigne\">@serigne</a> Are you using the original or reduced resolution jpegs? </p>",
      "rawMarkdown": "serigne Are you using the original or reduced resolution jpegs?",
      "votes": null
    },
    {
      "id": "869501",
      "postDate": "06/01/2020 03:54:43",
      "content": "<p>A slightly reduced resolution .  I am planning to use meta-data too but not yet</p>",
      "rawMarkdown": "A slightly reduced resolution .  I am planning to use meta-data too but not yet",
      "votes": null
    },
    {
      "id": "869754",
      "postDate": "06/01/2020 09:02:33",
      "content": "<p>OK sounds good! I wasn't sure if reading jpeg files was fast enough. Thanks for the tips!</p>",
      "rawMarkdown": "OK sounds good! I wasn't sure if reading jpeg files was fast enough. Thanks for the tips!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 865565,
      "author_name": "ovdnnest",
      "author_url": "",
      "post_date": "05/28/2020 17:45:04",
      "content": "<p>Sounds nice, what do these TFRecords specifically do?</p>",
      "votes": null,
      "replies": [
        {
          "id": 866390,
          "author_name": "juansensio",
          "author_url": "",
          "post_date": "05/29/2020 11:03:49",
          "content": "<p>They are protobuf objects that can store multiple images and labels at the same time so they can be loaded more efficently (instead of loading and decoding one image at a time). To maximize performance they should weight around 100MB. You can load this type of files with tf.data.TFRecordDataset (in tensorflow, obviously). It is usually a pain to create these objects, so it is nice that the organizers already provide them. Plus, it make it easier to run on TPUs. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 868117,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "05/31/2020 00:09:12",
      "content": "<p>The dataset is extremely imbalanced.  Unfortunately the TFRecords containers won't let you stratify it. </p>",
      "votes": null,
      "replies": [
        {
          "id": 869268,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "05/31/2020 22:03:54",
          "content": "<p>Yes that's what I've discovered. Quite frustrating...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869279,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/31/2020 22:21:38",
          "content": "<p>You can still use stratified splitted jpeg files for TPU.   That's what I am doing now.  </p>\n\n<p>The first epoch is relatively slow but it's fast for the rest</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869291,
          "author_name": "graf10a",
          "author_url": "",
          "post_date": "05/31/2020 22:46:44",
          "content": "<p><a href=\"/serigne\">@serigne</a> Are you using the original or reduced resolution jpegs? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869501,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "06/01/2020 03:54:43",
          "content": "<p>A slightly reduced resolution .  I am planning to use meta-data too but not yet</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869754,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "06/01/2020 09:02:33",
          "content": "<p>OK sounds good! I wasn't sure if reading jpeg files was fast enough. Thanks for the tips!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 869078,
      "author_name": "ovdnnest",
      "author_url": "",
      "post_date": "05/31/2020 18:14:20",
      "content": "<p>I suggest class weighting in this regard or programming a custom generator.</p>",
      "votes": null,
      "replies": [
        {
          "id": 869269,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "05/31/2020 22:04:50",
          "content": "<p>Unfortunately, Keras custom generator does not work well with TPU. You need to pass by tf.data.Dataset API.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "865231": "As I'm doing my EDA, I discovered with great joy that TFRecords file are provided. I'd like to thank the organizers for that since it will help us to build fast pipeline with TF and TPU. \n\nThat's the first time that I see it on computer vision competitions. I appreciate that!",
    "865565": "Sounds nice, what do these TFRecords specifically do?",
    "866390": "They are protobuf objects that can store multiple images and labels at the same time so they can be loaded more efficently (instead of loading and decoding one image at a time). To maximize performance they should weight around 100MB. You can load this type of files with tf.data.TFRecordDataset (in tensorflow, obviously). It is usually a pain to create these objects, so it is nice that the organizers already provide them. Plus, it make it easier to run on TPUs.",
    "868117": "The dataset is extremely imbalanced.  Unfortunately the TFRecords containers won't let you stratify it.",
    "869078": "I suggest class weighting in this regard or programming a custom generator.",
    "869268": "Yes that's what I've discovered. Quite frustrating...",
    "869269": "Unfortunately, Keras custom generator does not work well with TPU. You need to pass by tf.data.Dataset API.",
    "869279": "You can still use stratified splitted jpeg files for TPU.   That's what I am doing now.  \n\nThe first epoch is relatively slow but it's fast for the rest",
    "869291": "serigne Are you using the original or reduced resolution jpegs?",
    "869501": "A slightly reduced resolution .  I am planning to use meta-data too but not yet",
    "869754": "OK sounds good! I wasn't sure if reading jpeg files was fast enough. Thanks for the tips!"
  },
  "source": "meta"
}