{
  "id": 202332,
  "title": "HuBMAP tfrecords with images made by @Iafoss",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/202332",
  "author_name": "",
  "post_date": "2020-12-09T13:26:55.312639500Z",
  "votes": 17,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hello Kagglers!<br>\nFirst of all, I would like to thank the organizers for preparing this competion. Secondly, I thank <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> for preparing images on which you can train your models. One of the main tricks that occurs in this beautiful dataset is resize tiles from 1024 to 256. </p>\n<p>Therefore, I have prepared training datasets in tfrecords format in sizes 128, 256 and 512 (all downsized from 1024) and available for GPU/TPU here: <a href=\"https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-128\" target=\"_blank\">128</a>,  <a href=\"https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-256\" target=\"_blank\">256</a> and here <a href=\"https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-512\" target=\"_blank\">512</a>. </p>\n<p>All of them contains 8 training tfrecs - each for training image, with bytes structure:<br>\n<code>\nimage = tf.reshape( tf.io.decode_raw(single_example['image'],out_type=np.dtype('uint8')), (DIM,DIM, 3)); mask =  tf.reshape(tf.io.decode_raw(single_example['mask'],out_type='bool'),(DIM,DIM,1))</code></p>\n<p>Example notebook using these datasets is available <a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train\" target=\"_blank\">here</a> with full submission example <a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-subm\" target=\"_blank\">here</a> scoring: </p>\n<ul>\n<li>efficientunetb0: LB .834</li>\n<li>efficientunetb1: LB .830</li>\n<li>efficientunetb4: LB .839</li>\n</ul>\n<p>It is worth noting that these datasets have assigned to GCS_PATHs, so we can train on TPU also outside Kaggle, i.e. on Colab ;)<br>\nGreetings!</p>",
  "messages": [
    {
      "id": "1107193",
      "postDate": "12/09/2020 13:26:55",
      "content": "<p>Hello Kagglers!<br>\nFirst of all, I would like to thank the organizers for preparing this competion. Secondly, I thank <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> for preparing images on which you can train your models. One of the main tricks that occurs in this beautiful dataset is resize tiles from 1024 to 256. </p>\n<p>Therefore, I have prepared training datasets in tfrecords format in sizes 128, 256 and 512 (all downsized from 1024) and available for GPU/TPU here: <a href=\"https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-128\" target=\"_blank\">128</a>,  <a href=\"https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-256\" target=\"_blank\">256</a> and here <a href=\"https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-512\" target=\"_blank\">512</a>. </p>\n<p>All of them contains 8 training tfrecs - each for training image, with bytes structure:<br>\n<code>\nimage = tf.reshape( tf.io.decode_raw(single_example['image'],out_type=np.dtype('uint8')), (DIM,DIM, 3)); mask =  tf.reshape(tf.io.decode_raw(single_example['mask'],out_type='bool'),(DIM,DIM,1))</code></p>\n<p>Example notebook using these datasets is available <a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train\" target=\"_blank\">here</a> with full submission example <a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-subm\" target=\"_blank\">here</a> scoring: </p>\n<ul>\n<li>efficientunetb0: LB .834</li>\n<li>efficientunetb1: LB .830</li>\n<li>efficientunetb4: LB .839</li>\n</ul>\n<p>It is worth noting that these datasets have assigned to GCS_PATHs, so we can train on TPU also outside Kaggle, i.e. on Colab ;)<br>\nGreetings!</p>",
      "rawMarkdown": "Hello Kagglers!\nFirst of all, I would like to thank the organizers for preparing this competion. Secondly, I thank @iafoss for preparing images on which you can train your models. One of the main tricks that occurs in this beautiful dataset is resize tiles from 1024 to 256. \n\nTherefore, I have prepared training datasets in tfrecords format in sizes 128, 256 and 512 (all downsized from 1024) and available for GPU/TPU here: [128](https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-128),  [256](https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-256) and here [512](https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-512). \n\nAll of them contains 8 training tfrecs - each for training image, with bytes structure:\n`\nimage = tf.reshape( tf.io.decode_raw(single_example['image'],out_type=np.dtype('uint8')), (DIM,DIM, 3)); mask =  tf.reshape(tf.io.decode_raw(single_example['mask'],out_type='bool'),(DIM,DIM,1))`\n\nExample notebook using these datasets is available [here](https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train) with full submission example [here](https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-subm) scoring: \n- efficientunetb0: LB .834\n- efficientunetb1: LB .830\n- efficientunetb4: LB .839\n\nIt is worth noting that these datasets have assigned to GCS_PATHs, so we can train on TPU also outside Kaggle, i.e. on Colab ;)\nGreetings!",
      "votes": null
    },
    {
      "id": "1109090",
      "postDate": "12/11/2020 10:18:08",
      "content": "<p>Important update: in tfrec dataset I've moved files to train folder, so updated GCS_PATHS should look like this (fixed in version 11, already commited):<br>\n<code>ALL_TRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train/*.tfrec')</code></p>",
      "rawMarkdown": "Important update: in tfrec dataset I've moved files to train folder, so updated GCS_PATHS should look like this (fixed in version 11, already commited):\n`ALL_TRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train/*.tfrec')`",
      "votes": null
    },
    {
      "id": "1109103",
      "postDate": "12/11/2020 10:31:31",
      "content": "<p>Thank you. Can you share noutbook for make dataset (1024 to 512)?</p>",
      "rawMarkdown": "Thank you. Can you share noutbook for make dataset (1024 to 512)?",
      "votes": null
    },
    {
      "id": "1109110",
      "postDate": "12/11/2020 10:43:10",
      "content": "<p>Sure, I added you as a collaborator with a preview ability.<br>\nI didn't publish it, because it's heavily based on a great <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> notebook <a href=\"https://www.kaggle.com/iafoss/256x256-images/\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Sure, I added you as a collaborator with a preview ability.\nI didn't publish it, because it's heavily based on a great @iafoss notebook [here](https://www.kaggle.com/iafoss/256x256-images/)",
      "votes": null
    },
    {
      "id": "1109142",
      "postDate": "12/11/2020 11:11:57",
      "content": "<blockquote>\n  <p>Sure, I added you as a collaborator with a preview ability.<br>\n  Thank you.</p>\n</blockquote>",
      "rawMarkdown": "> Sure, I added you as a collaborator with a preview ability.\nThank you.",
      "votes": null
    },
    {
      "id": "1122609",
      "postDate": "12/22/2020 15:00:09",
      "content": "<p><a href=\"https://www.kaggle.com/wrrosa\" target=\"_blank\">@wrrosa</a> great work. But I can't find the notebook to create the dataset. Could you please share it?</p>",
      "rawMarkdown": "wrrosa great work. But I can't find the notebook to create the dataset. Could you please share it?",
      "votes": null
    },
    {
      "id": "1123009",
      "postDate": "12/22/2020 21:07:02",
      "content": "<p>It is in the commentary to the training notebook, when I find the moment I publish it</p>",
      "rawMarkdown": "It is in the commentary to the training notebook, when I find the moment I publish it",
      "votes": null
    },
    {
      "id": "1133156",
      "postDate": "12/31/2020 01:50:20",
      "content": "<p>Hi Wojtek, I have been using your notebook (HuBMAP: TPU with EfficientUNet) and notice that you also have the overlap image dataset. Question: how many pixels on the overlap?</p>",
      "rawMarkdown": "Hi Wojtek, I have been using your notebook (HuBMAP: TPU with EfficientUNet) and notice that you also have the overlap image dataset. Question: how many pixels on the overlap?",
      "votes": null
    },
    {
      "id": "1144575",
      "postDate": "01/08/2021 14:42:52",
      "content": "<p>Training: overlap = WINDOW//2<br>\nInference: min_overlap = 300 (WINDOW = 1024)<br>\nMore info here: <a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-tfrecs\" target=\"_blank\">https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-tfrecs</a></p>",
      "rawMarkdown": "Training: overlap = WINDOW//2\nInference: min_overlap = 300 (WINDOW = 1024)\nMore info here: https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-tfrecs",
      "votes": null
    },
    {
      "id": "1144578",
      "postDate": "01/08/2021 14:43:36",
      "content": "<p>See comment above, I've just published it</p>",
      "rawMarkdown": "See comment above, I've just published it",
      "votes": null
    },
    {
      "id": "1181586",
      "postDate": "02/02/2021 02:03:32",
      "content": "<p>Have you tried running this in colab?</p>",
      "rawMarkdown": "Have you tried running this in colab?",
      "votes": null
    },
    {
      "id": "1181687",
      "postDate": "02/02/2021 04:12:15",
      "content": "<p>Yes it works</p>",
      "rawMarkdown": "Yes it works",
      "votes": null
    },
    {
      "id": "1181878",
      "postDate": "02/02/2021 07:26:36",
      "content": "<p>I tried running your kernel in Colab, but I get an error. Do I need to make any changes to the kernel prior to running it on Colab?</p>",
      "rawMarkdown": "I tried running your kernel in Colab, but I get an error. Do I need to make any changes to the kernel prior to running it on Colab?",
      "votes": null
    },
    {
      "id": "1181898",
      "postDate": "02/02/2021 07:38:23",
      "content": "<p>Not likely, you may just need to reduce the batch_size (I think the RAM available on Colab is less stable than on Kaggle). It's worth adding that on Colab the current version of Tensorflow is 2.4 and on Kaggle it's 2.3, but I haven't noticed any problems with that.</p>",
      "rawMarkdown": "Not likely, you may just need to reduce the batch_size (I think the RAM available on Colab is less stable than on Kaggle). It's worth adding that on Colab the current version of Tensorflow is 2.4 and on Kaggle it's 2.3, but I haven't noticed any problems with that.",
      "votes": null
    },
    {
      "id": "1181905",
      "postDate": "02/02/2021 07:41:14",
      "content": "<p>BTW what kind of error is it?</p>",
      "rawMarkdown": "BTW what kind of error is it?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1109090,
      "author_name": "wrrosa",
      "author_url": "",
      "post_date": "12/11/2020 10:18:08",
      "content": "<p>Important update: in tfrec dataset I've moved files to train folder, so updated GCS_PATHS should look like this (fixed in version 11, already commited):<br>\n<code>ALL_TRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train/*.tfrec')</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1109103,
      "author_name": "aleksandrkruchinin",
      "author_url": "",
      "post_date": "12/11/2020 10:31:31",
      "content": "<p>Thank you. Can you share noutbook for make dataset (1024 to 512)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1109110,
          "author_name": "wrrosa",
          "author_url": "",
          "post_date": "12/11/2020 10:43:10",
          "content": "<p>Sure, I added you as a collaborator with a preview ability.<br>\nI didn't publish it, because it's heavily based on a great <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> notebook <a href=\"https://www.kaggle.com/iafoss/256x256-images/\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1109142,
      "author_name": "aleksandrkruchinin",
      "author_url": "",
      "post_date": "12/11/2020 11:11:57",
      "content": "<blockquote>\n  <p>Sure, I added you as a collaborator with a preview ability.<br>\n  Thank you.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1122609,
      "author_name": "awsaf49",
      "author_url": "",
      "post_date": "12/22/2020 15:00:09",
      "content": "<p><a href=\"https://www.kaggle.com/wrrosa\" target=\"_blank\">@wrrosa</a> great work. But I can't find the notebook to create the dataset. Could you please share it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123009,
          "author_name": "wrrosa",
          "author_url": "",
          "post_date": "12/22/2020 21:07:02",
          "content": "<p>It is in the commentary to the training notebook, when I find the moment I publish it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144578,
          "author_name": "wrrosa",
          "author_url": "",
          "post_date": "01/08/2021 14:43:36",
          "content": "<p>See comment above, I've just published it</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1133156,
      "author_name": "zungmann",
      "author_url": "",
      "post_date": "12/31/2020 01:50:20",
      "content": "<p>Hi Wojtek, I have been using your notebook (HuBMAP: TPU with EfficientUNet) and notice that you also have the overlap image dataset. Question: how many pixels on the overlap?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144575,
          "author_name": "wrrosa",
          "author_url": "",
          "post_date": "01/08/2021 14:42:52",
          "content": "<p>Training: overlap = WINDOW//2<br>\nInference: min_overlap = 300 (WINDOW = 1024)<br>\nMore info here: <a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-tfrecs\" target=\"_blank\">https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-tfrecs</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1181586,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "02/02/2021 02:03:32",
      "content": "<p>Have you tried running this in colab?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1181687,
          "author_name": "wrrosa",
          "author_url": "",
          "post_date": "02/02/2021 04:12:15",
          "content": "<p>Yes it works</p>",
          "votes": null,
          "replies": [
            {
              "id": 1181878,
              "author_name": "dskswu",
              "author_url": "",
              "post_date": "02/02/2021 07:26:36",
              "content": "<p>I tried running your kernel in Colab, but I get an error. Do I need to make any changes to the kernel prior to running it on Colab?</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1181898,
          "author_name": "wrrosa",
          "author_url": "",
          "post_date": "02/02/2021 07:38:23",
          "content": "<p>Not likely, you may just need to reduce the batch_size (I think the RAM available on Colab is less stable than on Kaggle). It's worth adding that on Colab the current version of Tensorflow is 2.4 and on Kaggle it's 2.3, but I haven't noticed any problems with that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1181905,
          "author_name": "wrrosa",
          "author_url": "",
          "post_date": "02/02/2021 07:41:14",
          "content": "<p>BTW what kind of error is it?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1107193": "Hello Kagglers!\nFirst of all, I would like to thank the organizers for preparing this competion. Secondly, I thank @iafoss for preparing images on which you can train your models. One of the main tricks that occurs in this beautiful dataset is resize tiles from 1024 to 256. \n\nTherefore, I have prepared training datasets in tfrecords format in sizes 128, 256 and 512 (all downsized from 1024) and available for GPU/TPU here: [128](https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-128),  [256](https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-256) and here [512](https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-512). \n\nAll of them contains 8 training tfrecs - each for training image, with bytes structure:\n`\nimage = tf.reshape( tf.io.decode_raw(single_example['image'],out_type=np.dtype('uint8')), (DIM,DIM, 3)); mask =  tf.reshape(tf.io.decode_raw(single_example['mask'],out_type='bool'),(DIM,DIM,1))`\n\nExample notebook using these datasets is available [here](https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train) with full submission example [here](https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-subm) scoring: \n- efficientunetb0: LB .834\n- efficientunetb1: LB .830\n- efficientunetb4: LB .839\n\nIt is worth noting that these datasets have assigned to GCS_PATHs, so we can train on TPU also outside Kaggle, i.e. on Colab ;)\nGreetings!",
    "1109090": "Important update: in tfrec dataset I've moved files to train folder, so updated GCS_PATHS should look like this (fixed in version 11, already commited):\n`ALL_TRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train/*.tfrec')`",
    "1109103": "Thank you. Can you share noutbook for make dataset (1024 to 512)?",
    "1109110": "Sure, I added you as a collaborator with a preview ability.\nI didn't publish it, because it's heavily based on a great @iafoss notebook [here](https://www.kaggle.com/iafoss/256x256-images/)",
    "1109142": "> Sure, I added you as a collaborator with a preview ability.\nThank you.",
    "1122609": "wrrosa great work. But I can't find the notebook to create the dataset. Could you please share it?",
    "1123009": "It is in the commentary to the training notebook, when I find the moment I publish it",
    "1133156": "Hi Wojtek, I have been using your notebook (HuBMAP: TPU with EfficientUNet) and notice that you also have the overlap image dataset. Question: how many pixels on the overlap?",
    "1144575": "Training: overlap = WINDOW//2\nInference: min_overlap = 300 (WINDOW = 1024)\nMore info here: https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-tfrecs",
    "1144578": "See comment above, I've just published it",
    "1181586": "Have you tried running this in colab?",
    "1181687": "Yes it works",
    "1181878": "I tried running your kernel in Colab, but I get an error. Do I need to make any changes to the kernel prior to running it on Colab?",
    "1181898": "Not likely, you may just need to reduce the batch_size (I think the RAM available on Colab is less stable than on Kaggle). It's worth adding that on Colab the current version of Tensorflow is 2.4 and on Kaggle it's 2.3, but I haven't noticed any problems with that.",
    "1181905": "BTW what kind of error is it?"
  },
  "source": "meta"
}