{
  "id": 347568,
  "title": "How to efficiently feed tiles to tensorflow dataset?",
  "url": "/competitions/hubmap-organ-segmentation/discussion/347568",
  "author_name": "",
  "post_date": "2022-08-24T16:55:41.022841300Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello! I'm using tensorflow 2 and trying to feed tiles to a Unet neural network (WIP). <br>\nI create a dataset with all relevant data from my dataframe, then map the preprocessing function on it, getting a crop of the image and the corresponding cropped mask.<br>\nI have not found a way to split data within the preprocessing function; meaning that I decode the entire image and entire mask, crop with the offset values, then decode the entire image again to get the next tile from the same .tiff… <br>\nThis looks extremely inefficient to me, does anyone know how to use the preprocessing function to return not the image and corresponding mask, but a \"list\" of these tiles from a single decode?<br>\nNotebook here:  <a href=\"https://www.kaggle.com/code/baptistejoseph/hubmap-hpa-notebook/\" target=\"_blank\">https://www.kaggle.com/code/baptistejoseph/hubmap-hpa-notebook/</a></p>",
  "messages": [
    {
      "id": "1912350",
      "postDate": "08/24/2022 16:55:41",
      "content": "<p>Hello! I'm using tensorflow 2 and trying to feed tiles to a Unet neural network (WIP). <br>\nI create a dataset with all relevant data from my dataframe, then map the preprocessing function on it, getting a crop of the image and the corresponding cropped mask.<br>\nI have not found a way to split data within the preprocessing function; meaning that I decode the entire image and entire mask, crop with the offset values, then decode the entire image again to get the next tile from the same .tiff… <br>\nThis looks extremely inefficient to me, does anyone know how to use the preprocessing function to return not the image and corresponding mask, but a \"list\" of these tiles from a single decode?<br>\nNotebook here:  <a href=\"https://www.kaggle.com/code/baptistejoseph/hubmap-hpa-notebook/\" target=\"_blank\">https://www.kaggle.com/code/baptistejoseph/hubmap-hpa-notebook/</a></p>",
      "rawMarkdown": "Hello! I'm using tensorflow 2 and trying to feed tiles to a Unet neural network (WIP). \nI create a dataset with all relevant data from my dataframe, then map the preprocessing function on it, getting a crop of the image and the corresponding cropped mask.\nI have not found a way to split data within the preprocessing function; meaning that I decode the entire image and entire mask, crop with the offset values, then decode the entire image again to get the next tile from the same .tiff... \nThis looks extremely inefficient to me, does anyone know how to use the preprocessing function to return not the image and corresponding mask, but a \"list\" of these tiles from a single decode?\nNotebook here:  https://www.kaggle.com/code/baptistejoseph/hubmap-hpa-notebook/",
      "votes": null
    },
    {
      "id": "1913648",
      "postDate": "08/25/2022 12:32:37",
      "content": "<p>Maybe one workaround is to generate multiple patches(crops) when reading a single image from .tiff file.<br>\nif you use tf.data API, you can make your tf.data.dataset object through the following process.</p>\n<p>1) read single tiff file (ex. tfio.experimental.image.decode_tiff)<br>\n2) generate N patches from the image (ex. tf.image.extract_patches or just multiple call of random_crop) -&gt; your dataset shape becomes (N,H,W,C)<br>\n3) tf.data.dataset.unbatch() -&gt; your dataset shape becomes (H,W,C)<br>\n4) tf.data.dataset.shuffle(N*m) -&gt; this is needed for random sampling (note that shuffling window size m will directly affect performance of the model)<br>\n5) tf.data.dataset.batch(batch_size)</p>\n<p>in 4), you may select large m for better training procedure, but it may lead to heavy overload in preprocessing step. the fastest way is always to use preprocessed dataset. so this is a tradeoff between easy/flexible implementation and performance/speed.</p>\n<p>hope this was helpful : )</p>",
      "rawMarkdown": "Maybe one workaround is to generate multiple patches(crops) when reading a single image from .tiff file.\nif you use tf.data API, you can make your tf.data.dataset object through the following process.\n\n1) read single tiff file (ex. tfio.experimental.image.decode_tiff)\n2) generate N patches from the image (ex. tf.image.extract_patches or just multiple call of random_crop) -> your dataset shape becomes (N,H,W,C)\n3) tf.data.dataset.unbatch() -> your dataset shape becomes (H,W,C)\n4) tf.data.dataset.shuffle(N*m) -> this is needed for random sampling (note that shuffling window size m will directly affect performance of the model)\n5) tf.data.dataset.batch(batch_size)\n\nin 4), you may select large m for better training procedure, but it may lead to heavy overload in preprocessing step. the fastest way is always to use preprocessed dataset. so this is a tradeoff between easy/flexible implementation and performance/speed.\n\nhope this was helpful : )",
      "votes": null
    },
    {
      "id": "1914803",
      "postDate": "08/26/2022 12:41:41",
      "content": "<p>This was very helpful! Exactly the information I was looking for, thank you very much!</p>",
      "rawMarkdown": "This was very helpful! Exactly the information I was looking for, thank you very much!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1913648,
      "author_name": "hoyso48",
      "author_url": "",
      "post_date": "08/25/2022 12:32:37",
      "content": "<p>Maybe one workaround is to generate multiple patches(crops) when reading a single image from .tiff file.<br>\nif you use tf.data API, you can make your tf.data.dataset object through the following process.</p>\n<p>1) read single tiff file (ex. tfio.experimental.image.decode_tiff)<br>\n2) generate N patches from the image (ex. tf.image.extract_patches or just multiple call of random_crop) -&gt; your dataset shape becomes (N,H,W,C)<br>\n3) tf.data.dataset.unbatch() -&gt; your dataset shape becomes (H,W,C)<br>\n4) tf.data.dataset.shuffle(N*m) -&gt; this is needed for random sampling (note that shuffling window size m will directly affect performance of the model)<br>\n5) tf.data.dataset.batch(batch_size)</p>\n<p>in 4), you may select large m for better training procedure, but it may lead to heavy overload in preprocessing step. the fastest way is always to use preprocessed dataset. so this is a tradeoff between easy/flexible implementation and performance/speed.</p>\n<p>hope this was helpful : )</p>",
      "votes": null,
      "replies": [
        {
          "id": 1914803,
          "author_name": "baptistejoseph",
          "author_url": "",
          "post_date": "08/26/2022 12:41:41",
          "content": "<p>This was very helpful! Exactly the information I was looking for, thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1912350": "Hello! I'm using tensorflow 2 and trying to feed tiles to a Unet neural network (WIP). \nI create a dataset with all relevant data from my dataframe, then map the preprocessing function on it, getting a crop of the image and the corresponding cropped mask.\nI have not found a way to split data within the preprocessing function; meaning that I decode the entire image and entire mask, crop with the offset values, then decode the entire image again to get the next tile from the same .tiff... \nThis looks extremely inefficient to me, does anyone know how to use the preprocessing function to return not the image and corresponding mask, but a \"list\" of these tiles from a single decode?\nNotebook here:  https://www.kaggle.com/code/baptistejoseph/hubmap-hpa-notebook/",
    "1913648": "Maybe one workaround is to generate multiple patches(crops) when reading a single image from .tiff file.\nif you use tf.data API, you can make your tf.data.dataset object through the following process.\n\n1) read single tiff file (ex. tfio.experimental.image.decode_tiff)\n2) generate N patches from the image (ex. tf.image.extract_patches or just multiple call of random_crop) -> your dataset shape becomes (N,H,W,C)\n3) tf.data.dataset.unbatch() -> your dataset shape becomes (H,W,C)\n4) tf.data.dataset.shuffle(N*m) -> this is needed for random sampling (note that shuffling window size m will directly affect performance of the model)\n5) tf.data.dataset.batch(batch_size)\n\nin 4), you may select large m for better training procedure, but it may lead to heavy overload in preprocessing step. the fastest way is always to use preprocessed dataset. so this is a tradeoff between easy/flexible implementation and performance/speed.\n\nhope this was helpful : )",
    "1914803": "This was very helpful! Exactly the information I was looking for, thank you very much!"
  },
  "source": "meta"
}