{
  "id": 343063,
  "title": "Tiling+Augmentations leads to exceeding disk space",
  "url": "/competitions/hubmap-organ-segmentation/discussion/343063",
  "author_name": "",
  "post_date": "2022-08-09T20:28:50.008740400Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,<br>\nI am currently trying to create the tiled augmented dataset but it apparently overshoots the disksize. The only option I am able to think of is to store different chunks of the data. Is there another workaround?<br>\nI divide the image into 36 512x512 Tiles.<br>\nI also tried storing them post resizing to 224x224 but the disk still became full.<br>\nThe process I follow in Augmenting the image and then creating tiles from that image.<br>\nNotebook link: <a href=\"https://www.kaggle.com/code/arnavchakravarthy/tiling-fast-augmentations-chunks-multithreading\" target=\"_blank\">https://www.kaggle.com/code/arnavchakravarthy/tiling-fast-augmentations-chunks-multithreading</a></p>",
  "messages": [
    {
      "id": "1892076",
      "postDate": "08/09/2022 20:28:50",
      "content": "<p>Hi,<br>\nI am currently trying to create the tiled augmented dataset but it apparently overshoots the disksize. The only option I am able to think of is to store different chunks of the data. Is there another workaround?<br>\nI divide the image into 36 512x512 Tiles.<br>\nI also tried storing them post resizing to 224x224 but the disk still became full.<br>\nThe process I follow in Augmenting the image and then creating tiles from that image.<br>\nNotebook link: <a href=\"https://www.kaggle.com/code/arnavchakravarthy/tiling-fast-augmentations-chunks-multithreading\" target=\"_blank\">https://www.kaggle.com/code/arnavchakravarthy/tiling-fast-augmentations-chunks-multithreading</a></p>",
      "rawMarkdown": "Hi,\nI am currently trying to create the tiled augmented dataset but it apparently overshoots the disksize. The only option I am able to think of is to store different chunks of the data. Is there another workaround?\nI divide the image into 36 512x512 Tiles.\nI also tried storing them post resizing to 224x224 but the disk still became full.\nThe process I follow in Augmenting the image and then creating tiles from that image.\nNotebook link: https://www.kaggle.com/code/arnavchakravarthy/tiling-fast-augmentations-chunks-multithreading",
      "votes": null
    },
    {
      "id": "1892098",
      "postDate": "08/09/2022 20:54:47",
      "content": "<p>I can see what you are doing in notebook :)</p>\n<pre><code>Each Image is being Augmented 79 times\n</code></pre>\n<p>The original image dataset is around 9Gb and you resize it on half so we can expect whole dataset will take around 4.5Gb * 80 which is 360Gb of storage data :)<br>\n<br><br>\nI would suggest that you convert TIFF images to more appropriate format which will not take so much size (JPEG's)</p>",
      "rawMarkdown": "I can see what you are doing in notebook :)\n```\nEach Image is being Augmented 79 times\n```\nThe original image dataset is around 9Gb and you resize it on half so we can expect whole dataset will take around 4.5Gb * 80 which is 360Gb of storage data :)\n<br>\nI would suggest that you convert TIFF images to more appropriate format which will not take so much size (JPEG's)",
      "votes": null
    },
    {
      "id": "1892100",
      "postDate": "08/09/2022 20:58:24",
      "content": "<p>One more tip: Do not add 80 augmentations on each image right away, slowly add one augmentation style and check how well the submission score responds on added augmentation, and try first with augmentations that you expect to occur in the private dataset.</p>\n<p>I wish you all the best!</p>",
      "rawMarkdown": "One more tip: Do not add 80 augmentations on each image right away, slowly add one augmentation style and check how well the submission score responds on added augmentation, and try first with augmentations that you expect to occur in the private dataset.\n\nI wish you all the best!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1892098,
      "author_name": "urosjarc",
      "author_url": "",
      "post_date": "08/09/2022 20:54:47",
      "content": "<p>I can see what you are doing in notebook :)</p>\n<pre><code>Each Image is being Augmented 79 times\n</code></pre>\n<p>The original image dataset is around 9Gb and you resize it on half so we can expect whole dataset will take around 4.5Gb * 80 which is 360Gb of storage data :)<br>\n<br><br>\nI would suggest that you convert TIFF images to more appropriate format which will not take so much size (JPEG's)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1892100,
      "author_name": "urosjarc",
      "author_url": "",
      "post_date": "08/09/2022 20:58:24",
      "content": "<p>One more tip: Do not add 80 augmentations on each image right away, slowly add one augmentation style and check how well the submission score responds on added augmentation, and try first with augmentations that you expect to occur in the private dataset.</p>\n<p>I wish you all the best!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1892076": "Hi,\nI am currently trying to create the tiled augmented dataset but it apparently overshoots the disksize. The only option I am able to think of is to store different chunks of the data. Is there another workaround?\nI divide the image into 36 512x512 Tiles.\nI also tried storing them post resizing to 224x224 but the disk still became full.\nThe process I follow in Augmenting the image and then creating tiles from that image.\nNotebook link: https://www.kaggle.com/code/arnavchakravarthy/tiling-fast-augmentations-chunks-multithreading",
    "1892098": "I can see what you are doing in notebook :)\n```\nEach Image is being Augmented 79 times\n```\nThe original image dataset is around 9Gb and you resize it on half so we can expect whole dataset will take around 4.5Gb * 80 which is 360Gb of storage data :)\n<br>\nI would suggest that you convert TIFF images to more appropriate format which will not take so much size (JPEG's)",
    "1892100": "One more tip: Do not add 80 augmentations on each image right away, slowly add one augmentation style and check how well the submission score responds on added augmentation, and try first with augmentations that you expect to occur in the private dataset.\n\nI wish you all the best!"
  },
  "source": "meta"
}