{
  "id": 313266,
  "title": "71 Gb ➡  14 Gb Dataset",
  "url": "/competitions/sorghum-id-fgvc-9/discussion/313266",
  "author_name": "",
  "post_date": "2022-03-16T09:47:07.638727Z",
  "votes": 13,
  "comment_count": 11,
  "views": 0,
  "content": "<p>This Comp data uses Png files which are very large . We can use JPEG to reduce the data size to just 14 gb<br>\n<strong>dataset link</strong><br>\n<a href=\"https://www.kaggle.com/datasets/mithilsalunkhe/small-jpegs-fgvc\" target=\"_blank\">https://www.kaggle.com/datasets/mithilsalunkhe/small-jpegs-fgvc</a></p>",
  "messages": [
    {
      "id": "1724523",
      "postDate": "03/16/2022 09:47:07",
      "content": "<p>This Comp data uses Png files which are very large . We can use JPEG to reduce the data size to just 14 gb<br>\n<strong>dataset link</strong><br>\n<a href=\"https://www.kaggle.com/datasets/mithilsalunkhe/small-jpegs-fgvc\" target=\"_blank\">https://www.kaggle.com/datasets/mithilsalunkhe/small-jpegs-fgvc</a></p>",
      "rawMarkdown": "This Comp data uses Png files which are very large . We can use JPEG to reduce the data size to just 14 gb\n**dataset link**\nhttps://www.kaggle.com/datasets/mithilsalunkhe/small-jpegs-fgvc",
      "votes": null
    },
    {
      "id": "1724872",
      "postDate": "03/16/2022 15:37:59",
      "content": "<p>Good work <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> !</p>",
      "rawMarkdown": "Good work @mithilsalunkhe !",
      "votes": null
    },
    {
      "id": "1724875",
      "postDate": "03/16/2022 15:40:27",
      "content": "<p>But please by doing it like this, won't we lose the quality of the images?</p>",
      "rawMarkdown": "But please by doing it like this, won't we lose the quality of the images?",
      "votes": null
    },
    {
      "id": "1724906",
      "postDate": "03/16/2022 16:08:45",
      "content": "<p>The data quality loss will be minimal for deep learning </p>",
      "rawMarkdown": "The data quality loss will be minimal for deep learning",
      "votes": null
    },
    {
      "id": "1724929",
      "postDate": "03/16/2022 16:29:28",
      "content": "<p>Okay I can see, so let's go and use it for the work!</p>",
      "rawMarkdown": "Okay I can see, so let's go and use it for the work!",
      "votes": null
    },
    {
      "id": "1725937",
      "postDate": "03/17/2022 14:26:06",
      "content": "<p>Hi, thanks for the data/work, but I think you made some mistake in the filename when saving to jpeg. It seens that the last character from all images are missing.</p>\n<p>Example of forst 3 files form each dataset:<br>\nOriginal Dataset:<br>\n../input/sorghum-id-fgvc-9/train_images/2017-06-01__10-26-27-47<strong>9</strong>.png<br>\n../input/sorghum-id-fgvc-9/train_images/2017-06-01__10-26-28-94<strong>4</strong>.png<br>\n../input/sorghum-id-fgvc-9/train_images/2017-06-01__10-26-30-47<strong>4</strong>.png</p>\n<p>JPEG Dataset:<br>\n../input/cultivar-identification-jpeg/train_img/2017-06-01__10-26-27-47.jpeg<br>\n../input/cultivar-identification-jpeg/train_img/2017-06-01__10-26-28-94.jpeg<br>\n../input/cultivar-identification-jpeg/train_img/2017-06-01__10-26-30-47.jpeg</p>\n<p>We can see that all 3 are missing the last digit.</p>\n<p>Cheers</p>",
      "rawMarkdown": "Hi, thanks for the data/work, but I think you made some mistake in the filename when saving to jpeg. It seens that the last character from all images are missing.\n\nExample of forst 3 files form each dataset:\nOriginal Dataset:\n../input/sorghum-id-fgvc-9/train_images/2017-06-01__10-26-27-47**9**.png\n../input/sorghum-id-fgvc-9/train_images/2017-06-01__10-26-28-94**4**.png\n../input/sorghum-id-fgvc-9/train_images/2017-06-01__10-26-30-47**4**.png\n\nJPEG Dataset:\n../input/cultivar-identification-jpeg/train_img/2017-06-01__10-26-27-47.jpeg\n../input/cultivar-identification-jpeg/train_img/2017-06-01__10-26-28-94.jpeg\n../input/cultivar-identification-jpeg/train_img/2017-06-01__10-26-30-47.jpeg\n\nWe can see that all 3 are missing the last digit.\n\n\nCheers",
      "votes": null
    },
    {
      "id": "1725944",
      "postDate": "03/17/2022 14:31:06",
      "content": "<p>I will try to fix this</p>",
      "rawMarkdown": "I will try to fix this",
      "votes": null
    },
    {
      "id": "1730728",
      "postDate": "03/21/2022 15:05:39",
      "content": "<p><a href=\"https://www.kaggle.com/ibombonato\" target=\"_blank\">@ibombonato</a> <a href=\"https://www.kaggle.com/datasets/mithilsalunkhe/small-jpegs-fgvc\" target=\"_blank\">https://www.kaggle.com/datasets/mithilsalunkhe/small-jpegs-fgvc</a> fixed version here</p>",
      "rawMarkdown": "ibombonato https://www.kaggle.com/datasets/mithilsalunkhe/small-jpegs-fgvc fixed version here",
      "votes": null
    },
    {
      "id": "1738269",
      "postDate": "03/29/2022 07:05:48",
      "content": "<p>Great jobs!</p>",
      "rawMarkdown": "Great jobs!",
      "votes": null
    },
    {
      "id": "1740551",
      "postDate": "03/31/2022 03:46:52",
      "content": "<p>Thank you for your work. </p>",
      "rawMarkdown": "Thank you for your work.",
      "votes": null
    },
    {
      "id": "1756563",
      "postDate": "04/15/2022 16:02:22",
      "content": "<p>I compared the peformance of a pretrained EfficientNet_b0 baseline with constant learning rate and w/o any augmentations on 128x128 images in png format and in jpg (this dataset): I found the difference in accuracy on the validation set to be about 2%. At the same time, training was a lot faster since there was no IO bottleneck anymore (slow storage). Not sure if those 2% are due to external factors or if it is because of the jpeg conversion.</p>",
      "rawMarkdown": "I compared the peformance of a pretrained EfficientNet_b0 baseline with constant learning rate and w/o any augmentations on 128x128 images in png format and in jpg (this dataset): I found the difference in accuracy on the validation set to be about 2%. At the same time, training was a lot faster since there was no IO bottleneck anymore (slow storage). Not sure if those 2% are due to external factors or if it is because of the jpeg conversion.",
      "votes": null
    },
    {
      "id": "1765233",
      "postDate": "04/23/2022 09:23:15",
      "content": "<p>Just a question ; DId you split the dataset into Kfolds and set a seed ?</p>",
      "rawMarkdown": "Just a question ; DId you split the dataset into Kfolds and set a seed ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1724872,
      "author_name": "prosperalikizang",
      "author_url": "",
      "post_date": "03/16/2022 15:37:59",
      "content": "<p>Good work <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1724875,
      "author_name": "prosperalikizang",
      "author_url": "",
      "post_date": "03/16/2022 15:40:27",
      "content": "<p>But please by doing it like this, won't we lose the quality of the images?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1724906,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "03/16/2022 16:08:45",
          "content": "<p>The data quality loss will be minimal for deep learning </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1724929,
          "author_name": "prosperalikizang",
          "author_url": "",
          "post_date": "03/16/2022 16:29:28",
          "content": "<p>Okay I can see, so let's go and use it for the work!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1756563,
          "author_name": "tlipss",
          "author_url": "",
          "post_date": "04/15/2022 16:02:22",
          "content": "<p>I compared the peformance of a pretrained EfficientNet_b0 baseline with constant learning rate and w/o any augmentations on 128x128 images in png format and in jpg (this dataset): I found the difference in accuracy on the validation set to be about 2%. At the same time, training was a lot faster since there was no IO bottleneck anymore (slow storage). Not sure if those 2% are due to external factors or if it is because of the jpeg conversion.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1765233,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "04/23/2022 09:23:15",
          "content": "<p>Just a question ; DId you split the dataset into Kfolds and set a seed ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1725937,
      "author_name": "ibombonato",
      "author_url": "",
      "post_date": "03/17/2022 14:26:06",
      "content": "<p>Hi, thanks for the data/work, but I think you made some mistake in the filename when saving to jpeg. It seens that the last character from all images are missing.</p>\n<p>Example of forst 3 files form each dataset:<br>\nOriginal Dataset:<br>\n../input/sorghum-id-fgvc-9/train_images/2017-06-01__10-26-27-47<strong>9</strong>.png<br>\n../input/sorghum-id-fgvc-9/train_images/2017-06-01__10-26-28-94<strong>4</strong>.png<br>\n../input/sorghum-id-fgvc-9/train_images/2017-06-01__10-26-30-47<strong>4</strong>.png</p>\n<p>JPEG Dataset:<br>\n../input/cultivar-identification-jpeg/train_img/2017-06-01__10-26-27-47.jpeg<br>\n../input/cultivar-identification-jpeg/train_img/2017-06-01__10-26-28-94.jpeg<br>\n../input/cultivar-identification-jpeg/train_img/2017-06-01__10-26-30-47.jpeg</p>\n<p>We can see that all 3 are missing the last digit.</p>\n<p>Cheers</p>",
      "votes": null,
      "replies": [
        {
          "id": 1725944,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "03/17/2022 14:31:06",
          "content": "<p>I will try to fix this</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1730728,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "03/21/2022 15:05:39",
          "content": "<p><a href=\"https://www.kaggle.com/ibombonato\" target=\"_blank\">@ibombonato</a> <a href=\"https://www.kaggle.com/datasets/mithilsalunkhe/small-jpegs-fgvc\" target=\"_blank\">https://www.kaggle.com/datasets/mithilsalunkhe/small-jpegs-fgvc</a> fixed version here</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1738269,
      "author_name": "leoooo333",
      "author_url": "",
      "post_date": "03/29/2022 07:05:48",
      "content": "<p>Great jobs!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1740551,
      "author_name": "zhenyuxie",
      "author_url": "",
      "post_date": "03/31/2022 03:46:52",
      "content": "<p>Thank you for your work. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1724523": "This Comp data uses Png files which are very large . We can use JPEG to reduce the data size to just 14 gb\n**dataset link**\nhttps://www.kaggle.com/datasets/mithilsalunkhe/small-jpegs-fgvc",
    "1724872": "Good work @mithilsalunkhe !",
    "1724875": "But please by doing it like this, won't we lose the quality of the images?",
    "1724906": "The data quality loss will be minimal for deep learning",
    "1724929": "Okay I can see, so let's go and use it for the work!",
    "1725937": "Hi, thanks for the data/work, but I think you made some mistake in the filename when saving to jpeg. It seens that the last character from all images are missing.\n\nExample of forst 3 files form each dataset:\nOriginal Dataset:\n../input/sorghum-id-fgvc-9/train_images/2017-06-01__10-26-27-47**9**.png\n../input/sorghum-id-fgvc-9/train_images/2017-06-01__10-26-28-94**4**.png\n../input/sorghum-id-fgvc-9/train_images/2017-06-01__10-26-30-47**4**.png\n\nJPEG Dataset:\n../input/cultivar-identification-jpeg/train_img/2017-06-01__10-26-27-47.jpeg\n../input/cultivar-identification-jpeg/train_img/2017-06-01__10-26-28-94.jpeg\n../input/cultivar-identification-jpeg/train_img/2017-06-01__10-26-30-47.jpeg\n\nWe can see that all 3 are missing the last digit.\n\n\nCheers",
    "1725944": "I will try to fix this",
    "1730728": "ibombonato https://www.kaggle.com/datasets/mithilsalunkhe/small-jpegs-fgvc fixed version here",
    "1738269": "Great jobs!",
    "1740551": "Thank you for your work.",
    "1756563": "I compared the peformance of a pretrained EfficientNet_b0 baseline with constant learning rate and w/o any augmentations on 128x128 images in png format and in jpg (this dataset): I found the difference in accuracy on the validation set to be about 2%. At the same time, training was a lot faster since there was no IO bottleneck anymore (slow storage). Not sure if those 2% are due to external factors or if it is because of the jpeg conversion.",
    "1765233": "Just a question ; DId you split the dataset into Kfolds and set a seed ?"
  },
  "source": "meta"
}