{
  "id": 672533,
  "title": "Request: Direct cloud access / bucketed hosting for 281 GB training tar",
  "url": "/competitions/plantclef-2026/discussion/672533",
  "author_name": "",
  "post_date": "2026-02-09T00:56:10.427577500Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have a question regarding the large external training tar archives, especially the <strong>~281 GB single‑plant training set (min side 800 px)</strong> and the other large complementary tar files hosted on the Pl@ntNet servers. For many participants, it is quite challenging to handle downloads of this size.</p>\n<ul>\n<li>Typical Colab / cloud notebook environments have limited local disk space, so streaming the 281 GB tar into a mounted Google Drive or similar often fails due to disk limits.  </li>\n<li>Downloading the tar to a local machine and then re‑uploading to cloud storage is impractical for many of us, both in terms of bandwidth and time.</li>\n</ul>\n<p>Would it be possible to provide an alternative “bucketed” access option for these large training archives? For example, hosting the training images in a public <strong>cloud storage bucket</strong> (e.g., GCS/AWS/S3/Zenodo‑style)?</p>\n<p>This kind of bucketed or chunked hosting would greatly lower the barrier for participants with modest hardware or those relying on Colab/Kaggle/other hosted environments, while also reducing the need for repeated huge downloads. </p>\n<p>If there are already recommended best practices or alternative endpoints to access the training images, please let me know. For example, is there some true streamline approach where I can use colab as just a link between the servers and my google drive such that it downloads the entire data directly to the drive without exceeding its local disk limit?</p>\n<p>Thanks a lot for considering this, and for all the work you’ve put into PlantCLEF and LifeCLEF in general!</p>",
  "messages": [
    {
      "id": "3403635",
      "postDate": "02/09/2026 00:56:10",
      "content": "<p>I have a question regarding the large external training tar archives, especially the <strong>~281 GB single‑plant training set (min side 800 px)</strong> and the other large complementary tar files hosted on the Pl@ntNet servers. For many participants, it is quite challenging to handle downloads of this size.</p>\n<ul>\n<li>Typical Colab / cloud notebook environments have limited local disk space, so streaming the 281 GB tar into a mounted Google Drive or similar often fails due to disk limits.  </li>\n<li>Downloading the tar to a local machine and then re‑uploading to cloud storage is impractical for many of us, both in terms of bandwidth and time.</li>\n</ul>\n<p>Would it be possible to provide an alternative “bucketed” access option for these large training archives? For example, hosting the training images in a public <strong>cloud storage bucket</strong> (e.g., GCS/AWS/S3/Zenodo‑style)?</p>\n<p>This kind of bucketed or chunked hosting would greatly lower the barrier for participants with modest hardware or those relying on Colab/Kaggle/other hosted environments, while also reducing the need for repeated huge downloads. </p>\n<p>If there are already recommended best practices or alternative endpoints to access the training images, please let me know. For example, is there some true streamline approach where I can use colab as just a link between the servers and my google drive such that it downloads the entire data directly to the drive without exceeding its local disk limit?</p>\n<p>Thanks a lot for considering this, and for all the work you’ve put into PlantCLEF and LifeCLEF in general!</p>",
      "rawMarkdown": "I have a question regarding the large external training tar archives, especially the **~281 GB single‑plant training set (min side 800 px)** and the other large complementary tar files hosted on the Pl@ntNet servers. For many participants, it is quite challenging to handle downloads of this size.\n\n- Typical Colab / cloud notebook environments have limited local disk space, so streaming the 281 GB tar into a mounted Google Drive or similar often fails due to disk limits.  \n- Downloading the tar to a local machine and then re‑uploading to cloud storage is impractical for many of us, both in terms of bandwidth and time.\n\nWould it be possible to provide an alternative “bucketed” access option for these large training archives? For example, hosting the training images in a public **cloud storage bucket** (e.g., GCS/AWS/S3/Zenodo‑style)?\n\nThis kind of bucketed or chunked hosting would greatly lower the barrier for participants with modest hardware or those relying on Colab/Kaggle/other hosted environments, while also reducing the need for repeated huge downloads. \n\nIf there are already recommended best practices or alternative endpoints to access the training images, please let me know. For example, is there some true streamline approach where I can use colab as just a link between the servers and my google drive such that it downloads the entire data directly to the drive without exceeding its local disk limit?\n\nThanks a lot for considering this, and for all the work you’ve put into PlantCLEF and LifeCLEF in general!",
      "votes": null
    },
    {
      "id": "3404218",
      "postDate": "02/10/2026 04:05:52",
      "content": "<p>Same here. 218 GB is too big for poor peasants like us 😭</p>",
      "rawMarkdown": "Same here. 218 GB is too big for poor peasants like us 😭",
      "votes": null
    },
    {
      "id": "3404225",
      "postDate": "02/10/2026 04:26:21",
      "content": "<p>Its 281, not 218, though both are out of bound to us. Also, even if we do manage to download it, what about extraction? Untar will take 281x2 (at least) = a system with 562 GB of free space because without complete extraction we cannot even delete the tar file either.</p>",
      "rawMarkdown": "Its 281, not 218, though both are out of bound to us. Also, even if we do manage to download it, what about extraction? Untar will take 281x2 (at least) = a system with 562 GB of free space because without complete extraction we cannot even delete the tar file either.",
      "votes": null
    },
    {
      "id": "3404307",
      "postDate": "02/10/2026 09:02:46",
      "content": "<p>Hey guys,</p>\n<p>Not sure if this resolves your issue, but you can use the \"max of 800px\" archive, which is 160 GB.</p>\n<p>The other option would be to create your own dataset by downloading data directly from GBIF.</p>\n<p>Best,\nLukas</p>",
      "rawMarkdown": "Hey guys,\n\nNot sure if this resolves your issue, but you can use the \"max of 800px\" archive, which is 160 GB.\n\nThe other option would be to create your own dataset by downloading data directly from GBIF.\n\nBest,\nLukas",
      "votes": null
    },
    {
      "id": "3405261",
      "postDate": "02/12/2026 15:03:44",
      "content": "<p>Hi,</p>\n<p>Regarding the storage and bandwidth constraints, here are two alternatives to downloading the full 281 GB tar archive:</p>\n<ol>\n<li>Custom Download via Metadata You can rely on the PlantCLEF2024_single_plant_training_metadata.csv file, which contains individual download URLs for the images.</li>\n</ol>\n<p>Selective Download: You can filter the CSV to create a subset of the data (e.g., by excluding species with very few training images to save space, fixing a max number of images per species, etc).</p>\n<p>On-the-fly Processing: You can write a script to download images individually, resize/compress them in memory to your target resolution, and save them directly to your drive. This bypasses the need to store the massive archive or perform a full extraction, effectively solving the \"double disk space\" issue (tar + extracted files).</p>\n<ol>\n<li>Alternative Hosting Proposal If there is strong demand, I can generate and host lighter versions of the dataset, such as:\nMax 600x600px, or even 256x256px  with a center crop</li>\n</ol>\n<p>Important Warning: Please be aware that using pre-resized or pre-cropped images (especially 256px) will severely limit your ability to use standard data augmentations (like RandomResizedCrop) effectively. This may also reinforce the domain gap between training and the high-resolution test images, which could negatively impact your leaderboard performance.</p>\n<p>Let me know if these smaller sets would be useful to you despite the trade-offs. Any other suggestions are welcome</p>",
      "rawMarkdown": "Hi,\n\nRegarding the storage and bandwidth constraints, here are two alternatives to downloading the full 281 GB tar archive:\n\n1. Custom Download via Metadata You can rely on the PlantCLEF2024_single_plant_training_metadata.csv file, which contains individual download URLs for the images.\n\nSelective Download: You can filter the CSV to create a subset of the data (e.g., by excluding species with very few training images to save space, fixing a max number of images per species, etc).\n\nOn-the-fly Processing: You can write a script to download images individually, resize/compress them in memory to your target resolution, and save them directly to your drive. This bypasses the need to store the massive archive or perform a full extraction, effectively solving the \"double disk space\" issue (tar + extracted files).\n\n2. Alternative Hosting Proposal If there is strong demand, I can generate and host lighter versions of the dataset, such as:\nMax 600x600px, or even 256x256px  with a center crop\n\nImportant Warning: Please be aware that using pre-resized or pre-cropped images (especially 256px) will severely limit your ability to use standard data augmentations (like RandomResizedCrop) effectively. This may also reinforce the domain gap between training and the high-resolution test images, which could negatively impact your leaderboard performance.\n\nLet me know if these smaller sets would be useful to you despite the trade-offs. Any other suggestions are welcome",
      "votes": null
    },
    {
      "id": "3405296",
      "postDate": "02/12/2026 17:11:14",
      "content": "<p>Thanks a lot for the detailed answer and the alternatives you suggested sir.</p>\n<p>Please note that after a huge amount of processing, I finally managed to download and extract the data, however for most of the competitors yet to join, this is still a huge problem and an actual barrier too that can stop many people from even participating. </p>\n<p>Hence, I think the proposed approaches are still somewhat suboptimal for the core problem, which is efficiently getting the <em>full</em> dataset onto limited storage (just slightly higher than size of the dataset, ~300GB let's say):</p>\n<ul>\n<li><p>Selective download via metadata is great if I’m willing to train on a reduced subset, but it doesn’t help if I actually want the full dataset. In that case, it just converts one large download into a huge number of small HTTP requests, which is typically even slower and more fragile, without reducing the total data volume.</p></li>\n<li><p>On‑the‑fly processing solves the “double disk space” issue (tar + extracted) only in the sense that I never store the monolithic tar. But if I want all images, I still have to download and process all 281 GB worth of content; I just change <em>where</em> the space is consumed and add some CPU overhead which will again slowdown the downloading. It doesn’t really make the end‑to‑end process faster, it mainly optimizes storage layout while sacrificing the downloading speed.</p></li>\n</ul>\n<p>A more direct solution for people who want the <strong>entire</strong> dataset but have limited free space (let's say 300GB would be a properly bucketed archive structure. For example, splitting the 281 GB into multiple archives of let's say around 40 GB each would let us:</p>\n<ul>\n<li>Download one 40 GB archive at a time.  </li>\n<li>Extract it, verify it, then delete that archive before moving to the next.  </li>\n</ul>\n<p>With this workflow, the peak space requirement is roughly “one chunk + its extracted files = ~80GB” instead of “full tar + full extraction.” Concretely, that means free space with a peak space requirement of ~340GB (which is reasonable given the size of the dataset) is enough to eventually obtain all data, instead of needing something close to 281×2=562 GB at once.</p>\n<p>So a bucketed archive version of the dataset would preserve all images, avoid the overhead of millions of small requests, and substantially reduce the peak disk requirement during download and extraction. I would hence suggest hosting a bucketed (chunked) version of the dataset for at-least the time frame of this competition to make it easy for the participants. </p>\n<p>Thanks a lot for considering this problem sir.</p>",
      "rawMarkdown": "Thanks a lot for the detailed answer and the alternatives you suggested sir.\n\nPlease note that after a huge amount of processing, I finally managed to download and extract the data, however for most of the competitors yet to join, this is still a huge problem and an actual barrier too that can stop many people from even participating. \n\nHence, I think the proposed approaches are still somewhat suboptimal for the core problem, which is efficiently getting the *full* dataset onto limited storage (just slightly higher than size of the dataset, ~300GB let's say):\n\n- Selective download via metadata is great if I’m willing to train on a reduced subset, but it doesn’t help if I actually want the full dataset. In that case, it just converts one large download into a huge number of small HTTP requests, which is typically even slower and more fragile, without reducing the total data volume.\n\n- On‑the‑fly processing solves the “double disk space” issue (tar + extracted) only in the sense that I never store the monolithic tar. But if I want all images, I still have to download and process all 281 GB worth of content; I just change *where* the space is consumed and add some CPU overhead which will again slowdown the downloading. It doesn’t really make the end‑to‑end process faster, it mainly optimizes storage layout while sacrificing the downloading speed.\n\nA more direct solution for people who want the **entire** dataset but have limited free space (let's say 300GB would be a properly bucketed archive structure. For example, splitting the 281 GB into multiple archives of let's say around 40 GB each would let us:\n\n- Download one 40 GB archive at a time.  \n- Extract it, verify it, then delete that archive before moving to the next.  \n\nWith this workflow, the peak space requirement is roughly “one chunk + its extracted files = ~80GB” instead of “full tar + full extraction.” Concretely, that means free space with a peak space requirement of ~340GB (which is reasonable given the size of the dataset) is enough to eventually obtain all data, instead of needing something close to 281×2=562 GB at once.\n\nSo a bucketed archive version of the dataset would preserve all images, avoid the overhead of millions of small requests, and substantially reduce the peak disk requirement during download and extraction. I would hence suggest hosting a bucketed (chunked) version of the dataset for at-least the time frame of this competition to make it easy for the participants. \n\nThanks a lot for considering this problem sir.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3404218,
      "author_name": "msaranyan",
      "author_url": "",
      "post_date": "02/10/2026 04:05:52",
      "content": "<p>Same here. 218 GB is too big for poor peasants like us 😭</p>",
      "votes": null,
      "replies": [
        {
          "id": 3404225,
          "author_name": "komilparmar",
          "author_url": "",
          "post_date": "02/10/2026 04:26:21",
          "content": "<p>Its 281, not 218, though both are out of bound to us. Also, even if we do manage to download it, what about extraction? Untar will take 281x2 (at least) = a system with 562 GB of free space because without complete extraction we cannot even delete the tar file either.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3405261,
              "author_name": "hgoeau",
              "author_url": "",
              "post_date": "02/12/2026 15:03:44",
              "content": "<p>Hi,</p>\n<p>Regarding the storage and bandwidth constraints, here are two alternatives to downloading the full 281 GB tar archive:</p>\n<ol>\n<li>Custom Download via Metadata You can rely on the PlantCLEF2024_single_plant_training_metadata.csv file, which contains individual download URLs for the images.</li>\n</ol>\n<p>Selective Download: You can filter the CSV to create a subset of the data (e.g., by excluding species with very few training images to save space, fixing a max number of images per species, etc).</p>\n<p>On-the-fly Processing: You can write a script to download images individually, resize/compress them in memory to your target resolution, and save them directly to your drive. This bypasses the need to store the massive archive or perform a full extraction, effectively solving the \"double disk space\" issue (tar + extracted files).</p>\n<ol>\n<li>Alternative Hosting Proposal If there is strong demand, I can generate and host lighter versions of the dataset, such as:\nMax 600x600px, or even 256x256px  with a center crop</li>\n</ol>\n<p>Important Warning: Please be aware that using pre-resized or pre-cropped images (especially 256px) will severely limit your ability to use standard data augmentations (like RandomResizedCrop) effectively. This may also reinforce the domain gap between training and the high-resolution test images, which could negatively impact your leaderboard performance.</p>\n<p>Let me know if these smaller sets would be useful to you despite the trade-offs. Any other suggestions are welcome</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3405296,
                  "author_name": "komilparmar",
                  "author_url": "",
                  "post_date": "02/12/2026 17:11:14",
                  "content": "<p>Thanks a lot for the detailed answer and the alternatives you suggested sir.</p>\n<p>Please note that after a huge amount of processing, I finally managed to download and extract the data, however for most of the competitors yet to join, this is still a huge problem and an actual barrier too that can stop many people from even participating. </p>\n<p>Hence, I think the proposed approaches are still somewhat suboptimal for the core problem, which is efficiently getting the <em>full</em> dataset onto limited storage (just slightly higher than size of the dataset, ~300GB let's say):</p>\n<ul>\n<li><p>Selective download via metadata is great if I’m willing to train on a reduced subset, but it doesn’t help if I actually want the full dataset. In that case, it just converts one large download into a huge number of small HTTP requests, which is typically even slower and more fragile, without reducing the total data volume.</p></li>\n<li><p>On‑the‑fly processing solves the “double disk space” issue (tar + extracted) only in the sense that I never store the monolithic tar. But if I want all images, I still have to download and process all 281 GB worth of content; I just change <em>where</em> the space is consumed and add some CPU overhead which will again slowdown the downloading. It doesn’t really make the end‑to‑end process faster, it mainly optimizes storage layout while sacrificing the downloading speed.</p></li>\n</ul>\n<p>A more direct solution for people who want the <strong>entire</strong> dataset but have limited free space (let's say 300GB would be a properly bucketed archive structure. For example, splitting the 281 GB into multiple archives of let's say around 40 GB each would let us:</p>\n<ul>\n<li>Download one 40 GB archive at a time.  </li>\n<li>Extract it, verify it, then delete that archive before moving to the next.  </li>\n</ul>\n<p>With this workflow, the peak space requirement is roughly “one chunk + its extracted files = ~80GB” instead of “full tar + full extraction.” Concretely, that means free space with a peak space requirement of ~340GB (which is reasonable given the size of the dataset) is enough to eventually obtain all data, instead of needing something close to 281×2=562 GB at once.</p>\n<p>So a bucketed archive version of the dataset would preserve all images, avoid the overhead of millions of small requests, and substantially reduce the peak disk requirement during download and extraction. I would hence suggest hosting a bucketed (chunked) version of the dataset for at-least the time frame of this competition to make it easy for the participants. </p>\n<p>Thanks a lot for considering this problem sir.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3404307,
      "author_name": "picekl",
      "author_url": "",
      "post_date": "02/10/2026 09:02:46",
      "content": "<p>Hey guys,</p>\n<p>Not sure if this resolves your issue, but you can use the \"max of 800px\" archive, which is 160 GB.</p>\n<p>The other option would be to create your own dataset by downloading data directly from GBIF.</p>\n<p>Best,\nLukas</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3403635": "I have a question regarding the large external training tar archives, especially the **~281 GB single‑plant training set (min side 800 px)** and the other large complementary tar files hosted on the Pl@ntNet servers. For many participants, it is quite challenging to handle downloads of this size.\n\n- Typical Colab / cloud notebook environments have limited local disk space, so streaming the 281 GB tar into a mounted Google Drive or similar often fails due to disk limits.  \n- Downloading the tar to a local machine and then re‑uploading to cloud storage is impractical for many of us, both in terms of bandwidth and time.\n\nWould it be possible to provide an alternative “bucketed” access option for these large training archives? For example, hosting the training images in a public **cloud storage bucket** (e.g., GCS/AWS/S3/Zenodo‑style)?\n\nThis kind of bucketed or chunked hosting would greatly lower the barrier for participants with modest hardware or those relying on Colab/Kaggle/other hosted environments, while also reducing the need for repeated huge downloads. \n\nIf there are already recommended best practices or alternative endpoints to access the training images, please let me know. For example, is there some true streamline approach where I can use colab as just a link between the servers and my google drive such that it downloads the entire data directly to the drive without exceeding its local disk limit?\n\nThanks a lot for considering this, and for all the work you’ve put into PlantCLEF and LifeCLEF in general!",
    "3404218": "Same here. 218 GB is too big for poor peasants like us 😭",
    "3404225": "Its 281, not 218, though both are out of bound to us. Also, even if we do manage to download it, what about extraction? Untar will take 281x2 (at least) = a system with 562 GB of free space because without complete extraction we cannot even delete the tar file either.",
    "3404307": "Hey guys,\n\nNot sure if this resolves your issue, but you can use the \"max of 800px\" archive, which is 160 GB.\n\nThe other option would be to create your own dataset by downloading data directly from GBIF.\n\nBest,\nLukas",
    "3405261": "Hi,\n\nRegarding the storage and bandwidth constraints, here are two alternatives to downloading the full 281 GB tar archive:\n\n1. Custom Download via Metadata You can rely on the PlantCLEF2024_single_plant_training_metadata.csv file, which contains individual download URLs for the images.\n\nSelective Download: You can filter the CSV to create a subset of the data (e.g., by excluding species with very few training images to save space, fixing a max number of images per species, etc).\n\nOn-the-fly Processing: You can write a script to download images individually, resize/compress them in memory to your target resolution, and save them directly to your drive. This bypasses the need to store the massive archive or perform a full extraction, effectively solving the \"double disk space\" issue (tar + extracted files).\n\n2. Alternative Hosting Proposal If there is strong demand, I can generate and host lighter versions of the dataset, such as:\nMax 600x600px, or even 256x256px  with a center crop\n\nImportant Warning: Please be aware that using pre-resized or pre-cropped images (especially 256px) will severely limit your ability to use standard data augmentations (like RandomResizedCrop) effectively. This may also reinforce the domain gap between training and the high-resolution test images, which could negatively impact your leaderboard performance.\n\nLet me know if these smaller sets would be useful to you despite the trade-offs. Any other suggestions are welcome",
    "3405296": "Thanks a lot for the detailed answer and the alternatives you suggested sir.\n\nPlease note that after a huge amount of processing, I finally managed to download and extract the data, however for most of the competitors yet to join, this is still a huge problem and an actual barrier too that can stop many people from even participating. \n\nHence, I think the proposed approaches are still somewhat suboptimal for the core problem, which is efficiently getting the *full* dataset onto limited storage (just slightly higher than size of the dataset, ~300GB let's say):\n\n- Selective download via metadata is great if I’m willing to train on a reduced subset, but it doesn’t help if I actually want the full dataset. In that case, it just converts one large download into a huge number of small HTTP requests, which is typically even slower and more fragile, without reducing the total data volume.\n\n- On‑the‑fly processing solves the “double disk space” issue (tar + extracted) only in the sense that I never store the monolithic tar. But if I want all images, I still have to download and process all 281 GB worth of content; I just change *where* the space is consumed and add some CPU overhead which will again slowdown the downloading. It doesn’t really make the end‑to‑end process faster, it mainly optimizes storage layout while sacrificing the downloading speed.\n\nA more direct solution for people who want the **entire** dataset but have limited free space (let's say 300GB would be a properly bucketed archive structure. For example, splitting the 281 GB into multiple archives of let's say around 40 GB each would let us:\n\n- Download one 40 GB archive at a time.  \n- Extract it, verify it, then delete that archive before moving to the next.  \n\nWith this workflow, the peak space requirement is roughly “one chunk + its extracted files = ~80GB” instead of “full tar + full extraction.” Concretely, that means free space with a peak space requirement of ~340GB (which is reasonable given the size of the dataset) is enough to eventually obtain all data, instead of needing something close to 281×2=562 GB at once.\n\nSo a bucketed archive version of the dataset would preserve all images, avoid the overhead of millions of small requests, and substantially reduce the peak disk requirement during download and extraction. I would hence suggest hosting a bucketed (chunked) version of the dataset for at-least the time frame of this competition to make it easy for the participants. \n\nThanks a lot for considering this problem sir."
  },
  "source": "meta"
}