{
  "id": 239106,
  "title": "[Feedback] Lack of Complete Dataset and my personal experience on this competition",
  "url": "/competitions/geolifeclef-2021/discussion/239106",
  "author_name": "",
  "post_date": "2021-05-14T16:41:27.757304900Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>For a future competition, would be so helpful if the complete dataset could be present inside the competition. It takes so much time to download/upload/decompress everything from a remote location. I will share my personal experience trying to compete on this competition:</p>\n<ul>\n<li>I joined very late (my own fault), like 7 days from the deadline.</li>\n<li>I then realized that the dataset provided inside the competition is just a tiny fraction of the whole image dataset.</li>\n<li>Kaggle doesn't allow remote datasets to be larger than 4GB, so I had to download it to my local computer.</li>\n<li>It took me over a day to download it to my local machine because somehow the host connection is not reliable and slow. I had to restart the download several times because the network would just give up and the download speeds were very slow.</li>\n<li>Then it took me another day to upload the dataset back to Kaggle. But that is my own fault for having an asymmetric internet connection (240MB/s download and 24MB/s upload).</li>\n<li>Ideally, I would like to run the training on google COLAB but for some strange reason, the public datasets I created would not work outside Kaggle (who knows why).</li>\n<li>I then decided to burn my precious GPU hours of Kaggle to train the model here because I just couldn't get the dataset to work outside of Kaggle.</li>\n<li>As we speak the network is training (for only 4 epochs at about 2h/epoch) because this is all I can afford right now due to Kaggle GPU limitations.</li>\n<li>If the notebook fails somehow (because of my own faults) I won't be able to submit it because it would have burned all the GPU time I have for this week. So, let's hope for the best!</li>\n</ul>\n<p>I don't know how difficult would have been for the competition creators to have provided the whole dataset up-front but that would have saved me at least 3 days of work. So, if that is possible, please do so in the next competition.</p>\n<p>Other than that, I enjoyed this comp as I think the subject is very interesting and challenging. Looking forward to the 2022 edition.</p>",
  "messages": [
    {
      "id": "1307764",
      "postDate": "05/14/2021 16:41:27",
      "content": "<p>For a future competition, would be so helpful if the complete dataset could be present inside the competition. It takes so much time to download/upload/decompress everything from a remote location. I will share my personal experience trying to compete on this competition:</p>\n<ul>\n<li>I joined very late (my own fault), like 7 days from the deadline.</li>\n<li>I then realized that the dataset provided inside the competition is just a tiny fraction of the whole image dataset.</li>\n<li>Kaggle doesn't allow remote datasets to be larger than 4GB, so I had to download it to my local computer.</li>\n<li>It took me over a day to download it to my local machine because somehow the host connection is not reliable and slow. I had to restart the download several times because the network would just give up and the download speeds were very slow.</li>\n<li>Then it took me another day to upload the dataset back to Kaggle. But that is my own fault for having an asymmetric internet connection (240MB/s download and 24MB/s upload).</li>\n<li>Ideally, I would like to run the training on google COLAB but for some strange reason, the public datasets I created would not work outside Kaggle (who knows why).</li>\n<li>I then decided to burn my precious GPU hours of Kaggle to train the model here because I just couldn't get the dataset to work outside of Kaggle.</li>\n<li>As we speak the network is training (for only 4 epochs at about 2h/epoch) because this is all I can afford right now due to Kaggle GPU limitations.</li>\n<li>If the notebook fails somehow (because of my own faults) I won't be able to submit it because it would have burned all the GPU time I have for this week. So, let's hope for the best!</li>\n</ul>\n<p>I don't know how difficult would have been for the competition creators to have provided the whole dataset up-front but that would have saved me at least 3 days of work. So, if that is possible, please do so in the next competition.</p>\n<p>Other than that, I enjoyed this comp as I think the subject is very interesting and challenging. Looking forward to the 2022 edition.</p>",
      "rawMarkdown": "For a future competition, would be so helpful if the complete dataset could be present inside the competition. It takes so much time to download/upload/decompress everything from a remote location. I will share my personal experience trying to compete on this competition:\n\n- I joined very late (my own fault), like 7 days from the deadline.\n- I then realized that the dataset provided inside the competition is just a tiny fraction of the whole image dataset.\n- Kaggle doesn't allow remote datasets to be larger than 4GB, so I had to download it to my local computer.\n- It took me over a day to download it to my local machine because somehow the host connection is not reliable and slow. I had to restart the download several times because the network would just give up and the download speeds were very slow.\n- Then it took me another day to upload the dataset back to Kaggle. But that is my own fault for having an asymmetric internet connection (240MB/s download and 24MB/s upload).\n- Ideally, I would like to run the training on google COLAB but for some strange reason, the public datasets I created would not work outside Kaggle (who knows why).\n- I then decided to burn my precious GPU hours of Kaggle to train the model here because I just couldn't get the dataset to work outside of Kaggle.\n- As we speak the network is training (for only 4 epochs at about 2h/epoch) because this is all I can afford right now due to Kaggle GPU limitations.\n- If the notebook fails somehow (because of my own faults) I won't be able to submit it because it would have burned all the GPU time I have for this week. So, let's hope for the best!\n\nI don't know how difficult would have been for the competition creators to have provided the whole dataset up-front but that would have saved me at least 3 days of work. So, if that is possible, please do so in the next competition.\n\nOther than that, I enjoyed this comp as I think the subject is very interesting and challenging. Looking forward to the 2022 edition.",
      "votes": null
    },
    {
      "id": "1312633",
      "postDate": "05/18/2021 06:51:11",
      "content": "<p>H Adriano, thanks a lot for your feedback, this is very useful. As you guess, we could not host the challenge on Kaggle because of its large size. We hosted it on <a href=\"http://lila.science/\" target=\"_blank\">http://lila.science/</a> which is repository for data sets related to biology and conservation intended as a ressource for ML researchers. This service is maintained by a working group that includes representatives from Zooniverse, the Evolving AI Lab, the University of Minnesota Lion Center, Snapshot Safari, and Microsoft AI for Earth. Hosting on Microsoft Azure is provided by Microsoft AI for Earth. Thus, you should theoretically not have connection issues from the host side. Several of us succeeded in downloading the full dataset in a few hours. We would be happy to have feedback from the other participants about that. Thanks a lot. </p>",
      "rawMarkdown": "H Adriano, thanks a lot for your feedback, this is very useful. As you guess, we could not host the challenge on Kaggle because of its large size. We hosted it on http://lila.science/ which is repository for data sets related to biology and conservation intended as a ressource for ML researchers. This service is maintained by a working group that includes representatives from Zooniverse, the Evolving AI Lab, the University of Minnesota Lion Center, Snapshot Safari, and Microsoft AI for Earth. Hosting on Microsoft Azure is provided by Microsoft AI for Earth. Thus, you should theoretically not have connection issues from the host side. Several of us succeeded in downloading the full dataset in a few hours. We would be happy to have feedback from the other participants about that. Thanks a lot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1312633,
      "author_name": "alxjoly",
      "author_url": "",
      "post_date": "05/18/2021 06:51:11",
      "content": "<p>H Adriano, thanks a lot for your feedback, this is very useful. As you guess, we could not host the challenge on Kaggle because of its large size. We hosted it on <a href=\"http://lila.science/\" target=\"_blank\">http://lila.science/</a> which is repository for data sets related to biology and conservation intended as a ressource for ML researchers. This service is maintained by a working group that includes representatives from Zooniverse, the Evolving AI Lab, the University of Minnesota Lion Center, Snapshot Safari, and Microsoft AI for Earth. Hosting on Microsoft Azure is provided by Microsoft AI for Earth. Thus, you should theoretically not have connection issues from the host side. Several of us succeeded in downloading the full dataset in a few hours. We would be happy to have feedback from the other participants about that. Thanks a lot. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1307764": "For a future competition, would be so helpful if the complete dataset could be present inside the competition. It takes so much time to download/upload/decompress everything from a remote location. I will share my personal experience trying to compete on this competition:\n\n- I joined very late (my own fault), like 7 days from the deadline.\n- I then realized that the dataset provided inside the competition is just a tiny fraction of the whole image dataset.\n- Kaggle doesn't allow remote datasets to be larger than 4GB, so I had to download it to my local computer.\n- It took me over a day to download it to my local machine because somehow the host connection is not reliable and slow. I had to restart the download several times because the network would just give up and the download speeds were very slow.\n- Then it took me another day to upload the dataset back to Kaggle. But that is my own fault for having an asymmetric internet connection (240MB/s download and 24MB/s upload).\n- Ideally, I would like to run the training on google COLAB but for some strange reason, the public datasets I created would not work outside Kaggle (who knows why).\n- I then decided to burn my precious GPU hours of Kaggle to train the model here because I just couldn't get the dataset to work outside of Kaggle.\n- As we speak the network is training (for only 4 epochs at about 2h/epoch) because this is all I can afford right now due to Kaggle GPU limitations.\n- If the notebook fails somehow (because of my own faults) I won't be able to submit it because it would have burned all the GPU time I have for this week. So, let's hope for the best!\n\nI don't know how difficult would have been for the competition creators to have provided the whole dataset up-front but that would have saved me at least 3 days of work. So, if that is possible, please do so in the next competition.\n\nOther than that, I enjoyed this comp as I think the subject is very interesting and challenging. Looking forward to the 2022 edition.",
    "1312633": "H Adriano, thanks a lot for your feedback, this is very useful. As you guess, we could not host the challenge on Kaggle because of its large size. We hosted it on http://lila.science/ which is repository for data sets related to biology and conservation intended as a ressource for ML researchers. This service is maintained by a working group that includes representatives from Zooniverse, the Evolving AI Lab, the University of Minnesota Lion Center, Snapshot Safari, and Microsoft AI for Earth. Hosting on Microsoft Azure is provided by Microsoft AI for Earth. Thus, you should theoretically not have connection issues from the host side. Several of us succeeded in downloading the full dataset in a few hours. We would be happy to have feedback from the other participants about that. Thanks a lot."
  },
  "source": "meta"
}