{
  "id": 95425,
  "title": "How to participate in the competition with large dataset?",
  "url": "/competitions/open-images-2019-object-detection/discussion/95425",
  "author_name": "",
  "post_date": "2019-06-12T06:33:43.666810600Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi, could anyone tell how are they handling the large dataset provided here.\n1) Local execution does not work, since data size&gt;available hard-disk space\n2) Tried Using Google Drive, and mounting on Google Collaboratory , but results in Input/Output error. \nCan anyone provide appropriately sized datasets to work here ?</p>",
  "messages": [
    {
      "id": "550931",
      "postDate": "06/12/2019 06:33:43",
      "content": "<p>Hi, could anyone tell how are they handling the large dataset provided here.\n1) Local execution does not work, since data size&gt;available hard-disk space\n2) Tried Using Google Drive, and mounting on Google Collaboratory , but results in Input/Output error. \nCan anyone provide appropriately sized datasets to work here ?</p>",
      "rawMarkdown": "Hi, could anyone tell how are they handling the large dataset provided here.\n1) Local execution does not work, since data size&gt;available hard-disk space\n2) Tried Using Google Drive, and mounting on Google Collaboratory , but results in Input/Output error. \nCan anyone provide appropriately sized datasets to work here ?",
      "votes": null
    },
    {
      "id": "550942",
      "postDate": "06/12/2019 06:46:38",
      "content": "<p>Meeting the same problems :(</p>",
      "rawMarkdown": "Meeting the same problems :(",
      "votes": null
    },
    {
      "id": "551029",
      "postDate": "06/12/2019 09:05:42",
      "content": "<p>No need to download entire dataset inorder to participate in the competition. You can probably download any one subset of training data and train your model on that. That would be possible on colab and google drive.</p>",
      "rawMarkdown": "No need to download entire dataset inorder to participate in the competition. You can probably download any one subset of training data and train your model on that. That would be possible on colab and google drive.",
      "votes": null
    },
    {
      "id": "551160",
      "postDate": "06/12/2019 12:06:56",
      "content": "<p>I put up part of the dataset on <a href=\"https://www.kaggle.com/c/open-images-2019-object-detection/discussion/94770#latest-550992\">Kaggle that was resized</a>. You should be able to use that to download into Google Collaboratory if you don't want to use Kaggle. </p>",
      "rawMarkdown": "I put up part of the dataset on [Kaggle that was resized](https://www.kaggle.com/c/open-images-2019-object-detection/discussion/94770#latest-550992). You should be able to use that to download into Google Collaboratory if you don't want to use Kaggle.",
      "votes": null
    },
    {
      "id": "558875",
      "postDate": "06/23/2019 04:48:19",
      "content": "<p>Even a single subset is of 58GB </p>",
      "rawMarkdown": "Even a single subset is of 58GB",
      "votes": null
    },
    {
      "id": "558930",
      "postDate": "06/23/2019 07:32:05",
      "content": "<p>You can use <a href=\"https://www.kaggle.com/mindtrinket/google-2019-30k-train/\">this dataset</a> which is resized to 256*256 and takes only 10 GB of space</p>\n\n<p>Thanks to <a href=\"/mindtrinket\">@mindtrinket</a> </p>",
      "rawMarkdown": "You can use [this dataset](https://www.kaggle.com/mindtrinket/google-2019-30k-train/) which is resized to 256*256 and takes only 10 GB of space\n\nThanks to @mindtrinket",
      "votes": null
    },
    {
      "id": "559471",
      "postDate": "06/24/2019 07:00:58",
      "content": "<p>hey use the reduced data set or download the data set using the wget method.</p>",
      "rawMarkdown": "hey use the reduced data set or download the data set using the wget method.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 550942,
      "author_name": "henryham",
      "author_url": "",
      "post_date": "06/12/2019 06:46:38",
      "content": "<p>Meeting the same problems :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 551029,
      "author_name": "chandanverma",
      "author_url": "",
      "post_date": "06/12/2019 09:05:42",
      "content": "<p>No need to download entire dataset inorder to participate in the competition. You can probably download any one subset of training data and train your model on that. That would be possible on colab and google drive.</p>",
      "votes": null,
      "replies": [
        {
          "id": 558875,
          "author_name": "vikashpathak",
          "author_url": "",
          "post_date": "06/23/2019 04:48:19",
          "content": "<p>Even a single subset is of 58GB </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 558930,
          "author_name": "chandanverma",
          "author_url": "",
          "post_date": "06/23/2019 07:32:05",
          "content": "<p>You can use <a href=\"https://www.kaggle.com/mindtrinket/google-2019-30k-train/\">this dataset</a> which is resized to 256*256 and takes only 10 GB of space</p>\n\n<p>Thanks to <a href=\"/mindtrinket\">@mindtrinket</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551160,
      "author_name": "mindtrinket",
      "author_url": "",
      "post_date": "06/12/2019 12:06:56",
      "content": "<p>I put up part of the dataset on <a href=\"https://www.kaggle.com/c/open-images-2019-object-detection/discussion/94770#latest-550992\">Kaggle that was resized</a>. You should be able to use that to download into Google Collaboratory if you don't want to use Kaggle. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 559471,
      "author_name": "krishnakatyal",
      "author_url": "",
      "post_date": "06/24/2019 07:00:58",
      "content": "<p>hey use the reduced data set or download the data set using the wget method.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "550931": "Hi, could anyone tell how are they handling the large dataset provided here.\n1) Local execution does not work, since data size&gt;available hard-disk space\n2) Tried Using Google Drive, and mounting on Google Collaboratory , but results in Input/Output error. \nCan anyone provide appropriately sized datasets to work here ?",
    "550942": "Meeting the same problems :(",
    "551029": "No need to download entire dataset inorder to participate in the competition. You can probably download any one subset of training data and train your model on that. That would be possible on colab and google drive.",
    "551160": "I put up part of the dataset on [Kaggle that was resized](https://www.kaggle.com/c/open-images-2019-object-detection/discussion/94770#latest-550992). You should be able to use that to download into Google Collaboratory if you don't want to use Kaggle.",
    "558875": "Even a single subset is of 58GB",
    "558930": "You can use [this dataset](https://www.kaggle.com/mindtrinket/google-2019-30k-train/) which is resized to 256*256 and takes only 10 GB of space\n\nThanks to @mindtrinket",
    "559471": "hey use the reduced data set or download the data set using the wget method."
  },
  "source": "meta"
}