{
  "id": 172955,
  "title": "How to handle such huge size dataset ??",
  "url": "/competitions/landmark-recognition-2020/discussion/172955",
  "author_name": "",
  "post_date": "2020-08-07T06:43:06.590730500Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I mostly work using Colab  , so for working with a dataset of size 98 gb . Can someone tell how are they handling this data . As in do you download the data onto your machine ,Or is their a resized format of the data ? . Cause inference on such large dataset would take a lot of time.</p>",
  "messages": [
    {
      "id": "961398",
      "postDate": "08/07/2020 06:43:06",
      "content": "<p>I mostly work using Colab  , so for working with a dataset of size 98 gb . Can someone tell how are they handling this data . As in do you download the data onto your machine ,Or is their a resized format of the data ? . Cause inference on such large dataset would take a lot of time.</p>",
      "rawMarkdown": "I mostly work using Colab  , so for working with a dataset of size 98 gb . Can someone tell how are they handling this data . As in do you download the data onto your machine ,Or is their a resized format of the data ? . Cause inference on such large dataset would take a lot of time.",
      "votes": null
    },
    {
      "id": "962001",
      "postDate": "08/07/2020 17:42:08",
      "content": "<p>If your data is in csv format, you can try converting them to .pkl files. Compression would not hurt too. And if I'm not wrong, colab gives you about 117gb, so unless your colab is crashing, you should have enough space.</p>",
      "rawMarkdown": "If your data is in csv format, you can try converting them to .pkl files. Compression would not hurt too. And if I'm not wrong, colab gives you about 117gb, so unless your colab is crashing, you should have enough space.",
      "votes": null
    },
    {
      "id": "965028",
      "postDate": "08/10/2020 10:37:13",
      "content": "<p>Hello,\nI guess this notebook will be helpful to you.\n<a href=\"https://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size\">https://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size</a></p>\n\n<p>Thank you</p>",
      "rawMarkdown": "Hello,\nI guess this notebook will be helpful to you.\nhttps://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size\n\nThank you",
      "votes": null
    },
    {
      "id": "965643",
      "postDate": "08/10/2020 18:33:28",
      "content": "<p>Here is my approach: </p>\n\n<ul>\n<li>donwsample the dataset to get a more balanced one, i.e. remove landmarks with a lot of images for example</li>\n<li>resize images (so that I can train on my desktop)</li>\n<li>run the network and wait :p </li>\n</ul>\n\n<p>You can also use these resized <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord\"><strong>TFRecords</strong></a>: <a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/172231\">https://www.kaggle.com/c/landmark-recognition-2020/discussion/172231</a></p>",
      "rawMarkdown": "Here is my approach: \n\n- donwsample the dataset to get a more balanced one, i.e. remove landmarks with a lot of images for example\n- resize images (so that I can train on my desktop)\n- run the network and wait :p \n\nYou can also use these resized [**TFRecords**](https://www.tensorflow.org/tutorials/load_data/tfrecord): https://www.kaggle.com/c/landmark-recognition-2020/discussion/172231",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 962001,
      "author_name": "ahmnazmussakib",
      "author_url": "",
      "post_date": "08/07/2020 17:42:08",
      "content": "<p>If your data is in csv format, you can try converting them to .pkl files. Compression would not hurt too. And if I'm not wrong, colab gives you about 117gb, so unless your colab is crashing, you should have enough space.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 965028,
      "author_name": "palaksood97",
      "author_url": "",
      "post_date": "08/10/2020 10:37:13",
      "content": "<p>Hello,\nI guess this notebook will be helpful to you.\n<a href=\"https://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size\">https://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size</a></p>\n\n<p>Thank you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 965643,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "08/10/2020 18:33:28",
      "content": "<p>Here is my approach: </p>\n\n<ul>\n<li>donwsample the dataset to get a more balanced one, i.e. remove landmarks with a lot of images for example</li>\n<li>resize images (so that I can train on my desktop)</li>\n<li>run the network and wait :p </li>\n</ul>\n\n<p>You can also use these resized <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord\"><strong>TFRecords</strong></a>: <a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/172231\">https://www.kaggle.com/c/landmark-recognition-2020/discussion/172231</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "961398": "I mostly work using Colab  , so for working with a dataset of size 98 gb . Can someone tell how are they handling this data . As in do you download the data onto your machine ,Or is their a resized format of the data ? . Cause inference on such large dataset would take a lot of time.",
    "962001": "If your data is in csv format, you can try converting them to .pkl files. Compression would not hurt too. And if I'm not wrong, colab gives you about 117gb, so unless your colab is crashing, you should have enough space.",
    "965028": "Hello,\nI guess this notebook will be helpful to you.\nhttps://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size\n\nThank you",
    "965643": "Here is my approach: \n\n- donwsample the dataset to get a more balanced one, i.e. remove landmarks with a lot of images for example\n- resize images (so that I can train on my desktop)\n- run the network and wait :p \n\nYou can also use these resized [**TFRecords**](https://www.tensorflow.org/tutorials/load_data/tfrecord): https://www.kaggle.com/c/landmark-recognition-2020/discussion/172231"
  },
  "source": "meta"
}