{
  "id": 178698,
  "title": "How to deal with large amount of data with only 13 GB RAM ??",
  "url": "/competitions/landmark-recognition-2020/discussion/178698",
  "author_name": "",
  "post_date": "2020-08-31T04:32:30.881537300Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I was working on another image classification problem where I have faced exhaust RAM issue ( notebook is trying to allocate more RAM that available). There the data set was smaller compare to this data set. I barely managed to complete the task there. Here the data set size is very large.</p>\n<p>Can anyone please guide me on how to deal with this issue ?  </p>\n<p>Thanks :)</p>",
  "messages": [
    {
      "id": "992245",
      "postDate": "08/31/2020 04:32:30",
      "content": "<p>I was working on another image classification problem where I have faced exhaust RAM issue ( notebook is trying to allocate more RAM that available). There the data set was smaller compare to this data set. I barely managed to complete the task there. Here the data set size is very large.</p>\n<p>Can anyone please guide me on how to deal with this issue ?  </p>\n<p>Thanks :)</p>",
      "rawMarkdown": "I was working on another image classification problem where I have faced exhaust RAM issue ( notebook is trying to allocate more RAM that available). There the data set was smaller compare to this data set. I barely managed to complete the task there. Here the data set size is very large.\n\nCan anyone please guide me on how to deal with this issue ?  \n\nThanks :)",
      "votes": null
    },
    {
      "id": "992290",
      "postDate": "08/31/2020 05:07:10",
      "content": "<p>Hey,<br>\nTry to reduce the batch size. You can also try to reduce the size of the image.</p>\n<p>Hope this helps!</p>",
      "rawMarkdown": "Hey,\nTry to reduce the batch size. You can also try to reduce the size of the image.\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "992899",
      "postDate": "08/31/2020 14:04:17",
      "content": "<p>Hi, try to use imagedatagenerator() in Keras. It'll be useful</p>",
      "rawMarkdown": "Hi, try to use imagedatagenerator() in Keras. It'll be useful",
      "votes": null
    },
    {
      "id": "992913",
      "postDate": "08/31/2020 14:09:56",
      "content": "<p>I would suggest using a cloud service.But in the meantime,you could use google colab scholar notebooks,they provide free GPU use upto a threshold.\n<a href=\"https://colab.research.google.com/\">https://colab.research.google.com/</a></p>",
      "rawMarkdown": "I would suggest using a cloud service.But in the meantime,you could use google colab scholar notebooks,they provide free GPU use upto a threshold.\nhttps://colab.research.google.com/",
      "votes": null
    },
    {
      "id": "992970",
      "postDate": "08/31/2020 14:51:16",
      "content": "<p>Hi Deepak. You could try one of the below. Many other strategies also exist.</p>\n\n<ul>\n<li>You can use online algorithms that only small amounts of the information into RAM at a time, such as here: <a href=\"https://scikit-learn.org/0.15/modules/scaling_strategies.html\">https://scikit-learn.org/0.15/modules/scaling_strategies.html</a></li>\n<li>You could also configure a virtual machine and use SDD or HDD as VRAM, but the configuration may be very slow.</li>\n<li>You could try using a cloud product, like Google Collab Pro (10USD/month though).</li>\n</ul>",
      "rawMarkdown": "Hi Deepak. You could try one of the below. Many other strategies also exist.\n\n- You can use online algorithms that only small amounts of the information into RAM at a time, such as here: https://scikit-learn.org/0.15/modules/scaling_strategies.html\n- You could also configure a virtual machine and use SDD or HDD as VRAM, but the configuration may be very slow.\n- You could try using a cloud product, like Google Collab Pro (10USD/month though).",
      "votes": null
    },
    {
      "id": "993031",
      "postDate": "08/31/2020 15:51:29",
      "content": "<p>I think we can go for colab pro. But, how to proceed with this competition on kaggle ? Because we have to put the whole notebook here. Please suggest 😀</p>",
      "rawMarkdown": "I think we can go for colab pro. But, how to proceed with this competition on kaggle ? Because we have to put the whole notebook here. Please suggest 😀",
      "votes": null
    },
    {
      "id": "993243",
      "postDate": "08/31/2020 18:48:26",
      "content": "<p>Hello, you can train in batches with gpu or tpu, this sould not give you any problems because normal ram will only store single prediction vectors. Nevertheless you can still have problems with the ram of gpu or tpu, their you have to configue correctly your data pipeline and model. For prediction you should also predict in batches (you can increase the size of the batch for prediction). Next, you can use this trained model for public and private test inference(prediction), you just need to upload your model as a dataset.</p>\n<p>Now, im running a notebook with colab that just use 3 gb ram and predicts 66k images in 30 minutes using a dataset with image size 384 x 384 (model = efficientnetB0)</p>\n<p>You can check Chris notebook in melanoma comp to check how to do it in tensorflow with tf records.</p>",
      "rawMarkdown": "Hello, you can train in batches with gpu or tpu, this sould not give you any problems because normal ram will only store single prediction vectors. Nevertheless you can still have problems with the ram of gpu or tpu, their you have to configue correctly your data pipeline and model. For prediction you should also predict in batches (you can increase the size of the batch for prediction). Next, you can use this trained model for public and private test inference(prediction), you just need to upload your model as a dataset.\n\nNow, im running a notebook with colab that just use 3 gb ram and predicts 66k images in 30 minutes using a dataset with image size 384 x 384 (model = efficientnetB0)\n\nYou can check Chris notebook in melanoma comp to check how to do it in tensorflow with tf records.",
      "votes": null
    },
    {
      "id": "999167",
      "postDate": "09/05/2020 12:31:45",
      "content": "<p>Hi Deepak. You could train the models separately to the notebook you submit and import them into the runtime of the notebook on Kaggle. You could also call cloud resources from the notebook on Kaggle.</p>",
      "rawMarkdown": "Hi Deepak. You could train the models separately to the notebook you submit and import them into the runtime of the notebook on Kaggle. You could also call cloud resources from the notebook on Kaggle.",
      "votes": null
    },
    {
      "id": "1008690",
      "postDate": "09/13/2020 10:12:44",
      "content": "<p>Hey, maybe it can help you: <a href=\"https://www.kaggle.com/pavansanagapati/14-simple-tips-to-save-ram-memory-for-1-gb-dataset\" target=\"_blank\">https://www.kaggle.com/pavansanagapati/14-simple-tips-to-save-ram-memory-for-1-gb-dataset</a></p>",
      "rawMarkdown": "Hey, maybe it can help you: https://www.kaggle.com/pavansanagapati/14-simple-tips-to-save-ram-memory-for-1-gb-dataset",
      "votes": null
    },
    {
      "id": "1024091",
      "postDate": "09/23/2020 16:32:46",
      "content": "<p><a href=\"https://www.kaggle.com/ragnar\" target=\"_blank\">@ragnar</a> : Thank you very much :)</p>",
      "rawMarkdown": "ragnar : Thank you very much :)",
      "votes": null
    },
    {
      "id": "1024594",
      "postDate": "09/24/2020 02:18:44",
      "content": "<p>Make use of the <a href=\"https://www.kaggle.com/jkreddy/gld20gb/\" target=\"_blank\">dataset</a> generated using the <a href=\"https://www.kaggle.com/jkreddy/kaggle-github-integration-to-work-on-large-dataset\" target=\"_blank\">notebook</a><br>\nIt is compressed dataset(~20GB). Images have been reduced to size (224,320) preserving aspect ratio. You can use this dataset to train on colab also.</p>",
      "rawMarkdown": "Make use of the [dataset](https://www.kaggle.com/jkreddy/gld20gb/) generated using the [notebook](https://www.kaggle.com/jkreddy/kaggle-github-integration-to-work-on-large-dataset)\nIt is compressed dataset(~20GB). Images have been reduced to size (224,320) preserving aspect ratio. You can use this dataset to train on colab also.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 992290,
      "author_name": "akshaychavan123",
      "author_url": "",
      "post_date": "08/31/2020 05:07:10",
      "content": "<p>Hey,<br>\nTry to reduce the batch size. You can also try to reduce the size of the image.</p>\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1008690,
      "author_name": "",
      "author_url": "",
      "post_date": "09/13/2020 10:12:44",
      "content": "<p>Hey, maybe it can help you: <a href=\"https://www.kaggle.com/pavansanagapati/14-simple-tips-to-save-ram-memory-for-1-gb-dataset\" target=\"_blank\">https://www.kaggle.com/pavansanagapati/14-simple-tips-to-save-ram-memory-for-1-gb-dataset</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1024594,
      "author_name": "jkreddy",
      "author_url": "",
      "post_date": "09/24/2020 02:18:44",
      "content": "<p>Make use of the <a href=\"https://www.kaggle.com/jkreddy/gld20gb/\" target=\"_blank\">dataset</a> generated using the <a href=\"https://www.kaggle.com/jkreddy/kaggle-github-integration-to-work-on-large-dataset\" target=\"_blank\">notebook</a><br>\nIt is compressed dataset(~20GB). Images have been reduced to size (224,320) preserving aspect ratio. You can use this dataset to train on colab also.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 992899,
      "author_name": "jason022085",
      "author_url": "",
      "post_date": "08/31/2020 14:04:17",
      "content": "<p>Hi, try to use imagedatagenerator() in Keras. It'll be useful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 992913,
      "author_name": "aldorain",
      "author_url": "",
      "post_date": "08/31/2020 14:09:56",
      "content": "<p>I would suggest using a cloud service.But in the meantime,you could use google colab scholar notebooks,they provide free GPU use upto a threshold.\n<a href=\"https://colab.research.google.com/\">https://colab.research.google.com/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 992970,
      "author_name": "caseyj2",
      "author_url": "",
      "post_date": "08/31/2020 14:51:16",
      "content": "<p>Hi Deepak. You could try one of the below. Many other strategies also exist.</p>\n\n<ul>\n<li>You can use online algorithms that only small amounts of the information into RAM at a time, such as here: <a href=\"https://scikit-learn.org/0.15/modules/scaling_strategies.html\">https://scikit-learn.org/0.15/modules/scaling_strategies.html</a></li>\n<li>You could also configure a virtual machine and use SDD or HDD as VRAM, but the configuration may be very slow.</li>\n<li>You could try using a cloud product, like Google Collab Pro (10USD/month though).</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 993031,
          "author_name": "deepakat002",
          "author_url": "",
          "post_date": "08/31/2020 15:51:29",
          "content": "<p>I think we can go for colab pro. But, how to proceed with this competition on kaggle ? Because we have to put the whole notebook here. Please suggest 😀</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 993243,
          "author_name": "ragnar123",
          "author_url": "",
          "post_date": "08/31/2020 18:48:26",
          "content": "<p>Hello, you can train in batches with gpu or tpu, this sould not give you any problems because normal ram will only store single prediction vectors. Nevertheless you can still have problems with the ram of gpu or tpu, their you have to configue correctly your data pipeline and model. For prediction you should also predict in batches (you can increase the size of the batch for prediction). Next, you can use this trained model for public and private test inference(prediction), you just need to upload your model as a dataset.</p>\n<p>Now, im running a notebook with colab that just use 3 gb ram and predicts 66k images in 30 minutes using a dataset with image size 384 x 384 (model = efficientnetB0)</p>\n<p>You can check Chris notebook in melanoma comp to check how to do it in tensorflow with tf records.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 999167,
          "author_name": "caseyj2",
          "author_url": "",
          "post_date": "09/05/2020 12:31:45",
          "content": "<p>Hi Deepak. You could train the models separately to the notebook you submit and import them into the runtime of the notebook on Kaggle. You could also call cloud resources from the notebook on Kaggle.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1024091,
          "author_name": "deepakat002",
          "author_url": "",
          "post_date": "09/23/2020 16:32:46",
          "content": "<p><a href=\"https://www.kaggle.com/ragnar\" target=\"_blank\">@ragnar</a> : Thank you very much :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "992245": "I was working on another image classification problem where I have faced exhaust RAM issue ( notebook is trying to allocate more RAM that available). There the data set was smaller compare to this data set. I barely managed to complete the task there. Here the data set size is very large.\n\nCan anyone please guide me on how to deal with this issue ?  \n\nThanks :)",
    "992290": "Hey,\nTry to reduce the batch size. You can also try to reduce the size of the image.\n\nHope this helps!",
    "992899": "Hi, try to use imagedatagenerator() in Keras. It'll be useful",
    "992913": "I would suggest using a cloud service.But in the meantime,you could use google colab scholar notebooks,they provide free GPU use upto a threshold.\nhttps://colab.research.google.com/",
    "992970": "Hi Deepak. You could try one of the below. Many other strategies also exist.\n\n- You can use online algorithms that only small amounts of the information into RAM at a time, such as here: https://scikit-learn.org/0.15/modules/scaling_strategies.html\n- You could also configure a virtual machine and use SDD or HDD as VRAM, but the configuration may be very slow.\n- You could try using a cloud product, like Google Collab Pro (10USD/month though).",
    "993031": "I think we can go for colab pro. But, how to proceed with this competition on kaggle ? Because we have to put the whole notebook here. Please suggest 😀",
    "993243": "Hello, you can train in batches with gpu or tpu, this sould not give you any problems because normal ram will only store single prediction vectors. Nevertheless you can still have problems with the ram of gpu or tpu, their you have to configue correctly your data pipeline and model. For prediction you should also predict in batches (you can increase the size of the batch for prediction). Next, you can use this trained model for public and private test inference(prediction), you just need to upload your model as a dataset.\n\nNow, im running a notebook with colab that just use 3 gb ram and predicts 66k images in 30 minutes using a dataset with image size 384 x 384 (model = efficientnetB0)\n\nYou can check Chris notebook in melanoma comp to check how to do it in tensorflow with tf records.",
    "999167": "Hi Deepak. You could train the models separately to the notebook you submit and import them into the runtime of the notebook on Kaggle. You could also call cloud resources from the notebook on Kaggle.",
    "1008690": "Hey, maybe it can help you: https://www.kaggle.com/pavansanagapati/14-simple-tips-to-save-ram-memory-for-1-gb-dataset",
    "1024091": "ragnar : Thank you very much :)",
    "1024594": "Make use of the [dataset](https://www.kaggle.com/jkreddy/gld20gb/) generated using the [notebook](https://www.kaggle.com/jkreddy/kaggle-github-integration-to-work-on-large-dataset)\nIt is compressed dataset(~20GB). Images have been reduced to size (224,320) preserving aspect ratio. You can use this dataset to train on colab also."
  },
  "source": "meta"
}