{
  "id": 124221,
  "title": "Saving Memory few tips",
  "url": "/competitions/bengaliai-cv19/discussion/124221",
  "author_name": "Amit",
  "post_date": "2020-01-02T17:23:25",
  "votes": 24,
  "comment_count": 12,
  "views": 0,
  "content": "<ol>\n<li><p>As the Data every image pixel is having values from 0-255 always use Datatype uint.</p>\n\n<p>For example \nreshaping the training images Data in parquet file\ntrainX = trainX.values.reshape().astype('uint8')</p>\n\n<p>In the above line unit8 consumes 1 byte where as int consumes 4 byte of memory.</p>\n\n<p>if we have taken int instead of uint it will be 50210 x 32332 cells in 1 parquet file\nwhich is equal to 50210 x 32332  x 3 ~ 4.87 GB extra space taken in RAM </p></li>\n<li><p>Delete Heavy used Dataframe and release the memory using gc.collect()</p>\n\n<p>For example\ntrainData = pd.read_parquet(stringpath  + r'/train_image_data_0.parquet')</p>\n\n<p>once the use of Dataframe  trainData  is over release the memory occupied  </p>\n\n<p>del trainData \ncollected = gc.collect()\nprint(collected) </p></li>\n<li><p>Test Data prediction:\nAlthough we have less number of images in Test Data you can iterate and predict on 1 file parquet file at a time, as during submission it will be tested on large data.\nSo your algo be like</p>\n\n<p>for i in range(4):\n TestDataframe = pd.read_parquet(stringpath  + r'/test_image_data_{}.parquet'.format(i))\n Predict output</p>\n\n<pre><code> Del TestDataframe \n</code></pre>\n\n<p>collected = gc.collect()\n print(collected) </p></li>\n</ol>\n\n<p>These are my observations you can add more in this thread.</p>",
  "messages": [
    {
      "id": 708786,
      "postDate": "2020-01-02T17:23:25Z",
      "content": "<ol>\n<li><p>As the Data every image pixel is having values from 0-255 always use Datatype uint.</p>\n\n<p>For example \nreshaping the training images Data in parquet file\ntrainX = trainX.values.reshape().astype('uint8')</p>\n\n<p>In the above line unit8 consumes 1 byte where as int consumes 4 byte of memory.</p>\n\n<p>if we have taken int instead of uint it will be 50210 x 32332 cells in 1 parquet file\nwhich is equal to 50210 x 32332  x 3 ~ 4.87 GB extra space taken in RAM </p></li>\n<li><p>Delete Heavy used Dataframe and release the memory using gc.collect()</p>\n\n<p>For example\ntrainData = pd.read_parquet(stringpath  + r'/train_image_data_0.parquet')</p>\n\n<p>once the use of Dataframe  trainData  is over release the memory occupied  </p>\n\n<p>del trainData \ncollected = gc.collect()\nprint(collected) </p></li>\n<li><p>Test Data prediction:\nAlthough we have less number of images in Test Data you can iterate and predict on 1 file parquet file at a time, as during submission it will be tested on large data.\nSo your algo be like</p>\n\n<p>for i in range(4):\n TestDataframe = pd.read_parquet(stringpath  + r'/test_image_data_{}.parquet'.format(i))\n Predict output</p>\n\n<pre><code> Del TestDataframe \n</code></pre>\n\n<p>collected = gc.collect()\n print(collected) </p></li>\n</ol>\n\n<p>These are my observations you can add more in this thread.</p>",
      "rawMarkdown": "1. As the Data every image pixel is having values from 0-255 always use Datatype uint.\n   \n   For example \n   reshaping the training images Data in parquet file\n   trainX = trainX.values.reshape().astype('uint8')\n\n   In the above line unit8 consumes 1 byte where as int consumes 4 byte of memory.\n   \n   if we have taken int instead of uint it will be 50210 x 32332 cells in 1 parquet file\n   which is equal to 50210 x 32332  x 3 ~ 4.87 GB extra space taken in RAM \n   \n\n2. Delete Heavy used Dataframe and release the memory using gc.collect()\n   \n   For example\n   trainData = pd.read_parquet(stringpath  + r'/train_image_data_0.parquet')\n   \n   once the use of Dataframe  trainData  is over release the memory occupied  \n   \n   del trainData \n   collected = gc.collect()\n   print(collected) \n\n3. Test Data prediction:\n   Although we have less number of images in Test Data you can iterate and predict on 1 file parquet file at a time, as during submission it will be tested on large data.\n   So your algo be like\n   \n   for i in range(4):\n  \t TestDataframe = pd.read_parquet(stringpath  + r'/test_image_data_{}.parquet'.format(i))\n   \t Predict output\n         \n         Del TestDataframe \n\t collected = gc.collect()\n  \t print(collected) \n         \n        \nThese are my observations you can add more in this thread.",
      "votes": 24
    },
    {
      "id": 737059,
      "postDate": "2020-02-04T21:34:58.293Z",
      "content": "<p>Thanks a lot for sharing <a href=\"/amit9484\">@amit9484</a>, really appreciate</p>",
      "rawMarkdown": "Thanks a lot for sharing @amit9484, really appreciate",
      "votes": 1,
      "replies": [
        {
          "id": 737156,
          "postDate": "2020-02-05T01:35:43.823Z",
          "content": "<p>Thanks <a href=\"/rohitagarwal\">@rohitagarwal</a> </p>",
          "rawMarkdown": "Thanks @rohitagarwal "
        }
      ]
    },
    {
      "id": 732714,
      "postDate": "2020-01-30T06:36:34.133Z",
      "content": "<p><a href=\"/amit9484\">@amit9484</a> I'm using keras pretrained model for my kernel which uses 3channels which can also cause some memory issue. Do you have any idea how to handle this? \nBtw thanks for sharing your observations.</p>",
      "rawMarkdown": "@amit9484 I'm using keras pretrained model for my kernel which uses 3channels which can also cause some memory issue. Do you have any idea how to handle this? \nBtw thanks for sharing your observations.",
      "votes": 1,
      "replies": [
        {
          "id": 732730,
          "postDate": "2020-01-30T07:00:05.517Z",
          "content": "<p><a href=\"/gaur128\">@gaur128</a>  Yes if your image size is large or the function that is resizing the image is not optimized than also you will get error.</p>\n\n<p>Try Reducing image size and if its already less than try replacing test images with train images and keep eye on RAM usage on kernel, by this you can find the areas that need optimization.</p>",
          "rawMarkdown": "@gaur128  Yes if your image size is large or the function that is resizing the image is not optimized than also you will get error.\n\nTry Reducing image size and if its already less than try replacing test images with train images and keep eye on RAM usage on kernel, by this you can find the areas that need optimization."
        }
      ]
    },
    {
      "id": 708799,
      "postDate": "2020-01-02T17:42:04.887Z",
      "content": "<p>Informative!!</p>",
      "rawMarkdown": "Informative!!",
      "votes": 1
    },
    {
      "id": 737183,
      "postDate": "2020-02-05T02:49:11.390Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 737196,
          "postDate": "2020-02-05T03:27:33.693Z",
          "content": "<p>Thanks <a href=\"/coderchava\">@coderchava</a> </p>",
          "rawMarkdown": "Thanks @coderchava "
        }
      ]
    },
    {
      "id": 734981,
      "postDate": "2020-02-02T10:48:56.877Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 3
    },
    {
      "id": 712564,
      "postDate": "2020-01-07T11:35:34.003Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.\n",
      "votes": 1
    },
    {
      "id": 709250,
      "postDate": "2020-01-03T08:53:45.847Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 709102,
      "postDate": "2020-01-03T03:29:53.690Z",
      "content": "<p>Thanks! Very helpful.</p>",
      "rawMarkdown": "Thanks! Very helpful.",
      "votes": 1
    },
    {
      "id": 708796,
      "postDate": "2020-01-02T17:35:56.433Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 737059,
      "author_name": "Rohit Agarwal",
      "author_url": "",
      "post_date": "2020-02-04T21:34:58.293000",
      "content": "<p>Thanks a lot for sharing <a href=\"/amit9484\">@amit9484</a>, really appreciate</p>",
      "votes": 1,
      "replies": [
        {
          "id": 737156,
          "author_name": "Amit",
          "author_url": "",
          "post_date": "2020-02-05T01:35:43.823000",
          "content": "<p>Thanks <a href=\"/rohitagarwal\">@rohitagarwal</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 732714,
      "author_name": "Gaurav Yadav",
      "author_url": "",
      "post_date": "2020-01-30T06:36:34.133000",
      "content": "<p><a href=\"/amit9484\">@amit9484</a> I'm using keras pretrained model for my kernel which uses 3channels which can also cause some memory issue. Do you have any idea how to handle this? \nBtw thanks for sharing your observations.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 732730,
          "author_name": "Amit",
          "author_url": "",
          "post_date": "2020-01-30T07:00:05.517000",
          "content": "<p><a href=\"/gaur128\">@gaur128</a>  Yes if your image size is large or the function that is resizing the image is not optimized than also you will get error.</p>\n\n<p>Try Reducing image size and if its already less than try replacing test images with train images and keep eye on RAM usage on kernel, by this you can find the areas that need optimization.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 708799,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-02T17:42:04.887000",
      "content": "<p>Informative!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 737183,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-05T02:49:11.390000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 737196,
          "author_name": "Amit",
          "author_url": "",
          "post_date": "2020-02-05T03:27:33.693000",
          "content": "<p>Thanks <a href=\"/coderchava\">@coderchava</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 734981,
      "author_name": "Sumit Mishra",
      "author_url": "",
      "post_date": "2020-02-02T10:48:56.877000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 712564,
      "author_name": "Anshul Tomar",
      "author_url": "",
      "post_date": "2020-01-07T11:35:34.003000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 709250,
      "author_name": "Sang Thieu",
      "author_url": "",
      "post_date": "2020-01-03T08:53:45.847000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 709102,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2020-01-03T03:29:53.690000",
      "content": "<p>Thanks! Very helpful.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 708796,
      "author_name": "DtneSEffct",
      "author_url": "",
      "post_date": "2020-01-02T17:35:56.433000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "708786": "1. As the Data every image pixel is having values from 0-255 always use Datatype uint.\n   \n   For example \n   reshaping the training images Data in parquet file\n   trainX = trainX.values.reshape().astype('uint8')\n\n   In the above line unit8 consumes 1 byte where as int consumes 4 byte of memory.\n   \n   if we have taken int instead of uint it will be 50210 x 32332 cells in 1 parquet file\n   which is equal to 50210 x 32332  x 3 ~ 4.87 GB extra space taken in RAM \n   \n\n2. Delete Heavy used Dataframe and release the memory using gc.collect()\n   \n   For example\n   trainData = pd.read_parquet(stringpath  + r'/train_image_data_0.parquet')\n   \n   once the use of Dataframe  trainData  is over release the memory occupied  \n   \n   del trainData \n   collected = gc.collect()\n   print(collected) \n\n3. Test Data prediction:\n   Although we have less number of images in Test Data you can iterate and predict on 1 file parquet file at a time, as during submission it will be tested on large data.\n   So your algo be like\n   \n   for i in range(4):\n  \t TestDataframe = pd.read_parquet(stringpath  + r'/test_image_data_{}.parquet'.format(i))\n   \t Predict output\n         \n         Del TestDataframe \n\t collected = gc.collect()\n  \t print(collected) \n         \n        \nThese are my observations you can add more in this thread.",
    "737059": "Thanks a lot for sharing @amit9484, really appreciate",
    "732714": "@amit9484 I'm using keras pretrained model for my kernel which uses 3channels which can also cause some memory issue. Do you have any idea how to handle this? \nBtw thanks for sharing your observations.",
    "708799": "Informative!!",
    "737183": "",
    "734981": "Thanks for sharing",
    "712564": "Thanks for sharing.\n",
    "709250": "Thanks for sharing!",
    "709102": "Thanks! Very helpful.",
    "708796": "Thanks for sharing!"
  }
}