{
  "id": 226500,
  "title": "Memory crash with more than 15K images",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/226500",
  "author_name": "GitMach",
  "post_date": "2021-03-16T16:36:28.958000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi, <br>\nI'm trying to train my model on a 20K image sample and the notebook seems to be overwhelmed with so many images. 😳<br>\nSo, my question, how have  you done to train a model with all RGBY images, which means 82K images approximately ?<br>\nBecause, this makes me remember about a story I read from a guy who wanted to start learning Machine Learning and Python in a Pentium4, 4MB Ram, 1Gb Hd, and Windows 98 <br>\nOr , I'm missing something… probably </p>",
  "messages": [
    {
      "id": 1240821,
      "postDate": "2021-03-16T16:36:28.960Z",
      "content": "<p>Hi, <br>\nI'm trying to train my model on a 20K image sample and the notebook seems to be overwhelmed with so many images. 😳<br>\nSo, my question, how have  you done to train a model with all RGBY images, which means 82K images approximately ?<br>\nBecause, this makes me remember about a story I read from a guy who wanted to start learning Machine Learning and Python in a Pentium4, 4MB Ram, 1Gb Hd, and Windows 98 <br>\nOr , I'm missing something… probably </p>",
      "rawMarkdown": "Hi, \nI'm trying to train my model on a 20K image sample and the notebook seems to be overwhelmed with so many images. 😳\nSo, my question, how have  you done to train a model with all RGBY images, which means 82K images approximately ?\nBecause, this makes me remember about a story I read from a guy who wanted to start learning Machine Learning and Python in a Pentium4, 4MB Ram, 1Gb Hd, and Windows 98 \nOr , I'm missing something... probably \n",
      "votes": 1
    },
    {
      "id": 1242273,
      "postDate": "2021-03-17T14:06:15.997Z",
      "content": "<p>Hi there, I think I can help.<br><br>\nThis could be a few different things. The two most obvious things (in my opinion) that could result in this are listed below… it might be something else though so don't take it as gospel.</p>\n<hr>\n<p><strong>1. You are plotting or displaying large amounts of data within the kernel. I have noticed this can cause the notebook to run INCREDIBLY slowly and freeze. i.e. plotting a large number of images or printing a large number of lines (10000+). This can be cumulative over many cells or a single cell with a large output.</strong></p>\n<hr>\n<p><strong>2. You are trying to load too many images into memory at once for training.</strong></p>\n<ul>\n<li>You cannot load tons of images into memory as you are limited by the CPU/GPU RAM (depending if you go OOM during training or before that)</li>\n<li>To overcome this you will need to stream the data somehow. i.e. graph based execution.</li>\n<li>In Tensorflow you would use something like <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset\" target=\"_blank\"><strong><code>tf.data</code></strong></a> to set up graph based execution for streaming data.</li>\n<li>Essentially you build a pipeline that does all the steps but only for a batch (or <strong><code>n</code></strong> batches) at a time.</li>\n</ul>\n<p><strong><em>Here's a basic example showing how you only need the memory to support <strong><code>m</code></strong> images in memory.</em></strong></p>\n<hr>\n<blockquote>\n  <ol>\n  <li>Load A Batch of Images Into Memory (<strong><code>m</code></strong> images)</li>\n  <li>Preprocess Batch of Images</li>\n  <li>Augment Batch of Images</li>\n  <li>Infer and Update Model Using Batch of Images</li>\n  <li>Clear Memory and Start Again</li>\n  <li>This Continues Until the Epoch is Complete and Then Usually the Data Is Shuffled (Or It's Done During The Epoch <strong><code>s</code></strong> Number of Examples Ahead)</li>\n  </ol>\n</blockquote>\n<hr>\n<p>Hope this helps! Alternatively, feel free to share a link to the kernel if you need more detailed troubleshooting.</p>",
      "rawMarkdown": "Hi there, I think I can help.<br>\nThis could be a few different things. The two most obvious things (in my opinion) that could result in this are listed below... it might be something else though so don't take it as gospel.\n\n---\n\n**1. You are plotting or displaying large amounts of data within the kernel. I have noticed this can cause the notebook to run INCREDIBLY slowly and freeze. i.e. plotting a large number of images or printing a large number of lines (10000+). This can be cumulative over many cells or a single cell with a large output.**\n\n---\n\n**2. You are trying to load too many images into memory at once for training.**\n* You cannot load tons of images into memory as you are limited by the CPU/GPU RAM (depending if you go OOM during training or before that)\n* To overcome this you will need to stream the data somehow. i.e. graph based execution.\n* In Tensorflow you would use something like [**`tf.data`**](https://www.tensorflow.org/api_docs/python/tf/data/Dataset) to set up graph based execution for streaming data.\n* Essentially you build a pipeline that does all the steps but only for a batch (or **`n`** batches) at a time.\n\n***Here's a basic example showing how you only need the memory to support **`m`** images in memory.***\n\n----\n\n>1. Load A Batch of Images Into Memory (**`m`** images)\n2. Preprocess Batch of Images\n3. Augment Batch of Images\n4. Infer and Update Model Using Batch of Images\n5. Clear Memory and Start Again\n6. This Continues Until the Epoch is Complete and Then Usually the Data Is Shuffled (Or It's Done During The Epoch **`s`** Number of Examples Ahead)\n\n---\n\nHope this helps! Alternatively, feel free to share a link to the kernel if you need more detailed troubleshooting.",
      "votes": 2,
      "replies": [
        {
          "id": 1243951,
          "postDate": "2021-03-18T15:53:10.590Z",
          "content": "<p>Hi Darien, <br>\nThank you for your feedback and explanations.  <br>\nI was reading about \"Batch and loading images\".<br>\nHere're methods and some tests I've been running so far:</p>\n<ol>\n<li>I've created a dataframe (train_data_rgby.csv) containing \"image_Path\" (Path and filenames to .png files\" and \"Labels\" in \"OneHotEncod\" format (19 Variables, 1 variable for each class) from all images in the train </li>\n<li>I took a random sample (10%) from this Dataframe and I build a 1st model with Keras (Flow_from_Dataframe) - PB - GPU doesn't work with \"Dataframe\" </li>\n<li>Using Same Random Sample I resize image with Opencv and convert result into 3D Tensor. I build a new model with keras using 3D Tensor numpy - GPU was working perfectly, train time was vey fast, about 2s per epoch - Inconvenients  : The array size it's huge : for this 10% Sample with 7k images the size of the array was about 5G, so we can easily imagine how much it would make all images from the dataset. Besides, when I added a validation set of 2000 Images the notebook crashed with the memory problem.<br>\n4.From same sample I did image augmentation with openCV (Rotation, HorizontalFlip) , same result, GPU works fine, training model was fast, but same issue with the size of Numpy Array.<br>\n<strong>## Conclusion :</strong> So far I understood, GPU doesn't work with \"Pandas Dataframe\" at least in \"Kaggle\" and extract features and convert into Numpy and train model it's faster but we get a problem of Memory because of Numpy Array Size<br>\nSo, the idea of loading batch images into memory seems to be a good solution as I don't need to load \"Numpy arrays into memory\".  The issue, I didn't found to much info about this type of function.<br>\nHere's the link to the Kernel and thank you for your help.<br>\n<a href=\"https://www.kaggle.com/centorit/hpa-model-train\" target=\"_blank\">https://www.kaggle.com/centorit/hpa-model-train</a></li>\n</ol>",
          "rawMarkdown": "Hi Darien, \nThank you for your feedback and explanations.  \nI was reading about \"Batch and loading images\".\nHere're methods and some tests I've been running so far:\n1. I've created a dataframe (train_data_rgby.csv) containing \"image_Path\" (Path and filenames to .png files\" and \"Labels\" in \"OneHotEncod\" format (19 Variables, 1 variable for each class) from all images in the train \n2. I took a random sample (10%) from this Dataframe and I build a 1st model with Keras (Flow_from_Dataframe) - PB - GPU doesn't work with \"Dataframe\" \n3. Using Same Random Sample I resize image with Opencv and convert result into 3D Tensor. I build a new model with keras using 3D Tensor numpy - GPU was working perfectly, train time was vey fast, about 2s per epoch - Inconvenients  : The array size it's huge : for this 10% Sample with 7k images the size of the array was about 5G, so we can easily imagine how much it would make all images from the dataset. Besides, when I added a validation set of 2000 Images the notebook crashed with the memory problem.\n4.From same sample I did image augmentation with openCV (Rotation, HorizontalFlip) , same result, GPU works fine, training model was fast, but same issue with the size of Numpy Array.\n**## Conclusion :** So far I understood, GPU doesn't work with \"Pandas Dataframe\" at least in \"Kaggle\" and extract features and convert into Numpy and train model it's faster but we get a problem of Memory because of Numpy Array Size\nSo, the idea of loading batch images into memory seems to be a good solution as I don't need to load \"Numpy arrays into memory\".  The issue, I didn't found to much info about this type of function.\nHere's the link to the Kernel and thank you for your help.\nhttps://www.kaggle.com/centorit/hpa-model-train\n "
        },
        {
          "id": 1244024,
          "postDate": "2021-03-18T16:52:02.957Z",
          "content": "<p>Thanks for all the info.</p>\n<p>I'll have a look later today or tomorrow and get back to you.</p>",
          "rawMarkdown": "Thanks for all the info.\n\nI'll have a look later today or tomorrow and get back to you."
        },
        {
          "id": 1246507,
          "postDate": "2021-03-20T19:35:09.620Z",
          "content": "<p><a href=\"https://www.kaggle.com/darien\" target=\"_blank\">@darien</a></p>\n<pre><code>Load A Batch of Images Into Memory (m images)\nPreprocess Batch of Images\nAugment Batch of Images\nInfer and Update Model Using Batch of Images\nClear Memory and Start Again  \nThis Continues Until the Epoch is Complete and Then Usually the Data Is Shuffled (Or It's Done During The Epoch s Number of Examples Ahead)\n</code></pre>\n<p>This \"continuous training\" approach allows to train much more images indeed. <br>\nNevertheless, memory slightly increases over each batch, but this time, with this \"training inside loop of batch images\" I was able to train 60K images.<br>\nThank you for you feedback </p>",
          "rawMarkdown": "@darien\n\n```\nLoad A Batch of Images Into Memory (m images)\nPreprocess Batch of Images\nAugment Batch of Images\nInfer and Update Model Using Batch of Images\nClear Memory and Start Again  \nThis Continues Until the Epoch is Complete and Then Usually the Data Is Shuffled (Or It's Done During The Epoch s Number of Examples Ahead)\n```\n\nThis \"continuous training\" approach allows to train much more images indeed. \nNevertheless, memory slightly increases over each batch, but this time, with this \"training inside loop of batch images\" I was able to train 60K images.\nThank you for you feedback \n\n"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1242273,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-03-17T14:06:15.997000",
      "content": "<p>Hi there, I think I can help.<br><br>\nThis could be a few different things. The two most obvious things (in my opinion) that could result in this are listed below… it might be something else though so don't take it as gospel.</p>\n<hr>\n<p><strong>1. You are plotting or displaying large amounts of data within the kernel. I have noticed this can cause the notebook to run INCREDIBLY slowly and freeze. i.e. plotting a large number of images or printing a large number of lines (10000+). This can be cumulative over many cells or a single cell with a large output.</strong></p>\n<hr>\n<p><strong>2. You are trying to load too many images into memory at once for training.</strong></p>\n<ul>\n<li>You cannot load tons of images into memory as you are limited by the CPU/GPU RAM (depending if you go OOM during training or before that)</li>\n<li>To overcome this you will need to stream the data somehow. i.e. graph based execution.</li>\n<li>In Tensorflow you would use something like <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset\" target=\"_blank\"><strong><code>tf.data</code></strong></a> to set up graph based execution for streaming data.</li>\n<li>Essentially you build a pipeline that does all the steps but only for a batch (or <strong><code>n</code></strong> batches) at a time.</li>\n</ul>\n<p><strong><em>Here's a basic example showing how you only need the memory to support <strong><code>m</code></strong> images in memory.</em></strong></p>\n<hr>\n<blockquote>\n  <ol>\n  <li>Load A Batch of Images Into Memory (<strong><code>m</code></strong> images)</li>\n  <li>Preprocess Batch of Images</li>\n  <li>Augment Batch of Images</li>\n  <li>Infer and Update Model Using Batch of Images</li>\n  <li>Clear Memory and Start Again</li>\n  <li>This Continues Until the Epoch is Complete and Then Usually the Data Is Shuffled (Or It's Done During The Epoch <strong><code>s</code></strong> Number of Examples Ahead)</li>\n  </ol>\n</blockquote>\n<hr>\n<p>Hope this helps! Alternatively, feel free to share a link to the kernel if you need more detailed troubleshooting.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1243951,
          "author_name": "GitMach",
          "author_url": "",
          "post_date": "2021-03-18T15:53:10.590000",
          "content": "<p>Hi Darien, <br>\nThank you for your feedback and explanations.  <br>\nI was reading about \"Batch and loading images\".<br>\nHere're methods and some tests I've been running so far:</p>\n<ol>\n<li>I've created a dataframe (train_data_rgby.csv) containing \"image_Path\" (Path and filenames to .png files\" and \"Labels\" in \"OneHotEncod\" format (19 Variables, 1 variable for each class) from all images in the train </li>\n<li>I took a random sample (10%) from this Dataframe and I build a 1st model with Keras (Flow_from_Dataframe) - PB - GPU doesn't work with \"Dataframe\" </li>\n<li>Using Same Random Sample I resize image with Opencv and convert result into 3D Tensor. I build a new model with keras using 3D Tensor numpy - GPU was working perfectly, train time was vey fast, about 2s per epoch - Inconvenients  : The array size it's huge : for this 10% Sample with 7k images the size of the array was about 5G, so we can easily imagine how much it would make all images from the dataset. Besides, when I added a validation set of 2000 Images the notebook crashed with the memory problem.<br>\n4.From same sample I did image augmentation with openCV (Rotation, HorizontalFlip) , same result, GPU works fine, training model was fast, but same issue with the size of Numpy Array.<br>\n<strong>## Conclusion :</strong> So far I understood, GPU doesn't work with \"Pandas Dataframe\" at least in \"Kaggle\" and extract features and convert into Numpy and train model it's faster but we get a problem of Memory because of Numpy Array Size<br>\nSo, the idea of loading batch images into memory seems to be a good solution as I don't need to load \"Numpy arrays into memory\".  The issue, I didn't found to much info about this type of function.<br>\nHere's the link to the Kernel and thank you for your help.<br>\n<a href=\"https://www.kaggle.com/centorit/hpa-model-train\" target=\"_blank\">https://www.kaggle.com/centorit/hpa-model-train</a></li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1244024,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-03-18T16:52:02.957000",
          "content": "<p>Thanks for all the info.</p>\n<p>I'll have a look later today or tomorrow and get back to you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1246507,
          "author_name": "GitMach",
          "author_url": "",
          "post_date": "2021-03-20T19:35:09.620000",
          "content": "<p><a href=\"https://www.kaggle.com/darien\" target=\"_blank\">@darien</a></p>\n<pre><code>Load A Batch of Images Into Memory (m images)\nPreprocess Batch of Images\nAugment Batch of Images\nInfer and Update Model Using Batch of Images\nClear Memory and Start Again  \nThis Continues Until the Epoch is Complete and Then Usually the Data Is Shuffled (Or It's Done During The Epoch s Number of Examples Ahead)\n</code></pre>\n<p>This \"continuous training\" approach allows to train much more images indeed. <br>\nNevertheless, memory slightly increases over each batch, but this time, with this \"training inside loop of batch images\" I was able to train 60K images.<br>\nThank you for you feedback </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1240821": "Hi, \nI'm trying to train my model on a 20K image sample and the notebook seems to be overwhelmed with so many images. 😳\nSo, my question, how have  you done to train a model with all RGBY images, which means 82K images approximately ?\nBecause, this makes me remember about a story I read from a guy who wanted to start learning Machine Learning and Python in a Pentium4, 4MB Ram, 1Gb Hd, and Windows 98 \nOr , I'm missing something... probably \n",
    "1242273": "Hi there, I think I can help.<br>\nThis could be a few different things. The two most obvious things (in my opinion) that could result in this are listed below... it might be something else though so don't take it as gospel.\n\n---\n\n**1. You are plotting or displaying large amounts of data within the kernel. I have noticed this can cause the notebook to run INCREDIBLY slowly and freeze. i.e. plotting a large number of images or printing a large number of lines (10000+). This can be cumulative over many cells or a single cell with a large output.**\n\n---\n\n**2. You are trying to load too many images into memory at once for training.**\n* You cannot load tons of images into memory as you are limited by the CPU/GPU RAM (depending if you go OOM during training or before that)\n* To overcome this you will need to stream the data somehow. i.e. graph based execution.\n* In Tensorflow you would use something like [**`tf.data`**](https://www.tensorflow.org/api_docs/python/tf/data/Dataset) to set up graph based execution for streaming data.\n* Essentially you build a pipeline that does all the steps but only for a batch (or **`n`** batches) at a time.\n\n***Here's a basic example showing how you only need the memory to support **`m`** images in memory.***\n\n----\n\n>1. Load A Batch of Images Into Memory (**`m`** images)\n2. Preprocess Batch of Images\n3. Augment Batch of Images\n4. Infer and Update Model Using Batch of Images\n5. Clear Memory and Start Again\n6. This Continues Until the Epoch is Complete and Then Usually the Data Is Shuffled (Or It's Done During The Epoch **`s`** Number of Examples Ahead)\n\n---\n\nHope this helps! Alternatively, feel free to share a link to the kernel if you need more detailed troubleshooting."
  }
}