{
  "id": 199198,
  "title": "Memory Constraints - Notebook Crashing when Opening Images",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/199198",
  "author_name": "",
  "post_date": "2020-11-24T20:04:01.946771600Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey all,</p>\n<p>I'm following Marcos N.'s excellent submission notebook example<br>\n(<a href=\"https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk\" target=\"_blank\">https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk</a>)<br>\nand I'm running into notebook memory limits - specifically appears to be when I open the image files.</p>\n<p>Originally I had made slight modifications to the notebook to match my inference pipeline, but nothing that seemed like it would affect this and I've pared back changes until I'm basically just running his notebook with my model - and I still can't execute the whole thing. It always crashes when opening a tiff, and not always the largest ones. I monitor the RAM usage and it hovers around 40% most of the time. It doesn't appear to spike before crashing, but there may be lag in the notebook monitoring system.</p>\n<p>Anyone else dealt with this? I saw some other folks on here using partial image opening by coords (<a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198050)\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198050)</a>, which I'm going to try later today - just surprised (and a little frustrated) that even opening the dataset (one image at a time) doesn't seem to be possible on the provided kernels. </p>",
  "messages": [
    {
      "id": "1089813",
      "postDate": "11/24/2020 20:04:01",
      "content": "<p>Hey all,</p>\n<p>I'm following Marcos N.'s excellent submission notebook example<br>\n(<a href=\"https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk\" target=\"_blank\">https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk</a>)<br>\nand I'm running into notebook memory limits - specifically appears to be when I open the image files.</p>\n<p>Originally I had made slight modifications to the notebook to match my inference pipeline, but nothing that seemed like it would affect this and I've pared back changes until I'm basically just running his notebook with my model - and I still can't execute the whole thing. It always crashes when opening a tiff, and not always the largest ones. I monitor the RAM usage and it hovers around 40% most of the time. It doesn't appear to spike before crashing, but there may be lag in the notebook monitoring system.</p>\n<p>Anyone else dealt with this? I saw some other folks on here using partial image opening by coords (<a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198050)\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198050)</a>, which I'm going to try later today - just surprised (and a little frustrated) that even opening the dataset (one image at a time) doesn't seem to be possible on the provided kernels. </p>",
      "rawMarkdown": "Hey all,\n\nI'm following Marcos N.'s excellent submission notebook example\n(https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk)\nand I'm running into notebook memory limits - specifically appears to be when I open the image files.\n\nOriginally I had made slight modifications to the notebook to match my inference pipeline, but nothing that seemed like it would affect this and I've pared back changes until I'm basically just running his notebook with my model - and I still can't execute the whole thing. It always crashes when opening a tiff, and not always the largest ones. I monitor the RAM usage and it hovers around 40% most of the time. It doesn't appear to spike before crashing, but there may be lag in the notebook monitoring system.\n\nAnyone else dealt with this? I saw some other folks on here using partial image opening by coords (https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198050), which I'm going to try later today - just surprised (and a little frustrated) that even opening the dataset (one image at a time) doesn't seem to be possible on the provided kernels.",
      "votes": null
    },
    {
      "id": "1090008",
      "postDate": "11/25/2020 01:27:33",
      "content": "<p><a href=\"https://www.kaggle.com/paulgamble\" target=\"_blank\">@paulgamble</a>,</p>\n<p>It is verified that tifffile can read all images in all datasets (train, test, private). Have a look at this notebook where I verify this:<br>\n<a href=\"https://www.kaggle.com/marcosnovaes/hubmap-read-data-and-build-tfrecords\" target=\"_blank\">https://www.kaggle.com/marcosnovaes/hubmap-read-data-and-build-tfrecords</a></p>\n<p>If you are reading them in a loop, I noticed I have to call garbage collect (gc.collect()) to make sure memory is freed before reading the next one. As explained in that notebook, some things to check:</p>\n<ul>\n<li>the order of shapes is different on a few images (channel first or last) </li>\n<li>use tifffile, some images are BIG tiff</li>\n<li>In my submission notebook, I had a plt.imshow statement just to show the last image. The plot.imgshow() function is a memory hog and it is leaky. It crashes the kernel consistently. To use imshow you do have to slice a section of the image, and it will still leak a lot. </li>\n<li>try running that notebook, as explained, the following should succeed:</li>\n</ul>\n<p>def verify_read(file_list):<br>\n    for file_name in file_list:<br>\n        baseimage = tifffile.imread(file_name)<br>\n        #baseimage = tif.series[0].asarray()<br>\n        print('img id = {}, shape = {}'.format(file_name,baseimage.shape))<br>\n        gc.collect()</p>\n<p>file_list = glob.glob('/kaggle/input/hubmap-kidney-segmentation/train/*.tiff')<br>\nverify_read(file_list)</p>",
      "rawMarkdown": "paulgamble,\n\nIt is verified that tifffile can read all images in all datasets (train, test, private). Have a look at this notebook where I verify this:\n[https://www.kaggle.com/marcosnovaes/hubmap-read-data-and-build-tfrecords](https://www.kaggle.com/marcosnovaes/hubmap-read-data-and-build-tfrecords)\n\nIf you are reading them in a loop, I noticed I have to call garbage collect (gc.collect()) to make sure memory is freed before reading the next one. As explained in that notebook, some things to check:\n- the order of shapes is different on a few images (channel first or last) \n- use tifffile, some images are BIG tiff\n- In my submission notebook, I had a plt.imshow statement just to show the last image. The plot.imgshow() function is a memory hog and it is leaky. It crashes the kernel consistently. To use imshow you do have to slice a section of the image, and it will still leak a lot. \n- try running that notebook, as explained, the following should succeed:\n\ndef verify_read(file_list):\n    for file_name in file_list:\n        baseimage = tifffile.imread(file_name)\n        #baseimage = tif.series[0].asarray()\n        print('img id = {}, shape = {}'.format(file_name,baseimage.shape))\n        gc.collect()\n        \nfile_list = glob.glob('/kaggle/input/hubmap-kidney-segmentation/train/*.tiff')\nverify_read(file_list)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1090008,
      "author_name": "marcosnovaes",
      "author_url": "",
      "post_date": "11/25/2020 01:27:33",
      "content": "<p><a href=\"https://www.kaggle.com/paulgamble\" target=\"_blank\">@paulgamble</a>,</p>\n<p>It is verified that tifffile can read all images in all datasets (train, test, private). Have a look at this notebook where I verify this:<br>\n<a href=\"https://www.kaggle.com/marcosnovaes/hubmap-read-data-and-build-tfrecords\" target=\"_blank\">https://www.kaggle.com/marcosnovaes/hubmap-read-data-and-build-tfrecords</a></p>\n<p>If you are reading them in a loop, I noticed I have to call garbage collect (gc.collect()) to make sure memory is freed before reading the next one. As explained in that notebook, some things to check:</p>\n<ul>\n<li>the order of shapes is different on a few images (channel first or last) </li>\n<li>use tifffile, some images are BIG tiff</li>\n<li>In my submission notebook, I had a plt.imshow statement just to show the last image. The plot.imgshow() function is a memory hog and it is leaky. It crashes the kernel consistently. To use imshow you do have to slice a section of the image, and it will still leak a lot. </li>\n<li>try running that notebook, as explained, the following should succeed:</li>\n</ul>\n<p>def verify_read(file_list):<br>\n    for file_name in file_list:<br>\n        baseimage = tifffile.imread(file_name)<br>\n        #baseimage = tif.series[0].asarray()<br>\n        print('img id = {}, shape = {}'.format(file_name,baseimage.shape))<br>\n        gc.collect()</p>\n<p>file_list = glob.glob('/kaggle/input/hubmap-kidney-segmentation/train/*.tiff')<br>\nverify_read(file_list)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1089813": "Hey all,\n\nI'm following Marcos N.'s excellent submission notebook example\n(https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk)\nand I'm running into notebook memory limits - specifically appears to be when I open the image files.\n\nOriginally I had made slight modifications to the notebook to match my inference pipeline, but nothing that seemed like it would affect this and I've pared back changes until I'm basically just running his notebook with my model - and I still can't execute the whole thing. It always crashes when opening a tiff, and not always the largest ones. I monitor the RAM usage and it hovers around 40% most of the time. It doesn't appear to spike before crashing, but there may be lag in the notebook monitoring system.\n\nAnyone else dealt with this? I saw some other folks on here using partial image opening by coords (https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198050), which I'm going to try later today - just surprised (and a little frustrated) that even opening the dataset (one image at a time) doesn't seem to be possible on the provided kernels.",
    "1090008": "paulgamble,\n\nIt is verified that tifffile can read all images in all datasets (train, test, private). Have a look at this notebook where I verify this:\n[https://www.kaggle.com/marcosnovaes/hubmap-read-data-and-build-tfrecords](https://www.kaggle.com/marcosnovaes/hubmap-read-data-and-build-tfrecords)\n\nIf you are reading them in a loop, I noticed I have to call garbage collect (gc.collect()) to make sure memory is freed before reading the next one. As explained in that notebook, some things to check:\n- the order of shapes is different on a few images (channel first or last) \n- use tifffile, some images are BIG tiff\n- In my submission notebook, I had a plt.imshow statement just to show the last image. The plot.imgshow() function is a memory hog and it is leaky. It crashes the kernel consistently. To use imshow you do have to slice a section of the image, and it will still leak a lot. \n- try running that notebook, as explained, the following should succeed:\n\ndef verify_read(file_list):\n    for file_name in file_list:\n        baseimage = tifffile.imread(file_name)\n        #baseimage = tif.series[0].asarray()\n        print('img id = {}, shape = {}'.format(file_name,baseimage.shape))\n        gc.collect()\n        \nfile_list = glob.glob('/kaggle/input/hubmap-kidney-segmentation/train/*.tiff')\nverify_read(file_list)"
  },
  "source": "meta"
}