{
  "id": 168278,
  "title": "Kernel restart (RAM) right after training",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/168278",
  "author_name": "",
  "post_date": "2020-07-20T01:05:40.555270500Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello all!</p>\n\n<p>This is not my first rodeo with OOM problems...</p>\n\n<p>Kernel is available <a href=\"https://www.kaggle.com/amneves/mobilenet-inception-v3-transfer-learning\">here</a>.</p>\n\n<p>If I keep everything as it is (batch sizes etc) the model finishes the one epoch and the kernel restarts. I tried dividing the steps per epoch by 2, to debug, but it still resets. My training seems to be blowing up the memory. I'm using a TFRecord input pipeline. </p>\n\n<p>What should we do in these cases? I even tried two networks to see if the problem was the size of the network. </p>",
  "messages": [
    {
      "id": "936092",
      "postDate": "07/20/2020 01:05:40",
      "content": "<p>Hello all!</p>\n\n<p>This is not my first rodeo with OOM problems...</p>\n\n<p>Kernel is available <a href=\"https://www.kaggle.com/amneves/mobilenet-inception-v3-transfer-learning\">here</a>.</p>\n\n<p>If I keep everything as it is (batch sizes etc) the model finishes the one epoch and the kernel restarts. I tried dividing the steps per epoch by 2, to debug, but it still resets. My training seems to be blowing up the memory. I'm using a TFRecord input pipeline. </p>\n\n<p>What should we do in these cases? I even tried two networks to see if the problem was the size of the network. </p>",
      "rawMarkdown": "Hello all!\n\nThis is not my first rodeo with OOM problems...\n\nKernel is available [here](https://www.kaggle.com/amneves/mobilenet-inception-v3-transfer-learning).\n\nIf I keep everything as it is (batch sizes etc) the model finishes the one epoch and the kernel restarts. I tried dividing the steps per epoch by 2, to debug, but it still resets. My training seems to be blowing up the memory. I'm using a TFRecord input pipeline. \n\nWhat should we do in these cases? I even tried two networks to see if the problem was the size of the network.",
      "votes": null
    },
    {
      "id": "936157",
      "postDate": "07/20/2020 03:01:14",
      "content": "<p>I would say try reducing the batch size to a really small number (maybe 4 or 8) and see if it trains properly. In most situations having a lower batch size will fix your problem</p>",
      "rawMarkdown": "I would say try reducing the batch size to a really small number (maybe 4 or 8) and see if it trains properly. In most situations having a lower batch size will fix your problem",
      "votes": null
    },
    {
      "id": "936213",
      "postDate": "07/20/2020 04:37:58",
      "content": "<p>Seems like you are resizing the images in the input pipeline.You shouldn't do that for 2 reasons\n1) Your data loading becomes very slow as it have resize every images and then feed into the network\n2) Maybe when resizing every image concurrently in GPU your GPU VRAM gets blown up before training\nSolution:\n1) Before feeding into the network use already resized images</p>",
      "rawMarkdown": "Seems like you are resizing the images in the input pipeline.You shouldn't do that for 2 reasons\n1) Your data loading becomes very slow as it have resize every images and then feed into the network\n2) Maybe when resizing every image concurrently in GPU your GPU VRAM gets blown up before training\nSolution:\n1) Before feeding into the network use already resized images",
      "votes": null
    },
    {
      "id": "936269",
      "postDate": "07/20/2020 05:42:23",
      "content": "<p>I see the same situation you describe on my local Ubuntu system when I hit borderline memory for tensorflow models with this data set.  </p>\n\n<p>I am resizing in the input pipeline in the same fashion as your code.  (I played with elimination of the pipeline resize and it gave me such a small improvement that I felt it was in the measurement error band rather than real - but that's on a system with a lot more punch than a kaggle kernel provides.)</p>\n\n<p>I play with batch size to get fastest possible execution of the code for a given image size.  OOM errors are of course the clue that I made the batch too big.  </p>\n\n<p>On occasion I can get a batch/image size that works for training but kills the PC after completion of the first epoch.   When I reboot the 1st epoch is done and the tensor board file was created for both train and validation.  </p>\n\n<p>This occurs most often when I am trying large image sizes - like 512 and I have needed to reduce batch size to 4 or less per GPU.    My assumption is that I tweaked the memory usage to the border and something in the very last part of the epoch takes the memory ugly.  Not sure why I don't get an OOM message in these cases, instead it kills the PC - only a machine reboot can bring it back.  I have a big RAM memory of 64GB and another 200GB swap so assuming that the issue is in the GPU memory.  </p>\n\n<p>I also can get just normal kernel death in these situations play with sizes, but the kernel deaths seem to occur almost always before the end of the epoch and no tensor board file for the validation generated.</p>\n\n<p>I run 4 machines that vary based on time I built them and all four have hit the ugly spot.  But the fastest and biggest is the most frequent machine only because that is the one I most often use for large image sizes.</p>\n\n<p>Fixes are (assuming I don't want to change my model)\n1.  Reduce batch size.\n2.  Reduce the image size.</p>\n\n<p>I have a model version that I will likely run tomorrow with large image size and will try MhdSharuk suggestion and not  resize in pipeline IF I get a crash.</p>",
      "rawMarkdown": "I see the same situation you describe on my local Ubuntu system when I hit borderline memory for tensorflow models with this data set.  \n\nI am resizing in the input pipeline in the same fashion as your code.  (I played with elimination of the pipeline resize and it gave me such a small improvement that I felt it was in the measurement error band rather than real - but that's on a system with a lot more punch than a kaggle kernel provides.)\n\nI play with batch size to get fastest possible execution of the code for a given image size.  OOM errors are of course the clue that I made the batch too big.  \n\nOn occasion I can get a batch/image size that works for training but kills the PC after completion of the first epoch.   When I reboot the 1st epoch is done and the tensor board file was created for both train and validation.  \n\nThis occurs most often when I am trying large image sizes - like 512 and I have needed to reduce batch size to 4 or less per GPU.    My assumption is that I tweaked the memory usage to the border and something in the very last part of the epoch takes the memory ugly.  Not sure why I don't get an OOM message in these cases, instead it kills the PC - only a machine reboot can bring it back.  I have a big RAM memory of 64GB and another 200GB swap so assuming that the issue is in the GPU memory.  \n\nI also can get just normal kernel death in these situations play with sizes, but the kernel deaths seem to occur almost always before the end of the epoch and no tensor board file for the validation generated.\n\nI run 4 machines that vary based on time I built them and all four have hit the ugly spot.  But the fastest and biggest is the most frequent machine only because that is the one I most often use for large image sizes.\n\nFixes are (assuming I don't want to change my model)\n1.  Reduce batch size.\n2.  Reduce the image size.\n\nI have a model version that I will likely run tomorrow with large image size and will try MhdSharuk suggestion and not  resize in pipeline IF I get a crash.",
      "votes": null
    },
    {
      "id": "936520",
      "postDate": "07/20/2020 09:38:16",
      "content": "<p>Are there any datasets you recommend with already resized images? I thought this could be the problem too.</p>\n\n<p>I'll probably just download the images and preprocess them locally, I might even just run everything locally!</p>",
      "rawMarkdown": "Are there any datasets you recommend with already resized images? I thought this could be the problem too.\n\nI'll probably just download the images and preprocess them locally, I might even just run everything locally!",
      "votes": null
    },
    {
      "id": "937704",
      "postDate": "07/21/2020 06:10:43",
      "content": "<p>Got to running my model with large image size on my local system.  It crashed and died near the end of epoch 1.</p>\n\n<p>Contrary to my belief that I had enough RAM - 64 in chips and 200 in an ssd swap file - the issue was RAM memory.  Death occurs as the system attempts to put stuff in the swap file after the hardware chips are full.  With a 200GB swap I NEVER have RAM memory crashes - until now.   I assume tf is responsible for the bug that prevents a smooth swap - almost looks like tf assumed the more standard swap of 2GB vs 200gb.</p>\n\n<p>Observation indicated that the RAM overload was occurring during the validation.  Each batch appeared to be adding to the RAM.</p>\n\n<p>The Fix - thank goodness I found one for my system :)\n Here was my validation dataset def (once again Kaggle craps all over any attempt to make this look like code!!!!</p>\n\n<p>`def get_validation_dataset(dataset):</p>\n\n<pre><code>dataset = dataset.batch(BATCH_SIZE)\n\ndataset = dataset.cache()\n\n# prefetch next batch while training (autotune prefetch buffer size)\n\ndataset = dataset.prefetch(AUTO)\n\nreturn dataset`\n</code></pre>\n\n<p>Removal of the dataset = dataset.cache() line is the winner.   My model built upon Alex <a href=\"https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512\">kernel</a> - cache works with small images but appears to fill the bucket on larger.</p>\n\n<p>Let me know if you have cache in your dataset build def.</p>",
      "rawMarkdown": "Got to running my model with large image size on my local system.  It crashed and died near the end of epoch 1.\n\nContrary to my belief that I had enough RAM - 64 in chips and 200 in an ssd swap file - the issue was RAM memory.  Death occurs as the system attempts to put stuff in the swap file after the hardware chips are full.  With a 200GB swap I NEVER have RAM memory crashes - until now.   I assume tf is responsible for the bug that prevents a smooth swap - almost looks like tf assumed the more standard swap of 2GB vs 200gb.\n\nObservation indicated that the RAM overload was occurring during the validation.  Each batch appeared to be adding to the RAM.\n\nThe Fix - thank goodness I found one for my system :)\n Here was my validation dataset def (once again Kaggle craps all over any attempt to make this look like code!!!!\n\n`def get_validation_dataset(dataset):\n\n    dataset = dataset.batch(BATCH_SIZE)\n\n    dataset = dataset.cache()\n\n    # prefetch next batch while training (autotune prefetch buffer size)\n\n    dataset = dataset.prefetch(AUTO)\n\n    return dataset`\n\nRemoval of the dataset = dataset.cache() line is the winner.   My model built upon Alex [kernel](https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512) - cache works with small images but appears to fill the bucket on larger.\n\nLet me know if you have cache in your dataset build def.",
      "votes": null
    },
    {
      "id": "937708",
      "postDate": "07/21/2020 06:12:00",
      "content": "<p>Realized you linked your kernel.  You do have cache line in your validation def.  </p>",
      "rawMarkdown": "Realized you linked your kernel.  You do have cache line in your validation def.",
      "votes": null
    },
    {
      "id": "937725",
      "postDate": "07/21/2020 06:21:19",
      "content": "<p>Most public notebooks are using the resized images <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\" target=\"_blank\">here</a>. You can download either JPEGs or TFRecords.</p>",
      "rawMarkdown": "Most public notebooks are using the resized images [here][1]. You can download either JPEGs or TFRecords.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 936157,
      "author_name": "aryanpandey1109",
      "author_url": "",
      "post_date": "07/20/2020 03:01:14",
      "content": "<p>I would say try reducing the batch size to a really small number (maybe 4 or 8) and see if it trains properly. In most situations having a lower batch size will fix your problem</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 936213,
      "author_name": "msharuk589",
      "author_url": "",
      "post_date": "07/20/2020 04:37:58",
      "content": "<p>Seems like you are resizing the images in the input pipeline.You shouldn't do that for 2 reasons\n1) Your data loading becomes very slow as it have resize every images and then feed into the network\n2) Maybe when resizing every image concurrently in GPU your GPU VRAM gets blown up before training\nSolution:\n1) Before feeding into the network use already resized images</p>",
      "votes": null,
      "replies": [
        {
          "id": 936520,
          "author_name": "amneves",
          "author_url": "",
          "post_date": "07/20/2020 09:38:16",
          "content": "<p>Are there any datasets you recommend with already resized images? I thought this could be the problem too.</p>\n\n<p>I'll probably just download the images and preprocess them locally, I might even just run everything locally!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937725,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/21/2020 06:21:19",
          "content": "<p>Most public notebooks are using the resized images <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\" target=\"_blank\">here</a>. You can download either JPEGs or TFRecords.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 936269,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "07/20/2020 05:42:23",
      "content": "<p>I see the same situation you describe on my local Ubuntu system when I hit borderline memory for tensorflow models with this data set.  </p>\n\n<p>I am resizing in the input pipeline in the same fashion as your code.  (I played with elimination of the pipeline resize and it gave me such a small improvement that I felt it was in the measurement error band rather than real - but that's on a system with a lot more punch than a kaggle kernel provides.)</p>\n\n<p>I play with batch size to get fastest possible execution of the code for a given image size.  OOM errors are of course the clue that I made the batch too big.  </p>\n\n<p>On occasion I can get a batch/image size that works for training but kills the PC after completion of the first epoch.   When I reboot the 1st epoch is done and the tensor board file was created for both train and validation.  </p>\n\n<p>This occurs most often when I am trying large image sizes - like 512 and I have needed to reduce batch size to 4 or less per GPU.    My assumption is that I tweaked the memory usage to the border and something in the very last part of the epoch takes the memory ugly.  Not sure why I don't get an OOM message in these cases, instead it kills the PC - only a machine reboot can bring it back.  I have a big RAM memory of 64GB and another 200GB swap so assuming that the issue is in the GPU memory.  </p>\n\n<p>I also can get just normal kernel death in these situations play with sizes, but the kernel deaths seem to occur almost always before the end of the epoch and no tensor board file for the validation generated.</p>\n\n<p>I run 4 machines that vary based on time I built them and all four have hit the ugly spot.  But the fastest and biggest is the most frequent machine only because that is the one I most often use for large image sizes.</p>\n\n<p>Fixes are (assuming I don't want to change my model)\n1.  Reduce batch size.\n2.  Reduce the image size.</p>\n\n<p>I have a model version that I will likely run tomorrow with large image size and will try MhdSharuk suggestion and not  resize in pipeline IF I get a crash.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937704,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "07/21/2020 06:10:43",
      "content": "<p>Got to running my model with large image size on my local system.  It crashed and died near the end of epoch 1.</p>\n\n<p>Contrary to my belief that I had enough RAM - 64 in chips and 200 in an ssd swap file - the issue was RAM memory.  Death occurs as the system attempts to put stuff in the swap file after the hardware chips are full.  With a 200GB swap I NEVER have RAM memory crashes - until now.   I assume tf is responsible for the bug that prevents a smooth swap - almost looks like tf assumed the more standard swap of 2GB vs 200gb.</p>\n\n<p>Observation indicated that the RAM overload was occurring during the validation.  Each batch appeared to be adding to the RAM.</p>\n\n<p>The Fix - thank goodness I found one for my system :)\n Here was my validation dataset def (once again Kaggle craps all over any attempt to make this look like code!!!!</p>\n\n<p>`def get_validation_dataset(dataset):</p>\n\n<pre><code>dataset = dataset.batch(BATCH_SIZE)\n\ndataset = dataset.cache()\n\n# prefetch next batch while training (autotune prefetch buffer size)\n\ndataset = dataset.prefetch(AUTO)\n\nreturn dataset`\n</code></pre>\n\n<p>Removal of the dataset = dataset.cache() line is the winner.   My model built upon Alex <a href=\"https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512\">kernel</a> - cache works with small images but appears to fill the bucket on larger.</p>\n\n<p>Let me know if you have cache in your dataset build def.</p>",
      "votes": null,
      "replies": [
        {
          "id": 937708,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/21/2020 06:12:00",
          "content": "<p>Realized you linked your kernel.  You do have cache line in your validation def.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "936092": "Hello all!\n\nThis is not my first rodeo with OOM problems...\n\nKernel is available [here](https://www.kaggle.com/amneves/mobilenet-inception-v3-transfer-learning).\n\nIf I keep everything as it is (batch sizes etc) the model finishes the one epoch and the kernel restarts. I tried dividing the steps per epoch by 2, to debug, but it still resets. My training seems to be blowing up the memory. I'm using a TFRecord input pipeline. \n\nWhat should we do in these cases? I even tried two networks to see if the problem was the size of the network.",
    "936157": "I would say try reducing the batch size to a really small number (maybe 4 or 8) and see if it trains properly. In most situations having a lower batch size will fix your problem",
    "936213": "Seems like you are resizing the images in the input pipeline.You shouldn't do that for 2 reasons\n1) Your data loading becomes very slow as it have resize every images and then feed into the network\n2) Maybe when resizing every image concurrently in GPU your GPU VRAM gets blown up before training\nSolution:\n1) Before feeding into the network use already resized images",
    "936269": "I see the same situation you describe on my local Ubuntu system when I hit borderline memory for tensorflow models with this data set.  \n\nI am resizing in the input pipeline in the same fashion as your code.  (I played with elimination of the pipeline resize and it gave me such a small improvement that I felt it was in the measurement error band rather than real - but that's on a system with a lot more punch than a kaggle kernel provides.)\n\nI play with batch size to get fastest possible execution of the code for a given image size.  OOM errors are of course the clue that I made the batch too big.  \n\nOn occasion I can get a batch/image size that works for training but kills the PC after completion of the first epoch.   When I reboot the 1st epoch is done and the tensor board file was created for both train and validation.  \n\nThis occurs most often when I am trying large image sizes - like 512 and I have needed to reduce batch size to 4 or less per GPU.    My assumption is that I tweaked the memory usage to the border and something in the very last part of the epoch takes the memory ugly.  Not sure why I don't get an OOM message in these cases, instead it kills the PC - only a machine reboot can bring it back.  I have a big RAM memory of 64GB and another 200GB swap so assuming that the issue is in the GPU memory.  \n\nI also can get just normal kernel death in these situations play with sizes, but the kernel deaths seem to occur almost always before the end of the epoch and no tensor board file for the validation generated.\n\nI run 4 machines that vary based on time I built them and all four have hit the ugly spot.  But the fastest and biggest is the most frequent machine only because that is the one I most often use for large image sizes.\n\nFixes are (assuming I don't want to change my model)\n1.  Reduce batch size.\n2.  Reduce the image size.\n\nI have a model version that I will likely run tomorrow with large image size and will try MhdSharuk suggestion and not  resize in pipeline IF I get a crash.",
    "936520": "Are there any datasets you recommend with already resized images? I thought this could be the problem too.\n\nI'll probably just download the images and preprocess them locally, I might even just run everything locally!",
    "937704": "Got to running my model with large image size on my local system.  It crashed and died near the end of epoch 1.\n\nContrary to my belief that I had enough RAM - 64 in chips and 200 in an ssd swap file - the issue was RAM memory.  Death occurs as the system attempts to put stuff in the swap file after the hardware chips are full.  With a 200GB swap I NEVER have RAM memory crashes - until now.   I assume tf is responsible for the bug that prevents a smooth swap - almost looks like tf assumed the more standard swap of 2GB vs 200gb.\n\nObservation indicated that the RAM overload was occurring during the validation.  Each batch appeared to be adding to the RAM.\n\nThe Fix - thank goodness I found one for my system :)\n Here was my validation dataset def (once again Kaggle craps all over any attempt to make this look like code!!!!\n\n`def get_validation_dataset(dataset):\n\n    dataset = dataset.batch(BATCH_SIZE)\n\n    dataset = dataset.cache()\n\n    # prefetch next batch while training (autotune prefetch buffer size)\n\n    dataset = dataset.prefetch(AUTO)\n\n    return dataset`\n\nRemoval of the dataset = dataset.cache() line is the winner.   My model built upon Alex [kernel](https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512) - cache works with small images but appears to fill the bucket on larger.\n\nLet me know if you have cache in your dataset build def.",
    "937708": "Realized you linked your kernel.  You do have cache line in your validation def.",
    "937725": "Most public notebooks are using the resized images [here][1]. You can download either JPEGs or TFRecords.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092"
  },
  "source": "meta"
}