{
  "id": 168313,
  "title": "memory issues with shape 331,331,3 and batch size 8 with EFnet 6",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/168313",
  "author_name": "",
  "post_date": "2020-07-20T05:58:00.047180Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>dataset = tf.data.Dataset.from_tensor_slices((train_filenames, labels))\ndef _parse_function(filename, label):\n    img = tf.io.read_file(filename)\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.image.convert_image_dtype(img, tf.float32)\n    img = tf.image.resize(img, [*size])\n    return img, label\nAUTOTUNE=tf.data.experimental.AUTOTUNE\ndataset = dataset.map(_parse_function,num_parallel_calls=AUTOTUNE)\ndataset = dataset.shuffle(buffer_size=10000,reshuffle_each_iteration=True)\ndataset = dataset.batch(8)\ndataset = dataset.repeat()</p>\n\n<p>137 error (used more than what you had got allocated memory) ?</p>\n\n<p>with batch size 6 it can able to run code with out issue , how do i increase batch size ? </p>",
  "messages": [
    {
      "id": "936283",
      "postDate": "07/20/2020 05:58:00",
      "content": "<p>dataset = tf.data.Dataset.from_tensor_slices((train_filenames, labels))\ndef _parse_function(filename, label):\n    img = tf.io.read_file(filename)\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.image.convert_image_dtype(img, tf.float32)\n    img = tf.image.resize(img, [*size])\n    return img, label\nAUTOTUNE=tf.data.experimental.AUTOTUNE\ndataset = dataset.map(_parse_function,num_parallel_calls=AUTOTUNE)\ndataset = dataset.shuffle(buffer_size=10000,reshuffle_each_iteration=True)\ndataset = dataset.batch(8)\ndataset = dataset.repeat()</p>\n\n<p>137 error (used more than what you had got allocated memory) ?</p>\n\n<p>with batch size 6 it can able to run code with out issue , how do i increase batch size ? </p>",
      "rawMarkdown": "dataset = tf.data.Dataset.from_tensor_slices((train_filenames, labels))\ndef _parse_function(filename, label):\n    img = tf.io.read_file(filename)\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.image.convert_image_dtype(img, tf.float32)\n    img = tf.image.resize(img, [*size])\n    return img, label\nAUTOTUNE=tf.data.experimental.AUTOTUNE\ndataset = dataset.map(_parse_function,num_parallel_calls=AUTOTUNE)\ndataset = dataset.shuffle(buffer_size=10000,reshuffle_each_iteration=True)\ndataset = dataset.batch(8)\ndataset = dataset.repeat()\n\n137 error (used more than what you had got allocated memory) ?\n\nwith batch size 6 it can able to run code with out issue , how do i increase batch size ?",
      "votes": null
    },
    {
      "id": "936292",
      "postDate": "07/20/2020 06:14:13",
      "content": "<p>Assume that your running in a Kaggle kernel and time is not your friend and you see the need for a few more epochs.  </p>\n\n<p>Hope someone has an answer for you - as I been there - done this - a lot of times.</p>\n\n<p>My workaround - run the model with the number of epochs that will fit the time slot.  Save the 5 folds of weights.  Make a private dataset of the fold weight files.  Make a new kernel that loads the weights and train some more.   A real bit of pain and time - but that's the price for using free computer resources.</p>",
      "rawMarkdown": "Assume that your running in a Kaggle kernel and time is not your friend and you see the need for a few more epochs.  \n\nHope someone has an answer for you - as I been there - done this - a lot of times.\n\nMy workaround - run the model with the number of epochs that will fit the time slot.  Save the 5 folds of weights.  Make a private dataset of the fold weight files.  Make a new kernel that loads the weights and train some more.   A real bit of pain and time - but that's the price for using free computer resources.",
      "votes": null
    },
    {
      "id": "936384",
      "postDate": "07/20/2020 07:19:54",
      "content": "<p>you mean to say that, take low size like (224,224,3) with 32 or 64 batch size and train with checkpoint and retrain to get good accuracy . did i understood correctly ? </p>",
      "rawMarkdown": "you mean to say that, take low size like (224,224,3) with 32 or 64 batch size and train with checkpoint and retrain to get good accuracy . did i understood correctly ?",
      "votes": null
    },
    {
      "id": "936408",
      "postDate": "07/20/2020 07:43:35",
      "content": "<p>Not exact.</p>\n\n<p>If I have 512x512 image I might only get 10 to 15 epochs in the 3 hours TPU or the 8 hours GPU time budget.  And I want 512x512 results!</p>\n\n<p>So train for the 10 epochs - save the model weights.  Lets pretend you are at 0.90 at the end of 10 epochs but the train/validation curves both suggest that the model was still learning and more epochs would further improve.</p>\n\n<p>Build a private data set with the weights - right now I am using 5 fold so I save the best weights for each fold.</p>\n\n<p>Now generate a new kernel and load the model with your weights.   Your first epoch will start near the 0.90 and continue. (use different learning curve).</p>\n\n<p>This is real easy to do on my local machines - 4 extra lines of code and I can restart a training model from where I left off.  So for experiments I might only run 10 epochs - if it looks good than I make a single line code change to load my saved weights and I can get another 10 epochs run.  </p>\n\n<p>It more painful on Kaggle since you need to build the data set, etc.  </p>",
      "rawMarkdown": "Not exact.\n\nIf I have 512x512 image I might only get 10 to 15 epochs in the 3 hours TPU or the 8 hours GPU time budget.  And I want 512x512 results!\n\nSo train for the 10 epochs - save the model weights.  Lets pretend you are at 0.90 at the end of 10 epochs but the train/validation curves both suggest that the model was still learning and more epochs would further improve.\n\nBuild a private data set with the weights - right now I am using 5 fold so I save the best weights for each fold.\n\nNow generate a new kernel and load the model with your weights.   Your first epoch will start near the 0.90 and continue. (use different learning curve).\n\nThis is real easy to do on my local machines - 4 extra lines of code and I can restart a training model from where I left off.  So for experiments I might only run 10 epochs - if it looks good than I make a single line code change to load my saved weights and I can get another 10 epochs run.  \n\nIt more painful on Kaggle since you need to build the data set, etc.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 936292,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "07/20/2020 06:14:13",
      "content": "<p>Assume that your running in a Kaggle kernel and time is not your friend and you see the need for a few more epochs.  </p>\n\n<p>Hope someone has an answer for you - as I been there - done this - a lot of times.</p>\n\n<p>My workaround - run the model with the number of epochs that will fit the time slot.  Save the 5 folds of weights.  Make a private dataset of the fold weight files.  Make a new kernel that loads the weights and train some more.   A real bit of pain and time - but that's the price for using free computer resources.</p>",
      "votes": null,
      "replies": [
        {
          "id": 936384,
          "author_name": "kunduruanil",
          "author_url": "",
          "post_date": "07/20/2020 07:19:54",
          "content": "<p>you mean to say that, take low size like (224,224,3) with 32 or 64 batch size and train with checkpoint and retrain to get good accuracy . did i understood correctly ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 936408,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/20/2020 07:43:35",
          "content": "<p>Not exact.</p>\n\n<p>If I have 512x512 image I might only get 10 to 15 epochs in the 3 hours TPU or the 8 hours GPU time budget.  And I want 512x512 results!</p>\n\n<p>So train for the 10 epochs - save the model weights.  Lets pretend you are at 0.90 at the end of 10 epochs but the train/validation curves both suggest that the model was still learning and more epochs would further improve.</p>\n\n<p>Build a private data set with the weights - right now I am using 5 fold so I save the best weights for each fold.</p>\n\n<p>Now generate a new kernel and load the model with your weights.   Your first epoch will start near the 0.90 and continue. (use different learning curve).</p>\n\n<p>This is real easy to do on my local machines - 4 extra lines of code and I can restart a training model from where I left off.  So for experiments I might only run 10 epochs - if it looks good than I make a single line code change to load my saved weights and I can get another 10 epochs run.  </p>\n\n<p>It more painful on Kaggle since you need to build the data set, etc.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "936283": "dataset = tf.data.Dataset.from_tensor_slices((train_filenames, labels))\ndef _parse_function(filename, label):\n    img = tf.io.read_file(filename)\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.image.convert_image_dtype(img, tf.float32)\n    img = tf.image.resize(img, [*size])\n    return img, label\nAUTOTUNE=tf.data.experimental.AUTOTUNE\ndataset = dataset.map(_parse_function,num_parallel_calls=AUTOTUNE)\ndataset = dataset.shuffle(buffer_size=10000,reshuffle_each_iteration=True)\ndataset = dataset.batch(8)\ndataset = dataset.repeat()\n\n137 error (used more than what you had got allocated memory) ?\n\nwith batch size 6 it can able to run code with out issue , how do i increase batch size ?",
    "936292": "Assume that your running in a Kaggle kernel and time is not your friend and you see the need for a few more epochs.  \n\nHope someone has an answer for you - as I been there - done this - a lot of times.\n\nMy workaround - run the model with the number of epochs that will fit the time slot.  Save the 5 folds of weights.  Make a private dataset of the fold weight files.  Make a new kernel that loads the weights and train some more.   A real bit of pain and time - but that's the price for using free computer resources.",
    "936384": "you mean to say that, take low size like (224,224,3) with 32 or 64 batch size and train with checkpoint and retrain to get good accuracy . did i understood correctly ?",
    "936408": "Not exact.\n\nIf I have 512x512 image I might only get 10 to 15 epochs in the 3 hours TPU or the 8 hours GPU time budget.  And I want 512x512 results!\n\nSo train for the 10 epochs - save the model weights.  Lets pretend you are at 0.90 at the end of 10 epochs but the train/validation curves both suggest that the model was still learning and more epochs would further improve.\n\nBuild a private data set with the weights - right now I am using 5 fold so I save the best weights for each fold.\n\nNow generate a new kernel and load the model with your weights.   Your first epoch will start near the 0.90 and continue. (use different learning curve).\n\nThis is real easy to do on my local machines - 4 extra lines of code and I can restart a training model from where I left off.  So for experiments I might only run 10 epochs - if it looks good than I make a single line code change to load my saved weights and I can get another 10 epochs run.  \n\nIt more painful on Kaggle since you need to build the data set, etc."
  },
  "source": "meta"
}