{
  "id": 307237,
  "title": "The problem of insufficient GPU memory.",
  "url": "/competitions/happy-whale-and-dolphin/discussion/307237",
  "author_name": "",
  "post_date": "2022-02-13T11:29:20.889145400Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Good evening.</p>\n<p>This is my first time to join kaggle competition.</p>\n<p>I'm using google colab to learn, but I don't have enough memory to learn. (GPU memory is 15GB).<br>\nThe batch size of the data is 18 and the resolution is 128x128.</p>\n<p>Can you tell me what I can do?</p>",
  "messages": [
    {
      "id": "1688090",
      "postDate": "02/13/2022 11:29:20",
      "content": "<p>Good evening.</p>\n<p>This is my first time to join kaggle competition.</p>\n<p>I'm using google colab to learn, but I don't have enough memory to learn. (GPU memory is 15GB).<br>\nThe batch size of the data is 18 and the resolution is 128x128.</p>\n<p>Can you tell me what I can do?</p>",
      "rawMarkdown": "Good evening.\n\nThis is my first time to join kaggle competition.\n\nI'm using google colab to learn, but I don't have enough memory to learn. (GPU memory is 15GB).\nThe batch size of the data is 18 and the resolution is 128x128.\n\nCan you tell me what I can do?",
      "votes": null
    },
    {
      "id": "1688175",
      "postDate": "02/13/2022 12:46:40",
      "content": "<p><a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/307135\" target=\"_blank\">https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/307135</a></p>\n<p>have same problem but in my case it's ram <br>\nI know with low batch size it's taking forever to run </p>",
      "rawMarkdown": "[https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/307135](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/307135)\n\nhave same problem but in my case it's ram \nI know with low batch size it's taking forever to run",
      "votes": null
    },
    {
      "id": "1688943",
      "postDate": "02/13/2022 22:48:56",
      "content": "<p>Hello! Somesh88.<br>\nThank you for answer.</p>",
      "rawMarkdown": "Hello! Somesh88.\nThank you for answer.",
      "votes": null
    },
    {
      "id": "1689193",
      "postDate": "02/14/2022 04:38:37",
      "content": "<p>maybe TPU will work but you have to check how can you spin up TPU </p>",
      "rawMarkdown": "maybe TPU will work but you have to check how can you spin up TPU",
      "votes": null
    },
    {
      "id": "1689718",
      "postDate": "02/14/2022 12:49:24",
      "content": "<p>Hi ! I'm facing this problem too, and the way I currently go is by performing an optimizer step over multiple batches (accumulate gradient). It's not faster but it allows to use virtually any batch size</p>",
      "rawMarkdown": "Hi ! I'm facing this problem too, and the way I currently go is by performing an optimizer step over multiple batches (accumulate gradient). It's not faster but it allows to use virtually any batch size",
      "votes": null
    },
    {
      "id": "1689731",
      "postDate": "02/14/2022 13:06:50",
      "content": "<p>but if we did accumulate gradient I think it'll affect model's overall performance there is always tradeback :( </p>",
      "rawMarkdown": "but if we did accumulate gradient I think it'll affect model's overall performance there is always tradeback :(",
      "votes": null
    },
    {
      "id": "1689735",
      "postDate": "02/14/2022 13:10:53",
      "content": "<p>I'm not sure I understand what you mean. To me accumulating gradient over 2 batches of size 16 for example is mathematically equivalent to a gradient obtained via a batch of 32 samples. <br>\nAnd actually, there is a tradeoff which is that performing one training step would require twice the time</p>",
      "rawMarkdown": "I'm not sure I understand what you mean. To me accumulating gradient over 2 batches of size 16 for example is mathematically equivalent to a gradient obtained via a batch of 32 samples. \nAnd actually, there is a tradeoff which is that performing one training step would require twice the time",
      "votes": null
    },
    {
      "id": "1689744",
      "postDate": "02/14/2022 13:15:42",
      "content": "<p>ohh I think I'll try to find some other approach or try something with more compute resources.</p>",
      "rawMarkdown": "ohh I think I'll try to find some other approach or try something with more compute resources.",
      "votes": null
    },
    {
      "id": "1689989",
      "postDate": "02/14/2022 16:22:51",
      "content": "<p>You can reduce batch size, or use less complex model (like EfficientNet B0 instead of EfficientNet B6).</p>",
      "rawMarkdown": "You can reduce batch size, or use less complex model (like EfficientNet B0 instead of EfficientNet B6).",
      "votes": null
    },
    {
      "id": "1690373",
      "postDate": "02/14/2022 22:49:34",
      "content": "<p>Théo Boyer!!<br>\nThank you.<br>\nI'll give it a try.</p>",
      "rawMarkdown": "Théo Boyer!!\nThank you.\nI'll give it a try.",
      "votes": null
    },
    {
      "id": "1690376",
      "postDate": "02/14/2022 22:52:12",
      "content": "<p>Hi, Araik Tamazian!<br>\nMy current model is a Resnet 18, so I'll give Efficientnet a try.<br>\nThanks.</p>",
      "rawMarkdown": "Hi, Araik Tamazian!\nMy current model is a Resnet 18, so I'll give Efficientnet a try.\nThanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1688175,
      "author_name": "somesh88",
      "author_url": "",
      "post_date": "02/13/2022 12:46:40",
      "content": "<p><a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/307135\" target=\"_blank\">https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/307135</a></p>\n<p>have same problem but in my case it's ram <br>\nI know with low batch size it's taking forever to run </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1688943,
      "author_name": "ishizakireo",
      "author_url": "",
      "post_date": "02/13/2022 22:48:56",
      "content": "<p>Hello! Somesh88.<br>\nThank you for answer.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1689193,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/14/2022 04:38:37",
          "content": "<p>maybe TPU will work but you have to check how can you spin up TPU </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1689718,
      "author_name": "wolfy73",
      "author_url": "",
      "post_date": "02/14/2022 12:49:24",
      "content": "<p>Hi ! I'm facing this problem too, and the way I currently go is by performing an optimizer step over multiple batches (accumulate gradient). It's not faster but it allows to use virtually any batch size</p>",
      "votes": null,
      "replies": [
        {
          "id": 1689731,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/14/2022 13:06:50",
          "content": "<p>but if we did accumulate gradient I think it'll affect model's overall performance there is always tradeback :( </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689735,
          "author_name": "wolfy73",
          "author_url": "",
          "post_date": "02/14/2022 13:10:53",
          "content": "<p>I'm not sure I understand what you mean. To me accumulating gradient over 2 batches of size 16 for example is mathematically equivalent to a gradient obtained via a batch of 32 samples. <br>\nAnd actually, there is a tradeoff which is that performing one training step would require twice the time</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689744,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/14/2022 13:15:42",
          "content": "<p>ohh I think I'll try to find some other approach or try something with more compute resources.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1689989,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "02/14/2022 16:22:51",
      "content": "<p>You can reduce batch size, or use less complex model (like EfficientNet B0 instead of EfficientNet B6).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1690373,
      "author_name": "ishizakireo",
      "author_url": "",
      "post_date": "02/14/2022 22:49:34",
      "content": "<p>Théo Boyer!!<br>\nThank you.<br>\nI'll give it a try.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1690376,
      "author_name": "ishizakireo",
      "author_url": "",
      "post_date": "02/14/2022 22:52:12",
      "content": "<p>Hi, Araik Tamazian!<br>\nMy current model is a Resnet 18, so I'll give Efficientnet a try.<br>\nThanks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1688090": "Good evening.\n\nThis is my first time to join kaggle competition.\n\nI'm using google colab to learn, but I don't have enough memory to learn. (GPU memory is 15GB).\nThe batch size of the data is 18 and the resolution is 128x128.\n\nCan you tell me what I can do?",
    "1688175": "[https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/307135](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/307135)\n\nhave same problem but in my case it's ram \nI know with low batch size it's taking forever to run",
    "1688943": "Hello! Somesh88.\nThank you for answer.",
    "1689193": "maybe TPU will work but you have to check how can you spin up TPU",
    "1689718": "Hi ! I'm facing this problem too, and the way I currently go is by performing an optimizer step over multiple batches (accumulate gradient). It's not faster but it allows to use virtually any batch size",
    "1689731": "but if we did accumulate gradient I think it'll affect model's overall performance there is always tradeback :(",
    "1689735": "I'm not sure I understand what you mean. To me accumulating gradient over 2 batches of size 16 for example is mathematically equivalent to a gradient obtained via a batch of 32 samples. \nAnd actually, there is a tradeoff which is that performing one training step would require twice the time",
    "1689744": "ohh I think I'll try to find some other approach or try something with more compute resources.",
    "1689989": "You can reduce batch size, or use less complex model (like EfficientNet B0 instead of EfficientNet B6).",
    "1690373": "Théo Boyer!!\nThank you.\nI'll give it a try.",
    "1690376": "Hi, Araik Tamazian!\nMy current model is a Resnet 18, so I'll give Efficientnet a try.\nThanks."
  },
  "source": "meta"
}