{
  "id": 122469,
  "title": "How to reduce amount of GPU needed?",
  "url": "/competitions/pku-autonomous-driving/discussion/122469",
  "author_name": "",
  "post_date": "2019-12-20T11:29:06.588178700Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I started using CenterNet baseline based on <a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">this Kernel</a>. I'm not familiar with PyTorch in which you need to explicitly move stuff to GPU device and then eventually detach them. In result, I ran out of memory on my local machine (4GB total). What are the ways to reduce GPU needed to train that model?</p>\n\n<p>Can you spot any places in hocop1's kernel where I could save some bytes?</p>\n\n<p>Or maybe this architecture has no chance to fit into 4GB and I should try another model?</p>",
  "messages": [
    {
      "id": "699387",
      "postDate": "12/20/2019 11:29:06",
      "content": "<p>I started using CenterNet baseline based on <a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">this Kernel</a>. I'm not familiar with PyTorch in which you need to explicitly move stuff to GPU device and then eventually detach them. In result, I ran out of memory on my local machine (4GB total). What are the ways to reduce GPU needed to train that model?</p>\n\n<p>Can you spot any places in hocop1's kernel where I could save some bytes?</p>\n\n<p>Or maybe this architecture has no chance to fit into 4GB and I should try another model?</p>",
      "rawMarkdown": "I started using CenterNet baseline based on [this Kernel](https://www.kaggle.com/hocop1/centernet-baseline). I'm not familiar with PyTorch in which you need to explicitly move stuff to GPU device and then eventually detach them. In result, I ran out of memory on my local machine (4GB total). What are the ways to reduce GPU needed to train that model?\n\nCan you spot any places in hocop1's kernel where I could save some bytes?\n\nOr maybe this architecture has no chance to fit into 4GB and I should try another model?",
      "votes": null
    },
    {
      "id": "699427",
      "postDate": "12/20/2019 12:52:20",
      "content": "<p>First of all, you can reduce the batch_size to 2.\nThen, you can reduce the input image size </p>\n\n<p>But I doubt, that you have a chance at this comp with only 4 GB of GPU RAM. You may want to look at Google Cloud Computing, google colab or AWS in addition to kernels.</p>",
      "rawMarkdown": "First of all, you can reduce the batch_size to 2.\nThen, you can reduce the input image size \n\nBut I doubt, that you have a chance at this comp with only 4 GB of GPU RAM. You may want to look at Google Cloud Computing, google colab or AWS in addition to kernels.",
      "votes": null
    },
    {
      "id": "699484",
      "postDate": "12/20/2019 14:07:02",
      "content": "<blockquote>\n  <p>But I doubt, that you have a chance at this comp with only 4 GB of GPU RAM</p>\n</blockquote>\n\n<p>Sure, I am aware of this. I'm just trying to learn new stuff. LB position doesn't matter that much now.</p>",
      "rawMarkdown": "&gt; But I doubt, that you have a chance at this comp with only 4 GB of GPU RAM\n\nSure, I am aware of this. I'm just trying to learn new stuff. LB position doesn't matter that much now.",
      "votes": null
    },
    {
      "id": "699486",
      "postDate": "12/20/2019 14:07:26",
      "content": "<p>I'll try reducing batch and image sizes, thanks!</p>",
      "rawMarkdown": "I'll try reducing batch and image sizes, thanks!",
      "votes": null
    },
    {
      "id": "699602",
      "postDate": "12/20/2019 16:18:07",
      "content": "<p>Don't forget an important point though: If your batch size is small, change your normalization from BatchNorm to Instance or GroupNorm!</p>",
      "rawMarkdown": "Don't forget an important point though: If your batch size is small, change your normalization from BatchNorm to Instance or GroupNorm!",
      "votes": null
    },
    {
      "id": "699844",
      "postDate": "12/21/2019 02:47:26",
      "content": "<p>Intersting, but how do u define \"small\"?</p>",
      "rawMarkdown": "Intersting, but how do u define \"small\"?",
      "votes": null
    },
    {
      "id": "700127",
      "postDate": "12/21/2019 13:52:06",
      "content": "<p>Anything less than 8, really. 8 works, but is not the most optimal.</p>",
      "rawMarkdown": "Anything less than 8, really. 8 works, but is not the most optimal.",
      "votes": null
    },
    {
      "id": "700622",
      "postDate": "12/22/2019 10:18:02",
      "content": "<p><a href=\"/mtszkw\">@mtszkw</a> you can always try <a href=\"https://github.com/NVIDIA/apex#quick-start\">apex</a> and train on <a href=\"https://nvlabs.github.io/iccv2019-mixed-precision-tutorial/files/dusan_stosic_intro_to_mixed_precision_training.pdf\">mixed-precision</a></p>",
      "rawMarkdown": "mtszkw you can always try [apex](https://github.com/NVIDIA/apex#quick-start) and train on [mixed-precision](https://nvlabs.github.io/iccv2019-mixed-precision-tutorial/files/dusan_stosic_intro_to_mixed_precision_training.pdf)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 699427,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "12/20/2019 12:52:20",
      "content": "<p>First of all, you can reduce the batch_size to 2.\nThen, you can reduce the input image size </p>\n\n<p>But I doubt, that you have a chance at this comp with only 4 GB of GPU RAM. You may want to look at Google Cloud Computing, google colab or AWS in addition to kernels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 699484,
          "author_name": "mtszkw",
          "author_url": "",
          "post_date": "12/20/2019 14:07:02",
          "content": "<blockquote>\n  <p>But I doubt, that you have a chance at this comp with only 4 GB of GPU RAM</p>\n</blockquote>\n\n<p>Sure, I am aware of this. I'm just trying to learn new stuff. LB position doesn't matter that much now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699486,
          "author_name": "mtszkw",
          "author_url": "",
          "post_date": "12/20/2019 14:07:26",
          "content": "<p>I'll try reducing batch and image sizes, thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 699602,
      "author_name": "chroteus",
      "author_url": "",
      "post_date": "12/20/2019 16:18:07",
      "content": "<p>Don't forget an important point though: If your batch size is small, change your normalization from BatchNorm to Instance or GroupNorm!</p>",
      "votes": null,
      "replies": [
        {
          "id": 699844,
          "author_name": "szuzhangzhi",
          "author_url": "",
          "post_date": "12/21/2019 02:47:26",
          "content": "<p>Intersting, but how do u define \"small\"?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 700127,
          "author_name": "chroteus",
          "author_url": "",
          "post_date": "12/21/2019 13:52:06",
          "content": "<p>Anything less than 8, really. 8 works, but is not the most optimal.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 700622,
      "author_name": "neongen",
      "author_url": "",
      "post_date": "12/22/2019 10:18:02",
      "content": "<p><a href=\"/mtszkw\">@mtszkw</a> you can always try <a href=\"https://github.com/NVIDIA/apex#quick-start\">apex</a> and train on <a href=\"https://nvlabs.github.io/iccv2019-mixed-precision-tutorial/files/dusan_stosic_intro_to_mixed_precision_training.pdf\">mixed-precision</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "699387": "I started using CenterNet baseline based on [this Kernel](https://www.kaggle.com/hocop1/centernet-baseline). I'm not familiar with PyTorch in which you need to explicitly move stuff to GPU device and then eventually detach them. In result, I ran out of memory on my local machine (4GB total). What are the ways to reduce GPU needed to train that model?\n\nCan you spot any places in hocop1's kernel where I could save some bytes?\n\nOr maybe this architecture has no chance to fit into 4GB and I should try another model?",
    "699427": "First of all, you can reduce the batch_size to 2.\nThen, you can reduce the input image size \n\nBut I doubt, that you have a chance at this comp with only 4 GB of GPU RAM. You may want to look at Google Cloud Computing, google colab or AWS in addition to kernels.",
    "699484": "&gt; But I doubt, that you have a chance at this comp with only 4 GB of GPU RAM\n\nSure, I am aware of this. I'm just trying to learn new stuff. LB position doesn't matter that much now.",
    "699486": "I'll try reducing batch and image sizes, thanks!",
    "699602": "Don't forget an important point though: If your batch size is small, change your normalization from BatchNorm to Instance or GroupNorm!",
    "699844": "Intersting, but how do u define \"small\"?",
    "700127": "Anything less than 8, really. 8 works, but is not the most optimal.",
    "700622": "mtszkw you can always try [apex](https://github.com/NVIDIA/apex#quick-start) and train on [mixed-precision](https://nvlabs.github.io/iccv2019-mixed-precision-tutorial/files/dusan_stosic_intro_to_mixed_precision_training.pdf)"
  },
  "source": "meta"
}