{
  "id": 107638,
  "title": "ResNet34 implementation of Unet works but ResNet 50 and 101 fails?",
  "url": "/competitions/understanding_cloud_organization/discussion/107638",
  "author_name": "",
  "post_date": "2019-09-05T13:37:45.722569800Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Implementing restNet-34 the Unet model runs fine, but why it fails when implementing resnet-50 and resNet 101? Any idea!</p>",
  "messages": [
    {
      "id": "618776",
      "postDate": "09/05/2019 13:37:45",
      "content": "<p>Implementing restNet-34 the Unet model runs fine, but why it fails when implementing resnet-50 and resNet 101? Any idea!</p>",
      "rawMarkdown": "Implementing restNet-34 the Unet model runs fine, but why it fails when implementing resnet-50 and resNet 101? Any idea!",
      "votes": null
    },
    {
      "id": "618795",
      "postDate": "09/05/2019 13:56:23",
      "content": "<p>I also get an OOM error. I think you should reduce the batch size.</p>",
      "rawMarkdown": "I also get an OOM error. I think you should reduce the batch size.",
      "votes": null
    },
    {
      "id": "621629",
      "postDate": "09/08/2019 18:56:46",
      "content": "<p>because resnet50 and 101 are bigger model and it asks for more computational resource than resnet34</p>",
      "rawMarkdown": "because resnet50 and 101 are bigger model and it asks for more computational resource than resnet34",
      "votes": null
    },
    {
      "id": "673688",
      "postDate": "11/15/2019 10:48:57",
      "content": "<p>Do you mean that I need to reduce the size of the batch size and the size of the epoch to use Lesnet 101? </p>",
      "rawMarkdown": "Do you mean that I need to reduce the size of the batch size and the size of the epoch to use Lesnet 101?",
      "votes": null
    },
    {
      "id": "673784",
      "postDate": "11/15/2019 13:31:15",
      "content": "<p>What does fail mean? Do you mean kernel crash or LB bad?</p>\n\n<p>If kernel crash, it is most likely a GPU memory error. When a network trains, it must save the forward activation of every layer. Layers are 4 dimension tensors like <code>(batch_size, height, width, channels)</code>. The deeper the network (i.e. resNet-50 has 50 layers whereas resNet-34 has 34), the more GPU memory you need even though the <code>batch_size</code> doesn't change. To solve the problem, consider reducing <code>batch_size</code>.</p>",
      "rawMarkdown": "What does fail mean? Do you mean kernel crash or LB bad?\n\nIf kernel crash, it is most likely a GPU memory error. When a network trains, it must save the forward activation of every layer. Layers are 4 dimension tensors like `(batch_size, height, width, channels)`. The deeper the network (i.e. resNet-50 has 50 layers whereas resNet-34 has 34), the more GPU memory you need even though the `batch_size` doesn't change. To solve the problem, consider reducing `batch_size`.",
      "votes": null
    },
    {
      "id": "673936",
      "postDate": "11/15/2019 17:13:31",
      "content": "<p>Kernel crash <a href=\"/cdeotte\">@cdeotte</a> </p>",
      "rawMarkdown": "Kernel crash @cdeotte",
      "votes": null
    },
    {
      "id": "674776",
      "postDate": "11/17/2019 03:56:41",
      "content": "<p>Isn't it because of the image size?</p>",
      "rawMarkdown": "Isn't it because of the image size?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 618795,
      "author_name": "gogo827jz",
      "author_url": "",
      "post_date": "09/05/2019 13:56:23",
      "content": "<p>I also get an OOM error. I think you should reduce the batch size.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 621629,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "09/08/2019 18:56:46",
      "content": "<p>because resnet50 and 101 are bigger model and it asks for more computational resource than resnet34</p>",
      "votes": null,
      "replies": [
        {
          "id": 673688,
          "author_name": "gwangmingo",
          "author_url": "",
          "post_date": "11/15/2019 10:48:57",
          "content": "<p>Do you mean that I need to reduce the size of the batch size and the size of the epoch to use Lesnet 101? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 673784,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/15/2019 13:31:15",
      "content": "<p>What does fail mean? Do you mean kernel crash or LB bad?</p>\n\n<p>If kernel crash, it is most likely a GPU memory error. When a network trains, it must save the forward activation of every layer. Layers are 4 dimension tensors like <code>(batch_size, height, width, channels)</code>. The deeper the network (i.e. resNet-50 has 50 layers whereas resNet-34 has 34), the more GPU memory you need even though the <code>batch_size</code> doesn't change. To solve the problem, consider reducing <code>batch_size</code>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 673936,
          "author_name": "chandraroy",
          "author_url": "",
          "post_date": "11/15/2019 17:13:31",
          "content": "<p>Kernel crash <a href=\"/cdeotte\">@cdeotte</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 674776,
      "author_name": "qkrwlsdn96",
      "author_url": "",
      "post_date": "11/17/2019 03:56:41",
      "content": "<p>Isn't it because of the image size?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "618776": "Implementing restNet-34 the Unet model runs fine, but why it fails when implementing resnet-50 and resNet 101? Any idea!",
    "618795": "I also get an OOM error. I think you should reduce the batch size.",
    "621629": "because resnet50 and 101 are bigger model and it asks for more computational resource than resnet34",
    "673688": "Do you mean that I need to reduce the size of the batch size and the size of the epoch to use Lesnet 101?",
    "673784": "What does fail mean? Do you mean kernel crash or LB bad?\n\nIf kernel crash, it is most likely a GPU memory error. When a network trains, it must save the forward activation of every layer. Layers are 4 dimension tensors like `(batch_size, height, width, channels)`. The deeper the network (i.e. resNet-50 has 50 layers whereas resNet-34 has 34), the more GPU memory you need even though the `batch_size` doesn't change. To solve the problem, consider reducing `batch_size`.",
    "673936": "Kernel crash @cdeotte",
    "674776": "Isn't it because of the image size?"
  },
  "source": "meta"
}