{
  "id": 463291,
  "title": "Unable to run custom model in public notebook",
  "url": "/competitions/blood-vessel-segmentation/discussion/463291",
  "author_name": "",
  "post_date": "2023-12-24T12:08:16.393727300Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Im using the 2.5d method public notebook<br>\nI trained a different model than the already given in the notebook<br>\nBut i struggle with the following error:</p>\n<blockquote>\n  <p>RuntimeError: Caught RuntimeError in replica 0 on device 0.<br>\n  Original Traceback (most recent call last):<br>\n    File \"/opt/conda/lib/python3.10/site-packages/torch/nn/parallel/parallel_apply.py\", line 64, in _worker<br>\n      output = module(*input, *<em>kwargs)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(</em>args, *<em>kwargs)\n  File \"/kaggle/input/segmentation-libs/segmentation_models_pytorch/base/model.py\", line 29, in forward\n    features = self.encoder(x)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(</em>args, *<em>kwargs)\n  File \"/kaggle/input/segmentation-libs/segmentation_models_pytorch/encoders/efficientnet.py\", line 66, in forward\n    x = stages<a href=\"x\" target=\"_blank\">i</a>\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(</em>args, *<em>kwargs)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/container.py\", line 217, in forward\n    input = module(input)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(</em>args, **kwargs)<br>\n    File \"/kaggle/input/segmentation-libs/efficientnet_pytorch/utils.py\", line 275, in forward<br>\n      x = F.conv2d(x, self.weight, self.bias, self.stride, self.padding, self.dilation, self.groups)<br>\n  RuntimeError: Input type (torch.cuda.ByteTensor) and weight type (torch.cuda.FloatTensor) should be the same</p>\n</blockquote>",
  "messages": [
    {
      "id": "2572623",
      "postDate": "12/24/2023 12:08:16",
      "content": "<p>Im using the 2.5d method public notebook<br>\nI trained a different model than the already given in the notebook<br>\nBut i struggle with the following error:</p>\n<blockquote>\n  <p>RuntimeError: Caught RuntimeError in replica 0 on device 0.<br>\n  Original Traceback (most recent call last):<br>\n    File \"/opt/conda/lib/python3.10/site-packages/torch/nn/parallel/parallel_apply.py\", line 64, in _worker<br>\n      output = module(*input, *<em>kwargs)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(</em>args, *<em>kwargs)\n  File \"/kaggle/input/segmentation-libs/segmentation_models_pytorch/base/model.py\", line 29, in forward\n    features = self.encoder(x)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(</em>args, *<em>kwargs)\n  File \"/kaggle/input/segmentation-libs/segmentation_models_pytorch/encoders/efficientnet.py\", line 66, in forward\n    x = stages<a href=\"x\" target=\"_blank\">i</a>\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(</em>args, *<em>kwargs)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/container.py\", line 217, in forward\n    input = module(input)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(</em>args, **kwargs)<br>\n    File \"/kaggle/input/segmentation-libs/efficientnet_pytorch/utils.py\", line 275, in forward<br>\n      x = F.conv2d(x, self.weight, self.bias, self.stride, self.padding, self.dilation, self.groups)<br>\n  RuntimeError: Input type (torch.cuda.ByteTensor) and weight type (torch.cuda.FloatTensor) should be the same</p>\n</blockquote>",
      "rawMarkdown": "Im using the 2.5d method public notebook\nI trained a different model than the already given in the notebook\nBut i struggle with the following error:\n\n\n>RuntimeError: Caught RuntimeError in replica 0 on device 0.\nOriginal Traceback (most recent call last):\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/parallel/parallel_apply.py\", line 64, in _worker\n    output = module(*input, **kwargs)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(*args, **kwargs)\n  File \"/kaggle/input/segmentation-libs/segmentation_models_pytorch/base/model.py\", line 29, in forward\n    features = self.encoder(x)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(*args, **kwargs)\n  File \"/kaggle/input/segmentation-libs/segmentation_models_pytorch/encoders/efficientnet.py\", line 66, in forward\n    x = stages[i](x)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(*args, **kwargs)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/container.py\", line 217, in forward\n    input = module(input)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(*args, **kwargs)\n  File \"/kaggle/input/segmentation-libs/efficientnet_pytorch/utils.py\", line 275, in forward\n    x = F.conv2d(x, self.weight, self.bias, self.stride, self.padding, self.dilation, self.groups)\nRuntimeError: Input type (torch.cuda.ByteTensor) and weight type (torch.cuda.FloatTensor) should be the same",
      "votes": null
    },
    {
      "id": "2572714",
      "postDate": "12/24/2023 13:26:41",
      "content": "<p>Before you feed the input to your network you need to normalize it or cast to float </p>",
      "rawMarkdown": "Before you feed the input to your network you need to normalize it or cast to float",
      "votes": null
    },
    {
      "id": "2572793",
      "postDate": "12/24/2023 14:27:41",
      "content": "<p>oh, that's the thing<br>\nnow i realise why was there separate custom model class</p>",
      "rawMarkdown": "oh, that's the thing\nnow i realise why was there separate custom model class",
      "votes": null
    },
    {
      "id": "2572795",
      "postDate": "12/24/2023 14:28:25",
      "content": "<p>But do you happen to understand why does simply converting to float32 makes it go OOM, even if the batch size is 4</p>",
      "rawMarkdown": "But do you happen to understand why does simply converting to float32 makes it go OOM, even if the batch size is 4",
      "votes": null
    },
    {
      "id": "2572806",
      "postDate": "12/24/2023 14:33:30",
      "content": "<p>What is the resolution you are using?</p>",
      "rawMarkdown": "What is the resolution you are using?",
      "votes": null
    },
    {
      "id": "2573067",
      "postDate": "12/24/2023 17:24:51",
      "content": "<p>The memory that needs a float is bigger than an integer. Depending on how much data you are loading at once yes, it can OOM.</p>",
      "rawMarkdown": "The memory that needs a float is bigger than an integer. Depending on how much data you are loading at once yes, it can OOM.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2572714,
      "author_name": "igorkrashenyi",
      "author_url": "",
      "post_date": "12/24/2023 13:26:41",
      "content": "<p>Before you feed the input to your network you need to normalize it or cast to float </p>",
      "votes": null,
      "replies": [
        {
          "id": 2572793,
          "author_name": "bhavesjain",
          "author_url": "",
          "post_date": "12/24/2023 14:27:41",
          "content": "<p>oh, that's the thing<br>\nnow i realise why was there separate custom model class</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2572795,
          "author_name": "bhavesjain",
          "author_url": "",
          "post_date": "12/24/2023 14:28:25",
          "content": "<p>But do you happen to understand why does simply converting to float32 makes it go OOM, even if the batch size is 4</p>",
          "votes": null,
          "replies": [
            {
              "id": 2572806,
              "author_name": "igorkrashenyi",
              "author_url": "",
              "post_date": "12/24/2023 14:33:30",
              "content": "<p>What is the resolution you are using?</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2573067,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "12/24/2023 17:24:51",
              "content": "<p>The memory that needs a float is bigger than an integer. Depending on how much data you are loading at once yes, it can OOM.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2572623": "Im using the 2.5d method public notebook\nI trained a different model than the already given in the notebook\nBut i struggle with the following error:\n\n\n>RuntimeError: Caught RuntimeError in replica 0 on device 0.\nOriginal Traceback (most recent call last):\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/parallel/parallel_apply.py\", line 64, in _worker\n    output = module(*input, **kwargs)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(*args, **kwargs)\n  File \"/kaggle/input/segmentation-libs/segmentation_models_pytorch/base/model.py\", line 29, in forward\n    features = self.encoder(x)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(*args, **kwargs)\n  File \"/kaggle/input/segmentation-libs/segmentation_models_pytorch/encoders/efficientnet.py\", line 66, in forward\n    x = stages[i](x)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(*args, **kwargs)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/container.py\", line 217, in forward\n    input = module(input)\n  File \"/opt/conda/lib/python3.10/site-packages/torch/nn/modules/module.py\", line 1501, in _call_impl\n    return forward_call(*args, **kwargs)\n  File \"/kaggle/input/segmentation-libs/efficientnet_pytorch/utils.py\", line 275, in forward\n    x = F.conv2d(x, self.weight, self.bias, self.stride, self.padding, self.dilation, self.groups)\nRuntimeError: Input type (torch.cuda.ByteTensor) and weight type (torch.cuda.FloatTensor) should be the same",
    "2572714": "Before you feed the input to your network you need to normalize it or cast to float",
    "2572793": "oh, that's the thing\nnow i realise why was there separate custom model class",
    "2572795": "But do you happen to understand why does simply converting to float32 makes it go OOM, even if the batch size is 4",
    "2572806": "What is the resolution you are using?",
    "2573067": "The memory that needs a float is bigger than an integer. Depending on how much data you are loading at once yes, it can OOM."
  },
  "source": "meta"
}