{
  "id": 205416,
  "title": "[!! HELP] Mixed Precision Problem",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/205416",
  "author_name": "",
  "post_date": "2020-12-20T04:57:35.735078900Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm new to pytorch. I was using tensorflow and keras but when I came to know about pytorch's stability and native AMP. I too have two RTX 2080ti and wanna try it. But when I tried it using pytorch's AMP with grad_accum =1 or 2. I'm getting cuda out of memory error. This also happens with Kaggle GPUs. The image resolution is [416 x 416]. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4730925%2F1e03b6611c77ae9f550651aa123ab967%2FScreenshot%20from%202020-12-20%2010-15-10.png?generation=1608440033707911&amp;alt=media\" alt=\"\"></p>\n<p>Am I coding the logic correctly? </p>\n<p>Thanks in Advance</p>",
  "messages": [
    {
      "id": "1119457",
      "postDate": "12/20/2020 04:57:35",
      "content": "<p>I'm new to pytorch. I was using tensorflow and keras but when I came to know about pytorch's stability and native AMP. I too have two RTX 2080ti and wanna try it. But when I tried it using pytorch's AMP with grad_accum =1 or 2. I'm getting cuda out of memory error. This also happens with Kaggle GPUs. The image resolution is [416 x 416]. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4730925%2F1e03b6611c77ae9f550651aa123ab967%2FScreenshot%20from%202020-12-20%2010-15-10.png?generation=1608440033707911&amp;alt=media\" alt=\"\"></p>\n<p>Am I coding the logic correctly? </p>\n<p>Thanks in Advance</p>",
      "rawMarkdown": "I'm new to pytorch. I was using tensorflow and keras but when I came to know about pytorch's stability and native AMP. I too have two RTX 2080ti and wanna try it. But when I tried it using pytorch's AMP with grad_accum =1 or 2. I'm getting cuda out of memory error. This also happens with Kaggle GPUs. The image resolution is [416 x 416]. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4730925%2F1e03b6611c77ae9f550651aa123ab967%2FScreenshot%20from%202020-12-20%2010-15-10.png?generation=1608440033707911&alt=media)\n\n\nAm I coding the logic correctly? \n\nThanks in Advance",
      "votes": null
    },
    {
      "id": "1122213",
      "postDate": "12/22/2020 09:12:20",
      "content": "<p>What is the training/validation set size and the batch size you're incorporating?? if the training set is large then the tensor value is not accumulated and gives the OOM error… try taking a smaller batch size and then see if the error is faced? hope this helps you!!</p>",
      "rawMarkdown": "What is the training/validation set size and the batch size you're incorporating?? if the training set is large then the tensor value is not accumulated and gives the OOM error... try taking a smaller batch size and then see if the error is faced? hope this helps you!!",
      "votes": null
    },
    {
      "id": "1122340",
      "postDate": "12/22/2020 11:11:43",
      "content": "<p>hi,</p>\n<p>Here is some points you might want to re-verify:<br>\n1) 'epoch' is passed in this function but unused in this visible code.<br>\n2) 'train_loader' is passed which is iterator itself but you are using another variable imbar to iterate. Looks like it might be taking entire length of train_loader which can cause memory error. Ideally, batch size is set in the dataloader and this data loader is iterated within epoch.</p>\n<pre><code>train_loader = torch.utils.data.DataLoader(customDataset, batch_size=128)\n...\n</code></pre>\n<p>for i in range(epoch):<br>\n    for batch_idx, (data, target) in enumerate(train_loader):<br>\n              data, target = data.to(device), target.to(device)<br>\n              …<br>\n3)If model is converted to cuda device as well.</p>\n<p>Somehow auto-cast does not seem to be the problem here.Good Luck!</p>",
      "rawMarkdown": "hi,\n\nHere is some points you might want to re-verify:\n1) 'epoch' is passed in this function but unused in this visible code.\n2) 'train_loader' is passed which is iterator itself but you are using another variable imbar to iterate. Looks like it might be taking entire length of train_loader which can cause memory error. Ideally, batch size is set in the dataloader and this data loader is iterated within epoch.\n\n    train_loader = torch.utils.data.DataLoader(customDataset, batch_size=128)\n    ...\nfor i in range(epoch):\n    for batch_idx, (data, target) in enumerate(train_loader):\n              data, target = data.to(device), target.to(device)\n              ...\n3)If model is converted to cuda device as well.\n\nSomehow auto-cast does not seem to be the problem here.Good Luck!",
      "votes": null
    },
    {
      "id": "1122360",
      "postDate": "12/22/2020 11:34:28",
      "content": "<p>Thank you so much, I have rectified the problem and my code is running now. It's nothing but the batch_size, I reduced the batch and increased the accumulation. That's it</p>",
      "rawMarkdown": "Thank you so much, I have rectified the problem and my code is running now. It's nothing but the batch_size, I reduced the batch and increased the accumulation. That's it",
      "votes": null
    },
    {
      "id": "1122361",
      "postDate": "12/22/2020 11:35:02",
      "content": "<p>I set the batch size to less amount and increased the accumulation</p>",
      "rawMarkdown": "I set the batch size to less amount and increased the accumulation",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1122213,
      "author_name": "shivamjohri",
      "author_url": "",
      "post_date": "12/22/2020 09:12:20",
      "content": "<p>What is the training/validation set size and the batch size you're incorporating?? if the training set is large then the tensor value is not accumulated and gives the OOM error… try taking a smaller batch size and then see if the error is faced? hope this helps you!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1122361,
          "author_name": "kingofarmy",
          "author_url": "",
          "post_date": "12/22/2020 11:35:02",
          "content": "<p>I set the batch size to less amount and increased the accumulation</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1122340,
      "author_name": "supplejade",
      "author_url": "",
      "post_date": "12/22/2020 11:11:43",
      "content": "<p>hi,</p>\n<p>Here is some points you might want to re-verify:<br>\n1) 'epoch' is passed in this function but unused in this visible code.<br>\n2) 'train_loader' is passed which is iterator itself but you are using another variable imbar to iterate. Looks like it might be taking entire length of train_loader which can cause memory error. Ideally, batch size is set in the dataloader and this data loader is iterated within epoch.</p>\n<pre><code>train_loader = torch.utils.data.DataLoader(customDataset, batch_size=128)\n...\n</code></pre>\n<p>for i in range(epoch):<br>\n    for batch_idx, (data, target) in enumerate(train_loader):<br>\n              data, target = data.to(device), target.to(device)<br>\n              …<br>\n3)If model is converted to cuda device as well.</p>\n<p>Somehow auto-cast does not seem to be the problem here.Good Luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1122360,
          "author_name": "kingofarmy",
          "author_url": "",
          "post_date": "12/22/2020 11:34:28",
          "content": "<p>Thank you so much, I have rectified the problem and my code is running now. It's nothing but the batch_size, I reduced the batch and increased the accumulation. That's it</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1119457": "I'm new to pytorch. I was using tensorflow and keras but when I came to know about pytorch's stability and native AMP. I too have two RTX 2080ti and wanna try it. But when I tried it using pytorch's AMP with grad_accum =1 or 2. I'm getting cuda out of memory error. This also happens with Kaggle GPUs. The image resolution is [416 x 416]. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4730925%2F1e03b6611c77ae9f550651aa123ab967%2FScreenshot%20from%202020-12-20%2010-15-10.png?generation=1608440033707911&alt=media)\n\n\nAm I coding the logic correctly? \n\nThanks in Advance",
    "1122213": "What is the training/validation set size and the batch size you're incorporating?? if the training set is large then the tensor value is not accumulated and gives the OOM error... try taking a smaller batch size and then see if the error is faced? hope this helps you!!",
    "1122340": "hi,\n\nHere is some points you might want to re-verify:\n1) 'epoch' is passed in this function but unused in this visible code.\n2) 'train_loader' is passed which is iterator itself but you are using another variable imbar to iterate. Looks like it might be taking entire length of train_loader which can cause memory error. Ideally, batch size is set in the dataloader and this data loader is iterated within epoch.\n\n    train_loader = torch.utils.data.DataLoader(customDataset, batch_size=128)\n    ...\nfor i in range(epoch):\n    for batch_idx, (data, target) in enumerate(train_loader):\n              data, target = data.to(device), target.to(device)\n              ...\n3)If model is converted to cuda device as well.\n\nSomehow auto-cast does not seem to be the problem here.Good Luck!",
    "1122360": "Thank you so much, I have rectified the problem and my code is running now. It's nothing but the batch_size, I reduced the batch and increased the accumulation. That's it",
    "1122361": "I set the batch size to less amount and increased the accumulation"
  },
  "source": "meta"
}