{
  "id": 172132,
  "title": "Why this competition is so difficult to train?",
  "url": "/competitions/landmark-recognition-2020/discussion/172132",
  "author_name": "Dracarys",
  "post_date": "2020-08-03T20:11:48.484000",
  "votes": 3,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I have an RTX2080 GPU(8GB), even with a smaller image size of 256x256, my efficientnetB0 (PyTorch with amp) model is taking around 4hrs per epoch to train. I guess with this speed it is too difficult to experiment. Does anyone have some better strategy to reduce training time. Any suggestion is welcome.</p>",
  "messages": [
    {
      "id": 956822,
      "postDate": "2020-08-03T20:11:48.483Z",
      "content": "<p>I have an RTX2080 GPU(8GB), even with a smaller image size of 256x256, my efficientnetB0 (PyTorch with amp) model is taking around 4hrs per epoch to train. I guess with this speed it is too difficult to experiment. Does anyone have some better strategy to reduce training time. Any suggestion is welcome.</p>",
      "rawMarkdown": "I have an RTX2080 GPU(8GB), even with a smaller image size of 256x256, my efficientnetB0 (PyTorch with amp) model is taking around 4hrs per epoch to train. I guess with this speed it is too difficult to experiment. Does anyone have some better strategy to reduce training time. Any suggestion is welcome.",
      "votes": 4
    },
    {
      "id": 956863,
      "postDate": "2020-08-03T21:01:13.353Z",
      "content": "<p>Maybe you could try to use TPUs from colab/kaggle, they provide serious speedups especially over a single GPU.\nsee: <a href=\"https://arxiv.org/pdf/1907.10701.pdf\">https://arxiv.org/pdf/1907.10701.pdf</a></p>",
      "rawMarkdown": "Maybe you could try to use TPUs from colab/kaggle, they provide serious speedups especially over a single GPU.\nsee: https://arxiv.org/pdf/1907.10701.pdf",
      "votes": 1,
      "replies": [
        {
          "id": 957168,
          "postDate": "2020-08-04T05:39:42.890Z",
          "content": "<p>That's right <a href=\"/riblidezso\">@riblidezso</a> , TPUs are faster for processing. But we have a constraint of 3hrs of usage time.\nI tried with TPUs and unable to complete and create output within 3hrs even with lower epochs and few optimizations. \nThus I shifted towards GPU.</p>",
          "rawMarkdown": "That's right @riblidezso , TPUs are faster for processing. But we have a constraint of 3hrs of usage time.\nI tried with TPUs and unable to complete and create output within 3hrs even with lower epochs and few optimizations. \nThus I shifted towards GPU.",
          "votes": 3,
          "replies": [
            {
              "id": 957227,
              "postDate": "2020-08-04T06:51:06.930Z",
              "content": "<p>I see. Colab has 12 hours runtime limit, and colab pro has 24 hours. Maybe that could be enough.</p>",
              "rawMarkdown": "I see. Colab has 12 hours runtime limit, and colab pro has 24 hours. Maybe that could be enough."
            }
          ]
        },
        {
          "id": 957237,
          "postDate": "2020-08-04T07:03:50.927Z",
          "content": "<p><a href=\"/riblidezso\">@riblidezso</a> Disk space of colab TPU config is around 50GB and size of our dataset is around 48GB(256x256). I don't think we can even train for smaller image size like 256 on colab.</p>",
          "rawMarkdown": "@riblidezso Disk space of colab TPU config is around 50GB and size of our dataset is around 48GB(256x256). I don't think we can even train for smaller image size like 256 on colab.",
          "votes": 1
        },
        {
          "id": 957269,
          "postDate": "2020-08-04T07:28:19.467Z",
          "content": "<p>That's right <a href=\"/rohitsingh9990\">@rohitsingh9990</a> . Additionally Colab Pro is available only in the US.\n<a href=\"https://colab.research.google.com/signup\">https://colab.research.google.com/signup</a></p>",
          "rawMarkdown": "That's right @rohitsingh9990 . Additionally Colab Pro is available only in the US.\nhttps://colab.research.google.com/signup",
          "votes": 1
        },
        {
          "id": 958221,
          "postDate": "2020-08-04T21:12:14.483Z",
          "content": "<p>FWIW, I think the „availability in the US” comes down to giving a zip code at registration - and I am not aware of any verification mechanism for what you input...</p>",
          "rawMarkdown": "FWIW, I think the „availability in the US” comes down to giving a zip code at registration - and I am not aware of any verification mechanism for what you input...",
          "votes": 2
        }
      ]
    },
    {
      "id": 957277,
      "postDate": "2020-08-04T07:34:21.043Z",
      "content": "<p>It shows me 116gb free, when I open a new high ram TPU runtime.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F468424%2Fb4bb0b0c7834b0dee22cac2338128266%2FScreenshot%202020-08-04%20at%209.32.32.png?generation=1596526383427602&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "It shows me 116gb free, when I open a new high ram TPU runtime.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F468424%2Fb4bb0b0c7834b0dee22cac2338128266%2FScreenshot%202020-08-04%20at%209.32.32.png?generation=1596526383427602&amp;alt=media)\n",
      "votes": 2,
      "replies": [
        {
          "id": 957292,
          "postDate": "2020-08-04T07:48:59.617Z",
          "content": "<p>Hi <a href=\"/riblidezso\">@riblidezso</a> \nAre you using Colab Pro?  Because Colab Pro has Faster GPUs, longer runtime as well as memory.\nUnfortunately Colab Pro is still not available in India.\nrefer: <a href=\"https://colab.research.google.com/signup\">https://colab.research.google.com/signup</a></p>",
          "rawMarkdown": "Hi @riblidezso \nAre you using Colab Pro?  Because Colab Pro has Faster GPUs, longer runtime as well as memory.\nUnfortunately Colab Pro is still not available in India.\nrefer: https://colab.research.google.com/signup",
          "replies": [
            {
              "id": 957314,
              "postDate": "2020-08-04T08:03:18.363Z",
              "content": "<p>Yes</p>",
              "rawMarkdown": "Yes"
            }
          ]
        },
        {
          "id": 957318,
          "postDate": "2020-08-04T08:08:18.090Z",
          "content": "<p><a href=\"/riblidezso\">@riblidezso</a> you are right, last time i checked it was around 50GB, currently it is showing <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3982638%2Fe5ef623768f5c693e98066d721bea8c1%2FScreenshot%202020-08-04%20at%201.35.13%20PM.png?generation=1596528424542679&amp;alt=media\" alt=\"\"></p>\n\n<p>so, i think we can train small models with small image size on colab.</p>",
          "rawMarkdown": "@riblidezso you are right, last time i checked it was around 50GB, currently it is showing ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3982638%2Fe5ef623768f5c693e98066d721bea8c1%2FScreenshot%202020-08-04%20at%201.35.13%20PM.png?generation=1596528424542679&amp;alt=media)\n\nso, i think we can train small models with small image size on colab.\n",
          "votes": 1
        },
        {
          "id": 985999,
          "postDate": "2020-08-26T06:29:14.757Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/riblidezso\" target=\"_blank\">@riblidezso</a> what's your model?  Delf?</p>",
          "rawMarkdown": "Hi @riblidezso what's your model?  Delf?"
        },
        {
          "id": 986000,
          "postDate": "2020-08-26T06:29:15.283Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 957698,
      "postDate": "2020-08-04T13:52:15.867Z",
      "content": "<p>What is your batch size?</p>",
      "rawMarkdown": "What is your batch size?",
      "replies": [
        {
          "id": 957723,
          "postDate": "2020-08-04T14:10:19.840Z",
          "content": "<p>batch_size=24 and num_workers=4</p>",
          "rawMarkdown": "batch_size=24 and num_workers=4\n",
          "votes": 1
        },
        {
          "id": 957752,
          "postDate": "2020-08-04T14:38:57.807Z",
          "content": "<p>keep the batch size as default(32), try to set numworkers=0 then everything will work fine, I also face similar issues, I think there might a few bus in pytorch when we set numworkers to 4 or 6 it not working fine..., it would run faster if you won't get out of memory exception.</p>",
          "rawMarkdown": "keep the batch size as default(32), try to set numworkers=0 then everything will work fine, I also face similar issues, I think there might a few bus in pytorch when we set numworkers to 4 or 6 it not working fine..., it would run faster if you won't get out of memory exception."
        },
        {
          "id": 957760,
          "postDate": "2020-08-04T14:47:45.023Z",
          "content": "<p><a href=\"/kiranbeethoju\">@kiranbeethoju</a> you are saying, you are able to train an epoch with around 12-14 lakh images with image size 256 in less than 3hr. Also i am not getting any out of memory issues, the whole trouble is there are 15lakh images, you won't be able to get reduce train time much until and unless you have access to cluster of GPU's.</p>",
          "rawMarkdown": "@kiranbeethoju you are saying, you are able to train an epoch with around 12-14 lakh images with image size 256 in less than 3hr. Also i am not getting any out of memory issues, the whole trouble is there are 15lakh images, you won't be able to get reduce train time much until and unless you have access to cluster of GPU's."
        },
        {
          "id": 957769,
          "postDate": "2020-08-04T14:51:10.440Z",
          "content": "<p>I think you totally misread this discussion.</p>",
          "rawMarkdown": "I think you totally misread this discussion.",
          "votes": 2
        }
      ]
    },
    {
      "id": 957412,
      "postDate": "2020-08-04T09:29:56.883Z",
      "content": "<p>Use Google colab</p>",
      "rawMarkdown": "Use Google colab"
    },
    {
      "id": 991108,
      "postDate": "2020-08-30T06:35:07.670Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 957786,
      "postDate": "2020-08-04T15:01:50.573Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 956863,
      "author_name": "Dezső Ribli",
      "author_url": "",
      "post_date": "2020-08-03T21:01:13.353000",
      "content": "<p>Maybe you could try to use TPUs from colab/kaggle, they provide serious speedups especially over a single GPU.\nsee: <a href=\"https://arxiv.org/pdf/1907.10701.pdf\">https://arxiv.org/pdf/1907.10701.pdf</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 957168,
          "author_name": "Jagadish Sivakumar",
          "author_url": "",
          "post_date": "2020-08-04T05:39:42.890000",
          "content": "<p>That's right <a href=\"/riblidezso\">@riblidezso</a> , TPUs are faster for processing. But we have a constraint of 3hrs of usage time.\nI tried with TPUs and unable to complete and create output within 3hrs even with lower epochs and few optimizations. \nThus I shifted towards GPU.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 957227,
              "author_name": "Dezső Ribli",
              "author_url": "",
              "post_date": "2020-08-04T06:51:06.930000",
              "content": "<p>I see. Colab has 12 hours runtime limit, and colab pro has 24 hours. Maybe that could be enough.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 957237,
          "author_name": "Dracarys",
          "author_url": "",
          "post_date": "2020-08-04T07:03:50.927000",
          "content": "<p><a href=\"/riblidezso\">@riblidezso</a> Disk space of colab TPU config is around 50GB and size of our dataset is around 48GB(256x256). I don't think we can even train for smaller image size like 256 on colab.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 957269,
          "author_name": "Jagadish Sivakumar",
          "author_url": "",
          "post_date": "2020-08-04T07:28:19.467000",
          "content": "<p>That's right <a href=\"/rohitsingh9990\">@rohitsingh9990</a> . Additionally Colab Pro is available only in the US.\n<a href=\"https://colab.research.google.com/signup\">https://colab.research.google.com/signup</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 958221,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2020-08-04T21:12:14.483000",
          "content": "<p>FWIW, I think the „availability in the US” comes down to giving a zip code at registration - and I am not aware of any verification mechanism for what you input...</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 957277,
      "author_name": "Dezső Ribli",
      "author_url": "",
      "post_date": "2020-08-04T07:34:21.043000",
      "content": "<p>It shows me 116gb free, when I open a new high ram TPU runtime.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F468424%2Fb4bb0b0c7834b0dee22cac2338128266%2FScreenshot%202020-08-04%20at%209.32.32.png?generation=1596526383427602&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 957292,
          "author_name": "Jagadish Sivakumar",
          "author_url": "",
          "post_date": "2020-08-04T07:48:59.617000",
          "content": "<p>Hi <a href=\"/riblidezso\">@riblidezso</a> \nAre you using Colab Pro?  Because Colab Pro has Faster GPUs, longer runtime as well as memory.\nUnfortunately Colab Pro is still not available in India.\nrefer: <a href=\"https://colab.research.google.com/signup\">https://colab.research.google.com/signup</a></p>",
          "votes": 0,
          "replies": [
            {
              "id": 957314,
              "author_name": "Dezső Ribli",
              "author_url": "",
              "post_date": "2020-08-04T08:03:18.363000",
              "content": "<p>Yes</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 957318,
          "author_name": "Dracarys",
          "author_url": "",
          "post_date": "2020-08-04T08:08:18.090000",
          "content": "<p><a href=\"/riblidezso\">@riblidezso</a> you are right, last time i checked it was around 50GB, currently it is showing <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3982638%2Fe5ef623768f5c693e98066d721bea8c1%2FScreenshot%202020-08-04%20at%201.35.13%20PM.png?generation=1596528424542679&amp;alt=media\" alt=\"\"></p>\n\n<p>so, i think we can train small models with small image size on colab.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 985999,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2020-08-26T06:29:14.757000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/riblidezso\" target=\"_blank\">@riblidezso</a> what's your model?  Delf?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 986000,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-26T06:29:15.283000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 957698,
      "author_name": "Kiranbeethoju",
      "author_url": "",
      "post_date": "2020-08-04T13:52:15.867000",
      "content": "<p>What is your batch size?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 957723,
          "author_name": "Dracarys",
          "author_url": "",
          "post_date": "2020-08-04T14:10:19.840000",
          "content": "<p>batch_size=24 and num_workers=4</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 957752,
          "author_name": "Kiranbeethoju",
          "author_url": "",
          "post_date": "2020-08-04T14:38:57.807000",
          "content": "<p>keep the batch size as default(32), try to set numworkers=0 then everything will work fine, I also face similar issues, I think there might a few bus in pytorch when we set numworkers to 4 or 6 it not working fine..., it would run faster if you won't get out of memory exception.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 957760,
          "author_name": "Dracarys",
          "author_url": "",
          "post_date": "2020-08-04T14:47:45.023000",
          "content": "<p><a href=\"/kiranbeethoju\">@kiranbeethoju</a> you are saying, you are able to train an epoch with around 12-14 lakh images with image size 256 in less than 3hr. Also i am not getting any out of memory issues, the whole trouble is there are 15lakh images, you won't be able to get reduce train time much until and unless you have access to cluster of GPU's.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 957769,
          "author_name": "Dracarys",
          "author_url": "",
          "post_date": "2020-08-04T14:51:10.440000",
          "content": "<p>I think you totally misread this discussion.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 957412,
      "author_name": "JUNAID KHAN",
      "author_url": "",
      "post_date": "2020-08-04T09:29:56.883000",
      "content": "<p>Use Google colab</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 991108,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-30T06:35:07.670000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 957786,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-04T15:01:50.573000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "956822": "I have an RTX2080 GPU(8GB), even with a smaller image size of 256x256, my efficientnetB0 (PyTorch with amp) model is taking around 4hrs per epoch to train. I guess with this speed it is too difficult to experiment. Does anyone have some better strategy to reduce training time. Any suggestion is welcome.",
    "956863": "Maybe you could try to use TPUs from colab/kaggle, they provide serious speedups especially over a single GPU.\nsee: https://arxiv.org/pdf/1907.10701.pdf",
    "957277": "It shows me 116gb free, when I open a new high ram TPU runtime.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F468424%2Fb4bb0b0c7834b0dee22cac2338128266%2FScreenshot%202020-08-04%20at%209.32.32.png?generation=1596526383427602&amp;alt=media)\n",
    "957698": "What is your batch size?",
    "957412": "Use Google colab",
    "991108": "",
    "957786": ""
  }
}