{
  "id": 263310,
  "title": "How long does it take for 1epoch?",
  "url": "/competitions/seti-breakthrough-listen/discussion/263310",
  "author_name": "kotatsu",
  "post_date": "2021-08-09T03:29:33.727000",
  "votes": 4,
  "comment_count": 12,
  "views": 0,
  "content": "<p>When I train a model, it takes about 4 ~ 5 hours for 1epoch(resnet18d, stratifiedKfold(n_splits=5). image_size=512) it takes too long to train in google colab(session expires within 24 hours)</p>\n<p>but as far as I can see this <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253385\" target=\"_blank\">discussion</a>, most of the participants can train more than 15epoch for 1 fold.</p>\n<p>I would greatly appreciate it if you could share how long does it take for 1 epoch and how to increase learning efficiency.</p>",
  "messages": [
    {
      "id": 1465263,
      "postDate": "2021-08-11T01:20:08.437Z",
      "content": "<p>In my case (Google colab pro, probably P100, image size 512 and 4 fold),<br>\nResnet34d: 820sec (13.5 mins)<br>\nEfficientnet_b0: 900sec (15 mins)</p>\n<p>My best guess is you put your training data in \"drive &gt; MyDrive\" not directly under \"contents\". Reading the data from google drive is SUPER slow in google colab. I faced this problem when I started to use Google colab for this competition</p>\n<p><a href=\"https://stackoverflow.com/questions/52929888/google-colab-very-slow-reading-data-images-from-google-drive\" target=\"_blank\">https://stackoverflow.com/questions/52929888/google-colab-very-slow-reading-data-images-from-google-drive</a></p>\n<p>If this is not the case, I have no idea for the potential solution. Might be better to check which part actually takes such a long time in your 1 epoch (data loading, forward, backward etc…)</p>",
      "rawMarkdown": "In my case (Google colab pro, probably P100, image size 512 and 4 fold),\nResnet34d: 820sec (13.5 mins)\nEfficientnet_b0: 900sec (15 mins)\n\nMy best guess is you put your training data in \"drive > MyDrive\" not directly under \"contents\". Reading the data from google drive is SUPER slow in google colab. I faced this problem when I started to use Google colab for this competition\n\nhttps://stackoverflow.com/questions/52929888/google-colab-very-slow-reading-data-images-from-google-drive\n\nIf this is not the case, I have no idea for the potential solution. Might be better to check which part actually takes such a long time in your 1 epoch (data loading, forward, backward etc...)",
      "votes": 5
    },
    {
      "id": 1460729,
      "postDate": "2021-08-09T03:48:58.483Z",
      "content": "<p>EfficientNetv2m<br>\n512 * 512<br>\n30min for 1epoch</p>",
      "rawMarkdown": "EfficientNetv2m\n512 * 512\n30min for 1epoch",
      "votes": 4,
      "replies": [
        {
          "id": 1460732,
          "postDate": "2021-08-09T03:53:06.913Z",
          "content": "<p>Thank you for sharing!<br>\nprobably there is a bug in my code…</p>",
          "rawMarkdown": "Thank you for sharing!\nprobably there is a bug in my code..."
        },
        {
          "id": 1462207,
          "postDate": "2021-08-09T18:18:22.287Z",
          "content": "<p>all data or only new train?</p>",
          "rawMarkdown": "all data or only new train?",
          "votes": -1
        },
        {
          "id": 1463118,
          "postDate": "2021-08-10T05:15:31.577Z",
          "content": "<p></p>",
          "rawMarkdown": "~~I use only new train~~",
          "votes": 1
        },
        {
          "id": 1467758,
          "postDate": "2021-08-12T06:12:28.573Z",
          "content": "<p>why there is a strikethrough 😆<br>\nDon't tell me we have to use all data…..I'm doomed </p>",
          "rawMarkdown": "why there is a strikethrough 😆\nDon't tell me we have to use all data.....I'm doomed "
        }
      ]
    },
    {
      "id": 1460705,
      "postDate": "2021-08-09T03:29:33.727Z",
      "content": "<p>When I train a model, it takes about 4 ~ 5 hours for 1epoch(resnet18d, stratifiedKfold(n_splits=5). image_size=512) it takes too long to train in google colab(session expires within 24 hours)</p>\n<p>but as far as I can see this <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253385\" target=\"_blank\">discussion</a>, most of the participants can train more than 15epoch for 1 fold.</p>\n<p>I would greatly appreciate it if you could share how long does it take for 1 epoch and how to increase learning efficiency.</p>",
      "rawMarkdown": "When I train a model, it takes about 4 ~ 5 hours for 1epoch(resnet18d, stratifiedKfold(n_splits=5). image_size=512) it takes too long to train in google colab(session expires within 24 hours)\n\nbut as far as I can see this [discussion](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253385), most of the participants can train more than 15epoch for 1 fold.\n\nI would greatly appreciate it if you could share how long does it take for 1 epoch and how to increase learning efficiency.\n",
      "votes": 4
    },
    {
      "id": 1467803,
      "postDate": "2021-08-12T06:37:50.987Z",
      "content": "<p><strong>458s</strong> per epoch for <code>nfnet_l0</code> with image size <code>819x256x3</code> and data parallel mode for <code>4x1080Ti</code>.</p>",
      "rawMarkdown": "**458s** per epoch for `nfnet_l0` with image size `819x256x3` and data parallel mode for `4x1080Ti`.",
      "votes": 1
    },
    {
      "id": 1461035,
      "postDate": "2021-08-09T07:21:14.613Z",
      "content": "<p>EfficientNetB2, original images (1638 x 256), TPU and TFRecords: 140s per epoch</p>",
      "rawMarkdown": "EfficientNetB2, original images (1638 x 256), TPU and TFRecords: 140s per epoch",
      "votes": 1,
      "replies": [
        {
          "id": 1462254,
          "postDate": "2021-08-09T18:50:49.920Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1463070,
      "postDate": "2021-08-10T04:46:26.963Z",
      "content": "<p>well, it depends actually. by random forking someone's model it takes me 1~2 hour per epoch when committing. But it only takes 20-30 min by using some simple model. If you want to improve the speed of training, you may can try to use TPU</p>",
      "rawMarkdown": "well, it depends actually. by random forking someone's model it takes me 1~2 hour per epoch when committing. But it only takes 20-30 min by using some simple model. If you want to improve the speed of training, you may can try to use TPU",
      "votes": 2
    },
    {
      "id": 1462381,
      "postDate": "2021-08-09T20:10:32.970Z",
      "content": "<p>It takes me about 20-30 min per epoch depending on the model in google colab</p>",
      "rawMarkdown": "It takes me about 20-30 min per epoch depending on the model in google colab",
      "votes": 2
    },
    {
      "id": 1460839,
      "postDate": "2021-08-09T05:13:23.957Z",
      "content": "<p>ssd or nvme harddisk 20min per epoch.</p>",
      "rawMarkdown": "ssd or nvme harddisk 20min per epoch."
    }
  ],
  "comments": [
    {
      "id": 1465263,
      "author_name": "yseeker",
      "author_url": "",
      "post_date": "2021-08-11T01:20:08.437000",
      "content": "<p>In my case (Google colab pro, probably P100, image size 512 and 4 fold),<br>\nResnet34d: 820sec (13.5 mins)<br>\nEfficientnet_b0: 900sec (15 mins)</p>\n<p>My best guess is you put your training data in \"drive &gt; MyDrive\" not directly under \"contents\". Reading the data from google drive is SUPER slow in google colab. I faced this problem when I started to use Google colab for this competition</p>\n<p><a href=\"https://stackoverflow.com/questions/52929888/google-colab-very-slow-reading-data-images-from-google-drive\" target=\"_blank\">https://stackoverflow.com/questions/52929888/google-colab-very-slow-reading-data-images-from-google-drive</a></p>\n<p>If this is not the case, I have no idea for the potential solution. Might be better to check which part actually takes such a long time in your 1 epoch (data loading, forward, backward etc…)</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1460729,
      "author_name": "Yamame🐟",
      "author_url": "",
      "post_date": "2021-08-09T03:48:58.483000",
      "content": "<p>EfficientNetv2m<br>\n512 * 512<br>\n30min for 1epoch</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1460732,
          "author_name": "kotatsu",
          "author_url": "",
          "post_date": "2021-08-09T03:53:06.913000",
          "content": "<p>Thank you for sharing!<br>\nprobably there is a bug in my code…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1462207,
          "author_name": "EnricRovira",
          "author_url": "",
          "post_date": "2021-08-09T18:18:22.287000",
          "content": "<p>all data or only new train?</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1463118,
          "author_name": "kotatsu",
          "author_url": "",
          "post_date": "2021-08-10T05:15:31.577000",
          "content": "<p></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1467758,
          "author_name": "yukiya",
          "author_url": "",
          "post_date": "2021-08-12T06:12:28.573000",
          "content": "<p>why there is a strikethrough 😆<br>\nDon't tell me we have to use all data…..I'm doomed </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1467803,
      "author_name": "Sergey Bryansky",
      "author_url": "",
      "post_date": "2021-08-12T06:37:50.987000",
      "content": "<p><strong>458s</strong> per epoch for <code>nfnet_l0</code> with image size <code>819x256x3</code> and data parallel mode for <code>4x1080Ti</code>.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1461035,
      "author_name": "Markus Frank",
      "author_url": "",
      "post_date": "2021-08-09T07:21:14.613000",
      "content": "<p>EfficientNetB2, original images (1638 x 256), TPU and TFRecords: 140s per epoch</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1462254,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-08-09T18:50:49.920000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1463070,
      "author_name": "lannds18",
      "author_url": "",
      "post_date": "2021-08-10T04:46:26.963000",
      "content": "<p>well, it depends actually. by random forking someone's model it takes me 1~2 hour per epoch when committing. But it only takes 20-30 min by using some simple model. If you want to improve the speed of training, you may can try to use TPU</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1462381,
      "author_name": "shiroe",
      "author_url": "",
      "post_date": "2021-08-09T20:10:32.970000",
      "content": "<p>It takes me about 20-30 min per epoch depending on the model in google colab</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1460839,
      "author_name": "shigengtian",
      "author_url": "",
      "post_date": "2021-08-09T05:13:23.957000",
      "content": "<p>ssd or nvme harddisk 20min per epoch.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1465263": "In my case (Google colab pro, probably P100, image size 512 and 4 fold),\nResnet34d: 820sec (13.5 mins)\nEfficientnet_b0: 900sec (15 mins)\n\nMy best guess is you put your training data in \"drive > MyDrive\" not directly under \"contents\". Reading the data from google drive is SUPER slow in google colab. I faced this problem when I started to use Google colab for this competition\n\nhttps://stackoverflow.com/questions/52929888/google-colab-very-slow-reading-data-images-from-google-drive\n\nIf this is not the case, I have no idea for the potential solution. Might be better to check which part actually takes such a long time in your 1 epoch (data loading, forward, backward etc...)",
    "1460729": "EfficientNetv2m\n512 * 512\n30min for 1epoch",
    "1460705": "When I train a model, it takes about 4 ~ 5 hours for 1epoch(resnet18d, stratifiedKfold(n_splits=5). image_size=512) it takes too long to train in google colab(session expires within 24 hours)\n\nbut as far as I can see this [discussion](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253385), most of the participants can train more than 15epoch for 1 fold.\n\nI would greatly appreciate it if you could share how long does it take for 1 epoch and how to increase learning efficiency.\n",
    "1467803": "**458s** per epoch for `nfnet_l0` with image size `819x256x3` and data parallel mode for `4x1080Ti`.",
    "1461035": "EfficientNetB2, original images (1638 x 256), TPU and TFRecords: 140s per epoch",
    "1463070": "well, it depends actually. by random forking someone's model it takes me 1~2 hour per epoch when committing. But it only takes 20-30 min by using some simple model. If you want to improve the speed of training, you may can try to use TPU",
    "1462381": "It takes me about 20-30 min per epoch depending on the model in google colab",
    "1460839": "ssd or nvme harddisk 20min per epoch."
  }
}