{
  "id": 174171,
  "title": "Where do you train?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174171",
  "author_name": "",
  "post_date": "2020-08-12T14:35:29.701989300Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I wonder where do you train in this competition? Do you train on Kaggle, locally or on Google Colab?</p>\n<p>I did everything on Kaggle directly, currently I'm trying to train on TPU following model:<br>\nModel: EfficientNet-B4<br>\nImage-Size: 384x384<br>\nFolds: 5-Folds<br>\nNo external-data<br>\nDownsampled Dataset (512x512-JPEG)<br>\nEpoch (atleast what I would like): 10</p>\n<p>An epoch with validation takes: ~18 Minutes, the 3 hours limitation of TPU (using all 8 Cores) on Kaggle does not allow me to train for 10 epochs.</p>\n<p>Roughly calculated: <br>\n10x18x5 = 900 Minutes -&gt; 15 hours</p>\n<p>Does anyone have any tips or tricks?</p>\n<p>Thank you in advance</p>",
  "messages": [
    {
      "id": "967848",
      "postDate": "08/12/2020 14:35:29",
      "content": "<p>Hi,</p>\n<p>I wonder where do you train in this competition? Do you train on Kaggle, locally or on Google Colab?</p>\n<p>I did everything on Kaggle directly, currently I'm trying to train on TPU following model:<br>\nModel: EfficientNet-B4<br>\nImage-Size: 384x384<br>\nFolds: 5-Folds<br>\nNo external-data<br>\nDownsampled Dataset (512x512-JPEG)<br>\nEpoch (atleast what I would like): 10</p>\n<p>An epoch with validation takes: ~18 Minutes, the 3 hours limitation of TPU (using all 8 Cores) on Kaggle does not allow me to train for 10 epochs.</p>\n<p>Roughly calculated: <br>\n10x18x5 = 900 Minutes -&gt; 15 hours</p>\n<p>Does anyone have any tips or tricks?</p>\n<p>Thank you in advance</p>",
      "rawMarkdown": "Hi,\n\nI wonder where do you train in this competition? Do you train on Kaggle, locally or on Google Colab?\n\nI did everything on Kaggle directly, currently I'm trying to train on TPU following model:\nModel: EfficientNet-B4\nImage-Size: 384x384\nFolds: 5-Folds\nNo external-data\nDownsampled Dataset (512x512-JPEG)\nEpoch (atleast what I would like): 10\n\nAn epoch with validation takes: ~18 Minutes, the 3 hours limitation of TPU (using all 8 Cores) on Kaggle does not allow me to train for 10 epochs.\n\nRoughly calculated: \n10x18x5 = 900 Minutes -&gt; 15 hours\n\nDoes anyone have any tips or tricks?\n\nThank you in advance",
      "votes": null
    },
    {
      "id": "967914",
      "postDate": "08/12/2020 15:28:11",
      "content": "<p>In Colab you could have a 12-hour continuous usage after which you will get a 12-hour cooldown. So a hack is to train using 2 google accounts alternatively and checkpoint your models in the drive.<br>\nOr you could train foldwise on Kaggle notebooks and ensemble all the submission files in another notebook.<br>\nNOTE: Kaggle TPUs have more RAM compared to Colab TPUs. So a good idea is to train small models on Colab and large ones on Kaggle.✌️</p>",
      "rawMarkdown": "In Colab you could have a 12-hour continuous usage after which you will get a 12-hour cooldown. So a hack is to train using 2 google accounts alternatively and checkpoint your models in the drive.\nOr you could train foldwise on Kaggle notebooks and ensemble all the submission files in another notebook.\nNOTE: Kaggle TPUs have more RAM compared to Colab TPUs. So a good idea is to train small models on Colab and large ones on Kaggle.✌️",
      "votes": null
    },
    {
      "id": "967976",
      "postDate": "08/12/2020 15:58:27",
      "content": "<p>Do you use TFRecs? For example <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\" target=\"_blank\">https://www.kaggle.com/cdeotte/melanoma-384x384</a>. It could speed up your learning process. I train models in Kaggle notebooks, of course have some problems (hard to work with big models and high resolutions), but I can train 512*512 with EfficientNetB4 + ISIC2019</p>",
      "rawMarkdown": "Do you use TFRecs? For example https://www.kaggle.com/cdeotte/melanoma-384x384. It could speed up your learning process. I train models in Kaggle notebooks, of course have some problems (hard to work with big models and high resolutions), but I can train 512*512 with EfficientNetB4 + ISIC2019",
      "votes": null
    },
    {
      "id": "967999",
      "postDate": "08/12/2020 16:14:17",
      "content": "<p>No I dont use TFRecs but I checked and he has 384x384 also for JPEG. Thank you for pointing that out.</p>",
      "rawMarkdown": "No I dont use TFRecs but I checked and he has 384x384 also for JPEG. Thank you for pointing that out.",
      "votes": null
    },
    {
      "id": "968002",
      "postDate": "08/12/2020 16:15:44",
      "content": "<p>Thank you! I did not think about that…but training Foldwise on Kaggle sounds like a great idea. Then 3 hours for each model would be more than enough for me! </p>",
      "rawMarkdown": "Thank you! I did not think about that...but training Foldwise on Kaggle sounds like a great idea. Then 3 hours for each model would be more than enough for me!",
      "votes": null
    },
    {
      "id": "968165",
      "postDate": "08/12/2020 18:31:03",
      "content": "<p>I use simple cloud solution - <a href=\"https://www.paperspace.com/\" target=\"_blank\">Paperspace</a>. Easy to use and affordable prices. Free GPU</p>",
      "rawMarkdown": "I use simple cloud solution - [Paperspace](https://www.paperspace.com/). Easy to use and affordable prices. Free GPU",
      "votes": null
    },
    {
      "id": "969602",
      "postDate": "08/13/2020 19:51:23",
      "content": "<p>Locally. Using pytorch and gpu.</p>",
      "rawMarkdown": "Locally. Using pytorch and gpu.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 967914,
      "author_name": "josealways123",
      "author_url": "",
      "post_date": "08/12/2020 15:28:11",
      "content": "<p>In Colab you could have a 12-hour continuous usage after which you will get a 12-hour cooldown. So a hack is to train using 2 google accounts alternatively and checkpoint your models in the drive.<br>\nOr you could train foldwise on Kaggle notebooks and ensemble all the submission files in another notebook.<br>\nNOTE: Kaggle TPUs have more RAM compared to Colab TPUs. So a good idea is to train small models on Colab and large ones on Kaggle.✌️</p>",
      "votes": null,
      "replies": [
        {
          "id": 968002,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "08/12/2020 16:15:44",
          "content": "<p>Thank you! I did not think about that…but training Foldwise on Kaggle sounds like a great idea. Then 3 hours for each model would be more than enough for me! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 967976,
      "author_name": "aybatov",
      "author_url": "",
      "post_date": "08/12/2020 15:58:27",
      "content": "<p>Do you use TFRecs? For example <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\" target=\"_blank\">https://www.kaggle.com/cdeotte/melanoma-384x384</a>. It could speed up your learning process. I train models in Kaggle notebooks, of course have some problems (hard to work with big models and high resolutions), but I can train 512*512 with EfficientNetB4 + ISIC2019</p>",
      "votes": null,
      "replies": [
        {
          "id": 967999,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "08/12/2020 16:14:17",
          "content": "<p>No I dont use TFRecs but I checked and he has 384x384 also for JPEG. Thank you for pointing that out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 968165,
      "author_name": "volcanoflash",
      "author_url": "",
      "post_date": "08/12/2020 18:31:03",
      "content": "<p>I use simple cloud solution - <a href=\"https://www.paperspace.com/\" target=\"_blank\">Paperspace</a>. Easy to use and affordable prices. Free GPU</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 969602,
      "author_name": "bfishh",
      "author_url": "",
      "post_date": "08/13/2020 19:51:23",
      "content": "<p>Locally. Using pytorch and gpu.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "967848": "Hi,\n\nI wonder where do you train in this competition? Do you train on Kaggle, locally or on Google Colab?\n\nI did everything on Kaggle directly, currently I'm trying to train on TPU following model:\nModel: EfficientNet-B4\nImage-Size: 384x384\nFolds: 5-Folds\nNo external-data\nDownsampled Dataset (512x512-JPEG)\nEpoch (atleast what I would like): 10\n\nAn epoch with validation takes: ~18 Minutes, the 3 hours limitation of TPU (using all 8 Cores) on Kaggle does not allow me to train for 10 epochs.\n\nRoughly calculated: \n10x18x5 = 900 Minutes -&gt; 15 hours\n\nDoes anyone have any tips or tricks?\n\nThank you in advance",
    "967914": "In Colab you could have a 12-hour continuous usage after which you will get a 12-hour cooldown. So a hack is to train using 2 google accounts alternatively and checkpoint your models in the drive.\nOr you could train foldwise on Kaggle notebooks and ensemble all the submission files in another notebook.\nNOTE: Kaggle TPUs have more RAM compared to Colab TPUs. So a good idea is to train small models on Colab and large ones on Kaggle.✌️",
    "967976": "Do you use TFRecs? For example https://www.kaggle.com/cdeotte/melanoma-384x384. It could speed up your learning process. I train models in Kaggle notebooks, of course have some problems (hard to work with big models and high resolutions), but I can train 512*512 with EfficientNetB4 + ISIC2019",
    "967999": "No I dont use TFRecs but I checked and he has 384x384 also for JPEG. Thank you for pointing that out.",
    "968002": "Thank you! I did not think about that...but training Foldwise on Kaggle sounds like a great idea. Then 3 hours for each model would be more than enough for me!",
    "968165": "I use simple cloud solution - [Paperspace](https://www.paperspace.com/). Easy to use and affordable prices. Free GPU",
    "969602": "Locally. Using pytorch and gpu."
  },
  "source": "meta"
}