{
  "id": 161878,
  "title": "What is your setup?",
  "url": "/competitions/alaska2-image-steganalysis/discussion/161878",
  "author_name": "",
  "post_date": "2020-06-26T14:15:55.856395500Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone, </p>\n\n<p>As I'm getting more and more addicted to Kaggle, I'm starting to consider taking a Google Colab Pro extension as well as a Google One extension for storage.\nI looked at configuration specs but couldn't find any clear information about which GPU Kaggle is using.\nRight now, an epoch with all the images using 4/5 of the images of the dataset is taking one hour to run (and I am using 100% of my GPU quota). Using a TPU, I guess I can significantly reduce the training time but still consider moving to Colab.</p>\n\n<p>Does anyone here use a similar setup as the one I'm considering? \nBesides, what is you average training time using a TPU/GPU for an epoch?</p>\n\n<p>Many thanks in advance, \nPierre-Antoine</p>",
  "messages": [
    {
      "id": "903004",
      "postDate": "06/26/2020 14:15:55",
      "content": "<p>Hello everyone, </p>\n\n<p>As I'm getting more and more addicted to Kaggle, I'm starting to consider taking a Google Colab Pro extension as well as a Google One extension for storage.\nI looked at configuration specs but couldn't find any clear information about which GPU Kaggle is using.\nRight now, an epoch with all the images using 4/5 of the images of the dataset is taking one hour to run (and I am using 100% of my GPU quota). Using a TPU, I guess I can significantly reduce the training time but still consider moving to Colab.</p>\n\n<p>Does anyone here use a similar setup as the one I'm considering? \nBesides, what is you average training time using a TPU/GPU for an epoch?</p>\n\n<p>Many thanks in advance, \nPierre-Antoine</p>",
      "rawMarkdown": "Hello everyone, \n\nAs I'm getting more and more addicted to Kaggle, I'm starting to consider taking a Google Colab Pro extension as well as a Google One extension for storage.\nI looked at configuration specs but couldn't find any clear information about which GPU Kaggle is using.\nRight now, an epoch with all the images using 4/5 of the images of the dataset is taking one hour to run (and I am using 100% of my GPU quota). Using a TPU, I guess I can significantly reduce the training time but still consider moving to Colab.\n\nDoes anyone here use a similar setup as the one I'm considering? \nBesides, what is you average training time using a TPU/GPU for an epoch?\n\nMany thanks in advance, \nPierre-Antoine",
      "votes": null
    },
    {
      "id": "903105",
      "postDate": "06/26/2020 15:18:03",
      "content": "<p>My first epoch took almost 30 minutes on TPU with 60K images. It will take more if I use all of the images at once. The second epoch took almost 12 minutes. I am training efficient net B3 and all of its parameters are trainable at the moment. Probably it will take less time if do not train the whole model. </p>\n\n<p>Regarding the setup, I think Kaggle uses Tesla P100 as far as I know. Not sure if they have changed it, </p>\n\n<p>My personal setup at the moment is Kaggle only and if the data is small in size then I run on my local machine with 2060 RTX GPU. But, that happens very less often. I wish I can do more training on the local machine because the notebook dies after some inactivity on Kaggle. </p>\n\n<p>I am not sure if colab pro is really great or not as I have not used that. But, I am pretty much sure it is helpful as many people here use that for training. </p>\n\n<p>Btw what is your training time for one epoch? </p>",
      "rawMarkdown": "My first epoch took almost 30 minutes on TPU with 60K images. It will take more if I use all of the images at once. The second epoch took almost 12 minutes. I am training efficient net B3 and all of its parameters are trainable at the moment. Probably it will take less time if do not train the whole model. \n\nRegarding the setup, I think Kaggle uses Tesla P100 as far as I know. Not sure if they have changed it, \n\nMy personal setup at the moment is Kaggle only and if the data is small in size then I run on my local machine with 2060 RTX GPU. But, that happens very less often. I wish I can do more training on the local machine because the notebook dies after some inactivity on Kaggle. \n\nI am not sure if colab pro is really great or not as I have not used that. But, I am pretty much sure it is helpful as many people here use that for training. \n\nBtw what is your training time for one epoch?",
      "votes": null
    },
    {
      "id": "903225",
      "postDate": "06/26/2020 16:48:55",
      "content": "<p>I've tried on GCP with Tesla T4 but taking to much time. Using Kaggle kernel for long time training is pretty bad idea as for the constraints time. Colab is not any way suitable either, I think. I've RTX 2070, just one epoch on complete dataset took around 4/5 hours, couldn't exceed batch size more than 8 with E0. Colab pro seems promising but unavailable in our region. Didn't try on multi GPU, maybe that could solve. This competition is somewhat like this: start training, and take a long trip for 2/4 days and return for your submission, see the results, make a decision and continue. </p>\n\n<p><a href=\"/wuliaokaola\">@wuliaokaola</a> come one, still, you need more power 😄 </p>",
      "rawMarkdown": "I've tried on GCP with Tesla T4 but taking to much time. Using Kaggle kernel for long time training is pretty bad idea as for the constraints time. Colab is not any way suitable either, I think. I've RTX 2070, just one epoch on complete dataset took around 4/5 hours, couldn't exceed batch size more than 8 with E0. Colab pro seems promising but unavailable in our region. Didn't try on multi GPU, maybe that could solve. This competition is somewhat like this: start training, and take a long trip for 2/4 days and return for your submission, see the results, make a decision and continue. \n\n@wuliaokaola come one, still, you need more power 😄",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 903105,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "06/26/2020 15:18:03",
      "content": "<p>My first epoch took almost 30 minutes on TPU with 60K images. It will take more if I use all of the images at once. The second epoch took almost 12 minutes. I am training efficient net B3 and all of its parameters are trainable at the moment. Probably it will take less time if do not train the whole model. </p>\n\n<p>Regarding the setup, I think Kaggle uses Tesla P100 as far as I know. Not sure if they have changed it, </p>\n\n<p>My personal setup at the moment is Kaggle only and if the data is small in size then I run on my local machine with 2060 RTX GPU. But, that happens very less often. I wish I can do more training on the local machine because the notebook dies after some inactivity on Kaggle. </p>\n\n<p>I am not sure if colab pro is really great or not as I have not used that. But, I am pretty much sure it is helpful as many people here use that for training. </p>\n\n<p>Btw what is your training time for one epoch? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 903225,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "06/26/2020 16:48:55",
      "content": "<p>I've tried on GCP with Tesla T4 but taking to much time. Using Kaggle kernel for long time training is pretty bad idea as for the constraints time. Colab is not any way suitable either, I think. I've RTX 2070, just one epoch on complete dataset took around 4/5 hours, couldn't exceed batch size more than 8 with E0. Colab pro seems promising but unavailable in our region. Didn't try on multi GPU, maybe that could solve. This competition is somewhat like this: start training, and take a long trip for 2/4 days and return for your submission, see the results, make a decision and continue. </p>\n\n<p><a href=\"/wuliaokaola\">@wuliaokaola</a> come one, still, you need more power 😄 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "903004": "Hello everyone, \n\nAs I'm getting more and more addicted to Kaggle, I'm starting to consider taking a Google Colab Pro extension as well as a Google One extension for storage.\nI looked at configuration specs but couldn't find any clear information about which GPU Kaggle is using.\nRight now, an epoch with all the images using 4/5 of the images of the dataset is taking one hour to run (and I am using 100% of my GPU quota). Using a TPU, I guess I can significantly reduce the training time but still consider moving to Colab.\n\nDoes anyone here use a similar setup as the one I'm considering? \nBesides, what is you average training time using a TPU/GPU for an epoch?\n\nMany thanks in advance, \nPierre-Antoine",
    "903105": "My first epoch took almost 30 minutes on TPU with 60K images. It will take more if I use all of the images at once. The second epoch took almost 12 minutes. I am training efficient net B3 and all of its parameters are trainable at the moment. Probably it will take less time if do not train the whole model. \n\nRegarding the setup, I think Kaggle uses Tesla P100 as far as I know. Not sure if they have changed it, \n\nMy personal setup at the moment is Kaggle only and if the data is small in size then I run on my local machine with 2060 RTX GPU. But, that happens very less often. I wish I can do more training on the local machine because the notebook dies after some inactivity on Kaggle. \n\nI am not sure if colab pro is really great or not as I have not used that. But, I am pretty much sure it is helpful as many people here use that for training. \n\nBtw what is your training time for one epoch?",
    "903225": "I've tried on GCP with Tesla T4 but taking to much time. Using Kaggle kernel for long time training is pretty bad idea as for the constraints time. Colab is not any way suitable either, I think. I've RTX 2070, just one epoch on complete dataset took around 4/5 hours, couldn't exceed batch size more than 8 with E0. Colab pro seems promising but unavailable in our region. Didn't try on multi GPU, maybe that could solve. This competition is somewhat like this: start training, and take a long trip for 2/4 days and return for your submission, see the results, make a decision and continue. \n\n@wuliaokaola come one, still, you need more power 😄"
  },
  "source": "meta"
}