{
  "id": 82119,
  "title": "Is it normal to take 30 minutes for one epoch?",
  "url": "/competitions/histopathologic-cancer-detection/discussion/82119",
  "author_name": "",
  "post_date": "2019-02-27T09:49:15.875110700Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>I am new to machine learning and this is my first Kaggle competition! I'm on my third attempt but my model takes around two and a half hours to run (5 epochs and 30 mins each). I have to wait a while to see if what I've implemented/changed has improved my model.</p>\n\n<p>I'm using Keras and the flow_from_dataframe function to load in my data. I'm also using the network architecture seen here: <a href=\"https://www.kaggle.com/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93\">https://www.kaggle.com/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93</a>\nIs this simply because there are 200k images or could I improve this time by using a different way to load in the data? I was thinking of changing the directories to let me use flow_from_directory. </p>\n\n<p>I'd be grateful for any advice. Thanks!</p>",
  "messages": [
    {
      "id": "479634",
      "postDate": "02/27/2019 09:49:15",
      "content": "<p>Hello,</p>\n\n<p>I am new to machine learning and this is my first Kaggle competition! I'm on my third attempt but my model takes around two and a half hours to run (5 epochs and 30 mins each). I have to wait a while to see if what I've implemented/changed has improved my model.</p>\n\n<p>I'm using Keras and the flow_from_dataframe function to load in my data. I'm also using the network architecture seen here: <a href=\"https://www.kaggle.com/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93\">https://www.kaggle.com/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93</a>\nIs this simply because there are 200k images or could I improve this time by using a different way to load in the data? I was thinking of changing the directories to let me use flow_from_directory. </p>\n\n<p>I'd be grateful for any advice. Thanks!</p>",
      "rawMarkdown": "Hello,\n\nI am new to machine learning and this is my first Kaggle competition! I'm on my third attempt but my model takes around two and a half hours to run (5 epochs and 30 mins each). I have to wait a while to see if what I've implemented/changed has improved my model.\n\nI'm using Keras and the flow_from_dataframe function to load in my data. I'm also using the network architecture seen here: [https://www.kaggle.com/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93][1]\nIs this simply because there are 200k images or could I improve this time by using a different way to load in the data? I was thinking of changing the directories to let me use flow_from_directory. \n\nI'd be grateful for any advice. Thanks!\n\n  [1]: https://www.kaggle.com/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93",
      "votes": null
    },
    {
      "id": "479848",
      "postDate": "02/27/2019 13:49:56",
      "content": "<p>Hi Jake,</p>\n\n<p>I would recommend checking on Google Colab/GCP where you can use TPU (it's free!) which is a game changer for such use cases. </p>\n\n<p>I have provided required info here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/79823#468260\">https://www.kaggle.com/c/microsoft-malware-prediction/discussion/79823#468260</a></p>\n\n<p>HTH</p>",
      "rawMarkdown": "Hi Jake,\n\nI would recommend checking on Google Colab/GCP where you can use TPU (it's free!) which is a game changer for such use cases. \n\nI have provided required info here:\n\nhttps://www.kaggle.com/c/microsoft-malware-prediction/discussion/79823#468260\n\nHTH",
      "votes": null
    },
    {
      "id": "479853",
      "postDate": "02/27/2019 13:56:48",
      "content": "<p>Hi Vibhor,</p>\n\n<p>Thank you for the reply! I will be sure to check that out.</p>",
      "rawMarkdown": "Hi Vibhor,\n\nThank you for the reply! I will be sure to check that out.",
      "votes": null
    },
    {
      "id": "483580",
      "postDate": "03/04/2019 20:56:08",
      "content": "<p>Hi Jake, 30 mins per epoch is not bad at all from my experience. With strong augmentation and batch size of 64, it takes me more than an hour per epoch. But Keras does print the current loss and accuracy on the training set. So that's a way to see if your metrics go into the right direction.</p>",
      "rawMarkdown": "Hi Jake, 30 mins per epoch is not bad at all from my experience. With strong augmentation and batch size of 64, it takes me more than an hour per epoch. But Keras does print the current loss and accuracy on the training set. So that's a way to see if your metrics go into the right direction.",
      "votes": null
    },
    {
      "id": "483581",
      "postDate": "03/04/2019 20:57:44",
      "content": "<p><a href=\"/vjcalling\">@vjcalling</a>  How do you keep your Colab instance alive during training? Mine always dies due to browser inactivity.</p>",
      "rawMarkdown": "vjcalling  How do you keep your Colab instance alive during training? Mine always dies due to browser inactivity.",
      "votes": null
    },
    {
      "id": "483727",
      "postDate": "03/05/2019 03:18:12",
      "content": "<p>Hi <a href=\"/franchini\">@franchini</a>!</p>\n\n<p>Is it because of OOM? Can you pls confirm?</p>",
      "rawMarkdown": "Hi @franchini!\n\nIs it because of OOM? Can you pls confirm?",
      "votes": null
    },
    {
      "id": "483894",
      "postDate": "03/05/2019 09:56:05",
      "content": "<p>Okay it's good to know that's the expected running time. Thanks for the reply.</p>",
      "rawMarkdown": "Okay it's good to know that's the expected running time. Thanks for the reply.",
      "votes": null
    },
    {
      "id": "484297",
      "postDate": "03/05/2019 19:37:06",
      "content": "<p>Hi Jake,\nYou might want to check your GPU RAM and the architecture you are using. You might want to increase the batch size to increase the training. I used transfer-learning from different Resnet architecture and none of them took 30 mins each. Approximately 5~15 mins each depending on the complexity of the architecture on a Tesla P4. Hopefully this helps. </p>",
      "rawMarkdown": "Hi Jake,\nYou might want to check your GPU RAM and the architecture you are using. You might want to increase the batch size to increase the training. I used transfer-learning from different Resnet architecture and none of them took 30 mins each. Approximately 5~15 mins each depending on the complexity of the architecture on a Tesla P4. Hopefully this helps.",
      "votes": null
    },
    {
      "id": "485484",
      "postDate": "03/07/2019 12:45:52",
      "content": "<p>@Franchini\nHere is answer,\n<a href=\"https://stackoverflow.com/questions/47474406/lifetime-of-a-colab-vm\">https://stackoverflow.com/questions/47474406/lifetime-of-a-colab-vm</a></p>",
      "rawMarkdown": "Franchini\nHere is answer,\nhttps://stackoverflow.com/questions/47474406/lifetime-of-a-colab-vm",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 479848,
      "author_name": "vjcalling",
      "author_url": "",
      "post_date": "02/27/2019 13:49:56",
      "content": "<p>Hi Jake,</p>\n\n<p>I would recommend checking on Google Colab/GCP where you can use TPU (it's free!) which is a game changer for such use cases. </p>\n\n<p>I have provided required info here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/79823#468260\">https://www.kaggle.com/c/microsoft-malware-prediction/discussion/79823#468260</a></p>\n\n<p>HTH</p>",
      "votes": null,
      "replies": [
        {
          "id": 479853,
          "author_name": "jtmurkz",
          "author_url": "",
          "post_date": "02/27/2019 13:56:48",
          "content": "<p>Hi Vibhor,</p>\n\n<p>Thank you for the reply! I will be sure to check that out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 483581,
          "author_name": "franchini",
          "author_url": "",
          "post_date": "03/04/2019 20:57:44",
          "content": "<p><a href=\"/vjcalling\">@vjcalling</a>  How do you keep your Colab instance alive during training? Mine always dies due to browser inactivity.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 483727,
          "author_name": "vjcalling",
          "author_url": "",
          "post_date": "03/05/2019 03:18:12",
          "content": "<p>Hi <a href=\"/franchini\">@franchini</a>!</p>\n\n<p>Is it because of OOM? Can you pls confirm?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 485484,
          "author_name": "deepblue1010",
          "author_url": "",
          "post_date": "03/07/2019 12:45:52",
          "content": "<p>@Franchini\nHere is answer,\n<a href=\"https://stackoverflow.com/questions/47474406/lifetime-of-a-colab-vm\">https://stackoverflow.com/questions/47474406/lifetime-of-a-colab-vm</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 483580,
      "author_name": "franchini",
      "author_url": "",
      "post_date": "03/04/2019 20:56:08",
      "content": "<p>Hi Jake, 30 mins per epoch is not bad at all from my experience. With strong augmentation and batch size of 64, it takes me more than an hour per epoch. But Keras does print the current loss and accuracy on the training set. So that's a way to see if your metrics go into the right direction.</p>",
      "votes": null,
      "replies": [
        {
          "id": 483894,
          "author_name": "jtmurkz",
          "author_url": "",
          "post_date": "03/05/2019 09:56:05",
          "content": "<p>Okay it's good to know that's the expected running time. Thanks for the reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 484297,
      "author_name": "davidhopezhao",
      "author_url": "",
      "post_date": "03/05/2019 19:37:06",
      "content": "<p>Hi Jake,\nYou might want to check your GPU RAM and the architecture you are using. You might want to increase the batch size to increase the training. I used transfer-learning from different Resnet architecture and none of them took 30 mins each. Approximately 5~15 mins each depending on the complexity of the architecture on a Tesla P4. Hopefully this helps. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "479634": "Hello,\n\nI am new to machine learning and this is my first Kaggle competition! I'm on my third attempt but my model takes around two and a half hours to run (5 epochs and 30 mins each). I have to wait a while to see if what I've implemented/changed has improved my model.\n\nI'm using Keras and the flow_from_dataframe function to load in my data. I'm also using the network architecture seen here: [https://www.kaggle.com/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93][1]\nIs this simply because there are 200k images or could I improve this time by using a different way to load in the data? I was thinking of changing the directories to let me use flow_from_directory. \n\nI'd be grateful for any advice. Thanks!\n\n  [1]: https://www.kaggle.com/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93",
    "479848": "Hi Jake,\n\nI would recommend checking on Google Colab/GCP where you can use TPU (it's free!) which is a game changer for such use cases. \n\nI have provided required info here:\n\nhttps://www.kaggle.com/c/microsoft-malware-prediction/discussion/79823#468260\n\nHTH",
    "479853": "Hi Vibhor,\n\nThank you for the reply! I will be sure to check that out.",
    "483580": "Hi Jake, 30 mins per epoch is not bad at all from my experience. With strong augmentation and batch size of 64, it takes me more than an hour per epoch. But Keras does print the current loss and accuracy on the training set. So that's a way to see if your metrics go into the right direction.",
    "483581": "vjcalling  How do you keep your Colab instance alive during training? Mine always dies due to browser inactivity.",
    "483727": "Hi @franchini!\n\nIs it because of OOM? Can you pls confirm?",
    "483894": "Okay it's good to know that's the expected running time. Thanks for the reply.",
    "484297": "Hi Jake,\nYou might want to check your GPU RAM and the architecture you are using. You might want to increase the batch size to increase the training. I used transfer-learning from different Resnet architecture and none of them took 30 mins each. Approximately 5~15 mins each depending on the complexity of the architecture on a Tesla P4. Hopefully this helps.",
    "485484": "Franchini\nHere is answer,\nhttps://stackoverflow.com/questions/47474406/lifetime-of-a-colab-vm"
  },
  "source": "meta"
}