{
  "id": 308707,
  "title": "Kaggle is too slow processing my model",
  "url": "/competitions/happy-whale-and-dolphin/discussion/308707",
  "author_name": "",
  "post_date": "2022-02-19T20:12:40.935194500Z",
  "votes": 5,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>I have created a CNN model for this competition but when I try to run it, it would take hours and practically impossible. Has anyone else run into this issue?<br>\nHow do you then use CNN on kaggle?</p>",
  "messages": [
    {
      "id": "1697756",
      "postDate": "02/19/2022 20:12:40",
      "content": "<p>Hello all,</p>\n<p>I have created a CNN model for this competition but when I try to run it, it would take hours and practically impossible. Has anyone else run into this issue?<br>\nHow do you then use CNN on kaggle?</p>",
      "rawMarkdown": "Hello all,\n\nI have created a CNN model for this competition but when I try to run it, it would take hours and practically impossible. Has anyone else run into this issue?\nHow do you then use CNN on kaggle?",
      "votes": null
    },
    {
      "id": "1697801",
      "postDate": "02/19/2022 21:07:29",
      "content": "<p>I think you should change attitude. Kaggle is enough you should simplify your pipeline. You will benefit:</p>\n<ul>\n<li>less parameters to optimize -&gt; faster to train and inference</li>\n<li>simpler solution will probably score better</li>\n</ul>\n<p>I suggest you reading this one article: <a href=\"https://machinelearningmastery.com/ensemble-learning-and-occams-razor/\" target=\"_blank\">https://machinelearningmastery.com/ensemble-learning-and-occams-razor/</a> It describes Occam’s razor in Machine Learning. Thinking such way you will benefit a lot. </p>",
      "rawMarkdown": "I think you should change attitude. Kaggle is enough you should simplify your pipeline. You will benefit:\n- less parameters to optimize -> faster to train and inference\n- simpler solution will probably score better\n\nI suggest you reading this one article: https://machinelearningmastery.com/ensemble-learning-and-occams-razor/ It describes Occam’s razor in Machine Learning. Thinking such way you will benefit a lot.",
      "votes": null
    },
    {
      "id": "1697868",
      "postDate": "02/19/2022 23:13:13",
      "content": "<p>I have used a pretrained CNN with the addition of a top layer. Nothing complicated. </p>\n<p>Really interested to know how others are using CNN for this competition while Kaggle runs this slowly.</p>",
      "rawMarkdown": "I have used a pretrained CNN with the addition of a top layer. Nothing complicated. \n\nReally interested to know how others are using CNN for this competition while Kaggle runs this slowly.",
      "votes": null
    },
    {
      "id": "1698331",
      "postDate": "02/20/2022 10:00:04",
      "content": "<p>Hey, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, thank you for sharing that article… but I'm in two minds. I understand the philosophy of Occam's razor, it definitely makes sense. But what about ensemble? From one point of view, it makes sense to avoid ensembles because of complexity. On another side - a lot of competitions end up with dozens of models in the ensemble.</p>\n<p>Where is the truth? <br>\nAnd one optional question - how can we evaluate ensemble? I mean we came up with two models (both KFold validated, OOF prediction saved). How to make sure that our ensemble of that two models is working well taking into account the fact that we utilized all the data on KFold validation). Are there any common use techniques? </p>\n<p>Would love to hear your thoughts about this. Thank you in advance</p>",
      "rawMarkdown": "Hey, @remekkinas, thank you for sharing that article... but I'm in two minds. I understand the philosophy of Occam's razor, it definitely makes sense. But what about ensemble? From one point of view, it makes sense to avoid ensembles because of complexity. On another side - a lot of competitions end up with dozens of models in the ensemble.\n\nWhere is the truth? \nAnd one optional question - how can we evaluate ensemble? I mean we came up with two models (both KFold validated, OOF prediction saved). How to make sure that our ensemble of that two models is working well taking into account the fact that we utilized all the data on KFold validation). Are there any common use techniques? \n\nWould love to hear your thoughts about this. Thank you in advance",
      "votes": null
    },
    {
      "id": "1698335",
      "postDate": "02/20/2022 10:07:50",
      "content": "<p>In the same article he wrote about Occam’s Razor and Ensemble Learning. </p>\n<p>\"The first razor remains an important heuristic in applied machine learning. The key aspect<br>\nof this razor is the predicate of all else being equal. That is, if two models are compared, they<br>\nmust be compared using their generalization error on a holdout dataset or estimated using<br>\nk-fold cross-validation. If their performance is equal under these circumstances, then the razor<br>\ncan come into effect and we can choose the simpler solution.<br>\nThis is not the only way to choose models. We may choose a simpler model because it is<br>\neasier to interpret, and this remains valid if model interpretability is a more important project<br>\nrequirement than predictive skill. Ensemble learning algorithms are unambiguously a more<br>\ncomplex type of model when the number of model parameters is considered the measure of<br>\ncomplexity. As such, an open problem in machine learning involves alternate measures of<br>\ncomplexit\"</p>\n<p>If you blend/ensemble many models you can validate them as well. Look here <a href=\"https://www.kaggle.com/c/tabular-playground-series-feb-2022/discussion/307367\" target=\"_blank\">https://www.kaggle.com/c/tabular-playground-series-feb-2022/discussion/307367</a> … I provided a lot of examples how to validate using sckikit learn function (it is quite good attitude) or custom cross validation loop (if you look for some customizations). Ultimately you can use one-leave method to validate models. BTW: if you blend using submission files validation is out of controll in my opinion this is why most blending provides LB overfitting only.</p>",
      "rawMarkdown": "In the same article he wrote about Occam’s Razor and Ensemble Learning. \n\n\"The first razor remains an important heuristic in applied machine learning. The key aspect\nof this razor is the predicate of all else being equal. That is, if two models are compared, they\nmust be compared using their generalization error on a holdout dataset or estimated using\nk-fold cross-validation. If their performance is equal under these circumstances, then the razor\ncan come into effect and we can choose the simpler solution.\nThis is not the only way to choose models. We may choose a simpler model because it is\neasier to interpret, and this remains valid if model interpretability is a more important project\nrequirement than predictive skill. Ensemble learning algorithms are unambiguously a more\ncomplex type of model when the number of model parameters is considered the measure of\ncomplexity. As such, an open problem in machine learning involves alternate measures of\ncomplexit\"\n\nIf you blend/ensemble many models you can validate them as well. Look here https://www.kaggle.com/c/tabular-playground-series-feb-2022/discussion/307367 ... I provided a lot of examples how to validate using sckikit learn function (it is quite good attitude) or custom cross validation loop (if you look for some customizations). Ultimately you can use one-leave method to validate models. BTW: if you blend using submission files validation is out of controll in my opinion this is why most blending provides LB overfitting only.",
      "votes": null
    },
    {
      "id": "1699026",
      "postDate": "02/20/2022 20:55:08",
      "content": "<p>It was the same for me as I am beginner too. It was taking around 3.5 hours to iterate through the whole data for one epoch while using VGG16(transfer learning) and changing the classification layers. I learned to use gpu which reduced the time to around 35 min. Then I converted all images to a fixed 256x256 size and uploaded the new dataset. This reduced the train time per epoch to 10-12 min. I know it leads to infomation loss, but I am still testing and trying to learn. And now, I learned to use TPU for pytorch model and the train time per epoch is around 2-3 minutes. Seeing to learn more, but I hope this might help.</p>",
      "rawMarkdown": "It was the same for me as I am beginner too. It was taking around 3.5 hours to iterate through the whole data for one epoch while using VGG16(transfer learning) and changing the classification layers. I learned to use gpu which reduced the time to around 35 min. Then I converted all images to a fixed 256x256 size and uploaded the new dataset. This reduced the train time per epoch to 10-12 min. I know it leads to infomation loss, but I am still testing and trying to learn. And now, I learned to use TPU for pytorch model and the train time per epoch is around 2-3 minutes. Seeing to learn more, but I hope this might help.",
      "votes": null
    },
    {
      "id": "1699033",
      "postDate": "02/20/2022 21:11:17",
      "content": "<p>Does it take 2-3 minutes to run one epoch using 256x256 dolphins data?</p>",
      "rawMarkdown": "Does it take 2-3 minutes to run one epoch using 256x256 dolphins data?",
      "votes": null
    },
    {
      "id": "1699035",
      "postDate": "02/20/2022 21:13:31",
      "content": "<p>For me yes, when I used VGG16.</p>",
      "rawMarkdown": "For me yes, when I used VGG16.",
      "votes": null
    },
    {
      "id": "1699107",
      "postDate": "02/20/2022 23:31:50",
      "content": "<p><a href=\"https://www.kaggle.com/beinggakash\" target=\"_blank\">@beinggakash</a> Hey, thanks for the info.<br>\nI have tried to use TPUs but still haven't managed to successfully run on TPU. No matter what I do, my code throws at me. I have followed the instructions and put my entire model within the TPU scope too but still running into issues.</p>\n<p>I have read the Kaggle instruction but it's awful and not clear.</p>\n<p>How did you learn to use TPUs on kaggle?<br>\nThanks.</p>",
      "rawMarkdown": "beinggakash Hey, thanks for the info.\nI have tried to use TPUs but still haven't managed to successfully run on TPU. No matter what I do, my code throws at me. I have followed the instructions and put my entire model within the TPU scope too but still running into issues.\n\nI have read the Kaggle instruction but it's awful and not clear.\n\nHow did you learn to use TPUs on kaggle?\nThanks.",
      "votes": null
    },
    {
      "id": "1699314",
      "postDate": "02/21/2022 05:22:03",
      "content": "<p>I referred to a few notebooks to write one. I shall try searching for that notebook, if I find, I will shall attach the link to the notebook.</p>",
      "rawMarkdown": "I referred to a few notebooks to write one. I shall try searching for that notebook, if I find, I will shall attach the link to the notebook.",
      "votes": null
    },
    {
      "id": "1699341",
      "postDate": "02/21/2022 05:50:03",
      "content": "<p><a href=\"https://www.kaggle.com/piantic/pytorch-tpu/notebook\" target=\"_blank\">https://www.kaggle.com/piantic/pytorch-tpu/notebook</a><br>\n<a href=\"https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r\" target=\"_blank\">https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r</a></p>\n<p>I referred to these two.</p>",
      "rawMarkdown": "https://www.kaggle.com/piantic/pytorch-tpu/notebook\nhttps://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r\n\nI referred to these two.",
      "votes": null
    },
    {
      "id": "1699996",
      "postDate": "02/21/2022 15:31:55",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!",
      "votes": null
    },
    {
      "id": "1701389",
      "postDate": "02/22/2022 18:07:45",
      "content": "<p>Unfortunately those are with Pytorch. I have used tensorflow. Can't get the TPUs to work and for only 20 epochs it takes about 20 hours to run. Not sure how on earth people are training their models when Kaggle resources are so limited.<br>\nMoving on to the next competition.</p>",
      "rawMarkdown": "Unfortunately those are with Pytorch. I have used tensorflow. Can't get the TPUs to work and for only 20 epochs it takes about 20 hours to run. Not sure how on earth people are training their models when Kaggle resources are so limited.\nMoving on to the next competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1697801,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "02/19/2022 21:07:29",
      "content": "<p>I think you should change attitude. Kaggle is enough you should simplify your pipeline. You will benefit:</p>\n<ul>\n<li>less parameters to optimize -&gt; faster to train and inference</li>\n<li>simpler solution will probably score better</li>\n</ul>\n<p>I suggest you reading this one article: <a href=\"https://machinelearningmastery.com/ensemble-learning-and-occams-razor/\" target=\"_blank\">https://machinelearningmastery.com/ensemble-learning-and-occams-razor/</a> It describes Occam’s razor in Machine Learning. Thinking such way you will benefit a lot. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1697868,
          "author_name": "amir6878",
          "author_url": "",
          "post_date": "02/19/2022 23:13:13",
          "content": "<p>I have used a pretrained CNN with the addition of a top layer. Nothing complicated. </p>\n<p>Really interested to know how others are using CNN for this competition while Kaggle runs this slowly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1698331,
          "author_name": "meowmeowmeowmeowmeow",
          "author_url": "",
          "post_date": "02/20/2022 10:00:04",
          "content": "<p>Hey, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, thank you for sharing that article… but I'm in two minds. I understand the philosophy of Occam's razor, it definitely makes sense. But what about ensemble? From one point of view, it makes sense to avoid ensembles because of complexity. On another side - a lot of competitions end up with dozens of models in the ensemble.</p>\n<p>Where is the truth? <br>\nAnd one optional question - how can we evaluate ensemble? I mean we came up with two models (both KFold validated, OOF prediction saved). How to make sure that our ensemble of that two models is working well taking into account the fact that we utilized all the data on KFold validation). Are there any common use techniques? </p>\n<p>Would love to hear your thoughts about this. Thank you in advance</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1698335,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/20/2022 10:07:50",
          "content": "<p>In the same article he wrote about Occam’s Razor and Ensemble Learning. </p>\n<p>\"The first razor remains an important heuristic in applied machine learning. The key aspect<br>\nof this razor is the predicate of all else being equal. That is, if two models are compared, they<br>\nmust be compared using their generalization error on a holdout dataset or estimated using<br>\nk-fold cross-validation. If their performance is equal under these circumstances, then the razor<br>\ncan come into effect and we can choose the simpler solution.<br>\nThis is not the only way to choose models. We may choose a simpler model because it is<br>\neasier to interpret, and this remains valid if model interpretability is a more important project<br>\nrequirement than predictive skill. Ensemble learning algorithms are unambiguously a more<br>\ncomplex type of model when the number of model parameters is considered the measure of<br>\ncomplexity. As such, an open problem in machine learning involves alternate measures of<br>\ncomplexit\"</p>\n<p>If you blend/ensemble many models you can validate them as well. Look here <a href=\"https://www.kaggle.com/c/tabular-playground-series-feb-2022/discussion/307367\" target=\"_blank\">https://www.kaggle.com/c/tabular-playground-series-feb-2022/discussion/307367</a> … I provided a lot of examples how to validate using sckikit learn function (it is quite good attitude) or custom cross validation loop (if you look for some customizations). Ultimately you can use one-leave method to validate models. BTW: if you blend using submission files validation is out of controll in my opinion this is why most blending provides LB overfitting only.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1699026,
      "author_name": "beinggakash",
      "author_url": "",
      "post_date": "02/20/2022 20:55:08",
      "content": "<p>It was the same for me as I am beginner too. It was taking around 3.5 hours to iterate through the whole data for one epoch while using VGG16(transfer learning) and changing the classification layers. I learned to use gpu which reduced the time to around 35 min. Then I converted all images to a fixed 256x256 size and uploaded the new dataset. This reduced the train time per epoch to 10-12 min. I know it leads to infomation loss, but I am still testing and trying to learn. And now, I learned to use TPU for pytorch model and the train time per epoch is around 2-3 minutes. Seeing to learn more, but I hope this might help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1699033,
          "author_name": "meowmeowmeowmeowmeow",
          "author_url": "",
          "post_date": "02/20/2022 21:11:17",
          "content": "<p>Does it take 2-3 minutes to run one epoch using 256x256 dolphins data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1699035,
          "author_name": "beinggakash",
          "author_url": "",
          "post_date": "02/20/2022 21:13:31",
          "content": "<p>For me yes, when I used VGG16.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1699107,
          "author_name": "amir6878",
          "author_url": "",
          "post_date": "02/20/2022 23:31:50",
          "content": "<p><a href=\"https://www.kaggle.com/beinggakash\" target=\"_blank\">@beinggakash</a> Hey, thanks for the info.<br>\nI have tried to use TPUs but still haven't managed to successfully run on TPU. No matter what I do, my code throws at me. I have followed the instructions and put my entire model within the TPU scope too but still running into issues.</p>\n<p>I have read the Kaggle instruction but it's awful and not clear.</p>\n<p>How did you learn to use TPUs on kaggle?<br>\nThanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1699314,
          "author_name": "beinggakash",
          "author_url": "",
          "post_date": "02/21/2022 05:22:03",
          "content": "<p>I referred to a few notebooks to write one. I shall try searching for that notebook, if I find, I will shall attach the link to the notebook.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1699341,
          "author_name": "beinggakash",
          "author_url": "",
          "post_date": "02/21/2022 05:50:03",
          "content": "<p><a href=\"https://www.kaggle.com/piantic/pytorch-tpu/notebook\" target=\"_blank\">https://www.kaggle.com/piantic/pytorch-tpu/notebook</a><br>\n<a href=\"https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r\" target=\"_blank\">https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r</a></p>\n<p>I referred to these two.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1699996,
          "author_name": "amir6878",
          "author_url": "",
          "post_date": "02/21/2022 15:31:55",
          "content": "<p>Thank you very much!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1701389,
          "author_name": "amir6878",
          "author_url": "",
          "post_date": "02/22/2022 18:07:45",
          "content": "<p>Unfortunately those are with Pytorch. I have used tensorflow. Can't get the TPUs to work and for only 20 epochs it takes about 20 hours to run. Not sure how on earth people are training their models when Kaggle resources are so limited.<br>\nMoving on to the next competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1697756": "Hello all,\n\nI have created a CNN model for this competition but when I try to run it, it would take hours and practically impossible. Has anyone else run into this issue?\nHow do you then use CNN on kaggle?",
    "1697801": "I think you should change attitude. Kaggle is enough you should simplify your pipeline. You will benefit:\n- less parameters to optimize -> faster to train and inference\n- simpler solution will probably score better\n\nI suggest you reading this one article: https://machinelearningmastery.com/ensemble-learning-and-occams-razor/ It describes Occam’s razor in Machine Learning. Thinking such way you will benefit a lot.",
    "1697868": "I have used a pretrained CNN with the addition of a top layer. Nothing complicated. \n\nReally interested to know how others are using CNN for this competition while Kaggle runs this slowly.",
    "1698331": "Hey, @remekkinas, thank you for sharing that article... but I'm in two minds. I understand the philosophy of Occam's razor, it definitely makes sense. But what about ensemble? From one point of view, it makes sense to avoid ensembles because of complexity. On another side - a lot of competitions end up with dozens of models in the ensemble.\n\nWhere is the truth? \nAnd one optional question - how can we evaluate ensemble? I mean we came up with two models (both KFold validated, OOF prediction saved). How to make sure that our ensemble of that two models is working well taking into account the fact that we utilized all the data on KFold validation). Are there any common use techniques? \n\nWould love to hear your thoughts about this. Thank you in advance",
    "1698335": "In the same article he wrote about Occam’s Razor and Ensemble Learning. \n\n\"The first razor remains an important heuristic in applied machine learning. The key aspect\nof this razor is the predicate of all else being equal. That is, if two models are compared, they\nmust be compared using their generalization error on a holdout dataset or estimated using\nk-fold cross-validation. If their performance is equal under these circumstances, then the razor\ncan come into effect and we can choose the simpler solution.\nThis is not the only way to choose models. We may choose a simpler model because it is\neasier to interpret, and this remains valid if model interpretability is a more important project\nrequirement than predictive skill. Ensemble learning algorithms are unambiguously a more\ncomplex type of model when the number of model parameters is considered the measure of\ncomplexity. As such, an open problem in machine learning involves alternate measures of\ncomplexit\"\n\nIf you blend/ensemble many models you can validate them as well. Look here https://www.kaggle.com/c/tabular-playground-series-feb-2022/discussion/307367 ... I provided a lot of examples how to validate using sckikit learn function (it is quite good attitude) or custom cross validation loop (if you look for some customizations). Ultimately you can use one-leave method to validate models. BTW: if you blend using submission files validation is out of controll in my opinion this is why most blending provides LB overfitting only.",
    "1699026": "It was the same for me as I am beginner too. It was taking around 3.5 hours to iterate through the whole data for one epoch while using VGG16(transfer learning) and changing the classification layers. I learned to use gpu which reduced the time to around 35 min. Then I converted all images to a fixed 256x256 size and uploaded the new dataset. This reduced the train time per epoch to 10-12 min. I know it leads to infomation loss, but I am still testing and trying to learn. And now, I learned to use TPU for pytorch model and the train time per epoch is around 2-3 minutes. Seeing to learn more, but I hope this might help.",
    "1699033": "Does it take 2-3 minutes to run one epoch using 256x256 dolphins data?",
    "1699035": "For me yes, when I used VGG16.",
    "1699107": "beinggakash Hey, thanks for the info.\nI have tried to use TPUs but still haven't managed to successfully run on TPU. No matter what I do, my code throws at me. I have followed the instructions and put my entire model within the TPU scope too but still running into issues.\n\nI have read the Kaggle instruction but it's awful and not clear.\n\nHow did you learn to use TPUs on kaggle?\nThanks.",
    "1699314": "I referred to a few notebooks to write one. I shall try searching for that notebook, if I find, I will shall attach the link to the notebook.",
    "1699341": "https://www.kaggle.com/piantic/pytorch-tpu/notebook\nhttps://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r\n\nI referred to these two.",
    "1699996": "Thank you very much!",
    "1701389": "Unfortunately those are with Pytorch. I have used tensorflow. Can't get the TPUs to work and for only 20 epochs it takes about 20 hours to run. Not sure how on earth people are training their models when Kaggle resources are so limited.\nMoving on to the next competition."
  },
  "source": "meta"
}