{
  "id": 265356,
  "title": "Infra and training tricks",
  "url": "/competitions/landmark-recognition-2021/discussion/265356",
  "author_name": "datafool",
  "post_date": "2021-08-15T15:30:09.043000",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi All, </p>\n<p>Given the size of data which we have, to me it seems its not possible to train any meaningful model on kaggle kernel, would be more than happy to be proven wrong. Reason for this deduction of mine is, I trained VGG19 with only last layer to be trainable, and for 1 epoch on 1000 randomly selected batches, with batch size of 8, training runs for 430 seconds. If i extrapolate this to entire dataset, 1 epoch would run for around 23 to 24 hours. </p>\n<p>Kaggle kernel can be run only for 9 hours on straight, which is not enough for even 1 epoch. So, question is what infra is being used by others for this competition, any cost benefit analysis would be good to discuss. Or, if there are any training tricks which others are using would be good to learn from. </p>\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": 1473528,
      "postDate": "2021-08-15T15:30:09.043Z",
      "content": "<p>Hi All, </p>\n<p>Given the size of data which we have, to me it seems its not possible to train any meaningful model on kaggle kernel, would be more than happy to be proven wrong. Reason for this deduction of mine is, I trained VGG19 with only last layer to be trainable, and for 1 epoch on 1000 randomly selected batches, with batch size of 8, training runs for 430 seconds. If i extrapolate this to entire dataset, 1 epoch would run for around 23 to 24 hours. </p>\n<p>Kaggle kernel can be run only for 9 hours on straight, which is not enough for even 1 epoch. So, question is what infra is being used by others for this competition, any cost benefit analysis would be good to discuss. Or, if there are any training tricks which others are using would be good to learn from. </p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Hi All, \n\nGiven the size of data which we have, to me it seems its not possible to train any meaningful model on kaggle kernel, would be more than happy to be proven wrong. Reason for this deduction of mine is, I trained VGG19 with only last layer to be trainable, and for 1 epoch on 1000 randomly selected batches, with batch size of 8, training runs for 430 seconds. If i extrapolate this to entire dataset, 1 epoch would run for around 23 to 24 hours. \n\nKaggle kernel can be run only for 9 hours on straight, which is not enough for even 1 epoch. So, question is what infra is being used by others for this competition, any cost benefit analysis would be good to discuss. Or, if there are any training tricks which others are using would be good to learn from. \n\nThanks in advance!",
      "votes": 3
    },
    {
      "id": 1475563,
      "postDate": "2021-08-16T17:56:22.560Z",
      "content": "<p>The discussion of Kaggle kernels being too limited in resources comes up a lot and the fact of the manner is: it's true. <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/250008#1371386\" target=\"_blank\">Getting into large scale Kaggle competitions requires a ton of computational resources</a>. However, you might be able to get around some of the limitations using the following solution:</p>\n<ul>\n<li>Use a notebook to exclusively train a single model then save the weights as outputs.</li>\n<li>With several notebooks you can create an ensemble.</li>\n<li>Upload the weights for your models to Kaggle as a dataset.</li>\n<li>In a new notebook you can build the models and load the weights from the notebooks.</li>\n<li>Using these pretrained models you can make predictions and do any postprocessing you need.</li>\n</ul>\n<p>Hope this gives you some idea of how to get around the resources needed 🙂 </p>",
      "rawMarkdown": "The discussion of Kaggle kernels being too limited in resources comes up a lot and the fact of the manner is: it's true. [Getting into large scale Kaggle competitions requires a ton of computational resources](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/250008#1371386). However, you might be able to get around some of the limitations using the following solution:\n\n- Use a notebook to exclusively train a single model then save the weights as outputs.\n- With several notebooks you can create an ensemble.\n- Upload the weights for your models to Kaggle as a dataset.\n- In a new notebook you can build the models and load the weights from the notebooks.\n- Using these pretrained models you can make predictions and do any postprocessing you need.\n\nHope this gives you some idea of how to get around the resources needed 🙂 \n\n",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1475563,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-16T17:56:22.560000",
      "content": "<p>The discussion of Kaggle kernels being too limited in resources comes up a lot and the fact of the manner is: it's true. <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/250008#1371386\" target=\"_blank\">Getting into large scale Kaggle competitions requires a ton of computational resources</a>. However, you might be able to get around some of the limitations using the following solution:</p>\n<ul>\n<li>Use a notebook to exclusively train a single model then save the weights as outputs.</li>\n<li>With several notebooks you can create an ensemble.</li>\n<li>Upload the weights for your models to Kaggle as a dataset.</li>\n<li>In a new notebook you can build the models and load the weights from the notebooks.</li>\n<li>Using these pretrained models you can make predictions and do any postprocessing you need.</li>\n</ul>\n<p>Hope this gives you some idea of how to get around the resources needed 🙂 </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1473528": "Hi All, \n\nGiven the size of data which we have, to me it seems its not possible to train any meaningful model on kaggle kernel, would be more than happy to be proven wrong. Reason for this deduction of mine is, I trained VGG19 with only last layer to be trainable, and for 1 epoch on 1000 randomly selected batches, with batch size of 8, training runs for 430 seconds. If i extrapolate this to entire dataset, 1 epoch would run for around 23 to 24 hours. \n\nKaggle kernel can be run only for 9 hours on straight, which is not enough for even 1 epoch. So, question is what infra is being used by others for this competition, any cost benefit analysis would be good to discuss. Or, if there are any training tricks which others are using would be good to learn from. \n\nThanks in advance!",
    "1475563": "The discussion of Kaggle kernels being too limited in resources comes up a lot and the fact of the manner is: it's true. [Getting into large scale Kaggle competitions requires a ton of computational resources](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/250008#1371386). However, you might be able to get around some of the limitations using the following solution:\n\n- Use a notebook to exclusively train a single model then save the weights as outputs.\n- With several notebooks you can create an ensemble.\n- Upload the weights for your models to Kaggle as a dataset.\n- In a new notebook you can build the models and load the weights from the notebooks.\n- Using these pretrained models you can make predictions and do any postprocessing you need.\n\nHope this gives you some idea of how to get around the resources needed 🙂 \n\n"
  }
}