{
  "id": 231800,
  "title": "How everyone is training such a big dataset within the kaggle timelimit?",
  "url": "/competitions/bms-molecular-translation/discussion/231800",
  "author_name": "Shalvin P Shaji",
  "post_date": "2021-04-10T11:49:37.996000",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hai,<br>\nI am pretty new to kaggle. The dataset provided is very big. Even by choosing a few subset (50k) of training examples it takes literally hours to train a model just below the kaggle hour limit. Can someone tell whether I am doing something wrong or is it the way it is supposed to be?<br>\nSorry if this question does not make any sense to you </p>",
  "messages": [
    {
      "id": 1269352,
      "postDate": "2021-04-10T13:16:59.503Z",
      "content": "<p><a href=\"https://www.kaggle.com/shalvinpshaji\" target=\"_blank\">@shalvinpshaji</a> The majority of the people aren't using kaggle kernels for this. They use either their own machines,  google colab or any other web platform. You can use kaggle kernels, by  running it for one epoch, saving the weights and optimizers/schedulers states, and then restart training by loading those values. And you keep on going until you are done with training.</p>",
      "rawMarkdown": "@shalvinpshaji The majority of the people aren't using kaggle kernels for this. They use either their own machines,  google colab or any other web platform. You can use kaggle kernels, by  running it for one epoch, saving the weights and optimizers/schedulers states, and then restart training by loading those values. And you keep on going until you are done with training.",
      "votes": 5
    },
    {
      "id": 1269285,
      "postDate": "2021-04-10T11:49:37.997Z",
      "content": "<p>Hai,<br>\nI am pretty new to kaggle. The dataset provided is very big. Even by choosing a few subset (50k) of training examples it takes literally hours to train a model just below the kaggle hour limit. Can someone tell whether I am doing something wrong or is it the way it is supposed to be?<br>\nSorry if this question does not make any sense to you </p>",
      "rawMarkdown": "Hai,\nI am pretty new to kaggle. The dataset provided is very big. Even by choosing a few subset (50k) of training examples it takes literally hours to train a model just below the kaggle hour limit. Can someone tell whether I am doing something wrong or is it the way it is supposed to be?\nSorry if this question does not make any sense to you ",
      "votes": 6
    },
    {
      "id": 1269561,
      "postDate": "2021-04-10T16:47:43.877Z",
      "content": "<p>I use Kaggle Kernels. </p>\n<p>I utilize free TPU. And am usually able to process/infer-on around 40 million images per 9 hour session.</p>\n<p>I will be releasing a very detailed training notebook soon… so stay tuned!</p>\n<p>Also feel free to share a link to your notebook, it will allow others to better offer suggestions/help.</p>",
      "rawMarkdown": "I use Kaggle Kernels. \n\nI utilize free TPU. And am usually able to process/infer-on around 40 million images per 9 hour session.\n\nI will be releasing a very detailed training notebook soon... so stay tuned!\n\nAlso feel free to share a link to your notebook, it will allow others to better offer suggestions/help.",
      "votes": 3,
      "replies": [
        {
          "id": 1273418,
          "postDate": "2021-04-14T10:29:13.323Z",
          "content": "<p>Is the TPU notebook runtime limit 9 hours? I thought it was 3 hours for TPU and 9 hours for GPU.</p>",
          "rawMarkdown": "Is the TPU notebook runtime limit 9 hours? I thought it was 3 hours for TPU and 9 hours for GPU."
        }
      ]
    },
    {
      "id": 1270060,
      "postDate": "2021-04-11T08:36:45.157Z",
      "content": "<p>Hi,<br>\nthis is also my first competition and I started with GPUs, which was terribly slow (probably also because of a bad model choice). <br>\nCheck out this <a href=\"https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92\" target=\"_blank\">notebook</a> on how to use TPUs in this competition. It was extremly helpful to me.<br>\nThen training is quite fast compared to before and reading through it I was able write an inference notebook that processes all test data in 10min (+20min for conversion to InChI strings).</p>",
      "rawMarkdown": "Hi,\nthis is also my first competition and I started with GPUs, which was terribly slow (probably also because of a bad model choice). \nCheck out this [notebook](https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92) on how to use TPUs in this competition. It was extremly helpful to me.\nThen training is quite fast compared to before and reading through it I was able write an inference notebook that processes all test data in 10min (+20min for conversion to InChI strings).",
      "votes": 1,
      "replies": [
        {
          "id": 1270667,
          "postDate": "2021-04-11T21:24:56.933Z",
          "content": "<p>will you share the inference part? because there is official example about TPU training but no inference on the whole test dataset. </p>",
          "rawMarkdown": "will you share the inference part? because there is official example about TPU training but no inference on the whole test dataset. ",
          "votes": 1
        },
        {
          "id": 1272366,
          "postDate": "2021-04-13T12:25:17.733Z",
          "content": "<p>Hi, the original notebook author decided against sharing the inference part and I would like to respect that. But the necessary code blocks are basically all there. Hint: the validation loop is very similar to the final inference loop. I learned a lot by having to write the inference by myself :)</p>",
          "rawMarkdown": "Hi, the original notebook author decided against sharing the inference part and I would like to respect that. But the necessary code blocks are basically all there. Hint: the validation loop is very similar to the final inference loop. I learned a lot by having to write the inference by myself :)"
        }
      ]
    },
    {
      "id": 1271073,
      "postDate": "2021-04-12T09:20:12.060Z",
      "content": "<p>(For now) I train 10 epochs, 1000 steps each with batch size 512 on the provided TPUs and get training + inference in under 2h, so I think with TPUs it's definitely possible</p>",
      "rawMarkdown": "(For now) I train 10 epochs, 1000 steps each with batch size 512 on the provided TPUs and get training + inference in under 2h, so I think with TPUs it's definitely possible",
      "votes": 2,
      "replies": [
        {
          "id": 1271920,
          "postDate": "2021-04-13T04:04:59.883Z",
          "content": "<p>Hai, are you using pytorch for training, if so could you provide some reference notebooks?</p>",
          "rawMarkdown": "Hai, are you using pytorch for training, if so could you provide some reference notebooks?",
          "votes": 1
        },
        {
          "id": 1288291,
          "postDate": "2021-04-29T20:41:31.013Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1269352,
      "author_name": "Nuno Ferreira",
      "author_url": "",
      "post_date": "2021-04-10T13:16:59.503000",
      "content": "<p><a href=\"https://www.kaggle.com/shalvinpshaji\" target=\"_blank\">@shalvinpshaji</a> The majority of the people aren't using kaggle kernels for this. They use either their own machines,  google colab or any other web platform. You can use kaggle kernels, by  running it for one epoch, saving the weights and optimizers/schedulers states, and then restart training by loading those values. And you keep on going until you are done with training.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1269561,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-04-10T16:47:43.877000",
      "content": "<p>I use Kaggle Kernels. </p>\n<p>I utilize free TPU. And am usually able to process/infer-on around 40 million images per 9 hour session.</p>\n<p>I will be releasing a very detailed training notebook soon… so stay tuned!</p>\n<p>Also feel free to share a link to your notebook, it will allow others to better offer suggestions/help.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1273418,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2021-04-14T10:29:13.323000",
          "content": "<p>Is the TPU notebook runtime limit 9 hours? I thought it was 3 hours for TPU and 9 hours for GPU.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1270060,
      "author_name": "Michael Wolff",
      "author_url": "",
      "post_date": "2021-04-11T08:36:45.157000",
      "content": "<p>Hi,<br>\nthis is also my first competition and I started with GPUs, which was terribly slow (probably also because of a bad model choice). <br>\nCheck out this <a href=\"https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92\" target=\"_blank\">notebook</a> on how to use TPUs in this competition. It was extremly helpful to me.<br>\nThen training is quite fast compared to before and reading through it I was able write an inference notebook that processes all test data in 10min (+20min for conversion to InChI strings).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1270667,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2021-04-11T21:24:56.933000",
          "content": "<p>will you share the inference part? because there is official example about TPU training but no inference on the whole test dataset. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1272366,
          "author_name": "Michael Wolff",
          "author_url": "",
          "post_date": "2021-04-13T12:25:17.733000",
          "content": "<p>Hi, the original notebook author decided against sharing the inference part and I would like to respect that. But the necessary code blocks are basically all there. Hint: the validation loop is very similar to the final inference loop. I learned a lot by having to write the inference by myself :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1271073,
      "author_name": "Max",
      "author_url": "",
      "post_date": "2021-04-12T09:20:12.060000",
      "content": "<p>(For now) I train 10 epochs, 1000 steps each with batch size 512 on the provided TPUs and get training + inference in under 2h, so I think with TPUs it's definitely possible</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1271920,
          "author_name": "Shalvin P Shaji",
          "author_url": "",
          "post_date": "2021-04-13T04:04:59.883000",
          "content": "<p>Hai, are you using pytorch for training, if so could you provide some reference notebooks?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1288291,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-29T20:41:31.013000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1269352": "@shalvinpshaji The majority of the people aren't using kaggle kernels for this. They use either their own machines,  google colab or any other web platform. You can use kaggle kernels, by  running it for one epoch, saving the weights and optimizers/schedulers states, and then restart training by loading those values. And you keep on going until you are done with training.",
    "1269285": "Hai,\nI am pretty new to kaggle. The dataset provided is very big. Even by choosing a few subset (50k) of training examples it takes literally hours to train a model just below the kaggle hour limit. Can someone tell whether I am doing something wrong or is it the way it is supposed to be?\nSorry if this question does not make any sense to you ",
    "1269561": "I use Kaggle Kernels. \n\nI utilize free TPU. And am usually able to process/infer-on around 40 million images per 9 hour session.\n\nI will be releasing a very detailed training notebook soon... so stay tuned!\n\nAlso feel free to share a link to your notebook, it will allow others to better offer suggestions/help.",
    "1270060": "Hi,\nthis is also my first competition and I started with GPUs, which was terribly slow (probably also because of a bad model choice). \nCheck out this [notebook](https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92) on how to use TPUs in this competition. It was extremly helpful to me.\nThen training is quite fast compared to before and reading through it I was able write an inference notebook that processes all test data in 10min (+20min for conversion to InChI strings).",
    "1271073": "(For now) I train 10 epochs, 1000 steps each with batch size 512 on the provided TPUs and get training + inference in under 2h, so I think with TPUs it's definitely possible"
  }
}