{
  "id": 270329,
  "title": " Question reg repeating dataset and replicas in Keras",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/270329",
  "author_name": "",
  "post_date": "2021-09-04T17:43:58.831237700Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>I'm more of a pytorch user but want to use Keras for this competition because there's an easy way to chain together different models using the functional API. I noticed that when doing distributed training, for example in this notebook here by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> : <a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-training\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-training</a> , there are two things that I would like some clarity on:</p>\n<ol>\n<li><p>What benefit does the <code>dataset.repeat()</code> option have? I also have TFRecords and would like to use a structure similar to this notebook. </p></li>\n<li><p>In order to use <code>dataset.repeat()</code>, Keras requires the <code>steps_per_epoch</code> option to be specified when using <code>model.fit()</code>. </p></li>\n</ol>\n<p>Normally, you would calculate this as:</p>\n<pre><code>steps_per_epoch = (num_tf_train_records * files_per_tf_record) // batch_size\n</code></pre>\n<p>If I understand correctly, when using a TPU or multi-gpu setup, you must divide this by the number of replicas, so that formula would be</p>\n<pre><code>steps_per_epoch = (num_tf_train_records * files_per_tf_record) // batch_size // replicas\n</code></pre>\n<p>However, I see that it has been specified as:</p>\n<pre><code>steps_per_epoch = (num_tf_train_records * files_per_tf_record) // batch_size // replicas // 4\n</code></pre>\n<p>What is the purpose of the additional <code>4</code> there? I am trying to replicate this setup in GPU and want to accurately calculate the <code>steps_per_epoch</code> for a 2 GPU setup. Would really appreciate some insight from the Keras experts here. Thank you</p>",
  "messages": [
    {
      "id": "1502855",
      "postDate": "09/04/2021 17:43:58",
      "content": "<p>Hello all,</p>\n<p>I'm more of a pytorch user but want to use Keras for this competition because there's an easy way to chain together different models using the functional API. I noticed that when doing distributed training, for example in this notebook here by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> : <a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-training\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-training</a> , there are two things that I would like some clarity on:</p>\n<ol>\n<li><p>What benefit does the <code>dataset.repeat()</code> option have? I also have TFRecords and would like to use a structure similar to this notebook. </p></li>\n<li><p>In order to use <code>dataset.repeat()</code>, Keras requires the <code>steps_per_epoch</code> option to be specified when using <code>model.fit()</code>. </p></li>\n</ol>\n<p>Normally, you would calculate this as:</p>\n<pre><code>steps_per_epoch = (num_tf_train_records * files_per_tf_record) // batch_size\n</code></pre>\n<p>If I understand correctly, when using a TPU or multi-gpu setup, you must divide this by the number of replicas, so that formula would be</p>\n<pre><code>steps_per_epoch = (num_tf_train_records * files_per_tf_record) // batch_size // replicas\n</code></pre>\n<p>However, I see that it has been specified as:</p>\n<pre><code>steps_per_epoch = (num_tf_train_records * files_per_tf_record) // batch_size // replicas // 4\n</code></pre>\n<p>What is the purpose of the additional <code>4</code> there? I am trying to replicate this setup in GPU and want to accurately calculate the <code>steps_per_epoch</code> for a 2 GPU setup. Would really appreciate some insight from the Keras experts here. Thank you</p>",
      "rawMarkdown": "Hello all,\n\nI'm more of a pytorch user but want to use Keras for this competition because there's an easy way to chain together different models using the functional API. I noticed that when doing distributed training, for example in this notebook here by @hidehisaarai1213 : https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-training , there are two things that I would like some clarity on:\n\n1. What benefit does the `dataset.repeat()` option have? I also have TFRecords and would like to use a structure similar to this notebook. \n\n2. In order to use `dataset.repeat()`, Keras requires the `steps_per_epoch` option to be specified when using `model.fit()`. \n\nNormally, you would calculate this as:\n```\nsteps_per_epoch = (num_tf_train_records * files_per_tf_record) // batch_size\n```\nIf I understand correctly, when using a TPU or multi-gpu setup, you must divide this by the number of replicas, so that formula would be\n\n```\nsteps_per_epoch = (num_tf_train_records * files_per_tf_record) // batch_size // replicas\n```\n\nHowever, I see that it has been specified as:\n\n```\nsteps_per_epoch = (num_tf_train_records * files_per_tf_record) // batch_size // replicas // 4\n```\n\nWhat is the purpose of the additional `4` there? I am trying to replicate this setup in GPU and want to accurately calculate the `steps_per_epoch` for a 2 GPU setup. Would really appreciate some insight from the Keras experts here. Thank you",
      "votes": null
    },
    {
      "id": "1503272",
      "postDate": "09/05/2021 08:04:30",
      "content": "<p>I was also surprised when I saw <em>repeat</em> and <em>//4</em> in this kernel. Looks like only 1/4 of training data is used each epoch. Maybe this is done because of time limitation of kaggle kernel. Hopefully  the author can comment on that.</p>\n<p>If you just want to use all training data each epoch - you don't need to use <em>repeat</em>. In that case you don't have to specify <em>steps_per_epoch</em> as well. Your formula for <em>steps_per_epoch</em> is correct.</p>",
      "rawMarkdown": "I was also surprised when I saw *repeat* and *//4* in this kernel. Looks like only 1/4 of training data is used each epoch. Maybe this is done because of time limitation of kaggle kernel. Hopefully  the author can comment on that.\n\nIf you just want to use all training data each epoch - you don't need to use *repeat*. In that case you don't have to specify *steps_per_epoch* as well. Your formula for *steps_per_epoch* is correct.",
      "votes": null
    },
    {
      "id": "1503724",
      "postDate": "09/05/2021 16:43:54",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/vzaguskin\" target=\"_blank\">@vzaguskin</a> . Using a subset of data for the kernel time limit makes sense. I will try avoiding using repeat and see what happens. :)</p>",
      "rawMarkdown": "Thanks @vzaguskin . Using a subset of data for the kernel time limit makes sense. I will try avoiding using repeat and see what happens. :)",
      "votes": null
    },
    {
      "id": "1508777",
      "postDate": "09/10/2021 14:43:51",
      "content": "<p>This has been clarified <a href=\"https://www.kaggle.com/ragnar123/g2net-effb7-100-seed-21\" target=\"_blank\">here</a> in the comment section.</p>\n<p>Indeed, it's 1/4 of training data each epoch, but training dataset doesn't reset each epoch, so for the period of 4 epochs the model sees 4 unique pieces of training data (1/4 of training set each epoch).</p>\n<p>Hence, real amount of times our model passes through training data is divided by 4 (relative to the number of epochs), but none of the training samples are getting thrown away during training.</p>",
      "rawMarkdown": "This has been clarified [here](https://www.kaggle.com/ragnar123/g2net-effb7-100-seed-21 ) in the comment section.\n\nIndeed, it's 1/4 of training data each epoch, but training dataset doesn't reset each epoch, so for the period of 4 epochs the model sees 4 unique pieces of training data (1/4 of training set each epoch).\n\nHence, real amount of times our model passes through training data is divided by 4 (relative to the number of epochs), but none of the training samples are getting thrown away during training.",
      "votes": null
    },
    {
      "id": "1559696",
      "postDate": "10/27/2021 07:07:03",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    },
    {
      "id": "1559985",
      "postDate": "10/27/2021 08:52:52",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1503272,
      "author_name": "vzaguskin",
      "author_url": "",
      "post_date": "09/05/2021 08:04:30",
      "content": "<p>I was also surprised when I saw <em>repeat</em> and <em>//4</em> in this kernel. Looks like only 1/4 of training data is used each epoch. Maybe this is done because of time limitation of kaggle kernel. Hopefully  the author can comment on that.</p>\n<p>If you just want to use all training data each epoch - you don't need to use <em>repeat</em>. In that case you don't have to specify <em>steps_per_epoch</em> as well. Your formula for <em>steps_per_epoch</em> is correct.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1503724,
          "author_name": "sairam6087",
          "author_url": "",
          "post_date": "09/05/2021 16:43:54",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/vzaguskin\" target=\"_blank\">@vzaguskin</a> . Using a subset of data for the kernel time limit makes sense. I will try avoiding using repeat and see what happens. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1508777,
          "author_name": "martynoveduard",
          "author_url": "",
          "post_date": "09/10/2021 14:43:51",
          "content": "<p>This has been clarified <a href=\"https://www.kaggle.com/ragnar123/g2net-effb7-100-seed-21\" target=\"_blank\">here</a> in the comment section.</p>\n<p>Indeed, it's 1/4 of training data each epoch, but training dataset doesn't reset each epoch, so for the period of 4 epochs the model sees 4 unique pieces of training data (1/4 of training set each epoch).</p>\n<p>Hence, real amount of times our model passes through training data is divided by 4 (relative to the number of epochs), but none of the training samples are getting thrown away during training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559696,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 07:07:03",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559985,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:52:52",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1502855": "Hello all,\n\nI'm more of a pytorch user but want to use Keras for this competition because there's an easy way to chain together different models using the functional API. I noticed that when doing distributed training, for example in this notebook here by @hidehisaarai1213 : https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-training , there are two things that I would like some clarity on:\n\n1. What benefit does the `dataset.repeat()` option have? I also have TFRecords and would like to use a structure similar to this notebook. \n\n2. In order to use `dataset.repeat()`, Keras requires the `steps_per_epoch` option to be specified when using `model.fit()`. \n\nNormally, you would calculate this as:\n```\nsteps_per_epoch = (num_tf_train_records * files_per_tf_record) // batch_size\n```\nIf I understand correctly, when using a TPU or multi-gpu setup, you must divide this by the number of replicas, so that formula would be\n\n```\nsteps_per_epoch = (num_tf_train_records * files_per_tf_record) // batch_size // replicas\n```\n\nHowever, I see that it has been specified as:\n\n```\nsteps_per_epoch = (num_tf_train_records * files_per_tf_record) // batch_size // replicas // 4\n```\n\nWhat is the purpose of the additional `4` there? I am trying to replicate this setup in GPU and want to accurately calculate the `steps_per_epoch` for a 2 GPU setup. Would really appreciate some insight from the Keras experts here. Thank you",
    "1503272": "I was also surprised when I saw *repeat* and *//4* in this kernel. Looks like only 1/4 of training data is used each epoch. Maybe this is done because of time limitation of kaggle kernel. Hopefully  the author can comment on that.\n\nIf you just want to use all training data each epoch - you don't need to use *repeat*. In that case you don't have to specify *steps_per_epoch* as well. Your formula for *steps_per_epoch* is correct.",
    "1503724": "Thanks @vzaguskin . Using a subset of data for the kernel time limit makes sense. I will try avoiding using repeat and see what happens. :)",
    "1508777": "This has been clarified [here](https://www.kaggle.com/ragnar123/g2net-effb7-100-seed-21 ) in the comment section.\n\nIndeed, it's 1/4 of training data each epoch, but training dataset doesn't reset each epoch, so for the period of 4 epochs the model sees 4 unique pieces of training data (1/4 of training set each epoch).\n\nHence, real amount of times our model passes through training data is divided by 4 (relative to the number of epochs), but none of the training samples are getting thrown away during training.",
    "1559696": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
    "1559985": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}