{
  "id": 220968,
  "title": "How to train all data and get checkpoints?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/220968",
  "author_name": "",
  "post_date": "2021-02-20T09:57:11.314159600Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>If I don’t want to split the data, how do I usually choose epochs? Or is it better to split the data than not?</p>",
  "messages": [
    {
      "id": "1211521",
      "postDate": "02/20/2021 09:57:11",
      "content": "<p>If I don’t want to split the data, how do I usually choose epochs? Or is it better to split the data than not?</p>",
      "rawMarkdown": "If I don’t want to split the data, how do I usually choose epochs? Or is it better to split the data than not?",
      "votes": null
    },
    {
      "id": "1211631",
      "postDate": "02/20/2021 11:41:19",
      "content": "<p>I'm not sure about your question.</p>\n<p>I think that you'd like to train with all the data, but instead of performing a single checkpoint at the end of the epoch, you'd like to have several checkpoint.</p>\n<p>In this case, you could modify your train loop to save a checkpoint every N batches.</p>\n<p>Another option is to split your dataset in N datasets. torch.utils.data comes with a random_split function for this purpose</p>\n<pre><code>DIVIDE_EPOCH_BY=10\n\nreduced_epoch_split_sizes=[len(train)//DIVIDE_EPOCH_BY]*(DIVIDE_EPOCH_BY-1)\nreduced_epoch_split_sizes+=[len(train)-np.sum(reduced_epoch_split_sizes)]\n\nreduced_epoch_train_datasets=torch.utils.data.random_split(train_dataset, reduced_epoch_split_sizes)\n\nfor i, ... in enumerate (...):\n    reduced_train_dataset=reduced_epoch_train_datasets[i%DIVIDE_EPOCH_BY]\n\n    if i%DIVIDE_EPOCH_BY == DIVIDE_EPOCH_BY-1:\n        # new random splits\n        reduced_epoch_train_datasets=torch.utils.data.random_split(train_dataset, reduced_epoch_split_sizes)\n</code></pre>",
      "rawMarkdown": "I'm not sure about your question.\n\nI think that you'd like to train with all the data, but instead of performing a single checkpoint at the end of the epoch, you'd like to have several checkpoint.\n\nIn this case, you could modify your train loop to save a checkpoint every N batches.\n\nAnother option is to split your dataset in N datasets. torch.utils.data comes with a random_split function for this purpose\n```\n\nDIVIDE_EPOCH_BY=10\n\nreduced_epoch_split_sizes=[len(train)//DIVIDE_EPOCH_BY]*(DIVIDE_EPOCH_BY-1)\nreduced_epoch_split_sizes+=[len(train)-np.sum(reduced_epoch_split_sizes)]\n\nreduced_epoch_train_datasets=torch.utils.data.random_split(train_dataset, reduced_epoch_split_sizes)\n\nfor i, ... in enumerate (...):\n\treduced_train_dataset=reduced_epoch_train_datasets[i%DIVIDE_EPOCH_BY]\n\n\tif i%DIVIDE_EPOCH_BY == DIVIDE_EPOCH_BY-1:\n\t\t# new random splits\n\t\treduced_epoch_train_datasets=torch.utils.data.random_split(train_dataset, reduced_epoch_split_sizes)\n```",
      "votes": null
    },
    {
      "id": "1211648",
      "postDate": "02/20/2021 12:05:22",
      "content": "<p>I don't want to split the data but I don't know when to stop training to avoid overfitting.</p>",
      "rawMarkdown": "I don't want to split the data but I don't know when to stop training to avoid overfitting.",
      "votes": null
    },
    {
      "id": "1211666",
      "postDate": "02/20/2021 12:27:47",
      "content": "<p>How about early stopping using a validation set?  And then check submit the results to Kaggle to see how does it fit in the public test set.</p>\n<p>Apart from that, you can measure and compare the error in validation set (green line) vs the error in train set:</p>\n<p><img src=\"https://i.stack.imgur.com/rpqa6.jpg\" alt=\"image from https://stats.stackexchange.com/questions/292283/general-question-regarding-over-fitting-vs-complexity-of-models\"></p>\n<p>Image from <a href=\"https://stats.stackexchange.com/questions/292283/general-question-regarding-over-fitting-vs-complexity-of-models\" target=\"_blank\">https://stats.stackexchange.com/questions/292283/general-question-regarding-over-fitting-vs-complexity-of-models</a></p>",
      "rawMarkdown": "How about early stopping using a validation set?  And then check submit the results to Kaggle to see how does it fit in the public test set.\n\nApart from that, you can measure and compare the error in validation set (green line) vs the error in train set:\n\n![image from https://stats.stackexchange.com/questions/292283/general-question-regarding-over-fitting-vs-complexity-of-models](https://i.stack.imgur.com/rpqa6.jpg)\n\nImage from https://stats.stackexchange.com/questions/292283/general-question-regarding-over-fitting-vs-complexity-of-models",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1211631,
      "author_name": "virilo",
      "author_url": "",
      "post_date": "02/20/2021 11:41:19",
      "content": "<p>I'm not sure about your question.</p>\n<p>I think that you'd like to train with all the data, but instead of performing a single checkpoint at the end of the epoch, you'd like to have several checkpoint.</p>\n<p>In this case, you could modify your train loop to save a checkpoint every N batches.</p>\n<p>Another option is to split your dataset in N datasets. torch.utils.data comes with a random_split function for this purpose</p>\n<pre><code>DIVIDE_EPOCH_BY=10\n\nreduced_epoch_split_sizes=[len(train)//DIVIDE_EPOCH_BY]*(DIVIDE_EPOCH_BY-1)\nreduced_epoch_split_sizes+=[len(train)-np.sum(reduced_epoch_split_sizes)]\n\nreduced_epoch_train_datasets=torch.utils.data.random_split(train_dataset, reduced_epoch_split_sizes)\n\nfor i, ... in enumerate (...):\n    reduced_train_dataset=reduced_epoch_train_datasets[i%DIVIDE_EPOCH_BY]\n\n    if i%DIVIDE_EPOCH_BY == DIVIDE_EPOCH_BY-1:\n        # new random splits\n        reduced_epoch_train_datasets=torch.utils.data.random_split(train_dataset, reduced_epoch_split_sizes)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1211648,
          "author_name": "h053473666",
          "author_url": "",
          "post_date": "02/20/2021 12:05:22",
          "content": "<p>I don't want to split the data but I don't know when to stop training to avoid overfitting.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1211666,
              "author_name": "virilo",
              "author_url": "",
              "post_date": "02/20/2021 12:27:47",
              "content": "<p>How about early stopping using a validation set?  And then check submit the results to Kaggle to see how does it fit in the public test set.</p>\n<p>Apart from that, you can measure and compare the error in validation set (green line) vs the error in train set:</p>\n<p><img src=\"https://i.stack.imgur.com/rpqa6.jpg\" alt=\"image from https://stats.stackexchange.com/questions/292283/general-question-regarding-over-fitting-vs-complexity-of-models\"></p>\n<p>Image from <a href=\"https://stats.stackexchange.com/questions/292283/general-question-regarding-over-fitting-vs-complexity-of-models\" target=\"_blank\">https://stats.stackexchange.com/questions/292283/general-question-regarding-over-fitting-vs-complexity-of-models</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1211521": "If I don’t want to split the data, how do I usually choose epochs? Or is it better to split the data than not?",
    "1211631": "I'm not sure about your question.\n\nI think that you'd like to train with all the data, but instead of performing a single checkpoint at the end of the epoch, you'd like to have several checkpoint.\n\nIn this case, you could modify your train loop to save a checkpoint every N batches.\n\nAnother option is to split your dataset in N datasets. torch.utils.data comes with a random_split function for this purpose\n```\n\nDIVIDE_EPOCH_BY=10\n\nreduced_epoch_split_sizes=[len(train)//DIVIDE_EPOCH_BY]*(DIVIDE_EPOCH_BY-1)\nreduced_epoch_split_sizes+=[len(train)-np.sum(reduced_epoch_split_sizes)]\n\nreduced_epoch_train_datasets=torch.utils.data.random_split(train_dataset, reduced_epoch_split_sizes)\n\nfor i, ... in enumerate (...):\n\treduced_train_dataset=reduced_epoch_train_datasets[i%DIVIDE_EPOCH_BY]\n\n\tif i%DIVIDE_EPOCH_BY == DIVIDE_EPOCH_BY-1:\n\t\t# new random splits\n\t\treduced_epoch_train_datasets=torch.utils.data.random_split(train_dataset, reduced_epoch_split_sizes)\n```",
    "1211648": "I don't want to split the data but I don't know when to stop training to avoid overfitting.",
    "1211666": "How about early stopping using a validation set?  And then check submit the results to Kaggle to see how does it fit in the public test set.\n\nApart from that, you can measure and compare the error in validation set (green line) vs the error in train set:\n\n![image from https://stats.stackexchange.com/questions/292283/general-question-regarding-over-fitting-vs-complexity-of-models](https://i.stack.imgur.com/rpqa6.jpg)\n\nImage from https://stats.stackexchange.com/questions/292283/general-question-regarding-over-fitting-vs-complexity-of-models"
  },
  "source": "meta"
}