{
  "id": 195234,
  "title": "How to avoid revisiting same data after resuming from a checkpoint?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/195234",
  "author_name": "",
  "post_date": "2020-11-04T05:53:21.929346300Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I trained my model for 12 hours and then got disconnected from Colab. I am now trying to resume training from my last checkpoint. However, after loading, the model starts training from the 0th data (And I get really low loss as the model has already trained over this before). Is there a way to ensure that after resuming from a checkpoint, the training does not happen on the same data again?</p>\n<p>Note: I am using shuffle=True in dataloader.</p>\n<p>For reference, some relevant code - </p>\n<pre><code>#Dataloader\ntrain_dataloader = DataLoader(train_dataset,\n                              shuffle=True,\n                              batch_size=16,\n                              num_workers=0)\n# Loading checkpoint\nmodel = LyftMultiModel(cfg)\nmodel.load_state_dict(torch.load(weight_path))\n\n#Training \ntr_it = iter(train_dataloader)\nprogress_bar = tqdm(range(cfg[\"train_params\"][\"max_num_steps\"]))\nfor i in progress_bar:\n    data = next(tr_it)\n    loss, _, _ = forward(data, model, device)\n</code></pre>",
  "messages": [
    {
      "id": "1069139",
      "postDate": "11/04/2020 05:53:21",
      "content": "<p>I trained my model for 12 hours and then got disconnected from Colab. I am now trying to resume training from my last checkpoint. However, after loading, the model starts training from the 0th data (And I get really low loss as the model has already trained over this before). Is there a way to ensure that after resuming from a checkpoint, the training does not happen on the same data again?</p>\n<p>Note: I am using shuffle=True in dataloader.</p>\n<p>For reference, some relevant code - </p>\n<pre><code>#Dataloader\ntrain_dataloader = DataLoader(train_dataset,\n                              shuffle=True,\n                              batch_size=16,\n                              num_workers=0)\n# Loading checkpoint\nmodel = LyftMultiModel(cfg)\nmodel.load_state_dict(torch.load(weight_path))\n\n#Training \ntr_it = iter(train_dataloader)\nprogress_bar = tqdm(range(cfg[\"train_params\"][\"max_num_steps\"]))\nfor i in progress_bar:\n    data = next(tr_it)\n    loss, _, _ = forward(data, model, device)\n</code></pre>",
      "rawMarkdown": "I trained my model for 12 hours and then got disconnected from Colab. I am now trying to resume training from my last checkpoint. However, after loading, the model starts training from the 0th data (And I get really low loss as the model has already trained over this before). Is there a way to ensure that after resuming from a checkpoint, the training does not happen on the same data again?\n\nNote: I am using shuffle=True in dataloader.\n\nFor reference, some relevant code - \n\n```\n#Dataloader\ntrain_dataloader = DataLoader(train_dataset,\n                              shuffle=True,\n                              batch_size=16,\n                              num_workers=0)\n# Loading checkpoint\nmodel = LyftMultiModel(cfg)\nmodel.load_state_dict(torch.load(weight_path))\n\n#Training \ntr_it = iter(train_dataloader)\nprogress_bar = tqdm(range(cfg[\"train_params\"][\"max_num_steps\"]))\nfor i in progress_bar:\n    data = next(tr_it)\n    loss, _, _ = forward(data, model, device)\n```",
      "votes": null
    },
    {
      "id": "1069162",
      "postDate": "11/04/2020 06:53:20",
      "content": "<p>write your own batch sampler which iterate over an input randomised index.<br>\njust reused this randomised index when you restart, begining from the last iteration number.</p>\n<pre><code>class IndexSampler(Sampler):\n    def __init__(self, index, is_shuffle=False):\n        #self.dataset = dataset\n        self.index = index\n        self.is_shuffle = is_shuffle\n        self.len = len(index)\n\n    def __iter__(self):\n        index = self.index.copy()\n        if self.is_shuffle:\n            random.shuffle(index)\n        return iter(index)\n\n    def __len__(self):\n        return self.len\n\n  index = ....\n  last_used = ....\n\n    loader  = DataLoader(\n        dataset,\n        sampler = IndexSampler(index[last_used:]),\n        batch_size = 16,\n        drop_last  = False,\n        num_workers = 8,\n        pin_memory  = True,\n    )\n</code></pre>",
      "rawMarkdown": "write your own batch sampler which iterate over an input randomised index.\njust reused this randomised index when you restart, begining from the last iteration number.\n\n\n```\nclass IndexSampler(Sampler):\n    def __init__(self, index, is_shuffle=False):\n        #self.dataset = dataset\n        self.index = index\n        self.is_shuffle = is_shuffle\n        self.len = len(index)\n\n    def __iter__(self):\n        index = self.index.copy()\n        if self.is_shuffle:\n            random.shuffle(index)\n        return iter(index)\n\n    def __len__(self):\n        return self.len\n\n  index = ....\n  last_used = ....\n\n    loader  = DataLoader(\n        dataset,\n        sampler = IndexSampler(index[last_used:]),\n        batch_size = 16,\n        drop_last  = False,\n        num_workers = 8,\n        pin_memory  = True,\n    )\n\n\n```",
      "votes": null
    },
    {
      "id": "1069175",
      "postDate": "11/04/2020 07:09:26",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>! I had the following question - </p>\n<ol>\n<li>Here, <strong>index</strong> is a list of integers having values from 0 to len(AgentDataset) and has been already randomized, right?</li>\n<li>If the above is true, that is the reason why you've kept is_shuffle=False as the list is already randomized, right?</li>\n</ol>",
      "rawMarkdown": "Thanks @hengck23! I had the following question - \n1. Here, **index** is a list of integers having values from 0 to len(AgentDataset) and has been already randomized, right?\n2. If the above is true, that is the reason why you've kept is_shuffle=False as the list is already randomized, right?",
      "votes": null
    },
    {
      "id": "1069262",
      "postDate": "11/04/2020 09:10:10",
      "content": "<p>it is up to you to design your sampler. my code is only an example.</p>\n<p>in pytorch, data loader will call sampler to produce a batch of index and each of this index is passed to __get_item__() function of the dataset</p>\n<p>another way is to overwrite the __get_item__() function of the dataset, so that it will not use certain data.</p>",
      "rawMarkdown": "it is up to you to design your sampler. my code is only an example.\n\nin pytorch, data loader will call sampler to produce a batch of index and each of this index is passed to \\__get\\_item\\__() function of the dataset\n\nanother way is to overwrite the \\__get\\_item\\__() function of the dataset, so that it will not use certain data.",
      "votes": null
    },
    {
      "id": "1069362",
      "postDate": "11/04/2020 10:55:07",
      "content": "<p>If you are using <code>shuffle=True</code>, you will be using the <a href=\"https://pytorch.org/docs/stable/_modules/torch/utils/data/sampler.html#RandomSampler\" target=\"_blank\">RandomSampler</a> by default with replacement=False.</p>\n<p>An easier workaround will be just set the pytorch random seed to different values rather than a fixed one for each round. <br>\n<code>torch.manual_seed(42 + n_round)</code> Though it won't guarantee no repeat but at least you won't train on the same set every time.</p>",
      "rawMarkdown": "If you are using `shuffle=True`, you will be using the [RandomSampler](https://pytorch.org/docs/stable/_modules/torch/utils/data/sampler.html#RandomSampler) by default with replacement=False.\n\nAn easier workaround will be just set the pytorch random seed to different values rather than a fixed one for each round. \n`torch.manual_seed(42 + n_round)` Though it won't guarantee no repeat but at least you won't train on the same set every time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1069162,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/04/2020 06:53:20",
      "content": "<p>write your own batch sampler which iterate over an input randomised index.<br>\njust reused this randomised index when you restart, begining from the last iteration number.</p>\n<pre><code>class IndexSampler(Sampler):\n    def __init__(self, index, is_shuffle=False):\n        #self.dataset = dataset\n        self.index = index\n        self.is_shuffle = is_shuffle\n        self.len = len(index)\n\n    def __iter__(self):\n        index = self.index.copy()\n        if self.is_shuffle:\n            random.shuffle(index)\n        return iter(index)\n\n    def __len__(self):\n        return self.len\n\n  index = ....\n  last_used = ....\n\n    loader  = DataLoader(\n        dataset,\n        sampler = IndexSampler(index[last_used:]),\n        batch_size = 16,\n        drop_last  = False,\n        num_workers = 8,\n        pin_memory  = True,\n    )\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1069175,
          "author_name": "arpitrf",
          "author_url": "",
          "post_date": "11/04/2020 07:09:26",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>! I had the following question - </p>\n<ol>\n<li>Here, <strong>index</strong> is a list of integers having values from 0 to len(AgentDataset) and has been already randomized, right?</li>\n<li>If the above is true, that is the reason why you've kept is_shuffle=False as the list is already randomized, right?</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069262,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/04/2020 09:10:10",
          "content": "<p>it is up to you to design your sampler. my code is only an example.</p>\n<p>in pytorch, data loader will call sampler to produce a batch of index and each of this index is passed to __get_item__() function of the dataset</p>\n<p>another way is to overwrite the __get_item__() function of the dataset, so that it will not use certain data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1069362,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/04/2020 10:55:07",
      "content": "<p>If you are using <code>shuffle=True</code>, you will be using the <a href=\"https://pytorch.org/docs/stable/_modules/torch/utils/data/sampler.html#RandomSampler\" target=\"_blank\">RandomSampler</a> by default with replacement=False.</p>\n<p>An easier workaround will be just set the pytorch random seed to different values rather than a fixed one for each round. <br>\n<code>torch.manual_seed(42 + n_round)</code> Though it won't guarantee no repeat but at least you won't train on the same set every time.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1069139": "I trained my model for 12 hours and then got disconnected from Colab. I am now trying to resume training from my last checkpoint. However, after loading, the model starts training from the 0th data (And I get really low loss as the model has already trained over this before). Is there a way to ensure that after resuming from a checkpoint, the training does not happen on the same data again?\n\nNote: I am using shuffle=True in dataloader.\n\nFor reference, some relevant code - \n\n```\n#Dataloader\ntrain_dataloader = DataLoader(train_dataset,\n                              shuffle=True,\n                              batch_size=16,\n                              num_workers=0)\n# Loading checkpoint\nmodel = LyftMultiModel(cfg)\nmodel.load_state_dict(torch.load(weight_path))\n\n#Training \ntr_it = iter(train_dataloader)\nprogress_bar = tqdm(range(cfg[\"train_params\"][\"max_num_steps\"]))\nfor i in progress_bar:\n    data = next(tr_it)\n    loss, _, _ = forward(data, model, device)\n```",
    "1069162": "write your own batch sampler which iterate over an input randomised index.\njust reused this randomised index when you restart, begining from the last iteration number.\n\n\n```\nclass IndexSampler(Sampler):\n    def __init__(self, index, is_shuffle=False):\n        #self.dataset = dataset\n        self.index = index\n        self.is_shuffle = is_shuffle\n        self.len = len(index)\n\n    def __iter__(self):\n        index = self.index.copy()\n        if self.is_shuffle:\n            random.shuffle(index)\n        return iter(index)\n\n    def __len__(self):\n        return self.len\n\n  index = ....\n  last_used = ....\n\n    loader  = DataLoader(\n        dataset,\n        sampler = IndexSampler(index[last_used:]),\n        batch_size = 16,\n        drop_last  = False,\n        num_workers = 8,\n        pin_memory  = True,\n    )\n\n\n```",
    "1069175": "Thanks @hengck23! I had the following question - \n1. Here, **index** is a list of integers having values from 0 to len(AgentDataset) and has been already randomized, right?\n2. If the above is true, that is the reason why you've kept is_shuffle=False as the list is already randomized, right?",
    "1069262": "it is up to you to design your sampler. my code is only an example.\n\nin pytorch, data loader will call sampler to produce a batch of index and each of this index is passed to \\__get\\_item\\__() function of the dataset\n\nanother way is to overwrite the \\__get\\_item\\__() function of the dataset, so that it will not use certain data.",
    "1069362": "If you are using `shuffle=True`, you will be using the [RandomSampler](https://pytorch.org/docs/stable/_modules/torch/utils/data/sampler.html#RandomSampler) by default with replacement=False.\n\nAn easier workaround will be just set the pytorch random seed to different values rather than a fixed one for each round. \n`torch.manual_seed(42 + n_round)` Though it won't guarantee no repeat but at least you won't train on the same set every time."
  },
  "source": "meta"
}