{
  "id": 135132,
  "title": "Be care about 'shuffle'!",
  "url": "/competitions/bengaliai-cv19/discussion/135132",
  "author_name": "",
  "post_date": "2020-03-12T06:58:54.926180300Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I changed the image size from 128 to 224(use original image to resize to 224), but I used the weights of 128 and start to train. <br>\nThe first 5 epochs behavior well and I jump to 0.97/0.98(lb/cv a big gap...). <br>\nHowever, when I save the weights(first 5) and used them for further 5 epochs(second 5) training. I found the first 3 epochs' behaviors are weird. And after I finished 5 epochs, I save them again and start another 5 epochs(third 5), the first 3 weird again.. Is it normal? It likes an endless loop...  </p>\n\n<hr>\n\n<p>second 5 epochs\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F31c8b3c3e5d4ab6e4800ee5ac69514f7%2Fsecond_5.png?generation=1583996022721880&amp;alt=media\" alt=\"\"></p>\n\n<hr>\n\n<p>third 5 epochs\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2Fff091d1dbd0d258bfa6bae865855e93e%2Fthird_5.png?generation=1583997735698651&amp;alt=media\" alt=\"\"></p>\n\n<p>It might be the small lr for early stage? I use Adam and lr is 3e-4. (The random seed is fixed)\nWhat's more, the loss whthin a epoch is shake, not from high to low...\n<strong>UPDATE 1:</strong> I use OneCycleScheduler(fastai fit one cycle), after I change the div_factor to default(25) the loss whthin first epoch is from high to low... But after, the loss is increasing again and the metric decreased... And behavior like the loop which I mentioned above. <br>\n<strong>UPDATE 2:</strong> I'm ready to chang fit one cycle to fit and add ReduceLROnPlateau, but now I was restricted by colab again so feedback might be tomorrow...</p>",
  "messages": [
    {
      "id": "769709",
      "postDate": "03/12/2020 06:58:54",
      "content": "<p>I changed the image size from 128 to 224(use original image to resize to 224), but I used the weights of 128 and start to train. <br>\nThe first 5 epochs behavior well and I jump to 0.97/0.98(lb/cv a big gap...). <br>\nHowever, when I save the weights(first 5) and used them for further 5 epochs(second 5) training. I found the first 3 epochs' behaviors are weird. And after I finished 5 epochs, I save them again and start another 5 epochs(third 5), the first 3 weird again.. Is it normal? It likes an endless loop...  </p>\n\n<hr>\n\n<p>second 5 epochs\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F31c8b3c3e5d4ab6e4800ee5ac69514f7%2Fsecond_5.png?generation=1583996022721880&amp;alt=media\" alt=\"\"></p>\n\n<hr>\n\n<p>third 5 epochs\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2Fff091d1dbd0d258bfa6bae865855e93e%2Fthird_5.png?generation=1583997735698651&amp;alt=media\" alt=\"\"></p>\n\n<p>It might be the small lr for early stage? I use Adam and lr is 3e-4. (The random seed is fixed)\nWhat's more, the loss whthin a epoch is shake, not from high to low...\n<strong>UPDATE 1:</strong> I use OneCycleScheduler(fastai fit one cycle), after I change the div_factor to default(25) the loss whthin first epoch is from high to low... But after, the loss is increasing again and the metric decreased... And behavior like the loop which I mentioned above. <br>\n<strong>UPDATE 2:</strong> I'm ready to chang fit one cycle to fit and add ReduceLROnPlateau, but now I was restricted by colab again so feedback might be tomorrow...</p>",
      "rawMarkdown": "I changed the image size from 128 to 224(use original image to resize to 224), but I used the weights of 128 and start to train.  \nThe first 5 epochs behavior well and I jump to 0.97/0.98(lb/cv a big gap...).  \nHowever, when I save the weights(first 5) and used them for further 5 epochs(second 5) training. I found the first 3 epochs' behaviors are weird. And after I finished 5 epochs, I save them again and start another 5 epochs(third 5), the first 3 weird again.. Is it normal? It likes an endless loop...  \n\n***\nsecond 5 epochs\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F31c8b3c3e5d4ab6e4800ee5ac69514f7%2Fsecond_5.png?generation=1583996022721880&amp;alt=media)\n***\nthird 5 epochs\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2Fff091d1dbd0d258bfa6bae865855e93e%2Fthird_5.png?generation=1583997735698651&amp;alt=media)\n\nIt might be the small lr for early stage? I use Adam and lr is 3e-4. (The random seed is fixed)\nWhat's more, the loss whthin a epoch is shake, not from high to low...\n**UPDATE 1:** I use OneCycleScheduler(fastai fit one cycle), after I change the div_factor to default(25) the loss whthin first epoch is from high to low... But after, the loss is increasing again and the metric decreased... And behavior like the loop which I mentioned above.  \n**UPDATE 2:** I'm ready to chang fit one cycle to fit and add ReduceLROnPlateau, but now I was restricted by colab again so feedback might be tomorrow...",
      "votes": null
    },
    {
      "id": "770517",
      "postDate": "03/13/2020 03:25:32",
      "content": "<p><strong>Result 1:</strong> <br>\n- ReduceLROnPlateau will cause the same suitation(first 5 epochs), if the second 5 epochs behavior the same, I will reduce the batch size from 128 to 64 as an answer from a blog. <br>\n- What's more, the ReduceLROnPlateau will make train loss decrease but valid loss shake...</p>",
      "rawMarkdown": "**Result 1:**  \n- ReduceLROnPlateau will cause the same suitation(first 5 epochs), if the second 5 epochs behavior the same, I will reduce the batch size from 128 to 64 as an answer from a blog.  \n- What's more, the ReduceLROnPlateau will make train loss decrease but valid loss shake...",
      "votes": null
    },
    {
      "id": "770592",
      "postDate": "03/13/2020 06:06:58",
      "content": "<p><strong>Solution:</strong> <br>\nJust before lunch, I found I set the shuffle=False in train dataloader by mistake and I fixed it then and start to train. Till now, everything works normally. <br>\nSorry for disturbing the discussion area. <br>\nIt might be a ridiculous mistake... ;)</p>",
      "rawMarkdown": "**Solution:**  \nJust before lunch, I found I set the shuffle=False in train dataloader by mistake and I fixed it then and start to train. Till now, everything works normally.  \nSorry for disturbing the discussion area.  \nIt might be a ridiculous mistake... ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 770517,
      "author_name": "cnzengshiyuan",
      "author_url": "",
      "post_date": "03/13/2020 03:25:32",
      "content": "<p><strong>Result 1:</strong> <br>\n- ReduceLROnPlateau will cause the same suitation(first 5 epochs), if the second 5 epochs behavior the same, I will reduce the batch size from 128 to 64 as an answer from a blog. <br>\n- What's more, the ReduceLROnPlateau will make train loss decrease but valid loss shake...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 770592,
      "author_name": "cnzengshiyuan",
      "author_url": "",
      "post_date": "03/13/2020 06:06:58",
      "content": "<p><strong>Solution:</strong> <br>\nJust before lunch, I found I set the shuffle=False in train dataloader by mistake and I fixed it then and start to train. Till now, everything works normally. <br>\nSorry for disturbing the discussion area. <br>\nIt might be a ridiculous mistake... ;)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "769709": "I changed the image size from 128 to 224(use original image to resize to 224), but I used the weights of 128 and start to train.  \nThe first 5 epochs behavior well and I jump to 0.97/0.98(lb/cv a big gap...).  \nHowever, when I save the weights(first 5) and used them for further 5 epochs(second 5) training. I found the first 3 epochs' behaviors are weird. And after I finished 5 epochs, I save them again and start another 5 epochs(third 5), the first 3 weird again.. Is it normal? It likes an endless loop...  \n\n***\nsecond 5 epochs\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2F31c8b3c3e5d4ab6e4800ee5ac69514f7%2Fsecond_5.png?generation=1583996022721880&amp;alt=media)\n***\nthird 5 epochs\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3149580%2Fff091d1dbd0d258bfa6bae865855e93e%2Fthird_5.png?generation=1583997735698651&amp;alt=media)\n\nIt might be the small lr for early stage? I use Adam and lr is 3e-4. (The random seed is fixed)\nWhat's more, the loss whthin a epoch is shake, not from high to low...\n**UPDATE 1:** I use OneCycleScheduler(fastai fit one cycle), after I change the div_factor to default(25) the loss whthin first epoch is from high to low... But after, the loss is increasing again and the metric decreased... And behavior like the loop which I mentioned above.  \n**UPDATE 2:** I'm ready to chang fit one cycle to fit and add ReduceLROnPlateau, but now I was restricted by colab again so feedback might be tomorrow...",
    "770517": "**Result 1:**  \n- ReduceLROnPlateau will cause the same suitation(first 5 epochs), if the second 5 epochs behavior the same, I will reduce the batch size from 128 to 64 as an answer from a blog.  \n- What's more, the ReduceLROnPlateau will make train loss decrease but valid loss shake...",
    "770592": "**Solution:**  \nJust before lunch, I found I set the shuffle=False in train dataloader by mistake and I fixed it then and start to train. Till now, everything works normally.  \nSorry for disturbing the discussion area.  \nIt might be a ridiculous mistake... ;)"
  },
  "source": "meta"
}