{
  "id": 126526,
  "title": "Dev-loss  steady decreased but training-loss fluctuated..why?",
  "url": "/competitions/pku-autonomous-driving/discussion/126526",
  "author_name": "",
  "post_date": "2020-01-18T04:44:00.472435300Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>My training loss has fluctuated widely. I assume this has happened because I have used dropout(0.7) twice in my Conv layers. Am i thinking right? I would appreciate your insights.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2718818%2F68b5a2ba9b86971239b22f986a043519%2Fpasted%20image%200%20(2\" alt=\"\">.png?generation=1579322623524904&amp;alt=media)</p>",
  "messages": [
    {
      "id": "722090",
      "postDate": "01/18/2020 04:44:00",
      "content": "<p>My training loss has fluctuated widely. I assume this has happened because I have used dropout(0.7) twice in my Conv layers. Am i thinking right? I would appreciate your insights.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2718818%2F68b5a2ba9b86971239b22f986a043519%2Fpasted%20image%200%20(2\" alt=\"\">.png?generation=1579322623524904&amp;alt=media)</p>",
      "rawMarkdown": "My training loss has fluctuated widely. I assume this has happened because I have used dropout(0.7) twice in my Conv layers. Am i thinking right? I would appreciate your insights.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2718818%2F68b5a2ba9b86971239b22f986a043519%2Fpasted%20image%200%20(2).png?generation=1579322623524904&amp;alt=media)",
      "votes": null
    },
    {
      "id": "722129",
      "postDate": "01/18/2020 06:23:08",
      "content": "<p>There can be several reasons. It can be related to the dataset and/or loss function. I think it's normal when we use a small batch size for training. However, if you look closely, fluctuation of loss is actually decreasing and if you smooth it, it will show a decreasing pattern. Dropout may not be directly related to the fluctuation of loss, however, higher dropout may result in overall higher loss and underfitting of a model. Another reason for the fluctuating loss is by using a high learning rate, but I see your learning rate is reasonably small. So LR might not be the reason here. </p>",
      "rawMarkdown": "There can be several reasons. It can be related to the dataset and/or loss function. I think it's normal when we use a small batch size for training. However, if you look closely, fluctuation of loss is actually decreasing and if you smooth it, it will show a decreasing pattern. Dropout may not be directly related to the fluctuation of loss, however, higher dropout may result in overall higher loss and underfitting of a model. Another reason for the fluctuating loss is by using a high learning rate, but I see your learning rate is reasonably small. So LR might not be the reason here.",
      "votes": null
    },
    {
      "id": "722162",
      "postDate": "01/18/2020 07:30:23",
      "content": "<p>i think the main reason here is small batch size and batchnormalization,,i faced similar issue here! his model will start diverging after few more epoches like it is happening here with me!</p>",
      "rawMarkdown": "i think the main reason here is small batch size and batchnormalization,,i faced similar issue here! his model will start diverging after few more epoches like it is happening here with me!",
      "votes": null
    },
    {
      "id": "722182",
      "postDate": "01/18/2020 08:24:27",
      "content": "<p>I think a simpler explanation could be if you are using hocop or my kernel then <strong>train_loss are not averaged</strong> across epoch at the end . But **Val_Loss is averaged for the full epoch **. Therefore the train_loss you see at the end is the loss for last   batch (if batch-size =1 , its the loss for one image)  alone . Which could vary greatly based on how good or how bad that particular image is and not representative of the performance of complete epoch.  </p>",
      "rawMarkdown": "I think a simpler explanation could be if you are using hocop or my kernel then **train_loss are not averaged** across epoch at the end . But **Val_Loss is averaged for the full epoch **. Therefore the train_loss you see at the end is the loss for last   batch (if batch-size =1 , its the loss for one image)  alone . Which could vary greatly based on how good or how bad that particular image is and not representative of the performance of complete epoch.",
      "votes": null
    },
    {
      "id": "722561",
      "postDate": "01/18/2020 18:04:34",
      "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> Thank you that makes lots of sense.</p>",
      "rawMarkdown": "phoenix9032 Thank you that makes lots of sense.",
      "votes": null
    },
    {
      "id": "722563",
      "postDate": "01/18/2020 18:06:06",
      "content": "<p><a href=\"/sgalib\">@sgalib</a> thank you for your very insightful comment.</p>",
      "rawMarkdown": "sgalib thank you for your very insightful comment.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 722129,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "01/18/2020 06:23:08",
      "content": "<p>There can be several reasons. It can be related to the dataset and/or loss function. I think it's normal when we use a small batch size for training. However, if you look closely, fluctuation of loss is actually decreasing and if you smooth it, it will show a decreasing pattern. Dropout may not be directly related to the fluctuation of loss, however, higher dropout may result in overall higher loss and underfitting of a model. Another reason for the fluctuating loss is by using a high learning rate, but I see your learning rate is reasonably small. So LR might not be the reason here. </p>",
      "votes": null,
      "replies": [
        {
          "id": 722162,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/18/2020 07:30:23",
          "content": "<p>i think the main reason here is small batch size and batchnormalization,,i faced similar issue here! his model will start diverging after few more epoches like it is happening here with me!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722563,
          "author_name": "subediaarjun",
          "author_url": "",
          "post_date": "01/18/2020 18:06:06",
          "content": "<p><a href=\"/sgalib\">@sgalib</a> thank you for your very insightful comment.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 722182,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "01/18/2020 08:24:27",
      "content": "<p>I think a simpler explanation could be if you are using hocop or my kernel then <strong>train_loss are not averaged</strong> across epoch at the end . But **Val_Loss is averaged for the full epoch **. Therefore the train_loss you see at the end is the loss for last   batch (if batch-size =1 , its the loss for one image)  alone . Which could vary greatly based on how good or how bad that particular image is and not representative of the performance of complete epoch.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 722561,
          "author_name": "subediaarjun",
          "author_url": "",
          "post_date": "01/18/2020 18:04:34",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> Thank you that makes lots of sense.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "722090": "My training loss has fluctuated widely. I assume this has happened because I have used dropout(0.7) twice in my Conv layers. Am i thinking right? I would appreciate your insights.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2718818%2F68b5a2ba9b86971239b22f986a043519%2Fpasted%20image%200%20(2).png?generation=1579322623524904&amp;alt=media)",
    "722129": "There can be several reasons. It can be related to the dataset and/or loss function. I think it's normal when we use a small batch size for training. However, if you look closely, fluctuation of loss is actually decreasing and if you smooth it, it will show a decreasing pattern. Dropout may not be directly related to the fluctuation of loss, however, higher dropout may result in overall higher loss and underfitting of a model. Another reason for the fluctuating loss is by using a high learning rate, but I see your learning rate is reasonably small. So LR might not be the reason here.",
    "722162": "i think the main reason here is small batch size and batchnormalization,,i faced similar issue here! his model will start diverging after few more epoches like it is happening here with me!",
    "722182": "I think a simpler explanation could be if you are using hocop or my kernel then **train_loss are not averaged** across epoch at the end . But **Val_Loss is averaged for the full epoch **. Therefore the train_loss you see at the end is the loss for last   batch (if batch-size =1 , its the loss for one image)  alone . Which could vary greatly based on how good or how bad that particular image is and not representative of the performance of complete epoch.",
    "722561": "phoenix9032 Thank you that makes lots of sense.",
    "722563": "sgalib thank you for your very insightful comment."
  },
  "source": "meta"
}