{
  "id": 124922,
  "title": "Why the validation loss don't decrease ?",
  "url": "/competitions/pku-autonomous-driving/discussion/124922",
  "author_name": "",
  "post_date": "2020-01-07T14:10:55.565100400Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<h3>How to get a lower validation loss, do you have any suggestion?</h3>\n\n<hr>\n\n<p>```\nJust after 3 epoches, the validation loss don't change - about 1.3, and train loss is about 0.6. </p>\n\n<p>Then I reduce the learning rate, validation loss decrease a little, but final mAP still be not good.</p>\n\n<p>I use the input encode way like ruslan's kernel, and initial lr is 0.01, optimizer is adamw.\n```</p>",
  "messages": [
    {
      "id": "712697",
      "postDate": "01/07/2020 14:10:55",
      "content": "<h3>How to get a lower validation loss, do you have any suggestion?</h3>\n\n<hr>\n\n<p>```\nJust after 3 epoches, the validation loss don't change - about 1.3, and train loss is about 0.6. </p>\n\n<p>Then I reduce the learning rate, validation loss decrease a little, but final mAP still be not good.</p>\n\n<p>I use the input encode way like ruslan's kernel, and initial lr is 0.01, optimizer is adamw.\n```</p>",
      "rawMarkdown": "### How to get a lower validation loss, do you have any suggestion? \n***\n\n```\nJust after 3 epoches, the validation loss don't change - about 1.3, and train loss is about 0.6. \n\nThen I reduce the learning rate, validation loss decrease a little, but final mAP still be not good.\n\nI use the input encode way like ruslan's kernel, and initial lr is 0.01, optimizer is adamw.\n```",
      "votes": null
    },
    {
      "id": "713137",
      "postDate": "01/08/2020 00:08:43",
      "content": "<p>Change loss function? I also have to face the \"mask loss\" decrease a little. At present, I try to use \"focal loss \" in CenterNet paper.\nlr = 0.004 adamW:decay = 0.0002\nepoch0: <br>\nDev loss: 3.2760\nMask loss: 1.8199\nRegr loss: 1.4561</p>\n\n<p>epoch1:\nDev loss: 3.0883\nMask loss: 1.7852\nRegr loss: 1.3032</p>\n\n<p>epoch2:\nDev loss: 3.0140\nMask loss: 1.8219\nRegr loss: 1.1921</p>\n\n<p>lr = 0.002\nepoch3:\nDev loss: 2.2477\nMask loss: 1.1969\nRegr loss: 1.0508</p>\n\n<p>epoch4:\nDev loss: 2.0596\nMask loss: 1.0325\nRegr loss: 1.0270\n...\nlr = 0.001\nepoch6:   map = 0.098\nDev loss: 1.8347\nMask loss: 0.9330\nRegr loss: 0.9017</p>\n\n<p>...Then becoming bad</p>\n\n<p>The loss decrease a little too. I try to change diffirent LR, Maybe  lr = 0.01 is a good solution?  </p>",
      "rawMarkdown": "Change loss function? I also have to face the \"mask loss\" decrease a little. At present, I try to use \"focal loss \" in CenterNet paper.\nlr = 0.004 adamW:decay = 0.0002\nepoch0:  \nDev loss: 3.2760\nMask loss: 1.8199\nRegr loss: 1.4561\n\nepoch1:\nDev loss: 3.0883\nMask loss: 1.7852\nRegr loss: 1.3032\n\nepoch2:\nDev loss: 3.0140\nMask loss: 1.8219\nRegr loss: 1.1921\n\nlr = 0.002\nepoch3:\nDev loss: 2.2477\nMask loss: 1.1969\nRegr loss: 1.0508\n\nepoch4:\nDev loss: 2.0596\nMask loss: 1.0325\nRegr loss: 1.0270\n...\nlr = 0.001\nepoch6:   map = 0.098\nDev loss: 1.8347\nMask loss: 0.9330\nRegr loss: 0.9017\n\n...Then becoming bad\n\nThe loss decrease a little too. I try to change diffirent LR, Maybe  lr = 0.01 is a good solution?",
      "votes": null
    },
    {
      "id": "713209",
      "postDate": "01/08/2020 02:54:57",
      "content": "<p>lr = 0.01 is too large for pretrained encoder. maybe around 0.0001 is good.\nthis can be irrelevant but one thing you should note is, when you use focal loss for CenterNet in heatmap prediction, increasing heatmap-loss doesn't always mean bad prediction. since for focal loss in original paper, it is ideal when only center-bit == 1 and others == 0, however trained CenterNet doesn't always return such peak in validation set, rather returns like-gaussian heatmap. and this can be fixed in post-processing like finding local peak. so you should rather care about xyz and rpy loss. (xyz and rpy is trained only on target == 1 pixels, so this converges more slowly than heatmap, i think.)\nin summary, you should keep training after total-loss increased. (until xyz and rpy loss increase.)\nLet's both do our best.</p>",
      "rawMarkdown": "lr = 0.01 is too large for pretrained encoder. maybe around 0.0001 is good.\nthis can be irrelevant but one thing you should note is, when you use focal loss for CenterNet in heatmap prediction, increasing heatmap-loss doesn't always mean bad prediction. since for focal loss in original paper, it is ideal when only center-bit == 1 and others == 0, however trained CenterNet doesn't always return such peak in validation set, rather returns like-gaussian heatmap. and this can be fixed in post-processing like finding local peak. so you should rather care about xyz and rpy loss. (xyz and rpy is trained only on target == 1 pixels, so this converges more slowly than heatmap, i think.)\nin summary, you should keep training after total-loss increased. (until xyz and rpy loss increase.)\nLet's both do our best.",
      "votes": null
    },
    {
      "id": "713616",
      "postDate": "01/08/2020 13:28:37",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/ryomak\"></a><a href=\"/ryomak\">@ryomak</a>. I did experiments with simple bce loss and \"center-bit == 1 , others == 0\" configuration, and got similar observation. The val loss first went down and then went up. The predicted values of center points became larger when the mask loss went up, for points surrounding center points increased as well. However, the surrounding points were still smaller than center points, and longer training actually increase the score.</p>",
      "rawMarkdown": "I agree with [@ryomak](https://www.kaggle.com/ryomak). I did experiments with simple bce loss and \"center-bit == 1 , others == 0\" configuration, and got similar observation. The val loss first went down and then went up. The predicted values of center points became larger when the mask loss went up, for points surrounding center points increased as well. However, the surrounding points were still smaller than center points, and longer training actually increase the score.",
      "votes": null
    },
    {
      "id": "716019",
      "postDate": "01/11/2020 06:45:33",
      "content": "<p>Thanks for your detailed explanation, it's very clear. <a href=\"/ryomak\">@ryomak</a> </p>",
      "rawMarkdown": "Thanks for your detailed explanation, it's very clear. @ryomak",
      "votes": null
    },
    {
      "id": "716026",
      "postDate": "01/11/2020 06:54:50",
      "content": "<p>thanks for your detailed information <a href=\"/guobaozi\">@guobaozi</a> </p>",
      "rawMarkdown": "thanks for your detailed information @guobaozi",
      "votes": null
    },
    {
      "id": "716028",
      "postDate": "01/11/2020 06:56:30",
      "content": "<p>thanks, I'll give it a try. <a href=\"/richardwong1994\">@richardwong1994</a> </p>",
      "rawMarkdown": "thanks, I'll give it a try. @richardwong1994",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 713137,
      "author_name": "guobaozi",
      "author_url": "",
      "post_date": "01/08/2020 00:08:43",
      "content": "<p>Change loss function? I also have to face the \"mask loss\" decrease a little. At present, I try to use \"focal loss \" in CenterNet paper.\nlr = 0.004 adamW:decay = 0.0002\nepoch0: <br>\nDev loss: 3.2760\nMask loss: 1.8199\nRegr loss: 1.4561</p>\n\n<p>epoch1:\nDev loss: 3.0883\nMask loss: 1.7852\nRegr loss: 1.3032</p>\n\n<p>epoch2:\nDev loss: 3.0140\nMask loss: 1.8219\nRegr loss: 1.1921</p>\n\n<p>lr = 0.002\nepoch3:\nDev loss: 2.2477\nMask loss: 1.1969\nRegr loss: 1.0508</p>\n\n<p>epoch4:\nDev loss: 2.0596\nMask loss: 1.0325\nRegr loss: 1.0270\n...\nlr = 0.001\nepoch6:   map = 0.098\nDev loss: 1.8347\nMask loss: 0.9330\nRegr loss: 0.9017</p>\n\n<p>...Then becoming bad</p>\n\n<p>The loss decrease a little too. I try to change diffirent LR, Maybe  lr = 0.01 is a good solution?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 716026,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "01/11/2020 06:54:50",
          "content": "<p>thanks for your detailed information <a href=\"/guobaozi\">@guobaozi</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 713209,
      "author_name": "ryomak",
      "author_url": "",
      "post_date": "01/08/2020 02:54:57",
      "content": "<p>lr = 0.01 is too large for pretrained encoder. maybe around 0.0001 is good.\nthis can be irrelevant but one thing you should note is, when you use focal loss for CenterNet in heatmap prediction, increasing heatmap-loss doesn't always mean bad prediction. since for focal loss in original paper, it is ideal when only center-bit == 1 and others == 0, however trained CenterNet doesn't always return such peak in validation set, rather returns like-gaussian heatmap. and this can be fixed in post-processing like finding local peak. so you should rather care about xyz and rpy loss. (xyz and rpy is trained only on target == 1 pixels, so this converges more slowly than heatmap, i think.)\nin summary, you should keep training after total-loss increased. (until xyz and rpy loss increase.)\nLet's both do our best.</p>",
      "votes": null,
      "replies": [
        {
          "id": 716019,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "01/11/2020 06:45:33",
          "content": "<p>Thanks for your detailed explanation, it's very clear. <a href=\"/ryomak\">@ryomak</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 713616,
      "author_name": "richardwong1994",
      "author_url": "",
      "post_date": "01/08/2020 13:28:37",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/ryomak\"></a><a href=\"/ryomak\">@ryomak</a>. I did experiments with simple bce loss and \"center-bit == 1 , others == 0\" configuration, and got similar observation. The val loss first went down and then went up. The predicted values of center points became larger when the mask loss went up, for points surrounding center points increased as well. However, the surrounding points were still smaller than center points, and longer training actually increase the score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 716028,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "01/11/2020 06:56:30",
          "content": "<p>thanks, I'll give it a try. <a href=\"/richardwong1994\">@richardwong1994</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "712697": "### How to get a lower validation loss, do you have any suggestion? \n***\n\n```\nJust after 3 epoches, the validation loss don't change - about 1.3, and train loss is about 0.6. \n\nThen I reduce the learning rate, validation loss decrease a little, but final mAP still be not good.\n\nI use the input encode way like ruslan's kernel, and initial lr is 0.01, optimizer is adamw.\n```",
    "713137": "Change loss function? I also have to face the \"mask loss\" decrease a little. At present, I try to use \"focal loss \" in CenterNet paper.\nlr = 0.004 adamW:decay = 0.0002\nepoch0:  \nDev loss: 3.2760\nMask loss: 1.8199\nRegr loss: 1.4561\n\nepoch1:\nDev loss: 3.0883\nMask loss: 1.7852\nRegr loss: 1.3032\n\nepoch2:\nDev loss: 3.0140\nMask loss: 1.8219\nRegr loss: 1.1921\n\nlr = 0.002\nepoch3:\nDev loss: 2.2477\nMask loss: 1.1969\nRegr loss: 1.0508\n\nepoch4:\nDev loss: 2.0596\nMask loss: 1.0325\nRegr loss: 1.0270\n...\nlr = 0.001\nepoch6:   map = 0.098\nDev loss: 1.8347\nMask loss: 0.9330\nRegr loss: 0.9017\n\n...Then becoming bad\n\nThe loss decrease a little too. I try to change diffirent LR, Maybe  lr = 0.01 is a good solution?",
    "713209": "lr = 0.01 is too large for pretrained encoder. maybe around 0.0001 is good.\nthis can be irrelevant but one thing you should note is, when you use focal loss for CenterNet in heatmap prediction, increasing heatmap-loss doesn't always mean bad prediction. since for focal loss in original paper, it is ideal when only center-bit == 1 and others == 0, however trained CenterNet doesn't always return such peak in validation set, rather returns like-gaussian heatmap. and this can be fixed in post-processing like finding local peak. so you should rather care about xyz and rpy loss. (xyz and rpy is trained only on target == 1 pixels, so this converges more slowly than heatmap, i think.)\nin summary, you should keep training after total-loss increased. (until xyz and rpy loss increase.)\nLet's both do our best.",
    "713616": "I agree with [@ryomak](https://www.kaggle.com/ryomak). I did experiments with simple bce loss and \"center-bit == 1 , others == 0\" configuration, and got similar observation. The val loss first went down and then went up. The predicted values of center points became larger when the mask loss went up, for points surrounding center points increased as well. However, the surrounding points were still smaller than center points, and longer training actually increase the score.",
    "716019": "Thanks for your detailed explanation, it's very clear. @ryomak",
    "716026": "thanks for your detailed information @guobaozi",
    "716028": "thanks, I'll give it a try. @richardwong1994"
  },
  "source": "meta"
}