{
  "id": 115673,
  "title": "Optimizing multiple losses",
  "url": "/competitions/pku-autonomous-driving/discussion/115673",
  "author_name": "",
  "post_date": "2019-11-04T14:18:38.897857700Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all, do you have any experience of optimizing multiple losses?</p>\n\n<p>For example, in <a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">https://www.kaggle.com/hocop1/centernet-baseline</a>, the centernet consists of mask and regression losses. From my observation, the mask loss will drop for a few epoches and then keep increasing (after epoch 5 or 7) while the regression loss can be minimized after epoch 5 or 7.</p>\n\n<p>Since the magnitude of mask loss (24) is larger than that of regression loss (e.g. 0.24), the mask loss dominates, which will result in early termination of the optimization process for centernet.  Do you have any experience of minimizing multiple losses at the same time?</p>\n\n<p>If this is not clear enough, please let me know and I will try my best to clarify. </p>",
  "messages": [
    {
      "id": "665006",
      "postDate": "11/04/2019 14:18:38",
      "content": "<p>Hi all, do you have any experience of optimizing multiple losses?</p>\n\n<p>For example, in <a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">https://www.kaggle.com/hocop1/centernet-baseline</a>, the centernet consists of mask and regression losses. From my observation, the mask loss will drop for a few epoches and then keep increasing (after epoch 5 or 7) while the regression loss can be minimized after epoch 5 or 7.</p>\n\n<p>Since the magnitude of mask loss (24) is larger than that of regression loss (e.g. 0.24), the mask loss dominates, which will result in early termination of the optimization process for centernet.  Do you have any experience of minimizing multiple losses at the same time?</p>\n\n<p>If this is not clear enough, please let me know and I will try my best to clarify. </p>",
      "rawMarkdown": "Hi all, do you have any experience of optimizing multiple losses?\n\nFor example, in https://www.kaggle.com/hocop1/centernet-baseline, the centernet consists of mask and regression losses. From my observation, the mask loss will drop for a few epoches and then keep increasing (after epoch 5 or 7) while the regression loss can be minimized after epoch 5 or 7.\n\nSince the magnitude of mask loss (24) is larger than that of regression loss (e.g. 0.24), the mask loss dominates, which will result in early termination of the optimization process for centernet.  Do you have any experience of minimizing multiple losses at the same time?\n\nIf this is not clear enough, please let me know and I will try my best to clarify.",
      "votes": null
    },
    {
      "id": "665056",
      "postDate": "11/04/2019 15:29:55",
      "content": "<p>Loss weights would help I think.</p>",
      "rawMarkdown": "Loss weights would help I think.",
      "votes": null
    },
    {
      "id": "678150",
      "postDate": "11/21/2019 03:08:36",
      "content": "<p>What about doing a log(maskloss) to reduce the scale ? .. I dont know if thats a thing .</p>",
      "rawMarkdown": "What about doing a log(maskloss) to reduce the scale ? .. I dont know if thats a thing .",
      "votes": null
    },
    {
      "id": "678239",
      "postDate": "11/21/2019 06:18:05",
      "content": "<p>I was just skimming through this , and found this paper </p>\n\n<p>\"Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics\"</p>\n\n<p><a href=\"https://arxiv.org/abs/1705.07115\">https://arxiv.org/abs/1705.07115</a></p>\n\n<p>Please see if its of any use . </p>\n\n<p>The paper specifically focuses on optimization for multitask loss .</p>",
      "rawMarkdown": "I was just skimming through this , and found this paper \n\n\"Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics\"\n\nhttps://arxiv.org/abs/1705.07115\n\nPlease see if its of any use . \n\nThe paper specifically focuses on optimization for multitask loss .",
      "votes": null
    },
    {
      "id": "678245",
      "postDate": "11/21/2019 06:28:18",
      "content": "<p>doing the same usually. \nJust fiddle around with the loss weights until they start to overfit roughly at the same time.</p>",
      "rawMarkdown": "doing the same usually. \nJust fiddle around with the loss weights until they start to overfit roughly at the same time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 665056,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "11/04/2019 15:29:55",
      "content": "<p>Loss weights would help I think.</p>",
      "votes": null,
      "replies": [
        {
          "id": 678245,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "11/21/2019 06:28:18",
          "content": "<p>doing the same usually. \nJust fiddle around with the loss weights until they start to overfit roughly at the same time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 678150,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "11/21/2019 03:08:36",
      "content": "<p>What about doing a log(maskloss) to reduce the scale ? .. I dont know if thats a thing .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 678239,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "11/21/2019 06:18:05",
      "content": "<p>I was just skimming through this , and found this paper </p>\n\n<p>\"Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics\"</p>\n\n<p><a href=\"https://arxiv.org/abs/1705.07115\">https://arxiv.org/abs/1705.07115</a></p>\n\n<p>Please see if its of any use . </p>\n\n<p>The paper specifically focuses on optimization for multitask loss .</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "665006": "Hi all, do you have any experience of optimizing multiple losses?\n\nFor example, in https://www.kaggle.com/hocop1/centernet-baseline, the centernet consists of mask and regression losses. From my observation, the mask loss will drop for a few epoches and then keep increasing (after epoch 5 or 7) while the regression loss can be minimized after epoch 5 or 7.\n\nSince the magnitude of mask loss (24) is larger than that of regression loss (e.g. 0.24), the mask loss dominates, which will result in early termination of the optimization process for centernet.  Do you have any experience of minimizing multiple losses at the same time?\n\nIf this is not clear enough, please let me know and I will try my best to clarify.",
    "665056": "Loss weights would help I think.",
    "678150": "What about doing a log(maskloss) to reduce the scale ? .. I dont know if thats a thing .",
    "678239": "I was just skimming through this , and found this paper \n\n\"Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics\"\n\nhttps://arxiv.org/abs/1705.07115\n\nPlease see if its of any use . \n\nThe paper specifically focuses on optimization for multitask loss .",
    "678245": "doing the same usually. \nJust fiddle around with the loss weights until they start to overfit roughly at the same time."
  },
  "source": "meta"
}