{
  "id": 300960,
  "title": "How to Handle When Learning Curve Become Unstable?",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/300960",
  "author_name": "",
  "post_date": "2022-01-15T07:00:42.064430200Z",
  "votes": 6,
  "comment_count": 16,
  "views": 0,
  "content": "<p><strong>Note for those who follows the original topic:</strong> I fully rewrote this topic based on the actual case I encountered. I think this is better compared to discussing about the general case.</p>\n<h1>Simptom</h1>\n<ul>\n<li>mAP suddenly drops at epoch 6 (see gray curve on Fig1)</li>\n<li>both train/val loss get higher abruptly at this epoch (gray curve of Fig2-3)</li>\n</ul>\n<h1>Assumed Cause</h1>\n<ul>\n<li>the gradient vanishes/explodes during training</li>\n</ul>\n<hr>\n<p>As <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> pointed out, the root cause might be data.<br>\nsee also this thread:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300960#1667201\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300960#1667201</a></p>\n<h1>Assumed Remedy</h1>\n<ul>\n<li>to decrease leaning rate</li>\n<li>to increase momentum weight when you use SGD, Adam, AdamW optimizer</li>\n</ul>\n<hr>\n<p><strong>Fig1:</strong><br>\n<a href=\"https://ibb.co/6gPcvcD\"><img src=\"https://i.ibb.co/9sTmwmy/Screen-Shot-2022-01-28-at-20-36-28.png\" alt=\"Screen-Shot-2022-01-28-at-20-36-28\"></a></p>\n<p><strong>Fig2:</strong><br>\n<a href=\"https://ibb.co/rmGYmSL\"><img src=\"https://i.ibb.co/dr0Frq3/Screen-Shot-2022-01-28-at-20-36-37.png\" alt=\"Screen-Shot-2022-01-28-at-20-36-37\"></a></p>\n<p><strong>Fig3:</strong><br>\n<a href=\"https://ibb.co/vBcm7qb\"><img src=\"https://i.ibb.co/Ctwvk8q/Screen-Shot-2022-01-28-at-20-36-45.png\" alt=\"Screen-Shot-2022-01-28-at-20-36-45\"></a></p>\n<hr>\n<h1>Update Note</h1>\n<ul>\n<li>2022/1/28: Fully updated topic based on the actual case</li>\n</ul>",
  "messages": [
    {
      "id": "1650606",
      "postDate": "01/15/2022 07:00:42",
      "content": "<p><strong>Note for those who follows the original topic:</strong> I fully rewrote this topic based on the actual case I encountered. I think this is better compared to discussing about the general case.</p>\n<h1>Simptom</h1>\n<ul>\n<li>mAP suddenly drops at epoch 6 (see gray curve on Fig1)</li>\n<li>both train/val loss get higher abruptly at this epoch (gray curve of Fig2-3)</li>\n</ul>\n<h1>Assumed Cause</h1>\n<ul>\n<li>the gradient vanishes/explodes during training</li>\n</ul>\n<hr>\n<p>As <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> pointed out, the root cause might be data.<br>\nsee also this thread:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300960#1667201\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300960#1667201</a></p>\n<h1>Assumed Remedy</h1>\n<ul>\n<li>to decrease leaning rate</li>\n<li>to increase momentum weight when you use SGD, Adam, AdamW optimizer</li>\n</ul>\n<hr>\n<p><strong>Fig1:</strong><br>\n<a href=\"https://ibb.co/6gPcvcD\"><img src=\"https://i.ibb.co/9sTmwmy/Screen-Shot-2022-01-28-at-20-36-28.png\" alt=\"Screen-Shot-2022-01-28-at-20-36-28\"></a></p>\n<p><strong>Fig2:</strong><br>\n<a href=\"https://ibb.co/rmGYmSL\"><img src=\"https://i.ibb.co/dr0Frq3/Screen-Shot-2022-01-28-at-20-36-37.png\" alt=\"Screen-Shot-2022-01-28-at-20-36-37\"></a></p>\n<p><strong>Fig3:</strong><br>\n<a href=\"https://ibb.co/vBcm7qb\"><img src=\"https://i.ibb.co/Ctwvk8q/Screen-Shot-2022-01-28-at-20-36-45.png\" alt=\"Screen-Shot-2022-01-28-at-20-36-45\"></a></p>\n<hr>\n<h1>Update Note</h1>\n<ul>\n<li>2022/1/28: Fully updated topic based on the actual case</li>\n</ul>",
      "rawMarkdown": "**Note for those who follows the original topic:** I fully rewrote this topic based on the actual case I encountered. I think this is better compared to discussing about the general case.\n\n# Simptom\n\n* mAP suddenly drops at epoch 6 (see gray curve on Fig1)\n* both train/val loss get higher abruptly at this epoch (gray curve of Fig2-3)\n\n# Assumed Cause\n\n* the gradient vanishes/explodes during training\n\n------\n\nAs @hengck23 pointed out, the root cause might be data.\nsee also this thread:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300960#1667201\n\n\n# Assumed Remedy\n\n* to decrease leaning rate\n* to increase momentum weight when you use SGD, Adam, AdamW optimizer\n\n---\n\n**Fig1:**\n<a href=\"https://ibb.co/6gPcvcD\"><img src=\"https://i.ibb.co/9sTmwmy/Screen-Shot-2022-01-28-at-20-36-28.png\" alt=\"Screen-Shot-2022-01-28-at-20-36-28\" border=\"0\"></a>\n\n**Fig2:**\n<a href=\"https://ibb.co/rmGYmSL\"><img src=\"https://i.ibb.co/dr0Frq3/Screen-Shot-2022-01-28-at-20-36-37.png\" alt=\"Screen-Shot-2022-01-28-at-20-36-37\" border=\"0\"></a>\n\n**Fig3:**\n<a href=\"https://ibb.co/vBcm7qb\"><img src=\"https://i.ibb.co/Ctwvk8q/Screen-Shot-2022-01-28-at-20-36-45.png\" alt=\"Screen-Shot-2022-01-28-at-20-36-45\" border=\"0\"></a>\n\n---\n\n# Update Note\n\n* 2022/1/28: Fully updated topic based on the actual case",
      "votes": null
    },
    {
      "id": "1653008",
      "postDate": "01/17/2022 07:00:40",
      "content": "<p>memo: from the forum topic for the related symptom</p>\n<blockquote>\n  <p>1) Look at training learning curve: How is the learning curve on train set? Does it learn the training set? If not, first work on that to make sure you can over fit on the training set.<br>\n  2) Check your data to make sure there is no NaN in it (training, validation, test)<br>\n  3) Check the gradients and the weights to make sure there is no NaN.<br>\n  4) Decrease the learning rate as you train to make sure it's not because of a sudden big update that stuck in a sharp minima.<br>\n  5) To make sure everything's right, check the predictions of your network so that your network is not making some constant, or repetitive predictions.<br>\n  6) Check if your data in your batch is balanced with respect to all classes.<br>\n  7) normalize your data to be zero mean unit variance. Initialize the weights likewise. It will assist the training.</p>\n</blockquote>\n<p>[1] <a href=\"https://stats.stackexchange.com/questions/228920/sudden-accuracy-drop-when-training-lstm-or-gru-in-keras\" target=\"_blank\">https://stats.stackexchange.com/questions/228920/sudden-accuracy-drop-when-training-lstm-or-gru-in-keras</a><br>\n[2] <a href=\"https://stackoverflow.com/questions/37044600/sudden-drop-in-accuracy-while-training-a-deep-neural-net\" target=\"_blank\">https://stackoverflow.com/questions/37044600/sudden-drop-in-accuracy-while-training-a-deep-neural-net</a></p>",
      "rawMarkdown": "memo: from the forum topic for the related symptom\n\n> 1) Look at training learning curve: How is the learning curve on train set? Does it learn the training set? If not, first work on that to make sure you can over fit on the training set.\n2) Check your data to make sure there is no NaN in it (training, validation, test)\n3) Check the gradients and the weights to make sure there is no NaN.\n4) Decrease the learning rate as you train to make sure it's not because of a sudden big update that stuck in a sharp minima.\n5) To make sure everything's right, check the predictions of your network so that your network is not making some constant, or repetitive predictions.\n6) Check if your data in your batch is balanced with respect to all classes.\n7) normalize your data to be zero mean unit variance. Initialize the weights likewise. It will assist the training.\n\n[1] https://stats.stackexchange.com/questions/228920/sudden-accuracy-drop-when-training-lstm-or-gru-in-keras\n[2] https://stackoverflow.com/questions/37044600/sudden-drop-in-accuracy-while-training-a-deep-neural-net",
      "votes": null
    },
    {
      "id": "1666969",
      "postDate": "01/28/2022 07:53:33",
      "content": "<h1>Tips: Decreasing Leaning Rate &amp; Increasing Momentum Parameter</h1>\n<p>In my case, the following parameter tuning was effective, at least in some cases:</p>\n<ul>\n<li>decreasing learning rate</li>\n<li>increase beta1 for AdamW optimizer (which is similar to momentum for SGD)</li>\n</ul>\n<p>One of the reason that the training process gets unstable is learning rate is too high.<br>\nIf it is the case, either decreasing learning rate or increasing momentum parameter (or both) will mitigate this problem.</p>",
      "rawMarkdown": "# Tips: Decreasing Leaning Rate & Increasing Momentum Parameter\n\nIn my case, the following parameter tuning was effective, at least in some cases:\n\n* decreasing learning rate\n* increase beta1 for AdamW optimizer (which is similar to momentum for SGD)\n\nOne of the reason that the training process gets unstable is learning rate is too high.\nIf it is the case, either decreasing learning rate or increasing momentum parameter (or both) will mitigate this problem.",
      "votes": null
    },
    {
      "id": "1666979",
      "postDate": "01/28/2022 08:07:07",
      "content": "<p>Look at the top image below.<br>\nThe gray curve shows a sudden drop in mAP at the sixth epoch.<br>\nAlso, in the middle figure, the learning loss is increased abruptly at this epoch.<br>\nThis indicates that something is wrong with the learning (e.g., the gradient vanished/exploded, the structure of the trained model is broken, etc.).<br>\nIn this case, you can try to reduce the learning rate or increase the momentum with your optimizer.<br>\nFor the blue curve, I reduced the learning rate (0.01 -&gt; 0.003) and increased the momentum (0.937 -&gt; 0.979).<br>\nCompared to the gray curve, the learning curve becomes more stable.</p>\n<p><a href=\"https://ibb.co/QfyHPBH\"><img src=\"https://i.ibb.co/47hS8yS/Screen-Shot-2022-01-28-at-16-54-02.png\" alt=\"Screen-Shot-2022-01-28-at-16-54-02\"></a><br>\n<a href=\"https://ibb.co/SyR1Vd7\"><img src=\"https://i.ibb.co/W5gdBPp/Screen-Shot-2022-01-28-at-16-54-17.png\" alt=\"Screen-Shot-2022-01-28-at-16-54-17\"></a><br>\n<a href=\"https://ibb.co/gP9FCTZ\"><img src=\"https://i.ibb.co/CM8sF7K/Screen-Shot-2022-01-28-at-16-54-30.png\" alt=\"Screen-Shot-2022-01-28-at-16-54-30\"></a></p>",
      "rawMarkdown": "Look at the top image below.\nThe gray curve shows a sudden drop in mAP at the sixth epoch.\nAlso, in the middle figure, the learning loss is increased abruptly at this epoch.\nThis indicates that something is wrong with the learning (e.g., the gradient vanished/exploded, the structure of the trained model is broken, etc.).\nIn this case, you can try to reduce the learning rate or increase the momentum with your optimizer.\nFor the blue curve, I reduced the learning rate (0.01 -> 0.003) and increased the momentum (0.937 -> 0.979).\nCompared to the gray curve, the learning curve becomes more stable.\n\n<a href=\"https://ibb.co/QfyHPBH\"><img src=\"https://i.ibb.co/47hS8yS/Screen-Shot-2022-01-28-at-16-54-02.png\" alt=\"Screen-Shot-2022-01-28-at-16-54-02\" border=\"0\"></a>\n<a href=\"https://ibb.co/SyR1Vd7\"><img src=\"https://i.ibb.co/W5gdBPp/Screen-Shot-2022-01-28-at-16-54-17.png\" alt=\"Screen-Shot-2022-01-28-at-16-54-17\" border=\"0\"></a>\n<a href=\"https://ibb.co/gP9FCTZ\"><img src=\"https://i.ibb.co/CM8sF7K/Screen-Shot-2022-01-28-at-16-54-30.png\" alt=\"Screen-Shot-2022-01-28-at-16-54-30\" border=\"0\"></a>",
      "votes": null
    },
    {
      "id": "1667080",
      "postDate": "01/28/2022 09:48:46",
      "content": "<p>how to get momentum value (i.e. 0.979)?</p>",
      "rawMarkdown": "how to get momentum value (i.e. 0.979)?",
      "votes": null
    },
    {
      "id": "1667129",
      "postDate": "01/28/2022 10:41:38",
      "content": "<p><a href=\"https://www.kaggle.com/yienngxiong\" target=\"_blank\">@yienngxiong</a> </p>\n<p>From this equation. This means the momentum is weighed x3 times.</p>\n<pre><code>&gt;&gt;&gt; 1 - ((1 - 0.937) / 3)\n0.979\n</code></pre>\n<h1>Reference</h1>\n<ul>\n<li>Adam: <a href=\"https://pytorch.org/docs/stable/generated/torch.optim.Adam.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.optim.Adam.html</a></li>\n<li>SGD: <a href=\"https://pytorch.org/docs/stable/generated/torch.optim.SGD.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.optim.SGD.html</a></li>\n</ul>",
      "rawMarkdown": "yienngxiong \n\nFrom this equation. This means the momentum is weighed x3 times.\n\n```\n>>> 1 - ((1 - 0.937) / 3)\n0.979\n```\n\n# Reference\n\n- Adam: https://pytorch.org/docs/stable/generated/torch.optim.Adam.html\n- SGD: https://pytorch.org/docs/stable/generated/torch.optim.SGD.html",
      "votes": null
    },
    {
      "id": "1667140",
      "postDate": "01/28/2022 10:46:31",
      "content": "<p>As I am using AdamW optimizer, I specified beta1, but it's equivalent to momentum in SGD.</p>",
      "rawMarkdown": "As I am using AdamW optimizer, I specified beta1, but it's equivalent to momentum in SGD.",
      "votes": null
    },
    {
      "id": "1667167",
      "postDate": "01/28/2022 11:11:00",
      "content": "<p>In other words, decreasing learning rate x1/3 and increasing momentum weight x3 means making the gradient of current step contributes x1/9 to the network parameter updates.</p>\n<p>This is effective to weaken “too much learning from just looking at the current batch”, and making your leaning curve smoother.</p>",
      "rawMarkdown": "In other words, decreasing learning rate x1/3 and increasing momentum weight x3 means making the gradient of current step contributes x1/9 to the network parameter updates.\n\nThis is effective to weaken “too much learning from just looking at the current batch”, and making your leaning curve smoother.",
      "votes": null
    },
    {
      "id": "1667177",
      "postDate": "01/28/2022 11:19:01",
      "content": "<p>I think this tuning is also effective to increase stability on small batch size. However because of high variance of BN statistics in small batch size, the effect must be limited.</p>",
      "rawMarkdown": "I think this tuning is also effective to increase stability on small batch size. However because of high variance of BN statistics in small batch size, the effect must be limited.",
      "votes": null
    },
    {
      "id": "1667201",
      "postDate": "01/28/2022 11:49:48",
      "content": "<p>the real reason for fluctuating loss curve is data.</p>\n<p>when the current batch of data is different from the previous one (or previous ones), the loss will also be different.<br>\nin short, it is effects of outliers.</p>\n<p>let me give you a few eamples. assume we are training a dog vs cat image classifier:</p>\n<ol>\n<li><p>if for one batch, there is only one class (e.g. batch of 256 images and all label are dog), then the model will learn to predict dog with much higher probability after the batch update. when the next batch contains cat images, the loss suddenly increases.</p></li>\n<li><p>for one batch, the label are 100% wrongly labelled., then the model will predict the opposite after the batch update.<br>\nwhen the next batch contains correct label, the loss suddenly increases.</p></li>\n</ol>\n<hr>\n<p>mometum, learning rate limiting gradient values prevents accidential wrong udpate, but they do do not solve the root cause of the problem. if the outlier ratio is low, you can use such methods.</p>\n<p>if you really want to find out the root cause, you have to think of a way to find out the outliers.</p>",
      "rawMarkdown": "the real reason for fluctuating loss curve is data.\n\nwhen the current batch of data is different from the previous one (or previous ones), the loss will also be different.\nin short, it is effects of outliers.\n\nlet me give you a few eamples. assume we are training a dog vs cat image classifier:\n1. if for one batch, there is only one class (e.g. batch of 256 images and all label are dog), then the model will learn to predict dog with much higher probability after the batch update. when the next batch contains cat images, the loss suddenly increases.\n\n2. for one batch, the label are 100% wrongly labelled., then the model will predict the opposite after the batch update.\nwhen the next batch contains correct label, the loss suddenly increases.\n\n----\n\nmometum, learning rate limiting gradient values prevents accidential wrong udpate, but they do do not solve the root cause of the problem. if the outlier ratio is low, you can use such methods.\n\nif you really want to find out the root cause, you have to think of a way to find out the outliers.",
      "votes": null
    },
    {
      "id": "1667223",
      "postDate": "01/28/2022 12:09:09",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks for the comment.</p>\n<p>Actually, the gray curve is obtained by adding 30% of negative samples (whereas the pink curve is 0% negative samples).</p>\n<p>The data is sampled with a certain probability where I get sample ratio neg:pos=0.3:1.0 (I used full image datasets, then sample per every epoch to make data of this ratio.).</p>\n<p>So, If I take it for granted what you say (the root cause is data), I guess the most probable cause of fluctuation is based on the randomness of sampling.</p>\n<p>I thought random sampling is better than fixed size sampling (because it can learns on better noise), but in this situation, do you think this strategy is bad?</p>\n<p>(Currently I don't know how to test if the randomness really matters, but let me think about this.)</p>\n<hr>\n<h2>Additional Note</h2>\n<p>on gray curve, both positive and negative data are sampled with certain probability, and they are sampled with <code>repeat=True</code> (namely, the same data can be sampled more than once).<br>\nOn the contrary, in the pink curve, each positive set is sampled only once (which is equivalent to <code>repeat=False</code>). So I guess the variance of sample is even bigger for the gray curve compared to pink one.</p>",
      "rawMarkdown": "hengck23 Thanks for the comment.\n\nActually, the gray curve is obtained by adding 30% of negative samples (whereas the pink curve is 0% negative samples).\n\nThe data is sampled with a certain probability where I get sample ratio neg:pos=0.3:1.0 (I used full image datasets, then sample per every epoch to make data of this ratio.).\n\nSo, If I take it for granted what you say (the root cause is data), I guess the most probable cause of fluctuation is based on the randomness of sampling.\n\nI thought random sampling is better than fixed size sampling (because it can learns on better noise), but in this situation, do you think this strategy is bad?\n\n(Currently I don't know how to test if the randomness really matters, but let me think about this.)\n\n---\n\n## Additional Note\n\non gray curve, both positive and negative data are sampled with certain probability, and they are sampled with `repeat=True` (namely, the same data can be sampled more than once).\nOn the contrary, in the pink curve, each positive set is sampled only once (which is equivalent to `repeat=False`). So I guess the variance of sample is even bigger for the gray curve compared to pink one.",
      "votes": null
    },
    {
      "id": "1667242",
      "postDate": "01/28/2022 12:26:41",
      "content": "<p>According to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , tuning the learning rate and momentum won't solve the root cause.</p>\n<p>See also this thread:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300960#1667201\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300960#1667201</a></p>",
      "rawMarkdown": "According to @hengck23 , tuning the learning rate and momentum won't solve the root cause.\n\nSee also this thread:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300960#1667201",
      "votes": null
    },
    {
      "id": "1667257",
      "postDate": "01/28/2022 12:35:04",
      "content": "<p>i think there are some COTS that looks like background. this cuases confusion when learning.<br>\n(also there are some blur small COTS in the ground truth, if you enlarge them in scale augmentation, there are may become bad samples. The larger COTS in the ground truth are not blur)</p>\n<p>when you do negative background sampling, you may sample background that looks like COTS.<br>\nthe ratio affects the frequency of such occurrence, etc</p>",
      "rawMarkdown": "i think there are some COTS that looks like background. this cuases confusion when learning.\n(also there are some blur small COTS in the ground truth, if you enlarge them in scale augmentation, there are may become bad samples. The larger COTS in the ground truth are not blur)\n\nwhen you do negative background sampling, you may sample background that looks like COTS.\nthe ratio affects the frequency of such occurrence, etc",
      "votes": null
    },
    {
      "id": "1667273",
      "postDate": "01/28/2022 12:42:05",
      "content": "<p>\"mAP suddenly drops at epoch 6 (see gray curve on Fig1)\"</p>\n<p>most kagglers will experience decrease up to 6,7,8 for using video 1 as validation.</p>\n<p>for me although the map raises again, it finally will fall back again as I uses more train iterations. I use specialised augmentation and EMA for this.</p>\n<hr>\n<p>this can be explained as follows:</p>\n<ul>\n<li>deep models learned general target object charateristics at first. (generalisation is good, loss curve falls rapidly)</li>\n<li>then it starts to learn (and sometimes wrongly overfitted to) noise in the data (the loss of the good data becomes very low. only bad data has high loss and contribute more to backprop gradient)</li>\n</ul>",
      "rawMarkdown": "\"mAP suddenly drops at epoch 6 (see gray curve on Fig1)\"\n\nmost kagglers will experience decrease up to 6,7,8 for using video 1 as validation.\n\nfor me although the map raises again, it finally will fall back again as I uses more train iterations. I use specialised augmentation and EMA for this.\n\n---\n\nthis can be explained as follows:\n- deep models learned general target object charateristics at first. (generalisation is good, loss curve falls rapidly)\n- then it starts to learn (and sometimes wrongly overfitted to) noise in the data (the loss of the good data becomes very low. only bad data has high loss and contribute more to backprop gradient)",
      "votes": null
    },
    {
      "id": "1667279",
      "postDate": "01/28/2022 12:49:42",
      "content": "<p>I see. I guess the un-labeled positive COTS (which is discussed in other discussions) might also harm learning.</p>",
      "rawMarkdown": "I see. I guess the un-labeled positive COTS (which is discussed in other discussions) might also harm learning.",
      "votes": null
    },
    {
      "id": "1667289",
      "postDate": "01/28/2022 13:01:38",
      "content": "<p>I see. It might be too fast judging the gray curve is failed.</p>\n<p>BTW, can I ask what EMA is?</p>",
      "rawMarkdown": "I see. It might be too fast judging the gray curve is failed.\n\nBTW, can I ask what EMA is?",
      "votes": null
    },
    {
      "id": "1667413",
      "postDate": "01/28/2022 14:58:18",
      "content": "<h1>Related Paper about Unlabeled Data/Noisy Labels</h1>\n<ul>\n<li><a href=\"https://openaccess.thecvf.com/content/CVPR2021/papers/Guo_Positive-Unlabeled_Data_Purification_in_the_Wild_for_Object_Detection_CVPR_2021_paper.pdf\" target=\"_blank\">Positive-Unlabeled Data Purification in the Wild for Object Detection</a></li>\n<li><a href=\"https://arxiv.org/pdf/2005.04757.pdf\" target=\"_blank\">A Simple Semi-Supervised Learning Framework for Object Detection</a></li>\n<li><a href=\"https://paperswithcode.com/paper/end-to-end-semi-supervised-object-detection\" target=\"_blank\">End-to-End Semi-Supervised Object Detection with Soft Teacher</a></li>\n<li><a href=\"https://paperswithcode.com/paper/unbiased-teacher-for-semi-supervised-object-1\" target=\"_blank\">Unbiased Teacher for Semi-Supervised Object Detection</a></li>\n<li><a href=\"https://arxiv.org/pdf/1706.08249.pdf\" target=\"_blank\">Few-Example Object Detection with Model Communication</a></li>\n</ul>",
      "rawMarkdown": "# Related Paper about Unlabeled Data/Noisy Labels\n\n- [Positive-Unlabeled Data Purification in the Wild for Object Detection](https://openaccess.thecvf.com/content/CVPR2021/papers/Guo_Positive-Unlabeled_Data_Purification_in_the_Wild_for_Object_Detection_CVPR_2021_paper.pdf)\n- [A Simple Semi-Supervised Learning Framework for Object Detection](https://arxiv.org/pdf/2005.04757.pdf)\n- [End-to-End Semi-Supervised Object Detection with Soft Teacher](https://paperswithcode.com/paper/end-to-end-semi-supervised-object-detection)\n- [Unbiased Teacher for Semi-Supervised Object Detection](https://paperswithcode.com/paper/unbiased-teacher-for-semi-supervised-object-1)\n- [Few-Example Object Detection with Model Communication](https://arxiv.org/pdf/1706.08249.pdf)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1653008,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "01/17/2022 07:00:40",
      "content": "<p>memo: from the forum topic for the related symptom</p>\n<blockquote>\n  <p>1) Look at training learning curve: How is the learning curve on train set? Does it learn the training set? If not, first work on that to make sure you can over fit on the training set.<br>\n  2) Check your data to make sure there is no NaN in it (training, validation, test)<br>\n  3) Check the gradients and the weights to make sure there is no NaN.<br>\n  4) Decrease the learning rate as you train to make sure it's not because of a sudden big update that stuck in a sharp minima.<br>\n  5) To make sure everything's right, check the predictions of your network so that your network is not making some constant, or repetitive predictions.<br>\n  6) Check if your data in your batch is balanced with respect to all classes.<br>\n  7) normalize your data to be zero mean unit variance. Initialize the weights likewise. It will assist the training.</p>\n</blockquote>\n<p>[1] <a href=\"https://stats.stackexchange.com/questions/228920/sudden-accuracy-drop-when-training-lstm-or-gru-in-keras\" target=\"_blank\">https://stats.stackexchange.com/questions/228920/sudden-accuracy-drop-when-training-lstm-or-gru-in-keras</a><br>\n[2] <a href=\"https://stackoverflow.com/questions/37044600/sudden-drop-in-accuracy-while-training-a-deep-neural-net\" target=\"_blank\">https://stackoverflow.com/questions/37044600/sudden-drop-in-accuracy-while-training-a-deep-neural-net</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1666969,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "01/28/2022 07:53:33",
      "content": "<h1>Tips: Decreasing Leaning Rate &amp; Increasing Momentum Parameter</h1>\n<p>In my case, the following parameter tuning was effective, at least in some cases:</p>\n<ul>\n<li>decreasing learning rate</li>\n<li>increase beta1 for AdamW optimizer (which is similar to momentum for SGD)</li>\n</ul>\n<p>One of the reason that the training process gets unstable is learning rate is too high.<br>\nIf it is the case, either decreasing learning rate or increasing momentum parameter (or both) will mitigate this problem.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1666979,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "01/28/2022 08:07:07",
          "content": "<p>Look at the top image below.<br>\nThe gray curve shows a sudden drop in mAP at the sixth epoch.<br>\nAlso, in the middle figure, the learning loss is increased abruptly at this epoch.<br>\nThis indicates that something is wrong with the learning (e.g., the gradient vanished/exploded, the structure of the trained model is broken, etc.).<br>\nIn this case, you can try to reduce the learning rate or increase the momentum with your optimizer.<br>\nFor the blue curve, I reduced the learning rate (0.01 -&gt; 0.003) and increased the momentum (0.937 -&gt; 0.979).<br>\nCompared to the gray curve, the learning curve becomes more stable.</p>\n<p><a href=\"https://ibb.co/QfyHPBH\"><img src=\"https://i.ibb.co/47hS8yS/Screen-Shot-2022-01-28-at-16-54-02.png\" alt=\"Screen-Shot-2022-01-28-at-16-54-02\"></a><br>\n<a href=\"https://ibb.co/SyR1Vd7\"><img src=\"https://i.ibb.co/W5gdBPp/Screen-Shot-2022-01-28-at-16-54-17.png\" alt=\"Screen-Shot-2022-01-28-at-16-54-17\"></a><br>\n<a href=\"https://ibb.co/gP9FCTZ\"><img src=\"https://i.ibb.co/CM8sF7K/Screen-Shot-2022-01-28-at-16-54-30.png\" alt=\"Screen-Shot-2022-01-28-at-16-54-30\"></a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667080,
          "author_name": "yienngxiong",
          "author_url": "",
          "post_date": "01/28/2022 09:48:46",
          "content": "<p>how to get momentum value (i.e. 0.979)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667129,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "01/28/2022 10:41:38",
          "content": "<p><a href=\"https://www.kaggle.com/yienngxiong\" target=\"_blank\">@yienngxiong</a> </p>\n<p>From this equation. This means the momentum is weighed x3 times.</p>\n<pre><code>&gt;&gt;&gt; 1 - ((1 - 0.937) / 3)\n0.979\n</code></pre>\n<h1>Reference</h1>\n<ul>\n<li>Adam: <a href=\"https://pytorch.org/docs/stable/generated/torch.optim.Adam.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.optim.Adam.html</a></li>\n<li>SGD: <a href=\"https://pytorch.org/docs/stable/generated/torch.optim.SGD.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.optim.SGD.html</a></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667140,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "01/28/2022 10:46:31",
          "content": "<p>As I am using AdamW optimizer, I specified beta1, but it's equivalent to momentum in SGD.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667167,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "01/28/2022 11:11:00",
          "content": "<p>In other words, decreasing learning rate x1/3 and increasing momentum weight x3 means making the gradient of current step contributes x1/9 to the network parameter updates.</p>\n<p>This is effective to weaken “too much learning from just looking at the current batch”, and making your leaning curve smoother.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667177,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "01/28/2022 11:19:01",
          "content": "<p>I think this tuning is also effective to increase stability on small batch size. However because of high variance of BN statistics in small batch size, the effect must be limited.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667242,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "01/28/2022 12:26:41",
          "content": "<p>According to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , tuning the learning rate and momentum won't solve the root cause.</p>\n<p>See also this thread:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300960#1667201\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300960#1667201</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1667201,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/28/2022 11:49:48",
      "content": "<p>the real reason for fluctuating loss curve is data.</p>\n<p>when the current batch of data is different from the previous one (or previous ones), the loss will also be different.<br>\nin short, it is effects of outliers.</p>\n<p>let me give you a few eamples. assume we are training a dog vs cat image classifier:</p>\n<ol>\n<li><p>if for one batch, there is only one class (e.g. batch of 256 images and all label are dog), then the model will learn to predict dog with much higher probability after the batch update. when the next batch contains cat images, the loss suddenly increases.</p></li>\n<li><p>for one batch, the label are 100% wrongly labelled., then the model will predict the opposite after the batch update.<br>\nwhen the next batch contains correct label, the loss suddenly increases.</p></li>\n</ol>\n<hr>\n<p>mometum, learning rate limiting gradient values prevents accidential wrong udpate, but they do do not solve the root cause of the problem. if the outlier ratio is low, you can use such methods.</p>\n<p>if you really want to find out the root cause, you have to think of a way to find out the outliers.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1667223,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "01/28/2022 12:09:09",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks for the comment.</p>\n<p>Actually, the gray curve is obtained by adding 30% of negative samples (whereas the pink curve is 0% negative samples).</p>\n<p>The data is sampled with a certain probability where I get sample ratio neg:pos=0.3:1.0 (I used full image datasets, then sample per every epoch to make data of this ratio.).</p>\n<p>So, If I take it for granted what you say (the root cause is data), I guess the most probable cause of fluctuation is based on the randomness of sampling.</p>\n<p>I thought random sampling is better than fixed size sampling (because it can learns on better noise), but in this situation, do you think this strategy is bad?</p>\n<p>(Currently I don't know how to test if the randomness really matters, but let me think about this.)</p>\n<hr>\n<h2>Additional Note</h2>\n<p>on gray curve, both positive and negative data are sampled with certain probability, and they are sampled with <code>repeat=True</code> (namely, the same data can be sampled more than once).<br>\nOn the contrary, in the pink curve, each positive set is sampled only once (which is equivalent to <code>repeat=False</code>). So I guess the variance of sample is even bigger for the gray curve compared to pink one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667257,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/28/2022 12:35:04",
          "content": "<p>i think there are some COTS that looks like background. this cuases confusion when learning.<br>\n(also there are some blur small COTS in the ground truth, if you enlarge them in scale augmentation, there are may become bad samples. The larger COTS in the ground truth are not blur)</p>\n<p>when you do negative background sampling, you may sample background that looks like COTS.<br>\nthe ratio affects the frequency of such occurrence, etc</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667279,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "01/28/2022 12:49:42",
          "content": "<p>I see. I guess the un-labeled positive COTS (which is discussed in other discussions) might also harm learning.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1667273,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/28/2022 12:42:05",
      "content": "<p>\"mAP suddenly drops at epoch 6 (see gray curve on Fig1)\"</p>\n<p>most kagglers will experience decrease up to 6,7,8 for using video 1 as validation.</p>\n<p>for me although the map raises again, it finally will fall back again as I uses more train iterations. I use specialised augmentation and EMA for this.</p>\n<hr>\n<p>this can be explained as follows:</p>\n<ul>\n<li>deep models learned general target object charateristics at first. (generalisation is good, loss curve falls rapidly)</li>\n<li>then it starts to learn (and sometimes wrongly overfitted to) noise in the data (the loss of the good data becomes very low. only bad data has high loss and contribute more to backprop gradient)</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1667289,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "01/28/2022 13:01:38",
          "content": "<p>I see. It might be too fast judging the gray curve is failed.</p>\n<p>BTW, can I ask what EMA is?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1667413,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "01/28/2022 14:58:18",
      "content": "<h1>Related Paper about Unlabeled Data/Noisy Labels</h1>\n<ul>\n<li><a href=\"https://openaccess.thecvf.com/content/CVPR2021/papers/Guo_Positive-Unlabeled_Data_Purification_in_the_Wild_for_Object_Detection_CVPR_2021_paper.pdf\" target=\"_blank\">Positive-Unlabeled Data Purification in the Wild for Object Detection</a></li>\n<li><a href=\"https://arxiv.org/pdf/2005.04757.pdf\" target=\"_blank\">A Simple Semi-Supervised Learning Framework for Object Detection</a></li>\n<li><a href=\"https://paperswithcode.com/paper/end-to-end-semi-supervised-object-detection\" target=\"_blank\">End-to-End Semi-Supervised Object Detection with Soft Teacher</a></li>\n<li><a href=\"https://paperswithcode.com/paper/unbiased-teacher-for-semi-supervised-object-1\" target=\"_blank\">Unbiased Teacher for Semi-Supervised Object Detection</a></li>\n<li><a href=\"https://arxiv.org/pdf/1706.08249.pdf\" target=\"_blank\">Few-Example Object Detection with Model Communication</a></li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1650606": "**Note for those who follows the original topic:** I fully rewrote this topic based on the actual case I encountered. I think this is better compared to discussing about the general case.\n\n# Simptom\n\n* mAP suddenly drops at epoch 6 (see gray curve on Fig1)\n* both train/val loss get higher abruptly at this epoch (gray curve of Fig2-3)\n\n# Assumed Cause\n\n* the gradient vanishes/explodes during training\n\n------\n\nAs @hengck23 pointed out, the root cause might be data.\nsee also this thread:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300960#1667201\n\n\n# Assumed Remedy\n\n* to decrease leaning rate\n* to increase momentum weight when you use SGD, Adam, AdamW optimizer\n\n---\n\n**Fig1:**\n<a href=\"https://ibb.co/6gPcvcD\"><img src=\"https://i.ibb.co/9sTmwmy/Screen-Shot-2022-01-28-at-20-36-28.png\" alt=\"Screen-Shot-2022-01-28-at-20-36-28\" border=\"0\"></a>\n\n**Fig2:**\n<a href=\"https://ibb.co/rmGYmSL\"><img src=\"https://i.ibb.co/dr0Frq3/Screen-Shot-2022-01-28-at-20-36-37.png\" alt=\"Screen-Shot-2022-01-28-at-20-36-37\" border=\"0\"></a>\n\n**Fig3:**\n<a href=\"https://ibb.co/vBcm7qb\"><img src=\"https://i.ibb.co/Ctwvk8q/Screen-Shot-2022-01-28-at-20-36-45.png\" alt=\"Screen-Shot-2022-01-28-at-20-36-45\" border=\"0\"></a>\n\n---\n\n# Update Note\n\n* 2022/1/28: Fully updated topic based on the actual case",
    "1653008": "memo: from the forum topic for the related symptom\n\n> 1) Look at training learning curve: How is the learning curve on train set? Does it learn the training set? If not, first work on that to make sure you can over fit on the training set.\n2) Check your data to make sure there is no NaN in it (training, validation, test)\n3) Check the gradients and the weights to make sure there is no NaN.\n4) Decrease the learning rate as you train to make sure it's not because of a sudden big update that stuck in a sharp minima.\n5) To make sure everything's right, check the predictions of your network so that your network is not making some constant, or repetitive predictions.\n6) Check if your data in your batch is balanced with respect to all classes.\n7) normalize your data to be zero mean unit variance. Initialize the weights likewise. It will assist the training.\n\n[1] https://stats.stackexchange.com/questions/228920/sudden-accuracy-drop-when-training-lstm-or-gru-in-keras\n[2] https://stackoverflow.com/questions/37044600/sudden-drop-in-accuracy-while-training-a-deep-neural-net",
    "1666969": "# Tips: Decreasing Leaning Rate & Increasing Momentum Parameter\n\nIn my case, the following parameter tuning was effective, at least in some cases:\n\n* decreasing learning rate\n* increase beta1 for AdamW optimizer (which is similar to momentum for SGD)\n\nOne of the reason that the training process gets unstable is learning rate is too high.\nIf it is the case, either decreasing learning rate or increasing momentum parameter (or both) will mitigate this problem.",
    "1666979": "Look at the top image below.\nThe gray curve shows a sudden drop in mAP at the sixth epoch.\nAlso, in the middle figure, the learning loss is increased abruptly at this epoch.\nThis indicates that something is wrong with the learning (e.g., the gradient vanished/exploded, the structure of the trained model is broken, etc.).\nIn this case, you can try to reduce the learning rate or increase the momentum with your optimizer.\nFor the blue curve, I reduced the learning rate (0.01 -> 0.003) and increased the momentum (0.937 -> 0.979).\nCompared to the gray curve, the learning curve becomes more stable.\n\n<a href=\"https://ibb.co/QfyHPBH\"><img src=\"https://i.ibb.co/47hS8yS/Screen-Shot-2022-01-28-at-16-54-02.png\" alt=\"Screen-Shot-2022-01-28-at-16-54-02\" border=\"0\"></a>\n<a href=\"https://ibb.co/SyR1Vd7\"><img src=\"https://i.ibb.co/W5gdBPp/Screen-Shot-2022-01-28-at-16-54-17.png\" alt=\"Screen-Shot-2022-01-28-at-16-54-17\" border=\"0\"></a>\n<a href=\"https://ibb.co/gP9FCTZ\"><img src=\"https://i.ibb.co/CM8sF7K/Screen-Shot-2022-01-28-at-16-54-30.png\" alt=\"Screen-Shot-2022-01-28-at-16-54-30\" border=\"0\"></a>",
    "1667080": "how to get momentum value (i.e. 0.979)?",
    "1667129": "yienngxiong \n\nFrom this equation. This means the momentum is weighed x3 times.\n\n```\n>>> 1 - ((1 - 0.937) / 3)\n0.979\n```\n\n# Reference\n\n- Adam: https://pytorch.org/docs/stable/generated/torch.optim.Adam.html\n- SGD: https://pytorch.org/docs/stable/generated/torch.optim.SGD.html",
    "1667140": "As I am using AdamW optimizer, I specified beta1, but it's equivalent to momentum in SGD.",
    "1667167": "In other words, decreasing learning rate x1/3 and increasing momentum weight x3 means making the gradient of current step contributes x1/9 to the network parameter updates.\n\nThis is effective to weaken “too much learning from just looking at the current batch”, and making your leaning curve smoother.",
    "1667177": "I think this tuning is also effective to increase stability on small batch size. However because of high variance of BN statistics in small batch size, the effect must be limited.",
    "1667201": "the real reason for fluctuating loss curve is data.\n\nwhen the current batch of data is different from the previous one (or previous ones), the loss will also be different.\nin short, it is effects of outliers.\n\nlet me give you a few eamples. assume we are training a dog vs cat image classifier:\n1. if for one batch, there is only one class (e.g. batch of 256 images and all label are dog), then the model will learn to predict dog with much higher probability after the batch update. when the next batch contains cat images, the loss suddenly increases.\n\n2. for one batch, the label are 100% wrongly labelled., then the model will predict the opposite after the batch update.\nwhen the next batch contains correct label, the loss suddenly increases.\n\n----\n\nmometum, learning rate limiting gradient values prevents accidential wrong udpate, but they do do not solve the root cause of the problem. if the outlier ratio is low, you can use such methods.\n\nif you really want to find out the root cause, you have to think of a way to find out the outliers.",
    "1667223": "hengck23 Thanks for the comment.\n\nActually, the gray curve is obtained by adding 30% of negative samples (whereas the pink curve is 0% negative samples).\n\nThe data is sampled with a certain probability where I get sample ratio neg:pos=0.3:1.0 (I used full image datasets, then sample per every epoch to make data of this ratio.).\n\nSo, If I take it for granted what you say (the root cause is data), I guess the most probable cause of fluctuation is based on the randomness of sampling.\n\nI thought random sampling is better than fixed size sampling (because it can learns on better noise), but in this situation, do you think this strategy is bad?\n\n(Currently I don't know how to test if the randomness really matters, but let me think about this.)\n\n---\n\n## Additional Note\n\non gray curve, both positive and negative data are sampled with certain probability, and they are sampled with `repeat=True` (namely, the same data can be sampled more than once).\nOn the contrary, in the pink curve, each positive set is sampled only once (which is equivalent to `repeat=False`). So I guess the variance of sample is even bigger for the gray curve compared to pink one.",
    "1667242": "According to @hengck23 , tuning the learning rate and momentum won't solve the root cause.\n\nSee also this thread:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300960#1667201",
    "1667257": "i think there are some COTS that looks like background. this cuases confusion when learning.\n(also there are some blur small COTS in the ground truth, if you enlarge them in scale augmentation, there are may become bad samples. The larger COTS in the ground truth are not blur)\n\nwhen you do negative background sampling, you may sample background that looks like COTS.\nthe ratio affects the frequency of such occurrence, etc",
    "1667273": "\"mAP suddenly drops at epoch 6 (see gray curve on Fig1)\"\n\nmost kagglers will experience decrease up to 6,7,8 for using video 1 as validation.\n\nfor me although the map raises again, it finally will fall back again as I uses more train iterations. I use specialised augmentation and EMA for this.\n\n---\n\nthis can be explained as follows:\n- deep models learned general target object charateristics at first. (generalisation is good, loss curve falls rapidly)\n- then it starts to learn (and sometimes wrongly overfitted to) noise in the data (the loss of the good data becomes very low. only bad data has high loss and contribute more to backprop gradient)",
    "1667279": "I see. I guess the un-labeled positive COTS (which is discussed in other discussions) might also harm learning.",
    "1667289": "I see. It might be too fast judging the gray curve is failed.\n\nBTW, can I ask what EMA is?",
    "1667413": "# Related Paper about Unlabeled Data/Noisy Labels\n\n- [Positive-Unlabeled Data Purification in the Wild for Object Detection](https://openaccess.thecvf.com/content/CVPR2021/papers/Guo_Positive-Unlabeled_Data_Purification_in_the_Wild_for_Object_Detection_CVPR_2021_paper.pdf)\n- [A Simple Semi-Supervised Learning Framework for Object Detection](https://arxiv.org/pdf/2005.04757.pdf)\n- [End-to-End Semi-Supervised Object Detection with Soft Teacher](https://paperswithcode.com/paper/end-to-end-semi-supervised-object-detection)\n- [Unbiased Teacher for Semi-Supervised Object Detection](https://paperswithcode.com/paper/unbiased-teacher-for-semi-supervised-object-1)\n- [Few-Example Object Detection with Model Communication](https://arxiv.org/pdf/1706.08249.pdf)"
  },
  "source": "meta"
}