{
  "id": 288677,
  "title": "Deep Double Decent: More epoch More better?",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/288677",
  "author_name": "",
  "post_date": "2021-11-18T06:40:35.319533600Z",
  "votes": 7,
  "comment_count": 11,
  "views": 0,
  "content": "<p>This is the general topic about the number of epoch we should choose.<br>\nI learnt the concept of double decent.<br>\nAccording to this, if the model has enough complexity, more epoch is more better.<br>\nIf you have knowledge about method of decision of it, let's discuss it:)</p>\n<p>Paper: <a href=\"https://arxiv.org/abs/1912.02292\" target=\"_blank\">https://arxiv.org/abs/1912.02292</a></p>\n<p>English blog: <a href=\"https://openai.com/blog/deep-double-descent/\" target=\"_blank\">https://openai.com/blog/deep-double-descent/</a></p>\n<p>Japanese blog: <a href=\"https://zenn.dev/hnishi/articles/20201015-deep-double-descent\" target=\"_blank\">https://zenn.dev/hnishi/articles/20201015-deep-double-descent</a></p>",
  "messages": [
    {
      "id": "1586606",
      "postDate": "11/18/2021 06:40:35",
      "content": "<p>This is the general topic about the number of epoch we should choose.<br>\nI learnt the concept of double decent.<br>\nAccording to this, if the model has enough complexity, more epoch is more better.<br>\nIf you have knowledge about method of decision of it, let's discuss it:)</p>\n<p>Paper: <a href=\"https://arxiv.org/abs/1912.02292\" target=\"_blank\">https://arxiv.org/abs/1912.02292</a></p>\n<p>English blog: <a href=\"https://openai.com/blog/deep-double-descent/\" target=\"_blank\">https://openai.com/blog/deep-double-descent/</a></p>\n<p>Japanese blog: <a href=\"https://zenn.dev/hnishi/articles/20201015-deep-double-descent\" target=\"_blank\">https://zenn.dev/hnishi/articles/20201015-deep-double-descent</a></p>",
      "rawMarkdown": "This is the general topic about the number of epoch we should choose.\nI learnt the concept of double decent.\nAccording to this, if the model has enough complexity, more epoch is more better.\nIf you have knowledge about method of decision of it, let's discuss it:)\n\nPaper: https://arxiv.org/abs/1912.02292\n\nEnglish blog: https://openai.com/blog/deep-double-descent/\n\nJapanese blog: https://zenn.dev/hnishi/articles/20201015-deep-double-descent",
      "votes": null
    },
    {
      "id": "1592057",
      "postDate": "11/22/2021 20:53:30",
      "content": "<p>I stop seeing improvements pretty quickly. This is 5 different folds trained for 10k iterations.<br>\n<img src=\"https://raw.githubusercontent.com/slawekslex/random/main/Screenshot%202021-11-22%20at%2021.29.36.png\" alt=\"\"></p>",
      "rawMarkdown": "I stop seeing improvements pretty quickly. This is 5 different folds trained for 10k iterations.\n![](https://raw.githubusercontent.com/slawekslex/random/main/Screenshot%202021-11-22%20at%2021.29.36.png)",
      "votes": null
    },
    {
      "id": "1592664",
      "postDate": "11/23/2021 08:41:31",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a>, can you tell are you using any kind of learning rate schedule?<br>\nI am experimenting with WarmupCosineLR and I observe the same behavior too (model stops improving after a certain number of iterations) </p>",
      "rawMarkdown": "Hi @slawekbiel, can you tell are you using any kind of learning rate schedule?\nI am experimenting with WarmupCosineLR and I observe the same behavior too (model stops improving after a certain number of iterations)",
      "votes": null
    },
    {
      "id": "1592676",
      "postDate": "11/23/2021 08:55:43",
      "content": "<p>This was detectron’s WarmupMultistepLR</p>",
      "rawMarkdown": "This was detectron’s WarmupMultistepLR",
      "votes": null
    },
    {
      "id": "1598139",
      "postDate": "11/28/2021 09:31:21",
      "content": "<p>What is your batch_size, I can only train on 4 or 2 due to memory constraint</p>",
      "rawMarkdown": "What is your batch_size, I can only train on 4 or 2 due to memory constraint",
      "votes": null
    },
    {
      "id": "1598179",
      "postDate": "11/28/2021 09:57:34",
      "content": "<p>Same here, batch size of 2</p>",
      "rawMarkdown": "Same here, batch size of 2",
      "votes": null
    },
    {
      "id": "1598283",
      "postDate": "11/28/2021 10:53:28",
      "content": "<p>the thing is probably that you got a better model which can detect more than labelled annotations. so if test dataset is not well labelled, then LB  make nonsense.  </p>\n<p>getting a good LB just an optimization for metric. </p>",
      "rawMarkdown": "the thing is probably that you got a better model which can detect more than labelled annotations. so if test dataset is not well labelled, then LB  make nonsense.  \n\ngetting a good LB just an optimization for metric.",
      "votes": null
    },
    {
      "id": "1601923",
      "postDate": "12/01/2021 16:41:44",
      "content": "<p>Hi, I see your 5 different folds validation score is very different. May I ask it is correlated to LB?</p>\n<p>I.e fold orange would have better score on LB than others</p>",
      "rawMarkdown": "Hi, I see your 5 different folds validation score is very different. May I ask it is correlated to LB?\n\nI.e fold orange would have better score on LB than others",
      "votes": null
    },
    {
      "id": "1602028",
      "postDate": "12/01/2021 17:52:14",
      "content": "<p>I recall reading this paper when it came out. My only comment is that x-axis on the figures in the paper represent model complexity, whereas hue = epochs. So it's not a regular loss curve…</p>",
      "rawMarkdown": "I recall reading this paper when it came out. My only comment is that x-axis on the figures in the paper represent model complexity, whereas hue = epochs. So it's not a regular loss curve...",
      "votes": null
    },
    {
      "id": "1602044",
      "postDate": "12/01/2021 18:01:37",
      "content": "<p><a href=\"https://www.kaggle.com/ptran1203\" target=\"_blank\">@ptran1203</a> the LB scores for five folds were in the span [.307, .316] with the orange one being the highest.</p>",
      "rawMarkdown": "ptran1203 the LB scores for five folds were in the span [.307, .316] with the orange one being the highest.",
      "votes": null
    },
    {
      "id": "1611677",
      "postDate": "12/08/2021 06:52:28",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a>, I see you are evaluating after every 1k iterations rather than each epoch. Of course, it reduces the complete experiment time, but was just curious to know was there any other reason behind this.</p>",
      "rawMarkdown": "Hi @slawekbiel, I see you are evaluating after every 1k iterations rather than each epoch. Of course, it reduces the complete experiment time, but was just curious to know was there any other reason behind this.",
      "votes": null
    },
    {
      "id": "1611850",
      "postDate": "12/08/2021 10:31:12",
      "content": "<p><a href=\"https://www.kaggle.com/atharvaingle\" target=\"_blank\">@atharvaingle</a> no particular reason, just seemed good enough.</p>",
      "rawMarkdown": "atharvaingle no particular reason, just seemed good enough.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1592057,
      "author_name": "slawekbiel",
      "author_url": "",
      "post_date": "11/22/2021 20:53:30",
      "content": "<p>I stop seeing improvements pretty quickly. This is 5 different folds trained for 10k iterations.<br>\n<img src=\"https://raw.githubusercontent.com/slawekslex/random/main/Screenshot%202021-11-22%20at%2021.29.36.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1592664,
          "author_name": "atharvaingle",
          "author_url": "",
          "post_date": "11/23/2021 08:41:31",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a>, can you tell are you using any kind of learning rate schedule?<br>\nI am experimenting with WarmupCosineLR and I observe the same behavior too (model stops improving after a certain number of iterations) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1592676,
          "author_name": "slawekbiel",
          "author_url": "",
          "post_date": "11/23/2021 08:55:43",
          "content": "<p>This was detectron’s WarmupMultistepLR</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1598139,
          "author_name": "ptran1203",
          "author_url": "",
          "post_date": "11/28/2021 09:31:21",
          "content": "<p>What is your batch_size, I can only train on 4 or 2 due to memory constraint</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1598179,
          "author_name": "slawekbiel",
          "author_url": "",
          "post_date": "11/28/2021 09:57:34",
          "content": "<p>Same here, batch size of 2</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1601923,
          "author_name": "ptran1203",
          "author_url": "",
          "post_date": "12/01/2021 16:41:44",
          "content": "<p>Hi, I see your 5 different folds validation score is very different. May I ask it is correlated to LB?</p>\n<p>I.e fold orange would have better score on LB than others</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1602044,
          "author_name": "slawekbiel",
          "author_url": "",
          "post_date": "12/01/2021 18:01:37",
          "content": "<p><a href=\"https://www.kaggle.com/ptran1203\" target=\"_blank\">@ptran1203</a> the LB scores for five folds were in the span [.307, .316] with the orange one being the highest.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1611677,
          "author_name": "atharvaingle",
          "author_url": "",
          "post_date": "12/08/2021 06:52:28",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a>, I see you are evaluating after every 1k iterations rather than each epoch. Of course, it reduces the complete experiment time, but was just curious to know was there any other reason behind this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1611850,
          "author_name": "slawekbiel",
          "author_url": "",
          "post_date": "12/08/2021 10:31:12",
          "content": "<p><a href=\"https://www.kaggle.com/atharvaingle\" target=\"_blank\">@atharvaingle</a> no particular reason, just seemed good enough.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1598283,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "11/28/2021 10:53:28",
      "content": "<p>the thing is probably that you got a better model which can detect more than labelled annotations. so if test dataset is not well labelled, then LB  make nonsense.  </p>\n<p>getting a good LB just an optimization for metric. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1602028,
      "author_name": "authman",
      "author_url": "",
      "post_date": "12/01/2021 17:52:14",
      "content": "<p>I recall reading this paper when it came out. My only comment is that x-axis on the figures in the paper represent model complexity, whereas hue = epochs. So it's not a regular loss curve…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1586606": "This is the general topic about the number of epoch we should choose.\nI learnt the concept of double decent.\nAccording to this, if the model has enough complexity, more epoch is more better.\nIf you have knowledge about method of decision of it, let's discuss it:)\n\nPaper: https://arxiv.org/abs/1912.02292\n\nEnglish blog: https://openai.com/blog/deep-double-descent/\n\nJapanese blog: https://zenn.dev/hnishi/articles/20201015-deep-double-descent",
    "1592057": "I stop seeing improvements pretty quickly. This is 5 different folds trained for 10k iterations.\n![](https://raw.githubusercontent.com/slawekslex/random/main/Screenshot%202021-11-22%20at%2021.29.36.png)",
    "1592664": "Hi @slawekbiel, can you tell are you using any kind of learning rate schedule?\nI am experimenting with WarmupCosineLR and I observe the same behavior too (model stops improving after a certain number of iterations)",
    "1592676": "This was detectron’s WarmupMultistepLR",
    "1598139": "What is your batch_size, I can only train on 4 or 2 due to memory constraint",
    "1598179": "Same here, batch size of 2",
    "1598283": "the thing is probably that you got a better model which can detect more than labelled annotations. so if test dataset is not well labelled, then LB  make nonsense.  \n\ngetting a good LB just an optimization for metric.",
    "1601923": "Hi, I see your 5 different folds validation score is very different. May I ask it is correlated to LB?\n\nI.e fold orange would have better score on LB than others",
    "1602028": "I recall reading this paper when it came out. My only comment is that x-axis on the figures in the paper represent model complexity, whereas hue = epochs. So it's not a regular loss curve...",
    "1602044": "ptran1203 the LB scores for five folds were in the span [.307, .316] with the orange one being the highest.",
    "1611677": "Hi @slawekbiel, I see you are evaluating after every 1k iterations rather than each epoch. Of course, it reduces the complete experiment time, but was just curious to know was there any other reason behind this.",
    "1611850": "atharvaingle no particular reason, just seemed good enough."
  },
  "source": "meta"
}