{
  "id": 231965,
  "title": "Need some advice for improve training params",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/231965",
  "author_name": "",
  "post_date": "2021-04-11T13:22:01.989922700Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>I made some config wrong ?</h1>\n<p>&nbsp;</p>\n<p>my best CV score is about  9.28  +-0.02 , LB is about 9.18 </p>\n<p>so <strong>I never reach CV score 9.33+</strong></p>\n<p>I read many efficientnet+unet notebook ,</p>\n<p>maybe you can take a look at my work </p>\n<p>and  tell me what I made mistake on architecture or params tuning ?  to get a higher CV score?</p>\n<p>&nbsp;</p>\n<h1>Here is my notebook:</h1>\n<p>&nbsp;</p>\n<p><strong>Preproccess tfrecord</strong>:<a href=\"https://www.kaggle.com/drzhuzhe/hubmap-tf-with-tpu-efficientunet-256-tfrecord\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/hubmap-tf-with-tpu-efficientunet-256-tfrecord</a><br>\n<strong>trainning</strong>: <a href=\"https://www.kaggle.com/drzhuzhe/hubmap-efficientnet-and-linknet-train\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/hubmap-efficientnet-and-linknet-train</a> thanks to wrrosa and vgarshin's work<br>\n<strong>infer</strong>: <a href=\"https://www.kaggle.com/drzhuzhe/efficientnet-linknet-or-unet\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/efficientnet-linknet-or-unet</a></p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<h2>Preproccess tfrecord:</h2>\n<p>&nbsp;</p>\n<ol>\n<li>reduce 1024 tile to 256 tiles</li>\n<li>shift every tile 512 </li>\n<li>create tfrecord for tpu</li>\n</ol>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<h2>Training :</h2>\n<p>&nbsp;</p>\n<ul>\n<li>what works</li>\n</ul>\n<ol>\n<li>add some augments like <strong>hue</strong> <strong>brightness</strong> <strong>saturation</strong> <strong>contrast</strong></li>\n<li><em>larger <strong>batchsize</strong>  my batchsize is <strong>1024</strong>  will significant speed up TPU trainning, but I don't know whether I set learning rate 5e-4 enough?</em></li>\n<li><strong>more epoch</strong>  some times training seem unstable suddenly more epoch may leads to an finally converge </li>\n</ol>\n<p>&nbsp;</p>\n<ul>\n<li>what doesn't work </li>\n</ul>\n<ol>\n<li>optimize.lookahead seems not improve </li>\n</ol>\n<p>&nbsp;</p>\n<ul>\n<li>what hurt </li>\n</ul>\n<ol>\n<li>jpeg_quality noise will hurt CV and LB score </li>\n<li><strong>smaller Learning rat</strong>  that 5e-4 like 5e-5 or <strong>minimal Learning rate</strong> after 1e-6 may hurt </li>\n</ol>",
  "messages": [
    {
      "id": "1270248",
      "postDate": "04/11/2021 13:22:01",
      "content": "<h1>I made some config wrong ?</h1>\n<p>&nbsp;</p>\n<p>my best CV score is about  9.28  +-0.02 , LB is about 9.18 </p>\n<p>so <strong>I never reach CV score 9.33+</strong></p>\n<p>I read many efficientnet+unet notebook ,</p>\n<p>maybe you can take a look at my work </p>\n<p>and  tell me what I made mistake on architecture or params tuning ?  to get a higher CV score?</p>\n<p>&nbsp;</p>\n<h1>Here is my notebook:</h1>\n<p>&nbsp;</p>\n<p><strong>Preproccess tfrecord</strong>:<a href=\"https://www.kaggle.com/drzhuzhe/hubmap-tf-with-tpu-efficientunet-256-tfrecord\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/hubmap-tf-with-tpu-efficientunet-256-tfrecord</a><br>\n<strong>trainning</strong>: <a href=\"https://www.kaggle.com/drzhuzhe/hubmap-efficientnet-and-linknet-train\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/hubmap-efficientnet-and-linknet-train</a> thanks to wrrosa and vgarshin's work<br>\n<strong>infer</strong>: <a href=\"https://www.kaggle.com/drzhuzhe/efficientnet-linknet-or-unet\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/efficientnet-linknet-or-unet</a></p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<h2>Preproccess tfrecord:</h2>\n<p>&nbsp;</p>\n<ol>\n<li>reduce 1024 tile to 256 tiles</li>\n<li>shift every tile 512 </li>\n<li>create tfrecord for tpu</li>\n</ol>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<h2>Training :</h2>\n<p>&nbsp;</p>\n<ul>\n<li>what works</li>\n</ul>\n<ol>\n<li>add some augments like <strong>hue</strong> <strong>brightness</strong> <strong>saturation</strong> <strong>contrast</strong></li>\n<li><em>larger <strong>batchsize</strong>  my batchsize is <strong>1024</strong>  will significant speed up TPU trainning, but I don't know whether I set learning rate 5e-4 enough?</em></li>\n<li><strong>more epoch</strong>  some times training seem unstable suddenly more epoch may leads to an finally converge </li>\n</ol>\n<p>&nbsp;</p>\n<ul>\n<li>what doesn't work </li>\n</ul>\n<ol>\n<li>optimize.lookahead seems not improve </li>\n</ol>\n<p>&nbsp;</p>\n<ul>\n<li>what hurt </li>\n</ul>\n<ol>\n<li>jpeg_quality noise will hurt CV and LB score </li>\n<li><strong>smaller Learning rat</strong>  that 5e-4 like 5e-5 or <strong>minimal Learning rate</strong> after 1e-6 may hurt </li>\n</ol>",
      "rawMarkdown": "# I made some config wrong ?\n\n&nbsp;\n\nmy best CV score is about  9.28  +-0.02 , LB is about 9.18 \n\nso **I never reach CV score 9.33+**\n\nI read many efficientnet+unet notebook ,\n\nmaybe you can take a look at my work \n\nand  tell me what I made mistake on architecture or params tuning ?  to get a higher CV score?\n\n&nbsp;\n\n# Here is my notebook: \n\n&nbsp;\n\n**Preproccess tfrecord**:https://www.kaggle.com/drzhuzhe/hubmap-tf-with-tpu-efficientunet-256-tfrecord\n**trainning**: https://www.kaggle.com/drzhuzhe/hubmap-efficientnet-and-linknet-train thanks to wrrosa and vgarshin's work\n**infer**: https://www.kaggle.com/drzhuzhe/efficientnet-linknet-or-unet\n\n&nbsp;\n\n&nbsp;\n\n## Preproccess tfrecord: \n\n&nbsp;\n\n1. reduce 1024 tile to 256 tiles\n2. shift every tile 512 \n3. create tfrecord for tpu\n\n&nbsp;\n\n&nbsp;\n\n## Training :\n\n&nbsp;\n\n- what works\n\n1. add some augments like **hue** **brightness** **saturation** **contrast**\n2. *larger **batchsize**  my batchsize is **1024**  will significant speed up TPU trainning, but I don't know whether I set learning rate 5e-4 enough?*\n3. **more epoch**  some times training seem unstable suddenly more epoch may leads to an finally converge \n\n&nbsp;\n\n- what doesn't work \n\n1. optimize.lookahead seems not improve \n\n&nbsp;\n\n- what hurt \n\n1. jpeg_quality noise will hurt CV and LB score \n2. **smaller Learning rat**  that 5e-4 like 5e-5 or **minimal Learning rate** after 1e-6 may hurt",
      "votes": null
    },
    {
      "id": "1270590",
      "postDate": "04/11/2021 19:20:43",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/drzhuzhe\" target=\"_blank\">@drzhuzhe</a> May'be try more and different things? Try out different losses and combinations of those. Try out different backbones. Try out may'be other frameworks than just Unet.</p>\n<p>Good luck.</p>",
      "rawMarkdown": "Hey @drzhuzhe May'be try more and different things? Try out different losses and combinations of those. Try out different backbones. Try out may'be other frameworks than just Unet.\n\nGood luck.",
      "votes": null
    },
    {
      "id": "1270794",
      "postDate": "04/12/2021 03:05:09",
      "content": "<ol>\n<li>[loss]: I try binaryCrossEntropy but seems not benefit</li>\n<li>[backbone]: I try efficientbet b4 but it make training 40% slower but not give an improve on CV too</li>\n<li>[framework]: I try linknet too, but I use as same optimize params as Unet seems  as Unet seems not <br>\nwork well…<br>\nwhat I most suspect is my Batch Size and learning rate doesn't config right </li>\n</ol>",
      "rawMarkdown": "1. [loss]: I try binaryCrossEntropy but seems not benefit\n\n2. [backbone]: I try efficientbet b4 but it make training 40% slower but not give an improve on CV too\n\n3. [framework]: I try linknet too, but I use as same optimize params as Unet seems  as Unet seems not \nwork well...\n\nwhat I most suspect is my Batch Size and learning rate doesn't config right",
      "votes": null
    },
    {
      "id": "1270919",
      "postDate": "04/12/2021 06:43:52",
      "content": "<p>About batchsize: From my opinion, 1024 batchsize is too large. Usually 32 or 64 is enough.<br>\nAbout backbone: efficientb3 can also get 0.92x results on LB. <br>\nAbout epochs: Try to use pretrain Encode weights from imagenet. Usually it will converge in 30 epochs. In my case, I use AdamW optimizer. And set initial lr to 1e-4.</p>",
      "rawMarkdown": "About batchsize: From my opinion, 1024 batchsize is too large. Usually 32 or 64 is enough.\nAbout backbone: efficientb3 can also get 0.92x results on LB. \nAbout epochs: Try to use pretrain Encode weights from imagenet. Usually it will converge in 30 epochs. In my case, I use AdamW optimizer. And set initial lr to 1e-4.",
      "votes": null
    },
    {
      "id": "1272162",
      "postDate": "04/13/2021 09:24:35",
      "content": "<p>Big batch size requires high learning rate. </p>\n<p>There is an anology  - imagine a pool table with 3 billiard balls randomly moving, and we will draw a center of mass of these balls. What trajectory these center of mass will have? With 3 balls that will be a trajectory with big oscillations. But what happends if there will be 1000 balls? Center of mass will not moving, because random directions of balls will mutually compensate each other. </p>\n<p>Same situation with big batch size - gradients of different samples will mutually compensate each other, so final gradient will be smaller, so higher learning rate is required</p>",
      "rawMarkdown": "Big batch size requires high learning rate. \n\nThere is an anology  - imagine a pool table with 3 billiard balls randomly moving, and we will draw a center of mass of these balls. What trajectory these center of mass will have? With 3 balls that will be a trajectory with big oscillations. But what happends if there will be 1000 balls? Center of mass will not moving, because random directions of balls will mutually compensate each other. \n\nSame situation with big batch size - gradients of different samples will mutually compensate each other, so final gradient will be smaller, so higher learning rate is required",
      "votes": null
    },
    {
      "id": "1272188",
      "postDate": "04/13/2021 09:48:47",
      "content": "<p>Really awesome explain, dude</p>\n<p>But  I still have two question make me very confuse </p>\n<ol>\n<li><p>I set my optimizer <strong>a learning rate reduce</strong></p>\n<p>with factor 0.1 and minimal: 0.01 * my initial learning rate</p>\n<p>does my <strong>minimal learning rate</strong> need scaler up with initial learning rate ?</p></li>\n<li><p>Does 1024 such a large batch size really hurt metrics score? </p>\n<p>some paper said <strong>small batch make more Regularization to optimizer</strong>, so larger batch may <br>\nin disadvantage</p>\n<p>but other paper said large batch make gradient change more slowly, so what I just need is only training more epoch till convergence  ?</p></li>\n</ol>\n<p>Due to my CV score is too close, I cannot figure out which change in my params make sense </p>",
      "rawMarkdown": "Really awesome explain, dude\n\nBut  I still have two question make me very confuse \n\n1. I set my optimizer **a learning rate reduce**\n\n with factor 0.1 and minimal: 0.01 * my initial learning rate\n\n does my **minimal learning rate** need scaler up with initial learning rate ?\n\n\n2. Does 1024 such a large batch size really hurt metrics score? \n\n some paper said **small batch make more Regularization to optimizer**, so larger batch may \nin disadvantage\n\n but other paper said large batch make gradient change more slowly, so what I just need is only training more epoch till convergence  ?\n\n\nDue to my CV score is too close, I cannot figure out which change in my params make sense",
      "votes": null
    },
    {
      "id": "1272245",
      "postDate": "04/13/2021 10:44:27",
      "content": "<p>Yes, small batch size can make regularization. With small batch size, network need to adapt to \"unseen\" samples on each new batch. So features, that badly generalize to unseen data, will not survive with small batch size (big gradient from new batch will destroy these features).</p>",
      "rawMarkdown": "Yes, small batch size can make regularization. With small batch size, network need to adapt to \"unseen\" samples on each new batch. So features, that badly generalize to unseen data, will not survive with small batch size (big gradient from new batch will destroy these features).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1270590,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "04/11/2021 19:20:43",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/drzhuzhe\" target=\"_blank\">@drzhuzhe</a> May'be try more and different things? Try out different losses and combinations of those. Try out different backbones. Try out may'be other frameworks than just Unet.</p>\n<p>Good luck.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1270794,
          "author_name": "drzhuzhe",
          "author_url": "",
          "post_date": "04/12/2021 03:05:09",
          "content": "<ol>\n<li>[loss]: I try binaryCrossEntropy but seems not benefit</li>\n<li>[backbone]: I try efficientbet b4 but it make training 40% slower but not give an improve on CV too</li>\n<li>[framework]: I try linknet too, but I use as same optimize params as Unet seems  as Unet seems not <br>\nwork well…<br>\nwhat I most suspect is my Batch Size and learning rate doesn't config right </li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1270919,
      "author_name": "shinewine",
      "author_url": "",
      "post_date": "04/12/2021 06:43:52",
      "content": "<p>About batchsize: From my opinion, 1024 batchsize is too large. Usually 32 or 64 is enough.<br>\nAbout backbone: efficientb3 can also get 0.92x results on LB. <br>\nAbout epochs: Try to use pretrain Encode weights from imagenet. Usually it will converge in 30 epochs. In my case, I use AdamW optimizer. And set initial lr to 1e-4.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1272162,
      "author_name": "zavodrobotov",
      "author_url": "",
      "post_date": "04/13/2021 09:24:35",
      "content": "<p>Big batch size requires high learning rate. </p>\n<p>There is an anology  - imagine a pool table with 3 billiard balls randomly moving, and we will draw a center of mass of these balls. What trajectory these center of mass will have? With 3 balls that will be a trajectory with big oscillations. But what happends if there will be 1000 balls? Center of mass will not moving, because random directions of balls will mutually compensate each other. </p>\n<p>Same situation with big batch size - gradients of different samples will mutually compensate each other, so final gradient will be smaller, so higher learning rate is required</p>",
      "votes": null,
      "replies": [
        {
          "id": 1272188,
          "author_name": "drzhuzhe",
          "author_url": "",
          "post_date": "04/13/2021 09:48:47",
          "content": "<p>Really awesome explain, dude</p>\n<p>But  I still have two question make me very confuse </p>\n<ol>\n<li><p>I set my optimizer <strong>a learning rate reduce</strong></p>\n<p>with factor 0.1 and minimal: 0.01 * my initial learning rate</p>\n<p>does my <strong>minimal learning rate</strong> need scaler up with initial learning rate ?</p></li>\n<li><p>Does 1024 such a large batch size really hurt metrics score? </p>\n<p>some paper said <strong>small batch make more Regularization to optimizer</strong>, so larger batch may <br>\nin disadvantage</p>\n<p>but other paper said large batch make gradient change more slowly, so what I just need is only training more epoch till convergence  ?</p></li>\n</ol>\n<p>Due to my CV score is too close, I cannot figure out which change in my params make sense </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1272245,
          "author_name": "zavodrobotov",
          "author_url": "",
          "post_date": "04/13/2021 10:44:27",
          "content": "<p>Yes, small batch size can make regularization. With small batch size, network need to adapt to \"unseen\" samples on each new batch. So features, that badly generalize to unseen data, will not survive with small batch size (big gradient from new batch will destroy these features).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1270248": "# I made some config wrong ?\n\n&nbsp;\n\nmy best CV score is about  9.28  +-0.02 , LB is about 9.18 \n\nso **I never reach CV score 9.33+**\n\nI read many efficientnet+unet notebook ,\n\nmaybe you can take a look at my work \n\nand  tell me what I made mistake on architecture or params tuning ?  to get a higher CV score?\n\n&nbsp;\n\n# Here is my notebook: \n\n&nbsp;\n\n**Preproccess tfrecord**:https://www.kaggle.com/drzhuzhe/hubmap-tf-with-tpu-efficientunet-256-tfrecord\n**trainning**: https://www.kaggle.com/drzhuzhe/hubmap-efficientnet-and-linknet-train thanks to wrrosa and vgarshin's work\n**infer**: https://www.kaggle.com/drzhuzhe/efficientnet-linknet-or-unet\n\n&nbsp;\n\n&nbsp;\n\n## Preproccess tfrecord: \n\n&nbsp;\n\n1. reduce 1024 tile to 256 tiles\n2. shift every tile 512 \n3. create tfrecord for tpu\n\n&nbsp;\n\n&nbsp;\n\n## Training :\n\n&nbsp;\n\n- what works\n\n1. add some augments like **hue** **brightness** **saturation** **contrast**\n2. *larger **batchsize**  my batchsize is **1024**  will significant speed up TPU trainning, but I don't know whether I set learning rate 5e-4 enough?*\n3. **more epoch**  some times training seem unstable suddenly more epoch may leads to an finally converge \n\n&nbsp;\n\n- what doesn't work \n\n1. optimize.lookahead seems not improve \n\n&nbsp;\n\n- what hurt \n\n1. jpeg_quality noise will hurt CV and LB score \n2. **smaller Learning rat**  that 5e-4 like 5e-5 or **minimal Learning rate** after 1e-6 may hurt",
    "1270590": "Hey @drzhuzhe May'be try more and different things? Try out different losses and combinations of those. Try out different backbones. Try out may'be other frameworks than just Unet.\n\nGood luck.",
    "1270794": "1. [loss]: I try binaryCrossEntropy but seems not benefit\n\n2. [backbone]: I try efficientbet b4 but it make training 40% slower but not give an improve on CV too\n\n3. [framework]: I try linknet too, but I use as same optimize params as Unet seems  as Unet seems not \nwork well...\n\nwhat I most suspect is my Batch Size and learning rate doesn't config right",
    "1270919": "About batchsize: From my opinion, 1024 batchsize is too large. Usually 32 or 64 is enough.\nAbout backbone: efficientb3 can also get 0.92x results on LB. \nAbout epochs: Try to use pretrain Encode weights from imagenet. Usually it will converge in 30 epochs. In my case, I use AdamW optimizer. And set initial lr to 1e-4.",
    "1272162": "Big batch size requires high learning rate. \n\nThere is an anology  - imagine a pool table with 3 billiard balls randomly moving, and we will draw a center of mass of these balls. What trajectory these center of mass will have? With 3 balls that will be a trajectory with big oscillations. But what happends if there will be 1000 balls? Center of mass will not moving, because random directions of balls will mutually compensate each other. \n\nSame situation with big batch size - gradients of different samples will mutually compensate each other, so final gradient will be smaller, so higher learning rate is required",
    "1272188": "Really awesome explain, dude\n\nBut  I still have two question make me very confuse \n\n1. I set my optimizer **a learning rate reduce**\n\n with factor 0.1 and minimal: 0.01 * my initial learning rate\n\n does my **minimal learning rate** need scaler up with initial learning rate ?\n\n\n2. Does 1024 such a large batch size really hurt metrics score? \n\n some paper said **small batch make more Regularization to optimizer**, so larger batch may \nin disadvantage\n\n but other paper said large batch make gradient change more slowly, so what I just need is only training more epoch till convergence  ?\n\n\nDue to my CV score is too close, I cannot figure out which change in my params make sense",
    "1272245": "Yes, small batch size can make regularization. With small batch size, network need to adapt to \"unseen\" samples on each new batch. So features, that badly generalize to unseen data, will not survive with small batch size (big gradient from new batch will destroy these features)."
  },
  "source": "meta"
}