{
  "id": 342159,
  "title": "Model training loss not decreasing, need tips!",
  "url": "/competitions/hubmap-organ-segmentation/discussion/342159",
  "author_name": "somuSan",
  "post_date": "2022-08-05T17:38:10.575000",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have seen people in the forum getting good scores with <code>Efficientnet b4</code> + <code>Unet</code> but when I'm submitting my inference notebook its giving me very low score [ ~0.1]. And im unable to figure out what are the reason. <br>\n<strong>It would be great if anyone can suggest some ways resolve this problem.</strong></p>\n<ul>\n<li><p>Trained with only <code>Dice loss</code>, No augmentations, <code>256x256</code> image size, <code>tiled data</code>, <code>batch size=8</code> , <code>lr=1e-4</code>, <code>LR scheduler = CosineAnnealingWarmRestarts</code>, <code>opt=Adam</code>, submitting after transposing.</p></li>\n<li><p><strong>One of the reason can be that, my model loss is not decreasing much, it is staying in <code>&gt;=0.95</code> range.</strong>  But I dont really understand why is that happening. It would be very helpful if anyone can suggest some tips to overcome this. I might be trying large batch size and augmentations.</p></li>\n</ul>\n<p>Below is the plot of training loss, validation loss, val dice score and val jaccard score over epoch.</p>\n<p><img src=\"https://i.imgur.com/4QOkx4T.png\"></p>\n<p>Below is my model config:</p>\n<pre><code>CONFIG = {\n    \"in_channels\" :3,\n    \"num_classes\": 1,\n    \"BATCH_SIZE\" : 8,\n    \"NUM_EPOCHS\" : 50,\n    \"n_accumulate\": 1,\n    \"competition\": \"HuBMAP-Kaggle\", # HuBMAP-Kaggle\n    \"model_name\": \"Efficientnet_B4_Unet\",\n    \"LEARNING_RATE\": 1e-4,\n    \"DEVICE\": \"cuda\" if torch.cuda.is_available() else \"cpu\", \n    \"AUG\": \"No\",\n    \"SEED\": 42,\n    \"opt\": 'Adam',\n    \"Normalization\": \"L2\",\n    \"img_size\": (256,256),\n    \"scheduler\": \"CosineAnnealingWarmRestarts\", #'CosineAnnealingLR',\n    \"warmup_epochs\": 10,\n    \"folds_to_run\": [0],\n    \"min_lr\": 1e-6,\n    \"T_max\": 500, \n    \"T_0\": 20\n}\n</code></pre>",
  "messages": [
    {
      "id": 1886321,
      "postDate": "2022-08-05T17:38:10.577Z",
      "content": "<p>I have seen people in the forum getting good scores with <code>Efficientnet b4</code> + <code>Unet</code> but when I'm submitting my inference notebook its giving me very low score [ ~0.1]. And im unable to figure out what are the reason. <br>\n<strong>It would be great if anyone can suggest some ways resolve this problem.</strong></p>\n<ul>\n<li><p>Trained with only <code>Dice loss</code>, No augmentations, <code>256x256</code> image size, <code>tiled data</code>, <code>batch size=8</code> , <code>lr=1e-4</code>, <code>LR scheduler = CosineAnnealingWarmRestarts</code>, <code>opt=Adam</code>, submitting after transposing.</p></li>\n<li><p><strong>One of the reason can be that, my model loss is not decreasing much, it is staying in <code>&gt;=0.95</code> range.</strong>  But I dont really understand why is that happening. It would be very helpful if anyone can suggest some tips to overcome this. I might be trying large batch size and augmentations.</p></li>\n</ul>\n<p>Below is the plot of training loss, validation loss, val dice score and val jaccard score over epoch.</p>\n<p><img src=\"https://i.imgur.com/4QOkx4T.png\"></p>\n<p>Below is my model config:</p>\n<pre><code>CONFIG = {\n    \"in_channels\" :3,\n    \"num_classes\": 1,\n    \"BATCH_SIZE\" : 8,\n    \"NUM_EPOCHS\" : 50,\n    \"n_accumulate\": 1,\n    \"competition\": \"HuBMAP-Kaggle\", # HuBMAP-Kaggle\n    \"model_name\": \"Efficientnet_B4_Unet\",\n    \"LEARNING_RATE\": 1e-4,\n    \"DEVICE\": \"cuda\" if torch.cuda.is_available() else \"cpu\", \n    \"AUG\": \"No\",\n    \"SEED\": 42,\n    \"opt\": 'Adam',\n    \"Normalization\": \"L2\",\n    \"img_size\": (256,256),\n    \"scheduler\": \"CosineAnnealingWarmRestarts\", #'CosineAnnealingLR',\n    \"warmup_epochs\": 10,\n    \"folds_to_run\": [0],\n    \"min_lr\": 1e-6,\n    \"T_max\": 500, \n    \"T_0\": 20\n}\n</code></pre>",
      "rawMarkdown": "I have seen people in the forum getting good scores with `Efficientnet b4` + `Unet` but when I'm submitting my inference notebook its giving me very low score [ ~0.1]. And im unable to figure out what are the reason. \n**It would be great if anyone can suggest some ways resolve this problem.**\n\n\n\n- Trained with only `Dice loss`, No augmentations, `256x256` image size, `tiled data`, `batch size=8` , `lr=1e-4`, `LR scheduler = CosineAnnealingWarmRestarts`, `opt=Adam`, submitting after transposing.\n\n- **One of the reason can be that, my model loss is not decreasing much, it is staying in `>=0.95` range.**  But I dont really understand why is that happening. It would be very helpful if anyone can suggest some tips to overcome this. I might be trying large batch size and augmentations.\n\nBelow is the plot of training loss, validation loss, val dice score and val jaccard score over epoch.\n\n<img src=\"https://i.imgur.com/4QOkx4T.png\">\n\nBelow is my model config:\n```\nCONFIG = {\n    \"in_channels\" :3,\n    \"num_classes\": 1,\n    \"BATCH_SIZE\" : 8,\n    \"NUM_EPOCHS\" : 50,\n    \"n_accumulate\": 1,\n    \"competition\": \"HuBMAP-Kaggle\", # HuBMAP-Kaggle\n    \"model_name\": \"Efficientnet_B4_Unet\",\n    \"LEARNING_RATE\": 1e-4,\n    \"DEVICE\": \"cuda\" if torch.cuda.is_available() else \"cpu\", \n    \"AUG\": \"No\",\n    \"SEED\": 42,\n    \"opt\": 'Adam',\n    \"Normalization\": \"L2\",\n    \"img_size\": (256,256),\n    \"scheduler\": \"CosineAnnealingWarmRestarts\", #'CosineAnnealingLR',\n    \"warmup_epochs\": 10,\n    \"folds_to_run\": [0],\n    \"min_lr\": 1e-6,\n    \"T_max\": 500, \n    \"T_0\": 20\n}\n```",
      "votes": 3
    },
    {
      "id": 1886334,
      "postDate": "2022-08-05T17:57:04.093Z",
      "content": "<p>discard dice loss and train with BCE first.</p>",
      "rawMarkdown": "discard dice loss and train with BCE first.",
      "votes": 1,
      "replies": [
        {
          "id": 1886409,
          "postDate": "2022-08-05T19:34:23.377Z",
          "content": "<p>Sure, I will try shifting to BCE.<br>\nI started with dice because people in the forum were suggesting its better to use Dice or Dice+BCE than only BCE . As,</p>\n<blockquote>\n  <p>this[BCE] still isn't the best loss since we have a lof of background pixels (0 pixels) in comparison with FTU pixels. Thus, we need a solution to fix this imbalance (otherwise the model will predict 0 all the time).</p>\n</blockquote>",
          "rawMarkdown": "Sure, I will try shifting to BCE.\nI started with dice because people in the forum were suggesting its better to use Dice or Dice+BCE than only BCE . As,\n> this[BCE] still isn't the best loss since we have a lof of background pixels (0 pixels) in comparison with FTU pixels. Thus, we need a solution to fix this imbalance (otherwise the model will predict 0 all the time)."
        },
        {
          "id": 1886666,
          "postDate": "2022-08-06T04:16:08.267Z",
          "content": "<p>there is not absoute conclusion in data science, everything will have to depends on the data. e.g.</p>\n<p>\"… Thus, we need a solution to fix this imbalance\". this may be true but you can always find data that don't have to do data balancing. for example, lets create synthetic data of size 10000x10000 of np.zeros(). then just random select one pixel and colored it with random color. the ground truth mask is 1 if pixel is non-zero, zero if otherwise. in this synthetic case, +ve and -ve ratio is highly imbalance. but signal strength is very strong. in this case just BCE will work.</p>\n<p>i suggest try BCE is because it is easier to work with. it may not give you best results, but you will get some results.<br>\nthe loss landscape of BCE is smoother and you are less likey to get stuck in poor solution.</p>\n<hr>\n<p>from the other forum discussion and public kernel, it is obvious that unet-efficientnet4 will at least get LB of 0.55, local validation should higher. if there is no bug in your code,data,model, then it must be your learning process. And when it comes to learning, things that can go wrong are:</p>\n<ul>\n<li>loss function</li>\n<li>optimizer and hyperparameters. </li>\n</ul>\n<p>always remember there is only 3 things that you need to consider when debugging algorithm (not debugging code):</p>\n<ol>\n<li>data</li>\n<li>model</li>\n<li>learning process (optimisation process)</li>\n</ol>\n<hr>\n<p>i always start with something simple, e.g. simple loss, remove scheduler, remove augmentation, ….. then advance to more complicated or enhanced methdos later.</p>",
          "rawMarkdown": "there is not absoute conclusion in data science, everything will have to depends on the data. e.g.\n\n\"... Thus, we need a solution to fix this imbalance\". this may be true but you can always find data that don't have to do data balancing. for example, lets create synthetic data of size 10000x10000 of np.zeros(). then just random select one pixel and colored it with random color. the ground truth mask is 1 if pixel is non-zero, zero if otherwise. in this synthetic case, +ve and -ve ratio is highly imbalance. but signal strength is very strong. in this case just BCE will work.\n\ni suggest try BCE is because it is easier to work with. it may not give you best results, but you will get some results.\nthe loss landscape of BCE is smoother and you are less likey to get stuck in poor solution.\n\n---\n\nfrom the other forum discussion and public kernel, it is obvious that unet-efficientnet4 will at least get LB of 0.55, local validation should higher. if there is no bug in your code,data,model, then it must be your learning process. And when it comes to learning, things that can go wrong are:\n\n- loss function\n- optimizer and hyperparameters. \n\nalways remember there is only 3 things that you need to consider when debugging algorithm (not debugging code):\n1. data\n2. model\n3. learning process (optimisation process)\n\n--- \n\ni always start with something simple, e.g. simple loss, remove scheduler, remove augmentation, ..... then advance to more complicated or enhanced methdos later.",
          "votes": 3
        },
        {
          "id": 1886759,
          "postDate": "2022-08-06T06:14:17.730Z",
          "content": "<p>Thank you so much for this great advice.<br>\nI will try to train my baseline using more minimal methods and use more advanced methods as I move along. I should have given BCE a try first then move on to trying something more complicated. Thank you so much for pointing that out.</p>",
          "rawMarkdown": "Thank you so much for this great advice.\nI will try to train my baseline using more minimal methods and use more advanced methods as I move along. I should have given BCE a try first then move on to trying something more complicated. Thank you so much for pointing that out."
        }
      ]
    },
    {
      "id": 1886360,
      "postDate": "2022-08-05T18:19:41.743Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1886334,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-08-05T17:57:04.093000",
      "content": "<p>discard dice loss and train with BCE first.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1886409,
          "author_name": "somuSan",
          "author_url": "",
          "post_date": "2022-08-05T19:34:23.377000",
          "content": "<p>Sure, I will try shifting to BCE.<br>\nI started with dice because people in the forum were suggesting its better to use Dice or Dice+BCE than only BCE . As,</p>\n<blockquote>\n  <p>this[BCE] still isn't the best loss since we have a lof of background pixels (0 pixels) in comparison with FTU pixels. Thus, we need a solution to fix this imbalance (otherwise the model will predict 0 all the time).</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1886666,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-08-06T04:16:08.267000",
          "content": "<p>there is not absoute conclusion in data science, everything will have to depends on the data. e.g.</p>\n<p>\"… Thus, we need a solution to fix this imbalance\". this may be true but you can always find data that don't have to do data balancing. for example, lets create synthetic data of size 10000x10000 of np.zeros(). then just random select one pixel and colored it with random color. the ground truth mask is 1 if pixel is non-zero, zero if otherwise. in this synthetic case, +ve and -ve ratio is highly imbalance. but signal strength is very strong. in this case just BCE will work.</p>\n<p>i suggest try BCE is because it is easier to work with. it may not give you best results, but you will get some results.<br>\nthe loss landscape of BCE is smoother and you are less likey to get stuck in poor solution.</p>\n<hr>\n<p>from the other forum discussion and public kernel, it is obvious that unet-efficientnet4 will at least get LB of 0.55, local validation should higher. if there is no bug in your code,data,model, then it must be your learning process. And when it comes to learning, things that can go wrong are:</p>\n<ul>\n<li>loss function</li>\n<li>optimizer and hyperparameters. </li>\n</ul>\n<p>always remember there is only 3 things that you need to consider when debugging algorithm (not debugging code):</p>\n<ol>\n<li>data</li>\n<li>model</li>\n<li>learning process (optimisation process)</li>\n</ol>\n<hr>\n<p>i always start with something simple, e.g. simple loss, remove scheduler, remove augmentation, ….. then advance to more complicated or enhanced methdos later.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1886759,
          "author_name": "somuSan",
          "author_url": "",
          "post_date": "2022-08-06T06:14:17.730000",
          "content": "<p>Thank you so much for this great advice.<br>\nI will try to train my baseline using more minimal methods and use more advanced methods as I move along. I should have given BCE a try first then move on to trying something more complicated. Thank you so much for pointing that out.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1886360,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-05T18:19:41.743000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1886321": "I have seen people in the forum getting good scores with `Efficientnet b4` + `Unet` but when I'm submitting my inference notebook its giving me very low score [ ~0.1]. And im unable to figure out what are the reason. \n**It would be great if anyone can suggest some ways resolve this problem.**\n\n\n\n- Trained with only `Dice loss`, No augmentations, `256x256` image size, `tiled data`, `batch size=8` , `lr=1e-4`, `LR scheduler = CosineAnnealingWarmRestarts`, `opt=Adam`, submitting after transposing.\n\n- **One of the reason can be that, my model loss is not decreasing much, it is staying in `>=0.95` range.**  But I dont really understand why is that happening. It would be very helpful if anyone can suggest some tips to overcome this. I might be trying large batch size and augmentations.\n\nBelow is the plot of training loss, validation loss, val dice score and val jaccard score over epoch.\n\n<img src=\"https://i.imgur.com/4QOkx4T.png\">\n\nBelow is my model config:\n```\nCONFIG = {\n    \"in_channels\" :3,\n    \"num_classes\": 1,\n    \"BATCH_SIZE\" : 8,\n    \"NUM_EPOCHS\" : 50,\n    \"n_accumulate\": 1,\n    \"competition\": \"HuBMAP-Kaggle\", # HuBMAP-Kaggle\n    \"model_name\": \"Efficientnet_B4_Unet\",\n    \"LEARNING_RATE\": 1e-4,\n    \"DEVICE\": \"cuda\" if torch.cuda.is_available() else \"cpu\", \n    \"AUG\": \"No\",\n    \"SEED\": 42,\n    \"opt\": 'Adam',\n    \"Normalization\": \"L2\",\n    \"img_size\": (256,256),\n    \"scheduler\": \"CosineAnnealingWarmRestarts\", #'CosineAnnealingLR',\n    \"warmup_epochs\": 10,\n    \"folds_to_run\": [0],\n    \"min_lr\": 1e-6,\n    \"T_max\": 500, \n    \"T_0\": 20\n}\n```",
    "1886334": "discard dice loss and train with BCE first.",
    "1886360": ""
  }
}