{
  "id": 351588,
  "title": "The problem that the validation set loss is inconsistent with the Dice performance of the validation set.",
  "url": "/competitions/hubmap-organ-segmentation/discussion/351588",
  "author_name": "",
  "post_date": "2022-09-11T03:36:34.415155800Z",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi, Our validation set loss (we tried bceloss and iou) is inconsistent with the validation set Dice. When the validation set loss is the lowest, our validation set Dice is not the lowest. When the validation set Dice is the lowest, the validation set loss has risen. . Shouldn't Dice be the lowest when the validation set loss is the lowest? Please give us some help, thanks a lot:)</p>",
  "messages": [
    {
      "id": "1933990",
      "postDate": "09/11/2022 03:36:34",
      "content": "<p>Hi, Our validation set loss (we tried bceloss and iou) is inconsistent with the validation set Dice. When the validation set loss is the lowest, our validation set Dice is not the lowest. When the validation set Dice is the lowest, the validation set loss has risen. . Shouldn't Dice be the lowest when the validation set loss is the lowest? Please give us some help, thanks a lot:)</p>",
      "rawMarkdown": "Hi, Our validation set loss (we tried bceloss and iou) is inconsistent with the validation set Dice. When the validation set loss is the lowest, our validation set Dice is not the lowest. When the validation set Dice is the lowest, the validation set loss has risen. . Shouldn't Dice be the lowest when the validation set loss is the lowest? Please give us some help, thanks a lot:)",
      "votes": null
    },
    {
      "id": "1934083",
      "postDate": "09/11/2022 06:29:17",
      "content": "<p>Hello, I have a few suggestions that may help you: </p>\n<ol>\n<li>You should ensure that the image of the input data corresponds to the mask one by one, otherwise the model cannot perform gradient descent. </li>\n<li>You may have cut the image and have so many empty masks in all the images that the ratio of positive and negative samples in the model is out of balance. The second reason may be that you cut the image without processing.</li>\n</ol>",
      "rawMarkdown": "Hello, I have a few suggestions that may help you: \n1. You should ensure that the image of the input data corresponds to the mask one by one, otherwise the model cannot perform gradient descent. \n2. You may have cut the image and have so many empty masks in all the images that the ratio of positive and negative samples in the model is out of balance. The second reason may be that you cut the image without processing.",
      "votes": null
    },
    {
      "id": "1934133",
      "postDate": "09/11/2022 07:02:55",
      "content": "<p><a href=\"https://www.kaggle.com/xiongzheli\" target=\"_blank\">@xiongzheli</a> Do you mean dice coefficient become worse or smaller as bce loss goes down? <br>\nIn a binary classification, If an example's logits or probabilities move in the direction towards ground truth, then its loss will decrease. But as long as its probability doesn't pass through 0.5, the example will still be classified as it is before by the metrics. So imagine a large number of examples, circumstances that total loss going down while metrics going worse indeed exist. But the general trend of loss and metrics should agree. That's my thought.</p>",
      "rawMarkdown": "xiongzheli Do you mean dice coefficient become worse or smaller as bce loss goes down? \nIn a binary classification, If an example's logits or probabilities move in the direction towards ground truth, then its loss will decrease. But as long as its probability doesn't pass through 0.5, the example will still be classified as it is before by the metrics. So imagine a large number of examples, circumstances that total loss going down while metrics going worse indeed exist. But the general trend of loss and metrics should agree. That's my thought.",
      "votes": null
    },
    {
      "id": "1934175",
      "postDate": "09/11/2022 07:48:40",
      "content": "<p>It is normal that dice coefficient fluctuates as the loss improves if you are calculating dice coefficient on a single threshold. Try to calculate dice coefficient on multiple thresholds and take the average.</p>",
      "rawMarkdown": "It is normal that dice coefficient fluctuates as the loss improves if you are calculating dice coefficient on a single threshold. Try to calculate dice coefficient on multiple thresholds and take the average.",
      "votes": null
    },
    {
      "id": "1934538",
      "postDate": "09/11/2022 12:43:59",
      "content": "<p>This is a very intuitive explanation, and the setting of the threshold may indeed lead to a gap between the loss and dice indicators. But overall it should remain the same. As mentioned by <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> and we'll average the multiple threshold to try to reduce this inconsistency. thank you very much.</p>",
      "rawMarkdown": "This is a very intuitive explanation, and the setting of the threshold may indeed lead to a gap between the loss and dice indicators. But overall it should remain the same. As mentioned by @gunesevitan and we'll average the multiple threshold to try to reduce this inconsistency. thank you very much.",
      "votes": null
    },
    {
      "id": "1934543",
      "postDate": "09/11/2022 12:46:40",
      "content": "<p>We believe this is a very important point of improvement and we will try it, thank you for your generous help</p>",
      "rawMarkdown": "We believe this is a very important point of improvement and we will try it, thank you for your generous help",
      "votes": null
    },
    {
      "id": "1934560",
      "postDate": "09/11/2022 13:00:56",
      "content": "<p>thank you Zhang. here's the detail we do.<br>\n1:we do checked the consistency of samples and labels and it may not disturb the gradient descent.<br>\n2:We divide the patches containing FTU and empty by 1:1. We guess that the model learns a broad range of low-probability policies due to mostly 0 masks. Maybe this 1:1 division is not a good way?</p>",
      "rawMarkdown": "thank you Zhang. here's the detail we do.\n1:we do checked the consistency of samples and labels and it may not disturb the gradient descent.\n2:We divide the patches containing FTU and empty by 1:1. We guess that the model learns a broad range of low-probability policies due to mostly 0 masks. Maybe this 1:1 division is not a good way?",
      "votes": null
    },
    {
      "id": "1934776",
      "postDate": "09/11/2022 15:33:25",
      "content": "<p>I have also encountered a large number of empty masks before, and the training cannot be performed. After that, I will delete all the empty masks, or keep a lot of non-mask parts in the sample, and the loss will look normal.</p>",
      "rawMarkdown": "I have also encountered a large number of empty masks before, and the training cannot be performed. After that, I will delete all the empty masks, or keep a lot of non-mask parts in the sample, and the loss will look normal.",
      "votes": null
    },
    {
      "id": "1934881",
      "postDate": "09/11/2022 16:32:37",
      "content": "<p>Think of a reduction in loss without a reduction in DICE as representing the model making the same predictions but with more appropriate confidence (either correct predictions are being made closer to 1 or 0 (as appropriate); or wrong predictions are being made with less confidence).</p>\n<p>Many aspects of training can influence the strength of prediction. For example, greater levels of augmentation result in harder work for the model. This often results in the model making less confident predictions, and so performing better in terms of DICE than loss (assuming decent accuracy) - because it has been trained to expect a harder dataset, where it is more likely to be wrong.</p>",
      "rawMarkdown": "Think of a reduction in loss without a reduction in DICE as representing the model making the same predictions but with more appropriate confidence (either correct predictions are being made closer to 1 or 0 (as appropriate); or wrong predictions are being made with less confidence).\n\nMany aspects of training can influence the strength of prediction. For example, greater levels of augmentation result in harder work for the model. This often results in the model making less confident predictions, and so performing better in terms of DICE than loss (assuming decent accuracy) - because it has been trained to expect a harder dataset, where it is more likely to be wrong.",
      "votes": null
    },
    {
      "id": "1941958",
      "postDate": "09/16/2022 11:03:51",
      "content": "<p>thanks for your sharing , we also tried your suggestion but did not get boost:(, after that we tried to reduce parameters of our model and get a more normal performance</p>",
      "rawMarkdown": "thanks for your sharing , we also tried your suggestion but did not get boost:(, after that we tried to reduce parameters of our model and get a more normal performance",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1934083,
      "author_name": "zhangasa",
      "author_url": "",
      "post_date": "09/11/2022 06:29:17",
      "content": "<p>Hello, I have a few suggestions that may help you: </p>\n<ol>\n<li>You should ensure that the image of the input data corresponds to the mask one by one, otherwise the model cannot perform gradient descent. </li>\n<li>You may have cut the image and have so many empty masks in all the images that the ratio of positive and negative samples in the model is out of balance. The second reason may be that you cut the image without processing.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1934560,
          "author_name": "xiongzheli",
          "author_url": "",
          "post_date": "09/11/2022 13:00:56",
          "content": "<p>thank you Zhang. here's the detail we do.<br>\n1:we do checked the consistency of samples and labels and it may not disturb the gradient descent.<br>\n2:We divide the patches containing FTU and empty by 1:1. We guess that the model learns a broad range of low-probability policies due to mostly 0 masks. Maybe this 1:1 division is not a good way?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1934776,
          "author_name": "zhangasa",
          "author_url": "",
          "post_date": "09/11/2022 15:33:25",
          "content": "<p>I have also encountered a large number of empty masks before, and the training cannot be performed. After that, I will delete all the empty masks, or keep a lot of non-mask parts in the sample, and the loss will look normal.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1941958,
          "author_name": "xiongzheli",
          "author_url": "",
          "post_date": "09/16/2022 11:03:51",
          "content": "<p>thanks for your sharing , we also tried your suggestion but did not get boost:(, after that we tried to reduce parameters of our model and get a more normal performance</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1934133,
      "author_name": "electro",
      "author_url": "",
      "post_date": "09/11/2022 07:02:55",
      "content": "<p><a href=\"https://www.kaggle.com/xiongzheli\" target=\"_blank\">@xiongzheli</a> Do you mean dice coefficient become worse or smaller as bce loss goes down? <br>\nIn a binary classification, If an example's logits or probabilities move in the direction towards ground truth, then its loss will decrease. But as long as its probability doesn't pass through 0.5, the example will still be classified as it is before by the metrics. So imagine a large number of examples, circumstances that total loss going down while metrics going worse indeed exist. But the general trend of loss and metrics should agree. That's my thought.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1934538,
          "author_name": "xiongzheli",
          "author_url": "",
          "post_date": "09/11/2022 12:43:59",
          "content": "<p>This is a very intuitive explanation, and the setting of the threshold may indeed lead to a gap between the loss and dice indicators. But overall it should remain the same. As mentioned by <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> and we'll average the multiple threshold to try to reduce this inconsistency. thank you very much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1934175,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "09/11/2022 07:48:40",
      "content": "<p>It is normal that dice coefficient fluctuates as the loss improves if you are calculating dice coefficient on a single threshold. Try to calculate dice coefficient on multiple thresholds and take the average.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1934543,
          "author_name": "xiongzheli",
          "author_url": "",
          "post_date": "09/11/2022 12:46:40",
          "content": "<p>We believe this is a very important point of improvement and we will try it, thank you for your generous help</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1934881,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "09/11/2022 16:32:37",
      "content": "<p>Think of a reduction in loss without a reduction in DICE as representing the model making the same predictions but with more appropriate confidence (either correct predictions are being made closer to 1 or 0 (as appropriate); or wrong predictions are being made with less confidence).</p>\n<p>Many aspects of training can influence the strength of prediction. For example, greater levels of augmentation result in harder work for the model. This often results in the model making less confident predictions, and so performing better in terms of DICE than loss (assuming decent accuracy) - because it has been trained to expect a harder dataset, where it is more likely to be wrong.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1933990": "Hi, Our validation set loss (we tried bceloss and iou) is inconsistent with the validation set Dice. When the validation set loss is the lowest, our validation set Dice is not the lowest. When the validation set Dice is the lowest, the validation set loss has risen. . Shouldn't Dice be the lowest when the validation set loss is the lowest? Please give us some help, thanks a lot:)",
    "1934083": "Hello, I have a few suggestions that may help you: \n1. You should ensure that the image of the input data corresponds to the mask one by one, otherwise the model cannot perform gradient descent. \n2. You may have cut the image and have so many empty masks in all the images that the ratio of positive and negative samples in the model is out of balance. The second reason may be that you cut the image without processing.",
    "1934133": "xiongzheli Do you mean dice coefficient become worse or smaller as bce loss goes down? \nIn a binary classification, If an example's logits or probabilities move in the direction towards ground truth, then its loss will decrease. But as long as its probability doesn't pass through 0.5, the example will still be classified as it is before by the metrics. So imagine a large number of examples, circumstances that total loss going down while metrics going worse indeed exist. But the general trend of loss and metrics should agree. That's my thought.",
    "1934175": "It is normal that dice coefficient fluctuates as the loss improves if you are calculating dice coefficient on a single threshold. Try to calculate dice coefficient on multiple thresholds and take the average.",
    "1934538": "This is a very intuitive explanation, and the setting of the threshold may indeed lead to a gap between the loss and dice indicators. But overall it should remain the same. As mentioned by @gunesevitan and we'll average the multiple threshold to try to reduce this inconsistency. thank you very much.",
    "1934543": "We believe this is a very important point of improvement and we will try it, thank you for your generous help",
    "1934560": "thank you Zhang. here's the detail we do.\n1:we do checked the consistency of samples and labels and it may not disturb the gradient descent.\n2:We divide the patches containing FTU and empty by 1:1. We guess that the model learns a broad range of low-probability policies due to mostly 0 masks. Maybe this 1:1 division is not a good way?",
    "1934776": "I have also encountered a large number of empty masks before, and the training cannot be performed. After that, I will delete all the empty masks, or keep a lot of non-mask parts in the sample, and the loss will look normal.",
    "1934881": "Think of a reduction in loss without a reduction in DICE as representing the model making the same predictions but with more appropriate confidence (either correct predictions are being made closer to 1 or 0 (as appropriate); or wrong predictions are being made with less confidence).\n\nMany aspects of training can influence the strength of prediction. For example, greater levels of augmentation result in harder work for the model. This often results in the model making less confident predictions, and so performing better in terms of DICE than loss (assuming decent accuracy) - because it has been trained to expect a harder dataset, where it is more likely to be wrong.",
    "1941958": "thanks for your sharing , we also tried your suggestion but did not get boost:(, after that we tried to reduce parameters of our model and get a more normal performance"
  },
  "source": "meta"
}