{
  "id": 456269,
  "title": "Segmentation covers just a small percentage of the entire slice.",
  "url": "/competitions/blood-vessel-segmentation/discussion/456269",
  "author_name": "",
  "post_date": "2023-11-19T04:29:56.742941700Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F974889%2Fa52df11d0e879044871b287c59fef16e%2FScreenshot%202023-11-18%20at%2011.27.20PM.png?generation=1700368120472541&amp;alt=media\" alt=\"\"></p>\n<p>The background constitutes 99% of the total area of the slice. Does this necessitate the use of a specific loss function?</p>",
  "messages": [
    {
      "id": "2530341",
      "postDate": "11/19/2023 04:29:56",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F974889%2Fa52df11d0e879044871b287c59fef16e%2FScreenshot%202023-11-18%20at%2011.27.20PM.png?generation=1700368120472541&amp;alt=media\" alt=\"\"></p>\n<p>The background constitutes 99% of the total area of the slice. Does this necessitate the use of a specific loss function?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F974889%2Fa52df11d0e879044871b287c59fef16e%2FScreenshot%202023-11-18%20at%2011.27.20PM.png?generation=1700368120472541&alt=media)\n\nThe background constitutes 99% of the total area of the slice. Does this necessitate the use of a specific loss function?",
      "votes": null
    },
    {
      "id": "2530611",
      "postDate": "11/19/2023 10:07:33",
      "content": "<p>not necessary.<br>\nand actually there is no problem at all.</p>\n<p>you should</p>\n<ol>\n<li>start off which normal binary cross entropy loss first.<br>\nif the problem exist, then think of data imbalance.</li>\n</ol>\n<hr>\n<p>data science is about proving/showing the problem exist first, then think of solution</p>\n<hr>\n<p>i am not saying that <strong>all</strong> data imbalance  will cause no problem but in this case, it doesn't.<br>\nIt doesn't because  the target blood vessel is distinctive enough even it is very very minority class.</p>\n<p>as an another example, you target is to build a neural net to detect/segment red(255,0,0) pixel from black pixel pixel(0,0,0).<br>\nThe pixels in an image is either red or black and nothing else. you have 100000 train images and the percentage of red pixel in training data is only 0.000001%</p>\n<p>do you need special imbalance loss?<br>\nprobably no. because the two class are completely separately.<br>\nbut you do need to ensure each training batch has both class. (training batch cannot contain single class)</p>",
      "rawMarkdown": "not necessary.\nand actually there is no problem at all.\n\nyou should\n1. start off which normal binary cross entropy loss first.\nif the problem exist, then think of data imbalance.\n\n---\ndata science is about proving/showing the problem exist first, then think of solution\n\n---\n\ni am not saying that **all** data imbalance  will cause no problem but in this case, it doesn't.\nIt doesn't because  the target blood vessel is distinctive enough even it is very very minority class.\n\nas an another example, you target is to build a neural net to detect/segment red(255,0,0) pixel from black pixel pixel(0,0,0).\nThe pixels in an image is either red or black and nothing else. you have 100000 train images and the percentage of red pixel in training data is only 0.000001%\n\ndo you need special imbalance loss?\nprobably no. because the two class are completely separately.\nbut you do need to ensure each training batch has both class. (training batch cannot contain single class)",
      "votes": null
    },
    {
      "id": "2530641",
      "postDate": "11/19/2023 11:14:15",
      "content": "<p>Very good insight on the data imbalance, could another method be to resample the training batch i.e. oversample from the minority class or undersample from the majority class, this should help the model generalize better.</p>",
      "rawMarkdown": "Very good insight on the data imbalance, could another method be to resample the training batch i.e. oversample from the minority class or undersample from the majority class, this should help the model generalize better.",
      "votes": null
    },
    {
      "id": "2546392",
      "postDate": "12/02/2023 11:28:23",
      "content": "<p>This is a really good question. If I understand correctly, in SenNet there is only one target variable, Y which is in [0, 1]. So hengck23 suggested using cross-entropy, but cross-entropy is a measure involving two variables, so my understanding there is no cross-entropy to be had. </p>\n<p>Perhaps hengck23 meant Entropy (aka self-entropy) as this is a measure of the uncertainty of a single variable. It is the number of bits necessary to represent it, i.e. the relative imbalance between the 0s and 1s.  So if the target variable was balanced 50/50 then self-entropy (using base 2) would be 1 indicating a single bit would encode the target variable. If the target variable was imbalanced (100000 to 1) the self-entropy would be near 0 indicating (theoretically) the data could be represented by \"fewer than\" 1 bits. So certainly self-entropy would show what we already know - the target variable is very imbalanced. (I found this to be helpful <a href=\"https://www.eecs.harvard.edu/cs286r/courses/fall10/papers/Chapter2.pdf\" target=\"_blank\">https://www.eecs.harvard.edu/cs286r/courses/fall10/papers/Chapter2.pdf</a>)</p>\n<p>But none of this answers the question which I will paraphrase as (what considerations should be taken when training on a heavily imbalanced training set?). Doing nothing we will be spending a lot of training cycles (time) on one of the class values, and little on the other. Should we bias our training set with more samples of the rare value, for example using scikit-learn GroupKFold or StratifiedKFold at <a href=\"https://scikit-learn.org/stable/modules/classes.html#module-sklearn.model_selection\" target=\"_blank\">https://scikit-learn.org/stable/modules/classes.html#module-sklearn.model_selection</a> ?   And if we do that - what is the impact to our model, and what can we do to accomodate that impact?</p>",
      "rawMarkdown": "This is a really good question. If I understand correctly, in SenNet there is only one target variable, Y which is in [0, 1]. So hengck23 suggested using cross-entropy, but cross-entropy is a measure involving two variables, so my understanding there is no cross-entropy to be had. \n\nPerhaps hengck23 meant Entropy (aka self-entropy) as this is a measure of the uncertainty of a single variable. It is the number of bits necessary to represent it, i.e. the relative imbalance between the 0s and 1s.  So if the target variable was balanced 50/50 then self-entropy (using base 2) would be 1 indicating a single bit would encode the target variable. If the target variable was imbalanced (100000 to 1) the self-entropy would be near 0 indicating (theoretically) the data could be represented by \"fewer than\" 1 bits. So certainly self-entropy would show what we already know - the target variable is very imbalanced. (I found this to be helpful https://www.eecs.harvard.edu/cs286r/courses/fall10/papers/Chapter2.pdf)\n\nBut none of this answers the question which I will paraphrase as (what considerations should be taken when training on a heavily imbalanced training set?). Doing nothing we will be spending a lot of training cycles (time) on one of the class values, and little on the other. Should we bias our training set with more samples of the rare value, for example using scikit-learn GroupKFold or StratifiedKFold at https://scikit-learn.org/stable/modules/classes.html#module-sklearn.model_selection ?   And if we do that - what is the impact to our model, and what can we do to accomodate that impact?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2530611,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/19/2023 10:07:33",
      "content": "<p>not necessary.<br>\nand actually there is no problem at all.</p>\n<p>you should</p>\n<ol>\n<li>start off which normal binary cross entropy loss first.<br>\nif the problem exist, then think of data imbalance.</li>\n</ol>\n<hr>\n<p>data science is about proving/showing the problem exist first, then think of solution</p>\n<hr>\n<p>i am not saying that <strong>all</strong> data imbalance  will cause no problem but in this case, it doesn't.<br>\nIt doesn't because  the target blood vessel is distinctive enough even it is very very minority class.</p>\n<p>as an another example, you target is to build a neural net to detect/segment red(255,0,0) pixel from black pixel pixel(0,0,0).<br>\nThe pixels in an image is either red or black and nothing else. you have 100000 train images and the percentage of red pixel in training data is only 0.000001%</p>\n<p>do you need special imbalance loss?<br>\nprobably no. because the two class are completely separately.<br>\nbut you do need to ensure each training batch has both class. (training batch cannot contain single class)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2530641,
          "author_name": "adebayo",
          "author_url": "",
          "post_date": "11/19/2023 11:14:15",
          "content": "<p>Very good insight on the data imbalance, could another method be to resample the training batch i.e. oversample from the minority class or undersample from the majority class, this should help the model generalize better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2546392,
          "author_name": "cwinsor",
          "author_url": "",
          "post_date": "12/02/2023 11:28:23",
          "content": "<p>This is a really good question. If I understand correctly, in SenNet there is only one target variable, Y which is in [0, 1]. So hengck23 suggested using cross-entropy, but cross-entropy is a measure involving two variables, so my understanding there is no cross-entropy to be had. </p>\n<p>Perhaps hengck23 meant Entropy (aka self-entropy) as this is a measure of the uncertainty of a single variable. It is the number of bits necessary to represent it, i.e. the relative imbalance between the 0s and 1s.  So if the target variable was balanced 50/50 then self-entropy (using base 2) would be 1 indicating a single bit would encode the target variable. If the target variable was imbalanced (100000 to 1) the self-entropy would be near 0 indicating (theoretically) the data could be represented by \"fewer than\" 1 bits. So certainly self-entropy would show what we already know - the target variable is very imbalanced. (I found this to be helpful <a href=\"https://www.eecs.harvard.edu/cs286r/courses/fall10/papers/Chapter2.pdf\" target=\"_blank\">https://www.eecs.harvard.edu/cs286r/courses/fall10/papers/Chapter2.pdf</a>)</p>\n<p>But none of this answers the question which I will paraphrase as (what considerations should be taken when training on a heavily imbalanced training set?). Doing nothing we will be spending a lot of training cycles (time) on one of the class values, and little on the other. Should we bias our training set with more samples of the rare value, for example using scikit-learn GroupKFold or StratifiedKFold at <a href=\"https://scikit-learn.org/stable/modules/classes.html#module-sklearn.model_selection\" target=\"_blank\">https://scikit-learn.org/stable/modules/classes.html#module-sklearn.model_selection</a> ?   And if we do that - what is the impact to our model, and what can we do to accomodate that impact?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2530341": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F974889%2Fa52df11d0e879044871b287c59fef16e%2FScreenshot%202023-11-18%20at%2011.27.20PM.png?generation=1700368120472541&alt=media)\n\nThe background constitutes 99% of the total area of the slice. Does this necessitate the use of a specific loss function?",
    "2530611": "not necessary.\nand actually there is no problem at all.\n\nyou should\n1. start off which normal binary cross entropy loss first.\nif the problem exist, then think of data imbalance.\n\n---\ndata science is about proving/showing the problem exist first, then think of solution\n\n---\n\ni am not saying that **all** data imbalance  will cause no problem but in this case, it doesn't.\nIt doesn't because  the target blood vessel is distinctive enough even it is very very minority class.\n\nas an another example, you target is to build a neural net to detect/segment red(255,0,0) pixel from black pixel pixel(0,0,0).\nThe pixels in an image is either red or black and nothing else. you have 100000 train images and the percentage of red pixel in training data is only 0.000001%\n\ndo you need special imbalance loss?\nprobably no. because the two class are completely separately.\nbut you do need to ensure each training batch has both class. (training batch cannot contain single class)",
    "2530641": "Very good insight on the data imbalance, could another method be to resample the training batch i.e. oversample from the minority class or undersample from the majority class, this should help the model generalize better.",
    "2546392": "This is a really good question. If I understand correctly, in SenNet there is only one target variable, Y which is in [0, 1]. So hengck23 suggested using cross-entropy, but cross-entropy is a measure involving two variables, so my understanding there is no cross-entropy to be had. \n\nPerhaps hengck23 meant Entropy (aka self-entropy) as this is a measure of the uncertainty of a single variable. It is the number of bits necessary to represent it, i.e. the relative imbalance between the 0s and 1s.  So if the target variable was balanced 50/50 then self-entropy (using base 2) would be 1 indicating a single bit would encode the target variable. If the target variable was imbalanced (100000 to 1) the self-entropy would be near 0 indicating (theoretically) the data could be represented by \"fewer than\" 1 bits. So certainly self-entropy would show what we already know - the target variable is very imbalanced. (I found this to be helpful https://www.eecs.harvard.edu/cs286r/courses/fall10/papers/Chapter2.pdf)\n\nBut none of this answers the question which I will paraphrase as (what considerations should be taken when training on a heavily imbalanced training set?). Doing nothing we will be spending a lot of training cycles (time) on one of the class values, and little on the other. Should we bias our training set with more samples of the rare value, for example using scikit-learn GroupKFold or StratifiedKFold at https://scikit-learn.org/stable/modules/classes.html#module-sklearn.model_selection ?   And if we do that - what is the impact to our model, and what can we do to accomodate that impact?"
  },
  "source": "meta"
}