{
  "id": 165213,
  "title": "binary_crossentropy vs categorical_crossentropy",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/165213",
  "author_name": "",
  "post_date": "2020-07-08T20:24:16.687930Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I've seen some notebooks that use loss function binary_crossentropy and  some notebooks use loss function categorical_crossentropy with 2 classes. What are the differences between them? What are the pros and cons of each approach?</p>",
  "messages": [
    {
      "id": "920827",
      "postDate": "07/08/2020 20:24:16",
      "content": "<p>I've seen some notebooks that use loss function binary_crossentropy and  some notebooks use loss function categorical_crossentropy with 2 classes. What are the differences between them? What are the pros and cons of each approach?</p>",
      "rawMarkdown": "I've seen some notebooks that use loss function binary_crossentropy and  some notebooks use loss function categorical_crossentropy with 2 classes. What are the differences between them? What are the pros and cons of each approach?",
      "votes": null
    },
    {
      "id": "920944",
      "postDate": "07/09/2020 00:28:02",
      "content": "<p>Categorical_crossentropy is typically used for multiclass, single-label problem types with softmax (last-layer activation) where the output is probability distribution of the class over \"n\" classes  and its sum is one. This means the probability for a class is not independent from the other class probabilities. This is good as long as we predict single label per sample because we will then choose the class with highest probability as our prediction. But if we have multiple labels per sample then we would have to set up some kind of threshold value to classify the labels as predictions which we do not want to do. </p>\n\n<p>On the other hand, for a binary classification problem (true/false types) or multiclass, multilabel classification problems, the loss function binary_crossentropy + sigmoid (last-layer activation) is used where the loss computed for every output class is independent of the other class i.e, the sample belonging to a certain class is not influenced by it belonging to some other class.</p>\n\n<p>There are several other resources online which explains this in more detail as well 😊 .</p>\n\n<p><a href=\"https://gombru.github.io/2018/05/23/cross_entropy_loss/\">Understanding Categorical Cross-Entropy Loss</a>\n<a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/51602\">SimilarKaggleDiscussion</a></p>\n\n<p>Hope this helps.</p>",
      "rawMarkdown": "Categorical_crossentropy is typically used for multiclass, single-label problem types with softmax (last-layer activation) where the output is probability distribution of the class over \"n\" classes  and its sum is one. This means the probability for a class is not independent from the other class probabilities. This is good as long as we predict single label per sample because we will then choose the class with highest probability as our prediction. But if we have multiple labels per sample then we would have to set up some kind of threshold value to classify the labels as predictions which we do not want to do. \n\nOn the other hand, for a binary classification problem (true/false types) or multiclass, multilabel classification problems, the loss function binary_crossentropy + sigmoid (last-layer activation) is used where the loss computed for every output class is independent of the other class i.e, the sample belonging to a certain class is not influenced by it belonging to some other class.\n\nThere are several other resources online which explains this in more detail as well 😊 .\n\n[Understanding Categorical Cross-Entropy Loss](https://gombru.github.io/2018/05/23/cross_entropy_loss/)\n[SimilarKaggleDiscussion](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/51602)\n\nHope this helps.",
      "votes": null
    },
    {
      "id": "921352",
      "postDate": "07/09/2020 08:24:01",
      "content": "<p><a href=\"/kckamojjala\">@kckamojjala</a> Thanks!\nWhat is better to use this competition, binary crossentropy or categorical crossentropy with 2 classes?</p>",
      "rawMarkdown": "kckamojjala Thanks!\nWhat is better to use this competition, binary crossentropy or categorical crossentropy with 2 classes?",
      "votes": null
    },
    {
      "id": "921630",
      "postDate": "07/09/2020 12:39:35",
      "content": "<p>I would use binary_crossentropy for this one as its true/false type problem.</p>",
      "rawMarkdown": "I would use binary_crossentropy for this one as its true/false type problem.",
      "votes": null
    },
    {
      "id": "925171",
      "postDate": "07/11/2020 21:26:38",
      "content": "<p>Who downvoted this and why? This is a ligit question!</p>",
      "rawMarkdown": "Who downvoted this and why? This is a ligit question!",
      "votes": null
    },
    {
      "id": "925221",
      "postDate": "07/11/2020 23:40:20",
      "content": "<p>Totally agree Jared, KC and Alon. Sometimes the doubt of any user is helpful for others. Mostly for those that are shy to ask.  Whenever we receive downvotes to have doubts, Kagglers are discouraged to make more topics asking for help. That isn't good even for the Platform too. We are suppose to give support and answer the other issues when and if we can, since Kteam can't deal with such huge community. In Science what's  meaningless/irrelevant for one, maybe is helpful for another person.</p>",
      "rawMarkdown": "Totally agree Jared, KC and Alon. Sometimes the doubt of any user is helpful for others. Mostly for those that are shy to ask.  Whenever we receive downvotes to have doubts, Kagglers are discouraged to make more topics asking for help. That isn't good even for the Platform too. We are suppose to give support and answer the other issues when and if we can, since Kteam can't deal with such huge community. In Science what's  meaningless/irrelevant for one, maybe is helpful for another person.",
      "votes": null
    },
    {
      "id": "925246",
      "postDate": "07/12/2020 00:12:31",
      "content": "<p>well said !</p>",
      "rawMarkdown": "well said !",
      "votes": null
    },
    {
      "id": "925937",
      "postDate": "07/12/2020 11:46:57",
      "content": "<p>binary_crossentropy is used when you have only two outputs(Say 0 or 1 ), whereas categorical_crossentropy is used when you have more than two output(Say 1,2,3 ,.....)but you can use categorical_crossentropy(Use - say you have to classify tweets in classes ranging from 1 to 45) only when the number of class is already known .There is one more famous loss function linear_regression that is used when you have multiple outputs(e.g. - house price prediction).\nhope you like it !</p>",
      "rawMarkdown": "binary_crossentropy is used when you have only two outputs(Say 0 or 1 ), whereas categorical_crossentropy is used when you have more than two output(Say 1,2,3 ,.....)but you can use categorical_crossentropy(Use - say you have to classify tweets in classes ranging from 1 to 45) only when the number of class is already known .There is one more famous loss function linear_regression that is used when you have multiple outputs(e.g. - house price prediction).\nhope you like it !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 920944,
      "author_name": "kckamojjala",
      "author_url": "",
      "post_date": "07/09/2020 00:28:02",
      "content": "<p>Categorical_crossentropy is typically used for multiclass, single-label problem types with softmax (last-layer activation) where the output is probability distribution of the class over \"n\" classes  and its sum is one. This means the probability for a class is not independent from the other class probabilities. This is good as long as we predict single label per sample because we will then choose the class with highest probability as our prediction. But if we have multiple labels per sample then we would have to set up some kind of threshold value to classify the labels as predictions which we do not want to do. </p>\n\n<p>On the other hand, for a binary classification problem (true/false types) or multiclass, multilabel classification problems, the loss function binary_crossentropy + sigmoid (last-layer activation) is used where the loss computed for every output class is independent of the other class i.e, the sample belonging to a certain class is not influenced by it belonging to some other class.</p>\n\n<p>There are several other resources online which explains this in more detail as well 😊 .</p>\n\n<p><a href=\"https://gombru.github.io/2018/05/23/cross_entropy_loss/\">Understanding Categorical Cross-Entropy Loss</a>\n<a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/51602\">SimilarKaggleDiscussion</a></p>\n\n<p>Hope this helps.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 921352,
      "author_name": "alongigi",
      "author_url": "",
      "post_date": "07/09/2020 08:24:01",
      "content": "<p><a href=\"/kckamojjala\">@kckamojjala</a> Thanks!\nWhat is better to use this competition, binary crossentropy or categorical crossentropy with 2 classes?</p>",
      "votes": null,
      "replies": [
        {
          "id": 921630,
          "author_name": "kckamojjala",
          "author_url": "",
          "post_date": "07/09/2020 12:39:35",
          "content": "<p>I would use binary_crossentropy for this one as its true/false type problem.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 925171,
      "author_name": "jaredsavage",
      "author_url": "",
      "post_date": "07/11/2020 21:26:38",
      "content": "<p>Who downvoted this and why? This is a ligit question!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 925221,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "07/11/2020 23:40:20",
      "content": "<p>Totally agree Jared, KC and Alon. Sometimes the doubt of any user is helpful for others. Mostly for those that are shy to ask.  Whenever we receive downvotes to have doubts, Kagglers are discouraged to make more topics asking for help. That isn't good even for the Platform too. We are suppose to give support and answer the other issues when and if we can, since Kteam can't deal with such huge community. In Science what's  meaningless/irrelevant for one, maybe is helpful for another person.</p>",
      "votes": null,
      "replies": [
        {
          "id": 925246,
          "author_name": "kckamojjala",
          "author_url": "",
          "post_date": "07/12/2020 00:12:31",
          "content": "<p>well said !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 925937,
      "author_name": "souravsamrat",
      "author_url": "",
      "post_date": "07/12/2020 11:46:57",
      "content": "<p>binary_crossentropy is used when you have only two outputs(Say 0 or 1 ), whereas categorical_crossentropy is used when you have more than two output(Say 1,2,3 ,.....)but you can use categorical_crossentropy(Use - say you have to classify tweets in classes ranging from 1 to 45) only when the number of class is already known .There is one more famous loss function linear_regression that is used when you have multiple outputs(e.g. - house price prediction).\nhope you like it !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "920827": "I've seen some notebooks that use loss function binary_crossentropy and  some notebooks use loss function categorical_crossentropy with 2 classes. What are the differences between them? What are the pros and cons of each approach?",
    "920944": "Categorical_crossentropy is typically used for multiclass, single-label problem types with softmax (last-layer activation) where the output is probability distribution of the class over \"n\" classes  and its sum is one. This means the probability for a class is not independent from the other class probabilities. This is good as long as we predict single label per sample because we will then choose the class with highest probability as our prediction. But if we have multiple labels per sample then we would have to set up some kind of threshold value to classify the labels as predictions which we do not want to do. \n\nOn the other hand, for a binary classification problem (true/false types) or multiclass, multilabel classification problems, the loss function binary_crossentropy + sigmoid (last-layer activation) is used where the loss computed for every output class is independent of the other class i.e, the sample belonging to a certain class is not influenced by it belonging to some other class.\n\nThere are several other resources online which explains this in more detail as well 😊 .\n\n[Understanding Categorical Cross-Entropy Loss](https://gombru.github.io/2018/05/23/cross_entropy_loss/)\n[SimilarKaggleDiscussion](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/51602)\n\nHope this helps.",
    "921352": "kckamojjala Thanks!\nWhat is better to use this competition, binary crossentropy or categorical crossentropy with 2 classes?",
    "921630": "I would use binary_crossentropy for this one as its true/false type problem.",
    "925171": "Who downvoted this and why? This is a ligit question!",
    "925221": "Totally agree Jared, KC and Alon. Sometimes the doubt of any user is helpful for others. Mostly for those that are shy to ask.  Whenever we receive downvotes to have doubts, Kagglers are discouraged to make more topics asking for help. That isn't good even for the Platform too. We are suppose to give support and answer the other issues when and if we can, since Kteam can't deal with such huge community. In Science what's  meaningless/irrelevant for one, maybe is helpful for another person.",
    "925246": "well said !",
    "925937": "binary_crossentropy is used when you have only two outputs(Say 0 or 1 ), whereas categorical_crossentropy is used when you have more than two output(Say 1,2,3 ,.....)but you can use categorical_crossentropy(Use - say you have to classify tweets in classes ranging from 1 to 45) only when the number of class is already known .There is one more famous loss function linear_regression that is used when you have multiple outputs(e.g. - house price prediction).\nhope you like it !"
  },
  "source": "meta"
}