{
  "id": 73918,
  "title": "Can someone tell me how to handle more than one target in train file because a simple cnn does not gives us multi target...",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/73918",
  "author_name": "Susnato Dhar",
  "post_date": "2018-12-06T17:18:41.791000",
  "votes": -2,
  "comment_count": 9,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": 436434,
      "postDate": "2018-12-10T09:58:54.557Z",
      "content": "<p>@GX Kok</p>\n\n<p>While you are right with one neuron output, the statement that sigmoid instead of softmax must be used is true, see kernels like:  <a href=\"https://www.kaggle.com/rejpalcz/cnn-128x128x4-keras-from-scratch-lb-0-328\">https://www.kaggle.com/rejpalcz/cnn-128x128x4-keras-from-scratch-lb-0-328</a> which uses sigmoid as anoutput and binary_crossentropy as a loss, same goes for: <a href=\"https://www.kaggle.com/byrachonok/pretrained-inceptionresnetv2-base-classifier\">https://www.kaggle.com/byrachonok/pretrained-inceptionresnetv2-base-classifier</a> and for <a href=\"https://www.kaggle.com/kwentar/two-branches-xception-lb-0-3\">https://www.kaggle.com/kwentar/two-branches-xception-lb-0-3</a>  . </p>\n\n<p>I am not really familiar with the mathematics, but I think that softmax only takes in account one label instead of multilabel cause you calculate the optimum distance to the other labels, when you multiple labels as true, you can not calculate the distance with sofmtax since more than one  are true</p>",
      "rawMarkdown": "@GX Kok\n\nWhile you are right with one neuron output, the statement that sigmoid instead of softmax must be used is true, see kernels like:  https://www.kaggle.com/rejpalcz/cnn-128x128x4-keras-from-scratch-lb-0-328 which uses sigmoid as anoutput and binary_crossentropy as a loss, same goes for: https://www.kaggle.com/byrachonok/pretrained-inceptionresnetv2-base-classifier and for https://www.kaggle.com/kwentar/two-branches-xception-lb-0-3  . \n\nI am not really familiar with the mathematics, but I think that softmax only takes in account one label instead of multilabel cause you calculate the optimum distance to the other labels, when you multiple labels as true, you can not calculate the distance with sofmtax since more than one  are true\n\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 438002,
          "postDate": "2018-12-12T23:30:11.470Z",
          "content": "<p>Perhaps the reason is that the output neurons are independent in that the sum of them don't sum to one. Softmax and categorical should be for one out of N class.</p>",
          "rawMarkdown": "Perhaps the reason is that the output neurons are independent in that the sum of them don't sum to one. Softmax and categorical should be for one out of N class."
        }
      ]
    },
    {
      "id": 436107,
      "postDate": "2018-12-09T15:51:18.010Z",
      "content": "<p>@GX Kok - I think the last layer should be sigmoid and the optimizer binary_crossentropy due to the fact that it is multi label instead of a single label, if I am not mistaken. </p>",
      "rawMarkdown": "@GX Kok - I think the last layer should be sigmoid and the optimizer binary_crossentropy due to the fact that it is multi label instead of a single label, if I am not mistaken. ",
      "replies": [
        {
          "id": 436242,
          "postDate": "2018-12-10T01:03:35.097Z",
          "content": "<p>Binary crossentropy is when you have a single output neuron and it takes a value in range [0, 1]. If there are multiple output neurons and each of them are binary and the sum of them equals 1, then you should use softmax activation and categorical crossentropy. </p>\n\n<p>The binary crossentropy formula basically assumes you have one neuron, and somehow <strong>converts/expands</strong> it to a two neuron output to compute the loss.</p>\n\n<pre><code>−(ylog(p)+(1−y)log(1−p))\n</code></pre>\n\n<p>See the y and (1-y) and p and (1-p) in the above equation.</p>",
          "rawMarkdown": "Binary crossentropy is when you have a single output neuron and it takes a value in range [0, 1]. If there are multiple output neurons and each of them are binary and the sum of them equals 1, then you should use softmax activation and categorical crossentropy. \n\nThe binary crossentropy formula basically assumes you have one neuron, and somehow **converts/expands** it to a two neuron output to compute the loss.\n\n    −(ylog(p)+(1−y)log(1−p))\n\nSee the y and (1-y) and p and (1-p) in the above equation."
        },
        {
          "id": 436504,
          "postDate": "2018-12-10T12:49:32.930Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 438001,
          "postDate": "2018-12-12T23:28:16.943Z",
          "content": "<p>You are right.</p>\n\n<p>One unit output we use sigmoid and binary crossentropy. For one out of N class use softmax and categorical crossentropy. And if multiple independent use sigmoid and binary crossentropy.</p>",
          "rawMarkdown": "You are right.\n\nOne unit output we use sigmoid and binary crossentropy. For one out of N class use softmax and categorical crossentropy. And if multiple independent use sigmoid and binary crossentropy.",
          "votes": 1
        }
      ]
    },
    {
      "id": 435430,
      "postDate": "2018-12-08T03:02:05.023Z",
      "content": "<p>At the final dense layer, set number of units or neurons to number of classes, use softmax activation for these units, and use categorical crossentropy for your optimizer loss function. Of course you also need labels to be in the correct form during training, in other words, they should be one hot encoded.</p>",
      "rawMarkdown": "At the final dense layer, set number of units or neurons to number of classes, use softmax activation for these units, and use categorical crossentropy for your optimizer loss function. Of course you also need labels to be in the correct form during training, in other words, they should be one hot encoded."
    },
    {
      "id": 434683,
      "postDate": "2018-12-06T19:37:19.823Z",
      "content": "<p>You may wish to look at the kernels section, there are examples in both Keras and PyTorch.</p>",
      "rawMarkdown": "You may wish to look at the kernels section, there are examples in both Keras and PyTorch."
    },
    {
      "id": 434614,
      "postDate": "2018-12-06T17:18:41.790Z",
      "rawMarkdown": "",
      "votes": -2
    },
    {
      "id": 436342,
      "postDate": "2018-12-10T06:37:01.853Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 436434,
      "author_name": "OxFEE1DEAD",
      "author_url": "",
      "post_date": "2018-12-10T09:58:54.557000",
      "content": "<p>@GX Kok</p>\n\n<p>While you are right with one neuron output, the statement that sigmoid instead of softmax must be used is true, see kernels like:  <a href=\"https://www.kaggle.com/rejpalcz/cnn-128x128x4-keras-from-scratch-lb-0-328\">https://www.kaggle.com/rejpalcz/cnn-128x128x4-keras-from-scratch-lb-0-328</a> which uses sigmoid as anoutput and binary_crossentropy as a loss, same goes for: <a href=\"https://www.kaggle.com/byrachonok/pretrained-inceptionresnetv2-base-classifier\">https://www.kaggle.com/byrachonok/pretrained-inceptionresnetv2-base-classifier</a> and for <a href=\"https://www.kaggle.com/kwentar/two-branches-xception-lb-0-3\">https://www.kaggle.com/kwentar/two-branches-xception-lb-0-3</a>  . </p>\n\n<p>I am not really familiar with the mathematics, but I think that softmax only takes in account one label instead of multilabel cause you calculate the optimum distance to the other labels, when you multiple labels as true, you can not calculate the distance with sofmtax since more than one  are true</p>",
      "votes": 1,
      "replies": [
        {
          "id": 438002,
          "author_name": "GX Kok",
          "author_url": "",
          "post_date": "2018-12-12T23:30:11.470000",
          "content": "<p>Perhaps the reason is that the output neurons are independent in that the sum of them don't sum to one. Softmax and categorical should be for one out of N class.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 436107,
      "author_name": "OxFEE1DEAD",
      "author_url": "",
      "post_date": "2018-12-09T15:51:18.010000",
      "content": "<p>@GX Kok - I think the last layer should be sigmoid and the optimizer binary_crossentropy due to the fact that it is multi label instead of a single label, if I am not mistaken. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 436242,
          "author_name": "GX Kok",
          "author_url": "",
          "post_date": "2018-12-10T01:03:35.097000",
          "content": "<p>Binary crossentropy is when you have a single output neuron and it takes a value in range [0, 1]. If there are multiple output neurons and each of them are binary and the sum of them equals 1, then you should use softmax activation and categorical crossentropy. </p>\n\n<p>The binary crossentropy formula basically assumes you have one neuron, and somehow <strong>converts/expands</strong> it to a two neuron output to compute the loss.</p>\n\n<pre><code>−(ylog(p)+(1−y)log(1−p))\n</code></pre>\n\n<p>See the y and (1-y) and p and (1-p) in the above equation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436504,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-10T12:49:32.930000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 438001,
          "author_name": "GX Kok",
          "author_url": "",
          "post_date": "2018-12-12T23:28:16.943000",
          "content": "<p>You are right.</p>\n\n<p>One unit output we use sigmoid and binary crossentropy. For one out of N class use softmax and categorical crossentropy. And if multiple independent use sigmoid and binary crossentropy.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 435430,
      "author_name": "GX Kok",
      "author_url": "",
      "post_date": "2018-12-08T03:02:05.023000",
      "content": "<p>At the final dense layer, set number of units or neurons to number of classes, use softmax activation for these units, and use categorical crossentropy for your optimizer loss function. Of course you also need labels to be in the correct form during training, in other words, they should be one hot encoded.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434683,
      "author_name": "Mark Worrall",
      "author_url": "",
      "post_date": "2018-12-06T19:37:19.823000",
      "content": "<p>You may wish to look at the kernels section, there are examples in both Keras and PyTorch.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 436342,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-10T06:37:01.853000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "436434": "@GX Kok\n\nWhile you are right with one neuron output, the statement that sigmoid instead of softmax must be used is true, see kernels like:  https://www.kaggle.com/rejpalcz/cnn-128x128x4-keras-from-scratch-lb-0-328 which uses sigmoid as anoutput and binary_crossentropy as a loss, same goes for: https://www.kaggle.com/byrachonok/pretrained-inceptionresnetv2-base-classifier and for https://www.kaggle.com/kwentar/two-branches-xception-lb-0-3  . \n\nI am not really familiar with the mathematics, but I think that softmax only takes in account one label instead of multilabel cause you calculate the optimum distance to the other labels, when you multiple labels as true, you can not calculate the distance with sofmtax since more than one  are true\n\n\n",
    "436107": "@GX Kok - I think the last layer should be sigmoid and the optimizer binary_crossentropy due to the fact that it is multi label instead of a single label, if I am not mistaken. ",
    "435430": "At the final dense layer, set number of units or neurons to number of classes, use softmax activation for these units, and use categorical crossentropy for your optimizer loss function. Of course you also need labels to be in the correct form during training, in other words, they should be one hot encoded.",
    "434683": "You may wish to look at the kernels section, there are examples in both Keras and PyTorch.",
    "434614": "",
    "436342": ""
  }
}