{
  "id": 134176,
  "title": "how about sigmoid instead of softmax - curiosity",
  "url": "/competitions/bengaliai-cv19/discussion/134176",
  "author_name": "",
  "post_date": "2020-03-06T12:16:06.385823900Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Most of public kernels approached this task using models with <code>3 heads + softmax</code>, @seesee used <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134161\"><code>1 head + softmax</code></a> and decoded predictions. </p>\n\n<p>Anyone tried to approach it as multilabel with a <code>1 head + sigmoid</code>.</p>\n\n<p>What would be pros and/or cons? if any....</p>",
  "messages": [
    {
      "id": "765246",
      "postDate": "03/06/2020 12:16:06",
      "content": "<p>Most of public kernels approached this task using models with <code>3 heads + softmax</code>, @seesee used <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134161\"><code>1 head + softmax</code></a> and decoded predictions. </p>\n\n<p>Anyone tried to approach it as multilabel with a <code>1 head + sigmoid</code>.</p>\n\n<p>What would be pros and/or cons? if any....</p>",
      "rawMarkdown": "Most of public kernels approached this task using models with `3 heads + softmax`, @seesee used [`1 head + softmax`](https://www.kaggle.com/c/bengaliai-cv19/discussion/134161) and decoded predictions. \n\nAnyone tried to approach it as multilabel with a `1 head + sigmoid`.\n\nWhat would be pros and/or cons? if any....",
      "votes": null
    },
    {
      "id": "765252",
      "postDate": "03/06/2020 12:20:24",
      "content": "<p>softmax is nothing but a more generalized version of the sigmoid. sigmoid will work only if there are two classes [0,1]. \nAs here minimum class is 7. even if you use one head still you have to go for softmax</p>",
      "rawMarkdown": "softmax is nothing but a more generalized version of the sigmoid. sigmoid will work only if there are two classes [0,1]. \nAs here minimum class is 7. even if you use one head still you have to go for softmax",
      "votes": null
    },
    {
      "id": "765262",
      "postDate": "03/06/2020 12:34:52",
      "content": "<p>Nope, you are wrong. You don't need to go with softmax and here is why.\nWith sigmoid function you predict absence-presence [0, 1] for each of the 186 (168+7+11) classes and each image has three positives.</p>\n\n<p>I know it can be used. Just wonder what more experienced guys think why this could or couldn't work.</p>\n\n<p>UPDATED: previous version had wrong number of classes thx <a href=\"/carlosmir\">@carlosmir</a> </p>",
      "rawMarkdown": "Nope, you are wrong. You don't need to go with softmax and here is why.\nWith sigmoid function you predict absence-presence [0, 1] for each of the 186 (168+7+11) classes and each image has three positives.\n\nI know it can be used. Just wonder what more experienced guys think why this could or couldn't work.\n\nUPDATED: previous version had wrong number of classes thx @carlosmir",
      "votes": null
    },
    {
      "id": "765265",
      "postDate": "03/06/2020 12:35:43",
      "content": "<p>well, in that case, i am not sure</p>",
      "rawMarkdown": "well, in that case, i am not sure",
      "votes": null
    },
    {
      "id": "765275",
      "postDate": "03/06/2020 12:53:53",
      "content": "<p><code>\nThere are roughly 10,000 possible graphemes, of which roughly 1,000 are represented in the training set. \n</code>\nRead the rules carefully, It means that you don't know exactly how much classes we have.\nBut maximum is 168 * 11 * 7 = 12936 classes.</p>",
      "rawMarkdown": "```\nThere are roughly 10,000 possible graphemes, of which roughly 1,000 are represented in the training set. \n```\nRead the rules carefully, It means that you don't know exactly how much classes we have.\nBut maximum is 168 * 11 * 7 = 12936 classes.",
      "votes": null
    },
    {
      "id": "765414",
      "postDate": "03/06/2020 15:42:27",
      "content": "<p>TL;DR: I got bad results using sigmoid where softmax can be used. Softmax helps convergence.</p>\n\n<hr>\n\n<p>Hi <a href=\"/valanm\">@valanm</a> \nFrom what I understand of <a href=\"/seesee\">@seesee</a>'s work, graphemes are identified by a unique triplet <code>(grapheme_root, vowel_diacritic, consonant_diacritic)</code>. Out of 12 936 possible combinations, there are only 1 292 present in the training data set.</p>\n\n<p>&gt; With sigmoid function you predict absence-presence [0, 1] for each of the 1292 classes and each image has three positives.</p>\n\n<p>I don't exactly get what you mean here, but as <a href=\"/mks2192\">@mks2192</a> suggested, using 1292 classes one can only think of using softmax, since we know there is a single correct answer for each label value. There are no <em>three positives</em> for each image in this setting.</p>\n\n<hr>\n\n<p>However, I guess there is a way to make it work. From the top of my head, the model would be identifying the three labels in a single vector, i.e. the labels would be 168+11+7=186 sized binary vectors and the model would have to set three of these 186 components to 1.</p>\n\n<p>(\nJust a little parenthesis. Single-head model === 3-head model. The number of layers and parameters (connections) is the same, i.e. in a 3-head model all the outputs from layer[-2] are connected to head0, head 1 and head 2 (so all the 186 outputs of the model), and in a single-head model all the outputs of layer[-2] are also connected to all the 186 outputs of the model. The 3 heads just allows us to specify groups of outputs. We can achieve the same by treating outputs 0 to 167, 168 to 179 and 180 to 186 independently (applying different losses, for example).\n)</p>\n\n<p>So, now, not only has the model to identify three things, but it is not constrained to choose one of each label... But you will have to constrain it :D What I mean by constraining is associating a segment of the output vector to a label (e.g. 0-168 are the first label, etc). And you have to do it in order to train it. Then, you would be back at 3-heads + sigmoid, which makes no sense in this scenario.</p>\n\n<p>I personally got quite bad results in another task using sigmoid where softmax could be used.</p>",
      "rawMarkdown": "TL;DR: I got bad results using sigmoid where softmax can be used. Softmax helps convergence.\n\n---\n\nHi @valanm \nFrom what I understand of @seesee's work, graphemes are identified by a unique triplet `(grapheme_root, vowel_diacritic, consonant_diacritic)`. Out of 12 936 possible combinations, there are only 1 292 present in the training data set.\n\n&gt; With sigmoid function you predict absence-presence [0, 1] for each of the 1292 classes and each image has three positives.\n\nI don't exactly get what you mean here, but as @mks2192 suggested, using 1292 classes one can only think of using softmax, since we know there is a single correct answer for each label value. There are no _three positives_ for each image in this setting.\n\n--- \n\nHowever, I guess there is a way to make it work. From the top of my head, the model would be identifying the three labels in a single vector, i.e. the labels would be 168+11+7=186 sized binary vectors and the model would have to set three of these 186 components to 1.\n\n(\nJust a little parenthesis. Single-head model === 3-head model. The number of layers and parameters (connections) is the same, i.e. in a 3-head model all the outputs from layer[-2] are connected to head0, head 1 and head 2 (so all the 186 outputs of the model), and in a single-head model all the outputs of layer[-2] are also connected to all the 186 outputs of the model. The 3 heads just allows us to specify groups of outputs. We can achieve the same by treating outputs 0 to 167, 168 to 179 and 180 to 186 independently (applying different losses, for example).\n)\n\nSo, now, not only has the model to identify three things, but it is not constrained to choose one of each label... But you will have to constrain it :D What I mean by constraining is associating a segment of the output vector to a label (e.g. 0-168 are the first label, etc). And you have to do it in order to train it. Then, you would be back at 3-heads + sigmoid, which makes no sense in this scenario.\n\nI personally got quite bad results in another task using sigmoid where softmax could be used.",
      "votes": null
    },
    {
      "id": "765432",
      "postDate": "03/06/2020 16:07:08",
      "content": "<p>From 186 sigmoid outputs you take max(0-167), max(7) and max(11) and you end up at the same place as with 3 heads + softmax.</p>",
      "rawMarkdown": "From 186 sigmoid outputs you take max(0-167), max(7) and max(11) and you end up at the same place as with 3 heads + softmax.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 765252,
      "author_name": "mks2192",
      "author_url": "",
      "post_date": "03/06/2020 12:20:24",
      "content": "<p>softmax is nothing but a more generalized version of the sigmoid. sigmoid will work only if there are two classes [0,1]. \nAs here minimum class is 7. even if you use one head still you have to go for softmax</p>",
      "votes": null,
      "replies": [
        {
          "id": 765262,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "03/06/2020 12:34:52",
          "content": "<p>Nope, you are wrong. You don't need to go with softmax and here is why.\nWith sigmoid function you predict absence-presence [0, 1] for each of the 186 (168+7+11) classes and each image has three positives.</p>\n\n<p>I know it can be used. Just wonder what more experienced guys think why this could or couldn't work.</p>\n\n<p>UPDATED: previous version had wrong number of classes thx <a href=\"/carlosmir\">@carlosmir</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765265,
          "author_name": "mks2192",
          "author_url": "",
          "post_date": "03/06/2020 12:35:43",
          "content": "<p>well, in that case, i am not sure</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765275,
          "author_name": "smivvla",
          "author_url": "",
          "post_date": "03/06/2020 12:53:53",
          "content": "<p><code>\nThere are roughly 10,000 possible graphemes, of which roughly 1,000 are represented in the training set. \n</code>\nRead the rules carefully, It means that you don't know exactly how much classes we have.\nBut maximum is 168 * 11 * 7 = 12936 classes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765414,
          "author_name": "carlosmir",
          "author_url": "",
          "post_date": "03/06/2020 15:42:27",
          "content": "<p>TL;DR: I got bad results using sigmoid where softmax can be used. Softmax helps convergence.</p>\n\n<hr>\n\n<p>Hi <a href=\"/valanm\">@valanm</a> \nFrom what I understand of <a href=\"/seesee\">@seesee</a>'s work, graphemes are identified by a unique triplet <code>(grapheme_root, vowel_diacritic, consonant_diacritic)</code>. Out of 12 936 possible combinations, there are only 1 292 present in the training data set.</p>\n\n<p>&gt; With sigmoid function you predict absence-presence [0, 1] for each of the 1292 classes and each image has three positives.</p>\n\n<p>I don't exactly get what you mean here, but as <a href=\"/mks2192\">@mks2192</a> suggested, using 1292 classes one can only think of using softmax, since we know there is a single correct answer for each label value. There are no <em>three positives</em> for each image in this setting.</p>\n\n<hr>\n\n<p>However, I guess there is a way to make it work. From the top of my head, the model would be identifying the three labels in a single vector, i.e. the labels would be 168+11+7=186 sized binary vectors and the model would have to set three of these 186 components to 1.</p>\n\n<p>(\nJust a little parenthesis. Single-head model === 3-head model. The number of layers and parameters (connections) is the same, i.e. in a 3-head model all the outputs from layer[-2] are connected to head0, head 1 and head 2 (so all the 186 outputs of the model), and in a single-head model all the outputs of layer[-2] are also connected to all the 186 outputs of the model. The 3 heads just allows us to specify groups of outputs. We can achieve the same by treating outputs 0 to 167, 168 to 179 and 180 to 186 independently (applying different losses, for example).\n)</p>\n\n<p>So, now, not only has the model to identify three things, but it is not constrained to choose one of each label... But you will have to constrain it :D What I mean by constraining is associating a segment of the output vector to a label (e.g. 0-168 are the first label, etc). And you have to do it in order to train it. Then, you would be back at 3-heads + sigmoid, which makes no sense in this scenario.</p>\n\n<p>I personally got quite bad results in another task using sigmoid where softmax could be used.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765432,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "03/06/2020 16:07:08",
          "content": "<p>From 186 sigmoid outputs you take max(0-167), max(7) and max(11) and you end up at the same place as with 3 heads + softmax.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "765246": "Most of public kernels approached this task using models with `3 heads + softmax`, @seesee used [`1 head + softmax`](https://www.kaggle.com/c/bengaliai-cv19/discussion/134161) and decoded predictions. \n\nAnyone tried to approach it as multilabel with a `1 head + sigmoid`.\n\nWhat would be pros and/or cons? if any....",
    "765252": "softmax is nothing but a more generalized version of the sigmoid. sigmoid will work only if there are two classes [0,1]. \nAs here minimum class is 7. even if you use one head still you have to go for softmax",
    "765262": "Nope, you are wrong. You don't need to go with softmax and here is why.\nWith sigmoid function you predict absence-presence [0, 1] for each of the 186 (168+7+11) classes and each image has three positives.\n\nI know it can be used. Just wonder what more experienced guys think why this could or couldn't work.\n\nUPDATED: previous version had wrong number of classes thx @carlosmir",
    "765265": "well, in that case, i am not sure",
    "765275": "```\nThere are roughly 10,000 possible graphemes, of which roughly 1,000 are represented in the training set. \n```\nRead the rules carefully, It means that you don't know exactly how much classes we have.\nBut maximum is 168 * 11 * 7 = 12936 classes.",
    "765414": "TL;DR: I got bad results using sigmoid where softmax can be used. Softmax helps convergence.\n\n---\n\nHi @valanm \nFrom what I understand of @seesee's work, graphemes are identified by a unique triplet `(grapheme_root, vowel_diacritic, consonant_diacritic)`. Out of 12 936 possible combinations, there are only 1 292 present in the training data set.\n\n&gt; With sigmoid function you predict absence-presence [0, 1] for each of the 1292 classes and each image has three positives.\n\nI don't exactly get what you mean here, but as @mks2192 suggested, using 1292 classes one can only think of using softmax, since we know there is a single correct answer for each label value. There are no _three positives_ for each image in this setting.\n\n--- \n\nHowever, I guess there is a way to make it work. From the top of my head, the model would be identifying the three labels in a single vector, i.e. the labels would be 168+11+7=186 sized binary vectors and the model would have to set three of these 186 components to 1.\n\n(\nJust a little parenthesis. Single-head model === 3-head model. The number of layers and parameters (connections) is the same, i.e. in a 3-head model all the outputs from layer[-2] are connected to head0, head 1 and head 2 (so all the 186 outputs of the model), and in a single-head model all the outputs of layer[-2] are also connected to all the 186 outputs of the model. The 3 heads just allows us to specify groups of outputs. We can achieve the same by treating outputs 0 to 167, 168 to 179 and 180 to 186 independently (applying different losses, for example).\n)\n\nSo, now, not only has the model to identify three things, but it is not constrained to choose one of each label... But you will have to constrain it :D What I mean by constraining is associating a segment of the output vector to a label (e.g. 0-168 are the first label, etc). And you have to do it in order to train it. Then, you would be back at 3-heads + sigmoid, which makes no sense in this scenario.\n\nI personally got quite bad results in another task using sigmoid where softmax could be used.",
    "765432": "From 186 sigmoid outputs you take max(0-167), max(7) and max(11) and you end up at the same place as with 3 heads + softmax."
  },
  "source": "meta"
}