{
  "id": 89823,
  "title": "Multi label Multi class classification",
  "url": "/competitions/imet-2019-fgvc6/discussion/89823",
  "author_name": "",
  "post_date": "2019-04-17T23:39:27.159518500Z",
  "votes": 18,
  "comment_count": 11,
  "views": 0,
  "content": "<p>There are many classes, and for each image, we should associate one or many labels</p>\n\n<p>As a newbie I have some questions that I can't answer:</p>\n\n<p>What is the appropriate, Activation to use <strong>Softmax</strong> or <strong>Sigmoid</strong> ?\nWhat is the appropriate Loss to use, <strong>BCE</strong> or <strong>Categorical-cross-entropy</strong> or <strong>Focal Loss</strong> ?\nWhat metric to use, <strong>F2-score</strong> as mentioned in the evaluation section, <strong>accuracy</strong> or <strong>top_k_categorical_accuracy</strong> ?</p>",
  "messages": [
    {
      "id": "518848",
      "postDate": "04/17/2019 23:39:27",
      "content": "<p>There are many classes, and for each image, we should associate one or many labels</p>\n\n<p>As a newbie I have some questions that I can't answer:</p>\n\n<p>What is the appropriate, Activation to use <strong>Softmax</strong> or <strong>Sigmoid</strong> ?\nWhat is the appropriate Loss to use, <strong>BCE</strong> or <strong>Categorical-cross-entropy</strong> or <strong>Focal Loss</strong> ?\nWhat metric to use, <strong>F2-score</strong> as mentioned in the evaluation section, <strong>accuracy</strong> or <strong>top_k_categorical_accuracy</strong> ?</p>",
      "rawMarkdown": "There are many classes, and for each image, we should associate one or many labels\n\nAs a newbie I have some questions that I can't answer:\n\nWhat is the appropriate, Activation to use **Softmax** or **Sigmoid** ?\nWhat is the appropriate Loss to use, **BCE** or **Categorical-cross-entropy** or **Focal Loss** ?\nWhat metric to use, **F2-score** as mentioned in the evaluation section, **accuracy** or **top_k_categorical_accuracy** ?",
      "votes": null
    },
    {
      "id": "518867",
      "postDate": "04/18/2019 01:14:32",
      "content": "<p>Sigmoid for multiclass. Focal for unbalanced labels sounds logical (but people said BCE can have the same performance). F2 as the competition metric.</p>",
      "rawMarkdown": "Sigmoid for multiclass. Focal for unbalanced labels sounds logical (but people said BCE can have the same performance). F2 as the competition metric.",
      "votes": null
    },
    {
      "id": "518872",
      "postDate": "04/18/2019 01:44:37",
      "content": "<p>It depends on how you reformulate the problem. Canonically we'd use SoftMax for multiclass classification (since SoftMax probabilities sums up to one) and Sigmoid for multi-label classification (2+ labels may be assigned to a sample, thus probabilities individually sum to one). But here's the catch: each dataset has different interesting properties that you can exploit to fit the metric you want.</p>\n\n<p>Here are two examples:\nIn the Amazon Rainforest competition, we are dealing with a classical multi-label classification with abundant data and F2 eval metric. Thus, it is easy to get a .90+ leaderboard score with a simple resnet and logloss. However, EDA reveals that logloss at the high .93 region does not line up with the F2 metric, which is what we care, and this became rather \"lethal\" as the public leaderboard continued to saturate. So some top-performers ended up using a differentiable version of soft-F2 loss to regularize the model further.</p>\n\n<p>In the recent Humpback Whale identification competition, we are presumably dealing with a multi-class classification (since whale can only have one id) with top-5 accuracy. In this case, however, EDA reveals two properties of the dataset: the heterogeneity of 'new0whale' class and noises in whale-ids. The first is really simple: neural networks are optimized with the assumption that samples with the same labels are inherently similar, but this is not the case with 'new-whale'. The second one has to do with faults during the labeling process: the same whale might be labeled with two different ids. So, instead of using SoftMax (which performs terribly already with the long-tailed feature of the dataset), our strategy was to use sigmoid loss without \"new-whale\" class. Thus, we are actually dealing with an one-vs-rest classification problem where we are making multiple decisions on whether the whale belongs to a particular id. Then we implemented an heuristic to determine the \"cutoff\", so that the label \"new-whale\" can be inserted (we are predicting top-5 labels for a sample, so we need to figure out where to insert the 'new-whale' label or if we should insert it at all). Of course we ended up exploiting the new_whale images by metric learning, which you can read off my teammate <a href=\"/qiaojian\">@qiaojian</a> 's fabulous post, but I think you've already gotten the gist here.</p>\n\n<p>How about iMet? Well, it's time for you to do experiments to figure it out :)</p>\n\n<p>Alex</p>",
      "rawMarkdown": "It depends on how you reformulate the problem. Canonically we'd use SoftMax for multiclass classification (since SoftMax probabilities sums up to one) and Sigmoid for multi-label classification (2+ labels may be assigned to a sample, thus probabilities individually sum to one). But here's the catch: each dataset has different interesting properties that you can exploit to fit the metric you want.\n\nHere are two examples:\nIn the Amazon Rainforest competition, we are dealing with a classical multi-label classification with abundant data and F2 eval metric. Thus, it is easy to get a .90+ leaderboard score with a simple resnet and logloss. However, EDA reveals that logloss at the high .93 region does not line up with the F2 metric, which is what we care, and this became rather \"lethal\" as the public leaderboard continued to saturate. So some top-performers ended up using a differentiable version of soft-F2 loss to regularize the model further.\n\nIn the recent Humpback Whale identification competition, we are presumably dealing with a multi-class classification (since whale can only have one id) with top-5 accuracy. In this case, however, EDA reveals two properties of the dataset: the heterogeneity of 'new0whale' class and noises in whale-ids. The first is really simple: neural networks are optimized with the assumption that samples with the same labels are inherently similar, but this is not the case with 'new-whale'. The second one has to do with faults during the labeling process: the same whale might be labeled with two different ids. So, instead of using SoftMax (which performs terribly already with the long-tailed feature of the dataset), our strategy was to use sigmoid loss without \"new-whale\" class. Thus, we are actually dealing with an one-vs-rest classification problem where we are making multiple decisions on whether the whale belongs to a particular id. Then we implemented an heuristic to determine the \"cutoff\", so that the label \"new-whale\" can be inserted (we are predicting top-5 labels for a sample, so we need to figure out where to insert the 'new-whale' label or if we should insert it at all). Of course we ended up exploiting the new_whale images by metric learning, which you can read off my teammate @qiaojian 's fabulous post, but I think you've already gotten the gist here.\n\nHow about iMet? Well, it's time for you to do experiments to figure it out :)\n\nAlex",
      "votes": null
    },
    {
      "id": "518899",
      "postDate": "04/18/2019 03:19:44",
      "content": "<p>Awesome Alex! Thank you for your post.</p>\n\n<p>I studied your Whale solution but at that time was quite not understand the logic behind the using of binary classifications in that competition. Now I understand much better, thanks again!</p>",
      "rawMarkdown": "Awesome Alex! Thank you for your post.\n\nI studied your Whale solution but at that time was quite not understand the logic behind the using of binary classifications in that competition. Now I understand much better, thanks again!",
      "votes": null
    },
    {
      "id": "518905",
      "postDate": "04/18/2019 03:29:21",
      "content": "<p>Apart from <a href=\"/alexanderliao\">@alexanderliao</a>  Alex's answer (which is already superb), I would like to add more explanation on vanilla multi-label vs. multi-classes formulation here which may useful for newcomers. Note that the following sentence looks at the problem as a vanilla (simple) formulation as oppose to Alex's answer below which is much deeper : </p>\n\n<h3>On sigmoid</h3>\n\n<p>In a simple formulation of a <strong>‘multi-labels’</strong> problem, by definition,  we want to be able to answer more than '1' label for each example. This is oppose to the <strong>‘multi-classes’</strong> problem where we only want a 'single' answer. </p>\n\n<p>In multi-label, <code>softmax</code> is not appropriate because <code>softmax</code> will prefer only 1 answer, e.g. in this problem, it can give one class as probability 0.99 and all other 1102 classes very tiny probabilities.</p>\n\n<p>In contrast, if we use <code>sigmoid</code> with 1103 outputs with binary crossentropy loss, we can treat this problem as 1103 binary classification, so the network can say <em>’yes’</em> to many labels at once (e.g. give 0.90++ to many labels)</p>\n\n<p>This is the goal of this competition — we can tag more than one labels into one example.</p>\n\n<h3>On BCE vs. Focal vs. Categorical CE</h3>\n\n<p>First, if we simply look at the problem as multi-label as stated above, we now understand why BCE is preferable to Categorical CE.\nSecondly, since BCE is a special case of a focal loss (gamma = 0), either one can be chosen based on validation performance.</p>\n\n<p>Hope this help!</p>",
      "rawMarkdown": "Apart from @alexanderliao  Alex's answer (which is already superb), I would like to add more explanation on vanilla multi-label vs. multi-classes formulation here which may useful for newcomers. Note that the following sentence looks at the problem as a vanilla (simple) formulation as oppose to Alex's answer below which is much deeper : \n\n### On sigmoid\nIn a simple formulation of a **‘multi-labels’** problem, by definition,  we want to be able to answer more than '1' label for each example. This is oppose to the **‘multi-classes’** problem where we only want a 'single' answer. \n\nIn multi-label, `softmax` is not appropriate because `softmax` will prefer only 1 answer, e.g. in this problem, it can give one class as probability 0.99 and all other 1102 classes very tiny probabilities.\n\nIn contrast, if we use `sigmoid` with 1103 outputs with binary crossentropy loss, we can treat this problem as 1103 binary classification, so the network can say *’yes’* to many labels at once (e.g. give 0.90++ to many labels)\n\nThis is the goal of this competition — we can tag more than one labels into one example.\n\n### On BCE vs. Focal vs. Categorical CE\nFirst, if we simply look at the problem as multi-label as stated above, we now understand why BCE is preferable to Categorical CE.\nSecondly, since BCE is a special case of a focal loss (gamma = 0), either one can be chosen based on validation performance.\n\nHope this help!",
      "votes": null
    },
    {
      "id": "518936",
      "postDate": "04/18/2019 05:20:55",
      "content": "<p>Thank you, Great explanation.</p>",
      "rawMarkdown": "Thank you, Great explanation.",
      "votes": null
    },
    {
      "id": "519269",
      "postDate": "04/18/2019 16:37:00",
      "content": "<p>Thanks for the great explanation of your whale solution! </p>\n\n<p>I’ve been going over my and others solutions for hpa to remind myself about dealing with multilabels. The difference between 28 and 1103 labels seems like a lot, but due to the nature of the objects I think it might not be. There are sets of easy labels, like dagger for example, that should be fairly easy to assign with high accuracy. And sets of more and more difficult labels, like is this dagger Egyptian or Chinese or Etruscan? </p>\n\n<p>So I think ultimately some combo of multilabel multicat could be a good way to go, though I have no idea how to do it efficiently yet...</p>",
      "rawMarkdown": "Thanks for the great explanation of your whale solution! \n\nI’ve been going over my and others solutions for hpa to remind myself about dealing with multilabels. The difference between 28 and 1103 labels seems like a lot, but due to the nature of the objects I think it might not be. There are sets of easy labels, like dagger for example, that should be fairly easy to assign with high accuracy. And sets of more and more difficult labels, like is this dagger Egyptian or Chinese or Etruscan? \n\nSo I think ultimately some combo of multilabel multicat could be a good way to go, though I have no idea how to do it efficiently yet...",
      "votes": null
    },
    {
      "id": "521093",
      "postDate": "04/22/2019 10:03:18",
      "content": "<p>Thanks for sharing! </p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "526154",
      "postDate": "05/02/2019 13:21:32",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a> have you used different thresholds for different classes or the same threshold for all the classes with sigmoid?</p>",
      "rawMarkdown": "ratthachat have you used different thresholds for different classes or the same threshold for all the classes with sigmoid?",
      "votes": null
    },
    {
      "id": "526166",
      "postDate": "05/02/2019 13:38:57",
      "content": "<p>Hi <a href=\"/rajeshbhat\">@rajeshbhat</a> , I did try a few ideas of using multi-thresholds; however, in my experiments, I found that optimal thresholds for each label are not quite different to the optimal single threshold, so I decide to use only single threshold so far.</p>",
      "rawMarkdown": "Hi @rajeshbhat , I did try a few ideas of using multi-thresholds; however, in my experiments, I found that optimal thresholds for each label are not quite different to the optimal single threshold, so I decide to use only single threshold so far.",
      "votes": null
    },
    {
      "id": "526206",
      "postDate": "05/02/2019 15:11:32",
      "content": "<p>Thanks for the quick reply <a href=\"/ratthachat\">@ratthachat</a> </p>",
      "rawMarkdown": "Thanks for the quick reply @ratthachat",
      "votes": null
    },
    {
      "id": "615954",
      "postDate": "09/02/2019 14:38:39",
      "content": "<p>Wow, these explanations help me a lot; thanks for sharing!</p>",
      "rawMarkdown": "Wow, these explanations help me a lot; thanks for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 518867,
      "author_name": "kokecacao",
      "author_url": "",
      "post_date": "04/18/2019 01:14:32",
      "content": "<p>Sigmoid for multiclass. Focal for unbalanced labels sounds logical (but people said BCE can have the same performance). F2 as the competition metric.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 518872,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "04/18/2019 01:44:37",
      "content": "<p>It depends on how you reformulate the problem. Canonically we'd use SoftMax for multiclass classification (since SoftMax probabilities sums up to one) and Sigmoid for multi-label classification (2+ labels may be assigned to a sample, thus probabilities individually sum to one). But here's the catch: each dataset has different interesting properties that you can exploit to fit the metric you want.</p>\n\n<p>Here are two examples:\nIn the Amazon Rainforest competition, we are dealing with a classical multi-label classification with abundant data and F2 eval metric. Thus, it is easy to get a .90+ leaderboard score with a simple resnet and logloss. However, EDA reveals that logloss at the high .93 region does not line up with the F2 metric, which is what we care, and this became rather \"lethal\" as the public leaderboard continued to saturate. So some top-performers ended up using a differentiable version of soft-F2 loss to regularize the model further.</p>\n\n<p>In the recent Humpback Whale identification competition, we are presumably dealing with a multi-class classification (since whale can only have one id) with top-5 accuracy. In this case, however, EDA reveals two properties of the dataset: the heterogeneity of 'new0whale' class and noises in whale-ids. The first is really simple: neural networks are optimized with the assumption that samples with the same labels are inherently similar, but this is not the case with 'new-whale'. The second one has to do with faults during the labeling process: the same whale might be labeled with two different ids. So, instead of using SoftMax (which performs terribly already with the long-tailed feature of the dataset), our strategy was to use sigmoid loss without \"new-whale\" class. Thus, we are actually dealing with an one-vs-rest classification problem where we are making multiple decisions on whether the whale belongs to a particular id. Then we implemented an heuristic to determine the \"cutoff\", so that the label \"new-whale\" can be inserted (we are predicting top-5 labels for a sample, so we need to figure out where to insert the 'new-whale' label or if we should insert it at all). Of course we ended up exploiting the new_whale images by metric learning, which you can read off my teammate <a href=\"/qiaojian\">@qiaojian</a> 's fabulous post, but I think you've already gotten the gist here.</p>\n\n<p>How about iMet? Well, it's time for you to do experiments to figure it out :)</p>\n\n<p>Alex</p>",
      "votes": null,
      "replies": [
        {
          "id": 518899,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "04/18/2019 03:19:44",
          "content": "<p>Awesome Alex! Thank you for your post.</p>\n\n<p>I studied your Whale solution but at that time was quite not understand the logic behind the using of binary classifications in that competition. Now I understand much better, thanks again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 519269,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "04/18/2019 16:37:00",
          "content": "<p>Thanks for the great explanation of your whale solution! </p>\n\n<p>I’ve been going over my and others solutions for hpa to remind myself about dealing with multilabels. The difference between 28 and 1103 labels seems like a lot, but due to the nature of the objects I think it might not be. There are sets of easy labels, like dagger for example, that should be fairly easy to assign with high accuracy. And sets of more and more difficult labels, like is this dagger Egyptian or Chinese or Etruscan? </p>\n\n<p>So I think ultimately some combo of multilabel multicat could be a good way to go, though I have no idea how to do it efficiently yet...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 521093,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "04/22/2019 10:03:18",
          "content": "<p>Thanks for sharing! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615954,
          "author_name": "phunghieu",
          "author_url": "",
          "post_date": "09/02/2019 14:38:39",
          "content": "<p>Wow, these explanations help me a lot; thanks for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 518905,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "04/18/2019 03:29:21",
      "content": "<p>Apart from <a href=\"/alexanderliao\">@alexanderliao</a>  Alex's answer (which is already superb), I would like to add more explanation on vanilla multi-label vs. multi-classes formulation here which may useful for newcomers. Note that the following sentence looks at the problem as a vanilla (simple) formulation as oppose to Alex's answer below which is much deeper : </p>\n\n<h3>On sigmoid</h3>\n\n<p>In a simple formulation of a <strong>‘multi-labels’</strong> problem, by definition,  we want to be able to answer more than '1' label for each example. This is oppose to the <strong>‘multi-classes’</strong> problem where we only want a 'single' answer. </p>\n\n<p>In multi-label, <code>softmax</code> is not appropriate because <code>softmax</code> will prefer only 1 answer, e.g. in this problem, it can give one class as probability 0.99 and all other 1102 classes very tiny probabilities.</p>\n\n<p>In contrast, if we use <code>sigmoid</code> with 1103 outputs with binary crossentropy loss, we can treat this problem as 1103 binary classification, so the network can say <em>’yes’</em> to many labels at once (e.g. give 0.90++ to many labels)</p>\n\n<p>This is the goal of this competition — we can tag more than one labels into one example.</p>\n\n<h3>On BCE vs. Focal vs. Categorical CE</h3>\n\n<p>First, if we simply look at the problem as multi-label as stated above, we now understand why BCE is preferable to Categorical CE.\nSecondly, since BCE is a special case of a focal loss (gamma = 0), either one can be chosen based on validation performance.</p>\n\n<p>Hope this help!</p>",
      "votes": null,
      "replies": [
        {
          "id": 518936,
          "author_name": "jmourad100",
          "author_url": "",
          "post_date": "04/18/2019 05:20:55",
          "content": "<p>Thank you, Great explanation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526154,
          "author_name": "rajeshbhat",
          "author_url": "",
          "post_date": "05/02/2019 13:21:32",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> have you used different thresholds for different classes or the same threshold for all the classes with sigmoid?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526166,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "05/02/2019 13:38:57",
          "content": "<p>Hi <a href=\"/rajeshbhat\">@rajeshbhat</a> , I did try a few ideas of using multi-thresholds; however, in my experiments, I found that optimal thresholds for each label are not quite different to the optimal single threshold, so I decide to use only single threshold so far.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526206,
          "author_name": "rajeshbhat",
          "author_url": "",
          "post_date": "05/02/2019 15:11:32",
          "content": "<p>Thanks for the quick reply <a href=\"/ratthachat\">@ratthachat</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "518848": "There are many classes, and for each image, we should associate one or many labels\n\nAs a newbie I have some questions that I can't answer:\n\nWhat is the appropriate, Activation to use **Softmax** or **Sigmoid** ?\nWhat is the appropriate Loss to use, **BCE** or **Categorical-cross-entropy** or **Focal Loss** ?\nWhat metric to use, **F2-score** as mentioned in the evaluation section, **accuracy** or **top_k_categorical_accuracy** ?",
    "518867": "Sigmoid for multiclass. Focal for unbalanced labels sounds logical (but people said BCE can have the same performance). F2 as the competition metric.",
    "518872": "It depends on how you reformulate the problem. Canonically we'd use SoftMax for multiclass classification (since SoftMax probabilities sums up to one) and Sigmoid for multi-label classification (2+ labels may be assigned to a sample, thus probabilities individually sum to one). But here's the catch: each dataset has different interesting properties that you can exploit to fit the metric you want.\n\nHere are two examples:\nIn the Amazon Rainforest competition, we are dealing with a classical multi-label classification with abundant data and F2 eval metric. Thus, it is easy to get a .90+ leaderboard score with a simple resnet and logloss. However, EDA reveals that logloss at the high .93 region does not line up with the F2 metric, which is what we care, and this became rather \"lethal\" as the public leaderboard continued to saturate. So some top-performers ended up using a differentiable version of soft-F2 loss to regularize the model further.\n\nIn the recent Humpback Whale identification competition, we are presumably dealing with a multi-class classification (since whale can only have one id) with top-5 accuracy. In this case, however, EDA reveals two properties of the dataset: the heterogeneity of 'new0whale' class and noises in whale-ids. The first is really simple: neural networks are optimized with the assumption that samples with the same labels are inherently similar, but this is not the case with 'new-whale'. The second one has to do with faults during the labeling process: the same whale might be labeled with two different ids. So, instead of using SoftMax (which performs terribly already with the long-tailed feature of the dataset), our strategy was to use sigmoid loss without \"new-whale\" class. Thus, we are actually dealing with an one-vs-rest classification problem where we are making multiple decisions on whether the whale belongs to a particular id. Then we implemented an heuristic to determine the \"cutoff\", so that the label \"new-whale\" can be inserted (we are predicting top-5 labels for a sample, so we need to figure out where to insert the 'new-whale' label or if we should insert it at all). Of course we ended up exploiting the new_whale images by metric learning, which you can read off my teammate @qiaojian 's fabulous post, but I think you've already gotten the gist here.\n\nHow about iMet? Well, it's time for you to do experiments to figure it out :)\n\nAlex",
    "518899": "Awesome Alex! Thank you for your post.\n\nI studied your Whale solution but at that time was quite not understand the logic behind the using of binary classifications in that competition. Now I understand much better, thanks again!",
    "518905": "Apart from @alexanderliao  Alex's answer (which is already superb), I would like to add more explanation on vanilla multi-label vs. multi-classes formulation here which may useful for newcomers. Note that the following sentence looks at the problem as a vanilla (simple) formulation as oppose to Alex's answer below which is much deeper : \n\n### On sigmoid\nIn a simple formulation of a **‘multi-labels’** problem, by definition,  we want to be able to answer more than '1' label for each example. This is oppose to the **‘multi-classes’** problem where we only want a 'single' answer. \n\nIn multi-label, `softmax` is not appropriate because `softmax` will prefer only 1 answer, e.g. in this problem, it can give one class as probability 0.99 and all other 1102 classes very tiny probabilities.\n\nIn contrast, if we use `sigmoid` with 1103 outputs with binary crossentropy loss, we can treat this problem as 1103 binary classification, so the network can say *’yes’* to many labels at once (e.g. give 0.90++ to many labels)\n\nThis is the goal of this competition — we can tag more than one labels into one example.\n\n### On BCE vs. Focal vs. Categorical CE\nFirst, if we simply look at the problem as multi-label as stated above, we now understand why BCE is preferable to Categorical CE.\nSecondly, since BCE is a special case of a focal loss (gamma = 0), either one can be chosen based on validation performance.\n\nHope this help!",
    "518936": "Thank you, Great explanation.",
    "519269": "Thanks for the great explanation of your whale solution! \n\nI’ve been going over my and others solutions for hpa to remind myself about dealing with multilabels. The difference between 28 and 1103 labels seems like a lot, but due to the nature of the objects I think it might not be. There are sets of easy labels, like dagger for example, that should be fairly easy to assign with high accuracy. And sets of more and more difficult labels, like is this dagger Egyptian or Chinese or Etruscan? \n\nSo I think ultimately some combo of multilabel multicat could be a good way to go, though I have no idea how to do it efficiently yet...",
    "521093": "Thanks for sharing!",
    "526154": "ratthachat have you used different thresholds for different classes or the same threshold for all the classes with sigmoid?",
    "526166": "Hi @rajeshbhat , I did try a few ideas of using multi-thresholds; however, in my experiments, I found that optimal thresholds for each label are not quite different to the optimal single threshold, so I decide to use only single threshold so far.",
    "526206": "Thanks for the quick reply @ratthachat",
    "615954": "Wow, these explanations help me a lot; thanks for sharing!"
  },
  "source": "meta"
}