{
  "id": 218141,
  "title": "Train 11 different classifiers or a single classifier with 11 labels",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/218141",
  "author_name": "",
  "post_date": "2021-02-09T12:16:48.738200200Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>On analyzing the data, I have come to see that there are some labels which are just way too imbalanced as compared to some other labels. Which is why making a single model with 11 targets would be tough w.r.t the metric (AUC since many a times there will be a case where all the labels of a batch belong to the same class) and also because class imbalance is just bad in general. <br>\nCome to making 11 different classifiers, the idea seems, well, clearly very inefficient. Can someone shed some light on what method they are using, if they are able to balance the labels in a way that the can use a classifier with 11 different targets rather than the second option.</p>",
  "messages": [
    {
      "id": "1193020",
      "postDate": "02/09/2021 12:16:48",
      "content": "<p>On analyzing the data, I have come to see that there are some labels which are just way too imbalanced as compared to some other labels. Which is why making a single model with 11 targets would be tough w.r.t the metric (AUC since many a times there will be a case where all the labels of a batch belong to the same class) and also because class imbalance is just bad in general. <br>\nCome to making 11 different classifiers, the idea seems, well, clearly very inefficient. Can someone shed some light on what method they are using, if they are able to balance the labels in a way that the can use a classifier with 11 different targets rather than the second option.</p>",
      "rawMarkdown": "On analyzing the data, I have come to see that there are some labels which are just way too imbalanced as compared to some other labels. Which is why making a single model with 11 targets would be tough w.r.t the metric (AUC since many a times there will be a case where all the labels of a batch belong to the same class) and also because class imbalance is just bad in general. \nCome to making 11 different classifiers, the idea seems, well, clearly very inefficient. Can someone shed some light on what method they are using, if they are able to balance the labels in a way that the can use a classifier with 11 different targets rather than the second option.",
      "votes": null
    },
    {
      "id": "1193354",
      "postDate": "02/09/2021 15:38:01",
      "content": "<p>There are a few misconceptions in your post:</p>\n<ul>\n<li><p>AUC metric is usually not calculated in a batch - differentiable loss like binary cross entropy is preferred. AUC is overall ranking metric and cannot be optimized directly by optimizers (there are a few proxy AUC loss functions… but this is not the question here)</p></li>\n<li><p>\"class imbalance is just bad in general\" - this is statement not everyone would fully agree - preserving real life example distributions can be beneficial. Dealing with class imbalance is a matter of experimentation with example weights during training, using weighted loss functions (i.e. focal loss), etc..</p></li>\n<li><p>having different models for different label does not necessarily yield better results (my personal experience). Having single multilabel model which can share similar feature maps may help model to better generalize between classes-for example in this competition <code>CVC - Normal</code> and <code>CVC - Borderline</code> classes are very similar and only slight differences can help tell them apart.</p></li>\n</ul>",
      "rawMarkdown": "There are a few misconceptions in your post:\n\n- AUC metric is usually not calculated in a batch - differentiable loss like binary cross entropy is preferred. AUC is overall ranking metric and cannot be optimized directly by optimizers (there are a few proxy AUC loss functions... but this is not the question here)\n\n- \"class imbalance is just bad in general\" - this is statement not everyone would fully agree - preserving real life example distributions can be beneficial. Dealing with class imbalance is a matter of experimentation with example weights during training, using weighted loss functions (i.e. focal loss), etc..\n\n- having different models for different label does not necessarily yield better results (my personal experience). Having single multilabel model which can share similar feature maps may help model to better generalize between classes-for example in this competition `CVC - Normal` and `CVC - Borderline` classes are very similar and only slight differences can help tell them apart.",
      "votes": null
    },
    {
      "id": "1193619",
      "postDate": "02/09/2021 18:42:31",
      "content": "<p>Thanks alot for the reply, it seems I am not conceptually clear about the metric but I still feel I may not have been able to get my point across.<br>\nAs far as I know auc is the area under the ROC curve which is basically the curve of true positive rate against the false positive rate.<br>\nMy question being if there is no true positive in the ground truth labels in a batch (arising due to the class imbalance), how can we calculate the true positive rate or even the false positive rate for that matter for the batch during validation. <br>\nIs the only way to calculate the metric, taking the entire validation dset and feeding it in as a single batch since that will have some positive examples in the same for sure?</p>",
      "rawMarkdown": "Thanks alot for the reply, it seems I am not conceptually clear about the metric but I still feel I may not have been able to get my point across.\nAs far as I know auc is the area under the ROC curve which is basically the curve of true positive rate against the false positive rate.\nMy question being if there is no true positive in the ground truth labels in a batch (arising due to the class imbalance), how can we calculate the true positive rate or even the false positive rate for that matter for the batch during validation. \nIs the only way to calculate the metric, taking the entire validation dset and feeding it in as a single batch since that will have some positive examples in the same for sure?",
      "votes": null
    },
    {
      "id": "1194418",
      "postDate": "02/10/2021 07:32:43",
      "content": "<pre><code>My question being if there is no true positive in the ground truth labels in a batch (arising due to the class imbalance), how can we calculate the true positive rate or even the false positive rate for that matter for the batch during validation.\n</code></pre>\n<p>Answer - you cannot. And there is no reason for you to calculate AUC in batch.</p>\n<pre><code>Is the only way to calculate the metric, taking the entire validation dset and feeding it in as a single batch since that will have some positive examples in the same for sure?\n</code></pre>\n<p>That is the idea. But you can do it in mini batches - run inference without batch metric calculations - concatenate all mini batch predictions and then calculate overall AUC. That is what everyone is doing.</p>",
      "rawMarkdown": "```\nMy question being if there is no true positive in the ground truth labels in a batch (arising due to the class imbalance), how can we calculate the true positive rate or even the false positive rate for that matter for the batch during validation.\n```\nAnswer - you cannot. And there is no reason for you to calculate AUC in batch.\n\n```\nIs the only way to calculate the metric, taking the entire validation dset and feeding it in as a single batch since that will have some positive examples in the same for sure?\n```\nThat is the idea. But you can do it in mini batches - run inference without batch metric calculations - concatenate all mini batch predictions and then calculate overall AUC. That is what everyone is doing.",
      "votes": null
    },
    {
      "id": "1195278",
      "postDate": "02/10/2021 16:49:10",
      "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> thanks for the reply, cleared my doubt completely</p>",
      "rawMarkdown": "raddar thanks for the reply, cleared my doubt completely",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1193354,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "02/09/2021 15:38:01",
      "content": "<p>There are a few misconceptions in your post:</p>\n<ul>\n<li><p>AUC metric is usually not calculated in a batch - differentiable loss like binary cross entropy is preferred. AUC is overall ranking metric and cannot be optimized directly by optimizers (there are a few proxy AUC loss functions… but this is not the question here)</p></li>\n<li><p>\"class imbalance is just bad in general\" - this is statement not everyone would fully agree - preserving real life example distributions can be beneficial. Dealing with class imbalance is a matter of experimentation with example weights during training, using weighted loss functions (i.e. focal loss), etc..</p></li>\n<li><p>having different models for different label does not necessarily yield better results (my personal experience). Having single multilabel model which can share similar feature maps may help model to better generalize between classes-for example in this competition <code>CVC - Normal</code> and <code>CVC - Borderline</code> classes are very similar and only slight differences can help tell them apart.</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1193619,
          "author_name": "aryaman1999",
          "author_url": "",
          "post_date": "02/09/2021 18:42:31",
          "content": "<p>Thanks alot for the reply, it seems I am not conceptually clear about the metric but I still feel I may not have been able to get my point across.<br>\nAs far as I know auc is the area under the ROC curve which is basically the curve of true positive rate against the false positive rate.<br>\nMy question being if there is no true positive in the ground truth labels in a batch (arising due to the class imbalance), how can we calculate the true positive rate or even the false positive rate for that matter for the batch during validation. <br>\nIs the only way to calculate the metric, taking the entire validation dset and feeding it in as a single batch since that will have some positive examples in the same for sure?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1194418,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "02/10/2021 07:32:43",
          "content": "<pre><code>My question being if there is no true positive in the ground truth labels in a batch (arising due to the class imbalance), how can we calculate the true positive rate or even the false positive rate for that matter for the batch during validation.\n</code></pre>\n<p>Answer - you cannot. And there is no reason for you to calculate AUC in batch.</p>\n<pre><code>Is the only way to calculate the metric, taking the entire validation dset and feeding it in as a single batch since that will have some positive examples in the same for sure?\n</code></pre>\n<p>That is the idea. But you can do it in mini batches - run inference without batch metric calculations - concatenate all mini batch predictions and then calculate overall AUC. That is what everyone is doing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1195278,
          "author_name": "aryaman1999",
          "author_url": "",
          "post_date": "02/10/2021 16:49:10",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> thanks for the reply, cleared my doubt completely</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1193020": "On analyzing the data, I have come to see that there are some labels which are just way too imbalanced as compared to some other labels. Which is why making a single model with 11 targets would be tough w.r.t the metric (AUC since many a times there will be a case where all the labels of a batch belong to the same class) and also because class imbalance is just bad in general. \nCome to making 11 different classifiers, the idea seems, well, clearly very inefficient. Can someone shed some light on what method they are using, if they are able to balance the labels in a way that the can use a classifier with 11 different targets rather than the second option.",
    "1193354": "There are a few misconceptions in your post:\n\n- AUC metric is usually not calculated in a batch - differentiable loss like binary cross entropy is preferred. AUC is overall ranking metric and cannot be optimized directly by optimizers (there are a few proxy AUC loss functions... but this is not the question here)\n\n- \"class imbalance is just bad in general\" - this is statement not everyone would fully agree - preserving real life example distributions can be beneficial. Dealing with class imbalance is a matter of experimentation with example weights during training, using weighted loss functions (i.e. focal loss), etc..\n\n- having different models for different label does not necessarily yield better results (my personal experience). Having single multilabel model which can share similar feature maps may help model to better generalize between classes-for example in this competition `CVC - Normal` and `CVC - Borderline` classes are very similar and only slight differences can help tell them apart.",
    "1193619": "Thanks alot for the reply, it seems I am not conceptually clear about the metric but I still feel I may not have been able to get my point across.\nAs far as I know auc is the area under the ROC curve which is basically the curve of true positive rate against the false positive rate.\nMy question being if there is no true positive in the ground truth labels in a batch (arising due to the class imbalance), how can we calculate the true positive rate or even the false positive rate for that matter for the batch during validation. \nIs the only way to calculate the metric, taking the entire validation dset and feeding it in as a single batch since that will have some positive examples in the same for sure?",
    "1194418": "```\nMy question being if there is no true positive in the ground truth labels in a batch (arising due to the class imbalance), how can we calculate the true positive rate or even the false positive rate for that matter for the batch during validation.\n```\nAnswer - you cannot. And there is no reason for you to calculate AUC in batch.\n\n```\nIs the only way to calculate the metric, taking the entire validation dset and feeding it in as a single batch since that will have some positive examples in the same for sure?\n```\nThat is the idea. But you can do it in mini batches - run inference without batch metric calculations - concatenate all mini batch predictions and then calculate overall AUC. That is what everyone is doing.",
    "1195278": "raddar thanks for the reply, cleared my doubt completely"
  },
  "source": "meta"
}