{
  "id": 170596,
  "title": "Why the probabilities are all under 0.5?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/170596",
  "author_name": "",
  "post_date": "2020-07-28T10:47:57.373116400Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I was wondering why both my results and some others' looks are all under 0.5. however, the scores are not bad--that may because most of the test targets are 0.  And if I want to calculate the accuracy using .round(), what I got is a bunch of zeros, which leads the accuracy to be around 0.98 which is about the percent of the zeros in the dataset.</p>\n\n<p>So how can we say a prediction is valid if it predicts all data with probabilities under 0.5? Or the point is not about the accuracy after .round() but the auc_score?</p>",
  "messages": [
    {
      "id": "948964",
      "postDate": "07/28/2020 10:47:57",
      "content": "<p>I was wondering why both my results and some others' looks are all under 0.5. however, the scores are not bad--that may because most of the test targets are 0.  And if I want to calculate the accuracy using .round(), what I got is a bunch of zeros, which leads the accuracy to be around 0.98 which is about the percent of the zeros in the dataset.</p>\n\n<p>So how can we say a prediction is valid if it predicts all data with probabilities under 0.5? Or the point is not about the accuracy after .round() but the auc_score?</p>",
      "rawMarkdown": "I was wondering why both my results and some others' looks are all under 0.5. however, the scores are not bad--that may because most of the test targets are 0.  And if I want to calculate the accuracy using .round(), what I got is a bunch of zeros, which leads the accuracy to be around 0.98 which is about the percent of the zeros in the dataset.\n\nSo how can we say a prediction is valid if it predicts all data with probabilities under 0.5? Or the point is not about the accuracy after .round() but the auc_score?",
      "votes": null
    },
    {
      "id": "948973",
      "postDate": "07/28/2020 10:57:54",
      "content": "<p>First of all, for such an imbalanced dataset you shouldn't look at accuracy but something like F1 score. 0.98 accuracy has no meaning here if 98% of the samples belong to one class. Also, you need to optimise threshold (not necessarily be 0.5).\nThat being said the metric of the competition is ROC-AUC score which takes into account the order/rank of soft probabilities instead of (0 or 1) absolute predictions. So, if your OOF AUC is great, you shouldn't worry about whether all or most of the probabilities is under 0.5. I hope that this was helpful. </p>",
      "rawMarkdown": "First of all, for such an imbalanced dataset you shouldn't look at accuracy but something like F1 score. 0.98 accuracy has no meaning here if 98% of the samples belong to one class. Also, you need to optimise threshold (not necessarily be 0.5).\nThat being said the metric of the competition is ROC-AUC score which takes into account the order/rank of soft probabilities instead of (0 or 1) absolute predictions. So, if your OOF AUC is great, you shouldn't worry about whether all or most of the probabilities is under 0.5. I hope that this was helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 948973,
      "author_name": "ks2019",
      "author_url": "",
      "post_date": "07/28/2020 10:57:54",
      "content": "<p>First of all, for such an imbalanced dataset you shouldn't look at accuracy but something like F1 score. 0.98 accuracy has no meaning here if 98% of the samples belong to one class. Also, you need to optimise threshold (not necessarily be 0.5).\nThat being said the metric of the competition is ROC-AUC score which takes into account the order/rank of soft probabilities instead of (0 or 1) absolute predictions. So, if your OOF AUC is great, you shouldn't worry about whether all or most of the probabilities is under 0.5. I hope that this was helpful. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "948964": "I was wondering why both my results and some others' looks are all under 0.5. however, the scores are not bad--that may because most of the test targets are 0.  And if I want to calculate the accuracy using .round(), what I got is a bunch of zeros, which leads the accuracy to be around 0.98 which is about the percent of the zeros in the dataset.\n\nSo how can we say a prediction is valid if it predicts all data with probabilities under 0.5? Or the point is not about the accuracy after .round() but the auc_score?",
    "948973": "First of all, for such an imbalanced dataset you shouldn't look at accuracy but something like F1 score. 0.98 accuracy has no meaning here if 98% of the samples belong to one class. Also, you need to optimise threshold (not necessarily be 0.5).\nThat being said the metric of the competition is ROC-AUC score which takes into account the order/rank of soft probabilities instead of (0 or 1) absolute predictions. So, if your OOF AUC is great, you shouldn't worry about whether all or most of the probabilities is under 0.5. I hope that this was helpful."
  },
  "source": "meta"
}