{
  "id": 156064,
  "title": "Beware of non-calibrated predictions",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/156064",
  "author_name": "Dmytro Danevskyi",
  "post_date": "2020-06-04T09:55:56.829000",
  "votes": 56,
  "comment_count": 10,
  "views": 0,
  "content": "<p>The target metric in this competition is based on ranks rather than on actual values. That means that as long as the order of your values is fixed, the metric will stay the same.</p>\n\n<p>To illustrate:</p>\n\n<p>```\ntarget = [1, 0, 1, 1, 0]\npreds = [0.5, 0.25, 0.2, 0.3, 0.1]</p>\n\n<p>metric = roc_auc_score(target, preds)  # 0.833</p>\n\n<p>target = [1, 0, 1, 1, 0]\npreds = [0.7, 0.15, 0.1, 0.2, 0.05]</p>\n\n<p>metric = roc_auc_score(target, preds)  # 0.833\n```</p>\n\n<p>That means that two different models that give the <strong>same</strong> score could actually output completely <strong>different</strong> values. They are not even required to be in (0, 1) range!</p>\n\n<p>```\ntarget = [1, 0, 1, 1, 0]\npreds = [100, 25, 20, 30, 10]</p>\n\n<p>metric = roc_auc_score(target, preds)  # 0.833\n```</p>\n\n<p>Then, if you will try to average the predictions of two non-calibrated models, you might observe that the score is not necessarily getting better and in some cases, it could become even worse! This happens because the prediction scales of these two models are not directly comparable because of the aforementioned issue.</p>\n\n<p>How this can be fixed?</p>\n\n<p>One simple solution is to bring the predictions to the same scale, e.g. with <code>scipy.stats.rankdata</code> function. This will turn scores into ranks, i.e. <code>[0.7, 0.15, 0.1, 0.2, 0.05]</code> will be turned into <code>[5, 3, 2, 4, 1]</code>. After this, the predictions could be blended.</p>\n\n<p>Note that this not always lead to better results and is highly dependent on which exactly models are blended, how strong is the bias, and so on.</p>\n\n<p>To illustrate, I decided to naively blend my best scoring model (ResNet18) that gives 0.914 on the public leaderboard with the best-scoring public kernel available (0.927) and I got 0.925. After I preprocessed both predictions with <code>rankdata</code>, my score improved to 0.933.</p>",
  "messages": [
    {
      "id": 873610,
      "postDate": "2020-06-04T09:55:56.830Z",
      "content": "<p>The target metric in this competition is based on ranks rather than on actual values. That means that as long as the order of your values is fixed, the metric will stay the same.</p>\n\n<p>To illustrate:</p>\n\n<p>```\ntarget = [1, 0, 1, 1, 0]\npreds = [0.5, 0.25, 0.2, 0.3, 0.1]</p>\n\n<p>metric = roc_auc_score(target, preds)  # 0.833</p>\n\n<p>target = [1, 0, 1, 1, 0]\npreds = [0.7, 0.15, 0.1, 0.2, 0.05]</p>\n\n<p>metric = roc_auc_score(target, preds)  # 0.833\n```</p>\n\n<p>That means that two different models that give the <strong>same</strong> score could actually output completely <strong>different</strong> values. They are not even required to be in (0, 1) range!</p>\n\n<p>```\ntarget = [1, 0, 1, 1, 0]\npreds = [100, 25, 20, 30, 10]</p>\n\n<p>metric = roc_auc_score(target, preds)  # 0.833\n```</p>\n\n<p>Then, if you will try to average the predictions of two non-calibrated models, you might observe that the score is not necessarily getting better and in some cases, it could become even worse! This happens because the prediction scales of these two models are not directly comparable because of the aforementioned issue.</p>\n\n<p>How this can be fixed?</p>\n\n<p>One simple solution is to bring the predictions to the same scale, e.g. with <code>scipy.stats.rankdata</code> function. This will turn scores into ranks, i.e. <code>[0.7, 0.15, 0.1, 0.2, 0.05]</code> will be turned into <code>[5, 3, 2, 4, 1]</code>. After this, the predictions could be blended.</p>\n\n<p>Note that this not always lead to better results and is highly dependent on which exactly models are blended, how strong is the bias, and so on.</p>\n\n<p>To illustrate, I decided to naively blend my best scoring model (ResNet18) that gives 0.914 on the public leaderboard with the best-scoring public kernel available (0.927) and I got 0.925. After I preprocessed both predictions with <code>rankdata</code>, my score improved to 0.933.</p>",
      "rawMarkdown": "The target metric in this competition is based on ranks rather than on actual values. That means that as long as the order of your values is fixed, the metric will stay the same.\n\nTo illustrate:\n\n```\ntarget = [1, 0, 1, 1, 0]\npreds = [0.5, 0.25, 0.2, 0.3, 0.1]\n\nmetric = roc_auc_score(target, preds)  # 0.833\n\ntarget = [1, 0, 1, 1, 0]\npreds = [0.7, 0.15, 0.1, 0.2, 0.05]\n\nmetric = roc_auc_score(target, preds)  # 0.833\n```\n\nThat means that two different models that give the **same** score could actually output completely **different** values. They are not even required to be in (0, 1) range!\n\n```\ntarget = [1, 0, 1, 1, 0]\npreds = [100, 25, 20, 30, 10]\n\nmetric = roc_auc_score(target, preds)  # 0.833\n```\n\nThen, if you will try to average the predictions of two non-calibrated models, you might observe that the score is not necessarily getting better and in some cases, it could become even worse! This happens because the prediction scales of these two models are not directly comparable because of the aforementioned issue.\n\nHow this can be fixed?\n\nOne simple solution is to bring the predictions to the same scale, e.g. with `scipy.stats.rankdata` function. This will turn scores into ranks, i.e. `[0.7, 0.15, 0.1, 0.2, 0.05]` will be turned into `[5, 3, 2, 4, 1]`. After this, the predictions could be blended.\n\nNote that this not always lead to better results and is highly dependent on which exactly models are blended, how strong is the bias, and so on.\n\nTo illustrate, I decided to naively blend my best scoring model (ResNet18) that gives 0.914 on the public leaderboard with the best-scoring public kernel available (0.927) and I got 0.925. After I preprocessed both predictions with `rankdata`, my score improved to 0.933.",
      "votes": 55
    },
    {
      "id": 874997,
      "postDate": "2020-06-05T12:39:39.750Z",
      "content": "<p>This guy knows his way around data science.</p>",
      "rawMarkdown": "This guy knows his way around data science.",
      "votes": 2
    },
    {
      "id": 873684,
      "postDate": "2020-06-04T11:07:50.113Z",
      "content": "<p>ensemble by rankdata:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F263ca2f402f619f6855c71cfef4f962c%2FSelection_087.png?generation=1591268852952242&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "ensemble by rankdata:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F263ca2f402f619f6855c71cfef4f962c%2FSelection_087.png?generation=1591268852952242&amp;alt=media)\n",
      "votes": 1,
      "replies": [
        {
          "id": 873774,
          "postDate": "2020-06-04T12:19:59.930Z",
          "content": "<p>Heng - can you please be kind to explain above ?</p>",
          "rawMarkdown": "Heng - can you please be kind to explain above ?",
          "votes": 1
        },
        {
          "id": 873823,
          "postDate": "2020-06-04T13:06:15.967Z",
          "content": "<p>How did it affect the score?</p>",
          "rawMarkdown": "How did it affect the score?"
        },
        {
          "id": 873881,
          "postDate": "2020-06-04T13:50:18.263Z",
          "content": "<ol>\n<li><p>the ideal distribution will be 2 Laplace peak at 0 and 1. </p></li>\n<li><p>datarank() assign near \"integer\" values will may destroy some information. So it ,may sometimes work and sometimes does not.</p></li>\n<li><p>you can see from the distribution datarank(), there is still some error (min(pos,neg) of each bar, i.e. overlap).  AUC is roughly correlated to this overlap error.</p></li>\n<li><p>because we may not have enough +ve samples (during local CV or public LB) to have a \"smooth distribution curve\", improvement in AUC values itself may not be reliable. you should check to see if the \"distribution curve\" is improved or not after certain post processing.</p></li>\n</ol>",
          "rawMarkdown": "1. the ideal distribution will be 2 Laplace peak at 0 and 1. \n\n2. datarank() assign near \"integer\" values will may destroy some information. So it ,may sometimes work and sometimes does not.\n\n3. you can see from the distribution datarank(), there is still some error (min(pos,neg) of each bar, i.e. overlap).  AUC is roughly correlated to this overlap error.\n\n4. because we may not have enough +ve samples (during local CV or public LB) to have a \"smooth distribution curve\", improvement in AUC values itself may not be reliable. you should check to see if the \"distribution curve\" is improved or not after certain post processing.\n",
          "votes": 3
        },
        {
          "id": 873891,
          "postDate": "2020-06-04T13:53:53.553Z",
          "content": "<p>but i do agree that \"calibrated predictions\" should improve results. we also need to see how the models are off-calibrated. calibration may be a non-trivial process.  </p>",
          "rawMarkdown": "but i do agree that \"calibrated predictions\" should improve results. we also need to see how the models are off-calibrated. calibration may be a non-trivial process.  ",
          "votes": 2
        }
      ]
    },
    {
      "id": 899340,
      "postDate": "2020-06-24T07:11:54.453Z",
      "content": "<p>Just tried it out. Actually increased my score. I'd recommend this.</p>",
      "rawMarkdown": "Just tried it out. Actually increased my score. I'd recommend this."
    },
    {
      "id": 893966,
      "postDate": "2020-06-20T05:13:49.287Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 893960,
      "postDate": "2020-06-20T05:10:21.520Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 893953,
      "postDate": "2020-06-20T05:02:12.393Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 874997,
      "author_name": "Roman",
      "author_url": "",
      "post_date": "2020-06-05T12:39:39.750000",
      "content": "<p>This guy knows his way around data science.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 873684,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-06-04T11:07:50.113000",
      "content": "<p>ensemble by rankdata:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F263ca2f402f619f6855c71cfef4f962c%2FSelection_087.png?generation=1591268852952242&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 873774,
          "author_name": "Pulkit Mehta",
          "author_url": "",
          "post_date": "2020-06-04T12:19:59.930000",
          "content": "<p>Heng - can you please be kind to explain above ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 873823,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-04T13:06:15.967000",
          "content": "<p>How did it affect the score?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 873881,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-04T13:50:18.263000",
          "content": "<ol>\n<li><p>the ideal distribution will be 2 Laplace peak at 0 and 1. </p></li>\n<li><p>datarank() assign near \"integer\" values will may destroy some information. So it ,may sometimes work and sometimes does not.</p></li>\n<li><p>you can see from the distribution datarank(), there is still some error (min(pos,neg) of each bar, i.e. overlap).  AUC is roughly correlated to this overlap error.</p></li>\n<li><p>because we may not have enough +ve samples (during local CV or public LB) to have a \"smooth distribution curve\", improvement in AUC values itself may not be reliable. you should check to see if the \"distribution curve\" is improved or not after certain post processing.</p></li>\n</ol>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 873891,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-04T13:53:53.553000",
          "content": "<p>but i do agree that \"calibrated predictions\" should improve results. we also need to see how the models are off-calibrated. calibration may be a non-trivial process.  </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 899340,
      "author_name": "Nitesh Chaudhry",
      "author_url": "",
      "post_date": "2020-06-24T07:11:54.453000",
      "content": "<p>Just tried it out. Actually increased my score. I'd recommend this.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 893966,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-20T05:13:49.287000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 893960,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-20T05:10:21.520000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 893953,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-20T05:02:12.393000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "873610": "The target metric in this competition is based on ranks rather than on actual values. That means that as long as the order of your values is fixed, the metric will stay the same.\n\nTo illustrate:\n\n```\ntarget = [1, 0, 1, 1, 0]\npreds = [0.5, 0.25, 0.2, 0.3, 0.1]\n\nmetric = roc_auc_score(target, preds)  # 0.833\n\ntarget = [1, 0, 1, 1, 0]\npreds = [0.7, 0.15, 0.1, 0.2, 0.05]\n\nmetric = roc_auc_score(target, preds)  # 0.833\n```\n\nThat means that two different models that give the **same** score could actually output completely **different** values. They are not even required to be in (0, 1) range!\n\n```\ntarget = [1, 0, 1, 1, 0]\npreds = [100, 25, 20, 30, 10]\n\nmetric = roc_auc_score(target, preds)  # 0.833\n```\n\nThen, if you will try to average the predictions of two non-calibrated models, you might observe that the score is not necessarily getting better and in some cases, it could become even worse! This happens because the prediction scales of these two models are not directly comparable because of the aforementioned issue.\n\nHow this can be fixed?\n\nOne simple solution is to bring the predictions to the same scale, e.g. with `scipy.stats.rankdata` function. This will turn scores into ranks, i.e. `[0.7, 0.15, 0.1, 0.2, 0.05]` will be turned into `[5, 3, 2, 4, 1]`. After this, the predictions could be blended.\n\nNote that this not always lead to better results and is highly dependent on which exactly models are blended, how strong is the bias, and so on.\n\nTo illustrate, I decided to naively blend my best scoring model (ResNet18) that gives 0.914 on the public leaderboard with the best-scoring public kernel available (0.927) and I got 0.925. After I preprocessed both predictions with `rankdata`, my score improved to 0.933.",
    "874997": "This guy knows his way around data science.",
    "873684": "ensemble by rankdata:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F263ca2f402f619f6855c71cfef4f962c%2FSelection_087.png?generation=1591268852952242&amp;alt=media)\n",
    "899340": "Just tried it out. Actually increased my score. I'd recommend this.",
    "893966": "",
    "893960": "",
    "893953": ""
  }
}