{
  "id": 205564,
  "title": "tips: rank averaging",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/205564",
  "author_name": "Tawara",
  "post_date": "2020-12-20T16:39:43.897000",
  "votes": 32,
  "comment_count": 12,
  "views": 0,
  "content": "<p>AUC considers only the order of predicted values. <br>\nFor example:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2F29aa6bdc4e9325b8711d866fda1b774d%2Froc_auc_example.png?generation=1608481018781222&amp;alt=media\"><br>\n<code>y_1</code> and <code>y_2</code> have different scale values but get the same AUC score for label <code>t</code>. </p>\n<p>So, we can average predicted <strong>ranks</strong> by models  instead of values  for ensemble.<br>\nThis method improves my LB by 0.001. (I used it for 5-fold averaging.)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fe4b59294d315b225ec6d45b33049e6b8%2Frank_averaging.png?generation=1608481765883075&amp;alt=media\"></p>\n<p>The improvement is only slight. I think it may be more effective when averaging <strong>different models</strong>.</p>",
  "messages": [
    {
      "id": 1120193,
      "postDate": "2020-12-20T16:39:43.897Z",
      "content": "<p>AUC considers only the order of predicted values. <br>\nFor example:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2F29aa6bdc4e9325b8711d866fda1b774d%2Froc_auc_example.png?generation=1608481018781222&amp;alt=media\"><br>\n<code>y_1</code> and <code>y_2</code> have different scale values but get the same AUC score for label <code>t</code>. </p>\n<p>So, we can average predicted <strong>ranks</strong> by models  instead of values  for ensemble.<br>\nThis method improves my LB by 0.001. (I used it for 5-fold averaging.)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fe4b59294d315b225ec6d45b33049e6b8%2Frank_averaging.png?generation=1608481765883075&amp;alt=media\"></p>\n<p>The improvement is only slight. I think it may be more effective when averaging <strong>different models</strong>.</p>",
      "rawMarkdown": "AUC considers only the order of predicted values. \nFor example:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2F29aa6bdc4e9325b8711d866fda1b774d%2Froc_auc_example.png?generation=1608481018781222&alt=media\" width=85%>\n`y_1` and `y_2` have different scale values but get the same AUC score for label `t`. \n\nSo, we can average predicted **ranks** by models  instead of values  for ensemble.\nThis method improves my LB by 0.001. (I used it for 5-fold averaging.)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fe4b59294d315b225ec6d45b33049e6b8%2Frank_averaging.png?generation=1608481765883075&alt=media\" width=80%>\n\nThe improvement is only slight. I think it may be more effective when averaging **different models**.",
      "votes": 32
    },
    {
      "id": 1198924,
      "postDate": "2021-02-13T12:14:24.533Z",
      "content": "<p>Hi! Nice notebook! I have one question. What do you mean \"The improvement is only slight. I think it may be more effective when averaging different models.\"? </p>",
      "rawMarkdown": "Hi! Nice notebook! I have one question. What do you mean \"The improvement is only slight. I think it may be more effective when averaging different models.\"? "
    },
    {
      "id": 1165942,
      "postDate": "2021-01-23T10:28:44.613Z",
      "content": "<p>There is a simpler way to do this. To handle the scale issue while averaging the predictions, <br>\nyou can apply Sigmoid activation on top of the model(last layer). Then all the predictions will be in the [0, 1]. <br>\nSince Sigmoid is a monotonic a function, it will map the every final score to a unique number between 0 to 1. </p>",
      "rawMarkdown": "There is a simpler way to do this. To handle the scale issue while averaging the predictions, \nyou can apply Sigmoid activation on top of the model(last layer). Then all the predictions will be in the [0, 1]. \nSince Sigmoid is a monotonic a function, it will map the every final score to a unique number between 0 to 1. "
    },
    {
      "id": 1162320,
      "postDate": "2021-01-21T05:21:19.790Z",
      "content": "<p>Hi,I just want to konw how to oprate with RankAveraging?Is there any examples?</p>",
      "rawMarkdown": "Hi,I just want to konw how to oprate with RankAveraging?Is there any examples?",
      "replies": [
        {
          "id": 1167190,
          "postDate": "2021-01-24T05:56:51.933Z",
          "content": "<p>Here I show a toy example:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fbf87ce8e8fe596d7366dfeb40aa6581f%2Fexample_rank_avg_1.png?generation=1611467345394928&amp;alt=media\" alt=\"\"><br>\nWe have two prediction arrays <code>y_0</code> and <code>y_1</code> (shape: (5, 3)) for 3 classes of 5 examples.<br>\n<code>y_avg</code> is output of simple averaging of  <code>y_0</code> and <code>y_1</code>.</p>\n<p>For rank averaging, we have to convert predicted values into ranks by <strong>classes</strong>(because ROC-AUC is calculated by classes).<br>\nAn easy way to do this is applying <code>argsort</code> <strong>twice</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fff1a4b39537bc1bcc41810cdece686c6%2Fexample_rank_avg_2.png?generation=1611467743292191&amp;alt=media\" alt=\"\"><br>\n<code>y_rank_avg</code> is output of simple averaging of  <code>y_rank_0</code> and <code>y_rank_1</code>.</p>\n<p><br><br>\nThis way have some problems when there are duplicated values.<br>\nyou can use <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.rankdata.html\" target=\"_blank\"><code>scipy.stats.rankdata</code></a> or <a href=\"https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.rank.html\" target=\"_blank\"><code>pandas.DataFrame.rank</code></a>.</p>",
          "rawMarkdown": "Here I show a toy example:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fbf87ce8e8fe596d7366dfeb40aa6581f%2Fexample_rank_avg_1.png?generation=1611467345394928&alt=media)\nWe have two prediction arrays `y_0` and `y_1` (shape: (5, 3)) for 3 classes of 5 examples.\n`y_avg` is output of simple averaging of  `y_0` and `y_1`.\n\nFor rank averaging, we have to convert predicted values into ranks by **classes**(because ROC-AUC is calculated by classes).\nAn easy way to do this is applying `argsort` **twice**.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fff1a4b39537bc1bcc41810cdece686c6%2Fexample_rank_avg_2.png?generation=1611467743292191&alt=media)\n`y_rank_avg` is output of simple averaging of  `y_rank_0` and `y_rank_1`.\n\n<br>\nThis way have some problems when there are duplicated values.\nyou can use [`scipy.stats.rankdata`](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.rankdata.html) or [`pandas.DataFrame.rank`](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.rank.html).",
          "votes": 1
        },
        {
          "id": 1167297,
          "postDate": "2021-01-24T06:55:09.197Z",
          "content": "<p>Hi,why should we sort the data twice?Thanks!</p>",
          "rawMarkdown": "Hi,why should we sort the data twice?Thanks!"
        },
        {
          "id": 1167481,
          "postDate": "2021-01-24T09:51:20Z",
          "content": "<p>To be very accurate, The best way is scipy.rankdata(method = average), and the worst way is scipy.rankdata(method = dense) when there are duplicated values. Of course np.argsort is good, but logically scipy.rankdata(method = average) should be better. (small difference, though.)</p>",
          "rawMarkdown": "To be very accurate, The best way is scipy.rankdata(method = average), and the worst way is scipy.rankdata(method = dense) when there are duplicated values. Of course np.argsort is good, but logically scipy.rankdata(method = average) should be better. (small difference, though.)",
          "votes": 2
        },
        {
          "id": 1167493,
          "postDate": "2021-01-24T10:03:27.500Z",
          "content": "<p>Hi，I test this method,but I find my result is not normal,the predict result value like '60.X',it is so large.I use the np.argsort twice,and then use np.mean to average the different folds result.I don't if I have some error? Thanks! <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> </p>",
          "rawMarkdown": "Hi，I test this method,but I find my result is not normal,the predict result value like '60.X',it is so large.I use the np.argsort twice,and then use np.mean to average the different folds result.I don't if I have some error? Thanks! @mamasinkgs @ttahara "
        },
        {
          "id": 1198917,
          "postDate": "2021-02-13T11:56:03.557Z",
          "content": "<p><a href=\"https://www.kaggle.com/bcwang\" target=\"_blank\">@bcwang</a> <br>\nNo? this is not a mistake. It should be so. Try to submit your result.</p>",
          "rawMarkdown": "@bcwang \nNo? this is not a mistake. It should be so. Try to submit your result."
        }
      ]
    },
    {
      "id": 1161052,
      "postDate": "2021-01-20T09:43:05.773Z",
      "content": "<p>Great idea,I have a question,what do you mean the 'ranks'?Thanks!</p>",
      "rawMarkdown": "Great idea,I have a question,what do you mean the 'ranks'?Thanks!",
      "replies": [
        {
          "id": 1163944,
          "postDate": "2021-01-22T04:22:04.147Z",
          "content": "<p>Here I use the word <code>rank</code> as <strong>the ascending order</strong> of values.</p>\n<p>For example:</p>\n<pre><code>y = [0.25, 0.95, 0.6, 0.97, 0.86]\n</code></pre>\n<p>The ascending order of <code>y</code> is as follows:</p>\n<pre><code>y_order = [0, 3, 1, 4, 2]\n</code></pre>",
          "rawMarkdown": "Here I use the word `rank` as **the ascending order** of values.\n\nFor example:\n```python\ny = [0.25, 0.95, 0.6, 0.97, 0.86]\n```\nThe ascending order of `y` is as follows:\n\n```\ny_order = [0, 3, 1, 4, 2]\n```"
        },
        {
          "id": 1164142,
          "postDate": "2021-01-22T07:25:47.607Z",
          "content": "<p>Thanks for you reply!you mean is we get the y_order=[0,3,1,4,2],and we average the y_order and get the average is y_average=[0.0, 0.3, 0.1, 0.4, 0.2]?Thanks!</p>",
          "rawMarkdown": "Thanks for you reply!you mean is we get the y_order=[0,3,1,4,2],and we average the y_order and get the average is y_average=[0.0, 0.3, 0.1, 0.4, 0.2]?Thanks!"
        },
        {
          "id": 1183507,
          "postDate": "2021-02-03T03:57:08.807Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1198924,
      "author_name": "Anton Makarenko",
      "author_url": "",
      "post_date": "2021-02-13T12:14:24.533000",
      "content": "<p>Hi! Nice notebook! I have one question. What do you mean \"The improvement is only slight. I think it may be more effective when averaging different models.\"? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1165942,
      "author_name": "Aravind P",
      "author_url": "",
      "post_date": "2021-01-23T10:28:44.613000",
      "content": "<p>There is a simpler way to do this. To handle the scale issue while averaging the predictions, <br>\nyou can apply Sigmoid activation on top of the model(last layer). Then all the predictions will be in the [0, 1]. <br>\nSince Sigmoid is a monotonic a function, it will map the every final score to a unique number between 0 to 1. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1162320,
      "author_name": "MaYang",
      "author_url": "",
      "post_date": "2021-01-21T05:21:19.790000",
      "content": "<p>Hi,I just want to konw how to oprate with RankAveraging?Is there any examples?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1167190,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-01-24T05:56:51.933000",
          "content": "<p>Here I show a toy example:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fbf87ce8e8fe596d7366dfeb40aa6581f%2Fexample_rank_avg_1.png?generation=1611467345394928&amp;alt=media\" alt=\"\"><br>\nWe have two prediction arrays <code>y_0</code> and <code>y_1</code> (shape: (5, 3)) for 3 classes of 5 examples.<br>\n<code>y_avg</code> is output of simple averaging of  <code>y_0</code> and <code>y_1</code>.</p>\n<p>For rank averaging, we have to convert predicted values into ranks by <strong>classes</strong>(because ROC-AUC is calculated by classes).<br>\nAn easy way to do this is applying <code>argsort</code> <strong>twice</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fff1a4b39537bc1bcc41810cdece686c6%2Fexample_rank_avg_2.png?generation=1611467743292191&amp;alt=media\" alt=\"\"><br>\n<code>y_rank_avg</code> is output of simple averaging of  <code>y_rank_0</code> and <code>y_rank_1</code>.</p>\n<p><br><br>\nThis way have some problems when there are duplicated values.<br>\nyou can use <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.rankdata.html\" target=\"_blank\"><code>scipy.stats.rankdata</code></a> or <a href=\"https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.rank.html\" target=\"_blank\"><code>pandas.DataFrame.rank</code></a>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1167297,
          "author_name": "Bcw93",
          "author_url": "",
          "post_date": "2021-01-24T06:55:09.197000",
          "content": "<p>Hi,why should we sort the data twice?Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1167481,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-24T09:51:20",
          "content": "<p>To be very accurate, The best way is scipy.rankdata(method = average), and the worst way is scipy.rankdata(method = dense) when there are duplicated values. Of course np.argsort is good, but logically scipy.rankdata(method = average) should be better. (small difference, though.)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1167493,
          "author_name": "Bcw93",
          "author_url": "",
          "post_date": "2021-01-24T10:03:27.500000",
          "content": "<p>Hi，I test this method,but I find my result is not normal,the predict result value like '60.X',it is so large.I use the np.argsort twice,and then use np.mean to average the different folds result.I don't if I have some error? Thanks! <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1198917,
          "author_name": "Anton Makarenko",
          "author_url": "",
          "post_date": "2021-02-13T11:56:03.557000",
          "content": "<p><a href=\"https://www.kaggle.com/bcwang\" target=\"_blank\">@bcwang</a> <br>\nNo? this is not a mistake. It should be so. Try to submit your result.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1161052,
      "author_name": "Bcw93",
      "author_url": "",
      "post_date": "2021-01-20T09:43:05.773000",
      "content": "<p>Great idea,I have a question,what do you mean the 'ranks'?Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1163944,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-01-22T04:22:04.147000",
          "content": "<p>Here I use the word <code>rank</code> as <strong>the ascending order</strong> of values.</p>\n<p>For example:</p>\n<pre><code>y = [0.25, 0.95, 0.6, 0.97, 0.86]\n</code></pre>\n<p>The ascending order of <code>y</code> is as follows:</p>\n<pre><code>y_order = [0, 3, 1, 4, 2]\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1164142,
          "author_name": "Bcw93",
          "author_url": "",
          "post_date": "2021-01-22T07:25:47.607000",
          "content": "<p>Thanks for you reply!you mean is we get the y_order=[0,3,1,4,2],and we average the y_order and get the average is y_average=[0.0, 0.3, 0.1, 0.4, 0.2]?Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1183507,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-03T03:57:08.807000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1120193": "AUC considers only the order of predicted values. \nFor example:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2F29aa6bdc4e9325b8711d866fda1b774d%2Froc_auc_example.png?generation=1608481018781222&alt=media\" width=85%>\n`y_1` and `y_2` have different scale values but get the same AUC score for label `t`. \n\nSo, we can average predicted **ranks** by models  instead of values  for ensemble.\nThis method improves my LB by 0.001. (I used it for 5-fold averaging.)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fe4b59294d315b225ec6d45b33049e6b8%2Frank_averaging.png?generation=1608481765883075&alt=media\" width=80%>\n\nThe improvement is only slight. I think it may be more effective when averaging **different models**.",
    "1198924": "Hi! Nice notebook! I have one question. What do you mean \"The improvement is only slight. I think it may be more effective when averaging different models.\"? ",
    "1165942": "There is a simpler way to do this. To handle the scale issue while averaging the predictions, \nyou can apply Sigmoid activation on top of the model(last layer). Then all the predictions will be in the [0, 1]. \nSince Sigmoid is a monotonic a function, it will map the every final score to a unique number between 0 to 1. ",
    "1162320": "Hi,I just want to konw how to oprate with RankAveraging?Is there any examples?",
    "1161052": "Great idea,I have a question,what do you mean the 'ranks'?Thanks!"
  }
}