{
  "id": 246783,
  "title": "Optimizing Metric Directly (AUC/mAP/Ranking)",
  "url": "/competitions/siim-covid19-detection/discussion/246783",
  "author_name": "Kerem Turgutlu",
  "post_date": "2021-06-17T02:18:47.969000",
  "votes": 24,
  "comment_count": 11,
  "views": 0,
  "content": "<p>mAP metric for classification of the 4 classes we have is actually a ranking problem, if you can rank the positive and negative cases for a given class correctly your AP would be 1. This also explains why many people use AUC as a proxy for training the classification problem. I was just wondering around the SETI competition discussions to see whether to enter that one as well and I realized both competitions are actually about ranking. Here is a nice post shared in those dicussions: <a href=\"https://towardsdatascience.com/explicit-auc-maximization-70beef6db14e\" target=\"_blank\">https://towardsdatascience.com/explicit-auc-maximization-70beef6db14e</a>. I wish I could remember the person who shared to also give them credit.</p>\n<p>Here is PyTorch implementation for ELO loss mentioned in that blog post, it will directly optimize the ranking. My intuition is that larger batch sizes will better as we are computing pairwise difference fo pos and neg. It didn't give me a boost in terms of model performance but I can tell that it provides huge stability in training.</p>\n<pre><code>def elo_loss(logits, y, reduction='mean'):\n    \"https://towardsdatascience.com/explicit-auc-maximization-70beef6db14e\"\n    # reduction mean, sum ...\n    losses = [] \n    class_ids = y.unique()    \n    for i in class_ids:\n        class_logits = logits[:,i.item()]\n        class_targs = (y == i).float()\n\n        mask = (class_targs.unsqueeze(1)*(1-class_targs.unsqueeze(0))).bool()\n        class_loss = -torch.sigmoid(class_logits.unsqueeze(1) - class_logits.unsqueeze(0))[mask].mean()\n        losses.append(class_loss)\n\n    loss = torch.stack(losses)\n    if reduction == 'mean': return torch.mean(loss)\n    if reduction == 'sum':  return torch.sum(loss)\n</code></pre>\n<p>Let me know if you find it helpful, but you don't have to 😃</p>\n<p><a href=\"https://www.kaggle.com/hengck\" target=\"_blank\">@hengck</a> also mentions a better loss function in this <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240233\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/240233</a>, not sure if this would be it.</p>",
  "messages": [
    {
      "id": 1353240,
      "postDate": "2021-06-17T02:18:47.970Z",
      "content": "<p>mAP metric for classification of the 4 classes we have is actually a ranking problem, if you can rank the positive and negative cases for a given class correctly your AP would be 1. This also explains why many people use AUC as a proxy for training the classification problem. I was just wondering around the SETI competition discussions to see whether to enter that one as well and I realized both competitions are actually about ranking. Here is a nice post shared in those dicussions: <a href=\"https://towardsdatascience.com/explicit-auc-maximization-70beef6db14e\" target=\"_blank\">https://towardsdatascience.com/explicit-auc-maximization-70beef6db14e</a>. I wish I could remember the person who shared to also give them credit.</p>\n<p>Here is PyTorch implementation for ELO loss mentioned in that blog post, it will directly optimize the ranking. My intuition is that larger batch sizes will better as we are computing pairwise difference fo pos and neg. It didn't give me a boost in terms of model performance but I can tell that it provides huge stability in training.</p>\n<pre><code>def elo_loss(logits, y, reduction='mean'):\n    \"https://towardsdatascience.com/explicit-auc-maximization-70beef6db14e\"\n    # reduction mean, sum ...\n    losses = [] \n    class_ids = y.unique()    \n    for i in class_ids:\n        class_logits = logits[:,i.item()]\n        class_targs = (y == i).float()\n\n        mask = (class_targs.unsqueeze(1)*(1-class_targs.unsqueeze(0))).bool()\n        class_loss = -torch.sigmoid(class_logits.unsqueeze(1) - class_logits.unsqueeze(0))[mask].mean()\n        losses.append(class_loss)\n\n    loss = torch.stack(losses)\n    if reduction == 'mean': return torch.mean(loss)\n    if reduction == 'sum':  return torch.sum(loss)\n</code></pre>\n<p>Let me know if you find it helpful, but you don't have to 😃</p>\n<p><a href=\"https://www.kaggle.com/hengck\" target=\"_blank\">@hengck</a> also mentions a better loss function in this <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240233\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/240233</a>, not sure if this would be it.</p>",
      "rawMarkdown": "mAP metric for classification of the 4 classes we have is actually a ranking problem, if you can rank the positive and negative cases for a given class correctly your AP would be 1. This also explains why many people use AUC as a proxy for training the classification problem. I was just wondering around the SETI competition discussions to see whether to enter that one as well and I realized both competitions are actually about ranking. Here is a nice post shared in those dicussions: https://towardsdatascience.com/explicit-auc-maximization-70beef6db14e. I wish I could remember the person who shared to also give them credit.\n\nHere is PyTorch implementation for ELO loss mentioned in that blog post, it will directly optimize the ranking. My intuition is that larger batch sizes will better as we are computing pairwise difference fo pos and neg. It didn't give me a boost in terms of model performance but I can tell that it provides huge stability in training.\n\n```\ndef elo_loss(logits, y, reduction='mean'):\n    \"https://towardsdatascience.com/explicit-auc-maximization-70beef6db14e\"\n    # reduction mean, sum ...\n    losses = [] \n    class_ids = y.unique()    \n    for i in class_ids:\n        class_logits = logits[:,i.item()]\n        class_targs = (y == i).float()\n\n        mask = (class_targs.unsqueeze(1)*(1-class_targs.unsqueeze(0))).bool()\n        class_loss = -torch.sigmoid(class_logits.unsqueeze(1) - class_logits.unsqueeze(0))[mask].mean()\n        losses.append(class_loss)\n        \n    loss = torch.stack(losses)\n    if reduction == 'mean': return torch.mean(loss)\n    if reduction == 'sum':  return torch.sum(loss)\n```\n\nLet me know if you find it helpful, but you don't have to 😃\n\n\n@hengck also mentions a better loss function in this https://www.kaggle.com/c/siim-covid19-detection/discussion/240233, not sure if this would be it.\n\n ",
      "votes": 23
    },
    {
      "id": 1356844,
      "postDate": "2021-06-19T08:41:31.973Z",
      "content": "<p><a href=\"https://github.com/yzhuoning/LibAUC/blob/main/examples/03_Optimizing_AUPRC_with_ResNet18_on_Imbalanced_CIFAR10.ipynb\" target=\"_blank\">https://github.com/yzhuoning/LibAUC/blob/main/examples/03_Optimizing_AUPRC_with_ResNet18_on_Imbalanced_CIFAR10.ipynb</a><br>\n<a href=\"https://libauc.org/\" target=\"_blank\">https://libauc.org/</a></p>",
      "rawMarkdown": "https://github.com/yzhuoning/LibAUC/blob/main/examples/03_Optimizing_AUPRC_with_ResNet18_on_Imbalanced_CIFAR10.ipynb\nhttps://libauc.org/",
      "votes": 3
    },
    {
      "id": 1353302,
      "postDate": "2021-06-17T03:30:34.973Z",
      "content": "<p>the key to winning is this one:</p>\n<ul>\n<li>BIMCV-COVID19 Data </li>\n<li>MIDRC-RICORD Data </li>\n</ul>\n<p>previous x-ray chest images competition has a large shakeup.</p>\n<p>you can google for paper to do directly AUC/mAP/Ranking using BOTH labeled and unlabeled set.</p>\n<p>but this competition is easier:</p>\n<h2>only one box class + five image level class for MAP</h2>",
      "rawMarkdown": "the key to winning is this one:\n\n-  BIMCV-COVID19 Data \n-  MIDRC-RICORD Data \n\nprevious x-ray chest images competition has a large shakeup.\n\nyou can google for paper to do directly AUC/mAP/Ranking using BOTH labeled and unlabeled set.\n\nbut this competition is easier:\nonly one box class + five image level class for MAP\n---\n",
      "votes": 2,
      "replies": [
        {
          "id": 1353334,
          "postDate": "2021-06-17T04:06:02.023Z",
          "content": "<p>Where is BIMCV-COVID19 Data available at? Yes, since metric cares about ranking any learning to rank  methods would also be fine here. </p>",
          "rawMarkdown": "Where is BIMCV-COVID19 Data available at? Yes, since metric cares about ranking any learning to rank  methods would also be fine here. "
        },
        {
          "id": 1353402,
          "postDate": "2021-06-17T04:57:43.373Z",
          "content": "<p><img src=\"https://i.ibb.co/BfLpSfD/Selection-283.png\" alt=\"\"></p>",
          "rawMarkdown": "![](https://i.ibb.co/BfLpSfD/Selection-283.png)",
          "votes": 1
        },
        {
          "id": 1353444,
          "postDate": "2021-06-17T05:37:12.970Z",
          "content": "<p><img src=\"https://i.ibb.co/sgKhJwF/Selection-289.png\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/keremt\" target=\"_blank\">@keremt</a> if you do download the external data, can you check if the above idea is feasible?</p>\n<ul>\n<li>there are indeed kaggle data found in  BIMCV-COVID19 set</li>\n<li>the kaggle  has indeed relabel them</li>\n<li>there exist indeed aux rich BIMCV-COVID19 labels that we may be able to use</li>\n</ul>\n<p>if it is feasible, maybe i can share starter kit on it</p>",
          "rawMarkdown": "![](https://i.ibb.co/sgKhJwF/Selection-289.png)\n\n@keremt if you do download the external data, can you check if the above idea is feasible?\n- there are indeed kaggle data found in  BIMCV-COVID19 set\n- the kaggle  has indeed relabel them\n- there exist indeed aux rich BIMCV-COVID19 labels that we may be able to use\n\nif it is feasible, maybe i can share starter kit on it",
          "votes": 1
        },
        {
          "id": 1353809,
          "postDate": "2021-06-17T08:52:14.857Z",
          "content": "<p>Hello,what is the link to that?</p>",
          "rawMarkdown": "Hello,what is the link to that?"
        },
        {
          "id": 1355061,
          "postDate": "2021-06-18T04:36:16.483Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> thanks for sharing it. I just downloaded the 2. iteration BIMCV (which includes both 1+2 iterations). I will check the data soon.  YOLOR is also something I glimpsed at the beginning of the competition. I also guess we are now free to use all external data (except for MIMIC-CXR) according to <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240080#1354578\" target=\"_blank\">dicussion</a>.</p>\n<p>FYI. BIMCV data is approx ~400GB.</p>",
          "rawMarkdown": "@hengck23 thanks for sharing it. I just downloaded the 2. iteration BIMCV (which includes both 1+2 iterations). I will check the data soon.  YOLOR is also something I glimpsed at the beginning of the competition. I also guess we are now free to use all external data (except for MIMIC-CXR) according to [dicussion](https://www.kaggle.com/c/siim-covid19-detection/discussion/240080#1354578).\n\nFYI. BIMCV data is approx ~400GB.",
          "votes": -1
        }
      ]
    },
    {
      "id": 1353246,
      "postDate": "2021-06-17T02:32:43.613Z",
      "content": "<p>Hello,Is only for classificatioon？</p>",
      "rawMarkdown": "Hello,Is only for classificatioon？",
      "replies": [
        {
          "id": 1353254,
          "postDate": "2021-06-17T02:36:45.883Z",
          "content": "<p>Yes, this is assuming you are training a classifier with n outputs, which outputs logits. In my case I used it for classifying <code>negative, typical, atypical, indeterminate</code>. logits is <code>bs x n_class</code> and y is a 1-d array of class ids, e.g. <code>[0,1,2,1,0,3]</code>.</p>",
          "rawMarkdown": "Yes, this is assuming you are training a classifier with n outputs, which outputs logits. In my case I used it for classifying `negative, typical, atypical, indeterminate`. logits is `bs x n_class` and y is a 1-d array of class ids, e.g. `[0,1,2,1,0,3]`."
        }
      ]
    },
    {
      "id": 1353479,
      "postDate": "2021-06-17T06:10:28.747Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1355117,
          "postDate": "2021-06-18T05:29:33.613Z",
          "content": "<p>If you mean metric here is what I use for classification mAP calculation: </p>\n<pre><code>def sklearn_mean_ap(preds, targs):\n    \"\"\"\n    Difference from COCO is precision is not interpolated\n    \"\"\"\n    return np.mean([average_precision_score(targs==i,preds[:,i]) for i in range(4)])*2/3\n</code></pre>\n<p>2/3 is there because I train classification and detection models separately.</p>",
          "rawMarkdown": "If you mean metric here is what I use for classification mAP calculation: \n```\ndef sklearn_mean_ap(preds, targs):\n    \"\"\"\n    Difference from COCO is precision is not interpolated\n    \"\"\"\n    return np.mean([average_precision_score(targs==i,preds[:,i]) for i in range(4)])*2/3\n```\n2/3 is there because I train classification and detection models separately.",
          "votes": 4
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1356844,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-06-19T08:41:31.973000",
      "content": "<p><a href=\"https://github.com/yzhuoning/LibAUC/blob/main/examples/03_Optimizing_AUPRC_with_ResNet18_on_Imbalanced_CIFAR10.ipynb\" target=\"_blank\">https://github.com/yzhuoning/LibAUC/blob/main/examples/03_Optimizing_AUPRC_with_ResNet18_on_Imbalanced_CIFAR10.ipynb</a><br>\n<a href=\"https://libauc.org/\" target=\"_blank\">https://libauc.org/</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1353302,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-06-17T03:30:34.973000",
      "content": "<p>the key to winning is this one:</p>\n<ul>\n<li>BIMCV-COVID19 Data </li>\n<li>MIDRC-RICORD Data </li>\n</ul>\n<p>previous x-ray chest images competition has a large shakeup.</p>\n<p>you can google for paper to do directly AUC/mAP/Ranking using BOTH labeled and unlabeled set.</p>\n<p>but this competition is easier:</p>\n<h2>only one box class + five image level class for MAP</h2>",
      "votes": 2,
      "replies": [
        {
          "id": 1353334,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2021-06-17T04:06:02.023000",
          "content": "<p>Where is BIMCV-COVID19 Data available at? Yes, since metric cares about ranking any learning to rank  methods would also be fine here. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1353402,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-06-17T04:57:43.373000",
          "content": "<p><img src=\"https://i.ibb.co/BfLpSfD/Selection-283.png\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1353444,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-06-17T05:37:12.970000",
          "content": "<p><img src=\"https://i.ibb.co/sgKhJwF/Selection-289.png\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/keremt\" target=\"_blank\">@keremt</a> if you do download the external data, can you check if the above idea is feasible?</p>\n<ul>\n<li>there are indeed kaggle data found in  BIMCV-COVID19 set</li>\n<li>the kaggle  has indeed relabel them</li>\n<li>there exist indeed aux rich BIMCV-COVID19 labels that we may be able to use</li>\n</ul>\n<p>if it is feasible, maybe i can share starter kit on it</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1353809,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-06-17T08:52:14.857000",
          "content": "<p>Hello,what is the link to that?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1355061,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2021-06-18T04:36:16.483000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> thanks for sharing it. I just downloaded the 2. iteration BIMCV (which includes both 1+2 iterations). I will check the data soon.  YOLOR is also something I glimpsed at the beginning of the competition. I also guess we are now free to use all external data (except for MIMIC-CXR) according to <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240080#1354578\" target=\"_blank\">dicussion</a>.</p>\n<p>FYI. BIMCV data is approx ~400GB.</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 1353246,
      "author_name": "Zekun",
      "author_url": "",
      "post_date": "2021-06-17T02:32:43.613000",
      "content": "<p>Hello,Is only for classificatioon？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1353254,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2021-06-17T02:36:45.883000",
          "content": "<p>Yes, this is assuming you are training a classifier with n outputs, which outputs logits. In my case I used it for classifying <code>negative, typical, atypical, indeterminate</code>. logits is <code>bs x n_class</code> and y is a 1-d array of class ids, e.g. <code>[0,1,2,1,0,3]</code>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1353479,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-17T06:10:28.747000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1355117,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2021-06-18T05:29:33.613000",
          "content": "<p>If you mean metric here is what I use for classification mAP calculation: </p>\n<pre><code>def sklearn_mean_ap(preds, targs):\n    \"\"\"\n    Difference from COCO is precision is not interpolated\n    \"\"\"\n    return np.mean([average_precision_score(targs==i,preds[:,i]) for i in range(4)])*2/3\n</code></pre>\n<p>2/3 is there because I train classification and detection models separately.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1353240": "mAP metric for classification of the 4 classes we have is actually a ranking problem, if you can rank the positive and negative cases for a given class correctly your AP would be 1. This also explains why many people use AUC as a proxy for training the classification problem. I was just wondering around the SETI competition discussions to see whether to enter that one as well and I realized both competitions are actually about ranking. Here is a nice post shared in those dicussions: https://towardsdatascience.com/explicit-auc-maximization-70beef6db14e. I wish I could remember the person who shared to also give them credit.\n\nHere is PyTorch implementation for ELO loss mentioned in that blog post, it will directly optimize the ranking. My intuition is that larger batch sizes will better as we are computing pairwise difference fo pos and neg. It didn't give me a boost in terms of model performance but I can tell that it provides huge stability in training.\n\n```\ndef elo_loss(logits, y, reduction='mean'):\n    \"https://towardsdatascience.com/explicit-auc-maximization-70beef6db14e\"\n    # reduction mean, sum ...\n    losses = [] \n    class_ids = y.unique()    \n    for i in class_ids:\n        class_logits = logits[:,i.item()]\n        class_targs = (y == i).float()\n\n        mask = (class_targs.unsqueeze(1)*(1-class_targs.unsqueeze(0))).bool()\n        class_loss = -torch.sigmoid(class_logits.unsqueeze(1) - class_logits.unsqueeze(0))[mask].mean()\n        losses.append(class_loss)\n        \n    loss = torch.stack(losses)\n    if reduction == 'mean': return torch.mean(loss)\n    if reduction == 'sum':  return torch.sum(loss)\n```\n\nLet me know if you find it helpful, but you don't have to 😃\n\n\n@hengck also mentions a better loss function in this https://www.kaggle.com/c/siim-covid19-detection/discussion/240233, not sure if this would be it.\n\n ",
    "1356844": "https://github.com/yzhuoning/LibAUC/blob/main/examples/03_Optimizing_AUPRC_with_ResNet18_on_Imbalanced_CIFAR10.ipynb\nhttps://libauc.org/",
    "1353302": "the key to winning is this one:\n\n-  BIMCV-COVID19 Data \n-  MIDRC-RICORD Data \n\nprevious x-ray chest images competition has a large shakeup.\n\nyou can google for paper to do directly AUC/mAP/Ranking using BOTH labeled and unlabeled set.\n\nbut this competition is easier:\nonly one box class + five image level class for MAP\n---\n",
    "1353246": "Hello,Is only for classificatioon？",
    "1353479": ""
  }
}