{
  "id": 214676,
  "title": "Blending/ensembling/stacking methods for label-weighted label-ranking average precision",
  "url": "/competitions/rfcx-species-audio-detection/discussion/214676",
  "author_name": "",
  "post_date": "2021-01-27T10:46:33.627088900Z",
  "votes": 12,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I've seen a lot less about blending/ensembling/stacking methods for label-weighted label-ranking average precision than for other competition metrics. Have I overlooked some obvious things that get used a lot or that should obviously be helpful?</p>\n<p>Since the competition metric is rank based, I assume a lot of the approaches often used for other rank based metrics like <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221\" target=\"_blank\">AUC</a> may make sense here. Those would be things like</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205564\" target=\"_blank\">Rank averaging</a> as also mentioned in the <a href=\"https://mlwave.com/kaggle-ensembling-guide/\" target=\"_blank\">Kaggle ensembling guide</a></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\" target=\"_blank\">Power averaging</a> (including e.g. <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211194\" target=\"_blank\">square root averaging</a>)</li>\n<li>The above can of course be weighted averages with the weights optimized using out-of-fold predictions. Finding weights that optimize out-of-fold predictions in terms of the competition metric sounds quite doable using <code>scipy.optimize</code>). In this context, presumably one should see whether it is a good idea to have seperate weight schemes for the different species?</li>\n</ul>\n<p>As always stacking may also be interesting, but when one fits a second level model, it's also not immediately obvious than some particular approach would be especially obviously good. E.g. is using a classification objective with binary cross-entropy a good training objective for the second stage model (given that we cannot easily use the non-differentiable competition metric as a loss function)?</p>\n<p>In the most similar competition with a similar metric (<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019\" target=\"_blank\">Freesound Audio Tagging 2019</a>) that I could find, I saw the following in top solutions:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97815\" target=\"_blank\">Stacking with neural net using binary_crossentropy</a></li>\n<li><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924\" target=\"_blank\">Geometric mean blending</a></li>\n<li><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/96440\" target=\"_blank\">Averaging weights based on OOF predictions</a></li>\n<li><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97812\" target=\"_blank\">Simple average</a></li>\n</ul>",
  "messages": [
    {
      "id": "1172312",
      "postDate": "01/27/2021 10:46:33",
      "content": "<p>I've seen a lot less about blending/ensembling/stacking methods for label-weighted label-ranking average precision than for other competition metrics. Have I overlooked some obvious things that get used a lot or that should obviously be helpful?</p>\n<p>Since the competition metric is rank based, I assume a lot of the approaches often used for other rank based metrics like <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221\" target=\"_blank\">AUC</a> may make sense here. Those would be things like</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205564\" target=\"_blank\">Rank averaging</a> as also mentioned in the <a href=\"https://mlwave.com/kaggle-ensembling-guide/\" target=\"_blank\">Kaggle ensembling guide</a></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\" target=\"_blank\">Power averaging</a> (including e.g. <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211194\" target=\"_blank\">square root averaging</a>)</li>\n<li>The above can of course be weighted averages with the weights optimized using out-of-fold predictions. Finding weights that optimize out-of-fold predictions in terms of the competition metric sounds quite doable using <code>scipy.optimize</code>). In this context, presumably one should see whether it is a good idea to have seperate weight schemes for the different species?</li>\n</ul>\n<p>As always stacking may also be interesting, but when one fits a second level model, it's also not immediately obvious than some particular approach would be especially obviously good. E.g. is using a classification objective with binary cross-entropy a good training objective for the second stage model (given that we cannot easily use the non-differentiable competition metric as a loss function)?</p>\n<p>In the most similar competition with a similar metric (<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019\" target=\"_blank\">Freesound Audio Tagging 2019</a>) that I could find, I saw the following in top solutions:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97815\" target=\"_blank\">Stacking with neural net using binary_crossentropy</a></li>\n<li><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924\" target=\"_blank\">Geometric mean blending</a></li>\n<li><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/96440\" target=\"_blank\">Averaging weights based on OOF predictions</a></li>\n<li><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97812\" target=\"_blank\">Simple average</a></li>\n</ul>",
      "rawMarkdown": "I've seen a lot less about blending/ensembling/stacking methods for label-weighted label-ranking average precision than for other competition metrics. Have I overlooked some obvious things that get used a lot or that should obviously be helpful?\n\nSince the competition metric is rank based, I assume a lot of the approaches often used for other rank based metrics like [AUC](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221) may make sense here. Those would be things like\n* [Rank averaging](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205564) as also mentioned in the [Kaggle ensembling guide](https://mlwave.com/kaggle-ensembling-guide/)\n* [Power averaging](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653) (including e.g. [square root averaging](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211194))\n* The above can of course be weighted averages with the weights optimized using out-of-fold predictions. Finding weights that optimize out-of-fold predictions in terms of the competition metric sounds quite doable using `scipy.optimize`). In this context, presumably one should see whether it is a good idea to have seperate weight schemes for the different species?\n\nAs always stacking may also be interesting, but when one fits a second level model, it's also not immediately obvious than some particular approach would be especially obviously good. E.g. is using a classification objective with binary cross-entropy a good training objective for the second stage model (given that we cannot easily use the non-differentiable competition metric as a loss function)?\n\nIn the most similar competition with a similar metric ([Freesound Audio Tagging 2019](https://www.kaggle.com/c/freesound-audio-tagging-2019)) that I could find, I saw the following in top solutions:\n* [Stacking with neural net using binary_crossentropy](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97815)\n* [Geometric mean blending](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924)\n* [Averaging weights based on OOF predictions](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/96440)\n* [Simple average](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97812)",
      "votes": null
    },
    {
      "id": "1172439",
      "postDate": "01/27/2021 11:40:03",
      "content": "<p>Simple average &gt;&gt;</p>\n<p>A lot of times, the simpler the better.</p>",
      "rawMarkdown": "Simple average >>\n\nA lot of times, the simpler the better.",
      "votes": null
    },
    {
      "id": "1172756",
      "postDate": "01/27/2021 14:38:25",
      "content": "<p>So far I just average predictions.</p>",
      "rawMarkdown": "So far I just average predictions.",
      "votes": null
    },
    {
      "id": "1172917",
      "postDate": "01/27/2021 15:43:42",
      "content": "<p>you can also combine them. It's risky but hey, you never know.</p>\n<p>Once I won a competition with the arithmetic mean for all categories but one where I used geometric. It boosted my CV and LB by 0.00002 enough to jump to the 1 place :) </p>",
      "rawMarkdown": "you can also combine them. It's risky but hey, you never know.\n\nOnce I won a competition with the arithmetic mean for all categories but one where I used geometric. It boosted my CV and LB by 0.00002 enough to jump to the 1 place :)",
      "votes": null
    },
    {
      "id": "1173641",
      "postDate": "01/28/2021 03:05:47",
      "content": "<p>I tried Stacking. But it didn't work. Maybe there is too little training data(about 1200 annotations) to use stacking.<br>\nSo I used <strong>Simple average</strong>. It is better performance and simple.</p>",
      "rawMarkdown": "I tried Stacking. But it didn't work. Maybe there is too little training data(about 1200 annotations) to use stacking.\nSo I used **Simple average**. It is better performance and simple.",
      "votes": null
    },
    {
      "id": "1173953",
      "postDate": "01/28/2021 07:43:22",
      "content": "<p>I just used Simple Averaging.</p>",
      "rawMarkdown": "I just used Simple Averaging.",
      "votes": null
    },
    {
      "id": "1189363",
      "postDate": "02/06/2021 23:48:20",
      "content": "<p>I am using simple average as well. In the cornell birdcall competition 2nd place solution was using square, average, and taking the square root again according to solution repo on github <a href=\"https://github.com/vlomme/Birdcall-Identification-competition\" target=\"_blank\">here</a></p>\n<p>It didnt help me but might be worth to try</p>",
      "rawMarkdown": "I am using simple average as well. In the cornell birdcall competition 2nd place solution was using square, average, and taking the square root again according to solution repo on github [here](https://github.com/vlomme/Birdcall-Identification-competition)\n\nIt didnt help me but might be worth to try",
      "votes": null
    },
    {
      "id": "1199672",
      "postDate": "02/14/2021 03:25:07",
      "content": "<p>Understand the metric. Check the score (and data) distribution. Check your error rate.<br>\nIt is a ranking metric so, \"move your prediction rank position\" carefully.<br>\nOne can definitely do better than simple average. … maybe much better </p>",
      "rawMarkdown": "Understand the metric. Check the score (and data) distribution. Check your error rate.\nIt is a ranking metric so, \"move your prediction rank position\" carefully.\nOne can definitely do better than simple average. ... maybe much better",
      "votes": null
    },
    {
      "id": "1200313",
      "postDate": "02/14/2021 15:21:48",
      "content": "<p>I agree with that, i.e. cheating with a single specie can bring LB0.01 boost.</p>",
      "rawMarkdown": "I agree with that, i.e. cheating with a single specie can bring LB0.01 boost.",
      "votes": null
    },
    {
      "id": "1200429",
      "postDate": "02/14/2021 16:47:08",
      "content": "<p>0.03 for me</p>",
      "rawMarkdown": "0.03 for me",
      "votes": null
    },
    {
      "id": "1200434",
      "postDate": "02/14/2021 16:51:02",
      "content": "<p>this is called \"manual data de-biasing\". </p>",
      "rawMarkdown": "this is called \"manual data de-biasing\".",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1172439,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "01/27/2021 11:40:03",
      "content": "<p>Simple average &gt;&gt;</p>\n<p>A lot of times, the simpler the better.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1172756,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "01/27/2021 14:38:25",
      "content": "<p>So far I just average predictions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1172917,
      "author_name": "valanm",
      "author_url": "",
      "post_date": "01/27/2021 15:43:42",
      "content": "<p>you can also combine them. It's risky but hey, you never know.</p>\n<p>Once I won a competition with the arithmetic mean for all categories but one where I used geometric. It boosted my CV and LB by 0.00002 enough to jump to the 1 place :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1173641,
      "author_name": "shinmurashinmura",
      "author_url": "",
      "post_date": "01/28/2021 03:05:47",
      "content": "<p>I tried Stacking. But it didn't work. Maybe there is too little training data(about 1200 annotations) to use stacking.<br>\nSo I used <strong>Simple average</strong>. It is better performance and simple.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1173953,
      "author_name": "kbh0287",
      "author_url": "",
      "post_date": "01/28/2021 07:43:22",
      "content": "<p>I just used Simple Averaging.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1189363,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/06/2021 23:48:20",
      "content": "<p>I am using simple average as well. In the cornell birdcall competition 2nd place solution was using square, average, and taking the square root again according to solution repo on github <a href=\"https://github.com/vlomme/Birdcall-Identification-competition\" target=\"_blank\">here</a></p>\n<p>It didnt help me but might be worth to try</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1199672,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/14/2021 03:25:07",
      "content": "<p>Understand the metric. Check the score (and data) distribution. Check your error rate.<br>\nIt is a ranking metric so, \"move your prediction rank position\" carefully.<br>\nOne can definitely do better than simple average. … maybe much better </p>",
      "votes": null,
      "replies": [
        {
          "id": 1200313,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "02/14/2021 15:21:48",
          "content": "<p>I agree with that, i.e. cheating with a single specie can bring LB0.01 boost.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1200429,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "02/14/2021 16:47:08",
          "content": "<p>0.03 for me</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1200434,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/14/2021 16:51:02",
          "content": "<p>this is called \"manual data de-biasing\". </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1172312": "I've seen a lot less about blending/ensembling/stacking methods for label-weighted label-ranking average precision than for other competition metrics. Have I overlooked some obvious things that get used a lot or that should obviously be helpful?\n\nSince the competition metric is rank based, I assume a lot of the approaches often used for other rank based metrics like [AUC](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221) may make sense here. Those would be things like\n* [Rank averaging](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205564) as also mentioned in the [Kaggle ensembling guide](https://mlwave.com/kaggle-ensembling-guide/)\n* [Power averaging](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653) (including e.g. [square root averaging](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211194))\n* The above can of course be weighted averages with the weights optimized using out-of-fold predictions. Finding weights that optimize out-of-fold predictions in terms of the competition metric sounds quite doable using `scipy.optimize`). In this context, presumably one should see whether it is a good idea to have seperate weight schemes for the different species?\n\nAs always stacking may also be interesting, but when one fits a second level model, it's also not immediately obvious than some particular approach would be especially obviously good. E.g. is using a classification objective with binary cross-entropy a good training objective for the second stage model (given that we cannot easily use the non-differentiable competition metric as a loss function)?\n\nIn the most similar competition with a similar metric ([Freesound Audio Tagging 2019](https://www.kaggle.com/c/freesound-audio-tagging-2019)) that I could find, I saw the following in top solutions:\n* [Stacking with neural net using binary_crossentropy](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97815)\n* [Geometric mean blending](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924)\n* [Averaging weights based on OOF predictions](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/96440)\n* [Simple average](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97812)",
    "1172439": "Simple average >>\n\nA lot of times, the simpler the better.",
    "1172756": "So far I just average predictions.",
    "1172917": "you can also combine them. It's risky but hey, you never know.\n\nOnce I won a competition with the arithmetic mean for all categories but one where I used geometric. It boosted my CV and LB by 0.00002 enough to jump to the 1 place :)",
    "1173641": "I tried Stacking. But it didn't work. Maybe there is too little training data(about 1200 annotations) to use stacking.\nSo I used **Simple average**. It is better performance and simple.",
    "1173953": "I just used Simple Averaging.",
    "1189363": "I am using simple average as well. In the cornell birdcall competition 2nd place solution was using square, average, and taking the square root again according to solution repo on github [here](https://github.com/vlomme/Birdcall-Identification-competition)\n\nIt didnt help me but might be worth to try",
    "1199672": "Understand the metric. Check the score (and data) distribution. Check your error rate.\nIt is a ranking metric so, \"move your prediction rank position\" carefully.\nOne can definitely do better than simple average. ... maybe much better",
    "1200313": "I agree with that, i.e. cheating with a single specie can bring LB0.01 boost.",
    "1200429": "0.03 for me",
    "1200434": "this is called \"manual data de-biasing\"."
  },
  "source": "meta"
}