{
  "id": 220314,
  "title": "s3 class postprocessing (+0.008 on private LB)",
  "url": "/competitions/rfcx-species-audio-detection/writeups/random-prediction-s3-class-postprocessing-0-008-on",
  "author_name": "",
  "post_date": "2021-02-18T00:37:46.621325400Z",
  "votes": 15,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As it was mentioned by others, s3 was a very interesting class. Looking more into this, our team figured out that we should scale up predictions for s3 class. We checked a number of ways to do this, but eventually a simple scale of 1.5x worked best for us (our predictions were all positive numbers, that scaling probably wouldn't work if your predictions are in logits).</p>\n<p>This scaling lifted our final ensemble from 0.933 to 0.941 (+0.008) on private LB.</p>\n<p>The intuition to convince ourselves that this was not a small public LB fluke (because it similarly helped on public LB) was as follows:</p>\n<ul>\n<li>Each class roughly has similar number of TP labels </li>\n<li>But you can see on OOF train predictions and on test prediction that classes are very imbalanced (let's say looking at frequencies of different classes being top1 or top3 predictions). And s3 class is actually clearly the most ubiquitous. </li>\n<li>Now because of specific labeling of this competition, we are in a situation where s3 is very common class, and hence very often present next to other labels, and model gets 0 target for s3 class (if you did not do any corrections, like loss masking)</li>\n<li>So model learns to be extra cautious at predicting s3, when in fact it should be quite aggressive in predicting it</li>\n</ul>\n<p>Hence we applied post-processing here. <br>\nWe could not leverage this on other classes, but s3 really stood out as a clear outlier for us. </p>",
  "messages": [
    {
      "id": "1207645",
      "postDate": "02/18/2021 00:37:46",
      "content": "<p>As it was mentioned by others, s3 was a very interesting class. Looking more into this, our team figured out that we should scale up predictions for s3 class. We checked a number of ways to do this, but eventually a simple scale of 1.5x worked best for us (our predictions were all positive numbers, that scaling probably wouldn't work if your predictions are in logits).</p>\n<p>This scaling lifted our final ensemble from 0.933 to 0.941 (+0.008) on private LB.</p>\n<p>The intuition to convince ourselves that this was not a small public LB fluke (because it similarly helped on public LB) was as follows:</p>\n<ul>\n<li>Each class roughly has similar number of TP labels </li>\n<li>But you can see on OOF train predictions and on test prediction that classes are very imbalanced (let's say looking at frequencies of different classes being top1 or top3 predictions). And s3 class is actually clearly the most ubiquitous. </li>\n<li>Now because of specific labeling of this competition, we are in a situation where s3 is very common class, and hence very often present next to other labels, and model gets 0 target for s3 class (if you did not do any corrections, like loss masking)</li>\n<li>So model learns to be extra cautious at predicting s3, when in fact it should be quite aggressive in predicting it</li>\n</ul>\n<p>Hence we applied post-processing here. <br>\nWe could not leverage this on other classes, but s3 really stood out as a clear outlier for us. </p>",
      "rawMarkdown": "As it was mentioned by others, s3 was a very interesting class. Looking more into this, our team figured out that we should scale up predictions for s3 class. We checked a number of ways to do this, but eventually a simple scale of 1.5x worked best for us (our predictions were all positive numbers, that scaling probably wouldn't work if your predictions are in logits).\n\nThis scaling lifted our final ensemble from 0.933 to 0.941 (+0.008) on private LB.\n\nThe intuition to convince ourselves that this was not a small public LB fluke (because it similarly helped on public LB) was as follows:\n- Each class roughly has similar number of TP labels \n- But you can see on OOF train predictions and on test prediction that classes are very imbalanced (let's say looking at frequencies of different classes being top1 or top3 predictions). And s3 class is actually clearly the most ubiquitous. \n- Now because of specific labeling of this competition, we are in a situation where s3 is very common class, and hence very often present next to other labels, and model gets 0 target for s3 class (if you did not do any corrections, like loss masking)\n- So model learns to be extra cautious at predicting s3, when in fact it should be quite aggressive in predicting it\n\nHence we applied post-processing here. \nWe could not leverage this on other classes, but s3 really stood out as a clear outlier for us.",
      "votes": null
    },
    {
      "id": "1208332",
      "postDate": "02/18/2021 08:49:20",
      "content": "<p>I did the same thing.<br>\nIn my model, just scaling up the prediction for s3 improved it by 0.019 (0.898 -&gt; 0.917) at private LB.<br>\nAnd scaling down s21, which was the next least accurate, improved the score by 0.001 (0.917 -&gt; 0.918).<br>\nI did not explore further because I thought it was not essential.<br>\nI was guessing that only s3 had more training label errors (not confirmed), but reading your discussion helped me understand better. Thank you for sharing.</p>",
      "rawMarkdown": "I did the same thing.\nIn my model, just scaling up the prediction for s3 improved it by 0.019 (0.898 -> 0.917) at private LB.\nAnd scaling down s21, which was the next least accurate, improved the score by 0.001 (0.917 -> 0.918).\nI did not explore further because I thought it was not essential.\nI was guessing that only s3 had more training label errors (not confirmed), but reading your discussion helped me understand better. Thank you for sharing.",
      "votes": null
    },
    {
      "id": "1208518",
      "postDate": "02/18/2021 10:11:08",
      "content": "<p>Today I understand the importance of postprocessing. Thank you friend for this writeup.</p>",
      "rawMarkdown": "Today I understand the importance of postprocessing. Thank you friend for this writeup.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1208332,
      "author_name": "naoism",
      "author_url": "",
      "post_date": "02/18/2021 08:49:20",
      "content": "<p>I did the same thing.<br>\nIn my model, just scaling up the prediction for s3 improved it by 0.019 (0.898 -&gt; 0.917) at private LB.<br>\nAnd scaling down s21, which was the next least accurate, improved the score by 0.001 (0.917 -&gt; 0.918).<br>\nI did not explore further because I thought it was not essential.<br>\nI was guessing that only s3 had more training label errors (not confirmed), but reading your discussion helped me understand better. Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208518,
      "author_name": "rajkumarl",
      "author_url": "",
      "post_date": "02/18/2021 10:11:08",
      "content": "<p>Today I understand the importance of postprocessing. Thank you friend for this writeup.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207645": "As it was mentioned by others, s3 was a very interesting class. Looking more into this, our team figured out that we should scale up predictions for s3 class. We checked a number of ways to do this, but eventually a simple scale of 1.5x worked best for us (our predictions were all positive numbers, that scaling probably wouldn't work if your predictions are in logits).\n\nThis scaling lifted our final ensemble from 0.933 to 0.941 (+0.008) on private LB.\n\nThe intuition to convince ourselves that this was not a small public LB fluke (because it similarly helped on public LB) was as follows:\n- Each class roughly has similar number of TP labels \n- But you can see on OOF train predictions and on test prediction that classes are very imbalanced (let's say looking at frequencies of different classes being top1 or top3 predictions). And s3 class is actually clearly the most ubiquitous. \n- Now because of specific labeling of this competition, we are in a situation where s3 is very common class, and hence very often present next to other labels, and model gets 0 target for s3 class (if you did not do any corrections, like loss masking)\n- So model learns to be extra cautious at predicting s3, when in fact it should be quite aggressive in predicting it\n\nHence we applied post-processing here. \nWe could not leverage this on other classes, but s3 really stood out as a clear outlier for us.",
    "1208332": "I did the same thing.\nIn my model, just scaling up the prediction for s3 improved it by 0.019 (0.898 -> 0.917) at private LB.\nAnd scaling down s21, which was the next least accurate, improved the score by 0.001 (0.917 -> 0.918).\nI did not explore further because I thought it was not essential.\nI was guessing that only s3 had more training label errors (not confirmed), but reading your discussion helped me understand better. Thank you for sharing.",
    "1208518": "Today I understand the importance of postprocessing. Thank you friend for this writeup."
  },
  "source": "meta"
}