{
  "id": 220389,
  "title": "9th Place - Post Process LB 0.926 to LB 0.963!",
  "url": "/competitions/rfcx-species-audio-detection/writeups/chris-deotte-9th-place-post-process-lb-0-926-to-lb",
  "author_name": "",
  "post_date": "2021-02-18T20:47:35.893Z",
  "votes": 142,
  "comment_count": 54,
  "views": 0,
  "content": "<p>Thanks Kaggle and RFCx for a fun competition. My final submission without post process achieves private LB 0.926 and with post process achieves <strong>private LB 0.963</strong>! That's +0.037 with post process!</p>\n<h1>How To Score LB 0.950+</h1>\n<p>The metric in this competition is different than other competitions. We are asked to provide <code>submission.csv</code> where each row is a test sample, and each column has a species prediction.</p>\n<p>In other competitions, the metric computes <strong>column wise AUC</strong>. In this competition, the metric is essentially <strong>row wise AUC</strong>. Therefore we need each column to represent probabilities. The columns of common species need to be large values and the columns of rare species need to be small values.</p>\n<h1>Train Distribution</h1>\n<p>The distribution of the train data has roughly 6x the number of false positives versus true positives for each species. If you train your model with that data, then when your model is unsure about a prediction, it will predict the mean that it observed in the train data which is 1/7. Therefore if it is unsure about species 3, it will predict 1/7 and if it is unsure about species 19, it will predict 1/7.</p>\n<p>This is a problem because species 3 appears in roughly 90% of test samples whereas species 19 appears in roughly 1% of test samples. Therefore when your model is unsure , it should predict 90% for species 3 and 1% for species 19.</p>\n<h1>Test Distribution - Post Process</h1>\n<p>In order to correct our model's predictions we scale the odds. Note that scaling odds doesn't affect predictions of 0 and 1. It only affects the unsure middle predictions. </p>\n<p>First we convert the <code>submission.csv</code> column of probabilities into odds with the formula<br>\n$$ \\text{odds} = \\frac{p}{1-p}$$<br>\nThen we scale the odds with<br>\n$$ \\text{new odds} = \\text{factor} * \\text{old odds}$$<br>\nAnd lastly we convert back to probabilities<br>\n$$ \\text{prob} = \\frac{\\text{new odds}}{1 + \\text{new odds}}$$</p>\n<h1>Sample Code</h1>\n<pre><code># CONVERT PROB TO ODDS, APPLY MULTIPLIER, CONVERT BACK TO PROB\ndef scale(probs, factor):\n    probs = probs.copy()\n    idx = np.where(probs!=1)[0]\n    odds = factor * probs[idx] / (1-probs[idx])\n    probs[idx] =  odds/(1+odds)\n    return probs\n\nfor k in range(24):\n    sub.iloc[:,1+k] = scale(sub.iloc[:,1+k].values, FACTORS[k])\n</code></pre>\n<h1>Increase LB by +0.040!</h1>\n<p>The only detail remaining is how to calculate the <code>FACTORS</code> above. There are at least 3 ways.</p>\n<ul>\n<li>Create a model with output layer sigmoid, not softmax. Train with BCE loss using both true and false positives. Predict the test data. Compute the mean of each column. Convert to odds and divide by the odds of training data.</li>\n<li>Probe the LB with <code>submission.csv</code> of all zeros and one column of ones. Then use math to compute <code>factor</code> for that species. UPDATE use random numbers less than 1 instead of 0s to avoid sorting uncertainty.</li>\n<li>Use the factors listed in RFCx's paper <a href=\"https://www.sciencedirect.com/science/article/pii/S1574954120300637\" target=\"_blank\">here</a>, Table 2 in Section 2.4</li>\n</ul>\n<p>Personally, i used the first option listed above. I didn't have 5 days to probe the LB 24 times for the second option. And I didn't find the paper until the last day of the competition for the third option.</p>\n<p>My single model scores private LB 0.921 without post process and scores private LB 0.958 with post process. Ensembling a variety of image sizes and backbones increased the LBs to 0.926 and LB 0.963 respectively.</p>\n<h1>Model Details</h1>\n<p>I converted each audio file into Mel Spectrogram with Librosa <code>feature.melspectrogram</code> and <code>power_to_db</code> using sampling rate <code>32_000</code>, <code>n_mels = 384</code>, <code>n_fft=2048</code>, <code>hop_length=512</code>, <code>win_length=2048</code>. This produced NumPy arrays of size <code>(384, 3751)</code>. I later normalized them with <code>img = (img+87)/128</code> and  trained with random crops of <code>384x384</code> which had frequency range 20Hz to 16_000Hz and time range 6.14 seconds. Each crop contained at least 75% of a true positive or false positive.</p>\n<p>I concatenated the true postive and false positive CSV files from Kaggle. My dataloader provided labels, masks, and one color images. The labels and masks were contained in the vector <code>y</code> which was 48 zeros where 2 were potentially altered. For species <code>k</code>, the <code>kth</code> element was 0 or 1 corresponding to false positive or true positive respectively. And the <code>k+24th</code> element was 1 indicating mask true to calculate loss for the <code>kth</code> species. I trained TF model with the following loss</p>\n<pre><code> def masked_loss(y_true, y_pred):\n\n     mask = y_true[:,24:]\n     y_true = y_true[:,:24]\n\n     y_pred = tf.convert_to_tensor(y_pred)\n     y_true = tf.cast(y_true, y_pred.dtype)\n     mask = tf.cast(mask, y_pred.dtype)\n     y_pred = tf.math.multiply(mask, y_pred, name=None)\n\n     return K.mean( K.binary_crossentropy(y_true, y_pred), axis=-1 )*24.0\n</code></pre>\n<p>I used EfficientNetB2 with <code>albu.CoarseDropout</code> and <code>albu.RandomBrightnessContrast</code>. The optimizer was Adam with learning rate <code>1e-3</code> and reduce on plateau <code>factor=0.3</code>, <code>patience=3</code>. The final layer of the model was <code>GlobalAveragePooling2D()</code> and <code>Dense(24, activation='sigmoid')</code>. I monitored <code>val_loss</code> to tune hyperameters.</p>\n<h1>Try PP on Your Sub</h1>\n<p>If you want to try post process on your <code>submission.csv</code> file, I posted a Kaggle notebook <a href=\"https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\" target=\"_blank\">here</a>. It uses the 3rd method above for computing <code>FACTORS</code>. It also has 3 <code>MODE</code> you can try to account for different ways that you may have used to train your models.</p>",
  "messages": [
    {
      "id": "1208127",
      "postDate": "02/18/2021 07:08:27",
      "content": "<p>Thanks Kaggle and RFCx for a fun competition. My final submission without post process achieves private LB 0.926 and with post process achieves <strong>private LB 0.963</strong>! That's +0.037 with post process!</p>\n<h1>How To Score LB 0.950+</h1>\n<p>The metric in this competition is different than other competitions. We are asked to provide <code>submission.csv</code> where each row is a test sample, and each column has a species prediction.</p>\n<p>In other competitions, the metric computes <strong>column wise AUC</strong>. In this competition, the metric is essentially <strong>row wise AUC</strong>. Therefore we need each column to represent probabilities. The columns of common species need to be large values and the columns of rare species need to be small values.</p>\n<h1>Train Distribution</h1>\n<p>The distribution of the train data has roughly 6x the number of false positives versus true positives for each species. If you train your model with that data, then when your model is unsure about a prediction, it will predict the mean that it observed in the train data which is 1/7. Therefore if it is unsure about species 3, it will predict 1/7 and if it is unsure about species 19, it will predict 1/7.</p>\n<p>This is a problem because species 3 appears in roughly 90% of test samples whereas species 19 appears in roughly 1% of test samples. Therefore when your model is unsure , it should predict 90% for species 3 and 1% for species 19.</p>\n<h1>Test Distribution - Post Process</h1>\n<p>In order to correct our model's predictions we scale the odds. Note that scaling odds doesn't affect predictions of 0 and 1. It only affects the unsure middle predictions. </p>\n<p>First we convert the <code>submission.csv</code> column of probabilities into odds with the formula<br>\n$$ \\text{odds} = \\frac{p}{1-p}$$<br>\nThen we scale the odds with<br>\n$$ \\text{new odds} = \\text{factor} * \\text{old odds}$$<br>\nAnd lastly we convert back to probabilities<br>\n$$ \\text{prob} = \\frac{\\text{new odds}}{1 + \\text{new odds}}$$</p>\n<h1>Sample Code</h1>\n<pre><code># CONVERT PROB TO ODDS, APPLY MULTIPLIER, CONVERT BACK TO PROB\ndef scale(probs, factor):\n    probs = probs.copy()\n    idx = np.where(probs!=1)[0]\n    odds = factor * probs[idx] / (1-probs[idx])\n    probs[idx] =  odds/(1+odds)\n    return probs\n\nfor k in range(24):\n    sub.iloc[:,1+k] = scale(sub.iloc[:,1+k].values, FACTORS[k])\n</code></pre>\n<h1>Increase LB by +0.040!</h1>\n<p>The only detail remaining is how to calculate the <code>FACTORS</code> above. There are at least 3 ways.</p>\n<ul>\n<li>Create a model with output layer sigmoid, not softmax. Train with BCE loss using both true and false positives. Predict the test data. Compute the mean of each column. Convert to odds and divide by the odds of training data.</li>\n<li>Probe the LB with <code>submission.csv</code> of all zeros and one column of ones. Then use math to compute <code>factor</code> for that species. UPDATE use random numbers less than 1 instead of 0s to avoid sorting uncertainty.</li>\n<li>Use the factors listed in RFCx's paper <a href=\"https://www.sciencedirect.com/science/article/pii/S1574954120300637\" target=\"_blank\">here</a>, Table 2 in Section 2.4</li>\n</ul>\n<p>Personally, i used the first option listed above. I didn't have 5 days to probe the LB 24 times for the second option. And I didn't find the paper until the last day of the competition for the third option.</p>\n<p>My single model scores private LB 0.921 without post process and scores private LB 0.958 with post process. Ensembling a variety of image sizes and backbones increased the LBs to 0.926 and LB 0.963 respectively.</p>\n<h1>Model Details</h1>\n<p>I converted each audio file into Mel Spectrogram with Librosa <code>feature.melspectrogram</code> and <code>power_to_db</code> using sampling rate <code>32_000</code>, <code>n_mels = 384</code>, <code>n_fft=2048</code>, <code>hop_length=512</code>, <code>win_length=2048</code>. This produced NumPy arrays of size <code>(384, 3751)</code>. I later normalized them with <code>img = (img+87)/128</code> and  trained with random crops of <code>384x384</code> which had frequency range 20Hz to 16_000Hz and time range 6.14 seconds. Each crop contained at least 75% of a true positive or false positive.</p>\n<p>I concatenated the true postive and false positive CSV files from Kaggle. My dataloader provided labels, masks, and one color images. The labels and masks were contained in the vector <code>y</code> which was 48 zeros where 2 were potentially altered. For species <code>k</code>, the <code>kth</code> element was 0 or 1 corresponding to false positive or true positive respectively. And the <code>k+24th</code> element was 1 indicating mask true to calculate loss for the <code>kth</code> species. I trained TF model with the following loss</p>\n<pre><code> def masked_loss(y_true, y_pred):\n\n     mask = y_true[:,24:]\n     y_true = y_true[:,:24]\n\n     y_pred = tf.convert_to_tensor(y_pred)\n     y_true = tf.cast(y_true, y_pred.dtype)\n     mask = tf.cast(mask, y_pred.dtype)\n     y_pred = tf.math.multiply(mask, y_pred, name=None)\n\n     return K.mean( K.binary_crossentropy(y_true, y_pred), axis=-1 )*24.0\n</code></pre>\n<p>I used EfficientNetB2 with <code>albu.CoarseDropout</code> and <code>albu.RandomBrightnessContrast</code>. The optimizer was Adam with learning rate <code>1e-3</code> and reduce on plateau <code>factor=0.3</code>, <code>patience=3</code>. The final layer of the model was <code>GlobalAveragePooling2D()</code> and <code>Dense(24, activation='sigmoid')</code>. I monitored <code>val_loss</code> to tune hyperameters.</p>\n<h1>Try PP on Your Sub</h1>\n<p>If you want to try post process on your <code>submission.csv</code> file, I posted a Kaggle notebook <a href=\"https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\" target=\"_blank\">here</a>. It uses the 3rd method above for computing <code>FACTORS</code>. It also has 3 <code>MODE</code> you can try to account for different ways that you may have used to train your models.</p>",
      "rawMarkdown": "Thanks Kaggle and RFCx for a fun competition. My final submission without post process achieves private LB 0.926 and with post process achieves **private LB 0.963**! That's +0.037 with post process!\n\n# How To Score LB 0.950+\nThe metric in this competition is different than other competitions. We are asked to provide `submission.csv` where each row is a test sample, and each column has a species prediction.\n\nIn other competitions, the metric computes **column wise AUC**. In this competition, the metric is essentially **row wise AUC**. Therefore we need each column to represent probabilities. The columns of common species need to be large values and the columns of rare species need to be small values.\n\n# Train Distribution\nThe distribution of the train data has roughly 6x the number of false positives versus true positives for each species. If you train your model with that data, then when your model is unsure about a prediction, it will predict the mean that it observed in the train data which is 1/7. Therefore if it is unsure about species 3, it will predict 1/7 and if it is unsure about species 19, it will predict 1/7.\n\nThis is a problem because species 3 appears in roughly 90% of test samples whereas species 19 appears in roughly 1% of test samples. Therefore when your model is unsure , it should predict 90% for species 3 and 1% for species 19.\n\n# Test Distribution - Post Process\nIn order to correct our model's predictions we scale the odds. Note that scaling odds doesn't affect predictions of 0 and 1. It only affects the unsure middle predictions. \n\nFirst we convert the `submission.csv` column of probabilities into odds with the formula\n$$ \\text{odds} = \\frac{p}{1-p}$$\nThen we scale the odds with\n$$ \\text{new odds} = \\text{factor} * \\text{old odds}$$\nAnd lastly we convert back to probabilities\n$$ \\text{prob} = \\frac{\\text{new odds}}{1 + \\text{new odds}}$$\n\n# Sample Code\n    # CONVERT PROB TO ODDS, APPLY MULTIPLIER, CONVERT BACK TO PROB\n    def scale(probs, factor):\n        probs = probs.copy()\n        idx = np.where(probs!=1)[0]\n        odds = factor * probs[idx] / (1-probs[idx])\n        probs[idx] =  odds/(1+odds)\n        return probs\n\n    for k in range(24):\n        sub.iloc[:,1+k] = scale(sub.iloc[:,1+k].values, FACTORS[k])\n\n# Increase LB by +0.040!\nThe only detail remaining is how to calculate the `FACTORS` above. There are at least 3 ways.\n* Create a model with output layer sigmoid, not softmax. Train with BCE loss using both true and false positives. Predict the test data. Compute the mean of each column. Convert to odds and divide by the odds of training data.\n* Probe the LB with `submission.csv` of all zeros and one column of ones. Then use math to compute `factor` for that species. UPDATE use random numbers less than 1 instead of 0s to avoid sorting uncertainty.\n* Use the factors listed in RFCx's paper [here][1], Table 2 in Section 2.4\n\nPersonally, i used the first option listed above. I didn't have 5 days to probe the LB 24 times for the second option. And I didn't find the paper until the last day of the competition for the third option.\n\nMy single model scores private LB 0.921 without post process and scores private LB 0.958 with post process. Ensembling a variety of image sizes and backbones increased the LBs to 0.926 and LB 0.963 respectively.\n\n# Model Details\nI converted each audio file into Mel Spectrogram with Librosa `feature.melspectrogram` and `power_to_db` using sampling rate `32_000`, `n_mels = 384`, `n_fft=2048`, `hop_length=512`, `win_length=2048`. This produced NumPy arrays of size `(384, 3751)`. I later normalized them with `img = (img+87)/128` and  trained with random crops of `384x384` which had frequency range 20Hz to 16_000Hz and time range 6.14 seconds. Each crop contained at least 75% of a true positive or false positive.\n\nI concatenated the true postive and false positive CSV files from Kaggle. My dataloader provided labels, masks, and one color images. The labels and masks were contained in the vector `y` which was 48 zeros where 2 were potentially altered. For species `k`, the `kth` element was 0 or 1 corresponding to false positive or true positive respectively. And the `k+24th` element was 1 indicating mask true to calculate loss for the `kth` species. I trained TF model with the following loss\n\n     def masked_loss(y_true, y_pred):\n    \n         mask = y_true[:,24:]\n         y_true = y_true[:,:24]\n      \n         y_pred = tf.convert_to_tensor(y_pred)\n         y_true = tf.cast(y_true, y_pred.dtype)\n         mask = tf.cast(mask, y_pred.dtype)\n         y_pred = tf.math.multiply(mask, y_pred, name=None)\n\n         return K.mean( K.binary_crossentropy(y_true, y_pred), axis=-1 )*24.0\n\nI used EfficientNetB2 with `albu.CoarseDropout` and `albu.RandomBrightnessContrast`. The optimizer was Adam with learning rate `1e-3` and reduce on plateau `factor=0.3`, `patience=3`. The final layer of the model was `GlobalAveragePooling2D()` and `Dense(24, activation='sigmoid')`. I monitored `val_loss` to tune hyperameters.\n\n# Try PP on Your Sub\nIf you want to try post process on your `submission.csv` file, I posted a Kaggle notebook [here][2]. It uses the 3rd method above for computing `FACTORS`. It also has 3 `MODE` you can try to account for different ways that you may have used to train your models.\n\n[1]: https://www.sciencedirect.com/science/article/pii/S1574954120300637\n[2]: https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970",
      "votes": null
    },
    {
      "id": "1208139",
      "postDate": "02/18/2021 07:19:02",
      "content": "<p>Amazing job in short time <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Your PP method is quite genius as always. I think doing the all-zeros probing is tricky as it depends on the sorting algorithm, better to set all random and the column to 1.</p>",
      "rawMarkdown": "Amazing job in short time @cdeotte. Your PP method is quite genius as always. I think doing the all-zeros probing is tricky as it depends on the sorting algorithm, better to set all random and the column to 1.",
      "votes": null
    },
    {
      "id": "1208150",
      "postDate": "02/18/2021 07:25:43",
      "content": "<p>Congratulations🎉🎉🎉<br>\ngreat job</p>",
      "rawMarkdown": "Congratulations🎉🎉🎉\ngreat job",
      "votes": null
    },
    {
      "id": "1208152",
      "postDate": "02/18/2021 07:27:58",
      "content": "<p>Thanks Psi. Congrats on another amazing 1st place finish. So impressive.</p>\n<p>Good point about the sorting algorithm. I probed one species and the result was a little weird, i think you just explained why. I will update my post with your idea.</p>",
      "rawMarkdown": "Thanks Psi. Congrats on another amazing 1st place finish. So impressive.\n\nGood point about the sorting algorithm. I probed one species and the result was a little weird, i think you just explained why. I will update my post with your idea.",
      "votes": null
    },
    {
      "id": "1208157",
      "postDate": "02/18/2021 07:30:13",
      "content": "<p>Congrats on solo gold medal, very strongly post processing method <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Congrats on solo gold medal, very strongly post processing method @cdeotte",
      "votes": null
    },
    {
      "id": "1208226",
      "postDate": "02/18/2021 08:00:20",
      "content": "<p>Your PP rocks, as often.  I only started looking at test prediction distributions last 2 hours of the competition and noticed a shift indeed.  But I couldn't devise your PP in time.</p>\n<p>I like to think in terms of logits (log of odds)  instead of odds.  Then your pp amounts to move the mean of test logits to the mean of train logits for each class.  Is that right?</p>",
      "rawMarkdown": "Your PP rocks, as often.  I only started looking at test prediction distributions last 2 hours of the competition and noticed a shift indeed.  But I couldn't devise your PP in time.\n\nI like to think in terms of logits (log of odds)  instead of odds.  Then your pp amounts to move the mean of test logits to the mean of train logits for each class.  Is that right?",
      "votes": null
    },
    {
      "id": "1208227",
      "postDate": "02/18/2021 08:00:47",
      "content": "<p>You seem to be the \"King\" of postprocessing. Just like Bengali, this is yet another impressive postprocessing from you. Congrats on that. </p>",
      "rawMarkdown": "You seem to be the \"King\" of postprocessing. Just like Bengali, this is yet another impressive postprocessing from you. Congrats on that.",
      "votes": null
    },
    {
      "id": "1208268",
      "postDate": "02/18/2021 08:22:20",
      "content": "<p>Thanks CPMP. Yes i think you can use logits. Then you \"move\" the mean with addition instead of multiplication as in <br>\n    $$\\text{new logits} = \\text{adjust} + \\text{old logits}$$<br>\nwhere <code>adjust = log(FACTOR)</code> in my formula.</p>\n<p>Note that i simplified my PP a bit. I actually did it repeatedly. I would start with my original submission.csv then calculate FACTOR, then apply it to scale the submission.csv. Then i would calculate new FACTOR from the modified submission.csv. Next, I would start with my orginal submission.csv again, then apply new FACTOR. Then calculate another FACTOR. I would do this repeatedly until the FACTOR did not change anymore. Then I would use that FACTOR to create my final submission from original submission.</p>\n<p><code>FACTOR = test odds / train odds</code> but this needs subtle adjustment too. Because we train crops with <code>odds = x</code>, but we must estimate full clip odds from crop odds. For example if we train crops with 50 / 50 true postive and false positive, that is <code>1:1</code> crop odds. But the odds of a species being present in the entire 60 seconds after taking max over sliding crops is more like <code>2:1</code> train odds because there is a higher probability of seeing a species in 30 crops versus 1 crop. So when computing <code>FACTOR</code>, we use the higher <code>2:1</code> for <code>train odds</code>.</p>",
      "rawMarkdown": "Thanks CPMP. Yes i think you can use logits. Then you \"move\" the mean with addition instead of multiplication as in \n    $$\\text{new logits} = \\text{adjust} + \\text{old logits}$$\nwhere `adjust = log(FACTOR)` in my formula.\n\nNote that i simplified my PP a bit. I actually did it repeatedly. I would start with my original submission.csv then calculate FACTOR, then apply it to scale the submission.csv. Then i would calculate new FACTOR from the modified submission.csv. Next, I would start with my orginal submission.csv again, then apply new FACTOR. Then calculate another FACTOR. I would do this repeatedly until the FACTOR did not change anymore. Then I would use that FACTOR to create my final submission from original submission.\n\n`FACTOR = test odds / train odds` but this needs subtle adjustment too. Because we train crops with `odds = x`, but we must estimate full clip odds from crop odds. For example if we train crops with 50 / 50 true postive and false positive, that is `1:1` crop odds. But the odds of a species being present in the entire 60 seconds after taking max over sliding crops is more like `2:1` train odds because there is a higher probability of seeing a species in 30 crops versus 1 crop. So when computing `FACTOR`, we use the higher `2:1` for `train odds`.",
      "votes": null
    },
    {
      "id": "1208352",
      "postDate": "02/18/2021 08:58:56",
      "content": "<p>I just post processed my final submission using the test distribution from Table 2 in the paper <a href=\"https://www.sciencedirect.com/science/article/pii/S1574954120300637\" target=\"_blank\">here</a>. It would have achieved private LB 0.970. So it looks like the paper's distribution is better than the one I calculate from my test predictions.</p>\n<p><img src=\"http://playagricola.com/Kaggle/paper_LB.png\" alt=\"image\"></p>",
      "rawMarkdown": "I just post processed my final submission using the test distribution from Table 2 in the paper [here][1]. It would have achieved private LB 0.970. So it looks like the paper's distribution is better than the one I calculate from my test predictions.\n\n![image](http://playagricola.com/Kaggle/paper_LB.png)\n\n[1]: https://www.sciencedirect.com/science/article/pii/S1574954120300637",
      "votes": null
    },
    {
      "id": "1208368",
      "postDate": "02/18/2021 09:09:59",
      "content": "<p>this means that the kaggle data is the same as the paper?<br>\ninteresting!<br>\nif so, I think the orangnizer has given us too little data! There are some many TP and FP in the paper.</p>\n<p>I am interested the results of using all TP and FP annotations in the paper. Has kagglers managed to achieve the same results using only 5% of available annotation?</p>\n<hr>\n<p>on another note, such post-processing are useful for distillation later. It forces the model to learn something new.</p>",
      "rawMarkdown": "this means that the kaggle data is the same as the paper?\ninteresting!\nif so, I think the orangnizer has given us too little data! There are some many TP and FP in the paper.\n\nI am interested the results of using all TP and FP annotations in the paper. Has kagglers managed to achieve the same results using only 5% of available annotation?\n\n---\non another note, such post-processing are useful for distillation later. It forces the model to learn something new.",
      "votes": null
    },
    {
      "id": "1208369",
      "postDate": "02/18/2021 09:12:08",
      "content": "<p>Brilliant Chris, as always.</p>\n<p>Congratz on the solo gold !</p>\n<p>Do you mind sharing your pp code ? I wanna see how well it works for us :) </p>",
      "rawMarkdown": "Brilliant Chris, as always.\n\nCongratz on the solo gold !\n\nDo you mind sharing your pp code ? I wanna see how well it works for us :)",
      "votes": null
    },
    {
      "id": "1208385",
      "postDate": "02/18/2021 09:25:17",
      "content": "<p>The paper says they stored all their audio at website ARBIMON <a href=\"https://arbimon.rfcx.org/\" target=\"_blank\">https://arbimon.rfcx.org/</a> . I didn't search nor use any external data in this comp. I wonder whether the paper's 512471 audio clips are publicly available on ARBIMON? That's a half million audio clips!</p>",
      "rawMarkdown": "The paper says they stored all their audio at website ARBIMON https://arbimon.rfcx.org/ . I didn't search nor use any external data in this comp. I wonder whether the paper's 512471 audio clips are publicly available on ARBIMON? That's a half million audio clips!",
      "votes": null
    },
    {
      "id": "1208386",
      "postDate": "02/18/2021 09:25:26",
      "content": "<p>The metric depends a lot on the scaling, so it is also difficult to compare with the results in the paper. So it seems from the paper that S3 is present in 90% of the samples, so you need to have it in top, otherwise the metric hurts a lot.</p>",
      "rawMarkdown": "The metric depends a lot on the scaling, so it is also difficult to compare with the results in the paper. So it seems from the paper that S3 is present in 90% of the samples, so you need to have it in top, otherwise the metric hurts a lot.",
      "votes": null
    },
    {
      "id": "1208433",
      "postDate": "02/18/2021 09:41:32",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for this explanation. I start understanding now that postprocessing is more important.</p>",
      "rawMarkdown": "Thanks @cdeotte for this explanation. I start understanding now that postprocessing is more important.",
      "votes": null
    },
    {
      "id": "1208523",
      "postDate": "02/18/2021 10:13:09",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Thanks Theo!</p>\n<p>Try the following. Make sure your <code>submission.csv</code> are probabilities between and including 0 and 1. Do not use logits. First try <code>MODE=1</code>, then if that isn't good, try <code>MODE=2</code>. If that isn't good try <code>MODE=3</code>. If that isn't good, try <code>MODE=1</code> with a different <code>FUDGE</code>. Perhaps try 0.5, 1, and 3.</p>\n<pre><code># USE MODE 1, 2, or 3\nMODE = 1\n\n# LOAD SUBMISSION\nimport pandas as pd, numpy as np\nFUDGE = 2.0\nFILE = 'submission.csv'\ndf = pd.read_csv(FILE)\nfor k in range(24):\n    df.iloc[:,1+k] -= df.iloc[:,1+k].min()\n    df.iloc[:,1+k] /= df.iloc[:,1+k].max()\n\n# CONVERT PROBS TO ODDS, APPLY MULTIPLIER, CONVERT BACK TO PROBS\ndef scale(probs, factor):\n    probs = probs.copy()\n    idx = np.where(probs!=1)[0]\n    odds = factor * probs[idx] / (1-probs[idx])\n    probs[idx] =  odds/(1+odds)\n    return probs\n\n# DIFFERENT DISTRIBUTIONS\nd1 = df.iloc[:,1:].mean().values\nd2 = np.array([113, 204, 44, 923, 53, 41, 3, 213, 44, 23, 26, 149, 255,  \n    14, 123, 222, 46, 6, 474, 4, 17, 18, 23, 72])/1000.\n\nfor k in range(24):\n    if MODE==1: d = FUDGE\n    if MODE==2: d = d1[k]/(1-d1[k])\n    if MODE==3: s = d2[k] / d1[k]\n    else: s = (d2[k]/(1-d2[k]))/d\n    df.iloc[:,k+1] = scale(df.iloc[:,k+1].values,s)\n\ndf.to_csv('submission_with_pp.csv',index=False)\n</code></pre>",
      "rawMarkdown": "theoviel Thanks Theo!\n\nTry the following. Make sure your `submission.csv` are probabilities between and including 0 and 1. Do not use logits. First try `MODE=1`, then if that isn't good, try `MODE=2`. If that isn't good try `MODE=3`. If that isn't good, try `MODE=1` with a different `FUDGE`. Perhaps try 0.5, 1, and 3.\n\n    # USE MODE 1, 2, or 3\n    MODE = 1\n\n    # LOAD SUBMISSION\n    import pandas as pd, numpy as np\n    FUDGE = 2.0\n    FILE = 'submission.csv'\n    df = pd.read_csv(FILE)\n    for k in range(24):\n        df.iloc[:,1+k] -= df.iloc[:,1+k].min()\n        df.iloc[:,1+k] /= df.iloc[:,1+k].max()\n\n    # CONVERT PROBS TO ODDS, APPLY MULTIPLIER, CONVERT BACK TO PROBS\n    def scale(probs, factor):\n        probs = probs.copy()\n        idx = np.where(probs!=1)[0]\n        odds = factor * probs[idx] / (1-probs[idx])\n        probs[idx] =  odds/(1+odds)\n        return probs\n\n    # DIFFERENT DISTRIBUTIONS\n    d1 = df.iloc[:,1:].mean().values\n    d2 = np.array([113, 204, 44, 923, 53, 41, 3, 213, 44, 23, 26, 149, 255,  \n        14, 123, 222, 46, 6, 474, 4, 17, 18, 23, 72])/1000.\n\n    for k in range(24):\n        if MODE==1: d = FUDGE\n        if MODE==2: d = d1[k]/(1-d1[k])\n        if MODE==3: s = d2[k] / d1[k]\n        else: s = (d2[k]/(1-d2[k]))/d\n        df.iloc[:,k+1] = scale(df.iloc[:,k+1].values,s)\n    \n    df.to_csv('submission_with_pp.csv',index=False)",
      "votes": null
    },
    {
      "id": "1208526",
      "postDate": "02/18/2021 10:15:45",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> CPMP, try the above PP and see if it helps your submission.csv file. Make sure your file contains probabilities, not logits. And follow the instructions above which suggests different modes and fudges in case it doesn't work.</p>",
      "rawMarkdown": "cpmpml CPMP, try the above PP and see if it helps your submission.csv file. Make sure your file contains probabilities, not logits. And follow the instructions above which suggests different modes and fudges in case it doesn't work.",
      "votes": null
    },
    {
      "id": "1208530",
      "postDate": "02/18/2021 10:19:06",
      "content": "<p>maybe some here:<br>\n<a href=\"https://arbimon.rfcx.org/project/perm-stations/audiodata/training-sets?set=2119\" target=\"_blank\">https://arbimon.rfcx.org/project/perm-stations/audiodata/training-sets?set=2119</a>?<br>\n<a href=\"https://arbimon.rfcx.org/project/perm-stations/audiodata/templates\" target=\"_blank\">https://arbimon.rfcx.org/project/perm-stations/audiodata/templates</a><br>\n<a href=\"https://arbimon.rfcx.org/project/rfcx-bird-frog/audiodata/templates\" target=\"_blank\">https://arbimon.rfcx.org/project/rfcx-bird-frog/audiodata/templates</a></p>\n<p><img src=\"https://i.ibb.co/C2vYzNK/Selection-051.png\" alt=\"\"></p>",
      "rawMarkdown": "maybe some here:\nhttps://arbimon.rfcx.org/project/perm-stations/audiodata/training-sets?set=2119?\nhttps://arbimon.rfcx.org/project/perm-stations/audiodata/templates\nhttps://arbimon.rfcx.org/project/rfcx-bird-frog/audiodata/templates\n\n![](https://i.ibb.co/C2vYzNK/Selection-051.png)",
      "votes": null
    },
    {
      "id": "1208553",
      "postDate": "02/18/2021 10:37:38",
      "content": "<p>For me it gives +0.03(0.927 - 0.959)</p>",
      "rawMarkdown": "For me it gives +0.03(0.927 - 0.959)",
      "votes": null
    },
    {
      "id": "1208561",
      "postDate": "02/18/2021 10:41:22",
      "content": "<p><a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> That's great Kramarenko. What <code>MODE</code> worked best for you?</p>",
      "rawMarkdown": "vlomme That's great Kramarenko. What `MODE` worked best for you?",
      "votes": null
    },
    {
      "id": "1208581",
      "postDate": "02/18/2021 10:47:43",
      "content": "<p>Mode 1.  Mode 2 gives -0.01</p>",
      "rawMarkdown": "Mode 1.  Mode 2 gives -0.01",
      "votes": null
    },
    {
      "id": "1208586",
      "postDate": "02/18/2021 10:50:21",
      "content": "<p>Works like a charm for me as well <br>\n950 -&gt; 978 <br>\n960 -&gt; 980<br>\n976 -&gt; 982</p>",
      "rawMarkdown": "Works like a charm for me as well \n950 -> 978 \n960 -> 980\n976 -> 982",
      "votes": null
    },
    {
      "id": "1208587",
      "postDate": "02/18/2021 10:51:01",
      "content": "<p>Thanks a lot ! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> :) </p>",
      "rawMarkdown": "Thanks a lot ! @cdeotte :)",
      "votes": null
    },
    {
      "id": "1208597",
      "postDate": "02/18/2021 10:58:53",
      "content": "<p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> Looks like you have very strong models, hope you will do a writeup.</p>",
      "rawMarkdown": "selimsef Looks like you have very strong models, hope you will do a writeup.",
      "votes": null
    },
    {
      "id": "1208636",
      "postDate": "02/18/2021 11:48:14",
      "content": "<p>Thanks Chris! <br>\nFor my ensemble, your pp is what I missed!<br>\nMode 3 <br>\n900 -&gt; 936 (SILVER ZONE)</p>",
      "rawMarkdown": "Thanks Chris! \nFor my ensemble, your pp is what I missed!\nMode 3 \n900 -> 936 (SILVER ZONE)",
      "votes": null
    },
    {
      "id": "1208649",
      "postDate": "02/18/2021 11:58:26",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! This is brilliant! My best submission score improved 0.940 -&gt; 0.950 with MODE=3.</p>",
      "rawMarkdown": "Thanks a lot @cdeotte! This is brilliant! My best submission score improved 0.940 -> 0.950 with MODE=3.",
      "votes": null
    },
    {
      "id": "1208681",
      "postDate": "02/18/2021 12:36:32",
      "content": "<p>I've just submitted our final selection with your postprocessing <code>mode 2</code>. Our score would have improved from 0.967 to 0.979 ( 2nd place ). One of our worst decision not to accept your merging proposal. Maybe next time. Anyway, it was a great competition, and congratulations for your solo gold.</p>\n<p><img src=\"https://i.ibb.co/98TWw0G/Untitled.png\" alt=\"untitled\"></p>",
      "rawMarkdown": "I've just submitted our final selection with your postprocessing `mode 2`. Our score would have improved from 0.967 to 0.979 ( 2nd place ). One of our worst decision not to accept your merging proposal. Maybe next time. Anyway, it was a great competition, and congratulations for your solo gold.\n\n![untitled](https://i.ibb.co/98TWw0G/Untitled.png)",
      "votes": null
    },
    {
      "id": "1208698",
      "postDate": "02/18/2021 12:52:59",
      "content": "<p>At least we don't need to retrain and reproduce the 80+ models in our final blend 😅</p>",
      "rawMarkdown": "At least we don't need to retrain and reproduce the 80+ models in our final blend 😅",
      "votes": null
    },
    {
      "id": "1208728",
      "postDate": "02/18/2021 13:09:08",
      "content": "<p>I wish it would be only 80 models :D</p>\n<p>I am surprised about private score being quite higher here on that sub.</p>",
      "rawMarkdown": "I wish it would be only 80 models :D\n\nI am surprised about private score being quite higher here on that sub.",
      "votes": null
    },
    {
      "id": "1208794",
      "postDate": "02/18/2021 13:45:46",
      "content": "<p>My submission would have reached 0.951 instead of 0.918 with mode 3.</p>\n<p>Ouch.</p>\n<p>So this is where all the magic was…</p>",
      "rawMarkdown": "My submission would have reached 0.951 instead of 0.918 with mode 3.\n\nOuch.\n\nSo this is where all the magic was...",
      "votes": null
    },
    {
      "id": "1209001",
      "postDate": "02/18/2021 16:02:35",
      "content": "<p>Unbelievable. Just like that, I have second place - masterful insights as usual, Chris!</p>",
      "rawMarkdown": "Unbelievable. Just like that, I have second place - masterful insights as usual, Chris!",
      "votes": null
    },
    {
      "id": "1209013",
      "postDate": "02/18/2021 16:19:04",
      "content": "<p>Unbelievable as usual ! </p>",
      "rawMarkdown": "Unbelievable as usual !",
      "votes": null
    },
    {
      "id": "1209038",
      "postDate": "02/18/2021 16:41:38",
      "content": "<p>Thanks for this info. Our team had a discussion when you moved up the leaderboard that there was probably some post-processing or probing opportunity because we knew that was in your bag of tricks from the previous other competitions. </p>\n<p>We did some slightly less confident rescaling. We saw in the paper the tp and fp and frequencies but I don't think any of us read too much into it. Ours was purely based on lb probing of s3 and s18</p>",
      "rawMarkdown": "Thanks for this info. Our team had a discussion when you moved up the leaderboard that there was probably some post-processing or probing opportunity because we knew that was in your bag of tricks from the previous other competitions. \n\nWe did some slightly less confident rescaling. We saw in the paper the tp and fp and frequencies but I don't think any of us read too much into it. Ours was purely based on lb probing of s3 and s18",
      "votes": null
    },
    {
      "id": "1209096",
      "postDate": "02/18/2021 17:21:32",
      "content": "<p>Scaling only s3 and s18 gives me almost the same result (0.949 vs 0.951 for all species scaling with baseline 0.918).</p>",
      "rawMarkdown": "Scaling only s3 and s18 gives me almost the same result (0.949 vs 0.951 for all species scaling with baseline 0.918).",
      "votes": null
    },
    {
      "id": "1209173",
      "postDate": "02/18/2021 18:28:28",
      "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> That is interesting. For me, i gain a lot of benefit by post processing all species. Perhaps you already downsample the other species in your training process. </p>\n<p>I train all species with 50% true positive and 50% false positive. So even my rare species get predicted with high probabilities and thus I need to PP them to make them smaller.</p>",
      "rawMarkdown": "fffrrt That is interesting. For me, i gain a lot of benefit by post processing all species. Perhaps you already downsample the other species in your training process. \n\nI train all species with 50% true positive and 50% false positive. So even my rare species get predicted with high probabilities and thus I need to PP them to make them smaller.",
      "votes": null
    },
    {
      "id": "1209205",
      "postDate": "02/18/2021 18:51:08",
      "content": "<p>Thank you and congratulations!</p>\n<p>Is it correct that this PP uses the fact that distribution on Private is similar to Public? I didn't check that, just see that the scores are pretty similar all the way for me…?</p>",
      "rawMarkdown": "Thank you and congratulations!\n\nIs it correct that this PP uses the fact that distribution on Private is similar to Public? I didn't check that, just see that the scores are pretty similar all the way for me...?",
      "votes": null
    },
    {
      "id": "1209228",
      "postDate": "02/18/2021 19:03:30",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I used false positives as an extra class in training, and just ignored it during prediction.</p>\n<p>This is how percentiles of logits on test data looks like pre-scaling - <br>\n<img src=\"https://i.imgur.com/Ryn4RhJ.png\" alt=\"https://i.imgur.com/Ryn4RhJ.png\"></p>\n<pre><code>submission[:,3] = submission[:,3] + 5\nsubmission[:,18] = submission[:,18] + 5\n</code></pre>\n<p>This two-liner would have made it a 0.949 one.</p>",
      "rawMarkdown": "cdeotte I used false positives as an extra class in training, and just ignored it during prediction.\n\nThis is how percentiles of logits on test data looks like pre-scaling - \n![https://i.imgur.com/Ryn4RhJ.png](https://i.imgur.com/Ryn4RhJ.png)\n\n```\nsubmission[:,3] = submission[:,3] + 5\nsubmission[:,18] = submission[:,18] + 5\n```\nThis two-liner would have made it a 0.949 one.",
      "votes": null
    },
    {
      "id": "1209260",
      "postDate": "02/18/2021 19:26:14",
      "content": "<p>Wow, simple adjustment for big gain. I see your predictions are logits, so adding 5 is like scaling, i.e. multiplying, the odds by <code>1.6 = log(5)</code>. </p>\n<p>This simple trick works in many comps. Whenever the metric is AUC comparing different targets, and this comp metric was AUC in disguise, then we can improve LB by making sure each group of predicted targets is properly calibrated against other groups. This same trick was used in Jigsaw Toxic Comp <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980\" target=\"_blank\">here</a>, described in section post process.</p>\n<p>And other metrics benefit from PP adjustments too. For example, Log Loss and MSE are sensitive to the mean as shown <a href=\"https://www.kaggle.com/cdeotte/moa-post-process-lb-1777\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Wow, simple adjustment for big gain. I see your predictions are logits, so adding 5 is like scaling, i.e. multiplying, the odds by `1.6 = log(5)`. \n\nThis simple trick works in many comps. Whenever the metric is AUC comparing different targets, and this comp metric was AUC in disguise, then we can improve LB by making sure each group of predicted targets is properly calibrated against other groups. This same trick was used in Jigsaw Toxic Comp [here][1], described in section post process.\n\nAnd other metrics benefit from PP adjustments too. For example, Log Loss and MSE are sensitive to the mean as shown [here][2]\n\n[1]: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980\n[2]: https://www.kaggle.com/cdeotte/moa-post-process-lb-1777",
      "votes": null
    },
    {
      "id": "1209292",
      "postDate": "02/18/2021 19:56:20",
      "content": "<p>This PP doesn't require that public and private are similar. This PP requires that we calculate <code>FACTORS</code>. When we calculate the <code>FACTORS</code> we need to make sure that they pertrain to private LB.</p>\n<p>I list 3 ways to calculate <code>FACTORS</code> above. If we compute <code>FACTORS</code> from probing public LB, that is risky. If we use the entire test data to compute <code>FACTORS</code> that is safest. The last method is using the <code>FACTORS</code> from the paper. That is risky too because we are not sure if they pertain to private.</p>",
      "rawMarkdown": "This PP doesn't require that public and private are similar. This PP requires that we calculate `FACTORS`. When we calculate the `FACTORS` we need to make sure that they pertrain to private LB.\n\nI list 3 ways to calculate `FACTORS` above. If we compute `FACTORS` from probing public LB, that is risky. If we use the entire test data to compute `FACTORS` that is safest. The last method is using the `FACTORS` from the paper. That is risky too because we are not sure if they pertain to private.",
      "votes": null
    },
    {
      "id": "1209338",
      "postDate": "02/18/2021 20:50:31",
      "content": "<p>I tried your pp on few subs, the best I selected gets to 0.9739 private with your pp.</p>\n<p>It means two things:</p>\n<ol>\n<li><p>my modeling is quite good, I didn't miss something key contrarily to what i thought.</p></li>\n<li><p>I need to pay attention to test prediction distribution.  It is not the first time that this makes a difference between my solution and top ones</p></li>\n</ol>\n<p>Thanks a lot for sharing.  FYI, MODE =1 is what worked best with the subs I tested.</p>",
      "rawMarkdown": "I tried your pp on few subs, the best I selected gets to 0.9739 private with your pp.\n\nIt means two things:\n\n1. my modeling is quite good, I didn't miss something key contrarily to what i thought.\n\n2. I need to pay attention to test prediction distribution.  It is not the first time that this makes a difference between my solution and top ones\n\nThanks a lot for sharing.  FYI, MODE =1 is what worked best with the subs I tested.",
      "votes": null
    },
    {
      "id": "1209377",
      "postDate": "02/18/2021 21:21:33",
      "content": "<p>I have used same strategy mode1 jumped from 0.909, 3 models ensemble, to 0.950 which was enough to get 12th place : gold? because a gold team was removed :)</p>",
      "rawMarkdown": "I have used same strategy mode1 jumped from 0.909, 3 models ensemble, to 0.950 which was enough to get 12th place : gold? because a gold team was removed :)",
      "votes": null
    },
    {
      "id": "1209383",
      "postDate": "02/18/2021 21:26:16",
      "content": "<p>Congratulations tugstugi. Great job achieving solo gold ! I saw you climbing quickly the last few days. Well done!</p>\n<p>It's interesting how PP increased your private LB score more than you public LB score. Do you have any idea why?</p>",
      "rawMarkdown": "Congratulations tugstugi. Great job achieving solo gold ! I saw you climbing quickly the last few days. Well done!\n\nIt's interesting how PP increased your private LB score more than you public LB score. Do you have any idea why?",
      "votes": null
    },
    {
      "id": "1209385",
      "postDate": "02/18/2021 21:27:08",
      "content": "<p>Wow this is amazing<br>\nMy private LB score is jumped to 0.94729 from 0.90990.</p>\n<p>Congrats on your solo gold. Your helps in the comment sections also helped me to improve my score as well. Thanks 🙏</p>",
      "rawMarkdown": "Wow this is amazing\nMy private LB score is jumped to 0.94729 from 0.90990.\n\nCongrats on your solo gold. Your helps in the comment sections also helped me to improve my score as well. Thanks 🙏",
      "votes": null
    },
    {
      "id": "1209395",
      "postDate": "02/18/2021 21:33:58",
      "content": "<p>no idea, probably luck.</p>",
      "rawMarkdown": "no idea, probably luck.",
      "votes": null
    },
    {
      "id": "1209463",
      "postDate": "02/18/2021 23:06:10",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! Great solution and nice PP trick.</p>",
      "rawMarkdown": "Congrats @cdeotte ! Great solution and nice PP trick.",
      "votes": null
    },
    {
      "id": "1210309",
      "postDate": "02/19/2021 10:42:19",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> great notebook. I was wondering what was the intuition behind the masked loss function you used, why wouldn't only using the true positives and training using simple binary cross entropy work?</p>",
      "rawMarkdown": "Hey @cdeotte great notebook. I was wondering what was the intuition behind the masked loss function you used, why wouldn't only using the true positives and training using simple binary cross entropy work?",
      "votes": null
    },
    {
      "id": "1210368",
      "postDate": "02/19/2021 11:34:10",
      "content": "<p>All top teams used masked loss. Why?  Because we are given TP and TN for species one at a time.  There is no point having a loss for unknown labels.</p>\n<p>It seems that many assumed that TP are positive only for the annotated class and that target for all other species should be 0.  This is a wrong assumption.</p>\n<p>The only known negative targets are the FPs.  They should have been called TN as they are the only true negatives we have access to,</p>",
      "rawMarkdown": "All top teams used masked loss. Why?  Because we are given TP and TN for species one at a time.  There is no point having a loss for unknown labels.\n\nIt seems that many assumed that TP are positive only for the annotated class and that target for all other species should be 0.  This is a wrong assumption.\n\nThe only known negative targets are the FPs.  They should have been called TN as they are the only true negatives we have access to,",
      "votes": null
    },
    {
      "id": "1210964",
      "postDate": "02/19/2021 20:47:42",
      "content": "<p>That's awesome <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> . You have a very accurate model ! </p>\n<p>When i use <code>MODE=1</code> with the code posted here including the <code>d2</code> vector from paper, my result is private LB 0.970. During comp i computed my own <code>d2</code> array and didn't use the paper and achieved private LB 0.963 so it seems the paper's test distribution is the correct one and better than i could calculate from my test predictions.</p>",
      "rawMarkdown": "That's awesome @cpmpml . You have a very accurate model ! \n\nWhen i use `MODE=1` with the code posted here including the `d2` vector from paper, my result is private LB 0.970. During comp i computed my own `d2` array and didn't use the paper and achieved private LB 0.963 so it seems the paper's test distribution is the correct one and better than i could calculate from my test predictions.",
      "votes": null
    },
    {
      "id": "1210966",
      "postDate": "02/19/2021 20:48:41",
      "content": "<p>Thanks Giba</p>",
      "rawMarkdown": "Thanks Giba",
      "votes": null
    },
    {
      "id": "1212057",
      "postDate": "02/20/2021 20:38:18",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>I would like to add an <strong>important point</strong> to this trick:</p>\n<p>Not all results are improved with this trick. For example, in <a href=\"https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\" target=\"_blank\">your notebook</a>, if only the results of 15 columns change with this trick, but the results of the other nine columns do not change at all, much better scores are obtained.</p>\n<p>To prove this, we wrote a notebook that you and all those interested can read at the following address.</p>\n<p><a href=\"https://www.kaggle.com/mehrankazeminia/lb-0-980-rainforest-comparative-method-part-a\" target=\"_blank\">https://www.kaggle.com/mehrankazeminia/lb-0-980-rainforest-comparative-method-part-a</a></p>",
      "rawMarkdown": "Thank you @cdeotte \n\nI would like to add an **important point** to this trick:\n\nNot all results are improved with this trick. For example, in [your notebook](https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970), if only the results of 15 columns change with this trick, but the results of the other nine columns do not change at all, much better scores are obtained.\n\nTo prove this, we wrote a notebook that you and all those interested can read at the following address.\n\nhttps://www.kaggle.com/mehrankazeminia/lb-0-980-rainforest-comparative-method-part-a",
      "votes": null
    },
    {
      "id": "1212083",
      "postDate": "02/20/2021 21:24:00",
      "content": "<p>Interesting. Nice analysis and discovery!</p>",
      "rawMarkdown": "Interesting. Nice analysis and discovery!",
      "votes": null
    },
    {
      "id": "1217321",
      "postDate": "02/25/2021 00:07:28",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nReally thank you for your nice sharing. I have one question though:</p>\n<pre><code>Probe the LB with submission.csv of all zeros and one column of ones. Then use math to compute factor for that species. UPDATE use random numbers less than 1 instead of 0s to avoid sorting uncertainty.\n</code></pre>\n<p>I thought probing LB only got the distribution for public LB ( from the public LB score ). So how can we compute the factor for the distribution which generalize both public and private LB well? And is that the reason why you trained a 2:1 model for ease of post-processing?</p>\n<p>Thank you so much!</p>",
      "rawMarkdown": "cdeotte \nReally thank you for your nice sharing. I have one question though:\n\n```\nProbe the LB with submission.csv of all zeros and one column of ones. Then use math to compute factor for that species. UPDATE use random numbers less than 1 instead of 0s to avoid sorting uncertainty.\n```\n\nI thought probing LB only got the distribution for public LB ( from the public LB score ). So how can we compute the factor for the distribution which generalize both public and private LB well? And is that the reason why you trained a 2:1 model for ease of post-processing?\n\nThank you so much!",
      "votes": null
    },
    {
      "id": "1217344",
      "postDate": "02/25/2021 01:07:05",
      "content": "<p>You are correct that probing public LB and using the result to post process for private LB is risky. However in many comps, public and private have a similar distribution and public is large enough to compute accurate enough means, so it works. In this comp, the 1st place team said they probed public LB and it worked on private LB.</p>\n<p>Yes, i trained <code>2:1</code> for ease of PP. By controlling the train distribution carefully, I could use Bayes Rule to adjust these odds for the private test dataset based on the distribution observed from making predictions on the full test dataset.</p>",
      "rawMarkdown": "You are correct that probing public LB and using the result to post process for private LB is risky. However in many comps, public and private have a similar distribution and public is large enough to compute accurate enough means, so it works. In this comp, the 1st place team said they probed public LB and it worked on private LB.\n\nYes, i trained `2:1` for ease of PP. By controlling the train distribution carefully, I could use Bayes Rule to adjust these odds for the private test dataset based on the distribution observed from making predictions on the full test dataset.",
      "votes": null
    },
    {
      "id": "1228414",
      "postDate": "03/06/2021 10:54:20",
      "content": "<p>Congrats for solo gold chirs <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !<br>\nI 'm sorry about my dumb question, but what are odds really mean? Why we need to convert the probabilities to odds first?</p>\n<blockquote>\n  <p>odds = p / (1-p)</p>\n</blockquote>",
      "rawMarkdown": "Congrats for solo gold chirs @cdeotte !\nI 'm sorry about my dumb question, but what are odds really mean? Why we need to convert the probabilities to odds first?\n\n> odds = p / (1-p)",
      "votes": null
    },
    {
      "id": "1240063",
      "postDate": "03/16/2021 08:03:48",
      "content": "<p><a href=\"https://www.kaggle.com/karlyukang\" target=\"_blank\">@karlyukang</a> if I may answer, odds are nothing but ratio of success to failure, so a 20% success works out to 1 in 4 odds (if we repeat trial 5 times, we get 1 success and 4 failures for a 20% probability). </p>\n<p>Odds are helpful when scaling. </p>\n<p>If we don't use odds when scaling, the probability goes beyond 1 and then stops making sense. So if I want to scale above probability 10 times, I get 200% which is not a probability. However by converting to odds, scaling and then converting back to probability, I get what I want -  it converts to a probability of around 70%. </p>\n<p>So when scaling 2 different species, I can use appropriate scaling factors with the assurance that the final p is between 0-1 and the relative scales are maintained. You can try scaling on just 'p' and you will find that you can't compare species effectively any more</p>",
      "rawMarkdown": "karlyukang if I may answer, odds are nothing but ratio of success to failure, so a 20% success works out to 1 in 4 odds (if we repeat trial 5 times, we get 1 success and 4 failures for a 20% probability). \n\nOdds are helpful when scaling. \n\nIf we don't use odds when scaling, the probability goes beyond 1 and then stops making sense. So if I want to scale above probability 10 times, I get 200% which is not a probability. However by converting to odds, scaling and then converting back to probability, I get what I want -  it converts to a probability of around 70%. \n\nSo when scaling 2 different species, I can use appropriate scaling factors with the assurance that the final p is between 0-1 and the relative scales are maintained. You can try scaling on just 'p' and you will find that you can't compare species effectively any more",
      "votes": null
    },
    {
      "id": "1360249",
      "postDate": "06/22/2021 00:21:41",
      "content": "<p>thanks for the excellent write-up! I made a simpler trick to utilize the train v.s. test distributional gap: replacing any class with prob &lt; 0.5 by its test set class weight (inferred by LB probing). This gives 0.06 - 0.01 perf gain (stronger model tends to get smaller gain, so I guess my trick is trying to alleviate the cost incurred by false-negative). Will definitely try out your trick and see the difference! </p>\n<p>Also I hope I am not too late to ask a few questions regarding your solution:</p>\n<ol>\n<li>how did you infer the test set distribution? I tried LB probing by doing submission for each class and infer its distribution by normalizing the LB score across classes. What I got is quite different from what you stated, though reaching the same conclusions where class 3 and 18 is majority. (green color is the inferred distribution, and the 3 other columns are the LB score by ranking targeting class to the top while randomly shuffle the other classes) <img src=\"https://imgur.com/boi63V0.png\" alt=\"\"></li>\n<li>u normalized input by <code>img = (img+87)/128</code>. Is it an empirical way to make your input to have 0 mean and 1 variance?</li>\n<li>If I understand correctly, u precomputed the mel-spectrogram. Did u sacrafice the audio augmentations by doing so, or u get some ways to take into account those augmentations in your cache?</li>\n</ol>\n<p>Many thanks!</p>",
      "rawMarkdown": "thanks for the excellent write-up! I made a simpler trick to utilize the train v.s. test distributional gap: replacing any class with prob < 0.5 by its test set class weight (inferred by LB probing). This gives 0.06 - 0.01 perf gain (stronger model tends to get smaller gain, so I guess my trick is trying to alleviate the cost incurred by false-negative). Will definitely try out your trick and see the difference! \n\nAlso I hope I am not too late to ask a few questions regarding your solution:\n1. how did you infer the test set distribution? I tried LB probing by doing submission for each class and infer its distribution by normalizing the LB score across classes. What I got is quite different from what you stated, though reaching the same conclusions where class 3 and 18 is majority. (green color is the inferred distribution, and the 3 other columns are the LB score by ranking targeting class to the top while randomly shuffle the other classes) ![](https://imgur.com/boi63V0.png)\n2. u normalized input by `img = (img+87)/128`. Is it an empirical way to make your input to have 0 mean and 1 variance?\n3. If I understand correctly, u precomputed the mel-spectrogram. Did u sacrafice the audio augmentations by doing so, or u get some ways to take into account those augmentations in your cache?\n\nMany thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1208139,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "02/18/2021 07:19:02",
      "content": "<p>Amazing job in short time <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Your PP method is quite genius as always. I think doing the all-zeros probing is tricky as it depends on the sorting algorithm, better to set all random and the column to 1.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208152,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2021 07:27:58",
          "content": "<p>Thanks Psi. Congrats on another amazing 1st place finish. So impressive.</p>\n<p>Good point about the sorting algorithm. I probed one species and the result was a little weird, i think you just explained why. I will update my post with your idea.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208352,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2021 08:58:56",
          "content": "<p>I just post processed my final submission using the test distribution from Table 2 in the paper <a href=\"https://www.sciencedirect.com/science/article/pii/S1574954120300637\" target=\"_blank\">here</a>. It would have achieved private LB 0.970. So it looks like the paper's distribution is better than the one I calculate from my test predictions.</p>\n<p><img src=\"http://playagricola.com/Kaggle/paper_LB.png\" alt=\"image\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208368,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/18/2021 09:09:59",
          "content": "<p>this means that the kaggle data is the same as the paper?<br>\ninteresting!<br>\nif so, I think the orangnizer has given us too little data! There are some many TP and FP in the paper.</p>\n<p>I am interested the results of using all TP and FP annotations in the paper. Has kagglers managed to achieve the same results using only 5% of available annotation?</p>\n<hr>\n<p>on another note, such post-processing are useful for distillation later. It forces the model to learn something new.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208385,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2021 09:25:17",
          "content": "<p>The paper says they stored all their audio at website ARBIMON <a href=\"https://arbimon.rfcx.org/\" target=\"_blank\">https://arbimon.rfcx.org/</a> . I didn't search nor use any external data in this comp. I wonder whether the paper's 512471 audio clips are publicly available on ARBIMON? That's a half million audio clips!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208386,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 09:25:26",
          "content": "<p>The metric depends a lot on the scaling, so it is also difficult to compare with the results in the paper. So it seems from the paper that S3 is present in 90% of the samples, so you need to have it in top, otherwise the metric hurts a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208530,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/18/2021 10:19:06",
          "content": "<p>maybe some here:<br>\n<a href=\"https://arbimon.rfcx.org/project/perm-stations/audiodata/training-sets?set=2119\" target=\"_blank\">https://arbimon.rfcx.org/project/perm-stations/audiodata/training-sets?set=2119</a>?<br>\n<a href=\"https://arbimon.rfcx.org/project/perm-stations/audiodata/templates\" target=\"_blank\">https://arbimon.rfcx.org/project/perm-stations/audiodata/templates</a><br>\n<a href=\"https://arbimon.rfcx.org/project/rfcx-bird-frog/audiodata/templates\" target=\"_blank\">https://arbimon.rfcx.org/project/rfcx-bird-frog/audiodata/templates</a></p>\n<p><img src=\"https://i.ibb.co/C2vYzNK/Selection-051.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208150,
      "author_name": "riadalmadani",
      "author_url": "",
      "post_date": "02/18/2021 07:25:43",
      "content": "<p>Congratulations🎉🎉🎉<br>\ngreat job</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208157,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 07:30:13",
      "content": "<p>Congrats on solo gold medal, very strongly post processing method <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208226,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/18/2021 08:00:20",
      "content": "<p>Your PP rocks, as often.  I only started looking at test prediction distributions last 2 hours of the competition and noticed a shift indeed.  But I couldn't devise your PP in time.</p>\n<p>I like to think in terms of logits (log of odds)  instead of odds.  Then your pp amounts to move the mean of test logits to the mean of train logits for each class.  Is that right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208268,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2021 08:22:20",
          "content": "<p>Thanks CPMP. Yes i think you can use logits. Then you \"move\" the mean with addition instead of multiplication as in <br>\n    $$\\text{new logits} = \\text{adjust} + \\text{old logits}$$<br>\nwhere <code>adjust = log(FACTOR)</code> in my formula.</p>\n<p>Note that i simplified my PP a bit. I actually did it repeatedly. I would start with my original submission.csv then calculate FACTOR, then apply it to scale the submission.csv. Then i would calculate new FACTOR from the modified submission.csv. Next, I would start with my orginal submission.csv again, then apply new FACTOR. Then calculate another FACTOR. I would do this repeatedly until the FACTOR did not change anymore. Then I would use that FACTOR to create my final submission from original submission.</p>\n<p><code>FACTOR = test odds / train odds</code> but this needs subtle adjustment too. Because we train crops with <code>odds = x</code>, but we must estimate full clip odds from crop odds. For example if we train crops with 50 / 50 true postive and false positive, that is <code>1:1</code> crop odds. But the odds of a species being present in the entire 60 seconds after taking max over sliding crops is more like <code>2:1</code> train odds because there is a higher probability of seeing a species in 30 crops versus 1 crop. So when computing <code>FACTOR</code>, we use the higher <code>2:1</code> for <code>train odds</code>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208227,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "02/18/2021 08:00:47",
      "content": "<p>You seem to be the \"King\" of postprocessing. Just like Bengali, this is yet another impressive postprocessing from you. Congrats on that. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208369,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/18/2021 09:12:08",
      "content": "<p>Brilliant Chris, as always.</p>\n<p>Congratz on the solo gold !</p>\n<p>Do you mind sharing your pp code ? I wanna see how well it works for us :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 1208523,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2021 10:13:09",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Thanks Theo!</p>\n<p>Try the following. Make sure your <code>submission.csv</code> are probabilities between and including 0 and 1. Do not use logits. First try <code>MODE=1</code>, then if that isn't good, try <code>MODE=2</code>. If that isn't good try <code>MODE=3</code>. If that isn't good, try <code>MODE=1</code> with a different <code>FUDGE</code>. Perhaps try 0.5, 1, and 3.</p>\n<pre><code># USE MODE 1, 2, or 3\nMODE = 1\n\n# LOAD SUBMISSION\nimport pandas as pd, numpy as np\nFUDGE = 2.0\nFILE = 'submission.csv'\ndf = pd.read_csv(FILE)\nfor k in range(24):\n    df.iloc[:,1+k] -= df.iloc[:,1+k].min()\n    df.iloc[:,1+k] /= df.iloc[:,1+k].max()\n\n# CONVERT PROBS TO ODDS, APPLY MULTIPLIER, CONVERT BACK TO PROBS\ndef scale(probs, factor):\n    probs = probs.copy()\n    idx = np.where(probs!=1)[0]\n    odds = factor * probs[idx] / (1-probs[idx])\n    probs[idx] =  odds/(1+odds)\n    return probs\n\n# DIFFERENT DISTRIBUTIONS\nd1 = df.iloc[:,1:].mean().values\nd2 = np.array([113, 204, 44, 923, 53, 41, 3, 213, 44, 23, 26, 149, 255,  \n    14, 123, 222, 46, 6, 474, 4, 17, 18, 23, 72])/1000.\n\nfor k in range(24):\n    if MODE==1: d = FUDGE\n    if MODE==2: d = d1[k]/(1-d1[k])\n    if MODE==3: s = d2[k] / d1[k]\n    else: s = (d2[k]/(1-d2[k]))/d\n    df.iloc[:,k+1] = scale(df.iloc[:,k+1].values,s)\n\ndf.to_csv('submission_with_pp.csv',index=False)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208526,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2021 10:15:45",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> CPMP, try the above PP and see if it helps your submission.csv file. Make sure your file contains probabilities, not logits. And follow the instructions above which suggests different modes and fudges in case it doesn't work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208553,
          "author_name": "vlomme",
          "author_url": "",
          "post_date": "02/18/2021 10:37:38",
          "content": "<p>For me it gives +0.03(0.927 - 0.959)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208561,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2021 10:41:22",
          "content": "<p><a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> That's great Kramarenko. What <code>MODE</code> worked best for you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208581,
          "author_name": "vlomme",
          "author_url": "",
          "post_date": "02/18/2021 10:47:43",
          "content": "<p>Mode 1.  Mode 2 gives -0.01</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208586,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/18/2021 10:50:21",
          "content": "<p>Works like a charm for me as well <br>\n950 -&gt; 978 <br>\n960 -&gt; 980<br>\n976 -&gt; 982</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208587,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/18/2021 10:51:01",
          "content": "<p>Thanks a lot ! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208597,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 10:58:53",
          "content": "<p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> Looks like you have very strong models, hope you will do a writeup.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208636,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "02/18/2021 11:48:14",
          "content": "<p>Thanks Chris! <br>\nFor my ensemble, your pp is what I missed!<br>\nMode 3 <br>\n900 -&gt; 936 (SILVER ZONE)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208433,
      "author_name": "rajkumarl",
      "author_url": "",
      "post_date": "02/18/2021 09:41:32",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for this explanation. I start understanding now that postprocessing is more important.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208649,
      "author_name": "mnpinto",
      "author_url": "",
      "post_date": "02/18/2021 11:58:26",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! This is brilliant! My best submission score improved 0.940 -&gt; 0.950 with MODE=3.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208681,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "02/18/2021 12:36:32",
      "content": "<p>I've just submitted our final selection with your postprocessing <code>mode 2</code>. Our score would have improved from 0.967 to 0.979 ( 2nd place ). One of our worst decision not to accept your merging proposal. Maybe next time. Anyway, it was a great competition, and congratulations for your solo gold.</p>\n<p><img src=\"https://i.ibb.co/98TWw0G/Untitled.png\" alt=\"untitled\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1208698,
          "author_name": "gaborfodor",
          "author_url": "",
          "post_date": "02/18/2021 12:52:59",
          "content": "<p>At least we don't need to retrain and reproduce the 80+ models in our final blend 😅</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208728,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 13:09:08",
          "content": "<p>I wish it would be only 80 models :D</p>\n<p>I am surprised about private score being quite higher here on that sub.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208794,
      "author_name": "fffrrt",
      "author_url": "",
      "post_date": "02/18/2021 13:45:46",
      "content": "<p>My submission would have reached 0.951 instead of 0.918 with mode 3.</p>\n<p>Ouch.</p>\n<p>So this is where all the magic was…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209001,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "02/18/2021 16:02:35",
      "content": "<p>Unbelievable. Just like that, I have second place - masterful insights as usual, Chris!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209013,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "02/18/2021 16:19:04",
      "content": "<p>Unbelievable as usual ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209038,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/18/2021 16:41:38",
      "content": "<p>Thanks for this info. Our team had a discussion when you moved up the leaderboard that there was probably some post-processing or probing opportunity because we knew that was in your bag of tricks from the previous other competitions. </p>\n<p>We did some slightly less confident rescaling. We saw in the paper the tp and fp and frequencies but I don't think any of us read too much into it. Ours was purely based on lb probing of s3 and s18</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209096,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "02/18/2021 17:21:32",
          "content": "<p>Scaling only s3 and s18 gives me almost the same result (0.949 vs 0.951 for all species scaling with baseline 0.918).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209173,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2021 18:28:28",
          "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> That is interesting. For me, i gain a lot of benefit by post processing all species. Perhaps you already downsample the other species in your training process. </p>\n<p>I train all species with 50% true positive and 50% false positive. So even my rare species get predicted with high probabilities and thus I need to PP them to make them smaller.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209228,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "02/18/2021 19:03:30",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I used false positives as an extra class in training, and just ignored it during prediction.</p>\n<p>This is how percentiles of logits on test data looks like pre-scaling - <br>\n<img src=\"https://i.imgur.com/Ryn4RhJ.png\" alt=\"https://i.imgur.com/Ryn4RhJ.png\"></p>\n<pre><code>submission[:,3] = submission[:,3] + 5\nsubmission[:,18] = submission[:,18] + 5\n</code></pre>\n<p>This two-liner would have made it a 0.949 one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209260,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2021 19:26:14",
          "content": "<p>Wow, simple adjustment for big gain. I see your predictions are logits, so adding 5 is like scaling, i.e. multiplying, the odds by <code>1.6 = log(5)</code>. </p>\n<p>This simple trick works in many comps. Whenever the metric is AUC comparing different targets, and this comp metric was AUC in disguise, then we can improve LB by making sure each group of predicted targets is properly calibrated against other groups. This same trick was used in Jigsaw Toxic Comp <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980\" target=\"_blank\">here</a>, described in section post process.</p>\n<p>And other metrics benefit from PP adjustments too. For example, Log Loss and MSE are sensitive to the mean as shown <a href=\"https://www.kaggle.com/cdeotte/moa-post-process-lb-1777\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209205,
      "author_name": "egm108",
      "author_url": "",
      "post_date": "02/18/2021 18:51:08",
      "content": "<p>Thank you and congratulations!</p>\n<p>Is it correct that this PP uses the fact that distribution on Private is similar to Public? I didn't check that, just see that the scores are pretty similar all the way for me…?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209292,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2021 19:56:20",
          "content": "<p>This PP doesn't require that public and private are similar. This PP requires that we calculate <code>FACTORS</code>. When we calculate the <code>FACTORS</code> we need to make sure that they pertrain to private LB.</p>\n<p>I list 3 ways to calculate <code>FACTORS</code> above. If we compute <code>FACTORS</code> from probing public LB, that is risky. If we use the entire test data to compute <code>FACTORS</code> that is safest. The last method is using the <code>FACTORS</code> from the paper. That is risky too because we are not sure if they pertain to private.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209338,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/18/2021 20:50:31",
      "content": "<p>I tried your pp on few subs, the best I selected gets to 0.9739 private with your pp.</p>\n<p>It means two things:</p>\n<ol>\n<li><p>my modeling is quite good, I didn't miss something key contrarily to what i thought.</p></li>\n<li><p>I need to pay attention to test prediction distribution.  It is not the first time that this makes a difference between my solution and top ones</p></li>\n</ol>\n<p>Thanks a lot for sharing.  FYI, MODE =1 is what worked best with the subs I tested.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210964,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/19/2021 20:47:42",
          "content": "<p>That's awesome <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> . You have a very accurate model ! </p>\n<p>When i use <code>MODE=1</code> with the code posted here including the <code>d2</code> vector from paper, my result is private LB 0.970. During comp i computed my own <code>d2</code> array and didn't use the paper and achieved private LB 0.963 so it seems the paper's test distribution is the correct one and better than i could calculate from my test predictions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209377,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "02/18/2021 21:21:33",
      "content": "<p>I have used same strategy mode1 jumped from 0.909, 3 models ensemble, to 0.950 which was enough to get 12th place : gold? because a gold team was removed :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209383,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2021 21:26:16",
          "content": "<p>Congratulations tugstugi. Great job achieving solo gold ! I saw you climbing quickly the last few days. Well done!</p>\n<p>It's interesting how PP increased your private LB score more than you public LB score. Do you have any idea why?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209395,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "02/18/2021 21:33:58",
          "content": "<p>no idea, probably luck.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209385,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/18/2021 21:27:08",
      "content": "<p>Wow this is amazing<br>\nMy private LB score is jumped to 0.94729 from 0.90990.</p>\n<p>Congrats on your solo gold. Your helps in the comment sections also helped me to improve my score as well. Thanks 🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209463,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "02/18/2021 23:06:10",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! Great solution and nice PP trick.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210966,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/19/2021 20:48:41",
          "content": "<p>Thanks Giba</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210309,
      "author_name": "aryaman1999",
      "author_url": "",
      "post_date": "02/19/2021 10:42:19",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> great notebook. I was wondering what was the intuition behind the masked loss function you used, why wouldn't only using the true positives and training using simple binary cross entropy work?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210368,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/19/2021 11:34:10",
          "content": "<p>All top teams used masked loss. Why?  Because we are given TP and TN for species one at a time.  There is no point having a loss for unknown labels.</p>\n<p>It seems that many assumed that TP are positive only for the annotated class and that target for all other species should be 0.  This is a wrong assumption.</p>\n<p>The only known negative targets are the FPs.  They should have been called TN as they are the only true negatives we have access to,</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212057,
      "author_name": "mehrankazeminia",
      "author_url": "",
      "post_date": "02/20/2021 20:38:18",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>I would like to add an <strong>important point</strong> to this trick:</p>\n<p>Not all results are improved with this trick. For example, in <a href=\"https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\" target=\"_blank\">your notebook</a>, if only the results of 15 columns change with this trick, but the results of the other nine columns do not change at all, much better scores are obtained.</p>\n<p>To prove this, we wrote a notebook that you and all those interested can read at the following address.</p>\n<p><a href=\"https://www.kaggle.com/mehrankazeminia/lb-0-980-rainforest-comparative-method-part-a\" target=\"_blank\">https://www.kaggle.com/mehrankazeminia/lb-0-980-rainforest-comparative-method-part-a</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1212083,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/20/2021 21:24:00",
          "content": "<p>Interesting. Nice analysis and discovery!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1217321,
      "author_name": "pukkinming",
      "author_url": "",
      "post_date": "02/25/2021 00:07:28",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nReally thank you for your nice sharing. I have one question though:</p>\n<pre><code>Probe the LB with submission.csv of all zeros and one column of ones. Then use math to compute factor for that species. UPDATE use random numbers less than 1 instead of 0s to avoid sorting uncertainty.\n</code></pre>\n<p>I thought probing LB only got the distribution for public LB ( from the public LB score ). So how can we compute the factor for the distribution which generalize both public and private LB well? And is that the reason why you trained a 2:1 model for ease of post-processing?</p>\n<p>Thank you so much!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1217344,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/25/2021 01:07:05",
          "content": "<p>You are correct that probing public LB and using the result to post process for private LB is risky. However in many comps, public and private have a similar distribution and public is large enough to compute accurate enough means, so it works. In this comp, the 1st place team said they probed public LB and it worked on private LB.</p>\n<p>Yes, i trained <code>2:1</code> for ease of PP. By controlling the train distribution carefully, I could use Bayes Rule to adjust these odds for the private test dataset based on the distribution observed from making predictions on the full test dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1228414,
      "author_name": "karlyukang",
      "author_url": "",
      "post_date": "03/06/2021 10:54:20",
      "content": "<p>Congrats for solo gold chirs <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !<br>\nI 'm sorry about my dumb question, but what are odds really mean? Why we need to convert the probabilities to odds first?</p>\n<blockquote>\n  <p>odds = p / (1-p)</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1240063,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "03/16/2021 08:03:48",
          "content": "<p><a href=\"https://www.kaggle.com/karlyukang\" target=\"_blank\">@karlyukang</a> if I may answer, odds are nothing but ratio of success to failure, so a 20% success works out to 1 in 4 odds (if we repeat trial 5 times, we get 1 success and 4 failures for a 20% probability). </p>\n<p>Odds are helpful when scaling. </p>\n<p>If we don't use odds when scaling, the probability goes beyond 1 and then stops making sense. So if I want to scale above probability 10 times, I get 200% which is not a probability. However by converting to odds, scaling and then converting back to probability, I get what I want -  it converts to a probability of around 70%. </p>\n<p>So when scaling 2 different species, I can use appropriate scaling factors with the assurance that the final p is between 0-1 and the relative scales are maintained. You can try scaling on just 'p' and you will find that you can't compare species effectively any more</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1360249,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "06/22/2021 00:21:41",
      "content": "<p>thanks for the excellent write-up! I made a simpler trick to utilize the train v.s. test distributional gap: replacing any class with prob &lt; 0.5 by its test set class weight (inferred by LB probing). This gives 0.06 - 0.01 perf gain (stronger model tends to get smaller gain, so I guess my trick is trying to alleviate the cost incurred by false-negative). Will definitely try out your trick and see the difference! </p>\n<p>Also I hope I am not too late to ask a few questions regarding your solution:</p>\n<ol>\n<li>how did you infer the test set distribution? I tried LB probing by doing submission for each class and infer its distribution by normalizing the LB score across classes. What I got is quite different from what you stated, though reaching the same conclusions where class 3 and 18 is majority. (green color is the inferred distribution, and the 3 other columns are the LB score by ranking targeting class to the top while randomly shuffle the other classes) <img src=\"https://imgur.com/boi63V0.png\" alt=\"\"></li>\n<li>u normalized input by <code>img = (img+87)/128</code>. Is it an empirical way to make your input to have 0 mean and 1 variance?</li>\n<li>If I understand correctly, u precomputed the mel-spectrogram. Did u sacrafice the audio augmentations by doing so, or u get some ways to take into account those augmentations in your cache?</li>\n</ol>\n<p>Many thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1208127": "Thanks Kaggle and RFCx for a fun competition. My final submission without post process achieves private LB 0.926 and with post process achieves **private LB 0.963**! That's +0.037 with post process!\n\n# How To Score LB 0.950+\nThe metric in this competition is different than other competitions. We are asked to provide `submission.csv` where each row is a test sample, and each column has a species prediction.\n\nIn other competitions, the metric computes **column wise AUC**. In this competition, the metric is essentially **row wise AUC**. Therefore we need each column to represent probabilities. The columns of common species need to be large values and the columns of rare species need to be small values.\n\n# Train Distribution\nThe distribution of the train data has roughly 6x the number of false positives versus true positives for each species. If you train your model with that data, then when your model is unsure about a prediction, it will predict the mean that it observed in the train data which is 1/7. Therefore if it is unsure about species 3, it will predict 1/7 and if it is unsure about species 19, it will predict 1/7.\n\nThis is a problem because species 3 appears in roughly 90% of test samples whereas species 19 appears in roughly 1% of test samples. Therefore when your model is unsure , it should predict 90% for species 3 and 1% for species 19.\n\n# Test Distribution - Post Process\nIn order to correct our model's predictions we scale the odds. Note that scaling odds doesn't affect predictions of 0 and 1. It only affects the unsure middle predictions. \n\nFirst we convert the `submission.csv` column of probabilities into odds with the formula\n$$ \\text{odds} = \\frac{p}{1-p}$$\nThen we scale the odds with\n$$ \\text{new odds} = \\text{factor} * \\text{old odds}$$\nAnd lastly we convert back to probabilities\n$$ \\text{prob} = \\frac{\\text{new odds}}{1 + \\text{new odds}}$$\n\n# Sample Code\n    # CONVERT PROB TO ODDS, APPLY MULTIPLIER, CONVERT BACK TO PROB\n    def scale(probs, factor):\n        probs = probs.copy()\n        idx = np.where(probs!=1)[0]\n        odds = factor * probs[idx] / (1-probs[idx])\n        probs[idx] =  odds/(1+odds)\n        return probs\n\n    for k in range(24):\n        sub.iloc[:,1+k] = scale(sub.iloc[:,1+k].values, FACTORS[k])\n\n# Increase LB by +0.040!\nThe only detail remaining is how to calculate the `FACTORS` above. There are at least 3 ways.\n* Create a model with output layer sigmoid, not softmax. Train with BCE loss using both true and false positives. Predict the test data. Compute the mean of each column. Convert to odds and divide by the odds of training data.\n* Probe the LB with `submission.csv` of all zeros and one column of ones. Then use math to compute `factor` for that species. UPDATE use random numbers less than 1 instead of 0s to avoid sorting uncertainty.\n* Use the factors listed in RFCx's paper [here][1], Table 2 in Section 2.4\n\nPersonally, i used the first option listed above. I didn't have 5 days to probe the LB 24 times for the second option. And I didn't find the paper until the last day of the competition for the third option.\n\nMy single model scores private LB 0.921 without post process and scores private LB 0.958 with post process. Ensembling a variety of image sizes and backbones increased the LBs to 0.926 and LB 0.963 respectively.\n\n# Model Details\nI converted each audio file into Mel Spectrogram with Librosa `feature.melspectrogram` and `power_to_db` using sampling rate `32_000`, `n_mels = 384`, `n_fft=2048`, `hop_length=512`, `win_length=2048`. This produced NumPy arrays of size `(384, 3751)`. I later normalized them with `img = (img+87)/128` and  trained with random crops of `384x384` which had frequency range 20Hz to 16_000Hz and time range 6.14 seconds. Each crop contained at least 75% of a true positive or false positive.\n\nI concatenated the true postive and false positive CSV files from Kaggle. My dataloader provided labels, masks, and one color images. The labels and masks were contained in the vector `y` which was 48 zeros where 2 were potentially altered. For species `k`, the `kth` element was 0 or 1 corresponding to false positive or true positive respectively. And the `k+24th` element was 1 indicating mask true to calculate loss for the `kth` species. I trained TF model with the following loss\n\n     def masked_loss(y_true, y_pred):\n    \n         mask = y_true[:,24:]\n         y_true = y_true[:,:24]\n      \n         y_pred = tf.convert_to_tensor(y_pred)\n         y_true = tf.cast(y_true, y_pred.dtype)\n         mask = tf.cast(mask, y_pred.dtype)\n         y_pred = tf.math.multiply(mask, y_pred, name=None)\n\n         return K.mean( K.binary_crossentropy(y_true, y_pred), axis=-1 )*24.0\n\nI used EfficientNetB2 with `albu.CoarseDropout` and `albu.RandomBrightnessContrast`. The optimizer was Adam with learning rate `1e-3` and reduce on plateau `factor=0.3`, `patience=3`. The final layer of the model was `GlobalAveragePooling2D()` and `Dense(24, activation='sigmoid')`. I monitored `val_loss` to tune hyperameters.\n\n# Try PP on Your Sub\nIf you want to try post process on your `submission.csv` file, I posted a Kaggle notebook [here][2]. It uses the 3rd method above for computing `FACTORS`. It also has 3 `MODE` you can try to account for different ways that you may have used to train your models.\n\n[1]: https://www.sciencedirect.com/science/article/pii/S1574954120300637\n[2]: https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970",
    "1208139": "Amazing job in short time @cdeotte. Your PP method is quite genius as always. I think doing the all-zeros probing is tricky as it depends on the sorting algorithm, better to set all random and the column to 1.",
    "1208150": "Congratulations🎉🎉🎉\ngreat job",
    "1208152": "Thanks Psi. Congrats on another amazing 1st place finish. So impressive.\n\nGood point about the sorting algorithm. I probed one species and the result was a little weird, i think you just explained why. I will update my post with your idea.",
    "1208157": "Congrats on solo gold medal, very strongly post processing method @cdeotte",
    "1208226": "Your PP rocks, as often.  I only started looking at test prediction distributions last 2 hours of the competition and noticed a shift indeed.  But I couldn't devise your PP in time.\n\nI like to think in terms of logits (log of odds)  instead of odds.  Then your pp amounts to move the mean of test logits to the mean of train logits for each class.  Is that right?",
    "1208227": "You seem to be the \"King\" of postprocessing. Just like Bengali, this is yet another impressive postprocessing from you. Congrats on that.",
    "1208268": "Thanks CPMP. Yes i think you can use logits. Then you \"move\" the mean with addition instead of multiplication as in \n    $$\\text{new logits} = \\text{adjust} + \\text{old logits}$$\nwhere `adjust = log(FACTOR)` in my formula.\n\nNote that i simplified my PP a bit. I actually did it repeatedly. I would start with my original submission.csv then calculate FACTOR, then apply it to scale the submission.csv. Then i would calculate new FACTOR from the modified submission.csv. Next, I would start with my orginal submission.csv again, then apply new FACTOR. Then calculate another FACTOR. I would do this repeatedly until the FACTOR did not change anymore. Then I would use that FACTOR to create my final submission from original submission.\n\n`FACTOR = test odds / train odds` but this needs subtle adjustment too. Because we train crops with `odds = x`, but we must estimate full clip odds from crop odds. For example if we train crops with 50 / 50 true postive and false positive, that is `1:1` crop odds. But the odds of a species being present in the entire 60 seconds after taking max over sliding crops is more like `2:1` train odds because there is a higher probability of seeing a species in 30 crops versus 1 crop. So when computing `FACTOR`, we use the higher `2:1` for `train odds`.",
    "1208352": "I just post processed my final submission using the test distribution from Table 2 in the paper [here][1]. It would have achieved private LB 0.970. So it looks like the paper's distribution is better than the one I calculate from my test predictions.\n\n![image](http://playagricola.com/Kaggle/paper_LB.png)\n\n[1]: https://www.sciencedirect.com/science/article/pii/S1574954120300637",
    "1208368": "this means that the kaggle data is the same as the paper?\ninteresting!\nif so, I think the orangnizer has given us too little data! There are some many TP and FP in the paper.\n\nI am interested the results of using all TP and FP annotations in the paper. Has kagglers managed to achieve the same results using only 5% of available annotation?\n\n---\non another note, such post-processing are useful for distillation later. It forces the model to learn something new.",
    "1208369": "Brilliant Chris, as always.\n\nCongratz on the solo gold !\n\nDo you mind sharing your pp code ? I wanna see how well it works for us :)",
    "1208385": "The paper says they stored all their audio at website ARBIMON https://arbimon.rfcx.org/ . I didn't search nor use any external data in this comp. I wonder whether the paper's 512471 audio clips are publicly available on ARBIMON? That's a half million audio clips!",
    "1208386": "The metric depends a lot on the scaling, so it is also difficult to compare with the results in the paper. So it seems from the paper that S3 is present in 90% of the samples, so you need to have it in top, otherwise the metric hurts a lot.",
    "1208433": "Thanks @cdeotte for this explanation. I start understanding now that postprocessing is more important.",
    "1208523": "theoviel Thanks Theo!\n\nTry the following. Make sure your `submission.csv` are probabilities between and including 0 and 1. Do not use logits. First try `MODE=1`, then if that isn't good, try `MODE=2`. If that isn't good try `MODE=3`. If that isn't good, try `MODE=1` with a different `FUDGE`. Perhaps try 0.5, 1, and 3.\n\n    # USE MODE 1, 2, or 3\n    MODE = 1\n\n    # LOAD SUBMISSION\n    import pandas as pd, numpy as np\n    FUDGE = 2.0\n    FILE = 'submission.csv'\n    df = pd.read_csv(FILE)\n    for k in range(24):\n        df.iloc[:,1+k] -= df.iloc[:,1+k].min()\n        df.iloc[:,1+k] /= df.iloc[:,1+k].max()\n\n    # CONVERT PROBS TO ODDS, APPLY MULTIPLIER, CONVERT BACK TO PROBS\n    def scale(probs, factor):\n        probs = probs.copy()\n        idx = np.where(probs!=1)[0]\n        odds = factor * probs[idx] / (1-probs[idx])\n        probs[idx] =  odds/(1+odds)\n        return probs\n\n    # DIFFERENT DISTRIBUTIONS\n    d1 = df.iloc[:,1:].mean().values\n    d2 = np.array([113, 204, 44, 923, 53, 41, 3, 213, 44, 23, 26, 149, 255,  \n        14, 123, 222, 46, 6, 474, 4, 17, 18, 23, 72])/1000.\n\n    for k in range(24):\n        if MODE==1: d = FUDGE\n        if MODE==2: d = d1[k]/(1-d1[k])\n        if MODE==3: s = d2[k] / d1[k]\n        else: s = (d2[k]/(1-d2[k]))/d\n        df.iloc[:,k+1] = scale(df.iloc[:,k+1].values,s)\n    \n    df.to_csv('submission_with_pp.csv',index=False)",
    "1208526": "cpmpml CPMP, try the above PP and see if it helps your submission.csv file. Make sure your file contains probabilities, not logits. And follow the instructions above which suggests different modes and fudges in case it doesn't work.",
    "1208530": "maybe some here:\nhttps://arbimon.rfcx.org/project/perm-stations/audiodata/training-sets?set=2119?\nhttps://arbimon.rfcx.org/project/perm-stations/audiodata/templates\nhttps://arbimon.rfcx.org/project/rfcx-bird-frog/audiodata/templates\n\n![](https://i.ibb.co/C2vYzNK/Selection-051.png)",
    "1208553": "For me it gives +0.03(0.927 - 0.959)",
    "1208561": "vlomme That's great Kramarenko. What `MODE` worked best for you?",
    "1208581": "Mode 1.  Mode 2 gives -0.01",
    "1208586": "Works like a charm for me as well \n950 -> 978 \n960 -> 980\n976 -> 982",
    "1208587": "Thanks a lot ! @cdeotte :)",
    "1208597": "selimsef Looks like you have very strong models, hope you will do a writeup.",
    "1208636": "Thanks Chris! \nFor my ensemble, your pp is what I missed!\nMode 3 \n900 -> 936 (SILVER ZONE)",
    "1208649": "Thanks a lot @cdeotte! This is brilliant! My best submission score improved 0.940 -> 0.950 with MODE=3.",
    "1208681": "I've just submitted our final selection with your postprocessing `mode 2`. Our score would have improved from 0.967 to 0.979 ( 2nd place ). One of our worst decision not to accept your merging proposal. Maybe next time. Anyway, it was a great competition, and congratulations for your solo gold.\n\n![untitled](https://i.ibb.co/98TWw0G/Untitled.png)",
    "1208698": "At least we don't need to retrain and reproduce the 80+ models in our final blend 😅",
    "1208728": "I wish it would be only 80 models :D\n\nI am surprised about private score being quite higher here on that sub.",
    "1208794": "My submission would have reached 0.951 instead of 0.918 with mode 3.\n\nOuch.\n\nSo this is where all the magic was...",
    "1209001": "Unbelievable. Just like that, I have second place - masterful insights as usual, Chris!",
    "1209013": "Unbelievable as usual !",
    "1209038": "Thanks for this info. Our team had a discussion when you moved up the leaderboard that there was probably some post-processing or probing opportunity because we knew that was in your bag of tricks from the previous other competitions. \n\nWe did some slightly less confident rescaling. We saw in the paper the tp and fp and frequencies but I don't think any of us read too much into it. Ours was purely based on lb probing of s3 and s18",
    "1209096": "Scaling only s3 and s18 gives me almost the same result (0.949 vs 0.951 for all species scaling with baseline 0.918).",
    "1209173": "fffrrt That is interesting. For me, i gain a lot of benefit by post processing all species. Perhaps you already downsample the other species in your training process. \n\nI train all species with 50% true positive and 50% false positive. So even my rare species get predicted with high probabilities and thus I need to PP them to make them smaller.",
    "1209205": "Thank you and congratulations!\n\nIs it correct that this PP uses the fact that distribution on Private is similar to Public? I didn't check that, just see that the scores are pretty similar all the way for me...?",
    "1209228": "cdeotte I used false positives as an extra class in training, and just ignored it during prediction.\n\nThis is how percentiles of logits on test data looks like pre-scaling - \n![https://i.imgur.com/Ryn4RhJ.png](https://i.imgur.com/Ryn4RhJ.png)\n\n```\nsubmission[:,3] = submission[:,3] + 5\nsubmission[:,18] = submission[:,18] + 5\n```\nThis two-liner would have made it a 0.949 one.",
    "1209260": "Wow, simple adjustment for big gain. I see your predictions are logits, so adding 5 is like scaling, i.e. multiplying, the odds by `1.6 = log(5)`. \n\nThis simple trick works in many comps. Whenever the metric is AUC comparing different targets, and this comp metric was AUC in disguise, then we can improve LB by making sure each group of predicted targets is properly calibrated against other groups. This same trick was used in Jigsaw Toxic Comp [here][1], described in section post process.\n\nAnd other metrics benefit from PP adjustments too. For example, Log Loss and MSE are sensitive to the mean as shown [here][2]\n\n[1]: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980\n[2]: https://www.kaggle.com/cdeotte/moa-post-process-lb-1777",
    "1209292": "This PP doesn't require that public and private are similar. This PP requires that we calculate `FACTORS`. When we calculate the `FACTORS` we need to make sure that they pertrain to private LB.\n\nI list 3 ways to calculate `FACTORS` above. If we compute `FACTORS` from probing public LB, that is risky. If we use the entire test data to compute `FACTORS` that is safest. The last method is using the `FACTORS` from the paper. That is risky too because we are not sure if they pertain to private.",
    "1209338": "I tried your pp on few subs, the best I selected gets to 0.9739 private with your pp.\n\nIt means two things:\n\n1. my modeling is quite good, I didn't miss something key contrarily to what i thought.\n\n2. I need to pay attention to test prediction distribution.  It is not the first time that this makes a difference between my solution and top ones\n\nThanks a lot for sharing.  FYI, MODE =1 is what worked best with the subs I tested.",
    "1209377": "I have used same strategy mode1 jumped from 0.909, 3 models ensemble, to 0.950 which was enough to get 12th place : gold? because a gold team was removed :)",
    "1209383": "Congratulations tugstugi. Great job achieving solo gold ! I saw you climbing quickly the last few days. Well done!\n\nIt's interesting how PP increased your private LB score more than you public LB score. Do you have any idea why?",
    "1209385": "Wow this is amazing\nMy private LB score is jumped to 0.94729 from 0.90990.\n\nCongrats on your solo gold. Your helps in the comment sections also helped me to improve my score as well. Thanks 🙏",
    "1209395": "no idea, probably luck.",
    "1209463": "Congrats @cdeotte ! Great solution and nice PP trick.",
    "1210309": "Hey @cdeotte great notebook. I was wondering what was the intuition behind the masked loss function you used, why wouldn't only using the true positives and training using simple binary cross entropy work?",
    "1210368": "All top teams used masked loss. Why?  Because we are given TP and TN for species one at a time.  There is no point having a loss for unknown labels.\n\nIt seems that many assumed that TP are positive only for the annotated class and that target for all other species should be 0.  This is a wrong assumption.\n\nThe only known negative targets are the FPs.  They should have been called TN as they are the only true negatives we have access to,",
    "1210964": "That's awesome @cpmpml . You have a very accurate model ! \n\nWhen i use `MODE=1` with the code posted here including the `d2` vector from paper, my result is private LB 0.970. During comp i computed my own `d2` array and didn't use the paper and achieved private LB 0.963 so it seems the paper's test distribution is the correct one and better than i could calculate from my test predictions.",
    "1210966": "Thanks Giba",
    "1212057": "Thank you @cdeotte \n\nI would like to add an **important point** to this trick:\n\nNot all results are improved with this trick. For example, in [your notebook](https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970), if only the results of 15 columns change with this trick, but the results of the other nine columns do not change at all, much better scores are obtained.\n\nTo prove this, we wrote a notebook that you and all those interested can read at the following address.\n\nhttps://www.kaggle.com/mehrankazeminia/lb-0-980-rainforest-comparative-method-part-a",
    "1212083": "Interesting. Nice analysis and discovery!",
    "1217321": "cdeotte \nReally thank you for your nice sharing. I have one question though:\n\n```\nProbe the LB with submission.csv of all zeros and one column of ones. Then use math to compute factor for that species. UPDATE use random numbers less than 1 instead of 0s to avoid sorting uncertainty.\n```\n\nI thought probing LB only got the distribution for public LB ( from the public LB score ). So how can we compute the factor for the distribution which generalize both public and private LB well? And is that the reason why you trained a 2:1 model for ease of post-processing?\n\nThank you so much!",
    "1217344": "You are correct that probing public LB and using the result to post process for private LB is risky. However in many comps, public and private have a similar distribution and public is large enough to compute accurate enough means, so it works. In this comp, the 1st place team said they probed public LB and it worked on private LB.\n\nYes, i trained `2:1` for ease of PP. By controlling the train distribution carefully, I could use Bayes Rule to adjust these odds for the private test dataset based on the distribution observed from making predictions on the full test dataset.",
    "1228414": "Congrats for solo gold chirs @cdeotte !\nI 'm sorry about my dumb question, but what are odds really mean? Why we need to convert the probabilities to odds first?\n\n> odds = p / (1-p)",
    "1240063": "karlyukang if I may answer, odds are nothing but ratio of success to failure, so a 20% success works out to 1 in 4 odds (if we repeat trial 5 times, we get 1 success and 4 failures for a 20% probability). \n\nOdds are helpful when scaling. \n\nIf we don't use odds when scaling, the probability goes beyond 1 and then stops making sense. So if I want to scale above probability 10 times, I get 200% which is not a probability. However by converting to odds, scaling and then converting back to probability, I get what I want -  it converts to a probability of around 70%. \n\nSo when scaling 2 different species, I can use appropriate scaling factors with the assurance that the final p is between 0-1 and the relative scales are maintained. You can try scaling on just 'p' and you will find that you can't compare species effectively any more",
    "1360249": "thanks for the excellent write-up! I made a simpler trick to utilize the train v.s. test distributional gap: replacing any class with prob < 0.5 by its test set class weight (inferred by LB probing). This gives 0.06 - 0.01 perf gain (stronger model tends to get smaller gain, so I guess my trick is trying to alleviate the cost incurred by false-negative). Will definitely try out your trick and see the difference! \n\nAlso I hope I am not too late to ask a few questions regarding your solution:\n1. how did you infer the test set distribution? I tried LB probing by doing submission for each class and infer its distribution by normalizing the LB score across classes. What I got is quite different from what you stated, though reaching the same conclusions where class 3 and 18 is majority. (green color is the inferred distribution, and the 3 other columns are the LB score by ranking targeting class to the top while randomly shuffle the other classes) ![](https://imgur.com/boi63V0.png)\n2. u normalized input by `img = (img+87)/128`. Is it an empirical way to make your input to have 0 mean and 1 variance?\n3. If I understand correctly, u precomputed the mel-spectrogram. Did u sacrafice the audio augmentations by doing so, or u get some ways to take into account those augmentations in your cache?\n\nMany thanks!"
  },
  "source": "meta"
}