{
  "id": 160986,
  "title": "1st place solution: post-processing",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/160986",
  "author_name": "",
  "post_date": "2020-06-23T11:10:13.482092700Z",
  "votes": 58,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Firstly, a big thanks to Kaggle for constantly delivering top-notch competitions. Competitions doesn't always end smooth (cf. Deepfakes 😄) but there is no other data science platform so rich in learning. Another thanks to Jigsaw for the interesting, multi-lingual, shakeup-free ride!  <br>\nAlso congrats to the other medalists, and to my talented team-mate <a href=\"https://www.kaggle.com/leecming\" target=\"_blank\">@leecming</a>!</p>\n<p>Since other competitors were asking about our post-processing technique, I dedicate this post to explain it in more detail. </p>\n<p>To see the post-processing in action, have a look at the related <a href=\"https://www.kaggle.com/rafiko1/1st-place-jigsaw-post-processing-example\" target=\"_blank\">notebook</a></p>\n<h2>Post-processing: intuition</h2>\n<p>The intuition of the post-processing is as following: we consider the <strong><em>trend</em></strong> of subsequent submissions of a specific language (e.g. Russian) for each example in the test dataset. If the trend of that example is positive i.e. going up, we <strong><em>nudge</em></strong> the example further in the positive direction. And vice versa - if the trend is negative i.e. going down, we nudge the example further in the negative direction.</p>\n<p>We measure the trend by taking the differences of all subsequent submissions for the specific language and averaging those differences. The nudge that we then give to the new submission is based on a predefined <strong><em>weight</em></strong>, typically we choose a weight of 1 or 1.5. </p>\n<h2>Post-processing: pseudo-code</h2>\n<p>In an attempt to pseudo-code the technique, given:</p>\n<pre><code>weight = predefined weight (typically 1 or 1.5)\npred_best = current best predictions on LB\ndiff_avg = average of differences of consecutive subs (trend)\n</code></pre>\n<p>Then for each example in test of the specific language (e.g. Turkish):</p>\n<pre><code>if diff_avg &lt; 0: # negative trend\n    pred_new = (1+weight*diff_avg)*pred_best  # nudge downwards\nelse: # positive trend\n    pred_new = (1-weight*diff_avg)*pred_best + weight*diff_avg # nudge upwards\n</code></pre>\n<p>Note: I'm not an expert on PP, so I assume this technique can be further optimized. The boost we got from it was relatively small (albeit significant) compared to the other methods we implemented.</p>",
  "messages": [
    {
      "id": "898183",
      "postDate": "06/23/2020 11:10:13",
      "content": "<p>Firstly, a big thanks to Kaggle for constantly delivering top-notch competitions. Competitions doesn't always end smooth (cf. Deepfakes 😄) but there is no other data science platform so rich in learning. Another thanks to Jigsaw for the interesting, multi-lingual, shakeup-free ride!  <br>\nAlso congrats to the other medalists, and to my talented team-mate <a href=\"https://www.kaggle.com/leecming\" target=\"_blank\">@leecming</a>!</p>\n<p>Since other competitors were asking about our post-processing technique, I dedicate this post to explain it in more detail. </p>\n<p>To see the post-processing in action, have a look at the related <a href=\"https://www.kaggle.com/rafiko1/1st-place-jigsaw-post-processing-example\" target=\"_blank\">notebook</a></p>\n<h2>Post-processing: intuition</h2>\n<p>The intuition of the post-processing is as following: we consider the <strong><em>trend</em></strong> of subsequent submissions of a specific language (e.g. Russian) for each example in the test dataset. If the trend of that example is positive i.e. going up, we <strong><em>nudge</em></strong> the example further in the positive direction. And vice versa - if the trend is negative i.e. going down, we nudge the example further in the negative direction.</p>\n<p>We measure the trend by taking the differences of all subsequent submissions for the specific language and averaging those differences. The nudge that we then give to the new submission is based on a predefined <strong><em>weight</em></strong>, typically we choose a weight of 1 or 1.5. </p>\n<h2>Post-processing: pseudo-code</h2>\n<p>In an attempt to pseudo-code the technique, given:</p>\n<pre><code>weight = predefined weight (typically 1 or 1.5)\npred_best = current best predictions on LB\ndiff_avg = average of differences of consecutive subs (trend)\n</code></pre>\n<p>Then for each example in test of the specific language (e.g. Turkish):</p>\n<pre><code>if diff_avg &lt; 0: # negative trend\n    pred_new = (1+weight*diff_avg)*pred_best  # nudge downwards\nelse: # positive trend\n    pred_new = (1-weight*diff_avg)*pred_best + weight*diff_avg # nudge upwards\n</code></pre>\n<p>Note: I'm not an expert on PP, so I assume this technique can be further optimized. The boost we got from it was relatively small (albeit significant) compared to the other methods we implemented.</p>",
      "rawMarkdown": "Firstly, a big thanks to Kaggle for constantly delivering top-notch competitions. Competitions doesn't always end smooth (cf. Deepfakes 😄) but there is no other data science platform so rich in learning. Another thanks to Jigsaw for the interesting, multi-lingual, shakeup-free ride!  \nAlso congrats to the other medalists, and to my talented team-mate @leecming!\n\nSince other competitors were asking about our post-processing technique, I dedicate this post to explain it in more detail. \n\nTo see the post-processing in action, have a look at the related [notebook](https://www.kaggle.com/rafiko1/1st-place-jigsaw-post-processing-example)\n## Post-processing: intuition\n\nThe intuition of the post-processing is as following: we consider the ***trend*** of subsequent submissions of a specific language (e.g. Russian) for each example in the test dataset. If the trend of that example is positive i.e. going up, we ***nudge*** the example further in the positive direction. And vice versa - if the trend is negative i.e. going down, we nudge the example further in the negative direction.\n\nWe measure the trend by taking the differences of all subsequent submissions for the specific language and averaging those differences. The nudge that we then give to the new submission is based on a predefined ***weight***, typically we choose a weight of 1 or 1.5. \n\n## Post-processing: pseudo-code\nIn an attempt to pseudo-code the technique, given:\n\n```\nweight = predefined weight (typically 1 or 1.5)\npred_best = current best predictions on LB\ndiff_avg = average of differences of consecutive subs (trend)\n```\n\nThen for each example in test of the specific language (e.g. Turkish):\n\n```\nif diff_avg &lt; 0: # negative trend\n    pred_new = (1+weight*diff_avg)*pred_best  # nudge downwards\nelse: # positive trend\n    pred_new = (1-weight*diff_avg)*pred_best + weight*diff_avg # nudge upwards\n```\n\nNote: I'm not an expert on PP, so I assume this technique can be further optimized. The boost we got from it was relatively small (albeit significant) compared to the other methods we implemented.",
      "votes": null
    },
    {
      "id": "898246",
      "postDate": "06/23/2020 11:54:18",
      "content": "<p><a href=\"/rafiko1\">@rafiko1</a> congrats again and thanks for sharing.</p>",
      "rawMarkdown": "rafiko1 congrats again and thanks for sharing.",
      "votes": null
    },
    {
      "id": "898260",
      "postDate": "06/23/2020 12:03:57",
      "content": "<p>Thank you <a href=\"/sheriytm\">@sheriytm</a>! </p>",
      "rawMarkdown": "Thank you @sheriytm!",
      "votes": null
    },
    {
      "id": "898477",
      "postDate": "06/23/2020 14:43:57",
      "content": "<p>Congratulations on being a master <a href=\"/rafiko1\">@rafiko1</a>!</p>",
      "rawMarkdown": "Congratulations on being a master @rafiko1!",
      "votes": null
    },
    {
      "id": "898536",
      "postDate": "06/23/2020 15:17:08",
      "content": "<p>Nice trick. How much did PP increase your public and private LB?</p>",
      "rawMarkdown": "Nice trick. How much did PP increase your public and private LB?",
      "votes": null
    },
    {
      "id": "898678",
      "postDate": "06/23/2020 17:07:38",
      "content": "<p>Thanks <a href=\"/duykhanh99\">@duykhanh99</a> :)</p>",
      "rawMarkdown": "Thanks @duykhanh99 :)",
      "votes": null
    },
    {
      "id": "898689",
      "postDate": "06/23/2020 17:12:55",
      "content": "<p>Congratulations and thanks for sharing the technique.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing the technique.",
      "votes": null
    },
    {
      "id": "898693",
      "postDate": "06/23/2020 17:16:04",
      "content": "<p>We were able to squeeze another ~0.001 on public (from 0.9546 to 0.9556) corresponding to 0.0006 on private (from 0.9530 to 0.9536). These are final results after using the new predictions as pseudolabels. </p>",
      "rawMarkdown": "We were able to squeeze another ~0.001 on public (from 0.9546 to 0.9556) corresponding to 0.0006 on private (from 0.9530 to 0.9536). These are final results after using the new predictions as pseudolabels.",
      "votes": null
    },
    {
      "id": "898712",
      "postDate": "06/23/2020 17:25:55",
      "content": "<p><a href=\"/rafiko1\">@rafiko1</a> congrats on being master</p>",
      "rawMarkdown": "rafiko1 congrats on being master",
      "votes": null
    },
    {
      "id": "898752",
      "postDate": "06/23/2020 17:56:48",
      "content": "<p>Thank you 🙏 </p>",
      "rawMarkdown": "Thank you 🙏",
      "votes": null
    },
    {
      "id": "899325",
      "postDate": "06/24/2020 06:56:56",
      "content": "<p>Congrats and thanks for sharing the PP technique! May I ask how the averaging was done? Is it like if you got the best prediction <code>sub_best</code> and several former and weaker (lower in LB), let's say <code>sub_a</code>, <code>sub_b</code>, and <code>sub_c</code>, then you calculated the diff between each weaker sub and the best one by data point/row (so in this case, there would be 3 diffs: <code>diff(sub_best, sub_a)</code>, <code>diff(sub_best, sub_b)</code>, and <code>diff(sub_best, sub_c))</code>, and then averaged the 3 diffs to get trend matrix in dim [63812, 1] that you discussed above, and plus since it looks like this PP was done row by row (each text has its own diff_avg) I guess the reason why you mentioned that you did it for a specific language is that you used one model for each language, isn‘t it? please kindly correct me if I am wrong, thanks again! Really loved your solution!!</p>",
      "rawMarkdown": "Congrats and thanks for sharing the PP technique! May I ask how the averaging was done? Is it like if you got the best prediction `sub_best` and several former and weaker (lower in LB), let's say `sub_a`, `sub_b`, and `sub_c`, then you calculated the diff between each weaker sub and the best one by data point/row (so in this case, there would be 3 diffs: `diff(sub_best, sub_a)`, `diff(sub_best, sub_b)`, and `diff(sub_best, sub_c))`, and then averaged the 3 diffs to get trend matrix in dim [63812, 1] that you discussed above, and plus since it looks like this PP was done row by row (each text has its own diff_avg) I guess the reason why you mentioned that you did it for a specific language is that you used one model for each language, isn‘t it? please kindly correct me if I am wrong, thanks again! Really loved your solution!!",
      "votes": null
    },
    {
      "id": "899629",
      "postDate": "06/24/2020 10:43:43",
      "content": "<p>Thanks, and good questions!\nThe difference is done consecutively for each sub improvement on the LB, so not with the best sub:\n<code>For sub_a &amp;lt; sub_b &amp;lt; sub_c, ... we take diff(sub_a, sub_b) and diff(sub_b, sub_c),... and compute the average of those differences</code> \nWe repeat the process for each language individually, mainly to account for our subs that are based on mono-lingual models. So in retrospect, we only change dim [N_test_language, 1] for the new sub, keeping all language other predictions as-is. We got a small improvement for es, tr, ru, fr.</p>\n\n<p>Let me know if you have other questions :)</p>",
      "rawMarkdown": "Thanks, and good questions!\nThe difference is done consecutively for each sub improvement on the LB, so not with the best sub:\n`For sub_a &lt; sub_b &lt; sub_c, ... we take diff(sub_a, sub_b) and diff(sub_b, sub_c),... and compute the average of those differences` \nWe repeat the process for each language individually, mainly to account for our subs that are based on mono-lingual models. So in retrospect, we only change dim [N\\_test\\_language, 1] for the new sub, keeping all language other predictions as-is. We got a small improvement for es, tr, ru, fr.\n\nLet me know if you have other questions :)",
      "votes": null
    },
    {
      "id": "900368",
      "postDate": "06/24/2020 19:10:12",
      "content": "<p>thanks for the detailed answers! Nice practice as a PP for classification problem lol.</p>",
      "rawMarkdown": "thanks for the detailed answers! Nice practice as a PP for classification problem lol.",
      "votes": null
    },
    {
      "id": "900431",
      "postDate": "06/24/2020 19:55:13",
      "content": "<p>Sure thing! Might come in handy :D</p>",
      "rawMarkdown": "Sure thing! Might come in handy :D",
      "votes": null
    },
    {
      "id": "900434",
      "postDate": "06/24/2020 19:57:53",
      "content": "<p>Congrats on becoming # 1 in this competition. Any plans to share a kernel close to the best solution? Really interested in seeing these PP techniques applied in code</p>",
      "rawMarkdown": "Congrats on becoming # 1 in this competition. Any plans to share a kernel close to the best solution? Really interested in seeing these PP techniques applied in code",
      "votes": null
    },
    {
      "id": "900879",
      "postDate": "06/25/2020 05:56:15",
      "content": "<p>Congratulations!!</p>\n\n<p>Thank You</p>",
      "rawMarkdown": "Congratulations!!\n\nThank You",
      "votes": null
    },
    {
      "id": "901020",
      "postDate": "06/25/2020 07:54:58",
      "content": "<p>Thanks! I'll discuss with my partner about what we'll share or not as kernels</p>",
      "rawMarkdown": "Thanks! I'll discuss with my partner about what we'll share or not as kernels",
      "votes": null
    },
    {
      "id": "902154",
      "postDate": "06/26/2020 00:45:29",
      "content": "<p>Great!</p>",
      "rawMarkdown": "Great!",
      "votes": null
    },
    {
      "id": "902360",
      "postDate": "06/26/2020 05:21:52",
      "content": "<p><a href=\"/rafiko1\">@rafiko1</a> am I getting this right that you basically finetuned your predictions on the public leaderboard? and more submissions =&gt; more reliable trend direction</p>",
      "rawMarkdown": "rafiko1 am I getting this right that you basically finetuned your predictions on the public leaderboard? and more submissions =&gt; more reliable trend direction",
      "votes": null
    },
    {
      "id": "902576",
      "postDate": "06/26/2020 08:31:29",
      "content": "<p>That's exactly right. The more submissions you have, the more you can fine-tune it to public LB.\nWorth to mention is that even subs that give a degradation in score (i.e. sub_c &lt; sub_b) can still be incorporated in a similar fashion</p>",
      "rawMarkdown": "That's exactly right. The more submissions you have, the more you can fine-tune it to public LB.\nWorth to mention is that even subs that give a degradation in score (i.e. sub\\_c &lt; sub\\_b) can still be incorporated in a similar fashion",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 898246,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "06/23/2020 11:54:18",
      "content": "<p><a href=\"/rafiko1\">@rafiko1</a> congrats again and thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 898260,
          "author_name": "rafiko1",
          "author_url": "",
          "post_date": "06/23/2020 12:03:57",
          "content": "<p>Thank you <a href=\"/sheriytm\">@sheriytm</a>! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898477,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "06/23/2020 14:43:57",
      "content": "<p>Congratulations on being a master <a href=\"/rafiko1\">@rafiko1</a>!</p>",
      "votes": null,
      "replies": [
        {
          "id": 898678,
          "author_name": "rafiko1",
          "author_url": "",
          "post_date": "06/23/2020 17:07:38",
          "content": "<p>Thanks <a href=\"/duykhanh99\">@duykhanh99</a> :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898536,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/23/2020 15:17:08",
      "content": "<p>Nice trick. How much did PP increase your public and private LB?</p>",
      "votes": null,
      "replies": [
        {
          "id": 898693,
          "author_name": "rafiko1",
          "author_url": "",
          "post_date": "06/23/2020 17:16:04",
          "content": "<p>We were able to squeeze another ~0.001 on public (from 0.9546 to 0.9556) corresponding to 0.0006 on private (from 0.9530 to 0.9536). These are final results after using the new predictions as pseudolabels. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898689,
      "author_name": "deepchatterjeevns",
      "author_url": "",
      "post_date": "06/23/2020 17:12:55",
      "content": "<p>Congratulations and thanks for sharing the technique.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898712,
      "author_name": "shubhamgoel1410",
      "author_url": "",
      "post_date": "06/23/2020 17:25:55",
      "content": "<p><a href=\"/rafiko1\">@rafiko1</a> congrats on being master</p>",
      "votes": null,
      "replies": [
        {
          "id": 898752,
          "author_name": "rafiko1",
          "author_url": "",
          "post_date": "06/23/2020 17:56:48",
          "content": "<p>Thank you 🙏 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 899325,
      "author_name": "",
      "author_url": "",
      "post_date": "06/24/2020 06:56:56",
      "content": "<p>Congrats and thanks for sharing the PP technique! May I ask how the averaging was done? Is it like if you got the best prediction <code>sub_best</code> and several former and weaker (lower in LB), let's say <code>sub_a</code>, <code>sub_b</code>, and <code>sub_c</code>, then you calculated the diff between each weaker sub and the best one by data point/row (so in this case, there would be 3 diffs: <code>diff(sub_best, sub_a)</code>, <code>diff(sub_best, sub_b)</code>, and <code>diff(sub_best, sub_c))</code>, and then averaged the 3 diffs to get trend matrix in dim [63812, 1] that you discussed above, and plus since it looks like this PP was done row by row (each text has its own diff_avg) I guess the reason why you mentioned that you did it for a specific language is that you used one model for each language, isn‘t it? please kindly correct me if I am wrong, thanks again! Really loved your solution!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 899629,
          "author_name": "rafiko1",
          "author_url": "",
          "post_date": "06/24/2020 10:43:43",
          "content": "<p>Thanks, and good questions!\nThe difference is done consecutively for each sub improvement on the LB, so not with the best sub:\n<code>For sub_a &amp;lt; sub_b &amp;lt; sub_c, ... we take diff(sub_a, sub_b) and diff(sub_b, sub_c),... and compute the average of those differences</code> \nWe repeat the process for each language individually, mainly to account for our subs that are based on mono-lingual models. So in retrospect, we only change dim [N_test_language, 1] for the new sub, keeping all language other predictions as-is. We got a small improvement for es, tr, ru, fr.</p>\n\n<p>Let me know if you have other questions :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 900368,
          "author_name": "",
          "author_url": "",
          "post_date": "06/24/2020 19:10:12",
          "content": "<p>thanks for the detailed answers! Nice practice as a PP for classification problem lol.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 900431,
          "author_name": "rafiko1",
          "author_url": "",
          "post_date": "06/24/2020 19:55:13",
          "content": "<p>Sure thing! Might come in handy :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 900434,
      "author_name": "dronych",
      "author_url": "",
      "post_date": "06/24/2020 19:57:53",
      "content": "<p>Congrats on becoming # 1 in this competition. Any plans to share a kernel close to the best solution? Really interested in seeing these PP techniques applied in code</p>",
      "votes": null,
      "replies": [
        {
          "id": 901020,
          "author_name": "rafiko1",
          "author_url": "",
          "post_date": "06/25/2020 07:54:58",
          "content": "<p>Thanks! I'll discuss with my partner about what we'll share or not as kernels</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 900879,
      "author_name": "mdselimreza",
      "author_url": "",
      "post_date": "06/25/2020 05:56:15",
      "content": "<p>Congratulations!!</p>\n\n<p>Thank You</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 902154,
      "author_name": "rashidulhasanhridoy",
      "author_url": "",
      "post_date": "06/26/2020 00:45:29",
      "content": "<p>Great!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 902360,
      "author_name": "dronych",
      "author_url": "",
      "post_date": "06/26/2020 05:21:52",
      "content": "<p><a href=\"/rafiko1\">@rafiko1</a> am I getting this right that you basically finetuned your predictions on the public leaderboard? and more submissions =&gt; more reliable trend direction</p>",
      "votes": null,
      "replies": [
        {
          "id": 902576,
          "author_name": "rafiko1",
          "author_url": "",
          "post_date": "06/26/2020 08:31:29",
          "content": "<p>That's exactly right. The more submissions you have, the more you can fine-tune it to public LB.\nWorth to mention is that even subs that give a degradation in score (i.e. sub_c &lt; sub_b) can still be incorporated in a similar fashion</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "898183": "Firstly, a big thanks to Kaggle for constantly delivering top-notch competitions. Competitions doesn't always end smooth (cf. Deepfakes 😄) but there is no other data science platform so rich in learning. Another thanks to Jigsaw for the interesting, multi-lingual, shakeup-free ride!  \nAlso congrats to the other medalists, and to my talented team-mate @leecming!\n\nSince other competitors were asking about our post-processing technique, I dedicate this post to explain it in more detail. \n\nTo see the post-processing in action, have a look at the related [notebook](https://www.kaggle.com/rafiko1/1st-place-jigsaw-post-processing-example)\n## Post-processing: intuition\n\nThe intuition of the post-processing is as following: we consider the ***trend*** of subsequent submissions of a specific language (e.g. Russian) for each example in the test dataset. If the trend of that example is positive i.e. going up, we ***nudge*** the example further in the positive direction. And vice versa - if the trend is negative i.e. going down, we nudge the example further in the negative direction.\n\nWe measure the trend by taking the differences of all subsequent submissions for the specific language and averaging those differences. The nudge that we then give to the new submission is based on a predefined ***weight***, typically we choose a weight of 1 or 1.5. \n\n## Post-processing: pseudo-code\nIn an attempt to pseudo-code the technique, given:\n\n```\nweight = predefined weight (typically 1 or 1.5)\npred_best = current best predictions on LB\ndiff_avg = average of differences of consecutive subs (trend)\n```\n\nThen for each example in test of the specific language (e.g. Turkish):\n\n```\nif diff_avg &lt; 0: # negative trend\n    pred_new = (1+weight*diff_avg)*pred_best  # nudge downwards\nelse: # positive trend\n    pred_new = (1-weight*diff_avg)*pred_best + weight*diff_avg # nudge upwards\n```\n\nNote: I'm not an expert on PP, so I assume this technique can be further optimized. The boost we got from it was relatively small (albeit significant) compared to the other methods we implemented.",
    "898246": "rafiko1 congrats again and thanks for sharing.",
    "898260": "Thank you @sheriytm!",
    "898477": "Congratulations on being a master @rafiko1!",
    "898536": "Nice trick. How much did PP increase your public and private LB?",
    "898678": "Thanks @duykhanh99 :)",
    "898689": "Congratulations and thanks for sharing the technique.",
    "898693": "We were able to squeeze another ~0.001 on public (from 0.9546 to 0.9556) corresponding to 0.0006 on private (from 0.9530 to 0.9536). These are final results after using the new predictions as pseudolabels.",
    "898712": "rafiko1 congrats on being master",
    "898752": "Thank you 🙏",
    "899325": "Congrats and thanks for sharing the PP technique! May I ask how the averaging was done? Is it like if you got the best prediction `sub_best` and several former and weaker (lower in LB), let's say `sub_a`, `sub_b`, and `sub_c`, then you calculated the diff between each weaker sub and the best one by data point/row (so in this case, there would be 3 diffs: `diff(sub_best, sub_a)`, `diff(sub_best, sub_b)`, and `diff(sub_best, sub_c))`, and then averaged the 3 diffs to get trend matrix in dim [63812, 1] that you discussed above, and plus since it looks like this PP was done row by row (each text has its own diff_avg) I guess the reason why you mentioned that you did it for a specific language is that you used one model for each language, isn‘t it? please kindly correct me if I am wrong, thanks again! Really loved your solution!!",
    "899629": "Thanks, and good questions!\nThe difference is done consecutively for each sub improvement on the LB, so not with the best sub:\n`For sub_a &lt; sub_b &lt; sub_c, ... we take diff(sub_a, sub_b) and diff(sub_b, sub_c),... and compute the average of those differences` \nWe repeat the process for each language individually, mainly to account for our subs that are based on mono-lingual models. So in retrospect, we only change dim [N\\_test\\_language, 1] for the new sub, keeping all language other predictions as-is. We got a small improvement for es, tr, ru, fr.\n\nLet me know if you have other questions :)",
    "900368": "thanks for the detailed answers! Nice practice as a PP for classification problem lol.",
    "900431": "Sure thing! Might come in handy :D",
    "900434": "Congrats on becoming # 1 in this competition. Any plans to share a kernel close to the best solution? Really interested in seeing these PP techniques applied in code",
    "900879": "Congratulations!!\n\nThank You",
    "901020": "Thanks! I'll discuss with my partner about what we'll share or not as kernels",
    "902154": "Great!",
    "902360": "rafiko1 am I getting this right that you basically finetuned your predictions on the public leaderboard? and more submissions =&gt; more reliable trend direction",
    "902576": "That's exactly right. The more submissions you have, the more you can fine-tune it to public LB.\nWorth to mention is that even subs that give a degradation in score (i.e. sub\\_c &lt; sub\\_b) can still be incorporated in a similar fashion"
  },
  "source": "meta"
}