{
  "id": 161974,
  "title": "13th Place Solution Part I. My first gold medal ",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/seed-42-13th-place-solution-part-i-my-first-gold-m",
  "author_name": "",
  "post_date": "2020-06-26T22:27:21.337Z",
  "votes": 17,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First of all, a huge thank you to my teammates <a href=\"/euclidean\">@euclidean</a> <a href=\"/nullrecurrent\">@nullrecurrent</a> <a href=\"/luohongchen1993\">@luohongchen1993</a> <a href=\"/sherryli94\">@sherryli94</a>, you guys are so creative and diligent, and I learned tremendously from all of you.  This gold medal is the first gold for all of us, and I feel so honored to share this experience with you. Cheers! </p>\n\n<p>And congratulations to all medal winners :)</p>\n\n<p>Most of the top winners have already covered many of our \"what worked/what didn't work\", I want to use this post to record some points that I think we did differently.  </p>\n\n<h3>Pipeline</h3>\n\n<p>Our final pipeline is:\n1) train on a selective part of training set\n2) train on another part of training set + PL on test set\n3) 5cv on validation + albumentation of validation (thanks to <a href=\"/shonenkov\">@shonenkov</a> 's kernel) </p>\n\n<p>In 5cv, using soft label (validation prediction + validation true label)/2 improves 5cv. Some very controversial labels become ~0.5 which is intuitive to me. </p>\n\n<p>Also, randomize the order of samples in 5cv while keeping the same comments (original and its augmentation) in the same batch was helping to add additional variability. Our number of augmentation vs original comments  was 1:1. More augmentation didn't help on LB for us. </p>\n\n<h3>Training Diversity</h3>\n\n<p>We added many variability in training using LR decay, Cosine annealing, SWA, class weights, sample weights (on augmentation/language), use pre-trained models, freeze embedding layers/encoding layers. Except for XLMR, we also tried LSTM, etc. These model may not work well alone, but the diversity helped in our stacking phase. </p>\n\n<h3>Stacking</h3>\n\n<p>We ended up with 1900+ first level features in stacking, which is quite amazing. We eventually passed 2 StackNet models and 1 forward selection model (all with lb 9494) to post processing phase. </p>\n\n<p>Initially greedy based forward selection ensemble overfits validation auc massively. But adding random initialization, and adding logloss criteria along with auc criteria helped to get a well balanced blending. </p>\n\n<h3>Post Processing</h3>\n\n<p>We scaled up the predicted probability of comments that contain cursed words of different languages. That improved about 3 bps for each stacking model, and rank mean of them finally sent us to the edge of gold zone. </p>\n\n<h3>Things that didn't work for me</h3>\n\n<ul>\n<li>Using PCA of embedding to construct similar training set compared to validation and test set. </li>\n<li>PCA selected train set to substitute validation set. </li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F834077%2Fb41ed87175b5b6ac45a779a7afc40ab1%2Fttagoutou.jpg?generation=1593209576045814&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "903496",
      "postDate": "06/26/2020 22:00:01",
      "content": "<p>First of all, a huge thank you to my teammates <a href=\"/euclidean\">@euclidean</a> <a href=\"/nullrecurrent\">@nullrecurrent</a> <a href=\"/luohongchen1993\">@luohongchen1993</a> <a href=\"/sherryli94\">@sherryli94</a>, you guys are so creative and diligent, and I learned tremendously from all of you.  This gold medal is the first gold for all of us, and I feel so honored to share this experience with you. Cheers! </p>\n\n<p>And congratulations to all medal winners :)</p>\n\n<p>Most of the top winners have already covered many of our \"what worked/what didn't work\", I want to use this post to record some points that I think we did differently.  </p>\n\n<h3>Pipeline</h3>\n\n<p>Our final pipeline is:\n1) train on a selective part of training set\n2) train on another part of training set + PL on test set\n3) 5cv on validation + albumentation of validation (thanks to <a href=\"/shonenkov\">@shonenkov</a> 's kernel) </p>\n\n<p>In 5cv, using soft label (validation prediction + validation true label)/2 improves 5cv. Some very controversial labels become ~0.5 which is intuitive to me. </p>\n\n<p>Also, randomize the order of samples in 5cv while keeping the same comments (original and its augmentation) in the same batch was helping to add additional variability. Our number of augmentation vs original comments  was 1:1. More augmentation didn't help on LB for us. </p>\n\n<h3>Training Diversity</h3>\n\n<p>We added many variability in training using LR decay, Cosine annealing, SWA, class weights, sample weights (on augmentation/language), use pre-trained models, freeze embedding layers/encoding layers. Except for XLMR, we also tried LSTM, etc. These model may not work well alone, but the diversity helped in our stacking phase. </p>\n\n<h3>Stacking</h3>\n\n<p>We ended up with 1900+ first level features in stacking, which is quite amazing. We eventually passed 2 StackNet models and 1 forward selection model (all with lb 9494) to post processing phase. </p>\n\n<p>Initially greedy based forward selection ensemble overfits validation auc massively. But adding random initialization, and adding logloss criteria along with auc criteria helped to get a well balanced blending. </p>\n\n<h3>Post Processing</h3>\n\n<p>We scaled up the predicted probability of comments that contain cursed words of different languages. That improved about 3 bps for each stacking model, and rank mean of them finally sent us to the edge of gold zone. </p>\n\n<h3>Things that didn't work for me</h3>\n\n<ul>\n<li>Using PCA of embedding to construct similar training set compared to validation and test set. </li>\n<li>PCA selected train set to substitute validation set. </li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F834077%2Fb41ed87175b5b6ac45a779a7afc40ab1%2Fttagoutou.jpg?generation=1593209576045814&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "First of all, a huge thank you to my teammates @euclidean @nullrecurrent @luohongchen1993 @sherryli94, you guys are so creative and diligent, and I learned tremendously from all of you.  This gold medal is the first gold for all of us, and I feel so honored to share this experience with you. Cheers! \n\nAnd congratulations to all medal winners :)\n\nMost of the top winners have already covered many of our \"what worked/what didn't work\", I want to use this post to record some points that I think we did differently.  \n\n### Pipeline\nOur final pipeline is:\n1) train on a selective part of training set\n2) train on another part of training set + PL on test set\n3) 5cv on validation + albumentation of validation (thanks to @shonenkov 's kernel) \n\nIn 5cv, using soft label (validation prediction + validation true label)/2 improves 5cv. Some very controversial labels become ~0.5 which is intuitive to me. \n\nAlso, randomize the order of samples in 5cv while keeping the same comments (original and its augmentation) in the same batch was helping to add additional variability. Our number of augmentation vs original comments  was 1:1. More augmentation didn't help on LB for us. \n\n### Training Diversity\nWe added many variability in training using LR decay, Cosine annealing, SWA, class weights, sample weights (on augmentation/language), use pre-trained models, freeze embedding layers/encoding layers. Except for XLMR, we also tried LSTM, etc. These model may not work well alone, but the diversity helped in our stacking phase. \n\n### Stacking\nWe ended up with 1900+ first level features in stacking, which is quite amazing. We eventually passed 2 StackNet models and 1 forward selection model (all with lb 9494) to post processing phase. \n\nInitially greedy based forward selection ensemble overfits validation auc massively. But adding random initialization, and adding logloss criteria along with auc criteria helped to get a well balanced blending. \n\n### Post Processing\nWe scaled up the predicted probability of comments that contain cursed words of different languages. That improved about 3 bps for each stacking model, and rank mean of them finally sent us to the edge of gold zone. \n\n### Things that didn't work for me\n- Using PCA of embedding to construct similar training set compared to validation and test set. \n- PCA selected train set to substitute validation set. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F834077%2Fb41ed87175b5b6ac45a779a7afc40ab1%2Fttagoutou.jpg?generation=1593209576045814&amp;alt=media)",
      "votes": null
    },
    {
      "id": "903514",
      "postDate": "06/26/2020 22:43:00",
      "content": "<p>Amazing work, well done!</p>",
      "rawMarkdown": "Amazing work, well done!",
      "votes": null
    },
    {
      "id": "903571",
      "postDate": "06/27/2020 00:43:25",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "903620",
      "postDate": "06/27/2020 01:59:38",
      "content": "<p>Huge congrats bro, well deserved 💯 !</p>",
      "rawMarkdown": "Huge congrats bro, well deserved 💯 !",
      "votes": null
    },
    {
      "id": "903624",
      "postDate": "06/27/2020 02:05:13",
      "content": "<p>嘿嘿多谢老大哥😊 </p>",
      "rawMarkdown": "嘿嘿多谢老大哥😊",
      "votes": null
    },
    {
      "id": "903830",
      "postDate": "06/27/2020 06:23:22",
      "content": "<p>congrats!</p>",
      "rawMarkdown": "congrats!",
      "votes": null
    },
    {
      "id": "903831",
      "postDate": "06/27/2020 06:26:55",
      "content": "<p>Kudos! Very well deserved! I was loosely following the updates and it seems like the post-processing stage was, as a matter of fact, quite crucial, wasn't it?</p>",
      "rawMarkdown": "Kudos! Very well deserved! I was loosely following the updates and it seems like the post-processing stage was, as a matter of fact, quite crucial, wasn't it?",
      "votes": null
    },
    {
      "id": "904625",
      "postDate": "06/27/2020 19:15:52",
      "content": "<p>Thank you Piyush! Yes as a matter of fact, we wouldn't have reached the gold range without PP, although the improvement was only 3-4 bps compared to some top teams' 10+/20+ bps. It was crucial to get to the top. </p>",
      "rawMarkdown": "Thank you Piyush! Yes as a matter of fact, we wouldn't have reached the gold range without PP, although the improvement was only 3-4 bps compared to some top teams' 10+/20+ bps. It was crucial to get to the top.",
      "votes": null
    },
    {
      "id": "904694",
      "postDate": "06/27/2020 20:20:55",
      "content": "<p>That just goes on to make it even more exciting imo. Congratulations!</p>",
      "rawMarkdown": "That just goes on to make it even more exciting imo. Congratulations!",
      "votes": null
    },
    {
      "id": "924296",
      "postDate": "07/11/2020 11:09:40",
      "content": "<p>Congrats for your first gold!\nSorry about late asking and my dumb question.</p>\n\n<blockquote>\n  <p>1 forward selection model (all with lb 9494)\n  Initially greedy based forward selection ensemble overfits validation auc massively.</p>\n</blockquote>\n\n<p>What is <code>forward selection model</code> and <code>greedy based forward selection ensemble</code>?</p>",
      "rawMarkdown": "Congrats for your first gold!\nSorry about late asking and my dumb question.\n\n&gt; 1 forward selection model (all with lb 9494)\n&gt; Initially greedy based forward selection ensemble overfits validation auc massively.\n\nWhat is `forward selection model` and `greedy based forward selection ensemble`?",
      "votes": null
    },
    {
      "id": "987763",
      "postDate": "08/27/2020 13:52:18",
      "content": "<p>Sorry about the late reply! There is no such question as dumb :) <br>\nForward selection basically is the method that first choose one model with the highest cv (in this case validation auc) as the ensemble set, and greedily/iteratively adding more models that increase the cv score the most to the current ensemble set, one model at a time. </p>",
      "rawMarkdown": "Sorry about the late reply! There is no such question as dumb :) \nForward selection basically is the method that first choose one model with the highest cv (in this case validation auc) as the ensemble set, and greedily/iteratively adding more models that increase the cv score the most to the current ensemble set, one model at a time.",
      "votes": null
    },
    {
      "id": "988039",
      "postDate": "08/27/2020 18:09:40",
      "content": "<p>Thanks for replying! </p>",
      "rawMarkdown": "Thanks for replying!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 903514,
      "author_name": "eward96",
      "author_url": "",
      "post_date": "06/26/2020 22:43:00",
      "content": "<p>Amazing work, well done!</p>",
      "votes": null,
      "replies": [
        {
          "id": 903571,
          "author_name": "strider1125",
          "author_url": "",
          "post_date": "06/27/2020 00:43:25",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 903620,
      "author_name": "benhzn07",
      "author_url": "",
      "post_date": "06/27/2020 01:59:38",
      "content": "<p>Huge congrats bro, well deserved 💯 !</p>",
      "votes": null,
      "replies": [
        {
          "id": 903624,
          "author_name": "strider1125",
          "author_url": "",
          "post_date": "06/27/2020 02:05:13",
          "content": "<p>嘿嘿多谢老大哥😊 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 903830,
      "author_name": "machinelp",
      "author_url": "",
      "post_date": "06/27/2020 06:23:22",
      "content": "<p>congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 903831,
      "author_name": "piyushmishra1999",
      "author_url": "",
      "post_date": "06/27/2020 06:26:55",
      "content": "<p>Kudos! Very well deserved! I was loosely following the updates and it seems like the post-processing stage was, as a matter of fact, quite crucial, wasn't it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 904625,
          "author_name": "strider1125",
          "author_url": "",
          "post_date": "06/27/2020 19:15:52",
          "content": "<p>Thank you Piyush! Yes as a matter of fact, we wouldn't have reached the gold range without PP, although the improvement was only 3-4 bps compared to some top teams' 10+/20+ bps. It was crucial to get to the top. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 904694,
          "author_name": "piyushmishra1999",
          "author_url": "",
          "post_date": "06/27/2020 20:20:55",
          "content": "<p>That just goes on to make it even more exciting imo. Congratulations!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 924296,
      "author_name": "karlyukang",
      "author_url": "",
      "post_date": "07/11/2020 11:09:40",
      "content": "<p>Congrats for your first gold!\nSorry about late asking and my dumb question.</p>\n\n<blockquote>\n  <p>1 forward selection model (all with lb 9494)\n  Initially greedy based forward selection ensemble overfits validation auc massively.</p>\n</blockquote>\n\n<p>What is <code>forward selection model</code> and <code>greedy based forward selection ensemble</code>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 987763,
          "author_name": "strider1125",
          "author_url": "",
          "post_date": "08/27/2020 13:52:18",
          "content": "<p>Sorry about the late reply! There is no such question as dumb :) <br>\nForward selection basically is the method that first choose one model with the highest cv (in this case validation auc) as the ensemble set, and greedily/iteratively adding more models that increase the cv score the most to the current ensemble set, one model at a time. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 988039,
          "author_name": "karlyukang",
          "author_url": "",
          "post_date": "08/27/2020 18:09:40",
          "content": "<p>Thanks for replying! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "903496": "First of all, a huge thank you to my teammates @euclidean @nullrecurrent @luohongchen1993 @sherryli94, you guys are so creative and diligent, and I learned tremendously from all of you.  This gold medal is the first gold for all of us, and I feel so honored to share this experience with you. Cheers! \n\nAnd congratulations to all medal winners :)\n\nMost of the top winners have already covered many of our \"what worked/what didn't work\", I want to use this post to record some points that I think we did differently.  \n\n### Pipeline\nOur final pipeline is:\n1) train on a selective part of training set\n2) train on another part of training set + PL on test set\n3) 5cv on validation + albumentation of validation (thanks to @shonenkov 's kernel) \n\nIn 5cv, using soft label (validation prediction + validation true label)/2 improves 5cv. Some very controversial labels become ~0.5 which is intuitive to me. \n\nAlso, randomize the order of samples in 5cv while keeping the same comments (original and its augmentation) in the same batch was helping to add additional variability. Our number of augmentation vs original comments  was 1:1. More augmentation didn't help on LB for us. \n\n### Training Diversity\nWe added many variability in training using LR decay, Cosine annealing, SWA, class weights, sample weights (on augmentation/language), use pre-trained models, freeze embedding layers/encoding layers. Except for XLMR, we also tried LSTM, etc. These model may not work well alone, but the diversity helped in our stacking phase. \n\n### Stacking\nWe ended up with 1900+ first level features in stacking, which is quite amazing. We eventually passed 2 StackNet models and 1 forward selection model (all with lb 9494) to post processing phase. \n\nInitially greedy based forward selection ensemble overfits validation auc massively. But adding random initialization, and adding logloss criteria along with auc criteria helped to get a well balanced blending. \n\n### Post Processing\nWe scaled up the predicted probability of comments that contain cursed words of different languages. That improved about 3 bps for each stacking model, and rank mean of them finally sent us to the edge of gold zone. \n\n### Things that didn't work for me\n- Using PCA of embedding to construct similar training set compared to validation and test set. \n- PCA selected train set to substitute validation set. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F834077%2Fb41ed87175b5b6ac45a779a7afc40ab1%2Fttagoutou.jpg?generation=1593209576045814&amp;alt=media)",
    "903514": "Amazing work, well done!",
    "903571": "Thank you!",
    "903620": "Huge congrats bro, well deserved 💯 !",
    "903624": "嘿嘿多谢老大哥😊",
    "903830": "congrats!",
    "903831": "Kudos! Very well deserved! I was loosely following the updates and it seems like the post-processing stage was, as a matter of fact, quite crucial, wasn't it?",
    "904625": "Thank you Piyush! Yes as a matter of fact, we wouldn't have reached the gold range without PP, although the improvement was only 3-4 bps compared to some top teams' 10+/20+ bps. It was crucial to get to the top.",
    "904694": "That just goes on to make it even more exciting imo. Congratulations!",
    "924296": "Congrats for your first gold!\nSorry about late asking and my dumb question.\n\n&gt; 1 forward selection model (all with lb 9494)\n&gt; Initially greedy based forward selection ensemble overfits validation auc massively.\n\nWhat is `forward selection model` and `greedy based forward selection ensemble`?",
    "987763": "Sorry about the late reply! There is no such question as dumb :) \nForward selection basically is the method that first choose one model with the highest cv (in this case validation auc) as the ensemble set, and greedily/iteratively adding more models that increase the cv score the most to the current ensemble set, one model at a time.",
    "988039": "Thanks for replying!"
  },
  "source": "meta"
}