{
  "id": 127266,
  "title": "47th Place Solution Write-Up (Ensembling)",
  "url": "/competitions/tensorflow2-question-answering/discussion/127266",
  "author_name": "Ram Ramrakhya",
  "post_date": "2020-01-23T05:45:45.540000",
  "votes": 17,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Thank you Kaggle and Kaggle community for this awesome competition. I learned a lot.\nWe tried a lot of new things with pytorch in the last week but it weren't able to get things working, but it has been really fun.</p>\n\n<h2><strong>Our Solution</strong></h2>\n\n<p>Our current solution is a TF 2.0 solution based on this great kernel <a href=\"https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models\">https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models</a> by <a href=\"/yihdarshieh\">@yihdarshieh</a>. Initially I started off with  finetuning the official bert joint baseline but it didn't give much improvements. In next few weeks I completed the TPU setup for the contest and had my pipeline ready for training and validation on dev set.</p>\n\n<p>All the experiments were done in Google colabs free tier TPUs. This is my first time seriously using TPUs and I have to say, it feels so good. Because they are so fast. Using TPUs dramatically reduced experimentation time in my case.</p>\n\n<p>Our solution is a simple ensemble of following 2 models:</p>\n\n<ol>\n<li>BERT-joint-large public lb score 0.6</li>\n<li>BERT-joint-base public lb score 0.58</li>\n<li>DistillBERT-joint public lb score 0.54</li>\n<li>Our final solution is an ensemble of 1 and 2 with some additions to postprocessing which scores 0.64 on public lb (also scored 0.64 on private LB but we didn't choose our best solution for final submission as it scored less on public LB with the new postprocessing)</li>\n</ol>\n\n<p>Here's our ensembling code, we use simple weighted average ensembling.</p>\n\n<p>```\n        nq_logits = bert_nq(nq_inputs, training=False)\n        base_nq_logits = base_bert_nq(nq_inputs, training=False)</p>\n\n<pre><code>    (start_pos_logits, end_pos_logits, answer_type_logits) = nq_logits\n    (base_start_pos_logits, base_end_pos_logits, baseanswer_type_logits) = base_nq_logits\n\n    start_pos_logits = (0.2 * start_pos_logits + 0.8 * base_start_pos_logits)\n    end_pos_logits = (0.2 * end_pos_logits + 0.8 * base_end_pos_logits)\n    answer_type_logits = (0.2 * answer_type_logits + 0.8 * baseanswer_type_logits)\n</code></pre>\n\n<p>```</p>\n\n<h2><strong>Postprocessing</strong></h2>\n\n<p>After experimenting with multiple single models I started focusing on postprocessing to improve on model performance. Initially I used postprocessing provided by this great kernel <a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook\">https://www.kaggle.com/prokaj/bert-joint-baseline-notebook</a> by @prvi which helped my single models score in range 0.56-0.58. To improve further on this I had an in-depth look at the predictions of the model and ground truths. Here I found that our model was predicting duplicate answer spans. So I added a duplicate removal logic to postprocessing which helped a score increase on 0.01 on public LB and 0.02 on dev set. I also observed a score improvement if my model doesn't predict any \"YES/NO\" answer so essentially my model was only outputting answer spans and null answers in my final solution.</p>\n\n<p>Deciding on thresholds was one of the important things to predict valid answers but I didn't play too much with answer thresholds. Initially I ran inference on validation set with 5 different answer thresholds [1.5, 3.0, 4.5, 6.0, 7.5] and saw the best validation score with a combination of 1.5 for long answer and 3.0 for short answer which I kept same for final solution.</p>\n\n<h2><strong>New Addition to postprocessing</strong></h2>\n\n<p>Postprocessing from @prvi's kernel only looks at the current <code>512</code> sequence to choose start and end indexes but as we are using strides of length <code>128</code> there's a overlap in every 2 consecutive sequences. So I decided to append top <code>k</code> start and end indexes from every 2 consecutive sequences to choose a pair of start and end indexes. In this case for overlapping start and end indexes I got 4 possible score values out of which I kept only the max score value and discarded the rest 3. This postprocessing didn't give us improvement on public LB for our current best model but it helped by a increment of 0.17 for a weak model.</p>\n\n<p>Our final solution didn't use the new postprocessing but our best solution on private set scored 0.64 (a 0.01 increment) with new postprocessing.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F887695%2Fae6bdb038215e5da57e9b68cc6de5242%2Fscore.png?generation=1579758032774744&amp;alt=media\" alt=\"\"></p>\n\n<h2>** Last Week of Competition**</h2>\n\n<p>In the last week of the competition I got an opportunity to team up with <a href=\"/abhishek\">@abhishek</a> and <a href=\"/rinnqd\">@rinnqd</a>. After working on BERT-large and BERT-base we wanted to try out ALBERT and RoBERTa in last few days so we started working on pytorch for these two models. Our GPU training pipeline was completed by <a href=\"/abhishek\">@abhishek</a> in just few hours. We tried to port it to TPU in next few days but weren't able to get it working in time. It was a great experience teaming up with these guys I learned a lot about how to easily prototype your training pipeline. How to start off with validation pipeline and how important it is. Learnt about how to write TPU code for pytorch. </p>\n\n<p>Thanks guys for teaming up. And thanks Kaggle for such a great contest.</p>\n\n<p>Lastly, here's our final solution kernel <a href=\"https://www.kaggle.com/axel81/inference-use-hugging-face-postprocess\">https://www.kaggle.com/axel81/inference-use-hugging-face-postprocess</a>.</p>\n\n<p>Happy Kaggling :)</p>",
  "messages": [
    {
      "id": 726599,
      "postDate": "2020-01-23T05:45:45.540Z",
      "content": "<p>Thank you Kaggle and Kaggle community for this awesome competition. I learned a lot.\nWe tried a lot of new things with pytorch in the last week but it weren't able to get things working, but it has been really fun.</p>\n\n<h2><strong>Our Solution</strong></h2>\n\n<p>Our current solution is a TF 2.0 solution based on this great kernel <a href=\"https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models\">https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models</a> by <a href=\"/yihdarshieh\">@yihdarshieh</a>. Initially I started off with  finetuning the official bert joint baseline but it didn't give much improvements. In next few weeks I completed the TPU setup for the contest and had my pipeline ready for training and validation on dev set.</p>\n\n<p>All the experiments were done in Google colabs free tier TPUs. This is my first time seriously using TPUs and I have to say, it feels so good. Because they are so fast. Using TPUs dramatically reduced experimentation time in my case.</p>\n\n<p>Our solution is a simple ensemble of following 2 models:</p>\n\n<ol>\n<li>BERT-joint-large public lb score 0.6</li>\n<li>BERT-joint-base public lb score 0.58</li>\n<li>DistillBERT-joint public lb score 0.54</li>\n<li>Our final solution is an ensemble of 1 and 2 with some additions to postprocessing which scores 0.64 on public lb (also scored 0.64 on private LB but we didn't choose our best solution for final submission as it scored less on public LB with the new postprocessing)</li>\n</ol>\n\n<p>Here's our ensembling code, we use simple weighted average ensembling.</p>\n\n<p>```\n        nq_logits = bert_nq(nq_inputs, training=False)\n        base_nq_logits = base_bert_nq(nq_inputs, training=False)</p>\n\n<pre><code>    (start_pos_logits, end_pos_logits, answer_type_logits) = nq_logits\n    (base_start_pos_logits, base_end_pos_logits, baseanswer_type_logits) = base_nq_logits\n\n    start_pos_logits = (0.2 * start_pos_logits + 0.8 * base_start_pos_logits)\n    end_pos_logits = (0.2 * end_pos_logits + 0.8 * base_end_pos_logits)\n    answer_type_logits = (0.2 * answer_type_logits + 0.8 * baseanswer_type_logits)\n</code></pre>\n\n<p>```</p>\n\n<h2><strong>Postprocessing</strong></h2>\n\n<p>After experimenting with multiple single models I started focusing on postprocessing to improve on model performance. Initially I used postprocessing provided by this great kernel <a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook\">https://www.kaggle.com/prokaj/bert-joint-baseline-notebook</a> by @prvi which helped my single models score in range 0.56-0.58. To improve further on this I had an in-depth look at the predictions of the model and ground truths. Here I found that our model was predicting duplicate answer spans. So I added a duplicate removal logic to postprocessing which helped a score increase on 0.01 on public LB and 0.02 on dev set. I also observed a score improvement if my model doesn't predict any \"YES/NO\" answer so essentially my model was only outputting answer spans and null answers in my final solution.</p>\n\n<p>Deciding on thresholds was one of the important things to predict valid answers but I didn't play too much with answer thresholds. Initially I ran inference on validation set with 5 different answer thresholds [1.5, 3.0, 4.5, 6.0, 7.5] and saw the best validation score with a combination of 1.5 for long answer and 3.0 for short answer which I kept same for final solution.</p>\n\n<h2><strong>New Addition to postprocessing</strong></h2>\n\n<p>Postprocessing from @prvi's kernel only looks at the current <code>512</code> sequence to choose start and end indexes but as we are using strides of length <code>128</code> there's a overlap in every 2 consecutive sequences. So I decided to append top <code>k</code> start and end indexes from every 2 consecutive sequences to choose a pair of start and end indexes. In this case for overlapping start and end indexes I got 4 possible score values out of which I kept only the max score value and discarded the rest 3. This postprocessing didn't give us improvement on public LB for our current best model but it helped by a increment of 0.17 for a weak model.</p>\n\n<p>Our final solution didn't use the new postprocessing but our best solution on private set scored 0.64 (a 0.01 increment) with new postprocessing.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F887695%2Fae6bdb038215e5da57e9b68cc6de5242%2Fscore.png?generation=1579758032774744&amp;alt=media\" alt=\"\"></p>\n\n<h2>** Last Week of Competition**</h2>\n\n<p>In the last week of the competition I got an opportunity to team up with <a href=\"/abhishek\">@abhishek</a> and <a href=\"/rinnqd\">@rinnqd</a>. After working on BERT-large and BERT-base we wanted to try out ALBERT and RoBERTa in last few days so we started working on pytorch for these two models. Our GPU training pipeline was completed by <a href=\"/abhishek\">@abhishek</a> in just few hours. We tried to port it to TPU in next few days but weren't able to get it working in time. It was a great experience teaming up with these guys I learned a lot about how to easily prototype your training pipeline. How to start off with validation pipeline and how important it is. Learnt about how to write TPU code for pytorch. </p>\n\n<p>Thanks guys for teaming up. And thanks Kaggle for such a great contest.</p>\n\n<p>Lastly, here's our final solution kernel <a href=\"https://www.kaggle.com/axel81/inference-use-hugging-face-postprocess\">https://www.kaggle.com/axel81/inference-use-hugging-face-postprocess</a>.</p>\n\n<p>Happy Kaggling :)</p>",
      "rawMarkdown": "Thank you Kaggle and Kaggle community for this awesome competition. I learned a lot.\nWe tried a lot of new things with pytorch in the last week but it weren't able to get things working, but it has been really fun.\n\n## **Our Solution**\n\nOur current solution is a TF 2.0 solution based on this great kernel https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models by @yihdarshieh. Initially I started off with  finetuning the official bert joint baseline but it didn't give much improvements. In next few weeks I completed the TPU setup for the contest and had my pipeline ready for training and validation on dev set.\n\nAll the experiments were done in Google colabs free tier TPUs. This is my first time seriously using TPUs and I have to say, it feels so good. Because they are so fast. Using TPUs dramatically reduced experimentation time in my case.\n\nOur solution is a simple ensemble of following 2 models:\n\n1. BERT-joint-large public lb score 0.6\n2. BERT-joint-base public lb score 0.58\n3. DistillBERT-joint public lb score 0.54\n3. Our final solution is an ensemble of 1 and 2 with some additions to postprocessing which scores 0.64 on public lb (also scored 0.64 on private LB but we didn't choose our best solution for final submission as it scored less on public LB with the new postprocessing)\n\nHere's our ensembling code, we use simple weighted average ensembling.\n\n```\n        nq_logits = bert_nq(nq_inputs, training=False)\n        base_nq_logits = base_bert_nq(nq_inputs, training=False)\n\n        (start_pos_logits, end_pos_logits, answer_type_logits) = nq_logits\n        (base_start_pos_logits, base_end_pos_logits, baseanswer_type_logits) = base_nq_logits\n        \n        start_pos_logits = (0.2 * start_pos_logits + 0.8 * base_start_pos_logits)\n        end_pos_logits = (0.2 * end_pos_logits + 0.8 * base_end_pos_logits)\n        answer_type_logits = (0.2 * answer_type_logits + 0.8 * baseanswer_type_logits)\n```\n\n## **Postprocessing**\n\nAfter experimenting with multiple single models I started focusing on postprocessing to improve on model performance. Initially I used postprocessing provided by this great kernel https://www.kaggle.com/prokaj/bert-joint-baseline-notebook by @prvi which helped my single models score in range 0.56-0.58. To improve further on this I had an in-depth look at the predictions of the model and ground truths. Here I found that our model was predicting duplicate answer spans. So I added a duplicate removal logic to postprocessing which helped a score increase on 0.01 on public LB and 0.02 on dev set. I also observed a score improvement if my model doesn't predict any \"YES/NO\" answer so essentially my model was only outputting answer spans and null answers in my final solution.\n\nDeciding on thresholds was one of the important things to predict valid answers but I didn't play too much with answer thresholds. Initially I ran inference on validation set with 5 different answer thresholds [1.5, 3.0, 4.5, 6.0, 7.5] and saw the best validation score with a combination of 1.5 for long answer and 3.0 for short answer which I kept same for final solution.\n\n## **New Addition to postprocessing**\n\nPostprocessing from @prvi's kernel only looks at the current `512` sequence to choose start and end indexes but as we are using strides of length `128` there's a overlap in every 2 consecutive sequences. So I decided to append top `k` start and end indexes from every 2 consecutive sequences to choose a pair of start and end indexes. In this case for overlapping start and end indexes I got 4 possible score values out of which I kept only the max score value and discarded the rest 3. This postprocessing didn't give us improvement on public LB for our current best model but it helped by a increment of 0.17 for a weak model.\n\nOur final solution didn't use the new postprocessing but our best solution on private set scored 0.64 (a 0.01 increment) with new postprocessing.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F887695%2Fae6bdb038215e5da57e9b68cc6de5242%2Fscore.png?generation=1579758032774744&amp;alt=media)\n\n## ** Last Week of Competition**\n\nIn the last week of the competition I got an opportunity to team up with @abhishek and @rinnqd. After working on BERT-large and BERT-base we wanted to try out ALBERT and RoBERTa in last few days so we started working on pytorch for these two models. Our GPU training pipeline was completed by @abhishek in just few hours. We tried to port it to TPU in next few days but weren't able to get it working in time. It was a great experience teaming up with these guys I learned a lot about how to easily prototype your training pipeline. How to start off with validation pipeline and how important it is. Learnt about how to write TPU code for pytorch. \n\nThanks guys for teaming up. And thanks Kaggle for such a great contest.\n\nLastly, here's our final solution kernel https://www.kaggle.com/axel81/inference-use-hugging-face-postprocess.\n\nHappy Kaggling :)",
      "votes": 17
    },
    {
      "id": 726627,
      "postDate": "2020-01-23T06:15:24.133Z",
      "content": "<p>Congrats <a href=\"/axel81\">@axel81</a> and thanks for sharing. \nAlso thanks for being very active in this competition.\nYou motivated me a lot.</p>",
      "rawMarkdown": "Congrats @axel81 and thanks for sharing. \nAlso thanks for being very active in this competition.\nYou motivated me a lot.",
      "votes": 1,
      "replies": [
        {
          "id": 726645,
          "postDate": "2020-01-23T06:28:06.740Z",
          "content": "<p>Thanks <a href=\"/higepon\">@higepon</a> . Same for me, it was great a contest and I enjoyed competing with you too :)</p>",
          "rawMarkdown": "Thanks @higepon . Same for me, it was great a contest and I enjoyed competing with you too :)"
        }
      ]
    },
    {
      "id": 726621,
      "postDate": "2020-01-23T06:09:30.823Z",
      "content": "<p><a href=\"/axel81\">@axel81</a> , Congratulation!</p>\n\n<p>As you predicted in another thread, I paid a lot for playing too much with thresholds :)</p>\n\n<p>I wonder that is your <code>BERT-joint-large public lb score 0.6</code> the original bert joint with some basic post-processing (and thresholds)? Thanks.</p>",
      "rawMarkdown": "@axel81 , Congratulation!\n\nAs you predicted in another thread, I paid a lot for playing too much with thresholds :)\n\nI wonder that is your `BERT-joint-large public lb score 0.6` the original bert joint with some basic post-processing (and thresholds)? Thanks.",
      "votes": 1,
      "replies": [
        {
          "id": 726639,
          "postDate": "2020-01-23T06:26:34.390Z",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> my BERT-joint-large which got 0.6 is the one I trained on TPU. It got 0.58 without changes is postprocessing but with some modifications mentioned above it score 0.6. And as I said I kept constant thresholds after deciding from dev set throughout the experiments.</p>\n\n<p>Your work was really great, if you had focused on model training and ensembling you would've got a really good score on private set. As I said my solution is mostly based on your kernel :)</p>",
          "rawMarkdown": "@yihdarshieh my BERT-joint-large which got 0.6 is the one I trained on TPU. It got 0.58 without changes is postprocessing but with some modifications mentioned above it score 0.6. And as I said I kept constant thresholds after deciding from dev set throughout the experiments.\n\nYour work was really great, if you had focused on model training and ensembling you would've got a really good score on private set. As I said my solution is mostly based on your kernel :)",
          "votes": 1
        },
        {
          "id": 726655,
          "postDate": "2020-01-23T06:34:38.357Z",
          "content": "<p>Good to know I helped :) </p>\n\n<p>One more question, how many epochs for the model scoring 0.58?\nAnd can you tell us more about how did you do ensembling? I am not familar  with it.</p>",
          "rawMarkdown": "Good to know I helped :) \n\nOne more question, how many epochs for the model scoring 0.58?\nAnd can you tell us more about how did you do ensembling? I am not familar  with it."
        },
        {
          "id": 726663,
          "postDate": "2020-01-23T06:38:37.047Z",
          "content": "<p>2 epochs. Sure, I'll update the post for ensembling logic too. If you check out my kernel you'll see how I do ensembling. I used simple weighted average ensembling, here's the code.</p>\n\n<p>```\n        nq_logits = bert_nq(nq_inputs, training=False)\n        base_nq_logits = base_bert_nq(nq_inputs, training=False)</p>\n\n<pre><code>    (start_pos_logits, end_pos_logits, answer_type_logits) = nq_logits\n    (base_start_pos_logits, base_end_pos_logits, baseanswer_type_logits) = base_nq_logits\n\n    start_pos_logits = (0.2 * start_pos_logits + 0.8 * base_start_pos_logits)\n    end_pos_logits = (0.2 * end_pos_logits + 0.8 * base_end_pos_logits)\n    answer_type_logits = (0.2 * answer_type_logits + 0.8 * baseanswer_type_logits)\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "2 epochs. Sure, I'll update the post for ensembling logic too. If you check out my kernel you'll see how I do ensembling. I used simple weighted average ensembling, here's the code.\n\n```\n        nq_logits = bert_nq(nq_inputs, training=False)\n        base_nq_logits = base_bert_nq(nq_inputs, training=False)\n\n        (start_pos_logits, end_pos_logits, answer_type_logits) = nq_logits\n        (base_start_pos_logits, base_end_pos_logits, baseanswer_type_logits) = base_nq_logits\n        \n        start_pos_logits = (0.2 * start_pos_logits + 0.8 * base_start_pos_logits)\n        end_pos_logits = (0.2 * end_pos_logits + 0.8 * base_end_pos_logits)\n        answer_type_logits = (0.2 * answer_type_logits + 0.8 * baseanswer_type_logits)\n```\n\n",
          "votes": 1
        },
        {
          "id": 726677,
          "postDate": "2020-01-23T06:49:03.600Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        },
        {
          "id": 726789,
          "postDate": "2020-01-23T08:36:38.827Z",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> Thank you very much for your public kernel, your work is of great help for me,  based on you training and inference kernel, I got 0.67 using RoBERTa-large.</p>",
          "rawMarkdown": "@yihdarshieh Thank you very much for your public kernel, your work is of great help for me,  based on you training and inference kernel, I got 0.67 using RoBERTa-large.",
          "votes": 1
        },
        {
          "id": 726875,
          "postDate": "2020-01-23T09:45:30.107Z",
          "content": "<p>You are welcome. Happy to know, at least I can imagine I were close to medals 😂</p>",
          "rawMarkdown": "You are welcome. Happy to know, at least I can imagine I were close to medals 😂"
        }
      ]
    },
    {
      "id": 732950,
      "postDate": "2020-01-30T13:35:36.360Z",
      "content": "<p>Congrats</p>",
      "rawMarkdown": "Congrats"
    },
    {
      "id": 727368,
      "postDate": "2020-01-23T16:58:26.777Z",
      "content": "<p>Congrats and another silver for you. Thanks for sharing your approach <a href=\"/axel81\">@axel81</a> </p>",
      "rawMarkdown": "Congrats and another silver for you. Thanks for sharing your approach @axel81 "
    },
    {
      "id": 726903,
      "postDate": "2020-01-23T10:07:33.960Z",
      "content": "<p>Congrats &amp; Thanks for sharing!!🎉 😄 👍 </p>",
      "rawMarkdown": "Congrats &amp; Thanks for sharing!!🎉 😄 👍 "
    },
    {
      "id": 726618,
      "postDate": "2020-01-23T06:07:13.037Z",
      "content": "<p>Congratulations and thank you for sharing the kernel!</p>",
      "rawMarkdown": "Congratulations and thank you for sharing the kernel!"
    },
    {
      "id": 726609,
      "postDate": "2020-01-23T05:58:40.387Z",
      "content": "<p><a href=\"/axel81\">@axel81</a>  Congratulations buddy!\nSo how did the roberta and albert performed, any significant learning from there, if you can share ? \nOne more thing, did you try to manipulate score calculation? Like rather than having sum of logits trying harmonic means or something like that?\nThanks for sharing the kernal. :)</p>",
      "rawMarkdown": "@axel81  Congratulations buddy!\nSo how did the roberta and albert performed, any significant learning from there, if you can share ? \nOne more thing, did you try to manipulate score calculation? Like rather than having sum of logits trying harmonic means or something like that?\nThanks for sharing the kernal. :)",
      "replies": [
        {
          "id": 726614,
          "postDate": "2020-01-23T06:03:09.173Z",
          "content": "<p><a href=\"/rajraviprajapat\">@rajraviprajapat</a> we weren't able to get it working. But as I read in other discussions RoBERTa performed better compared to bert large</p>",
          "rawMarkdown": "@rajraviprajapat we weren't able to get it working. But as I read in other discussions RoBERTa performed better compared to bert large"
        }
      ]
    },
    {
      "id": 726647,
      "postDate": "2020-01-23T06:28:20.380Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 726627,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "2020-01-23T06:15:24.133000",
      "content": "<p>Congrats <a href=\"/axel81\">@axel81</a> and thanks for sharing. \nAlso thanks for being very active in this competition.\nYou motivated me a lot.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 726645,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2020-01-23T06:28:06.740000",
          "content": "<p>Thanks <a href=\"/higepon\">@higepon</a> . Same for me, it was great a contest and I enjoyed competing with you too :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 726621,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2020-01-23T06:09:30.823000",
      "content": "<p><a href=\"/axel81\">@axel81</a> , Congratulation!</p>\n\n<p>As you predicted in another thread, I paid a lot for playing too much with thresholds :)</p>\n\n<p>I wonder that is your <code>BERT-joint-large public lb score 0.6</code> the original bert joint with some basic post-processing (and thresholds)? Thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 726639,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2020-01-23T06:26:34.390000",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> my BERT-joint-large which got 0.6 is the one I trained on TPU. It got 0.58 without changes is postprocessing but with some modifications mentioned above it score 0.6. And as I said I kept constant thresholds after deciding from dev set throughout the experiments.</p>\n\n<p>Your work was really great, if you had focused on model training and ensembling you would've got a really good score on private set. As I said my solution is mostly based on your kernel :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 726655,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-01-23T06:34:38.357000",
          "content": "<p>Good to know I helped :) </p>\n\n<p>One more question, how many epochs for the model scoring 0.58?\nAnd can you tell us more about how did you do ensembling? I am not familar  with it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 726663,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2020-01-23T06:38:37.047000",
          "content": "<p>2 epochs. Sure, I'll update the post for ensembling logic too. If you check out my kernel you'll see how I do ensembling. I used simple weighted average ensembling, here's the code.</p>\n\n<p>```\n        nq_logits = bert_nq(nq_inputs, training=False)\n        base_nq_logits = base_bert_nq(nq_inputs, training=False)</p>\n\n<pre><code>    (start_pos_logits, end_pos_logits, answer_type_logits) = nq_logits\n    (base_start_pos_logits, base_end_pos_logits, baseanswer_type_logits) = base_nq_logits\n\n    start_pos_logits = (0.2 * start_pos_logits + 0.8 * base_start_pos_logits)\n    end_pos_logits = (0.2 * end_pos_logits + 0.8 * base_end_pos_logits)\n    answer_type_logits = (0.2 * answer_type_logits + 0.8 * baseanswer_type_logits)\n</code></pre>\n\n<p>```</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 726677,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-01-23T06:49:03.600000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 726789,
          "author_name": "Zhiyu Guo",
          "author_url": "",
          "post_date": "2020-01-23T08:36:38.827000",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> Thank you very much for your public kernel, your work is of great help for me,  based on you training and inference kernel, I got 0.67 using RoBERTa-large.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 726875,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-01-23T09:45:30.107000",
          "content": "<p>You are welcome. Happy to know, at least I can imagine I were close to medals 😂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 732950,
      "author_name": "哈尔的移动城堡",
      "author_url": "",
      "post_date": "2020-01-30T13:35:36.360000",
      "content": "<p>Congrats</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727368,
      "author_name": "Kurian Benoy",
      "author_url": "",
      "post_date": "2020-01-23T16:58:26.777000",
      "content": "<p>Congrats and another silver for you. Thanks for sharing your approach <a href=\"/axel81\">@axel81</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 726903,
      "author_name": "Miyabon",
      "author_url": "",
      "post_date": "2020-01-23T10:07:33.960000",
      "content": "<p>Congrats &amp; Thanks for sharing!!🎉 😄 👍 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 726618,
      "author_name": "DtneSEffct",
      "author_url": "",
      "post_date": "2020-01-23T06:07:13.037000",
      "content": "<p>Congratulations and thank you for sharing the kernel!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 726609,
      "author_name": "Rraj",
      "author_url": "",
      "post_date": "2020-01-23T05:58:40.387000",
      "content": "<p><a href=\"/axel81\">@axel81</a>  Congratulations buddy!\nSo how did the roberta and albert performed, any significant learning from there, if you can share ? \nOne more thing, did you try to manipulate score calculation? Like rather than having sum of logits trying harmonic means or something like that?\nThanks for sharing the kernal. :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 726614,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2020-01-23T06:03:09.173000",
          "content": "<p><a href=\"/rajraviprajapat\">@rajraviprajapat</a> we weren't able to get it working. But as I read in other discussions RoBERTa performed better compared to bert large</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 726647,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-23T06:28:20.380000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "726599": "Thank you Kaggle and Kaggle community for this awesome competition. I learned a lot.\nWe tried a lot of new things with pytorch in the last week but it weren't able to get things working, but it has been really fun.\n\n## **Our Solution**\n\nOur current solution is a TF 2.0 solution based on this great kernel https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models by @yihdarshieh. Initially I started off with  finetuning the official bert joint baseline but it didn't give much improvements. In next few weeks I completed the TPU setup for the contest and had my pipeline ready for training and validation on dev set.\n\nAll the experiments were done in Google colabs free tier TPUs. This is my first time seriously using TPUs and I have to say, it feels so good. Because they are so fast. Using TPUs dramatically reduced experimentation time in my case.\n\nOur solution is a simple ensemble of following 2 models:\n\n1. BERT-joint-large public lb score 0.6\n2. BERT-joint-base public lb score 0.58\n3. DistillBERT-joint public lb score 0.54\n3. Our final solution is an ensemble of 1 and 2 with some additions to postprocessing which scores 0.64 on public lb (also scored 0.64 on private LB but we didn't choose our best solution for final submission as it scored less on public LB with the new postprocessing)\n\nHere's our ensembling code, we use simple weighted average ensembling.\n\n```\n        nq_logits = bert_nq(nq_inputs, training=False)\n        base_nq_logits = base_bert_nq(nq_inputs, training=False)\n\n        (start_pos_logits, end_pos_logits, answer_type_logits) = nq_logits\n        (base_start_pos_logits, base_end_pos_logits, baseanswer_type_logits) = base_nq_logits\n        \n        start_pos_logits = (0.2 * start_pos_logits + 0.8 * base_start_pos_logits)\n        end_pos_logits = (0.2 * end_pos_logits + 0.8 * base_end_pos_logits)\n        answer_type_logits = (0.2 * answer_type_logits + 0.8 * baseanswer_type_logits)\n```\n\n## **Postprocessing**\n\nAfter experimenting with multiple single models I started focusing on postprocessing to improve on model performance. Initially I used postprocessing provided by this great kernel https://www.kaggle.com/prokaj/bert-joint-baseline-notebook by @prvi which helped my single models score in range 0.56-0.58. To improve further on this I had an in-depth look at the predictions of the model and ground truths. Here I found that our model was predicting duplicate answer spans. So I added a duplicate removal logic to postprocessing which helped a score increase on 0.01 on public LB and 0.02 on dev set. I also observed a score improvement if my model doesn't predict any \"YES/NO\" answer so essentially my model was only outputting answer spans and null answers in my final solution.\n\nDeciding on thresholds was one of the important things to predict valid answers but I didn't play too much with answer thresholds. Initially I ran inference on validation set with 5 different answer thresholds [1.5, 3.0, 4.5, 6.0, 7.5] and saw the best validation score with a combination of 1.5 for long answer and 3.0 for short answer which I kept same for final solution.\n\n## **New Addition to postprocessing**\n\nPostprocessing from @prvi's kernel only looks at the current `512` sequence to choose start and end indexes but as we are using strides of length `128` there's a overlap in every 2 consecutive sequences. So I decided to append top `k` start and end indexes from every 2 consecutive sequences to choose a pair of start and end indexes. In this case for overlapping start and end indexes I got 4 possible score values out of which I kept only the max score value and discarded the rest 3. This postprocessing didn't give us improvement on public LB for our current best model but it helped by a increment of 0.17 for a weak model.\n\nOur final solution didn't use the new postprocessing but our best solution on private set scored 0.64 (a 0.01 increment) with new postprocessing.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F887695%2Fae6bdb038215e5da57e9b68cc6de5242%2Fscore.png?generation=1579758032774744&amp;alt=media)\n\n## ** Last Week of Competition**\n\nIn the last week of the competition I got an opportunity to team up with @abhishek and @rinnqd. After working on BERT-large and BERT-base we wanted to try out ALBERT and RoBERTa in last few days so we started working on pytorch for these two models. Our GPU training pipeline was completed by @abhishek in just few hours. We tried to port it to TPU in next few days but weren't able to get it working in time. It was a great experience teaming up with these guys I learned a lot about how to easily prototype your training pipeline. How to start off with validation pipeline and how important it is. Learnt about how to write TPU code for pytorch. \n\nThanks guys for teaming up. And thanks Kaggle for such a great contest.\n\nLastly, here's our final solution kernel https://www.kaggle.com/axel81/inference-use-hugging-face-postprocess.\n\nHappy Kaggling :)",
    "726627": "Congrats @axel81 and thanks for sharing. \nAlso thanks for being very active in this competition.\nYou motivated me a lot.",
    "726621": "@axel81 , Congratulation!\n\nAs you predicted in another thread, I paid a lot for playing too much with thresholds :)\n\nI wonder that is your `BERT-joint-large public lb score 0.6` the original bert joint with some basic post-processing (and thresholds)? Thanks.",
    "732950": "Congrats",
    "727368": "Congrats and another silver for you. Thanks for sharing your approach @axel81 ",
    "726903": "Congrats &amp; Thanks for sharing!!🎉 😄 👍 ",
    "726618": "Congratulations and thank you for sharing the kernel!",
    "726609": "@axel81  Congratulations buddy!\nSo how did the roberta and albert performed, any significant learning from there, if you can share ? \nOne more thing, did you try to manipulate score calculation? Like rather than having sum of logits trying harmonic means or something like that?\nThanks for sharing the kernal. :)",
    "726647": ""
  }
}