{
  "id": 127551,
  "title": "1st place solution with code",
  "url": "/competitions/tensorflow2-question-answering/discussion/127551",
  "author_name": "Guanshuo Xu",
  "post_date": "2020-01-24T17:06:54.047000",
  "votes": 151,
  "comment_count": 55,
  "views": 0,
  "content": "<h1><strong>04/01/2020: Source code is attached below.</strong></h1>\n\n<p>Thanks to the Kaggle and Tensorflow team for holding this competition. I was new to question answering, it took me more than 5 weeks to make my first real submission, and I have learnt a lot during the journey. My initial plan before joining the competition was to learn both QA and TF2.0, but in the end I didn't have time to touch TF2.0, so my solution stays in pure pytorch. Thanks to <a href=\"/sakami\">@sakami</a> for the great kernel <a href=\"https://www.kaggle.com/sakami/tfqa-pytorch-baseline\">https://www.kaggle.com/sakami/tfqa-pytorch-baseline</a>. Your kernel was the starting point of my journey. And of course thanks to huggingface (<a href=\"https://github.com/huggingface/transformers\">https://github.com/huggingface/transformers</a>), NLP finetuning is made much easier. </p>\n\n<p>My solution is described below.</p>\n\n<h3><strong>- Overview</strong></h3>\n\n<p>I trained on the provided candidates instead of sampling from the original documents (examples) as done in the baseline paper (<a href=\"https://arxiv.org/abs/1901.08634\">https://arxiv.org/abs/1901.08634</a>). Since there are a total of 40 million candidates in the training data, for each epoch, I sampled only one negative candidate from each document. For more efficient training, hard negative sampling was used to replace uniform random sampling. The final submission was an ensemble of five models. </p>\n\n<h3><strong>- Sampling Strategy</strong></h3>\n\n<p>Initially, I tried uniform sampling on the negative candidates, but the result was unsatisfactory. The reason might be that most of the negative candidates are \"too easy\", the model might only need to learn some \"basic\" patterns for good candidate-level classification performance. But in the testing stage our actual goal is to predict the most probable positive candidate from each document, and this document-level classification is a more difficult task. So I replaced the uniform sampling by hard negative sampling to increase the difficulty of the candidate-level training, as expected, the performance was greatly improved. To perform hard negative sampling in the following models, I firstly trained a model with uniform sampling, and predicted on the whole training data, and stored the answer probability for each negative candidate. The last step was to normalize the probabilities of negative candidates within documents to form a distribution. For the following model training the negative candidates could be sampled from the probability distribution.</p>\n\n<h3><strong>- New Tokens</strong></h3>\n\n<p>According to the baseline paper, I added html tags as new tokens for better model performance. All the 9 tags from the Data Statistics Section of <a href=\"https://github.com/google-research-datasets/natural-questions\">https://github.com/google-research-datasets/natural-questions</a> was added. For html tags that are not in the 9 added tokens, I replaced them with a unique token in the tokenization dictionary or simply addedanother new token to represent them. I did not have time to try adding paragraph or table number similar to what the baseline paper does.</p>\n\n<h3><strong>- Model Architecture, Training and Evaluation</strong></h3>\n\n<p>Overall, the model architecture was the same as the baseline paper (a 5 class classification branch + 2 span classification branch). The five classes was \"no_answer\", \"long_answer_only\", \"short_answer\", \"yes\", \"no\". In my case there was no span prediction for answers without a short answer span because I directly used candidates. The loss update of the span prediction branch was simply ignored if no short answer span exist during training. In testing stage, for each document, I used 1.0-prob(no_answer) as the long answer score (confidence) for each candidate, and the candidate with the highest confidence was chosen to represent the document. Short answer spans were forced to be within the highest score long answer candidate (not sure if this is necessary). I used prob(short_answer)+prob(yes)+prob(no) as the short answer score. The exact class of the short answer was determined by the maximum of the three prob values. For span prediction, the output token-level probabilities were mapped to the word-level (white space tokenized) probabilities for easier ensembling of models with different tokenizers. </p>\n\n<h3><strong>- Models and Results</strong></h3>\n\n<p>My final submission was an ensemble of one Bert-base, two Bert-large (WWM), and two Albert-xxl (v2) models, all uncased. The Bert large and Albert models had been tuned on the SQUAD data before training. Below list their validation performance on the dev set using the code <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py\">https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py</a>. I did not try to implement the competition metric.</p>\n\n<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; long-best-threshold-f1&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; short-best-threshold-f1\nBert-base&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;0.618&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.457\nBert-large &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;0.679&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.541\nAlbert-xxl&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;0.700&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.555\nensemble&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; 0.731&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.582</p>\n\n<h3><strong>- Final LB results</strong></h3>\n\n<p>My best ensemble only achieved 0.66 public LB (0.69 private) performance using the optimized thresholds. At that time I had already lost most of my hope to win. In my last 2-3 submission, I arbitrarily played with the thresholds. One of the submissions scored 0.71 (both public and private LB), and I chose it and won the competition. Unbelievable.</p>",
  "messages": [
    {
      "id": 728384,
      "postDate": "2020-01-24T17:06:54.047Z",
      "content": "<h1><strong>04/01/2020: Source code is attached below.</strong></h1>\n\n<p>Thanks to the Kaggle and Tensorflow team for holding this competition. I was new to question answering, it took me more than 5 weeks to make my first real submission, and I have learnt a lot during the journey. My initial plan before joining the competition was to learn both QA and TF2.0, but in the end I didn't have time to touch TF2.0, so my solution stays in pure pytorch. Thanks to <a href=\"/sakami\">@sakami</a> for the great kernel <a href=\"https://www.kaggle.com/sakami/tfqa-pytorch-baseline\">https://www.kaggle.com/sakami/tfqa-pytorch-baseline</a>. Your kernel was the starting point of my journey. And of course thanks to huggingface (<a href=\"https://github.com/huggingface/transformers\">https://github.com/huggingface/transformers</a>), NLP finetuning is made much easier. </p>\n\n<p>My solution is described below.</p>\n\n<h3><strong>- Overview</strong></h3>\n\n<p>I trained on the provided candidates instead of sampling from the original documents (examples) as done in the baseline paper (<a href=\"https://arxiv.org/abs/1901.08634\">https://arxiv.org/abs/1901.08634</a>). Since there are a total of 40 million candidates in the training data, for each epoch, I sampled only one negative candidate from each document. For more efficient training, hard negative sampling was used to replace uniform random sampling. The final submission was an ensemble of five models. </p>\n\n<h3><strong>- Sampling Strategy</strong></h3>\n\n<p>Initially, I tried uniform sampling on the negative candidates, but the result was unsatisfactory. The reason might be that most of the negative candidates are \"too easy\", the model might only need to learn some \"basic\" patterns for good candidate-level classification performance. But in the testing stage our actual goal is to predict the most probable positive candidate from each document, and this document-level classification is a more difficult task. So I replaced the uniform sampling by hard negative sampling to increase the difficulty of the candidate-level training, as expected, the performance was greatly improved. To perform hard negative sampling in the following models, I firstly trained a model with uniform sampling, and predicted on the whole training data, and stored the answer probability for each negative candidate. The last step was to normalize the probabilities of negative candidates within documents to form a distribution. For the following model training the negative candidates could be sampled from the probability distribution.</p>\n\n<h3><strong>- New Tokens</strong></h3>\n\n<p>According to the baseline paper, I added html tags as new tokens for better model performance. All the 9 tags from the Data Statistics Section of <a href=\"https://github.com/google-research-datasets/natural-questions\">https://github.com/google-research-datasets/natural-questions</a> was added. For html tags that are not in the 9 added tokens, I replaced them with a unique token in the tokenization dictionary or simply addedanother new token to represent them. I did not have time to try adding paragraph or table number similar to what the baseline paper does.</p>\n\n<h3><strong>- Model Architecture, Training and Evaluation</strong></h3>\n\n<p>Overall, the model architecture was the same as the baseline paper (a 5 class classification branch + 2 span classification branch). The five classes was \"no_answer\", \"long_answer_only\", \"short_answer\", \"yes\", \"no\". In my case there was no span prediction for answers without a short answer span because I directly used candidates. The loss update of the span prediction branch was simply ignored if no short answer span exist during training. In testing stage, for each document, I used 1.0-prob(no_answer) as the long answer score (confidence) for each candidate, and the candidate with the highest confidence was chosen to represent the document. Short answer spans were forced to be within the highest score long answer candidate (not sure if this is necessary). I used prob(short_answer)+prob(yes)+prob(no) as the short answer score. The exact class of the short answer was determined by the maximum of the three prob values. For span prediction, the output token-level probabilities were mapped to the word-level (white space tokenized) probabilities for easier ensembling of models with different tokenizers. </p>\n\n<h3><strong>- Models and Results</strong></h3>\n\n<p>My final submission was an ensemble of one Bert-base, two Bert-large (WWM), and two Albert-xxl (v2) models, all uncased. The Bert large and Albert models had been tuned on the SQUAD data before training. Below list their validation performance on the dev set using the code <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py\">https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py</a>. I did not try to implement the competition metric.</p>\n\n<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; long-best-threshold-f1&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; short-best-threshold-f1\nBert-base&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;0.618&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.457\nBert-large &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;0.679&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.541\nAlbert-xxl&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;0.700&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.555\nensemble&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; 0.731&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.582</p>\n\n<h3><strong>- Final LB results</strong></h3>\n\n<p>My best ensemble only achieved 0.66 public LB (0.69 private) performance using the optimized thresholds. At that time I had already lost most of my hope to win. In my last 2-3 submission, I arbitrarily played with the thresholds. One of the submissions scored 0.71 (both public and private LB), and I chose it and won the competition. Unbelievable.</p>",
      "rawMarkdown": "# **04/01/2020: Source code is attached below.**\n\nThanks to the Kaggle and Tensorflow team for holding this competition. I was new to question answering, it took me more than 5 weeks to make my first real submission, and I have learnt a lot during the journey. My initial plan before joining the competition was to learn both QA and TF2.0, but in the end I didn't have time to touch TF2.0, so my solution stays in pure pytorch. Thanks to @sakami for the great kernel https://www.kaggle.com/sakami/tfqa-pytorch-baseline. Your kernel was the starting point of my journey. And of course thanks to huggingface (https://github.com/huggingface/transformers), NLP finetuning is made much easier. \n\nMy solution is described below.\n\n### **- Overview**\nI trained on the provided candidates instead of sampling from the original documents (examples) as done in the baseline paper (https://arxiv.org/abs/1901.08634). Since there are a total of 40 million candidates in the training data, for each epoch, I sampled only one negative candidate from each document. For more efficient training, hard negative sampling was used to replace uniform random sampling. The final submission was an ensemble of five models. \n\n### **- Sampling Strategy**\nInitially, I tried uniform sampling on the negative candidates, but the result was unsatisfactory. The reason might be that most of the negative candidates are \"too easy\", the model might only need to learn some \"basic\" patterns for good candidate-level classification performance. But in the testing stage our actual goal is to predict the most probable positive candidate from each document, and this document-level classification is a more difficult task. So I replaced the uniform sampling by hard negative sampling to increase the difficulty of the candidate-level training, as expected, the performance was greatly improved. To perform hard negative sampling in the following models, I firstly trained a model with uniform sampling, and predicted on the whole training data, and stored the answer probability for each negative candidate. The last step was to normalize the probabilities of negative candidates within documents to form a distribution. For the following model training the negative candidates could be sampled from the probability distribution.\n\n### **- New Tokens**\nAccording to the baseline paper, I added html tags as new tokens for better model performance. All the 9 tags from the Data Statistics Section of https://github.com/google-research-datasets/natural-questions was added. For html tags that are not in the 9 added tokens, I replaced them with a unique token in the tokenization dictionary or simply addedanother new token to represent them. I did not have time to try adding paragraph or table number similar to what the baseline paper does.\n\n### **- Model Architecture, Training and Evaluation**\nOverall, the model architecture was the same as the baseline paper (a 5 class classification branch + 2 span classification branch). The five classes was \"no_answer\", \"long_answer_only\", \"short_answer\", \"yes\", \"no\". In my case there was no span prediction for answers without a short answer span because I directly used candidates. The loss update of the span prediction branch was simply ignored if no short answer span exist during training. In testing stage, for each document, I used 1.0-prob(no_answer) as the long answer score (confidence) for each candidate, and the candidate with the highest confidence was chosen to represent the document. Short answer spans were forced to be within the highest score long answer candidate (not sure if this is necessary). I used prob(short_answer)+prob(yes)+prob(no) as the short answer score. The exact class of the short answer was determined by the maximum of the three prob values. For span prediction, the output token-level probabilities were mapped to the word-level (white space tokenized) probabilities for easier ensembling of models with different tokenizers. \n\n### **- Models and Results**\nMy final submission was an ensemble of one Bert-base, two Bert-large (WWM), and two Albert-xxl (v2) models, all uncased. The Bert large and Albert models had been tuned on the SQUAD data before training. Below list their validation performance on the dev set using the code https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py. I did not try to implement the competition metric.\n\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; long-best-threshold-f1&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; short-best-threshold-f1\nBert-base&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;0.618&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.457\nBert-large &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;0.679&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.541\nAlbert-xxl&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;0.700&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.555\nensemble&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; 0.731&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.582\n\n### **- Final LB results**\nMy best ensemble only achieved 0.66 public LB (0.69 private) performance using the optimized thresholds. At that time I had already lost most of my hope to win. In my last 2-3 submission, I arbitrarily played with the thresholds. One of the submissions scored 0.71 (both public and private LB), and I chose it and won the competition. Unbelievable.",
      "votes": 151
    },
    {
      "id": 728421,
      "postDate": "2020-01-24T17:46:19.160Z",
      "content": "<p>Amazing job, congrats! You are again showing that tuning models on public LB can help, I have to re-think some of my beliefs 😅 </p>",
      "rawMarkdown": "Amazing job, congrats! You are again showing that tuning models on public LB can help, I have to re-think some of my beliefs 😅 ",
      "votes": 5,
      "replies": [
        {
          "id": 728508,
          "postDate": "2020-01-24T20:20:33.590Z",
          "content": "<p>I followed the old school rule of thumb for choosing submissions: pick one with best local score and the other one with best LB score💯 </p>",
          "rawMarkdown": "I followed the old school rule of thumb for choosing submissions: pick one with best local score and the other one with best LB score💯 ",
          "votes": 10
        }
      ]
    },
    {
      "id": 729002,
      "postDate": "2020-01-25T15:05:21.890Z",
      "content": "<p>Congrats! I'm glad that I could help you. ;)</p>",
      "rawMarkdown": "Congrats! I'm glad that I could help you. ;)",
      "votes": 3,
      "replies": [
        {
          "id": 729009,
          "postDate": "2020-01-25T15:11:07.680Z",
          "content": "<p>👍 👍 👍 </p>",
          "rawMarkdown": "👍 👍 👍 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 728429,
      "postDate": "2020-01-24T18:00:26.900Z",
      "content": "<p>This is really interesting - thanks for the writeup. I like the \"hard negative sampling\" idea - making the training harder to get better results. How much of a bump did this provide, if you've kept track?</p>\n\n<p>Can you say some more about your fine tuning process?\n- Same for each component in the ensemble?\n- What were the stopping conditions?\n- What was the validation set? Why that one?</p>\n\n<p>Why no cased models?</p>\n\n<p>How did you fit all those models into the kernel runtime limit? Sounds like a few other competitors got timeouts. Anything special to speed up inference?</p>\n\n<p>How did your ensembling work? \n- Sounds like you averaged word-level logits from each model (pre softmax)?\n- Have you tried other ensembling methods, eg post-softmax, or even post- compute_pred_dict?\n- Equal weight?\n- How did you pick the final ensemble components among all the models you trained?</p>\n\n<p>Thanks again - hope you don't mind the many questions, but learning from you guys is one of the big perks of participating in these.</p>\n\n<p>Regards,\nAlon</p>",
      "rawMarkdown": "This is really interesting - thanks for the writeup. I like the \"hard negative sampling\" idea - making the training harder to get better results. How much of a bump did this provide, if you've kept track?\n\nCan you say some more about your fine tuning process?\n- Same for each component in the ensemble?\n- What were the stopping conditions?\n- What was the validation set? Why that one?\n\nWhy no cased models?\n\nHow did you fit all those models into the kernel runtime limit? Sounds like a few other competitors got timeouts. Anything special to speed up inference?\n\nHow did your ensembling work? \n- Sounds like you averaged word-level logits from each model (pre softmax)?\n- Have you tried other ensembling methods, eg post-softmax, or even post- compute_pred_dict?\n- Equal weight?\n- How did you pick the final ensemble components among all the models you trained?\n\n\nThanks again - hope you don't mind the many questions, but learning from you guys is one of the big perks of participating in these.\n\nRegards,\nAlon",
      "votes": 3,
      "replies": [
        {
          "id": 728519,
          "postDate": "2020-01-24T20:51:51.793Z",
          "content": "<p><strong>I like the \"hard negative sampling\" idea - making the training harder to get better results. How much of a bump did this provide, if you've kept track?</strong>\nI don't, unfortunately.</p>\n\n<p><strong>Can you say some more about your fine tuning process?</strong>\nThe tuning process of the three types of models are generally same. They were tuning for 3-4 epochs. Early stopping was based on validation performance. The validation set (the dev) is the standard dataset provided in the NQ dataset along with the code for F1 score calculation. </p>\n\n<p><strong>Why no cased models?</strong>\nI remember I saw the advise to not use cased model unless there is a strong reason to (maybe in the bert paper?)</p>\n\n<p><strong>How did you fit all those models into the kernel runtime limit? Sounds like a few other competitors got timeouts. Anything special to speed up inference?</strong>\nGood question. I used the bert-base model for candidate proposal. From each document, the bert-base only proposed a small set of most probable candidates, and the bigger models only predicted on those candidates. I saw the total kernel inference time was only about an hour.</p>\n\n<p><strong>How did your ensembling work?</strong>\nIt was all after softmax. Weighted average. Albert has slightly higher weight. The exact weights are not important.\nI tried roberta and xlnet. Performance of roberta was much worse, there might be something wrong in my implementation. For Xlnet, I couldn't even run the provided huggingface squad tuning code. I'm sure these two could make even better ensemble together with bert and albert but I just didn't have time and energy to make them work correctly.</p>",
          "rawMarkdown": "**I like the \"hard negative sampling\" idea - making the training harder to get better results. How much of a bump did this provide, if you've kept track?**\nI don't, unfortunately.\n\n**Can you say some more about your fine tuning process?**\nThe tuning process of the three types of models are generally same. They were tuning for 3-4 epochs. Early stopping was based on validation performance. The validation set (the dev) is the standard dataset provided in the NQ dataset along with the code for F1 score calculation. \n\n**Why no cased models?**\nI remember I saw the advise to not use cased model unless there is a strong reason to (maybe in the bert paper?)\n\n**How did you fit all those models into the kernel runtime limit? Sounds like a few other competitors got timeouts. Anything special to speed up inference?**\nGood question. I used the bert-base model for candidate proposal. From each document, the bert-base only proposed a small set of most probable candidates, and the bigger models only predicted on those candidates. I saw the total kernel inference time was only about an hour.\n\n**How did your ensembling work?**\nIt was all after softmax. Weighted average. Albert has slightly higher weight. The exact weights are not important.\nI tried roberta and xlnet. Performance of roberta was much worse, there might be something wrong in my implementation. For Xlnet, I couldn't even run the provided huggingface squad tuning code. I'm sure these two could make even better ensemble together with bert and albert but I just didn't have time and energy to make them work correctly.",
          "votes": 14
        },
        {
          "id": 728531,
          "postDate": "2020-01-24T21:33:02.783Z",
          "content": "<p>\"I used the bert-base model for candidate proposal.\" - awesome!</p>",
          "rawMarkdown": "\"I used the bert-base model for candidate proposal.\" - awesome!"
        },
        {
          "id": 728808,
          "postDate": "2020-01-25T10:01:11.500Z",
          "content": "<p>I want to clarify if understanding of your ensembling is right: Bert-base to predict \"has long answer or not\". if at least one of  the spans of a example has long answer, then use big models to predict all spans of this example, right?</p>",
          "rawMarkdown": "I want to clarify if understanding of your ensembling is right: Bert-base to predict \"has long answer or not\". if at least one of  the spans of a example has long answer, then use big models to predict all spans of this example, right?"
        },
        {
          "id": 729008,
          "postDate": "2020-01-25T15:10:01.720Z",
          "content": "<blockquote>\n  <p>I want to clarify if understanding of your ensembling is right: Bert-base to predict \"has long answer or not\". if at least one of the spans of a example has long answer, then use big models to predict all spans of this example, right?</p>\n</blockquote>\n\n<p>The proposal (pre-selection) works on candidates, not examples.\nAssume an example has 100 candidates, the bert-base model will predict on all the 100 candidates, rank them with the long answer probabilities,  then, choose the most probable  5 candidates for the bert-large to predict on. This way the big models will reduce the workload by 95%.</p>",
          "rawMarkdown": "&gt; I want to clarify if understanding of your ensembling is right: Bert-base to predict \"has long answer or not\". if at least one of the spans of a example has long answer, then use big models to predict all spans of this example, right?\n\nThe proposal (pre-selection) works on candidates, not examples.\nAssume an example has 100 candidates, the bert-base model will predict on all the 100 candidates, rank them with the long answer probabilities,  then, choose the most probable  5 candidates for the bert-large to predict on. This way the big models will reduce the workload by 95%.",
          "votes": 3
        },
        {
          "id": 729017,
          "postDate": "2020-01-25T15:31:26.830Z",
          "content": "<p>Thank you for your correction. </p>",
          "rawMarkdown": "Thank you for your correction. "
        },
        {
          "id": 749419,
          "postDate": "2020-02-18T17:14:28.310Z",
          "content": "<p>If the same tokenization method is used, then this makes sense. Bert-base and bert-large both have the same tokenization and they can work as described. But for different models, say Albert and Bert, they have different tokenization, which might generate different candidates. How did you do ensemble as you said for such cases?</p>",
          "rawMarkdown": "If the same tokenization method is used, then this makes sense. Bert-base and bert-large both have the same tokenization and they can work as described. But for different models, say Albert and Bert, they have different tokenization, which might generate different candidates. How did you do ensemble as you said for such cases?"
        }
      ]
    },
    {
      "id": 730206,
      "postDate": "2020-01-27T08:05:07.693Z",
      "content": "<p>Congratss!! Thanks for sharing thats useful kernel !</p>",
      "rawMarkdown": "Congratss!! Thanks for sharing thats useful kernel !",
      "votes": 4
    },
    {
      "id": 728415,
      "postDate": "2020-01-24T17:37:58.707Z",
      "content": "<p>Amazing indeed <a href=\"/wowfattie\">@wowfattie</a> . Would you be releasing your pytorch source code? </p>",
      "rawMarkdown": "Amazing indeed @wowfattie . Would you be releasing your pytorch source code? ",
      "votes": 4,
      "replies": [
        {
          "id": 729426,
          "postDate": "2020-01-26T07:27:22.960Z",
          "content": "<p>I second this.</p>",
          "rawMarkdown": "I second this.",
          "votes": 1
        },
        {
          "id": 732948,
          "postDate": "2020-01-30T13:33:15.167Z",
          "content": "<p>Second again  :)</p>",
          "rawMarkdown": "Second again  :)",
          "votes": 1
        },
        {
          "id": 734537,
          "postDate": "2020-02-01T16:19:37.230Z",
          "content": "<p>Would love to see the training code as well. ;)</p>",
          "rawMarkdown": "Would love to see the training code as well. ;)"
        },
        {
          "id": 794162,
          "postDate": "2020-04-01T16:08:16.620Z",
          "content": "<p>Source code is attached.</p>",
          "rawMarkdown": "Source code is attached.",
          "votes": 2
        }
      ]
    },
    {
      "id": 790222,
      "postDate": "2020-03-29T12:20:35.077Z",
      "content": "<p>Congats </p>",
      "rawMarkdown": "Congats ",
      "votes": 1
    },
    {
      "id": 728416,
      "postDate": "2020-01-24T17:39:12.513Z",
      "content": "<p>Thanks for sharing and congrats on solo win!</p>",
      "rawMarkdown": "Thanks for sharing and congrats on solo win!",
      "votes": 1
    },
    {
      "id": 730415,
      "postDate": "2020-01-27T13:19:24.223Z",
      "content": "<p>Good Try Keep it up</p>",
      "rawMarkdown": "Good Try Keep it up",
      "votes": 2
    },
    {
      "id": 729745,
      "postDate": "2020-01-26T15:49:30.603Z",
      "content": "<p>Awesome solution and a motivating journey. Thanks for sharing.\nJust one minor detail, the <strong>bert</strong> paper link is broken, there is one extra <code>)</code> at the end. </p>",
      "rawMarkdown": "Awesome solution and a motivating journey. Thanks for sharing.\nJust one minor detail, the **bert** paper link is broken, there is one extra `)` at the end. ",
      "votes": 2
    },
    {
      "id": 729425,
      "postDate": "2020-01-26T07:26:58.647Z",
      "content": "<p>Congratulations! Any plans to release your pytorch training code? Given that there are not much pytorch solutions it would be great to see how the winner used pytorch in the solution! </p>",
      "rawMarkdown": "Congratulations! Any plans to release your pytorch training code? Given that there are not much pytorch solutions it would be great to see how the winner used pytorch in the solution! ",
      "votes": 2
    },
    {
      "id": 739420,
      "postDate": "2020-02-07T20:23:43.250Z",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>) </p>",
      "rawMarkdown": "@wowfattie, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques) "
    },
    {
      "id": 940206,
      "postDate": "2020-07-22T18:28:49.313Z",
      "content": "<p>Thanks for sharing! To understand your workflow better, did you implement your solutions in Notebooks first and then migrated it to script or do you prefer starting with script code right away?</p>",
      "rawMarkdown": "Thanks for sharing! To understand your workflow better, did you implement your solutions in Notebooks first and then migrated it to script or do you prefer starting with script code right away?"
    },
    {
      "id": 736816,
      "postDate": "2020-02-04T15:39:57.407Z",
      "content": "<p>That's a brilliant solution. Congratulations and thanks for sharing!!</p>",
      "rawMarkdown": "That's a brilliant solution. Congratulations and thanks for sharing!!"
    },
    {
      "id": 733642,
      "postDate": "2020-01-31T11:34:32.093Z",
      "content": "<p>unbelievable. victory. congraturation ^^/ good job. </p>",
      "rawMarkdown": "unbelievable. victory. congraturation ^^/ good job. "
    },
    {
      "id": 733176,
      "postDate": "2020-01-30T19:19:02.923Z",
      "content": "<p>Ty, gj :)</p>",
      "rawMarkdown": "Ty, gj :)"
    },
    {
      "id": 732978,
      "postDate": "2020-01-30T14:14:08.637Z",
      "content": "<p>Congratulations! <a href=\"/wowfattie\">@wowfattie</a> </p>",
      "rawMarkdown": "Congratulations! @wowfattie "
    },
    {
      "id": 732854,
      "postDate": "2020-01-30T11:25:09.870Z",
      "content": "<p>Congrats.. Winning streak continues.. <a href=\"/wowfattie\">@wowfattie</a> </p>\n\n<p>Kaggle - Review of Year 2019</p>\n\n<p><a href=\"https://youtu.be/pL96IPZZ-88\">https://youtu.be/pL96IPZZ-88</a></p>\n\n<p>Check out the video above.</p>",
      "rawMarkdown": "Congrats.. Winning streak continues.. @wowfattie \n\nKaggle - Review of Year 2019\n\n[https://youtu.be/pL96IPZZ-88](https://youtu.be/pL96IPZZ-88)\n\nCheck out the video above.\n\n"
    },
    {
      "id": 728687,
      "postDate": "2020-01-25T05:00:30.963Z",
      "content": "<blockquote>\n  <p>In my last 2-3 submission, I arbitrarily played with the thresholds. One of the submissions scored 0.71 (both public and private LB), and I chose it and won the competition. Unbelievable.</p>\n</blockquote>\n\n<p>Could you share the heuristics of tuning threshold? Is it grid search, random search or something else?</p>",
      "rawMarkdown": "&gt;  In my last 2-3 submission, I arbitrarily played with the thresholds. One of the submissions scored 0.71 (both public and private LB), and I chose it and won the competition. Unbelievable.\n\nCould you share the heuristics of tuning threshold? Is it grid search, random search or something else?",
      "replies": [
        {
          "id": 728714,
          "postDate": "2020-01-25T06:04:01.810Z",
          "content": "<p>There was no actual thresholds tuning. After I submitted my final ensemble (0.66 public LB) with locally optimized thresholds, I only adjusted the thresholds two times, reducing both long and short thresholds by 0.1 (0.7 public LB) and reducing both long and short thresholds by 0.2 (0.71 public LB). And I went to shopping, came back 30 mins before the competition deadline, and selected the 0.71 as my second selection.😁 </p>",
          "rawMarkdown": "There was no actual thresholds tuning. After I submitted my final ensemble (0.66 public LB) with locally optimized thresholds, I only adjusted the thresholds two times, reducing both long and short thresholds by 0.1 (0.7 public LB) and reducing both long and short thresholds by 0.2 (0.71 public LB). And I went to shopping, came back 30 mins before the competition deadline, and selected the 0.71 as my second selection.😁 ",
          "votes": 7
        },
        {
          "id": 729105,
          "postDate": "2020-01-25T19:08:22.967Z",
          "content": "<p>Is it just a random thought to lower the threshold or you observe some distribution difference between train and test data?</p>",
          "rawMarkdown": "Is it just a random thought to lower the threshold or you observe some distribution difference between train and test data?",
          "votes": 1
        },
        {
          "id": 729117,
          "postDate": "2020-01-25T19:27:19.447Z",
          "content": "<p>I'm not sure if there is (label) distribution difference.\nIndeed the decision to try lowering the thresholds were driven by my submission public LB scores. Before I submitted the 5-model ensemble, I had a 3-model ensemble (bert-base, bert-large, albert-xxl) which scored 0.68 public LB. The 5-model ensemble had an extra bert-large and an extra albert-xxl, so it must be better than the 3-model ensemble, but only scored 0.66. I double-checked my code and found that the 3-model ensemble used lower thresholds than the 5 model ensemble. That's why I tried lowering thresholds.\nThe funny part is the 5-model ensemble scored 0.69 in private LB whereas the 3-model ensemble still scored 0.68 private.</p>",
          "rawMarkdown": "I'm not sure if there is (label) distribution difference.\nIndeed the decision to try lowering the thresholds were driven by my submission public LB scores. Before I submitted the 5-model ensemble, I had a 3-model ensemble (bert-base, bert-large, albert-xxl) which scored 0.68 public LB. The 5-model ensemble had an extra bert-large and an extra albert-xxl, so it must be better than the 3-model ensemble, but only scored 0.66. I double-checked my code and found that the 3-model ensemble used lower thresholds than the 5 model ensemble. That's why I tried lowering thresholds.\nThe funny part is the 5-model ensemble scored 0.69 in private LB whereas the 3-model ensemble still scored 0.68 private.",
          "votes": 4
        },
        {
          "id": 729130,
          "postDate": "2020-01-25T19:36:35.487Z",
          "content": "<p><a href=\"/wowfattie\">@wowfattie</a> I am glad you give these additional details.  Your score isn't due to luck as your write up conclusion seemed to say.  I would have hated to lose to luck ;)  </p>",
          "rawMarkdown": "@wowfattie I am glad you give these additional details.  Your score isn't due to luck as your write up conclusion seemed to say.  I would have hated to lose to luck ;)  ",
          "votes": 3
        },
        {
          "id": 729236,
          "postDate": "2020-01-25T23:47:55.710Z",
          "content": "<p>&gt; Your score isn't due to luck</p>\n\n<p>Definitely. The 5-model ensemble is the key and lowering threshold is a nice touch. Thank you so much for sharing these details.</p>",
          "rawMarkdown": "&gt; Your score isn't due to luck\n\nDefinitely. The 5-model ensemble is the key and lowering threshold is a nice touch. Thank you so much for sharing these details.",
          "votes": 2
        }
      ]
    },
    {
      "id": 728538,
      "postDate": "2020-01-24T21:46:16.807Z",
      "content": "<p>I''m sorry for a noob question, but what is \"hard negative sampling\"?</p>",
      "rawMarkdown": "I''m sorry for a noob question, but what is \"hard negative sampling\"?",
      "replies": [
        {
          "id": 728565,
          "postDate": "2020-01-24T22:52:03.843Z",
          "content": "<p>It's similar to hard negative mining used in computer vision object detection</p>",
          "rawMarkdown": "It's similar to hard negative mining used in computer vision object detection",
          "votes": 1
        },
        {
          "id": 728693,
          "postDate": "2020-01-25T05:08:25.373Z",
          "content": "<p>Basically you train a baseline model and use it to predict all candidates in your validation data. Then we can look at the wrong predictions made by the baseline model, especially the negative candidate which have no answers within them but get high scores by the model. They have huge logloss indicating that it is hard for the model to discern them from positive candidates. Hence these candidates are \"hard negative candidates\". So in the second pass of training, we can just train the model with these hard negative candidates instead of randomly sampled negative candidates.</p>\n\n<p>Please correct me if I'm wrong.</p>",
          "rawMarkdown": "Basically you train a baseline model and use it to predict all candidates in your validation data. Then we can look at the wrong predictions made by the baseline model, especially the negative candidate which have no answers within them but get high scores by the model. They have huge logloss indicating that it is hard for the model to discern them from positive candidates. Hence these candidates are \"hard negative candidates\". So in the second pass of training, we can just train the model with these hard negative candidates instead of randomly sampled negative candidates.\n\nPlease correct me if I'm wrong.",
          "votes": 2
        },
        {
          "id": 728719,
          "postDate": "2020-01-25T06:24:50.657Z",
          "content": "<p><strong>Basically you train a baseline model and use it to predict all candidates in your validation data.</strong>\nWhat I did was I trained the model on the training set and used the trained model to predict on the same training set. </p>",
          "rawMarkdown": "**Basically you train a baseline model and use it to predict all candidates in your validation data.**\nWhat I did was I trained the model on the training set and used the trained model to predict on the same training set. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 728512,
      "postDate": "2020-01-24T20:31:41.723Z",
      "content": "<p>Hey <a href=\"/wowfattie\">@wowfattie</a>, big thank you for sharing this, I am really impressed that you used ALBERT-XXL when lots of people said that it timed out? May I know how did you go around that?</p>",
      "rawMarkdown": "Hey @wowfattie, big thank you for sharing this, I am really impressed that you used ALBERT-XXL when lots of people said that it timed out? May I know how did you go around that?",
      "replies": [
        {
          "id": 728522,
          "postDate": "2020-01-24T20:56:36.460Z",
          "content": "<p>Please see my answers here\n<a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/127551#728519\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/127551#728519</a></p>",
          "rawMarkdown": "Please see my answers here\nhttps://www.kaggle.com/c/tensorflow2-question-answering/discussion/127551#728519"
        },
        {
          "id": 728536,
          "postDate": "2020-01-24T21:40:52.037Z",
          "content": "<p>Thanks plenty <a href=\"/wowfattie\">@wowfattie</a>!</p>",
          "rawMarkdown": "Thanks plenty @wowfattie!"
        }
      ]
    },
    {
      "id": 728447,
      "postDate": "2020-01-24T18:21:17.717Z",
      "content": "<p>Congrats! Thanks for sharing insights and thought process.</p>",
      "rawMarkdown": "Congrats! Thanks for sharing insights and thought process."
    },
    {
      "id": 728425,
      "postDate": "2020-01-24T17:51:37.893Z",
      "content": "<p>Congrats! That is indeed an amazing solution! This is the best gift to celebrate Chinese New Year!</p>",
      "rawMarkdown": "Congrats! That is indeed an amazing solution! This is the best gift to celebrate Chinese New Year!"
    },
    {
      "id": 728417,
      "postDate": "2020-01-24T17:40:50.307Z",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a> Woah! Mindblown!🙌 🙌 🙌 </p>",
      "rawMarkdown": "@wowfattie Woah! Mindblown!🙌 🙌 🙌 "
    },
    {
      "id": 728394,
      "postDate": "2020-01-24T17:17:14.620Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 730363,
      "postDate": "2020-01-27T11:42:24.713Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 730343,
      "postDate": "2020-01-27T11:19:49.557Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 734228,
      "postDate": "2020-02-01T05:56:18.597Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 733293,
      "postDate": "2020-01-30T23:46:08.330Z",
      "content": "<p>Congrats and thank you for writing up.</p>",
      "rawMarkdown": "Congrats and thank you for writing up."
    },
    {
      "id": 730396,
      "postDate": "2020-01-27T12:40:52.403Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing\n"
    },
    {
      "id": 730076,
      "postDate": "2020-01-27T04:26:20.767Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing\n"
    },
    {
      "id": 729064,
      "postDate": "2020-01-25T17:24:39.330Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing."
    },
    {
      "id": 728498,
      "postDate": "2020-01-24T19:57:46.377Z",
      "content": "<p>Congrats! Thanks for sharing! </p>",
      "rawMarkdown": "Congrats! Thanks for sharing! "
    },
    {
      "id": 730004,
      "postDate": "2020-01-27T00:50:39.093Z",
      "content": "<p>Congrats! Thanks for sharing!</p>",
      "rawMarkdown": "Congrats! Thanks for sharing!",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 728421,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-01-24T17:46:19.160000",
      "content": "<p>Amazing job, congrats! You are again showing that tuning models on public LB can help, I have to re-think some of my beliefs 😅 </p>",
      "votes": 5,
      "replies": [
        {
          "id": 728508,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-01-24T20:20:33.590000",
          "content": "<p>I followed the old school rule of thumb for choosing submissions: pick one with best local score and the other one with best LB score💯 </p>",
          "votes": 10,
          "replies": []
        }
      ]
    },
    {
      "id": 729002,
      "author_name": "sakami",
      "author_url": "",
      "post_date": "2020-01-25T15:05:21.890000",
      "content": "<p>Congrats! I'm glad that I could help you. ;)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 729009,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-01-25T15:11:07.680000",
          "content": "<p>👍 👍 👍 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 728429,
      "author_name": "Alon Bochman",
      "author_url": "",
      "post_date": "2020-01-24T18:00:26.900000",
      "content": "<p>This is really interesting - thanks for the writeup. I like the \"hard negative sampling\" idea - making the training harder to get better results. How much of a bump did this provide, if you've kept track?</p>\n\n<p>Can you say some more about your fine tuning process?\n- Same for each component in the ensemble?\n- What were the stopping conditions?\n- What was the validation set? Why that one?</p>\n\n<p>Why no cased models?</p>\n\n<p>How did you fit all those models into the kernel runtime limit? Sounds like a few other competitors got timeouts. Anything special to speed up inference?</p>\n\n<p>How did your ensembling work? \n- Sounds like you averaged word-level logits from each model (pre softmax)?\n- Have you tried other ensembling methods, eg post-softmax, or even post- compute_pred_dict?\n- Equal weight?\n- How did you pick the final ensemble components among all the models you trained?</p>\n\n<p>Thanks again - hope you don't mind the many questions, but learning from you guys is one of the big perks of participating in these.</p>\n\n<p>Regards,\nAlon</p>",
      "votes": 3,
      "replies": [
        {
          "id": 728519,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-01-24T20:51:51.793000",
          "content": "<p><strong>I like the \"hard negative sampling\" idea - making the training harder to get better results. How much of a bump did this provide, if you've kept track?</strong>\nI don't, unfortunately.</p>\n\n<p><strong>Can you say some more about your fine tuning process?</strong>\nThe tuning process of the three types of models are generally same. They were tuning for 3-4 epochs. Early stopping was based on validation performance. The validation set (the dev) is the standard dataset provided in the NQ dataset along with the code for F1 score calculation. </p>\n\n<p><strong>Why no cased models?</strong>\nI remember I saw the advise to not use cased model unless there is a strong reason to (maybe in the bert paper?)</p>\n\n<p><strong>How did you fit all those models into the kernel runtime limit? Sounds like a few other competitors got timeouts. Anything special to speed up inference?</strong>\nGood question. I used the bert-base model for candidate proposal. From each document, the bert-base only proposed a small set of most probable candidates, and the bigger models only predicted on those candidates. I saw the total kernel inference time was only about an hour.</p>\n\n<p><strong>How did your ensembling work?</strong>\nIt was all after softmax. Weighted average. Albert has slightly higher weight. The exact weights are not important.\nI tried roberta and xlnet. Performance of roberta was much worse, there might be something wrong in my implementation. For Xlnet, I couldn't even run the provided huggingface squad tuning code. I'm sure these two could make even better ensemble together with bert and albert but I just didn't have time and energy to make them work correctly.</p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 728531,
          "author_name": "Alon Bochman",
          "author_url": "",
          "post_date": "2020-01-24T21:33:02.783000",
          "content": "<p>\"I used the bert-base model for candidate proposal.\" - awesome!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 728808,
          "author_name": "SchenbergZ",
          "author_url": "",
          "post_date": "2020-01-25T10:01:11.500000",
          "content": "<p>I want to clarify if understanding of your ensembling is right: Bert-base to predict \"has long answer or not\". if at least one of  the spans of a example has long answer, then use big models to predict all spans of this example, right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 729008,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-01-25T15:10:01.720000",
          "content": "<blockquote>\n  <p>I want to clarify if understanding of your ensembling is right: Bert-base to predict \"has long answer or not\". if at least one of the spans of a example has long answer, then use big models to predict all spans of this example, right?</p>\n</blockquote>\n\n<p>The proposal (pre-selection) works on candidates, not examples.\nAssume an example has 100 candidates, the bert-base model will predict on all the 100 candidates, rank them with the long answer probabilities,  then, choose the most probable  5 candidates for the bert-large to predict on. This way the big models will reduce the workload by 95%.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 729017,
          "author_name": "SchenbergZ",
          "author_url": "",
          "post_date": "2020-01-25T15:31:26.830000",
          "content": "<p>Thank you for your correction. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 749419,
          "author_name": "nzholmes",
          "author_url": "",
          "post_date": "2020-02-18T17:14:28.310000",
          "content": "<p>If the same tokenization method is used, then this makes sense. Bert-base and bert-large both have the same tokenization and they can work as described. But for different models, say Albert and Bert, they have different tokenization, which might generate different candidates. How did you do ensemble as you said for such cases?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 730206,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-27T08:05:07.693000",
      "content": "<p>Congratss!! Thanks for sharing thats useful kernel !</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 728415,
      "author_name": "Abhishek Thakur",
      "author_url": "",
      "post_date": "2020-01-24T17:37:58.707000",
      "content": "<p>Amazing indeed <a href=\"/wowfattie\">@wowfattie</a> . Would you be releasing your pytorch source code? </p>",
      "votes": 4,
      "replies": [
        {
          "id": 729426,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-01-26T07:27:22.960000",
          "content": "<p>I second this.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 732948,
          "author_name": "哈尔的移动城堡",
          "author_url": "",
          "post_date": "2020-01-30T13:33:15.167000",
          "content": "<p>Second again  :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 734537,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2020-02-01T16:19:37.230000",
          "content": "<p>Would love to see the training code as well. ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794162,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-04-01T16:08:16.620000",
          "content": "<p>Source code is attached.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 790222,
      "author_name": "podsyp",
      "author_url": "",
      "post_date": "2020-03-29T12:20:35.077000",
      "content": "<p>Congats </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 728416,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-01-24T17:39:12.513000",
      "content": "<p>Thanks for sharing and congrats on solo win!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 730415,
      "author_name": "Tayyab A",
      "author_url": "",
      "post_date": "2020-01-27T13:19:24.223000",
      "content": "<p>Good Try Keep it up</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 729745,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-01-26T15:49:30.603000",
      "content": "<p>Awesome solution and a motivating journey. Thanks for sharing.\nJust one minor detail, the <strong>bert</strong> paper link is broken, there is one extra <code>)</code> at the end. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 729425,
      "author_name": "ilovescience",
      "author_url": "",
      "post_date": "2020-01-26T07:26:58.647000",
      "content": "<p>Congratulations! Any plans to release your pytorch training code? Given that there are not much pytorch solutions it would be great to see how the winner used pytorch in the solution! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 739420,
      "author_name": "Vitalii Mokin",
      "author_url": "",
      "post_date": "2020-02-07T20:23:43.250000",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 940206,
      "author_name": "Muennighoff",
      "author_url": "",
      "post_date": "2020-07-22T18:28:49.313000",
      "content": "<p>Thanks for sharing! To understand your workflow better, did you implement your solutions in Notebooks first and then migrated it to script or do you prefer starting with script code right away?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 736816,
      "author_name": "zaymo",
      "author_url": "",
      "post_date": "2020-02-04T15:39:57.407000",
      "content": "<p>That's a brilliant solution. Congratulations and thanks for sharing!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 733642,
      "author_name": "HELLOWORLD",
      "author_url": "",
      "post_date": "2020-01-31T11:34:32.093000",
      "content": "<p>unbelievable. victory. congraturation ^^/ good job. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 733176,
      "author_name": "Dustin",
      "author_url": "",
      "post_date": "2020-01-30T19:19:02.923000",
      "content": "<p>Ty, gj :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 732978,
      "author_name": "gulshan kumar",
      "author_url": "",
      "post_date": "2020-01-30T14:14:08.637000",
      "content": "<p>Congratulations! <a href=\"/wowfattie\">@wowfattie</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 732854,
      "author_name": "Nandanam",
      "author_url": "",
      "post_date": "2020-01-30T11:25:09.870000",
      "content": "<p>Congrats.. Winning streak continues.. <a href=\"/wowfattie\">@wowfattie</a> </p>\n\n<p>Kaggle - Review of Year 2019</p>\n\n<p><a href=\"https://youtu.be/pL96IPZZ-88\">https://youtu.be/pL96IPZZ-88</a></p>\n\n<p>Check out the video above.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728687,
      "author_name": "Jiwei Liu",
      "author_url": "",
      "post_date": "2020-01-25T05:00:30.963000",
      "content": "<blockquote>\n  <p>In my last 2-3 submission, I arbitrarily played with the thresholds. One of the submissions scored 0.71 (both public and private LB), and I chose it and won the competition. Unbelievable.</p>\n</blockquote>\n\n<p>Could you share the heuristics of tuning threshold? Is it grid search, random search or something else?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 728714,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-01-25T06:04:01.810000",
          "content": "<p>There was no actual thresholds tuning. After I submitted my final ensemble (0.66 public LB) with locally optimized thresholds, I only adjusted the thresholds two times, reducing both long and short thresholds by 0.1 (0.7 public LB) and reducing both long and short thresholds by 0.2 (0.71 public LB). And I went to shopping, came back 30 mins before the competition deadline, and selected the 0.71 as my second selection.😁 </p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 729105,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2020-01-25T19:08:22.967000",
          "content": "<p>Is it just a random thought to lower the threshold or you observe some distribution difference between train and test data?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 729117,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-01-25T19:27:19.447000",
          "content": "<p>I'm not sure if there is (label) distribution difference.\nIndeed the decision to try lowering the thresholds were driven by my submission public LB scores. Before I submitted the 5-model ensemble, I had a 3-model ensemble (bert-base, bert-large, albert-xxl) which scored 0.68 public LB. The 5-model ensemble had an extra bert-large and an extra albert-xxl, so it must be better than the 3-model ensemble, but only scored 0.66. I double-checked my code and found that the 3-model ensemble used lower thresholds than the 5 model ensemble. That's why I tried lowering thresholds.\nThe funny part is the 5-model ensemble scored 0.69 in private LB whereas the 3-model ensemble still scored 0.68 private.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 729130,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-01-25T19:36:35.487000",
          "content": "<p><a href=\"/wowfattie\">@wowfattie</a> I am glad you give these additional details.  Your score isn't due to luck as your write up conclusion seemed to say.  I would have hated to lose to luck ;)  </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 729236,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2020-01-25T23:47:55.710000",
          "content": "<p>&gt; Your score isn't due to luck</p>\n\n<p>Definitely. The 5-model ensemble is the key and lowering threshold is a nice touch. Thank you so much for sharing these details.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 728538,
      "author_name": "MaximKa",
      "author_url": "",
      "post_date": "2020-01-24T21:46:16.807000",
      "content": "<p>I''m sorry for a noob question, but what is \"hard negative sampling\"?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 728565,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-01-24T22:52:03.843000",
          "content": "<p>It's similar to hard negative mining used in computer vision object detection</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 728693,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2020-01-25T05:08:25.373000",
          "content": "<p>Basically you train a baseline model and use it to predict all candidates in your validation data. Then we can look at the wrong predictions made by the baseline model, especially the negative candidate which have no answers within them but get high scores by the model. They have huge logloss indicating that it is hard for the model to discern them from positive candidates. Hence these candidates are \"hard negative candidates\". So in the second pass of training, we can just train the model with these hard negative candidates instead of randomly sampled negative candidates.</p>\n\n<p>Please correct me if I'm wrong.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 728719,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-01-25T06:24:50.657000",
          "content": "<p><strong>Basically you train a baseline model and use it to predict all candidates in your validation data.</strong>\nWhat I did was I trained the model on the training set and used the trained model to predict on the same training set. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 728512,
      "author_name": "Mourad",
      "author_url": "",
      "post_date": "2020-01-24T20:31:41.723000",
      "content": "<p>Hey <a href=\"/wowfattie\">@wowfattie</a>, big thank you for sharing this, I am really impressed that you used ALBERT-XXL when lots of people said that it timed out? May I know how did you go around that?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 728522,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-01-24T20:56:36.460000",
          "content": "<p>Please see my answers here\n<a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/127551#728519\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/127551#728519</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 728536,
          "author_name": "Mourad",
          "author_url": "",
          "post_date": "2020-01-24T21:40:52.037000",
          "content": "<p>Thanks plenty <a href=\"/wowfattie\">@wowfattie</a>!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 728447,
      "author_name": "James Ingram",
      "author_url": "",
      "post_date": "2020-01-24T18:21:17.717000",
      "content": "<p>Congrats! Thanks for sharing insights and thought process.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728425,
      "author_name": "xiao-xiao",
      "author_url": "",
      "post_date": "2020-01-24T17:51:37.893000",
      "content": "<p>Congrats! That is indeed an amazing solution! This is the best gift to celebrate Chinese New Year!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728417,
      "author_name": "Akhilesh",
      "author_url": "",
      "post_date": "2020-01-24T17:40:50.307000",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a> Woah! Mindblown!🙌 🙌 🙌 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728394,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-24T17:17:14.620000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 730363,
      "author_name": "lior perlmutter shoshany",
      "author_url": "",
      "post_date": "2020-01-27T11:42:24.713000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 730343,
      "author_name": "Aravind",
      "author_url": "",
      "post_date": "2020-01-27T11:19:49.557000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 734228,
      "author_name": "Victor Butoi",
      "author_url": "",
      "post_date": "2020-02-01T05:56:18.597000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 733293,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-01-30T23:46:08.330000",
      "content": "<p>Congrats and thank you for writing up.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 730396,
      "author_name": "Oğuzhan Duran",
      "author_url": "",
      "post_date": "2020-01-27T12:40:52.403000",
      "content": "<p>thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 730076,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-27T04:26:20.767000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 729064,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-25T17:24:39.330000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728498,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-24T19:57:46.377000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 730004,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-27T00:50:39.093000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "728384": "# **04/01/2020: Source code is attached below.**\n\nThanks to the Kaggle and Tensorflow team for holding this competition. I was new to question answering, it took me more than 5 weeks to make my first real submission, and I have learnt a lot during the journey. My initial plan before joining the competition was to learn both QA and TF2.0, but in the end I didn't have time to touch TF2.0, so my solution stays in pure pytorch. Thanks to @sakami for the great kernel https://www.kaggle.com/sakami/tfqa-pytorch-baseline. Your kernel was the starting point of my journey. And of course thanks to huggingface (https://github.com/huggingface/transformers), NLP finetuning is made much easier. \n\nMy solution is described below.\n\n### **- Overview**\nI trained on the provided candidates instead of sampling from the original documents (examples) as done in the baseline paper (https://arxiv.org/abs/1901.08634). Since there are a total of 40 million candidates in the training data, for each epoch, I sampled only one negative candidate from each document. For more efficient training, hard negative sampling was used to replace uniform random sampling. The final submission was an ensemble of five models. \n\n### **- Sampling Strategy**\nInitially, I tried uniform sampling on the negative candidates, but the result was unsatisfactory. The reason might be that most of the negative candidates are \"too easy\", the model might only need to learn some \"basic\" patterns for good candidate-level classification performance. But in the testing stage our actual goal is to predict the most probable positive candidate from each document, and this document-level classification is a more difficult task. So I replaced the uniform sampling by hard negative sampling to increase the difficulty of the candidate-level training, as expected, the performance was greatly improved. To perform hard negative sampling in the following models, I firstly trained a model with uniform sampling, and predicted on the whole training data, and stored the answer probability for each negative candidate. The last step was to normalize the probabilities of negative candidates within documents to form a distribution. For the following model training the negative candidates could be sampled from the probability distribution.\n\n### **- New Tokens**\nAccording to the baseline paper, I added html tags as new tokens for better model performance. All the 9 tags from the Data Statistics Section of https://github.com/google-research-datasets/natural-questions was added. For html tags that are not in the 9 added tokens, I replaced them with a unique token in the tokenization dictionary or simply addedanother new token to represent them. I did not have time to try adding paragraph or table number similar to what the baseline paper does.\n\n### **- Model Architecture, Training and Evaluation**\nOverall, the model architecture was the same as the baseline paper (a 5 class classification branch + 2 span classification branch). The five classes was \"no_answer\", \"long_answer_only\", \"short_answer\", \"yes\", \"no\". In my case there was no span prediction for answers without a short answer span because I directly used candidates. The loss update of the span prediction branch was simply ignored if no short answer span exist during training. In testing stage, for each document, I used 1.0-prob(no_answer) as the long answer score (confidence) for each candidate, and the candidate with the highest confidence was chosen to represent the document. Short answer spans were forced to be within the highest score long answer candidate (not sure if this is necessary). I used prob(short_answer)+prob(yes)+prob(no) as the short answer score. The exact class of the short answer was determined by the maximum of the three prob values. For span prediction, the output token-level probabilities were mapped to the word-level (white space tokenized) probabilities for easier ensembling of models with different tokenizers. \n\n### **- Models and Results**\nMy final submission was an ensemble of one Bert-base, two Bert-large (WWM), and two Albert-xxl (v2) models, all uncased. The Bert large and Albert models had been tuned on the SQUAD data before training. Below list their validation performance on the dev set using the code https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py. I did not try to implement the competition metric.\n\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; long-best-threshold-f1&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; short-best-threshold-f1\nBert-base&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;0.618&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.457\nBert-large &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;0.679&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.541\nAlbert-xxl&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;0.700&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.555\nensemble&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; 0.731&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.582\n\n### **- Final LB results**\nMy best ensemble only achieved 0.66 public LB (0.69 private) performance using the optimized thresholds. At that time I had already lost most of my hope to win. In my last 2-3 submission, I arbitrarily played with the thresholds. One of the submissions scored 0.71 (both public and private LB), and I chose it and won the competition. Unbelievable.",
    "728421": "Amazing job, congrats! You are again showing that tuning models on public LB can help, I have to re-think some of my beliefs 😅 ",
    "729002": "Congrats! I'm glad that I could help you. ;)",
    "728429": "This is really interesting - thanks for the writeup. I like the \"hard negative sampling\" idea - making the training harder to get better results. How much of a bump did this provide, if you've kept track?\n\nCan you say some more about your fine tuning process?\n- Same for each component in the ensemble?\n- What were the stopping conditions?\n- What was the validation set? Why that one?\n\nWhy no cased models?\n\nHow did you fit all those models into the kernel runtime limit? Sounds like a few other competitors got timeouts. Anything special to speed up inference?\n\nHow did your ensembling work? \n- Sounds like you averaged word-level logits from each model (pre softmax)?\n- Have you tried other ensembling methods, eg post-softmax, or even post- compute_pred_dict?\n- Equal weight?\n- How did you pick the final ensemble components among all the models you trained?\n\n\nThanks again - hope you don't mind the many questions, but learning from you guys is one of the big perks of participating in these.\n\nRegards,\nAlon",
    "730206": "Congratss!! Thanks for sharing thats useful kernel !",
    "728415": "Amazing indeed @wowfattie . Would you be releasing your pytorch source code? ",
    "790222": "Congats ",
    "728416": "Thanks for sharing and congrats on solo win!",
    "730415": "Good Try Keep it up",
    "729745": "Awesome solution and a motivating journey. Thanks for sharing.\nJust one minor detail, the **bert** paper link is broken, there is one extra `)` at the end. ",
    "729425": "Congratulations! Any plans to release your pytorch training code? Given that there are not much pytorch solutions it would be great to see how the winner used pytorch in the solution! ",
    "739420": "@wowfattie, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques) ",
    "940206": "Thanks for sharing! To understand your workflow better, did you implement your solutions in Notebooks first and then migrated it to script or do you prefer starting with script code right away?",
    "736816": "That's a brilliant solution. Congratulations and thanks for sharing!!",
    "733642": "unbelievable. victory. congraturation ^^/ good job. ",
    "733176": "Ty, gj :)",
    "732978": "Congratulations! @wowfattie ",
    "732854": "Congrats.. Winning streak continues.. @wowfattie \n\nKaggle - Review of Year 2019\n\n[https://youtu.be/pL96IPZZ-88](https://youtu.be/pL96IPZZ-88)\n\nCheck out the video above.\n\n",
    "728687": "&gt;  In my last 2-3 submission, I arbitrarily played with the thresholds. One of the submissions scored 0.71 (both public and private LB), and I chose it and won the competition. Unbelievable.\n\nCould you share the heuristics of tuning threshold? Is it grid search, random search or something else?",
    "728538": "I''m sorry for a noob question, but what is \"hard negative sampling\"?",
    "728512": "Hey @wowfattie, big thank you for sharing this, I am really impressed that you used ALBERT-XXL when lots of people said that it timed out? May I know how did you go around that?",
    "728447": "Congrats! Thanks for sharing insights and thought process.",
    "728425": "Congrats! That is indeed an amazing solution! This is the best gift to celebrate Chinese New Year!",
    "728417": "@wowfattie Woah! Mindblown!🙌 🙌 🙌 ",
    "728394": "",
    "730363": "Thanks for sharing",
    "730343": "Thanks for sharing",
    "734228": "Thanks for sharing!",
    "733293": "Congrats and thank you for writing up.",
    "730396": "thanks for sharing\n",
    "730076": "Thanks for sharing\n",
    "729064": "Thank you for sharing.",
    "728498": "Congrats! Thanks for sharing! ",
    "730004": "Congrats! Thanks for sharing!"
  }
}