{
  "id": 127545,
  "title": "8th place solution",
  "url": "/competitions/tensorflow2-question-answering/discussion/127545",
  "author_name": "Oleg Platonov",
  "post_date": "2020-01-24T16:12:27.269000",
  "votes": 30,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, I'd like to thank the organizers for this interesting competition with a not-so-standard dataset. Question answering is one of the most fun tasks in NLP and it is great to finally see it on Kaggle. Also, thanks for providing GCP credits and TPU quota to the participants. I'm sure it greatly widened the range of models and ideas that were tested in this competition.</p>\n\n<p>My work started as an attempt to reimplement BERT-joint in PyTorch using RoBERTa as the backbone. However, I added quite a few tricks and tweaks along the way and ended up with a model and pipeline significantly different from the original BERT-joint. Here I'm going to describe the main changes.</p>\n\n<h2>Preprocessing</h2>\n\n<p>Instead of sliding a window over the entire Wikipedia article, I processed each top level long answer candidate separately. For each candidate, I either converted it into one training example if its length permitted it or split it into several training examples if the candidate was too long. I also added some of the surrounding context to those candidates that were particularly short.</p>\n\n<p>The above preprocessing resulted in approximately 152k positive and more than 12 million negative (not containing any answer) examples, so I decreased the number of negative examples to 160k by random sampling. I used a kind of hard negative mining strategy by sampling more of those negative examples that have high TF-IDF similarity between the question and the candidate. I also sampled several non-overlapping subsets of negative examples to use for different epochs of training thus increasing the diversity of my training data.</p>\n\n<h2>Model</h2>\n\n<p>My model is just RoBERTa-large (I use the implementation from Transformers library) with a new output layer on top of it. In addition to a token-level span predictor for short answers, I use a binary classifier to determine whether a candidate is a long answer or not. The combination of an answerability classifier and a span predictor is a standard approach for SQUAD2.0 (XLNet, RoBERTa, ALBERT all use it). NQ dataset differs from SQuAD2.0 in that a question can be considered answerable even when the correct short answer span is empty (this happens when a question has a long answer, but no short answer).</p>\n\n<p>For span predictor, I use a trick from XLNet: instead of predicting start and end tokens independently, I first predict the start token, then concatenate its representation from the final encoder layer to representations of all the tokens and pass these concatenated representations as input to the end token predictor. This means that the prediction of the end token is conditioned on the start token, which significantly improves the quality of span prediction. </p>\n\n<p>I did not find a way to include YES/NO answers in my predictions without a decrease in the total score so I chose not to predict such answers.</p>\n\n<p>During inference, I first find the long answer candidate that has the highest answerability score. If this score is above a certain threshold, I choose this candidate as my long answer prediction and predict a short span for this candidate. If this span's score is also above a certain threshold, I choose it as my short answer prediction. I used the official NQ dev set to find the best thresholds for both long and short answers.</p>\n\n<h2>Training hyperparameters</h2>\n\n<p>I used AdamW optimizer with weight decay of 0.01 and a linearly decaying learning rate with warmup for all experiments. I had neither time nor computational resources to try a wide range of hyperparameters so the ones I've chosen can be far from optimal. I got the best results on the dev set with a model trained for 5 epochs with a batch size of 48 and a maximum learning rate of 2e-5. I used this model for one of my final submissions. I also had two other models with good results that I later used for ensembling: one was trained for 3 epochs with a batch size of 24 and a maximum learning rate of 3e-5 and the other was trained for 2 epochs with a batch size of 15 and a maximum learning rate of 3e-5.</p>\n\n<p>Training RoBERTa-large for 1 epoch (312k training examples) takes approximately 4 hours on a single V100 GPU using mixed precision.</p>\n\n<h2>Ensembling</h2>\n\n<p>For my second final submission, I ensembled three models by simply summing their output layer logits. This approach led to a significant improvement on the dev set, but it could not fit in the submission time limit. In order to fix it, I decided to limit the number of long answer candidates per question by taking only the first N candidates (most answers are found in the first few paragraphs anyway). However, when my final models and ensembling code were ready, I only had five hours before the deadline and two submissions left so I did not have a chance to select the maximum value of N that will allow my submission to fit within the time limit. I ended up choosing too small of a value for N which probably harmed the performance of my ensemble. In hindsight, it seems a better approach could be to score all candidates with just one model and then use the other two models only for several candidates that got the highest answerability scores from the first model.</p>\n\n<p>In the end, all three of my main models, as well as the ensemble, got a score of 0.68 on the private test set. Well, at least I got stable results.</p>\n\n<h2>Some ideas that did not quite work</h2>\n\n<ul>\n<li><p>While SQuAD2.0 pretraining seemed beneficial in my early experiments, it harmed the performance of my final models so I ended up not using it. I suspect that while changing the output layer architecture I might have introduced some bugs in the SQuAD pretraining code. It explains why many other participants, as well as several papers about the NQ dataset, report improvements from SQuAD pretraining.</p></li>\n<li><p>I tried adding one more binary classifier to determine whether a candidate contains a short answer or not, but it did not lead to an improvement on the dev set. Now I can see that my early submission with this additional classifier got a slightly higher score on the private test set than a similar submission without it, so it might have been a useful idea after all. </p></li>\n</ul>",
  "messages": [
    {
      "id": 728332,
      "postDate": "2020-01-24T16:12:27.270Z",
      "content": "<p>First of all, I'd like to thank the organizers for this interesting competition with a not-so-standard dataset. Question answering is one of the most fun tasks in NLP and it is great to finally see it on Kaggle. Also, thanks for providing GCP credits and TPU quota to the participants. I'm sure it greatly widened the range of models and ideas that were tested in this competition.</p>\n\n<p>My work started as an attempt to reimplement BERT-joint in PyTorch using RoBERTa as the backbone. However, I added quite a few tricks and tweaks along the way and ended up with a model and pipeline significantly different from the original BERT-joint. Here I'm going to describe the main changes.</p>\n\n<h2>Preprocessing</h2>\n\n<p>Instead of sliding a window over the entire Wikipedia article, I processed each top level long answer candidate separately. For each candidate, I either converted it into one training example if its length permitted it or split it into several training examples if the candidate was too long. I also added some of the surrounding context to those candidates that were particularly short.</p>\n\n<p>The above preprocessing resulted in approximately 152k positive and more than 12 million negative (not containing any answer) examples, so I decreased the number of negative examples to 160k by random sampling. I used a kind of hard negative mining strategy by sampling more of those negative examples that have high TF-IDF similarity between the question and the candidate. I also sampled several non-overlapping subsets of negative examples to use for different epochs of training thus increasing the diversity of my training data.</p>\n\n<h2>Model</h2>\n\n<p>My model is just RoBERTa-large (I use the implementation from Transformers library) with a new output layer on top of it. In addition to a token-level span predictor for short answers, I use a binary classifier to determine whether a candidate is a long answer or not. The combination of an answerability classifier and a span predictor is a standard approach for SQUAD2.0 (XLNet, RoBERTa, ALBERT all use it). NQ dataset differs from SQuAD2.0 in that a question can be considered answerable even when the correct short answer span is empty (this happens when a question has a long answer, but no short answer).</p>\n\n<p>For span predictor, I use a trick from XLNet: instead of predicting start and end tokens independently, I first predict the start token, then concatenate its representation from the final encoder layer to representations of all the tokens and pass these concatenated representations as input to the end token predictor. This means that the prediction of the end token is conditioned on the start token, which significantly improves the quality of span prediction. </p>\n\n<p>I did not find a way to include YES/NO answers in my predictions without a decrease in the total score so I chose not to predict such answers.</p>\n\n<p>During inference, I first find the long answer candidate that has the highest answerability score. If this score is above a certain threshold, I choose this candidate as my long answer prediction and predict a short span for this candidate. If this span's score is also above a certain threshold, I choose it as my short answer prediction. I used the official NQ dev set to find the best thresholds for both long and short answers.</p>\n\n<h2>Training hyperparameters</h2>\n\n<p>I used AdamW optimizer with weight decay of 0.01 and a linearly decaying learning rate with warmup for all experiments. I had neither time nor computational resources to try a wide range of hyperparameters so the ones I've chosen can be far from optimal. I got the best results on the dev set with a model trained for 5 epochs with a batch size of 48 and a maximum learning rate of 2e-5. I used this model for one of my final submissions. I also had two other models with good results that I later used for ensembling: one was trained for 3 epochs with a batch size of 24 and a maximum learning rate of 3e-5 and the other was trained for 2 epochs with a batch size of 15 and a maximum learning rate of 3e-5.</p>\n\n<p>Training RoBERTa-large for 1 epoch (312k training examples) takes approximately 4 hours on a single V100 GPU using mixed precision.</p>\n\n<h2>Ensembling</h2>\n\n<p>For my second final submission, I ensembled three models by simply summing their output layer logits. This approach led to a significant improvement on the dev set, but it could not fit in the submission time limit. In order to fix it, I decided to limit the number of long answer candidates per question by taking only the first N candidates (most answers are found in the first few paragraphs anyway). However, when my final models and ensembling code were ready, I only had five hours before the deadline and two submissions left so I did not have a chance to select the maximum value of N that will allow my submission to fit within the time limit. I ended up choosing too small of a value for N which probably harmed the performance of my ensemble. In hindsight, it seems a better approach could be to score all candidates with just one model and then use the other two models only for several candidates that got the highest answerability scores from the first model.</p>\n\n<p>In the end, all three of my main models, as well as the ensemble, got a score of 0.68 on the private test set. Well, at least I got stable results.</p>\n\n<h2>Some ideas that did not quite work</h2>\n\n<ul>\n<li><p>While SQuAD2.0 pretraining seemed beneficial in my early experiments, it harmed the performance of my final models so I ended up not using it. I suspect that while changing the output layer architecture I might have introduced some bugs in the SQuAD pretraining code. It explains why many other participants, as well as several papers about the NQ dataset, report improvements from SQuAD pretraining.</p></li>\n<li><p>I tried adding one more binary classifier to determine whether a candidate contains a short answer or not, but it did not lead to an improvement on the dev set. Now I can see that my early submission with this additional classifier got a slightly higher score on the private test set than a similar submission without it, so it might have been a useful idea after all. </p></li>\n</ul>",
      "rawMarkdown": "First of all, I'd like to thank the organizers for this interesting competition with a not-so-standard dataset. Question answering is one of the most fun tasks in NLP and it is great to finally see it on Kaggle. Also, thanks for providing GCP credits and TPU quota to the participants. I'm sure it greatly widened the range of models and ideas that were tested in this competition.\n\nMy work started as an attempt to reimplement BERT-joint in PyTorch using RoBERTa as the backbone. However, I added quite a few tricks and tweaks along the way and ended up with a model and pipeline significantly different from the original BERT-joint. Here I'm going to describe the main changes.\n\n\n## Preprocessing\n\nInstead of sliding a window over the entire Wikipedia article, I processed each top level long answer candidate separately. For each candidate, I either converted it into one training example if its length permitted it or split it into several training examples if the candidate was too long. I also added some of the surrounding context to those candidates that were particularly short.\n\nThe above preprocessing resulted in approximately 152k positive and more than 12 million negative (not containing any answer) examples, so I decreased the number of negative examples to 160k by random sampling. I used a kind of hard negative mining strategy by sampling more of those negative examples that have high TF-IDF similarity between the question and the candidate. I also sampled several non-overlapping subsets of negative examples to use for different epochs of training thus increasing the diversity of my training data.\n\n\n## Model\n\nMy model is just RoBERTa-large (I use the implementation from Transformers library) with a new output layer on top of it. In addition to a token-level span predictor for short answers, I use a binary classifier to determine whether a candidate is a long answer or not. The combination of an answerability classifier and a span predictor is a standard approach for SQUAD2.0 (XLNet, RoBERTa, ALBERT all use it). NQ dataset differs from SQuAD2.0 in that a question can be considered answerable even when the correct short answer span is empty (this happens when a question has a long answer, but no short answer).\n\nFor span predictor, I use a trick from XLNet: instead of predicting start and end tokens independently, I first predict the start token, then concatenate its representation from the final encoder layer to representations of all the tokens and pass these concatenated representations as input to the end token predictor. This means that the prediction of the end token is conditioned on the start token, which significantly improves the quality of span prediction. \n\nI did not find a way to include YES/NO answers in my predictions without a decrease in the total score so I chose not to predict such answers.\n\nDuring inference, I first find the long answer candidate that has the highest answerability score. If this score is above a certain threshold, I choose this candidate as my long answer prediction and predict a short span for this candidate. If this span's score is also above a certain threshold, I choose it as my short answer prediction. I used the official NQ dev set to find the best thresholds for both long and short answers.\n\n\n## Training hyperparameters\n\nI used AdamW optimizer with weight decay of 0.01 and a linearly decaying learning rate with warmup for all experiments. I had neither time nor computational resources to try a wide range of hyperparameters so the ones I've chosen can be far from optimal. I got the best results on the dev set with a model trained for 5 epochs with a batch size of 48 and a maximum learning rate of 2e-5. I used this model for one of my final submissions. I also had two other models with good results that I later used for ensembling: one was trained for 3 epochs with a batch size of 24 and a maximum learning rate of 3e-5 and the other was trained for 2 epochs with a batch size of 15 and a maximum learning rate of 3e-5.\n\nTraining RoBERTa-large for 1 epoch (312k training examples) takes approximately 4 hours on a single V100 GPU using mixed precision.\n\n\n## Ensembling\n\nFor my second final submission, I ensembled three models by simply summing their output layer logits. This approach led to a significant improvement on the dev set, but it could not fit in the submission time limit. In order to fix it, I decided to limit the number of long answer candidates per question by taking only the first N candidates (most answers are found in the first few paragraphs anyway). However, when my final models and ensembling code were ready, I only had five hours before the deadline and two submissions left so I did not have a chance to select the maximum value of N that will allow my submission to fit within the time limit. I ended up choosing too small of a value for N which probably harmed the performance of my ensemble. In hindsight, it seems a better approach could be to score all candidates with just one model and then use the other two models only for several candidates that got the highest answerability scores from the first model.\n\nIn the end, all three of my main models, as well as the ensemble, got a score of 0.68 on the private test set. Well, at least I got stable results.\n\n\n## Some ideas that did not quite work\n\n- While SQuAD2.0 pretraining seemed beneficial in my early experiments, it harmed the performance of my final models so I ended up not using it. I suspect that while changing the output layer architecture I might have introduced some bugs in the SQuAD pretraining code. It explains why many other participants, as well as several papers about the NQ dataset, report improvements from SQuAD pretraining.\n\n- I tried adding one more binary classifier to determine whether a candidate contains a short answer or not, but it did not lead to an improvement on the dev set. Now I can see that my early submission with this additional classifier got a slightly higher score on the private test set than a similar submission without it, so it might have been a useful idea after all. ",
      "votes": 30
    },
    {
      "id": 739418,
      "postDate": "2020-02-07T20:16:23.627Z",
      "content": "<p><a href=\"/olegplatonov\">@olegplatonov</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>)</p>",
      "rawMarkdown": "@olegplatonov, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques)"
    },
    {
      "id": 728355,
      "postDate": "2020-01-24T16:32:23.520Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 731252,
      "postDate": "2020-01-28T13:11:22.600Z",
      "content": "<p>Congrats!! Thanks for sharing!</p>",
      "rawMarkdown": "Congrats!! Thanks for sharing!",
      "votes": 4
    },
    {
      "id": 733306,
      "postDate": "2020-01-31T00:15:10.953Z",
      "content": "<p>Congrats 🎉 Thank you for sharing.</p>",
      "rawMarkdown": "Congrats 🎉 Thank you for sharing."
    }
  ],
  "comments": [
    {
      "id": 739418,
      "author_name": "Vitalii Mokin",
      "author_url": "",
      "post_date": "2020-02-07T20:16:23.627000",
      "content": "<p><a href=\"/olegplatonov\">@olegplatonov</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728355,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-24T16:32:23.520000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 731252,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-28T13:11:22.600000",
      "content": "<p>Congrats!! Thanks for sharing!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 733306,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-01-31T00:15:10.953000",
      "content": "<p>Congrats 🎉 Thank you for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "728332": "First of all, I'd like to thank the organizers for this interesting competition with a not-so-standard dataset. Question answering is one of the most fun tasks in NLP and it is great to finally see it on Kaggle. Also, thanks for providing GCP credits and TPU quota to the participants. I'm sure it greatly widened the range of models and ideas that were tested in this competition.\n\nMy work started as an attempt to reimplement BERT-joint in PyTorch using RoBERTa as the backbone. However, I added quite a few tricks and tweaks along the way and ended up with a model and pipeline significantly different from the original BERT-joint. Here I'm going to describe the main changes.\n\n\n## Preprocessing\n\nInstead of sliding a window over the entire Wikipedia article, I processed each top level long answer candidate separately. For each candidate, I either converted it into one training example if its length permitted it or split it into several training examples if the candidate was too long. I also added some of the surrounding context to those candidates that were particularly short.\n\nThe above preprocessing resulted in approximately 152k positive and more than 12 million negative (not containing any answer) examples, so I decreased the number of negative examples to 160k by random sampling. I used a kind of hard negative mining strategy by sampling more of those negative examples that have high TF-IDF similarity between the question and the candidate. I also sampled several non-overlapping subsets of negative examples to use for different epochs of training thus increasing the diversity of my training data.\n\n\n## Model\n\nMy model is just RoBERTa-large (I use the implementation from Transformers library) with a new output layer on top of it. In addition to a token-level span predictor for short answers, I use a binary classifier to determine whether a candidate is a long answer or not. The combination of an answerability classifier and a span predictor is a standard approach for SQUAD2.0 (XLNet, RoBERTa, ALBERT all use it). NQ dataset differs from SQuAD2.0 in that a question can be considered answerable even when the correct short answer span is empty (this happens when a question has a long answer, but no short answer).\n\nFor span predictor, I use a trick from XLNet: instead of predicting start and end tokens independently, I first predict the start token, then concatenate its representation from the final encoder layer to representations of all the tokens and pass these concatenated representations as input to the end token predictor. This means that the prediction of the end token is conditioned on the start token, which significantly improves the quality of span prediction. \n\nI did not find a way to include YES/NO answers in my predictions without a decrease in the total score so I chose not to predict such answers.\n\nDuring inference, I first find the long answer candidate that has the highest answerability score. If this score is above a certain threshold, I choose this candidate as my long answer prediction and predict a short span for this candidate. If this span's score is also above a certain threshold, I choose it as my short answer prediction. I used the official NQ dev set to find the best thresholds for both long and short answers.\n\n\n## Training hyperparameters\n\nI used AdamW optimizer with weight decay of 0.01 and a linearly decaying learning rate with warmup for all experiments. I had neither time nor computational resources to try a wide range of hyperparameters so the ones I've chosen can be far from optimal. I got the best results on the dev set with a model trained for 5 epochs with a batch size of 48 and a maximum learning rate of 2e-5. I used this model for one of my final submissions. I also had two other models with good results that I later used for ensembling: one was trained for 3 epochs with a batch size of 24 and a maximum learning rate of 3e-5 and the other was trained for 2 epochs with a batch size of 15 and a maximum learning rate of 3e-5.\n\nTraining RoBERTa-large for 1 epoch (312k training examples) takes approximately 4 hours on a single V100 GPU using mixed precision.\n\n\n## Ensembling\n\nFor my second final submission, I ensembled three models by simply summing their output layer logits. This approach led to a significant improvement on the dev set, but it could not fit in the submission time limit. In order to fix it, I decided to limit the number of long answer candidates per question by taking only the first N candidates (most answers are found in the first few paragraphs anyway). However, when my final models and ensembling code were ready, I only had five hours before the deadline and two submissions left so I did not have a chance to select the maximum value of N that will allow my submission to fit within the time limit. I ended up choosing too small of a value for N which probably harmed the performance of my ensemble. In hindsight, it seems a better approach could be to score all candidates with just one model and then use the other two models only for several candidates that got the highest answerability scores from the first model.\n\nIn the end, all three of my main models, as well as the ensemble, got a score of 0.68 on the private test set. Well, at least I got stable results.\n\n\n## Some ideas that did not quite work\n\n- While SQuAD2.0 pretraining seemed beneficial in my early experiments, it harmed the performance of my final models so I ended up not using it. I suspect that while changing the output layer architecture I might have introduced some bugs in the SQuAD pretraining code. It explains why many other participants, as well as several papers about the NQ dataset, report improvements from SQuAD pretraining.\n\n- I tried adding one more binary classifier to determine whether a candidate contains a short answer or not, but it did not lead to an improvement on the dev set. Now I can see that my early submission with this additional classifier got a slightly higher score on the private test set than a similar submission without it, so it might have been a useful idea after all. ",
    "739418": "@olegplatonov, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques)",
    "728355": "",
    "731252": "Congrats!! Thanks for sharing!",
    "733306": "Congrats 🎉 Thank you for sharing."
  }
}