{
  "id": 127374,
  "title": "23rd place solution: ensemble, rank passage and predict span",
  "url": "/competitions/tensorflow2-question-answering/writeups/were-it-so-easy-23rd-place-solution-ensemble-rank-",
  "author_name": "",
  "post_date": "2020-01-23T15:42:54.600Z",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Firstly, thanks Kaggle for great challenging competition, and congrats winners.\nThis is my first time to handle QA task, so I could learn a lot of things from you.</p>\n\n<p>Secondly, thanks you all kagglers those had many discussions and ideas here, especially <a href=\"/christofhenkel\">@christofhenkel</a>, <a href=\"/boliu0\">@boliu0</a>, <a href=\"/higepon\">@higepon</a> and <a href=\"/kashnitsky\">@kashnitsky</a> .</p>\n\n<p>This is my brief solution. I welcome questions and advises.</p>\n\n<h1>My Solution</h1>\n\n<p>Combining 1 ranker model &amp; 4 span prediction ensemble model.</p>\n\n<h2>Whole Prediction</h2>\n\n<p>1) compute passage score for all long answer candidates on test dataset\n2) select top 10 score passages for each record\n3) feed selected passage into span prediction models\n4) get averaged score by each model</p>\n\n<h2>Ranker Model</h2>\n\n<p>One of problems is that NQ dataset has so many candidates for long answer. These include obviously negative passages and takes much time to predict for all.</p>\n\n<p>I used <code>bert-base-uncased</code> pre-trained and construct binary classification model to predict the passage is including long/short answer or not.\nIt get abount 0.98 recall@10 score on my validation dataset.\nIt takes about 5 minutes for public test dataset.\nOther settings are same as span prediction model.\nBy this model, I make <code>ranker-selected</code> dataset for train and test data by selecting top 10 score candidate for each record.</p>\n\n<h2>Span Prediction Model</h2>\n\n<p>I used 4 models for ensemble:</p>\n\n<ol>\n<li><code>bert-large-uncased-squad1</code> pre-trained + 1 epoch on NQ dataset</li>\n<li><code>bert-large-uncased-squad2</code> pre-trained + 1 epoch on NQ dataset</li>\n<li><code>spanbert-large-cased-squad2</code> pre-trained + 1 epoch on NQ dataset</li>\n<li><code>bert-large-uncased-squad2</code> pre-trained + 1 epoch on <code>ranker-selected</code> NQ dataset</li>\n</ol>\n\n<p>All models are bert-joint based model.</p>\n\n<h3>Training</h3>\n\n<p>For 1st~3rd models, I use whole NQ dataset.\nIn training, as reported in <a href=\"https://arxiv.org/abs/1909.05286\">Frustratingly Easy Natural Question Answering</a>, I use 196 as stride, different down sampling rate for answerable and non-answerable question (each 0.01, 0.04).\nBatch size is 32, max learning rate is 3e-5.\nFor 4th model, I used <code>ranker-selected</code> NQ training dataset and adjust sampling rate to 0.03 for answerable and 0.12 for unanswerable.</p>\n\n<h3>Inference</h3>\n\n<p>I used only <code>ranker-selected</code> test dataset. This makes slight improvement on val score than predicting all candidates, and what is more important, this makes faster prediction.\nIt takes about 3 minute for each model prediction.\nI get 0.65 private LB score by single model, 0.67 private LB score by ensemble model.</p>\n\n<h1>Trials which didn't works for me</h1>\n\n<ul>\n<li>using albert, xlnet didn't improve scores. maybe I need more tuning.</li>\n<li>Attention over Attention didn't affect positively. But I don't have confidence for implementation.</li>\n<li>BERT layer combination on last 2, 4, 8, 12 layers. It slightly improve but pre-trained by squad was better.</li>\n<li>kinds of dropout on dense layer.</li>\n<li>label smoothing on start and end position.</li>\n<li>combine ranker model score into span prediction lead worse result.</li>\n<li>dividing short and long span prediction, or only predict short spans get worse result.</li>\n<li>kinds of preprocessing\n<ul><li>no special token</li>\n<li>partly use special token in BERT-joint</li></ul></li>\n<li>kinds of postprocessing\n<ul><li>use only max context position as score</li>\n<li>get all logits score and obtain top k candidates</li></ul></li>\n</ul>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "727216",
      "postDate": "01/23/2020 14:34:35",
      "content": "<p>Firstly, thanks Kaggle for great challenging competition, and congrats winners.\nThis is my first time to handle QA task, so I could learn a lot of things from you.</p>\n\n<p>Secondly, thanks you all kagglers those had many discussions and ideas here, especially <a href=\"/christofhenkel\">@christofhenkel</a>, <a href=\"/boliu0\">@boliu0</a>, <a href=\"/higepon\">@higepon</a> and <a href=\"/kashnitsky\">@kashnitsky</a> .</p>\n\n<p>This is my brief solution. I welcome questions and advises.</p>\n\n<h1>My Solution</h1>\n\n<p>Combining 1 ranker model &amp; 4 span prediction ensemble model.</p>\n\n<h2>Whole Prediction</h2>\n\n<p>1) compute passage score for all long answer candidates on test dataset\n2) select top 10 score passages for each record\n3) feed selected passage into span prediction models\n4) get averaged score by each model</p>\n\n<h2>Ranker Model</h2>\n\n<p>One of problems is that NQ dataset has so many candidates for long answer. These include obviously negative passages and takes much time to predict for all.</p>\n\n<p>I used <code>bert-base-uncased</code> pre-trained and construct binary classification model to predict the passage is including long/short answer or not.\nIt get abount 0.98 recall@10 score on my validation dataset.\nIt takes about 5 minutes for public test dataset.\nOther settings are same as span prediction model.\nBy this model, I make <code>ranker-selected</code> dataset for train and test data by selecting top 10 score candidate for each record.</p>\n\n<h2>Span Prediction Model</h2>\n\n<p>I used 4 models for ensemble:</p>\n\n<ol>\n<li><code>bert-large-uncased-squad1</code> pre-trained + 1 epoch on NQ dataset</li>\n<li><code>bert-large-uncased-squad2</code> pre-trained + 1 epoch on NQ dataset</li>\n<li><code>spanbert-large-cased-squad2</code> pre-trained + 1 epoch on NQ dataset</li>\n<li><code>bert-large-uncased-squad2</code> pre-trained + 1 epoch on <code>ranker-selected</code> NQ dataset</li>\n</ol>\n\n<p>All models are bert-joint based model.</p>\n\n<h3>Training</h3>\n\n<p>For 1st~3rd models, I use whole NQ dataset.\nIn training, as reported in <a href=\"https://arxiv.org/abs/1909.05286\">Frustratingly Easy Natural Question Answering</a>, I use 196 as stride, different down sampling rate for answerable and non-answerable question (each 0.01, 0.04).\nBatch size is 32, max learning rate is 3e-5.\nFor 4th model, I used <code>ranker-selected</code> NQ training dataset and adjust sampling rate to 0.03 for answerable and 0.12 for unanswerable.</p>\n\n<h3>Inference</h3>\n\n<p>I used only <code>ranker-selected</code> test dataset. This makes slight improvement on val score than predicting all candidates, and what is more important, this makes faster prediction.\nIt takes about 3 minute for each model prediction.\nI get 0.65 private LB score by single model, 0.67 private LB score by ensemble model.</p>\n\n<h1>Trials which didn't works for me</h1>\n\n<ul>\n<li>using albert, xlnet didn't improve scores. maybe I need more tuning.</li>\n<li>Attention over Attention didn't affect positively. But I don't have confidence for implementation.</li>\n<li>BERT layer combination on last 2, 4, 8, 12 layers. It slightly improve but pre-trained by squad was better.</li>\n<li>kinds of dropout on dense layer.</li>\n<li>label smoothing on start and end position.</li>\n<li>combine ranker model score into span prediction lead worse result.</li>\n<li>dividing short and long span prediction, or only predict short spans get worse result.</li>\n<li>kinds of preprocessing\n<ul><li>no special token</li>\n<li>partly use special token in BERT-joint</li></ul></li>\n<li>kinds of postprocessing\n<ul><li>use only max context position as score</li>\n<li>get all logits score and obtain top k candidates</li></ul></li>\n</ul>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Firstly, thanks Kaggle for great challenging competition, and congrats winners.\nThis is my first time to handle QA task, so I could learn a lot of things from you.\n\nSecondly, thanks you all kagglers those had many discussions and ideas here, especially @christofhenkel, @boliu0, @higepon and @kashnitsky .\n\nThis is my brief solution. I welcome questions and advises.\n\n# My Solution\n\nCombining 1 ranker model &amp; 4 span prediction ensemble model.\n\n## Whole Prediction\n\n1) compute passage score for all long answer candidates on test dataset\n2) select top 10 score passages for each record\n3) feed selected passage into span prediction models\n4) get averaged score by each model\n\n## Ranker Model\n\nOne of problems is that NQ dataset has so many candidates for long answer. These include obviously negative passages and takes much time to predict for all.\n\nI used `bert-base-uncased` pre-trained and construct binary classification model to predict the passage is including long/short answer or not.\nIt get abount 0.98 recall@10 score on my validation dataset.\nIt takes about 5 minutes for public test dataset.\nOther settings are same as span prediction model.\nBy this model, I make `ranker-selected` dataset for train and test data by selecting top 10 score candidate for each record.\n\n## Span Prediction Model\n\nI used 4 models for ensemble:\n\n1. `bert-large-uncased-squad1` pre-trained + 1 epoch on NQ dataset\n2. `bert-large-uncased-squad2` pre-trained + 1 epoch on NQ dataset\n3. `spanbert-large-cased-squad2` pre-trained + 1 epoch on NQ dataset\n4. `bert-large-uncased-squad2` pre-trained + 1 epoch on `ranker-selected` NQ dataset\n\nAll models are bert-joint based model.\n\n### Training\n\nFor 1st~3rd models, I use whole NQ dataset.\nIn training, as reported in [Frustratingly Easy Natural Question Answering](https://arxiv.org/abs/1909.05286), I use 196 as stride, different down sampling rate for answerable and non-answerable question (each 0.01, 0.04).\nBatch size is 32, max learning rate is 3e-5.\nFor 4th model, I used `ranker-selected` NQ training dataset and adjust sampling rate to 0.03 for answerable and 0.12 for unanswerable.\n\n### Inference\n\nI used only `ranker-selected` test dataset. This makes slight improvement on val score than predicting all candidates, and what is more important, this makes faster prediction.\nIt takes about 3 minute for each model prediction.\nI get 0.65 private LB score by single model, 0.67 private LB score by ensemble model.\n\n# Trials which didn't works for me\n\n- using albert, xlnet didn't improve scores. maybe I need more tuning.\n- Attention over Attention didn't affect positively. But I don't have confidence for implementation.\n- BERT layer combination on last 2, 4, 8, 12 layers. It slightly improve but pre-trained by squad was better.\n- kinds of dropout on dense layer.\n- label smoothing on start and end position.\n- combine ranker model score into span prediction lead worse result.\n- dividing short and long span prediction, or only predict short spans get worse result.\n- kinds of preprocessing\n    - no special token\n    - partly use special token in BERT-joint\n- kinds of postprocessing\n    - use only max context position as score\n    - get all logits score and obtain top k candidates\n\nThanks.",
      "votes": null
    },
    {
      "id": "727752",
      "postDate": "01/24/2020 02:22:05",
      "content": "<p>Congratulations\nThanks for Sharing your Approach &amp; Insights!! <a href=\"/kentaronakanishi\">@kentaronakanishi</a> </p>",
      "rawMarkdown": "Congratulations\nThanks for Sharing your Approach &amp; Insights!! @kentaronakanishi",
      "votes": null
    },
    {
      "id": "728288",
      "postDate": "01/24/2020 15:10:33",
      "content": "<p>Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 </p>",
      "rawMarkdown": "Congrats &amp; Thanks for sharing your solutions🎉 😄 👍",
      "votes": null
    },
    {
      "id": "733312",
      "postDate": "01/31/2020 00:23:03",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 727752,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "01/24/2020 02:22:05",
      "content": "<p>Congratulations\nThanks for Sharing your Approach &amp; Insights!! <a href=\"/kentaronakanishi\">@kentaronakanishi</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 728288,
      "author_name": "mashlyn",
      "author_url": "",
      "post_date": "01/24/2020 15:10:33",
      "content": "<p>Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 733312,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "01/31/2020 00:23:03",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "727216": "Firstly, thanks Kaggle for great challenging competition, and congrats winners.\nThis is my first time to handle QA task, so I could learn a lot of things from you.\n\nSecondly, thanks you all kagglers those had many discussions and ideas here, especially @christofhenkel, @boliu0, @higepon and @kashnitsky .\n\nThis is my brief solution. I welcome questions and advises.\n\n# My Solution\n\nCombining 1 ranker model &amp; 4 span prediction ensemble model.\n\n## Whole Prediction\n\n1) compute passage score for all long answer candidates on test dataset\n2) select top 10 score passages for each record\n3) feed selected passage into span prediction models\n4) get averaged score by each model\n\n## Ranker Model\n\nOne of problems is that NQ dataset has so many candidates for long answer. These include obviously negative passages and takes much time to predict for all.\n\nI used `bert-base-uncased` pre-trained and construct binary classification model to predict the passage is including long/short answer or not.\nIt get abount 0.98 recall@10 score on my validation dataset.\nIt takes about 5 minutes for public test dataset.\nOther settings are same as span prediction model.\nBy this model, I make `ranker-selected` dataset for train and test data by selecting top 10 score candidate for each record.\n\n## Span Prediction Model\n\nI used 4 models for ensemble:\n\n1. `bert-large-uncased-squad1` pre-trained + 1 epoch on NQ dataset\n2. `bert-large-uncased-squad2` pre-trained + 1 epoch on NQ dataset\n3. `spanbert-large-cased-squad2` pre-trained + 1 epoch on NQ dataset\n4. `bert-large-uncased-squad2` pre-trained + 1 epoch on `ranker-selected` NQ dataset\n\nAll models are bert-joint based model.\n\n### Training\n\nFor 1st~3rd models, I use whole NQ dataset.\nIn training, as reported in [Frustratingly Easy Natural Question Answering](https://arxiv.org/abs/1909.05286), I use 196 as stride, different down sampling rate for answerable and non-answerable question (each 0.01, 0.04).\nBatch size is 32, max learning rate is 3e-5.\nFor 4th model, I used `ranker-selected` NQ training dataset and adjust sampling rate to 0.03 for answerable and 0.12 for unanswerable.\n\n### Inference\n\nI used only `ranker-selected` test dataset. This makes slight improvement on val score than predicting all candidates, and what is more important, this makes faster prediction.\nIt takes about 3 minute for each model prediction.\nI get 0.65 private LB score by single model, 0.67 private LB score by ensemble model.\n\n# Trials which didn't works for me\n\n- using albert, xlnet didn't improve scores. maybe I need more tuning.\n- Attention over Attention didn't affect positively. But I don't have confidence for implementation.\n- BERT layer combination on last 2, 4, 8, 12 layers. It slightly improve but pre-trained by squad was better.\n- kinds of dropout on dense layer.\n- label smoothing on start and end position.\n- combine ranker model score into span prediction lead worse result.\n- dividing short and long span prediction, or only predict short spans get worse result.\n- kinds of preprocessing\n    - no special token\n    - partly use special token in BERT-joint\n- kinds of postprocessing\n    - use only max context position as score\n    - get all logits score and obtain top k candidates\n\nThanks.",
    "727752": "Congratulations\nThanks for Sharing your Approach &amp; Insights!! @kentaronakanishi",
    "728288": "Congrats &amp; Thanks for sharing your solutions🎉 😄 👍",
    "733312": "Thanks for sharing"
  },
  "source": "meta"
}