{
  "id": 127339,
  "title": "3rd place solution",
  "url": "/competitions/tensorflow2-question-answering/discussion/127339",
  "author_name": "Dieter",
  "post_date": "2020-01-23T11:41:12.130000",
  "votes": 77,
  "comment_count": 30,
  "views": 0,
  "content": "<p>Thanks to the sponsors and kaggle for hosting such an interesting and challenging competition. Also a big thank you to <a href=\"/cpmpml\">@cpmpml</a> for being my teammate.</p>\n\n<h2>Brief Summay</h2>\n\n<p>As our teamname suggests, we did everything with pytorch. In summary, we used 3 roberta-large models which were ensembled by voting. In general input features of our models are very close to bertjoint baseline. We used a learning rate of 1-e5, a batchsize of 16 and simple Adam optimizer with no schedule. All models were trained for 1 epoch.</p>\n\n<p>Roberta 1:\n- initialized with roberta-large weights\n- stride 128\n- prediction of span &amp; 5 answer types (unknown, yes, no, short , long)</p>\n\n<p>Roberta 2:\n- initialized with roberta-large weights, then pretrained on Squad2.0\n- stride 192\n- prediction of span &amp; 2 answer types (short , long)</p>\n\n<p>Roberta 3:\n- initialized with roberta-large weights, then pretrained on Squad2.0\n- additional linear layer (768→768 + relu) before predicting start, respectively end token \n- stride 192\n- prediction of span &amp; 2 answer types (short , long)</p>\n\n<p>We optimized thresholds for each of the models and set predictios below threshold to blank. Then we used majority voting to ensemble the 3 models. Besides some smaller tricks, we predicted test set with a stride of 224 to fit inference of 3 models into the kernel.</p>\n\n<h2>Longer Summary</h2>\n\n<h3>Validation scheme</h3>\n\n<p>As always, I start with setting up a solid validation scheme, which ideally has a high correlation to leaderboard. It turned out harder than anticipated, since organisers did not share enough information on the intended metric as well as implemented it wrongly. This first phase was very frustrating and I spent quite some time reverse engineering their mistake in order to reconcile leaderboard scores. After I figured out the metric and shared in forum, organizers changed the metric. Imagine my face in that moment… and believe it or not, it took me another 6 weeks to figure out the new one. At the end we used the dev set of the original NQ dataset as our validation set and had a very high lb correlation.</p>\n\n<h3>Software</h3>\n\n<p>I reused a lot of preprocessing scripts from bertjoint baseline shared by organisers and did all training with pytorch relying on huggingface for transformer weights and code + pytorch-lightning for writing training pipeline.</p>\n\n<h3>Hardware</h3>\n\n<p>I did all training on my home desktop pc (3 GTX1080Ti) and <a href=\"/cpmpml\">@cpmpml</a> on his pc (2 GTX 1080Ti). Training one epoch took quite a while, hence we did not spent much time on hyper-parameter tuning. The training time for Roberta1 was 35h. Finetuning roberta-large on SQuAD2.0 took 30h and finetuning the resulting model to the data of this competition took about 24h when using a stride of 192.</p>\n\n<h3>Architectures and pretrained models</h3>\n\n<p>I fully agree with <a href=\"/boliu0\">@boliu0</a> that is was frustratingly hard to beat the bertjoint baseline. I did a lot of experiments on different preprocessing as well using different (in my opinion more suited) targets. But 99% of what I did was worse than the baseline. So at the end we kept the preprocessing and only adjusted the answer type targets sightly. I used distilbert for a lot of those experiments because due to its size it helps to iterate fast while giving reasonable indication if an idea works or not. <br>\nIn general the bertjoint baseline suggest to stride over the full answer with a windowing approach and concatenating those windows with the question in order to find if the short answer is contained in the window. One major interesting question is how to aggregate the resulting predictions. Thats where we spent some time because we saw a lot of room for improvement. So what we did is to map the start and end token predictions of each window back to the original answer and create a answer length x answer length heatmap. We then apply some restrictions, like e.g. short span length should be less than 30 tokens, and get the following result.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2F15f5fac1de79c9e7d6d67adb3b45b1de%2FScreenshot%20from%202019-12-21%2009-26-43.png?generation=1579778810842946&amp;alt=media\" alt=\"\"></p>\n\n<p>The argmax of this matrix then gives start and end token (here 957:973). Nice thing of this approach is that you can easily blend these matrices over different models. So after we figured that out we tried different model architectures, including all popular ones from the huggingface repo (albert, gpt2, bert, roberta, xlnet) as well as less popular ones like Spanbert. For us roberta-large worked best with some distance to the second best which was spanbert. Considering the time of preprocessing we thought that ensembling 2 or 3 versions of the same model type will be better than ensembling different model types as you need to do preprocessing only once. So we continued training slightly different versions of roberta-large, including pretraining it on SquAD2.0 first, while working on probably the most important part of this competition, namely thresholding of when to set a blank prediction.</p>\n\n<h3>Thresholding:</h3>\n\n<p>Thresholding when using f1 is challenging. Its super important for your overall score but at the same time has high variance, and might not relate to test set. We used different schemes and at the end a 4-way thresholding worked best. We build thresholds for long and short answer type as well as logits of start + end tokens. We determined the thresholds by simple 4d grid search, which was improved by <a href=\"/cpmpml\">@cpmpml</a>  using scipy.optimize.minimize. Instead of using the thresholds found by fitting on the dev set directly, we also experimented with using the corresponding quantiles. Our best submission uses that approach.</p>\n\n<h3>Ensembling:</h3>\n\n<p>We elaborated different ensemble methods and chose 2 different ones for our final sub:</p>\n\n<ol>\n<li>Apply postprocessing and threshold to model prediction and majority vote between the results</li>\n<li>blend model predictions and apply thresholding</li>\n</ol>\n\n<p>While 2. preformed better on our val set, 1. performed better on public and private LB</p>\n\n<h3>Wrapping things up and putting into kernel:</h3>\n\n<p>We used several things to speed up the final kernel in order to fit the inference of 3 models in.\n- use stride of 224 for test data \n- convert model to fp16 for predictions\n- use multiprocessing for preprocessing and postprocessing</p>\n\n<p>Thanks for reading. </p>",
  "messages": [
    {
      "id": 727020,
      "postDate": "2020-01-23T11:41:12.130Z",
      "content": "<p>Thanks to the sponsors and kaggle for hosting such an interesting and challenging competition. Also a big thank you to <a href=\"/cpmpml\">@cpmpml</a> for being my teammate.</p>\n\n<h2>Brief Summay</h2>\n\n<p>As our teamname suggests, we did everything with pytorch. In summary, we used 3 roberta-large models which were ensembled by voting. In general input features of our models are very close to bertjoint baseline. We used a learning rate of 1-e5, a batchsize of 16 and simple Adam optimizer with no schedule. All models were trained for 1 epoch.</p>\n\n<p>Roberta 1:\n- initialized with roberta-large weights\n- stride 128\n- prediction of span &amp; 5 answer types (unknown, yes, no, short , long)</p>\n\n<p>Roberta 2:\n- initialized with roberta-large weights, then pretrained on Squad2.0\n- stride 192\n- prediction of span &amp; 2 answer types (short , long)</p>\n\n<p>Roberta 3:\n- initialized with roberta-large weights, then pretrained on Squad2.0\n- additional linear layer (768→768 + relu) before predicting start, respectively end token \n- stride 192\n- prediction of span &amp; 2 answer types (short , long)</p>\n\n<p>We optimized thresholds for each of the models and set predictios below threshold to blank. Then we used majority voting to ensemble the 3 models. Besides some smaller tricks, we predicted test set with a stride of 224 to fit inference of 3 models into the kernel.</p>\n\n<h2>Longer Summary</h2>\n\n<h3>Validation scheme</h3>\n\n<p>As always, I start with setting up a solid validation scheme, which ideally has a high correlation to leaderboard. It turned out harder than anticipated, since organisers did not share enough information on the intended metric as well as implemented it wrongly. This first phase was very frustrating and I spent quite some time reverse engineering their mistake in order to reconcile leaderboard scores. After I figured out the metric and shared in forum, organizers changed the metric. Imagine my face in that moment… and believe it or not, it took me another 6 weeks to figure out the new one. At the end we used the dev set of the original NQ dataset as our validation set and had a very high lb correlation.</p>\n\n<h3>Software</h3>\n\n<p>I reused a lot of preprocessing scripts from bertjoint baseline shared by organisers and did all training with pytorch relying on huggingface for transformer weights and code + pytorch-lightning for writing training pipeline.</p>\n\n<h3>Hardware</h3>\n\n<p>I did all training on my home desktop pc (3 GTX1080Ti) and <a href=\"/cpmpml\">@cpmpml</a> on his pc (2 GTX 1080Ti). Training one epoch took quite a while, hence we did not spent much time on hyper-parameter tuning. The training time for Roberta1 was 35h. Finetuning roberta-large on SQuAD2.0 took 30h and finetuning the resulting model to the data of this competition took about 24h when using a stride of 192.</p>\n\n<h3>Architectures and pretrained models</h3>\n\n<p>I fully agree with <a href=\"/boliu0\">@boliu0</a> that is was frustratingly hard to beat the bertjoint baseline. I did a lot of experiments on different preprocessing as well using different (in my opinion more suited) targets. But 99% of what I did was worse than the baseline. So at the end we kept the preprocessing and only adjusted the answer type targets sightly. I used distilbert for a lot of those experiments because due to its size it helps to iterate fast while giving reasonable indication if an idea works or not. <br>\nIn general the bertjoint baseline suggest to stride over the full answer with a windowing approach and concatenating those windows with the question in order to find if the short answer is contained in the window. One major interesting question is how to aggregate the resulting predictions. Thats where we spent some time because we saw a lot of room for improvement. So what we did is to map the start and end token predictions of each window back to the original answer and create a answer length x answer length heatmap. We then apply some restrictions, like e.g. short span length should be less than 30 tokens, and get the following result.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2F15f5fac1de79c9e7d6d67adb3b45b1de%2FScreenshot%20from%202019-12-21%2009-26-43.png?generation=1579778810842946&amp;alt=media\" alt=\"\"></p>\n\n<p>The argmax of this matrix then gives start and end token (here 957:973). Nice thing of this approach is that you can easily blend these matrices over different models. So after we figured that out we tried different model architectures, including all popular ones from the huggingface repo (albert, gpt2, bert, roberta, xlnet) as well as less popular ones like Spanbert. For us roberta-large worked best with some distance to the second best which was spanbert. Considering the time of preprocessing we thought that ensembling 2 or 3 versions of the same model type will be better than ensembling different model types as you need to do preprocessing only once. So we continued training slightly different versions of roberta-large, including pretraining it on SquAD2.0 first, while working on probably the most important part of this competition, namely thresholding of when to set a blank prediction.</p>\n\n<h3>Thresholding:</h3>\n\n<p>Thresholding when using f1 is challenging. Its super important for your overall score but at the same time has high variance, and might not relate to test set. We used different schemes and at the end a 4-way thresholding worked best. We build thresholds for long and short answer type as well as logits of start + end tokens. We determined the thresholds by simple 4d grid search, which was improved by <a href=\"/cpmpml\">@cpmpml</a>  using scipy.optimize.minimize. Instead of using the thresholds found by fitting on the dev set directly, we also experimented with using the corresponding quantiles. Our best submission uses that approach.</p>\n\n<h3>Ensembling:</h3>\n\n<p>We elaborated different ensemble methods and chose 2 different ones for our final sub:</p>\n\n<ol>\n<li>Apply postprocessing and threshold to model prediction and majority vote between the results</li>\n<li>blend model predictions and apply thresholding</li>\n</ol>\n\n<p>While 2. preformed better on our val set, 1. performed better on public and private LB</p>\n\n<h3>Wrapping things up and putting into kernel:</h3>\n\n<p>We used several things to speed up the final kernel in order to fit the inference of 3 models in.\n- use stride of 224 for test data \n- convert model to fp16 for predictions\n- use multiprocessing for preprocessing and postprocessing</p>\n\n<p>Thanks for reading. </p>",
      "rawMarkdown": "Thanks to the sponsors and kaggle for hosting such an interesting and challenging competition. Also a big thank you to @cpmpml for being my teammate.\n\n\n## Brief Summay\n\nAs our teamname suggests, we did everything with pytorch. In summary, we used 3 roberta-large models which were ensembled by voting. In general input features of our models are very close to bertjoint baseline. We used a learning rate of 1-e5, a batchsize of 16 and simple Adam optimizer with no schedule. All models were trained for 1 epoch.\n\nRoberta 1:\n- initialized with roberta-large weights\n- stride 128\n- prediction of span &amp; 5 answer types (unknown, yes, no, short , long)\n\nRoberta 2:\n- initialized with roberta-large weights, then pretrained on Squad2.0\n- stride 192\n- prediction of span &amp; 2 answer types (short , long)\n\nRoberta 3:\n- initialized with roberta-large weights, then pretrained on Squad2.0\n- additional linear layer (768→768 + relu) before predicting start, respectively end token \n- stride 192\n- prediction of span &amp; 2 answer types (short , long)\n\nWe optimized thresholds for each of the models and set predictios below threshold to blank. Then we used majority voting to ensemble the 3 models. Besides some smaller tricks, we predicted test set with a stride of 224 to fit inference of 3 models into the kernel.\n\n## Longer Summary\n\n### Validation scheme\nAs always, I start with setting up a solid validation scheme, which ideally has a high correlation to leaderboard. It turned out harder than anticipated, since organisers did not share enough information on the intended metric as well as implemented it wrongly. This first phase was very frustrating and I spent quite some time reverse engineering their mistake in order to reconcile leaderboard scores. After I figured out the metric and shared in forum, organizers changed the metric. Imagine my face in that moment… and believe it or not, it took me another 6 weeks to figure out the new one. At the end we used the dev set of the original NQ dataset as our validation set and had a very high lb correlation.\n\n### Software\nI reused a lot of preprocessing scripts from bertjoint baseline shared by organisers and did all training with pytorch relying on huggingface for transformer weights and code + pytorch-lightning for writing training pipeline.\n\n### Hardware\nI did all training on my home desktop pc (3 GTX1080Ti) and @cpmpml on his pc (2 GTX 1080Ti). Training one epoch took quite a while, hence we did not spent much time on hyper-parameter tuning. The training time for Roberta1 was 35h. Finetuning roberta-large on SQuAD2.0 took 30h and finetuning the resulting model to the data of this competition took about 24h when using a stride of 192.\n\n### Architectures and pretrained models\nI fully agree with @boliu0 that is was frustratingly hard to beat the bertjoint baseline. I did a lot of experiments on different preprocessing as well using different (in my opinion more suited) targets. But 99% of what I did was worse than the baseline. So at the end we kept the preprocessing and only adjusted the answer type targets sightly. I used distilbert for a lot of those experiments because due to its size it helps to iterate fast while giving reasonable indication if an idea works or not.  \nIn general the bertjoint baseline suggest to stride over the full answer with a windowing approach and concatenating those windows with the question in order to find if the short answer is contained in the window. One major interesting question is how to aggregate the resulting predictions. Thats where we spent some time because we saw a lot of room for improvement. So what we did is to map the start and end token predictions of each window back to the original answer and create a answer length x answer length heatmap. We then apply some restrictions, like e.g. short span length should be less than 30 tokens, and get the following result.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2F15f5fac1de79c9e7d6d67adb3b45b1de%2FScreenshot%20from%202019-12-21%2009-26-43.png?generation=1579778810842946&amp;alt=media)\n\n\nThe argmax of this matrix then gives start and end token (here 957:973). Nice thing of this approach is that you can easily blend these matrices over different models. So after we figured that out we tried different model architectures, including all popular ones from the huggingface repo (albert, gpt2, bert, roberta, xlnet) as well as less popular ones like Spanbert. For us roberta-large worked best with some distance to the second best which was spanbert. Considering the time of preprocessing we thought that ensembling 2 or 3 versions of the same model type will be better than ensembling different model types as you need to do preprocessing only once. So we continued training slightly different versions of roberta-large, including pretraining it on SquAD2.0 first, while working on probably the most important part of this competition, namely thresholding of when to set a blank prediction.\n\n\n### Thresholding:\n\nThresholding when using f1 is challenging. Its super important for your overall score but at the same time has high variance, and might not relate to test set. We used different schemes and at the end a 4-way thresholding worked best. We build thresholds for long and short answer type as well as logits of start + end tokens. We determined the thresholds by simple 4d grid search, which was improved by @cpmpml  using scipy.optimize.minimize. Instead of using the thresholds found by fitting on the dev set directly, we also experimented with using the corresponding quantiles. Our best submission uses that approach.\n\n### Ensembling:\nWe elaborated different ensemble methods and chose 2 different ones for our final sub:\n\n1. Apply postprocessing and threshold to model prediction and majority vote between the results\n2. blend model predictions and apply thresholding\n \nWhile 2. preformed better on our val set, 1. performed better on public and private LB\n\n### Wrapping things up and putting into kernel:\n\nWe used several things to speed up the final kernel in order to fit the inference of 3 models in.\n- use stride of 224 for test data \n- convert model to fp16 for predictions\n- use multiprocessing for preprocessing and postprocessing\n\nThanks for reading. ",
      "votes": 77
    },
    {
      "id": 727305,
      "postDate": "2020-01-23T16:09:25.090Z",
      "content": "<p>Thanks <a href=\"/christofhenkel\">@christofhenkel</a> for being such a good team mate.  You're not only smart, but also fun to work with! </p>\n\n<p>The write up is perfect, I just want to add few little missing pieces.  </p>\n\n<p>When aggregating the predictions over all the windows for a given sample we averaged the start and end logit instead of taking the max as in the joint baseline paper.</p>\n\n<p>Another difference with bert baseline was to use the first short answer span instead of the convex hull of all short answer spans.  We tried to use all short answer spans for training instead of the first one, either by creating one window for each span, or by using BCE on start and end logits to accomodate  the presence of several 1s.  None of these improve, to the contrary.  I think there is room for improvement here.</p>\n\n<p>For ensembling we also tried, at the last minute, to blend predicted logits before thresholding, instead of voting after thresholding.  This looked better on dev set cross validation, at 0.725, but scored 0.69 on private LB.</p>\n\n<p>The only thing we could have done better IMHO (There aren't many because Dieter had done so many things right already) is to deal with empty window sampling.  We kept the baseline paper way.  This sampling has a side effect: the longer the context (wikipedia text) for an example, the larger the proportion of windows without short answers.  Said differently, the probability that a window contains a short answer decreases with context size.  There was no way our model could learn that.  </p>\n\n<p>To conclude, I'm not only extremely happy with the result, but also very happy to have learned a lot about pytorch, huggingfaces transformers, and NLP state of the art.  </p>",
      "rawMarkdown": "Thanks @christofhenkel for being such a good team mate.  You're not only smart, but also fun to work with! \n\nThe write up is perfect, I just want to add few little missing pieces.  \n\nWhen aggregating the predictions over all the windows for a given sample we averaged the start and end logit instead of taking the max as in the joint baseline paper.\n\nAnother difference with bert baseline was to use the first short answer span instead of the convex hull of all short answer spans.  We tried to use all short answer spans for training instead of the first one, either by creating one window for each span, or by using BCE on start and end logits to accomodate  the presence of several 1s.  None of these improve, to the contrary.  I think there is room for improvement here.\n\nFor ensembling we also tried, at the last minute, to blend predicted logits before thresholding, instead of voting after thresholding.  This looked better on dev set cross validation, at 0.725, but scored 0.69 on private LB.\n\nThe only thing we could have done better IMHO (There aren't many because Dieter had done so many things right already) is to deal with empty window sampling.  We kept the baseline paper way.  This sampling has a side effect: the longer the context (wikipedia text) for an example, the larger the proportion of windows without short answers.  Said differently, the probability that a window contains a short answer decreases with context size.  There was no way our model could learn that.  \n\nTo conclude, I'm not only extremely happy with the result, but also very happy to have learned a lot about pytorch, huggingfaces transformers, and NLP state of the art.  ",
      "votes": 7
    },
    {
      "id": 727434,
      "postDate": "2020-01-23T18:08:55.890Z",
      "content": "<p>Well done <a href=\"/christofhenkel\">@christofhenkel</a> <a href=\"/cpmpml\">@cpmpml</a>  .. and doing it locally on your own machines is even more impressive.\nQuick question, <code>convert model to fp16 for predictions</code>; was this with apex ? or how did you do it ? I thought <code>apex.initialze(...)</code> only does fp16 in training, does it also do it out of the box for predictions ?</p>",
      "rawMarkdown": "Well done @christofhenkel @cpmpml  .. and doing it locally on your own machines is even more impressive.\nQuick question, `convert model to fp16 for predictions`; was this with apex ? or how did you do it ? I thought `apex.initialze(...)` only does fp16 in training, does it also do it out of the box for predictions ?",
      "votes": 3,
      "replies": [
        {
          "id": 727457,
          "postDate": "2020-01-23T18:25:20.047Z",
          "content": "<p>you can simply do something like</p>\n\n<p><code>\nmodel = TFQARoberta()\nmodel.half().cuda()\n</code></p>\n\n<p>no need for apex</p>",
          "rawMarkdown": "you can simply do something like\n\n```\nmodel = TFQARoberta()\nmodel.half().cuda()\n```\n\nno need for apex",
          "votes": 7
        },
        {
          "id": 727484,
          "postDate": "2020-01-23T18:40:32.003Z",
          "content": "<p>Haha :) clever</p>",
          "rawMarkdown": "Haha :) clever"
        }
      ]
    },
    {
      "id": 727456,
      "postDate": "2020-01-23T18:24:16.033Z",
      "content": "<p>Glad that third place went to Munich :-)</p>",
      "rawMarkdown": "Glad that third place went to Munich :-)",
      "votes": 1
    },
    {
      "id": 727400,
      "postDate": "2020-01-23T17:35:39.043Z",
      "content": "<p>Congrats Dieter and thanks for sharing! Nice work with only GPUs.</p>\n\n<blockquote>\n  <p>After I figured out the metric and shared in forum, organizers changed the metric. Imagine my face in that moment… and believe it or not, it took me another 6 weeks to figure out the new one.</p>\n</blockquote>\n\n<p>Hey it worked out in the end. But I'm a little surprised it took you that long for the new one. We just used the official NQ's eval script without any issue.</p>\n\n<blockquote>\n  <p>I did a lot of experiments on different preprocessing as well using different (in my opinion more suited) targets. But 99% of what I did was worse than the baseline. </p>\n</blockquote>\n\n<p>Same here. What I've been telling myself and <a href=\"/jiweiliu\">@jiweiliu</a> to make us feel better is that, joint-bert's co-author Kenton Lee is also the co-author for both Elmo and Bert, so his intuition is probably much better than us.</p>",
      "rawMarkdown": "Congrats Dieter and thanks for sharing! Nice work with only GPUs.\n\n&gt; After I figured out the metric and shared in forum, organizers changed the metric. Imagine my face in that moment… and believe it or not, it took me another 6 weeks to figure out the new one.\n\nHey it worked out in the end. But I'm a little surprised it took you that long for the new one. We just used the official NQ's eval script without any issue.\n\n&gt; I did a lot of experiments on different preprocessing as well using different (in my opinion more suited) targets. But 99% of what I did was worse than the baseline. \n\nSame here. What I've been telling myself and @jiweiliu to make us feel better is that, joint-bert's co-author Kenton Lee is also the co-author for both Elmo and Bert, so his intuition is probably much better than us.",
      "votes": 1,
      "replies": [
        {
          "id": 727403,
          "postDate": "2020-01-23T17:40:35.130Z",
          "content": "<blockquote>\n  <p>But I'm a little surprised it took you that long for the new one.</p>\n</blockquote>\n\n<p>It took me so long to find out that the dev examples on NQ had multiple annotations and a precise description how to label an example. Without the NQ dev set, say you take a hold out of this competitions training data instead, you won't be able to reconcile the LB score. That fact was hidden from me. Could have been made more transparent by organizers...</p>",
          "rawMarkdown": "&gt; But I'm a little surprised it took you that long for the new one.\n\nIt took me so long to find out that the dev examples on NQ had multiple annotations and a precise description how to label an example. Without the NQ dev set, say you take a hold out of this competitions training data instead, you won't be able to reconcile the LB score. That fact was hidden from me. Could have been made more transparent by organizers...\n",
          "votes": 3
        },
        {
          "id": 728002,
          "postDate": "2020-01-24T10:19:20.007Z",
          "content": "<p>There was a competition inside the competition: find how to compute the metric.</p>",
          "rawMarkdown": "There was a competition inside the competition: find how to compute the metric.",
          "votes": 1
        }
      ]
    },
    {
      "id": 727044,
      "postDate": "2020-01-23T12:04:43.683Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": 1
    },
    {
      "id": 727202,
      "postDate": "2020-01-23T14:28:18.113Z",
      "content": "<p>Are you planning to release the pytorch training code ?</p>",
      "rawMarkdown": "Are you planning to release the pytorch training code ?",
      "votes": 2,
      "replies": [
        {
          "id": 730402,
          "postDate": "2020-01-27T12:50:52.463Z",
          "content": "<p>HI, we will wait after QUEST competition end in case one of us wants to enter it.</p>",
          "rawMarkdown": "HI, we will wait after QUEST competition end in case one of us wants to enter it."
        }
      ]
    },
    {
      "id": 739423,
      "postDate": "2020-02-07T20:27:54.953Z",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>)</p>",
      "rawMarkdown": "@christofhenkel, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques)"
    },
    {
      "id": 2190450,
      "postDate": "2023-03-21T08:41:36.267Z",
      "content": "<p>Nice work！</p>",
      "rawMarkdown": "Nice work！"
    },
    {
      "id": 728121,
      "postDate": "2020-01-24T12:20:07.563Z",
      "content": "<p>Interesting, thank you. Using bert-large, I was able to fit in 3 inferences with float32 and only 200 stride length. I would have thought roberta was a similar speed.</p>",
      "rawMarkdown": "Interesting, thank you. Using bert-large, I was able to fit in 3 inferences with float32 and only 200 stride length. I would have thought roberta was a similar speed.",
      "replies": [
        {
          "id": 731290,
          "postDate": "2020-01-28T13:31:01.023Z",
          "content": "<p>might also as well be that is would have worked. I did not try much different strides. Just 128, 192, and 224. I just put fp16, because its a bit faster, and then 192 and checked if its under 20min for commit time. It was at around 22 min. So I changed stride to 224 and voila it was below 20 min.</p>",
          "rawMarkdown": "might also as well be that is would have worked. I did not try much different strides. Just 128, 192, and 224. I just put fp16, because its a bit faster, and then 192 and checked if its under 20min for commit time. It was at around 22 min. So I changed stride to 224 and voila it was below 20 min."
        }
      ]
    },
    {
      "id": 727588,
      "postDate": "2020-01-23T21:06:20.473Z",
      "content": "<p>Big congrats on the top 3 finish. Very nice writeup.</p>",
      "rawMarkdown": "Big congrats on the top 3 finish. Very nice writeup."
    },
    {
      "id": 727301,
      "postDate": "2020-01-23T16:06:00.767Z",
      "content": "<p>Thank you for you sharing. Want to know how you ensemble models with different structure of answer-type labels.</p>",
      "rawMarkdown": "Thank you for you sharing. Want to know how you ensemble models with different structure of answer-type labels."
    },
    {
      "id": 727175,
      "postDate": "2020-01-23T14:07:29.200Z",
      "content": "<p>Great write-up. Congrats!</p>",
      "rawMarkdown": "Great write-up. Congrats!"
    },
    {
      "id": 727169,
      "postDate": "2020-01-23T14:03:51.393Z",
      "content": "<p>Thank you for for a very detailed write-up and congratulations on third place!</p>",
      "rawMarkdown": "Thank you for for a very detailed write-up and congratulations on third place!"
    },
    {
      "id": 727111,
      "postDate": "2020-01-23T13:18:11.037Z",
      "content": "<p>Nice . any plans to make the kernel public ? :-p</p>",
      "rawMarkdown": "Nice . any plans to make the kernel public ? :-p",
      "replies": [
        {
          "id": 727398,
          "postDate": "2020-01-23T17:31:13.253Z",
          "content": "<p>I would like to wait for publishing our inference kernel after the Google Quest competition.</p>",
          "rawMarkdown": "I would like to wait for publishing our inference kernel after the Google Quest competition.",
          "votes": 2
        },
        {
          "id": 738490,
          "postDate": "2020-02-06T15:44:13.307Z",
          "content": "<p><a href=\"https://www.kaggle.com/christofhenkel/inference-v3-224\">https://www.kaggle.com/christofhenkel/inference-v3-224</a></p>",
          "rawMarkdown": "https://www.kaggle.com/christofhenkel/inference-v3-224",
          "votes": 2
        },
        {
          "id": 860328,
          "postDate": "2020-05-25T07:46:13.437Z",
          "content": "<p>Hi, the link for the solution kernel you posted for Google NQ competition is broken (404), may I know if the solution kernel (both training and inference) is posted elsewhere?</p>",
          "rawMarkdown": "Hi, the link for the solution kernel you posted for Google NQ competition is broken (404), may I know if the solution kernel (both training and inference) is posted elsewhere?"
        }
      ]
    },
    {
      "id": 731263,
      "postDate": "2020-01-28T13:15:39.207Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 731286,
          "postDate": "2020-01-28T13:27:30.787Z",
          "content": "<p>stride doesn't skip anything as long as its less then 512 - (question len). Its rather the overlap of the windows. As for the number 224 exactly, we just increased from 192 and checked if our models fit into the kernel time</p>",
          "rawMarkdown": "stride doesn't skip anything as long as its less then 512 - (question len). Its rather the overlap of the windows. As for the number 224 exactly, we just increased from 192 and checked if our models fit into the kernel time"
        }
      ]
    },
    {
      "id": 727039,
      "postDate": "2020-01-23T12:01:04.027Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 727021,
      "postDate": "2020-01-23T11:42:14.463Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 727490,
      "postDate": "2020-01-23T18:43:35.760Z",
      "content": "<p>Congratss!! Thanks for sharing !</p>",
      "rawMarkdown": "Congratss!! Thanks for sharing !",
      "votes": 4
    },
    {
      "id": 733300,
      "postDate": "2020-01-31T00:02:17.633Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!"
    },
    {
      "id": 727830,
      "postDate": "2020-01-24T05:08:34.513Z",
      "content": "<p>Thanks for sharing!  </p>",
      "rawMarkdown": "Thanks for sharing!  "
    }
  ],
  "comments": [
    {
      "id": 727305,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-01-23T16:09:25.090000",
      "content": "<p>Thanks <a href=\"/christofhenkel\">@christofhenkel</a> for being such a good team mate.  You're not only smart, but also fun to work with! </p>\n\n<p>The write up is perfect, I just want to add few little missing pieces.  </p>\n\n<p>When aggregating the predictions over all the windows for a given sample we averaged the start and end logit instead of taking the max as in the joint baseline paper.</p>\n\n<p>Another difference with bert baseline was to use the first short answer span instead of the convex hull of all short answer spans.  We tried to use all short answer spans for training instead of the first one, either by creating one window for each span, or by using BCE on start and end logits to accomodate  the presence of several 1s.  None of these improve, to the contrary.  I think there is room for improvement here.</p>\n\n<p>For ensembling we also tried, at the last minute, to blend predicted logits before thresholding, instead of voting after thresholding.  This looked better on dev set cross validation, at 0.725, but scored 0.69 on private LB.</p>\n\n<p>The only thing we could have done better IMHO (There aren't many because Dieter had done so many things right already) is to deal with empty window sampling.  We kept the baseline paper way.  This sampling has a side effect: the longer the context (wikipedia text) for an example, the larger the proportion of windows without short answers.  Said differently, the probability that a window contains a short answer decreases with context size.  There was no way our model could learn that.  </p>\n\n<p>To conclude, I'm not only extremely happy with the result, but also very happy to have learned a lot about pytorch, huggingfaces transformers, and NLP state of the art.  </p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 727434,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2020-01-23T18:08:55.890000",
      "content": "<p>Well done <a href=\"/christofhenkel\">@christofhenkel</a> <a href=\"/cpmpml\">@cpmpml</a>  .. and doing it locally on your own machines is even more impressive.\nQuick question, <code>convert model to fp16 for predictions</code>; was this with apex ? or how did you do it ? I thought <code>apex.initialze(...)</code> only does fp16 in training, does it also do it out of the box for predictions ?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 727457,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-01-23T18:25:20.047000",
          "content": "<p>you can simply do something like</p>\n\n<p><code>\nmodel = TFQARoberta()\nmodel.half().cuda()\n</code></p>\n\n<p>no need for apex</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 727484,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2020-01-23T18:40:32.003000",
          "content": "<p>Haha :) clever</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 727456,
      "author_name": "David",
      "author_url": "",
      "post_date": "2020-01-23T18:24:16.033000",
      "content": "<p>Glad that third place went to Munich :-)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 727400,
      "author_name": "Bo",
      "author_url": "",
      "post_date": "2020-01-23T17:35:39.043000",
      "content": "<p>Congrats Dieter and thanks for sharing! Nice work with only GPUs.</p>\n\n<blockquote>\n  <p>After I figured out the metric and shared in forum, organizers changed the metric. Imagine my face in that moment… and believe it or not, it took me another 6 weeks to figure out the new one.</p>\n</blockquote>\n\n<p>Hey it worked out in the end. But I'm a little surprised it took you that long for the new one. We just used the official NQ's eval script without any issue.</p>\n\n<blockquote>\n  <p>I did a lot of experiments on different preprocessing as well using different (in my opinion more suited) targets. But 99% of what I did was worse than the baseline. </p>\n</blockquote>\n\n<p>Same here. What I've been telling myself and <a href=\"/jiweiliu\">@jiweiliu</a> to make us feel better is that, joint-bert's co-author Kenton Lee is also the co-author for both Elmo and Bert, so his intuition is probably much better than us.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 727403,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-01-23T17:40:35.130000",
          "content": "<blockquote>\n  <p>But I'm a little surprised it took you that long for the new one.</p>\n</blockquote>\n\n<p>It took me so long to find out that the dev examples on NQ had multiple annotations and a precise description how to label an example. Without the NQ dev set, say you take a hold out of this competitions training data instead, you won't be able to reconcile the LB score. That fact was hidden from me. Could have been made more transparent by organizers...</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 728002,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-01-24T10:19:20.007000",
          "content": "<p>There was a competition inside the competition: find how to compute the metric.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 727044,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-01-23T12:04:43.683000",
      "content": "<p>Congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 727202,
      "author_name": "Abhishek Thakur",
      "author_url": "",
      "post_date": "2020-01-23T14:28:18.113000",
      "content": "<p>Are you planning to release the pytorch training code ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 730402,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-01-27T12:50:52.463000",
          "content": "<p>HI, we will wait after QUEST competition end in case one of us wants to enter it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 739423,
      "author_name": "Vitalii Mokin",
      "author_url": "",
      "post_date": "2020-02-07T20:27:54.953000",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2190450,
      "author_name": "IHSINGC",
      "author_url": "",
      "post_date": "2023-03-21T08:41:36.267000",
      "content": "<p>Nice work！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728121,
      "author_name": "Ken Krige",
      "author_url": "",
      "post_date": "2020-01-24T12:20:07.563000",
      "content": "<p>Interesting, thank you. Using bert-large, I was able to fit in 3 inferences with float32 and only 200 stride length. I would have thought roberta was a similar speed.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 731290,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-01-28T13:31:01.023000",
          "content": "<p>might also as well be that is would have worked. I did not try much different strides. Just 128, 192, and 224. I just put fp16, because its a bit faster, and then 192 and checked if its under 20min for commit time. It was at around 22 min. So I changed stride to 224 and voila it was below 20 min.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 727588,
      "author_name": "Jiwei Liu",
      "author_url": "",
      "post_date": "2020-01-23T21:06:20.473000",
      "content": "<p>Big congrats on the top 3 finish. Very nice writeup.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727301,
      "author_name": "SchenbergZ",
      "author_url": "",
      "post_date": "2020-01-23T16:06:00.767000",
      "content": "<p>Thank you for you sharing. Want to know how you ensemble models with different structure of answer-type labels.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727175,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2020-01-23T14:07:29.200000",
      "content": "<p>Great write-up. Congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727169,
      "author_name": "David",
      "author_url": "",
      "post_date": "2020-01-23T14:03:51.393000",
      "content": "<p>Thank you for for a very detailed write-up and congratulations on third place!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727111,
      "author_name": "aintnosunshine",
      "author_url": "",
      "post_date": "2020-01-23T13:18:11.037000",
      "content": "<p>Nice . any plans to make the kernel public ? :-p</p>",
      "votes": 0,
      "replies": [
        {
          "id": 727398,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-01-23T17:31:13.253000",
          "content": "<p>I would like to wait for publishing our inference kernel after the Google Quest competition.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 738490,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-02-06T15:44:13.307000",
          "content": "<p><a href=\"https://www.kaggle.com/christofhenkel/inference-v3-224\">https://www.kaggle.com/christofhenkel/inference-v3-224</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 860328,
          "author_name": "Chew Kok Wah",
          "author_url": "",
          "post_date": "2020-05-25T07:46:13.437000",
          "content": "<p>Hi, the link for the solution kernel you posted for Google NQ competition is broken (404), may I know if the solution kernel (both training and inference) is posted elsewhere?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 731263,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-28T13:15:39.207000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 731286,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-01-28T13:27:30.787000",
          "content": "<p>stride doesn't skip anything as long as its less then 512 - (question len). Its rather the overlap of the windows. As for the number 224 exactly, we just increased from 192 and checked if our models fit into the kernel time</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 727039,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-23T12:01:04.027000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727021,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-23T11:42:14.463000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727490,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-23T18:43:35.760000",
      "content": "<p>Congratss!! Thanks for sharing !</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 733300,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-01-31T00:02:17.633000",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727830,
      "author_name": "xiao-xiao",
      "author_url": "",
      "post_date": "2020-01-24T05:08:34.513000",
      "content": "<p>Thanks for sharing!  </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "727020": "Thanks to the sponsors and kaggle for hosting such an interesting and challenging competition. Also a big thank you to @cpmpml for being my teammate.\n\n\n## Brief Summay\n\nAs our teamname suggests, we did everything with pytorch. In summary, we used 3 roberta-large models which were ensembled by voting. In general input features of our models are very close to bertjoint baseline. We used a learning rate of 1-e5, a batchsize of 16 and simple Adam optimizer with no schedule. All models were trained for 1 epoch.\n\nRoberta 1:\n- initialized with roberta-large weights\n- stride 128\n- prediction of span &amp; 5 answer types (unknown, yes, no, short , long)\n\nRoberta 2:\n- initialized with roberta-large weights, then pretrained on Squad2.0\n- stride 192\n- prediction of span &amp; 2 answer types (short , long)\n\nRoberta 3:\n- initialized with roberta-large weights, then pretrained on Squad2.0\n- additional linear layer (768→768 + relu) before predicting start, respectively end token \n- stride 192\n- prediction of span &amp; 2 answer types (short , long)\n\nWe optimized thresholds for each of the models and set predictios below threshold to blank. Then we used majority voting to ensemble the 3 models. Besides some smaller tricks, we predicted test set with a stride of 224 to fit inference of 3 models into the kernel.\n\n## Longer Summary\n\n### Validation scheme\nAs always, I start with setting up a solid validation scheme, which ideally has a high correlation to leaderboard. It turned out harder than anticipated, since organisers did not share enough information on the intended metric as well as implemented it wrongly. This first phase was very frustrating and I spent quite some time reverse engineering their mistake in order to reconcile leaderboard scores. After I figured out the metric and shared in forum, organizers changed the metric. Imagine my face in that moment… and believe it or not, it took me another 6 weeks to figure out the new one. At the end we used the dev set of the original NQ dataset as our validation set and had a very high lb correlation.\n\n### Software\nI reused a lot of preprocessing scripts from bertjoint baseline shared by organisers and did all training with pytorch relying on huggingface for transformer weights and code + pytorch-lightning for writing training pipeline.\n\n### Hardware\nI did all training on my home desktop pc (3 GTX1080Ti) and @cpmpml on his pc (2 GTX 1080Ti). Training one epoch took quite a while, hence we did not spent much time on hyper-parameter tuning. The training time for Roberta1 was 35h. Finetuning roberta-large on SQuAD2.0 took 30h and finetuning the resulting model to the data of this competition took about 24h when using a stride of 192.\n\n### Architectures and pretrained models\nI fully agree with @boliu0 that is was frustratingly hard to beat the bertjoint baseline. I did a lot of experiments on different preprocessing as well using different (in my opinion more suited) targets. But 99% of what I did was worse than the baseline. So at the end we kept the preprocessing and only adjusted the answer type targets sightly. I used distilbert for a lot of those experiments because due to its size it helps to iterate fast while giving reasonable indication if an idea works or not.  \nIn general the bertjoint baseline suggest to stride over the full answer with a windowing approach and concatenating those windows with the question in order to find if the short answer is contained in the window. One major interesting question is how to aggregate the resulting predictions. Thats where we spent some time because we saw a lot of room for improvement. So what we did is to map the start and end token predictions of each window back to the original answer and create a answer length x answer length heatmap. We then apply some restrictions, like e.g. short span length should be less than 30 tokens, and get the following result.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2F15f5fac1de79c9e7d6d67adb3b45b1de%2FScreenshot%20from%202019-12-21%2009-26-43.png?generation=1579778810842946&amp;alt=media)\n\n\nThe argmax of this matrix then gives start and end token (here 957:973). Nice thing of this approach is that you can easily blend these matrices over different models. So after we figured that out we tried different model architectures, including all popular ones from the huggingface repo (albert, gpt2, bert, roberta, xlnet) as well as less popular ones like Spanbert. For us roberta-large worked best with some distance to the second best which was spanbert. Considering the time of preprocessing we thought that ensembling 2 or 3 versions of the same model type will be better than ensembling different model types as you need to do preprocessing only once. So we continued training slightly different versions of roberta-large, including pretraining it on SquAD2.0 first, while working on probably the most important part of this competition, namely thresholding of when to set a blank prediction.\n\n\n### Thresholding:\n\nThresholding when using f1 is challenging. Its super important for your overall score but at the same time has high variance, and might not relate to test set. We used different schemes and at the end a 4-way thresholding worked best. We build thresholds for long and short answer type as well as logits of start + end tokens. We determined the thresholds by simple 4d grid search, which was improved by @cpmpml  using scipy.optimize.minimize. Instead of using the thresholds found by fitting on the dev set directly, we also experimented with using the corresponding quantiles. Our best submission uses that approach.\n\n### Ensembling:\nWe elaborated different ensemble methods and chose 2 different ones for our final sub:\n\n1. Apply postprocessing and threshold to model prediction and majority vote between the results\n2. blend model predictions and apply thresholding\n \nWhile 2. preformed better on our val set, 1. performed better on public and private LB\n\n### Wrapping things up and putting into kernel:\n\nWe used several things to speed up the final kernel in order to fit the inference of 3 models in.\n- use stride of 224 for test data \n- convert model to fp16 for predictions\n- use multiprocessing for preprocessing and postprocessing\n\nThanks for reading. ",
    "727305": "Thanks @christofhenkel for being such a good team mate.  You're not only smart, but also fun to work with! \n\nThe write up is perfect, I just want to add few little missing pieces.  \n\nWhen aggregating the predictions over all the windows for a given sample we averaged the start and end logit instead of taking the max as in the joint baseline paper.\n\nAnother difference with bert baseline was to use the first short answer span instead of the convex hull of all short answer spans.  We tried to use all short answer spans for training instead of the first one, either by creating one window for each span, or by using BCE on start and end logits to accomodate  the presence of several 1s.  None of these improve, to the contrary.  I think there is room for improvement here.\n\nFor ensembling we also tried, at the last minute, to blend predicted logits before thresholding, instead of voting after thresholding.  This looked better on dev set cross validation, at 0.725, but scored 0.69 on private LB.\n\nThe only thing we could have done better IMHO (There aren't many because Dieter had done so many things right already) is to deal with empty window sampling.  We kept the baseline paper way.  This sampling has a side effect: the longer the context (wikipedia text) for an example, the larger the proportion of windows without short answers.  Said differently, the probability that a window contains a short answer decreases with context size.  There was no way our model could learn that.  \n\nTo conclude, I'm not only extremely happy with the result, but also very happy to have learned a lot about pytorch, huggingfaces transformers, and NLP state of the art.  ",
    "727434": "Well done @christofhenkel @cpmpml  .. and doing it locally on your own machines is even more impressive.\nQuick question, `convert model to fp16 for predictions`; was this with apex ? or how did you do it ? I thought `apex.initialze(...)` only does fp16 in training, does it also do it out of the box for predictions ?",
    "727456": "Glad that third place went to Munich :-)",
    "727400": "Congrats Dieter and thanks for sharing! Nice work with only GPUs.\n\n&gt; After I figured out the metric and shared in forum, organizers changed the metric. Imagine my face in that moment… and believe it or not, it took me another 6 weeks to figure out the new one.\n\nHey it worked out in the end. But I'm a little surprised it took you that long for the new one. We just used the official NQ's eval script without any issue.\n\n&gt; I did a lot of experiments on different preprocessing as well using different (in my opinion more suited) targets. But 99% of what I did was worse than the baseline. \n\nSame here. What I've been telling myself and @jiweiliu to make us feel better is that, joint-bert's co-author Kenton Lee is also the co-author for both Elmo and Bert, so his intuition is probably much better than us.",
    "727044": "Congrats!",
    "727202": "Are you planning to release the pytorch training code ?",
    "739423": "@christofhenkel, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques)",
    "2190450": "Nice work！",
    "728121": "Interesting, thank you. Using bert-large, I was able to fit in 3 inferences with float32 and only 200 stride length. I would have thought roberta was a similar speed.",
    "727588": "Big congrats on the top 3 finish. Very nice writeup.",
    "727301": "Thank you for you sharing. Want to know how you ensemble models with different structure of answer-type labels.",
    "727175": "Great write-up. Congrats!",
    "727169": "Thank you for for a very detailed write-up and congratulations on third place!",
    "727111": "Nice . any plans to make the kernel public ? :-p",
    "731263": "",
    "727039": "",
    "727021": "",
    "727490": "Congratss!! Thanks for sharing !",
    "733300": "Congrats and thanks for sharing!",
    "727830": "Thanks for sharing!  "
  }
}