{
  "id": 123345,
  "title": "My TF2 Inference + Validation kernel is done.",
  "url": "/competitions/tensorflow2-question-answering/discussion/123345",
  "author_name": "Yih-Dar SHIEH",
  "post_date": "2019-12-26T23:49:27.270000",
  "votes": 23,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I finally finished my <a href=\"https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models\">TF2 inference + validation kernel</a>.</p>\n\n<p>(Please see remarks at the end also, please)</p>\n\n<ul>\n<li><p>I published it earlier for some reason, but since then, a lot of bugs or logic improvements are fixed / done. In particular, 2 places of tf.compat.v1 are removed.</p></li>\n<li><p>It can run validation on nq-dev (simplified) dataset, or a smaller subset (The first1000 examples).</p></li>\n<li><p>For my few tests, the public LB score is about 0.03 ~ 0.05 lower (depends on using full nq-dev or smaller one) than the validation score. I have values CV / LB like (0.36, 0.33), (0.47, 0.45) or (0.55, 0.52)</p></li>\n<li><p>For validation code, I copied the logic from <a href=\"https://github.com/google-research-datasets/natural-questions\">nq_eval.py</a>, in particular, the vote for labels among 5 annotations.</p></li>\n<li><p>I am only using distilled bert for now. The best LB score I can get is 0.52 (so sad...). I tried to copy some code from <a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook\">prokaj's kernel(0.48)</a> and <a href=\"https://www.kaggle.com/mmmarchetti/tensorflow-2-0-bert-yes-no-answers\">mmmarchetti's kernel (0.57)</a>.</p></li>\n<li><p>I got OOM issue on collecting results (model logits, flat to python lists) during predicting, so I only save partial results, thus I need to make some logic change to the original baseline kernel.</p></li>\n</ul>\n\n<p>I hope some of you find it useful or helpful. For me, I am done working on kernels. I won't publish any new kernel in this competition.</p>\n\n<p>It's good to have something working, but I am at 315th place now, and I have no clear idea how to go up (other than working on Bert Large, which I need to use TPU and have no experience at all...).</p>\n\n<p>Good lucks, everyone!</p>\n\n<hr>\n\n<p>Some remarks:</p>\n\n<ul>\n<li><p>In this inference + validation kernel, I forgot to<code>tf.function</code> to speed up model calculation. For distilled bert, this is not a big problem. If you go for larger models, please use <code>tf.function</code> at a proper place. You can find an usage in my <a href=\"https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models\">TF2 training kernel</a>.</p></li>\n<li><p>In <code>compute_predictions(example)</code> method, which savs the prediction json file,  it returns, unlike in the starter kernel, a list of ScoreSummary objects instead of a single one. And in , <code>compute_pred_dict()</code>,  We have</p></li>\n</ul>\n\n<p><code>nq_pred_dict[e.example_id] = [summary.predicted_label for summary in all_summaries]</code>.</p>\n\n<p>However, when doing validation / submission, only <code>preds[0]</code> is used.</p>",
  "messages": [
    {
      "id": 704016,
      "postDate": "2019-12-26T23:49:27.270Z",
      "content": "<p>I finally finished my <a href=\"https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models\">TF2 inference + validation kernel</a>.</p>\n\n<p>(Please see remarks at the end also, please)</p>\n\n<ul>\n<li><p>I published it earlier for some reason, but since then, a lot of bugs or logic improvements are fixed / done. In particular, 2 places of tf.compat.v1 are removed.</p></li>\n<li><p>It can run validation on nq-dev (simplified) dataset, or a smaller subset (The first1000 examples).</p></li>\n<li><p>For my few tests, the public LB score is about 0.03 ~ 0.05 lower (depends on using full nq-dev or smaller one) than the validation score. I have values CV / LB like (0.36, 0.33), (0.47, 0.45) or (0.55, 0.52)</p></li>\n<li><p>For validation code, I copied the logic from <a href=\"https://github.com/google-research-datasets/natural-questions\">nq_eval.py</a>, in particular, the vote for labels among 5 annotations.</p></li>\n<li><p>I am only using distilled bert for now. The best LB score I can get is 0.52 (so sad...). I tried to copy some code from <a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook\">prokaj's kernel(0.48)</a> and <a href=\"https://www.kaggle.com/mmmarchetti/tensorflow-2-0-bert-yes-no-answers\">mmmarchetti's kernel (0.57)</a>.</p></li>\n<li><p>I got OOM issue on collecting results (model logits, flat to python lists) during predicting, so I only save partial results, thus I need to make some logic change to the original baseline kernel.</p></li>\n</ul>\n\n<p>I hope some of you find it useful or helpful. For me, I am done working on kernels. I won't publish any new kernel in this competition.</p>\n\n<p>It's good to have something working, but I am at 315th place now, and I have no clear idea how to go up (other than working on Bert Large, which I need to use TPU and have no experience at all...).</p>\n\n<p>Good lucks, everyone!</p>\n\n<hr>\n\n<p>Some remarks:</p>\n\n<ul>\n<li><p>In this inference + validation kernel, I forgot to<code>tf.function</code> to speed up model calculation. For distilled bert, this is not a big problem. If you go for larger models, please use <code>tf.function</code> at a proper place. You can find an usage in my <a href=\"https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models\">TF2 training kernel</a>.</p></li>\n<li><p>In <code>compute_predictions(example)</code> method, which savs the prediction json file,  it returns, unlike in the starter kernel, a list of ScoreSummary objects instead of a single one. And in , <code>compute_pred_dict()</code>,  We have</p></li>\n</ul>\n\n<p><code>nq_pred_dict[e.example_id] = [summary.predicted_label for summary in all_summaries]</code>.</p>\n\n<p>However, when doing validation / submission, only <code>preds[0]</code> is used.</p>",
      "rawMarkdown": "I finally finished my [TF2 inference + validation kernel](https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models).\n\n(Please see remarks at the end also, please)\n\n- I published it earlier for some reason, but since then, a lot of bugs or logic improvements are fixed / done. In particular, 2 places of tf.compat.v1 are removed.\n\n- It can run validation on nq-dev (simplified) dataset, or a smaller subset (The first1000 examples).\n\n- For my few tests, the public LB score is about 0.03 ~ 0.05 lower (depends on using full nq-dev or smaller one) than the validation score. I have values CV / LB like (0.36, 0.33), (0.47, 0.45) or (0.55, 0.52)\n\n- For validation code, I copied the logic from [nq_eval.py](https://github.com/google-research-datasets/natural-questions), in particular, the vote for labels among 5 annotations.\n\n- I am only using distilled bert for now. The best LB score I can get is 0.52 (so sad...). I tried to copy some code from [prokaj's kernel(0.48)](https://www.kaggle.com/prokaj/bert-joint-baseline-notebook) and [mmmarchetti's kernel (0.57)](https://www.kaggle.com/mmmarchetti/tensorflow-2-0-bert-yes-no-answers).\n\n- I got OOM issue on collecting results (model logits, flat to python lists) during predicting, so I only save partial results, thus I need to make some logic change to the original baseline kernel.\n\nI hope some of you find it useful or helpful. For me, I am done working on kernels. I won't publish any new kernel in this competition.\n\nIt's good to have something working, but I am at 315th place now, and I have no clear idea how to go up (other than working on Bert Large, which I need to use TPU and have no experience at all...).\n\nGood lucks, everyone!\n\n--------------------------------------------------\n\nSome remarks:\n\n- In this inference + validation kernel, I forgot to`tf.function` to speed up model calculation. For distilled bert, this is not a big problem. If you go for larger models, please use `tf.function` at a proper place. You can find an usage in my [TF2 training kernel](https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models).\n\n- In `compute_predictions(example)` method, which savs the prediction json file,  it returns, unlike in the starter kernel, a list of ScoreSummary objects instead of a single one. And in , `compute_pred_dict()`,  We have\n\n`nq_pred_dict[e.example_id] = [summary.predicted_label for summary in all_summaries]`.\n\nHowever, when doing validation / submission, only `preds[0]` is used.\n",
      "votes": 23
    },
    {
      "id": 704632,
      "postDate": "2019-12-27T18:52:06.930Z",
      "content": "<p>There is a new Google tutorial using TPU to finetune Bert based on latest TF 2.1, not sure if it is helpful to you.\n<a href=\"https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md\">https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md</a></p>",
      "rawMarkdown": "There is a new Google tutorial using TPU to finetune Bert based on latest TF 2.1, not sure if it is helpful to you.\nhttps://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md",
      "votes": 1,
      "replies": [
        {
          "id": 704637,
          "postDate": "2019-12-27T18:55:27.250Z",
          "content": "<p>Thank you</p>",
          "rawMarkdown": " Thank you"
        }
      ]
    },
    {
      "id": 704165,
      "postDate": "2019-12-27T05:57:02.843Z",
      "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> great work, your kernel is a great starter kernel</p>",
      "rawMarkdown": "@yihdarshieh great work, your kernel is a great starter kernel",
      "votes": 1
    },
    {
      "id": 704653,
      "postDate": "2019-12-27T19:36:11.407Z",
      "content": "<p>Wow, a great job.\nI have a few suggestions, but I'll try them out first and merge them with some parts of my code. If I can get a higher score with your code, I'll put suggestions here for you (and everyone).\nAgain, a great job! helped a lot.</p>",
      "rawMarkdown": "Wow, a great job.\nI have a few suggestions, but I'll try them out first and merge them with some parts of my code. If I can get a higher score with your code, I'll put suggestions here for you (and everyone).\nAgain, a great job! helped a lot.",
      "votes": 2,
      "replies": [
        {
          "id": 704670,
          "postDate": "2019-12-27T20:04:42.503Z",
          "content": "<p>Looking forward for feedbacks/suggestions :)</p>",
          "rawMarkdown": "Looking forward for feedbacks/suggestions :)"
        }
      ]
    },
    {
      "id": 705398,
      "postDate": "2019-12-28T21:35:21.973Z",
      "content": "<p>There is one user ask me privately for clarification, and I understand his intention is not for his own benefit, but for not making people feel my kernel having some potential issue.</p>\n\n<p>But to respect the competition rule (hmm, even I am still in 330 place...), I decided to make the communication public without mentioning the identity.</p>\n\n<p>The original question:</p>\n\n<p>&gt; ... how did you handle the custom tokens created by the bert_joint code, such as [Paragraph=1] and [Q]? Did you add them to the pretrained vocab file from HuggingFace? If not, I suspect that is hurting your CV score compared with the top-scoring public kernels.</p>\n\n<p>My answer:</p>\n\n<p>&gt; Thanks for the information. If you go to the <code>nq-competition</code> dataset (created by me), and go to Hugging Face model dirs like <code>bert-base-uncased</code>, <code>bert-large-uncased-whole-word-masking-finetuned-squad</code> or <code>distilbert-base-uncased-distilled-squad</code>, each dir has a <code>vocab.txt</code> file. If you look the content, you can see</p>\n\n<p>...\n[unused97]\n[unused98]\n[UNK]\n[CLS]\n[SEP]\n[MASK]\n[Q]\n[YES]\n[NO]\n[NoLongAnswer]\n[NoShortAnswer]\n[Paragraph=1]\n[Paragraph=2]\n...</p>\n\n<p>&gt; which is just the vocab-nq.txt at the top level. I replace the original Hugging Face vocab.txt by vocab-nq.txt, so there shouldn't be any problem. For people who use other models downloaded by themselves, it's their own responsibility to make sure the correct files being used. </p>",
      "rawMarkdown": "There is one user ask me privately for clarification, and I understand his intention is not for his own benefit, but for not making people feel my kernel having some potential issue.\n\nBut to respect the competition rule (hmm, even I am still in 330 place...), I decided to make the communication public without mentioning the identity.\n\nThe original question:\n\n&gt; ... how did you handle the custom tokens created by the bert_joint code, such as [Paragraph=1] and [Q]? Did you add them to the pretrained vocab file from HuggingFace? If not, I suspect that is hurting your CV score compared with the top-scoring public kernels.\n\nMy answer:\n\n&gt; Thanks for the information. If you go to the `nq-competition` dataset (created by me), and go to Hugging Face model dirs like `bert-base-uncased`, `bert-large-uncased-whole-word-masking-finetuned-squad` or `distilbert-base-uncased-distilled-squad`, each dir has a `vocab.txt` file. If you look the content, you can see\n\n...\n[unused97]\n[unused98]\n[UNK]\n[CLS]\n[SEP]\n[MASK]\n[Q]\n[YES]\n[NO]\n[NoLongAnswer]\n[NoShortAnswer]\n[Paragraph=1]\n[Paragraph=2]\n...\n\n&gt; which is just the vocab-nq.txt at the top level. I replace the original Hugging Face vocab.txt by vocab-nq.txt, so there shouldn't be any problem. For people who use other models downloaded by themselves, it's their own responsibility to make sure the correct files being used. "
    },
    {
      "id": 705092,
      "postDate": "2019-12-28T11:55:32.927Z",
      "content": "<p>For those who successfully integrate my validation code into your own kernel, I would like to know your CV score vs LB score, thanks! I want to know if it also gives expected results for you, since I didn't test it thoroughly.</p>",
      "rawMarkdown": "For those who successfully integrate my validation code into your own kernel, I would like to know your CV score vs LB score, thanks! I want to know if it also gives expected results for you, since I didn't test it thoroughly.",
      "replies": [
        {
          "id": 705178,
          "postDate": "2019-12-28T15:59:17.817Z",
          "content": "<p>Thank you for your wonderful kernels!\nI employ your training kernel with \"bert-base-uncased\" model. I have trained the model for 12 hours and give the result as follows. However, I use my trained model and your inference kernel with LB score is 0.10. Would you like to share your training loss? I will try to correct my training. 🙈 </p>\n\n<p>Loss 1.388112 | Loss_S 1.755851 | Loss_E 1.790547 | Loss_T 0.617944</p>\n\n<p>Acc 0.618678 |  Acc_S 0.551588 |  Acc_E 0.552556 |  Acc_T 0.751891</p>",
          "rawMarkdown": "Thank you for your wonderful kernels!\nI employ your training kernel with \"bert-base-uncased\" model. I have trained the model for 12 hours and give the result as follows. However, I use my trained model and your inference kernel with LB score is 0.10. Would you like to share your training loss? I will try to correct my training. 🙈 \n\nLoss 1.388112 | Loss_S 1.755851 | Loss_E 1.790547 | Loss_T 0.617944\n\n\n Acc 0.618678 |  Acc_S 0.551588 |  Acc_E 0.552556 |  Acc_T 0.751891",
          "votes": 1
        },
        {
          "id": 705181,
          "postDate": "2019-12-28T16:06:09.270Z",
          "content": "<p>I noticed that in your training kernel, the training loss is:</p>\n\n<p>Epoch 1 | Batch 900 | Elapsed Time 16.982293</p>\n\n<p>Loss 0.674758 | Loss_S 0.831863 | Loss_E 0.850084 | Loss_T 0.342327</p>\n\n<p>Acc 0.793434 | Acc_S 0.752204 | Acc_E 0.759684 | Acc_T 0.868413</p>\n\n<p>However, my training loss is quite abnormal.</p>",
          "rawMarkdown": "I noticed that in your training kernel, the training loss is:\n\nEpoch 1 | Batch 900 | Elapsed Time 16.982293\n\n\nLoss 0.674758 | Loss_S 0.831863 | Loss_E 0.850084 | Loss_T 0.342327\n\n\nAcc 0.793434 | Acc_S 0.752204 | Acc_E 0.759684 | Acc_T 0.868413\n\nHowever, my training loss is quite abnormal."
        },
        {
          "id": 705182,
          "postDate": "2019-12-28T16:10:38.550Z",
          "content": "<p>For distilled bert, my 1st epoch gives</p>\n\n<p><code>Epoch 1</code>\n<code>Loss 1.406846 | Loss_S 1.776818 | Loss_E 1.799931 | Loss_T 0.643786</code>\n<code>Acc 0.612386 |  Acc_S 0.546348 |  Acc_E 0.549019 |  Acc_T 0.741791</code></p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models?scriptVersionId=25468430\">https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models?scriptVersionId=25468430</a></p>\n\n<p>The result you found is probably epoch 10 (latest commit). But I think it's already too much, since I got the same public score for epoch 4 and epoch 10.</p>\n\n<p>I don't know if you used my previous version training kernel. If yes, please check</p>\n\n<p><a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/122855\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/122855</a></p>\n\n<p>Also, my inference kernel is updated, you should use the latest version, which is already the case if you use the link in this post.</p>",
          "rawMarkdown": "For distilled bert, my 1st epoch gives\n\n`Epoch 1`\n`Loss 1.406846 | Loss_S 1.776818 | Loss_E 1.799931 | Loss_T 0.643786`\n`Acc 0.612386 |  Acc_S 0.546348 |  Acc_E 0.549019 |  Acc_T 0.741791`\n\n[https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models?scriptVersionId=25468430](https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models?scriptVersionId=25468430)\n\nThe result you found is probably epoch 10 (latest commit). But I think it's already too much, since I got the same public score for epoch 4 and epoch 10.\n\nI don't know if you used my previous version training kernel. If yes, please check\n\n[https://www.kaggle.com/c/tensorflow2-question-answering/discussion/122855](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/122855)\n\nAlso, my inference kernel is updated, you should use the latest version, which is already the case if you use the link in this post.\n",
          "votes": 2
        },
        {
          "id": 705188,
          "postDate": "2019-12-28T16:18:18.847Z",
          "content": "<p>Thank you very much~\nDid you get a 0.52 LB score with epoch 4?</p>",
          "rawMarkdown": "Thank you very much~\nDid you get a 0.52 LB score with epoch 4?"
        },
        {
          "id": 705194,
          "postDate": "2019-12-28T16:26:20.487Z",
          "content": "<p>I have done similar QA tasks before. I think it will be better to make a good classifier, such as textcnn or capsule network. I will try them later. Thank you again for your wonderful kernels~</p>",
          "rawMarkdown": "I have done similar QA tasks before. I think it will be better to make a good classifier, such as textcnn or capsule network. I will try them later. Thank you again for your wonderful kernels~"
        },
        {
          "id": 705196,
          "postDate": "2019-12-28T16:27:42.137Z",
          "content": "<p>With the buggy inference kernel, I got the same 0.45 for epoch 4 and epoch 10, see</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models\">epoch 4</a></p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models\">epoch 10</a></p>\n\n<p>After fixing inference kernel, I only run with epoch 10, not for epoch 4. I spent too much time working on training/inference/validation kernel, and I also had no more GPU quota ...</p>\n\n<p>By the way, when I said distilled bert, it's actually <code>distilbert-base-uncased-distilled-squad</code>, not just <code>distilbert-base-uncased</code>. However, for <code>bert-base-uncased</code>, there is no version of fine tuned on squad. For <code>bert-large-uncased</code>, there is.</p>",
          "rawMarkdown": "With the buggy inference kernel, I got the same 0.45 for epoch 4 and epoch 10, see\n\n[epoch 4](https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models)\n\n[epoch 10](https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models)\n\nAfter fixing inference kernel, I only run with epoch 10, not for epoch 4. I spent too much time working on training/inference/validation kernel, and I also had no more GPU quota ...\n\nBy the way, when I said distilled bert, it's actually `distilbert-base-uncased-distilled-squad`, not just `distilbert-base-uncased`. However, for `bert-base-uncased`, there is no version of fine tuned on squad. For `bert-large-uncased`, there is.",
          "votes": 1
        },
        {
          "id": 705201,
          "postDate": "2019-12-28T16:32:30.857Z",
          "content": "<p>Thank you very much~</p>",
          "rawMarkdown": "Thank you very much~"
        },
        {
          "id": 705298,
          "postDate": "2019-12-28T18:28:53.330Z",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> can you share what was the bug in your inference kernel? My bert-base-uncased  gives same LB after 2 epochs using your inference kernel</p>",
          "rawMarkdown": "@yihdarshieh can you share what was the bug in your inference kernel? My bert-base-uncased  gives same LB after 2 epochs using your inference kernel",
          "votes": 1
        },
        {
          "id": 705313,
          "postDate": "2019-12-28T19:03:59.387Z",
          "content": "<p><a href=\"/axel81\">@axel81</a> , \nThere are quite a lot things involved. For example:</p>\n\n<ul>\n<li><p>See the last part of this reply (about <code>MY_OWN_NQ_DIR</code>)</p></li>\n<li><p>In a <a href=\"https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models?scriptVersionId=25660348\">previous version</a>, the method <code>is_any_pred_ok()</code> doesn't use <code>a vote among 5 annotations</code> to create correct labels. In the latest version, <code>is_pred_ok</code> use vote. But this is validation part, not inference part.</p></li>\n<li><p>In <a href=\"https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models?scriptVersionId=25529516\">version</a>, the <code>compute_pred_dict()</code>, I had</p>\n\n<p><code>example = examples_by_id[example_id]</code>\n    <code>examples.append(EvalExample(example[0], example[1]))</code>\n    <code>examples[-1].features[example_id] = features_by_id[example_id]</code>\n    <code>examples[-1].results[example_id] = raw_results_by_id[example_id]</code></p></li>\n</ul>\n\n<p>So for an nq-example to predict, there was only one <code>raw_result</code> and <code>feature</code> are used, although we have splitted the document into several chunk of length 512. I copied some code  from a file I found online <code>tf2-0-baseline-w-bert-translated-to-tf2-0.py</code>, and I didn't realized it had some problems.</p>\n\n<ul>\n<li>I had an issue in my model definition, which is explained in <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/122855\">this post</a>, but I guess you already use or read the latest corrected version?</li>\n</ul>\n\n<p>These are just some examples I remembered. I can't recall all the thing I fixed.</p>\n\n<p>By the way, <code>bert-base-uncased</code> is not fine-tuned with SQuAD, while I used <code>distilbert-base-uncased-distilled-squad</code> is fine-tuned with SQuAD (by Hugging Face). </p>\n\n<p>By <code>... gives same LB ...</code>, what do you mean? You mean it scores 0.52? I don't know, but also make sure you upload your checkpoint, and you changed accordingly the following lines </p>\n\n<p><code># If you want to use your own .tfrecord or new trained checkpoints,</code>\n<code>you can put them under you own nq dir (MYOWNNQ_DIR)</code></p>\n\n<p><code># Default to NQ_DIR. You have to change it to the dir containing your own working files.</code></p>\n\n<p><code>MY_OWN_NQ_DIR = NQ_DIR</code></p>\n\n<p>You can look the output to see if if your checkpoint is correctly loaded. For me, it looks like</p>\n\n<p>&gt; Latest BertNQ checkpoint restored -- Model trained for 4 epochs`</p>\n\n<p>&gt; /kaggle/input/nq-competition/checkpoints/distilbert-base-uncased-distilled-squad\ncheckpoints/distilbert-base-uncased-distilled-squad</p>\n\n<p>You should see <code>for 2 epochs</code> and <code>bert-base-uncased</code> in your output</p>",
          "rawMarkdown": "@axel81 , \nThere are quite a lot things involved. For example:\n\n- See the last part of this reply (about `MY_OWN_NQ_DIR`)\n\n- In a [previous version](https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models?scriptVersionId=25660348), the method `is_any_pred_ok()` doesn't use `a vote among 5 annotations` to create correct labels. In the latest version, `is_pred_ok` use vote. But this is validation part, not inference part.\n\n- In [version](https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models?scriptVersionId=25529516), the `compute_pred_dict()`, I had\n\n       ` example = examples_by_id[example_id]`\n        `examples.append(EvalExample(example[0], example[1]))`\n        `examples[-1].features[example_id] = features_by_id[example_id]`\n        `examples[-1].results[example_id] = raw_results_by_id[example_id]`\n\nSo for an nq-example to predict, there was only one `raw_result` and `feature` are used, although we have splitted the document into several chunk of length 512. I copied some code  from a file I found online `tf2-0-baseline-w-bert-translated-to-tf2-0.py`, and I didn't realized it had some problems.\n\n- I had an issue in my model definition, which is explained in [this post](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/122855), but I guess you already use or read the latest corrected version?\n\nThese are just some examples I remembered. I can't recall all the thing I fixed.\n\nBy the way, ` bert-base-uncased` is not fine-tuned with SQuAD, while I used `distilbert-base-uncased-distilled-squad` is fine-tuned with SQuAD (by Hugging Face). \n\nBy `... gives same LB ...`, what do you mean? You mean it scores 0.52? I don't know, but also make sure you upload your checkpoint, and you changed accordingly the following lines \n\n`# If you want to use your own .tfrecord or new trained checkpoints,`\n`you can put them under you own nq dir (MYOWNNQ_DIR)`\n\n`# Default to NQ_DIR. You have to change it to the dir containing your own working files.`\n\n`MY_OWN_NQ_DIR = NQ_DIR`\n\nYou can look the output to see if if your checkpoint is correctly loaded. For me, it looks like\n\n&gt; Latest BertNQ checkpoint restored -- Model trained for 4 epochs`\n\n&gt; /kaggle/input/nq-competition/checkpoints/distilbert-base-uncased-distilled-squad\ncheckpoints/distilbert-base-uncased-distilled-squad\n\nYou should see `for 2 epochs` and `bert-base-uncased` in your output",
          "votes": 1
        },
        {
          "id": 705608,
          "postDate": "2019-12-29T06:21:08.010Z",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> thanks for the explanation, it helps.</p>",
          "rawMarkdown": "@yihdarshieh thanks for the explanation, it helps."
        }
      ]
    },
    {
      "id": 705177,
      "postDate": "2019-12-28T15:58:49.170Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 704632,
      "author_name": "Chew Kok Wah",
      "author_url": "",
      "post_date": "2019-12-27T18:52:06.930000",
      "content": "<p>There is a new Google tutorial using TPU to finetune Bert based on latest TF 2.1, not sure if it is helpful to you.\n<a href=\"https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md\">https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 704637,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2019-12-27T18:55:27.250000",
          "content": "<p>Thank you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 704165,
      "author_name": "Ram Ramrakhya",
      "author_url": "",
      "post_date": "2019-12-27T05:57:02.843000",
      "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> great work, your kernel is a great starter kernel</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 704653,
      "author_name": "Marcos Martins Marchetti",
      "author_url": "",
      "post_date": "2019-12-27T19:36:11.407000",
      "content": "<p>Wow, a great job.\nI have a few suggestions, but I'll try them out first and merge them with some parts of my code. If I can get a higher score with your code, I'll put suggestions here for you (and everyone).\nAgain, a great job! helped a lot.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 704670,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2019-12-27T20:04:42.503000",
          "content": "<p>Looking forward for feedbacks/suggestions :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 705398,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2019-12-28T21:35:21.973000",
      "content": "<p>There is one user ask me privately for clarification, and I understand his intention is not for his own benefit, but for not making people feel my kernel having some potential issue.</p>\n\n<p>But to respect the competition rule (hmm, even I am still in 330 place...), I decided to make the communication public without mentioning the identity.</p>\n\n<p>The original question:</p>\n\n<p>&gt; ... how did you handle the custom tokens created by the bert_joint code, such as [Paragraph=1] and [Q]? Did you add them to the pretrained vocab file from HuggingFace? If not, I suspect that is hurting your CV score compared with the top-scoring public kernels.</p>\n\n<p>My answer:</p>\n\n<p>&gt; Thanks for the information. If you go to the <code>nq-competition</code> dataset (created by me), and go to Hugging Face model dirs like <code>bert-base-uncased</code>, <code>bert-large-uncased-whole-word-masking-finetuned-squad</code> or <code>distilbert-base-uncased-distilled-squad</code>, each dir has a <code>vocab.txt</code> file. If you look the content, you can see</p>\n\n<p>...\n[unused97]\n[unused98]\n[UNK]\n[CLS]\n[SEP]\n[MASK]\n[Q]\n[YES]\n[NO]\n[NoLongAnswer]\n[NoShortAnswer]\n[Paragraph=1]\n[Paragraph=2]\n...</p>\n\n<p>&gt; which is just the vocab-nq.txt at the top level. I replace the original Hugging Face vocab.txt by vocab-nq.txt, so there shouldn't be any problem. For people who use other models downloaded by themselves, it's their own responsibility to make sure the correct files being used. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 705092,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2019-12-28T11:55:32.927000",
      "content": "<p>For those who successfully integrate my validation code into your own kernel, I would like to know your CV score vs LB score, thanks! I want to know if it also gives expected results for you, since I didn't test it thoroughly.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 705178,
          "author_name": "Rainbow Cat",
          "author_url": "",
          "post_date": "2019-12-28T15:59:17.817000",
          "content": "<p>Thank you for your wonderful kernels!\nI employ your training kernel with \"bert-base-uncased\" model. I have trained the model for 12 hours and give the result as follows. However, I use my trained model and your inference kernel with LB score is 0.10. Would you like to share your training loss? I will try to correct my training. 🙈 </p>\n\n<p>Loss 1.388112 | Loss_S 1.755851 | Loss_E 1.790547 | Loss_T 0.617944</p>\n\n<p>Acc 0.618678 |  Acc_S 0.551588 |  Acc_E 0.552556 |  Acc_T 0.751891</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 705181,
          "author_name": "Rainbow Cat",
          "author_url": "",
          "post_date": "2019-12-28T16:06:09.270000",
          "content": "<p>I noticed that in your training kernel, the training loss is:</p>\n\n<p>Epoch 1 | Batch 900 | Elapsed Time 16.982293</p>\n\n<p>Loss 0.674758 | Loss_S 0.831863 | Loss_E 0.850084 | Loss_T 0.342327</p>\n\n<p>Acc 0.793434 | Acc_S 0.752204 | Acc_E 0.759684 | Acc_T 0.868413</p>\n\n<p>However, my training loss is quite abnormal.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 705182,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2019-12-28T16:10:38.550000",
          "content": "<p>For distilled bert, my 1st epoch gives</p>\n\n<p><code>Epoch 1</code>\n<code>Loss 1.406846 | Loss_S 1.776818 | Loss_E 1.799931 | Loss_T 0.643786</code>\n<code>Acc 0.612386 |  Acc_S 0.546348 |  Acc_E 0.549019 |  Acc_T 0.741791</code></p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models?scriptVersionId=25468430\">https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models?scriptVersionId=25468430</a></p>\n\n<p>The result you found is probably epoch 10 (latest commit). But I think it's already too much, since I got the same public score for epoch 4 and epoch 10.</p>\n\n<p>I don't know if you used my previous version training kernel. If yes, please check</p>\n\n<p><a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/122855\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/122855</a></p>\n\n<p>Also, my inference kernel is updated, you should use the latest version, which is already the case if you use the link in this post.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 705188,
          "author_name": "Rainbow Cat",
          "author_url": "",
          "post_date": "2019-12-28T16:18:18.847000",
          "content": "<p>Thank you very much~\nDid you get a 0.52 LB score with epoch 4?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 705194,
          "author_name": "Rainbow Cat",
          "author_url": "",
          "post_date": "2019-12-28T16:26:20.487000",
          "content": "<p>I have done similar QA tasks before. I think it will be better to make a good classifier, such as textcnn or capsule network. I will try them later. Thank you again for your wonderful kernels~</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 705196,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2019-12-28T16:27:42.137000",
          "content": "<p>With the buggy inference kernel, I got the same 0.45 for epoch 4 and epoch 10, see</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models\">epoch 4</a></p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models\">epoch 10</a></p>\n\n<p>After fixing inference kernel, I only run with epoch 10, not for epoch 4. I spent too much time working on training/inference/validation kernel, and I also had no more GPU quota ...</p>\n\n<p>By the way, when I said distilled bert, it's actually <code>distilbert-base-uncased-distilled-squad</code>, not just <code>distilbert-base-uncased</code>. However, for <code>bert-base-uncased</code>, there is no version of fine tuned on squad. For <code>bert-large-uncased</code>, there is.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 705201,
          "author_name": "Rainbow Cat",
          "author_url": "",
          "post_date": "2019-12-28T16:32:30.857000",
          "content": "<p>Thank you very much~</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 705298,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2019-12-28T18:28:53.330000",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> can you share what was the bug in your inference kernel? My bert-base-uncased  gives same LB after 2 epochs using your inference kernel</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 705313,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2019-12-28T19:03:59.387000",
          "content": "<p><a href=\"/axel81\">@axel81</a> , \nThere are quite a lot things involved. For example:</p>\n\n<ul>\n<li><p>See the last part of this reply (about <code>MY_OWN_NQ_DIR</code>)</p></li>\n<li><p>In a <a href=\"https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models?scriptVersionId=25660348\">previous version</a>, the method <code>is_any_pred_ok()</code> doesn't use <code>a vote among 5 annotations</code> to create correct labels. In the latest version, <code>is_pred_ok</code> use vote. But this is validation part, not inference part.</p></li>\n<li><p>In <a href=\"https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models?scriptVersionId=25529516\">version</a>, the <code>compute_pred_dict()</code>, I had</p>\n\n<p><code>example = examples_by_id[example_id]</code>\n    <code>examples.append(EvalExample(example[0], example[1]))</code>\n    <code>examples[-1].features[example_id] = features_by_id[example_id]</code>\n    <code>examples[-1].results[example_id] = raw_results_by_id[example_id]</code></p></li>\n</ul>\n\n<p>So for an nq-example to predict, there was only one <code>raw_result</code> and <code>feature</code> are used, although we have splitted the document into several chunk of length 512. I copied some code  from a file I found online <code>tf2-0-baseline-w-bert-translated-to-tf2-0.py</code>, and I didn't realized it had some problems.</p>\n\n<ul>\n<li>I had an issue in my model definition, which is explained in <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/122855\">this post</a>, but I guess you already use or read the latest corrected version?</li>\n</ul>\n\n<p>These are just some examples I remembered. I can't recall all the thing I fixed.</p>\n\n<p>By the way, <code>bert-base-uncased</code> is not fine-tuned with SQuAD, while I used <code>distilbert-base-uncased-distilled-squad</code> is fine-tuned with SQuAD (by Hugging Face). </p>\n\n<p>By <code>... gives same LB ...</code>, what do you mean? You mean it scores 0.52? I don't know, but also make sure you upload your checkpoint, and you changed accordingly the following lines </p>\n\n<p><code># If you want to use your own .tfrecord or new trained checkpoints,</code>\n<code>you can put them under you own nq dir (MYOWNNQ_DIR)</code></p>\n\n<p><code># Default to NQ_DIR. You have to change it to the dir containing your own working files.</code></p>\n\n<p><code>MY_OWN_NQ_DIR = NQ_DIR</code></p>\n\n<p>You can look the output to see if if your checkpoint is correctly loaded. For me, it looks like</p>\n\n<p>&gt; Latest BertNQ checkpoint restored -- Model trained for 4 epochs`</p>\n\n<p>&gt; /kaggle/input/nq-competition/checkpoints/distilbert-base-uncased-distilled-squad\ncheckpoints/distilbert-base-uncased-distilled-squad</p>\n\n<p>You should see <code>for 2 epochs</code> and <code>bert-base-uncased</code> in your output</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 705608,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2019-12-29T06:21:08.010000",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> thanks for the explanation, it helps.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 705177,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-28T15:58:49.170000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "704016": "I finally finished my [TF2 inference + validation kernel](https://www.kaggle.com/yihdarshieh/inference-use-hugging-face-models).\n\n(Please see remarks at the end also, please)\n\n- I published it earlier for some reason, but since then, a lot of bugs or logic improvements are fixed / done. In particular, 2 places of tf.compat.v1 are removed.\n\n- It can run validation on nq-dev (simplified) dataset, or a smaller subset (The first1000 examples).\n\n- For my few tests, the public LB score is about 0.03 ~ 0.05 lower (depends on using full nq-dev or smaller one) than the validation score. I have values CV / LB like (0.36, 0.33), (0.47, 0.45) or (0.55, 0.52)\n\n- For validation code, I copied the logic from [nq_eval.py](https://github.com/google-research-datasets/natural-questions), in particular, the vote for labels among 5 annotations.\n\n- I am only using distilled bert for now. The best LB score I can get is 0.52 (so sad...). I tried to copy some code from [prokaj's kernel(0.48)](https://www.kaggle.com/prokaj/bert-joint-baseline-notebook) and [mmmarchetti's kernel (0.57)](https://www.kaggle.com/mmmarchetti/tensorflow-2-0-bert-yes-no-answers).\n\n- I got OOM issue on collecting results (model logits, flat to python lists) during predicting, so I only save partial results, thus I need to make some logic change to the original baseline kernel.\n\nI hope some of you find it useful or helpful. For me, I am done working on kernels. I won't publish any new kernel in this competition.\n\nIt's good to have something working, but I am at 315th place now, and I have no clear idea how to go up (other than working on Bert Large, which I need to use TPU and have no experience at all...).\n\nGood lucks, everyone!\n\n--------------------------------------------------\n\nSome remarks:\n\n- In this inference + validation kernel, I forgot to`tf.function` to speed up model calculation. For distilled bert, this is not a big problem. If you go for larger models, please use `tf.function` at a proper place. You can find an usage in my [TF2 training kernel](https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models).\n\n- In `compute_predictions(example)` method, which savs the prediction json file,  it returns, unlike in the starter kernel, a list of ScoreSummary objects instead of a single one. And in , `compute_pred_dict()`,  We have\n\n`nq_pred_dict[e.example_id] = [summary.predicted_label for summary in all_summaries]`.\n\nHowever, when doing validation / submission, only `preds[0]` is used.\n",
    "704632": "There is a new Google tutorial using TPU to finetune Bert based on latest TF 2.1, not sure if it is helpful to you.\nhttps://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md",
    "704165": "@yihdarshieh great work, your kernel is a great starter kernel",
    "704653": "Wow, a great job.\nI have a few suggestions, but I'll try them out first and merge them with some parts of my code. If I can get a higher score with your code, I'll put suggestions here for you (and everyone).\nAgain, a great job! helped a lot.",
    "705398": "There is one user ask me privately for clarification, and I understand his intention is not for his own benefit, but for not making people feel my kernel having some potential issue.\n\nBut to respect the competition rule (hmm, even I am still in 330 place...), I decided to make the communication public without mentioning the identity.\n\nThe original question:\n\n&gt; ... how did you handle the custom tokens created by the bert_joint code, such as [Paragraph=1] and [Q]? Did you add them to the pretrained vocab file from HuggingFace? If not, I suspect that is hurting your CV score compared with the top-scoring public kernels.\n\nMy answer:\n\n&gt; Thanks for the information. If you go to the `nq-competition` dataset (created by me), and go to Hugging Face model dirs like `bert-base-uncased`, `bert-large-uncased-whole-word-masking-finetuned-squad` or `distilbert-base-uncased-distilled-squad`, each dir has a `vocab.txt` file. If you look the content, you can see\n\n...\n[unused97]\n[unused98]\n[UNK]\n[CLS]\n[SEP]\n[MASK]\n[Q]\n[YES]\n[NO]\n[NoLongAnswer]\n[NoShortAnswer]\n[Paragraph=1]\n[Paragraph=2]\n...\n\n&gt; which is just the vocab-nq.txt at the top level. I replace the original Hugging Face vocab.txt by vocab-nq.txt, so there shouldn't be any problem. For people who use other models downloaded by themselves, it's their own responsibility to make sure the correct files being used. ",
    "705092": "For those who successfully integrate my validation code into your own kernel, I would like to know your CV score vs LB score, thanks! I want to know if it also gives expected results for you, since I didn't test it thoroughly.",
    "705177": ""
  }
}