{
  "id": 122293,
  "title": "[PyTorch] Question about pre-trained Bert specific to this competition",
  "url": "/competitions/tensorflow2-question-answering/discussion/122293",
  "author_name": "",
  "post_date": "2019-12-19T05:27:30.374153800Z",
  "votes": 4,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hi everyone,\nI was able to run my inference kernel written in <strong>PyTorch</strong> and scored 0.18 and realized I maybe wrong because I'm using 'bert-base-uncased' model from transformer module and fine tuning it.</p>\n\n<h2>Question</h2>\n\n<p>Should I use pre-trained model trained for specifically this Q&amp;A dataset instead of generic bert-base-uncased? If so how can I do that in PyTorch?</p>\n\n<h2>My understanding</h2>\n\n<p>Typical flow of using Bert for your problem is</p>\n\n<ol>\n<li>Choose pre-trained Bert model and load it. (Weights for the Bert model is frozen)</li>\n<li>Add some layers on top of  the Bert model output(s) so that the entire model can serve for your purpose.\n<ul><li>For this competition, we need start, end and ans_type as output.</li></ul></li>\n<li>So your model consists of Bert(frozen) +  your layers (to be trained)</li>\n<li>Fine-tune your model using training data and update weights in your layers.</li>\n<li>Make predictions using the entire mode.</li>\n</ol>\n\n<p><strong>BUT</strong>, in this competition there's <a href=\"https://www.kaggle.com/philculliton/using-tensorflow-2-0-w-bert-on-nq\">starter kernel</a> which is\n- (Probably)  Pre-trained for Q&amp;A dataset.  (Check vocabulary file in <a href=\"https://www.kaggle.com/philculliton/bertjointbaseline\">bertjointbaseline</a>. You see some tokens specific to this competition. YES, NO [Paragraph=1] and etc).\n- Already finished fine tuning part. Because of this people can get better score by improving post-process part w/o any fine tuning by themselves.</p>\n\n<p>I'm new to Bert so I maybe missing something obvious, but I appreciate your thoughts here.\nthanks!</p>",
  "messages": [
    {
      "id": "698348",
      "postDate": "12/19/2019 05:27:30",
      "content": "<p>Hi everyone,\nI was able to run my inference kernel written in <strong>PyTorch</strong> and scored 0.18 and realized I maybe wrong because I'm using 'bert-base-uncased' model from transformer module and fine tuning it.</p>\n\n<h2>Question</h2>\n\n<p>Should I use pre-trained model trained for specifically this Q&amp;A dataset instead of generic bert-base-uncased? If so how can I do that in PyTorch?</p>\n\n<h2>My understanding</h2>\n\n<p>Typical flow of using Bert for your problem is</p>\n\n<ol>\n<li>Choose pre-trained Bert model and load it. (Weights for the Bert model is frozen)</li>\n<li>Add some layers on top of  the Bert model output(s) so that the entire model can serve for your purpose.\n<ul><li>For this competition, we need start, end and ans_type as output.</li></ul></li>\n<li>So your model consists of Bert(frozen) +  your layers (to be trained)</li>\n<li>Fine-tune your model using training data and update weights in your layers.</li>\n<li>Make predictions using the entire mode.</li>\n</ol>\n\n<p><strong>BUT</strong>, in this competition there's <a href=\"https://www.kaggle.com/philculliton/using-tensorflow-2-0-w-bert-on-nq\">starter kernel</a> which is\n- (Probably)  Pre-trained for Q&amp;A dataset.  (Check vocabulary file in <a href=\"https://www.kaggle.com/philculliton/bertjointbaseline\">bertjointbaseline</a>. You see some tokens specific to this competition. YES, NO [Paragraph=1] and etc).\n- Already finished fine tuning part. Because of this people can get better score by improving post-process part w/o any fine tuning by themselves.</p>\n\n<p>I'm new to Bert so I maybe missing something obvious, but I appreciate your thoughts here.\nthanks!</p>",
      "rawMarkdown": "Hi everyone,\nI was able to run my inference kernel written in __PyTorch__ and scored 0.18 and realized I maybe wrong because I'm using 'bert-base-uncased' model from transformer module and fine tuning it.\n\n## Question\nShould I use pre-trained model trained for specifically this Q&amp;A dataset instead of generic bert-base-uncased? If so how can I do that in PyTorch?\n\n## My understanding\nTypical flow of using Bert for your problem is\n\n1.  Choose pre-trained Bert model and load it. (Weights for the Bert model is frozen)\n1. Add some layers on top of  the Bert model output(s) so that the entire model can serve for your purpose.\n  - For this competition, we need start, end and ans_type as output.\n1. So your model consists of Bert(frozen) +  your layers (to be trained)\n1. Fine-tune your model using training data and update weights in your layers.\n1. Make predictions using the entire mode.\n\n\n__BUT__, in this competition there's [starter kernel](https://www.kaggle.com/philculliton/using-tensorflow-2-0-w-bert-on-nq) which is\n- (Probably)  Pre-trained for Q&amp;A dataset.  (Check vocabulary file in [bertjointbaseline](https://www.kaggle.com/philculliton/bertjointbaseline). You see some tokens specific to this competition. YES, NO [Paragraph=1] and etc).\n- Already finished fine tuning part. Because of this people can get better score by improving post-process part w/o any fine tuning by themselves.\n\nI'm new to Bert so I maybe missing something obvious, but I appreciate your thoughts here.\nthanks!",
      "votes": null
    },
    {
      "id": "698705",
      "postDate": "12/19/2019 15:48:31",
      "content": "<p>&gt; Should I use pre-trained model trained for specifically this Q&amp;A dataset instead of generic bert-base-uncased? If so how can I do that in PyTorch?</p>\n\n<p>If you have parameters pre-trained by QA dataset, it might be better to use it. For example, <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#training-our-model\">bert-joint baseline use model initialized by BERT and pre-trained with SQuAD 2.0</a>.</p>\n\n<p>As for my case, I'm using original BERT and just finetuned it.</p>\n\n<p>&gt; BUT, in this competition there's starter kernel which is\n- (Probably) Pre-trained for Q&amp;A dataset. (Check vocabulary file in bertjointbaseline. You see some tokens specific to this competition. YES, NO [Paragraph=1] and etc).\n- Already finished fine tuning part. Because of this people can get better score by improving post-process part w/o any fine tuning by themselves.</p>\n\n<p>Yes. Starter kernel is using trained parameters by <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#evaluating-our-pretrained-model\">bert-joint</a>, which is already trained by NQ dataset. You don't need more finetune if you use this.</p>\n\n<p>Your understanding about BERT flow is almost correct IMO, and there are many options to improve a model (e.g. not/partly freeze bert parameters, add more layers, add other targets, ...).</p>",
      "rawMarkdown": "&gt; Should I use pre-trained model trained for specifically this Q&amp;A dataset instead of generic bert-base-uncased? If so how can I do that in PyTorch?\n\nIf you have parameters pre-trained by QA dataset, it might be better to use it. For example, [bert-joint baseline use model initialized by BERT and pre-trained with SQuAD 2.0](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#training-our-model).\n\nAs for my case, I'm using original BERT and just finetuned it.\n\n&gt; BUT, in this competition there's starter kernel which is\n- (Probably) Pre-trained for Q&amp;A dataset. (Check vocabulary file in bertjointbaseline. You see some tokens specific to this competition. YES, NO [Paragraph=1] and etc).\n- Already finished fine tuning part. Because of this people can get better score by improving post-process part w/o any fine tuning by themselves.\n\nYes. Starter kernel is using trained parameters by [bert-joint](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#evaluating-our-pretrained-model), which is already trained by NQ dataset. You don't need more finetune if you use this.\n\nYour understanding about BERT flow is almost correct IMO, and there are many options to improve a model (e.g. not/partly freeze bert parameters, add more layers, add other targets, ...).",
      "votes": null
    },
    {
      "id": "698937",
      "postDate": "12/19/2019 23:12:51",
      "content": "<p>Thank you so much for your answer. I'm glad that I'm in the right track.\nDid you get your best score 0.60 by \"I'm using original BERT and just finetuned it.\"? Then I have tons of room to improve :)</p>\n\n<p>thank you!</p>",
      "rawMarkdown": "Thank you so much for your answer. I'm glad that I'm in the right track.\nDid you get your best score 0.60 by \"I'm using original BERT and just finetuned it.\"? Then I have tons of room to improve :)\n\nthank you!",
      "votes": null
    },
    {
      "id": "699081",
      "postDate": "12/20/2019 03:18:01",
      "content": "<p>Yes, but I think this is just overfit to public test dataset 😭 </p>",
      "rawMarkdown": "Yes, but I think this is just overfit to public test dataset 😭",
      "votes": null
    },
    {
      "id": "699138",
      "postDate": "12/20/2019 04:51:29",
      "content": "<p>Thank you. We'll see soon if it's overfit.</p>",
      "rawMarkdown": "Thank you. We'll see soon if it's overfit.",
      "votes": null
    },
    {
      "id": "704928",
      "postDate": "12/28/2019 07:01:11",
      "content": "<p><a href=\"/kentaronakanishi\">@kentaronakanishi</a> are you using <code>bert-base</code> or <code>bert-large</code> I am trying to finetune <code>bert-base</code> but I am not able to get a score above 0.5 with postprocessing. Any suggestions?</p>",
      "rawMarkdown": "kentaronakanishi are you using `bert-base` or `bert-large` I am trying to finetune `bert-base` but I am not able to get a score above 0.5 with postprocessing. Any suggestions?",
      "votes": null
    },
    {
      "id": "704939",
      "postDate": "12/28/2019 07:25:56",
      "content": "<p>I'd like to know it as well.\nBut FYI I'm getting 0.52 LB by finetuning 'bert-base-uncased'.</p>",
      "rawMarkdown": "I'd like to know it as well.\nBut FYI I'm getting 0.52 LB by finetuning 'bert-base-uncased'.",
      "votes": null
    },
    {
      "id": "705621",
      "postDate": "12/29/2019 07:06:29",
      "content": "<p>I used <code>bert-large</code> with word whole masking pretrained model.</p>",
      "rawMarkdown": "I used `bert-large` with word whole masking pretrained model.",
      "votes": null
    },
    {
      "id": "708361",
      "postDate": "01/02/2020 08:14:33",
      "content": "<p>hello, I'm new here. I also have this question. The original bert doesn't have these special tokens([context=-1], [Table=1]). So how to process these special tokens? I see the answer above is to pretrain bert on the NQ dataset, but how to add these special tokens in original bert? Thanks!</p>",
      "rawMarkdown": "hello, I'm new here. I also have this question. The original bert doesn't have these special tokens([context=-1], [Table=1]). So how to process these special tokens? I see the answer above is to pretrain bert on the NQ dataset, but how to add these special tokens in original bert? Thanks!",
      "votes": null
    },
    {
      "id": "708408",
      "postDate": "01/02/2020 09:25:28",
      "content": "<p><a href=\"/kscp123\">@kscp123</a> you can find the new vocabulary with these special tokens in the dataset shared by organizers of contest. You can find it in their public kernel</p>",
      "rawMarkdown": "kscp123 you can find the new vocabulary with these special tokens in the dataset shared by organizers of contest. You can find it in their public kernel",
      "votes": null
    },
    {
      "id": "708414",
      "postDate": "01/02/2020 09:29:58",
      "content": "<p><a href=\"/axel81\">@axel81</a> I wonder how you can train 'bert-uncased' with the vocabulary.txt. I guess you'll have to train bert from scratch as it needs to learn new embeddings?</p>",
      "rawMarkdown": "axel81 I wonder how you can train 'bert-uncased' with the vocabulary.txt. I guess you'll have to train bert from scratch as it needs to learn new embeddings?",
      "votes": null
    },
    {
      "id": "708440",
      "postDate": "01/02/2020 09:57:48",
      "content": "<p><a href=\"/higepon\">@higepon</a> you can do that or you can add new tokens in the vocab.txt for bert-uncased and then finetune only on train data</p>",
      "rawMarkdown": "higepon you can do that or you can add new tokens in the vocab.txt for bert-uncased and then finetune only on train data",
      "votes": null
    },
    {
      "id": "708453",
      "postDate": "01/02/2020 10:09:38",
      "content": "<blockquote>\n  <p>you can add new tokens in the vocab.txt for bert-uncased and then finetune only on train data</p>\n</blockquote>\n\n<p>How the embedding works in this case? Bert will learn new embeddings through finetune? Does it mean you un-freeze Bert layers?</p>",
      "rawMarkdown": "&gt; you can add new tokens in the vocab.txt for bert-uncased and then finetune only on train data\n\nHow the embedding works in this case? Bert will learn new embeddings through finetune? Does it mean you un-freeze Bert layers?",
      "votes": null
    },
    {
      "id": "708460",
      "postDate": "01/02/2020 10:13:44",
      "content": "<p>yes, but I haven't trained bert base so I am not sure how will it perform</p>",
      "rawMarkdown": "yes, but I haven't trained bert base so I am not sure how will it perform",
      "votes": null
    },
    {
      "id": "709058",
      "postDate": "01/03/2020 01:39:53",
      "content": "<p><a href=\"/axel81\">@axel81</a> If use original bert, these special tokens are not in the original bert vocab. Doesn't bert see them as unk? Though you add them in the vocab.txt. </p>",
      "rawMarkdown": "axel81 If use original bert, these special tokens are not in the original bert vocab. Doesn't bert see them as unk? Though you add them in the vocab.txt.",
      "votes": null
    },
    {
      "id": "709129",
      "postDate": "01/03/2020 04:24:05",
      "content": "<p><a href=\"/kscp123\">@kscp123</a> you have to use the <code>nq-vocab.txt</code> for your bert tokenizer</p>",
      "rawMarkdown": "kscp123 you have to use the `nq-vocab.txt` for your bert tokenizer",
      "votes": null
    },
    {
      "id": "709170",
      "postDate": "01/03/2020 05:46:01",
      "content": "<p><a href=\"/axel81\">@axel81</a>  I know use nq-vocab.txt and bert tokenizer will cut the special tokens correctly. But the pretrained bert has not seen these special tokens when it was trained. How could bert give right embedding for these special tokens? Or I got wrong understanding about bert. Please correct me.</p>",
      "rawMarkdown": "axel81  I know use nq-vocab.txt and bert tokenizer will cut the special tokens correctly. But the pretrained bert has not seen these special tokens when it was trained. How could bert give right embedding for these special tokens? Or I got wrong understanding about bert. Please correct me.",
      "votes": null
    },
    {
      "id": "709178",
      "postDate": "01/03/2020 05:56:58",
      "content": "<p>As pretrain bert has not seen this tokens you will finetune it to learn that.</p>",
      "rawMarkdown": "As pretrain bert has not seen this tokens you will finetune it to learn that.",
      "votes": null
    },
    {
      "id": "709258",
      "postDate": "01/03/2020 09:17:22",
      "content": "<p><a href=\"/axel81\">@axel81</a> Thanks for your explanation, it helps a lot.</p>",
      "rawMarkdown": "axel81 Thanks for your explanation, it helps a lot.",
      "votes": null
    },
    {
      "id": "710715",
      "postDate": "01/05/2020 05:52:08",
      "content": "<p><a href=\"/higepon\">@higepon</a> It's amazing to see you went up so high in the LB. I entered the competition late and trying to decide whether I should put my effort into this competition right now. I just have two questions:\n1. Did you get to your current position using Pytorch\n2. What kind of hardware did you use? Was it TPU</p>",
      "rawMarkdown": "higepon It's amazing to see you went up so high in the LB. I entered the competition late and trying to decide whether I should put my effort into this competition right now. I just have two questions:\n1. Did you get to your current position using Pytorch\n2. What kind of hardware did you use? Was it TPU",
      "votes": null
    },
    {
      "id": "710814",
      "postDate": "01/05/2020 08:57:16",
      "content": "<p>Hi,\nI’m using PyTorch and GPU on GCP.</p>",
      "rawMarkdown": "Hi,\nI’m using PyTorch and GPU on GCP.",
      "votes": null
    },
    {
      "id": "710872",
      "postDate": "01/05/2020 11:11:26",
      "content": "<p><a href=\"/kentaronakanishi\">@kentaronakanishi</a> </p>\n\n<blockquote>\n  <p>Your understanding about BERT flow is almost correct IMO, and there are many options to improve a model (e.g. not/partly freeze bert parameters, add more layers, add other targets, …).&gt; </p>\n</blockquote>\n\n<p>What kind of additional layers can I possibly add for fine-tuning without unfreezing BERT original? It'd be great if you can put more light on that.</p>",
      "rawMarkdown": "kentaronakanishi \n&gt; Your understanding about BERT flow is almost correct IMO, and there are many options to improve a model (e.g. not/partly freeze bert parameters, add more layers, add other targets, …).&gt; \n\nWhat kind of additional layers can I possibly add for fine-tuning without unfreezing BERT original? It'd be great if you can put more light on that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 698705,
      "author_name": "kentaronakanishi",
      "author_url": "",
      "post_date": "12/19/2019 15:48:31",
      "content": "<p>&gt; Should I use pre-trained model trained for specifically this Q&amp;A dataset instead of generic bert-base-uncased? If so how can I do that in PyTorch?</p>\n\n<p>If you have parameters pre-trained by QA dataset, it might be better to use it. For example, <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#training-our-model\">bert-joint baseline use model initialized by BERT and pre-trained with SQuAD 2.0</a>.</p>\n\n<p>As for my case, I'm using original BERT and just finetuned it.</p>\n\n<p>&gt; BUT, in this competition there's starter kernel which is\n- (Probably) Pre-trained for Q&amp;A dataset. (Check vocabulary file in bertjointbaseline. You see some tokens specific to this competition. YES, NO [Paragraph=1] and etc).\n- Already finished fine tuning part. Because of this people can get better score by improving post-process part w/o any fine tuning by themselves.</p>\n\n<p>Yes. Starter kernel is using trained parameters by <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#evaluating-our-pretrained-model\">bert-joint</a>, which is already trained by NQ dataset. You don't need more finetune if you use this.</p>\n\n<p>Your understanding about BERT flow is almost correct IMO, and there are many options to improve a model (e.g. not/partly freeze bert parameters, add more layers, add other targets, ...).</p>",
      "votes": null,
      "replies": [
        {
          "id": 698937,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/19/2019 23:12:51",
          "content": "<p>Thank you so much for your answer. I'm glad that I'm in the right track.\nDid you get your best score 0.60 by \"I'm using original BERT and just finetuned it.\"? Then I have tons of room to improve :)</p>\n\n<p>thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699081,
          "author_name": "kentaronakanishi",
          "author_url": "",
          "post_date": "12/20/2019 03:18:01",
          "content": "<p>Yes, but I think this is just overfit to public test dataset 😭 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699138,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/20/2019 04:51:29",
          "content": "<p>Thank you. We'll see soon if it's overfit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 704928,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "12/28/2019 07:01:11",
          "content": "<p><a href=\"/kentaronakanishi\">@kentaronakanishi</a> are you using <code>bert-base</code> or <code>bert-large</code> I am trying to finetune <code>bert-base</code> but I am not able to get a score above 0.5 with postprocessing. Any suggestions?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 704939,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/28/2019 07:25:56",
          "content": "<p>I'd like to know it as well.\nBut FYI I'm getting 0.52 LB by finetuning 'bert-base-uncased'.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 705621,
          "author_name": "kentaronakanishi",
          "author_url": "",
          "post_date": "12/29/2019 07:06:29",
          "content": "<p>I used <code>bert-large</code> with word whole masking pretrained model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710872,
          "author_name": "rajraviprajapat",
          "author_url": "",
          "post_date": "01/05/2020 11:11:26",
          "content": "<p><a href=\"/kentaronakanishi\">@kentaronakanishi</a> </p>\n\n<blockquote>\n  <p>Your understanding about BERT flow is almost correct IMO, and there are many options to improve a model (e.g. not/partly freeze bert parameters, add more layers, add other targets, …).&gt; </p>\n</blockquote>\n\n<p>What kind of additional layers can I possibly add for fine-tuning without unfreezing BERT original? It'd be great if you can put more light on that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 708361,
      "author_name": "",
      "author_url": "",
      "post_date": "01/02/2020 08:14:33",
      "content": "<p>hello, I'm new here. I also have this question. The original bert doesn't have these special tokens([context=-1], [Table=1]). So how to process these special tokens? I see the answer above is to pretrain bert on the NQ dataset, but how to add these special tokens in original bert? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 708408,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/02/2020 09:25:28",
          "content": "<p><a href=\"/kscp123\">@kscp123</a> you can find the new vocabulary with these special tokens in the dataset shared by organizers of contest. You can find it in their public kernel</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708414,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/02/2020 09:29:58",
          "content": "<p><a href=\"/axel81\">@axel81</a> I wonder how you can train 'bert-uncased' with the vocabulary.txt. I guess you'll have to train bert from scratch as it needs to learn new embeddings?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708440,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/02/2020 09:57:48",
          "content": "<p><a href=\"/higepon\">@higepon</a> you can do that or you can add new tokens in the vocab.txt for bert-uncased and then finetune only on train data</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708453,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/02/2020 10:09:38",
          "content": "<blockquote>\n  <p>you can add new tokens in the vocab.txt for bert-uncased and then finetune only on train data</p>\n</blockquote>\n\n<p>How the embedding works in this case? Bert will learn new embeddings through finetune? Does it mean you un-freeze Bert layers?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708460,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/02/2020 10:13:44",
          "content": "<p>yes, but I haven't trained bert base so I am not sure how will it perform</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709058,
          "author_name": "",
          "author_url": "",
          "post_date": "01/03/2020 01:39:53",
          "content": "<p><a href=\"/axel81\">@axel81</a> If use original bert, these special tokens are not in the original bert vocab. Doesn't bert see them as unk? Though you add them in the vocab.txt. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709129,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/03/2020 04:24:05",
          "content": "<p><a href=\"/kscp123\">@kscp123</a> you have to use the <code>nq-vocab.txt</code> for your bert tokenizer</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709170,
          "author_name": "",
          "author_url": "",
          "post_date": "01/03/2020 05:46:01",
          "content": "<p><a href=\"/axel81\">@axel81</a>  I know use nq-vocab.txt and bert tokenizer will cut the special tokens correctly. But the pretrained bert has not seen these special tokens when it was trained. How could bert give right embedding for these special tokens? Or I got wrong understanding about bert. Please correct me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709178,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/03/2020 05:56:58",
          "content": "<p>As pretrain bert has not seen this tokens you will finetune it to learn that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709258,
          "author_name": "",
          "author_url": "",
          "post_date": "01/03/2020 09:17:22",
          "content": "<p><a href=\"/axel81\">@axel81</a> Thanks for your explanation, it helps a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 710715,
      "author_name": "wenrui29",
      "author_url": "",
      "post_date": "01/05/2020 05:52:08",
      "content": "<p><a href=\"/higepon\">@higepon</a> It's amazing to see you went up so high in the LB. I entered the competition late and trying to decide whether I should put my effort into this competition right now. I just have two questions:\n1. Did you get to your current position using Pytorch\n2. What kind of hardware did you use? Was it TPU</p>",
      "votes": null,
      "replies": [
        {
          "id": 710814,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/05/2020 08:57:16",
          "content": "<p>Hi,\nI’m using PyTorch and GPU on GCP.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "698348": "Hi everyone,\nI was able to run my inference kernel written in __PyTorch__ and scored 0.18 and realized I maybe wrong because I'm using 'bert-base-uncased' model from transformer module and fine tuning it.\n\n## Question\nShould I use pre-trained model trained for specifically this Q&amp;A dataset instead of generic bert-base-uncased? If so how can I do that in PyTorch?\n\n## My understanding\nTypical flow of using Bert for your problem is\n\n1.  Choose pre-trained Bert model and load it. (Weights for the Bert model is frozen)\n1. Add some layers on top of  the Bert model output(s) so that the entire model can serve for your purpose.\n  - For this competition, we need start, end and ans_type as output.\n1. So your model consists of Bert(frozen) +  your layers (to be trained)\n1. Fine-tune your model using training data and update weights in your layers.\n1. Make predictions using the entire mode.\n\n\n__BUT__, in this competition there's [starter kernel](https://www.kaggle.com/philculliton/using-tensorflow-2-0-w-bert-on-nq) which is\n- (Probably)  Pre-trained for Q&amp;A dataset.  (Check vocabulary file in [bertjointbaseline](https://www.kaggle.com/philculliton/bertjointbaseline). You see some tokens specific to this competition. YES, NO [Paragraph=1] and etc).\n- Already finished fine tuning part. Because of this people can get better score by improving post-process part w/o any fine tuning by themselves.\n\nI'm new to Bert so I maybe missing something obvious, but I appreciate your thoughts here.\nthanks!",
    "698705": "&gt; Should I use pre-trained model trained for specifically this Q&amp;A dataset instead of generic bert-base-uncased? If so how can I do that in PyTorch?\n\nIf you have parameters pre-trained by QA dataset, it might be better to use it. For example, [bert-joint baseline use model initialized by BERT and pre-trained with SQuAD 2.0](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#training-our-model).\n\nAs for my case, I'm using original BERT and just finetuned it.\n\n&gt; BUT, in this competition there's starter kernel which is\n- (Probably) Pre-trained for Q&amp;A dataset. (Check vocabulary file in bertjointbaseline. You see some tokens specific to this competition. YES, NO [Paragraph=1] and etc).\n- Already finished fine tuning part. Because of this people can get better score by improving post-process part w/o any fine tuning by themselves.\n\nYes. Starter kernel is using trained parameters by [bert-joint](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#evaluating-our-pretrained-model), which is already trained by NQ dataset. You don't need more finetune if you use this.\n\nYour understanding about BERT flow is almost correct IMO, and there are many options to improve a model (e.g. not/partly freeze bert parameters, add more layers, add other targets, ...).",
    "698937": "Thank you so much for your answer. I'm glad that I'm in the right track.\nDid you get your best score 0.60 by \"I'm using original BERT and just finetuned it.\"? Then I have tons of room to improve :)\n\nthank you!",
    "699081": "Yes, but I think this is just overfit to public test dataset 😭",
    "699138": "Thank you. We'll see soon if it's overfit.",
    "704928": "kentaronakanishi are you using `bert-base` or `bert-large` I am trying to finetune `bert-base` but I am not able to get a score above 0.5 with postprocessing. Any suggestions?",
    "704939": "I'd like to know it as well.\nBut FYI I'm getting 0.52 LB by finetuning 'bert-base-uncased'.",
    "705621": "I used `bert-large` with word whole masking pretrained model.",
    "708361": "hello, I'm new here. I also have this question. The original bert doesn't have these special tokens([context=-1], [Table=1]). So how to process these special tokens? I see the answer above is to pretrain bert on the NQ dataset, but how to add these special tokens in original bert? Thanks!",
    "708408": "kscp123 you can find the new vocabulary with these special tokens in the dataset shared by organizers of contest. You can find it in their public kernel",
    "708414": "axel81 I wonder how you can train 'bert-uncased' with the vocabulary.txt. I guess you'll have to train bert from scratch as it needs to learn new embeddings?",
    "708440": "higepon you can do that or you can add new tokens in the vocab.txt for bert-uncased and then finetune only on train data",
    "708453": "&gt; you can add new tokens in the vocab.txt for bert-uncased and then finetune only on train data\n\nHow the embedding works in this case? Bert will learn new embeddings through finetune? Does it mean you un-freeze Bert layers?",
    "708460": "yes, but I haven't trained bert base so I am not sure how will it perform",
    "709058": "axel81 If use original bert, these special tokens are not in the original bert vocab. Doesn't bert see them as unk? Though you add them in the vocab.txt.",
    "709129": "kscp123 you have to use the `nq-vocab.txt` for your bert tokenizer",
    "709170": "axel81  I know use nq-vocab.txt and bert tokenizer will cut the special tokens correctly. But the pretrained bert has not seen these special tokens when it was trained. How could bert give right embedding for these special tokens? Or I got wrong understanding about bert. Please correct me.",
    "709178": "As pretrain bert has not seen this tokens you will finetune it to learn that.",
    "709258": "axel81 Thanks for your explanation, it helps a lot.",
    "710715": "higepon It's amazing to see you went up so high in the LB. I entered the competition late and trying to decide whether I should put my effort into this competition right now. I just have two questions:\n1. Did you get to your current position using Pytorch\n2. What kind of hardware did you use? Was it TPU",
    "710814": "Hi,\nI’m using PyTorch and GPU on GCP.",
    "710872": "kentaronakanishi \n&gt; Your understanding about BERT flow is almost correct IMO, and there are many options to improve a model (e.g. not/partly freeze bert parameters, add more layers, add other targets, …).&gt; \n\nWhat kind of additional layers can I possibly add for fine-tuning without unfreezing BERT original? It'd be great if you can put more light on that."
  },
  "source": "meta"
}