{
  "id": 122799,
  "title": "BERT Base vs/ BERT Large",
  "url": "/competitions/tensorflow2-question-answering/discussion/122799",
  "author_name": "",
  "post_date": "2019-12-23T00:56:21.411230400Z",
  "votes": 3,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Has anyone seen any improvements by using BERT Large instead of Base?\nFor example we might get some gain by using \"bert-large-uncased-whole-word-masking-finetuned-squad\".</p>\n\n<p>Related topic\n<a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/92977\">BERT Base vs/ BERT Large - Jigsaw Unintended Bias in Toxicity Classification</a></p>",
  "messages": [
    {
      "id": "701003",
      "postDate": "12/23/2019 00:56:21",
      "content": "<p>Has anyone seen any improvements by using BERT Large instead of Base?\nFor example we might get some gain by using \"bert-large-uncased-whole-word-masking-finetuned-squad\".</p>\n\n<p>Related topic\n<a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/92977\">BERT Base vs/ BERT Large - Jigsaw Unintended Bias in Toxicity Classification</a></p>",
      "rawMarkdown": "Has anyone seen any improvements by using BERT Large instead of Base?\nFor example we might get some gain by using \"bert-large-uncased-whole-word-masking-finetuned-squad\".\n\nRelated topic\n[BERT Base vs/ BERT Large - Jigsaw Unintended Bias in Toxicity Classification](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/92977)",
      "votes": null
    },
    {
      "id": "701159",
      "postDate": "12/23/2019 06:59:11",
      "content": "<p>I was running training using \"bert-large-uncased-whole-word-masking-finetuned-squad\". But my GCP instance died lol.  (batch_size = 4 and accumulation_steps = 8).</p>",
      "rawMarkdown": "I was running training using \"bert-large-uncased-whole-word-masking-finetuned-squad\". But my GCP instance died lol.  (batch_size = 4 and accumulation_steps = 8).",
      "votes": null
    },
    {
      "id": "701165",
      "postDate": "12/23/2019 07:07:20",
      "content": "<p>how were your initial results? I was planning on trying bert-large</p>",
      "rawMarkdown": "how were your initial results? I was planning on trying bert-large",
      "votes": null
    },
    {
      "id": "701186",
      "postDate": "12/23/2019 07:46:40",
      "content": "<p>I was getting 0.20 ish public score for bert_uncased. And I think I found a bug on my code and I'm re-retraining it right now.</p>",
      "rawMarkdown": "I was getting 0.20 ish public score for bert_uncased. And I think I found a bug on my code and I'm re-retraining it right now.",
      "votes": null
    },
    {
      "id": "701586",
      "postDate": "12/23/2019 16:30:19",
      "content": "<p>I tried using bert large but it doesn't fit in P100 GPU memory, we need TPUs</p>",
      "rawMarkdown": "I tried using bert large but it doesn't fit in P100 GPU memory, we need TPUs",
      "votes": null
    },
    {
      "id": "703502",
      "postDate": "12/26/2019 09:07:08",
      "content": "<p>I wonder how to get \"bert-large-uncased-whole-word-masking-finetuned-squad\" ? I didn't find this model on original bert github page. Can you help me?</p>",
      "rawMarkdown": "I wonder how to get \"bert-large-uncased-whole-word-masking-finetuned-squad\" ? I didn't find this model on original bert github page. Can you help me?",
      "votes": null
    },
    {
      "id": "703556",
      "postDate": "12/26/2019 09:42:12",
      "content": "<p>You can find it in <a href=\"https://huggingface.co/transformers/pretrained_models.html\">https://huggingface.co/transformers/pretrained_models.html</a> :)</p>",
      "rawMarkdown": "You can find it in https://huggingface.co/transformers/pretrained_models.html :)",
      "votes": null
    },
    {
      "id": "703633",
      "postDate": "12/26/2019 11:54:47",
      "content": "<p>Thanks for your reply!\nBTW, I think the bert-joint model was trained on bert-large. The <code>bert_config.json</code> file in <a href=\"https://www.kaggle.com/philculliton/bertjointbaseline\">bert-joint-baseline</a> dataset was described as following:\n<code>\nattention_probs_dropout_prob:0.1\nhidden_act:gelu\nhidden_dropout_prob:0.1\nhidden_size:1024\ninitializer_range:0.02\nintermediate_size:4096\nmax_position_embeddings:512\nnum_attention_heads:16\nnum_hidden_layers:24\ntype_vocab_size:2\nvocab_size:30522\n</code></p>",
      "rawMarkdown": "Thanks for your reply!\nBTW, I think the bert-joint model was trained on bert-large. The `bert_config.json` file in [bert-joint-baseline](https://www.kaggle.com/philculliton/bertjointbaseline) dataset was described as following:\n```\nattention_probs_dropout_prob:0.1\nhidden_act:gelu\nhidden_dropout_prob:0.1\nhidden_size:1024\ninitializer_range:0.02\nintermediate_size:4096\nmax_position_embeddings:512\nnum_attention_heads:16\nnum_hidden_layers:24\ntype_vocab_size:2\nvocab_size:30522\n```",
      "votes": null
    },
    {
      "id": "703638",
      "postDate": "12/26/2019 12:09:30",
      "content": "<p>Wow. thanks for the info. </p>",
      "rawMarkdown": "Wow. thanks for the info.",
      "votes": null
    },
    {
      "id": "707115",
      "postDate": "12/31/2019 09:09:57",
      "content": "<p>Update. I got 0.55 with large model when it was 0.52 for bert base.</p>",
      "rawMarkdown": "Update. I got 0.55 with large model when it was 0.52 for bert base.",
      "votes": null
    },
    {
      "id": "707206",
      "postDate": "12/31/2019 12:14:24",
      "content": "<p><a href=\"/higepon\">@higepon</a> that's great, my first version gives 0.58. I think there is a lot of room to improve. I think my postprocessing needs improvements.</p>",
      "rawMarkdown": "higepon that's great, my first version gives 0.58. I think there is a lot of room to improve. I think my postprocessing needs improvements.",
      "votes": null
    },
    {
      "id": "707771",
      "postDate": "01/01/2020 13:09:58",
      "content": "<p>I'm training a bert-large(float16)with batch_size32, accumulationsteps=4 with 44G GPU. cost 26hours for 1 epoch.</p>",
      "rawMarkdown": "I'm training a bert-large(float16)with batch_size32, accumulationsteps=4 with 44G GPU. cost 26hours for 1 epoch.",
      "votes": null
    },
    {
      "id": "708209",
      "postDate": "01/02/2020 04:34:12",
      "content": "<p><a href=\"/mcggood\">@mcggood</a> you have access to great resources</p>",
      "rawMarkdown": "mcggood you have access to great resources",
      "votes": null
    },
    {
      "id": "708293",
      "postDate": "01/02/2020 06:39:00",
      "content": "<p>yeah, but I failed. my 1 epoch model is predicting same value in matrix. \nI see top teams using TPU with 1 hour per epoch.\nmaybe TPU is better than GPU</p>",
      "rawMarkdown": "yeah, but I failed. my 1 epoch model is predicting same value in matrix. \nI see top teams using TPU with 1 hour per epoch.\nmaybe TPU is better than GPU",
      "votes": null
    },
    {
      "id": "708318",
      "postDate": "01/02/2020 07:19:56",
      "content": "<p>What? 1 hour per epoch with TPU...</p>",
      "rawMarkdown": "What? 1 hour per epoch with TPU...",
      "votes": null
    },
    {
      "id": "708466",
      "postDate": "01/02/2020 10:18:37",
      "content": "<p>Both the large and base models are pre-trained? Or you trained them by yourself?</p>",
      "rawMarkdown": "Both the large and base models are pre-trained? Or you trained them by yourself?",
      "votes": null
    },
    {
      "id": "708470",
      "postDate": "01/02/2020 10:23:05",
      "content": "<p>Both of them are pre-trained. </p>",
      "rawMarkdown": "Both of them are pre-trained.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 701159,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "12/23/2019 06:59:11",
      "content": "<p>I was running training using \"bert-large-uncased-whole-word-masking-finetuned-squad\". But my GCP instance died lol.  (batch_size = 4 and accumulation_steps = 8).</p>",
      "votes": null,
      "replies": [
        {
          "id": 701165,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "12/23/2019 07:07:20",
          "content": "<p>how were your initial results? I was planning on trying bert-large</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 701186,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/23/2019 07:46:40",
          "content": "<p>I was getting 0.20 ish public score for bert_uncased. And I think I found a bug on my code and I'm re-retraining it right now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 701586,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "12/23/2019 16:30:19",
          "content": "<p>I tried using bert large but it doesn't fit in P100 GPU memory, we need TPUs</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707771,
          "author_name": "mcggood",
          "author_url": "",
          "post_date": "01/01/2020 13:09:58",
          "content": "<p>I'm training a bert-large(float16)with batch_size32, accumulationsteps=4 with 44G GPU. cost 26hours for 1 epoch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708209,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/02/2020 04:34:12",
          "content": "<p><a href=\"/mcggood\">@mcggood</a> you have access to great resources</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708293,
          "author_name": "mcggood",
          "author_url": "",
          "post_date": "01/02/2020 06:39:00",
          "content": "<p>yeah, but I failed. my 1 epoch model is predicting same value in matrix. \nI see top teams using TPU with 1 hour per epoch.\nmaybe TPU is better than GPU</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708318,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/02/2020 07:19:56",
          "content": "<p>What? 1 hour per epoch with TPU...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 703502,
      "author_name": "likaiwu",
      "author_url": "",
      "post_date": "12/26/2019 09:07:08",
      "content": "<p>I wonder how to get \"bert-large-uncased-whole-word-masking-finetuned-squad\" ? I didn't find this model on original bert github page. Can you help me?</p>",
      "votes": null,
      "replies": [
        {
          "id": 703556,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/26/2019 09:42:12",
          "content": "<p>You can find it in <a href=\"https://huggingface.co/transformers/pretrained_models.html\">https://huggingface.co/transformers/pretrained_models.html</a> :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 703633,
          "author_name": "likaiwu",
          "author_url": "",
          "post_date": "12/26/2019 11:54:47",
          "content": "<p>Thanks for your reply!\nBTW, I think the bert-joint model was trained on bert-large. The <code>bert_config.json</code> file in <a href=\"https://www.kaggle.com/philculliton/bertjointbaseline\">bert-joint-baseline</a> dataset was described as following:\n<code>\nattention_probs_dropout_prob:0.1\nhidden_act:gelu\nhidden_dropout_prob:0.1\nhidden_size:1024\ninitializer_range:0.02\nintermediate_size:4096\nmax_position_embeddings:512\nnum_attention_heads:16\nnum_hidden_layers:24\ntype_vocab_size:2\nvocab_size:30522\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 703638,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/26/2019 12:09:30",
          "content": "<p>Wow. thanks for the info. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 707115,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "12/31/2019 09:09:57",
      "content": "<p>Update. I got 0.55 with large model when it was 0.52 for bert base.</p>",
      "votes": null,
      "replies": [
        {
          "id": 707206,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "12/31/2019 12:14:24",
          "content": "<p><a href=\"/higepon\">@higepon</a> that's great, my first version gives 0.58. I think there is a lot of room to improve. I think my postprocessing needs improvements.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708466,
          "author_name": "likaiwu",
          "author_url": "",
          "post_date": "01/02/2020 10:18:37",
          "content": "<p>Both the large and base models are pre-trained? Or you trained them by yourself?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708470,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/02/2020 10:23:05",
          "content": "<p>Both of them are pre-trained. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "701003": "Has anyone seen any improvements by using BERT Large instead of Base?\nFor example we might get some gain by using \"bert-large-uncased-whole-word-masking-finetuned-squad\".\n\nRelated topic\n[BERT Base vs/ BERT Large - Jigsaw Unintended Bias in Toxicity Classification](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/92977)",
    "701159": "I was running training using \"bert-large-uncased-whole-word-masking-finetuned-squad\". But my GCP instance died lol.  (batch_size = 4 and accumulation_steps = 8).",
    "701165": "how were your initial results? I was planning on trying bert-large",
    "701186": "I was getting 0.20 ish public score for bert_uncased. And I think I found a bug on my code and I'm re-retraining it right now.",
    "701586": "I tried using bert large but it doesn't fit in P100 GPU memory, we need TPUs",
    "703502": "I wonder how to get \"bert-large-uncased-whole-word-masking-finetuned-squad\" ? I didn't find this model on original bert github page. Can you help me?",
    "703556": "You can find it in https://huggingface.co/transformers/pretrained_models.html :)",
    "703633": "Thanks for your reply!\nBTW, I think the bert-joint model was trained on bert-large. The `bert_config.json` file in [bert-joint-baseline](https://www.kaggle.com/philculliton/bertjointbaseline) dataset was described as following:\n```\nattention_probs_dropout_prob:0.1\nhidden_act:gelu\nhidden_dropout_prob:0.1\nhidden_size:1024\ninitializer_range:0.02\nintermediate_size:4096\nmax_position_embeddings:512\nnum_attention_heads:16\nnum_hidden_layers:24\ntype_vocab_size:2\nvocab_size:30522\n```",
    "703638": "Wow. thanks for the info.",
    "707115": "Update. I got 0.55 with large model when it was 0.52 for bert base.",
    "707206": "higepon that's great, my first version gives 0.58. I think there is a lot of room to improve. I think my postprocessing needs improvements.",
    "707771": "I'm training a bert-large(float16)with batch_size32, accumulationsteps=4 with 44G GPU. cost 26hours for 1 epoch.",
    "708209": "mcggood you have access to great resources",
    "708293": "yeah, but I failed. my 1 epoch model is predicting same value in matrix. \nI see top teams using TPU with 1 hour per epoch.\nmaybe TPU is better than GPU",
    "708318": "What? 1 hour per epoch with TPU...",
    "708466": "Both the large and base models are pre-trained? Or you trained them by yourself?",
    "708470": "Both of them are pre-trained."
  },
  "source": "meta"
}