{
  "id": 118868,
  "title": "Is anyone able to do fine tuning Bert model using TF2.0 or Keras?",
  "url": "/competitions/tensorflow2-question-answering/discussion/118868",
  "author_name": "",
  "post_date": "2019-11-25T03:36:07.185360100Z",
  "votes": 8,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I have been facing OOM as describe in <a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook/comments#680085\">here</a>.</p>\n\n<p>I'm curious if anyone can successfully do fine tuning in TF2.0. I may switch to Torch if it's hard.</p>",
  "messages": [
    {
      "id": "680668",
      "postDate": "11/25/2019 03:36:07",
      "content": "<p>I have been facing OOM as describe in <a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook/comments#680085\">here</a>.</p>\n\n<p>I'm curious if anyone can successfully do fine tuning in TF2.0. I may switch to Torch if it's hard.</p>",
      "rawMarkdown": "I have been facing OOM as describe in [here](https://www.kaggle.com/prokaj/bert-joint-baseline-notebook/comments#680085).\n\nI'm curious if anyone can successfully do fine tuning in TF2.0. I may switch to Torch if it's hard.",
      "votes": null
    },
    {
      "id": "680734",
      "postDate": "11/25/2019 06:35:49",
      "content": "<p><a href=\"/higepon\">@higepon</a> are you using a P100 on colab? I think you can finetune it on a P100 in ~5 hours</p>",
      "rawMarkdown": "higepon are you using a P100 on colab? I think you can finetune it on a P100 in ~5 hours",
      "votes": null
    },
    {
      "id": "680899",
      "postDate": "11/25/2019 11:33:36",
      "content": "<p>Thank you <a href=\"/axel81\">@axel81</a>!\nI'm using P100 (16GB) on GCP. I was able to load pre-trained BERT, but I'm getting OOM (can't allocate tensor) when I call fit_generator. Are you using baseline kernel? I'm wondering what's different from your model.</p>",
      "rawMarkdown": "Thank you @axel81!\nI'm using P100 (16GB) on GCP. I was able to load pre-trained BERT, but I'm getting OOM (can't allocate tensor) when I call fit_generator. Are you using baseline kernel? I'm wondering what's different from your model.",
      "votes": null
    },
    {
      "id": "680901",
      "postDate": "11/25/2019 11:34:50",
      "content": "<p>One possibility I can think of is loss function. I might be using completely wrong loss function.</p>",
      "rawMarkdown": "One possibility I can think of is loss function. I might be using completely wrong loss function.",
      "votes": null
    },
    {
      "id": "680937",
      "postDate": "11/25/2019 12:42:33",
      "content": "<p><a href=\"/higepon\">@higepon</a> did you unfreeze the whole model?</p>",
      "rawMarkdown": "higepon did you unfreeze the whole model?",
      "votes": null
    },
    {
      "id": "680947",
      "postDate": "11/25/2019 12:53:57",
      "content": "<p>Yes. Thanks for the kind hint. Now I think I have something try out.</p>",
      "rawMarkdown": "Yes. Thanks for the kind hint. Now I think I have something try out.",
      "votes": null
    },
    {
      "id": "681014",
      "postDate": "11/25/2019 14:33:57",
      "content": "<p>Happy to help :)</p>",
      "rawMarkdown": "Happy to help :)",
      "votes": null
    },
    {
      "id": "681192",
      "postDate": "11/25/2019 20:02:37",
      "content": "<p><a href=\"/higepon\">@higepon</a> I am trying to find a kernel that uses keras and TF for loading pre-trained BERT model and take it from there. Does the kernel you have here do the job?</p>",
      "rawMarkdown": "higepon I am trying to find a kernel that uses keras and TF for loading pre-trained BERT model and take it from there. Does the kernel you have here do the job?",
      "votes": null
    },
    {
      "id": "681260",
      "postDate": "11/25/2019 23:07:13",
      "content": "<p>It didn't work, but I think I needed it anyway. I still see exact same OOM when allocating self attention.\n<code>\nOOM when allocating tensor with shape[4,16,512,512] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc [Op:BatchMatMulV2] name: bert-baseline/bert/encoder/layer_12/self_attention/einsum/MatMul/\n</code></p>\n\n<p>IIUC, you meant I should freeze Bert pre-trained model.\nI did it as <code>model.layers[3].trainable = False</code> and was able to see that I have far fewer trainable variable.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2F9a9446fd08f7501cc578dd58ad1d19fa%2F2019-11-26%208.02.57.png?generation=1574723010064520&amp;alt=media\" alt=\"\"></p>\n\n<p>Any thoughts?</p>",
      "rawMarkdown": "It didn't work, but I think I needed it anyway. I still see exact same OOM when allocating self attention.\n```\nOOM when allocating tensor with shape[4,16,512,512] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc [Op:BatchMatMulV2] name: bert-baseline/bert/encoder/layer_12/self_attention/einsum/MatMul/\n```\n\nIIUC, you meant I should freeze Bert pre-trained model.\nI did it as ```model.layers[3].trainable = False``` and was able to see that I have far fewer trainable variable.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2F9a9446fd08f7501cc578dd58ad1d19fa%2F2019-11-26%208.02.57.png?generation=1574723010064520&amp;alt=media)\n\nAny thoughts?",
      "votes": null
    },
    {
      "id": "681262",
      "postDate": "11/25/2019 23:08:17",
      "content": "<p>I'm using <a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook\">this awesome kernel</a> and it works perfectly for loading model weights and prediction.</p>",
      "rawMarkdown": "I'm using [this awesome kernel](https://www.kaggle.com/prokaj/bert-joint-baseline-notebook) and it works perfectly for loading model weights and prediction.",
      "votes": null
    },
    {
      "id": "681267",
      "postDate": "11/25/2019 23:13:23",
      "content": "<p>Strange thing is that I'm using <a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook\">this kernel</a> which works perfectly for prediction. The OOM is not happening when prediction. Maybe I'm doing something wrong in fit_generator.</p>\n\n<p>```\nlosses = {\n    \"tf_op_layer_start_squeeze\": \"categorical_crossentropy\",\n    \"tf_op_layer_end_squeeze\": \"categorical_crossentropy\",\n    \"ans_type\": \"categorical_crossentropy\",\n}\nloss_weights = {\"ans_type\": 1.0, \"tf_op_layer_start_squeeze\": 1.0, \"tf_op_layer_end_squeeze\": 1.0}</p>\n\n<p>model.compile(optimizer=op, loss=losses, loss_weights=loss_weights,\n                  metrics=['accuracy'])</p>\n\n<p>model.fit_generator(ds,verbose=2)\n```</p>",
      "rawMarkdown": "Strange thing is that I'm using [this kernel](https://www.kaggle.com/prokaj/bert-joint-baseline-notebook) which works perfectly for prediction. The OOM is not happening when prediction. Maybe I'm doing something wrong in fit_generator.\n\n\n```\nlosses = {\n    \"tf_op_layer_start_squeeze\": \"categorical_crossentropy\",\n    \"tf_op_layer_end_squeeze\": \"categorical_crossentropy\",\n    \"ans_type\": \"categorical_crossentropy\",\n}\nloss_weights = {\"ans_type\": 1.0, \"tf_op_layer_start_squeeze\": 1.0, \"tf_op_layer_end_squeeze\": 1.0}\n\nmodel.compile(optimizer=op, loss=losses, loss_weights=loss_weights,\n                  metrics=['accuracy'])\n\nmodel.fit_generator(ds,verbose=2)\n```",
      "votes": null
    },
    {
      "id": "681577",
      "postDate": "11/26/2019 09:21:45",
      "content": "<p><a href=\"/higepon\">@higepon</a> did you try reducing <code>batch_size</code>? For starters check if you can use <code>batch_size=1</code>. If it works that means you may have to change <code>max_sequence_length</code> to reduce input to your model. But I think <code>max_sequence_length=512</code> should work as I saw in paper or on github that you can finetune the model on a P100. I haven't started finetuning yet, I am still working on my pytorch version. I'll let you know if I figure out the issue.</p>",
      "rawMarkdown": "higepon did you try reducing `batch_size`? For starters check if you can use `batch_size=1`. If it works that means you may have to change `max_sequence_length` to reduce input to your model. But I think `max_sequence_length=512` should work as I saw in paper or on github that you can finetune the model on a P100. I haven't started finetuning yet, I am still working on my pytorch version. I'll let you know if I figure out the issue.",
      "votes": null
    },
    {
      "id": "681744",
      "postDate": "11/26/2019 13:46:44",
      "content": "<p><a href=\"/higepon\">@higepon</a> Thanks. The comments say this is a pretrained BERT model, I can't find where this pretrained model is imported?</p>",
      "rawMarkdown": "higepon Thanks. The comments say this is a pretrained BERT model, I can't find where this pretrained model is imported?",
      "votes": null
    },
    {
      "id": "681892",
      "postDate": "11/26/2019 16:45:14",
      "content": "<p><a href=\"/higepon\">@higepon</a> I tried to run the kernel,  and was at below point for a couple of hours before stopping the kernel run. Is it going to take a long time or am I going to run out of memory eventually?\nresult=model.predict_generator(ds,verbose=1 if not on_kaggle_server else 0)</p>",
      "rawMarkdown": "higepon I tried to run the kernel,  and was at below point for a couple of hours before stopping the kernel run. Is it going to take a long time or am I going to run out of memory eventually?\nresult=model.predict_generator(ds,verbose=1 if not on_kaggle_server else 0)",
      "votes": null
    },
    {
      "id": "682310",
      "postDate": "11/27/2019 08:31:58",
      "content": "<p>Yeah I tried batch_size=1 and got another error. I think it's related to my loss function and I'm looking into it. Thank you!\nI should probably switch to PyTorch..</p>",
      "rawMarkdown": "Yeah I tried batch_size=1 and got another error. I think it's related to my loss function and I'm looking into it. Thank you!\nI should probably switch to PyTorch..",
      "votes": null
    },
    {
      "id": "682311",
      "postDate": "11/27/2019 08:33:15",
      "content": "<p>I took only about 15 min for me. Maybe you want to debug a bit.</p>",
      "rawMarkdown": "I took only about 15 min for me. Maybe you want to debug a bit.",
      "votes": null
    },
    {
      "id": "683338",
      "postDate": "11/28/2019 10:14:34",
      "content": "<p>Update: I was able to run fit_generator for batch_size=1 and 2. But batch_size=4 throws OOM.</p>\n\n<p>I may switch to PyTorch.</p>",
      "rawMarkdown": "Update: I was able to run fit_generator for batch_size=1 and 2. But batch_size=4 throws OOM.\n\nI may switch to PyTorch.",
      "votes": null
    },
    {
      "id": "689718",
      "postDate": "12/07/2019 10:35:09",
      "content": "<p><a href=\"/higepon\">@higepon</a> were you facing this issue with the <code>model.fit_generator</code></p>\n\n<h2>```</h2>\n\n<p>KeyError                                  Traceback (most recent call last)\n/usr/local/lib/python3.6/dist-packages/tensorflow_core/python/keras/engine/training_utils.py in standardize_input_data(data, names, shapes, check_batch_axis, exception_prefix)\n    498           if data[x].<strong>class</strong>._<em>name</em>_ == 'DataFrame' else data[x]\n--&gt; 499           for x in names\n    500       ]</p>\n\n<p>7 frames\nKeyError: 'tf_op_layer_start_squeeze'</p>\n\n<p>During handling of the above exception, another exception occurred:</p>\n\n<p>ValueError                                Traceback (most recent call last)\n/usr/local/lib/python3.6/dist-packages/tensorflow_core/python/keras/engine/training_utils.py in standardize_input_data(data, names, shapes, check_batch_axis, exception_prefix)\n    501     except KeyError as e:\n    502       raise ValueError('No data provided for \"' + e.args[0] + '\". Need data '\n--&gt; 503                        'for each key in: ' + str(names))\n    504   elif isinstance(data, (list, tuple)):\n    505     if isinstance(data[0], (list, tuple)):</p>\n\n<p>ValueError: No data provided for \"tf_op_layer_start_squeeze\". Need data for each key in: ['tf_op_layer_start_squeeze', 'tf_op_layer_end_squeeze', 'ans_type']\n```</p>",
      "rawMarkdown": "higepon were you facing this issue with the `model.fit_generator`\n\n```\n---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\n/usr/local/lib/python3.6/dist-packages/tensorflow_core/python/keras/engine/training_utils.py in standardize_input_data(data, names, shapes, check_batch_axis, exception_prefix)\n    498           if data[x].__class__.__name__ == 'DataFrame' else data[x]\n--&gt; 499           for x in names\n    500       ]\n\n7 frames\nKeyError: 'tf_op_layer_start_squeeze'\n\nDuring handling of the above exception, another exception occurred:\n\nValueError                                Traceback (most recent call last)\n/usr/local/lib/python3.6/dist-packages/tensorflow_core/python/keras/engine/training_utils.py in standardize_input_data(data, names, shapes, check_batch_axis, exception_prefix)\n    501     except KeyError as e:\n    502       raise ValueError('No data provided for \"' + e.args[0] + '\". Need data '\n--&gt; 503                        'for each key in: ' + str(names))\n    504   elif isinstance(data, (list, tuple)):\n    505     if isinstance(data[0], (list, tuple)):\n\nValueError: No data provided for \"tf_op_layer_start_squeeze\". Need data for each key in: ['tf_op_layer_start_squeeze', 'tf_op_layer_end_squeeze', 'ans_type']\n```",
      "votes": null
    },
    {
      "id": "690360",
      "postDate": "12/08/2019 12:48:53",
      "content": "<p>yeah.\nIf I remember correctly. I did something like the following to resolve.</p>\n\n<p><code>\n    y['tf_op_layer_start_squeeze'] = example['start_positions']\n    y['tf_op_layer_end_squeeze'] = example['end_positions']\n    y['ans_type'] = example['answer_types']\n    y['unique_id'] = example['unique_id']\n</code></p>",
      "rawMarkdown": "yeah.\nIf I remember correctly. I did something like the following to resolve.\n\n```\n    y['tf_op_layer_start_squeeze'] = example['start_positions']\n    y['tf_op_layer_end_squeeze'] = example['end_positions']\n    y['ans_type'] = example['answer_types']\n    y['unique_id'] = example['unique_id']\n```",
      "votes": null
    },
    {
      "id": "698025",
      "postDate": "12/18/2019 17:35:24",
      "content": "<p><a href=\"/higepon\">@higepon</a> Were you able to finally fine-tune the model on TF2.0 or did you switch to PyTorch?</p>",
      "rawMarkdown": "higepon Were you able to finally fine-tune the model on TF2.0 or did you switch to PyTorch?",
      "votes": null
    },
    {
      "id": "698206",
      "postDate": "12/18/2019 23:37:42",
      "content": "<p>Hey! I switched to PyTorch.</p>",
      "rawMarkdown": "Hey! I switched to PyTorch.",
      "votes": null
    },
    {
      "id": "802595",
      "postDate": "04/09/2020 16:35:26",
      "content": "<p><a href=\"/higepon\">@higepon</a> you may find this kernel useful.\n<a href=\"https://www.kaggle.com/ashoksrinivas/bert-implementation-in-keras\">https://www.kaggle.com/ashoksrinivas/bert-implementation-in-keras</a></p>",
      "rawMarkdown": "higepon you may find this kernel useful.\nhttps://www.kaggle.com/ashoksrinivas/bert-implementation-in-keras",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 680734,
      "author_name": "axel81",
      "author_url": "",
      "post_date": "11/25/2019 06:35:49",
      "content": "<p><a href=\"/higepon\">@higepon</a> are you using a P100 on colab? I think you can finetune it on a P100 in ~5 hours</p>",
      "votes": null,
      "replies": [
        {
          "id": 680899,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "11/25/2019 11:33:36",
          "content": "<p>Thank you <a href=\"/axel81\">@axel81</a>!\nI'm using P100 (16GB) on GCP. I was able to load pre-trained BERT, but I'm getting OOM (can't allocate tensor) when I call fit_generator. Are you using baseline kernel? I'm wondering what's different from your model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 680901,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "11/25/2019 11:34:50",
          "content": "<p>One possibility I can think of is loss function. I might be using completely wrong loss function.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 680937,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "11/25/2019 12:42:33",
          "content": "<p><a href=\"/higepon\">@higepon</a> did you unfreeze the whole model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 680947,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "11/25/2019 12:53:57",
          "content": "<p>Yes. Thanks for the kind hint. Now I think I have something try out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 681014,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "11/25/2019 14:33:57",
          "content": "<p>Happy to help :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 681260,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "11/25/2019 23:07:13",
          "content": "<p>It didn't work, but I think I needed it anyway. I still see exact same OOM when allocating self attention.\n<code>\nOOM when allocating tensor with shape[4,16,512,512] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc [Op:BatchMatMulV2] name: bert-baseline/bert/encoder/layer_12/self_attention/einsum/MatMul/\n</code></p>\n\n<p>IIUC, you meant I should freeze Bert pre-trained model.\nI did it as <code>model.layers[3].trainable = False</code> and was able to see that I have far fewer trainable variable.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2F9a9446fd08f7501cc578dd58ad1d19fa%2F2019-11-26%208.02.57.png?generation=1574723010064520&amp;alt=media\" alt=\"\"></p>\n\n<p>Any thoughts?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 681267,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "11/25/2019 23:13:23",
          "content": "<p>Strange thing is that I'm using <a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook\">this kernel</a> which works perfectly for prediction. The OOM is not happening when prediction. Maybe I'm doing something wrong in fit_generator.</p>\n\n<p>```\nlosses = {\n    \"tf_op_layer_start_squeeze\": \"categorical_crossentropy\",\n    \"tf_op_layer_end_squeeze\": \"categorical_crossentropy\",\n    \"ans_type\": \"categorical_crossentropy\",\n}\nloss_weights = {\"ans_type\": 1.0, \"tf_op_layer_start_squeeze\": 1.0, \"tf_op_layer_end_squeeze\": 1.0}</p>\n\n<p>model.compile(optimizer=op, loss=losses, loss_weights=loss_weights,\n                  metrics=['accuracy'])</p>\n\n<p>model.fit_generator(ds,verbose=2)\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 681577,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "11/26/2019 09:21:45",
          "content": "<p><a href=\"/higepon\">@higepon</a> did you try reducing <code>batch_size</code>? For starters check if you can use <code>batch_size=1</code>. If it works that means you may have to change <code>max_sequence_length</code> to reduce input to your model. But I think <code>max_sequence_length=512</code> should work as I saw in paper or on github that you can finetune the model on a P100. I haven't started finetuning yet, I am still working on my pytorch version. I'll let you know if I figure out the issue.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 682310,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "11/27/2019 08:31:58",
          "content": "<p>Yeah I tried batch_size=1 and got another error. I think it's related to my loss function and I'm looking into it. Thank you!\nI should probably switch to PyTorch..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 683338,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "11/28/2019 10:14:34",
          "content": "<p>Update: I was able to run fit_generator for batch_size=1 and 2. But batch_size=4 throws OOM.</p>\n\n<p>I may switch to PyTorch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 689718,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "12/07/2019 10:35:09",
          "content": "<p><a href=\"/higepon\">@higepon</a> were you facing this issue with the <code>model.fit_generator</code></p>\n\n<h2>```</h2>\n\n<p>KeyError                                  Traceback (most recent call last)\n/usr/local/lib/python3.6/dist-packages/tensorflow_core/python/keras/engine/training_utils.py in standardize_input_data(data, names, shapes, check_batch_axis, exception_prefix)\n    498           if data[x].<strong>class</strong>._<em>name</em>_ == 'DataFrame' else data[x]\n--&gt; 499           for x in names\n    500       ]</p>\n\n<p>7 frames\nKeyError: 'tf_op_layer_start_squeeze'</p>\n\n<p>During handling of the above exception, another exception occurred:</p>\n\n<p>ValueError                                Traceback (most recent call last)\n/usr/local/lib/python3.6/dist-packages/tensorflow_core/python/keras/engine/training_utils.py in standardize_input_data(data, names, shapes, check_batch_axis, exception_prefix)\n    501     except KeyError as e:\n    502       raise ValueError('No data provided for \"' + e.args[0] + '\". Need data '\n--&gt; 503                        'for each key in: ' + str(names))\n    504   elif isinstance(data, (list, tuple)):\n    505     if isinstance(data[0], (list, tuple)):</p>\n\n<p>ValueError: No data provided for \"tf_op_layer_start_squeeze\". Need data for each key in: ['tf_op_layer_start_squeeze', 'tf_op_layer_end_squeeze', 'ans_type']\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 690360,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/08/2019 12:48:53",
          "content": "<p>yeah.\nIf I remember correctly. I did something like the following to resolve.</p>\n\n<p><code>\n    y['tf_op_layer_start_squeeze'] = example['start_positions']\n    y['tf_op_layer_end_squeeze'] = example['end_positions']\n    y['ans_type'] = example['answer_types']\n    y['unique_id'] = example['unique_id']\n</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 681192,
      "author_name": "kevinehsani",
      "author_url": "",
      "post_date": "11/25/2019 20:02:37",
      "content": "<p><a href=\"/higepon\">@higepon</a> I am trying to find a kernel that uses keras and TF for loading pre-trained BERT model and take it from there. Does the kernel you have here do the job?</p>",
      "votes": null,
      "replies": [
        {
          "id": 681262,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "11/25/2019 23:08:17",
          "content": "<p>I'm using <a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook\">this awesome kernel</a> and it works perfectly for loading model weights and prediction.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 681744,
          "author_name": "kevinehsani",
          "author_url": "",
          "post_date": "11/26/2019 13:46:44",
          "content": "<p><a href=\"/higepon\">@higepon</a> Thanks. The comments say this is a pretrained BERT model, I can't find where this pretrained model is imported?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 681892,
          "author_name": "kevinehsani",
          "author_url": "",
          "post_date": "11/26/2019 16:45:14",
          "content": "<p><a href=\"/higepon\">@higepon</a> I tried to run the kernel,  and was at below point for a couple of hours before stopping the kernel run. Is it going to take a long time or am I going to run out of memory eventually?\nresult=model.predict_generator(ds,verbose=1 if not on_kaggle_server else 0)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 682311,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "11/27/2019 08:33:15",
          "content": "<p>I took only about 15 min for me. Maybe you want to debug a bit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 698025,
      "author_name": "rohitagarwal",
      "author_url": "",
      "post_date": "12/18/2019 17:35:24",
      "content": "<p><a href=\"/higepon\">@higepon</a> Were you able to finally fine-tune the model on TF2.0 or did you switch to PyTorch?</p>",
      "votes": null,
      "replies": [
        {
          "id": 698206,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/18/2019 23:37:42",
          "content": "<p>Hey! I switched to PyTorch.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 802595,
      "author_name": "ashoksrinivas",
      "author_url": "",
      "post_date": "04/09/2020 16:35:26",
      "content": "<p><a href=\"/higepon\">@higepon</a> you may find this kernel useful.\n<a href=\"https://www.kaggle.com/ashoksrinivas/bert-implementation-in-keras\">https://www.kaggle.com/ashoksrinivas/bert-implementation-in-keras</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "680668": "I have been facing OOM as describe in [here](https://www.kaggle.com/prokaj/bert-joint-baseline-notebook/comments#680085).\n\nI'm curious if anyone can successfully do fine tuning in TF2.0. I may switch to Torch if it's hard.",
    "680734": "higepon are you using a P100 on colab? I think you can finetune it on a P100 in ~5 hours",
    "680899": "Thank you @axel81!\nI'm using P100 (16GB) on GCP. I was able to load pre-trained BERT, but I'm getting OOM (can't allocate tensor) when I call fit_generator. Are you using baseline kernel? I'm wondering what's different from your model.",
    "680901": "One possibility I can think of is loss function. I might be using completely wrong loss function.",
    "680937": "higepon did you unfreeze the whole model?",
    "680947": "Yes. Thanks for the kind hint. Now I think I have something try out.",
    "681014": "Happy to help :)",
    "681192": "higepon I am trying to find a kernel that uses keras and TF for loading pre-trained BERT model and take it from there. Does the kernel you have here do the job?",
    "681260": "It didn't work, but I think I needed it anyway. I still see exact same OOM when allocating self attention.\n```\nOOM when allocating tensor with shape[4,16,512,512] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc [Op:BatchMatMulV2] name: bert-baseline/bert/encoder/layer_12/self_attention/einsum/MatMul/\n```\n\nIIUC, you meant I should freeze Bert pre-trained model.\nI did it as ```model.layers[3].trainable = False``` and was able to see that I have far fewer trainable variable.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2F9a9446fd08f7501cc578dd58ad1d19fa%2F2019-11-26%208.02.57.png?generation=1574723010064520&amp;alt=media)\n\nAny thoughts?",
    "681262": "I'm using [this awesome kernel](https://www.kaggle.com/prokaj/bert-joint-baseline-notebook) and it works perfectly for loading model weights and prediction.",
    "681267": "Strange thing is that I'm using [this kernel](https://www.kaggle.com/prokaj/bert-joint-baseline-notebook) which works perfectly for prediction. The OOM is not happening when prediction. Maybe I'm doing something wrong in fit_generator.\n\n\n```\nlosses = {\n    \"tf_op_layer_start_squeeze\": \"categorical_crossentropy\",\n    \"tf_op_layer_end_squeeze\": \"categorical_crossentropy\",\n    \"ans_type\": \"categorical_crossentropy\",\n}\nloss_weights = {\"ans_type\": 1.0, \"tf_op_layer_start_squeeze\": 1.0, \"tf_op_layer_end_squeeze\": 1.0}\n\nmodel.compile(optimizer=op, loss=losses, loss_weights=loss_weights,\n                  metrics=['accuracy'])\n\nmodel.fit_generator(ds,verbose=2)\n```",
    "681577": "higepon did you try reducing `batch_size`? For starters check if you can use `batch_size=1`. If it works that means you may have to change `max_sequence_length` to reduce input to your model. But I think `max_sequence_length=512` should work as I saw in paper or on github that you can finetune the model on a P100. I haven't started finetuning yet, I am still working on my pytorch version. I'll let you know if I figure out the issue.",
    "681744": "higepon Thanks. The comments say this is a pretrained BERT model, I can't find where this pretrained model is imported?",
    "681892": "higepon I tried to run the kernel,  and was at below point for a couple of hours before stopping the kernel run. Is it going to take a long time or am I going to run out of memory eventually?\nresult=model.predict_generator(ds,verbose=1 if not on_kaggle_server else 0)",
    "682310": "Yeah I tried batch_size=1 and got another error. I think it's related to my loss function and I'm looking into it. Thank you!\nI should probably switch to PyTorch..",
    "682311": "I took only about 15 min for me. Maybe you want to debug a bit.",
    "683338": "Update: I was able to run fit_generator for batch_size=1 and 2. But batch_size=4 throws OOM.\n\nI may switch to PyTorch.",
    "689718": "higepon were you facing this issue with the `model.fit_generator`\n\n```\n---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\n/usr/local/lib/python3.6/dist-packages/tensorflow_core/python/keras/engine/training_utils.py in standardize_input_data(data, names, shapes, check_batch_axis, exception_prefix)\n    498           if data[x].__class__.__name__ == 'DataFrame' else data[x]\n--&gt; 499           for x in names\n    500       ]\n\n7 frames\nKeyError: 'tf_op_layer_start_squeeze'\n\nDuring handling of the above exception, another exception occurred:\n\nValueError                                Traceback (most recent call last)\n/usr/local/lib/python3.6/dist-packages/tensorflow_core/python/keras/engine/training_utils.py in standardize_input_data(data, names, shapes, check_batch_axis, exception_prefix)\n    501     except KeyError as e:\n    502       raise ValueError('No data provided for \"' + e.args[0] + '\". Need data '\n--&gt; 503                        'for each key in: ' + str(names))\n    504   elif isinstance(data, (list, tuple)):\n    505     if isinstance(data[0], (list, tuple)):\n\nValueError: No data provided for \"tf_op_layer_start_squeeze\". Need data for each key in: ['tf_op_layer_start_squeeze', 'tf_op_layer_end_squeeze', 'ans_type']\n```",
    "690360": "yeah.\nIf I remember correctly. I did something like the following to resolve.\n\n```\n    y['tf_op_layer_start_squeeze'] = example['start_positions']\n    y['tf_op_layer_end_squeeze'] = example['end_positions']\n    y['ans_type'] = example['answer_types']\n    y['unique_id'] = example['unique_id']\n```",
    "698025": "higepon Were you able to finally fine-tune the model on TF2.0 or did you switch to PyTorch?",
    "698206": "Hey! I switched to PyTorch.",
    "802595": "higepon you may find this kernel useful.\nhttps://www.kaggle.com/ashoksrinivas/bert-implementation-in-keras"
  },
  "source": "meta"
}