{
  "id": 122680,
  "title": "Customizing BERT",
  "url": "/competitions/tensorflow2-question-answering/discussion/122680",
  "author_name": "",
  "post_date": "2019-12-22T05:31:43.681935100Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In case you are new to BERT, I wanted to share my understandings about how to use it on the cloud. The main steps are explained here: <a href=\"https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md\">https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md</a></p>\n\n<p>What I want to explain is how to customize BERT for your needs. This is basically done on run-classifier.py. There is also the NQ version of this related to this competition, called run-nq, but that one is not flexible enough. You will need run_classifier.</p>\n\n<p>Once you run run-classifier.py, it will execute the main function. Then:</p>\n\n<ol>\n<li><p>BERT will be configured using the bert-config-file. This is tower you will be building on.</p></li>\n<li><p>Then comes model-fn function. That is where you specify the model you will be attaching at the end of BERT.</p></li>\n<li><p>Then comes estimator function. That one is built through another python file, but uses the config and model files from prev steps.</p></li>\n<li><p>Next is get-train-examples. In run-nq, there is no such step because all files come preprocessed in tfrecord format. \na) Reads the file and iterates through the lines.\nb) Returns examples in unicode.</p></li>\n<li><p>Next comes file-based-covert-examples-to-features. This is where it gets convoluted:\na) First it gets the examples from get_train_examples and iterates through them.\nb) Runs the convert_single_example function. THIS is the KEY to changing the labels and adding additional inputs. Here you specify the labels to your wish. In this competition, there will be tokens-a and tokens-b.\nc) Writes them into a file in tfrecord format.</p></li>\n<li><p>Next is file-based-input-fn-builder. This one basically takes the tfrecord files and creates an input-fn, which the estimator uses to train on. Once the estimator is trained, you can use it to get predictions. Estimator is basically your NN at this point. Happy Kaggling! </p></li>\n</ol>",
  "messages": [
    {
      "id": "700498",
      "postDate": "12/22/2019 05:31:43",
      "content": "<p>In case you are new to BERT, I wanted to share my understandings about how to use it on the cloud. The main steps are explained here: <a href=\"https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md\">https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md</a></p>\n\n<p>What I want to explain is how to customize BERT for your needs. This is basically done on run-classifier.py. There is also the NQ version of this related to this competition, called run-nq, but that one is not flexible enough. You will need run_classifier.</p>\n\n<p>Once you run run-classifier.py, it will execute the main function. Then:</p>\n\n<ol>\n<li><p>BERT will be configured using the bert-config-file. This is tower you will be building on.</p></li>\n<li><p>Then comes model-fn function. That is where you specify the model you will be attaching at the end of BERT.</p></li>\n<li><p>Then comes estimator function. That one is built through another python file, but uses the config and model files from prev steps.</p></li>\n<li><p>Next is get-train-examples. In run-nq, there is no such step because all files come preprocessed in tfrecord format. \na) Reads the file and iterates through the lines.\nb) Returns examples in unicode.</p></li>\n<li><p>Next comes file-based-covert-examples-to-features. This is where it gets convoluted:\na) First it gets the examples from get_train_examples and iterates through them.\nb) Runs the convert_single_example function. THIS is the KEY to changing the labels and adding additional inputs. Here you specify the labels to your wish. In this competition, there will be tokens-a and tokens-b.\nc) Writes them into a file in tfrecord format.</p></li>\n<li><p>Next is file-based-input-fn-builder. This one basically takes the tfrecord files and creates an input-fn, which the estimator uses to train on. Once the estimator is trained, you can use it to get predictions. Estimator is basically your NN at this point. Happy Kaggling! </p></li>\n</ol>",
      "rawMarkdown": "In case you are new to BERT, I wanted to share my understandings about how to use it on the cloud. The main steps are explained here: https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md\n\nWhat I want to explain is how to customize BERT for your needs. This is basically done on run-classifier.py. There is also the NQ version of this related to this competition, called run-nq, but that one is not flexible enough. You will need run_classifier.\n\nOnce you run run-classifier.py, it will execute the main function. Then:\n\n1. BERT will be configured using the bert-config-file. This is tower you will be building on.\n\n2. Then comes model-fn function. That is where you specify the model you will be attaching at the end of BERT.\n\n3. Then comes estimator function. That one is built through another python file, but uses the config and model files from prev steps.\n\n4. Next is get-train-examples. In run-nq, there is no such step because all files come preprocessed in tfrecord format. \na) Reads the file and iterates through the lines.\nb) Returns examples in unicode.\n\n5. Next comes file-based-covert-examples-to-features. This is where it gets convoluted:\na) First it gets the examples from get_train_examples and iterates through them.\nb) Runs the convert_single_example function. THIS is the KEY to changing the labels and adding additional inputs. Here you specify the labels to your wish. In this competition, there will be tokens-a and tokens-b.\nc) Writes them into a file in tfrecord format.\n\n6. Next is file-based-input-fn-builder. This one basically takes the tfrecord files and creates an input-fn, which the estimator uses to train on. Once the estimator is trained, you can use it to get predictions. Estimator is basically your NN at this point. Happy Kaggling!",
      "votes": null
    },
    {
      "id": "700835",
      "postDate": "12/22/2019 17:26:59",
      "content": "<p>Thank you for taking the time to put this together. Bookmarked it to look at it again once I learn a bit (well, quite a lot 😃) more.</p>",
      "rawMarkdown": "Thank you for taking the time to put this together. Bookmarked it to look at it again once I learn a bit (well, quite a lot 😃) more.",
      "votes": null
    },
    {
      "id": "700918",
      "postDate": "12/22/2019 20:16:16",
      "content": "<p>You are welcome 🙂</p>",
      "rawMarkdown": "You are welcome 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 700835,
      "author_name": "eghove",
      "author_url": "",
      "post_date": "12/22/2019 17:26:59",
      "content": "<p>Thank you for taking the time to put this together. Bookmarked it to look at it again once I learn a bit (well, quite a lot 😃) more.</p>",
      "votes": null,
      "replies": [
        {
          "id": 700918,
          "author_name": "isikkuntay",
          "author_url": "",
          "post_date": "12/22/2019 20:16:16",
          "content": "<p>You are welcome 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "700498": "In case you are new to BERT, I wanted to share my understandings about how to use it on the cloud. The main steps are explained here: https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md\n\nWhat I want to explain is how to customize BERT for your needs. This is basically done on run-classifier.py. There is also the NQ version of this related to this competition, called run-nq, but that one is not flexible enough. You will need run_classifier.\n\nOnce you run run-classifier.py, it will execute the main function. Then:\n\n1. BERT will be configured using the bert-config-file. This is tower you will be building on.\n\n2. Then comes model-fn function. That is where you specify the model you will be attaching at the end of BERT.\n\n3. Then comes estimator function. That one is built through another python file, but uses the config and model files from prev steps.\n\n4. Next is get-train-examples. In run-nq, there is no such step because all files come preprocessed in tfrecord format. \na) Reads the file and iterates through the lines.\nb) Returns examples in unicode.\n\n5. Next comes file-based-covert-examples-to-features. This is where it gets convoluted:\na) First it gets the examples from get_train_examples and iterates through them.\nb) Runs the convert_single_example function. THIS is the KEY to changing the labels and adding additional inputs. Here you specify the labels to your wish. In this competition, there will be tokens-a and tokens-b.\nc) Writes them into a file in tfrecord format.\n\n6. Next is file-based-input-fn-builder. This one basically takes the tfrecord files and creates an input-fn, which the estimator uses to train on. Once the estimator is trained, you can use it to get predictions. Estimator is basically your NN at this point. Happy Kaggling!",
    "700835": "Thank you for taking the time to put this together. Bookmarked it to look at it again once I learn a bit (well, quite a lot 😃) more.",
    "700918": "You are welcome 🙂"
  },
  "source": "meta"
}