{
  "id": 121972,
  "title": "Fine tuning",
  "url": "/competitions/tensorflow2-question-answering/discussion/121972",
  "author_name": "AlexanderCh",
  "post_date": "2019-12-16T21:16:02.251000",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Has anyone tried BERT fine tuning?\nI have tried to go through these <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#data-preparation\">steps</a>\nbut there is problem with Data Preparation step\n```\npython -m language.question_answering.bert_joint.prepare_nq_data \\\n  --logtostderr \\\n  --input_jsonl simplified-nq-train.jsonl \\\n  --output_tfrecord ~/output_dir/nq-train.tfrecords-00000-of-00001 \\\n  --max_seq_length=512 \\\n  --include_unknowns=0.02 \\\n  --vocab_file=bert-joint-baseline/vocab-nq.txt</p>\n\n<p>```</p>\n\n<p>Error:\n<code>\n    for example in get_examples(FLAGS.input_jsonl):\n  File \"/data1/achaptykov/model/googlebert/language/question_answering/bert_joint/prepare_nq_data.py\", line 59, in get_examples\n    for line in input_file:\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 374, in readline\n    return self._buffer.readline(size)\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/_compression.py\", line 68, in readinto\n    data = self.read(len(byte_view))\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 463, in read\n    if not self._read_gzip_header():\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 406, in _read_gzip_header\n    magic = self._fp.read(2)\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 91, in read\n    self.file.read(size-self._length+read)\nTypeError: can't concat str to bytes\n</code>\nShould I put json from competition data? Or shall I modified it.</p>",
  "messages": [
    {
      "id": 696626,
      "postDate": "2019-12-16T21:16:02.250Z",
      "content": "<p>Has anyone tried BERT fine tuning?\nI have tried to go through these <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#data-preparation\">steps</a>\nbut there is problem with Data Preparation step\n```\npython -m language.question_answering.bert_joint.prepare_nq_data \\\n  --logtostderr \\\n  --input_jsonl simplified-nq-train.jsonl \\\n  --output_tfrecord ~/output_dir/nq-train.tfrecords-00000-of-00001 \\\n  --max_seq_length=512 \\\n  --include_unknowns=0.02 \\\n  --vocab_file=bert-joint-baseline/vocab-nq.txt</p>\n\n<p>```</p>\n\n<p>Error:\n<code>\n    for example in get_examples(FLAGS.input_jsonl):\n  File \"/data1/achaptykov/model/googlebert/language/question_answering/bert_joint/prepare_nq_data.py\", line 59, in get_examples\n    for line in input_file:\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 374, in readline\n    return self._buffer.readline(size)\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/_compression.py\", line 68, in readinto\n    data = self.read(len(byte_view))\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 463, in read\n    if not self._read_gzip_header():\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 406, in _read_gzip_header\n    magic = self._fp.read(2)\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 91, in read\n    self.file.read(size-self._length+read)\nTypeError: can't concat str to bytes\n</code>\nShould I put json from competition data? Or shall I modified it.</p>",
      "rawMarkdown": "Has anyone tried BERT fine tuning?\nI have tried to go through these [steps](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#data-preparation)\nbut there is problem with Data Preparation step\n```\npython -m language.question_answering.bert_joint.prepare_nq_data \\\n  --logtostderr \\\n  --input_jsonl simplified-nq-train.jsonl \\\n  --output_tfrecord ~/output_dir/nq-train.tfrecords-00000-of-00001 \\\n  --max_seq_length=512 \\\n  --include_unknowns=0.02 \\\n  --vocab_file=bert-joint-baseline/vocab-nq.txt\n\n```\n\nError:\n```\n    for example in get_examples(FLAGS.input_jsonl):\n  File \"/data1/achaptykov/model/googlebert/language/question_answering/bert_joint/prepare_nq_data.py\", line 59, in get_examples\n    for line in input_file:\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 374, in readline\n    return self._buffer.readline(size)\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/_compression.py\", line 68, in readinto\n    data = self.read(len(byte_view))\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 463, in read\n    if not self._read_gzip_header():\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 406, in _read_gzip_header\n    magic = self._fp.read(2)\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 91, in read\n    self.file.read(size-self._length+read)\nTypeError: can't concat str to bytes\n```\nShould I put json from competition data? Or shall I modified it.",
      "votes": 3
    },
    {
      "id": 697426,
      "postDate": "2019-12-17T22:34:06.137Z",
      "content": "<p>I had the same problem. Turned out to be a problem of Python 3 not supporting implicit conversion from str to bytes according to this <a href=\"https://stackoverflow.com/a/21916982/2070636\">SO answer</a>.</p>\n\n<p>Here is the workaround I used to fix it for now. I changed <a href=\"https://github.com/google-research/language/blob/623e016a76eaeb054d09a4173db1026685bc5b08/language/question_answering/bert_joint/prepare_nq_data.py#L58\">this line</a> in <a href=\"https://github.com/google-research/language/blob/master/language/question_answering/bert_joint/prepare_nq_data.py\"><code>prepare_nq_data</code></a>:</p>\n\n<p>```\n with gzip.GzipFile(fileobj=tf.gfile.Open(input_path)) as input_file:</p>\n\n<p>```\nTo</p>\n\n<p><code>\n with tf.io.gfile.GFile(input_path, mode='rb') as input_file:\n</code></p>\n\n<p>Not sure if that's the correct/best solution, but it seems to work so far.</p>",
      "rawMarkdown": "I had the same problem. Turned out to be a problem of Python 3 not supporting implicit conversion from str to bytes according to this [SO answer](https://stackoverflow.com/a/21916982/2070636).\n\n\nHere is the workaround I used to fix it for now. I changed [this line](https://github.com/google-research/language/blob/623e016a76eaeb054d09a4173db1026685bc5b08/language/question_answering/bert_joint/prepare_nq_data.py#L58) in [`prepare_nq_data`](https://github.com/google-research/language/blob/master/language/question_answering/bert_joint/prepare_nq_data.py):\n\n```\n with gzip.GzipFile(fileobj=tf.gfile.Open(input_path)) as input_file:\n\n```\nTo\n\n```\n with tf.io.gfile.GFile(input_path, mode='rb') as input_file:\n```\n\nNot sure if that's the correct/best solution, but it seems to work so far.",
      "votes": 1,
      "replies": [
        {
          "id": 697662,
          "postDate": "2019-12-18T08:44:11.047Z",
          "content": "<p>with Python2 no problems</p>",
          "rawMarkdown": "with Python2 no problems",
          "votes": 1
        }
      ]
    },
    {
      "id": 696844,
      "postDate": "2019-12-17T06:10:42.103Z",
      "content": "<p>This is what I read:</p>\n\n<p>You should then download our model and preprocessed training set with:</p>\n\n<p>gsutil cp -R gs://bert-nq/bert-joint-baseline .</p>",
      "rawMarkdown": "This is what I read:\n\nYou should then download our model and preprocessed training set with:\n\ngsutil cp -R gs://bert-nq/bert-joint-baseline .",
      "votes": 1,
      "replies": [
        {
          "id": 696899,
          "postDate": "2019-12-17T08:01:31.897Z",
          "content": "<p>Yes, I did it. But there is no any json with training data here:</p>\n\n<p>```\nbert-joint-baseline/nq-train.tfrecords-00000-of-00001\nbert-joint-baseline/bert_config.json\nbert-joint-baseline/vocab-nq.txt\nbert-joint-baseline/bert_joint.ckpt.data-00000-of-00001\nbert-joint-baseline/bert_joint.ckpt.index</p>\n\n<p>```</p>",
          "rawMarkdown": "Yes, I did it. But there is no any json with training data here:\n\n```\nbert-joint-baseline/nq-train.tfrecords-00000-of-00001\nbert-joint-baseline/bert_config.json\nbert-joint-baseline/vocab-nq.txt\nbert-joint-baseline/bert_joint.ckpt.data-00000-of-00001\nbert-joint-baseline/bert_joint.ckpt.index\n\n```",
          "votes": 1
        },
        {
          "id": 697385,
          "postDate": "2019-12-17T20:31:35.230Z",
          "content": "<p>These are the same files already included in the <a href=\"https://www.kaggle.com/philculliton/using-tensorflow-2-0-w-bert-on-nq\">starter notebook</a> under <code>input/bert-joint-baseline</code>. </p>\n\n<p>Here is how I see it. You will need to use the checkpoint as the init checkpoint for your training, but you should use the jsonl file from the competition dataset to precompute the tf-records to be passed to the training phase. </p>",
          "rawMarkdown": "These are the same files already included in the [starter notebook](https://www.kaggle.com/philculliton/using-tensorflow-2-0-w-bert-on-nq) under `input/bert-joint-baseline`. \n\nHere is how I see it. You will need to use the checkpoint as the init checkpoint for your training, but you should use the jsonl file from the competition dataset to precompute the tf-records to be passed to the training phase. "
        },
        {
          "id": 697407,
          "postDate": "2019-12-17T21:15:29.600Z",
          "content": "<p>The train data supposed to be tfrecords, doesn’t it?</p>",
          "rawMarkdown": "The train data supposed to be tfrecords, doesn’t it?"
        },
        {
          "id": 697429,
          "postDate": "2019-12-17T22:41:54.417Z",
          "content": "<p>Yes, according to the description <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#data-preparation\">here</a>:</p>\n\n<blockquote>\n  <p>The training set for the Natural Question is quite large so we precompute all the features for it as tensorflow examples.</p>\n</blockquote>\n\n<p>You then pass the precomputed tf-records as the value of <code>train_precomputed</code> flag as described <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#training-our-model\">here</a>.</p>",
          "rawMarkdown": "Yes, according to the description [here](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#data-preparation):\n\n&gt; The training set for the Natural Question is quite large so we precompute all the features for it as tensorflow examples.\n\nYou then pass the precomputed tf-records as the value of `train_precomputed` flag as described [here](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#training-our-model).\n\n"
        },
        {
          "id": 712280,
          "postDate": "2020-01-07T02:58:15.020Z",
          "content": "<p>But the precomputd train-set is prepared with a different dataset not with the competition simplified-nq-train.jsonl data. Do we need to retrain for the same. </p>",
          "rawMarkdown": "But the precomputd train-set is prepared with a different dataset not with the competition simplified-nq-train.jsonl data. Do we need to retrain for the same. "
        }
      ]
    },
    {
      "id": 696692,
      "postDate": "2019-12-17T00:31:12.680Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 697426,
      "author_name": "Miro",
      "author_url": "",
      "post_date": "2019-12-17T22:34:06.137000",
      "content": "<p>I had the same problem. Turned out to be a problem of Python 3 not supporting implicit conversion from str to bytes according to this <a href=\"https://stackoverflow.com/a/21916982/2070636\">SO answer</a>.</p>\n\n<p>Here is the workaround I used to fix it for now. I changed <a href=\"https://github.com/google-research/language/blob/623e016a76eaeb054d09a4173db1026685bc5b08/language/question_answering/bert_joint/prepare_nq_data.py#L58\">this line</a> in <a href=\"https://github.com/google-research/language/blob/master/language/question_answering/bert_joint/prepare_nq_data.py\"><code>prepare_nq_data</code></a>:</p>\n\n<p>```\n with gzip.GzipFile(fileobj=tf.gfile.Open(input_path)) as input_file:</p>\n\n<p>```\nTo</p>\n\n<p><code>\n with tf.io.gfile.GFile(input_path, mode='rb') as input_file:\n</code></p>\n\n<p>Not sure if that's the correct/best solution, but it seems to work so far.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 697662,
          "author_name": "AlexanderCh",
          "author_url": "",
          "post_date": "2019-12-18T08:44:11.047000",
          "content": "<p>with Python2 no problems</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 696844,
      "author_name": "Kepler456b",
      "author_url": "",
      "post_date": "2019-12-17T06:10:42.103000",
      "content": "<p>This is what I read:</p>\n\n<p>You should then download our model and preprocessed training set with:</p>\n\n<p>gsutil cp -R gs://bert-nq/bert-joint-baseline .</p>",
      "votes": 1,
      "replies": [
        {
          "id": 696899,
          "author_name": "AlexanderCh",
          "author_url": "",
          "post_date": "2019-12-17T08:01:31.897000",
          "content": "<p>Yes, I did it. But there is no any json with training data here:</p>\n\n<p>```\nbert-joint-baseline/nq-train.tfrecords-00000-of-00001\nbert-joint-baseline/bert_config.json\nbert-joint-baseline/vocab-nq.txt\nbert-joint-baseline/bert_joint.ckpt.data-00000-of-00001\nbert-joint-baseline/bert_joint.ckpt.index</p>\n\n<p>```</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 697385,
          "author_name": "Miro",
          "author_url": "",
          "post_date": "2019-12-17T20:31:35.230000",
          "content": "<p>These are the same files already included in the <a href=\"https://www.kaggle.com/philculliton/using-tensorflow-2-0-w-bert-on-nq\">starter notebook</a> under <code>input/bert-joint-baseline</code>. </p>\n\n<p>Here is how I see it. You will need to use the checkpoint as the init checkpoint for your training, but you should use the jsonl file from the competition dataset to precompute the tf-records to be passed to the training phase. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 697407,
          "author_name": "Kepler456b",
          "author_url": "",
          "post_date": "2019-12-17T21:15:29.600000",
          "content": "<p>The train data supposed to be tfrecords, doesn’t it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 697429,
          "author_name": "Miro",
          "author_url": "",
          "post_date": "2019-12-17T22:41:54.417000",
          "content": "<p>Yes, according to the description <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#data-preparation\">here</a>:</p>\n\n<blockquote>\n  <p>The training set for the Natural Question is quite large so we precompute all the features for it as tensorflow examples.</p>\n</blockquote>\n\n<p>You then pass the precomputed tf-records as the value of <code>train_precomputed</code> flag as described <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#training-our-model\">here</a>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 712280,
          "author_name": "Rraj",
          "author_url": "",
          "post_date": "2020-01-07T02:58:15.020000",
          "content": "<p>But the precomputd train-set is prepared with a different dataset not with the competition simplified-nq-train.jsonl data. Do we need to retrain for the same. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 696692,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-17T00:31:12.680000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "696626": "Has anyone tried BERT fine tuning?\nI have tried to go through these [steps](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint#data-preparation)\nbut there is problem with Data Preparation step\n```\npython -m language.question_answering.bert_joint.prepare_nq_data \\\n  --logtostderr \\\n  --input_jsonl simplified-nq-train.jsonl \\\n  --output_tfrecord ~/output_dir/nq-train.tfrecords-00000-of-00001 \\\n  --max_seq_length=512 \\\n  --include_unknowns=0.02 \\\n  --vocab_file=bert-joint-baseline/vocab-nq.txt\n\n```\n\nError:\n```\n    for example in get_examples(FLAGS.input_jsonl):\n  File \"/data1/achaptykov/model/googlebert/language/question_answering/bert_joint/prepare_nq_data.py\", line 59, in get_examples\n    for line in input_file:\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 374, in readline\n    return self._buffer.readline(size)\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/_compression.py\", line 68, in readinto\n    data = self.read(len(byte_view))\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 463, in read\n    if not self._read_gzip_header():\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 406, in _read_gzip_header\n    magic = self._fp.read(2)\n  File \"/home/achaptykov/anaconda3/envs/py36-test/lib/python3.6/gzip.py\", line 91, in read\n    self.file.read(size-self._length+read)\nTypeError: can't concat str to bytes\n```\nShould I put json from competition data? Or shall I modified it.",
    "697426": "I had the same problem. Turned out to be a problem of Python 3 not supporting implicit conversion from str to bytes according to this [SO answer](https://stackoverflow.com/a/21916982/2070636).\n\n\nHere is the workaround I used to fix it for now. I changed [this line](https://github.com/google-research/language/blob/623e016a76eaeb054d09a4173db1026685bc5b08/language/question_answering/bert_joint/prepare_nq_data.py#L58) in [`prepare_nq_data`](https://github.com/google-research/language/blob/master/language/question_answering/bert_joint/prepare_nq_data.py):\n\n```\n with gzip.GzipFile(fileobj=tf.gfile.Open(input_path)) as input_file:\n\n```\nTo\n\n```\n with tf.io.gfile.GFile(input_path, mode='rb') as input_file:\n```\n\nNot sure if that's the correct/best solution, but it seems to work so far.",
    "696844": "This is what I read:\n\nYou should then download our model and preprocessed training set with:\n\ngsutil cp -R gs://bert-nq/bert-joint-baseline .",
    "696692": ""
  }
}