{
  "id": 122855,
  "title": "[Warning] -- Wrong model in my TF2 traning kernel",
  "url": "/competitions/tensorflow2-question-answering/discussion/122855",
  "author_name": "Yih-Dar SHIEH",
  "post_date": "2019-12-23T09:03:31.648000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I previously published <a href=\"https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models/comments\">TF2 training kernel</a>. When it was published, it was in <code>Commit Version 4</code>.</p>\n\n<p>After working on metric and inference kernel, then I realized that the model definition in my training kernel is WRONG! -- I don't know how many people are using my kernel, but if any, please   check the latest version! I am really sorry for the time, money or GCP credit you wasted due to my mistake...</p>\n\n<p>The mistake is:</p>\n\n<ul>\n<li><p>In the previous version, I only use the output of [CLS] token (so shape = (batch_size, hidden_dim)),  passed it to a <code>Dense layer of dimension seq_len</code>, to get an output of shape (batch_size, seq_len) which is interpreted as <code>start_pos_logit</code>. Similarly for <code>end_pos_logit</code>.</p></li>\n<li><p>In the original bert joint baseline model, the full sequence output from bert (so shape = (batch_size, seq_len, hidden_dim)) was passed to a <code>Dense layer of dimension 2</code>, to get an output of shape (batch_size, seq_len, 2). Then it is reshaped to (2, batch_size, seq_len). Then it is interpreted as <code>start_pos_logit</code> and <code>end_pos_logit</code>.</p></li>\n</ul>\n\n<p>As you can see, my model only use [CLS] token output, while the original model use the full information from each token in the sequence.</p>\n\n<p>I get really low score with the wrong model, and I get much better result with the fixed model. (I am working on distilled bert for now, so the score is not as good as the bert large).</p>\n\n<p>Conclusion: If you used or want to use my kernel, use the latest version (Version 11).</p>",
  "messages": [
    {
      "id": 701228,
      "postDate": "2019-12-23T09:03:31.650Z",
      "content": "<p>I previously published <a href=\"https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models/comments\">TF2 training kernel</a>. When it was published, it was in <code>Commit Version 4</code>.</p>\n\n<p>After working on metric and inference kernel, then I realized that the model definition in my training kernel is WRONG! -- I don't know how many people are using my kernel, but if any, please   check the latest version! I am really sorry for the time, money or GCP credit you wasted due to my mistake...</p>\n\n<p>The mistake is:</p>\n\n<ul>\n<li><p>In the previous version, I only use the output of [CLS] token (so shape = (batch_size, hidden_dim)),  passed it to a <code>Dense layer of dimension seq_len</code>, to get an output of shape (batch_size, seq_len) which is interpreted as <code>start_pos_logit</code>. Similarly for <code>end_pos_logit</code>.</p></li>\n<li><p>In the original bert joint baseline model, the full sequence output from bert (so shape = (batch_size, seq_len, hidden_dim)) was passed to a <code>Dense layer of dimension 2</code>, to get an output of shape (batch_size, seq_len, 2). Then it is reshaped to (2, batch_size, seq_len). Then it is interpreted as <code>start_pos_logit</code> and <code>end_pos_logit</code>.</p></li>\n</ul>\n\n<p>As you can see, my model only use [CLS] token output, while the original model use the full information from each token in the sequence.</p>\n\n<p>I get really low score with the wrong model, and I get much better result with the fixed model. (I am working on distilled bert for now, so the score is not as good as the bert large).</p>\n\n<p>Conclusion: If you used or want to use my kernel, use the latest version (Version 11).</p>",
      "rawMarkdown": "I previously published [TF2 training kernel](https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models/comments). When it was published, it was in `Commit Version 4`.\n\nAfter working on metric and inference kernel, then I realized that the model definition in my training kernel is WRONG! -- I don't know how many people are using my kernel, but if any, please   check the latest version! I am really sorry for the time, money or GCP credit you wasted due to my mistake...\n\nThe mistake is:\n\n- In the previous version, I only use the output of [CLS] token (so shape = (batch\\_size, hidden\\_dim)),  passed it to a `Dense layer of dimension seq_len`, to get an output of shape (batch\\_size, seq\\_len) which is interpreted as `start_pos_logit`. Similarly for `end_pos_logit`.\n\n- In the original bert joint baseline model, the full sequence output from bert (so shape = (batch\\_size, seq_len, hidden\\_dim)) was passed to a `Dense layer of dimension 2`, to get an output of shape (batch\\_size, seq\\_len, 2). Then it is reshaped to (2, batch\\_size, seq\\_len). Then it is interpreted as `start_pos_logit` and `end_pos_logit`.\n\nAs you can see, my model only use [CLS] token output, while the original model use the full information from each token in the sequence.\n\nI get really low score with the wrong model, and I get much better result with the fixed model. (I am working on distilled bert for now, so the score is not as good as the bert large).\n\nConclusion: If you used or want to use my kernel, use the latest version (Version 11).\n\n",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "701228": "I previously published [TF2 training kernel](https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models/comments). When it was published, it was in `Commit Version 4`.\n\nAfter working on metric and inference kernel, then I realized that the model definition in my training kernel is WRONG! -- I don't know how many people are using my kernel, but if any, please   check the latest version! I am really sorry for the time, money or GCP credit you wasted due to my mistake...\n\nThe mistake is:\n\n- In the previous version, I only use the output of [CLS] token (so shape = (batch\\_size, hidden\\_dim)),  passed it to a `Dense layer of dimension seq_len`, to get an output of shape (batch\\_size, seq\\_len) which is interpreted as `start_pos_logit`. Similarly for `end_pos_logit`.\n\n- In the original bert joint baseline model, the full sequence output from bert (so shape = (batch\\_size, seq_len, hidden\\_dim)) was passed to a `Dense layer of dimension 2`, to get an output of shape (batch\\_size, seq\\_len, 2). Then it is reshaped to (2, batch\\_size, seq\\_len). Then it is interpreted as `start_pos_logit` and `end_pos_logit`.\n\nAs you can see, my model only use [CLS] token output, while the original model use the full information from each token in the sequence.\n\nI get really low score with the wrong model, and I get much better result with the fixed model. (I am working on distilled bert for now, so the score is not as good as the bert large).\n\nConclusion: If you used or want to use my kernel, use the latest version (Version 11).\n\n"
  }
}