{
  "id": 125594,
  "title": "question (or my problem) about CV and LB",
  "url": "/competitions/tensorflow2-question-answering/discussion/125594",
  "author_name": "",
  "post_date": "2020-01-12T05:12:01.178837600Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>it seems that in public test data majority of the short answers (200+) are in slices (class 3) instead of blank, which is far distant from the training data distribution  (155225 class-0 vs 106926 class-3). After I changed my f1-eval function from @christofhenkel to @kentaronakanishi consider False Positive my CV that has LB 0.62 decreases from 0.68 to 0.46. \nSo, I  want to know how to think about the behavior？Are you consider private data will be more real distributed like training data?</p>\n\n<p>(metric: <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061</a>)</p>\n\n<p>It is very nice to find my models overfit the public LB at the end of the competition. And all my experiment did before has no much sense and information😂 </p>\n\n<p>Btw, if anyone want to achieve for tensorflow prize, you can team me as I already convert pytorch huggingface model to keras-bert, keras-xlnet, keras-roberta etc. (Although I have questions whether calling tf.keras is treated as using tensorflow 2.0 technics.)</p>",
  "messages": [
    {
      "id": "716684",
      "postDate": "01/12/2020 05:12:01",
      "content": "<p>it seems that in public test data majority of the short answers (200+) are in slices (class 3) instead of blank, which is far distant from the training data distribution  (155225 class-0 vs 106926 class-3). After I changed my f1-eval function from @christofhenkel to @kentaronakanishi consider False Positive my CV that has LB 0.62 decreases from 0.68 to 0.46. \nSo, I  want to know how to think about the behavior？Are you consider private data will be more real distributed like training data?</p>\n\n<p>(metric: <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061</a>)</p>\n\n<p>It is very nice to find my models overfit the public LB at the end of the competition. And all my experiment did before has no much sense and information😂 </p>\n\n<p>Btw, if anyone want to achieve for tensorflow prize, you can team me as I already convert pytorch huggingface model to keras-bert, keras-xlnet, keras-roberta etc. (Although I have questions whether calling tf.keras is treated as using tensorflow 2.0 technics.)</p>",
      "rawMarkdown": "it seems that in public test data majority of the short answers (200+) are in slices (class 3) instead of blank, which is far distant from the training data distribution  (155225 class\\-0 vs 106926 class\\-3). After I changed my f1-eval function from @christofhenkel to @kentaronakanishi consider False Positive my CV that has LB 0.62 decreases from 0.68 to 0.46. \nSo, I  want to know how to think about the behavior？Are you consider private data will be more real distributed like training data?\n\n(metric: https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061)\n\nIt is very nice to find my models overfit the public LB at the end of the competition. And all my experiment did before has no much sense and information😂 \n\nBtw, if anyone want to achieve for tensorflow prize, you can team me as I already convert pytorch huggingface model to keras-bert, keras-xlnet, keras-roberta etc. (Although I have questions whether calling tf.keras is treated as using tensorflow 2.0 technics.)",
      "votes": null
    },
    {
      "id": "725183",
      "postDate": "01/21/2020 22:08:17",
      "content": "<p>same similar problem I have , for my 2000 example validation set , f1 score about 0.49 which is very different with my public score(0.61).\nI just try to raise my public score. </p>",
      "rawMarkdown": "same similar problem I have , for my 2000 example validation set , f1 score about 0.49 which is very different with my public score(0.61).\nI just try to raise my public score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 725183,
      "author_name": "michealkim",
      "author_url": "",
      "post_date": "01/21/2020 22:08:17",
      "content": "<p>same similar problem I have , for my 2000 example validation set , f1 score about 0.49 which is very different with my public score(0.61).\nI just try to raise my public score. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "716684": "it seems that in public test data majority of the short answers (200+) are in slices (class 3) instead of blank, which is far distant from the training data distribution  (155225 class\\-0 vs 106926 class\\-3). After I changed my f1-eval function from @christofhenkel to @kentaronakanishi consider False Positive my CV that has LB 0.62 decreases from 0.68 to 0.46. \nSo, I  want to know how to think about the behavior？Are you consider private data will be more real distributed like training data?\n\n(metric: https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061)\n\nIt is very nice to find my models overfit the public LB at the end of the competition. And all my experiment did before has no much sense and information😂 \n\nBtw, if anyone want to achieve for tensorflow prize, you can team me as I already convert pytorch huggingface model to keras-bert, keras-xlnet, keras-roberta etc. (Although I have questions whether calling tf.keras is treated as using tensorflow 2.0 technics.)",
    "725183": "same similar problem I have , for my 2000 example validation set , f1 score about 0.49 which is very different with my public score(0.61).\nI just try to raise my public score."
  },
  "source": "meta"
}