{
  "id": 119462,
  "title": "CV vs LB",
  "url": "/competitions/tensorflow2-question-answering/discussion/119462",
  "author_name": "",
  "post_date": "2019-11-28T19:37:51.646852900Z",
  "votes": 6,
  "comment_count": 10,
  "views": 0,
  "content": "<p>So I am still struggling implementing the metric. I guess I implemented it correctly but still have a cv of 0.4 compared to a LB of 0.73. Any similar experiences?</p>",
  "messages": [
    {
      "id": "683772",
      "postDate": "11/28/2019 19:37:51",
      "content": "<p>So I am still struggling implementing the metric. I guess I implemented it correctly but still have a cv of 0.4 compared to a LB of 0.73. Any similar experiences?</p>",
      "rawMarkdown": "So I am still struggling implementing the metric. I guess I implemented it correctly but still have a cv of 0.4 compared to a LB of 0.73. Any similar experiences?",
      "votes": null
    },
    {
      "id": "683773",
      "postDate": "11/28/2019 19:39:45",
      "content": "<p>For more information: As a first step I just take a part set of train which is approx same size as test and apply baseline jointbert model</p>",
      "rawMarkdown": "For more information: As a first step I just take a part set of train which is approx same size as test and apply baseline jointbert model",
      "votes": null
    },
    {
      "id": "683814",
      "postDate": "11/28/2019 21:07:34",
      "content": "<p>There's also a dev set there. \nI haven't yet implemented the metric, probably it will be shared soon :)</p>",
      "rawMarkdown": "There's also a dev set there. \nI haven't yet implemented the metric, probably it will be shared soon :)",
      "votes": null
    },
    {
      "id": "684014",
      "postDate": "11/29/2019 06:15:19",
      "content": "<p>Hi，what does cv means？ cross validate？ Or a single model</p>",
      "rawMarkdown": "Hi，what does cv means？ cross validate？ Or a single model",
      "votes": null
    },
    {
      "id": "684042",
      "postDate": "11/29/2019 06:48:45",
      "content": "<p>cross validation</p>",
      "rawMarkdown": "cross validation",
      "votes": null
    },
    {
      "id": "684092",
      "postDate": "11/29/2019 08:20:33",
      "content": "<p>By dev set you mean one available on <a href=\"https://ai.google.com/research/NaturalQuestions/download\">https://ai.google.com/research/NaturalQuestions/download</a> ? Does it have compatible enough format?</p>",
      "rawMarkdown": "By dev set you mean one available on https://ai.google.com/research/NaturalQuestions/download ? Does it have compatible enough format?",
      "votes": null
    },
    {
      "id": "684095",
      "postDate": "11/29/2019 08:22:27",
      "content": "<p>I have a pytorch model similar to the baseline (most likely a bit worse), and I have a metric implementation which ignores YES/NO answers and allows multiple short answers, it gets me 0.52 overall F1, 0.73 F1 on long answers and 0.40 on short answers (on a tiny holdout). Obviously it's quite different from what is currently used on LB. I didn't submit the model yet (need to fix a few things) but I expect it to be well below 0.7 on the LB. (EDIT: this metrics are with a bug, real metrics are much lower unfortunately)</p>",
      "rawMarkdown": "I have a pytorch model similar to the baseline (most likely a bit worse), and I have a metric implementation which ignores YES/NO answers and allows multiple short answers, it gets me 0.52 overall F1, 0.73 F1 on long answers and 0.40 on short answers (on a tiny holdout). Obviously it's quite different from what is currently used on LB. I didn't submit the model yet (need to fix a few things) but I expect it to be well below 0.7 on the LB. (EDIT: this metrics are with a bug, real metrics are much lower unfortunately)",
      "votes": null
    },
    {
      "id": "684096",
      "postDate": "11/29/2019 08:22:31",
      "content": "<p>Yes, this one. You can adapt the format, it’s almost the same. </p>",
      "rawMarkdown": "Yes, this one. You can adapt the format, it’s almost the same.",
      "votes": null
    },
    {
      "id": "684706",
      "postDate": "11/30/2019 10:21:08",
      "content": "<p>Turns out my metric was not correct. Now cv and LB matches perfectly. Also LB probing confirms distribution of Yes/ No of test is same as in train set.</p>",
      "rawMarkdown": "Turns out my metric was not correct. Now cv and LB matches perfectly. Also LB probing confirms distribution of Yes/ No of test is same as in train set.",
      "votes": null
    },
    {
      "id": "684712",
      "postDate": "11/30/2019 10:43:23",
      "content": "<p>Thanks! Good to know. </p>",
      "rawMarkdown": "Thanks! Good to know.",
      "votes": null
    },
    {
      "id": "685917",
      "postDate": "12/02/2019 14:24:17",
      "content": "<p>One question from your experience: for true possitive, should the slice be definetely equal?</p>",
      "rawMarkdown": "One question from your experience: for true possitive, should the slice be definetely equal?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 683773,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "11/28/2019 19:39:45",
      "content": "<p>For more information: As a first step I just take a part set of train which is approx same size as test and apply baseline jointbert model</p>",
      "votes": null,
      "replies": [
        {
          "id": 683814,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "11/28/2019 21:07:34",
          "content": "<p>There's also a dev set there. \nI haven't yet implemented the metric, probably it will be shared soon :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 684092,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "11/29/2019 08:20:33",
          "content": "<p>By dev set you mean one available on <a href=\"https://ai.google.com/research/NaturalQuestions/download\">https://ai.google.com/research/NaturalQuestions/download</a> ? Does it have compatible enough format?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 684096,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "11/29/2019 08:22:31",
          "content": "<p>Yes, this one. You can adapt the format, it’s almost the same. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 684014,
      "author_name": "zhaomeng1126",
      "author_url": "",
      "post_date": "11/29/2019 06:15:19",
      "content": "<p>Hi，what does cv means？ cross validate？ Or a single model</p>",
      "votes": null,
      "replies": [
        {
          "id": 684042,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "11/29/2019 06:48:45",
          "content": "<p>cross validation</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 684095,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "11/29/2019 08:22:27",
      "content": "<p>I have a pytorch model similar to the baseline (most likely a bit worse), and I have a metric implementation which ignores YES/NO answers and allows multiple short answers, it gets me 0.52 overall F1, 0.73 F1 on long answers and 0.40 on short answers (on a tiny holdout). Obviously it's quite different from what is currently used on LB. I didn't submit the model yet (need to fix a few things) but I expect it to be well below 0.7 on the LB. (EDIT: this metrics are with a bug, real metrics are much lower unfortunately)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 684706,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "11/30/2019 10:21:08",
      "content": "<p>Turns out my metric was not correct. Now cv and LB matches perfectly. Also LB probing confirms distribution of Yes/ No of test is same as in train set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 684712,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "11/30/2019 10:43:23",
          "content": "<p>Thanks! Good to know. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 685917,
      "author_name": "httpwwwfszyc",
      "author_url": "",
      "post_date": "12/02/2019 14:24:17",
      "content": "<p>One question from your experience: for true possitive, should the slice be definetely equal?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "683772": "So I am still struggling implementing the metric. I guess I implemented it correctly but still have a cv of 0.4 compared to a LB of 0.73. Any similar experiences?",
    "683773": "For more information: As a first step I just take a part set of train which is approx same size as test and apply baseline jointbert model",
    "683814": "There's also a dev set there. \nI haven't yet implemented the metric, probably it will be shared soon :)",
    "684014": "Hi，what does cv means？ cross validate？ Or a single model",
    "684042": "cross validation",
    "684092": "By dev set you mean one available on https://ai.google.com/research/NaturalQuestions/download ? Does it have compatible enough format?",
    "684095": "I have a pytorch model similar to the baseline (most likely a bit worse), and I have a metric implementation which ignores YES/NO answers and allows multiple short answers, it gets me 0.52 overall F1, 0.73 F1 on long answers and 0.40 on short answers (on a tiny holdout). Obviously it's quite different from what is currently used on LB. I didn't submit the model yet (need to fix a few things) but I expect it to be well below 0.7 on the LB. (EDIT: this metrics are with a bug, real metrics are much lower unfortunately)",
    "684096": "Yes, this one. You can adapt the format, it’s almost the same.",
    "684706": "Turns out my metric was not correct. Now cv and LB matches perfectly. Also LB probing confirms distribution of Yes/ No of test is same as in train set.",
    "684712": "Thanks! Good to know.",
    "685917": "One question from your experience: for true possitive, should the slice be definetely equal?"
  },
  "source": "meta"
}