{
  "id": 127512,
  "title": "The actual competition metric (unofficial)",
  "url": "/competitions/tensorflow2-question-answering/discussion/127512",
  "author_name": "",
  "post_date": "2020-01-24T11:33:16.322273Z",
  "votes": 12,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Already mentioned the crucial change w.r.t. to <code>nq_eval</code> in a <a href=\"https://www.kaggle.com/kashnitsky/solid-bert-joint-baseline-0-65-0-66\">shared Notebook</a> with our best single model (yes, BERT-joint again, on steroids), but to elaborate a bit, what worked ideally for us is the following:\n<code>nq_eval_yorko</code> is a small modification of the <code>nq_eval</code> script, uploaded to <a href=\"https://github.com/Yorko/natural-questions\">my fork</a> of the natural-questions repo. The key difference is just one line:\n<img src=\"https://habrastorage.org/webt/ar/12/g0/ar12g0cea9fnojk_ghhjs2wyyae.png\"></p>\n\n<p>So that's exactly what stated on the Evaluation page (and Phil fixed the metric correctly):\n&gt; The metric in this competition diverges from the original metric in two key respects: 1) short and long answer formats do not receive separate scores, but are instead combined into a micro F1 score across both formats, and 2) this competition's metric does not use confidence scores to find an optimal threshold for predictions.</p>\n\n<p>There in the repo you also see <code>nq-dev-all-annotations.jsonl.gz</code> - it's the same official dev set from <a href=\"https://ai.google.com/research/NaturalQuestions/download\">Natural Questions</a> but I left only annotations for ~100x faster metric computation (no need to read all these document tokens). </p>\n\n<p>When you have your <code>pred.json</code> file, evaluation is the following:\n- <code>python nq_eval_yorko.py --gold_path=nq-dev-all-annotations.jsonl.gz --predictions_path=pred.json</code> - thus you get Long and Short F1 scores and optimal thresholds\n- <code>python nq_eval_yorko.py --gold_path=nq-dev-all-annotations.jsonl.gz --predictions_path=pred.json --score_thres_long $LONG_THRES --score_thres_short $SHORT_THRES</code> - this will show only Long, Short, and All F1 scores. This \"All\" is our competition metric. </p>\n\n<p>For sure, this can be refactored into just one script. And preserving threshold tuning is also nice (<code>nq_eval_yorko</code> only applies the provided thresholds) but to be honest, I'm in such a grief with our 13th place that good enough (my brother gave me a nickname \"Kaggle Bachelor\").</p>",
  "messages": [
    {
      "id": "728076",
      "postDate": "01/24/2020 11:33:16",
      "content": "<p>Already mentioned the crucial change w.r.t. to <code>nq_eval</code> in a <a href=\"https://www.kaggle.com/kashnitsky/solid-bert-joint-baseline-0-65-0-66\">shared Notebook</a> with our best single model (yes, BERT-joint again, on steroids), but to elaborate a bit, what worked ideally for us is the following:\n<code>nq_eval_yorko</code> is a small modification of the <code>nq_eval</code> script, uploaded to <a href=\"https://github.com/Yorko/natural-questions\">my fork</a> of the natural-questions repo. The key difference is just one line:\n<img src=\"https://habrastorage.org/webt/ar/12/g0/ar12g0cea9fnojk_ghhjs2wyyae.png\"></p>\n\n<p>So that's exactly what stated on the Evaluation page (and Phil fixed the metric correctly):\n&gt; The metric in this competition diverges from the original metric in two key respects: 1) short and long answer formats do not receive separate scores, but are instead combined into a micro F1 score across both formats, and 2) this competition's metric does not use confidence scores to find an optimal threshold for predictions.</p>\n\n<p>There in the repo you also see <code>nq-dev-all-annotations.jsonl.gz</code> - it's the same official dev set from <a href=\"https://ai.google.com/research/NaturalQuestions/download\">Natural Questions</a> but I left only annotations for ~100x faster metric computation (no need to read all these document tokens). </p>\n\n<p>When you have your <code>pred.json</code> file, evaluation is the following:\n- <code>python nq_eval_yorko.py --gold_path=nq-dev-all-annotations.jsonl.gz --predictions_path=pred.json</code> - thus you get Long and Short F1 scores and optimal thresholds\n- <code>python nq_eval_yorko.py --gold_path=nq-dev-all-annotations.jsonl.gz --predictions_path=pred.json --score_thres_long $LONG_THRES --score_thres_short $SHORT_THRES</code> - this will show only Long, Short, and All F1 scores. This \"All\" is our competition metric. </p>\n\n<p>For sure, this can be refactored into just one script. And preserving threshold tuning is also nice (<code>nq_eval_yorko</code> only applies the provided thresholds) but to be honest, I'm in such a grief with our 13th place that good enough (my brother gave me a nickname \"Kaggle Bachelor\").</p>",
      "rawMarkdown": "Already mentioned the crucial change w.r.t. to `nq_eval` in a [shared Notebook](https://www.kaggle.com/kashnitsky/solid-bert-joint-baseline-0-65-0-66) with our best single model (yes, BERT-joint again, on steroids), but to elaborate a bit, what worked ideally for us is the following:\n`nq_eval_yorko` is a small modification of the `nq_eval` script, uploaded to [my fork](https://github.com/Yorko/natural-questions) of the natural-questions repo. The key difference is just one line:\n<img src=\"https://habrastorage.org/webt/ar/12/g0/ar12g0cea9fnojk_ghhjs2wyyae.png\">\n\nSo that's exactly what stated on the Evaluation page (and Phil fixed the metric correctly):\n&gt; The metric in this competition diverges from the original metric in two key respects: 1) short and long answer formats do not receive separate scores, but are instead combined into a micro F1 score across both formats, and 2) this competition's metric does not use confidence scores to find an optimal threshold for predictions.\n\nThere in the repo you also see `nq-dev-all-annotations.jsonl.gz` - it's the same official dev set from [Natural Questions](https://ai.google.com/research/NaturalQuestions/download) but I left only annotations for ~100x faster metric computation (no need to read all these document tokens). \n\nWhen you have your `pred.json` file, evaluation is the following:\n- `python nq_eval_yorko.py --gold_path=nq-dev-all-annotations.jsonl.gz --predictions_path=pred.json` - thus you get Long and Short F1 scores and optimal thresholds\n- `python nq_eval_yorko.py --gold_path=nq-dev-all-annotations.jsonl.gz --predictions_path=pred.json --score_thres_long $LONG_THRES --score_thres_short $SHORT_THRES` - this will show only Long, Short, and All F1 scores. This \"All\" is our competition metric. \n\nFor sure, this can be refactored into just one script. And preserving threshold tuning is also nice (`nq_eval_yorko` only applies the provided thresholds) but to be honest, I'm in such a grief with our 13th place that good enough (my brother gave me a nickname \"Kaggle Bachelor\").",
      "votes": null
    },
    {
      "id": "728088",
      "postDate": "01/24/2020 11:42:20",
      "content": "<p>Another key observation that wasn't really shared by organizers is that the NQ dev test contains more than one annotation per example, like test, when train only had one.  Then scoring depends on the number of annotations, see  part of the NQ dataset paper (<a href=\"https://storage.googleapis.com/pub-tools-public-publication-data/pdf/1f7b46b5378d757553d3e92ead36bda2e4254244.pdf\">https://storage.googleapis.com/pub-tools-public-publication-data/pdf/1f7b46b5378d757553d3e92ead36bda2e4254244.pdf</a>):</p>\n\n<blockquote>\n  <p>If at least2out of 5 annotators have given a non-null long answer on the example, then the system is required to output a non-null answer that is seen at least once in the 5 annotations; conversely if fewer than 2 annotators give a non-null long answer, the system is required to return NULL as its output</p>\n</blockquote>\n\n<p>Without this it is impossible to have a good CV - LB correlation IMHO.</p>\n\n<p>We found about this quite late, maybe two weeks before end.  I feel for all who didn't saw this.  Without proper CV then this becomes more of a lottery.</p>",
      "rawMarkdown": "Another key observation that wasn't really shared by organizers is that the NQ dev test contains more than one annotation per example, like test, when train only had one.  Then scoring depends on the number of annotations, see  part of the NQ dataset paper (https://storage.googleapis.com/pub-tools-public-publication-data/pdf/1f7b46b5378d757553d3e92ead36bda2e4254244.pdf):\n\n&gt; If at least2out of 5 annotators have given a non-null long answer on the example, then the system is required to output a non-null answer that is seen at least once in the 5 annotations; conversely if fewer than 2 annotators give a non-null long answer, the system is required to return NULL as its output\n\nWithout this it is impossible to have a good CV - LB correlation IMHO.\n\nWe found about this quite late, maybe two weeks before end.  I feel for all who didn't saw this.  Without proper CV then this becomes more of a lottery.",
      "votes": null
    },
    {
      "id": "728107",
      "postDate": "01/24/2020 12:08:30",
      "content": "<p>Aha! I missed the \"fewer than 2\". I tuned my models to a metric that returned null only if fewer than 1 (ie zero) annotators gave non-null output.</p>",
      "rawMarkdown": "Aha! I missed the \"fewer than 2\". I tuned my models to a metric that returned null only if fewer than 1 (ie zero) annotators gave non-null output.",
      "votes": null
    },
    {
      "id": "728109",
      "postDate": "01/24/2020 12:10:43",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> thanks for this. Sorry to see how painfully close you were to gold. That must be frustrating!</p>",
      "rawMarkdown": "kashnitsky thanks for this. Sorry to see how painfully close you were to gold. That must be frustrating!",
      "votes": null
    },
    {
      "id": "728157",
      "postDate": "01/24/2020 13:16:11",
      "content": "<p>True. That’s why we validated only with the dev set once the metric was fixed. </p>",
      "rawMarkdown": "True. That’s why we validated only with the dev set once the metric was fixed.",
      "votes": null
    },
    {
      "id": "728160",
      "postDate": "01/24/2020 13:17:17",
      "content": "<p>Well, actually we also submitted once to Natural Questions but with one sub per week it’s not a good source of feedback :)</p>",
      "rawMarkdown": "Well, actually we also submitted once to Natural Questions but with one sub per week it’s not a good source of feedback :)",
      "votes": null
    },
    {
      "id": "728163",
      "postDate": "01/24/2020 13:20:10",
      "content": "<p>Thanks, Ken! Yeah, I also had a girl born in December. So it’s also stealing time from her :(</p>\n\n<p>But the good news is that Dmitry and Oleg are the real masters, cool experience working with them. So hopefully, Trimorph will be back soon :)</p>",
      "rawMarkdown": "Thanks, Ken! Yeah, I also had a girl born in December. So it’s also stealing time from her :(\n\nBut the good news is that Dmitry and Oleg are the real masters, cool experience working with them. So hopefully, Trimorph will be back soon :)",
      "votes": null
    },
    {
      "id": "728205",
      "postDate": "01/24/2020 13:57:02",
      "content": "<p>Did you submit to the nq leaderboard: <a href=\"https://ai.google.com/research/NaturalQuestions/leaderboard\">https://ai.google.com/research/NaturalQuestions/leaderboard</a>? If you did, what's the score compared to the metric here?</p>",
      "rawMarkdown": "Did you submit to the nq leaderboard: https://ai.google.com/research/NaturalQuestions/leaderboard? If you did, what's the score compared to the metric here?",
      "votes": null
    },
    {
      "id": "728213",
      "postDate": "01/24/2020 14:01:57",
      "content": "<p>Yes. That was pretty straightforward technically. And I was surprised to realize that such a simple idea of having one more big test set came to me so late. Anyway, only 1 sun per week. </p>\n\n<p>The score is a bit too optimistic there due to automatic threshold tuning. So it was some 1.5-2 points higher. </p>",
      "rawMarkdown": "Yes. That was pretty straightforward technically. And I was surprised to realize that such a simple idea of having one more big test set came to me so late. Anyway, only 1 sun per week. \n\nThe score is a bit too optimistic there due to automatic threshold tuning. So it was some 1.5-2 points higher.",
      "votes": null
    },
    {
      "id": "728215",
      "postDate": "01/24/2020 14:06:01",
      "content": "<p>I’m there on that LB, yorko. 69.4/57.1 long/short. While locally it was 67.8/56.2 for that version of Bert-large WWM, squad-pretrained. 63 LB in our competition </p>",
      "rawMarkdown": "I’m there on that LB, yorko. 69.4/57.1 long/short. While locally it was 67.8/56.2 for that version of Bert-large WWM, squad-pretrained. 63 LB in our competition",
      "votes": null
    },
    {
      "id": "728226",
      "postDate": "01/24/2020 14:15:19",
      "content": "<p>That's great, thanks. I hope I get some time to adjust my code and submit as well.</p>",
      "rawMarkdown": "That's great, thanks. I hope I get some time to adjust my code and submit as well.",
      "votes": null
    },
    {
      "id": "728236",
      "postDate": "01/24/2020 14:19:40",
      "content": "<p>Yes, just follow the quick-start, it went smooth for me. And there you’re given 24 hour slot (2x P100 afair) - so unfortunately, crazy ensembles are more than welcome there. </p>",
      "rawMarkdown": "Yes, just follow the quick-start, it went smooth for me. And there you’re given 24 hour slot (2x P100 afair) - so unfortunately, crazy ensembles are more than welcome there.",
      "votes": null
    },
    {
      "id": "728271",
      "postDate": "01/24/2020 14:51:39",
      "content": "<p>Congratulations and good luck with being a father. Big job. I just became a grandfather this week, but that won't be as much of a lifestyle change as being a father.</p>",
      "rawMarkdown": "Congratulations and good luck with being a father. Big job. I just became a grandfather this week, but that won't be as much of a lifestyle change as being a father.",
      "votes": null
    },
    {
      "id": "728278",
      "postDate": "01/24/2020 14:57:30",
      "content": "<p>Cool! Congrats as well!</p>",
      "rawMarkdown": "Cool! Congrats as well!",
      "votes": null
    },
    {
      "id": "728309",
      "postDate": "01/24/2020 15:38:33",
      "content": "<p>I am still uncertain as to what the metric was. The <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation\">evaluation page</a> says:</p>\n\n<blockquote>\n  <p>There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.</p>\n</blockquote>\n\n<p>By saying it should be left blank if NO answer applies, they directly contradict the <code>fewer than 2</code> correctly quoted above from the original paper by <a href=\"/cpmpml\">@cpmpml</a> which they point us to on the same page.</p>\n\n<p>I went by the competition page, which clearly states <strong>none</strong>, meaning <code>fewer than 1</code></p>",
      "rawMarkdown": "I am still uncertain as to what the metric was. The [evaluation page](https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation) says:\n&gt; There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.\n\nBy saying it should be left blank if NO answer applies, they directly contradict the `fewer than 2` correctly quoted above from the original paper by @cpmpml which they point us to on the same page.\n\nI went by the competition page, which clearly states **none**, meaning `fewer than 1`",
      "votes": null
    },
    {
      "id": "728314",
      "postDate": "01/24/2020 15:43:58",
      "content": "<p>A better choice is to delve into <code>nq_eval</code> code. There it's clearly seen that an answer is considered correct if &gt;= 2 annotators marked it as correct. </p>",
      "rawMarkdown": "A better choice is to delve into `nq_eval` code. There it's clearly seen that an answer is considered correct if &gt;= 2 annotators marked it as correct.",
      "votes": null
    },
    {
      "id": "728329",
      "postDate": "01/24/2020 16:05:47",
      "content": "<p>But that contradicts what is very clearly stated on the official competition <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation\">evaluation page</a>. If no code is provided, surely we must go by the competition description, which says only leave it blank if <code>no answer</code> applies.</p>\n\n<p><a href=\"/philculliton\">@philculliton</a> ?</p>",
      "rawMarkdown": "But that contradicts what is very clearly stated on the official competition [evaluation page](https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation). If no code is provided, surely we must go by the competition description, which says only leave it blank if `no answer` applies.\n\n@philculliton ?",
      "votes": null
    },
    {
      "id": "728347",
      "postDate": "01/24/2020 16:29:03",
      "content": "<p>Ok, I see. Indeed, yet another fail. </p>\n\n<p>At some point I actually ignored everything written in the competition description and digged into code and paper. </p>\n\n<p>Had enough troubles with paper/code controversy, let alone other surprises from organizers. </p>",
      "rawMarkdown": "Ok, I see. Indeed, yet another fail. \n\nAt some point I actually ignored everything written in the competition description and digged into code and paper. \n\nHad enough troubles with paper/code controversy, let alone other surprises from organizers.",
      "votes": null
    },
    {
      "id": "728370",
      "postDate": "01/24/2020 16:45:31",
      "content": "<p>You make a good point but I think mistakes and controversies will be inevitable with a complex process like this. Running these competitions is a very tough job. So I believe the organizers do a great job and are always willing to listen to questions and sometimes make changes during a competition.  I do like the kaggle platorm and get a lot of personal benefit at no financial cost.</p>\n\n<p>Maybe they should always publish the metric code. I can't think of a disadvantage of that, if they publish the testing code but not the data. To me, that would be the same as publishing the rules before a sports match.</p>\n\n<p>Yury, I think you (or someone) did already make this point in another forum on this comp.</p>",
      "rawMarkdown": "You make a good point but I think mistakes and controversies will be inevitable with a complex process like this. Running these competitions is a very tough job. So I believe the organizers do a great job and are always willing to listen to questions and sometimes make changes during a competition.  I do like the kaggle platorm and get a lot of personal benefit at no financial cost.\n\nMaybe they should always publish the metric code. I can't think of a disadvantage of that, if they publish the testing code but not the data. To me, that would be the same as publishing the rules before a sports match.\n\nYury, I think you (or someone) did already make this point in another forum on this comp.",
      "votes": null
    },
    {
      "id": "728788",
      "postDate": "01/25/2020 09:04:19",
      "content": "<p>Yes, Metric was a hot topic rising in many threads. I had to chat in private with Phil. Dieter also spend much time helping organizers to fix the metric. \nFor sure - metric is obliged to be shared. Especially if it’s only one line of code away from <code>nq_eval</code>. </p>\n\n<p>You’re right about Kaggle competitions in general. But still I think they need to learn from mistakes. This time the beginning was absolutely spoilt. And it’s not the first time - motivates to enter a competition at least a month before the end or even later. </p>",
      "rawMarkdown": "Yes, Metric was a hot topic rising in many threads. I had to chat in private with Phil. Dieter also spend much time helping organizers to fix the metric. \nFor sure - metric is obliged to be shared. Especially if it’s only one line of code away from `nq_eval`. \n\nYou’re right about Kaggle competitions in general. But still I think they need to learn from mistakes. This time the beginning was absolutely spoilt. And it’s not the first time - motivates to enter a competition at least a month before the end or even later.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 728088,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "01/24/2020 11:42:20",
      "content": "<p>Another key observation that wasn't really shared by organizers is that the NQ dev test contains more than one annotation per example, like test, when train only had one.  Then scoring depends on the number of annotations, see  part of the NQ dataset paper (<a href=\"https://storage.googleapis.com/pub-tools-public-publication-data/pdf/1f7b46b5378d757553d3e92ead36bda2e4254244.pdf\">https://storage.googleapis.com/pub-tools-public-publication-data/pdf/1f7b46b5378d757553d3e92ead36bda2e4254244.pdf</a>):</p>\n\n<blockquote>\n  <p>If at least2out of 5 annotators have given a non-null long answer on the example, then the system is required to output a non-null answer that is seen at least once in the 5 annotations; conversely if fewer than 2 annotators give a non-null long answer, the system is required to return NULL as its output</p>\n</blockquote>\n\n<p>Without this it is impossible to have a good CV - LB correlation IMHO.</p>\n\n<p>We found about this quite late, maybe two weeks before end.  I feel for all who didn't saw this.  Without proper CV then this becomes more of a lottery.</p>",
      "votes": null,
      "replies": [
        {
          "id": 728107,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "01/24/2020 12:08:30",
          "content": "<p>Aha! I missed the \"fewer than 2\". I tuned my models to a metric that returned null only if fewer than 1 (ie zero) annotators gave non-null output.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728157,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "01/24/2020 13:16:11",
          "content": "<p>True. That’s why we validated only with the dev set once the metric was fixed. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728309,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "01/24/2020 15:38:33",
          "content": "<p>I am still uncertain as to what the metric was. The <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation\">evaluation page</a> says:</p>\n\n<blockquote>\n  <p>There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.</p>\n</blockquote>\n\n<p>By saying it should be left blank if NO answer applies, they directly contradict the <code>fewer than 2</code> correctly quoted above from the original paper by <a href=\"/cpmpml\">@cpmpml</a> which they point us to on the same page.</p>\n\n<p>I went by the competition page, which clearly states <strong>none</strong>, meaning <code>fewer than 1</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728314,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "01/24/2020 15:43:58",
          "content": "<p>A better choice is to delve into <code>nq_eval</code> code. There it's clearly seen that an answer is considered correct if &gt;= 2 annotators marked it as correct. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728329,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "01/24/2020 16:05:47",
          "content": "<p>But that contradicts what is very clearly stated on the official competition <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation\">evaluation page</a>. If no code is provided, surely we must go by the competition description, which says only leave it blank if <code>no answer</code> applies.</p>\n\n<p><a href=\"/philculliton\">@philculliton</a> ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728347,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "01/24/2020 16:29:03",
          "content": "<p>Ok, I see. Indeed, yet another fail. </p>\n\n<p>At some point I actually ignored everything written in the competition description and digged into code and paper. </p>\n\n<p>Had enough troubles with paper/code controversy, let alone other surprises from organizers. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728370,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "01/24/2020 16:45:31",
          "content": "<p>You make a good point but I think mistakes and controversies will be inevitable with a complex process like this. Running these competitions is a very tough job. So I believe the organizers do a great job and are always willing to listen to questions and sometimes make changes during a competition.  I do like the kaggle platorm and get a lot of personal benefit at no financial cost.</p>\n\n<p>Maybe they should always publish the metric code. I can't think of a disadvantage of that, if they publish the testing code but not the data. To me, that would be the same as publishing the rules before a sports match.</p>\n\n<p>Yury, I think you (or someone) did already make this point in another forum on this comp.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728788,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "01/25/2020 09:04:19",
          "content": "<p>Yes, Metric was a hot topic rising in many threads. I had to chat in private with Phil. Dieter also spend much time helping organizers to fix the metric. \nFor sure - metric is obliged to be shared. Especially if it’s only one line of code away from <code>nq_eval</code>. </p>\n\n<p>You’re right about Kaggle competitions in general. But still I think they need to learn from mistakes. This time the beginning was absolutely spoilt. And it’s not the first time - motivates to enter a competition at least a month before the end or even later. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 728109,
      "author_name": "kenkrige",
      "author_url": "",
      "post_date": "01/24/2020 12:10:43",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> thanks for this. Sorry to see how painfully close you were to gold. That must be frustrating!</p>",
      "votes": null,
      "replies": [
        {
          "id": 728163,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "01/24/2020 13:20:10",
          "content": "<p>Thanks, Ken! Yeah, I also had a girl born in December. So it’s also stealing time from her :(</p>\n\n<p>But the good news is that Dmitry and Oleg are the real masters, cool experience working with them. So hopefully, Trimorph will be back soon :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728271,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "01/24/2020 14:51:39",
          "content": "<p>Congratulations and good luck with being a father. Big job. I just became a grandfather this week, but that won't be as much of a lifestyle change as being a father.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728278,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "01/24/2020 14:57:30",
          "content": "<p>Cool! Congrats as well!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 728160,
      "author_name": "kashnitsky",
      "author_url": "",
      "post_date": "01/24/2020 13:17:17",
      "content": "<p>Well, actually we also submitted once to Natural Questions but with one sub per week it’s not a good source of feedback :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 728205,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "01/24/2020 13:57:02",
          "content": "<p>Did you submit to the nq leaderboard: <a href=\"https://ai.google.com/research/NaturalQuestions/leaderboard\">https://ai.google.com/research/NaturalQuestions/leaderboard</a>? If you did, what's the score compared to the metric here?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728213,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "01/24/2020 14:01:57",
          "content": "<p>Yes. That was pretty straightforward technically. And I was surprised to realize that such a simple idea of having one more big test set came to me so late. Anyway, only 1 sun per week. </p>\n\n<p>The score is a bit too optimistic there due to automatic threshold tuning. So it was some 1.5-2 points higher. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728215,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "01/24/2020 14:06:01",
          "content": "<p>I’m there on that LB, yorko. 69.4/57.1 long/short. While locally it was 67.8/56.2 for that version of Bert-large WWM, squad-pretrained. 63 LB in our competition </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728226,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "01/24/2020 14:15:19",
          "content": "<p>That's great, thanks. I hope I get some time to adjust my code and submit as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728236,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "01/24/2020 14:19:40",
          "content": "<p>Yes, just follow the quick-start, it went smooth for me. And there you’re given 24 hour slot (2x P100 afair) - so unfortunately, crazy ensembles are more than welcome there. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "728076": "Already mentioned the crucial change w.r.t. to `nq_eval` in a [shared Notebook](https://www.kaggle.com/kashnitsky/solid-bert-joint-baseline-0-65-0-66) with our best single model (yes, BERT-joint again, on steroids), but to elaborate a bit, what worked ideally for us is the following:\n`nq_eval_yorko` is a small modification of the `nq_eval` script, uploaded to [my fork](https://github.com/Yorko/natural-questions) of the natural-questions repo. The key difference is just one line:\n<img src=\"https://habrastorage.org/webt/ar/12/g0/ar12g0cea9fnojk_ghhjs2wyyae.png\">\n\nSo that's exactly what stated on the Evaluation page (and Phil fixed the metric correctly):\n&gt; The metric in this competition diverges from the original metric in two key respects: 1) short and long answer formats do not receive separate scores, but are instead combined into a micro F1 score across both formats, and 2) this competition's metric does not use confidence scores to find an optimal threshold for predictions.\n\nThere in the repo you also see `nq-dev-all-annotations.jsonl.gz` - it's the same official dev set from [Natural Questions](https://ai.google.com/research/NaturalQuestions/download) but I left only annotations for ~100x faster metric computation (no need to read all these document tokens). \n\nWhen you have your `pred.json` file, evaluation is the following:\n- `python nq_eval_yorko.py --gold_path=nq-dev-all-annotations.jsonl.gz --predictions_path=pred.json` - thus you get Long and Short F1 scores and optimal thresholds\n- `python nq_eval_yorko.py --gold_path=nq-dev-all-annotations.jsonl.gz --predictions_path=pred.json --score_thres_long $LONG_THRES --score_thres_short $SHORT_THRES` - this will show only Long, Short, and All F1 scores. This \"All\" is our competition metric. \n\nFor sure, this can be refactored into just one script. And preserving threshold tuning is also nice (`nq_eval_yorko` only applies the provided thresholds) but to be honest, I'm in such a grief with our 13th place that good enough (my brother gave me a nickname \"Kaggle Bachelor\").",
    "728088": "Another key observation that wasn't really shared by organizers is that the NQ dev test contains more than one annotation per example, like test, when train only had one.  Then scoring depends on the number of annotations, see  part of the NQ dataset paper (https://storage.googleapis.com/pub-tools-public-publication-data/pdf/1f7b46b5378d757553d3e92ead36bda2e4254244.pdf):\n\n&gt; If at least2out of 5 annotators have given a non-null long answer on the example, then the system is required to output a non-null answer that is seen at least once in the 5 annotations; conversely if fewer than 2 annotators give a non-null long answer, the system is required to return NULL as its output\n\nWithout this it is impossible to have a good CV - LB correlation IMHO.\n\nWe found about this quite late, maybe two weeks before end.  I feel for all who didn't saw this.  Without proper CV then this becomes more of a lottery.",
    "728107": "Aha! I missed the \"fewer than 2\". I tuned my models to a metric that returned null only if fewer than 1 (ie zero) annotators gave non-null output.",
    "728109": "kashnitsky thanks for this. Sorry to see how painfully close you were to gold. That must be frustrating!",
    "728157": "True. That’s why we validated only with the dev set once the metric was fixed.",
    "728160": "Well, actually we also submitted once to Natural Questions but with one sub per week it’s not a good source of feedback :)",
    "728163": "Thanks, Ken! Yeah, I also had a girl born in December. So it’s also stealing time from her :(\n\nBut the good news is that Dmitry and Oleg are the real masters, cool experience working with them. So hopefully, Trimorph will be back soon :)",
    "728205": "Did you submit to the nq leaderboard: https://ai.google.com/research/NaturalQuestions/leaderboard? If you did, what's the score compared to the metric here?",
    "728213": "Yes. That was pretty straightforward technically. And I was surprised to realize that such a simple idea of having one more big test set came to me so late. Anyway, only 1 sun per week. \n\nThe score is a bit too optimistic there due to automatic threshold tuning. So it was some 1.5-2 points higher.",
    "728215": "I’m there on that LB, yorko. 69.4/57.1 long/short. While locally it was 67.8/56.2 for that version of Bert-large WWM, squad-pretrained. 63 LB in our competition",
    "728226": "That's great, thanks. I hope I get some time to adjust my code and submit as well.",
    "728236": "Yes, just follow the quick-start, it went smooth for me. And there you’re given 24 hour slot (2x P100 afair) - so unfortunately, crazy ensembles are more than welcome there.",
    "728271": "Congratulations and good luck with being a father. Big job. I just became a grandfather this week, but that won't be as much of a lifestyle change as being a father.",
    "728278": "Cool! Congrats as well!",
    "728309": "I am still uncertain as to what the metric was. The [evaluation page](https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation) says:\n&gt; There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.\n\nBy saying it should be left blank if NO answer applies, they directly contradict the `fewer than 2` correctly quoted above from the original paper by @cpmpml which they point us to on the same page.\n\nI went by the competition page, which clearly states **none**, meaning `fewer than 1`",
    "728314": "A better choice is to delve into `nq_eval` code. There it's clearly seen that an answer is considered correct if &gt;= 2 annotators marked it as correct.",
    "728329": "But that contradicts what is very clearly stated on the official competition [evaluation page](https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation). If no code is provided, surely we must go by the competition description, which says only leave it blank if `no answer` applies.\n\n@philculliton ?",
    "728347": "Ok, I see. Indeed, yet another fail. \n\nAt some point I actually ignored everything written in the competition description and digged into code and paper. \n\nHad enough troubles with paper/code controversy, let alone other surprises from organizers.",
    "728370": "You make a good point but I think mistakes and controversies will be inevitable with a complex process like this. Running these competitions is a very tough job. So I believe the organizers do a great job and are always willing to listen to questions and sometimes make changes during a competition.  I do like the kaggle platorm and get a lot of personal benefit at no financial cost.\n\nMaybe they should always publish the metric code. I can't think of a disadvantage of that, if they publish the testing code but not the data. To me, that would be the same as publishing the rules before a sports match.\n\nYury, I think you (or someone) did already make this point in another forum on this comp.",
    "728788": "Yes, Metric was a hot topic rising in many threads. I had to chat in private with Phil. Dieter also spend much time helping organizers to fix the metric. \nFor sure - metric is obliged to be shared. Especially if it’s only one line of code away from `nq_eval`. \n\nYou’re right about Kaggle competitions in general. But still I think they need to learn from mistakes. This time the beginning was absolutely spoilt. And it’s not the first time - motivates to enter a competition at least a month before the end or even later."
  },
  "source": "meta"
}