{
  "id": 119280,
  "title": "A story about 2 submission files",
  "url": "/competitions/tensorflow2-question-answering/discussion/119280",
  "author_name": "",
  "post_date": "2019-11-27T21:30:00.259026700Z",
  "votes": 7,
  "comment_count": 14,
  "views": 0,
  "content": "<p>So I have 2 submission files. Indexes are fine, just like in the sample submission file. No empty strings, each <code>PredictionString</code> is a string like <code>x:y</code>, i.e. spans are predicted for both long and short answers. </p>\n\n<p>As for predictions, these two are 57% overlapping (<code>PredictionString</code> coincides for 394 rows out of 692). \nThe first submission file is worth 0.75 Public LB.\nThe second submission file is worth (you know what?) 0.00 Public LB.</p>\n\n<p>And so I ponder whether it's in principle realistic with any reasonable metric. And whether I need to make submissions at all or I need to wait until it's all fixed by organizers. </p>\n\n<p><a href=\"/philculliton\">@philculliton</a> 100% sure that the evaluation function is implemented correctly? How many more hours do we all spend before we see the <em>real</em> metric?</p>\n\n<p><img src=\"https://habrastorage.org/webt/6j/rg/cd/6jrgcdhn9sxiui6px39d6sasx5i.jpeg\"></p>",
  "messages": [
    {
      "id": "682796",
      "postDate": "11/27/2019 21:30:00",
      "content": "<p>So I have 2 submission files. Indexes are fine, just like in the sample submission file. No empty strings, each <code>PredictionString</code> is a string like <code>x:y</code>, i.e. spans are predicted for both long and short answers. </p>\n\n<p>As for predictions, these two are 57% overlapping (<code>PredictionString</code> coincides for 394 rows out of 692). \nThe first submission file is worth 0.75 Public LB.\nThe second submission file is worth (you know what?) 0.00 Public LB.</p>\n\n<p>And so I ponder whether it's in principle realistic with any reasonable metric. And whether I need to make submissions at all or I need to wait until it's all fixed by organizers. </p>\n\n<p><a href=\"/philculliton\">@philculliton</a> 100% sure that the evaluation function is implemented correctly? How many more hours do we all spend before we see the <em>real</em> metric?</p>\n\n<p><img src=\"https://habrastorage.org/webt/6j/rg/cd/6jrgcdhn9sxiui6px39d6sasx5i.jpeg\"></p>",
      "rawMarkdown": "So I have 2 submission files. Indexes are fine, just like in the sample submission file. No empty strings, each `PredictionString` is a string like `x:y`, i.e. spans are predicted for both long and short answers. \n\nAs for predictions, these two are 57% overlapping (`PredictionString` coincides for 394 rows out of 692). \nThe first submission file is worth 0.75 Public LB.\nThe second submission file is worth (you know what?) 0.00 Public LB.\n\nAnd so I ponder whether it's in principle realistic with any reasonable metric. And whether I need to make submissions at all or I need to wait until it's all fixed by organizers. \n\n@philculliton 100% sure that the evaluation function is implemented correctly? How many more hours do we all spend before we see the *real* metric?\n\n<img src=\"https://habrastorage.org/webt/6j/rg/cd/6jrgcdhn9sxiui6px39d6sasx5i.jpeg\">",
      "votes": null
    },
    {
      "id": "682812",
      "postDate": "11/27/2019 22:37:38",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> - thanks for your feedback. Please double check that your 0.00 Public LB submission file contains predictions. If you still think it’s correct, please upload the two submission files in question to a private dataset and share it with me, and I’ll take a look.</p>",
      "rawMarkdown": "kashnitsky - thanks for your feedback. Please double check that your 0.00 Public LB submission file contains predictions. If you still think it’s correct, please upload the two submission files in question to a private dataset and share it with me, and I’ll take a look.",
      "votes": null
    },
    {
      "id": "682851",
      "postDate": "11/27/2019 23:27:50",
      "content": "<p>Thanks, <a href=\"/philculliton\">@philculliton</a> I've shared a Kernel with you, its version 3 scores 0, but version 6 scores 0.75. No idea why it shows ('No public score') for version 6, in my submission history I see that it's 0.75. The overlap between their PredictionStrings is even higher - 482/692. <br>\nIf needed, tomorrow I can run two more notebooks from scratch to make it crystal clear, without versioning. \nPS. Sorry, I'm being a bit emotional, need to rest :) Still true that currently participants mostly struggle with the evaluation metric. </p>",
      "rawMarkdown": "Thanks, @philculliton I've shared a Kernel with you, its version 3 scores 0, but version 6 scores 0.75. No idea why it shows ('No public score') for version 6, in my submission history I see that it's 0.75. The overlap between their PredictionStrings is even higher - 482/692.  \nIf needed, tomorrow I can run two more notebooks from scratch to make it crystal clear, without versioning. \nPS. Sorry, I'm being a bit emotional, need to rest :) Still true that currently participants mostly struggle with the evaluation metric.",
      "votes": null
    },
    {
      "id": "682866",
      "postDate": "11/27/2019 23:42:58",
      "content": "<blockquote>\n  <p>I can run two more notebooks from scratch to make it crystal clear</p>\n</blockquote>\n\n<p>Not sure about it now. For some Kernels I got \"Notebook Threw Exception\" even though they were committed and submitted fine before (producing some non-zero LB score). </p>",
      "rawMarkdown": "&gt; I can run two more notebooks from scratch to make it crystal clear\n\nNot sure about it now. For some Kernels I got \"Notebook Threw Exception\" even though they were committed and submitted fine before (producing some non-zero LB score).",
      "votes": null
    },
    {
      "id": "682913",
      "postDate": "11/28/2019 00:24:10",
      "content": "<p>Thanks very much <a href=\"/kashnitsky\">@kashnitsky</a>! I appreciate your sharing the kernels and explaining your problem. The evaluation metric is definitely a little mind-bending, and I'm thinking through how to better explain it. I don't want people to feel like they're struggling.</p>\n\n<p>I'm looking at your kernels - thanks for your responding comments. There doesn't seem to be a problem with the metric code itself (that scores appropriately when I feed your submissions to it directly) so there may be something odd going on kernel-wise. I'm looking at it with our engineers.</p>",
      "rawMarkdown": "Thanks very much @kashnitsky! I appreciate your sharing the kernels and explaining your problem. The evaluation metric is definitely a little mind-bending, and I'm thinking through how to better explain it. I don't want people to feel like they're struggling.\n\nI'm looking at your kernels - thanks for your responding comments. There doesn't seem to be a problem with the metric code itself (that scores appropriately when I feed your submissions to it directly) so there may be something odd going on kernel-wise. I'm looking at it with our engineers.",
      "votes": null
    },
    {
      "id": "682966",
      "postDate": "11/28/2019 01:33:32",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> - I've left some feedback in one of the kernels you shared. Thanks again for your help in digging into this.</p>\n\n<p>For everyone else: the issue described here is not related to the metric. I confirmed that the metric is producing correct scores. I really appreciate everyone bringing potential issues to my attention!</p>",
      "rawMarkdown": "kashnitsky - I've left some feedback in one of the kernels you shared. Thanks again for your help in digging into this.\n\nFor everyone else: the issue described here is not related to the metric. I confirmed that the metric is producing correct scores. I really appreciate everyone bringing potential issues to my attention!",
      "votes": null
    },
    {
      "id": "683028",
      "postDate": "11/28/2019 02:52:36",
      "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a> , if the metric is so complicated that it \"is definitely a little mind-bending\", and even you need to \"think through how to better explain it\", maybe not try to explain it, and just release the evaluation code? The code explains itself. Is there any potential problem of doing that?</p>\n\n<p>I appreciate your patience looking into <a href=\"/kashnitsky\">@kashnitsky</a> 's kernels, but if everyone who questions the metric makes private kernels for you to manually examine, isn't that counter-productive? </p>\n\n<p>In the meanwhile, do you mind responding to my questions in <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118758\">this</a> thread? </p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi @philculliton , if the metric is so complicated that it \"is definitely a little mind-bending\", and even you need to \"think through how to better explain it\", maybe not try to explain it, and just release the evaluation code? The code explains itself. Is there any potential problem of doing that?\n\nI appreciate your patience looking into @kashnitsky 's kernels, but if everyone who questions the metric makes private kernels for you to manually examine, isn't that counter-productive? \n\nIn the meanwhile, do you mind responding to my questions in [this](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118758) thread? \n\nThanks",
      "votes": null
    },
    {
      "id": "683053",
      "postDate": "11/28/2019 03:22:26",
      "content": "<p>Hi <a href=\"/boliu0\">@boliu0</a> - thanks for your questions!</p>\n\n<blockquote>\n  <p>Hi <a href=\"/philculliton\">@philculliton</a> , if the metric is so complicated that it \"is definitely a little mind-bending\", and even you need to \"think through how to better explain it\", maybe not try to explain it, and just release the evaluation code? The code explains itself. Is there any potential problem of doing that?</p>\n</blockquote>\n\n<p>The C# code is not especially clearer than <code>nq_eval.py</code> - in fact a number of critical operations have been offloaded to places that are not clear from looking at the code. With the exceptions noted on the Evaluation page, the two do line up nicely, though. I'd suggest taking a look at the original Python metric for a readable picture of what's going on.</p>\n\n<p>I wouldn't say the metric is especially complicated on its own, but the <em>problem</em> handled by the metric is, which is why it's mind-bending: it's a simple metric with lots of caveats / constraints on what constitutes true positives, false positives, etc.</p>\n\n<blockquote>\n  <p>I appreciate your patience looking into <a href=\"/kashnitsky\">@kashnitsky</a> 's kernels, but if everyone who questions the metric makes private kernels for you to manually examine, isn't that counter-productive?</p>\n</blockquote>\n\n<p>It definitely would be! I looked into these kernels because I was concerned there might be a kernel bug happening. I don't intend to look at every metric question like this.</p>\n\n<blockquote>\n  <p>In the meanwhile, do you mind responding to my questions in this thread?</p>\n</blockquote>\n\n<p>I'll be responding to your questions ASAP.</p>",
      "rawMarkdown": "Hi @boliu0 - thanks for your questions!\n\n&gt; Hi @philculliton , if the metric is so complicated that it \"is definitely a little mind-bending\", and even you need to \"think through how to better explain it\", maybe not try to explain it, and just release the evaluation code? The code explains itself. Is there any potential problem of doing that?\n\nThe C# code is not especially clearer than `nq_eval.py` - in fact a number of critical operations have been offloaded to places that are not clear from looking at the code. With the exceptions noted on the Evaluation page, the two do line up nicely, though. I'd suggest taking a look at the original Python metric for a readable picture of what's going on.\n\nI wouldn't say the metric is especially complicated on its own, but the _problem_ handled by the metric is, which is why it's mind-bending: it's a simple metric with lots of caveats / constraints on what constitutes true positives, false positives, etc.\n\n&gt; I appreciate your patience looking into @kashnitsky 's kernels, but if everyone who questions the metric makes private kernels for you to manually examine, isn't that counter-productive?\n\nIt definitely would be! I looked into these kernels because I was concerned there might be a kernel bug happening. I don't intend to look at every metric question like this.\n\n&gt; In the meanwhile, do you mind responding to my questions in this thread?\n\nI'll be responding to your questions ASAP.",
      "votes": null
    },
    {
      "id": "683066",
      "postDate": "11/28/2019 03:50:41",
      "content": "<p>Thanks for the quick reply.</p>\n\n<p>I did look at the official <code>nq_eval.py</code> script, and fully understand its quirks by running it step by step on dev set. That's why I believe LB behaves differently. Other folks have reported similar experience. Anyways, look forward to your response in the other thread.</p>",
      "rawMarkdown": "Thanks for the quick reply.\n\nI did look at the official `nq_eval.py` script, and fully understand its quirks by running it step by step on dev set. That's why I believe LB behaves differently. Other folks have reported similar experience. Anyways, look forward to your response in the other thread.",
      "votes": null
    },
    {
      "id": "683168",
      "postDate": "11/28/2019 06:52:30",
      "content": "<p>Thanks for taking your time to clarify the metric, this is really important.</p>\n\n<blockquote>\n  <p>With the exceptions noted on the Evaluation page, the two do line up nicely, though. </p>\n</blockquote>\n\n<p>Maybe it's possible to check that by doing some modifications to <code>nq_eval.py</code> it achieves the same score as on LB? If yes, maybe it's possible to publish these modifications to <code>nq_eval</code> to bring extra clarity?\nIf it's not possible to check this, then we can't be sure that the quirks of actual scoring code are the same as in <code>nq_eval.py</code>. If publishing the code is not possible, and describing it seems also very hard and prone to missing details, maybe it's not too late to switch to <code>nq_eval</code> metric, taking an average of long and short F1s?</p>",
      "rawMarkdown": "Thanks for taking your time to clarify the metric, this is really important.\n\n&gt; With the exceptions noted on the Evaluation page, the two do line up nicely, though. \n\nMaybe it's possible to check that by doing some modifications to `nq_eval.py` it achieves the same score as on LB? If yes, maybe it's possible to publish these modifications to `nq_eval` to bring extra clarity?\nIf it's not possible to check this, then we can't be sure that the quirks of actual scoring code are the same as in `nq_eval.py`. If publishing the code is not possible, and describing it seems also very hard and prone to missing details, maybe it's not too late to switch to `nq_eval` metric, taking an average of long and short F1s?",
      "votes": null
    },
    {
      "id": "683358",
      "postDate": "11/28/2019 10:49:20",
      "content": "<p>To share what Phil told me: one of the problems occurred with <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118129\">this trick</a>. The script actually threw an Exception, and to submission file not found, it picked up sample submission. Therefore LB score 0. Otherwise, without the trick, it shall be \"Notebook Threw Exception\" in My Submissions. </p>\n\n<p>This, however, doesn't answer my original question yet (with 2 submission files) - I'm currently spending 2 more submissions to reproduce the exact problem. </p>",
      "rawMarkdown": "To share what Phil told me: one of the problems occurred with [this trick](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118129). The script actually threw an Exception, and to submission file not found, it picked up sample submission. Therefore LB score 0. Otherwise, without the trick, it shall be \"Notebook Threw Exception\" in My Submissions. \n\nThis, however, doesn't answer my original question yet (with 2 submission files) - I'm currently spending 2 more submissions to reproduce the exact problem.",
      "votes": null
    },
    {
      "id": "683550",
      "postDate": "11/28/2019 14:14:11",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> thanks for readiness to help. The initial question is still the same. I've shared 2 more Kernels with you (actually the same as before, but no mess anymore with versioning):\n - ####_ Ver. 3 - scores 0.00 Public LB\n - ####_ Ver. 6 -  scores 0.75 Public LB</p>\n\n<p>No tricks applied, both Kernels produce submission files different from a sample submission. No empty strings in any of the 2 submission files. 482 rows out of 692 (70%) coincide. </p>",
      "rawMarkdown": "philculliton thanks for readiness to help. The initial question is still the same. I've shared 2 more Kernels with you (actually the same as before, but no mess anymore with versioning):\n - ####_ Ver. 3 - scores 0.00 Public LB\n - ####_ Ver. 6 -  scores 0.75 Public LB\n\nNo tricks applied, both Kernels produce submission files different from a sample submission. No empty strings in any of the 2 submission files. 482 rows out of 692 (70%) coincide.",
      "votes": null
    },
    {
      "id": "684443",
      "postDate": "11/29/2019 18:55:22",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> share with you notebook with 0.00 score without execution exceptions, help please</p>",
      "rawMarkdown": "philculliton share with you notebook with 0.00 score without execution exceptions, help please",
      "votes": null
    },
    {
      "id": "685949",
      "postDate": "12/02/2019 14:54:16",
      "content": "<p>Hi <a href=\"/kashnitsky\">@kashnitsky</a> - thanks for sharing your notebooks! Ver. 3 is also throwing an error partway through and submitting the sample submission. Ver. 6 I'm checking with the back end devs.</p>",
      "rawMarkdown": "Hi @kashnitsky - thanks for sharing your notebooks! Ver. 3 is also throwing an error partway through and submitting the sample submission. Ver. 6 I'm checking with the back end devs.",
      "votes": null
    },
    {
      "id": "685991",
      "postDate": "12/02/2019 16:00:57",
      "content": "<p>Hi, <a href=\"/philculliton\">@philculliton</a>! Thanks! Considering that the diff between these two versions is only minimal, I can only conclude that there's <em>something</em> in the hidden private test set that makes fail the code that works fine with public test set. Thanks for checking this, now I actually see a possibility that metics/scoring are fine, and it's all about data and our code. </p>",
      "rawMarkdown": "Hi, @philculliton! Thanks! Considering that the diff between these two versions is only minimal, I can only conclude that there's *something* in the hidden private test set that makes fail the code that works fine with public test set. Thanks for checking this, now I actually see a possibility that metics/scoring are fine, and it's all about data and our code.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 682812,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "11/27/2019 22:37:38",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> - thanks for your feedback. Please double check that your 0.00 Public LB submission file contains predictions. If you still think it’s correct, please upload the two submission files in question to a private dataset and share it with me, and I’ll take a look.</p>",
      "votes": null,
      "replies": [
        {
          "id": 682851,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "11/27/2019 23:27:50",
          "content": "<p>Thanks, <a href=\"/philculliton\">@philculliton</a> I've shared a Kernel with you, its version 3 scores 0, but version 6 scores 0.75. No idea why it shows ('No public score') for version 6, in my submission history I see that it's 0.75. The overlap between their PredictionStrings is even higher - 482/692. <br>\nIf needed, tomorrow I can run two more notebooks from scratch to make it crystal clear, without versioning. \nPS. Sorry, I'm being a bit emotional, need to rest :) Still true that currently participants mostly struggle with the evaluation metric. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 682866,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "11/27/2019 23:42:58",
          "content": "<blockquote>\n  <p>I can run two more notebooks from scratch to make it crystal clear</p>\n</blockquote>\n\n<p>Not sure about it now. For some Kernels I got \"Notebook Threw Exception\" even though they were committed and submitted fine before (producing some non-zero LB score). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 682913,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "11/28/2019 00:24:10",
          "content": "<p>Thanks very much <a href=\"/kashnitsky\">@kashnitsky</a>! I appreciate your sharing the kernels and explaining your problem. The evaluation metric is definitely a little mind-bending, and I'm thinking through how to better explain it. I don't want people to feel like they're struggling.</p>\n\n<p>I'm looking at your kernels - thanks for your responding comments. There doesn't seem to be a problem with the metric code itself (that scores appropriately when I feed your submissions to it directly) so there may be something odd going on kernel-wise. I'm looking at it with our engineers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 682966,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "11/28/2019 01:33:32",
          "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> - I've left some feedback in one of the kernels you shared. Thanks again for your help in digging into this.</p>\n\n<p>For everyone else: the issue described here is not related to the metric. I confirmed that the metric is producing correct scores. I really appreciate everyone bringing potential issues to my attention!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 683028,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "11/28/2019 02:52:36",
          "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a> , if the metric is so complicated that it \"is definitely a little mind-bending\", and even you need to \"think through how to better explain it\", maybe not try to explain it, and just release the evaluation code? The code explains itself. Is there any potential problem of doing that?</p>\n\n<p>I appreciate your patience looking into <a href=\"/kashnitsky\">@kashnitsky</a> 's kernels, but if everyone who questions the metric makes private kernels for you to manually examine, isn't that counter-productive? </p>\n\n<p>In the meanwhile, do you mind responding to my questions in <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118758\">this</a> thread? </p>\n\n<p>Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 683053,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "11/28/2019 03:22:26",
          "content": "<p>Hi <a href=\"/boliu0\">@boliu0</a> - thanks for your questions!</p>\n\n<blockquote>\n  <p>Hi <a href=\"/philculliton\">@philculliton</a> , if the metric is so complicated that it \"is definitely a little mind-bending\", and even you need to \"think through how to better explain it\", maybe not try to explain it, and just release the evaluation code? The code explains itself. Is there any potential problem of doing that?</p>\n</blockquote>\n\n<p>The C# code is not especially clearer than <code>nq_eval.py</code> - in fact a number of critical operations have been offloaded to places that are not clear from looking at the code. With the exceptions noted on the Evaluation page, the two do line up nicely, though. I'd suggest taking a look at the original Python metric for a readable picture of what's going on.</p>\n\n<p>I wouldn't say the metric is especially complicated on its own, but the <em>problem</em> handled by the metric is, which is why it's mind-bending: it's a simple metric with lots of caveats / constraints on what constitutes true positives, false positives, etc.</p>\n\n<blockquote>\n  <p>I appreciate your patience looking into <a href=\"/kashnitsky\">@kashnitsky</a> 's kernels, but if everyone who questions the metric makes private kernels for you to manually examine, isn't that counter-productive?</p>\n</blockquote>\n\n<p>It definitely would be! I looked into these kernels because I was concerned there might be a kernel bug happening. I don't intend to look at every metric question like this.</p>\n\n<blockquote>\n  <p>In the meanwhile, do you mind responding to my questions in this thread?</p>\n</blockquote>\n\n<p>I'll be responding to your questions ASAP.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 683066,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "11/28/2019 03:50:41",
          "content": "<p>Thanks for the quick reply.</p>\n\n<p>I did look at the official <code>nq_eval.py</code> script, and fully understand its quirks by running it step by step on dev set. That's why I believe LB behaves differently. Other folks have reported similar experience. Anyways, look forward to your response in the other thread.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 683168,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "11/28/2019 06:52:30",
          "content": "<p>Thanks for taking your time to clarify the metric, this is really important.</p>\n\n<blockquote>\n  <p>With the exceptions noted on the Evaluation page, the two do line up nicely, though. </p>\n</blockquote>\n\n<p>Maybe it's possible to check that by doing some modifications to <code>nq_eval.py</code> it achieves the same score as on LB? If yes, maybe it's possible to publish these modifications to <code>nq_eval</code> to bring extra clarity?\nIf it's not possible to check this, then we can't be sure that the quirks of actual scoring code are the same as in <code>nq_eval.py</code>. If publishing the code is not possible, and describing it seems also very hard and prone to missing details, maybe it's not too late to switch to <code>nq_eval</code> metric, taking an average of long and short F1s?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 683358,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "11/28/2019 10:49:20",
          "content": "<p>To share what Phil told me: one of the problems occurred with <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118129\">this trick</a>. The script actually threw an Exception, and to submission file not found, it picked up sample submission. Therefore LB score 0. Otherwise, without the trick, it shall be \"Notebook Threw Exception\" in My Submissions. </p>\n\n<p>This, however, doesn't answer my original question yet (with 2 submission files) - I'm currently spending 2 more submissions to reproduce the exact problem. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 683550,
      "author_name": "kashnitsky",
      "author_url": "",
      "post_date": "11/28/2019 14:14:11",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> thanks for readiness to help. The initial question is still the same. I've shared 2 more Kernels with you (actually the same as before, but no mess anymore with versioning):\n - ####_ Ver. 3 - scores 0.00 Public LB\n - ####_ Ver. 6 -  scores 0.75 Public LB</p>\n\n<p>No tricks applied, both Kernels produce submission files different from a sample submission. No empty strings in any of the 2 submission files. 482 rows out of 692 (70%) coincide. </p>",
      "votes": null,
      "replies": [
        {
          "id": 685949,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "12/02/2019 14:54:16",
          "content": "<p>Hi <a href=\"/kashnitsky\">@kashnitsky</a> - thanks for sharing your notebooks! Ver. 3 is also throwing an error partway through and submitting the sample submission. Ver. 6 I'm checking with the back end devs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 685991,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/02/2019 16:00:57",
          "content": "<p>Hi, <a href=\"/philculliton\">@philculliton</a>! Thanks! Considering that the diff between these two versions is only minimal, I can only conclude that there's <em>something</em> in the hidden private test set that makes fail the code that works fine with public test set. Thanks for checking this, now I actually see a possibility that metics/scoring are fine, and it's all about data and our code. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 684443,
      "author_name": "mizinovmv",
      "author_url": "",
      "post_date": "11/29/2019 18:55:22",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> share with you notebook with 0.00 score without execution exceptions, help please</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "682796": "So I have 2 submission files. Indexes are fine, just like in the sample submission file. No empty strings, each `PredictionString` is a string like `x:y`, i.e. spans are predicted for both long and short answers. \n\nAs for predictions, these two are 57% overlapping (`PredictionString` coincides for 394 rows out of 692). \nThe first submission file is worth 0.75 Public LB.\nThe second submission file is worth (you know what?) 0.00 Public LB.\n\nAnd so I ponder whether it's in principle realistic with any reasonable metric. And whether I need to make submissions at all or I need to wait until it's all fixed by organizers. \n\n@philculliton 100% sure that the evaluation function is implemented correctly? How many more hours do we all spend before we see the *real* metric?\n\n<img src=\"https://habrastorage.org/webt/6j/rg/cd/6jrgcdhn9sxiui6px39d6sasx5i.jpeg\">",
    "682812": "kashnitsky - thanks for your feedback. Please double check that your 0.00 Public LB submission file contains predictions. If you still think it’s correct, please upload the two submission files in question to a private dataset and share it with me, and I’ll take a look.",
    "682851": "Thanks, @philculliton I've shared a Kernel with you, its version 3 scores 0, but version 6 scores 0.75. No idea why it shows ('No public score') for version 6, in my submission history I see that it's 0.75. The overlap between their PredictionStrings is even higher - 482/692.  \nIf needed, tomorrow I can run two more notebooks from scratch to make it crystal clear, without versioning. \nPS. Sorry, I'm being a bit emotional, need to rest :) Still true that currently participants mostly struggle with the evaluation metric.",
    "682866": "&gt; I can run two more notebooks from scratch to make it crystal clear\n\nNot sure about it now. For some Kernels I got \"Notebook Threw Exception\" even though they were committed and submitted fine before (producing some non-zero LB score).",
    "682913": "Thanks very much @kashnitsky! I appreciate your sharing the kernels and explaining your problem. The evaluation metric is definitely a little mind-bending, and I'm thinking through how to better explain it. I don't want people to feel like they're struggling.\n\nI'm looking at your kernels - thanks for your responding comments. There doesn't seem to be a problem with the metric code itself (that scores appropriately when I feed your submissions to it directly) so there may be something odd going on kernel-wise. I'm looking at it with our engineers.",
    "682966": "kashnitsky - I've left some feedback in one of the kernels you shared. Thanks again for your help in digging into this.\n\nFor everyone else: the issue described here is not related to the metric. I confirmed that the metric is producing correct scores. I really appreciate everyone bringing potential issues to my attention!",
    "683028": "Hi @philculliton , if the metric is so complicated that it \"is definitely a little mind-bending\", and even you need to \"think through how to better explain it\", maybe not try to explain it, and just release the evaluation code? The code explains itself. Is there any potential problem of doing that?\n\nI appreciate your patience looking into @kashnitsky 's kernels, but if everyone who questions the metric makes private kernels for you to manually examine, isn't that counter-productive? \n\nIn the meanwhile, do you mind responding to my questions in [this](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118758) thread? \n\nThanks",
    "683053": "Hi @boliu0 - thanks for your questions!\n\n&gt; Hi @philculliton , if the metric is so complicated that it \"is definitely a little mind-bending\", and even you need to \"think through how to better explain it\", maybe not try to explain it, and just release the evaluation code? The code explains itself. Is there any potential problem of doing that?\n\nThe C# code is not especially clearer than `nq_eval.py` - in fact a number of critical operations have been offloaded to places that are not clear from looking at the code. With the exceptions noted on the Evaluation page, the two do line up nicely, though. I'd suggest taking a look at the original Python metric for a readable picture of what's going on.\n\nI wouldn't say the metric is especially complicated on its own, but the _problem_ handled by the metric is, which is why it's mind-bending: it's a simple metric with lots of caveats / constraints on what constitutes true positives, false positives, etc.\n\n&gt; I appreciate your patience looking into @kashnitsky 's kernels, but if everyone who questions the metric makes private kernels for you to manually examine, isn't that counter-productive?\n\nIt definitely would be! I looked into these kernels because I was concerned there might be a kernel bug happening. I don't intend to look at every metric question like this.\n\n&gt; In the meanwhile, do you mind responding to my questions in this thread?\n\nI'll be responding to your questions ASAP.",
    "683066": "Thanks for the quick reply.\n\nI did look at the official `nq_eval.py` script, and fully understand its quirks by running it step by step on dev set. That's why I believe LB behaves differently. Other folks have reported similar experience. Anyways, look forward to your response in the other thread.",
    "683168": "Thanks for taking your time to clarify the metric, this is really important.\n\n&gt; With the exceptions noted on the Evaluation page, the two do line up nicely, though. \n\nMaybe it's possible to check that by doing some modifications to `nq_eval.py` it achieves the same score as on LB? If yes, maybe it's possible to publish these modifications to `nq_eval` to bring extra clarity?\nIf it's not possible to check this, then we can't be sure that the quirks of actual scoring code are the same as in `nq_eval.py`. If publishing the code is not possible, and describing it seems also very hard and prone to missing details, maybe it's not too late to switch to `nq_eval` metric, taking an average of long and short F1s?",
    "683358": "To share what Phil told me: one of the problems occurred with [this trick](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118129). The script actually threw an Exception, and to submission file not found, it picked up sample submission. Therefore LB score 0. Otherwise, without the trick, it shall be \"Notebook Threw Exception\" in My Submissions. \n\nThis, however, doesn't answer my original question yet (with 2 submission files) - I'm currently spending 2 more submissions to reproduce the exact problem.",
    "683550": "philculliton thanks for readiness to help. The initial question is still the same. I've shared 2 more Kernels with you (actually the same as before, but no mess anymore with versioning):\n - ####_ Ver. 3 - scores 0.00 Public LB\n - ####_ Ver. 6 -  scores 0.75 Public LB\n\nNo tricks applied, both Kernels produce submission files different from a sample submission. No empty strings in any of the 2 submission files. 482 rows out of 692 (70%) coincide.",
    "684443": "philculliton share with you notebook with 0.00 score without execution exceptions, help please",
    "685949": "Hi @kashnitsky - thanks for sharing your notebooks! Ver. 3 is also throwing an error partway through and submitting the sample submission. Ver. 6 I'm checking with the back end devs.",
    "685991": "Hi, @philculliton! Thanks! Considering that the diff between these two versions is only minimal, I can only conclude that there's *something* in the hidden private test set that makes fail the code that works fine with public test set. Thanks for checking this, now I actually see a possibility that metics/scoring are fine, and it's all about data and our code."
  },
  "source": "meta"
}