{
  "id": 121286,
  "title": "Please provide implementation of the competition metric",
  "url": "/competitions/tensorflow2-question-answering/discussion/121286",
  "author_name": "Dmytro Danevskyi",
  "post_date": "2019-12-12T08:38:30.094000",
  "votes": 45,
  "comment_count": 10,
  "views": 0,
  "content": "<p>The current state of this competition looks like it's a really bad kind of joke. Now it doesn't look just like someone's oversight, but rather a plain disregard.</p>\n\n<p>I wonder, how many hours have been wasted just because the metric was not released? Probably now it's time to finally release the metric and stop this madness? </p>",
  "messages": [
    {
      "id": 693311,
      "postDate": "2019-12-12T08:38:30.093Z",
      "content": "<p>The current state of this competition looks like it's a really bad kind of joke. Now it doesn't look just like someone's oversight, but rather a plain disregard.</p>\n\n<p>I wonder, how many hours have been wasted just because the metric was not released? Probably now it's time to finally release the metric and stop this madness? </p>",
      "rawMarkdown": "The current state of this competition looks like it's a really bad kind of joke. Now it doesn't look just like someone's oversight, but rather a plain disregard.\n\nI wonder, how many hours have been wasted just because the metric was not released? Probably now it's time to finally release the metric and stop this madness? ",
      "votes": 45
    },
    {
      "id": 693513,
      "postDate": "2019-12-12T13:15:44.873Z",
      "content": "<p>Second this. <a href=\"/philculliton\">@philculliton</a> I first tried to help you fixing stuff (by sharing/reproducing my solutions), but I don’t think I’ll go on with this “gradient evaluation metric ascent”. The competition is spoiled. It can be fixed only if you share the metric implementation, even if it’s C# and spread over several files. </p>",
      "rawMarkdown": "Second this. @philculliton I first tried to help you fixing stuff (by sharing/reproducing my solutions), but I don’t think I’ll go on with this “gradient evaluation metric ascent”. The competition is spoiled. It can be fixed only if you share the metric implementation, even if it’s C# and spread over several files. ",
      "votes": 10
    },
    {
      "id": 693648,
      "postDate": "2019-12-12T16:06:52.170Z",
      "content": "<p>Hi all! Thanks for the feedback. Thanks also for the ping, Yury.</p>\n\n<p>I'm sorry that not releasing the C# metric is still causing so much confusion. I thought between the description update and the fix that we were on the right path. Releasing the C# metric is a problem - it is spread over several files, yes, but also some of the metric's decisions regarding the TP/FP/FN calculations take place outside the metric code. This metric is not complex in itself, but the conditions around what constitutes a TP or FP is certainly so.</p>\n\n<p>I think it's very clear that this is not a sufficient situation for anyone. Also <code>nq_eval.py</code> seemed like a useful proxy but apparently is not clear at all. I think the most efficient way forward would be for me to provide simplified Python code - essentially to do some testing and bring <a href=\"/christofhenkel\">@christofhenkel</a>'s (assuming that's all right with <a href=\"/christofhenkel\">@christofhenkel</a>) code in line with the actuality of the C# metric if it does, in fact, diverge.</p>\n\n<p>Please do let me know if this seems insufficient. Sorry again for the continuing confusion, and thanks again for the feedback.</p>",
      "rawMarkdown": "Hi all! Thanks for the feedback. Thanks also for the ping, Yury.\n\nI'm sorry that not releasing the C# metric is still causing so much confusion. I thought between the description update and the fix that we were on the right path. Releasing the C# metric is a problem - it is spread over several files, yes, but also some of the metric's decisions regarding the TP/FP/FN calculations take place outside the metric code. This metric is not complex in itself, but the conditions around what constitutes a TP or FP is certainly so.\n\nI think it's very clear that this is not a sufficient situation for anyone. Also `nq_eval.py` seemed like a useful proxy but apparently is not clear at all. I think the most efficient way forward would be for me to provide simplified Python code - essentially to do some testing and bring @christofhenkel's (assuming that's all right with @christofhenkel) code in line with the actuality of the C# metric if it does, in fact, diverge.\n\nPlease do let me know if this seems insufficient. Sorry again for the continuing confusion, and thanks again for the feedback.",
      "replies": [
        {
          "id": 693657,
          "postDate": "2019-12-12T16:30:29.340Z",
          "content": "<p>If your C# code is so convoluted, then it’s not a surprise that things are easily screwed. Sorry. </p>\n\n<p>Yes, providing a simplified Python implementation will be very nice. </p>",
          "rawMarkdown": "If your C# code is so convoluted, then it’s not a surprise that things are easily screwed. Sorry. \n\nYes, providing a simplified Python implementation will be very nice. \n\n",
          "votes": 4
        },
        {
          "id": 693677,
          "postDate": "2019-12-12T17:00:15.930Z",
          "content": "<p>After playing around, I’m starting to get convinced that the new score is being implemented correctly.\n(still exploring)</p>\n\n<p>My first version eval script gives a lower score of ~0.4 using baseline.</p>",
          "rawMarkdown": "After playing around, I’m starting to get convinced that the new score is being implemented correctly.\n(still exploring)\n\nMy first version eval script gives a lower score of ~0.4 using baseline."
        },
        {
          "id": 693840,
          "postDate": "2019-12-12T20:51:57.520Z",
          "content": "<ul>\n<li>providing a simplified Python implementation will be very nice. </li>\n<li><strong>providing an implementation should be mandatory for any competition.</strong></li>\n</ul>",
          "rawMarkdown": "- providing a simplified Python implementation will be very nice. \n- **providing an implementation should be mandatory for any competition.**",
          "votes": 2
        },
        {
          "id": 693865,
          "postDate": "2019-12-12T21:41:57.077Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> sure thats alright with me. I think <a href=\"/kentaronakanishi\">@kentaronakanishi</a> implemented a version here: <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061#691933\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061#691933</a></p>",
          "rawMarkdown": "@philculliton sure thats alright with me. I think @kentaronakanishi implemented a version here: https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061#691933",
          "votes": 2
        }
      ]
    },
    {
      "id": 701411,
      "postDate": "2019-12-23T13:19:54.997Z",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> Hello! How are you? I wonder if there are any updates on your side?</p>",
      "rawMarkdown": "@philculliton Hello! How are you? I wonder if there are any updates on your side?",
      "votes": 1,
      "replies": [
        {
          "id": 704560,
          "postDate": "2019-12-27T16:08:05.597Z",
          "content": "<p>You can check my kernel </p>\n\n<p><a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/123345\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/123345</a></p>\n\n<p>It's not official, and I didn't test it thoroughly, but probably helpful</p>",
          "rawMarkdown": "You can check my kernel \n\n[https://www.kaggle.com/c/tensorflow2-question-answering/discussion/123345](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/123345)\n \nIt's not official, and I didn't test it thoroughly, but probably helpful",
          "votes": 2
        },
        {
          "id": 704677,
          "postDate": "2019-12-27T20:27:56.600Z",
          "content": "<p>Thanks a lot!</p>",
          "rawMarkdown": "Thanks a lot!"
        }
      ]
    },
    {
      "id": 698344,
      "postDate": "2019-12-19T05:17:41.327Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 693513,
      "author_name": "Yury Kashnitsky",
      "author_url": "",
      "post_date": "2019-12-12T13:15:44.873000",
      "content": "<p>Second this. <a href=\"/philculliton\">@philculliton</a> I first tried to help you fixing stuff (by sharing/reproducing my solutions), but I don’t think I’ll go on with this “gradient evaluation metric ascent”. The competition is spoiled. It can be fixed only if you share the metric implementation, even if it’s C# and spread over several files. </p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 693648,
      "author_name": "Phil Culliton",
      "author_url": "",
      "post_date": "2019-12-12T16:06:52.170000",
      "content": "<p>Hi all! Thanks for the feedback. Thanks also for the ping, Yury.</p>\n\n<p>I'm sorry that not releasing the C# metric is still causing so much confusion. I thought between the description update and the fix that we were on the right path. Releasing the C# metric is a problem - it is spread over several files, yes, but also some of the metric's decisions regarding the TP/FP/FN calculations take place outside the metric code. This metric is not complex in itself, but the conditions around what constitutes a TP or FP is certainly so.</p>\n\n<p>I think it's very clear that this is not a sufficient situation for anyone. Also <code>nq_eval.py</code> seemed like a useful proxy but apparently is not clear at all. I think the most efficient way forward would be for me to provide simplified Python code - essentially to do some testing and bring <a href=\"/christofhenkel\">@christofhenkel</a>'s (assuming that's all right with <a href=\"/christofhenkel\">@christofhenkel</a>) code in line with the actuality of the C# metric if it does, in fact, diverge.</p>\n\n<p>Please do let me know if this seems insufficient. Sorry again for the continuing confusion, and thanks again for the feedback.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 693657,
          "author_name": "Yury Kashnitsky",
          "author_url": "",
          "post_date": "2019-12-12T16:30:29.340000",
          "content": "<p>If your C# code is so convoluted, then it’s not a surprise that things are easily screwed. Sorry. </p>\n\n<p>Yes, providing a simplified Python implementation will be very nice. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 693677,
          "author_name": "Shane",
          "author_url": "",
          "post_date": "2019-12-12T17:00:15.930000",
          "content": "<p>After playing around, I’m starting to get convinced that the new score is being implemented correctly.\n(still exploring)</p>\n\n<p>My first version eval script gives a lower score of ~0.4 using baseline.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 693840,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2019-12-12T20:51:57.520000",
          "content": "<ul>\n<li>providing a simplified Python implementation will be very nice. </li>\n<li><strong>providing an implementation should be mandatory for any competition.</strong></li>\n</ul>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 693865,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2019-12-12T21:41:57.077000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> sure thats alright with me. I think <a href=\"/kentaronakanishi\">@kentaronakanishi</a> implemented a version here: <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061#691933\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061#691933</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 701411,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2019-12-23T13:19:54.997000",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> Hello! How are you? I wonder if there are any updates on your side?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 704560,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2019-12-27T16:08:05.597000",
          "content": "<p>You can check my kernel </p>\n\n<p><a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/123345\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/123345</a></p>\n\n<p>It's not official, and I didn't test it thoroughly, but probably helpful</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 704677,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-12-27T20:27:56.600000",
          "content": "<p>Thanks a lot!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 698344,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-19T05:17:41.327000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "693311": "The current state of this competition looks like it's a really bad kind of joke. Now it doesn't look just like someone's oversight, but rather a plain disregard.\n\nI wonder, how many hours have been wasted just because the metric was not released? Probably now it's time to finally release the metric and stop this madness? ",
    "693513": "Second this. @philculliton I first tried to help you fixing stuff (by sharing/reproducing my solutions), but I don’t think I’ll go on with this “gradient evaluation metric ascent”. The competition is spoiled. It can be fixed only if you share the metric implementation, even if it’s C# and spread over several files. ",
    "693648": "Hi all! Thanks for the feedback. Thanks also for the ping, Yury.\n\nI'm sorry that not releasing the C# metric is still causing so much confusion. I thought between the description update and the fix that we were on the right path. Releasing the C# metric is a problem - it is spread over several files, yes, but also some of the metric's decisions regarding the TP/FP/FN calculations take place outside the metric code. This metric is not complex in itself, but the conditions around what constitutes a TP or FP is certainly so.\n\nI think it's very clear that this is not a sufficient situation for anyone. Also `nq_eval.py` seemed like a useful proxy but apparently is not clear at all. I think the most efficient way forward would be for me to provide simplified Python code - essentially to do some testing and bring @christofhenkel's (assuming that's all right with @christofhenkel) code in line with the actuality of the C# metric if it does, in fact, diverge.\n\nPlease do let me know if this seems insufficient. Sorry again for the continuing confusion, and thanks again for the feedback.",
    "701411": "@philculliton Hello! How are you? I wonder if there are any updates on your side?",
    "698344": ""
  }
}