{
  "id": 121080,
  "title": "Wrong metric waste too much time.Is current metric correct now? ",
  "url": "/competitions/tensorflow2-question-answering/discussion/121080",
  "author_name": "sakuranew",
  "post_date": "2019-12-11T01:17:02.647000",
  "votes": 10,
  "comment_count": 13,
  "views": 0,
  "content": "<p>The goal of optimization was wrong. The time of train model was wasted</p>",
  "messages": [
    {
      "id": 692159,
      "postDate": "2019-12-11T01:17:02.647Z",
      "content": "<p>The goal of optimization was wrong. The time of train model was wasted</p>",
      "rawMarkdown": "The goal of optimization was wrong. The time of train model was wasted",
      "votes": 10
    },
    {
      "id": 694271,
      "postDate": "2019-12-13T11:24:08.053Z",
      "content": "<p>Update:\nThe LB match my offline score well,so the metric online is reliable.</p>",
      "rawMarkdown": "Update:\nThe LB match my offline score well,so the metric online is reliable.\n",
      "votes": 3,
      "replies": [
        {
          "id": 694273,
          "postDate": "2019-12-13T11:27:47.083Z",
          "content": "<p>Can you tell a bit more how you calculate the offline score? Which val set for example</p>",
          "rawMarkdown": "Can you tell a bit more how you calculate the offline score? Which val set for example\n",
          "votes": 1
        },
        {
          "id": 694279,
          "postDate": "2019-12-13T11:37:30.877Z",
          "content": "<p>Just like <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120030#686536\">last discussion</a>\nLB will be higher because one prediction has many labels</p>",
          "rawMarkdown": "Just like [last discussion](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120030#686536)\nLB will be higher because one prediction has many labels",
          "votes": 2
        },
        {
          "id": 694319,
          "postDate": "2019-12-13T12:25:23.053Z",
          "content": "<p>but only many labels for short answers not for long answers, or?</p>",
          "rawMarkdown": "but only many labels for short answers not for long answers, or?"
        },
        {
          "id": 694322,
          "postDate": "2019-12-13T12:27:45.730Z",
          "content": "<p>no, long answers also have many</p>",
          "rawMarkdown": "no, long answers also have many",
          "votes": 1
        },
        {
          "id": 694333,
          "postDate": "2019-12-13T12:45:46.570Z",
          "content": "<p>Could you lead me to any resource (kaggle discussion or external) which is the reason for your statement? I could not find anything on kaggle or NQ repo indicating there are multiple long answers. Many thanks</p>",
          "rawMarkdown": "Could you lead me to any resource (kaggle discussion or external) which is the reason for your statement? I could not find anything on kaggle or NQ repo indicating there are multiple long answers. Many thanks"
        },
        {
          "id": 694334,
          "postDate": "2019-12-13T12:48:54.213Z",
          "content": "<p>&gt; There may be up to five labels for long answers, and more for short</p>\n\n<p>In the Evaluation page</p>",
          "rawMarkdown": "&gt; There may be up to five labels for long answers, and more for short\n\nIn the Evaluation page",
          "votes": 3
        },
        {
          "id": 694428,
          "postDate": "2019-12-13T15:27:15.763Z",
          "content": "<p>Yes, and Phil commented that you get a credit for an NQ example if you predict at least one of the answers correctly. </p>",
          "rawMarkdown": "Yes, and Phil commented that you get a credit for an NQ example if you predict at least one of the answers correctly. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 692296,
      "postDate": "2019-12-11T05:39:22.613Z",
      "content": "<p>Honestly, I don't know. I implemented what I thought is the new metric and get 0.42 cv vs 0.50 LB. That seems off.</p>",
      "rawMarkdown": "Honestly, I don't know. I implemented what I thought is the new metric and get 0.42 cv vs 0.50 LB. That seems off.",
      "votes": 3,
      "replies": [
        {
          "id": 693031,
          "postDate": "2019-12-12T01:05:59.647Z",
          "content": "<p>On my side, I see a similar story. My long answer f1 (higher ~45%) is closer to public LB... (However full set local CV 38%-&gt; public LB 51%) When using the previous score (thanks to your implementation) it gives a local CV of 73%, public LB 75%. I'm not sure about whether they change the  distribution of the eval data-set or maybe my score is not implemented correctly.</p>",
          "rawMarkdown": "On my side, I see a similar story. My long answer f1 (higher ~45%) is closer to public LB... (However full set local CV 38%-&gt; public LB 51%) When using the previous score (thanks to your implementation) it gives a local CV of 73%, public LB 75%. I'm not sure about whether they change the  distribution of the eval data-set or maybe my score is not implemented correctly.",
          "votes": 1
        }
      ]
    },
    {
      "id": 692302,
      "postDate": "2019-12-11T05:51:30.170Z",
      "content": "<p>Gradient descent to get to the correct metric is very slow ......</p>",
      "rawMarkdown": "Gradient descent to get to the correct metric is very slow ......",
      "votes": 1,
      "replies": [
        {
          "id": 692303,
          "postDate": "2019-12-11T05:55:01.007Z",
          "content": "<p>Especially since you don't see the code you want to fix.</p>",
          "rawMarkdown": "Especially since you don't see the code you want to fix.",
          "votes": 2
        }
      ]
    },
    {
      "id": 692254,
      "postDate": "2019-12-11T04:33:29.430Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 694271,
      "author_name": "sakuranew",
      "author_url": "",
      "post_date": "2019-12-13T11:24:08.053000",
      "content": "<p>Update:\nThe LB match my offline score well,so the metric online is reliable.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 694273,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2019-12-13T11:27:47.083000",
          "content": "<p>Can you tell a bit more how you calculate the offline score? Which val set for example</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 694279,
          "author_name": "sakuranew",
          "author_url": "",
          "post_date": "2019-12-13T11:37:30.877000",
          "content": "<p>Just like <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120030#686536\">last discussion</a>\nLB will be higher because one prediction has many labels</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 694319,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2019-12-13T12:25:23.053000",
          "content": "<p>but only many labels for short answers not for long answers, or?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 694322,
          "author_name": "sakuranew",
          "author_url": "",
          "post_date": "2019-12-13T12:27:45.730000",
          "content": "<p>no, long answers also have many</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 694333,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2019-12-13T12:45:46.570000",
          "content": "<p>Could you lead me to any resource (kaggle discussion or external) which is the reason for your statement? I could not find anything on kaggle or NQ repo indicating there are multiple long answers. Many thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 694334,
          "author_name": "sakuranew",
          "author_url": "",
          "post_date": "2019-12-13T12:48:54.213000",
          "content": "<p>&gt; There may be up to five labels for long answers, and more for short</p>\n\n<p>In the Evaluation page</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 694428,
          "author_name": "Yury Kashnitsky",
          "author_url": "",
          "post_date": "2019-12-13T15:27:15.763000",
          "content": "<p>Yes, and Phil commented that you get a credit for an NQ example if you predict at least one of the answers correctly. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 692296,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2019-12-11T05:39:22.613000",
      "content": "<p>Honestly, I don't know. I implemented what I thought is the new metric and get 0.42 cv vs 0.50 LB. That seems off.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 693031,
          "author_name": "Shane",
          "author_url": "",
          "post_date": "2019-12-12T01:05:59.647000",
          "content": "<p>On my side, I see a similar story. My long answer f1 (higher ~45%) is closer to public LB... (However full set local CV 38%-&gt; public LB 51%) When using the previous score (thanks to your implementation) it gives a local CV of 73%, public LB 75%. I'm not sure about whether they change the  distribution of the eval data-set or maybe my score is not implemented correctly.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 692302,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2019-12-11T05:51:30.170000",
      "content": "<p>Gradient descent to get to the correct metric is very slow ......</p>",
      "votes": 1,
      "replies": [
        {
          "id": 692303,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2019-12-11T05:55:01.007000",
          "content": "<p>Especially since you don't see the code you want to fix.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 692254,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-11T04:33:29.430000",
      "content": "",
      "votes": -2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "692159": "The goal of optimization was wrong. The time of train model was wasted",
    "694271": "Update:\nThe LB match my offline score well,so the metric online is reliable.\n",
    "692296": "Honestly, I don't know. I implemented what I thought is the new metric and get 0.42 cv vs 0.50 LB. That seems off.",
    "692302": "Gradient descent to get to the correct metric is very slow ......",
    "692254": ""
  }
}