{
  "id": 125750,
  "title": "Eval Metric - my guess in a kernel",
  "url": "/competitions/tensorflow2-question-answering/discussion/125750",
  "author_name": "Ken Krige",
  "post_date": "2020-01-13T10:07:02.152000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>There has a fair amount of forum discussion on the evaluation metric. <a href=\"https://www.kaggle.com/kenkrige/possible-evaluation-metric\">This kernel</a> is my guess at how it works. It is not a definitive answer as I am just as unclear as many other participants. However, it does seem to achieve a similar score on the <code>tiny-dev</code> set to what the same model achieves on the public test set on the LB. So it could be something close to the metric being used for the LB.</p>",
  "messages": [
    {
      "id": 717578,
      "postDate": "2020-01-13T10:07:02.153Z",
      "content": "<p>There has a fair amount of forum discussion on the evaluation metric. <a href=\"https://www.kaggle.com/kenkrige/possible-evaluation-metric\">This kernel</a> is my guess at how it works. It is not a definitive answer as I am just as unclear as many other participants. However, it does seem to achieve a similar score on the <code>tiny-dev</code> set to what the same model achieves on the public test set on the LB. So it could be something close to the metric being used for the LB.</p>",
      "rawMarkdown": "There has a fair amount of forum discussion on the evaluation metric. [This kernel](https://www.kaggle.com/kenkrige/possible-evaluation-metric) is my guess at how it works. It is not a definitive answer as I am just as unclear as many other participants. However, it does seem to achieve a similar score on the `tiny-dev` set to what the same model achieves on the public test set on the LB. So it could be something close to the metric being used for the LB.",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "717578": "There has a fair amount of forum discussion on the evaluation metric. [This kernel](https://www.kaggle.com/kenkrige/possible-evaluation-metric) is my guess at how it works. It is not a definitive answer as I am just as unclear as many other participants. However, it does seem to achieve a similar score on the `tiny-dev` set to what the same model achieves on the public test set on the LB. So it could be something close to the metric being used for the LB."
  }
}