{
  "id": 24220,
  "title": "Confidence level",
  "url": "/competitions/outbrain-click-prediction/discussion/24220",
  "author_name": "",
  "post_date": "2016-10-09T15:39:50.553Z",
  "votes": null,
  "comment_count": 1,
  "views": 251,
  "content": "<p>Hi!</p>\n\n<p>Can someone explain to me what is the meaning of confidence levels in <code>documents_[entities|categories|topics].csv</code> files. I don't get it completely form the description provided on the data page. Especially in <code>documents_categories.csv</code> file.</p>\n\n<p>Thanks! :)</p>",
  "messages": [
    {
      "id": "138562",
      "postDate": "10/09/2016 15:39:50",
      "content": "<p>Hi!</p>\n\n<p>Can someone explain to me what is the meaning of confidence levels in <code>documents_[entities|categories|topics].csv</code> files. I don't get it completely form the description provided on the data page. Especially in <code>documents_categories.csv</code> file.</p>\n\n<p>Thanks! :)</p>",
      "rawMarkdown": "Hi!\r\n\r\nCan someone explain to me what is the meaning of confidence levels in `documents_[entities|categories|topics].csv` files. I don't get it completely form the description provided on the data page. Especially in `documents_categories.csv` file.\r\n\r\nThanks! :)",
      "votes": null
    },
    {
      "id": "138798",
      "postDate": "10/10/2016 22:24:27",
      "content": "<p>While actual algorithm is unknown, the number is to be interpreted as the amount of certainty OB has over its classification tag for a document. Higher numbers are better of course. Usually, you would expect the confidence scores to be a number between [0, 1], with sum of all confidence scores across all possible topics/entities/categories for a document to sum up to 1, but this is not the case here. I have seen such examples elsewhere in real world as well. </p>\n\n<p>The confidence score will need to be bucketed for feature extraction. </p>\n\n<p>Also, an interesting side note could be the accuracy of the content classification algorithms used by OB - that will also impact the prediction. But this out of scope. </p>",
      "rawMarkdown": "While actual algorithm is unknown, the number is to be interpreted as the amount of certainty OB has over its classification tag for a document. Higher numbers are better of course. Usually, you would expect the confidence scores to be a number between [0, 1], with sum of all confidence scores across all possible topics/entities/categories for a document to sum up to 1, but this is not the case here. I have seen such examples elsewhere in real world as well. \r\n\r\nThe confidence score will need to be bucketed for feature extraction. \r\n\r\nAlso, an interesting side note could be the accuracy of the content classification algorithms used by OB - that will also impact the prediction. But this out of scope.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 138798,
      "author_name": "ceeker",
      "author_url": "",
      "post_date": "10/10/2016 22:24:27",
      "content": "<p>While actual algorithm is unknown, the number is to be interpreted as the amount of certainty OB has over its classification tag for a document. Higher numbers are better of course. Usually, you would expect the confidence scores to be a number between [0, 1], with sum of all confidence scores across all possible topics/entities/categories for a document to sum up to 1, but this is not the case here. I have seen such examples elsewhere in real world as well. </p>\n\n<p>The confidence score will need to be bucketed for feature extraction. </p>\n\n<p>Also, an interesting side note could be the accuracy of the content classification algorithms used by OB - that will also impact the prediction. But this out of scope. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "138562": "Hi!\r\n\r\nCan someone explain to me what is the meaning of confidence levels in `documents_[entities|categories|topics].csv` files. I don't get it completely form the description provided on the data page. Especially in `documents_categories.csv` file.\r\n\r\nThanks! :)",
    "138798": "While actual algorithm is unknown, the number is to be interpreted as the amount of certainty OB has over its classification tag for a document. Higher numbers are better of course. Usually, you would expect the confidence scores to be a number between [0, 1], with sum of all confidence scores across all possible topics/entities/categories for a document to sum up to 1, but this is not the case here. I have seen such examples elsewhere in real world as well. \r\n\r\nThe confidence score will need to be bucketed for feature extraction. \r\n\r\nAlso, an interesting side note could be the accuracy of the content classification algorithms used by OB - that will also impact the prediction. But this out of scope."
  },
  "source": "meta"
}