{
  "id": 51848,
  "title": "Submission decimal or 0/1",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51848",
  "author_name": "",
  "post_date": "2018-03-13T16:06:55.463961200Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi,\nI was wondering...\nI trained a model and predicted is_attributed, if i submitted it, i got a result of 0.92~~\nIf I took the prediction and did If(Is_attributed &gt;= 0.5 , 1, 0) and submitted it, the score decreased to about 0.75~~</p>\n\n<p>Why is that?\nHow is the score calculated if not the same way as taking the value and deciding if it's 1 or 0 based if it's &gt;=0.5?</p>\n\n<p>Any ideas?\nThanks!</p>",
  "messages": [
    {
      "id": "295409",
      "postDate": "03/13/2018 16:06:55",
      "content": "<p>Hi,\nI was wondering...\nI trained a model and predicted is_attributed, if i submitted it, i got a result of 0.92~~\nIf I took the prediction and did If(Is_attributed &gt;= 0.5 , 1, 0) and submitted it, the score decreased to about 0.75~~</p>\n\n<p>Why is that?\nHow is the score calculated if not the same way as taking the value and deciding if it's 1 or 0 based if it's &gt;=0.5?</p>\n\n<p>Any ideas?\nThanks!</p>",
      "rawMarkdown": "Hi,\nI was wondering...\nI trained a model and predicted is_attributed, if i submitted it, i got a result of 0.92~~\nIf I took the prediction and did If(Is_attributed &gt;= 0.5 , 1, 0) and submitted it, the score decreased to about 0.75~~\n\nWhy is that?\nHow is the score calculated if not the same way as taking the value and deciding if it's 1 or 0 based if it's &gt;=0.5?\n\nAny ideas?\nThanks!",
      "votes": null
    },
    {
      "id": "295544",
      "postDate": "03/13/2018 20:33:45",
      "content": "<p>It's based on your ROC score, which takes into account prediction probability. If it's a 1, then you get a higher score for having a prediction of .99 vs .51 even though both of these will round to 1. </p>",
      "rawMarkdown": "It's based on your ROC score, which takes into account prediction probability. If it's a 1, then you get a higher score for having a prediction of .99 vs .51 even though both of these will round to 1.",
      "votes": null
    },
    {
      "id": "295573",
      "postDate": "03/13/2018 21:27:42",
      "content": "<p>Thanks!!</p>",
      "rawMarkdown": "Thanks!!",
      "votes": null
    },
    {
      "id": "295692",
      "postDate": "03/14/2018 02:48:40",
      "content": "<p>Because in calculating AUC we sort the data first. \nPlease see this answer : <a href=\"https://stats.stackexchange.com/questions/145566/how-to-calculate-area-under-the-curve-auc-or-the-c-statistic-by-hand/146259#146259\">https://stats.stackexchange.com/questions/145566/how-to-calculate-area-under-the-curve-auc-or-the-c-statistic-by-hand/146259#146259</a></p>\n\n<p>Moreover we don't need to care about the actual probabilities, because its the \"rank\" that matters. Simply tresholding the actual value might destroy the rank data!.\nTake care when doing assembling or blending.  Simple probabilities average will not works well, its the \"rank\" that matter!</p>",
      "rawMarkdown": "Because in calculating AUC we sort the data first. \nPlease see this answer : https://stats.stackexchange.com/questions/145566/how-to-calculate-area-under-the-curve-auc-or-the-c-statistic-by-hand/146259#146259\n\nMoreover we don't need to care about the actual probabilities, because its the \"rank\" that matters. Simply tresholding the actual value might destroy the rank data!.\nTake care when doing assembling or blending.  Simple probabilities average will not works well, its the \"rank\" that matter!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 295544,
      "author_name": "bendodson6",
      "author_url": "",
      "post_date": "03/13/2018 20:33:45",
      "content": "<p>It's based on your ROC score, which takes into account prediction probability. If it's a 1, then you get a higher score for having a prediction of .99 vs .51 even though both of these will round to 1. </p>",
      "votes": null,
      "replies": [
        {
          "id": 295573,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "03/13/2018 21:27:42",
          "content": "<p>Thanks!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 295692,
      "author_name": "muhammadalfiansyah",
      "author_url": "",
      "post_date": "03/14/2018 02:48:40",
      "content": "<p>Because in calculating AUC we sort the data first. \nPlease see this answer : <a href=\"https://stats.stackexchange.com/questions/145566/how-to-calculate-area-under-the-curve-auc-or-the-c-statistic-by-hand/146259#146259\">https://stats.stackexchange.com/questions/145566/how-to-calculate-area-under-the-curve-auc-or-the-c-statistic-by-hand/146259#146259</a></p>\n\n<p>Moreover we don't need to care about the actual probabilities, because its the \"rank\" that matters. Simply tresholding the actual value might destroy the rank data!.\nTake care when doing assembling or blending.  Simple probabilities average will not works well, its the \"rank\" that matter!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "295409": "Hi,\nI was wondering...\nI trained a model and predicted is_attributed, if i submitted it, i got a result of 0.92~~\nIf I took the prediction and did If(Is_attributed &gt;= 0.5 , 1, 0) and submitted it, the score decreased to about 0.75~~\n\nWhy is that?\nHow is the score calculated if not the same way as taking the value and deciding if it's 1 or 0 based if it's &gt;=0.5?\n\nAny ideas?\nThanks!",
    "295544": "It's based on your ROC score, which takes into account prediction probability. If it's a 1, then you get a higher score for having a prediction of .99 vs .51 even though both of these will round to 1.",
    "295573": "Thanks!!",
    "295692": "Because in calculating AUC we sort the data first. \nPlease see this answer : https://stats.stackexchange.com/questions/145566/how-to-calculate-area-under-the-curve-auc-or-the-c-statistic-by-hand/146259#146259\n\nMoreover we don't need to care about the actual probabilities, because its the \"rank\" that matters. Simply tresholding the actual value might destroy the rank data!.\nTake care when doing assembling or blending.  Simple probabilities average will not works well, its the \"rank\" that matter!"
  },
  "source": "meta"
}