{
  "id": 240083,
  "title": "Why not submit 0 and 1 as targets?",
  "url": "/competitions/seti-breakthrough-listen/discussion/240083",
  "author_name": "Arkaitz",
  "post_date": "2021-05-18T14:10:54.928000",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have been observing some notebooks on the SETI Breakthrough Listen competition, and I see that every submission uses probabilities as output target. Why isn't 0 and 1 used as output targets like in the train set?  Is there any rule I missed which states that the output must be the probability of being a needle? Maybe I do not understand the area under the ROC curve score well enough.</p>",
  "messages": [
    {
      "id": 1313356,
      "postDate": "2021-05-18T14:16:11.080Z",
      "content": "<p>The beauty of the ROC score is it doesn't require you to pick a \"cutoff\" to use to discriminate between 0 and 1. If there is a cutoff point in your predictions where all the normals (true 0s) are below, and all the abnormals (true 1s) are above it, then you will score 1 -  you don't need to choose that point yourself. If your model just randomly picks numbers, then you will score 0.5 If your model is perfectly wrong, you will score 0.</p>\n<p>So you are ALWAYS better off not dichotomising, unless you know the perfect threshold for your model on the testing sets (in which case you will be no worse off, just the same).</p>",
      "rawMarkdown": "The beauty of the ROC score is it doesn't require you to pick a \"cutoff\" to use to discriminate between 0 and 1. If there is a cutoff point in your predictions where all the normals (true 0s) are below, and all the abnormals (true 1s) are above it, then you will score 1 -  you don't need to choose that point yourself. If your model just randomly picks numbers, then you will score 0.5 If your model is perfectly wrong, you will score 0.\n\nSo you are ALWAYS better off not dichotomising, unless you know the perfect threshold for your model on the testing sets (in which case you will be no worse off, just the same).",
      "votes": 10,
      "replies": [
        {
          "id": 1314463,
          "postDate": "2021-05-19T07:22:58.360Z",
          "content": "<p>All right, thank you very much!</p>",
          "rawMarkdown": "All right, thank you very much!",
          "votes": 1
        },
        {
          "id": 1359120,
          "postDate": "2021-06-21T05:10:40.043Z",
          "content": "<p>Hey James, thanks for your comment! Unfortunately, I have a lot of leftover questions after reading it:</p>\n<p>What is meant by \"cutoff\" points? These are just the values in the <code>target</code> column, right?<br>\nIn other words, aren't there supposed to be multiple \"cutoff\" points in our <code>submission.csv</code> file since each row in our <code>submission.csv</code> file has a \"cutoff\" point in the <code>target</code> column?</p>\n<p>If there is only one \"cutoff\" point in range <code>(0,1)</code>, isn't it impossible to <em>not</em> get a score of 1? Yet, you say it's possible to get a score of <code>0.5</code> which makes sense or else this competition wouldn't be necessary.</p>\n<p>How can a true <code>0</code> be <em>above</em> a cutoff point located in the range <code>(0,1)</code> when <code>0</code> is the greatest lower bound of <code>(0,1)</code>? Similar question for a true <code>1</code>… <code>1</code> is the least upper bound of <code>(0,1)</code>.</p>\n<p>Isn't a perfectly wrong model perfectly right after reversing all classifications?</p>",
          "rawMarkdown": "Hey James, thanks for your comment! Unfortunately, I have a lot of leftover questions after reading it:\n\nWhat is meant by \"cutoff\" points? These are just the values in the `target` column, right?\nIn other words, aren't there supposed to be multiple \"cutoff\" points in our `submission.csv` file since each row in our `submission.csv` file has a \"cutoff\" point in the `target` column?\n\nIf there is only one \"cutoff\" point in range `(0,1)`, isn't it impossible to *not* get a score of 1? Yet, you say it's possible to get a score of `0.5` which makes sense or else this competition wouldn't be necessary.\n\nHow can a true `0` be *above* a cutoff point located in the range `(0,1)` when `0` is the greatest lower bound of `(0,1)`? Similar question for a true `1`... `1` is the least upper bound of `(0,1)`.\n\nIsn't a perfectly wrong model perfectly right after reversing all classifications?"
        },
        {
          "id": 1367732,
          "postDate": "2021-06-28T04:03:53.327Z",
          "content": "<p>Never mind, <a href=\"https://www.youtube.com/watch?v=4jRBRDbJemM&amp;ab_channel=StatQuestwithJoshStarmer\" target=\"_blank\">this statquest video</a> cleared things up for me.</p>",
          "rawMarkdown": "Never mind, [this statquest video](https://www.youtube.com/watch?v=4jRBRDbJemM&ab_channel=StatQuestwithJoshStarmer) cleared things up for me."
        }
      ]
    },
    {
      "id": 1313331,
      "postDate": "2021-05-18T14:10:54.930Z",
      "content": "<p>I have been observing some notebooks on the SETI Breakthrough Listen competition, and I see that every submission uses probabilities as output target. Why isn't 0 and 1 used as output targets like in the train set?  Is there any rule I missed which states that the output must be the probability of being a needle? Maybe I do not understand the area under the ROC curve score well enough.</p>",
      "rawMarkdown": "I have been observing some notebooks on the SETI Breakthrough Listen competition, and I see that every submission uses probabilities as output target. Why isn't 0 and 1 used as output targets like in the train set?  Is there any rule I missed which states that the output must be the probability of being a needle? Maybe I do not understand the area under the ROC curve score well enough.",
      "votes": 4
    },
    {
      "id": 1313353,
      "postDate": "2021-05-18T14:15:54.527Z",
      "content": "<p>there has big gap in quantity of pos and neg,AUC performance better than precision</p>",
      "rawMarkdown": "there has big gap in quantity of pos and neg,AUC performance better than precision"
    }
  ],
  "comments": [
    {
      "id": 1313356,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2021-05-18T14:16:11.080000",
      "content": "<p>The beauty of the ROC score is it doesn't require you to pick a \"cutoff\" to use to discriminate between 0 and 1. If there is a cutoff point in your predictions where all the normals (true 0s) are below, and all the abnormals (true 1s) are above it, then you will score 1 -  you don't need to choose that point yourself. If your model just randomly picks numbers, then you will score 0.5 If your model is perfectly wrong, you will score 0.</p>\n<p>So you are ALWAYS better off not dichotomising, unless you know the perfect threshold for your model on the testing sets (in which case you will be no worse off, just the same).</p>",
      "votes": 10,
      "replies": [
        {
          "id": 1314463,
          "author_name": "Arkaitz",
          "author_url": "",
          "post_date": "2021-05-19T07:22:58.360000",
          "content": "<p>All right, thank you very much!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1359120,
          "author_name": "Erik Kaufman",
          "author_url": "",
          "post_date": "2021-06-21T05:10:40.043000",
          "content": "<p>Hey James, thanks for your comment! Unfortunately, I have a lot of leftover questions after reading it:</p>\n<p>What is meant by \"cutoff\" points? These are just the values in the <code>target</code> column, right?<br>\nIn other words, aren't there supposed to be multiple \"cutoff\" points in our <code>submission.csv</code> file since each row in our <code>submission.csv</code> file has a \"cutoff\" point in the <code>target</code> column?</p>\n<p>If there is only one \"cutoff\" point in range <code>(0,1)</code>, isn't it impossible to <em>not</em> get a score of 1? Yet, you say it's possible to get a score of <code>0.5</code> which makes sense or else this competition wouldn't be necessary.</p>\n<p>How can a true <code>0</code> be <em>above</em> a cutoff point located in the range <code>(0,1)</code> when <code>0</code> is the greatest lower bound of <code>(0,1)</code>? Similar question for a true <code>1</code>… <code>1</code> is the least upper bound of <code>(0,1)</code>.</p>\n<p>Isn't a perfectly wrong model perfectly right after reversing all classifications?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1367732,
          "author_name": "Erik Kaufman",
          "author_url": "",
          "post_date": "2021-06-28T04:03:53.327000",
          "content": "<p>Never mind, <a href=\"https://www.youtube.com/watch?v=4jRBRDbJemM&amp;ab_channel=StatQuestwithJoshStarmer\" target=\"_blank\">this statquest video</a> cleared things up for me.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1313353,
      "author_name": "FatLiuyun",
      "author_url": "",
      "post_date": "2021-05-18T14:15:54.527000",
      "content": "<p>there has big gap in quantity of pos and neg,AUC performance better than precision</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1313356": "The beauty of the ROC score is it doesn't require you to pick a \"cutoff\" to use to discriminate between 0 and 1. If there is a cutoff point in your predictions where all the normals (true 0s) are below, and all the abnormals (true 1s) are above it, then you will score 1 -  you don't need to choose that point yourself. If your model just randomly picks numbers, then you will score 0.5 If your model is perfectly wrong, you will score 0.\n\nSo you are ALWAYS better off not dichotomising, unless you know the perfect threshold for your model on the testing sets (in which case you will be no worse off, just the same).",
    "1313331": "I have been observing some notebooks on the SETI Breakthrough Listen competition, and I see that every submission uses probabilities as output target. Why isn't 0 and 1 used as output targets like in the train set?  Is there any rule I missed which states that the output must be the probability of being a needle? Maybe I do not understand the area under the ROC curve score well enough.",
    "1313353": "there has big gap in quantity of pos and neg,AUC performance better than precision"
  }
}