{
  "id": 44198,
  "title": "How evaluation is made",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/44198",
  "author_name": "",
  "post_date": "2017-11-25T04:19:20.579131Z",
  "votes": -3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hai \nCan anyone help me to know what is meant by the points in the leaderboard. </p>",
  "messages": [
    {
      "id": "248160",
      "postDate": "11/25/2017 04:19:20",
      "content": "<p>Hai \nCan anyone help me to know what is meant by the points in the leaderboard. </p>",
      "rawMarkdown": "Hai \nCan anyone help me to know what is meant by the points in the leaderboard.",
      "votes": null
    },
    {
      "id": "248902",
      "postDate": "11/27/2017 08:30:52",
      "content": "<p><a href=\"https://www.kaggle.com/c/passenger-screening-algorithm-challenge#evaluation\">https://www.kaggle.com/c/passenger-screening-algorithm-challenge#evaluation</a></p>",
      "rawMarkdown": "https://www.kaggle.com/c/passenger-screening-algorithm-challenge#evaluation",
      "votes": null
    },
    {
      "id": "249191",
      "postDate": "11/27/2017 22:55:50",
      "content": "<p>It's basically the log loss; unlike accuracy, which is either <em>correct</em> or <em>incorrect</em>, the score penalizes wrong answers exponentially.  If you guess that there's a 100% chance of a threat present and there isn't, your score goes through the roof; it's exponentially worse than saying there's a 75% chance of a threat present when there isn't.  Lower scores are better.  Anything under 0.01 is arguably perfect.</p>\n\n<p>The numbers themselves are the average of the 100 scans in the test set (or validation set; frankly I'm not sure which is more accurate to say.)  So if your predictions score 0.1 on one test scan and 0.3 on another, your leaderboard score would be 0.2 (ignoring the other 98 scans in the scoring set.)</p>\n\n<p>I hope that helps :)</p>",
      "rawMarkdown": "It's basically the log loss; unlike accuracy, which is either *correct* or *incorrect*, the score penalizes wrong answers exponentially.  If you guess that there's a 100% chance of a threat present and there isn't, your score goes through the roof; it's exponentially worse than saying there's a 75% chance of a threat present when there isn't.  Lower scores are better.  Anything under 0.01 is arguably perfect.\n\nThe numbers themselves are the average of the 100 scans in the test set (or validation set; frankly I'm not sure which is more accurate to say.)  So if your predictions score 0.1 on one test scan and 0.3 on another, your leaderboard score would be 0.2 (ignoring the other 98 scans in the scoring set.)\n\nI hope that helps :)",
      "votes": null
    },
    {
      "id": "249201",
      "postDate": "11/27/2017 23:31:16",
      "content": "<p>under 0.01 is far from perfect, You can be under 0.01 and still have very confident false positives/negatives in your submission. We know we have some at 99% confidence not sure if they are miss labeled in the test set or our model just can't see. We found this out by simply thresholding our very confident predictions to 0 and 1, and our lb score got worse not better.</p>",
      "rawMarkdown": "under 0.01 is far from perfect, You can be under 0.01 and still have very confident false positives/negatives in your submission. We know we have some at 99% confidence not sure if they are miss labeled in the test set or our model just can't see. We found this out by simply thresholding our very confident predictions to 0 and 1, and our lb score got worse not better.",
      "votes": null
    },
    {
      "id": "249255",
      "postDate": "11/28/2017 02:21:49",
      "content": "<p>Well, say there's a threat and you predict a 0.5 (50%) chance of a threat in that zone.  Your score for that area is 0.010158698821276728.  If you predict a 0.51 chance, your score is 0.0099555248448511945.  Hence my statement; I may be forgetting to divide by 17 though... amusingly enough, arithmetic is not my strong suit, haha.</p>\n\n<p>Your point's taken though.  Like I said, arguably :)</p>",
      "rawMarkdown": "Well, say there's a threat and you predict a 0.5 (50%) chance of a threat in that zone.  Your score for that area is 0.010158698821276728.  If you predict a 0.51 chance, your score is 0.0099555248448511945.  Hence my statement; I may be forgetting to divide by 17 though... amusingly enough, arithmetic is not my strong suit, haha.\n\nYour point's taken though.  Like I said, arguably :)",
      "votes": null
    },
    {
      "id": "249271",
      "postDate": "11/28/2017 03:46:01",
      "content": "<p>Point taken as well, my 2 cents to kaggle is that in stage 2, it will be every easy for them to identify potentially miss-labeled samples in the stage 2 data by simply looking for zones that are miss-classified by a majority of the top teams at a high error rate, and manually reviewing those labels. What they choose to do or what their policy is, I don't know, but it should be trivial to Identify the labeling errors. I can see where some teams that benefits from the label errors gets upset, if changes are made to the leaderboard post stage 2 start, because there will be teams that benefit from the label errors.</p>",
      "rawMarkdown": "Point taken as well, my 2 cents to kaggle is that in stage 2, it will be every easy for them to identify potentially miss-labeled samples in the stage 2 data by simply looking for zones that are miss-classified by a majority of the top teams at a high error rate, and manually reviewing those labels. What they choose to do or what their policy is, I don't know, but it should be trivial to Identify the labeling errors. I can see where some teams that benefits from the label errors gets upset, if changes are made to the leaderboard post stage 2 start, because there will be teams that benefit from the label errors.",
      "votes": null
    },
    {
      "id": "249286",
      "postDate": "11/28/2017 03:54:44",
      "content": "<p>Each prediction should actually be weighted as 1/1700 since you have 17 zones per passenger x 100 passengers in stage 1 test set.</p>\n\n<blockquote>\n  <p><strong>Murray wrote</strong></p>\n  \n  <blockquote>\n    <p>Well, say there's a threat and you predict a 0.5 (50%) chance of a threat in that zone.  Your score for that area is 0.010158698821276728.  If you predict a 0.51 chance, your score is 0.0099555248448511945.  Hence my statement; I may be forgetting to divide by 17 though... amusingly enough, arithmetic is not my strong suit, haha.</p>\n  </blockquote>\n  \n  <p>Your point's taken though.  Like I said, arguably :)</p>\n</blockquote>",
      "rawMarkdown": "Each prediction should actually be weighted as 1/1700 since you have 17 zones per passenger x 100 passengers in stage 1 test set.\n\n&gt; **Murray wrote**\n&gt; \n&gt; &gt; Well, say there's a threat and you predict a 0.5 (50%) chance of a threat in that zone.  Your score for that area is 0.010158698821276728.  If you predict a 0.51 chance, your score is 0.0099555248448511945.  Hence my statement; I may be forgetting to divide by 17 though... amusingly enough, arithmetic is not my strong suit, haha.\n&gt; \n&gt; Your point's taken though.  Like I said, arguably :)",
      "votes": null
    },
    {
      "id": "249300",
      "postDate": "11/28/2017 05:33:47",
      "content": "<p>@DavidGbodiOdaibo</p>\n\n<p>I've been doing something similar; rounding my current submissions to the nearest class, doing log-loss against those fake 'labels' and then compare it to the LB score of the same submission.  If the fake log-loss is lower than the LB, you know you have some misclassifications.  So far, I've never had a perfect classification yet, its probably one big mislabel or a few small ones, not sure.</p>\n\n<p>As far as hedging against mislabel issues, you get 2 submissions for this competition.  What I'm doing is assuming perfect labels for one of them, and the other will limit how confident any zone can be assuming an equal percentage of mislabels.  I feel like this is the safest strategy.</p>",
      "rawMarkdown": "DavidGbodiOdaibo\n\nI've been doing something similar; rounding my current submissions to the nearest class, doing log-loss against those fake 'labels' and then compare it to the LB score of the same submission.  If the fake log-loss is lower than the LB, you know you have some misclassifications.  So far, I've never had a perfect classification yet, its probably one big mislabel or a few small ones, not sure.\n\nAs far as hedging against mislabel issues, you get 2 submissions for this competition.  What I'm doing is assuming perfect labels for one of them, and the other will limit how confident any zone can be assuming an equal percentage of mislabels.  I feel like this is the safest strategy.",
      "votes": null
    },
    {
      "id": "253831",
      "postDate": "12/05/2017 18:04:06",
      "content": "<p>do we need to submit the log loss score for the test set in addition to the predicted probability?</p>",
      "rawMarkdown": "do we need to submit the log loss score for the test set in addition to the predicted probability?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 248902,
      "author_name": "jduffy",
      "author_url": "",
      "post_date": "11/27/2017 08:30:52",
      "content": "<p><a href=\"https://www.kaggle.com/c/passenger-screening-algorithm-challenge#evaluation\">https://www.kaggle.com/c/passenger-screening-algorithm-challenge#evaluation</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 249191,
      "author_name": "mmiron",
      "author_url": "",
      "post_date": "11/27/2017 22:55:50",
      "content": "<p>It's basically the log loss; unlike accuracy, which is either <em>correct</em> or <em>incorrect</em>, the score penalizes wrong answers exponentially.  If you guess that there's a 100% chance of a threat present and there isn't, your score goes through the roof; it's exponentially worse than saying there's a 75% chance of a threat present when there isn't.  Lower scores are better.  Anything under 0.01 is arguably perfect.</p>\n\n<p>The numbers themselves are the average of the 100 scans in the test set (or validation set; frankly I'm not sure which is more accurate to say.)  So if your predictions score 0.1 on one test scan and 0.3 on another, your leaderboard score would be 0.2 (ignoring the other 98 scans in the scoring set.)</p>\n\n<p>I hope that helps :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 249201,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "11/27/2017 23:31:16",
      "content": "<p>under 0.01 is far from perfect, You can be under 0.01 and still have very confident false positives/negatives in your submission. We know we have some at 99% confidence not sure if they are miss labeled in the test set or our model just can't see. We found this out by simply thresholding our very confident predictions to 0 and 1, and our lb score got worse not better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 249255,
          "author_name": "mmiron",
          "author_url": "",
          "post_date": "11/28/2017 02:21:49",
          "content": "<p>Well, say there's a threat and you predict a 0.5 (50%) chance of a threat in that zone.  Your score for that area is 0.010158698821276728.  If you predict a 0.51 chance, your score is 0.0099555248448511945.  Hence my statement; I may be forgetting to divide by 17 though... amusingly enough, arithmetic is not my strong suit, haha.</p>\n\n<p>Your point's taken though.  Like I said, arguably :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 249271,
          "author_name": "godaibo",
          "author_url": "",
          "post_date": "11/28/2017 03:46:01",
          "content": "<p>Point taken as well, my 2 cents to kaggle is that in stage 2, it will be every easy for them to identify potentially miss-labeled samples in the stage 2 data by simply looking for zones that are miss-classified by a majority of the top teams at a high error rate, and manually reviewing those labels. What they choose to do or what their policy is, I don't know, but it should be trivial to Identify the labeling errors. I can see where some teams that benefits from the label errors gets upset, if changes are made to the leaderboard post stage 2 start, because there will be teams that benefit from the label errors.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 249286,
          "author_name": "jamesrequa",
          "author_url": "",
          "post_date": "11/28/2017 03:54:44",
          "content": "<p>Each prediction should actually be weighted as 1/1700 since you have 17 zones per passenger x 100 passengers in stage 1 test set.</p>\n\n<blockquote>\n  <p><strong>Murray wrote</strong></p>\n  \n  <blockquote>\n    <p>Well, say there's a threat and you predict a 0.5 (50%) chance of a threat in that zone.  Your score for that area is 0.010158698821276728.  If you predict a 0.51 chance, your score is 0.0099555248448511945.  Hence my statement; I may be forgetting to divide by 17 though... amusingly enough, arithmetic is not my strong suit, haha.</p>\n  </blockquote>\n  \n  <p>Your point's taken though.  Like I said, arguably :)</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 249300,
          "author_name": "hackerpoet",
          "author_url": "",
          "post_date": "11/28/2017 05:33:47",
          "content": "<p>@DavidGbodiOdaibo</p>\n\n<p>I've been doing something similar; rounding my current submissions to the nearest class, doing log-loss against those fake 'labels' and then compare it to the LB score of the same submission.  If the fake log-loss is lower than the LB, you know you have some misclassifications.  So far, I've never had a perfect classification yet, its probably one big mislabel or a few small ones, not sure.</p>\n\n<p>As far as hedging against mislabel issues, you get 2 submissions for this competition.  What I'm doing is assuming perfect labels for one of them, and the other will limit how confident any zone can be assuming an equal percentage of mislabels.  I feel like this is the safest strategy.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 253831,
      "author_name": "w9wang2",
      "author_url": "",
      "post_date": "12/05/2017 18:04:06",
      "content": "<p>do we need to submit the log loss score for the test set in addition to the predicted probability?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "248160": "Hai \nCan anyone help me to know what is meant by the points in the leaderboard.",
    "248902": "https://www.kaggle.com/c/passenger-screening-algorithm-challenge#evaluation",
    "249191": "It's basically the log loss; unlike accuracy, which is either *correct* or *incorrect*, the score penalizes wrong answers exponentially.  If you guess that there's a 100% chance of a threat present and there isn't, your score goes through the roof; it's exponentially worse than saying there's a 75% chance of a threat present when there isn't.  Lower scores are better.  Anything under 0.01 is arguably perfect.\n\nThe numbers themselves are the average of the 100 scans in the test set (or validation set; frankly I'm not sure which is more accurate to say.)  So if your predictions score 0.1 on one test scan and 0.3 on another, your leaderboard score would be 0.2 (ignoring the other 98 scans in the scoring set.)\n\nI hope that helps :)",
    "249201": "under 0.01 is far from perfect, You can be under 0.01 and still have very confident false positives/negatives in your submission. We know we have some at 99% confidence not sure if they are miss labeled in the test set or our model just can't see. We found this out by simply thresholding our very confident predictions to 0 and 1, and our lb score got worse not better.",
    "249255": "Well, say there's a threat and you predict a 0.5 (50%) chance of a threat in that zone.  Your score for that area is 0.010158698821276728.  If you predict a 0.51 chance, your score is 0.0099555248448511945.  Hence my statement; I may be forgetting to divide by 17 though... amusingly enough, arithmetic is not my strong suit, haha.\n\nYour point's taken though.  Like I said, arguably :)",
    "249271": "Point taken as well, my 2 cents to kaggle is that in stage 2, it will be every easy for them to identify potentially miss-labeled samples in the stage 2 data by simply looking for zones that are miss-classified by a majority of the top teams at a high error rate, and manually reviewing those labels. What they choose to do or what their policy is, I don't know, but it should be trivial to Identify the labeling errors. I can see where some teams that benefits from the label errors gets upset, if changes are made to the leaderboard post stage 2 start, because there will be teams that benefit from the label errors.",
    "249286": "Each prediction should actually be weighted as 1/1700 since you have 17 zones per passenger x 100 passengers in stage 1 test set.\n\n&gt; **Murray wrote**\n&gt; \n&gt; &gt; Well, say there's a threat and you predict a 0.5 (50%) chance of a threat in that zone.  Your score for that area is 0.010158698821276728.  If you predict a 0.51 chance, your score is 0.0099555248448511945.  Hence my statement; I may be forgetting to divide by 17 though... amusingly enough, arithmetic is not my strong suit, haha.\n&gt; \n&gt; Your point's taken though.  Like I said, arguably :)",
    "249300": "DavidGbodiOdaibo\n\nI've been doing something similar; rounding my current submissions to the nearest class, doing log-loss against those fake 'labels' and then compare it to the LB score of the same submission.  If the fake log-loss is lower than the LB, you know you have some misclassifications.  So far, I've never had a perfect classification yet, its probably one big mislabel or a few small ones, not sure.\n\nAs far as hedging against mislabel issues, you get 2 submissions for this competition.  What I'm doing is assuming perfect labels for one of them, and the other will limit how confident any zone can be assuming an equal percentage of mislabels.  I feel like this is the safest strategy.",
    "253831": "do we need to submit the log loss score for the test set in addition to the predicted probability?"
  },
  "source": "meta"
}