{
  "id": 43548,
  "title": "Guesstimate on the winning log loss?",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/43548",
  "author_name": "",
  "post_date": "2017-11-16T07:18:34.217764Z",
  "votes": 2,
  "comment_count": 34,
  "views": 0,
  "content": "<p>The top LB scores are obviously a result of probing / hand labeling...\nAny guesses on what the winning log loss will be?</p>",
  "messages": [
    {
      "id": "244398",
      "postDate": "11/16/2017 07:18:34",
      "content": "<p>The top LB scores are obviously a result of probing / hand labeling...\nAny guesses on what the winning log loss will be?</p>",
      "rawMarkdown": "The top LB scores are obviously a result of probing / hand labeling...\nAny guesses on what the winning log loss will be?",
      "votes": null
    },
    {
      "id": "244406",
      "postDate": "11/16/2017 07:37:06",
      "content": "<p>I'm going with 0.09384</p>",
      "rawMarkdown": "I'm going with 0.09384",
      "votes": null
    },
    {
      "id": "244408",
      "postDate": "11/16/2017 07:39:24",
      "content": "<p>Is that your local CV?  ;)\nMy guess is slightly below 0.10 as well.</p>",
      "rawMarkdown": "Is that your local CV?  ;)\nMy guess is slightly below 0.10 as well.",
      "votes": null
    },
    {
      "id": "244412",
      "postDate": "11/16/2017 07:52:17",
      "content": "<p>No, that's just based off the observation that scores (excluding the probers/hand-labelers) in quite a few other image classification competitions seem to drop off quite a bit from public to private LB. Taking it to the fifth decimal place is just to break any ties with other guessers.</p>",
      "rawMarkdown": "No, that's just based off the observation that scores (excluding the probers/hand-labelers) in quite a few other image classification competitions seem to drop off quite a bit from public to private LB. Taking it to the fifth decimal place is just to break any ties with other guessers.",
      "votes": null
    },
    {
      "id": "244768",
      "postDate": "11/16/2017 21:48:48",
      "content": "<p>I'd be very surprised if it was above 0.01, but I guess I haven't really experienced similar competitions before...</p>\n\n<p>Even using your own current public score, what makes you think it would be almost 4 times higher with the private set?  Or more generally, how could it be <em>that</em> over-fit in other competitions?  I feel like you'd have to be using some seriously bad statistical practices to be off by that much.  I'm curious, what other competitions did you see this phenomenon?</p>\n\n<p>Or are you worried more specifically that the private set will have images and threats that are significantly different from the public set (body types, threat types, threat distribution, etc) or even mislabels?  Mislabels are still my biggest fear.  I feel like that can add another ~0.02 penalty if there's as many in the private set as there are in the public one.</p>",
      "rawMarkdown": "I'd be very surprised if it was above 0.01, but I guess I haven't really experienced similar competitions before...\n\nEven using your own current public score, what makes you think it would be almost 4 times higher with the private set?  Or more generally, how could it be *that* over-fit in other competitions?  I feel like you'd have to be using some seriously bad statistical practices to be off by that much.  I'm curious, what other competitions did you see this phenomenon?\n\nOr are you worried more specifically that the private set will have images and threats that are significantly different from the public set (body types, threat types, threat distribution, etc) or even mislabels?  Mislabels are still my biggest fear.  I feel like that can add another ~0.02 penalty if there's as many in the private set as there are in the public one.",
      "votes": null
    },
    {
      "id": "245135",
      "postDate": "11/17/2017 17:39:46",
      "content": "<p>Are you really scoring below 0.01 without training on the validation set?  If that's the case (and if you don't win), please put the design up somewhere after the contest.  I'd love to study it.  At any rate, I'm guessing around 0.1 too.</p>",
      "rawMarkdown": "Are you really scoring below 0.01 without training on the validation set?  If that's the case (and if you don't win), please put the design up somewhere after the contest.  I'd love to study it.  At any rate, I'm guessing around 0.1 too.",
      "votes": null
    },
    {
      "id": "245254",
      "postDate": "11/17/2017 21:43:25",
      "content": "<p>I don't know the validation labels, and even if I had them, I wouldn't use them in my training because otherwise I wouldn't have anything else to validate against.  My validation score does agree with my test/train split score so I'm not worried about over-fitting in that sense, unless both the training and validation are not representative of stage2.</p>",
      "rawMarkdown": "I don't know the validation labels, and even if I had them, I wouldn't use them in my training because otherwise I wouldn't have anything else to validate against.  My validation score does agree with my test/train split score so I'm not worried about over-fitting in that sense, unless both the training and validation are not representative of stage2.",
      "votes": null
    },
    {
      "id": "245561",
      "postDate": "11/18/2017 20:49:46",
      "content": "<p>Your score is absolutely astounding, Kevin; it borders on literal perfection.  Apparently your model is capable of generalizing to almost any similar image and classification task.  Since you need to be off by less than half a percent -- on average -- to score 0.005, I'd personally call that as perfect as it's going to get.</p>\n\n<p>I mean it: I'd love to study your design, but since you're obviously going to win, that may be tricky.  Still, I hope you'd be open to it later on :)</p>",
      "rawMarkdown": "Your score is absolutely astounding, Kevin; it borders on literal perfection.  Apparently your model is capable of generalizing to almost any similar image and classification task.  Since you need to be off by less than half a percent -- on average -- to score 0.005, I'd personally call that as perfect as it's going to get.\n\nI mean it: I'd love to study your design, but since you're obviously going to win, that may be tricky.  Still, I hope you'd be open to it later on :)",
      "votes": null
    },
    {
      "id": "245567",
      "postDate": "11/18/2017 21:21:08",
      "content": "<p>Thank you!  There are other competitors that have close scores and I'm assuming their models are legit too.  And given the uncertainty with mislabels, its definitely not assured at all, and there may be even more submissions coming in last minute!</p>\n\n<p>Regardless of ranking, I'll definitely post a write-up of my model once the competition ends.</p>",
      "rawMarkdown": "Thank you!  There are other competitors that have close scores and I'm assuming their models are legit too.  And given the uncertainty with mislabels, its definitely not assured at all, and there may be even more submissions coming in last minute!\n\nRegardless of ranking, I'll definitely post a write-up of my model once the competition ends.",
      "votes": null
    },
    {
      "id": "254506",
      "postDate": "12/07/2017 04:13:44",
      "content": "<blockquote>\n  <p>Even using your own current public score, what makes you think it would be almost 4 times higher with the private set? Or more generally, how could it be that over-fit in other competitions? I feel like you'd have to be using some seriously bad statistical practices to be off by that much. I'm curious, what other competitions did you see this phenomenon?</p>\n  \n  <p>Or are you worried more specifically that the private set will have images and threats that are significantly different from the public set (body types, threat types, threat distribution, etc) or even mislabels? Mislabels are still my biggest fear. I feel like that can add another ~0.02 penalty if there's as many in the private set as there are in the public one.</p>\n</blockquote>\n\n<p>I don't trust my public score, I think I'm overfitting. <a href=\"https://www.kaggle.com/c/the-nature-conservancy-fisheries-monitoring/leaderboard\">The Nature Conservancy Fisheries Monitoring</a> comes to mind as one of the competitions where the Stage 1 scores were much better than the Stage 2 scores. The winners of that competition were 49th on the Public LB. The two stages in that competition used different boats and the two stages in this competition will have different volunteers and some new threats will appear in Stage 2, so they are similar in that way. </p>",
      "rawMarkdown": "&gt; Even using your own current public score, what makes you think it would be almost 4 times higher with the private set? Or more generally, how could it be that over-fit in other competitions? I feel like you'd have to be using some seriously bad statistical practices to be off by that much. I'm curious, what other competitions did you see this phenomenon?\n\n&gt; Or are you worried more specifically that the private set will have images and threats that are significantly different from the public set (body types, threat types, threat distribution, etc) or even mislabels? Mislabels are still my biggest fear. I feel like that can add another ~0.02 penalty if there's as many in the private set as there are in the public one.\n\nI don't trust my public score, I think I'm overfitting. [The Nature Conservancy Fisheries Monitoring][1] comes to mind as one of the competitions where the Stage 1 scores were much better than the Stage 2 scores. The winners of that competition were 49th on the Public LB. The two stages in that competition used different boats and the two stages in this competition will have different volunteers and some new threats will appear in Stage 2, so they are similar in that way. \n\n\n  [1]: https://www.kaggle.com/c/the-nature-conservancy-fisheries-monitoring/leaderboard",
      "votes": null
    },
    {
      "id": "254545",
      "postDate": "12/07/2017 07:00:41",
      "content": "<p>Thanks a bunch for the reference!</p>\n\n<p>Yeah I can see your point about that now.  That competition seems so harsh with the second stage being so different!  I guess we'll have to wait and see with this one how different stage 2 will \n be... But it was interesting reading some of the discussions there, hopefully I'll gain some insight :)</p>",
      "rawMarkdown": "Thanks a bunch for the reference!\n\nYeah I can see your point about that now.  That competition seems so harsh with the second stage being so different!  I guess we'll have to wait and see with this one how different stage 2 will \n be... But it was interesting reading some of the discussions there, hopefully I'll gain some insight :)",
      "votes": null
    },
    {
      "id": "255363",
      "postDate": "12/08/2017 21:54:05",
      "content": "<p><a href=\"https://www.kaggle.com/c/data-science-bowl-2017/leaderboard\">Lung Cancer Detection</a> was another image competition with a pretty massive shakeup and a decent sized gap between the top score on the Public LB and Private LB. Most of the Top 10 moved up hundreds of positions from the Public LB. Not sure what caused that yet though.</p>",
      "rawMarkdown": "[Lung Cancer Detection][1] was another image competition with a pretty massive shakeup and a decent sized gap between the top score on the Public LB and Private LB. Most of the Top 10 moved up hundreds of positions from the Public LB. Not sure what caused that yet though.\n\n\n  [1]: https://www.kaggle.com/c/data-science-bowl-2017/leaderboard",
      "votes": null
    },
    {
      "id": "255378",
      "postDate": "12/08/2017 22:31:13",
      "content": "<p>Actually the lung shakeup wasn't too strange. Our final results were quite consistent with CV scores. The reason for the big swing in LB rankings was just that there were a lot of manual labelers and overfitters on the public LB which was only 100 samples. Current LB is 1700 (albeit 100 scans and even fewer \"subjects\") but should be more reliable. I wouldn't expect a huge shake up here unless lots of our methods fail to generalize to new subjects. </p>",
      "rawMarkdown": "Actually the lung shakeup wasn't too strange. Our final results were quite consistent with CV scores. The reason for the big swing in LB rankings was just that there were a lot of manual labelers and overfitters on the public LB which was only 100 samples. Current LB is 1700 (albeit 100 scans and even fewer \"subjects\") but should be more reliable. I wouldn't expect a huge shake up here unless lots of our methods fail to generalize to new subjects.",
      "votes": null
    },
    {
      "id": "255437",
      "postDate": "12/09/2017 02:44:08",
      "content": "<p>I feel the same way as Branden. When they said the stage 2 data would be similar to stage 1 I assumed they meant it would be the same people, but then I recently learned they're purposely using new volunteers in the stage 2 data.</p>\n\n<p>I'm pretty sure there's going to be big swings in scores especially for people who didn't take into consideration the above.</p>",
      "rawMarkdown": "I feel the same way as Branden. When they said the stage 2 data would be similar to stage 1 I assumed they meant it would be the same people, but then I recently learned they're purposely using new volunteers in the stage 2 data.\n\nI'm pretty sure there's going to be big swings in scores especially for people who didn't take into consideration the above.",
      "votes": null
    },
    {
      "id": "255438",
      "postDate": "12/09/2017 02:53:57",
      "content": "<p>When/where did they announce the details regarding stage 2 test data?  </p>",
      "rawMarkdown": "When/where did they announce the details regarding stage 2 test data?",
      "votes": null
    },
    {
      "id": "255446",
      "postDate": "12/09/2017 03:36:03",
      "content": "<p>It's in the 3rd paragraph of the data tab. I feel like I would have noticed that detail if it was there from the start. Maybe it was snuck in?</p>",
      "rawMarkdown": "It's in the 3rd paragraph of the data tab. I feel like I would have noticed that detail if it was there from the start. Maybe it was snuck in?",
      "votes": null
    },
    {
      "id": "255449",
      "postDate": "12/09/2017 03:41:48",
      "content": "<p>Oh wow.  Thanks @Moejoe...  I did not notice that, either.  Getting a low logloss on local CV / public LB doesn't mean crap!</p>",
      "rawMarkdown": "Oh wow.  Thanks @Moejoe...  I did not notice that, either.  Getting a low logloss on local CV / public LB doesn't mean crap!",
      "votes": null
    },
    {
      "id": "255452",
      "postDate": "12/09/2017 03:53:59",
      "content": "<p>Glad I'm not going crazy. I think it's really unfair they didn't make that clear from the start and even sort of misled us to believe the opposite (the test data not being indicative of what we would actually be scored on).</p>",
      "rawMarkdown": "Glad I'm not going crazy. I think it's really unfair they didn't make that clear from the start and even sort of misled us to believe the opposite (the test data not being indicative of what we would actually be scored on).",
      "votes": null
    },
    {
      "id": "255454",
      "postDate": "12/09/2017 04:06:48",
      "content": "<p>Oh well, I guess it is still <em>fair</em> as everyone has access to the same info.\nFrom the standpoint of evaluating which models are usable in real life, it only makes sense that stage 2 test data be of different people.  But if that's actually the case for this competition, then perhaps stage 1 test data could have been representative of this and should have had different subjects/threat types than what's in the training set.  I suppose one can still manually separate out subjects/threat types when creating folds, etc., which in hindsight is what I should have done to get a more reliable estimate of how well the models generalize (lessons learned!)  I'm hoping that the preprocessing that I'm doing would somewhat take care of individual differences.</p>",
      "rawMarkdown": "Oh well, I guess it is still *fair* as everyone has access to the same info.\nFrom the standpoint of evaluating which models are usable in real life, it only makes sense that stage 2 test data be of different people.  But if that's actually the case for this competition, then perhaps stage 1 test data could have been representative of this and should have had different subjects/threat types than what's in the training set.  I suppose one can still manually separate out subjects/threat types when creating folds, etc., which in hindsight is what I should have done to get a more reliable estimate of how well the models generalize (lessons learned!)  I'm hoping that the preprocessing that I'm doing would somewhat take care of individual differences.",
      "votes": null
    },
    {
      "id": "255457",
      "postDate": "12/09/2017 04:22:37",
      "content": "<p>I recall seeing this info from the start... I'm pretty sure it was discussed in one of the threads.</p>",
      "rawMarkdown": "I recall seeing this info from the start... I'm pretty sure it was discussed in one of the threads.",
      "votes": null
    },
    {
      "id": "255461",
      "postDate": "12/09/2017 04:41:50",
      "content": "<p>The fact that the Stage 2 data will have different \"subjects\" than Stage 1 has been in the description from the start, as far as I can remember.</p>",
      "rawMarkdown": "The fact that the Stage 2 data will have different \"subjects\" than Stage 1 has been in the description from the start, as far as I can remember.",
      "votes": null
    },
    {
      "id": "255464",
      "postDate": "12/09/2017 04:46:12",
      "content": "<p>Ah I misread what was being said in the thread I linked. Oh well.</p>",
      "rawMarkdown": "Ah I misread what was being said in the thread I linked. Oh well.",
      "votes": null
    },
    {
      "id": "255467",
      "postDate": "12/09/2017 05:10:25",
      "content": "<p>Well, with only a few days remaining...</p>\n\n<p>GOOD LUCK EVERYONE!</p>",
      "rawMarkdown": "Well, with only a few days remaining...\n\nGOOD LUCK EVERYONE!",
      "votes": null
    },
    {
      "id": "256013",
      "postDate": "12/10/2017 23:38:05",
      "content": "<p>My guesstimate is an ensemble of your guesstimates: <code>sqrt(0.09384 * 0.01)</code></p>",
      "rawMarkdown": "My guesstimate is an ensemble of your guesstimates: `sqrt(0.09384 * 0.01)`",
      "votes": null
    },
    {
      "id": "256063",
      "postDate": "12/11/2017 02:43:19",
      "content": "<p>My 'under 0.01' guesstimate was assuming there were no mislabels :P  But I'm going to guess there will be ~6 mislabels in stage 2 and this will make my actual guesstimate around 0.03</p>",
      "rawMarkdown": "My 'under 0.01' guesstimate was assuming there were no mislabels :P  But I'm going to guess there will be ~6 mislabels in stage 2 and this will make my actual guesstimate around 0.03",
      "votes": null
    },
    {
      "id": "256066",
      "postDate": "12/11/2017 03:04:35",
      "content": "<p>How does the math for that work out? Would't it be something closer to 0.01*(1-6/(1147*17))-ln(1-exp(-0.01))*6/(1147*17) = 0.0114, assuming all your predictions have the same loss (which isn't correct, but maybe is a reasonable approximation)? I think even if you predict with maximum confidence, the additional loss shouldn't be more than -ln(1e-15)*6/(1147*17) = 0.0106</p>",
      "rawMarkdown": "How does the math for that work out? Would't it be something closer to 0.01*(1-6/(1147*17))-ln(1-exp(-0.01))*6/(1147*17) = 0.0114, assuming all your predictions have the same loss (which isn't correct, but maybe is a reasonable approximation)? I think even if you predict with maximum confidence, the additional loss shouldn't be more than -ln(1e-15)*6/(1147*17) = 0.0106",
      "votes": null
    },
    {
      "id": "256067",
      "postDate": "12/11/2017 03:06:12",
      "content": "<p>.02 assuming mislabels aren't super common in stage2 (unfortunately they are quite likely).</p>\n\n<p>Hopefully there isn't too much degradation when applying our methods to new people!</p>",
      "rawMarkdown": ".02 assuming mislabels aren't super common in stage2 (unfortunately they are quite likely).\n\nHopefully there isn't too much degradation when applying our methods to new people!",
      "votes": null
    },
    {
      "id": "256072",
      "postDate": "12/11/2017 03:17:48",
      "content": "<p>0.02 would be <em>super</em> impressive.  I wonder if there will be new types of threats in the stage 2 test set in addition to new subjects.  Can't wait to find out.</p>",
      "rawMarkdown": "0.02 would be *super* impressive.  I wonder if there will be new types of threats in the stage 2 test set in addition to new subjects.  Can't wait to find out.",
      "votes": null
    },
    {
      "id": "256740",
      "postDate": "12/12/2017 16:38:52",
      "content": "<blockquote>\n  <p><strong>Yusaku Sako wrote</strong></p>\n  \n  <blockquote>\n    <p>I suppose one can still manually separate out subjects/threat types when creating folds, etc., which in hindsight is what I should have done to get a more reliable estimate of how well the models generalize (lessons learned!) </p>\n  </blockquote>\n</blockquote>\n\n<p>I'm curious to see how many other teams did this</p>",
      "rawMarkdown": "&gt; **Yusaku Sako wrote**\n&gt; \n&gt;\n&gt; &gt; I suppose one can still manually separate out subjects/threat types when creating folds, etc., which in hindsight is what I should have done to get a more reliable estimate of how well the models generalize (lessons learned!) \n\nI'm curious to see how many other teams did this",
      "votes": null
    },
    {
      "id": "256768",
      "postDate": "12/12/2017 17:31:04",
      "content": "<p>PCA works well for sorting individual subjects, any idea what technique works for sorting individual threats?</p>",
      "rawMarkdown": "PCA works well for sorting individual subjects, any idea what technique works for sorting individual threats?",
      "votes": null
    },
    {
      "id": "256996",
      "postDate": "12/13/2017 05:18:16",
      "content": "<p>I used k-means on my annotation data to sort the subjects. That got me about 70% of the way and then I refined it by hand. I didn't bother sorting the threats, but in hindsight I wish I would've taken the time to do it.</p>",
      "rawMarkdown": "I used k-means on my annotation data to sort the subjects. That got me about 70% of the way and then I refined it by hand. I didn't bother sorting the threats, but in hindsight I wish I would've taken the time to do it.",
      "votes": null
    },
    {
      "id": "256999",
      "postDate": "12/13/2017 05:21:56",
      "content": "<p>Wonder what you mean with \"annotation data\", either way, PCA seems to perform better on this task.</p>",
      "rawMarkdown": "Wonder what you mean with \"annotation data\", either way, PCA seems to perform better on this task.",
      "votes": null
    },
    {
      "id": "257005",
      "postDate": "12/13/2017 05:40:21",
      "content": "<p>I also used PCA to reduce the dimensions of the aps images, then ran k-means and refined clusters by hand. I didn't sort threats either, but I'm not yet sure how much I regret not doing it :P</p>",
      "rawMarkdown": "I also used PCA to reduce the dimensions of the aps images, then ran k-means and refined clusters by hand. I didn't sort threats either, but I'm not yet sure how much I regret not doing it :P",
      "votes": null
    },
    {
      "id": "257087",
      "postDate": "12/13/2017 11:28:37",
      "content": "<p>I used rmse on my zone detection (1 ~ 17 body zone) model's output to find the same volunteers. It achieved nearly 100% accuracy since the outputs of nn model are more stable than my clumsy hand labeling and there are only around 20 volunteers. </p>",
      "rawMarkdown": "I used rmse on my zone detection (1 ~ 17 body zone) model's output to find the same volunteers. It achieved nearly 100% accuracy since the outputs of nn model are more stable than my clumsy hand labeling and there are only around 20 volunteers.",
      "votes": null
    },
    {
      "id": "258348",
      "postDate": "12/16/2017 00:36:15",
      "content": "<p>Your guess was pretty darn close!</p>",
      "rawMarkdown": "Your guess was pretty darn close!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 244406,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "11/16/2017 07:37:06",
      "content": "<p>I'm going with 0.09384</p>",
      "votes": null,
      "replies": [
        {
          "id": 244408,
          "author_name": "u39kun",
          "author_url": "",
          "post_date": "11/16/2017 07:39:24",
          "content": "<p>Is that your local CV?  ;)\nMy guess is slightly below 0.10 as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 244412,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "11/16/2017 07:52:17",
          "content": "<p>No, that's just based off the observation that scores (excluding the probers/hand-labelers) in quite a few other image classification competitions seem to drop off quite a bit from public to private LB. Taking it to the fifth decimal place is just to break any ties with other guessers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 244768,
          "author_name": "hackerpoet",
          "author_url": "",
          "post_date": "11/16/2017 21:48:48",
          "content": "<p>I'd be very surprised if it was above 0.01, but I guess I haven't really experienced similar competitions before...</p>\n\n<p>Even using your own current public score, what makes you think it would be almost 4 times higher with the private set?  Or more generally, how could it be <em>that</em> over-fit in other competitions?  I feel like you'd have to be using some seriously bad statistical practices to be off by that much.  I'm curious, what other competitions did you see this phenomenon?</p>\n\n<p>Or are you worried more specifically that the private set will have images and threats that are significantly different from the public set (body types, threat types, threat distribution, etc) or even mislabels?  Mislabels are still my biggest fear.  I feel like that can add another ~0.02 penalty if there's as many in the private set as there are in the public one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 245135,
          "author_name": "mmiron",
          "author_url": "",
          "post_date": "11/17/2017 17:39:46",
          "content": "<p>Are you really scoring below 0.01 without training on the validation set?  If that's the case (and if you don't win), please put the design up somewhere after the contest.  I'd love to study it.  At any rate, I'm guessing around 0.1 too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 245254,
          "author_name": "hackerpoet",
          "author_url": "",
          "post_date": "11/17/2017 21:43:25",
          "content": "<p>I don't know the validation labels, and even if I had them, I wouldn't use them in my training because otherwise I wouldn't have anything else to validate against.  My validation score does agree with my test/train split score so I'm not worried about over-fitting in that sense, unless both the training and validation are not representative of stage2.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 245561,
          "author_name": "mmiron",
          "author_url": "",
          "post_date": "11/18/2017 20:49:46",
          "content": "<p>Your score is absolutely astounding, Kevin; it borders on literal perfection.  Apparently your model is capable of generalizing to almost any similar image and classification task.  Since you need to be off by less than half a percent -- on average -- to score 0.005, I'd personally call that as perfect as it's going to get.</p>\n\n<p>I mean it: I'd love to study your design, but since you're obviously going to win, that may be tricky.  Still, I hope you'd be open to it later on :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 245567,
          "author_name": "hackerpoet",
          "author_url": "",
          "post_date": "11/18/2017 21:21:08",
          "content": "<p>Thank you!  There are other competitors that have close scores and I'm assuming their models are legit too.  And given the uncertainty with mislabels, its definitely not assured at all, and there may be even more submissions coming in last minute!</p>\n\n<p>Regardless of ranking, I'll definitely post a write-up of my model once the competition ends.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 254506,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "12/07/2017 04:13:44",
          "content": "<blockquote>\n  <p>Even using your own current public score, what makes you think it would be almost 4 times higher with the private set? Or more generally, how could it be that over-fit in other competitions? I feel like you'd have to be using some seriously bad statistical practices to be off by that much. I'm curious, what other competitions did you see this phenomenon?</p>\n  \n  <p>Or are you worried more specifically that the private set will have images and threats that are significantly different from the public set (body types, threat types, threat distribution, etc) or even mislabels? Mislabels are still my biggest fear. I feel like that can add another ~0.02 penalty if there's as many in the private set as there are in the public one.</p>\n</blockquote>\n\n<p>I don't trust my public score, I think I'm overfitting. <a href=\"https://www.kaggle.com/c/the-nature-conservancy-fisheries-monitoring/leaderboard\">The Nature Conservancy Fisheries Monitoring</a> comes to mind as one of the competitions where the Stage 1 scores were much better than the Stage 2 scores. The winners of that competition were 49th on the Public LB. The two stages in that competition used different boats and the two stages in this competition will have different volunteers and some new threats will appear in Stage 2, so they are similar in that way. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 254545,
          "author_name": "hackerpoet",
          "author_url": "",
          "post_date": "12/07/2017 07:00:41",
          "content": "<p>Thanks a bunch for the reference!</p>\n\n<p>Yeah I can see your point about that now.  That competition seems so harsh with the second stage being so different!  I guess we'll have to wait and see with this one how different stage 2 will \n be... But it was interesting reading some of the discussions there, hopefully I'll gain some insight :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255363,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "12/08/2017 21:54:05",
          "content": "<p><a href=\"https://www.kaggle.com/c/data-science-bowl-2017/leaderboard\">Lung Cancer Detection</a> was another image competition with a pretty massive shakeup and a decent sized gap between the top score on the Public LB and Private LB. Most of the Top 10 moved up hundreds of positions from the Public LB. Not sure what caused that yet though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255378,
          "author_name": "dhammack",
          "author_url": "",
          "post_date": "12/08/2017 22:31:13",
          "content": "<p>Actually the lung shakeup wasn't too strange. Our final results were quite consistent with CV scores. The reason for the big swing in LB rankings was just that there were a lot of manual labelers and overfitters on the public LB which was only 100 samples. Current LB is 1700 (albeit 100 scans and even fewer \"subjects\") but should be more reliable. I wouldn't expect a huge shake up here unless lots of our methods fail to generalize to new subjects. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255437,
          "author_name": "moejoe",
          "author_url": "",
          "post_date": "12/09/2017 02:44:08",
          "content": "<p>I feel the same way as Branden. When they said the stage 2 data would be similar to stage 1 I assumed they meant it would be the same people, but then I recently learned they're purposely using new volunteers in the stage 2 data.</p>\n\n<p>I'm pretty sure there's going to be big swings in scores especially for people who didn't take into consideration the above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255438,
          "author_name": "u39kun",
          "author_url": "",
          "post_date": "12/09/2017 02:53:57",
          "content": "<p>When/where did they announce the details regarding stage 2 test data?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255446,
          "author_name": "moejoe",
          "author_url": "",
          "post_date": "12/09/2017 03:36:03",
          "content": "<p>It's in the 3rd paragraph of the data tab. I feel like I would have noticed that detail if it was there from the start. Maybe it was snuck in?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255449,
          "author_name": "u39kun",
          "author_url": "",
          "post_date": "12/09/2017 03:41:48",
          "content": "<p>Oh wow.  Thanks @Moejoe...  I did not notice that, either.  Getting a low logloss on local CV / public LB doesn't mean crap!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255452,
          "author_name": "moejoe",
          "author_url": "",
          "post_date": "12/09/2017 03:53:59",
          "content": "<p>Glad I'm not going crazy. I think it's really unfair they didn't make that clear from the start and even sort of misled us to believe the opposite (the test data not being indicative of what we would actually be scored on).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255454,
          "author_name": "u39kun",
          "author_url": "",
          "post_date": "12/09/2017 04:06:48",
          "content": "<p>Oh well, I guess it is still <em>fair</em> as everyone has access to the same info.\nFrom the standpoint of evaluating which models are usable in real life, it only makes sense that stage 2 test data be of different people.  But if that's actually the case for this competition, then perhaps stage 1 test data could have been representative of this and should have had different subjects/threat types than what's in the training set.  I suppose one can still manually separate out subjects/threat types when creating folds, etc., which in hindsight is what I should have done to get a more reliable estimate of how well the models generalize (lessons learned!)  I'm hoping that the preprocessing that I'm doing would somewhat take care of individual differences.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255457,
          "author_name": "hackerpoet",
          "author_url": "",
          "post_date": "12/09/2017 04:22:37",
          "content": "<p>I recall seeing this info from the start... I'm pretty sure it was discussed in one of the threads.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255461,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "12/09/2017 04:41:50",
          "content": "<p>The fact that the Stage 2 data will have different \"subjects\" than Stage 1 has been in the description from the start, as far as I can remember.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255464,
          "author_name": "moejoe",
          "author_url": "",
          "post_date": "12/09/2017 04:46:12",
          "content": "<p>Ah I misread what was being said in the thread I linked. Oh well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255467,
          "author_name": "u39kun",
          "author_url": "",
          "post_date": "12/09/2017 05:10:25",
          "content": "<p>Well, with only a few days remaining...</p>\n\n<p>GOOD LUCK EVERYONE!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256740,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "12/12/2017 16:38:52",
          "content": "<blockquote>\n  <p><strong>Yusaku Sako wrote</strong></p>\n  \n  <blockquote>\n    <p>I suppose one can still manually separate out subjects/threat types when creating folds, etc., which in hindsight is what I should have done to get a more reliable estimate of how well the models generalize (lessons learned!) </p>\n  </blockquote>\n</blockquote>\n\n<p>I'm curious to see how many other teams did this</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256768,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/12/2017 17:31:04",
          "content": "<p>PCA works well for sorting individual subjects, any idea what technique works for sorting individual threats?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256996,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "12/13/2017 05:18:16",
          "content": "<p>I used k-means on my annotation data to sort the subjects. That got me about 70% of the way and then I refined it by hand. I didn't bother sorting the threats, but in hindsight I wish I would've taken the time to do it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256999,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/13/2017 05:21:56",
          "content": "<p>Wonder what you mean with \"annotation data\", either way, PCA seems to perform better on this task.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 257005,
          "author_name": "suchir",
          "author_url": "",
          "post_date": "12/13/2017 05:40:21",
          "content": "<p>I also used PCA to reduce the dimensions of the aps images, then ran k-means and refined clusters by hand. I didn't sort threats either, but I'm not yet sure how much I regret not doing it :P</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 257087,
          "author_name": "ckomaki",
          "author_url": "",
          "post_date": "12/13/2017 11:28:37",
          "content": "<p>I used rmse on my zone detection (1 ~ 17 body zone) model's output to find the same volunteers. It achieved nearly 100% accuracy since the outputs of nn model are more stable than my clumsy hand labeling and there are only around 20 volunteers. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 256013,
      "author_name": "olegtrott",
      "author_url": "",
      "post_date": "12/10/2017 23:38:05",
      "content": "<p>My guesstimate is an ensemble of your guesstimates: <code>sqrt(0.09384 * 0.01)</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 256063,
          "author_name": "hackerpoet",
          "author_url": "",
          "post_date": "12/11/2017 02:43:19",
          "content": "<p>My 'under 0.01' guesstimate was assuming there were no mislabels :P  But I'm going to guess there will be ~6 mislabels in stage 2 and this will make my actual guesstimate around 0.03</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256066,
          "author_name": "suchir",
          "author_url": "",
          "post_date": "12/11/2017 03:04:35",
          "content": "<p>How does the math for that work out? Would't it be something closer to 0.01*(1-6/(1147*17))-ln(1-exp(-0.01))*6/(1147*17) = 0.0114, assuming all your predictions have the same loss (which isn't correct, but maybe is a reasonable approximation)? I think even if you predict with maximum confidence, the additional loss shouldn't be more than -ln(1e-15)*6/(1147*17) = 0.0106</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 256067,
      "author_name": "dhammack",
      "author_url": "",
      "post_date": "12/11/2017 03:06:12",
      "content": "<p>.02 assuming mislabels aren't super common in stage2 (unfortunately they are quite likely).</p>\n\n<p>Hopefully there isn't too much degradation when applying our methods to new people!</p>",
      "votes": null,
      "replies": [
        {
          "id": 256072,
          "author_name": "u39kun",
          "author_url": "",
          "post_date": "12/11/2017 03:17:48",
          "content": "<p>0.02 would be <em>super</em> impressive.  I wonder if there will be new types of threats in the stage 2 test set in addition to new subjects.  Can't wait to find out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258348,
          "author_name": "u39kun",
          "author_url": "",
          "post_date": "12/16/2017 00:36:15",
          "content": "<p>Your guess was pretty darn close!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "244398": "The top LB scores are obviously a result of probing / hand labeling...\nAny guesses on what the winning log loss will be?",
    "244406": "I'm going with 0.09384",
    "244408": "Is that your local CV?  ;)\nMy guess is slightly below 0.10 as well.",
    "244412": "No, that's just based off the observation that scores (excluding the probers/hand-labelers) in quite a few other image classification competitions seem to drop off quite a bit from public to private LB. Taking it to the fifth decimal place is just to break any ties with other guessers.",
    "244768": "I'd be very surprised if it was above 0.01, but I guess I haven't really experienced similar competitions before...\n\nEven using your own current public score, what makes you think it would be almost 4 times higher with the private set?  Or more generally, how could it be *that* over-fit in other competitions?  I feel like you'd have to be using some seriously bad statistical practices to be off by that much.  I'm curious, what other competitions did you see this phenomenon?\n\nOr are you worried more specifically that the private set will have images and threats that are significantly different from the public set (body types, threat types, threat distribution, etc) or even mislabels?  Mislabels are still my biggest fear.  I feel like that can add another ~0.02 penalty if there's as many in the private set as there are in the public one.",
    "245135": "Are you really scoring below 0.01 without training on the validation set?  If that's the case (and if you don't win), please put the design up somewhere after the contest.  I'd love to study it.  At any rate, I'm guessing around 0.1 too.",
    "245254": "I don't know the validation labels, and even if I had them, I wouldn't use them in my training because otherwise I wouldn't have anything else to validate against.  My validation score does agree with my test/train split score so I'm not worried about over-fitting in that sense, unless both the training and validation are not representative of stage2.",
    "245561": "Your score is absolutely astounding, Kevin; it borders on literal perfection.  Apparently your model is capable of generalizing to almost any similar image and classification task.  Since you need to be off by less than half a percent -- on average -- to score 0.005, I'd personally call that as perfect as it's going to get.\n\nI mean it: I'd love to study your design, but since you're obviously going to win, that may be tricky.  Still, I hope you'd be open to it later on :)",
    "245567": "Thank you!  There are other competitors that have close scores and I'm assuming their models are legit too.  And given the uncertainty with mislabels, its definitely not assured at all, and there may be even more submissions coming in last minute!\n\nRegardless of ranking, I'll definitely post a write-up of my model once the competition ends.",
    "254506": "&gt; Even using your own current public score, what makes you think it would be almost 4 times higher with the private set? Or more generally, how could it be that over-fit in other competitions? I feel like you'd have to be using some seriously bad statistical practices to be off by that much. I'm curious, what other competitions did you see this phenomenon?\n\n&gt; Or are you worried more specifically that the private set will have images and threats that are significantly different from the public set (body types, threat types, threat distribution, etc) or even mislabels? Mislabels are still my biggest fear. I feel like that can add another ~0.02 penalty if there's as many in the private set as there are in the public one.\n\nI don't trust my public score, I think I'm overfitting. [The Nature Conservancy Fisheries Monitoring][1] comes to mind as one of the competitions where the Stage 1 scores were much better than the Stage 2 scores. The winners of that competition were 49th on the Public LB. The two stages in that competition used different boats and the two stages in this competition will have different volunteers and some new threats will appear in Stage 2, so they are similar in that way. \n\n\n  [1]: https://www.kaggle.com/c/the-nature-conservancy-fisheries-monitoring/leaderboard",
    "254545": "Thanks a bunch for the reference!\n\nYeah I can see your point about that now.  That competition seems so harsh with the second stage being so different!  I guess we'll have to wait and see with this one how different stage 2 will \n be... But it was interesting reading some of the discussions there, hopefully I'll gain some insight :)",
    "255363": "[Lung Cancer Detection][1] was another image competition with a pretty massive shakeup and a decent sized gap between the top score on the Public LB and Private LB. Most of the Top 10 moved up hundreds of positions from the Public LB. Not sure what caused that yet though.\n\n\n  [1]: https://www.kaggle.com/c/data-science-bowl-2017/leaderboard",
    "255378": "Actually the lung shakeup wasn't too strange. Our final results were quite consistent with CV scores. The reason for the big swing in LB rankings was just that there were a lot of manual labelers and overfitters on the public LB which was only 100 samples. Current LB is 1700 (albeit 100 scans and even fewer \"subjects\") but should be more reliable. I wouldn't expect a huge shake up here unless lots of our methods fail to generalize to new subjects.",
    "255437": "I feel the same way as Branden. When they said the stage 2 data would be similar to stage 1 I assumed they meant it would be the same people, but then I recently learned they're purposely using new volunteers in the stage 2 data.\n\nI'm pretty sure there's going to be big swings in scores especially for people who didn't take into consideration the above.",
    "255438": "When/where did they announce the details regarding stage 2 test data?",
    "255446": "It's in the 3rd paragraph of the data tab. I feel like I would have noticed that detail if it was there from the start. Maybe it was snuck in?",
    "255449": "Oh wow.  Thanks @Moejoe...  I did not notice that, either.  Getting a low logloss on local CV / public LB doesn't mean crap!",
    "255452": "Glad I'm not going crazy. I think it's really unfair they didn't make that clear from the start and even sort of misled us to believe the opposite (the test data not being indicative of what we would actually be scored on).",
    "255454": "Oh well, I guess it is still *fair* as everyone has access to the same info.\nFrom the standpoint of evaluating which models are usable in real life, it only makes sense that stage 2 test data be of different people.  But if that's actually the case for this competition, then perhaps stage 1 test data could have been representative of this and should have had different subjects/threat types than what's in the training set.  I suppose one can still manually separate out subjects/threat types when creating folds, etc., which in hindsight is what I should have done to get a more reliable estimate of how well the models generalize (lessons learned!)  I'm hoping that the preprocessing that I'm doing would somewhat take care of individual differences.",
    "255457": "I recall seeing this info from the start... I'm pretty sure it was discussed in one of the threads.",
    "255461": "The fact that the Stage 2 data will have different \"subjects\" than Stage 1 has been in the description from the start, as far as I can remember.",
    "255464": "Ah I misread what was being said in the thread I linked. Oh well.",
    "255467": "Well, with only a few days remaining...\n\nGOOD LUCK EVERYONE!",
    "256013": "My guesstimate is an ensemble of your guesstimates: `sqrt(0.09384 * 0.01)`",
    "256063": "My 'under 0.01' guesstimate was assuming there were no mislabels :P  But I'm going to guess there will be ~6 mislabels in stage 2 and this will make my actual guesstimate around 0.03",
    "256066": "How does the math for that work out? Would't it be something closer to 0.01*(1-6/(1147*17))-ln(1-exp(-0.01))*6/(1147*17) = 0.0114, assuming all your predictions have the same loss (which isn't correct, but maybe is a reasonable approximation)? I think even if you predict with maximum confidence, the additional loss shouldn't be more than -ln(1e-15)*6/(1147*17) = 0.0106",
    "256067": ".02 assuming mislabels aren't super common in stage2 (unfortunately they are quite likely).\n\nHopefully there isn't too much degradation when applying our methods to new people!",
    "256072": "0.02 would be *super* impressive.  I wonder if there will be new types of threats in the stage 2 test set in addition to new subjects.  Can't wait to find out.",
    "256740": "&gt; **Yusaku Sako wrote**\n&gt; \n&gt;\n&gt; &gt; I suppose one can still manually separate out subjects/threat types when creating folds, etc., which in hindsight is what I should have done to get a more reliable estimate of how well the models generalize (lessons learned!) \n\nI'm curious to see how many other teams did this",
    "256768": "PCA works well for sorting individual subjects, any idea what technique works for sorting individual threats?",
    "256996": "I used k-means on my annotation data to sort the subjects. That got me about 70% of the way and then I refined it by hand. I didn't bother sorting the threats, but in hindsight I wish I would've taken the time to do it.",
    "256999": "Wonder what you mean with \"annotation data\", either way, PCA seems to perform better on this task.",
    "257005": "I also used PCA to reduce the dimensions of the aps images, then ran k-means and refined clusters by hand. I didn't sort threats either, but I'm not yet sure how much I regret not doing it :P",
    "257087": "I used rmse on my zone detection (1 ~ 17 body zone) model's output to find the same volunteers. It achieved nearly 100% accuracy since the outputs of nn model are more stable than my clumsy hand labeling and there are only around 20 volunteers.",
    "258348": "Your guess was pretty darn close!"
  },
  "source": "meta"
}