{
  "id": 66323,
  "title": "Inconsistencies between train dataset and test dataset  ",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/66323",
  "author_name": "",
  "post_date": "2018-09-20T13:18:55.089922600Z",
  "votes": 18,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Hi all - and especially the organizers,\nAfter a month or so working in this competition and submitting a lot of submissions it is clear that the statistics of the test set is very different than the statistics of the training set. </p>\n\n<p>During our work we randomly split the training dataset to train, valid and test sets and see that when we calculate the submission score based on the LB metric (as given by <a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">https://www.kaggle.com/chenyc15/mean-average-precision-metric</a>) our results are consistent between the sets.</p>\n\n<p>BUT - when submitting we get different scores between our internal test set and the competition test set. Therefore it is clear that the test statistics is very different than the train statistics. It is very difficult to train models like this since we actually start to use the submission for model tuning - which is something you should never do.</p>\n\n<p>It will be very helpful if the organizers can comment about the source of this discrepancy. One such source is the fact that three radiologists annotated each image in the test set, while each image in the train set was annotated by only one radiologist (as explained in <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723</a>). Are there additional sources of inconsistency?</p>\n\n<p>Can we get the list of images in the train set (there should be 1500 of these) that were triple read?</p>\n\n<p>Any other comments will be helpful. </p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "390585",
      "postDate": "09/20/2018 13:18:55",
      "content": "<p>Hi all - and especially the organizers,\nAfter a month or so working in this competition and submitting a lot of submissions it is clear that the statistics of the test set is very different than the statistics of the training set. </p>\n\n<p>During our work we randomly split the training dataset to train, valid and test sets and see that when we calculate the submission score based on the LB metric (as given by <a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">https://www.kaggle.com/chenyc15/mean-average-precision-metric</a>) our results are consistent between the sets.</p>\n\n<p>BUT - when submitting we get different scores between our internal test set and the competition test set. Therefore it is clear that the test statistics is very different than the train statistics. It is very difficult to train models like this since we actually start to use the submission for model tuning - which is something you should never do.</p>\n\n<p>It will be very helpful if the organizers can comment about the source of this discrepancy. One such source is the fact that three radiologists annotated each image in the test set, while each image in the train set was annotated by only one radiologist (as explained in <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723</a>). Are there additional sources of inconsistency?</p>\n\n<p>Can we get the list of images in the train set (there should be 1500 of these) that were triple read?</p>\n\n<p>Any other comments will be helpful. </p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Hi all - and especially the organizers,\nAfter a month or so working in this competition and submitting a lot of submissions it is clear that the statistics of the test set is very different than the statistics of the training set. \n\nDuring our work we randomly split the training dataset to train, valid and test sets and see that when we calculate the submission score based on the LB metric (as given by https://www.kaggle.com/chenyc15/mean-average-precision-metric) our results are consistent between the sets.\n\nBUT - when submitting we get different scores between our internal test set and the competition test set. Therefore it is clear that the test statistics is very different than the train statistics. It is very difficult to train models like this since we actually start to use the submission for model tuning - which is something you should never do.\n\nIt will be very helpful if the organizers can comment about the source of this discrepancy. One such source is the fact that three radiologists annotated each image in the test set, while each image in the train set was annotated by only one radiologist (as explained in https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723). Are there additional sources of inconsistency?\n\nCan we get the list of images in the train set (there should be 1500 of these) that were triple read?\n\nAny other comments will be helpful. \n\nThanks.",
      "votes": null
    },
    {
      "id": "390709",
      "postDate": "09/20/2018 17:19:48",
      "content": "<p>What kind of differences are you seeing between your local validation and your LB scores? Mine are within 0.01-0.03 which seems reasonable for a somewhat small test set.</p>",
      "rawMarkdown": "What kind of differences are you seeing between your local validation and your LB scores? Mine are within 0.01-0.03 which seems reasonable for a somewhat small test set.",
      "votes": null
    },
    {
      "id": "390742",
      "postDate": "09/20/2018 18:51:25",
      "content": "<p>I get 0.27 on validation, 0.15 on test. Consistently, across several models and architectures</p>",
      "rawMarkdown": "I get 0.27 on validation, 0.15 on test. Consistently, across several models and architectures",
      "votes": null
    },
    {
      "id": "390995",
      "postDate": "09/21/2018 06:03:41",
      "content": "<p>We get also above 0.21 on validation and anywhere from 0.1 to 0.16 on test depending on parameter tuning. For example, one of our models is a pneumonia classifier after which we run another model to extract bounding boxes. On validation we get the maximal score when we set the pneumonia threshold to 0.4. On test, when we lower the pneumonia threshold to 0.1 we get the maximal value with a variation of 0.05 between the 0.4 and and 0.1 threshold. Lowering the threshold to 0.1 obviously introduces more images of the test set on which the bounding box extractor acts, however since the thresholds are different between validation and test it seems to imply the statistics are also different between them. In a sense it seems like overfitting, however our validation and internal test were never learnt so the overfitting is with regards to the statistics of the train set.</p>",
      "rawMarkdown": "We get also above 0.21 on validation and anywhere from 0.1 to 0.16 on test depending on parameter tuning. For example, one of our models is a pneumonia classifier after which we run another model to extract bounding boxes. On validation we get the maximal score when we set the pneumonia threshold to 0.4. On test, when we lower the pneumonia threshold to 0.1 we get the maximal value with a variation of 0.05 between the 0.4 and and 0.1 threshold. Lowering the threshold to 0.1 obviously introduces more images of the test set on which the bounding box extractor acts, however since the thresholds are different between validation and test it seems to imply the statistics are also different between them. In a sense it seems like overfitting, however our validation and internal test were never learnt so the overfitting is with regards to the statistics of the train set.",
      "votes": null
    },
    {
      "id": "391022",
      "postDate": "09/21/2018 06:37:45",
      "content": "<p>Yes, couldn't phrase it better myself. I actually triple checked that nothing from training sneaked into the validation, it was so pronounced</p>",
      "rawMarkdown": "Yes, couldn't phrase it better myself. I actually triple checked that nothing from training sneaked into the validation, it was so pronounced",
      "votes": null
    },
    {
      "id": "391056",
      "postDate": "09/21/2018 07:35:15",
      "content": "<p>Thank you for your comment,\nAs @GuyE said, It is not only the large difference in the score (can get event up to 0.07 difference in the score), it is also the threshold taken on our classifier output. The maximal score is achieved  using different thresholds. This is why we suspect the statistics is different in the test set and the training set.</p>\n\n<p>I am actually a little bit surprised by what you are saying, since we never saw an agreement on the 0.01 level. We tested several different model and approaches, and always we got the large differences where the score in the LB is significantly lower than our internal score. </p>",
      "rawMarkdown": "Thank you for your comment,\nAs @GuyE said, It is not only the large difference in the score (can get event up to 0.07 difference in the score), it is also the threshold taken on our classifier output. The maximal score is achieved  using different thresholds. This is why we suspect the statistics is different in the test set and the training set.\n\nI am actually a little bit surprised by what you are saying, since we never saw an agreement on the 0.01 level. We tested several different model and approaches, and always we got the large differences where the score in the LB is significantly lower than our internal score.",
      "votes": null
    },
    {
      "id": "391107",
      "postDate": "09/21/2018 09:20:37",
      "content": "<p>I trained 3 different architectures. All showed the same problem. 0.20-0.27 on the validation set, 0.10-0.16 on the test. I have also checked the results on different checkpoints for all models - the results were correlated (got to a peak together and started going down on overfitting together), so it's definitively not overfitting but big difference between the training set and the test. I feel that there is an inherent problem here of data that came from two different sources.</p>",
      "rawMarkdown": "I trained 3 different architectures. All showed the same problem. 0.20-0.27 on the validation set, 0.10-0.16 on the test. I have also checked the results on different checkpoints for all models - the results were correlated (got to a peak together and started going down on overfitting together), so it's definitively not overfitting but big difference between the training set and the test. I feel that there is an inherent problem here of data that came from two different sources.",
      "votes": null
    },
    {
      "id": "391275",
      "postDate": "09/21/2018 13:57:32",
      "content": "<p><a href=\"/moshel\">@moshel</a> Are you using a holdout validation? Or something like k-fold? Is it the same validation set for each run?</p>",
      "rawMarkdown": "moshel Are you using a holdout validation? Or something like k-fold? Is it the same validation set for each run?",
      "votes": null
    },
    {
      "id": "391276",
      "postDate": "09/21/2018 14:01:28",
      "content": "<p>Given the nature of what we're asked to predict and the small size of the test set, I would not trust it to be representative of anything.  Could \"inconsistencies\" be mostly the result of sampling error (and maybe leptokurtosis) on the test set rather than underlying differences from the training data?</p>",
      "rawMarkdown": "Given the nature of what we're asked to predict and the small size of the test set, I would not trust it to be representative of anything.  Could \"inconsistencies\" be mostly the result of sampling error (and maybe leptokurtosis) on the test set rather than underlying differences from the training data?",
      "votes": null
    },
    {
      "id": "391356",
      "postDate": "09/21/2018 17:14:13",
      "content": "<p>I have the same problem:\nLS: ~0.28 LB: ~0.13-0.17</p>\n\n<p>I suspect it's because train and test set was annotated in different time by different people. Or there is some problem/missunderstanding with scoring formula. The third possible problem is that there are many duplicates of images or patients in train set.</p>",
      "rawMarkdown": "I have the same problem:\nLS: ~0.28 LB: ~0.13-0.17\n\nI suspect it's because train and test set was annotated in different time by different people. Or there is some problem/missunderstanding with scoring formula. The third possible problem is that there are many duplicates of images or patients in train set.",
      "votes": null
    },
    {
      "id": "391429",
      "postDate": "09/21/2018 19:56:14",
      "content": "<p>I am not sure what half the things you say means, but when 4 different people that probably took very different approach report the same thing, i would say there is more than a sampling error here. Either this or we should really buy lottery tickets. </p>",
      "rawMarkdown": "I am not sure what half the things you say means, but when 4 different people that probably took very different approach report the same thing, i would say there is more than a sampling error here. Either this or we should really buy lottery tickets.",
      "votes": null
    },
    {
      "id": "391433",
      "postDate": "09/21/2018 20:04:15",
      "content": "<p>I feel that such a gap can occur from one of two reasons:</p>\n\n<ol>\n<li><p>Different annotation methodology. For example, between marking every little sign of pneumonia or marking just the most prominent area, sometimes encapsulating several \"clouds\" together.</p></li>\n<li><p>Different difficulty level. For example, in the Google object detection competition the test set was much much harder.</p></li>\n</ol>\n\n<p>The second option is fair game, the first is not. </p>",
      "rawMarkdown": "I feel that such a gap can occur from one of two reasons:\n\n1. Different annotation methodology. For example, between marking every little sign of pneumonia or marking just the most prominent area, sometimes encapsulating several \"clouds\" together.\n\n2. Different difficulty level. For example, in the Google object detection competition the test set was much much harder.\n\nThe second option is fair game, the first is not.",
      "votes": null
    },
    {
      "id": "391453",
      "postDate": "09/21/2018 21:06:36",
      "content": "<p>There will undoubtedly be some difference due to having multiple annotators on the test set, but I'm not convinced that's the main source of differences that people are reporting in their scores.  A couple of things that would be useful in sorting this out:</p>\n\n<ol>\n<li>A comparison of results from different folds (and sub-folds) to see if the variance among folds (and sub-folds) is comparable to the variance between validation and test data.  (I say \"sub-folds\" because the public test set has only 1000 images, whereas even a 10-fold cross-validation will give you around 2500 images per validation fold.  I'm thinking maybe use 5 folds, divide each into 5 sub-folds, and look at variation with the exact same fitted model across sub-folds, as well as the different fits across folds.)</li>\n<li>A model to distinguish test images from training images.  If such a model can be developed and does well, that shows that there are predictable differences between training and test sets.  And the model can also be used to choose which training cases will be most useful for validation.</li>\n</ol>\n\n<p>FWIW \"leptokurtosis\" = long tails in a statistical distribution, so that you get a lot of cases that are similar and a few that are very different, but not so many that are moderately different (like with the stock market, you have lots of days with small moves and a few with crashes or spikes, but not so many with moderate-sized moves). That makes it hard to draw conclusions from a small sample, because any small sample might be missing the extreme cases (\"black swans\"), which are common enough to matter but rare enough that you won't necessarily see them unless you have a very large sample.  (Also, a particular small sample might <em>contain</em> some extreme cases and be overly influenced by them, leading to wrong conclusions.  With financial data, for example, you'll often get completely different conclusions depending on whether your sample includes 2008.)</p>",
      "rawMarkdown": "There will undoubtedly be some difference due to having multiple annotators on the test set, but I'm not convinced that's the main source of differences that people are reporting in their scores.  A couple of things that would be useful in sorting this out:\n\n 1. A comparison of results from different folds (and sub-folds) to see if the variance among folds (and sub-folds) is comparable to the variance between validation and test data.  (I say \"sub-folds\" because the public test set has only 1000 images, whereas even a 10-fold cross-validation will give you around 2500 images per validation fold.  I'm thinking maybe use 5 folds, divide each into 5 sub-folds, and look at variation with the exact same fitted model across sub-folds, as well as the different fits across folds.)\n 2. A model to distinguish test images from training images.  If such a model can be developed and does well, that shows that there are predictable differences between training and test sets.  And the model can also be used to choose which training cases will be most useful for validation.\n\nFWIW \"leptokurtosis\" = long tails in a statistical distribution, so that you get a lot of cases that are similar and a few that are very different, but not so many that are moderately different (like with the stock market, you have lots of days with small moves and a few with crashes or spikes, but not so many with moderate-sized moves). That makes it hard to draw conclusions from a small sample, because any small sample might be missing the extreme cases (\"black swans\"), which are common enough to matter but rare enough that you won't necessarily see them unless you have a very large sample.  (Also, a particular small sample might _contain_ some extreme cases and be overly influenced by them, leading to wrong conclusions.  With financial data, for example, you'll often get completely different conclusions depending on whether your sample includes 2008.)",
      "votes": null
    },
    {
      "id": "391459",
      "postDate": "09/21/2018 21:25:37",
      "content": "<p>I agree that we CAN develop methods to bypass the way the sets were divided. I am just saying this is not in the spirit of the challenge.\nThe difference is not within any standard deviation. 0.27 compared to 0.15 is not reasonable. its almost twice. Again, I trained three different models, on two splits, with about the same results. The models all \"see\" things differently (yolo, retinanet and frcnn) so getting consistent results across all models points out to test and train being annotated by different people with different methodology in the annotation. I thing the hosts didn't do a very good job at making the sets, and it makes our life much more difficult for no ML reason.</p>\n\n<p>Having no \"valid\" validation just makes us use the test scoring as our indication. meaning we have to submit a lot and adjust. Using this strategy does not make a good general prediction model. It does not even make a good model. So whats the point?</p>",
      "rawMarkdown": "I agree that we CAN develop methods to bypass the way the sets were divided. I am just saying this is not in the spirit of the challenge.\nThe difference is not within any standard deviation. 0.27 compared to 0.15 is not reasonable. its almost twice. Again, I trained three different models, on two splits, with about the same results. The models all \"see\" things differently (yolo, retinanet and frcnn) so getting consistent results across all models points out to test and train being annotated by different people with different methodology in the annotation. I thing the hosts didn't do a very good job at making the sets, and it makes our life much more difficult for no ML reason.\n\nHaving no \"valid\" validation just makes us use the test scoring as our indication. meaning we have to submit a lot and adjust. Using this strategy does not make a good general prediction model. It does not even make a good model. So whats the point?",
      "votes": null
    },
    {
      "id": "391470",
      "postDate": "09/21/2018 22:01:21",
      "content": "<p>I get LB scores ranging from 0.07 to 0.14 from running almost identical models (even most of that from running exactly identical models without setting seeds), which suggests to me that test set sampling error is a big issue.  I agree that 0.27 vs. 0.15 is not reasonable, but I still think a lot of the 0.15 can potentially be explained by bad luck in which images were chosen as stage 1 test images, rather than systematic train-vs-test differences.</p>",
      "rawMarkdown": "I get LB scores ranging from 0.07 to 0.14 from running almost identical models (even most of that from running exactly identical models without setting seeds), which suggests to me that test set sampling error is a big issue.  I agree that 0.27 vs. 0.15 is not reasonable, but I still think a lot of the 0.15 can potentially be explained by bad luck in which images were chosen as stage 1 test images, rather than systematic train-vs-test differences.",
      "votes": null
    },
    {
      "id": "391471",
      "postDate": "09/21/2018 22:08:28",
      "content": "<p>@Andy, I have the same experience as you i.e. getting 0.100 to 0.124 from running the same model and could not figure out a way to make my results consistent for any of my models. Even fixing the seed did not help in my case. Anybody have an idea other than fixing the seeds?</p>",
      "rawMarkdown": "Andy, I have the same experience as you i.e. getting 0.100 to 0.124 from running the same model and could not figure out a way to make my results consistent for any of my models. Even fixing the seed did not help in my case. Anybody have an idea other than fixing the seeds?",
      "votes": null
    },
    {
      "id": "391472",
      "postDate": "09/21/2018 22:13:59",
      "content": "<p>I would not worry about stage 1 LB. Since the set is so small it can't be representative of the stage 2 set(hopefully much larger). Part of the challenge might be trusting CV instead of the test set?</p>",
      "rawMarkdown": "I would not worry about stage 1 LB. Since the set is so small it can't be representative of the stage 2 set(hopefully much larger). Part of the challenge might be trusting CV instead of the test set?",
      "votes": null
    },
    {
      "id": "391473",
      "postDate": "09/21/2018 22:25:23",
      "content": "<p><a href=\"/sheriytm\">@sheriytm</a> I tend to agree with <a href=\"/arpandhatt\">@arpandhatt</a> that we shouldn't worry a whole lot about stage 1 LB scores. But even getting a robust CV is expensive.  I get a lot of variation among folds and among different fold assignments and even from running the same model on the same folds different times.  I think a reliable score for model selection needs to be a robust average (maybe trimmed mean) from multiple runs of the same model. That involves a lot of GPU hours.</p>",
      "rawMarkdown": "sheriytm I tend to agree with @arpandhatt that we shouldn't worry a whole lot about stage 1 LB scores. But even getting a robust CV is expensive.  I get a lot of variation among folds and among different fold assignments and even from running the same model on the same folds different times.  I think a reliable score for model selection needs to be a robust average (maybe trimmed mean) from multiple runs of the same model. That involves a lot of GPU hours.",
      "votes": null
    },
    {
      "id": "391477",
      "postDate": "09/21/2018 22:34:44",
      "content": "<p>These discrepancies between train, validation and test are not surprising given the differences in annotation is not surprising.  <a href=\"/anoukstein\">@anoukstein</a> nicely summarized the dataset annotation here: <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723</a>. Hopefully things will even out in the next phases of the competition. Thank you all for your engaging commentary and questions.  </p>",
      "rawMarkdown": "These discrepancies between train, validation and test are not surprising given the differences in annotation is not surprising.  @anoukstein nicely summarized the dataset annotation here: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723. Hopefully things will even out in the next phases of the competition. Thank you all for your engaging commentary and questions.",
      "votes": null
    },
    {
      "id": "391478",
      "postDate": "09/21/2018 22:37:27",
      "content": "<p>I really don't follow your reasoning. 4 people get CONSISTENT gap of about 0.15 between LB and validation across different models and different splits. It has nothing to do with your fluctuation of score due to seed, model difference or model training regime difference. This is not statistically possible. it DOES NOT stem from seeds etc. Multiple runs of the same (properly trained) model will result in (maybe, if you are lucky) 1% difference. If the stage 2 set is similar to stage 1, how come we should trust our CV? I truly do not follow.</p>",
      "rawMarkdown": "I really don't follow your reasoning. 4 people get CONSISTENT gap of about 0.15 between LB and validation across different models and different splits. It has nothing to do with your fluctuation of score due to seed, model difference or model training regime difference. This is not statistically possible. it DOES NOT stem from seeds etc. Multiple runs of the same (properly trained) model will result in (maybe, if you are lucky) 1% difference. If the stage 2 set is similar to stage 1, how come we should trust our CV? I truly do not follow.",
      "votes": null
    },
    {
      "id": "391479",
      "postDate": "09/21/2018 22:41:25",
      "content": "<p>Thanks @Andy Harless. I have not tried averaging several runs of the same model yet but that may be my next move. Like you said that would involve a lot of hours hence expensive, so I am not sure I should do that. However, that approach worked to stabilize my CV in the TGS contest.</p>",
      "rawMarkdown": "Thanks @Andy Harless. I have not tried averaging several runs of the same model yet but that may be my next move. Like you said that would involve a lot of hours hence expensive, so I am not sure I should do that. However, that approach worked to stabilize my CV in the TGS contest.",
      "votes": null
    },
    {
      "id": "391482",
      "postDate": "09/21/2018 22:49:26",
      "content": "<p>I have retrained one of my models from scratch 3 times. Got the same results. 2 times were with different split. </p>",
      "rawMarkdown": "I have retrained one of my models from scratch 3 times. Got the same results. 2 times were with different split.",
      "votes": null
    },
    {
      "id": "391497",
      "postDate": "09/21/2018 23:31:28",
      "content": "<p>If small differences in models produce large differences in LB scores, that indicates that the test set has very particular characteristics that are sensitive to small differences in predictions and could well be the result of random sample selection.  Those same particular characteristics, possibly resulting from random sample selection, could have the general effect that most models produce lower scores on this particular test set than on a more general test set.  (That could be considered a \"different difficulty level\" but not because the images were chosen to be more difficult, just because they randomly happened to select more \"difficult\" images.)  In general I would think the more stringent annotation procedure would make the test set less difficult, in that it would reduce the amount of noise.  And I don't think standard statistical rules of thumb (typically derived from normal distributions) work with this kind of data, where the chaotic nature of the annotation process and the will result in too many outliers.</p>",
      "rawMarkdown": "If small differences in models produce large differences in LB scores, that indicates that the test set has very particular characteristics that are sensitive to small differences in predictions and could well be the result of random sample selection.  Those same particular characteristics, possibly resulting from random sample selection, could have the general effect that most models produce lower scores on this particular test set than on a more general test set.  (That could be considered a \"different difficulty level\" but not because the images were chosen to be more difficult, just because they randomly happened to select more \"difficult\" images.)  In general I would think the more stringent annotation procedure would make the test set less difficult, in that it would reduce the amount of noise.  And I don't think standard statistical rules of thumb (typically derived from normal distributions) work with this kind of data, where the chaotic nature of the annotation process and the will result in too many outliers.",
      "votes": null
    },
    {
      "id": "392218",
      "postDate": "09/23/2018 09:29:39",
      "content": "<p>If this is so, it gets to the heart of the challenge: this is a challenge in Machine <strong>Learning</strong> and in order to learn it is crucial to have the same statistics on the training and on the test data. Perhaps the host could do something to fix that? For example publish the list of the 1500 studies in the training set that were annotated in the same methodology as the test data set. Otherwise we waste our time a little bit</p>",
      "rawMarkdown": "If this is so, it gets to the heart of the challenge: this is a challenge in Machine **Learning** and in order to learn it is crucial to have the same statistics on the training and on the test data. Perhaps the host could do something to fix that? For example publish the list of the 1500 studies in the training set that were annotated in the same methodology as the test data set. Otherwise we waste our time a little bit",
      "votes": null
    },
    {
      "id": "396298",
      "postDate": "09/30/2018 12:19:08",
      "content": "<p>Like Branden, my local results are reasonably consistent with the test set when using <a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">https://www.kaggle.com/chenyc15/mean-average-precision-metric</a>. Difference of approx 0.04 for most recent.</p>\n\n<p>Make sure that you are correctly penalizing for false positives when performing local validation: if you predict anything for an instance with no ground truth labels, you score a zero which must be included in the final mean.</p>\n\n<p>Also make sure you don't somehow give yourself points for correctly detecting nothing: if there are no ground truths, and no objects detected, the image is not included in the final mean calculation.</p>",
      "rawMarkdown": "Like Branden, my local results are reasonably consistent with the test set when using https://www.kaggle.com/chenyc15/mean-average-precision-metric. Difference of approx 0.04 for most recent.\n\nMake sure that you are correctly penalizing for false positives when performing local validation: if you predict anything for an instance with no ground truth labels, you score a zero which must be included in the final mean.\n\nAlso make sure you don't somehow give yourself points for correctly detecting nothing: if there are no ground truths, and no objects detected, the image is not included in the final mean calculation.",
      "votes": null
    },
    {
      "id": "403694",
      "postDate": "10/14/2018 10:58:43",
      "content": "<p>Hey, I tried different sets of validation set but still my result are inconsistent with LB score. I trained three models which gave validation score stating model A better than Model B better than Model C. However, after submitting results, It turn out to be model B better than Model A better than Model C. Also there was great difference between validation score of model A and model B. Model A was supposed to outperform but still model B gave better result.\nIs it anything wrong with my validation set ? - I tried to use different validation set but still getting same order of score</p>",
      "rawMarkdown": "Hey, I tried different sets of validation set but still my result are inconsistent with LB score. I trained three models which gave validation score stating model A better than Model B better than Model C. However, after submitting results, It turn out to be model B better than Model A better than Model C. Also there was great difference between validation score of model A and model B. Model A was supposed to outperform but still model B gave better result.\nIs it anything wrong with my validation set ? - I tried to use different validation set but still getting same order of score",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 390709,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "09/20/2018 17:19:48",
      "content": "<p>What kind of differences are you seeing between your local validation and your LB scores? Mine are within 0.01-0.03 which seems reasonable for a somewhat small test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 391056,
          "author_name": "hggshntr",
          "author_url": "",
          "post_date": "09/21/2018 07:35:15",
          "content": "<p>Thank you for your comment,\nAs @GuyE said, It is not only the large difference in the score (can get event up to 0.07 difference in the score), it is also the threshold taken on our classifier output. The maximal score is achieved  using different thresholds. This is why we suspect the statistics is different in the test set and the training set.</p>\n\n<p>I am actually a little bit surprised by what you are saying, since we never saw an agreement on the 0.01 level. We tested several different model and approaches, and always we got the large differences where the score in the LB is significantly lower than our internal score. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391107,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/21/2018 09:20:37",
          "content": "<p>I trained 3 different architectures. All showed the same problem. 0.20-0.27 on the validation set, 0.10-0.16 on the test. I have also checked the results on different checkpoints for all models - the results were correlated (got to a peak together and started going down on overfitting together), so it's definitively not overfitting but big difference between the training set and the test. I feel that there is an inherent problem here of data that came from two different sources.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391275,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "09/21/2018 13:57:32",
          "content": "<p><a href=\"/moshel\">@moshel</a> Are you using a holdout validation? Or something like k-fold? Is it the same validation set for each run?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 390742,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "09/20/2018 18:51:25",
      "content": "<p>I get 0.27 on validation, 0.15 on test. Consistently, across several models and architectures</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 390995,
      "author_name": "guyeng",
      "author_url": "",
      "post_date": "09/21/2018 06:03:41",
      "content": "<p>We get also above 0.21 on validation and anywhere from 0.1 to 0.16 on test depending on parameter tuning. For example, one of our models is a pneumonia classifier after which we run another model to extract bounding boxes. On validation we get the maximal score when we set the pneumonia threshold to 0.4. On test, when we lower the pneumonia threshold to 0.1 we get the maximal value with a variation of 0.05 between the 0.4 and and 0.1 threshold. Lowering the threshold to 0.1 obviously introduces more images of the test set on which the bounding box extractor acts, however since the thresholds are different between validation and test it seems to imply the statistics are also different between them. In a sense it seems like overfitting, however our validation and internal test were never learnt so the overfitting is with regards to the statistics of the train set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 391022,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/21/2018 06:37:45",
          "content": "<p>Yes, couldn't phrase it better myself. I actually triple checked that nothing from training sneaked into the validation, it was so pronounced</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 391276,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "09/21/2018 14:01:28",
      "content": "<p>Given the nature of what we're asked to predict and the small size of the test set, I would not trust it to be representative of anything.  Could \"inconsistencies\" be mostly the result of sampling error (and maybe leptokurtosis) on the test set rather than underlying differences from the training data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 391429,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/21/2018 19:56:14",
          "content": "<p>I am not sure what half the things you say means, but when 4 different people that probably took very different approach report the same thing, i would say there is more than a sampling error here. Either this or we should really buy lottery tickets. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391453,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "09/21/2018 21:06:36",
          "content": "<p>There will undoubtedly be some difference due to having multiple annotators on the test set, but I'm not convinced that's the main source of differences that people are reporting in their scores.  A couple of things that would be useful in sorting this out:</p>\n\n<ol>\n<li>A comparison of results from different folds (and sub-folds) to see if the variance among folds (and sub-folds) is comparable to the variance between validation and test data.  (I say \"sub-folds\" because the public test set has only 1000 images, whereas even a 10-fold cross-validation will give you around 2500 images per validation fold.  I'm thinking maybe use 5 folds, divide each into 5 sub-folds, and look at variation with the exact same fitted model across sub-folds, as well as the different fits across folds.)</li>\n<li>A model to distinguish test images from training images.  If such a model can be developed and does well, that shows that there are predictable differences between training and test sets.  And the model can also be used to choose which training cases will be most useful for validation.</li>\n</ol>\n\n<p>FWIW \"leptokurtosis\" = long tails in a statistical distribution, so that you get a lot of cases that are similar and a few that are very different, but not so many that are moderately different (like with the stock market, you have lots of days with small moves and a few with crashes or spikes, but not so many with moderate-sized moves). That makes it hard to draw conclusions from a small sample, because any small sample might be missing the extreme cases (\"black swans\"), which are common enough to matter but rare enough that you won't necessarily see them unless you have a very large sample.  (Also, a particular small sample might <em>contain</em> some extreme cases and be overly influenced by them, leading to wrong conclusions.  With financial data, for example, you'll often get completely different conclusions depending on whether your sample includes 2008.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391459,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/21/2018 21:25:37",
          "content": "<p>I agree that we CAN develop methods to bypass the way the sets were divided. I am just saying this is not in the spirit of the challenge.\nThe difference is not within any standard deviation. 0.27 compared to 0.15 is not reasonable. its almost twice. Again, I trained three different models, on two splits, with about the same results. The models all \"see\" things differently (yolo, retinanet and frcnn) so getting consistent results across all models points out to test and train being annotated by different people with different methodology in the annotation. I thing the hosts didn't do a very good job at making the sets, and it makes our life much more difficult for no ML reason.</p>\n\n<p>Having no \"valid\" validation just makes us use the test scoring as our indication. meaning we have to submit a lot and adjust. Using this strategy does not make a good general prediction model. It does not even make a good model. So whats the point?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391470,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "09/21/2018 22:01:21",
          "content": "<p>I get LB scores ranging from 0.07 to 0.14 from running almost identical models (even most of that from running exactly identical models without setting seeds), which suggests to me that test set sampling error is a big issue.  I agree that 0.27 vs. 0.15 is not reasonable, but I still think a lot of the 0.15 can potentially be explained by bad luck in which images were chosen as stage 1 test images, rather than systematic train-vs-test differences.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391471,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "09/21/2018 22:08:28",
          "content": "<p>@Andy, I have the same experience as you i.e. getting 0.100 to 0.124 from running the same model and could not figure out a way to make my results consistent for any of my models. Even fixing the seed did not help in my case. Anybody have an idea other than fixing the seeds?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391473,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "09/21/2018 22:25:23",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a> I tend to agree with <a href=\"/arpandhatt\">@arpandhatt</a> that we shouldn't worry a whole lot about stage 1 LB scores. But even getting a robust CV is expensive.  I get a lot of variation among folds and among different fold assignments and even from running the same model on the same folds different times.  I think a reliable score for model selection needs to be a robust average (maybe trimmed mean) from multiple runs of the same model. That involves a lot of GPU hours.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391478,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/21/2018 22:37:27",
          "content": "<p>I really don't follow your reasoning. 4 people get CONSISTENT gap of about 0.15 between LB and validation across different models and different splits. It has nothing to do with your fluctuation of score due to seed, model difference or model training regime difference. This is not statistically possible. it DOES NOT stem from seeds etc. Multiple runs of the same (properly trained) model will result in (maybe, if you are lucky) 1% difference. If the stage 2 set is similar to stage 1, how come we should trust our CV? I truly do not follow.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391479,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "09/21/2018 22:41:25",
          "content": "<p>Thanks @Andy Harless. I have not tried averaging several runs of the same model yet but that may be my next move. Like you said that would involve a lot of hours hence expensive, so I am not sure I should do that. However, that approach worked to stabilize my CV in the TGS contest.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391482,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/21/2018 22:49:26",
          "content": "<p>I have retrained one of my models from scratch 3 times. Got the same results. 2 times were with different split. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391497,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "09/21/2018 23:31:28",
          "content": "<p>If small differences in models produce large differences in LB scores, that indicates that the test set has very particular characteristics that are sensitive to small differences in predictions and could well be the result of random sample selection.  Those same particular characteristics, possibly resulting from random sample selection, could have the general effect that most models produce lower scores on this particular test set than on a more general test set.  (That could be considered a \"different difficulty level\" but not because the images were chosen to be more difficult, just because they randomly happened to select more \"difficult\" images.)  In general I would think the more stringent annotation procedure would make the test set less difficult, in that it would reduce the amount of noise.  And I don't think standard statistical rules of thumb (typically derived from normal distributions) work with this kind of data, where the chaotic nature of the annotation process and the will result in too many outliers.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 391356,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "09/21/2018 17:14:13",
      "content": "<p>I have the same problem:\nLS: ~0.28 LB: ~0.13-0.17</p>\n\n<p>I suspect it's because train and test set was annotated in different time by different people. Or there is some problem/missunderstanding with scoring formula. The third possible problem is that there are many duplicates of images or patients in train set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 391433,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/21/2018 20:04:15",
          "content": "<p>I feel that such a gap can occur from one of two reasons:</p>\n\n<ol>\n<li><p>Different annotation methodology. For example, between marking every little sign of pneumonia or marking just the most prominent area, sometimes encapsulating several \"clouds\" together.</p></li>\n<li><p>Different difficulty level. For example, in the Google object detection competition the test set was much much harder.</p></li>\n</ol>\n\n<p>The second option is fair game, the first is not. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 391472,
      "author_name": "arpandhatt",
      "author_url": "",
      "post_date": "09/21/2018 22:13:59",
      "content": "<p>I would not worry about stage 1 LB. Since the set is so small it can't be representative of the stage 2 set(hopefully much larger). Part of the challenge might be trusting CV instead of the test set?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 391477,
      "author_name": "safwanhalabi",
      "author_url": "",
      "post_date": "09/21/2018 22:34:44",
      "content": "<p>These discrepancies between train, validation and test are not surprising given the differences in annotation is not surprising.  <a href=\"/anoukstein\">@anoukstein</a> nicely summarized the dataset annotation here: <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723</a>. Hopefully things will even out in the next phases of the competition. Thank you all for your engaging commentary and questions.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 392218,
          "author_name": "hadarpo",
          "author_url": "",
          "post_date": "09/23/2018 09:29:39",
          "content": "<p>If this is so, it gets to the heart of the challenge: this is a challenge in Machine <strong>Learning</strong> and in order to learn it is crucial to have the same statistics on the training and on the test data. Perhaps the host could do something to fix that? For example publish the list of the 1500 studies in the training set that were annotated in the same methodology as the test data set. Otherwise we waste our time a little bit</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 396298,
      "author_name": "taindow",
      "author_url": "",
      "post_date": "09/30/2018 12:19:08",
      "content": "<p>Like Branden, my local results are reasonably consistent with the test set when using <a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">https://www.kaggle.com/chenyc15/mean-average-precision-metric</a>. Difference of approx 0.04 for most recent.</p>\n\n<p>Make sure that you are correctly penalizing for false positives when performing local validation: if you predict anything for an instance with no ground truth labels, you score a zero which must be included in the final mean.</p>\n\n<p>Also make sure you don't somehow give yourself points for correctly detecting nothing: if there are no ground truths, and no objects detected, the image is not included in the final mean calculation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 403694,
          "author_name": "ronak555",
          "author_url": "",
          "post_date": "10/14/2018 10:58:43",
          "content": "<p>Hey, I tried different sets of validation set but still my result are inconsistent with LB score. I trained three models which gave validation score stating model A better than Model B better than Model C. However, after submitting results, It turn out to be model B better than Model A better than Model C. Also there was great difference between validation score of model A and model B. Model A was supposed to outperform but still model B gave better result.\nIs it anything wrong with my validation set ? - I tried to use different validation set but still getting same order of score</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "390585": "Hi all - and especially the organizers,\nAfter a month or so working in this competition and submitting a lot of submissions it is clear that the statistics of the test set is very different than the statistics of the training set. \n\nDuring our work we randomly split the training dataset to train, valid and test sets and see that when we calculate the submission score based on the LB metric (as given by https://www.kaggle.com/chenyc15/mean-average-precision-metric) our results are consistent between the sets.\n\nBUT - when submitting we get different scores between our internal test set and the competition test set. Therefore it is clear that the test statistics is very different than the train statistics. It is very difficult to train models like this since we actually start to use the submission for model tuning - which is something you should never do.\n\nIt will be very helpful if the organizers can comment about the source of this discrepancy. One such source is the fact that three radiologists annotated each image in the test set, while each image in the train set was annotated by only one radiologist (as explained in https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723). Are there additional sources of inconsistency?\n\nCan we get the list of images in the train set (there should be 1500 of these) that were triple read?\n\nAny other comments will be helpful. \n\nThanks.",
    "390709": "What kind of differences are you seeing between your local validation and your LB scores? Mine are within 0.01-0.03 which seems reasonable for a somewhat small test set.",
    "390742": "I get 0.27 on validation, 0.15 on test. Consistently, across several models and architectures",
    "390995": "We get also above 0.21 on validation and anywhere from 0.1 to 0.16 on test depending on parameter tuning. For example, one of our models is a pneumonia classifier after which we run another model to extract bounding boxes. On validation we get the maximal score when we set the pneumonia threshold to 0.4. On test, when we lower the pneumonia threshold to 0.1 we get the maximal value with a variation of 0.05 between the 0.4 and and 0.1 threshold. Lowering the threshold to 0.1 obviously introduces more images of the test set on which the bounding box extractor acts, however since the thresholds are different between validation and test it seems to imply the statistics are also different between them. In a sense it seems like overfitting, however our validation and internal test were never learnt so the overfitting is with regards to the statistics of the train set.",
    "391022": "Yes, couldn't phrase it better myself. I actually triple checked that nothing from training sneaked into the validation, it was so pronounced",
    "391056": "Thank you for your comment,\nAs @GuyE said, It is not only the large difference in the score (can get event up to 0.07 difference in the score), it is also the threshold taken on our classifier output. The maximal score is achieved  using different thresholds. This is why we suspect the statistics is different in the test set and the training set.\n\nI am actually a little bit surprised by what you are saying, since we never saw an agreement on the 0.01 level. We tested several different model and approaches, and always we got the large differences where the score in the LB is significantly lower than our internal score.",
    "391107": "I trained 3 different architectures. All showed the same problem. 0.20-0.27 on the validation set, 0.10-0.16 on the test. I have also checked the results on different checkpoints for all models - the results were correlated (got to a peak together and started going down on overfitting together), so it's definitively not overfitting but big difference between the training set and the test. I feel that there is an inherent problem here of data that came from two different sources.",
    "391275": "moshel Are you using a holdout validation? Or something like k-fold? Is it the same validation set for each run?",
    "391276": "Given the nature of what we're asked to predict and the small size of the test set, I would not trust it to be representative of anything.  Could \"inconsistencies\" be mostly the result of sampling error (and maybe leptokurtosis) on the test set rather than underlying differences from the training data?",
    "391356": "I have the same problem:\nLS: ~0.28 LB: ~0.13-0.17\n\nI suspect it's because train and test set was annotated in different time by different people. Or there is some problem/missunderstanding with scoring formula. The third possible problem is that there are many duplicates of images or patients in train set.",
    "391429": "I am not sure what half the things you say means, but when 4 different people that probably took very different approach report the same thing, i would say there is more than a sampling error here. Either this or we should really buy lottery tickets.",
    "391433": "I feel that such a gap can occur from one of two reasons:\n\n1. Different annotation methodology. For example, between marking every little sign of pneumonia or marking just the most prominent area, sometimes encapsulating several \"clouds\" together.\n\n2. Different difficulty level. For example, in the Google object detection competition the test set was much much harder.\n\nThe second option is fair game, the first is not.",
    "391453": "There will undoubtedly be some difference due to having multiple annotators on the test set, but I'm not convinced that's the main source of differences that people are reporting in their scores.  A couple of things that would be useful in sorting this out:\n\n 1. A comparison of results from different folds (and sub-folds) to see if the variance among folds (and sub-folds) is comparable to the variance between validation and test data.  (I say \"sub-folds\" because the public test set has only 1000 images, whereas even a 10-fold cross-validation will give you around 2500 images per validation fold.  I'm thinking maybe use 5 folds, divide each into 5 sub-folds, and look at variation with the exact same fitted model across sub-folds, as well as the different fits across folds.)\n 2. A model to distinguish test images from training images.  If such a model can be developed and does well, that shows that there are predictable differences between training and test sets.  And the model can also be used to choose which training cases will be most useful for validation.\n\nFWIW \"leptokurtosis\" = long tails in a statistical distribution, so that you get a lot of cases that are similar and a few that are very different, but not so many that are moderately different (like with the stock market, you have lots of days with small moves and a few with crashes or spikes, but not so many with moderate-sized moves). That makes it hard to draw conclusions from a small sample, because any small sample might be missing the extreme cases (\"black swans\"), which are common enough to matter but rare enough that you won't necessarily see them unless you have a very large sample.  (Also, a particular small sample might _contain_ some extreme cases and be overly influenced by them, leading to wrong conclusions.  With financial data, for example, you'll often get completely different conclusions depending on whether your sample includes 2008.)",
    "391459": "I agree that we CAN develop methods to bypass the way the sets were divided. I am just saying this is not in the spirit of the challenge.\nThe difference is not within any standard deviation. 0.27 compared to 0.15 is not reasonable. its almost twice. Again, I trained three different models, on two splits, with about the same results. The models all \"see\" things differently (yolo, retinanet and frcnn) so getting consistent results across all models points out to test and train being annotated by different people with different methodology in the annotation. I thing the hosts didn't do a very good job at making the sets, and it makes our life much more difficult for no ML reason.\n\nHaving no \"valid\" validation just makes us use the test scoring as our indication. meaning we have to submit a lot and adjust. Using this strategy does not make a good general prediction model. It does not even make a good model. So whats the point?",
    "391470": "I get LB scores ranging from 0.07 to 0.14 from running almost identical models (even most of that from running exactly identical models without setting seeds), which suggests to me that test set sampling error is a big issue.  I agree that 0.27 vs. 0.15 is not reasonable, but I still think a lot of the 0.15 can potentially be explained by bad luck in which images were chosen as stage 1 test images, rather than systematic train-vs-test differences.",
    "391471": "Andy, I have the same experience as you i.e. getting 0.100 to 0.124 from running the same model and could not figure out a way to make my results consistent for any of my models. Even fixing the seed did not help in my case. Anybody have an idea other than fixing the seeds?",
    "391472": "I would not worry about stage 1 LB. Since the set is so small it can't be representative of the stage 2 set(hopefully much larger). Part of the challenge might be trusting CV instead of the test set?",
    "391473": "sheriytm I tend to agree with @arpandhatt that we shouldn't worry a whole lot about stage 1 LB scores. But even getting a robust CV is expensive.  I get a lot of variation among folds and among different fold assignments and even from running the same model on the same folds different times.  I think a reliable score for model selection needs to be a robust average (maybe trimmed mean) from multiple runs of the same model. That involves a lot of GPU hours.",
    "391477": "These discrepancies between train, validation and test are not surprising given the differences in annotation is not surprising.  @anoukstein nicely summarized the dataset annotation here: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723. Hopefully things will even out in the next phases of the competition. Thank you all for your engaging commentary and questions.",
    "391478": "I really don't follow your reasoning. 4 people get CONSISTENT gap of about 0.15 between LB and validation across different models and different splits. It has nothing to do with your fluctuation of score due to seed, model difference or model training regime difference. This is not statistically possible. it DOES NOT stem from seeds etc. Multiple runs of the same (properly trained) model will result in (maybe, if you are lucky) 1% difference. If the stage 2 set is similar to stage 1, how come we should trust our CV? I truly do not follow.",
    "391479": "Thanks @Andy Harless. I have not tried averaging several runs of the same model yet but that may be my next move. Like you said that would involve a lot of hours hence expensive, so I am not sure I should do that. However, that approach worked to stabilize my CV in the TGS contest.",
    "391482": "I have retrained one of my models from scratch 3 times. Got the same results. 2 times were with different split.",
    "391497": "If small differences in models produce large differences in LB scores, that indicates that the test set has very particular characteristics that are sensitive to small differences in predictions and could well be the result of random sample selection.  Those same particular characteristics, possibly resulting from random sample selection, could have the general effect that most models produce lower scores on this particular test set than on a more general test set.  (That could be considered a \"different difficulty level\" but not because the images were chosen to be more difficult, just because they randomly happened to select more \"difficult\" images.)  In general I would think the more stringent annotation procedure would make the test set less difficult, in that it would reduce the amount of noise.  And I don't think standard statistical rules of thumb (typically derived from normal distributions) work with this kind of data, where the chaotic nature of the annotation process and the will result in too many outliers.",
    "392218": "If this is so, it gets to the heart of the challenge: this is a challenge in Machine **Learning** and in order to learn it is crucial to have the same statistics on the training and on the test data. Perhaps the host could do something to fix that? For example publish the list of the 1500 studies in the training set that were annotated in the same methodology as the test data set. Otherwise we waste our time a little bit",
    "396298": "Like Branden, my local results are reasonably consistent with the test set when using https://www.kaggle.com/chenyc15/mean-average-precision-metric. Difference of approx 0.04 for most recent.\n\nMake sure that you are correctly penalizing for false positives when performing local validation: if you predict anything for an instance with no ground truth labels, you score a zero which must be included in the final mean.\n\nAlso make sure you don't somehow give yourself points for correctly detecting nothing: if there are no ground truths, and no objects detected, the image is not included in the final mean calculation.",
    "403694": "Hey, I tried different sets of validation set but still my result are inconsistent with LB score. I trained three models which gave validation score stating model A better than Model B better than Model C. However, after submitting results, It turn out to be model B better than Model A better than Model C. Also there was great difference between validation score of model A and model B. Model A was supposed to outperform but still model B gave better result.\nIs it anything wrong with my validation set ? - I tried to use different validation set but still getting same order of score"
  },
  "source": "meta"
}