{
  "id": 69678,
  "title": "I knew the train & test data were different, but...",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/69678",
  "author_name": "",
  "post_date": "2018-10-26T00:45:14.518646Z",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Dude.</p>\n\n<pre><code>import pandas as pd\nimport numpy as np\ndf1 = pd.read_csv('../input/stage_1_detailed_class_info.csv').drop_duplicates(keep='first')\ndf2 = pd.read_csv('../input/stage_2_detailed_class_info.csv').drop_duplicates(keep='first')\nstage_1_train_image_positives = (df1['class']=='Lung Opacity').sum()\nstage_1_train_image_count = df1.shape[0]\nstage_2_train_image_positives = (df2['class']=='Lung Opacity').sum()\nstage_2_train_image_count = df2.shape[0]\nstage_1_test_image_positives = stage_2_train_image_positives - stage_1_train_image_positives\nstage_1_test_image_count = stage_2_train_image_count - stage_1_train_image_count\nstage_1_train_image_postiive_rate = stage_1_train_image_positives / stage_1_train_image_count\nstage_1_test_image_postiive_rate = stage_1_test_image_positives / stage_1_test_image_count\nprint( \"Stage 1 train positive rate:  {0:.3f}\\nStage 1 test positive rate:   {1:.3f}\".format(\n      stage_1_train_image_postiive_rate, stage_1_test_image_postiive_rate) )\n</code></pre>\n\n<pre>Stage 1 train positive rate:  0.220\nStage 1 test positive rate:   0.353</pre>",
  "messages": [
    {
      "id": "410384",
      "postDate": "10/26/2018 00:45:14",
      "content": "<p>Dude.</p>\n\n<pre><code>import pandas as pd\nimport numpy as np\ndf1 = pd.read_csv('../input/stage_1_detailed_class_info.csv').drop_duplicates(keep='first')\ndf2 = pd.read_csv('../input/stage_2_detailed_class_info.csv').drop_duplicates(keep='first')\nstage_1_train_image_positives = (df1['class']=='Lung Opacity').sum()\nstage_1_train_image_count = df1.shape[0]\nstage_2_train_image_positives = (df2['class']=='Lung Opacity').sum()\nstage_2_train_image_count = df2.shape[0]\nstage_1_test_image_positives = stage_2_train_image_positives - stage_1_train_image_positives\nstage_1_test_image_count = stage_2_train_image_count - stage_1_train_image_count\nstage_1_train_image_postiive_rate = stage_1_train_image_positives / stage_1_train_image_count\nstage_1_test_image_postiive_rate = stage_1_test_image_positives / stage_1_test_image_count\nprint( \"Stage 1 train positive rate:  {0:.3f}\\nStage 1 test positive rate:   {1:.3f}\".format(\n      stage_1_train_image_postiive_rate, stage_1_test_image_postiive_rate) )\n</code></pre>\n\n<pre>Stage 1 train positive rate:  0.220\nStage 1 test positive rate:   0.353</pre>",
      "rawMarkdown": "Dude.\n\n    import pandas as pd\n    import numpy as np\n    df1 = pd.read_csv('../input/stage_1_detailed_class_info.csv').drop_duplicates(keep='first')\n    df2 = pd.read_csv('../input/stage_2_detailed_class_info.csv').drop_duplicates(keep='first')\n    stage_1_train_image_positives = (df1['class']=='Lung Opacity').sum()\n    stage_1_train_image_count = df1.shape[0]\n    stage_2_train_image_positives = (df2['class']=='Lung Opacity').sum()\n    stage_2_train_image_count = df2.shape[0]\n    stage_1_test_image_positives = stage_2_train_image_positives - stage_1_train_image_positives\n    stage_1_test_image_count = stage_2_train_image_count - stage_1_train_image_count\n    stage_1_train_image_postiive_rate = stage_1_train_image_positives / stage_1_train_image_count\n    stage_1_test_image_postiive_rate = stage_1_test_image_positives / stage_1_test_image_count\n    print( \"Stage 1 train positive rate:  {0:.3f}\\nStage 1 test positive rate:   {1:.3f}\".format(\n          stage_1_train_image_postiive_rate, stage_1_test_image_postiive_rate) )\n\n<pre>Stage 1 train positive rate:  0.220\nStage 1 test positive rate:   0.353</pre>",
      "votes": null
    },
    {
      "id": "410401",
      "postDate": "10/26/2018 01:59:54",
      "content": "<p>Weren't you the one arguing it was our mistake? Repenting? </p>",
      "rawMarkdown": "Weren't you the one arguing it was our mistake? Repenting?",
      "votes": null
    },
    {
      "id": "410408",
      "postDate": "10/26/2018 02:07:26",
      "content": "<p>Btw, this is a meaningless difference. For all we know all the test set could have been positives or negwtives. No one expects the test set to have similar composition to the train set. </p>",
      "rawMarkdown": "Btw, this is a meaningless difference. For all we know all the test set could have been positives or negwtives. No one expects the test set to have similar composition to the train set.",
      "votes": null
    },
    {
      "id": "410410",
      "postDate": "10/26/2018 02:15:29",
      "content": "<p>You were right. The difference is too large to be explained by sampling variation. I started to suspect this when I started looking at classification rather than segmentation. The statistics for classification are more straightforward, and most of my models were suggesting the rate of positives was much higher in the test data, to a degree that was probably statistically significant. Even then, I didn't realize how large.</p>",
      "rawMarkdown": "You were right. The difference is too large to be explained by sampling variation. I started to suspect this when I started looking at classification rather than segmentation. The statistics for classification are more straightforward, and most of my models were suggesting the rate of positives was much higher in the test data, to a degree that was probably statistically significant. Even then, I didn't realize how large.",
      "votes": null
    },
    {
      "id": "410411",
      "postDate": "10/26/2018 02:26:35",
      "content": "<p>Well, the test set could have had a biased selection of images (though I don't think that's a good practice for competition setups), but my models didn't find anything obviously special about the selection of images. They just found that higher probability thresholds worked better on validation sets than on the test set.  Which leaves us still uncertain, because we don't know how much of the difference was due to chance, how much was due to differences in annotation procedures, how much was due to biased selections when our models just weren't good enough to find the difference, and how much was due to other as yet unknown factors.</p>",
      "rawMarkdown": "Well, the test set could have had a biased selection of images (though I don't think that's a good practice for competition setups), but my models didn't find anything obviously special about the selection of images. They just found that higher probability thresholds worked better on validation sets than on the test set.  Which leaves us still uncertain, because we don't know how much of the difference was due to chance, how much was due to differences in annotation procedures, how much was due to biased selections when our models just weren't good enough to find the difference, and how much was due to other as yet unknown factors.",
      "votes": null
    },
    {
      "id": "410452",
      "postDate": "10/26/2018 04:57:51",
      "content": "<p>Ouch, LB probing must have hurt.</p>",
      "rawMarkdown": "Ouch, LB probing must have hurt.",
      "votes": null
    },
    {
      "id": "410577",
      "postDate": "10/26/2018 09:42:10",
      "content": "<p>Simple model landmark 0.353 has been proven to be a leak of this rate. LOL</p>",
      "rawMarkdown": "Simple model landmark 0.353 has been proven to be a leak of this rate. LOL",
      "votes": null
    },
    {
      "id": "410649",
      "postDate": "10/26/2018 12:00:26",
      "content": "<p>It seems that the Stage2 test set is like the Stage1 test set. So I think LB probing (at Stage 1) is ok.</p>",
      "rawMarkdown": "It seems that the Stage2 test set is like the Stage1 test set. So I think LB probing (at Stage 1) is ok.",
      "votes": null
    },
    {
      "id": "410790",
      "postDate": "10/26/2018 17:01:13",
      "content": "<p>how do you know this?</p>",
      "rawMarkdown": "how do you know this?",
      "votes": null
    },
    {
      "id": "414183",
      "postDate": "11/02/2018 09:23:49",
      "content": "<p>I don't know how one would have known this before getting the final results, but it does appear that the stage 2 test set was like stage 1. The classification thresholds used in our final submission (for both stages) were based on making a subjective adjustment to our highest scoring submission on the stage 1 LB, to account for the likelihood of being overfit (a partial concession to the CV, which wanted much higher thresholds).  In the end, it turns out that the thresholds used in the original highest-scoring submission would have produced a slightly higher score in stage 2.  It seems Ohad Silbert was quite right <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/66323\">to complain</a> about inconsistencies between the train and test set.</p>",
      "rawMarkdown": "I don't know how one would have known this before getting the final results, but it does appear that the stage 2 test set was like stage 1. The classification thresholds used in our final submission (for both stages) were based on making a subjective adjustment to our highest scoring submission on the stage 1 LB, to account for the likelihood of being overfit (a partial concession to the CV, which wanted much higher thresholds).  In the end, it turns out that the thresholds used in the original highest-scoring submission would have produced a slightly higher score in stage 2.  It seems Ohad Silbert was quite right [to complain][1] about inconsistencies between the train and test set.\n\n [1]: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/66323",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 410401,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "10/26/2018 01:59:54",
      "content": "<p>Weren't you the one arguing it was our mistake? Repenting? </p>",
      "votes": null,
      "replies": [
        {
          "id": 410410,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "10/26/2018 02:15:29",
          "content": "<p>You were right. The difference is too large to be explained by sampling variation. I started to suspect this when I started looking at classification rather than segmentation. The statistics for classification are more straightforward, and most of my models were suggesting the rate of positives was much higher in the test data, to a degree that was probably statistically significant. Even then, I didn't realize how large.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 410408,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "10/26/2018 02:07:26",
      "content": "<p>Btw, this is a meaningless difference. For all we know all the test set could have been positives or negwtives. No one expects the test set to have similar composition to the train set. </p>",
      "votes": null,
      "replies": [
        {
          "id": 410411,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "10/26/2018 02:26:35",
          "content": "<p>Well, the test set could have had a biased selection of images (though I don't think that's a good practice for competition setups), but my models didn't find anything obviously special about the selection of images. They just found that higher probability thresholds worked better on validation sets than on the test set.  Which leaves us still uncertain, because we don't know how much of the difference was due to chance, how much was due to differences in annotation procedures, how much was due to biased selections when our models just weren't good enough to find the difference, and how much was due to other as yet unknown factors.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 410452,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "10/26/2018 04:57:51",
      "content": "<p>Ouch, LB probing must have hurt.</p>",
      "votes": null,
      "replies": [
        {
          "id": 410649,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "10/26/2018 12:00:26",
          "content": "<p>It seems that the Stage2 test set is like the Stage1 test set. So I think LB probing (at Stage 1) is ok.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410790,
          "author_name": "arvindmvepa",
          "author_url": "",
          "post_date": "10/26/2018 17:01:13",
          "content": "<p>how do you know this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414183,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "11/02/2018 09:23:49",
          "content": "<p>I don't know how one would have known this before getting the final results, but it does appear that the stage 2 test set was like stage 1. The classification thresholds used in our final submission (for both stages) were based on making a subjective adjustment to our highest scoring submission on the stage 1 LB, to account for the likelihood of being overfit (a partial concession to the CV, which wanted much higher thresholds).  In the end, it turns out that the thresholds used in the original highest-scoring submission would have produced a slightly higher score in stage 2.  It seems Ohad Silbert was quite right <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/66323\">to complain</a> about inconsistencies between the train and test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 410577,
      "author_name": "yhirano",
      "author_url": "",
      "post_date": "10/26/2018 09:42:10",
      "content": "<p>Simple model landmark 0.353 has been proven to be a leak of this rate. LOL</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "410384": "Dude.\n\n    import pandas as pd\n    import numpy as np\n    df1 = pd.read_csv('../input/stage_1_detailed_class_info.csv').drop_duplicates(keep='first')\n    df2 = pd.read_csv('../input/stage_2_detailed_class_info.csv').drop_duplicates(keep='first')\n    stage_1_train_image_positives = (df1['class']=='Lung Opacity').sum()\n    stage_1_train_image_count = df1.shape[0]\n    stage_2_train_image_positives = (df2['class']=='Lung Opacity').sum()\n    stage_2_train_image_count = df2.shape[0]\n    stage_1_test_image_positives = stage_2_train_image_positives - stage_1_train_image_positives\n    stage_1_test_image_count = stage_2_train_image_count - stage_1_train_image_count\n    stage_1_train_image_postiive_rate = stage_1_train_image_positives / stage_1_train_image_count\n    stage_1_test_image_postiive_rate = stage_1_test_image_positives / stage_1_test_image_count\n    print( \"Stage 1 train positive rate:  {0:.3f}\\nStage 1 test positive rate:   {1:.3f}\".format(\n          stage_1_train_image_postiive_rate, stage_1_test_image_postiive_rate) )\n\n<pre>Stage 1 train positive rate:  0.220\nStage 1 test positive rate:   0.353</pre>",
    "410401": "Weren't you the one arguing it was our mistake? Repenting?",
    "410408": "Btw, this is a meaningless difference. For all we know all the test set could have been positives or negwtives. No one expects the test set to have similar composition to the train set.",
    "410410": "You were right. The difference is too large to be explained by sampling variation. I started to suspect this when I started looking at classification rather than segmentation. The statistics for classification are more straightforward, and most of my models were suggesting the rate of positives was much higher in the test data, to a degree that was probably statistically significant. Even then, I didn't realize how large.",
    "410411": "Well, the test set could have had a biased selection of images (though I don't think that's a good practice for competition setups), but my models didn't find anything obviously special about the selection of images. They just found that higher probability thresholds worked better on validation sets than on the test set.  Which leaves us still uncertain, because we don't know how much of the difference was due to chance, how much was due to differences in annotation procedures, how much was due to biased selections when our models just weren't good enough to find the difference, and how much was due to other as yet unknown factors.",
    "410452": "Ouch, LB probing must have hurt.",
    "410577": "Simple model landmark 0.353 has been proven to be a leak of this rate. LOL",
    "410649": "It seems that the Stage2 test set is like the Stage1 test set. So I think LB probing (at Stage 1) is ok.",
    "410790": "how do you know this?",
    "414183": "I don't know how one would have known this before getting the final results, but it does appear that the stage 2 test set was like stage 1. The classification thresholds used in our final submission (for both stages) were based on making a subjective adjustment to our highest scoring submission on the stage 1 LB, to account for the likelihood of being overfit (a partial concession to the CV, which wanted much higher thresholds).  In the end, it turns out that the thresholds used in the original highest-scoring submission would have produced a slightly higher score in stage 2.  It seems Ohad Silbert was quite right [to complain][1] about inconsistencies between the train and test set.\n\n [1]: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/66323"
  },
  "source": "meta"
}