{
  "id": 69689,
  "title": "Parameter search using the stage 1 test labels must be promising",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/69689",
  "author_name": "",
  "post_date": "2018-10-26T03:03:32.592619400Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As many of you know, label distribution between train data and test data is different (or my scoring implementation is wrong). So searching the best parameters for post-processing must be promising. After Julia announced that the stage 1 test label will provided, I should have notice this idea in remained few hours . But unfortunately I had not and in my upload files, there is no parameter-search scripts. </p>\n\n<p>OMG I shouldn’t lose focus up to the end of the competition. The last day is the most important day in kaggle competitions. </p>\n\n<p>As described <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">here</a>, the most of stage 1 train labels are verified by two doctors and test data of stage 1 and 2 are verified by three doctor including certificated radiologist. Actually my local CV and public score shows big difference. Note that some high placed competitor said <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/66323\">here</a> that there is no big difference between train and test, so I may miss something important. Anyway, in my case, I use segmentation approach. The most difference is the threshold of there is opacity or not. It looks like many images regarded as no opacity in train label criteria are regarded as opacity in test label criteria. I use max values in predicted segmentation as confidence score. This features classified opacity existence in image with AUC 0.90 in local CV. I set threshold on this confidence score and mask candidates with confidence score under this threshold is removed. In local CV, the high threshold get good score but public score is low. Bounding box size may be also different, I doubt. </p>\n\n<p>Many other parameter search must be promissing using stage 1 test label (eg. blending rate of models). I will try this idea, but of course I can’t choose it as final submission. </p>\n\n<p>I will update this post with actual values of my experiment. </p>",
  "messages": [
    {
      "id": "410420",
      "postDate": "10/26/2018 03:03:32",
      "content": "<p>As many of you know, label distribution between train data and test data is different (or my scoring implementation is wrong). So searching the best parameters for post-processing must be promising. After Julia announced that the stage 1 test label will provided, I should have notice this idea in remained few hours . But unfortunately I had not and in my upload files, there is no parameter-search scripts. </p>\n\n<p>OMG I shouldn’t lose focus up to the end of the competition. The last day is the most important day in kaggle competitions. </p>\n\n<p>As described <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">here</a>, the most of stage 1 train labels are verified by two doctors and test data of stage 1 and 2 are verified by three doctor including certificated radiologist. Actually my local CV and public score shows big difference. Note that some high placed competitor said <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/66323\">here</a> that there is no big difference between train and test, so I may miss something important. Anyway, in my case, I use segmentation approach. The most difference is the threshold of there is opacity or not. It looks like many images regarded as no opacity in train label criteria are regarded as opacity in test label criteria. I use max values in predicted segmentation as confidence score. This features classified opacity existence in image with AUC 0.90 in local CV. I set threshold on this confidence score and mask candidates with confidence score under this threshold is removed. In local CV, the high threshold get good score but public score is low. Bounding box size may be also different, I doubt. </p>\n\n<p>Many other parameter search must be promissing using stage 1 test label (eg. blending rate of models). I will try this idea, but of course I can’t choose it as final submission. </p>\n\n<p>I will update this post with actual values of my experiment. </p>",
      "rawMarkdown": "As many of you know, label distribution between train data and test data is different (or my scoring implementation is wrong). So searching the best parameters for post-processing must be promising. After Julia announced that the stage 1 test label will provided, I should have notice this idea in remained few hours . But unfortunately I had not and in my upload files, there is no parameter-search scripts. \n\nOMG I shouldn’t lose focus up to the end of the competition. The last day is the most important day in kaggle competitions. \n\nAs described [here](https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723), the most of stage 1 train labels are verified by two doctors and test data of stage 1 and 2 are verified by three doctor including certificated radiologist. Actually my local CV and public score shows big difference. Note that some high placed competitor said [here](https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/66323) that there is no big difference between train and test, so I may miss something important. Anyway, in my case, I use segmentation approach. The most difference is the threshold of there is opacity or not. It looks like many images regarded as no opacity in train label criteria are regarded as opacity in test label criteria. I use max values in predicted segmentation as confidence score. This features classified opacity existence in image with AUC 0.90 in local CV. I set threshold on this confidence score and mask candidates with confidence score under this threshold is removed. In local CV, the high threshold get good score but public score is low. Bounding box size may be also different, I doubt. \n\nMany other parameter search must be promissing using stage 1 test label (eg. blending rate of models). I will try this idea, but of course I can’t choose it as final submission. \n\nI will update this post with actual values of my experiment.",
      "votes": null
    },
    {
      "id": "410428",
      "postDate": "10/26/2018 03:35:47",
      "content": "<p>I'm not sure that this would be allowed under the current rules.</p>\n\n<blockquote>\n  <p>We expect you may need to make some \"non scientific\" alterations, such\n  as changes to path names, in order to create your submissions for the\n  second stage. You are allowed to re-train your model (including the\n  stage one data), but your code should not change. You should not be\n  doing any hyper parameter tuning in the second stage. Parameter tuning\n  is permitted as long as it is fully automated.</p>\n</blockquote>\n\n<p>Searching for the best post-processing parameters (e.g., threshold for labeling an image as opacity) falls under hyperparameters for me, which we're not allowed to change. I might also be misunderstanding your post. </p>",
      "rawMarkdown": "I'm not sure that this would be allowed under the current rules.\n\n&gt; We expect you may need to make some \"non scientific\" alterations, such\n&gt; as changes to path names, in order to create your submissions for the\n&gt; second stage. You are allowed to re-train your model (including the\n&gt; stage one data), but your code should not change. You should not be\n&gt; doing any hyper parameter tuning in the second stage. Parameter tuning\n&gt; is permitted as long as it is fully automated.\n\nSearching for the best post-processing parameters (e.g., threshold for labeling an image as opacity) falls under hyperparameters for me, which we're not allowed to change. I might also be misunderstanding your post.",
      "votes": null
    },
    {
      "id": "410430",
      "postDate": "10/26/2018 03:46:27",
      "content": "<p>Leaving some train data and treating as pseudo test data for parameter search in stage 1 and replacing them with stage 1 test data later is reasonable approach, I think. </p>",
      "rawMarkdown": "Leaving some train data and treating as pseudo test data for parameter search in stage 1 and replacing them with stage 1 test data later is reasonable approach, I think.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 410428,
      "author_name": "vaillant",
      "author_url": "",
      "post_date": "10/26/2018 03:35:47",
      "content": "<p>I'm not sure that this would be allowed under the current rules.</p>\n\n<blockquote>\n  <p>We expect you may need to make some \"non scientific\" alterations, such\n  as changes to path names, in order to create your submissions for the\n  second stage. You are allowed to re-train your model (including the\n  stage one data), but your code should not change. You should not be\n  doing any hyper parameter tuning in the second stage. Parameter tuning\n  is permitted as long as it is fully automated.</p>\n</blockquote>\n\n<p>Searching for the best post-processing parameters (e.g., threshold for labeling an image as opacity) falls under hyperparameters for me, which we're not allowed to change. I might also be misunderstanding your post. </p>",
      "votes": null,
      "replies": [
        {
          "id": 410430,
          "author_name": "osciiart",
          "author_url": "",
          "post_date": "10/26/2018 03:46:27",
          "content": "<p>Leaving some train data and treating as pseudo test data for parameter search in stage 1 and replacing them with stage 1 test data later is reasonable approach, I think. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "410420": "As many of you know, label distribution between train data and test data is different (or my scoring implementation is wrong). So searching the best parameters for post-processing must be promising. After Julia announced that the stage 1 test label will provided, I should have notice this idea in remained few hours . But unfortunately I had not and in my upload files, there is no parameter-search scripts. \n\nOMG I shouldn’t lose focus up to the end of the competition. The last day is the most important day in kaggle competitions. \n\nAs described [here](https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723), the most of stage 1 train labels are verified by two doctors and test data of stage 1 and 2 are verified by three doctor including certificated radiologist. Actually my local CV and public score shows big difference. Note that some high placed competitor said [here](https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/66323) that there is no big difference between train and test, so I may miss something important. Anyway, in my case, I use segmentation approach. The most difference is the threshold of there is opacity or not. It looks like many images regarded as no opacity in train label criteria are regarded as opacity in test label criteria. I use max values in predicted segmentation as confidence score. This features classified opacity existence in image with AUC 0.90 in local CV. I set threshold on this confidence score and mask candidates with confidence score under this threshold is removed. In local CV, the high threshold get good score but public score is low. Bounding box size may be also different, I doubt. \n\nMany other parameter search must be promissing using stage 1 test label (eg. blending rate of models). I will try this idea, but of course I can’t choose it as final submission. \n\nI will update this post with actual values of my experiment.",
    "410428": "I'm not sure that this would be allowed under the current rules.\n\n&gt; We expect you may need to make some \"non scientific\" alterations, such\n&gt; as changes to path names, in order to create your submissions for the\n&gt; second stage. You are allowed to re-train your model (including the\n&gt; stage one data), but your code should not change. You should not be\n&gt; doing any hyper parameter tuning in the second stage. Parameter tuning\n&gt; is permitted as long as it is fully automated.\n\nSearching for the best post-processing parameters (e.g., threshold for labeling an image as opacity) falls under hyperparameters for me, which we're not allowed to change. I might also be misunderstanding your post.",
    "410430": "Leaving some train data and treating as pseudo test data for parameter search in stage 1 and replacing them with stage 1 test data later is reasonable approach, I think."
  },
  "source": "meta"
}