{
  "id": 75752,
  "title": "Get worse result when using cross-validation ensemble",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/75752",
  "author_name": "ChienYiChi",
  "post_date": "2018-12-26T04:24:29.079000",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I use multi-label stratification and 5 fold cross-validation.  I use the lowest validation loss model to make prediction and then  average the probabilty output of those 5 fold models before thresholding. I try it   on both resnet18 and resnet50. And both of them got worse Public LB score when using cross-validataion ensemble, worse than single fold model of them. \nAlso , I have try majority vote on it , still no good result\nBut I know most of teams got help from this cross-validation ensemble\nwhy is this weird situation happen? </p>",
  "messages": [
    {
      "id": 447509,
      "postDate": "2018-12-30T01:20:13.767Z",
      "content": "<p>Some tips:</p>\n\n<ol>\n<li>Threshold value for ensemble is usually lower than a single fold. This is because you have averaged from 5 models. So prediction confidence usually gets lower and your threshold should be lower too.</li>\n<li>For majority voting, try different voting option also. like 2/5, 3/5 etc. </li>\n</ol>\n\n<p>Hope this helps.</p>",
      "rawMarkdown": "Some tips:\n\n1. Threshold value for ensemble is usually lower than a single fold. This is because you have averaged from 5 models. So prediction confidence usually gets lower and your threshold should be lower too.\n2. For majority voting, try different voting option also. like 2/5, 3/5 etc. \n\nHope this helps.",
      "votes": 5
    },
    {
      "id": 445261,
      "postDate": "2018-12-26T04:24:29.080Z",
      "content": "<p>I use multi-label stratification and 5 fold cross-validation.  I use the lowest validation loss model to make prediction and then  average the probabilty output of those 5 fold models before thresholding. I try it   on both resnet18 and resnet50. And both of them got worse Public LB score when using cross-validataion ensemble, worse than single fold model of them. \nAlso , I have try majority vote on it , still no good result\nBut I know most of teams got help from this cross-validation ensemble\nwhy is this weird situation happen? </p>",
      "rawMarkdown": "I use multi-label stratification and 5 fold cross-validation.  I use the lowest validation loss model to make prediction and then  average the probabilty output of those 5 fold models before thresholding. I try it   on both resnet18 and resnet50. And both of them got worse Public LB score when using cross-validataion ensemble, worse than single fold model of them. \nAlso , I have try majority vote on it , still no good result\nBut I know most of teams got help from this cross-validation ensemble\nwhy is this weird situation happen? ",
      "votes": 1
    },
    {
      "id": 448201,
      "postDate": "2018-12-31T12:40:46.370Z",
      "content": "<p>Do you use a fixed threshold? Or do you probe the optimal threshold for each class in each fold? For my current setup this works very good...I average the probability output of each of the folds..... and I average the optimal thresholds found for each fold. I get an increase on the LB of 0.025 - 0.030. Not sure how that will work out on the Private Board...but if we average the probabilities to get a more generalized output then why not also average the thresholds?</p>",
      "rawMarkdown": "Do you use a fixed threshold? Or do you probe the optimal threshold for each class in each fold? For my current setup this works very good...I average the probability output of each of the folds..... and I average the optimal thresholds found for each fold. I get an increase on the LB of 0.025 - 0.030. Not sure how that will work out on the Private Board...but if we average the probabilities to get a more generalized output then why not also average the thresholds?"
    },
    {
      "id": 446246,
      "postDate": "2018-12-27T18:58:36.273Z",
      "content": "<p>After you average the outputs, the proper threshold value may be changed.</p>",
      "rawMarkdown": "After you average the outputs, the proper threshold value may be changed."
    },
    {
      "id": 446088,
      "postDate": "2018-12-27T13:30:17.367Z",
      "content": "<p>This has happened with me as well. As of now, my single model predictions are performing best on LB and CV-ensemble couldn't even reach 0.4</p>",
      "rawMarkdown": "This has happened with me as well. As of now, my single model predictions are performing best on LB and CV-ensemble couldn't even reach 0.4"
    },
    {
      "id": 448218,
      "postDate": "2018-12-31T13:51:38.047Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 447509,
      "author_name": "Shai",
      "author_url": "",
      "post_date": "2018-12-30T01:20:13.767000",
      "content": "<p>Some tips:</p>\n\n<ol>\n<li>Threshold value for ensemble is usually lower than a single fold. This is because you have averaged from 5 models. So prediction confidence usually gets lower and your threshold should be lower too.</li>\n<li>For majority voting, try different voting option also. like 2/5, 3/5 etc. </li>\n</ol>\n\n<p>Hope this helps.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 448201,
      "author_name": "Robin Smits",
      "author_url": "",
      "post_date": "2018-12-31T12:40:46.370000",
      "content": "<p>Do you use a fixed threshold? Or do you probe the optimal threshold for each class in each fold? For my current setup this works very good...I average the probability output of each of the folds..... and I average the optimal thresholds found for each fold. I get an increase on the LB of 0.025 - 0.030. Not sure how that will work out on the Private Board...but if we average the probabilities to get a more generalized output then why not also average the thresholds?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446246,
      "author_name": "Ildoo Kim",
      "author_url": "",
      "post_date": "2018-12-27T18:58:36.273000",
      "content": "<p>After you average the outputs, the proper threshold value may be changed.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446088,
      "author_name": "Criminal Mind",
      "author_url": "",
      "post_date": "2018-12-27T13:30:17.367000",
      "content": "<p>This has happened with me as well. As of now, my single model predictions are performing best on LB and CV-ensemble couldn't even reach 0.4</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 448218,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-31T13:51:38.047000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "447509": "Some tips:\n\n1. Threshold value for ensemble is usually lower than a single fold. This is because you have averaged from 5 models. So prediction confidence usually gets lower and your threshold should be lower too.\n2. For majority voting, try different voting option also. like 2/5, 3/5 etc. \n\nHope this helps.",
    "445261": "I use multi-label stratification and 5 fold cross-validation.  I use the lowest validation loss model to make prediction and then  average the probabilty output of those 5 fold models before thresholding. I try it   on both resnet18 and resnet50. And both of them got worse Public LB score when using cross-validataion ensemble, worse than single fold model of them. \nAlso , I have try majority vote on it , still no good result\nBut I know most of teams got help from this cross-validation ensemble\nwhy is this weird situation happen? ",
    "448201": "Do you use a fixed threshold? Or do you probe the optimal threshold for each class in each fold? For my current setup this works very good...I average the probability output of each of the folds..... and I average the optimal thresholds found for each fold. I get an increase on the LB of 0.025 - 0.030. Not sure how that will work out on the Private Board...but if we average the probabilities to get a more generalized output then why not also average the thresholds?",
    "446246": "After you average the outputs, the proper threshold value may be changed.",
    "446088": "This has happened with me as well. As of now, my single model predictions are performing best on LB and CV-ensemble couldn't even reach 0.4",
    "448218": ""
  }
}