{
  "id": 117126,
  "title": "Wired ensemble boost",
  "url": "/competitions/understanding_cloud_organization/discussion/117126",
  "author_name": "",
  "post_date": "2019-11-13T13:47:11.080941800Z",
  "votes": 2,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Today I ensembled three of my models which trained long time ago(about 3 weeks). And the LB and CV score of these 3 models are less than 0.65, which is quite bad. \n(model 1 -&gt; 0.648, model 2 -&gt; 0.645, model 3 -&gt; 0.647) </p>\n\n<p>But after ensemble them, with 0.6 pixel threshold and 15000 mask size threshold. The LB score is 0.664..almost 0.02 boost.</p>\n\n<p>Do you guys have the similar situation? </p>",
  "messages": [
    {
      "id": "672075",
      "postDate": "11/13/2019 13:47:11",
      "content": "<p>Today I ensembled three of my models which trained long time ago(about 3 weeks). And the LB and CV score of these 3 models are less than 0.65, which is quite bad. \n(model 1 -&gt; 0.648, model 2 -&gt; 0.645, model 3 -&gt; 0.647) </p>\n\n<p>But after ensemble them, with 0.6 pixel threshold and 15000 mask size threshold. The LB score is 0.664..almost 0.02 boost.</p>\n\n<p>Do you guys have the similar situation? </p>",
      "rawMarkdown": "Today I ensembled three of my models which trained long time ago(about 3 weeks). And the LB and CV score of these 3 models are less than 0.65, which is quite bad. \n(model 1 -&gt; 0.648, model 2 -&gt; 0.645, model 3 -&gt; 0.647) \n\nBut after ensemble them, with 0.6 pixel threshold and 15000 mask size threshold. The LB score is 0.664..almost 0.02 boost.\n\nDo you guys have the similar situation?",
      "votes": null
    },
    {
      "id": "672090",
      "postDate": "11/13/2019 14:05:21",
      "content": "<p>But after ensemble them with my current best models, the score decrease.</p>",
      "rawMarkdown": "But after ensemble them with my current best models, the score decrease.",
      "votes": null
    },
    {
      "id": "672116",
      "postDate": "11/13/2019 14:23:47",
      "content": "<p>Something similar happened, my choice is to believe those models with better CV&amp;LB scores.</p>",
      "rawMarkdown": "Something similar happened, my choice is to believe those models with better CV&amp;LB scores.",
      "votes": null
    },
    {
      "id": "672196",
      "postDate": "11/13/2019 15:52:59",
      "content": "<p>same model, I could get .645~.655 with different thresholds...so weird</p>",
      "rawMarkdown": "same model, I could get .645~.655 with different thresholds...so weird",
      "votes": null
    },
    {
      "id": "672204",
      "postDate": "11/13/2019 16:11:47",
      "content": "<p>\"But after ensemble them, with 0.6 pixel threshold and 15000 mask size threshold. The LB score is 0.664..almost 0.02 boost.\"</p>\n\n<p>after ensemble and post processing by size, you end up with a list of negative (no mask) and positive prediction (mask).\nfix the negative prediction and leave it unchanged.</p>\n\n<p>Then, try to increase of decrease (mask erode or dilate) the positive prediction mask. maybe there will improvement. e.g. if the ground truth box is intersection of different  annotators, conservative prediction like erosion may be better. if the ground truth box is union, maybe dilation is better.</p>\n\n<p>try in on validation set first.</p>\n\n<p>since this operation does not change you negative prediction, you will not increase false positive</p>",
      "rawMarkdown": "\"But after ensemble them, with 0.6 pixel threshold and 15000 mask size threshold. The LB score is 0.664..almost 0.02 boost.\"\n\n\nafter ensemble and post processing by size, you end up with a list of negative (no mask) and positive prediction (mask).\nfix the negative prediction and leave it unchanged.\n\nThen, try to increase of decrease (mask erode or dilate) the positive prediction mask. maybe there will improvement. e.g. if the ground truth box is intersection of different  annotators, conservative prediction like erosion may be better. if the ground truth box is union, maybe dilation is better.\n\ntry in on validation set first.\n\nsince this operation does not change you negative prediction, you will not increase false positive",
      "votes": null
    },
    {
      "id": "672477",
      "postDate": "11/13/2019 23:47:09",
      "content": "<p>Since I did't spilt the data well at that time, so the CV score of these three models ensemble is unreliable. This make me hard to trust the ensemble result. But this just my case, I am sure yours is fine. Good luck with you!</p>",
      "rawMarkdown": "Since I did't spilt the data well at that time, so the CV score of these three models ensemble is unreliable. This make me hard to trust the ensemble result. But this just my case, I am sure yours is fine. Good luck with you!",
      "votes": null
    },
    {
      "id": "672480",
      "postDate": "11/13/2019 23:51:34",
      "content": "<p>Yes, I have the same problem.. \nThe score is impacted by the threshold so much. This make me really worry about the shake up..\nSo I tend to fix the threshold and just try the result by different model ensemble now.</p>",
      "rawMarkdown": "Yes, I have the same problem.. \nThe score is impacted by the threshold so much. This make me really worry about the shake up..\nSo I tend to fix the threshold and just try the result by different model ensemble now.",
      "votes": null
    },
    {
      "id": "672489",
      "postDate": "11/14/2019 00:06:34",
      "content": "<p>Thanks for the advice! \nI understand that this can remove some false positive for me. But 0.02 boost still make me hard to trust these three models. 😂 \nAlso since I did't fix data split on these 3 models(trained long time age), so I don't have the validation set to verify. I might just give up these 3 models..But they do make me worry about the shake up.</p>",
      "rawMarkdown": "Thanks for the advice! \nI understand that this can remove some false positive for me. But 0.02 boost still make me hard to trust these three models. 😂 \nAlso since I did't fix data split on these 3 models(trained long time age), so I don't have the validation set to verify. I might just give up these 3 models..But they do make me worry about the shake up.",
      "votes": null
    },
    {
      "id": "672519",
      "postDate": "11/14/2019 00:44:49",
      "content": "<p>\"But they do make me worry about the shake up.\"</p>\n\n<p>this is something i haven't think of how to deal with yet for now. it is also part of the challenge</p>\n\n<p>\"so I don't have the validation set to verify\"</p>\n\n<p>you can do without a validation set.</p>\n\n<ol>\n<li>assume you have train a classifier on a train A set and validation set to verify its performance.</li>\n<li>you have another train classifier B without validation set.</li>\n<li>what you need is a third set S not used to trained A or B. It is ok that this third set is unlabeled.</li>\n<li>now test A and B on S.</li>\n<li>perturb S. test  A and B again on S.</li>\n<li>you can measure the consistency of the results between A and B and consistency of their behavior on perturbation (i.e. how stable or sensitive their behavior is).</li>\n</ol>\n\n<hr>\n\n<p>in doubt, you can still try this:</p>\n\n<ol>\n<li>use the pseudo label of B of set S to pretrain a classifier C.</li>\n<li>then use the original train data (does not contains S) to finetune C.</li>\n</ol>\n\n<p>in step.1 the knowledge (unknown accuracy) of B is transferred to C. in step.2 good knowledge will be retained and bad knowledge will be corrected with the use of good data (i.e. the original train set with label)</p>\n\n<hr>\n\n<p>note: you can google for validation without validation set. you can also borrow the ideas from semi-supervised learning</p>",
      "rawMarkdown": "\"But they do make me worry about the shake up.\"\n\nthis is something i haven't think of how to deal with yet for now. it is also part of the challenge\n\n\"so I don't have the validation set to verify\"\n\nyou can do without a validation set.\n\n1.  assume you have train a classifier on a train A set and validation set to verify its performance.\n2. you have another train classifier B without validation set.\n3. what you need is a third set S not used to trained A or B. It is ok that this third set is unlabeled.\n4. now test A and B on S.\n5. perturb S. test  A and B again on S.\n6. you can measure the consistency of the results between A and B and consistency of their behavior on perturbation (i.e. how stable or sensitive their behavior is).\n\n----\n\nin doubt, you can still try this:\n\n1. use the pseudo label of B of set S to pretrain a classifier C.\n2. then use the original train data (does not contains S) to finetune C.\n\nin step.1 the knowledge (unknown accuracy) of B is transferred to C. in step.2 good knowledge will be retained and bad knowledge will be corrected with the use of good data (i.e. the original train set with label)\n\n---- \n\nnote: you can google for validation without validation set. you can also borrow the ideas from semi-supervised learning",
      "votes": null
    },
    {
      "id": "672536",
      "postDate": "11/14/2019 01:09:11",
      "content": "<p>Thanks a lot for the advice and knowledge, this is something I never knew before! It is really great to learn something new.\nAlthough I don't have enough time and hardware resource to verify them with your awesome advice, since I have something close to my final submission now. But I definitely gonna try it after the competition! Many thanks for the good advice, good luck with you!</p>",
      "rawMarkdown": "Thanks a lot for the advice and knowledge, this is something I never knew before! It is really great to learn something new.\nAlthough I don't have enough time and hardware resource to verify them with your awesome advice, since I have something close to my final submission now. But I definitely gonna try it after the competition! Many thanks for the good advice, good luck with you!",
      "votes": null
    },
    {
      "id": "672570",
      "postDate": "11/14/2019 02:00:07",
      "content": "<p>what's you ensemble method, vote, average on probability, or weighted sum?\nMy CV is not linear with LB, I'm also worry about the shake up.</p>",
      "rawMarkdown": "what's you ensemble method, vote, average on probability, or weighted sum?\nMy CV is not linear with LB, I'm also worry about the shake up.",
      "votes": null
    },
    {
      "id": "672586",
      "postDate": "11/14/2019 02:16:18",
      "content": "<p>I am using average on raw prediction as ensemble method. My local CV has about 0.002-0.004 difference with LB score. And also nonlinear with LB.\nSo I might make 2 final submission with \n1. Best LB \n2. Best CV</p>",
      "rawMarkdown": "I am using average on raw prediction as ensemble method. My local CV has about 0.002-0.004 difference with LB score. And also nonlinear with LB.\nSo I might make 2 final submission with \n1. Best LB \n2. Best CV",
      "votes": null
    },
    {
      "id": "672595",
      "postDate": "11/14/2019 02:24:43",
      "content": "<p>But I wont expect I will remain silver zone in PB.. I bet my score definitely  going to shake....</p>",
      "rawMarkdown": "But I wont expect I will remain silver zone in PB.. I bet my score definitely  going to shake....",
      "votes": null
    },
    {
      "id": "672604",
      "postDate": "11/14/2019 02:30:55",
      "content": "<p>My CV is also nonlinear with LB. 😂 \nSo I try to add something when CV and LB increase together...</p>",
      "rawMarkdown": "My CV is also nonlinear with LB. 😂 \nSo I try to add something when CV and LB increase together...",
      "votes": null
    },
    {
      "id": "672625",
      "postDate": "11/14/2019 02:54:24",
      "content": "<p>Thanks for the share, seems I need to find my important factor for increasing both CV and LB either..</p>",
      "rawMarkdown": "Thanks for the share, seems I need to find my important factor for increasing both CV and LB either..",
      "votes": null
    },
    {
      "id": "672634",
      "postDate": "11/14/2019 03:09:50",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a> Hi, do you get CV 0.66+? my best cv is about .655</p>",
      "rawMarkdown": "xiejialun Hi, do you get CV 0.66+? my best cv is about .655",
      "votes": null
    },
    {
      "id": "672645",
      "postDate": "11/14/2019 03:31:41",
      "content": "<p>Yes, my best CV is 0.663x.  But with best CV setting(pixel threshold, minsize threshold), LB score drop..\nAlso all the scores I said are ensemble with classifier.</p>",
      "rawMarkdown": "Yes, my best CV is 0.663x.  But with best CV setting(pixel threshold, minsize threshold), LB score drop..\nAlso all the scores I said are ensemble with classifier.",
      "votes": null
    },
    {
      "id": "672707",
      "postDate": "11/14/2019 04:36:51",
      "content": "<p>while this cannot stablize your  cv and lb score, but it can make a prediction</p>\n\n<p>```\nfor a chosen trained model, we record performance metric on validation set (and train set)\nx1 = bce loss\nx2 = num of mask1\nx3 = num of mask2\nx4 = num of no mask prediction \nx5 .....\nx6 ....</p>\n\n<p>after we make a submission, we have kaggle score s.</p>\n\n<p>with enough submission samples, can we:</p>\n\n<p>predict s = xgboost(x1,x2,x3 .....)</p>\n\n<p>``` </p>",
      "rawMarkdown": "while this cannot stablize your  cv and lb score, but it can make a prediction\n\n```\nfor a chosen trained model, we record performance metric on validation set (and train set)\nx1 = bce loss\nx2 = num of mask1\nx3 = num of mask2\nx4 = num of no mask prediction \nx5 .....\nx6 ....\n\nafter we make a submission, we have kaggle score s.\n\nwith enough submission samples, can we:\n\npredict s = xgboost(x1,x2,x3 .....)\n\n\n```",
      "votes": null
    },
    {
      "id": "672712",
      "postDate": "11/14/2019 04:45:06",
      "content": "<p>thanks, I found my local cv has a bug...</p>",
      "rawMarkdown": "thanks, I found my local cv has a bug...",
      "votes": null
    },
    {
      "id": "672838",
      "postDate": "11/14/2019 08:12:24",
      "content": "<p>This is great! I used to write a notebook to record the results, but in the end, the tons of numbers just confuse me more.. \nI really like this idea!  Thanks Heng, this is very helpful.</p>",
      "rawMarkdown": "This is great! I used to write a notebook to record the results, but in the end, the tons of numbers just confuse me more.. \nI really like this idea!  Thanks Heng, this is very helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 672090,
      "author_name": "xiejialun",
      "author_url": "",
      "post_date": "11/13/2019 14:05:21",
      "content": "<p>But after ensemble them with my current best models, the score decrease.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 672116,
      "author_name": "dandingclam",
      "author_url": "",
      "post_date": "11/13/2019 14:23:47",
      "content": "<p>Something similar happened, my choice is to believe those models with better CV&amp;LB scores.</p>",
      "votes": null,
      "replies": [
        {
          "id": 672477,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/13/2019 23:47:09",
          "content": "<p>Since I did't spilt the data well at that time, so the CV score of these three models ensemble is unreliable. This make me hard to trust the ensemble result. But this just my case, I am sure yours is fine. Good luck with you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 672196,
      "author_name": "niuddd",
      "author_url": "",
      "post_date": "11/13/2019 15:52:59",
      "content": "<p>same model, I could get .645~.655 with different thresholds...so weird</p>",
      "votes": null,
      "replies": [
        {
          "id": 672480,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/13/2019 23:51:34",
          "content": "<p>Yes, I have the same problem.. \nThe score is impacted by the threshold so much. This make me really worry about the shake up..\nSo I tend to fix the threshold and just try the result by different model ensemble now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 672204,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/13/2019 16:11:47",
      "content": "<p>\"But after ensemble them, with 0.6 pixel threshold and 15000 mask size threshold. The LB score is 0.664..almost 0.02 boost.\"</p>\n\n<p>after ensemble and post processing by size, you end up with a list of negative (no mask) and positive prediction (mask).\nfix the negative prediction and leave it unchanged.</p>\n\n<p>Then, try to increase of decrease (mask erode or dilate) the positive prediction mask. maybe there will improvement. e.g. if the ground truth box is intersection of different  annotators, conservative prediction like erosion may be better. if the ground truth box is union, maybe dilation is better.</p>\n\n<p>try in on validation set first.</p>\n\n<p>since this operation does not change you negative prediction, you will not increase false positive</p>",
      "votes": null,
      "replies": [
        {
          "id": 672489,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/14/2019 00:06:34",
          "content": "<p>Thanks for the advice! \nI understand that this can remove some false positive for me. But 0.02 boost still make me hard to trust these three models. 😂 \nAlso since I did't fix data split on these 3 models(trained long time age), so I don't have the validation set to verify. I might just give up these 3 models..But they do make me worry about the shake up.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 672519,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/14/2019 00:44:49",
          "content": "<p>\"But they do make me worry about the shake up.\"</p>\n\n<p>this is something i haven't think of how to deal with yet for now. it is also part of the challenge</p>\n\n<p>\"so I don't have the validation set to verify\"</p>\n\n<p>you can do without a validation set.</p>\n\n<ol>\n<li>assume you have train a classifier on a train A set and validation set to verify its performance.</li>\n<li>you have another train classifier B without validation set.</li>\n<li>what you need is a third set S not used to trained A or B. It is ok that this third set is unlabeled.</li>\n<li>now test A and B on S.</li>\n<li>perturb S. test  A and B again on S.</li>\n<li>you can measure the consistency of the results between A and B and consistency of their behavior on perturbation (i.e. how stable or sensitive their behavior is).</li>\n</ol>\n\n<hr>\n\n<p>in doubt, you can still try this:</p>\n\n<ol>\n<li>use the pseudo label of B of set S to pretrain a classifier C.</li>\n<li>then use the original train data (does not contains S) to finetune C.</li>\n</ol>\n\n<p>in step.1 the knowledge (unknown accuracy) of B is transferred to C. in step.2 good knowledge will be retained and bad knowledge will be corrected with the use of good data (i.e. the original train set with label)</p>\n\n<hr>\n\n<p>note: you can google for validation without validation set. you can also borrow the ideas from semi-supervised learning</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 672536,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/14/2019 01:09:11",
          "content": "<p>Thanks a lot for the advice and knowledge, this is something I never knew before! It is really great to learn something new.\nAlthough I don't have enough time and hardware resource to verify them with your awesome advice, since I have something close to my final submission now. But I definitely gonna try it after the competition! Many thanks for the good advice, good luck with you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 672570,
      "author_name": "rentiansky",
      "author_url": "",
      "post_date": "11/14/2019 02:00:07",
      "content": "<p>what's you ensemble method, vote, average on probability, or weighted sum?\nMy CV is not linear with LB, I'm also worry about the shake up.</p>",
      "votes": null,
      "replies": [
        {
          "id": 672586,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/14/2019 02:16:18",
          "content": "<p>I am using average on raw prediction as ensemble method. My local CV has about 0.002-0.004 difference with LB score. And also nonlinear with LB.\nSo I might make 2 final submission with \n1. Best LB \n2. Best CV</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 672595,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/14/2019 02:24:43",
          "content": "<p>But I wont expect I will remain silver zone in PB.. I bet my score definitely  going to shake....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 672604,
          "author_name": "limerobot",
          "author_url": "",
          "post_date": "11/14/2019 02:30:55",
          "content": "<p>My CV is also nonlinear with LB. 😂 \nSo I try to add something when CV and LB increase together...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 672625,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/14/2019 02:54:24",
          "content": "<p>Thanks for the share, seems I need to find my important factor for increasing both CV and LB either..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 672634,
          "author_name": "niuddd",
          "author_url": "",
          "post_date": "11/14/2019 03:09:50",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a> Hi, do you get CV 0.66+? my best cv is about .655</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 672645,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/14/2019 03:31:41",
          "content": "<p>Yes, my best CV is 0.663x.  But with best CV setting(pixel threshold, minsize threshold), LB score drop..\nAlso all the scores I said are ensemble with classifier.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 672712,
          "author_name": "niuddd",
          "author_url": "",
          "post_date": "11/14/2019 04:45:06",
          "content": "<p>thanks, I found my local cv has a bug...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 672707,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/14/2019 04:36:51",
      "content": "<p>while this cannot stablize your  cv and lb score, but it can make a prediction</p>\n\n<p>```\nfor a chosen trained model, we record performance metric on validation set (and train set)\nx1 = bce loss\nx2 = num of mask1\nx3 = num of mask2\nx4 = num of no mask prediction \nx5 .....\nx6 ....</p>\n\n<p>after we make a submission, we have kaggle score s.</p>\n\n<p>with enough submission samples, can we:</p>\n\n<p>predict s = xgboost(x1,x2,x3 .....)</p>\n\n<p>``` </p>",
      "votes": null,
      "replies": [
        {
          "id": 672838,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/14/2019 08:12:24",
          "content": "<p>This is great! I used to write a notebook to record the results, but in the end, the tons of numbers just confuse me more.. \nI really like this idea!  Thanks Heng, this is very helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "672075": "Today I ensembled three of my models which trained long time ago(about 3 weeks). And the LB and CV score of these 3 models are less than 0.65, which is quite bad. \n(model 1 -&gt; 0.648, model 2 -&gt; 0.645, model 3 -&gt; 0.647) \n\nBut after ensemble them, with 0.6 pixel threshold and 15000 mask size threshold. The LB score is 0.664..almost 0.02 boost.\n\nDo you guys have the similar situation?",
    "672090": "But after ensemble them with my current best models, the score decrease.",
    "672116": "Something similar happened, my choice is to believe those models with better CV&amp;LB scores.",
    "672196": "same model, I could get .645~.655 with different thresholds...so weird",
    "672204": "\"But after ensemble them, with 0.6 pixel threshold and 15000 mask size threshold. The LB score is 0.664..almost 0.02 boost.\"\n\n\nafter ensemble and post processing by size, you end up with a list of negative (no mask) and positive prediction (mask).\nfix the negative prediction and leave it unchanged.\n\nThen, try to increase of decrease (mask erode or dilate) the positive prediction mask. maybe there will improvement. e.g. if the ground truth box is intersection of different  annotators, conservative prediction like erosion may be better. if the ground truth box is union, maybe dilation is better.\n\ntry in on validation set first.\n\nsince this operation does not change you negative prediction, you will not increase false positive",
    "672477": "Since I did't spilt the data well at that time, so the CV score of these three models ensemble is unreliable. This make me hard to trust the ensemble result. But this just my case, I am sure yours is fine. Good luck with you!",
    "672480": "Yes, I have the same problem.. \nThe score is impacted by the threshold so much. This make me really worry about the shake up..\nSo I tend to fix the threshold and just try the result by different model ensemble now.",
    "672489": "Thanks for the advice! \nI understand that this can remove some false positive for me. But 0.02 boost still make me hard to trust these three models. 😂 \nAlso since I did't fix data split on these 3 models(trained long time age), so I don't have the validation set to verify. I might just give up these 3 models..But they do make me worry about the shake up.",
    "672519": "\"But they do make me worry about the shake up.\"\n\nthis is something i haven't think of how to deal with yet for now. it is also part of the challenge\n\n\"so I don't have the validation set to verify\"\n\nyou can do without a validation set.\n\n1.  assume you have train a classifier on a train A set and validation set to verify its performance.\n2. you have another train classifier B without validation set.\n3. what you need is a third set S not used to trained A or B. It is ok that this third set is unlabeled.\n4. now test A and B on S.\n5. perturb S. test  A and B again on S.\n6. you can measure the consistency of the results between A and B and consistency of their behavior on perturbation (i.e. how stable or sensitive their behavior is).\n\n----\n\nin doubt, you can still try this:\n\n1. use the pseudo label of B of set S to pretrain a classifier C.\n2. then use the original train data (does not contains S) to finetune C.\n\nin step.1 the knowledge (unknown accuracy) of B is transferred to C. in step.2 good knowledge will be retained and bad knowledge will be corrected with the use of good data (i.e. the original train set with label)\n\n---- \n\nnote: you can google for validation without validation set. you can also borrow the ideas from semi-supervised learning",
    "672536": "Thanks a lot for the advice and knowledge, this is something I never knew before! It is really great to learn something new.\nAlthough I don't have enough time and hardware resource to verify them with your awesome advice, since I have something close to my final submission now. But I definitely gonna try it after the competition! Many thanks for the good advice, good luck with you!",
    "672570": "what's you ensemble method, vote, average on probability, or weighted sum?\nMy CV is not linear with LB, I'm also worry about the shake up.",
    "672586": "I am using average on raw prediction as ensemble method. My local CV has about 0.002-0.004 difference with LB score. And also nonlinear with LB.\nSo I might make 2 final submission with \n1. Best LB \n2. Best CV",
    "672595": "But I wont expect I will remain silver zone in PB.. I bet my score definitely  going to shake....",
    "672604": "My CV is also nonlinear with LB. 😂 \nSo I try to add something when CV and LB increase together...",
    "672625": "Thanks for the share, seems I need to find my important factor for increasing both CV and LB either..",
    "672634": "xiejialun Hi, do you get CV 0.66+? my best cv is about .655",
    "672645": "Yes, my best CV is 0.663x.  But with best CV setting(pixel threshold, minsize threshold), LB score drop..\nAlso all the scores I said are ensemble with classifier.",
    "672707": "while this cannot stablize your  cv and lb score, but it can make a prediction\n\n```\nfor a chosen trained model, we record performance metric on validation set (and train set)\nx1 = bce loss\nx2 = num of mask1\nx3 = num of mask2\nx4 = num of no mask prediction \nx5 .....\nx6 ....\n\nafter we make a submission, we have kaggle score s.\n\nwith enough submission samples, can we:\n\npredict s = xgboost(x1,x2,x3 .....)\n\n\n```",
    "672712": "thanks, I found my local cv has a bug...",
    "672838": "This is great! I used to write a notebook to record the results, but in the end, the tons of numbers just confuse me more.. \nI really like this idea!  Thanks Heng, this is very helpful."
  },
  "source": "meta"
}