{
  "id": 76213,
  "title": "What will be your final submission choice?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76213",
  "author_name": "",
  "post_date": "2018-12-30T14:36:03.255894700Z",
  "votes": 4,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I see that the LB is very sensitive to very small changes. One thing I am pondering over right now is which Kernel I should submit:\n1. One with the best LB = 0.698,  Local CV=0.6806 ? or \n2. The one with best Local CV = 0.697 and LB = 0.692</p>\n\n<p>I am feeling like I should go for the second choice, but if I go with that I am going blind in the final stage of the competition. What if the test and train data are not distributed identically and somehow the kernel that performs best on LB is actually the best? </p>\n\n<p>I know there are 2 submissions and I can have both submissions to check. But I really want to use submissions wisely for I have other ideas for the second submission.</p>\n\n<p>What are your thoughts? </p>",
  "messages": [
    {
      "id": "447776",
      "postDate": "12/30/2018 14:36:03",
      "content": "<p>I see that the LB is very sensitive to very small changes. One thing I am pondering over right now is which Kernel I should submit:\n1. One with the best LB = 0.698,  Local CV=0.6806 ? or \n2. The one with best Local CV = 0.697 and LB = 0.692</p>\n\n<p>I am feeling like I should go for the second choice, but if I go with that I am going blind in the final stage of the competition. What if the test and train data are not distributed identically and somehow the kernel that performs best on LB is actually the best? </p>\n\n<p>I know there are 2 submissions and I can have both submissions to check. But I really want to use submissions wisely for I have other ideas for the second submission.</p>\n\n<p>What are your thoughts? </p>",
      "rawMarkdown": "I see that the LB is very sensitive to very small changes. One thing I am pondering over right now is which Kernel I should submit:\n1. One with the best LB = 0.698,  Local CV=0.6806 ? or \n2. The one with best Local CV = 0.697 and LB = 0.692\n\nI am feeling like I should go for the second choice, but if I go with that I am going blind in the final stage of the competition. What if the test and train data are not distributed identically and somehow the kernel that performs best on LB is actually the best? \n\nI know there are 2 submissions and I can have both submissions to check. But I really want to use submissions wisely for I have other ideas for the second submission.\n\nWhat are your thoughts?",
      "votes": null
    },
    {
      "id": "447794",
      "postDate": "12/30/2018 15:24:44",
      "content": "<p>Depends on how you do your CV, I'll trust a 4-fold 0.697 (which is absurdly high to me btw), but a 5+ fold may not have the diversity needed. </p>",
      "rawMarkdown": "Depends on how you do your CV, I'll trust a 4-fold 0.697 (which is absurdly high to me btw), but a 5+ fold may not have the diversity needed.",
      "votes": null
    },
    {
      "id": "447797",
      "postDate": "12/30/2018 15:27:08",
      "content": "<p>What I do right now to get the CV is :</p>\n\n<p>get oof predictions on train data while doing 5 fold cv. Get the F1 score using these predictions on the train data. Is this the right approach? </p>\n\n<p>Should I try out 4 fold cv with the same approach?</p>",
      "rawMarkdown": "What I do right now to get the CV is :\n\nget oof predictions on train data while doing 5 fold cv. Get the F1 score using these predictions on the train data. Is this the right approach? \n\nShould I try out 4 fold cv with the same approach?",
      "votes": null
    },
    {
      "id": "447799",
      "postDate": "12/30/2018 15:31:09",
      "content": "<p>Also the local CV of .697 is a stacked model CV in which the best model has a CV of 0.684</p>",
      "rawMarkdown": "Also the local CV of .697 is a stacked model CV in which the best model has a CV of 0.684",
      "votes": null
    },
    {
      "id": "447808",
      "postDate": "12/30/2018 15:54:25",
      "content": "<p>I think you should try 4 fold, will probably give lower CV but LB may stay the same or even go up. Even if you have the same LB it's still a win since you get more time. </p>",
      "rawMarkdown": "I think you should try 4 fold, will probably give lower CV but LB may stay the same or even go up. Even if you have the same LB it's still a win since you get more time.",
      "votes": null
    },
    {
      "id": "447889",
      "postDate": "12/30/2018 19:17:57",
      "content": "<p>Can I ask how many models you used for your stacking? and what were models like is it RNN's or linear + RNN</p>",
      "rawMarkdown": "Can I ask how many models you used for your stacking? and what were models like is it RNN's or linear + RNN",
      "votes": null
    },
    {
      "id": "447893",
      "postDate": "12/30/2018 19:31:49",
      "content": "<p>5 models. Used mostly GRU+LSTM+Attention architecture in the models</p>",
      "rawMarkdown": "5 models. Used mostly GRU+LSTM+Attention architecture in the models",
      "votes": null
    },
    {
      "id": "447911",
      "postDate": "12/30/2018 20:24:34",
      "content": "<p>5 models. Did it get trained in 2hrs? Any tricks for time optimization ?</p>",
      "rawMarkdown": "5 models. Did it get trained in 2hrs? Any tricks for time optimization ?",
      "votes": null
    },
    {
      "id": "448031",
      "postDate": "12/31/2018 05:28:53",
      "content": "<p>If it is my own result, I choose cv high, lb low. But I am not sure how you do it.</p>",
      "rawMarkdown": "If it is my own result, I choose cv high, lb low. But I am not sure how you do it.",
      "votes": null
    },
    {
      "id": "448056",
      "postDate": "12/31/2018 06:41:16",
      "content": "<p>I remember that we can select 2 kernels for the second stage? For me I will select one best LB score and one best CV score.</p>",
      "rawMarkdown": "I remember that we can select 2 kernels for the second stage? For me I will select one best LB score and one best CV score.",
      "votes": null
    },
    {
      "id": "448064",
      "postDate": "12/31/2018 06:54:50",
      "content": "<p>clever boy.</p>",
      "rawMarkdown": "clever boy.",
      "votes": null
    },
    {
      "id": "448067",
      "postDate": "12/31/2018 07:01:45",
      "content": "<p>And then your best score comes from another kernel with decent CV and decent LB.</p>",
      "rawMarkdown": "And then your best score comes from another kernel with decent CV and decent LB.",
      "votes": null
    },
    {
      "id": "448075",
      "postDate": "12/31/2018 07:23:49",
      "content": "<p>I am running the model on full training set once I get the CV results from another kernel.</p>",
      "rawMarkdown": "I am running the model on full training set once I get the CV results from another kernel.",
      "votes": null
    },
    {
      "id": "448076",
      "postDate": "12/31/2018 07:25:18",
      "content": "<p>Makes sense. But it's hard not to follow the leaderboard. </p>",
      "rawMarkdown": "Makes sense. But it's hard not to follow the leaderboard.",
      "votes": null
    },
    {
      "id": "448083",
      "postDate": "12/31/2018 07:52:07",
      "content": "<p>Okay, got a definitive answer sort of from @Tilli in one of the threads in the toxic comments competition. Sorry the images are not there but the answer is explanator without them too..</p>\n\n<p>All of us have faced this issue: how do we best combine our predictions?</p>\n\n<p>As everyone knows, the simplest way to do that is by averaging predictions:</p>\n\n<p>(pred_01 + pred_02 + pred_03) / 3</p>\n\n<p>Or we could get fancy, because we know that pred_01 is better than the other two:</p>\n\n<p>( (pred_01 * 0.7) + (pred_01 * 0.2) + (pred_01 * 0.1) )</p>\n\n<p>or its variant:</p>\n\n<p>( (pred_01 * 7) + (pred_01 * 2) + (pred_01 * 1) ) / 10</p>\n\n<p>What we are doing here is multiplying our final predictions by some weights that are semi-arbitrary at first, and then we adjust them based on LB feedback. We keep adjusting them until we get the best LB score, and then publish a kernel that instantly earns 50+ votes and a gold medal.</p>\n\n<p>But there is an iceberg-sized problem with that approach - it may not look big on surface, but it usually ends up sinking lots of predictions. That problem is that we are making several assumptions that may or may not be correct, which can easily lead to over-fitting our weights based on LB feedback. The assumptions: 1) we assume that the LB score is a true measure of our model quality, while it is only a measure of our model’s quality on 25% of data; 2) we assume that the public LB data (25%) and the private LB data (75%) have the same distributions; 3) we adjust our initial weights based on the public LB feedback, which again comes from 25% of data; 4) we settle on our final weight based on the public LB feedback, which – for the last time – comes from 25% of data. That’s a lot of assumptions.</p>\n\n<p>Now, let’s assume that public and private LB data distributions are just slightly off, say 1% (even though that’s not the way to measure different distributions). How many people here think that 1% difference is insignificant when we compare 25% of data we used for all assumptions listed above and 75% that ends up determining the final ranking? Let me remind you that differences at 3rd decimal place of AUC scores will likely end up separating hundreds of competitors.</p>\n\n<p>But what can we do about this other than praying that public and private LB data are similarly distributed? If only there was a way to determine those blending weights in unbiased fashion …</p>\n\n<p>That’s what out-of-fold (OOF) data is about. It provides a dataset for which we know the distribution and the final class labels – it is based on train data. The only thing we have to make sure is that those predictions are unbiased, which is easily done during cross-validation. We train on 4/5 (or 9/10) of data and predict on the remaining 1/5, and repeat that another 4 times until all of train data has been excluded once from training and instead used for predictions. All the while we are doing these 5 training sessions, we will make predictions on test data and average them. Now we have a single averaged prediction on test data, and a single OOF prediction on train data.</p>\n\n<p>Still, how will that help us find blending weights any better than combining them by fitting the LB score?</p>\n\n<p>Let’s say we have three OOF predictions called model1, model2 and model3. Four data points are shown below for simplicity for each of them, along with the real target values (we know them because OOF data is based on train data). Let’s say that we initially give each of them a weight of 1/3, and that all of our weights must add up to 1. Finally, we decide to minimize the function which is the absolute difference between the true target values and (model1weight1 + model2weight2 + model3*weight3). We could have chosen to maximize the AUC value, but I didn’t do that in this example.</p>\n\n<p>This is a simple number-crunching exercise, which can be solved many different ways. By eye-balling these 3 OOF models, we could guess that model1 is best on its own. We can prove that formally by giving model1 weight of 1. That will give us a baseline value of 0.3792 that we need to make smaller.</p>\n\n<p>This particular combination gives larger function value, so it is not good. This exercise is not something you want to do by guessing, even if you have only 4 data points. Check out this kernel for how to do this either by straight-up numerical minimization or by MCMC sampling.</p>\n\n<p>Now we use these weights obtained from OOF models to multiply our predictions. Wait, but what if that gives us a score that is worse than finding weights by intuition?</p>\n\n<p>Keep in mind that OOF predictions are unbiased, and you know the target value for all of them. Therefore, that is likely a much larger dataset than 25% of public LB data that gives you your other piece of information. Let’s assume that your CV scores are consistent with LB scores, which is the only assumption you need to make to conclude that your OOF data is reliable. So here is your choice:</p>\n\n<ul>\n<li><p>Trust the initial CV results and subsequent blending weights from an\nOOF dataset that you know and understand, even if it gives you a\nlower score</p></li>\n<li><p>Trust the initial LB scores and subsequent blending weights based on a\nfeedback from 25% of data which you do not fully understand, yet may\ngive you a higher score</p></li>\n</ul>\n\n<p><a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52224\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52224</a></p>",
      "rawMarkdown": "Okay, got a definitive answer sort of from @Tilli in one of the threads in the toxic comments competition. Sorry the images are not there but the answer is explanator without them too..\n\nAll of us have faced this issue: how do we best combine our predictions?\n\nAs everyone knows, the simplest way to do that is by averaging predictions:\n\n(pred_01 + pred_02 + pred_03) / 3\n\nOr we could get fancy, because we know that pred_01 is better than the other two:\n\n( (pred_01 * 0.7) + (pred_01 * 0.2) + (pred_01 * 0.1) )\n\nor its variant:\n\n( (pred_01 * 7) + (pred_01 * 2) + (pred_01 * 1) ) / 10\n\nWhat we are doing here is multiplying our final predictions by some weights that are semi-arbitrary at first, and then we adjust them based on LB feedback. We keep adjusting them until we get the best LB score, and then publish a kernel that instantly earns 50+ votes and a gold medal.\n\nBut there is an iceberg-sized problem with that approach - it may not look big on surface, but it usually ends up sinking lots of predictions. That problem is that we are making several assumptions that may or may not be correct, which can easily lead to over-fitting our weights based on LB feedback. The assumptions: 1) we assume that the LB score is a true measure of our model quality, while it is only a measure of our model’s quality on 25% of data; 2) we assume that the public LB data (25%) and the private LB data (75%) have the same distributions; 3) we adjust our initial weights based on the public LB feedback, which again comes from 25% of data; 4) we settle on our final weight based on the public LB feedback, which – for the last time – comes from 25% of data. That’s a lot of assumptions.\n\nNow, let’s assume that public and private LB data distributions are just slightly off, say 1% (even though that’s not the way to measure different distributions). How many people here think that 1% difference is insignificant when we compare 25% of data we used for all assumptions listed above and 75% that ends up determining the final ranking? Let me remind you that differences at 3rd decimal place of AUC scores will likely end up separating hundreds of competitors.\n\nBut what can we do about this other than praying that public and private LB data are similarly distributed? If only there was a way to determine those blending weights in unbiased fashion …\n\nThat’s what out-of-fold (OOF) data is about. It provides a dataset for which we know the distribution and the final class labels – it is based on train data. The only thing we have to make sure is that those predictions are unbiased, which is easily done during cross-validation. We train on 4/5 (or 9/10) of data and predict on the remaining 1/5, and repeat that another 4 times until all of train data has been excluded once from training and instead used for predictions. All the while we are doing these 5 training sessions, we will make predictions on test data and average them. Now we have a single averaged prediction on test data, and a single OOF prediction on train data.\n\nStill, how will that help us find blending weights any better than combining them by fitting the LB score?\n\nLet’s say we have three OOF predictions called model1, model2 and model3. Four data points are shown below for simplicity for each of them, along with the real target values (we know them because OOF data is based on train data). Let’s say that we initially give each of them a weight of 1/3, and that all of our weights must add up to 1. Finally, we decide to minimize the function which is the absolute difference between the true target values and (model1weight1 + model2weight2 + model3*weight3). We could have chosen to maximize the AUC value, but I didn’t do that in this example.\n\nThis is a simple number-crunching exercise, which can be solved many different ways. By eye-balling these 3 OOF models, we could guess that model1 is best on its own. We can prove that formally by giving model1 weight of 1. That will give us a baseline value of 0.3792 that we need to make smaller.\n\nThis particular combination gives larger function value, so it is not good. This exercise is not something you want to do by guessing, even if you have only 4 data points. Check out this kernel for how to do this either by straight-up numerical minimization or by MCMC sampling.\n\nNow we use these weights obtained from OOF models to multiply our predictions. Wait, but what if that gives us a score that is worse than finding weights by intuition?\n\nKeep in mind that OOF predictions are unbiased, and you know the target value for all of them. Therefore, that is likely a much larger dataset than 25% of public LB data that gives you your other piece of information. Let’s assume that your CV scores are consistent with LB scores, which is the only assumption you need to make to conclude that your OOF data is reliable. So here is your choice:\n\n- Trust the initial CV results and subsequent blending weights from an\nOOF dataset that you know and understand, even if it gives you a\nlower score\n\n- Trust the initial LB scores and subsequent blending weights based on a\nfeedback from 25% of data which you do not fully understand, yet may\ngive you a higher score\n\nhttps://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52224",
      "votes": null
    },
    {
      "id": "448522",
      "postDate": "01/01/2019 12:10:44",
      "content": "<p><a href=\"/mlwhiz\">@mlwhiz</a>\nI have an amazing guess that all models will get a higher score than the current one in the new test data. \nI hope not to hit my face, haha~~</p>",
      "rawMarkdown": "mlwhiz\nI have an amazing guess that all models will get a higher score than the current one in the new test data. \nI hope not to hit my face, haha~~",
      "votes": null
    },
    {
      "id": "448794",
      "postDate": "01/02/2019 05:16:25",
      "content": "<p>Same here. I am not even thinking of submitting my best model on the LB for final submission.</p>",
      "rawMarkdown": "Same here. I am not even thinking of submitting my best model on the LB for final submission.",
      "votes": null
    },
    {
      "id": "448846",
      "postDate": "01/02/2019 08:09:59",
      "content": "<p>That's a very definite possibility but I have nothing to do with this.</p>",
      "rawMarkdown": "That's a very definite possibility but I have nothing to do with this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 447794,
      "author_name": "suicaokhoailang",
      "author_url": "",
      "post_date": "12/30/2018 15:24:44",
      "content": "<p>Depends on how you do your CV, I'll trust a 4-fold 0.697 (which is absurdly high to me btw), but a 5+ fold may not have the diversity needed. </p>",
      "votes": null,
      "replies": [
        {
          "id": 447797,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "12/30/2018 15:27:08",
          "content": "<p>What I do right now to get the CV is :</p>\n\n<p>get oof predictions on train data while doing 5 fold cv. Get the F1 score using these predictions on the train data. Is this the right approach? </p>\n\n<p>Should I try out 4 fold cv with the same approach?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447799,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "12/30/2018 15:31:09",
          "content": "<p>Also the local CV of .697 is a stacked model CV in which the best model has a CV of 0.684</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447808,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "12/30/2018 15:54:25",
          "content": "<p>I think you should try 4 fold, will probably give lower CV but LB may stay the same or even go up. Even if you have the same LB it's still a win since you get more time. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447889,
          "author_name": "suchith0312",
          "author_url": "",
          "post_date": "12/30/2018 19:17:57",
          "content": "<p>Can I ask how many models you used for your stacking? and what were models like is it RNN's or linear + RNN</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447893,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "12/30/2018 19:31:49",
          "content": "<p>5 models. Used mostly GRU+LSTM+Attention architecture in the models</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447911,
          "author_name": "suchith0312",
          "author_url": "",
          "post_date": "12/30/2018 20:24:34",
          "content": "<p>5 models. Did it get trained in 2hrs? Any tricks for time optimization ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448075,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "12/31/2018 07:23:49",
          "content": "<p>I am running the model on full training set once I get the CV results from another kernel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 448031,
      "author_name": "xiaobai1123q",
      "author_url": "",
      "post_date": "12/31/2018 05:28:53",
      "content": "<p>If it is my own result, I choose cv high, lb low. But I am not sure how you do it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 448076,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "12/31/2018 07:25:18",
          "content": "<p>Makes sense. But it's hard not to follow the leaderboard. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448522,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "01/01/2019 12:10:44",
          "content": "<p><a href=\"/mlwhiz\">@mlwhiz</a>\nI have an amazing guess that all models will get a higher score than the current one in the new test data. \nI hope not to hit my face, haha~~</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448794,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "01/02/2019 05:16:25",
          "content": "<p>Same here. I am not even thinking of submitting my best model on the LB for final submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 448056,
      "author_name": "kaggleczs",
      "author_url": "",
      "post_date": "12/31/2018 06:41:16",
      "content": "<p>I remember that we can select 2 kernels for the second stage? For me I will select one best LB score and one best CV score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 448064,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "12/31/2018 06:54:50",
          "content": "<p>clever boy.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448067,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "12/31/2018 07:01:45",
          "content": "<p>And then your best score comes from another kernel with decent CV and decent LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448846,
          "author_name": "kaggleczs",
          "author_url": "",
          "post_date": "01/02/2019 08:09:59",
          "content": "<p>That's a very definite possibility but I have nothing to do with this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 448083,
      "author_name": "mlwhiz",
      "author_url": "",
      "post_date": "12/31/2018 07:52:07",
      "content": "<p>Okay, got a definitive answer sort of from @Tilli in one of the threads in the toxic comments competition. Sorry the images are not there but the answer is explanator without them too..</p>\n\n<p>All of us have faced this issue: how do we best combine our predictions?</p>\n\n<p>As everyone knows, the simplest way to do that is by averaging predictions:</p>\n\n<p>(pred_01 + pred_02 + pred_03) / 3</p>\n\n<p>Or we could get fancy, because we know that pred_01 is better than the other two:</p>\n\n<p>( (pred_01 * 0.7) + (pred_01 * 0.2) + (pred_01 * 0.1) )</p>\n\n<p>or its variant:</p>\n\n<p>( (pred_01 * 7) + (pred_01 * 2) + (pred_01 * 1) ) / 10</p>\n\n<p>What we are doing here is multiplying our final predictions by some weights that are semi-arbitrary at first, and then we adjust them based on LB feedback. We keep adjusting them until we get the best LB score, and then publish a kernel that instantly earns 50+ votes and a gold medal.</p>\n\n<p>But there is an iceberg-sized problem with that approach - it may not look big on surface, but it usually ends up sinking lots of predictions. That problem is that we are making several assumptions that may or may not be correct, which can easily lead to over-fitting our weights based on LB feedback. The assumptions: 1) we assume that the LB score is a true measure of our model quality, while it is only a measure of our model’s quality on 25% of data; 2) we assume that the public LB data (25%) and the private LB data (75%) have the same distributions; 3) we adjust our initial weights based on the public LB feedback, which again comes from 25% of data; 4) we settle on our final weight based on the public LB feedback, which – for the last time – comes from 25% of data. That’s a lot of assumptions.</p>\n\n<p>Now, let’s assume that public and private LB data distributions are just slightly off, say 1% (even though that’s not the way to measure different distributions). How many people here think that 1% difference is insignificant when we compare 25% of data we used for all assumptions listed above and 75% that ends up determining the final ranking? Let me remind you that differences at 3rd decimal place of AUC scores will likely end up separating hundreds of competitors.</p>\n\n<p>But what can we do about this other than praying that public and private LB data are similarly distributed? If only there was a way to determine those blending weights in unbiased fashion …</p>\n\n<p>That’s what out-of-fold (OOF) data is about. It provides a dataset for which we know the distribution and the final class labels – it is based on train data. The only thing we have to make sure is that those predictions are unbiased, which is easily done during cross-validation. We train on 4/5 (or 9/10) of data and predict on the remaining 1/5, and repeat that another 4 times until all of train data has been excluded once from training and instead used for predictions. All the while we are doing these 5 training sessions, we will make predictions on test data and average them. Now we have a single averaged prediction on test data, and a single OOF prediction on train data.</p>\n\n<p>Still, how will that help us find blending weights any better than combining them by fitting the LB score?</p>\n\n<p>Let’s say we have three OOF predictions called model1, model2 and model3. Four data points are shown below for simplicity for each of them, along with the real target values (we know them because OOF data is based on train data). Let’s say that we initially give each of them a weight of 1/3, and that all of our weights must add up to 1. Finally, we decide to minimize the function which is the absolute difference between the true target values and (model1weight1 + model2weight2 + model3*weight3). We could have chosen to maximize the AUC value, but I didn’t do that in this example.</p>\n\n<p>This is a simple number-crunching exercise, which can be solved many different ways. By eye-balling these 3 OOF models, we could guess that model1 is best on its own. We can prove that formally by giving model1 weight of 1. That will give us a baseline value of 0.3792 that we need to make smaller.</p>\n\n<p>This particular combination gives larger function value, so it is not good. This exercise is not something you want to do by guessing, even if you have only 4 data points. Check out this kernel for how to do this either by straight-up numerical minimization or by MCMC sampling.</p>\n\n<p>Now we use these weights obtained from OOF models to multiply our predictions. Wait, but what if that gives us a score that is worse than finding weights by intuition?</p>\n\n<p>Keep in mind that OOF predictions are unbiased, and you know the target value for all of them. Therefore, that is likely a much larger dataset than 25% of public LB data that gives you your other piece of information. Let’s assume that your CV scores are consistent with LB scores, which is the only assumption you need to make to conclude that your OOF data is reliable. So here is your choice:</p>\n\n<ul>\n<li><p>Trust the initial CV results and subsequent blending weights from an\nOOF dataset that you know and understand, even if it gives you a\nlower score</p></li>\n<li><p>Trust the initial LB scores and subsequent blending weights based on a\nfeedback from 25% of data which you do not fully understand, yet may\ngive you a higher score</p></li>\n</ul>\n\n<p><a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52224\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52224</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "447776": "I see that the LB is very sensitive to very small changes. One thing I am pondering over right now is which Kernel I should submit:\n1. One with the best LB = 0.698,  Local CV=0.6806 ? or \n2. The one with best Local CV = 0.697 and LB = 0.692\n\nI am feeling like I should go for the second choice, but if I go with that I am going blind in the final stage of the competition. What if the test and train data are not distributed identically and somehow the kernel that performs best on LB is actually the best? \n\nI know there are 2 submissions and I can have both submissions to check. But I really want to use submissions wisely for I have other ideas for the second submission.\n\nWhat are your thoughts?",
    "447794": "Depends on how you do your CV, I'll trust a 4-fold 0.697 (which is absurdly high to me btw), but a 5+ fold may not have the diversity needed.",
    "447797": "What I do right now to get the CV is :\n\nget oof predictions on train data while doing 5 fold cv. Get the F1 score using these predictions on the train data. Is this the right approach? \n\nShould I try out 4 fold cv with the same approach?",
    "447799": "Also the local CV of .697 is a stacked model CV in which the best model has a CV of 0.684",
    "447808": "I think you should try 4 fold, will probably give lower CV but LB may stay the same or even go up. Even if you have the same LB it's still a win since you get more time.",
    "447889": "Can I ask how many models you used for your stacking? and what were models like is it RNN's or linear + RNN",
    "447893": "5 models. Used mostly GRU+LSTM+Attention architecture in the models",
    "447911": "5 models. Did it get trained in 2hrs? Any tricks for time optimization ?",
    "448031": "If it is my own result, I choose cv high, lb low. But I am not sure how you do it.",
    "448056": "I remember that we can select 2 kernels for the second stage? For me I will select one best LB score and one best CV score.",
    "448064": "clever boy.",
    "448067": "And then your best score comes from another kernel with decent CV and decent LB.",
    "448075": "I am running the model on full training set once I get the CV results from another kernel.",
    "448076": "Makes sense. But it's hard not to follow the leaderboard.",
    "448083": "Okay, got a definitive answer sort of from @Tilli in one of the threads in the toxic comments competition. Sorry the images are not there but the answer is explanator without them too..\n\nAll of us have faced this issue: how do we best combine our predictions?\n\nAs everyone knows, the simplest way to do that is by averaging predictions:\n\n(pred_01 + pred_02 + pred_03) / 3\n\nOr we could get fancy, because we know that pred_01 is better than the other two:\n\n( (pred_01 * 0.7) + (pred_01 * 0.2) + (pred_01 * 0.1) )\n\nor its variant:\n\n( (pred_01 * 7) + (pred_01 * 2) + (pred_01 * 1) ) / 10\n\nWhat we are doing here is multiplying our final predictions by some weights that are semi-arbitrary at first, and then we adjust them based on LB feedback. We keep adjusting them until we get the best LB score, and then publish a kernel that instantly earns 50+ votes and a gold medal.\n\nBut there is an iceberg-sized problem with that approach - it may not look big on surface, but it usually ends up sinking lots of predictions. That problem is that we are making several assumptions that may or may not be correct, which can easily lead to over-fitting our weights based on LB feedback. The assumptions: 1) we assume that the LB score is a true measure of our model quality, while it is only a measure of our model’s quality on 25% of data; 2) we assume that the public LB data (25%) and the private LB data (75%) have the same distributions; 3) we adjust our initial weights based on the public LB feedback, which again comes from 25% of data; 4) we settle on our final weight based on the public LB feedback, which – for the last time – comes from 25% of data. That’s a lot of assumptions.\n\nNow, let’s assume that public and private LB data distributions are just slightly off, say 1% (even though that’s not the way to measure different distributions). How many people here think that 1% difference is insignificant when we compare 25% of data we used for all assumptions listed above and 75% that ends up determining the final ranking? Let me remind you that differences at 3rd decimal place of AUC scores will likely end up separating hundreds of competitors.\n\nBut what can we do about this other than praying that public and private LB data are similarly distributed? If only there was a way to determine those blending weights in unbiased fashion …\n\nThat’s what out-of-fold (OOF) data is about. It provides a dataset for which we know the distribution and the final class labels – it is based on train data. The only thing we have to make sure is that those predictions are unbiased, which is easily done during cross-validation. We train on 4/5 (or 9/10) of data and predict on the remaining 1/5, and repeat that another 4 times until all of train data has been excluded once from training and instead used for predictions. All the while we are doing these 5 training sessions, we will make predictions on test data and average them. Now we have a single averaged prediction on test data, and a single OOF prediction on train data.\n\nStill, how will that help us find blending weights any better than combining them by fitting the LB score?\n\nLet’s say we have three OOF predictions called model1, model2 and model3. Four data points are shown below for simplicity for each of them, along with the real target values (we know them because OOF data is based on train data). Let’s say that we initially give each of them a weight of 1/3, and that all of our weights must add up to 1. Finally, we decide to minimize the function which is the absolute difference between the true target values and (model1weight1 + model2weight2 + model3*weight3). We could have chosen to maximize the AUC value, but I didn’t do that in this example.\n\nThis is a simple number-crunching exercise, which can be solved many different ways. By eye-balling these 3 OOF models, we could guess that model1 is best on its own. We can prove that formally by giving model1 weight of 1. That will give us a baseline value of 0.3792 that we need to make smaller.\n\nThis particular combination gives larger function value, so it is not good. This exercise is not something you want to do by guessing, even if you have only 4 data points. Check out this kernel for how to do this either by straight-up numerical minimization or by MCMC sampling.\n\nNow we use these weights obtained from OOF models to multiply our predictions. Wait, but what if that gives us a score that is worse than finding weights by intuition?\n\nKeep in mind that OOF predictions are unbiased, and you know the target value for all of them. Therefore, that is likely a much larger dataset than 25% of public LB data that gives you your other piece of information. Let’s assume that your CV scores are consistent with LB scores, which is the only assumption you need to make to conclude that your OOF data is reliable. So here is your choice:\n\n- Trust the initial CV results and subsequent blending weights from an\nOOF dataset that you know and understand, even if it gives you a\nlower score\n\n- Trust the initial LB scores and subsequent blending weights based on a\nfeedback from 25% of data which you do not fully understand, yet may\ngive you a higher score\n\nhttps://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52224",
    "448522": "mlwhiz\nI have an amazing guess that all models will get a higher score than the current one in the new test data. \nI hope not to hit my face, haha~~",
    "448794": "Same here. I am not even thinking of submitting my best model on the LB for final submission.",
    "448846": "That's a very definite possibility but I have nothing to do with this."
  },
  "source": "meta"
}