{
  "id": 55821,
  "title": "Choosing entries for final submission",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55821",
  "author_name": "",
  "post_date": "2018-05-02T07:01:26.093086300Z",
  "votes": 2,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Let me start by saying that this is my first Kaggle competition. Yesterday we received an email with a reminder about the choice of  entries for final submission. It means that slowly but inevitably we are approaching the end of the competition. I have question for you. In particular to experienced Kagglers. How do you choose entry?</p>\n\n<p>I read posts like <a href=\"https://www.kaggle.com/c/mercedes-benz-greener-manufacturing/discussion/36136\">in CV you must trust</a> and I'm scared. People write there that they fell by more than 1000 places down. Will you make a decision based only on your CV? What other factors do you take into account? What other forum posts do you recommend reading?</p>",
  "messages": [
    {
      "id": "321927",
      "postDate": "05/02/2018 07:01:26",
      "content": "<p>Let me start by saying that this is my first Kaggle competition. Yesterday we received an email with a reminder about the choice of  entries for final submission. It means that slowly but inevitably we are approaching the end of the competition. I have question for you. In particular to experienced Kagglers. How do you choose entry?</p>\n\n<p>I read posts like <a href=\"https://www.kaggle.com/c/mercedes-benz-greener-manufacturing/discussion/36136\">in CV you must trust</a> and I'm scared. People write there that they fell by more than 1000 places down. Will you make a decision based only on your CV? What other factors do you take into account? What other forum posts do you recommend reading?</p>",
      "rawMarkdown": "Let me start by saying that this is my first Kaggle competition. Yesterday we received an email with a reminder about the choice of  entries for final submission. It means that slowly but inevitably we are approaching the end of the competition. I have question for you. In particular to experienced Kagglers. How do you choose entry?\n\nI read posts like [in CV you must trust][1] and I'm scared. People write there that they fell by more than 1000 places down. Will you make a decision based only on your CV? What other factors do you take into account? What other forum posts do you recommend reading?\n\n  [1]: https://www.kaggle.com/c/mercedes-benz-greener-manufacturing/discussion/36136",
      "votes": null
    },
    {
      "id": "321934",
      "postDate": "05/02/2018 07:11:40",
      "content": "<p>Tried a few competitions on kaggle, learned some lessons like:  </p>\n\n<p>Rule No.1: trust your local cv;</p>\n\n<p>Rule No.x: follow the rule No.(x-1)</p>",
      "rawMarkdown": "Tried a few competitions on kaggle, learned some lessons like:  \n\nRule No.1: trust your local cv;\n\nRule No.x: follow the rule No.(x-1)",
      "votes": null
    },
    {
      "id": "321943",
      "postDate": "05/02/2018 07:22:36",
      "content": "<p>Hah, I understand. And what about the situation when I have models trained on different data sets? One time for 25M another time for 50M and so on. In this situation, it is impossible to compare their local CV effectively. CV is comparable only when it compares two models that were trained on the same data set and the same validation set. I choose the one that was trained on a larger collection? Is the one whose CV is closer to LB?</p>",
      "rawMarkdown": "Hah, I understand. And what about the situation when I have models trained on different data sets? One time for 25M another time for 50M and so on. In this situation, it is impossible to compare their local CV effectively. CV is comparable only when it compares two models that were trained on the same data set and the same validation set. I choose the one that was trained on a larger collection? Is the one whose CV is closer to LB?",
      "votes": null
    },
    {
      "id": "321964",
      "postDate": "05/02/2018 08:04:04",
      "content": "<p>I would prefer the result with more data.</p>",
      "rawMarkdown": "I would prefer the result with more data.",
      "votes": null
    },
    {
      "id": "321968",
      "postDate": "05/02/2018 08:13:36",
      "content": "<p>Thank you so much for your help and tips.</p>",
      "rawMarkdown": "Thank you so much for your help and tips.",
      "votes": null
    },
    {
      "id": "321974",
      "postDate": "05/02/2018 08:29:03",
      "content": "<p>It may depend sometimes on luck.. </p>\n\n<p>On Statoil  competition , there were people who just submitted a crazy blend of blends of blends kernel and finished above me with silver medal.</p>\n\n<p>Note that the dataset was very small on mercedes. So the 2 are not necessarily comparable. </p>\n\n<p>We are doing true CV for some of our models but I guess some are just validating on a subset due to the size of the whole data.  </p>",
      "rawMarkdown": "It may depend sometimes on luck.. \n\nOn Statoil  competition , there were people who just submitted a crazy blend of blends of blends kernel and finished above me with silver medal.\n\nNote that the dataset was very small on mercedes. So the 2 are not necessarily comparable. \n\nWe are doing true CV for some of our models but I guess some are just validating on a subset due to the size of the whole data.",
      "votes": null
    },
    {
      "id": "321978",
      "postDate": "05/02/2018 08:43:57",
      "content": "<p>In my case, I could not achieve any success using other models than LightGBM. My single best LGBM model is 0.9799 on LB score. So it seems to me that I do not even have anything to blend. Of course, I have a lot of LGBM models and I paln to average them.</p>",
      "rawMarkdown": "In my case, I could not achieve any success using other models than LightGBM. My single best LGBM model is 0.9799 on LB score. So it seems to me that I do not even have anything to blend. Of course, I have a lot of LGBM models and I paln to average them.",
      "votes": null
    },
    {
      "id": "321979",
      "postDate": "05/02/2018 08:45:59",
      "content": "<p>I've tried this strategy in other competitions, blend the results from same model(+ different params), sometime it works.</p>",
      "rawMarkdown": "I've tried this strategy in other competitions, blend the results from same model(+ different params), sometime it works.",
      "votes": null
    },
    {
      "id": "321985",
      "postDate": "05/02/2018 08:58:43",
      "content": "<blockquote>\n  <p><strong>SkalskiP wrote</strong></p>\n  \n  <blockquote>\n    <p>In my case, I could not achieve any success using other models than LightGBM. My single best LGBM model is 0.9799 on LB score. So it seems to me that I do not even have anything to blend. Of course, I have a lot of LGBM models and I paln to average them.</p>\n  </blockquote>\n</blockquote>\n\n<p>As said by yyqing,  you can use differents strategies, differents params , differents features for the same model.   That's what we do with LGBM ( our best single LGBM is scoring 0.9809 with 13 features ) </p>\n\n<p>If it's any consolation, we have also a NN model scoring 0.9802 on LB but does not blend well with our LGBM  ( or we don't know how to do it correctly for now ^^) </p>",
      "rawMarkdown": "&gt; **SkalskiP wrote**\n&gt; \n&gt; &gt; In my case, I could not achieve any success using other models than LightGBM. My single best LGBM model is 0.9799 on LB score. So it seems to me that I do not even have anything to blend. Of course, I have a lot of LGBM models and I paln to average them.\n\nAs said by yyqing,  you can use differents strategies, differents params , differents features for the same model.   That's what we do with LGBM ( our best single LGBM is scoring 0.9809 with 13 features ) \n\nIf it's any consolation, we have also a NN model scoring 0.9802 on LB but does not blend well with our LGBM  ( or we don't know how to do it correctly for now ^^)",
      "votes": null
    },
    {
      "id": "321987",
      "postDate": "05/02/2018 09:04:12",
      "content": "<p>From what I read on the forum, LGBM is influenced by the order of the columns in DataFrame. And of course, random seed. So you say that it is a good idea to take my best model and pruba to put it on different parameters? And the average of the results from these few models?</p>",
      "rawMarkdown": "From what I read on the forum, LGBM is influenced by the order of the columns in DataFrame. And of course, random seed. So you say that it is a good idea to take my best model and pruba to put it on different parameters? And the average of the results from these few models?",
      "votes": null
    },
    {
      "id": "321991",
      "postDate": "05/02/2018 09:12:20",
      "content": "<p>Can anyone confirm the dependency of LGB on column order and has  an explanation for the reason? I find it a bit disturbing, as from my understanding column_sample should take columns at random.</p>\n\n<p>Another strategy can be to build LGB on different subset of the data and average the results, makes even more sense here with different days that most likely have slightly different pattern as well</p>",
      "rawMarkdown": "Can anyone confirm the dependency of LGB on column order and has  an explanation for the reason? I find it a bit disturbing, as from my understanding column_sample should take columns at random.\n\nAnother strategy can be to build LGB on different subset of the data and average the results, makes even more sense here with different days that most likely have slightly different pattern as well",
      "votes": null
    },
    {
      "id": "321993",
      "postDate": "05/02/2018 09:19:57",
      "content": "<blockquote>\n  <p><strong>Serigne  wrote</strong>\n  As said by yyqing,  you can use differents strategies, differents params , differents features for the same model.   That's what we do with LGBM ( our best single LGBM is scoring 0.9809 with 13 features ) </p>\n</blockquote>\n\n<p>BTW: My God. I'm jealous of your such a strong model with so few features. In my case, I have about 25 features. I have a question by the way. Is it possible that my CV will improve when, for example, I remove one of my current features? In formulating this differently: Is it possible that any of my featurów worsens my result? </p>",
      "rawMarkdown": "&gt; **Serigne  wrote**\n&gt; As said by yyqing,  you can use differents strategies, differents params , differents features for the same model.   That's what we do with LGBM ( our best single LGBM is scoring 0.9809 with 13 features ) \n\nBTW: My God. I'm jealous of your such a strong model with so few features. In my case, I have about 25 features. I have a question by the way. Is it possible that my CV will improve when, for example, I remove one of my current features? In formulating this differently: Is it possible that any of my featurów worsens my result?",
      "votes": null
    },
    {
      "id": "321996",
      "postDate": "05/02/2018 09:23:36",
      "content": "<p>Yes some features definately worsen the results, due to overfitting. For this reason most people in this competition use greedy stepwise procedures: start with an empty set of features and add 1 or more features. If the CV scores improve, then keep them, otherwise discard.</p>",
      "rawMarkdown": "Yes some features definately worsen the results, due to overfitting. For this reason most people in this competition use greedy stepwise procedures: start with an empty set of features and add 1 or more features. If the CV scores improve, then keep them, otherwise discard.",
      "votes": null
    },
    {
      "id": "322004",
      "postDate": "05/02/2018 09:37:45",
      "content": "<p>Thank you very much for your answer. I realize that my questions are probably trivial to most people who will read this. But I have three more and it would be great if someone could answer it:\n1. How can I interpret the feature importance obtained from the model? Does it mean to some extent whether the feature is good or not?\n2. Is getting rid of some basic feature such as 'app' or 'os' should I also take into account? Is the basic set from which I should start is 'app', 'os', 'channel' and 'devise' (I understand that 'ip' should not be taken into account, so I read in earlier posts). And from this set I start and add new features.\n3. Can I test models with new features on a smaller data set or on a smaller larning rate and be reasonably sure that the correct CV for such a model will translate into an improvement for my main model? My main model is tening on the 7th and 8th day and validation on 9. This calculation takes a lot of time.</p>",
      "rawMarkdown": "Thank you very much for your answer. I realize that my questions are probably trivial to most people who will read this. But I have three more and it would be great if someone could answer it:\n1. How can I interpret the feature importance obtained from the model? Does it mean to some extent whether the feature is good or not?\n2. Is getting rid of some basic feature such as 'app' or 'os' should I also take into account? Is the basic set from which I should start is 'app', 'os', 'channel' and 'devise' (I understand that 'ip' should not be taken into account, so I read in earlier posts). And from this set I start and add new features.\n3. Can I test models with new features on a smaller data set or on a smaller larning rate and be reasonably sure that the correct CV for such a model will translate into an improvement for my main model? My main model is tening on the 7th and 8th day and validation on 9. This calculation takes a lot of time.",
      "votes": null
    },
    {
      "id": "322223",
      "postDate": "05/02/2018 15:19:03",
      "content": "<p>Here's a very common, sensible heuristic for choosing the two submissions:</p>\n\n<ol>\n<li>Model/ensemble with highest local validation</li>\n<li>Model/ensemble with highest public LB score</li>\n</ol>\n\n<p>Basically you want to be thinking about how to diversify your risk - if your submissions are essentially identical, there's no point in having more than 1. Local validation should typically be the best guide assuming you have a good setup and that the train data is similar to the test data, but the highest PLB submission is a way to partially hedge against the risk of overfitting to your local validation scheme.  </p>\n\n<p>Sometimes 1 and 2 are the same, in which case it's harder to pick a second one. Again thinking about risk, I would typically try to choose a submission that's substantively different from the first. Maybe it's a single model, or leaves out some components of an ensemble or ideas that you're less confident in. This way you can get a submission that's \"safer\" from a complexity perspective (both in the sense of model complexity and entire pipeline complexity).</p>\n\n<p>In this competition in particular, I think it's key to remember that the private test hours may be statistically different from the public test hours. In my opinion, your validation scheme should try to mirror those hours, and then should give you more confidence than the public leaderboard.   </p>",
      "rawMarkdown": "Here's a very common, sensible heuristic for choosing the two submissions:\n\n 1. Model/ensemble with highest local validation\n 2. Model/ensemble with highest public LB score\n\nBasically you want to be thinking about how to diversify your risk - if your submissions are essentially identical, there's no point in having more than 1. Local validation should typically be the best guide assuming you have a good setup and that the train data is similar to the test data, but the highest PLB submission is a way to partially hedge against the risk of overfitting to your local validation scheme.  \n\nSometimes 1 and 2 are the same, in which case it's harder to pick a second one. Again thinking about risk, I would typically try to choose a submission that's substantively different from the first. Maybe it's a single model, or leaves out some components of an ensemble or ideas that you're less confident in. This way you can get a submission that's \"safer\" from a complexity perspective (both in the sense of model complexity and entire pipeline complexity).\n\nIn this competition in particular, I think it's key to remember that the private test hours may be statistically different from the public test hours. In my opinion, your validation scheme should try to mirror those hours, and then should give you more confidence than the public leaderboard.",
      "votes": null
    },
    {
      "id": "322358",
      "postDate": "05/02/2018 20:02:11",
      "content": "<p>I would drop the model with the highest LB score (if is paired by a smaller CV than others), if I have few different models with higher CV and consistent LB score to choose from.</p>",
      "rawMarkdown": "I would drop the model with the highest LB score (if is paired by a smaller CV than others), if I have few different models with higher CV and consistent LB score to choose from.",
      "votes": null
    },
    {
      "id": "322375",
      "postDate": "05/02/2018 20:35:04",
      "content": "<p>I think in practice the decision should depend on a lot of factors - how big are the differences, how similar are the models you're choosing from, how big is the public LB, how much do you expect the public LB to differ from private, how much do you expect train to differ from test, etc. Can become a very complex decision process. </p>\n\n<p>I wouldn't advise to always go highest LB + highest local validation, but I think it's a good, relatively safe baseline strategy to choose if you're unsure what to do.</p>",
      "rawMarkdown": "I think in practice the decision should depend on a lot of factors - how big are the differences, how similar are the models you're choosing from, how big is the public LB, how much do you expect the public LB to differ from private, how much do you expect train to differ from test, etc. Can become a very complex decision process. \n\nI wouldn't advise to always go highest LB + highest local validation, but I think it's a good, relatively safe baseline strategy to choose if you're unsure what to do.",
      "votes": null
    },
    {
      "id": "322427",
      "postDate": "05/02/2018 22:44:54",
      "content": "<p>Thank you very much for your valuable comments. I will certainly use them at the end of the final models. Thank you again for your time and knowledge.</p>",
      "rawMarkdown": "Thank you very much for your valuable comments. I will certainly use them at the end of the final models. Thank you again for your time and knowledge.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 321934,
      "author_name": "yyqing",
      "author_url": "",
      "post_date": "05/02/2018 07:11:40",
      "content": "<p>Tried a few competitions on kaggle, learned some lessons like:  </p>\n\n<p>Rule No.1: trust your local cv;</p>\n\n<p>Rule No.x: follow the rule No.(x-1)</p>",
      "votes": null,
      "replies": [
        {
          "id": 321943,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/02/2018 07:22:36",
          "content": "<p>Hah, I understand. And what about the situation when I have models trained on different data sets? One time for 25M another time for 50M and so on. In this situation, it is impossible to compare their local CV effectively. CV is comparable only when it compares two models that were trained on the same data set and the same validation set. I choose the one that was trained on a larger collection? Is the one whose CV is closer to LB?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321964,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "05/02/2018 08:04:04",
          "content": "<p>I would prefer the result with more data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321968,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/02/2018 08:13:36",
          "content": "<p>Thank you so much for your help and tips.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321974,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "05/02/2018 08:29:03",
      "content": "<p>It may depend sometimes on luck.. </p>\n\n<p>On Statoil  competition , there were people who just submitted a crazy blend of blends of blends kernel and finished above me with silver medal.</p>\n\n<p>Note that the dataset was very small on mercedes. So the 2 are not necessarily comparable. </p>\n\n<p>We are doing true CV for some of our models but I guess some are just validating on a subset due to the size of the whole data.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 321978,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/02/2018 08:43:57",
          "content": "<p>In my case, I could not achieve any success using other models than LightGBM. My single best LGBM model is 0.9799 on LB score. So it seems to me that I do not even have anything to blend. Of course, I have a lot of LGBM models and I paln to average them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321979,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "05/02/2018 08:45:59",
          "content": "<p>I've tried this strategy in other competitions, blend the results from same model(+ different params), sometime it works.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321985,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/02/2018 08:58:43",
          "content": "<blockquote>\n  <p><strong>SkalskiP wrote</strong></p>\n  \n  <blockquote>\n    <p>In my case, I could not achieve any success using other models than LightGBM. My single best LGBM model is 0.9799 on LB score. So it seems to me that I do not even have anything to blend. Of course, I have a lot of LGBM models and I paln to average them.</p>\n  </blockquote>\n</blockquote>\n\n<p>As said by yyqing,  you can use differents strategies, differents params , differents features for the same model.   That's what we do with LGBM ( our best single LGBM is scoring 0.9809 with 13 features ) </p>\n\n<p>If it's any consolation, we have also a NN model scoring 0.9802 on LB but does not blend well with our LGBM  ( or we don't know how to do it correctly for now ^^) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321987,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/02/2018 09:04:12",
          "content": "<p>From what I read on the forum, LGBM is influenced by the order of the columns in DataFrame. And of course, random seed. So you say that it is a good idea to take my best model and pruba to put it on different parameters? And the average of the results from these few models?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321991,
          "author_name": "malten",
          "author_url": "",
          "post_date": "05/02/2018 09:12:20",
          "content": "<p>Can anyone confirm the dependency of LGB on column order and has  an explanation for the reason? I find it a bit disturbing, as from my understanding column_sample should take columns at random.</p>\n\n<p>Another strategy can be to build LGB on different subset of the data and average the results, makes even more sense here with different days that most likely have slightly different pattern as well</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321993,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/02/2018 09:19:57",
          "content": "<blockquote>\n  <p><strong>Serigne  wrote</strong>\n  As said by yyqing,  you can use differents strategies, differents params , differents features for the same model.   That's what we do with LGBM ( our best single LGBM is scoring 0.9809 with 13 features ) </p>\n</blockquote>\n\n<p>BTW: My God. I'm jealous of your such a strong model with so few features. In my case, I have about 25 features. I have a question by the way. Is it possible that my CV will improve when, for example, I remove one of my current features? In formulating this differently: Is it possible that any of my featurów worsens my result? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321996,
          "author_name": "malten",
          "author_url": "",
          "post_date": "05/02/2018 09:23:36",
          "content": "<p>Yes some features definately worsen the results, due to overfitting. For this reason most people in this competition use greedy stepwise procedures: start with an empty set of features and add 1 or more features. If the CV scores improve, then keep them, otherwise discard.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 322004,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/02/2018 09:37:45",
          "content": "<p>Thank you very much for your answer. I realize that my questions are probably trivial to most people who will read this. But I have three more and it would be great if someone could answer it:\n1. How can I interpret the feature importance obtained from the model? Does it mean to some extent whether the feature is good or not?\n2. Is getting rid of some basic feature such as 'app' or 'os' should I also take into account? Is the basic set from which I should start is 'app', 'os', 'channel' and 'devise' (I understand that 'ip' should not be taken into account, so I read in earlier posts). And from this set I start and add new features.\n3. Can I test models with new features on a smaller data set or on a smaller larning rate and be reasonably sure that the correct CV for such a model will translate into an improvement for my main model? My main model is tening on the 7th and 8th day and validation on 9. This calculation takes a lot of time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 322223,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "05/02/2018 15:19:03",
      "content": "<p>Here's a very common, sensible heuristic for choosing the two submissions:</p>\n\n<ol>\n<li>Model/ensemble with highest local validation</li>\n<li>Model/ensemble with highest public LB score</li>\n</ol>\n\n<p>Basically you want to be thinking about how to diversify your risk - if your submissions are essentially identical, there's no point in having more than 1. Local validation should typically be the best guide assuming you have a good setup and that the train data is similar to the test data, but the highest PLB submission is a way to partially hedge against the risk of overfitting to your local validation scheme.  </p>\n\n<p>Sometimes 1 and 2 are the same, in which case it's harder to pick a second one. Again thinking about risk, I would typically try to choose a submission that's substantively different from the first. Maybe it's a single model, or leaves out some components of an ensemble or ideas that you're less confident in. This way you can get a submission that's \"safer\" from a complexity perspective (both in the sense of model complexity and entire pipeline complexity).</p>\n\n<p>In this competition in particular, I think it's key to remember that the private test hours may be statistically different from the public test hours. In my opinion, your validation scheme should try to mirror those hours, and then should give you more confidence than the public leaderboard.   </p>",
      "votes": null,
      "replies": [
        {
          "id": 322358,
          "author_name": "gpreda",
          "author_url": "",
          "post_date": "05/02/2018 20:02:11",
          "content": "<p>I would drop the model with the highest LB score (if is paired by a smaller CV than others), if I have few different models with higher CV and consistent LB score to choose from.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 322375,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "05/02/2018 20:35:04",
          "content": "<p>I think in practice the decision should depend on a lot of factors - how big are the differences, how similar are the models you're choosing from, how big is the public LB, how much do you expect the public LB to differ from private, how much do you expect train to differ from test, etc. Can become a very complex decision process. </p>\n\n<p>I wouldn't advise to always go highest LB + highest local validation, but I think it's a good, relatively safe baseline strategy to choose if you're unsure what to do.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 322427,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/02/2018 22:44:54",
          "content": "<p>Thank you very much for your valuable comments. I will certainly use them at the end of the final models. Thank you again for your time and knowledge.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "321927": "Let me start by saying that this is my first Kaggle competition. Yesterday we received an email with a reminder about the choice of  entries for final submission. It means that slowly but inevitably we are approaching the end of the competition. I have question for you. In particular to experienced Kagglers. How do you choose entry?\n\nI read posts like [in CV you must trust][1] and I'm scared. People write there that they fell by more than 1000 places down. Will you make a decision based only on your CV? What other factors do you take into account? What other forum posts do you recommend reading?\n\n  [1]: https://www.kaggle.com/c/mercedes-benz-greener-manufacturing/discussion/36136",
    "321934": "Tried a few competitions on kaggle, learned some lessons like:  \n\nRule No.1: trust your local cv;\n\nRule No.x: follow the rule No.(x-1)",
    "321943": "Hah, I understand. And what about the situation when I have models trained on different data sets? One time for 25M another time for 50M and so on. In this situation, it is impossible to compare their local CV effectively. CV is comparable only when it compares two models that were trained on the same data set and the same validation set. I choose the one that was trained on a larger collection? Is the one whose CV is closer to LB?",
    "321964": "I would prefer the result with more data.",
    "321968": "Thank you so much for your help and tips.",
    "321974": "It may depend sometimes on luck.. \n\nOn Statoil  competition , there were people who just submitted a crazy blend of blends of blends kernel and finished above me with silver medal.\n\nNote that the dataset was very small on mercedes. So the 2 are not necessarily comparable. \n\nWe are doing true CV for some of our models but I guess some are just validating on a subset due to the size of the whole data.",
    "321978": "In my case, I could not achieve any success using other models than LightGBM. My single best LGBM model is 0.9799 on LB score. So it seems to me that I do not even have anything to blend. Of course, I have a lot of LGBM models and I paln to average them.",
    "321979": "I've tried this strategy in other competitions, blend the results from same model(+ different params), sometime it works.",
    "321985": "&gt; **SkalskiP wrote**\n&gt; \n&gt; &gt; In my case, I could not achieve any success using other models than LightGBM. My single best LGBM model is 0.9799 on LB score. So it seems to me that I do not even have anything to blend. Of course, I have a lot of LGBM models and I paln to average them.\n\nAs said by yyqing,  you can use differents strategies, differents params , differents features for the same model.   That's what we do with LGBM ( our best single LGBM is scoring 0.9809 with 13 features ) \n\nIf it's any consolation, we have also a NN model scoring 0.9802 on LB but does not blend well with our LGBM  ( or we don't know how to do it correctly for now ^^)",
    "321987": "From what I read on the forum, LGBM is influenced by the order of the columns in DataFrame. And of course, random seed. So you say that it is a good idea to take my best model and pruba to put it on different parameters? And the average of the results from these few models?",
    "321991": "Can anyone confirm the dependency of LGB on column order and has  an explanation for the reason? I find it a bit disturbing, as from my understanding column_sample should take columns at random.\n\nAnother strategy can be to build LGB on different subset of the data and average the results, makes even more sense here with different days that most likely have slightly different pattern as well",
    "321993": "&gt; **Serigne  wrote**\n&gt; As said by yyqing,  you can use differents strategies, differents params , differents features for the same model.   That's what we do with LGBM ( our best single LGBM is scoring 0.9809 with 13 features ) \n\nBTW: My God. I'm jealous of your such a strong model with so few features. In my case, I have about 25 features. I have a question by the way. Is it possible that my CV will improve when, for example, I remove one of my current features? In formulating this differently: Is it possible that any of my featurów worsens my result?",
    "321996": "Yes some features definately worsen the results, due to overfitting. For this reason most people in this competition use greedy stepwise procedures: start with an empty set of features and add 1 or more features. If the CV scores improve, then keep them, otherwise discard.",
    "322004": "Thank you very much for your answer. I realize that my questions are probably trivial to most people who will read this. But I have three more and it would be great if someone could answer it:\n1. How can I interpret the feature importance obtained from the model? Does it mean to some extent whether the feature is good or not?\n2. Is getting rid of some basic feature such as 'app' or 'os' should I also take into account? Is the basic set from which I should start is 'app', 'os', 'channel' and 'devise' (I understand that 'ip' should not be taken into account, so I read in earlier posts). And from this set I start and add new features.\n3. Can I test models with new features on a smaller data set or on a smaller larning rate and be reasonably sure that the correct CV for such a model will translate into an improvement for my main model? My main model is tening on the 7th and 8th day and validation on 9. This calculation takes a lot of time.",
    "322223": "Here's a very common, sensible heuristic for choosing the two submissions:\n\n 1. Model/ensemble with highest local validation\n 2. Model/ensemble with highest public LB score\n\nBasically you want to be thinking about how to diversify your risk - if your submissions are essentially identical, there's no point in having more than 1. Local validation should typically be the best guide assuming you have a good setup and that the train data is similar to the test data, but the highest PLB submission is a way to partially hedge against the risk of overfitting to your local validation scheme.  \n\nSometimes 1 and 2 are the same, in which case it's harder to pick a second one. Again thinking about risk, I would typically try to choose a submission that's substantively different from the first. Maybe it's a single model, or leaves out some components of an ensemble or ideas that you're less confident in. This way you can get a submission that's \"safer\" from a complexity perspective (both in the sense of model complexity and entire pipeline complexity).\n\nIn this competition in particular, I think it's key to remember that the private test hours may be statistically different from the public test hours. In my opinion, your validation scheme should try to mirror those hours, and then should give you more confidence than the public leaderboard.",
    "322358": "I would drop the model with the highest LB score (if is paired by a smaller CV than others), if I have few different models with higher CV and consistent LB score to choose from.",
    "322375": "I think in practice the decision should depend on a lot of factors - how big are the differences, how similar are the models you're choosing from, how big is the public LB, how much do you expect the public LB to differ from private, how much do you expect train to differ from test, etc. Can become a very complex decision process. \n\nI wouldn't advise to always go highest LB + highest local validation, but I think it's a good, relatively safe baseline strategy to choose if you're unsure what to do.",
    "322427": "Thank you very much for your valuable comments. I will certainly use them at the end of the final models. Thank you again for your time and knowledge."
  },
  "source": "meta"
}