{
  "id": 73953,
  "title": "variance in solutions for test vs train",
  "url": "/competitions/PLAsTiCC-2018/discussion/73953",
  "author_name": "",
  "post_date": "2018-12-07T01:02:08.282704300Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I'm back with another less than useful thread, but i hope you at least find it entertaining and maybe have an answer for me.  (whoever you are :) ) after my last thread about class 53, I've been plugging away at re-examining my Genetic Algorithm fitness evaluation process. I want to see if there are any good ways i can improve it. Thus far it really hasnt done anything that is any better than what GBM might do by itself (solutions used have given me a hair of an improvement or none at all). I've had no evolutions in to singular great features to use (without over fitting which then shows up as bad on the LB)</p>\n\n<p>So, that said, I dont want my GA evolving to a training-centric solution (which as i said it loves to do). i've come to a conclusion that i'm sure many people have previously come up. (with this and other data sets and contests.)  If I can check the test output to make sure its matches the training output statistically than at least i know i have a good solution. Assuming the test set's output is the same as the training on the whole.</p>\n\n<p>That is, is it reasonable in general to assume that if 10% of the solutions have class 53 as the answer in the training set, you can expect that in the test set (i made that up, i dont know what it is off hand. this is more of a question about data/contests in general anyway). Or is that a terrible assumption. Do people routinely probe the test set/leader board to get an idea of what the right percentages are?  (ie make sure their output lines up statistically with what they know to be true)</p>\n\n<p>It seems like that would be a great way to keep my GA program from overfitting (or at least another check, which isn't a bad thing). And it seems like it would be a HUGE trick for people engineering their own features. IE if class 53 is a 10% odds in the training set... but is supposed to be a 20% in the test set and i see a feature that strongly represents that... well i want to use that.</p>\n\n<p>thoughts? </p>",
  "messages": [
    {
      "id": "434805",
      "postDate": "12/07/2018 01:02:08",
      "content": "<p>I'm back with another less than useful thread, but i hope you at least find it entertaining and maybe have an answer for me.  (whoever you are :) ) after my last thread about class 53, I've been plugging away at re-examining my Genetic Algorithm fitness evaluation process. I want to see if there are any good ways i can improve it. Thus far it really hasnt done anything that is any better than what GBM might do by itself (solutions used have given me a hair of an improvement or none at all). I've had no evolutions in to singular great features to use (without over fitting which then shows up as bad on the LB)</p>\n\n<p>So, that said, I dont want my GA evolving to a training-centric solution (which as i said it loves to do). i've come to a conclusion that i'm sure many people have previously come up. (with this and other data sets and contests.)  If I can check the test output to make sure its matches the training output statistically than at least i know i have a good solution. Assuming the test set's output is the same as the training on the whole.</p>\n\n<p>That is, is it reasonable in general to assume that if 10% of the solutions have class 53 as the answer in the training set, you can expect that in the test set (i made that up, i dont know what it is off hand. this is more of a question about data/contests in general anyway). Or is that a terrible assumption. Do people routinely probe the test set/leader board to get an idea of what the right percentages are?  (ie make sure their output lines up statistically with what they know to be true)</p>\n\n<p>It seems like that would be a great way to keep my GA program from overfitting (or at least another check, which isn't a bad thing). And it seems like it would be a HUGE trick for people engineering their own features. IE if class 53 is a 10% odds in the training set... but is supposed to be a 20% in the test set and i see a feature that strongly represents that... well i want to use that.</p>\n\n<p>thoughts? </p>",
      "rawMarkdown": "I'm back with another less than useful thread, but i hope you at least find it entertaining and maybe have an answer for me.  (whoever you are :) ) after my last thread about class 53, I've been plugging away at re-examining my Genetic Algorithm fitness evaluation process. I want to see if there are any good ways i can improve it. Thus far it really hasnt done anything that is any better than what GBM might do by itself (solutions used have given me a hair of an improvement or none at all). I've had no evolutions in to singular great features to use (without over fitting which then shows up as bad on the LB)\n\nSo, that said, I dont want my GA evolving to a training-centric solution (which as i said it loves to do). i've come to a conclusion that i'm sure many people have previously come up. (with this and other data sets and contests.)  If I can check the test output to make sure its matches the training output statistically than at least i know i have a good solution. Assuming the test set's output is the same as the training on the whole.\n\nThat is, is it reasonable in general to assume that if 10% of the solutions have class 53 as the answer in the training set, you can expect that in the test set (i made that up, i dont know what it is off hand. this is more of a question about data/contests in general anyway). Or is that a terrible assumption. Do people routinely probe the test set/leader board to get an idea of what the right percentages are?  (ie make sure their output lines up statistically with what they know to be true)\n\nIt seems like that would be a great way to keep my GA program from overfitting (or at least another check, which isn't a bad thing). And it seems like it would be a HUGE trick for people engineering their own features. IE if class 53 is a 10% odds in the training set... but is supposed to be a 20% in the test set and i see a feature that strongly represents that... well i want to use that.\n\nthoughts?",
      "votes": null
    },
    {
      "id": "434929",
      "postDate": "12/07/2018 07:06:00",
      "content": "<blockquote>\n  <p>Do people routinely probe the test set/leader board to get an idea of what the right percentages are? </p>\n</blockquote>\n\n<p>The evaluation metric makes this impossible AFAIK.</p>",
      "rawMarkdown": "&gt; Do people routinely probe the test set/leader board to get an idea of what the right percentages are? \n\nThe evaluation metric makes this impossible AFAIK.",
      "votes": null
    },
    {
      "id": "435051",
      "postDate": "12/07/2018 11:48:25",
      "content": "<p>As long as your are not looking at extreme top scores you could also run one of the known public kernels and make a class count from its test results. (f.ex. <a href=\"https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss\">https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss</a> scores 1.08)</p>\n\n<p>Under any circumstance, you should be very carefull about using the meta-info \"hostgal_specz\" during training. \"hostgal_specz\" holds significant info and is present on <em>all training</em> data but only on 4% of the test data. Using it can therefore easily lead to very skewed models. </p>",
      "rawMarkdown": "As long as your are not looking at extreme top scores you could also run one of the known public kernels and make a class count from its test results. (f.ex. https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss scores 1.08)\n\nUnder any circumstance, you should be very carefull about using the meta-info \"hostgal_specz\" during training. \"hostgal_specz\" holds significant info and is present on *all training* data but only on 4% of the test data. Using it can therefore easily lead to very skewed models.",
      "votes": null
    },
    {
      "id": "435310",
      "postDate": "12/07/2018 20:44:31",
      "content": "<p>this is not really kernel related, its all custom code running on my home machine. but yeah, essentially it is meta info. so presumably its not a terrible idea just an easy trap to fall in to. (over committing to it)</p>",
      "rawMarkdown": "this is not really kernel related, its all custom code running on my home machine. but yeah, essentially it is meta info. so presumably its not a terrible idea just an easy trap to fall in to. (over committing to it)",
      "votes": null
    },
    {
      "id": "435318",
      "postDate": "12/07/2018 20:59:26",
      "content": "<p>probably right. (with the weighting and the logarithmic nature of what they did.) or at least its far more difficult than plugging in 15 equations from 15 test submissions  and solve the linear algebra to get a least squared fit for the the appropriate  percentage chance.</p>",
      "rawMarkdown": "probably right. (with the weighting and the logarithmic nature of what they did.) or at least its far more difficult than plugging in 15 equations from 15 test submissions  and solve the linear algebra to get a least squared fit for the the appropriate  percentage chance.",
      "votes": null
    },
    {
      "id": "435338",
      "postDate": "12/07/2018 21:52:30",
      "content": "<p>it's worth adding i took this idea a bit further and ran with it. instead of using simple averages i think maybe using a normal curve and predicting the likely hood the two normal curves are the same (there is a mathy way to do this) it gives you a Probability likelyhood. assuming i crossed all by t's and dotted all my i's correctly i can compare a sample of the test results vs the training results and the normal curves for a given solution should be similiar adn if they are not i have an exact probability that 1 could happen given the other. i use that probability to adjust my fitness. so far.... i mean it does stuff. does it work ? maybe, i'll know in a few hours when this run is done</p>",
      "rawMarkdown": "it's worth adding i took this idea a bit further and ran with it. instead of using simple averages i think maybe using a normal curve and predicting the likely hood the two normal curves are the same (there is a mathy way to do this) it gives you a Probability likelyhood. assuming i crossed all by t's and dotted all my i's correctly i can compare a sample of the test results vs the training results and the normal curves for a given solution should be similiar adn if they are not i have an exact probability that 1 could happen given the other. i use that probability to adjust my fitness. so far.... i mean it does stuff. does it work ? maybe, i'll know in a few hours when this run is done",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 434929,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/07/2018 07:06:00",
      "content": "<blockquote>\n  <p>Do people routinely probe the test set/leader board to get an idea of what the right percentages are? </p>\n</blockquote>\n\n<p>The evaluation metric makes this impossible AFAIK.</p>",
      "votes": null,
      "replies": [
        {
          "id": 435318,
          "author_name": "jscheibel",
          "author_url": "",
          "post_date": "12/07/2018 20:59:26",
          "content": "<p>probably right. (with the weighting and the logarithmic nature of what they did.) or at least its far more difficult than plugging in 15 equations from 15 test submissions  and solve the linear algebra to get a least squared fit for the the appropriate  percentage chance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 435051,
      "author_name": "petersorensen360",
      "author_url": "",
      "post_date": "12/07/2018 11:48:25",
      "content": "<p>As long as your are not looking at extreme top scores you could also run one of the known public kernels and make a class count from its test results. (f.ex. <a href=\"https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss\">https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss</a> scores 1.08)</p>\n\n<p>Under any circumstance, you should be very carefull about using the meta-info \"hostgal_specz\" during training. \"hostgal_specz\" holds significant info and is present on <em>all training</em> data but only on 4% of the test data. Using it can therefore easily lead to very skewed models. </p>",
      "votes": null,
      "replies": [
        {
          "id": 435310,
          "author_name": "jscheibel",
          "author_url": "",
          "post_date": "12/07/2018 20:44:31",
          "content": "<p>this is not really kernel related, its all custom code running on my home machine. but yeah, essentially it is meta info. so presumably its not a terrible idea just an easy trap to fall in to. (over committing to it)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 435338,
      "author_name": "jscheibel",
      "author_url": "",
      "post_date": "12/07/2018 21:52:30",
      "content": "<p>it's worth adding i took this idea a bit further and ran with it. instead of using simple averages i think maybe using a normal curve and predicting the likely hood the two normal curves are the same (there is a mathy way to do this) it gives you a Probability likelyhood. assuming i crossed all by t's and dotted all my i's correctly i can compare a sample of the test results vs the training results and the normal curves for a given solution should be similiar adn if they are not i have an exact probability that 1 could happen given the other. i use that probability to adjust my fitness. so far.... i mean it does stuff. does it work ? maybe, i'll know in a few hours when this run is done</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "434805": "I'm back with another less than useful thread, but i hope you at least find it entertaining and maybe have an answer for me.  (whoever you are :) ) after my last thread about class 53, I've been plugging away at re-examining my Genetic Algorithm fitness evaluation process. I want to see if there are any good ways i can improve it. Thus far it really hasnt done anything that is any better than what GBM might do by itself (solutions used have given me a hair of an improvement or none at all). I've had no evolutions in to singular great features to use (without over fitting which then shows up as bad on the LB)\n\nSo, that said, I dont want my GA evolving to a training-centric solution (which as i said it loves to do). i've come to a conclusion that i'm sure many people have previously come up. (with this and other data sets and contests.)  If I can check the test output to make sure its matches the training output statistically than at least i know i have a good solution. Assuming the test set's output is the same as the training on the whole.\n\nThat is, is it reasonable in general to assume that if 10% of the solutions have class 53 as the answer in the training set, you can expect that in the test set (i made that up, i dont know what it is off hand. this is more of a question about data/contests in general anyway). Or is that a terrible assumption. Do people routinely probe the test set/leader board to get an idea of what the right percentages are?  (ie make sure their output lines up statistically with what they know to be true)\n\nIt seems like that would be a great way to keep my GA program from overfitting (or at least another check, which isn't a bad thing). And it seems like it would be a HUGE trick for people engineering their own features. IE if class 53 is a 10% odds in the training set... but is supposed to be a 20% in the test set and i see a feature that strongly represents that... well i want to use that.\n\nthoughts?",
    "434929": "&gt; Do people routinely probe the test set/leader board to get an idea of what the right percentages are? \n\nThe evaluation metric makes this impossible AFAIK.",
    "435051": "As long as your are not looking at extreme top scores you could also run one of the known public kernels and make a class count from its test results. (f.ex. https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss scores 1.08)\n\nUnder any circumstance, you should be very carefull about using the meta-info \"hostgal_specz\" during training. \"hostgal_specz\" holds significant info and is present on *all training* data but only on 4% of the test data. Using it can therefore easily lead to very skewed models.",
    "435310": "this is not really kernel related, its all custom code running on my home machine. but yeah, essentially it is meta info. so presumably its not a terrible idea just an easy trap to fall in to. (over committing to it)",
    "435318": "probably right. (with the weighting and the logarithmic nature of what they did.) or at least its far more difficult than plugging in 15 equations from 15 test submissions  and solve the linear algebra to get a least squared fit for the the appropriate  percentage chance.",
    "435338": "it's worth adding i took this idea a bit further and ran with it. instead of using simple averages i think maybe using a normal curve and predicting the likely hood the two normal curves are the same (there is a mathy way to do this) it gives you a Probability likelyhood. assuming i crossed all by t's and dotted all my i's correctly i can compare a sample of the test results vs the training results and the normal curves for a given solution should be similiar adn if they are not i have an exact probability that 1 could happen given the other. i use that probability to adjust my fitness. so far.... i mean it does stuff. does it work ? maybe, i'll know in a few hours when this run is done"
  },
  "source": "meta"
}