{
  "id": 72724,
  "title": "question about the difficulty of class 53",
  "url": "/competitions/PLAsTiCC-2018/discussion/72724",
  "author_name": "",
  "post_date": "2018-11-26T17:53:15.817872700Z",
  "votes": 1,
  "comment_count": 19,
  "views": 0,
  "content": "<p>i haven't really been working too much on analyzing the data (i never do, that's my old story). </p>\n\n<p>That being said I left my genetic algorithm working on the data over the last 4 days. And 1 class in particular (class 53 or rather class 6  of our 15) it was able to predict with 100% accuracy against the training data. I wondered if this was just a serious case of over-fitting or if that particular class is actually really easy to predict. I did it twice using two different fitness tests and one was 100% and the other was all but.</p>\n\n<p>basically, i was wondering if any of you humans analysis people :) (vs my crazy machine analysis) came up with \"yeah, class 53 is really easy.\"</p>",
  "messages": [
    {
      "id": "428073",
      "postDate": "11/26/2018 17:53:15",
      "content": "<p>i haven't really been working too much on analyzing the data (i never do, that's my old story). </p>\n\n<p>That being said I left my genetic algorithm working on the data over the last 4 days. And 1 class in particular (class 53 or rather class 6  of our 15) it was able to predict with 100% accuracy against the training data. I wondered if this was just a serious case of over-fitting or if that particular class is actually really easy to predict. I did it twice using two different fitness tests and one was 100% and the other was all but.</p>\n\n<p>basically, i was wondering if any of you humans analysis people :) (vs my crazy machine analysis) came up with \"yeah, class 53 is really easy.\"</p>",
      "rawMarkdown": "i haven't really been working too much on analyzing the data (i never do, that's my old story). \n\nThat being said I left my genetic algorithm working on the data over the last 4 days. And 1 class in particular (class 53 or rather class 6  of our 15) it was able to predict with 100% accuracy against the training data. I wondered if this was just a serious case of over-fitting or if that particular class is actually really easy to predict. I did it twice using two different fitness tests and one was 100% and the other was all but.\n\nbasically, i was wondering if any of you humans analysis people :) (vs my crazy machine analysis) came up with \"yeah, class 53 is really easy.\"",
      "votes": null
    },
    {
      "id": "428077",
      "postDate": "11/26/2018 18:05:04",
      "content": "<p>Reading the forum can be a proxy for analyzing the data ;)</p>\n\n<p>Class 53 is very easy, as shown by the confusion matrix in these posts:</p>\n\n<p><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70669\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70669</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72613\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72613</a></p>",
      "rawMarkdown": "Reading the forum can be a proxy for analyzing the data ;)\n\nClass 53 is very easy, as shown by the confusion matrix in these posts:\n\nhttps://www.kaggle.com/c/PLAsTiCC-2018/discussion/70669\n\nhttps://www.kaggle.com/c/PLAsTiCC-2018/discussion/72613",
      "votes": null
    },
    {
      "id": "428092",
      "postDate": "11/26/2018 18:32:23",
      "content": "<p>cool thanks! you can understand my initial skepticism. when i saw it i was like \"100% ... sounds to good to be true. what did i do wrong.\" </p>",
      "rawMarkdown": "cool thanks! you can understand my initial skepticism. when i saw it i was like \"100% ... sounds to good to be true. what did i do wrong.\"",
      "votes": null
    },
    {
      "id": "428100",
      "postDate": "11/26/2018 18:50:55",
      "content": "<p>this is actually fascinating enough to share. it lines up pretty well with your matrix.</p>\n\n<p>I rank my final genetic predictions against the training data as a whole.  I get the following</p>\n\n<p>class score1  score2\n6   0.9234418 0.9569862\n15  0.8641432 0.8884696\n16  0.9643756 0.970057\n42  0.7077606 0.7526829\n52  0.8414314 0.8528698\n53  0.9940292 1\n62  0.8336724 0.8529119\n64  0.9847192 -still running-\n65  0.9588371 -still running-\n67  0.8774132 -still running-\n88  0.9793245 -still running-\n90  0.6686415 -still running-\n92  0.9801163 -still running-\n95  0.9696037 -still running-</p>\n\n<p>score1 was using a different fitness test. I made some changes and now i'm getting better results as you can see in fitness score2. (still running though)</p>\n\n<p>unfortunately the number i'm using isn't an apples to apples with the confusion matrix you made but it seems they are very similar.</p>",
      "rawMarkdown": "this is actually fascinating enough to share. it lines up pretty well with your matrix.\n\nI rank my final genetic predictions against the training data as a whole.  I get the following\n\nclass score1  score2\n6\t0.9234418 0.9569862\n15\t0.8641432 0.8884696\n16\t0.9643756 0.970057\n42\t0.7077606 0.7526829\n52\t0.8414314 0.8528698\n53\t0.9940292 1\n62\t0.8336724 0.8529119\n64\t0.9847192 -still running-\n65\t0.9588371 -still running-\n67\t0.8774132 -still running-\n88\t0.9793245 -still running-\n90\t0.6686415 -still running-\n92\t0.9801163 -still running-\n95\t0.9696037 -still running-\n\nscore1 was using a different fitness test. I made some changes and now i'm getting better results as you can see in fitness score2. (still running though)\n\nunfortunately the number i'm using isn't an apples to apples with the confusion matrix you made but it seems they are very similar.",
      "votes": null
    },
    {
      "id": "428138",
      "postDate": "11/26/2018 20:25:52",
      "content": "<p>that metrics are calculated using a validation fold?</p>",
      "rawMarkdown": "that metrics are calculated using a validation fold?",
      "votes": null
    },
    {
      "id": "428213",
      "postDate": "11/26/2018 23:36:15",
      "content": "<p>i'm not sure i understand the question, or was that directed at cpmp?</p>",
      "rawMarkdown": "i'm not sure i understand the question, or was that directed at cpmp?",
      "votes": null
    },
    {
      "id": "428292",
      "postDate": "11/27/2018 03:17:56",
      "content": "<p>Genetic algorithm for generating features?  that's cool</p>",
      "rawMarkdown": "Genetic algorithm for generating features?  that's cool",
      "votes": null
    },
    {
      "id": "428340",
      "postDate": "11/27/2018 05:34:17",
      "content": "<p>yeah, i used to use it to solve whole problems, but for various reasons i think its much better suited at building features.</p>",
      "rawMarkdown": "yeah, i used to use it to solve whole problems, but for various reasons i think its much better suited at building features.",
      "votes": null
    },
    {
      "id": "428409",
      "postDate": "11/27/2018 08:26:57",
      "content": "<p>If it is not too much trouble, how do you set up the fitness function and chromosomes of GA?</p>",
      "rawMarkdown": "If it is not too much trouble, how do you set up the fitness function and chromosomes of GA?",
      "votes": null
    },
    {
      "id": "428419",
      "postDate": "11/27/2018 08:51:26",
      "content": "<p>I guess you answered the question if you don't understand it ;)  The question is for you, not me ;)</p>\n\n<p>The question is this: when you say you have 100% accuracy predicting class 53, how do you compute the 100%  Ideally you run your genetic algorithm on some of the data, then see what it yields on another part of the data.  This is called a train/test split validation.  A n even better way is k folds cross validation.  </p>\n\n<p>SO, what is your way of evaluating the quality of your GA ?</p>",
      "rawMarkdown": "I guess you answered the question if you don't understand it ;)  The question is for you, not me ;)\n\nThe question is this: when you say you have 100% accuracy predicting class 53, how do you compute the 100%  Ideally you run your genetic algorithm on some of the data, then see what it yields on another part of the data.  This is called a train/test split validation.  A n even better way is k folds cross validation.  \n\nSO, what is your way of evaluating the quality of your GA ?",
      "votes": null
    },
    {
      "id": "428554",
      "postDate": "11/27/2018 13:15:38",
      "content": "<p>@j_scheibel, @CPMP  is right, the question is for you. You must evaluate the metrics in a validation or test set and not in the same train set you fitted the GA. My train set scores are pretty close to 1.0 for all classes, but its not the same in the validation set ;)</p>",
      "rawMarkdown": "j_scheibel, @CPMP  is right, the question is for you. You must evaluate the metrics in a validation or test set and not in the same train set you fitted the GA. My train set scores are pretty close to 1.0 for all classes, but its not the same in the validation set ;)",
      "votes": null
    },
    {
      "id": "428581",
      "postDate": "11/27/2018 14:21:05",
      "content": "<blockquote>\n  <p>but its not the same in the validation set ;)</p>\n</blockquote>\n\n<p>And it is even worse on the public LB data ;)</p>",
      "rawMarkdown": "&gt; but its not the same in the validation set ;)\n\nAnd it is even worse on the public LB data ;)",
      "votes": null
    },
    {
      "id": "428650",
      "postDate": "11/27/2018 16:58:38",
      "content": "<p>I know what cross validation is :) .  I swear i read giba's post differently yesterday and it didn't make sense to me then. It doesn't matter though, because yes... I've been using either 3, 5 or 9 fold validation on anything i've ever done here that is conventional data mining. </p>\n\n<p>That being said this is not conventional data mining.  </p>\n\n<p>The problem is, each iteration will grow to fit whatever you are using. It <em>will</em> over fit. that's what it does.  If I gave it all the data to work with  (and i have before, so I've seen this first hand) it fits to it. if i hold out some of the data for a while but then use it in evaluation later, whenever i start using it things become contaminated. so you have to keep it out forever. </p>\n\n<p>the thing is if you do 3 fold cross validation. (to keep it simple) You have a need to do 3 runs . each with 2/3 of the data (since you can never cross contaminate them)... so those genes in each pile, what? never interact?  if you do that you will end up with 3 completely different pictures. each that have nothing to do with each other. also each will completely overfit to it's 2/3 of the data. </p>\n\n<p>you could ensemble them but that's merging 3 over fitted things and is not what we are going for (that path leads to random forests .. with genes instead of trees. ). we want something that actually produces a good result with 1 solution.  in short cross validation as a  scoring approach doesn't work here. so what does?</p>\n\n<p>to answer the question originally asked by giba.  no there is no folded cross validation in those scores (there cannot be for the reason i tried to explain above  (hrumph i say! hrumph! - hehehe) )  I do have a solution to the problem and it does seem to work. (it has nothing to do with this data persay) a solution where i get 1 set of instructions that i evolved iteratively that produces a result that is fitted to the data that while completely biased (it has to be! all data mining is) is not an over fit. It takes some of explanation... I really wasn't going for any of that here, i would not want to defend the things i've done till i have a reason to  ... </p>\n\n<p>i just wanted to a sanity check on the 100% i got on class 53. :)  </p>\n\n<p><em>edited this some to be clear</em>\nthe scores i listed actually happen to be the same algorithm i'm using as the fitness test mechanism's scoring tool. it's a measure of accuracy. I've actually tried lots of scoring mechanisms for fitness RMSE , RMSLE, Correlation, AUC etc  I was using correlation before now (that's actually what score 1 is using under the covers) the columns value though comes from something I created a few years ago when i needed a way to measure accuracy of things that ... well didn't really have accuracy. It uses a normal curve from the answers and figures out the odds (percentage) of a particular event and compares it with odds of the predicted value. so really it a 0 is  a perfect match and 100% is completely wrong, i just invert it. it turns out that for making features its seems better as a fitness fucntion as well. i do like correlation a lot though. </p>",
      "rawMarkdown": "I know what cross validation is :) .  I swear i read giba's post differently yesterday and it didn't make sense to me then. It doesn't matter though, because yes... I've been using either 3, 5 or 9 fold validation on anything i've ever done here that is conventional data mining. \n\nThat being said this is not conventional data mining.  \n\nThe problem is, each iteration will grow to fit whatever you are using. It _will_ over fit. that's what it does.  If I gave it all the data to work with  (and i have before, so I've seen this first hand) it fits to it. if i hold out some of the data for a while but then use it in evaluation later, whenever i start using it things become contaminated. so you have to keep it out forever. \n\nthe thing is if you do 3 fold cross validation. (to keep it simple) You have a need to do 3 runs . each with 2/3 of the data (since you can never cross contaminate them)... so those genes in each pile, what? never interact?  if you do that you will end up with 3 completely different pictures. each that have nothing to do with each other. also each will completely overfit to it's 2/3 of the data. \n\nyou could ensemble them but that's merging 3 over fitted things and is not what we are going for (that path leads to random forests .. with genes instead of trees. ). we want something that actually produces a good result with 1 solution.  in short cross validation as a  scoring approach doesn't work here. so what does?\n\nto answer the question originally asked by giba.  no there is no folded cross validation in those scores (there cannot be for the reason i tried to explain above  (hrumph i say! hrumph! - hehehe) )  I do have a solution to the problem and it does seem to work. (it has nothing to do with this data persay) a solution where i get 1 set of instructions that i evolved iteratively that produces a result that is fitted to the data that while completely biased (it has to be! all data mining is) is not an over fit. It takes some of explanation... I really wasn't going for any of that here, i would not want to defend the things i've done till i have a reason to  ... \n\ni just wanted to a sanity check on the 100% i got on class 53. :)  \n\n*edited this some to be clear*\nthe scores i listed actually happen to be the same algorithm i'm using as the fitness test mechanism's scoring tool. it's a measure of accuracy. I've actually tried lots of scoring mechanisms for fitness RMSE , RMSLE, Correlation, AUC etc  I was using correlation before now (that's actually what score 1 is using under the covers) the columns value though comes from something I created a few years ago when i needed a way to measure accuracy of things that ... well didn't really have accuracy. It uses a normal curve from the answers and figures out the odds (percentage) of a particular event and compares it with odds of the predicted value. so really it a 0 is  a perfect match and 100% is completely wrong, i just invert it. it turns out that for making features its seems better as a fitness fucntion as well. i do like correlation a lot though.",
      "votes": null
    },
    {
      "id": "428655",
      "postDate": "11/27/2018 17:04:23",
      "content": "<blockquote>\n  <p>if you do that you will end up with 3 completely different pictures. each that have nothing to do with each other. also each will completely overfit to it's 2/3 of the data. </p>\n</blockquote>\n\n<p>Exactly, and you keep these 3 totally separate.  Then you average their validation scores, and that's it.</p>\n\n<p>It is clearly not what you are doing, and I must say I am not sure about how you do things, it is still very confusing to me.  I'm probably not smart enough to follow you, but I'll make an effort if this translates to good LB score.</p>",
      "rawMarkdown": "&gt;  if you do that you will end up with 3 completely different pictures. each that have nothing to do with each other. also each will completely overfit to it's 2/3 of the data. \n\nExactly, and you keep these 3 totally separate.  Then you average their validation scores, and that's it.\n\nIt is clearly not what you are doing, and I must say I am not sure about how you do things, it is still very confusing to me.  I'm probably not smart enough to follow you, but I'll make an effort if this translates to good LB score.",
      "votes": null
    },
    {
      "id": "428659",
      "postDate": "11/27/2018 17:09:44",
      "content": "<p>I evolve my chromosomes (you know, i always call them genes in my head. rolls of the tongue easier but i digress) to get the most accurate answer i can (in isolation from the other classes). And by most accurate answer i can, i mean i'm using a normal curve based on the answer set which I then compare to the expected result of each answer vs the predicted answer on said curve. i turn those points in to percentage odds and measure the gap. i'm trying to minimize the over all gap. (see the paragraph above in cpmp's thread) is that what you are asking? </p>",
      "rawMarkdown": "I evolve my chromosomes (you know, i always call them genes in my head. rolls of the tongue easier but i digress) to get the most accurate answer i can (in isolation from the other classes). And by most accurate answer i can, i mean i'm using a normal curve based on the answer set which I then compare to the expected result of each answer vs the predicted answer on said curve. i turn those points in to percentage odds and measure the gap. i'm trying to minimize the over all gap. (see the paragraph above in cpmp's thread) is that what you are asking?",
      "votes": null
    },
    {
      "id": "428666",
      "postDate": "11/27/2018 17:24:15",
      "content": "<p>lol says the guy in 1st....</p>\n\n<p>it's all good. i worked on that post for like an hour. i really didn't want to say anything that could be construed as \"you're an idiot\" i know what i'm doing but i don't want to defend it to my peers right now.</p>\n\n<p>i've spent a lot ... i mean A LOT of time working on algorithms and trying to conceptualize what needs to happen, and then even more time writing them and trying things (just to see).  Ironically, this whole thread isn't even related to the main goal. i wasn't even going to run my GA stuff on this contest but i wanted an apples to apples comparison with a new expert system i'm working on. </p>\n\n<p>That one is going to process streamed data, i thought it could try and deal with the weird timeseries we have here. so I figured i'd fire up the GA get some features, run them through my data minnig tool and see what they add to my score before i changed over. The score i currently have doesn't use anything from this post. its just my own personal gradient boosting implementation using a handful of features. the ones provided and from the time series, averages, min, max and i actually bothered to implement the period finding fft (what a pain). i couldnt find a c# implementation so i converted a java one. it worked well-ish but it doesnt use the error margins like the scikit one does.</p>\n\n<p>i have a lot of fun. :) </p>",
      "rawMarkdown": "lol says the guy in 1st....\n\nit's all good. i worked on that post for like an hour. i really didn't want to say anything that could be construed as \"you're an idiot\" i know what i'm doing but i don't want to defend it to my peers right now.\n\n i've spent a lot ... i mean A LOT of time working on algorithms and trying to conceptualize what needs to happen, and then even more time writing them and trying things (just to see).  Ironically, this whole thread isn't even related to the main goal. i wasn't even going to run my GA stuff on this contest but i wanted an apples to apples comparison with a new expert system i'm working on. \n\nThat one is going to process streamed data, i thought it could try and deal with the weird timeseries we have here. so I figured i'd fire up the GA get some features, run them through my data minnig tool and see what they add to my score before i changed over. The score i currently have doesn't use anything from this post. its just my own personal gradient boosting implementation using a handful of features. the ones provided and from the time series, averages, min, max and i actually bothered to implement the period finding fft (what a pain). i couldnt find a c# implementation so i converted a java one. it worked well-ish but it doesnt use the error margins like the scikit one does.\n\ni have a lot of fun. :)",
      "votes": null
    },
    {
      "id": "428688",
      "postDate": "11/27/2018 18:07:01",
      "content": "<p>I'm serious, and I never thought you were an idiot.  I'll reread your posts with a fresh head tomorrow ;)</p>",
      "rawMarkdown": "I'm serious, and I never thought you were an idiot.  I'll reread your posts with a fresh head tomorrow ;)",
      "votes": null
    },
    {
      "id": "428872",
      "postDate": "11/28/2018 02:20:25",
      "content": "<p>thx for reply.</p>",
      "rawMarkdown": "thx for reply.",
      "votes": null
    },
    {
      "id": "430591",
      "postDate": "11/30/2018 16:58:09",
      "content": "<p>Are you worried you aren't normalizing all the targets (Softmax for example).  Doing these individually will probably end in tears - I hope not considering the amount of time you have spent training - good luck! ;)</p>",
      "rawMarkdown": "Are you worried you aren't normalizing all the targets (Softmax for example).  Doing these individually will probably end in tears - I hope not considering the amount of time you have spent training - good luck! ;)",
      "votes": null
    },
    {
      "id": "430635",
      "postDate": "11/30/2018 18:52:01",
      "content": "<p>well i'm just making features here so they don't really need to be normalized. the results i get for each group go in as new columns and are classically mined using a version of gbm. but I do normalize them. i've skipped a lot of the details. 1 important one is i do have a hold out that i use for that exact normalization. each feature is only calculated on 63.2% of the data when the run is done i use linear algebra using the hold out data to scale and offset the results appropriately. </p>\n\n<p>As for combining them up. I'm sure i could work on that some more, there may be a better way mathematically, but I'll tell you what i do as it works \"okay\". (this of course has nothing to do with anything other than this contest) if a class predicts 72% i take the 28% it's unsure of and distribute it among the other classes. except of course the other classes have their own prediction except class 99. so that would get 2%... do that for all 14 classes and take the average for class 99.  the final prediction is scaled to add up to 1 using the predictions as weights. the up shot is, it works just like a human would do it naively. if its not any of the other 14 classes it must be unknown.</p>\n\n<p>here's a simple example using 4 classes 1 is unknown</p>\n\n<p>a:15%\nb:40%\nc:1%\nd:unknown</p>\n\n<p>unknown = (0.85/3 + .6/3 + .99/3)/3.0\nunknown = (.283 + 0.2 + .33)/3.0\nunknown = 27.067%</p>\n\n<p>so final values then need to add up to 1 they currently add up to 0.83067 so divide each answer by .83067 you get</p>\n\n<p>a:18.06%\nb:48.15%\nc:1.2%\nd:32.58%</p>\n\n<p>.... an improvement on this would probably be to see if the other classes do add up to less than 100% to maybe add more in to unknown.  i don't think i've tried that yet. the thing is you want an apples to apples and the final normalization tends to take care of that. you also dont want negative values. I played around with it a bunch originally but nothing seemed to work well other than the above. I should probably see if there is more to be done... but i've been  busy playing rimworld. :)</p>",
      "rawMarkdown": "well i'm just making features here so they don't really need to be normalized. the results i get for each group go in as new columns and are classically mined using a version of gbm. but I do normalize them. i've skipped a lot of the details. 1 important one is i do have a hold out that i use for that exact normalization. each feature is only calculated on 63.2% of the data when the run is done i use linear algebra using the hold out data to scale and offset the results appropriately. \n\nAs for combining them up. I'm sure i could work on that some more, there may be a better way mathematically, but I'll tell you what i do as it works \"okay\". (this of course has nothing to do with anything other than this contest) if a class predicts 72% i take the 28% it's unsure of and distribute it among the other classes. except of course the other classes have their own prediction except class 99. so that would get 2%... do that for all 14 classes and take the average for class 99.  the final prediction is scaled to add up to 1 using the predictions as weights. the up shot is, it works just like a human would do it naively. if its not any of the other 14 classes it must be unknown.\n\nhere's a simple example using 4 classes 1 is unknown\n\na:15%\nb:40%\nc:1%\nd:unknown\n\nunknown = (0.85/3 + .6/3 + .99/3)/3.0\nunknown = (.283 + 0.2 + .33)/3.0\nunknown = 27.067%\n\nso final values then need to add up to 1 they currently add up to 0.83067 so divide each answer by .83067 you get\n\na:18.06%\nb:48.15%\nc:1.2%\nd:32.58%\n\n.... an improvement on this would probably be to see if the other classes do add up to less than 100% to maybe add more in to unknown.  i don't think i've tried that yet. the thing is you want an apples to apples and the final normalization tends to take care of that. you also dont want negative values. I played around with it a bunch originally but nothing seemed to work well other than the above. I should probably see if there is more to be done... but i've been  busy playing rimworld. :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 428077,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/26/2018 18:05:04",
      "content": "<p>Reading the forum can be a proxy for analyzing the data ;)</p>\n\n<p>Class 53 is very easy, as shown by the confusion matrix in these posts:</p>\n\n<p><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70669\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70669</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72613\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72613</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 428092,
          "author_name": "jscheibel",
          "author_url": "",
          "post_date": "11/26/2018 18:32:23",
          "content": "<p>cool thanks! you can understand my initial skepticism. when i saw it i was like \"100% ... sounds to good to be true. what did i do wrong.\" </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428100,
          "author_name": "jscheibel",
          "author_url": "",
          "post_date": "11/26/2018 18:50:55",
          "content": "<p>this is actually fascinating enough to share. it lines up pretty well with your matrix.</p>\n\n<p>I rank my final genetic predictions against the training data as a whole.  I get the following</p>\n\n<p>class score1  score2\n6   0.9234418 0.9569862\n15  0.8641432 0.8884696\n16  0.9643756 0.970057\n42  0.7077606 0.7526829\n52  0.8414314 0.8528698\n53  0.9940292 1\n62  0.8336724 0.8529119\n64  0.9847192 -still running-\n65  0.9588371 -still running-\n67  0.8774132 -still running-\n88  0.9793245 -still running-\n90  0.6686415 -still running-\n92  0.9801163 -still running-\n95  0.9696037 -still running-</p>\n\n<p>score1 was using a different fitness test. I made some changes and now i'm getting better results as you can see in fitness score2. (still running though)</p>\n\n<p>unfortunately the number i'm using isn't an apples to apples with the confusion matrix you made but it seems they are very similar.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428138,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "11/26/2018 20:25:52",
          "content": "<p>that metrics are calculated using a validation fold?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428213,
          "author_name": "jscheibel",
          "author_url": "",
          "post_date": "11/26/2018 23:36:15",
          "content": "<p>i'm not sure i understand the question, or was that directed at cpmp?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428419,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/27/2018 08:51:26",
          "content": "<p>I guess you answered the question if you don't understand it ;)  The question is for you, not me ;)</p>\n\n<p>The question is this: when you say you have 100% accuracy predicting class 53, how do you compute the 100%  Ideally you run your genetic algorithm on some of the data, then see what it yields on another part of the data.  This is called a train/test split validation.  A n even better way is k folds cross validation.  </p>\n\n<p>SO, what is your way of evaluating the quality of your GA ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428554,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "11/27/2018 13:15:38",
          "content": "<p>@j_scheibel, @CPMP  is right, the question is for you. You must evaluate the metrics in a validation or test set and not in the same train set you fitted the GA. My train set scores are pretty close to 1.0 for all classes, but its not the same in the validation set ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428581,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/27/2018 14:21:05",
          "content": "<blockquote>\n  <p>but its not the same in the validation set ;)</p>\n</blockquote>\n\n<p>And it is even worse on the public LB data ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428650,
          "author_name": "jscheibel",
          "author_url": "",
          "post_date": "11/27/2018 16:58:38",
          "content": "<p>I know what cross validation is :) .  I swear i read giba's post differently yesterday and it didn't make sense to me then. It doesn't matter though, because yes... I've been using either 3, 5 or 9 fold validation on anything i've ever done here that is conventional data mining. </p>\n\n<p>That being said this is not conventional data mining.  </p>\n\n<p>The problem is, each iteration will grow to fit whatever you are using. It <em>will</em> over fit. that's what it does.  If I gave it all the data to work with  (and i have before, so I've seen this first hand) it fits to it. if i hold out some of the data for a while but then use it in evaluation later, whenever i start using it things become contaminated. so you have to keep it out forever. </p>\n\n<p>the thing is if you do 3 fold cross validation. (to keep it simple) You have a need to do 3 runs . each with 2/3 of the data (since you can never cross contaminate them)... so those genes in each pile, what? never interact?  if you do that you will end up with 3 completely different pictures. each that have nothing to do with each other. also each will completely overfit to it's 2/3 of the data. </p>\n\n<p>you could ensemble them but that's merging 3 over fitted things and is not what we are going for (that path leads to random forests .. with genes instead of trees. ). we want something that actually produces a good result with 1 solution.  in short cross validation as a  scoring approach doesn't work here. so what does?</p>\n\n<p>to answer the question originally asked by giba.  no there is no folded cross validation in those scores (there cannot be for the reason i tried to explain above  (hrumph i say! hrumph! - hehehe) )  I do have a solution to the problem and it does seem to work. (it has nothing to do with this data persay) a solution where i get 1 set of instructions that i evolved iteratively that produces a result that is fitted to the data that while completely biased (it has to be! all data mining is) is not an over fit. It takes some of explanation... I really wasn't going for any of that here, i would not want to defend the things i've done till i have a reason to  ... </p>\n\n<p>i just wanted to a sanity check on the 100% i got on class 53. :)  </p>\n\n<p><em>edited this some to be clear</em>\nthe scores i listed actually happen to be the same algorithm i'm using as the fitness test mechanism's scoring tool. it's a measure of accuracy. I've actually tried lots of scoring mechanisms for fitness RMSE , RMSLE, Correlation, AUC etc  I was using correlation before now (that's actually what score 1 is using under the covers) the columns value though comes from something I created a few years ago when i needed a way to measure accuracy of things that ... well didn't really have accuracy. It uses a normal curve from the answers and figures out the odds (percentage) of a particular event and compares it with odds of the predicted value. so really it a 0 is  a perfect match and 100% is completely wrong, i just invert it. it turns out that for making features its seems better as a fitness fucntion as well. i do like correlation a lot though. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428655,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/27/2018 17:04:23",
          "content": "<blockquote>\n  <p>if you do that you will end up with 3 completely different pictures. each that have nothing to do with each other. also each will completely overfit to it's 2/3 of the data. </p>\n</blockquote>\n\n<p>Exactly, and you keep these 3 totally separate.  Then you average their validation scores, and that's it.</p>\n\n<p>It is clearly not what you are doing, and I must say I am not sure about how you do things, it is still very confusing to me.  I'm probably not smart enough to follow you, but I'll make an effort if this translates to good LB score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428666,
          "author_name": "jscheibel",
          "author_url": "",
          "post_date": "11/27/2018 17:24:15",
          "content": "<p>lol says the guy in 1st....</p>\n\n<p>it's all good. i worked on that post for like an hour. i really didn't want to say anything that could be construed as \"you're an idiot\" i know what i'm doing but i don't want to defend it to my peers right now.</p>\n\n<p>i've spent a lot ... i mean A LOT of time working on algorithms and trying to conceptualize what needs to happen, and then even more time writing them and trying things (just to see).  Ironically, this whole thread isn't even related to the main goal. i wasn't even going to run my GA stuff on this contest but i wanted an apples to apples comparison with a new expert system i'm working on. </p>\n\n<p>That one is going to process streamed data, i thought it could try and deal with the weird timeseries we have here. so I figured i'd fire up the GA get some features, run them through my data minnig tool and see what they add to my score before i changed over. The score i currently have doesn't use anything from this post. its just my own personal gradient boosting implementation using a handful of features. the ones provided and from the time series, averages, min, max and i actually bothered to implement the period finding fft (what a pain). i couldnt find a c# implementation so i converted a java one. it worked well-ish but it doesnt use the error margins like the scikit one does.</p>\n\n<p>i have a lot of fun. :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428688,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/27/2018 18:07:01",
          "content": "<p>I'm serious, and I never thought you were an idiot.  I'll reread your posts with a fresh head tomorrow ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 430591,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "11/30/2018 16:58:09",
          "content": "<p>Are you worried you aren't normalizing all the targets (Softmax for example).  Doing these individually will probably end in tears - I hope not considering the amount of time you have spent training - good luck! ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 430635,
          "author_name": "jscheibel",
          "author_url": "",
          "post_date": "11/30/2018 18:52:01",
          "content": "<p>well i'm just making features here so they don't really need to be normalized. the results i get for each group go in as new columns and are classically mined using a version of gbm. but I do normalize them. i've skipped a lot of the details. 1 important one is i do have a hold out that i use for that exact normalization. each feature is only calculated on 63.2% of the data when the run is done i use linear algebra using the hold out data to scale and offset the results appropriately. </p>\n\n<p>As for combining them up. I'm sure i could work on that some more, there may be a better way mathematically, but I'll tell you what i do as it works \"okay\". (this of course has nothing to do with anything other than this contest) if a class predicts 72% i take the 28% it's unsure of and distribute it among the other classes. except of course the other classes have their own prediction except class 99. so that would get 2%... do that for all 14 classes and take the average for class 99.  the final prediction is scaled to add up to 1 using the predictions as weights. the up shot is, it works just like a human would do it naively. if its not any of the other 14 classes it must be unknown.</p>\n\n<p>here's a simple example using 4 classes 1 is unknown</p>\n\n<p>a:15%\nb:40%\nc:1%\nd:unknown</p>\n\n<p>unknown = (0.85/3 + .6/3 + .99/3)/3.0\nunknown = (.283 + 0.2 + .33)/3.0\nunknown = 27.067%</p>\n\n<p>so final values then need to add up to 1 they currently add up to 0.83067 so divide each answer by .83067 you get</p>\n\n<p>a:18.06%\nb:48.15%\nc:1.2%\nd:32.58%</p>\n\n<p>.... an improvement on this would probably be to see if the other classes do add up to less than 100% to maybe add more in to unknown.  i don't think i've tried that yet. the thing is you want an apples to apples and the final normalization tends to take care of that. you also dont want negative values. I played around with it a bunch originally but nothing seemed to work well other than the above. I should probably see if there is more to be done... but i've been  busy playing rimworld. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 428292,
      "author_name": "marcuslin",
      "author_url": "",
      "post_date": "11/27/2018 03:17:56",
      "content": "<p>Genetic algorithm for generating features?  that's cool</p>",
      "votes": null,
      "replies": [
        {
          "id": 428340,
          "author_name": "jscheibel",
          "author_url": "",
          "post_date": "11/27/2018 05:34:17",
          "content": "<p>yeah, i used to use it to solve whole problems, but for various reasons i think its much better suited at building features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428409,
          "author_name": "marcuslin",
          "author_url": "",
          "post_date": "11/27/2018 08:26:57",
          "content": "<p>If it is not too much trouble, how do you set up the fitness function and chromosomes of GA?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428659,
          "author_name": "jscheibel",
          "author_url": "",
          "post_date": "11/27/2018 17:09:44",
          "content": "<p>I evolve my chromosomes (you know, i always call them genes in my head. rolls of the tongue easier but i digress) to get the most accurate answer i can (in isolation from the other classes). And by most accurate answer i can, i mean i'm using a normal curve based on the answer set which I then compare to the expected result of each answer vs the predicted answer on said curve. i turn those points in to percentage odds and measure the gap. i'm trying to minimize the over all gap. (see the paragraph above in cpmp's thread) is that what you are asking? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428872,
          "author_name": "marcuslin",
          "author_url": "",
          "post_date": "11/28/2018 02:20:25",
          "content": "<p>thx for reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "428073": "i haven't really been working too much on analyzing the data (i never do, that's my old story). \n\nThat being said I left my genetic algorithm working on the data over the last 4 days. And 1 class in particular (class 53 or rather class 6  of our 15) it was able to predict with 100% accuracy against the training data. I wondered if this was just a serious case of over-fitting or if that particular class is actually really easy to predict. I did it twice using two different fitness tests and one was 100% and the other was all but.\n\nbasically, i was wondering if any of you humans analysis people :) (vs my crazy machine analysis) came up with \"yeah, class 53 is really easy.\"",
    "428077": "Reading the forum can be a proxy for analyzing the data ;)\n\nClass 53 is very easy, as shown by the confusion matrix in these posts:\n\nhttps://www.kaggle.com/c/PLAsTiCC-2018/discussion/70669\n\nhttps://www.kaggle.com/c/PLAsTiCC-2018/discussion/72613",
    "428092": "cool thanks! you can understand my initial skepticism. when i saw it i was like \"100% ... sounds to good to be true. what did i do wrong.\"",
    "428100": "this is actually fascinating enough to share. it lines up pretty well with your matrix.\n\nI rank my final genetic predictions against the training data as a whole.  I get the following\n\nclass score1  score2\n6\t0.9234418 0.9569862\n15\t0.8641432 0.8884696\n16\t0.9643756 0.970057\n42\t0.7077606 0.7526829\n52\t0.8414314 0.8528698\n53\t0.9940292 1\n62\t0.8336724 0.8529119\n64\t0.9847192 -still running-\n65\t0.9588371 -still running-\n67\t0.8774132 -still running-\n88\t0.9793245 -still running-\n90\t0.6686415 -still running-\n92\t0.9801163 -still running-\n95\t0.9696037 -still running-\n\nscore1 was using a different fitness test. I made some changes and now i'm getting better results as you can see in fitness score2. (still running though)\n\nunfortunately the number i'm using isn't an apples to apples with the confusion matrix you made but it seems they are very similar.",
    "428138": "that metrics are calculated using a validation fold?",
    "428213": "i'm not sure i understand the question, or was that directed at cpmp?",
    "428292": "Genetic algorithm for generating features?  that's cool",
    "428340": "yeah, i used to use it to solve whole problems, but for various reasons i think its much better suited at building features.",
    "428409": "If it is not too much trouble, how do you set up the fitness function and chromosomes of GA?",
    "428419": "I guess you answered the question if you don't understand it ;)  The question is for you, not me ;)\n\nThe question is this: when you say you have 100% accuracy predicting class 53, how do you compute the 100%  Ideally you run your genetic algorithm on some of the data, then see what it yields on another part of the data.  This is called a train/test split validation.  A n even better way is k folds cross validation.  \n\nSO, what is your way of evaluating the quality of your GA ?",
    "428554": "j_scheibel, @CPMP  is right, the question is for you. You must evaluate the metrics in a validation or test set and not in the same train set you fitted the GA. My train set scores are pretty close to 1.0 for all classes, but its not the same in the validation set ;)",
    "428581": "&gt; but its not the same in the validation set ;)\n\nAnd it is even worse on the public LB data ;)",
    "428650": "I know what cross validation is :) .  I swear i read giba's post differently yesterday and it didn't make sense to me then. It doesn't matter though, because yes... I've been using either 3, 5 or 9 fold validation on anything i've ever done here that is conventional data mining. \n\nThat being said this is not conventional data mining.  \n\nThe problem is, each iteration will grow to fit whatever you are using. It _will_ over fit. that's what it does.  If I gave it all the data to work with  (and i have before, so I've seen this first hand) it fits to it. if i hold out some of the data for a while but then use it in evaluation later, whenever i start using it things become contaminated. so you have to keep it out forever. \n\nthe thing is if you do 3 fold cross validation. (to keep it simple) You have a need to do 3 runs . each with 2/3 of the data (since you can never cross contaminate them)... so those genes in each pile, what? never interact?  if you do that you will end up with 3 completely different pictures. each that have nothing to do with each other. also each will completely overfit to it's 2/3 of the data. \n\nyou could ensemble them but that's merging 3 over fitted things and is not what we are going for (that path leads to random forests .. with genes instead of trees. ). we want something that actually produces a good result with 1 solution.  in short cross validation as a  scoring approach doesn't work here. so what does?\n\nto answer the question originally asked by giba.  no there is no folded cross validation in those scores (there cannot be for the reason i tried to explain above  (hrumph i say! hrumph! - hehehe) )  I do have a solution to the problem and it does seem to work. (it has nothing to do with this data persay) a solution where i get 1 set of instructions that i evolved iteratively that produces a result that is fitted to the data that while completely biased (it has to be! all data mining is) is not an over fit. It takes some of explanation... I really wasn't going for any of that here, i would not want to defend the things i've done till i have a reason to  ... \n\ni just wanted to a sanity check on the 100% i got on class 53. :)  \n\n*edited this some to be clear*\nthe scores i listed actually happen to be the same algorithm i'm using as the fitness test mechanism's scoring tool. it's a measure of accuracy. I've actually tried lots of scoring mechanisms for fitness RMSE , RMSLE, Correlation, AUC etc  I was using correlation before now (that's actually what score 1 is using under the covers) the columns value though comes from something I created a few years ago when i needed a way to measure accuracy of things that ... well didn't really have accuracy. It uses a normal curve from the answers and figures out the odds (percentage) of a particular event and compares it with odds of the predicted value. so really it a 0 is  a perfect match and 100% is completely wrong, i just invert it. it turns out that for making features its seems better as a fitness fucntion as well. i do like correlation a lot though.",
    "428655": "&gt;  if you do that you will end up with 3 completely different pictures. each that have nothing to do with each other. also each will completely overfit to it's 2/3 of the data. \n\nExactly, and you keep these 3 totally separate.  Then you average their validation scores, and that's it.\n\nIt is clearly not what you are doing, and I must say I am not sure about how you do things, it is still very confusing to me.  I'm probably not smart enough to follow you, but I'll make an effort if this translates to good LB score.",
    "428659": "I evolve my chromosomes (you know, i always call them genes in my head. rolls of the tongue easier but i digress) to get the most accurate answer i can (in isolation from the other classes). And by most accurate answer i can, i mean i'm using a normal curve based on the answer set which I then compare to the expected result of each answer vs the predicted answer on said curve. i turn those points in to percentage odds and measure the gap. i'm trying to minimize the over all gap. (see the paragraph above in cpmp's thread) is that what you are asking?",
    "428666": "lol says the guy in 1st....\n\nit's all good. i worked on that post for like an hour. i really didn't want to say anything that could be construed as \"you're an idiot\" i know what i'm doing but i don't want to defend it to my peers right now.\n\n i've spent a lot ... i mean A LOT of time working on algorithms and trying to conceptualize what needs to happen, and then even more time writing them and trying things (just to see).  Ironically, this whole thread isn't even related to the main goal. i wasn't even going to run my GA stuff on this contest but i wanted an apples to apples comparison with a new expert system i'm working on. \n\nThat one is going to process streamed data, i thought it could try and deal with the weird timeseries we have here. so I figured i'd fire up the GA get some features, run them through my data minnig tool and see what they add to my score before i changed over. The score i currently have doesn't use anything from this post. its just my own personal gradient boosting implementation using a handful of features. the ones provided and from the time series, averages, min, max and i actually bothered to implement the period finding fft (what a pain). i couldnt find a c# implementation so i converted a java one. it worked well-ish but it doesnt use the error margins like the scikit one does.\n\ni have a lot of fun. :)",
    "428688": "I'm serious, and I never thought you were an idiot.  I'll reread your posts with a fresh head tomorrow ;)",
    "428872": "thx for reply.",
    "430591": "Are you worried you aren't normalizing all the targets (Softmax for example).  Doing these individually will probably end in tears - I hope not considering the amount of time you have spent training - good luck! ;)",
    "430635": "well i'm just making features here so they don't really need to be normalized. the results i get for each group go in as new columns and are classically mined using a version of gbm. but I do normalize them. i've skipped a lot of the details. 1 important one is i do have a hold out that i use for that exact normalization. each feature is only calculated on 63.2% of the data when the run is done i use linear algebra using the hold out data to scale and offset the results appropriately. \n\nAs for combining them up. I'm sure i could work on that some more, there may be a better way mathematically, but I'll tell you what i do as it works \"okay\". (this of course has nothing to do with anything other than this contest) if a class predicts 72% i take the 28% it's unsure of and distribute it among the other classes. except of course the other classes have their own prediction except class 99. so that would get 2%... do that for all 14 classes and take the average for class 99.  the final prediction is scaled to add up to 1 using the predictions as weights. the up shot is, it works just like a human would do it naively. if its not any of the other 14 classes it must be unknown.\n\nhere's a simple example using 4 classes 1 is unknown\n\na:15%\nb:40%\nc:1%\nd:unknown\n\nunknown = (0.85/3 + .6/3 + .99/3)/3.0\nunknown = (.283 + 0.2 + .33)/3.0\nunknown = 27.067%\n\nso final values then need to add up to 1 they currently add up to 0.83067 so divide each answer by .83067 you get\n\na:18.06%\nb:48.15%\nc:1.2%\nd:32.58%\n\n.... an improvement on this would probably be to see if the other classes do add up to less than 100% to maybe add more in to unknown.  i don't think i've tried that yet. the thing is you want an apples to apples and the final normalization tends to take care of that. you also dont want negative values. I played around with it a bunch originally but nothing seemed to work well other than the above. I should probably see if there is more to be done... but i've been  busy playing rimworld. :)"
  },
  "source": "meta"
}