{
  "id": 68845,
  "title": "trouble with Cross Validation",
  "url": "/competitions/PLAsTiCC-2018/discussion/68845",
  "author_name": "",
  "post_date": "2018-10-17T20:20:21.757494700Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>It feels weird even asking for help with cross validation but at this point I'm really just scratching my head. I do everything in C# so I have to write my own ... well everything. here's the code i'm using for my evaluation</p>\n\n<pre><code>//weighted multi-class logarithmic loss aka wmcll\n            groupResultAvg = 0;\n\n            for (int group = 0; group &lt; groups.Length; group++) {\n                int rowsCount = 0;\n                for (count = 0; count &lt; realValues[group].Length; count++) {\n                    if (realValues[group][count] &gt; 0) {\n                        rowsCount = rowsCount+1;\n                    }\n                }\n\n                double work = 0;\n                for (count = 0; count &lt; realValues[group].Length; count++) {\n\n                    if (realValues[group][count] &gt; 0) {\n                        //predictions are already % at this point\n                        work = work + Math.Log(predictions[group][count]) / rowsCount;\n                    }\n\n                }\n\n                groupResultAvg = groupResultAvg + work/ groups.Length;\n            }\n\n            groupResultAvg = -groupResultAvg; //note even weights on each class.. \n</code></pre>\n\n<p>And while it seems to work. I get values that seem \"reasonable\" for instance if i put a hard coded value of 0.0666 in place of  predictions[group][count] i get 2.52844 against the training set.</p>\n\n<p>the problem i'm having is i get good (say 1.5 or 1.3 or something )results locally but the LB always gives me a 2.3+ score. I've checked that my result columns are ordered right (something that could easily make a mess of things) and i've got non-zero values in the 99 class. I even tried forcing that classes percentage to around 0.0666 specifically just to make sure it wasnt my issue. but it changed basically nothing. </p>\n\n<p>Is this just a case of over-fitting? I'm using a version of GBM I wrote, but its not new or anything I've used it for many contests and it worked fine there. I even set it up with a small gradient (at least for the last few tests) and shallow tree depths (as of late have only been like 3) not to mention i'm only layering about 10 trees... so it seems like i should be getting close to if not exactly the same result locally. </p>\n\n<p>Am I missing something?</p>",
  "messages": [
    {
      "id": "405616",
      "postDate": "10/17/2018 20:20:21",
      "content": "<p>It feels weird even asking for help with cross validation but at this point I'm really just scratching my head. I do everything in C# so I have to write my own ... well everything. here's the code i'm using for my evaluation</p>\n\n<pre><code>//weighted multi-class logarithmic loss aka wmcll\n            groupResultAvg = 0;\n\n            for (int group = 0; group &lt; groups.Length; group++) {\n                int rowsCount = 0;\n                for (count = 0; count &lt; realValues[group].Length; count++) {\n                    if (realValues[group][count] &gt; 0) {\n                        rowsCount = rowsCount+1;\n                    }\n                }\n\n                double work = 0;\n                for (count = 0; count &lt; realValues[group].Length; count++) {\n\n                    if (realValues[group][count] &gt; 0) {\n                        //predictions are already % at this point\n                        work = work + Math.Log(predictions[group][count]) / rowsCount;\n                    }\n\n                }\n\n                groupResultAvg = groupResultAvg + work/ groups.Length;\n            }\n\n            groupResultAvg = -groupResultAvg; //note even weights on each class.. \n</code></pre>\n\n<p>And while it seems to work. I get values that seem \"reasonable\" for instance if i put a hard coded value of 0.0666 in place of  predictions[group][count] i get 2.52844 against the training set.</p>\n\n<p>the problem i'm having is i get good (say 1.5 or 1.3 or something )results locally but the LB always gives me a 2.3+ score. I've checked that my result columns are ordered right (something that could easily make a mess of things) and i've got non-zero values in the 99 class. I even tried forcing that classes percentage to around 0.0666 specifically just to make sure it wasnt my issue. but it changed basically nothing. </p>\n\n<p>Is this just a case of over-fitting? I'm using a version of GBM I wrote, but its not new or anything I've used it for many contests and it worked fine there. I even set it up with a small gradient (at least for the last few tests) and shallow tree depths (as of late have only been like 3) not to mention i'm only layering about 10 trees... so it seems like i should be getting close to if not exactly the same result locally. </p>\n\n<p>Am I missing something?</p>",
      "rawMarkdown": "It feels weird even asking for help with cross validation but at this point I'm really just scratching my head. I do everything in C# so I have to write my own ... well everything. here's the code i'm using for my evaluation\n\n    //weighted multi-class logarithmic loss aka wmcll\n    \t\t\tgroupResultAvg = 0;\n    \n    \t\t\tfor (int group = 0; group &lt; groups.Length; group++) {\n    \t\t\t\tint rowsCount = 0;\n    \t\t\t\tfor (count = 0; count &lt; realValues[group].Length; count++) {\n    \t\t\t\t\tif (realValues[group][count] &gt; 0) {\n    \t\t\t\t\t\trowsCount = rowsCount+1;\n    \t\t\t\t\t}\n    \t\t\t\t}\n    \n    \t\t\t\tdouble work = 0;\n    \t\t\t\tfor (count = 0; count &lt; realValues[group].Length; count++) {\n    \n    \t\t\t\t\tif (realValues[group][count] &gt; 0) {\n    \t\t\t\t\t\t//predictions are already % at this point\n    \t\t\t\t\t\twork = work + Math.Log(predictions[group][count]) / rowsCount;\n    \t\t\t\t\t}\n    \n    \t\t\t\t}\n    \n    \t\t\t\tgroupResultAvg = groupResultAvg + work/ groups.Length;\n    \t\t\t}\n    \n    \t\t\tgroupResultAvg = -groupResultAvg; //note even weights on each class.. \n\nAnd while it seems to work. I get values that seem \"reasonable\" for instance if i put a hard coded value of 0.0666 in place of  predictions[group][count] i get 2.52844 against the training set.\n\nthe problem i'm having is i get good (say 1.5 or 1.3 or something )results locally but the LB always gives me a 2.3+ score. I've checked that my result columns are ordered right (something that could easily make a mess of things) and i've got non-zero values in the 99 class. I even tried forcing that classes percentage to around 0.0666 specifically just to make sure it wasnt my issue. but it changed basically nothing. \n\nIs this just a case of over-fitting? I'm using a version of GBM I wrote, but its not new or anything I've used it for many contests and it worked fine there. I even set it up with a small gradient (at least for the last few tests) and shallow tree depths (as of late have only been like 3) not to mention i'm only layering about 10 trees... so it seems like i should be getting close to if not exactly the same result locally. \n\nAm I missing something?",
      "votes": null
    },
    {
      "id": "405631",
      "postDate": "10/17/2018 20:36:54",
      "content": "<p>when the daily limit rolls over for me. i'm going to put a little more aggressive floor on the values I'm submitting. maybe not let any percentage be lower than 1% maybe that'll help. seems unlikely but at this its the only thing I can figure.</p>\n\n<p><em>edit</em> I submitted what CVed at 1.76674 only 3 fold i'll move it back to 9 i was testing larger folds to see if they made any difference. that includes the 1% lower end cap. the result was 2.16 ... pretty terrible by comparison.</p>",
      "rawMarkdown": "when the daily limit rolls over for me. i'm going to put a little more aggressive floor on the values I'm submitting. maybe not let any percentage be lower than 1% maybe that'll help. seems unlikely but at this its the only thing I can figure.\n\n*edit* I submitted what CVed at 1.76674 only 3 fold i'll move it back to 9 i was testing larger folds to see if they made any difference. that includes the 1% lower end cap. the result was 2.16 ... pretty terrible by comparison.",
      "votes": null
    },
    {
      "id": "405752",
      "postDate": "10/18/2018 03:49:31",
      "content": "<p>I think your CV is okay and the problem is the predictions for class_99.</p>\n\n<p>My CV is around 1.30 and my score is around 1.7x, but if I change the predictions for class_99 my score have some changes.</p>\n\n<p>Maybe the great question is this competition is how to make predictions for class_99.</p>",
      "rawMarkdown": "I think your CV is okay and the problem is the predictions for class_99.\n\nMy CV is around 1.30 and my score is around 1.7x, but if I change the predictions for class_99 my score have some changes.\n\nMaybe the great question is this competition is how to make predictions for class_99.",
      "votes": null
    },
    {
      "id": "405770",
      "postDate": "10/18/2018 04:51:31",
      "content": "<p>i mean, if you think i did it right... then maybe that class 99 is the problem like you say. I mean it must be. Either that or the statistics i'm building to do my initial runs are garbage (nothing fancy yet, just sums and averages over the data).... or the test data is just nothing like the train data in other ways. like deliberately over representing fringe cases so that train averages don't work for predictions. </p>\n\n<p>i tried forcibly giving the 99 feature 1/15 th (6.67%) chance across the board leaving the other 93% of predictions and it didnt seem to impact results significantly =/ . So I dunno, seems fishy that its just that, but maybe.</p>\n\n<p>I've tweaked my results further by moving the floor to .5% instead of a whole percent, I also fiddled a little with the algorithm (general improvement unrelated to anything here) and those dropped it a little further but i'm still 2+ vs the 1.72 i'm seeing locally.</p>\n\n<p>As for actually calculating the 99 class. i'm predicting each class by itself as a binary. so if i get back i dunno 40% chance for a particular class the other 60% gets divided on all the other classes. i need to double check my math to make sure i got the weights set right, but for now it seems decent. In general it seems naively like the best way to find class 99. in essence if everything comes back with a 2% chance of being it, then the class 99 will automatically become inflated (to 72%). solution by contradiction i suppose.</p>",
      "rawMarkdown": "i mean, if you think i did it right... then maybe that class 99 is the problem like you say. I mean it must be. Either that or the statistics i'm building to do my initial runs are garbage (nothing fancy yet, just sums and averages over the data).... or the test data is just nothing like the train data in other ways. like deliberately over representing fringe cases so that train averages don't work for predictions. \n\ni tried forcibly giving the 99 feature 1/15 th (6.67%) chance across the board leaving the other 93% of predictions and it didnt seem to impact results significantly =/ . So I dunno, seems fishy that its just that, but maybe.\n\nI've tweaked my results further by moving the floor to .5% instead of a whole percent, I also fiddled a little with the algorithm (general improvement unrelated to anything here) and those dropped it a little further but i'm still 2+ vs the 1.72 i'm seeing locally.\n\nAs for actually calculating the 99 class. i'm predicting each class by itself as a binary. so if i get back i dunno 40% chance for a particular class the other 60% gets divided on all the other classes. i need to double check my math to make sure i got the weights set right, but for now it seems decent. In general it seems naively like the best way to find class 99. in essence if everything comes back with a 2% chance of being it, then the class 99 will automatically become inflated (to 72%). solution by contradiction i suppose.",
      "votes": null
    },
    {
      "id": "405797",
      "postDate": "10/18/2018 06:01:49",
      "content": "<p>Perhaps? Classes 99, 64 &amp; 15 have double the weight?</p>\n\n<p><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194#397153\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194#397153</a></p>",
      "rawMarkdown": "Perhaps? Classes 99, 64 &amp; 15 have double the weight?\n\nhttps://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194#397153",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 405631,
      "author_name": "jscheibel",
      "author_url": "",
      "post_date": "10/17/2018 20:36:54",
      "content": "<p>when the daily limit rolls over for me. i'm going to put a little more aggressive floor on the values I'm submitting. maybe not let any percentage be lower than 1% maybe that'll help. seems unlikely but at this its the only thing I can figure.</p>\n\n<p><em>edit</em> I submitted what CVed at 1.76674 only 3 fold i'll move it back to 9 i was testing larger folds to see if they made any difference. that includes the 1% lower end cap. the result was 2.16 ... pretty terrible by comparison.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 405752,
      "author_name": "joaopmpeinado",
      "author_url": "",
      "post_date": "10/18/2018 03:49:31",
      "content": "<p>I think your CV is okay and the problem is the predictions for class_99.</p>\n\n<p>My CV is around 1.30 and my score is around 1.7x, but if I change the predictions for class_99 my score have some changes.</p>\n\n<p>Maybe the great question is this competition is how to make predictions for class_99.</p>",
      "votes": null,
      "replies": [
        {
          "id": 405770,
          "author_name": "jscheibel",
          "author_url": "",
          "post_date": "10/18/2018 04:51:31",
          "content": "<p>i mean, if you think i did it right... then maybe that class 99 is the problem like you say. I mean it must be. Either that or the statistics i'm building to do my initial runs are garbage (nothing fancy yet, just sums and averages over the data).... or the test data is just nothing like the train data in other ways. like deliberately over representing fringe cases so that train averages don't work for predictions. </p>\n\n<p>i tried forcibly giving the 99 feature 1/15 th (6.67%) chance across the board leaving the other 93% of predictions and it didnt seem to impact results significantly =/ . So I dunno, seems fishy that its just that, but maybe.</p>\n\n<p>I've tweaked my results further by moving the floor to .5% instead of a whole percent, I also fiddled a little with the algorithm (general improvement unrelated to anything here) and those dropped it a little further but i'm still 2+ vs the 1.72 i'm seeing locally.</p>\n\n<p>As for actually calculating the 99 class. i'm predicting each class by itself as a binary. so if i get back i dunno 40% chance for a particular class the other 60% gets divided on all the other classes. i need to double check my math to make sure i got the weights set right, but for now it seems decent. In general it seems naively like the best way to find class 99. in essence if everything comes back with a 2% chance of being it, then the class 99 will automatically become inflated (to 72%). solution by contradiction i suppose.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 405797,
      "author_name": "glimmung",
      "author_url": "",
      "post_date": "10/18/2018 06:01:49",
      "content": "<p>Perhaps? Classes 99, 64 &amp; 15 have double the weight?</p>\n\n<p><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194#397153\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194#397153</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "405616": "It feels weird even asking for help with cross validation but at this point I'm really just scratching my head. I do everything in C# so I have to write my own ... well everything. here's the code i'm using for my evaluation\n\n    //weighted multi-class logarithmic loss aka wmcll\n    \t\t\tgroupResultAvg = 0;\n    \n    \t\t\tfor (int group = 0; group &lt; groups.Length; group++) {\n    \t\t\t\tint rowsCount = 0;\n    \t\t\t\tfor (count = 0; count &lt; realValues[group].Length; count++) {\n    \t\t\t\t\tif (realValues[group][count] &gt; 0) {\n    \t\t\t\t\t\trowsCount = rowsCount+1;\n    \t\t\t\t\t}\n    \t\t\t\t}\n    \n    \t\t\t\tdouble work = 0;\n    \t\t\t\tfor (count = 0; count &lt; realValues[group].Length; count++) {\n    \n    \t\t\t\t\tif (realValues[group][count] &gt; 0) {\n    \t\t\t\t\t\t//predictions are already % at this point\n    \t\t\t\t\t\twork = work + Math.Log(predictions[group][count]) / rowsCount;\n    \t\t\t\t\t}\n    \n    \t\t\t\t}\n    \n    \t\t\t\tgroupResultAvg = groupResultAvg + work/ groups.Length;\n    \t\t\t}\n    \n    \t\t\tgroupResultAvg = -groupResultAvg; //note even weights on each class.. \n\nAnd while it seems to work. I get values that seem \"reasonable\" for instance if i put a hard coded value of 0.0666 in place of  predictions[group][count] i get 2.52844 against the training set.\n\nthe problem i'm having is i get good (say 1.5 or 1.3 or something )results locally but the LB always gives me a 2.3+ score. I've checked that my result columns are ordered right (something that could easily make a mess of things) and i've got non-zero values in the 99 class. I even tried forcing that classes percentage to around 0.0666 specifically just to make sure it wasnt my issue. but it changed basically nothing. \n\nIs this just a case of over-fitting? I'm using a version of GBM I wrote, but its not new or anything I've used it for many contests and it worked fine there. I even set it up with a small gradient (at least for the last few tests) and shallow tree depths (as of late have only been like 3) not to mention i'm only layering about 10 trees... so it seems like i should be getting close to if not exactly the same result locally. \n\nAm I missing something?",
    "405631": "when the daily limit rolls over for me. i'm going to put a little more aggressive floor on the values I'm submitting. maybe not let any percentage be lower than 1% maybe that'll help. seems unlikely but at this its the only thing I can figure.\n\n*edit* I submitted what CVed at 1.76674 only 3 fold i'll move it back to 9 i was testing larger folds to see if they made any difference. that includes the 1% lower end cap. the result was 2.16 ... pretty terrible by comparison.",
    "405752": "I think your CV is okay and the problem is the predictions for class_99.\n\nMy CV is around 1.30 and my score is around 1.7x, but if I change the predictions for class_99 my score have some changes.\n\nMaybe the great question is this competition is how to make predictions for class_99.",
    "405770": "i mean, if you think i did it right... then maybe that class 99 is the problem like you say. I mean it must be. Either that or the statistics i'm building to do my initial runs are garbage (nothing fancy yet, just sums and averages over the data).... or the test data is just nothing like the train data in other ways. like deliberately over representing fringe cases so that train averages don't work for predictions. \n\ni tried forcibly giving the 99 feature 1/15 th (6.67%) chance across the board leaving the other 93% of predictions and it didnt seem to impact results significantly =/ . So I dunno, seems fishy that its just that, but maybe.\n\nI've tweaked my results further by moving the floor to .5% instead of a whole percent, I also fiddled a little with the algorithm (general improvement unrelated to anything here) and those dropped it a little further but i'm still 2+ vs the 1.72 i'm seeing locally.\n\nAs for actually calculating the 99 class. i'm predicting each class by itself as a binary. so if i get back i dunno 40% chance for a particular class the other 60% gets divided on all the other classes. i need to double check my math to make sure i got the weights set right, but for now it seems decent. In general it seems naively like the best way to find class 99. in essence if everything comes back with a 2% chance of being it, then the class 99 will automatically become inflated (to 72%). solution by contradiction i suppose.",
    "405797": "Perhaps? Classes 99, 64 &amp; 15 have double the weight?\n\nhttps://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194#397153"
  },
  "source": "meta"
}