{
  "id": 17009,
  "title": "Congratulations to winners!",
  "url": "/competitions/dato-native/discussion/17009",
  "author_name": "",
  "post_date": "2015-10-15T00:07:14.923Z",
  "votes": 6,
  "comment_count": 56,
  "views": 8600,
  "content": "<p>Congratulations, Mad professors :P</p>",
  "messages": [
    {
      "id": "96171",
      "postDate": "10/15/2015 00:07:14",
      "content": "<p>Congratulations, Mad professors :P</p>",
      "rawMarkdown": "Congratulations, Mad professors :P",
      "votes": null
    },
    {
      "id": "96172",
      "postDate": "10/15/2015 00:10:49",
      "content": "<p>I also want to thank my teammates. We have made it to our limit : )</p>",
      "rawMarkdown": "I also want to thank my teammates. We have made it to our limit : )",
      "votes": null
    },
    {
      "id": "96173",
      "postDate": "10/15/2015 00:11:47",
      "content": "<p>Ok... that was unexpected XD. Great work guys.</p>",
      "rawMarkdown": "Ok... that was unexpected XD. Great work guys.",
      "votes": null
    },
    {
      "id": "96180",
      "postDate": "10/15/2015 00:38:05",
      "content": "<p>haha, thx. We are not that mad about it anymore!</p>\n\n<p>This was a tough competition (given the reset and the big size of the data) , but a very interesting problem. </p>\n\n<p>Well done to my teammates (Faron and Triskelion alphabetically) for their hard work. </p>\n\n<p>Also big credit to mortehu (that also soloed this) and bibaze for their great finishes. </p>\n\n<p>P.S. mortehu , I hope you don't hate us! Luck likes to play funny games, it happened to favor us this time, I hope next time to favor you :). Again well done on getting your master badge :)</p>\n\n<p>Big thank you to the organizers for their hard work and the Dato-graphlab guys for sponsoring this. </p>",
      "rawMarkdown": "haha, thx. We are not that mad about it anymore!\r\n\r\nThis was a tough competition (given the reset and the big size of the data) , but a very interesting problem. \r\n\r\nWell done to my teammates (Faron and Triskelion alphabetically) for their hard work. \r\n\r\nAlso big credit to mortehu (that also soloed this) and bibaze for their great finishes. \r\n\r\n\r\nP.S. mortehu , I hope you don't hate us! Luck likes to play funny games, it happened to favor us this time, I hope next time to favor you :). Again well done on getting your master badge :)\r\n\r\nBig thank you to the organizers for their hard work and the Dato-graphlab guys for sponsoring this.",
      "votes": null
    },
    {
      "id": "96181",
      "postDate": "10/15/2015 00:42:55",
      "content": "<p>@KazAnova, congratulations!</p>\n\n<p>Also please confirm that you will get Date Create Prize so we don't need to think about it :P</p>",
      "rawMarkdown": "KazAnova, congratulations!\r\n\r\nAlso please confirm that you will get Date Create Prize so we don't need to think about it :P",
      "votes": null
    },
    {
      "id": "96182",
      "postDate": "10/15/2015 00:44:55",
      "content": "<p>Please share your golden features :)</p>",
      "rawMarkdown": "Please share your golden features :)",
      "votes": null
    },
    {
      "id": "96183",
      "postDate": "10/15/2015 00:46:02",
      "content": "<p>@KazAnova or anyone else who got 0.98 with Random Forest, how did you do that? We could barely make 0.93 with multiple attempts.</p>",
      "rawMarkdown": "KazAnova or anyone else who got 0.98 with Random Forest, how did you do that? We could barely make 0.93 with multiple attempts.",
      "votes": null
    },
    {
      "id": "96184",
      "postDate": "10/15/2015 00:47:53",
      "content": "<p>Our best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.</p>\n\n<p>There was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.</p>\n\n<p>We also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.</p>\n\n<p>some models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . </p>\n\n<p>I think we reached Level 4 meta-modelling in the end :)</p>",
      "rawMarkdown": "Our best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.\r\n\r\nThere was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.\r\n\r\nWe also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.\r\n\r\nsome models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . \r\n\r\nI think we reached Level 4 meta-modelling in the end :)",
      "votes": null
    },
    {
      "id": "96185",
      "postDate": "10/15/2015 00:54:35",
      "content": "<p>Well, my hats off to David Shinn. He entered very late in the competition, eliminated the leaderboard very quickly and finished 10th, just dethroning us from our precious position. So I truly natively hate him, but also greatly, absolutely respect him for doing this on his own. Hats off to you sir!</p>\n\n<p>Other shouts go out to:\n- Andrew Bell: we competed a while back head to head, before you hit turbo mode and went out of my sight. Great effort.\n- bibaze team: excellent progress makers.\n- mortehu, orchid, student_2012, Glen: are you guys aliens or humans?</p>\n\n<p>and a thumbs up for NxGTR. My best model only got to 0.9803, so not even close to his 0.984. What he lacked were ensembling techniques and variety in models or data.</p>",
      "rawMarkdown": "Well, my hats off to David Shinn. He entered very late in the competition, eliminated the leaderboard very quickly and finished 10th, just dethroning us from our precious position. So I truly natively hate him, but also greatly, absolutely respect him for doing this on his own. Hats off to you sir!\r\n\r\nOther shouts go out to:\r\n- Andrew Bell: we competed a while back head to head, before you hit turbo mode and went out of my sight. Great effort.\r\n- bibaze team: excellent progress makers.\r\n- mortehu, orchid, student_2012, Glen: are you guys aliens or humans?\r\n\r\nand a thumbs up for NxGTR. My best model only got to 0.9803, so not even close to his 0.984. What he lacked were ensembling techniques and variety in models or data.",
      "votes": null
    },
    {
      "id": "96186",
      "postDate": "10/15/2015 00:55:09",
      "content": "<p>For us, there were no golden features. Our method was an ensemble of a stacked xgboost, keras, simple xgboost and ftrl. We used binary counts as the features for the last 2 models.</p>",
      "rawMarkdown": "For us, there were no golden features. Our method was an ensemble of a stacked xgboost, keras, simple xgboost and ftrl. We used binary counts as the features for the last 2 models.",
      "votes": null
    },
    {
      "id": "96187",
      "postDate": "10/15/2015 00:57:30",
      "content": "<p>I attempted FTRL in this competition and got to 0.960 auc. Is that about what you achieved?</p>",
      "rawMarkdown": "I attempted FTRL in this competition and got to 0.960 auc. Is that about what you achieved?",
      "votes": null
    },
    {
      "id": "96189",
      "postDate": "10/15/2015 00:58:54",
      "content": "<p>We achieved 0.979 with FTRL, trained it for 5 epochs.</p>",
      "rawMarkdown": "We achieved 0.979 with FTRL, trained it for 5 epochs.",
      "votes": null
    },
    {
      "id": "96190",
      "postDate": "10/15/2015 01:09:15",
      "content": "<p>[quote=Subhajit Mandal;96183]</p>\n\n<p>@KazAnova or anyone else who got 0.98 with Random Forest, how did you do that? We could barely make 0.93 with multiple attempts.</p>\n\n<p>[/quote]\nAdd these count features to the benchmark code, you will get to at least 0.95 with Random Forest</p>\n\n<pre><code>    row['spaces'] = text.count(' ')\n    row['tabs'] = text.count('\\t')\n    row['braces'] = text.count('{')\n    row['brackets'] = text.count('[')\n    row['words'] = len(re.split('\\s+', text))\n    row['length'] = len(text)\n    row['bracket_2'] = text.count('&lt;')\n    row['bracket_3'] = text.count('(')\n    row['re_w'] = len(re.split('\\w+', text))\n    row['new_par'] = text.count('\\n\\t')\n    row['dbslash'] = text.count('//')\n    row['bslash'] = text.count('/')\n    row['bol_0'] = text.count('&lt;!')\n    row['bol_1'] = text.count('/&gt;')\n    row['bol_2'] = text.count('&amp;')\n    row['bol_3'] = text.count(';')\n    row['bol_4'] = text.count('==')\n    row['bol_5'] = text.count('===')\n    row['bol_6'] = text.count('css')\n    row['bol_7'] = text.count('#')\n    row['bol_8'] = text.count('@')\n    row['bol_9'] = text.count('$')\n    row['bol_10'] = text.count('%')\n    row['bol_11'] = text.count('^')\n    row['bol_12'] = text.count('+')\n    row['bol_13'] = text.count('?')\n    row['bol_14'] = text.count('|')\n    row['bol_15'] = text.count('\\\\')\n    row['bol_16'] = text.count('*')\n    row['bol_17'] = text.count('||')\n    row['bol_18'] = text.count('\\t\\t')\n    row['bol_19'] = text.count('\\t\\t\\t')\n</code></pre>",
      "rawMarkdown": "[quote=Subhajit Mandal;96183]\r\n\r\n@KazAnova or anyone else who got 0.98 with Random Forest, how did you do that? We could barely make 0.93 with multiple attempts.\r\n\r\n[/quote]\r\nAdd these count features to the benchmark code, you will get to at least 0.95 with Random Forest\r\n\r\n        row['spaces'] = text.count(' ')\r\n        row['tabs'] = text.count('\\t')\r\n        row['braces'] = text.count('{')\r\n        row['brackets'] = text.count('[')\r\n        row['words'] = len(re.split('\\s+', text))\r\n        row['length'] = len(text)\r\n        row['bracket_2'] = text.count('<')\r\n        row['bracket_3'] = text.count('(')\r\n        row['re_w'] = len(re.split('\\w+', text))\r\n        row['new_par'] = text.count('\\n\\t')\r\n        row['dbslash'] = text.count('//')\r\n        row['bslash'] = text.count('/')\r\n        row['bol_0'] = text.count('<!')\r\n        row['bol_1'] = text.count('/>')\r\n        row['bol_2'] = text.count('&')\r\n        row['bol_3'] = text.count(';')\r\n        row['bol_4'] = text.count('==')\r\n        row['bol_5'] = text.count('===')\r\n        row['bol_6'] = text.count('css')\r\n        row['bol_7'] = text.count('#')\r\n        row['bol_8'] = text.count('@')\r\n        row['bol_9'] = text.count('$')\r\n        row['bol_10'] = text.count('%')\r\n        row['bol_11'] = text.count('^')\r\n        row['bol_12'] = text.count('+')\r\n        row['bol_13'] = text.count('?')\r\n        row['bol_14'] = text.count('|')\r\n        row['bol_15'] = text.count('\\\\')\r\n        row['bol_16'] = text.count('*')\r\n        row['bol_17'] = text.count('||')\r\n        row['bol_18'] = text.count('\\t\\t')\r\n        row['bol_19'] = text.count('\\t\\t\\t')",
      "votes": null
    },
    {
      "id": "96191",
      "postDate": "10/15/2015 01:11:21",
      "content": "<p>Horrifying. What were the features? I trained FTRL on 4-5 epochs on 1800 raw features, reduced from 960k.  My features were the most prevalent 1800 word features with optimized L1/L2 regularization.</p>\n\n<p>I went fruity-loopy in this competition and made a custom C parser, a bigram/trigram processor in python and in the end I was processing pretty much portions of matrices to get things going (scipy.sparse.hstack). All of this to reduce the memory requirements for processing the data and decreasing the time before training and model verification would occur.</p>",
      "rawMarkdown": "Horrifying. What were the features? I trained FTRL on 4-5 epochs on 1800 raw features, reduced from 960k.  My features were the most prevalent 1800 word features with optimized L1/L2 regularization.\r\n\r\nI went fruity-loopy in this competition and made a custom C parser, a bigram/trigram processor in python and in the end I was processing pretty much portions of matrices to get things going (scipy.sparse.hstack). All of this to reduce the memory requirements for processing the data and decreasing the time before training and model verification would occur.",
      "votes": null
    },
    {
      "id": "96192",
      "postDate": "10/15/2015 01:11:25",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;96184]</p>\n\n<p>Our best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.</p>\n\n<p>There was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.</p>\n\n<p>We also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.</p>\n\n<p>some models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . </p>\n\n<p>I think we reached Level 4 meta-modelling in the end :)\n[/quote]\nHow much RAM it requires? </p>",
      "rawMarkdown": "[quote=Μαριος Μιχαηλιδης KazAnova;96184]\r\n\r\nOur best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.\r\n\r\nThere was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.\r\n\r\nWe also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.\r\n\r\nsome models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . \r\n\r\nI think we reached Level 4 meta-modelling in the end :)\r\n[/quote]\r\nHow much RAM it requires?",
      "votes": null
    },
    {
      "id": "96193",
      "postDate": "10/15/2015 01:13:12",
      "content": "<p>For us it was like this:</p>\n\n<p>Xgboost (single): 0.98422</p>\n\n<p>Xgboost (dual ensemble): 0.98459</p>\n\n<p>RandomForest (single): 0.97455</p>\n\n<p>ExtraTrees (single): 0.97628</p>\n\n<p>Done :D, our final submission is an ensemble of pretty much those 4. </p>",
      "rawMarkdown": "For us it was like this:\r\n\r\nXgboost (single): 0.98422\r\n\r\nXgboost (dual ensemble): 0.98459\r\n\r\nRandomForest (single): 0.97455\r\n\r\nExtraTrees (single): 0.97628\r\n\r\nDone :D, our final submission is an ensemble of pretty much those 4.",
      "votes": null
    },
    {
      "id": "96195",
      "postDate": "10/15/2015 01:17:34",
      "content": "<p>[quote=Artem;96192]</p>\n\n<p>How much RAM it requires? </p>\n\n<p>[/quote]</p>\n\n<p>A lot!</p>\n\n<p>128GB should be ok!</p>\n\n<p>But to be honest, we were reckless (as we had 256GB available). I cannot tell you what is the most optimum you can run it on. 64GB is possibly enough too.</p>",
      "rawMarkdown": "[quote=Artem;96192]\r\n\r\nHow much RAM it requires? \r\n\r\n[/quote]\r\n\r\nA lot!\r\n\r\n128GB should be ok!\r\n\r\nBut to be honest, we were reckless (as we had 256GB available). I cannot tell you what is the most optimum you can run it on. 64GB is possibly enough too.",
      "votes": null
    },
    {
      "id": "96196",
      "postDate": "10/15/2015 01:18:07",
      "content": "<p>[quote=Gerard Toonstra;96191]</p>\n\n<p>Horrifying. What were the features? I trained FTRL on 4-5 epochs on 1800 raw features, reduced from 960k.  My features were the most prevalent 1800 word features with optimized L1/L2 regularization.</p>\n\n<p>I went fruity-loopy in this competition and made a custom C parser, a bigram/trigram processor in python and in the end I was processing pretty much portions of matrices to get things going (scipy.sparse.hstack). All of this to reduce the memory requirements for processing the data and decreasing the time before training and model verification would occur.</p>\n\n<p>[/quote]</p>\n\n<p>For FTRL, we used simple binary counts (is it there or not?) of unigrams and bigrams as features. We tried many more feature engineering techniques, but the simple was the best.</p>",
      "rawMarkdown": "[quote=Gerard Toonstra;96191]\r\n\r\nHorrifying. What were the features? I trained FTRL on 4-5 epochs on 1800 raw features, reduced from 960k.  My features were the most prevalent 1800 word features with optimized L1/L2 regularization.\r\n\r\nI went fruity-loopy in this competition and made a custom C parser, a bigram/trigram processor in python and in the end I was processing pretty much portions of matrices to get things going (scipy.sparse.hstack). All of this to reduce the memory requirements for processing the data and decreasing the time before training and model verification would occur.\r\n\r\n\r\n[/quote]\r\n\r\nFor FTRL, we used simple binary counts (is it there or not?) of unigrams and bigrams as features. We tried many more feature engineering techniques, but the simple was the best.",
      "votes": null
    },
    {
      "id": "96197",
      "postDate": "10/15/2015 01:18:39",
      "content": "<p>[quote=Subhajit Mandal;96189]</p>\n\n<p>We achieved 0.979 with FTRL, trained it for 5 epochs.</p>\n\n<p>[/quote]</p>\n\n<p>Along with FTRL, we had following scores with other models:</p>\n\n<p>xgboost: 0.983</p>\n\n<p>stacked xgboost: 0.985</p>\n\n<p>MLP with 3 layerS: 0.952</p>\n\n<p>I agree with @rcarson and @NxGTR, bigger teams and heavy ensembing is the way to do well in these competitions.</p>",
      "rawMarkdown": "[quote=Subhajit Mandal;96189]\r\n\r\nWe achieved 0.979 with FTRL, trained it for 5 epochs.\r\n\r\n[/quote]\r\n\r\nAlong with FTRL, we had following scores with other models:\r\n\r\nxgboost: 0.983\r\n\r\nstacked xgboost: 0.985\r\n\r\nMLP with 3 layerS: 0.952\r\n\r\nI agree with @rcarson and @NxGTR, bigger teams and heavy ensembing is the way to do well in these competitions.",
      "votes": null
    },
    {
      "id": "96198",
      "postDate": "10/15/2015 01:21:46",
      "content": "<p>[quote=Little Boat;96190]</p>\n\n<p>[quote=Subhajit Mandal;96183]</p>\n\n<p>@KazAnova or anyone else who got 0.98 with Random Forest, how did you do that? We could barely make 0.93 with multiple attempts.</p>\n\n<p>[/quote]\nAdd these count features to the benchmark code, you will get to at least 0.95 with Random Forest</p>\n\n<pre><code>    row['spaces'] = text.count(' ')\n    row['tabs'] = text.count('\\t')\n    row['braces'] = text.count('{')\n    row['brackets'] = text.count('[')\n    row['words'] = len(re.split('\\s+', text))\n    row['length'] = len(text)\n    row['bracket_2'] = text.count('&lt;')\n    row['bracket_3'] = text.count('(')\n    row['re_w'] = len(re.split('\\w+', text))\n    row['new_par'] = text.count('\\n\\t')\n    row['dbslash'] = text.count('//')\n    row['bslash'] = text.count('/')\n    row['bol_0'] = text.count('&lt;!')\n    row['bol_1'] = text.count('/&gt;')\n    row['bol_2'] = text.count('&amp;')\n    row['bol_3'] = text.count(';')\n    row['bol_4'] = text.count('==')\n    row['bol_5'] = text.count('===')\n    row['bol_6'] = text.count('css')\n    row['bol_7'] = text.count('#')\n    row['bol_8'] = text.count('@')\n    row['bol_9'] = text.count('$')\n    row['bol_10'] = text.count('%')\n    row['bol_11'] = text.count('^')\n    row['bol_12'] = text.count('+')\n    row['bol_13'] = text.count('?')\n    row['bol_14'] = text.count('|')\n    row['bol_15'] = text.count('\\\\')\n    row['bol_16'] = text.count('*')\n    row['bol_17'] = text.count('||')\n    row['bol_18'] = text.count('\\t\\t')\n    row['bol_19'] = text.count('\\t\\t\\t')\n</code></pre>\n\n<p>[/quote]</p>\n\n<p>I did not see that coming ;-) should not have removed the punctuation.</p>",
      "rawMarkdown": "[quote=Little Boat;96190]\r\n\r\n[quote=Subhajit Mandal;96183]\r\n\r\n@KazAnova or anyone else who got 0.98 with Random Forest, how did you do that? We could barely make 0.93 with multiple attempts.\r\n\r\n[/quote]\r\nAdd these count features to the benchmark code, you will get to at least 0.95 with Random Forest\r\n\r\n        row['spaces'] = text.count(' ')\r\n        row['tabs'] = text.count('\\t')\r\n        row['braces'] = text.count('{')\r\n        row['brackets'] = text.count('[')\r\n        row['words'] = len(re.split('\\s+', text))\r\n        row['length'] = len(text)\r\n        row['bracket_2'] = text.count('<')\r\n        row['bracket_3'] = text.count('(')\r\n        row['re_w'] = len(re.split('\\w+', text))\r\n        row['new_par'] = text.count('\\n\\t')\r\n        row['dbslash'] = text.count('//')\r\n        row['bslash'] = text.count('/')\r\n        row['bol_0'] = text.count('<!')\r\n        row['bol_1'] = text.count('/>')\r\n        row['bol_2'] = text.count('&')\r\n        row['bol_3'] = text.count(';')\r\n        row['bol_4'] = text.count('==')\r\n        row['bol_5'] = text.count('===')\r\n        row['bol_6'] = text.count('css')\r\n        row['bol_7'] = text.count('#')\r\n        row['bol_8'] = text.count('@')\r\n        row['bol_9'] = text.count('$')\r\n        row['bol_10'] = text.count('%')\r\n        row['bol_11'] = text.count('^')\r\n        row['bol_12'] = text.count('+')\r\n        row['bol_13'] = text.count('?')\r\n        row['bol_14'] = text.count('|')\r\n        row['bol_15'] = text.count('\\\\')\r\n        row['bol_16'] = text.count('*')\r\n        row['bol_17'] = text.count('||')\r\n        row['bol_18'] = text.count('\\t\\t')\r\n        row['bol_19'] = text.count('\\t\\t\\t')\r\n\r\n[/quote]\r\n\r\nI did not see that coming ;-) should not have removed the punctuation.",
      "votes": null
    },
    {
      "id": "96199",
      "postDate": "10/15/2015 01:23:31",
      "content": "<p>[quote=Artem;96192]</p>\n\n<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;96184]</p>\n\n<p>Our best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.</p>\n\n<p>There was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.</p>\n\n<p>We also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.</p>\n\n<p>some models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . </p>\n\n<p>I think we reached Level 4 meta-modelling in the end :)\n[/quote]\nHow much RAM it requires? </p>\n\n<p>[/quote]</p>\n\n<p>I managed to run a large 500k single word + 232k bigram + 178k trigram model in 34G space (close to 0.9/1M features). Basically externalizing the entire tf-idf vectorizer. Then you notice how little those classes contribute vs. the cost and how much easier it would be if you just take 2 passes at 2-3x the time... But if memory is lacking...</p>",
      "rawMarkdown": "[quote=Artem;96192]\r\n\r\n[quote=Μαριος Μιχαηλιδης KazAnova;96184]\r\n\r\nOur best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.\r\n\r\nThere was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.\r\n\r\nWe also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.\r\n\r\nsome models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . \r\n\r\nI think we reached Level 4 meta-modelling in the end :)\r\n[/quote]\r\nHow much RAM it requires? \r\n\r\n[/quote]\r\n\r\nI managed to run a large 500k single word + 232k bigram + 178k trigram model in 34G space (close to 0.9/1M features). Basically externalizing the entire tf-idf vectorizer. Then you notice how little those classes contribute vs. the cost and how much easier it would be if you just take 2 passes at 2-3x the time... But if memory is lacking...",
      "votes": null
    },
    {
      "id": "96200",
      "postDate": "10/15/2015 01:24:13",
      "content": "<p>[quote=Gerard Toonstra;96185]</p>\n\n<p>Well, my hats off to David Shinn. He entered very late in the competition, eliminated the leaderboard very quickly and finished 10th, just dethroning us from our precious position. So I truly natively hate him, but also greatly, absolutely respect him for doing this on his own. Hats off to you sir!</p>\n\n<p>[/quote]</p>\n\n<p>Thank @Gerard and @Redo, your constant improvement until the end, even a little, kept me motivated to keep pushing till the end.  I congratulate all of the top winners, I fight tooth and nail just to break top 10.  This was a crazy exhausting and crazy fun competition.  Thanks to orchid, Bang Nguyen, Ashtericz, Redo, and NxGTR &amp; Sky for making just cracking top 10 a nail biting experience.</p>",
      "rawMarkdown": "[quote=Gerard Toonstra;96185]\r\n\r\nWell, my hats off to David Shinn. He entered very late in the competition, eliminated the leaderboard very quickly and finished 10th, just dethroning us from our precious position. So I truly natively hate him, but also greatly, absolutely respect him for doing this on his own. Hats off to you sir!\r\n\r\n[/quote]\r\n\r\nThank @Gerard and @Redo, your constant improvement until the end, even a little, kept me motivated to keep pushing till the end.  I congratulate all of the top winners, I fight tooth and nail just to break top 10.  This was a crazy exhausting and crazy fun competition.  Thanks to orchid, Bang Nguyen, Ashtericz, Redo, and NxGTR & Sky for making just cracking top 10 a nail biting experience.",
      "votes": null
    },
    {
      "id": "96201",
      "postDate": "10/15/2015 01:25:32",
      "content": "<p>[quote=Subhajit Mandal;96196]</p>\n\n<p>[quote=Gerard Toonstra;96191]</p>\n\n<p>Horrifying. What were the features? I trained FTRL on 4-5 epochs on 1800 raw features, reduced from 960k.  My features were the most prevalent 1800 word features with optimized L1/L2 regularization.</p>\n\n<p>I went fruity-loopy in this competition and made a custom C parser, a bigram/trigram processor in python and in the end I was processing pretty much portions of matrices to get things going (scipy.sparse.hstack). All of this to reduce the memory requirements for processing the data and decreasing the time before training and model verification would occur.</p>\n\n<p>[/quote]</p>\n\n<p>For FTRL, we used simple binary counts (is it there or not?) of unigrams and bigrams as features. We tried many more feature engineering techniques, but the simple was the best.</p>\n\n<p>[/quote]</p>\n\n<p>Ah, I only included the positive presence 1, but not the 0 -presence... This being FTRL, it's basically unrolling your matrix to disk...</p>",
      "rawMarkdown": "[quote=Subhajit Mandal;96196]\r\n\r\n[quote=Gerard Toonstra;96191]\r\n\r\nHorrifying. What were the features? I trained FTRL on 4-5 epochs on 1800 raw features, reduced from 960k.  My features were the most prevalent 1800 word features with optimized L1/L2 regularization.\r\n\r\nI went fruity-loopy in this competition and made a custom C parser, a bigram/trigram processor in python and in the end I was processing pretty much portions of matrices to get things going (scipy.sparse.hstack). All of this to reduce the memory requirements for processing the data and decreasing the time before training and model verification would occur.\r\n\r\n\r\n[/quote]\r\n\r\nFor FTRL, we used simple binary counts (is it there or not?) of unigrams and bigrams as features. We tried many more feature engineering techniques, but the simple was the best.\r\n\r\n[/quote]\r\n\r\nAh, I only included the positive presence 1, but not the 0 -presence... This being FTRL, it's basically unrolling your matrix to disk...",
      "votes": null
    },
    {
      "id": "96202",
      "postDate": "10/15/2015 01:45:03",
      "content": "<p>[quote=Gerard Toonstra;96185]</p>\n\n<p>What we lacked were ensembling techniques and variety in models or data.</p>\n\n<p>[/quote]</p>\n\n<p>Well, if you had done what I did, you sure would have bested me easily.  My single best model only barely reaches 0.980 on the Private LB.  What I lacked in a single model I made up for in my Frankenstein stacked model.  I just counted, I had 80 models as features in my second layer (plus around 20 or so other basic features).  I used Random Forest, Extra Trees, XGBoost, Vowpal Wabbit, KNN, SVM, and logistic regression in my first set of models.  Each of those trained on different subset of features: character counts, html tags, domain urls, html attributes, file extensions, and 123gram tfidf and 123gram counts.  I even tried some crazy things like the character and html tag counts from the first half of the file and the second half of the file as different training feature sets (sort of a data augmentation).  In my second layer, I used a boosted random forest and xgboost models.  My third layer was a hyperopt optimized power and linear transformation of those 2 predictions from the second layer.  This was my first stab at serious stacking, so it was a great learning experience.</p>",
      "rawMarkdown": "[quote=Gerard Toonstra;96185]\r\n\r\nWhat we lacked were ensembling techniques and variety in models or data.\r\n\r\n[/quote]\r\n\r\nWell, if you had done what I did, you sure would have bested me easily.  My single best model only barely reaches 0.980 on the Private LB.  What I lacked in a single model I made up for in my Frankenstein stacked model.  I just counted, I had 80 models as features in my second layer (plus around 20 or so other basic features).  I used Random Forest, Extra Trees, XGBoost, Vowpal Wabbit, KNN, SVM, and logistic regression in my first set of models.  Each of those trained on different subset of features: character counts, html tags, domain urls, html attributes, file extensions, and 123gram tfidf and 123gram counts.  I even tried some crazy things like the character and html tag counts from the first half of the file and the second half of the file as different training feature sets (sort of a data augmentation).  In my second layer, I used a boosted random forest and xgboost models.  My third layer was a hyperopt optimized power and linear transformation of those 2 predictions from the second layer.  This was my first stab at serious stacking, so it was a great learning experience.",
      "votes": null
    },
    {
      "id": "96205",
      "postDate": "10/15/2015 02:36:56",
      "content": "<p>I didn't do nearly as well as you guys, but almost all of my models were just variants of TINRTGU's code in pure Python. I just greedily averaged these FTRLP with different orders of files, different one hot splitting techniques, etc. I wish I had time for better models, but I ran out of time in terms of creating meta features to be fed into XG.</p>",
      "rawMarkdown": "I didn't do nearly as well as you guys, but almost all of my models were just variants of TINRTGU's code in pure Python. I just greedily averaged these FTRLP with different orders of files, different one hot splitting techniques, etc. I wish I had time for better models, but I ran out of time in terms of creating meta features to be fed into XG.",
      "votes": null
    },
    {
      "id": "96206",
      "postDate": "10/15/2015 02:38:39",
      "content": "<p>@Mad Professors: Congratulations, and good job!  Don't worry, I knew our positions could easily reverse.  It was fun to have a real shot at the top position for a while.</p>\n\n<p>My submission was just the scaled sum of two models, a linear SVM of skip-grams and element/attribute combinations (79 million unique features, ~45 minutes training time), and an XGBoost model using substrings as features (30,000 features, 2000 trees, ~5 hours training time).</p>\n\n<p>The substrings were automatically selected using this tool: <a href=\"https://github.com/mortehu/substring-frequencies\">https://github.com/mortehu/substring-frequencies</a></p>\n\n<p>These are the substrings used as features in the XGBoost model, with the most predictive being listed first: <a href=\"https://gist.github.com/mortehu/07f0c59dc30bf495853d\">https://gist.github.com/mortehu/07f0c59dc30bf495853d</a></p>",
      "rawMarkdown": "Mad Professors: Congratulations, and good job!  Don't worry, I knew our positions could easily reverse.  It was fun to have a real shot at the top position for a while.\r\n\r\nMy submission was just the scaled sum of two models, a linear SVM of skip-grams and element/attribute combinations (79 million unique features, ~45 minutes training time), and an XGBoost model using substrings as features (30,000 features, 2000 trees, ~5 hours training time).\r\n\r\nThe substrings were automatically selected using this tool: https://github.com/mortehu/substring-frequencies\r\n\r\nThese are the substrings used as features in the XGBoost model, with the most predictive being listed first: https://gist.github.com/mortehu/07f0c59dc30bf495853d",
      "votes": null
    },
    {
      "id": "96209",
      "postDate": "10/15/2015 03:02:50",
      "content": "<p>Our feature set is huge, so it is painful to try stacking or other complicated stuff. But if I don't  go into the Physics one, I will probably have time to finish tfidf and two stage stacking. I was too optimistic to think we can perform well on two competitions at the same time. With all the other works, I just don't have enough energy and computers to try lots of things in two competitions.</p>\n\n<p>[quote=NxGTR;96193]</p>\n\n<p>For us it was like this:</p>\n\n<p>Xgboost (single): 0.98422</p>\n\n<p>Xgboost (dual ensemble): 0.98459</p>\n\n<p>RandomForest (single): 0.97455</p>\n\n<p>ExtraTrees (single): 0.97628</p>\n\n<p>Done :D, our final submission is an ensemble of pretty much those 4. </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Our feature set is huge, so it is painful to try stacking or other complicated stuff. But if I don't  go into the Physics one, I will probably have time to finish tfidf and two stage stacking. I was too optimistic to think we can perform well on two competitions at the same time. With all the other works, I just don't have enough energy and computers to try lots of things in two competitions.\r\n\r\n[quote=NxGTR;96193]\r\n\r\nFor us it was like this:\r\n\r\nXgboost (single): 0.98422\r\n\r\nXgboost (dual ensemble): 0.98459\r\n\r\nRandomForest (single): 0.97455\r\n\r\nExtraTrees (single): 0.97628\r\n\r\nDone :D, our final submission is an ensemble of pretty much those 4. \r\n\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "96210",
      "postDate": "10/15/2015 03:08:45",
      "content": "<p>You <strong>have</strong> performed well on both of the competitions. In Physics competition it was a really bad luck for you that high scoring scripts got posted at the last moment. Luckily, there was no scripts in this competition.</p>",
      "rawMarkdown": "You **have** performed well on both of the competitions. In Physics competition it was a really bad luck for you that high scoring scripts got posted at the last moment. Luckily, there was no scripts in this competition.",
      "votes": null
    },
    {
      "id": "96213",
      "postDate": "10/15/2015 03:15:11",
      "content": "<p>It is amazing you can get such good results with a simple ensemble of two models. \nBut I guess your two models are probably very strong by themselves. </p>\n\n<p>May I ask how are you able to train a linear SVM with 79 million features in 45 minutes,  I always find svm to be very slow. Another thing is did you use some software package to extract skip gram features?</p>\n\n<p>[quote=mortehu;96206]</p>\n\n<p>@Mad Professors: Congratulations, and good job!  Don't worry, I knew our positions could easily reverse.  It was fun to have a real shot at the top position for a while.</p>\n\n<p>My submission was just the scaled sum of two models, a linear SVM of skip-grams and element/attribute combinations (79 million unique features, ~45 minutes training time), and an XGBoost model using substrings as features (30,000 features, 2000 trees, ~5 hours training time).</p>\n\n<p>The substrings were automatically selected using this tool: <a href=\"https://github.com/mortehu/substring-frequencies\">https://github.com/mortehu/substring-frequencies</a></p>\n\n<p>These are the substrings used as features in the XGBoost model, with the most predictive being listed first: <a href=\"https://gist.github.com/mortehu/07f0c59dc30bf495853d\">https://gist.github.com/mortehu/07f0c59dc30bf495853d</a></p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "It is amazing you can get such good results with a simple ensemble of two models. \r\nBut I guess your two models are probably very strong by themselves. \r\n\r\nMay I ask how are you able to train a linear SVM with 79 million features in 45 minutes,  I always find svm to be very slow. Another thing is did you use some software package to extract skip gram features?\r\n\r\n[quote=mortehu;96206]\r\n\r\n@Mad Professors: Congratulations, and good job!  Don't worry, I knew our positions could easily reverse.  It was fun to have a real shot at the top position for a while.\r\n\r\nMy submission was just the scaled sum of two models, a linear SVM of skip-grams and element/attribute combinations (79 million unique features, ~45 minutes training time), and an XGBoost model using substrings as features (30,000 features, 2000 trees, ~5 hours training time).\r\n\r\nThe substrings were automatically selected using this tool: https://github.com/mortehu/substring-frequencies\r\n\r\nThese are the substrings used as features in the XGBoost model, with the most predictive being listed first: https://gist.github.com/mortehu/07f0c59dc30bf495853d\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "96214",
      "postDate": "10/15/2015 03:15:20",
      "content": "<p>[quote=SkyLibrary;96209]\nI was too optimistic to think we can perform well on two competitions at the same time...\n[/quote]</p>\n\n<p>Hehehe, Top Kagglers issues: when two &quot;top 13th&quot; is bad! :P</p>\n\n<p>It was OK, could be better but heck!, pretty much always can be better :D</p>",
      "rawMarkdown": "[quote=SkyLibrary;96209]\r\nI was too optimistic to think we can perform well on two competitions at the same time...\r\n[/quote]\r\n\r\nHehehe, Top Kagglers issues: when two \"top 13th\" is bad! :P\r\n\r\nIt was OK, could be better but heck!, pretty much always can be better :D",
      "votes": null
    },
    {
      "id": "96217",
      "postDate": "10/15/2015 03:26:12",
      "content": "<p>[quote=SkyLibrary;96213]</p>\n\n<p>May I ask how are you able to train a linear SVM with 79 million features in 45 minutes,  I always find svm to be very slow. Another thing is did you use some software package to extract skip gram features?</p>\n\n<p>[/quote]</p>\n\n<p>mortehu is an amazing programmer according to his repo. He probably just coded it from scratch in c++</p>",
      "rawMarkdown": "[quote=SkyLibrary;96213]\r\n\r\nMay I ask how are you able to train a linear SVM with 79 million features in 45 minutes,  I always find svm to be very slow. Another thing is did you use some software package to extract skip gram features?\r\n\r\n[/quote]\r\n\r\nmortehu is an amazing programmer according to his repo. He probably just coded it from scratch in c++",
      "votes": null
    },
    {
      "id": "96218",
      "postDate": "10/15/2015 03:27:29",
      "content": "<p>lol. Two 13th is definitely better than two 26th. But wait a minutes, does 13 suppose to be a unlucky number in some countries, luckily I don't come from these countries so I don't have hard feelings for it.\n[quote=NxGTR;96214]</p>\n\n<p>[quote=SkyLibrary;96209]\nI was too optimistic to think we can perform well on two competitions at the same time...\n[/quote]</p>\n\n<p>Hehehe, Top Kagglers issues: when two &quot;top 13th&quot; is bad! :P</p>\n\n<p>It was OK, could be better but heck!, pretty much always can be better :D</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "lol. Two 13th is definitely better than two 26th. But wait a minutes, does 13 suppose to be a unlucky number in some countries, luckily I don't come from these countries so I don't have hard feelings for it.\r\n[quote=NxGTR;96214]\r\n\r\n[quote=SkyLibrary;96209]\r\nI was too optimistic to think we can perform well on two competitions at the same time...\r\n[/quote]\r\n\r\nHehehe, Top Kagglers issues: when two \"top 13th\" is bad! :P\r\n\r\nIt was OK, could be better but heck!, pretty much always can be better :D\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "96220",
      "postDate": "10/15/2015 03:45:28",
      "content": "<p>Congrats Mad Professors, mortehu and bibaze.! Special thanks to @David shinn for his excellent starter code which helped us enter this competition since we started very late. Thanks everyone for posting their approaches. </p>\n\n<p>@Marios: 4-level meta models.! You are pushing the boundaries further. Looking forward to your team's code soon. \n@rcarson: One more top 10 finish to your list. You pretty much get there almost every time :)</p>",
      "rawMarkdown": "Congrats Mad Professors, mortehu and bibaze.! Special thanks to @David shinn for his excellent starter code which helped us enter this competition since we started very late. Thanks everyone for posting their approaches. \r\n\r\n@Marios: 4-level meta models.! You are pushing the boundaries further. Looking forward to your team's code soon. \r\n@rcarson: One more top 10 finish to your list. You pretty much get there almost every time :)",
      "votes": null
    },
    {
      "id": "96222",
      "postDate": "10/15/2015 03:57:02",
      "content": "<p>[quote=mortehu;96206]</p>\n\n<p>@Mad Professors: Congratulations, and good job!  Don't worry, I knew our positions could easily reverse.  It was fun to have a real shot at the top position for a while.</p>\n\n<p>My submission was just the scaled sum of two models, a linear SVM of skip-grams and element/attribute combinations (79 million unique features, ~45 minutes training time), and an XGBoost model using substrings as features (30,000 features, 2000 trees, ~5 hours training time).</p>\n\n<p>The substrings were automatically selected using this tool: <a href=\"https://github.com/mortehu/substring-frequencies\">https://github.com/mortehu/substring-frequencies</a></p>\n\n<p>These are the substrings used as features in the XGBoost model, with the most predictive being listed first: <a href=\"https://gist.github.com/mortehu/07f0c59dc30bf495853d\">https://gist.github.com/mortehu/07f0c59dc30bf495853d</a></p>\n\n<p>[/quote]\nThese codes are awesome.. I wish I can code anywhere near that level someday.</p>\n\n<p>Something interesting I noticed in the substrings list is that there are many &quot;dates&quot; strings that coincide with the leak found mid competition.. I hadn't figured out but perhaps most of the pages have a 'current time' recorded in the html file. Maybe we could call that a 'nerfed' version of the leak. Clueless about what that '1529' would mean, though. edit: maybe week number?</p>\n\n<p>Also, anybody has an idea as to why these sponsored ads seems so predictable through features even though it doesn't have any apparent correlation? After all, these sponsored pages are just pages people paid StumbleUpon to feature in their system. Never thought background features like number of tags or chars could be so effective at predicting that.</p>\n\n<p>Congrats to everyone and I'm specially thankful to David Shinn, whose code got me started in the competition as well.</p>",
      "rawMarkdown": "[quote=mortehu;96206]\r\n\r\n@Mad Professors: Congratulations, and good job!  Don't worry, I knew our positions could easily reverse.  It was fun to have a real shot at the top position for a while.\r\n\r\nMy submission was just the scaled sum of two models, a linear SVM of skip-grams and element/attribute combinations (79 million unique features, ~45 minutes training time), and an XGBoost model using substrings as features (30,000 features, 2000 trees, ~5 hours training time).\r\n\r\nThe substrings were automatically selected using this tool: https://github.com/mortehu/substring-frequencies\r\n\r\nThese are the substrings used as features in the XGBoost model, with the most predictive being listed first: https://gist.github.com/mortehu/07f0c59dc30bf495853d\r\n\r\n[/quote]\r\nThese codes are awesome.. I wish I can code anywhere near that level someday.\r\n\r\nSomething interesting I noticed in the substrings list is that there are many \"dates\" strings that coincide with the leak found mid competition.. I hadn't figured out but perhaps most of the pages have a 'current time' recorded in the html file. Maybe we could call that a 'nerfed' version of the leak. Clueless about what that '1529' would mean, though. edit: maybe week number?\r\n\r\nAlso, anybody has an idea as to why these sponsored ads seems so predictable through features even though it doesn't have any apparent correlation? After all, these sponsored pages are just pages people paid StumbleUpon to feature in their system. Never thought background features like number of tags or chars could be so effective at predicting that.\r\n\r\nCongrats to everyone and I'm specially thankful to David Shinn, whose code got me started in the competition as well.",
      "votes": null
    },
    {
      "id": "96223",
      "postDate": "10/15/2015 04:01:39",
      "content": "<p>As far as I can tell, this competition was more about determining the date each page was downloaded, rather than whether or not something was advertising.  Negative examples were mostly downloaded towards the end, for example.  Current events, like &quot;Pluto&quot; (New Horizons passed Pluto on July 24th), &quot;Donald Trump&quot;, and &quot;El Chapo&quot; were all strong features, with &quot;infographic&quot; appearing pretty far down on the list.  Some of the other number strings seen in my feature lists are versions of JavaScript libraries that were released after the positive samples had been downloaded.</p>\n\n<p>I'll post the source code for my linear SVM training and HTML tokenization program tomorrow.</p>",
      "rawMarkdown": "As far as I can tell, this competition was more about determining the date each page was downloaded, rather than whether or not something was advertising.  Negative examples were mostly downloaded towards the end, for example.  Current events, like \"Pluto\" (New Horizons passed Pluto on July 24th), \"Donald Trump\", and \"El Chapo\" were all strong features, with \"infographic\" appearing pretty far down on the list.  Some of the other number strings seen in my feature lists are versions of JavaScript libraries that were released after the positive samples had been downloaded.\r\n\r\nI'll post the source code for my linear SVM training and HTML tokenization program tomorrow.",
      "votes": null
    },
    {
      "id": "96239",
      "postDate": "10/15/2015 07:38:53",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;96184]</p>\n\n<p>Our best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.</p>\n\n<p>There was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.</p>\n\n<p>We also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.</p>\n\n<p>some models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . </p>\n\n<p>I think we reached Level 4 meta-modelling in the end :)</p>\n\n<p>[/quote]\nI am curious about your 3-grams tf-idf, do you include the tags and scripts of HTML files? Seems that the tags and scripts are only helpful to classify HTML files from the same websites(they share the same HTML structure.) I would be appreciate if you can share more about features and how you clean HTML files, thanks a lot!</p>",
      "rawMarkdown": "[quote=Μαριος Μιχαηλιδης KazAnova;96184]\r\n\r\nOur best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.\r\n\r\nThere was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.\r\n\r\nWe also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.\r\n\r\nsome models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . \r\n\r\nI think we reached Level 4 meta-modelling in the end :)\r\n\r\n\r\n\r\n[/quote]\r\nI am curious about your 3-grams tf-idf, do you include the tags and scripts of HTML files? Seems that the tags and scripts are only helpful to classify HTML files from the same websites(they share the same HTML structure.) I would be appreciate if you can share more about features and how you clean HTML files, thanks a lot!",
      "votes": null
    },
    {
      "id": "96249",
      "postDate": "10/15/2015 09:13:53",
      "content": "<p>One doubt I had... I started with tf-idf on xgb, but noticed that IDF itself had no impact on end results. Also, the official application of term frequency seemed to slightly reduce the results as well, probably because the size of the documents varies quite a bit. So my XGB would perform 0.96/0.97 on TF calculated as : f(t) / n(t,d)  , whereas a simple f(t) (just a basic count) would get me 0.98 performance. Can you guys confirm you got better results with simple counts as well?</p>\n\n<p>My max_depth for xgb was optimized at 50 with mostly default settings, except colsample_by_tree set to 0.46. subsampling made no difference and neither did gamma. &quot;scale_pos_weight&quot; sounds like a great opportunity to tweak, because 1's were less than 0's and it seemed to work in the first few iterations, but then it leveled off very quickly.</p>\n\n<p>I also did a full minhash doc-doc comparison on 150 hashes. I calculated the hashes in python, but wrote a custom C app to do the comparison. That got me 0.89 on a 20% similarity threshold setting. A combination of 80/50/20% similarity results got me 0.929.</p>",
      "rawMarkdown": "One doubt I had... I started with tf-idf on xgb, but noticed that IDF itself had no impact on end results. Also, the official application of term frequency seemed to slightly reduce the results as well, probably because the size of the documents varies quite a bit. So my XGB would perform 0.96/0.97 on TF calculated as : f(t) / n(t,d)  , whereas a simple f(t) (just a basic count) would get me 0.98 performance. Can you guys confirm you got better results with simple counts as well?\r\n\r\nMy max_depth for xgb was optimized at 50 with mostly default settings, except colsample_by_tree set to 0.46. subsampling made no difference and neither did gamma. \"scale_pos_weight\" sounds like a great opportunity to tweak, because 1's were less than 0's and it seemed to work in the first few iterations, but then it leveled off very quickly.\r\n\r\nI also did a full minhash doc-doc comparison on 150 hashes. I calculated the hashes in python, but wrote a custom C app to do the comparison. That got me 0.89 on a 20% similarity threshold setting. A combination of 80/50/20% similarity results got me 0.929.",
      "votes": null
    },
    {
      "id": "96252",
      "postDate": "10/15/2015 09:44:04",
      "content": "<p>[quote=SRK;96220]</p>\n\n<p>Congrats Mad Professors, mortehu and bibaze.! Special thanks to @David shinn for his excellent starter code which helped us enter this competition since we started very late. Thanks everyone for posting their approaches. </p>\n\n<p>@Marios: 4-level meta models.! You are pushing the boundaries further. Looking forward to your team's code soon. \n@rcarson: One more top 10 finish to your list. You pretty much get there almost every time :)</p>\n\n<p>[/quote]</p>\n\n<p>Congratulations Mad Professors - awesome finish @mortehu - your repo looks intimidating!\nThanks to @DavidShinn for the starter code without which I wouldn't have even thought of attempting this daunting challenge.</p>\n\n<p>@SRK, special thanks to U for considering to team up with a Kaggler, despite being a Master yourself! </p>\n\n<p>@Everybody else, SRK is an awesome teammate. <em>Endorsing him</em></p>\n\n<p>We got late into the competition (~17 days to go), ignored the competition for a few more days and actually worked only for the last 3 days to be very honest and ending up at #33 is no mean feat. Very happy to finish well and it could only happen with having a teammate like SRK who pretty much guided me throughout the last 3 days where we tried what we could! :-)</p>\n\n<p>I have learnt a lot of things and continue to learn more and do well :-)</p>\n\n<p>@rcarson - got to know a lot about you. SRK spoke very highly of you. I hope I team up with you someday. For that, I'll work hard, earn the right and ask U :-)</p>",
      "rawMarkdown": "[quote=SRK;96220]\r\n\r\nCongrats Mad Professors, mortehu and bibaze.! Special thanks to @David shinn for his excellent starter code which helped us enter this competition since we started very late. Thanks everyone for posting their approaches. \r\n\r\n@Marios: 4-level meta models.! You are pushing the boundaries further. Looking forward to your team's code soon. \r\n@rcarson: One more top 10 finish to your list. You pretty much get there almost every time :)\r\n\r\n\r\n[/quote]\r\n\r\nCongratulations Mad Professors - awesome finish @mortehu - your repo looks intimidating!\r\nThanks to @DavidShinn for the starter code without which I wouldn't have even thought of attempting this daunting challenge.\r\n\r\n@SRK, special thanks to U for considering to team up with a Kaggler, despite being a Master yourself! \r\n\r\n@Everybody else, SRK is an awesome teammate. *Endorsing him*\r\n\r\nWe got late into the competition (~17 days to go), ignored the competition for a few more days and actually worked only for the last 3 days to be very honest and ending up at #33 is no mean feat. Very happy to finish well and it could only happen with having a teammate like SRK who pretty much guided me throughout the last 3 days where we tried what we could! :-)\r\n\r\nI have learnt a lot of things and continue to learn more and do well :-)\r\n\r\n@rcarson - got to know a lot about you. SRK spoke very highly of you. I hope I team up with you someday. For that, I'll work hard, earn the right and ask U :-)",
      "votes": null
    },
    {
      "id": "96273",
      "postDate": "10/15/2015 12:47:51",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;96184]</p>\n\n<p>Our best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.</p>\n\n<p>There was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.</p>\n\n<p>We also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.</p>\n\n<p>some models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . </p>\n\n<p>I think we reached Level 4 meta-modelling in the end :)</p>\n\n<p>[/quote]</p>\n\n<p>As a n00b, can I ask how exactly did you preprocess those files? Or in general, how should I deal with such noisy data in the future, if I'd like to extract some usefull TF-IFDs?</p>",
      "rawMarkdown": "[quote=Μαριος Μιχαηλιδης KazAnova;96184]\r\n\r\nOur best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.\r\n\r\nThere was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.\r\n\r\nWe also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.\r\n\r\nsome models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . \r\n\r\nI think we reached Level 4 meta-modelling in the end :)\r\n\r\n[/quote]\r\n\r\n\r\nAs a n00b, can I ask how exactly did you preprocess those files? Or in general, how should I deal with such noisy data in the future, if I'd like to extract some usefull TF-IFDs?",
      "votes": null
    },
    {
      "id": "96277",
      "postDate": "10/15/2015 13:02:27",
      "content": "<p>Actually,  I just loaded each file as a big string (lets call it mline)  and then cleaned with this:</p>\n\n<p>import re</p>\n\n<p>mline=re.sub(&quot;[^a-zA-Z0-9]&quot;,&quot; &quot;, mline) # remove non-alphanumeric</p>\n\n<p>then I re-Wrote this to a file. </p>\n\n<p>The test is up to the to the tf-idf parameters!</p>\n\n<p>tfv=TfidfVectorizer(min_df=2, max_features=None, strip_accents='unicode',lowercase =True,\n                        analyzer='word', dtype=np.float64, token_pattern=r'\\w{2,}', ngram_range=(1, 3), use_idf=True,smooth_idf=True, \n    sublinear_tf=True, stop_words = 'english')   </p>",
      "rawMarkdown": "Actually,  I just loaded each file as a big string (lets call it mline)  and then cleaned with this:\r\n\r\nimport re\r\n \r\nmline=re.sub(\"[^a-zA-Z0-9]\",\" \", mline) # remove non-alphanumeric\r\n \r\nthen I re-Wrote this to a file. \r\n\r\nThe test is up to the to the tf-idf parameters!\r\n\r\n tfv=TfidfVectorizer(min_df=2, max_features=None, strip_accents='unicode',lowercase =True,\r\n                        analyzer='word', dtype=np.float64, token_pattern=r'\\w{2,}', ngram_range=(1, 3), use_idf=True,smooth_idf=True, \r\n    sublinear_tf=True, stop_words = 'english')",
      "votes": null
    },
    {
      "id": "96282",
      "postDate": "10/15/2015 14:36:58",
      "content": "<p>Thanks Dato, Kaggle, team mates and all competitors! Grats all winners and new Kaggle Masters.</p>\n\n<p>My focus was on online learning and generating non-alphanumeric features.</p>\n\n<p>I could not extract the entire dataset on my HD, so I worked with the compressed archives only, unzipping and streaming the documents into the algo's VW, FTRL, SGD and perceptron. The scores were not as high as I've seen reported in this thread (well done!), but I think the models/vectorizers were diverse enough to contribute to the ensemble.</p>\n\n<p>We also pursued three crazy ideas that happened to work (a tiny bit, but enough to be exciting for further research):</p>\n\n<ul>\n<li><p>Training a model on a sentiment analysis data set unrelated to this competition, then using the predictions for the train and test sets for further modeling. Such a &quot;sentiment&quot;-feature was informative.</p></li>\n<li><p>Training a model on the 20 newsgroups data set and creating multi-class predictions for train and test sets. I think these vectors added a form of categorization/topic modeling. (&quot;this document talks about 'SciMed' and 'CompSci', but not 'Religion'&quot;).</p></li>\n<li><p>Using Normalized Compression Distance with the fast Snappy compressor to calculate distance between every document and 4 randomly created anchor corpora.</p></li>\n</ul>\n\n<p>Again, learned a lot, for instance about geometric mean for model averaging, feature and meta-model interactions with <a href=\"https://github.com/far0n/xgbfi\">XGBfi</a>, proper competition management/iteration, patience, and stacking. I am particularly excited about our 4 levels of stacking, with 4th level models using the predictions from the 1st-3rd level models (fully connected stacknet?).</p>",
      "rawMarkdown": "Thanks Dato, Kaggle, team mates and all competitors! Grats all winners and new Kaggle Masters.\r\n\r\nMy focus was on online learning and generating non-alphanumeric features.\r\n\r\nI could not extract the entire dataset on my HD, so I worked with the compressed archives only, unzipping and streaming the documents into the algo's VW, FTRL, SGD and perceptron. The scores were not as high as I've seen reported in this thread (well done!), but I think the models/vectorizers were diverse enough to contribute to the ensemble.\r\n\r\nWe also pursued three crazy ideas that happened to work (a tiny bit, but enough to be exciting for further research):\r\n\r\n- Training a model on a sentiment analysis data set unrelated to this competition, then using the predictions for the train and test sets for further modeling. Such a \"sentiment\"-feature was informative.\r\n\r\n- Training a model on the 20 newsgroups data set and creating multi-class predictions for train and test sets. I think these vectors added a form of categorization/topic modeling. (\"this document talks about 'SciMed' and 'CompSci', but not 'Religion'\").\r\n\r\n- Using Normalized Compression Distance with the fast Snappy compressor to calculate distance between every document and 4 randomly created anchor corpora.\r\n\r\nAgain, learned a lot, for instance about geometric mean for model averaging, feature and meta-model interactions with [XGBfi][1], proper competition management/iteration, patience, and stacking. I am particularly excited about our 4 levels of stacking, with 4th level models using the predictions from the 1st-3rd level models (fully connected stacknet?).\r\n\r\n  [1]: https://github.com/far0n/xgbfi",
      "votes": null
    },
    {
      "id": "96289",
      "postDate": "10/15/2015 15:19:15",
      "content": "<p>[quote=Triskelion;96282]\nAgain, learned a lot, for instance about geometric mean for model averaging, feature and meta-model interactions with XGBfi.\n[/quote]</p>\n\n<p>That looks very interesting Thanks!</p>\n\n<p>I gave it a quick read, but, is that Windows-Only (XGBfi)?</p>",
      "rawMarkdown": "[quote=Triskelion;96282]\r\nAgain, learned a lot, for instance about geometric mean for model averaging, feature and meta-model interactions with XGBfi.\r\n[/quote]\r\n\r\nThat looks very interesting Thanks!\r\n\r\nI gave it a quick read, but, is that Windows-Only (XGBfi)?",
      "votes": null
    },
    {
      "id": "96297",
      "postDate": "10/15/2015 16:01:11",
      "content": "<p>I think so, though maybe you can run on Linux Mono. Far0n is considering porting to C++.</p>\n\n<p>I think the nicest would be to have it added to XGBoost source. It's really useful. Adding high-ranking interactions as features for a linear algorithm boosts its score.</p>",
      "rawMarkdown": "I think so, though maybe you can run on Linux Mono. Far0n is considering porting to C++.\r\n\r\nI think the nicest would be to have it added to XGBoost source. It's really useful. Adding high-ranking interactions as features for a linear algorithm boosts its score.",
      "votes": null
    },
    {
      "id": "96301",
      "postDate": "10/15/2015 16:07:43",
      "content": "<p>Here's the text-classifier I used for this competition released as open source: <a href=\"https://github.com/mortehu/text-classifier\">https://github.com/mortehu/text-classifier</a></p>\n\n<p>I hope this can be used to set a high baseline in any future text classification competition, or to be more generally useful.</p>",
      "rawMarkdown": "Here's the text-classifier I used for this competition released as open source: https://github.com/mortehu/text-classifier\r\n\r\nI hope this can be used to set a high baseline in any future text classification competition, or to be more generally useful.",
      "votes": null
    },
    {
      "id": "96307",
      "postDate": "10/15/2015 16:17:40",
      "content": "<p>So who used graphlab?</p>",
      "rawMarkdown": "So who used graphlab?",
      "votes": null
    },
    {
      "id": "96324",
      "postDate": "10/15/2015 17:39:19",
      "content": "<p>[quote=NxGTR;96289]</p>\n\n<p>[quote=Triskelion;96282]\nAgain, learned a lot, for instance about geometric mean for model averaging, feature and meta-model interactions with XGBfi.\n[/quote]</p>\n\n<p>That looks very interesting Thanks!</p>\n\n<p>I gave it a quick read, but, is that Windows-Only (XGBfi)?</p>\n\n<p>[/quote]</p>\n\n<p>So I just checked and it runs fine under linux using mono. </p>\n\n<p>I know .Net is a bit odd, but I'm planning to build a GUI interfacing XGB with features like training graphs, comparing trainings graphs of different parameter sets, querying feature split values, feature importance based on hold out data performance (which I expect to be useful to detect features which rank high but hurt performance) and so on. But I'm only familar with WPF for creating GUIs. That is why xgbfi is currently written in .Net.</p>",
      "rawMarkdown": "[quote=NxGTR;96289]\r\n\r\n[quote=Triskelion;96282]\r\nAgain, learned a lot, for instance about geometric mean for model averaging, feature and meta-model interactions with XGBfi.\r\n[/quote]\r\n\r\nThat looks very interesting Thanks!\r\n\r\nI gave it a quick read, but, is that Windows-Only (XGBfi)?\r\n\r\n\r\n[/quote]\r\n\r\nSo I just checked and it runs fine under linux using mono. \r\n\r\nI know .Net is a bit odd, but I'm planning to build a GUI interfacing XGB with features like training graphs, comparing trainings graphs of different parameter sets, querying feature split values, feature importance based on hold out data performance (which I expect to be useful to detect features which rank high but hurt performance) and so on. But I'm only familar with WPF for creating GUIs. That is why xgbfi is currently written in .Net.",
      "votes": null
    },
    {
      "id": "96325",
      "postDate": "10/15/2015 17:41:04",
      "content": "<p>[quote=mortehu;96301]</p>\n\n<p>Here's the text-classifier I used for this competition released as open source: <a href=\"https://github.com/mortehu/text-classifier\">https://github.com/mortehu/text-classifier</a></p>\n\n<p>I hope this can be used to set a high baseline in any future text classification competition, or to be more generally useful.</p>\n\n<p>[/quote]</p>\n\n<p>just wow!</p>",
      "rawMarkdown": "[quote=mortehu;96301]\r\n\r\nHere's the text-classifier I used for this competition released as open source: https://github.com/mortehu/text-classifier\r\n\r\nI hope this can be used to set a high baseline in any future text classification competition, or to be more generally useful.\r\n\r\n[/quote]\r\n\r\njust wow!",
      "votes": null
    },
    {
      "id": "96369",
      "postDate": "10/15/2015 22:25:35",
      "content": "<p>Dumb question, what is &quot;FTRL&quot;?</p>\n\n<p>Also, what 2 files did you put into this tool to select substrings? <a href=\"https://github.com/mortehu/substring-frequencies\">https://github.com/mortehu/substring-frequencies</a></p>",
      "rawMarkdown": "Dumb question, what is \"FTRL\"?\r\n\r\nAlso, what 2 files did you put into this tool to select substrings? https://github.com/mortehu/substring-frequencies",
      "votes": null
    },
    {
      "id": "96374",
      "postDate": "10/15/2015 22:34:30",
      "content": "<p>[quote=Zach;96369]\nAlso, what 2 files did you put into this tool to select substrings? <a href=\"https://github.com/mortehu/substring-frequencies\">https://github.com/mortehu/substring-frequencies</a>\n[/quote]</p>\n\n<p>I picked a random sample of 512 MB of positive examples, and 512 MB of negative examples.  I used the <code>--document</code> mode.</p>",
      "rawMarkdown": "[quote=Zach;96369]\r\nAlso, what 2 files did you put into this tool to select substrings? https://github.com/mortehu/substring-frequencies\r\n[/quote]\r\n\r\nI picked a random sample of 512 MB of positive examples, and 512 MB of negative examples.  I used the `--document` mode.",
      "votes": null
    },
    {
      "id": "96376",
      "postDate": "10/15/2015 22:39:30",
      "content": "<p>[quote=Zach;96369]</p>\n\n<p>Dumb question, what is &quot;FTRL&quot;?</p>\n\n<p>[/quote]</p>\n\n<p>FTRL stands for the online learning algorithm &quot;Follow The Regularized Leader&quot;.</p>",
      "rawMarkdown": "[quote=Zach;96369]\r\n\r\nDumb question, what is \"FTRL\"?\r\n\r\n[/quote]\r\n\r\nFTRL stands for the online learning algorithm \"Follow The Regularized Leader\".",
      "votes": null
    },
    {
      "id": "96387",
      "postDate": "10/16/2015 00:42:31",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mortehu\">@mortehu</a>,</p>\n\n<p>I'm intrigued and really admiring your way. Just curious, would it be possible to have your substring frequency more (or less) accurate to instruct the suffix array to discount overlapped things? (In case my description sounds vague and weird or this method isn't popular... it's kinda trick to eliminate stop words without knowing then and sometimes emphasize critical information. For example from the earliest paper I can find, when encountered a bigram &quot;rail enquiries&quot; in BNC, it's actually useless because it was always part of either &quot;national rail enquiries&quot; or &quot;British british rail enquiries.&quot;)</p>\n\n<p>I didn't use it just because I didn't get to that phase yet. :p</p>",
      "rawMarkdown": "Hi [@mortehu][1],\r\n\r\nI'm intrigued and really admiring your way. Just curious, would it be possible to have your substring frequency more (or less) accurate to instruct the suffix array to discount overlapped things? (In case my description sounds vague and weird or this method isn't popular... it's kinda trick to eliminate stop words without knowing then and sometimes emphasize critical information. For example from the earliest paper I can find, when encountered a bigram \"rail enquiries\" in BNC, it's actually useless because it was always part of either \"national rail enquiries\" or \"British british rail enquiries.\")\r\n\r\nI didn't use it just because I didn't get to that phase yet. :p\r\n\r\n  [1]: https://www.kaggle.com/mortehu",
      "votes": null
    },
    {
      "id": "96393",
      "postDate": "10/16/2015 02:54:47",
      "content": "<p>@Barabbas, even though it might not look like it, a lot of redundant substrings are already removed. However, the code currently keeps overlapping strings if they have different occurrence counts, even if that is stupid in most cases. It's hard to know without building a model whether a substring of another feature is truly redundant, so I didn't even try. You can see the duplicate detection code on line 249 in substrings.cc.</p>",
      "rawMarkdown": "Barabbas, even though it might not look like it, a lot of redundant substrings are already removed. However, the code currently keeps overlapping strings if they have different occurrence counts, even if that is stupid in most cases. It's hard to know without building a model whether a substring of another feature is truly redundant, so I didn't even try. You can see the duplicate detection code on line 249 in substrings.cc.",
      "votes": null
    },
    {
      "id": "96400",
      "postDate": "10/16/2015 05:44:16",
      "content": "<p><a href=\"https://www.kaggle.com/mortehu\">@mortehu</a>: I see, thank you. :D</p>",
      "rawMarkdown": "[@mortehu][1]: I see, thank you. :D\r\n\r\n\r\n  [1]: https://www.kaggle.com/mortehu",
      "votes": null
    },
    {
      "id": "96414",
      "postDate": "10/16/2015 12:57:25",
      "content": "<p>[quote=Faron;96376]</p>\n\n<p>FTRL stands for the online learning algorithm &quot;Follow The Regularized Leader&quot;.</p>\n\n<p>[/quote]</p>\n\n<p>How does that differ from other online algorithms, e.g. vowpal wabbit?  Where can I get the software to try it out?</p>\n\n<p>Actually it looks like VW has a <a href=\"https://github.com/JohnLangford/vowpal_wabbit/wiki/Command-line-arguments\">--ftrl option</a>.  I'll have to try that out sometime.</p>",
      "rawMarkdown": "[quote=Faron;96376]\r\n\r\nFTRL stands for the online learning algorithm \"Follow The Regularized Leader\".\r\n\r\n[/quote]\r\n\r\nHow does that differ from other online algorithms, e.g. vowpal wabbit?  Where can I get the software to try it out?\r\n\r\nActually it looks like VW has a [--ftrl option][1].  I'll have to try that out sometime.\r\n\r\n\r\n  [1]: https://github.com/JohnLangford/vowpal_wabbit/wiki/Command-line-arguments",
      "votes": null
    },
    {
      "id": "96418",
      "postDate": "10/16/2015 13:05:13",
      "content": "<p>[quote=Zach;96414]</p>\n\n<p>How does that differ from other online algorithms, e.g. vowpal wabbit?  Where can I get the software to try it out?</p>\n\n<p>[/quote]</p>\n\n<p>For me it is like an SGD, with the difference that in each step the coefficients are being copied and converted to &quot;new values&quot; based on the regularization terms. </p>\n\n<p>For l2 regul, this should not have much difference with sgd, but for l1 it makes quite a difference  because the coefficients will always remain zero (0.0) if they never pass the the C value (something not really achievable with simple sgd)  and will never be copied across . </p>\n\n<p>In other words, only if something becomes significant (after the updates's phase) will start being copied to further boost itself (hence the follow the leading features after regularization is applied!) </p>\n\n<p>Also, most of the times it uses passed gradients as a form of adaptive learning rate so generally as an algorithm is more independent. </p>\n\n<p>At least, that is my interpretation!</p>",
      "rawMarkdown": "[quote=Zach;96414]\r\n\r\nHow does that differ from other online algorithms, e.g. vowpal wabbit?  Where can I get the software to try it out?\r\n\r\n[/quote]\r\n\r\nFor me it is like an SGD, with the difference that in each step the coefficients are being copied and converted to \"new values\" based on the regularization terms. \r\n\r\nFor l2 regul, this should not have much difference with sgd, but for l1 it makes quite a difference  because the coefficients will always remain zero (0.0) if they never pass the the C value (something not really achievable with simple sgd)  and will never be copied across . \r\n\r\nIn other words, only if something becomes significant (after the updates's phase) will start being copied to further boost itself (hence the follow the leading features after regularization is applied!) \r\n\r\nAlso, most of the times it uses passed gradients as a form of adaptive learning rate so generally as an algorithm is more independent. \r\n\r\nAt least, that is my interpretation!",
      "votes": null
    },
    {
      "id": "96425",
      "postDate": "10/16/2015 14:52:25",
      "content": "<p>Hi! I want to thank David Shinn for his very nice script, it motivated me to continue working on the comp, since it was so easy to improve :) </p>\n\n<p>After dabbling a while with it and improving it manually, I calculated how much popular (top50k) individual words separated the classes by abs(freq_pos - freq_neg), and took the top3k and used them in RF and XGB as features with no tuning of parameters, then simple average. I did not expect that simple solution to earn us a 19th place. It was a really fun competition due to the size of data!</p>\n\n<p>Also thanks to my teammate</p>",
      "rawMarkdown": "Hi! I want to thank David Shinn for his very nice script, it motivated me to continue working on the comp, since it was so easy to improve :) \r\n\r\nAfter dabbling a while with it and improving it manually, I calculated how much popular (top50k) individual words separated the classes by abs(freq_pos - freq_neg), and took the top3k and used them in RF and XGB as features with no tuning of parameters, then simple average. I did not expect that simple solution to earn us a 19th place. It was a really fun competition due to the size of data!\r\n\r\nAlso thanks to my teammate",
      "votes": null
    },
    {
      "id": "96614",
      "postDate": "10/19/2015 05:29:08",
      "content": "<p>Congratulation to winner. </p>\n\n<p>I ended up at 14th using liblinear and xgboost. I built three 1-gram models: 2 liblinear models with Bernoulli and multinominal feature weights  and 1 xgboost model with Bernoulli feature weights. On the 2nd level I blended those three model using xgboost and linear regression. Then at 3rd level, I used h mean to merge two models at 2nd level.</p>\n\n<p>According to me, preprocessing is the most important thing. I used a regular expression to split the html string to tokens. This regular expression can catch domain names, javascript function calls, html attributes, date, time... Some important features from my xgboost model are:</p>\n\n<pre><code>(u'/2015/07/0', 255), (u'2013%20', 251), (u'keywords!', 246), (u'2014%20fifa%20world%20cup', 244), (u'2015-07-07+21', 214), (u'blog!', 209), (u'/blog%206', 198), (u'other!', 195), (u'july!', 194), (u'copyright!', 192), (u'2015-07-14+00', 191), (u'rss!', 181), (u'14%', 180), (u'20%', 178), (u'web!', 177), (u'100%%', 176), (u'our!', 173), (u'news!', 173), (u'199.302', 172), (u'article!', 170), (u'author!', 170), (u'site_name.indexof(', 170), (u'30%', 168), (u'/plus.google.com.com', 168), (u'border?', 167), (u'terms!', 167), (u'en!', 167), (u'tag!', 166), (u'block!', 166), (u'/favicon.ico$$3000.629-18', 165), (u'clear!', 165), (u'/assets-0.housingcdn.com', 165), (u'pinterest!', 165)\n</code></pre>\n\n<p>I wanted to try 2 grams models but 3 days before the deadline, one of two fans in my Macbook broke, so I can never submit another submission before deadline again. But it was my false for depending too much on something and don't have a backup plan. </p>",
      "rawMarkdown": "Congratulation to winner. \r\n\r\nI ended up at 14th using liblinear and xgboost. I built three 1-gram models: 2 liblinear models with Bernoulli and multinominal feature weights  and 1 xgboost model with Bernoulli feature weights. On the 2nd level I blended those three model using xgboost and linear regression. Then at 3rd level, I used h mean to merge two models at 2nd level.\r\n\r\nAccording to me, preprocessing is the most important thing. I used a regular expression to split the html string to tokens. This regular expression can catch domain names, javascript function calls, html attributes, date, time... Some important features from my xgboost model are:\r\n\r\n    (u'/2015/07/0', 255), (u'2013%20', 251), (u'keywords!', 246), (u'2014%20fifa%20world%20cup', 244), (u'2015-07-07+21', 214), (u'blog!', 209), (u'/blog%206', 198), (u'other!', 195), (u'july!', 194), (u'copyright!', 192), (u'2015-07-14+00', 191), (u'rss!', 181), (u'14%', 180), (u'20%', 178), (u'web!', 177), (u'100%%', 176), (u'our!', 173), (u'news!', 173), (u'199.302', 172), (u'article!', 170), (u'author!', 170), (u'site_name.indexof(', 170), (u'30%', 168), (u'/plus.google.com.com', 168), (u'border?', 167), (u'terms!', 167), (u'en!', 167), (u'tag!', 166), (u'block!', 166), (u'/favicon.ico$$3000.629-18', 165), (u'clear!', 165), (u'/assets-0.housingcdn.com', 165), (u'pinterest!', 165)\r\n\r\nI wanted to try 2 grams models but 3 days before the deadline, one of two fans in my Macbook broke, so I can never submit another submission before deadline again. But it was my false for depending too much on something and don't have a backup plan.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 96172,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "10/15/2015 00:10:49",
      "content": "<p>I also want to thank my teammates. We have made it to our limit : )</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96173,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "10/15/2015 00:11:47",
      "content": "<p>Ok... that was unexpected XD. Great work guys.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96180,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "10/15/2015 00:38:05",
      "content": "<p>haha, thx. We are not that mad about it anymore!</p>\n\n<p>This was a tough competition (given the reset and the big size of the data) , but a very interesting problem. </p>\n\n<p>Well done to my teammates (Faron and Triskelion alphabetically) for their hard work. </p>\n\n<p>Also big credit to mortehu (that also soloed this) and bibaze for their great finishes. </p>\n\n<p>P.S. mortehu , I hope you don't hate us! Luck likes to play funny games, it happened to favor us this time, I hope next time to favor you :). Again well done on getting your master badge :)</p>\n\n<p>Big thank you to the organizers for their hard work and the Dato-graphlab guys for sponsoring this. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96181,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "10/15/2015 00:42:55",
      "content": "<p>@KazAnova, congratulations!</p>\n\n<p>Also please confirm that you will get Date Create Prize so we don't need to think about it :P</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96182,
      "author_name": "rushter",
      "author_url": "",
      "post_date": "10/15/2015 00:44:55",
      "content": "<p>Please share your golden features :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96183,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "10/15/2015 00:46:02",
      "content": "<p>@KazAnova or anyone else who got 0.98 with Random Forest, how did you do that? We could barely make 0.93 with multiple attempts.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96184,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "10/15/2015 00:47:53",
      "content": "<p>Our best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.</p>\n\n<p>There was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.</p>\n\n<p>We also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.</p>\n\n<p>some models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . </p>\n\n<p>I think we reached Level 4 meta-modelling in the end :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96185,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "10/15/2015 00:54:35",
      "content": "<p>Well, my hats off to David Shinn. He entered very late in the competition, eliminated the leaderboard very quickly and finished 10th, just dethroning us from our precious position. So I truly natively hate him, but also greatly, absolutely respect him for doing this on his own. Hats off to you sir!</p>\n\n<p>Other shouts go out to:\n- Andrew Bell: we competed a while back head to head, before you hit turbo mode and went out of my sight. Great effort.\n- bibaze team: excellent progress makers.\n- mortehu, orchid, student_2012, Glen: are you guys aliens or humans?</p>\n\n<p>and a thumbs up for NxGTR. My best model only got to 0.9803, so not even close to his 0.984. What he lacked were ensembling techniques and variety in models or data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96186,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "10/15/2015 00:55:09",
      "content": "<p>For us, there were no golden features. Our method was an ensemble of a stacked xgboost, keras, simple xgboost and ftrl. We used binary counts as the features for the last 2 models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96187,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "10/15/2015 00:57:30",
      "content": "<p>I attempted FTRL in this competition and got to 0.960 auc. Is that about what you achieved?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96189,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "10/15/2015 00:58:54",
      "content": "<p>We achieved 0.979 with FTRL, trained it for 5 epochs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96190,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "10/15/2015 01:09:15",
      "content": "<p>[quote=Subhajit Mandal;96183]</p>\n\n<p>@KazAnova or anyone else who got 0.98 with Random Forest, how did you do that? We could barely make 0.93 with multiple attempts.</p>\n\n<p>[/quote]\nAdd these count features to the benchmark code, you will get to at least 0.95 with Random Forest</p>\n\n<pre><code>    row['spaces'] = text.count(' ')\n    row['tabs'] = text.count('\\t')\n    row['braces'] = text.count('{')\n    row['brackets'] = text.count('[')\n    row['words'] = len(re.split('\\s+', text))\n    row['length'] = len(text)\n    row['bracket_2'] = text.count('&lt;')\n    row['bracket_3'] = text.count('(')\n    row['re_w'] = len(re.split('\\w+', text))\n    row['new_par'] = text.count('\\n\\t')\n    row['dbslash'] = text.count('//')\n    row['bslash'] = text.count('/')\n    row['bol_0'] = text.count('&lt;!')\n    row['bol_1'] = text.count('/&gt;')\n    row['bol_2'] = text.count('&amp;')\n    row['bol_3'] = text.count(';')\n    row['bol_4'] = text.count('==')\n    row['bol_5'] = text.count('===')\n    row['bol_6'] = text.count('css')\n    row['bol_7'] = text.count('#')\n    row['bol_8'] = text.count('@')\n    row['bol_9'] = text.count('$')\n    row['bol_10'] = text.count('%')\n    row['bol_11'] = text.count('^')\n    row['bol_12'] = text.count('+')\n    row['bol_13'] = text.count('?')\n    row['bol_14'] = text.count('|')\n    row['bol_15'] = text.count('\\\\')\n    row['bol_16'] = text.count('*')\n    row['bol_17'] = text.count('||')\n    row['bol_18'] = text.count('\\t\\t')\n    row['bol_19'] = text.count('\\t\\t\\t')\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96191,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "10/15/2015 01:11:21",
      "content": "<p>Horrifying. What were the features? I trained FTRL on 4-5 epochs on 1800 raw features, reduced from 960k.  My features were the most prevalent 1800 word features with optimized L1/L2 regularization.</p>\n\n<p>I went fruity-loopy in this competition and made a custom C parser, a bigram/trigram processor in python and in the end I was processing pretty much portions of matrices to get things going (scipy.sparse.hstack). All of this to reduce the memory requirements for processing the data and decreasing the time before training and model verification would occur.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96192,
      "author_name": "rushter",
      "author_url": "",
      "post_date": "10/15/2015 01:11:25",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;96184]</p>\n\n<p>Our best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.</p>\n\n<p>There was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.</p>\n\n<p>We also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.</p>\n\n<p>some models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . </p>\n\n<p>I think we reached Level 4 meta-modelling in the end :)\n[/quote]\nHow much RAM it requires? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96193,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "10/15/2015 01:13:12",
      "content": "<p>For us it was like this:</p>\n\n<p>Xgboost (single): 0.98422</p>\n\n<p>Xgboost (dual ensemble): 0.98459</p>\n\n<p>RandomForest (single): 0.97455</p>\n\n<p>ExtraTrees (single): 0.97628</p>\n\n<p>Done :D, our final submission is an ensemble of pretty much those 4. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96195,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "10/15/2015 01:17:34",
      "content": "<p>[quote=Artem;96192]</p>\n\n<p>How much RAM it requires? </p>\n\n<p>[/quote]</p>\n\n<p>A lot!</p>\n\n<p>128GB should be ok!</p>\n\n<p>But to be honest, we were reckless (as we had 256GB available). I cannot tell you what is the most optimum you can run it on. 64GB is possibly enough too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96196,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "10/15/2015 01:18:07",
      "content": "<p>[quote=Gerard Toonstra;96191]</p>\n\n<p>Horrifying. What were the features? I trained FTRL on 4-5 epochs on 1800 raw features, reduced from 960k.  My features were the most prevalent 1800 word features with optimized L1/L2 regularization.</p>\n\n<p>I went fruity-loopy in this competition and made a custom C parser, a bigram/trigram processor in python and in the end I was processing pretty much portions of matrices to get things going (scipy.sparse.hstack). All of this to reduce the memory requirements for processing the data and decreasing the time before training and model verification would occur.</p>\n\n<p>[/quote]</p>\n\n<p>For FTRL, we used simple binary counts (is it there or not?) of unigrams and bigrams as features. We tried many more feature engineering techniques, but the simple was the best.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96197,
      "author_name": "sjuvekar",
      "author_url": "",
      "post_date": "10/15/2015 01:18:39",
      "content": "<p>[quote=Subhajit Mandal;96189]</p>\n\n<p>We achieved 0.979 with FTRL, trained it for 5 epochs.</p>\n\n<p>[/quote]</p>\n\n<p>Along with FTRL, we had following scores with other models:</p>\n\n<p>xgboost: 0.983</p>\n\n<p>stacked xgboost: 0.985</p>\n\n<p>MLP with 3 layerS: 0.952</p>\n\n<p>I agree with @rcarson and @NxGTR, bigger teams and heavy ensembing is the way to do well in these competitions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96198,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "10/15/2015 01:21:46",
      "content": "<p>[quote=Little Boat;96190]</p>\n\n<p>[quote=Subhajit Mandal;96183]</p>\n\n<p>@KazAnova or anyone else who got 0.98 with Random Forest, how did you do that? We could barely make 0.93 with multiple attempts.</p>\n\n<p>[/quote]\nAdd these count features to the benchmark code, you will get to at least 0.95 with Random Forest</p>\n\n<pre><code>    row['spaces'] = text.count(' ')\n    row['tabs'] = text.count('\\t')\n    row['braces'] = text.count('{')\n    row['brackets'] = text.count('[')\n    row['words'] = len(re.split('\\s+', text))\n    row['length'] = len(text)\n    row['bracket_2'] = text.count('&lt;')\n    row['bracket_3'] = text.count('(')\n    row['re_w'] = len(re.split('\\w+', text))\n    row['new_par'] = text.count('\\n\\t')\n    row['dbslash'] = text.count('//')\n    row['bslash'] = text.count('/')\n    row['bol_0'] = text.count('&lt;!')\n    row['bol_1'] = text.count('/&gt;')\n    row['bol_2'] = text.count('&amp;')\n    row['bol_3'] = text.count(';')\n    row['bol_4'] = text.count('==')\n    row['bol_5'] = text.count('===')\n    row['bol_6'] = text.count('css')\n    row['bol_7'] = text.count('#')\n    row['bol_8'] = text.count('@')\n    row['bol_9'] = text.count('$')\n    row['bol_10'] = text.count('%')\n    row['bol_11'] = text.count('^')\n    row['bol_12'] = text.count('+')\n    row['bol_13'] = text.count('?')\n    row['bol_14'] = text.count('|')\n    row['bol_15'] = text.count('\\\\')\n    row['bol_16'] = text.count('*')\n    row['bol_17'] = text.count('||')\n    row['bol_18'] = text.count('\\t\\t')\n    row['bol_19'] = text.count('\\t\\t\\t')\n</code></pre>\n\n<p>[/quote]</p>\n\n<p>I did not see that coming ;-) should not have removed the punctuation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96199,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "10/15/2015 01:23:31",
      "content": "<p>[quote=Artem;96192]</p>\n\n<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;96184]</p>\n\n<p>Our best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.</p>\n\n<p>There was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.</p>\n\n<p>We also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.</p>\n\n<p>some models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . </p>\n\n<p>I think we reached Level 4 meta-modelling in the end :)\n[/quote]\nHow much RAM it requires? </p>\n\n<p>[/quote]</p>\n\n<p>I managed to run a large 500k single word + 232k bigram + 178k trigram model in 34G space (close to 0.9/1M features). Basically externalizing the entire tf-idf vectorizer. Then you notice how little those classes contribute vs. the cost and how much easier it would be if you just take 2 passes at 2-3x the time... But if memory is lacking...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96200,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/15/2015 01:24:13",
      "content": "<p>[quote=Gerard Toonstra;96185]</p>\n\n<p>Well, my hats off to David Shinn. He entered very late in the competition, eliminated the leaderboard very quickly and finished 10th, just dethroning us from our precious position. So I truly natively hate him, but also greatly, absolutely respect him for doing this on his own. Hats off to you sir!</p>\n\n<p>[/quote]</p>\n\n<p>Thank @Gerard and @Redo, your constant improvement until the end, even a little, kept me motivated to keep pushing till the end.  I congratulate all of the top winners, I fight tooth and nail just to break top 10.  This was a crazy exhausting and crazy fun competition.  Thanks to orchid, Bang Nguyen, Ashtericz, Redo, and NxGTR &amp; Sky for making just cracking top 10 a nail biting experience.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96201,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "10/15/2015 01:25:32",
      "content": "<p>[quote=Subhajit Mandal;96196]</p>\n\n<p>[quote=Gerard Toonstra;96191]</p>\n\n<p>Horrifying. What were the features? I trained FTRL on 4-5 epochs on 1800 raw features, reduced from 960k.  My features were the most prevalent 1800 word features with optimized L1/L2 regularization.</p>\n\n<p>I went fruity-loopy in this competition and made a custom C parser, a bigram/trigram processor in python and in the end I was processing pretty much portions of matrices to get things going (scipy.sparse.hstack). All of this to reduce the memory requirements for processing the data and decreasing the time before training and model verification would occur.</p>\n\n<p>[/quote]</p>\n\n<p>For FTRL, we used simple binary counts (is it there or not?) of unigrams and bigrams as features. We tried many more feature engineering techniques, but the simple was the best.</p>\n\n<p>[/quote]</p>\n\n<p>Ah, I only included the positive presence 1, but not the 0 -presence... This being FTRL, it's basically unrolling your matrix to disk...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96202,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/15/2015 01:45:03",
      "content": "<p>[quote=Gerard Toonstra;96185]</p>\n\n<p>What we lacked were ensembling techniques and variety in models or data.</p>\n\n<p>[/quote]</p>\n\n<p>Well, if you had done what I did, you sure would have bested me easily.  My single best model only barely reaches 0.980 on the Private LB.  What I lacked in a single model I made up for in my Frankenstein stacked model.  I just counted, I had 80 models as features in my second layer (plus around 20 or so other basic features).  I used Random Forest, Extra Trees, XGBoost, Vowpal Wabbit, KNN, SVM, and logistic regression in my first set of models.  Each of those trained on different subset of features: character counts, html tags, domain urls, html attributes, file extensions, and 123gram tfidf and 123gram counts.  I even tried some crazy things like the character and html tag counts from the first half of the file and the second half of the file as different training feature sets (sort of a data augmentation).  In my second layer, I used a boosted random forest and xgboost models.  My third layer was a hyperopt optimized power and linear transformation of those 2 predictions from the second layer.  This was my first stab at serious stacking, so it was a great learning experience.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96205,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "10/15/2015 02:36:56",
      "content": "<p>I didn't do nearly as well as you guys, but almost all of my models were just variants of TINRTGU's code in pure Python. I just greedily averaged these FTRLP with different orders of files, different one hot splitting techniques, etc. I wish I had time for better models, but I ran out of time in terms of creating meta features to be fed into XG.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96206,
      "author_name": "mortehu",
      "author_url": "",
      "post_date": "10/15/2015 02:38:39",
      "content": "<p>@Mad Professors: Congratulations, and good job!  Don't worry, I knew our positions could easily reverse.  It was fun to have a real shot at the top position for a while.</p>\n\n<p>My submission was just the scaled sum of two models, a linear SVM of skip-grams and element/attribute combinations (79 million unique features, ~45 minutes training time), and an XGBoost model using substrings as features (30,000 features, 2000 trees, ~5 hours training time).</p>\n\n<p>The substrings were automatically selected using this tool: <a href=\"https://github.com/mortehu/substring-frequencies\">https://github.com/mortehu/substring-frequencies</a></p>\n\n<p>These are the substrings used as features in the XGBoost model, with the most predictive being listed first: <a href=\"https://gist.github.com/mortehu/07f0c59dc30bf495853d\">https://gist.github.com/mortehu/07f0c59dc30bf495853d</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96209,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "10/15/2015 03:02:50",
      "content": "<p>Our feature set is huge, so it is painful to try stacking or other complicated stuff. But if I don't  go into the Physics one, I will probably have time to finish tfidf and two stage stacking. I was too optimistic to think we can perform well on two competitions at the same time. With all the other works, I just don't have enough energy and computers to try lots of things in two competitions.</p>\n\n<p>[quote=NxGTR;96193]</p>\n\n<p>For us it was like this:</p>\n\n<p>Xgboost (single): 0.98422</p>\n\n<p>Xgboost (dual ensemble): 0.98459</p>\n\n<p>RandomForest (single): 0.97455</p>\n\n<p>ExtraTrees (single): 0.97628</p>\n\n<p>Done :D, our final submission is an ensemble of pretty much those 4. </p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96210,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "10/15/2015 03:08:45",
      "content": "<p>You <strong>have</strong> performed well on both of the competitions. In Physics competition it was a really bad luck for you that high scoring scripts got posted at the last moment. Luckily, there was no scripts in this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96213,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "10/15/2015 03:15:11",
      "content": "<p>It is amazing you can get such good results with a simple ensemble of two models. \nBut I guess your two models are probably very strong by themselves. </p>\n\n<p>May I ask how are you able to train a linear SVM with 79 million features in 45 minutes,  I always find svm to be very slow. Another thing is did you use some software package to extract skip gram features?</p>\n\n<p>[quote=mortehu;96206]</p>\n\n<p>@Mad Professors: Congratulations, and good job!  Don't worry, I knew our positions could easily reverse.  It was fun to have a real shot at the top position for a while.</p>\n\n<p>My submission was just the scaled sum of two models, a linear SVM of skip-grams and element/attribute combinations (79 million unique features, ~45 minutes training time), and an XGBoost model using substrings as features (30,000 features, 2000 trees, ~5 hours training time).</p>\n\n<p>The substrings were automatically selected using this tool: <a href=\"https://github.com/mortehu/substring-frequencies\">https://github.com/mortehu/substring-frequencies</a></p>\n\n<p>These are the substrings used as features in the XGBoost model, with the most predictive being listed first: <a href=\"https://gist.github.com/mortehu/07f0c59dc30bf495853d\">https://gist.github.com/mortehu/07f0c59dc30bf495853d</a></p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96214,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "10/15/2015 03:15:20",
      "content": "<p>[quote=SkyLibrary;96209]\nI was too optimistic to think we can perform well on two competitions at the same time...\n[/quote]</p>\n\n<p>Hehehe, Top Kagglers issues: when two &quot;top 13th&quot; is bad! :P</p>\n\n<p>It was OK, could be better but heck!, pretty much always can be better :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96217,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "10/15/2015 03:26:12",
      "content": "<p>[quote=SkyLibrary;96213]</p>\n\n<p>May I ask how are you able to train a linear SVM with 79 million features in 45 minutes,  I always find svm to be very slow. Another thing is did you use some software package to extract skip gram features?</p>\n\n<p>[/quote]</p>\n\n<p>mortehu is an amazing programmer according to his repo. He probably just coded it from scratch in c++</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96218,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "10/15/2015 03:27:29",
      "content": "<p>lol. Two 13th is definitely better than two 26th. But wait a minutes, does 13 suppose to be a unlucky number in some countries, luckily I don't come from these countries so I don't have hard feelings for it.\n[quote=NxGTR;96214]</p>\n\n<p>[quote=SkyLibrary;96209]\nI was too optimistic to think we can perform well on two competitions at the same time...\n[/quote]</p>\n\n<p>Hehehe, Top Kagglers issues: when two &quot;top 13th&quot; is bad! :P</p>\n\n<p>It was OK, could be better but heck!, pretty much always can be better :D</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96220,
      "author_name": "sudalairajkumar",
      "author_url": "",
      "post_date": "10/15/2015 03:45:28",
      "content": "<p>Congrats Mad Professors, mortehu and bibaze.! Special thanks to @David shinn for his excellent starter code which helped us enter this competition since we started very late. Thanks everyone for posting their approaches. </p>\n\n<p>@Marios: 4-level meta models.! You are pushing the boundaries further. Looking forward to your team's code soon. \n@rcarson: One more top 10 finish to your list. You pretty much get there almost every time :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96222,
      "author_name": "fernandoprocy",
      "author_url": "",
      "post_date": "10/15/2015 03:57:02",
      "content": "<p>[quote=mortehu;96206]</p>\n\n<p>@Mad Professors: Congratulations, and good job!  Don't worry, I knew our positions could easily reverse.  It was fun to have a real shot at the top position for a while.</p>\n\n<p>My submission was just the scaled sum of two models, a linear SVM of skip-grams and element/attribute combinations (79 million unique features, ~45 minutes training time), and an XGBoost model using substrings as features (30,000 features, 2000 trees, ~5 hours training time).</p>\n\n<p>The substrings were automatically selected using this tool: <a href=\"https://github.com/mortehu/substring-frequencies\">https://github.com/mortehu/substring-frequencies</a></p>\n\n<p>These are the substrings used as features in the XGBoost model, with the most predictive being listed first: <a href=\"https://gist.github.com/mortehu/07f0c59dc30bf495853d\">https://gist.github.com/mortehu/07f0c59dc30bf495853d</a></p>\n\n<p>[/quote]\nThese codes are awesome.. I wish I can code anywhere near that level someday.</p>\n\n<p>Something interesting I noticed in the substrings list is that there are many &quot;dates&quot; strings that coincide with the leak found mid competition.. I hadn't figured out but perhaps most of the pages have a 'current time' recorded in the html file. Maybe we could call that a 'nerfed' version of the leak. Clueless about what that '1529' would mean, though. edit: maybe week number?</p>\n\n<p>Also, anybody has an idea as to why these sponsored ads seems so predictable through features even though it doesn't have any apparent correlation? After all, these sponsored pages are just pages people paid StumbleUpon to feature in their system. Never thought background features like number of tags or chars could be so effective at predicting that.</p>\n\n<p>Congrats to everyone and I'm specially thankful to David Shinn, whose code got me started in the competition as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96223,
      "author_name": "mortehu",
      "author_url": "",
      "post_date": "10/15/2015 04:01:39",
      "content": "<p>As far as I can tell, this competition was more about determining the date each page was downloaded, rather than whether or not something was advertising.  Negative examples were mostly downloaded towards the end, for example.  Current events, like &quot;Pluto&quot; (New Horizons passed Pluto on July 24th), &quot;Donald Trump&quot;, and &quot;El Chapo&quot; were all strong features, with &quot;infographic&quot; appearing pretty far down on the list.  Some of the other number strings seen in my feature lists are versions of JavaScript libraries that were released after the positive samples had been downloaded.</p>\n\n<p>I'll post the source code for my linear SVM training and HTML tokenization program tomorrow.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96239,
      "author_name": "gardenia22",
      "author_url": "",
      "post_date": "10/15/2015 07:38:53",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;96184]</p>\n\n<p>Our best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.</p>\n\n<p>There was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.</p>\n\n<p>We also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.</p>\n\n<p>some models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . </p>\n\n<p>I think we reached Level 4 meta-modelling in the end :)</p>\n\n<p>[/quote]\nI am curious about your 3-grams tf-idf, do you include the tags and scripts of HTML files? Seems that the tags and scripts are only helpful to classify HTML files from the same websites(they share the same HTML structure.) I would be appreciate if you can share more about features and how you clean HTML files, thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96249,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "10/15/2015 09:13:53",
      "content": "<p>One doubt I had... I started with tf-idf on xgb, but noticed that IDF itself had no impact on end results. Also, the official application of term frequency seemed to slightly reduce the results as well, probably because the size of the documents varies quite a bit. So my XGB would perform 0.96/0.97 on TF calculated as : f(t) / n(t,d)  , whereas a simple f(t) (just a basic count) would get me 0.98 performance. Can you guys confirm you got better results with simple counts as well?</p>\n\n<p>My max_depth for xgb was optimized at 50 with mostly default settings, except colsample_by_tree set to 0.46. subsampling made no difference and neither did gamma. &quot;scale_pos_weight&quot; sounds like a great opportunity to tweak, because 1's were less than 0's and it seemed to work in the first few iterations, but then it leveled off very quickly.</p>\n\n<p>I also did a full minhash doc-doc comparison on 150 hashes. I calculated the hashes in python, but wrote a custom C app to do the comparison. That got me 0.89 on a 20% similarity threshold setting. A combination of 80/50/20% similarity results got me 0.929.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96252,
      "author_name": "phanisrikanth",
      "author_url": "",
      "post_date": "10/15/2015 09:44:04",
      "content": "<p>[quote=SRK;96220]</p>\n\n<p>Congrats Mad Professors, mortehu and bibaze.! Special thanks to @David shinn for his excellent starter code which helped us enter this competition since we started very late. Thanks everyone for posting their approaches. </p>\n\n<p>@Marios: 4-level meta models.! You are pushing the boundaries further. Looking forward to your team's code soon. \n@rcarson: One more top 10 finish to your list. You pretty much get there almost every time :)</p>\n\n<p>[/quote]</p>\n\n<p>Congratulations Mad Professors - awesome finish @mortehu - your repo looks intimidating!\nThanks to @DavidShinn for the starter code without which I wouldn't have even thought of attempting this daunting challenge.</p>\n\n<p>@SRK, special thanks to U for considering to team up with a Kaggler, despite being a Master yourself! </p>\n\n<p>@Everybody else, SRK is an awesome teammate. <em>Endorsing him</em></p>\n\n<p>We got late into the competition (~17 days to go), ignored the competition for a few more days and actually worked only for the last 3 days to be very honest and ending up at #33 is no mean feat. Very happy to finish well and it could only happen with having a teammate like SRK who pretty much guided me throughout the last 3 days where we tried what we could! :-)</p>\n\n<p>I have learnt a lot of things and continue to learn more and do well :-)</p>\n\n<p>@rcarson - got to know a lot about you. SRK spoke very highly of you. I hope I team up with you someday. For that, I'll work hard, earn the right and ask U :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96273,
      "author_name": "",
      "author_url": "",
      "post_date": "10/15/2015 12:47:51",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;96184]</p>\n\n<p>Our best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.</p>\n\n<p>There was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.</p>\n\n<p>We also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.</p>\n\n<p>some models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . </p>\n\n<p>I think we reached Level 4 meta-modelling in the end :)</p>\n\n<p>[/quote]</p>\n\n<p>As a n00b, can I ask how exactly did you preprocess those files? Or in general, how should I deal with such noisy data in the future, if I'd like to extract some usefull TF-IFDs?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96277,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "10/15/2015 13:02:27",
      "content": "<p>Actually,  I just loaded each file as a big string (lets call it mline)  and then cleaned with this:</p>\n\n<p>import re</p>\n\n<p>mline=re.sub(&quot;[^a-zA-Z0-9]&quot;,&quot; &quot;, mline) # remove non-alphanumeric</p>\n\n<p>then I re-Wrote this to a file. </p>\n\n<p>The test is up to the to the tf-idf parameters!</p>\n\n<p>tfv=TfidfVectorizer(min_df=2, max_features=None, strip_accents='unicode',lowercase =True,\n                        analyzer='word', dtype=np.float64, token_pattern=r'\\w{2,}', ngram_range=(1, 3), use_idf=True,smooth_idf=True, \n    sublinear_tf=True, stop_words = 'english')   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96282,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "10/15/2015 14:36:58",
      "content": "<p>Thanks Dato, Kaggle, team mates and all competitors! Grats all winners and new Kaggle Masters.</p>\n\n<p>My focus was on online learning and generating non-alphanumeric features.</p>\n\n<p>I could not extract the entire dataset on my HD, so I worked with the compressed archives only, unzipping and streaming the documents into the algo's VW, FTRL, SGD and perceptron. The scores were not as high as I've seen reported in this thread (well done!), but I think the models/vectorizers were diverse enough to contribute to the ensemble.</p>\n\n<p>We also pursued three crazy ideas that happened to work (a tiny bit, but enough to be exciting for further research):</p>\n\n<ul>\n<li><p>Training a model on a sentiment analysis data set unrelated to this competition, then using the predictions for the train and test sets for further modeling. Such a &quot;sentiment&quot;-feature was informative.</p></li>\n<li><p>Training a model on the 20 newsgroups data set and creating multi-class predictions for train and test sets. I think these vectors added a form of categorization/topic modeling. (&quot;this document talks about 'SciMed' and 'CompSci', but not 'Religion'&quot;).</p></li>\n<li><p>Using Normalized Compression Distance with the fast Snappy compressor to calculate distance between every document and 4 randomly created anchor corpora.</p></li>\n</ul>\n\n<p>Again, learned a lot, for instance about geometric mean for model averaging, feature and meta-model interactions with <a href=\"https://github.com/far0n/xgbfi\">XGBfi</a>, proper competition management/iteration, patience, and stacking. I am particularly excited about our 4 levels of stacking, with 4th level models using the predictions from the 1st-3rd level models (fully connected stacknet?).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96289,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "10/15/2015 15:19:15",
      "content": "<p>[quote=Triskelion;96282]\nAgain, learned a lot, for instance about geometric mean for model averaging, feature and meta-model interactions with XGBfi.\n[/quote]</p>\n\n<p>That looks very interesting Thanks!</p>\n\n<p>I gave it a quick read, but, is that Windows-Only (XGBfi)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96297,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "10/15/2015 16:01:11",
      "content": "<p>I think so, though maybe you can run on Linux Mono. Far0n is considering porting to C++.</p>\n\n<p>I think the nicest would be to have it added to XGBoost source. It's really useful. Adding high-ranking interactions as features for a linear algorithm boosts its score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96301,
      "author_name": "mortehu",
      "author_url": "",
      "post_date": "10/15/2015 16:07:43",
      "content": "<p>Here's the text-classifier I used for this competition released as open source: <a href=\"https://github.com/mortehu/text-classifier\">https://github.com/mortehu/text-classifier</a></p>\n\n<p>I hope this can be used to set a high baseline in any future text classification competition, or to be more generally useful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96307,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "10/15/2015 16:17:40",
      "content": "<p>So who used graphlab?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96324,
      "author_name": "mmueller",
      "author_url": "",
      "post_date": "10/15/2015 17:39:19",
      "content": "<p>[quote=NxGTR;96289]</p>\n\n<p>[quote=Triskelion;96282]\nAgain, learned a lot, for instance about geometric mean for model averaging, feature and meta-model interactions with XGBfi.\n[/quote]</p>\n\n<p>That looks very interesting Thanks!</p>\n\n<p>I gave it a quick read, but, is that Windows-Only (XGBfi)?</p>\n\n<p>[/quote]</p>\n\n<p>So I just checked and it runs fine under linux using mono. </p>\n\n<p>I know .Net is a bit odd, but I'm planning to build a GUI interfacing XGB with features like training graphs, comparing trainings graphs of different parameter sets, querying feature split values, feature importance based on hold out data performance (which I expect to be useful to detect features which rank high but hurt performance) and so on. But I'm only familar with WPF for creating GUIs. That is why xgbfi is currently written in .Net.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96325,
      "author_name": "mmueller",
      "author_url": "",
      "post_date": "10/15/2015 17:41:04",
      "content": "<p>[quote=mortehu;96301]</p>\n\n<p>Here's the text-classifier I used for this competition released as open source: <a href=\"https://github.com/mortehu/text-classifier\">https://github.com/mortehu/text-classifier</a></p>\n\n<p>I hope this can be used to set a high baseline in any future text classification competition, or to be more generally useful.</p>\n\n<p>[/quote]</p>\n\n<p>just wow!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96369,
      "author_name": "zachmayer",
      "author_url": "",
      "post_date": "10/15/2015 22:25:35",
      "content": "<p>Dumb question, what is &quot;FTRL&quot;?</p>\n\n<p>Also, what 2 files did you put into this tool to select substrings? <a href=\"https://github.com/mortehu/substring-frequencies\">https://github.com/mortehu/substring-frequencies</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96374,
      "author_name": "mortehu",
      "author_url": "",
      "post_date": "10/15/2015 22:34:30",
      "content": "<p>[quote=Zach;96369]\nAlso, what 2 files did you put into this tool to select substrings? <a href=\"https://github.com/mortehu/substring-frequencies\">https://github.com/mortehu/substring-frequencies</a>\n[/quote]</p>\n\n<p>I picked a random sample of 512 MB of positive examples, and 512 MB of negative examples.  I used the <code>--document</code> mode.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96376,
      "author_name": "mmueller",
      "author_url": "",
      "post_date": "10/15/2015 22:39:30",
      "content": "<p>[quote=Zach;96369]</p>\n\n<p>Dumb question, what is &quot;FTRL&quot;?</p>\n\n<p>[/quote]</p>\n\n<p>FTRL stands for the online learning algorithm &quot;Follow The Regularized Leader&quot;.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96387,
      "author_name": "tmjiang",
      "author_url": "",
      "post_date": "10/16/2015 00:42:31",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mortehu\">@mortehu</a>,</p>\n\n<p>I'm intrigued and really admiring your way. Just curious, would it be possible to have your substring frequency more (or less) accurate to instruct the suffix array to discount overlapped things? (In case my description sounds vague and weird or this method isn't popular... it's kinda trick to eliminate stop words without knowing then and sometimes emphasize critical information. For example from the earliest paper I can find, when encountered a bigram &quot;rail enquiries&quot; in BNC, it's actually useless because it was always part of either &quot;national rail enquiries&quot; or &quot;British british rail enquiries.&quot;)</p>\n\n<p>I didn't use it just because I didn't get to that phase yet. :p</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96393,
      "author_name": "mortehu",
      "author_url": "",
      "post_date": "10/16/2015 02:54:47",
      "content": "<p>@Barabbas, even though it might not look like it, a lot of redundant substrings are already removed. However, the code currently keeps overlapping strings if they have different occurrence counts, even if that is stupid in most cases. It's hard to know without building a model whether a substring of another feature is truly redundant, so I didn't even try. You can see the duplicate detection code on line 249 in substrings.cc.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96400,
      "author_name": "tmjiang",
      "author_url": "",
      "post_date": "10/16/2015 05:44:16",
      "content": "<p><a href=\"https://www.kaggle.com/mortehu\">@mortehu</a>: I see, thank you. :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96414,
      "author_name": "zachmayer",
      "author_url": "",
      "post_date": "10/16/2015 12:57:25",
      "content": "<p>[quote=Faron;96376]</p>\n\n<p>FTRL stands for the online learning algorithm &quot;Follow The Regularized Leader&quot;.</p>\n\n<p>[/quote]</p>\n\n<p>How does that differ from other online algorithms, e.g. vowpal wabbit?  Where can I get the software to try it out?</p>\n\n<p>Actually it looks like VW has a <a href=\"https://github.com/JohnLangford/vowpal_wabbit/wiki/Command-line-arguments\">--ftrl option</a>.  I'll have to try that out sometime.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96418,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "10/16/2015 13:05:13",
      "content": "<p>[quote=Zach;96414]</p>\n\n<p>How does that differ from other online algorithms, e.g. vowpal wabbit?  Where can I get the software to try it out?</p>\n\n<p>[/quote]</p>\n\n<p>For me it is like an SGD, with the difference that in each step the coefficients are being copied and converted to &quot;new values&quot; based on the regularization terms. </p>\n\n<p>For l2 regul, this should not have much difference with sgd, but for l1 it makes quite a difference  because the coefficients will always remain zero (0.0) if they never pass the the C value (something not really achievable with simple sgd)  and will never be copied across . </p>\n\n<p>In other words, only if something becomes significant (after the updates's phase) will start being copied to further boost itself (hence the follow the leading features after regularization is applied!) </p>\n\n<p>Also, most of the times it uses passed gradients as a form of adaptive learning rate so generally as an algorithm is more independent. </p>\n\n<p>At least, that is my interpretation!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96425,
      "author_name": "grigorydymov",
      "author_url": "",
      "post_date": "10/16/2015 14:52:25",
      "content": "<p>Hi! I want to thank David Shinn for his very nice script, it motivated me to continue working on the comp, since it was so easy to improve :) </p>\n\n<p>After dabbling a while with it and improving it manually, I calculated how much popular (top50k) individual words separated the classes by abs(freq_pos - freq_neg), and took the top3k and used them in RF and XGB as features with no tuning of parameters, then simple average. I did not expect that simple solution to earn us a 19th place. It was a really fun competition due to the size of data!</p>\n\n<p>Also thanks to my teammate</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96614,
      "author_name": "",
      "author_url": "",
      "post_date": "10/19/2015 05:29:08",
      "content": "<p>Congratulation to winner. </p>\n\n<p>I ended up at 14th using liblinear and xgboost. I built three 1-gram models: 2 liblinear models with Bernoulli and multinominal feature weights  and 1 xgboost model with Bernoulli feature weights. On the 2nd level I blended those three model using xgboost and linear regression. Then at 3rd level, I used h mean to merge two models at 2nd level.</p>\n\n<p>According to me, preprocessing is the most important thing. I used a regular expression to split the html string to tokens. This regular expression can catch domain names, javascript function calls, html attributes, date, time... Some important features from my xgboost model are:</p>\n\n<pre><code>(u'/2015/07/0', 255), (u'2013%20', 251), (u'keywords!', 246), (u'2014%20fifa%20world%20cup', 244), (u'2015-07-07+21', 214), (u'blog!', 209), (u'/blog%206', 198), (u'other!', 195), (u'july!', 194), (u'copyright!', 192), (u'2015-07-14+00', 191), (u'rss!', 181), (u'14%', 180), (u'20%', 178), (u'web!', 177), (u'100%%', 176), (u'our!', 173), (u'news!', 173), (u'199.302', 172), (u'article!', 170), (u'author!', 170), (u'site_name.indexof(', 170), (u'30%', 168), (u'/plus.google.com.com', 168), (u'border?', 167), (u'terms!', 167), (u'en!', 167), (u'tag!', 166), (u'block!', 166), (u'/favicon.ico$$3000.629-18', 165), (u'clear!', 165), (u'/assets-0.housingcdn.com', 165), (u'pinterest!', 165)\n</code></pre>\n\n<p>I wanted to try 2 grams models but 3 days before the deadline, one of two fans in my Macbook broke, so I can never submit another submission before deadline again. But it was my false for depending too much on something and don't have a backup plan. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "96171": "Congratulations, Mad professors :P",
    "96172": "I also want to thank my teammates. We have made it to our limit : )",
    "96173": "Ok... that was unexpected XD. Great work guys.",
    "96180": "haha, thx. We are not that mad about it anymore!\r\n\r\nThis was a tough competition (given the reset and the big size of the data) , but a very interesting problem. \r\n\r\nWell done to my teammates (Faron and Triskelion alphabetically) for their hard work. \r\n\r\nAlso big credit to mortehu (that also soloed this) and bibaze for their great finishes. \r\n\r\n\r\nP.S. mortehu , I hope you don't hate us! Luck likes to play funny games, it happened to favor us this time, I hope next time to favor you :). Again well done on getting your master badge :)\r\n\r\nBig thank you to the organizers for their hard work and the Dato-graphlab guys for sponsoring this.",
    "96181": "KazAnova, congratulations!\r\n\r\nAlso please confirm that you will get Date Create Prize so we don't need to think about it :P",
    "96182": "Please share your golden features :)",
    "96183": "KazAnova or anyone else who got 0.98 with Random Forest, how did you do that? We could barely make 0.93 with multiple attempts.",
    "96184": "Our best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.\r\n\r\nThere was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.\r\n\r\nWe also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.\r\n\r\nsome models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . \r\n\r\nI think we reached Level 4 meta-modelling in the end :)",
    "96185": "Well, my hats off to David Shinn. He entered very late in the competition, eliminated the leaderboard very quickly and finished 10th, just dethroning us from our precious position. So I truly natively hate him, but also greatly, absolutely respect him for doing this on his own. Hats off to you sir!\r\n\r\nOther shouts go out to:\r\n- Andrew Bell: we competed a while back head to head, before you hit turbo mode and went out of my sight. Great effort.\r\n- bibaze team: excellent progress makers.\r\n- mortehu, orchid, student_2012, Glen: are you guys aliens or humans?\r\n\r\nand a thumbs up for NxGTR. My best model only got to 0.9803, so not even close to his 0.984. What he lacked were ensembling techniques and variety in models or data.",
    "96186": "For us, there were no golden features. Our method was an ensemble of a stacked xgboost, keras, simple xgboost and ftrl. We used binary counts as the features for the last 2 models.",
    "96187": "I attempted FTRL in this competition and got to 0.960 auc. Is that about what you achieved?",
    "96189": "We achieved 0.979 with FTRL, trained it for 5 epochs.",
    "96190": "[quote=Subhajit Mandal;96183]\r\n\r\n@KazAnova or anyone else who got 0.98 with Random Forest, how did you do that? We could barely make 0.93 with multiple attempts.\r\n\r\n[/quote]\r\nAdd these count features to the benchmark code, you will get to at least 0.95 with Random Forest\r\n\r\n        row['spaces'] = text.count(' ')\r\n        row['tabs'] = text.count('\\t')\r\n        row['braces'] = text.count('{')\r\n        row['brackets'] = text.count('[')\r\n        row['words'] = len(re.split('\\s+', text))\r\n        row['length'] = len(text)\r\n        row['bracket_2'] = text.count('<')\r\n        row['bracket_3'] = text.count('(')\r\n        row['re_w'] = len(re.split('\\w+', text))\r\n        row['new_par'] = text.count('\\n\\t')\r\n        row['dbslash'] = text.count('//')\r\n        row['bslash'] = text.count('/')\r\n        row['bol_0'] = text.count('<!')\r\n        row['bol_1'] = text.count('/>')\r\n        row['bol_2'] = text.count('&')\r\n        row['bol_3'] = text.count(';')\r\n        row['bol_4'] = text.count('==')\r\n        row['bol_5'] = text.count('===')\r\n        row['bol_6'] = text.count('css')\r\n        row['bol_7'] = text.count('#')\r\n        row['bol_8'] = text.count('@')\r\n        row['bol_9'] = text.count('$')\r\n        row['bol_10'] = text.count('%')\r\n        row['bol_11'] = text.count('^')\r\n        row['bol_12'] = text.count('+')\r\n        row['bol_13'] = text.count('?')\r\n        row['bol_14'] = text.count('|')\r\n        row['bol_15'] = text.count('\\\\')\r\n        row['bol_16'] = text.count('*')\r\n        row['bol_17'] = text.count('||')\r\n        row['bol_18'] = text.count('\\t\\t')\r\n        row['bol_19'] = text.count('\\t\\t\\t')",
    "96191": "Horrifying. What were the features? I trained FTRL on 4-5 epochs on 1800 raw features, reduced from 960k.  My features were the most prevalent 1800 word features with optimized L1/L2 regularization.\r\n\r\nI went fruity-loopy in this competition and made a custom C parser, a bigram/trigram processor in python and in the end I was processing pretty much portions of matrices to get things going (scipy.sparse.hstack). All of this to reduce the memory requirements for processing the data and decreasing the time before training and model verification would occur.",
    "96192": "[quote=Μαριος Μιχαηλιδης KazAnova;96184]\r\n\r\nOur best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.\r\n\r\nThere was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.\r\n\r\nWe also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.\r\n\r\nsome models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . \r\n\r\nI think we reached Level 4 meta-modelling in the end :)\r\n[/quote]\r\nHow much RAM it requires?",
    "96193": "For us it was like this:\r\n\r\nXgboost (single): 0.98422\r\n\r\nXgboost (dual ensemble): 0.98459\r\n\r\nRandomForest (single): 0.97455\r\n\r\nExtraTrees (single): 0.97628\r\n\r\nDone :D, our final submission is an ensemble of pretty much those 4.",
    "96195": "[quote=Artem;96192]\r\n\r\nHow much RAM it requires? \r\n\r\n[/quote]\r\n\r\nA lot!\r\n\r\n128GB should be ok!\r\n\r\nBut to be honest, we were reckless (as we had 256GB available). I cannot tell you what is the most optimum you can run it on. 64GB is possibly enough too.",
    "96196": "[quote=Gerard Toonstra;96191]\r\n\r\nHorrifying. What were the features? I trained FTRL on 4-5 epochs on 1800 raw features, reduced from 960k.  My features were the most prevalent 1800 word features with optimized L1/L2 regularization.\r\n\r\nI went fruity-loopy in this competition and made a custom C parser, a bigram/trigram processor in python and in the end I was processing pretty much portions of matrices to get things going (scipy.sparse.hstack). All of this to reduce the memory requirements for processing the data and decreasing the time before training and model verification would occur.\r\n\r\n\r\n[/quote]\r\n\r\nFor FTRL, we used simple binary counts (is it there or not?) of unigrams and bigrams as features. We tried many more feature engineering techniques, but the simple was the best.",
    "96197": "[quote=Subhajit Mandal;96189]\r\n\r\nWe achieved 0.979 with FTRL, trained it for 5 epochs.\r\n\r\n[/quote]\r\n\r\nAlong with FTRL, we had following scores with other models:\r\n\r\nxgboost: 0.983\r\n\r\nstacked xgboost: 0.985\r\n\r\nMLP with 3 layerS: 0.952\r\n\r\nI agree with @rcarson and @NxGTR, bigger teams and heavy ensembing is the way to do well in these competitions.",
    "96198": "[quote=Little Boat;96190]\r\n\r\n[quote=Subhajit Mandal;96183]\r\n\r\n@KazAnova or anyone else who got 0.98 with Random Forest, how did you do that? We could barely make 0.93 with multiple attempts.\r\n\r\n[/quote]\r\nAdd these count features to the benchmark code, you will get to at least 0.95 with Random Forest\r\n\r\n        row['spaces'] = text.count(' ')\r\n        row['tabs'] = text.count('\\t')\r\n        row['braces'] = text.count('{')\r\n        row['brackets'] = text.count('[')\r\n        row['words'] = len(re.split('\\s+', text))\r\n        row['length'] = len(text)\r\n        row['bracket_2'] = text.count('<')\r\n        row['bracket_3'] = text.count('(')\r\n        row['re_w'] = len(re.split('\\w+', text))\r\n        row['new_par'] = text.count('\\n\\t')\r\n        row['dbslash'] = text.count('//')\r\n        row['bslash'] = text.count('/')\r\n        row['bol_0'] = text.count('<!')\r\n        row['bol_1'] = text.count('/>')\r\n        row['bol_2'] = text.count('&')\r\n        row['bol_3'] = text.count(';')\r\n        row['bol_4'] = text.count('==')\r\n        row['bol_5'] = text.count('===')\r\n        row['bol_6'] = text.count('css')\r\n        row['bol_7'] = text.count('#')\r\n        row['bol_8'] = text.count('@')\r\n        row['bol_9'] = text.count('$')\r\n        row['bol_10'] = text.count('%')\r\n        row['bol_11'] = text.count('^')\r\n        row['bol_12'] = text.count('+')\r\n        row['bol_13'] = text.count('?')\r\n        row['bol_14'] = text.count('|')\r\n        row['bol_15'] = text.count('\\\\')\r\n        row['bol_16'] = text.count('*')\r\n        row['bol_17'] = text.count('||')\r\n        row['bol_18'] = text.count('\\t\\t')\r\n        row['bol_19'] = text.count('\\t\\t\\t')\r\n\r\n[/quote]\r\n\r\nI did not see that coming ;-) should not have removed the punctuation.",
    "96199": "[quote=Artem;96192]\r\n\r\n[quote=Μαριος Μιχαηλιδης KazAnova;96184]\r\n\r\nOur best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.\r\n\r\nThere was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.\r\n\r\nWe also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.\r\n\r\nsome models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . \r\n\r\nI think we reached Level 4 meta-modelling in the end :)\r\n[/quote]\r\nHow much RAM it requires? \r\n\r\n[/quote]\r\n\r\nI managed to run a large 500k single word + 232k bigram + 178k trigram model in 34G space (close to 0.9/1M features). Basically externalizing the entire tf-idf vectorizer. Then you notice how little those classes contribute vs. the cost and how much easier it would be if you just take 2 passes at 2-3x the time... But if memory is lacking...",
    "96200": "[quote=Gerard Toonstra;96185]\r\n\r\nWell, my hats off to David Shinn. He entered very late in the competition, eliminated the leaderboard very quickly and finished 10th, just dethroning us from our precious position. So I truly natively hate him, but also greatly, absolutely respect him for doing this on his own. Hats off to you sir!\r\n\r\n[/quote]\r\n\r\nThank @Gerard and @Redo, your constant improvement until the end, even a little, kept me motivated to keep pushing till the end.  I congratulate all of the top winners, I fight tooth and nail just to break top 10.  This was a crazy exhausting and crazy fun competition.  Thanks to orchid, Bang Nguyen, Ashtericz, Redo, and NxGTR & Sky for making just cracking top 10 a nail biting experience.",
    "96201": "[quote=Subhajit Mandal;96196]\r\n\r\n[quote=Gerard Toonstra;96191]\r\n\r\nHorrifying. What were the features? I trained FTRL on 4-5 epochs on 1800 raw features, reduced from 960k.  My features were the most prevalent 1800 word features with optimized L1/L2 regularization.\r\n\r\nI went fruity-loopy in this competition and made a custom C parser, a bigram/trigram processor in python and in the end I was processing pretty much portions of matrices to get things going (scipy.sparse.hstack). All of this to reduce the memory requirements for processing the data and decreasing the time before training and model verification would occur.\r\n\r\n\r\n[/quote]\r\n\r\nFor FTRL, we used simple binary counts (is it there or not?) of unigrams and bigrams as features. We tried many more feature engineering techniques, but the simple was the best.\r\n\r\n[/quote]\r\n\r\nAh, I only included the positive presence 1, but not the 0 -presence... This being FTRL, it's basically unrolling your matrix to disk...",
    "96202": "[quote=Gerard Toonstra;96185]\r\n\r\nWhat we lacked were ensembling techniques and variety in models or data.\r\n\r\n[/quote]\r\n\r\nWell, if you had done what I did, you sure would have bested me easily.  My single best model only barely reaches 0.980 on the Private LB.  What I lacked in a single model I made up for in my Frankenstein stacked model.  I just counted, I had 80 models as features in my second layer (plus around 20 or so other basic features).  I used Random Forest, Extra Trees, XGBoost, Vowpal Wabbit, KNN, SVM, and logistic regression in my first set of models.  Each of those trained on different subset of features: character counts, html tags, domain urls, html attributes, file extensions, and 123gram tfidf and 123gram counts.  I even tried some crazy things like the character and html tag counts from the first half of the file and the second half of the file as different training feature sets (sort of a data augmentation).  In my second layer, I used a boosted random forest and xgboost models.  My third layer was a hyperopt optimized power and linear transformation of those 2 predictions from the second layer.  This was my first stab at serious stacking, so it was a great learning experience.",
    "96205": "I didn't do nearly as well as you guys, but almost all of my models were just variants of TINRTGU's code in pure Python. I just greedily averaged these FTRLP with different orders of files, different one hot splitting techniques, etc. I wish I had time for better models, but I ran out of time in terms of creating meta features to be fed into XG.",
    "96206": "Mad Professors: Congratulations, and good job!  Don't worry, I knew our positions could easily reverse.  It was fun to have a real shot at the top position for a while.\r\n\r\nMy submission was just the scaled sum of two models, a linear SVM of skip-grams and element/attribute combinations (79 million unique features, ~45 minutes training time), and an XGBoost model using substrings as features (30,000 features, 2000 trees, ~5 hours training time).\r\n\r\nThe substrings were automatically selected using this tool: https://github.com/mortehu/substring-frequencies\r\n\r\nThese are the substrings used as features in the XGBoost model, with the most predictive being listed first: https://gist.github.com/mortehu/07f0c59dc30bf495853d",
    "96209": "Our feature set is huge, so it is painful to try stacking or other complicated stuff. But if I don't  go into the Physics one, I will probably have time to finish tfidf and two stage stacking. I was too optimistic to think we can perform well on two competitions at the same time. With all the other works, I just don't have enough energy and computers to try lots of things in two competitions.\r\n\r\n[quote=NxGTR;96193]\r\n\r\nFor us it was like this:\r\n\r\nXgboost (single): 0.98422\r\n\r\nXgboost (dual ensemble): 0.98459\r\n\r\nRandomForest (single): 0.97455\r\n\r\nExtraTrees (single): 0.97628\r\n\r\nDone :D, our final submission is an ensemble of pretty much those 4. \r\n\r\n\r\n[/quote]",
    "96210": "You **have** performed well on both of the competitions. In Physics competition it was a really bad luck for you that high scoring scripts got posted at the last moment. Luckily, there was no scripts in this competition.",
    "96213": "It is amazing you can get such good results with a simple ensemble of two models. \r\nBut I guess your two models are probably very strong by themselves. \r\n\r\nMay I ask how are you able to train a linear SVM with 79 million features in 45 minutes,  I always find svm to be very slow. Another thing is did you use some software package to extract skip gram features?\r\n\r\n[quote=mortehu;96206]\r\n\r\n@Mad Professors: Congratulations, and good job!  Don't worry, I knew our positions could easily reverse.  It was fun to have a real shot at the top position for a while.\r\n\r\nMy submission was just the scaled sum of two models, a linear SVM of skip-grams and element/attribute combinations (79 million unique features, ~45 minutes training time), and an XGBoost model using substrings as features (30,000 features, 2000 trees, ~5 hours training time).\r\n\r\nThe substrings were automatically selected using this tool: https://github.com/mortehu/substring-frequencies\r\n\r\nThese are the substrings used as features in the XGBoost model, with the most predictive being listed first: https://gist.github.com/mortehu/07f0c59dc30bf495853d\r\n\r\n[/quote]",
    "96214": "[quote=SkyLibrary;96209]\r\nI was too optimistic to think we can perform well on two competitions at the same time...\r\n[/quote]\r\n\r\nHehehe, Top Kagglers issues: when two \"top 13th\" is bad! :P\r\n\r\nIt was OK, could be better but heck!, pretty much always can be better :D",
    "96217": "[quote=SkyLibrary;96213]\r\n\r\nMay I ask how are you able to train a linear SVM with 79 million features in 45 minutes,  I always find svm to be very slow. Another thing is did you use some software package to extract skip gram features?\r\n\r\n[/quote]\r\n\r\nmortehu is an amazing programmer according to his repo. He probably just coded it from scratch in c++",
    "96218": "lol. Two 13th is definitely better than two 26th. But wait a minutes, does 13 suppose to be a unlucky number in some countries, luckily I don't come from these countries so I don't have hard feelings for it.\r\n[quote=NxGTR;96214]\r\n\r\n[quote=SkyLibrary;96209]\r\nI was too optimistic to think we can perform well on two competitions at the same time...\r\n[/quote]\r\n\r\nHehehe, Top Kagglers issues: when two \"top 13th\" is bad! :P\r\n\r\nIt was OK, could be better but heck!, pretty much always can be better :D\r\n\r\n[/quote]",
    "96220": "Congrats Mad Professors, mortehu and bibaze.! Special thanks to @David shinn for his excellent starter code which helped us enter this competition since we started very late. Thanks everyone for posting their approaches. \r\n\r\n@Marios: 4-level meta models.! You are pushing the boundaries further. Looking forward to your team's code soon. \r\n@rcarson: One more top 10 finish to your list. You pretty much get there almost every time :)",
    "96222": "[quote=mortehu;96206]\r\n\r\n@Mad Professors: Congratulations, and good job!  Don't worry, I knew our positions could easily reverse.  It was fun to have a real shot at the top position for a while.\r\n\r\nMy submission was just the scaled sum of two models, a linear SVM of skip-grams and element/attribute combinations (79 million unique features, ~45 minutes training time), and an XGBoost model using substrings as features (30,000 features, 2000 trees, ~5 hours training time).\r\n\r\nThe substrings were automatically selected using this tool: https://github.com/mortehu/substring-frequencies\r\n\r\nThese are the substrings used as features in the XGBoost model, with the most predictive being listed first: https://gist.github.com/mortehu/07f0c59dc30bf495853d\r\n\r\n[/quote]\r\nThese codes are awesome.. I wish I can code anywhere near that level someday.\r\n\r\nSomething interesting I noticed in the substrings list is that there are many \"dates\" strings that coincide with the leak found mid competition.. I hadn't figured out but perhaps most of the pages have a 'current time' recorded in the html file. Maybe we could call that a 'nerfed' version of the leak. Clueless about what that '1529' would mean, though. edit: maybe week number?\r\n\r\nAlso, anybody has an idea as to why these sponsored ads seems so predictable through features even though it doesn't have any apparent correlation? After all, these sponsored pages are just pages people paid StumbleUpon to feature in their system. Never thought background features like number of tags or chars could be so effective at predicting that.\r\n\r\nCongrats to everyone and I'm specially thankful to David Shinn, whose code got me started in the competition as well.",
    "96223": "As far as I can tell, this competition was more about determining the date each page was downloaded, rather than whether or not something was advertising.  Negative examples were mostly downloaded towards the end, for example.  Current events, like \"Pluto\" (New Horizons passed Pluto on July 24th), \"Donald Trump\", and \"El Chapo\" were all strong features, with \"infographic\" appearing pretty far down on the list.  Some of the other number strings seen in my feature lists are versions of JavaScript libraries that were released after the positive samples had been downloaded.\r\n\r\nI'll post the source code for my linear SVM training and HTML tokenization program tomorrow.",
    "96239": "[quote=Μαριος Μιχαηλιδης KazAnova;96184]\r\n\r\nOur best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.\r\n\r\nThere was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.\r\n\r\nWe also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.\r\n\r\nsome models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . \r\n\r\nI think we reached Level 4 meta-modelling in the end :)\r\n\r\n\r\n\r\n[/quote]\r\nI am curious about your 3-grams tf-idf, do you include the tags and scripts of HTML files? Seems that the tags and scripts are only helpful to classify HTML files from the same websites(they share the same HTML structure.) I would be appreciate if you can share more about features and how you clean HTML files, thanks a lot!",
    "96249": "One doubt I had... I started with tf-idf on xgb, but noticed that IDF itself had no impact on end results. Also, the official application of term frequency seemed to slightly reduce the results as well, probably because the size of the documents varies quite a bit. So my XGB would perform 0.96/0.97 on TF calculated as : f(t) / n(t,d)  , whereas a simple f(t) (just a basic count) would get me 0.98 performance. Can you guys confirm you got better results with simple counts as well?\r\n\r\nMy max_depth for xgb was optimized at 50 with mostly default settings, except colsample_by_tree set to 0.46. subsampling made no difference and neither did gamma. \"scale_pos_weight\" sounds like a great opportunity to tweak, because 1's were less than 0's and it seemed to work in the first few iterations, but then it leveled off very quickly.\r\n\r\nI also did a full minhash doc-doc comparison on 150 hashes. I calculated the hashes in python, but wrote a custom C app to do the comparison. That got me 0.89 on a 20% similarity threshold setting. A combination of 80/50/20% similarity results got me 0.929.",
    "96252": "[quote=SRK;96220]\r\n\r\nCongrats Mad Professors, mortehu and bibaze.! Special thanks to @David shinn for his excellent starter code which helped us enter this competition since we started very late. Thanks everyone for posting their approaches. \r\n\r\n@Marios: 4-level meta models.! You are pushing the boundaries further. Looking forward to your team's code soon. \r\n@rcarson: One more top 10 finish to your list. You pretty much get there almost every time :)\r\n\r\n\r\n[/quote]\r\n\r\nCongratulations Mad Professors - awesome finish @mortehu - your repo looks intimidating!\r\nThanks to @DavidShinn for the starter code without which I wouldn't have even thought of attempting this daunting challenge.\r\n\r\n@SRK, special thanks to U for considering to team up with a Kaggler, despite being a Master yourself! \r\n\r\n@Everybody else, SRK is an awesome teammate. *Endorsing him*\r\n\r\nWe got late into the competition (~17 days to go), ignored the competition for a few more days and actually worked only for the last 3 days to be very honest and ending up at #33 is no mean feat. Very happy to finish well and it could only happen with having a teammate like SRK who pretty much guided me throughout the last 3 days where we tried what we could! :-)\r\n\r\nI have learnt a lot of things and continue to learn more and do well :-)\r\n\r\n@rcarson - got to know a lot about you. SRK spoke very highly of you. I hope I team up with you someday. For that, I'll work hard, earn the right and ask U :-)",
    "96273": "[quote=Μαριος Μιχαηλιδης KazAnova;96184]\r\n\r\nOur best single model was an xgboost trained on 3-grams  tf idf. it score 0.9873 by itself! I think it would make it top 10.\r\n\r\nThere was some pre-processing with removing stopwords and some clean-up of the files, but generally we used all the content to create tf-idf.\r\n\r\nWe also created many other features, sentiment-based, count of urls, dots and other non alphanumeric. I think our final solution includes around 50 models.\r\n\r\nsome models trained on normal input data (normally tf-idf matrices or svd componets) or meta models (e.g. models that used other model's predictions as input) . \r\n\r\nI think we reached Level 4 meta-modelling in the end :)\r\n\r\n[/quote]\r\n\r\n\r\nAs a n00b, can I ask how exactly did you preprocess those files? Or in general, how should I deal with such noisy data in the future, if I'd like to extract some usefull TF-IFDs?",
    "96277": "Actually,  I just loaded each file as a big string (lets call it mline)  and then cleaned with this:\r\n\r\nimport re\r\n \r\nmline=re.sub(\"[^a-zA-Z0-9]\",\" \", mline) # remove non-alphanumeric\r\n \r\nthen I re-Wrote this to a file. \r\n\r\nThe test is up to the to the tf-idf parameters!\r\n\r\n tfv=TfidfVectorizer(min_df=2, max_features=None, strip_accents='unicode',lowercase =True,\r\n                        analyzer='word', dtype=np.float64, token_pattern=r'\\w{2,}', ngram_range=(1, 3), use_idf=True,smooth_idf=True, \r\n    sublinear_tf=True, stop_words = 'english')",
    "96282": "Thanks Dato, Kaggle, team mates and all competitors! Grats all winners and new Kaggle Masters.\r\n\r\nMy focus was on online learning and generating non-alphanumeric features.\r\n\r\nI could not extract the entire dataset on my HD, so I worked with the compressed archives only, unzipping and streaming the documents into the algo's VW, FTRL, SGD and perceptron. The scores were not as high as I've seen reported in this thread (well done!), but I think the models/vectorizers were diverse enough to contribute to the ensemble.\r\n\r\nWe also pursued three crazy ideas that happened to work (a tiny bit, but enough to be exciting for further research):\r\n\r\n- Training a model on a sentiment analysis data set unrelated to this competition, then using the predictions for the train and test sets for further modeling. Such a \"sentiment\"-feature was informative.\r\n\r\n- Training a model on the 20 newsgroups data set and creating multi-class predictions for train and test sets. I think these vectors added a form of categorization/topic modeling. (\"this document talks about 'SciMed' and 'CompSci', but not 'Religion'\").\r\n\r\n- Using Normalized Compression Distance with the fast Snappy compressor to calculate distance between every document and 4 randomly created anchor corpora.\r\n\r\nAgain, learned a lot, for instance about geometric mean for model averaging, feature and meta-model interactions with [XGBfi][1], proper competition management/iteration, patience, and stacking. I am particularly excited about our 4 levels of stacking, with 4th level models using the predictions from the 1st-3rd level models (fully connected stacknet?).\r\n\r\n  [1]: https://github.com/far0n/xgbfi",
    "96289": "[quote=Triskelion;96282]\r\nAgain, learned a lot, for instance about geometric mean for model averaging, feature and meta-model interactions with XGBfi.\r\n[/quote]\r\n\r\nThat looks very interesting Thanks!\r\n\r\nI gave it a quick read, but, is that Windows-Only (XGBfi)?",
    "96297": "I think so, though maybe you can run on Linux Mono. Far0n is considering porting to C++.\r\n\r\nI think the nicest would be to have it added to XGBoost source. It's really useful. Adding high-ranking interactions as features for a linear algorithm boosts its score.",
    "96301": "Here's the text-classifier I used for this competition released as open source: https://github.com/mortehu/text-classifier\r\n\r\nI hope this can be used to set a high baseline in any future text classification competition, or to be more generally useful.",
    "96307": "So who used graphlab?",
    "96324": "[quote=NxGTR;96289]\r\n\r\n[quote=Triskelion;96282]\r\nAgain, learned a lot, for instance about geometric mean for model averaging, feature and meta-model interactions with XGBfi.\r\n[/quote]\r\n\r\nThat looks very interesting Thanks!\r\n\r\nI gave it a quick read, but, is that Windows-Only (XGBfi)?\r\n\r\n\r\n[/quote]\r\n\r\nSo I just checked and it runs fine under linux using mono. \r\n\r\nI know .Net is a bit odd, but I'm planning to build a GUI interfacing XGB with features like training graphs, comparing trainings graphs of different parameter sets, querying feature split values, feature importance based on hold out data performance (which I expect to be useful to detect features which rank high but hurt performance) and so on. But I'm only familar with WPF for creating GUIs. That is why xgbfi is currently written in .Net.",
    "96325": "[quote=mortehu;96301]\r\n\r\nHere's the text-classifier I used for this competition released as open source: https://github.com/mortehu/text-classifier\r\n\r\nI hope this can be used to set a high baseline in any future text classification competition, or to be more generally useful.\r\n\r\n[/quote]\r\n\r\njust wow!",
    "96369": "Dumb question, what is \"FTRL\"?\r\n\r\nAlso, what 2 files did you put into this tool to select substrings? https://github.com/mortehu/substring-frequencies",
    "96374": "[quote=Zach;96369]\r\nAlso, what 2 files did you put into this tool to select substrings? https://github.com/mortehu/substring-frequencies\r\n[/quote]\r\n\r\nI picked a random sample of 512 MB of positive examples, and 512 MB of negative examples.  I used the `--document` mode.",
    "96376": "[quote=Zach;96369]\r\n\r\nDumb question, what is \"FTRL\"?\r\n\r\n[/quote]\r\n\r\nFTRL stands for the online learning algorithm \"Follow The Regularized Leader\".",
    "96387": "Hi [@mortehu][1],\r\n\r\nI'm intrigued and really admiring your way. Just curious, would it be possible to have your substring frequency more (or less) accurate to instruct the suffix array to discount overlapped things? (In case my description sounds vague and weird or this method isn't popular... it's kinda trick to eliminate stop words without knowing then and sometimes emphasize critical information. For example from the earliest paper I can find, when encountered a bigram \"rail enquiries\" in BNC, it's actually useless because it was always part of either \"national rail enquiries\" or \"British british rail enquiries.\")\r\n\r\nI didn't use it just because I didn't get to that phase yet. :p\r\n\r\n  [1]: https://www.kaggle.com/mortehu",
    "96393": "Barabbas, even though it might not look like it, a lot of redundant substrings are already removed. However, the code currently keeps overlapping strings if they have different occurrence counts, even if that is stupid in most cases. It's hard to know without building a model whether a substring of another feature is truly redundant, so I didn't even try. You can see the duplicate detection code on line 249 in substrings.cc.",
    "96400": "[@mortehu][1]: I see, thank you. :D\r\n\r\n\r\n  [1]: https://www.kaggle.com/mortehu",
    "96414": "[quote=Faron;96376]\r\n\r\nFTRL stands for the online learning algorithm \"Follow The Regularized Leader\".\r\n\r\n[/quote]\r\n\r\nHow does that differ from other online algorithms, e.g. vowpal wabbit?  Where can I get the software to try it out?\r\n\r\nActually it looks like VW has a [--ftrl option][1].  I'll have to try that out sometime.\r\n\r\n\r\n  [1]: https://github.com/JohnLangford/vowpal_wabbit/wiki/Command-line-arguments",
    "96418": "[quote=Zach;96414]\r\n\r\nHow does that differ from other online algorithms, e.g. vowpal wabbit?  Where can I get the software to try it out?\r\n\r\n[/quote]\r\n\r\nFor me it is like an SGD, with the difference that in each step the coefficients are being copied and converted to \"new values\" based on the regularization terms. \r\n\r\nFor l2 regul, this should not have much difference with sgd, but for l1 it makes quite a difference  because the coefficients will always remain zero (0.0) if they never pass the the C value (something not really achievable with simple sgd)  and will never be copied across . \r\n\r\nIn other words, only if something becomes significant (after the updates's phase) will start being copied to further boost itself (hence the follow the leading features after regularization is applied!) \r\n\r\nAlso, most of the times it uses passed gradients as a form of adaptive learning rate so generally as an algorithm is more independent. \r\n\r\nAt least, that is my interpretation!",
    "96425": "Hi! I want to thank David Shinn for his very nice script, it motivated me to continue working on the comp, since it was so easy to improve :) \r\n\r\nAfter dabbling a while with it and improving it manually, I calculated how much popular (top50k) individual words separated the classes by abs(freq_pos - freq_neg), and took the top3k and used them in RF and XGB as features with no tuning of parameters, then simple average. I did not expect that simple solution to earn us a 19th place. It was a really fun competition due to the size of data!\r\n\r\nAlso thanks to my teammate",
    "96614": "Congratulation to winner. \r\n\r\nI ended up at 14th using liblinear and xgboost. I built three 1-gram models: 2 liblinear models with Bernoulli and multinominal feature weights  and 1 xgboost model with Bernoulli feature weights. On the 2nd level I blended those three model using xgboost and linear regression. Then at 3rd level, I used h mean to merge two models at 2nd level.\r\n\r\nAccording to me, preprocessing is the most important thing. I used a regular expression to split the html string to tokens. This regular expression can catch domain names, javascript function calls, html attributes, date, time... Some important features from my xgboost model are:\r\n\r\n    (u'/2015/07/0', 255), (u'2013%20', 251), (u'keywords!', 246), (u'2014%20fifa%20world%20cup', 244), (u'2015-07-07+21', 214), (u'blog!', 209), (u'/blog%206', 198), (u'other!', 195), (u'july!', 194), (u'copyright!', 192), (u'2015-07-14+00', 191), (u'rss!', 181), (u'14%', 180), (u'20%', 178), (u'web!', 177), (u'100%%', 176), (u'our!', 173), (u'news!', 173), (u'199.302', 172), (u'article!', 170), (u'author!', 170), (u'site_name.indexof(', 170), (u'30%', 168), (u'/plus.google.com.com', 168), (u'border?', 167), (u'terms!', 167), (u'en!', 167), (u'tag!', 166), (u'block!', 166), (u'/favicon.ico$$3000.629-18', 165), (u'clear!', 165), (u'/assets-0.housingcdn.com', 165), (u'pinterest!', 165)\r\n\r\nI wanted to try 2 grams models but 3 days before the deadline, one of two fans in my Macbook broke, so I can never submit another submission before deadline again. But it was my false for depending too much on something and don't have a backup plan."
  },
  "source": "meta"
}