{
  "id": 16991,
  "title": "Thanks Everybody!",
  "url": "/competitions/dato-native/discussion/16991",
  "author_name": "",
  "post_date": "2015-10-14T02:18:36.280Z",
  "votes": 15,
  "comment_count": 13,
  "views": 2080,
  "content": "<p>Before competition ends and everybody disappear, I just wanted to have the chance to thank all the incredible and unexpected support that for some reason I received in this competition. Several people contacted me by private message with the sole purpose to wish me luck, something I really appreciate.</p>\n\n<p>Unless something weird happen to the LB I feel that I let some people down, but do not worry, there is more from NxGTR, just wait for it :P</p>\n\n<p>I feel very sad, not for me or the competition performance, but for the poor guys at Deloitte hahaha :D, I am sorry guys, I guess you will have to miss my super skills :D</p>\n\n<p>Great work from my team mate!, and you guys as always with outstanding performance.</p>\n\n<p>See you around!, I always mess in everything :D</p>",
  "messages": [
    {
      "id": "96058",
      "postDate": "10/14/2015 02:18:36",
      "content": "<p>Before competition ends and everybody disappear, I just wanted to have the chance to thank all the incredible and unexpected support that for some reason I received in this competition. Several people contacted me by private message with the sole purpose to wish me luck, something I really appreciate.</p>\n\n<p>Unless something weird happen to the LB I feel that I let some people down, but do not worry, there is more from NxGTR, just wait for it :P</p>\n\n<p>I feel very sad, not for me or the competition performance, but for the poor guys at Deloitte hahaha :D, I am sorry guys, I guess you will have to miss my super skills :D</p>\n\n<p>Great work from my team mate!, and you guys as always with outstanding performance.</p>\n\n<p>See you around!, I always mess in everything :D</p>",
      "rawMarkdown": "Before competition ends and everybody disappear, I just wanted to have the chance to thank all the incredible and unexpected support that for some reason I received in this competition. Several people contacted me by private message with the sole purpose to wish me luck, something I really appreciate.\r\n\r\nUnless something weird happen to the LB I feel that I let some people down, but do not worry, there is more from NxGTR, just wait for it :P\r\n\r\nI feel very sad, not for me or the competition performance, but for the poor guys at Deloitte hahaha :D, I am sorry guys, I guess you will have to miss my super skills :D\r\n\r\nGreat work from my team mate!, and you guys as always with outstanding performance.\r\n\r\nSee you around!, I always mess in everything :D",
      "votes": null
    },
    {
      "id": "96062",
      "postDate": "10/14/2015 03:01:03",
      "content": "<p>Thank you NXGTR. And yeah let's have some weird thing happen!</p>",
      "rawMarkdown": "Thank you NXGTR. And yeah let's have some weird thing happen!",
      "votes": null
    },
    {
      "id": "96065",
      "postDate": "10/14/2015 03:29:37",
      "content": "<p>I thought I could get us a top 10 in physics, but that one is so weird, we droped from 7 to 13 during the last day even though we improved the score a lot... </p>\n\n<p>For this competition, NxGTR really did great job showing how much a single model with simple feature engineering can do, some times simple methods are more valuable.  One disapointment is that GLC boost tree keep crushing because of memory error even though we have more than enough memory, it is basically the same as xgboost, not sure what happened. </p>",
      "rawMarkdown": "I thought I could get us a top 10 in physics, but that one is so weird, we droped from 7 to 13 during the last day even though we improved the score a lot... \r\n\r\nFor this competition, NxGTR really did great job showing how much a single model with simple feature engineering can do, some times simple methods are more valuable.  One disapointment is that GLC boost tree keep crushing because of memory error even though we have more than enough memory, it is basically the same as xgboost, not sure what happened.",
      "votes": null
    },
    {
      "id": "96129",
      "postDate": "10/14/2015 17:33:01",
      "content": "<p>I feel your pain. I'm having spacy crash as well... which means my NLP-based features won't make it into the model. Of course it's my own fault that I have started so late and now run out of time, but still.</p>",
      "rawMarkdown": "I feel your pain. I'm having spacy crash as well... which means my NLP-based features won't make it into the model. Of course it's my own fault that I have started so late and now run out of time, but still.",
      "votes": null
    },
    {
      "id": "96131",
      "postDate": "10/14/2015 18:19:09",
      "content": "<p>[quote=Toby Cheese;96129]\nI feel your pain. I'm having spacy crash as well... which means my NLP-based features won't make it into the model. Of course it's my own fault that I have started so late and now run out of time, but still.\n[/quote]</p>\n\n<p>Good thing is you can always submit after competition ends, just keep on the good work :D</p>",
      "rawMarkdown": "[quote=Toby Cheese;96129]\r\nI feel your pain. I'm having spacy crash as well... which means my NLP-based features won't make it into the model. Of course it's my own fault that I have started so late and now run out of time, but still.\r\n[/quote]\r\n\r\nGood thing is you can always submit after competition ends, just keep on the good work :D",
      "votes": null
    },
    {
      "id": "96151",
      "postDate": "10/14/2015 20:31:44",
      "content": "<p>[quote=NxGTR;96131]</p>\n\n<p>[quote=Toby Cheese;96129]\nI feel your pain. I'm having spacy crash as well... which means my NLP-based features won't make it into the model. Of course it's my own fault that I have started so late and now run out of time, but still.\n[/quote]</p>\n\n<p>Good thing is you can always submit after competition ends, just keep on the good work :D</p>\n\n<p>[/quote]</p>\n\n<p>@NxGTR, if our team had invited yours for a merge before the deadline, would you guys say yes?  </p>\n\n<p>It's more and more difficult to end up in top 10 now without a large team (certainly not the case for mortehu). I am considering joining a big team of more than 5 people from now on in case we are competing for money..</p>",
      "rawMarkdown": "[quote=NxGTR;96131]\r\n\r\n[quote=Toby Cheese;96129]\r\nI feel your pain. I'm having spacy crash as well... which means my NLP-based features won't make it into the model. Of course it's my own fault that I have started so late and now run out of time, but still.\r\n[/quote]\r\n\r\nGood thing is you can always submit after competition ends, just keep on the good work :D\r\n\r\n[/quote]\r\n\r\n@NxGTR, if our team had invited yours for a merge before the deadline, would you guys say yes?  \r\n\r\nIt's more and more difficult to end up in top 10 now without a large team (certainly not the case for mortehu). I am considering joining a big team of more than 5 people from now on in case we are competing for money..",
      "votes": null
    },
    {
      "id": "96153",
      "postDate": "10/14/2015 20:43:38",
      "content": "<p>[quote=rcarson;96151]\n@NxGTR, if our team had invited yours for a merge before the deadline, would you guys say yes? <br>\n[/quote]</p>\n\n<p>I would have asked my teammate, but I would say yes, for sure.\nThis is my first Kaggle team up experience, and it was great, looking forward more teamups if chances arise.</p>\n\n<p>Just for fun, I can post a &quot;steal my entry&quot; (after competition finish, I think it breaks no rules), so we can find out how well we could have done.</p>",
      "rawMarkdown": "[quote=rcarson;96151]\r\n@NxGTR, if our team had invited yours for a merge before the deadline, would you guys say yes?  \r\n[/quote]\r\n\r\nI would have asked my teammate, but I would say yes, for sure.\r\nThis is my first Kaggle team up experience, and it was great, looking forward more teamups if chances arise.\r\n\r\nJust for fun, I can post a \"steal my entry\" (after competition finish, I think it breaks no rules), so we can find out how well we could have done.",
      "votes": null
    },
    {
      "id": "96175",
      "postDate": "10/15/2015 00:21:22",
      "content": "<p>That was a really fun competition!  I still consider myself a complete novice and it was enjoyable to work with a more raw data set than some of the other competitions I have tried.  It was also semi refreshing to escape from the scripts from a bit.  As much as I enjoy seeing what others are doing, it was fun to just compete without worrying about people copying a well performing script at the end of the competition.</p>\n\n<p>I look forward to seeing some of the innovative approaches of the top performers!</p>",
      "rawMarkdown": "That was a really fun competition!  I still consider myself a complete novice and it was enjoyable to work with a more raw data set than some of the other competitions I have tried.  It was also semi refreshing to escape from the scripts from a bit.  As much as I enjoy seeing what others are doing, it was fun to just compete without worrying about people copying a well performing script at the end of the competition.\r\n\r\nI look forward to seeing some of the innovative approaches of the top performers!",
      "votes": null
    },
    {
      "id": "96250",
      "postDate": "10/15/2015 09:17:03",
      "content": "<p>Hey, NxGTR, I remember You achieving .98-ish scores right at the start of the competition. Wouldn't You like to share some details of Your approach ;)?</p>",
      "rawMarkdown": "Hey, NxGTR, I remember You achieving .98-ish scores right at the start of the competition. Wouldn't You like to share some details of Your approach ;)?",
      "votes": null
    },
    {
      "id": "96258",
      "postDate": "10/15/2015 10:59:28",
      "content": "<p>You get so close. Hope we can both make top 10 in next match. :)</p>",
      "rawMarkdown": "You get so close. Hope we can both make top 10 in next match. :)",
      "votes": null
    },
    {
      "id": "96288",
      "postDate": "10/15/2015 15:10:18",
      "content": "<p>[quote=M.E.;96250]\nHey, NxGTR, I remember You achieving .98-ish scores right at the start of the competition. Wouldn't You like to share some details of Your approach ;)?\n[/quote]</p>\n\n<p>It was nothing fancy really (pretty much what everybody did, count and count things), probably thats the reason my score pretty much remained the same from the start to finish hehehe, the only &quot;semi-fancy&quot; is that I have very limited RAM (8GB), so all the feature engineering is done online with multiple passes.</p>\n\n<p>Features were divided in:</p>\n\n<p>Shallow: Counts of sections, e.g. instead of just count individual words/chars as features, I included specific things like what kind of metadata was included in the page, and so on.</p>\n\n<p>Deep(?): I tried to build some kick-ass features such as &quot;Does this webpage uses new JS standards?&quot;, very specific to this context.</p>\n\n<p>Feature number: Before the restart, 10M, After the restart, 8M, I got lended a machine with 16GB RAM, but never managed to do tfid, so they are just plain counts.</p>\n\n<p>Training:\nPretty much the winner was Xgboost, with deep=19, it took I think over 20 Hrs to train, I never managed to fit it memory, so I chunked the data or remove to as few as 250K features to fit it.</p>\n\n<p>Memory was always an issue, and after the restart, it was much worse for me, it was just too much data.</p>",
      "rawMarkdown": "[quote=M.E.;96250]\r\nHey, NxGTR, I remember You achieving .98-ish scores right at the start of the competition. Wouldn't You like to share some details of Your approach ;)?\r\n[/quote]\r\n\r\nIt was nothing fancy really (pretty much what everybody did, count and count things), probably thats the reason my score pretty much remained the same from the start to finish hehehe, the only \"semi-fancy\" is that I have very limited RAM (8GB), so all the feature engineering is done online with multiple passes.\r\n\r\nFeatures were divided in:\r\n\r\nShallow: Counts of sections, e.g. instead of just count individual words/chars as features, I included specific things like what kind of metadata was included in the page, and so on.\r\n\r\nDeep(?): I tried to build some kick-ass features such as \"Does this webpage uses new JS standards?\", very specific to this context.\r\n\r\nFeature number: Before the restart, 10M, After the restart, 8M, I got lended a machine with 16GB RAM, but never managed to do tfid, so they are just plain counts.\r\n\r\nTraining:\r\nPretty much the winner was Xgboost, with deep=19, it took I think over 20 Hrs to train, I never managed to fit it memory, so I chunked the data or remove to as few as 250K features to fit it.\r\n\r\nMemory was always an issue, and after the restart, it was much worse for me, it was just too much data.",
      "votes": null
    },
    {
      "id": "96321",
      "postDate": "10/15/2015 17:30:47",
      "content": "<p>When I started running up against memory issues around the string manipulation (or rather mapping in RAM), I switched to Java and murmur3 hashing - there were still collisions (I don't have my notes in front of me, but I think it was around 1 collision in every 330k hashes, and some papers note that the collisions may even be useful - although targeted collisions like MinHash would obviously more useful).\nThat allowed both easy MP and also typed structures (using Trove) so that I could fit it all in memory on a 16GB machine (I don't recall how much, but I think 9GB).</p>\n\n<p>That didn't help if I wanted to output to CSV and load it into R of course.</p>\n\n<p>I need to go back through my code and see where I went wrong - I was excited to see what the winners were doing - but it was what I was doing... but obviously I wasn't if my score is so bad, so I wonder if my ngram code was buggy or the tf-idf side...</p>\n\n<p>Towards the end I even wrote a genetic algorithm layer to take all of the training output (40+ x their variations in RF/XGB/VW/etc) and that worked well - but I never had single models in the .98 range, so was clearly missing out.</p>\n\n<p>There is also a paper out there talking about the value in classification tasks using no alphanumeric chars at all and keeping only the &quot;punctuation&quot; (in HTML/JS that would be not strictly punctuation), but I never saw that get past .94ish.</p>\n\n<p>This was a fun one, but I am disappointed in my performance, in that my background should have helped here - oh well.</p>",
      "rawMarkdown": "When I started running up against memory issues around the string manipulation (or rather mapping in RAM), I switched to Java and murmur3 hashing - there were still collisions (I don't have my notes in front of me, but I think it was around 1 collision in every 330k hashes, and some papers note that the collisions may even be useful - although targeted collisions like MinHash would obviously more useful).\r\nThat allowed both easy MP and also typed structures (using Trove) so that I could fit it all in memory on a 16GB machine (I don't recall how much, but I think 9GB).\r\n\r\nThat didn't help if I wanted to output to CSV and load it into R of course.\r\n\r\nI need to go back through my code and see where I went wrong - I was excited to see what the winners were doing - but it was what I was doing... but obviously I wasn't if my score is so bad, so I wonder if my ngram code was buggy or the tf-idf side...\r\n\r\nTowards the end I even wrote a genetic algorithm layer to take all of the training output (40+ x their variations in RF/XGB/VW/etc) and that worked well - but I never had single models in the .98 range, so was clearly missing out.\r\n\r\n\r\nThere is also a paper out there talking about the value in classification tasks using no alphanumeric chars at all and keeping only the \"punctuation\" (in HTML/JS that would be not strictly punctuation), but I never saw that get past .94ish.\r\n\r\nThis was a fun one, but I am disappointed in my performance, in that my background should have helped here - oh well.",
      "votes": null
    },
    {
      "id": "96334",
      "postDate": "10/15/2015 18:49:02",
      "content": "<p>[quote=NxGTR;96131]</p>\n\n<p>Good thing is you can always submit after competition ends, just keep on the good work :D</p>\n\n<p>[/quote]</p>\n\n<p>You too, I am amazed how well you consistently do, especially for multiple competitions in parallel! That masters badge isn't far off I'm sure of it.</p>\n\n<p>And yes, I am already planning to finish my solution and a write-up, before I have a peek at what everybody else did.</p>",
      "rawMarkdown": "[quote=NxGTR;96131]\r\n\r\nGood thing is you can always submit after competition ends, just keep on the good work :D\r\n\r\n[/quote]\r\n\r\nYou too, I am amazed how well you consistently do, especially for multiple competitions in parallel! That masters badge isn't far off I'm sure of it.\r\n\r\nAnd yes, I am already planning to finish my solution and a write-up, before I have a peek at what everybody else did.",
      "votes": null
    },
    {
      "id": "96391",
      "postDate": "10/16/2015 01:21:15",
      "content": "<p>At the start of the match, I only have a 4GB machine. Then I bought a 16GB machine, but still cannot fit the batch data after restart. I can't do SVD after tfidf, so I gave up xgboost and used LR to fit the sparse matrix, which may be a wrong choice.</p>",
      "rawMarkdown": "At the start of the match, I only have a 4GB machine. Then I bought a 16GB machine, but still cannot fit the batch data after restart. I can't do SVD after tfidf, so I gave up xgboost and used LR to fit the sparse matrix, which may be a wrong choice.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 96062,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "10/14/2015 03:01:03",
      "content": "<p>Thank you NXGTR. And yeah let's have some weird thing happen!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96065,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "10/14/2015 03:29:37",
      "content": "<p>I thought I could get us a top 10 in physics, but that one is so weird, we droped from 7 to 13 during the last day even though we improved the score a lot... </p>\n\n<p>For this competition, NxGTR really did great job showing how much a single model with simple feature engineering can do, some times simple methods are more valuable.  One disapointment is that GLC boost tree keep crushing because of memory error even though we have more than enough memory, it is basically the same as xgboost, not sure what happened. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96129,
      "author_name": "tobycheese",
      "author_url": "",
      "post_date": "10/14/2015 17:33:01",
      "content": "<p>I feel your pain. I'm having spacy crash as well... which means my NLP-based features won't make it into the model. Of course it's my own fault that I have started so late and now run out of time, but still.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96131,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "10/14/2015 18:19:09",
      "content": "<p>[quote=Toby Cheese;96129]\nI feel your pain. I'm having spacy crash as well... which means my NLP-based features won't make it into the model. Of course it's my own fault that I have started so late and now run out of time, but still.\n[/quote]</p>\n\n<p>Good thing is you can always submit after competition ends, just keep on the good work :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96151,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "10/14/2015 20:31:44",
      "content": "<p>[quote=NxGTR;96131]</p>\n\n<p>[quote=Toby Cheese;96129]\nI feel your pain. I'm having spacy crash as well... which means my NLP-based features won't make it into the model. Of course it's my own fault that I have started so late and now run out of time, but still.\n[/quote]</p>\n\n<p>Good thing is you can always submit after competition ends, just keep on the good work :D</p>\n\n<p>[/quote]</p>\n\n<p>@NxGTR, if our team had invited yours for a merge before the deadline, would you guys say yes?  </p>\n\n<p>It's more and more difficult to end up in top 10 now without a large team (certainly not the case for mortehu). I am considering joining a big team of more than 5 people from now on in case we are competing for money..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96153,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "10/14/2015 20:43:38",
      "content": "<p>[quote=rcarson;96151]\n@NxGTR, if our team had invited yours for a merge before the deadline, would you guys say yes? <br>\n[/quote]</p>\n\n<p>I would have asked my teammate, but I would say yes, for sure.\nThis is my first Kaggle team up experience, and it was great, looking forward more teamups if chances arise.</p>\n\n<p>Just for fun, I can post a &quot;steal my entry&quot; (after competition finish, I think it breaks no rules), so we can find out how well we could have done.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96175,
      "author_name": "sunsoutfunsout",
      "author_url": "",
      "post_date": "10/15/2015 00:21:22",
      "content": "<p>That was a really fun competition!  I still consider myself a complete novice and it was enjoyable to work with a more raw data set than some of the other competitions I have tried.  It was also semi refreshing to escape from the scripts from a bit.  As much as I enjoy seeing what others are doing, it was fun to just compete without worrying about people copying a well performing script at the end of the competition.</p>\n\n<p>I look forward to seeing some of the innovative approaches of the top performers!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96250,
      "author_name": "",
      "author_url": "",
      "post_date": "10/15/2015 09:17:03",
      "content": "<p>Hey, NxGTR, I remember You achieving .98-ish scores right at the start of the competition. Wouldn't You like to share some details of Your approach ;)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96258,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "10/15/2015 10:59:28",
      "content": "<p>You get so close. Hope we can both make top 10 in next match. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96288,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "10/15/2015 15:10:18",
      "content": "<p>[quote=M.E.;96250]\nHey, NxGTR, I remember You achieving .98-ish scores right at the start of the competition. Wouldn't You like to share some details of Your approach ;)?\n[/quote]</p>\n\n<p>It was nothing fancy really (pretty much what everybody did, count and count things), probably thats the reason my score pretty much remained the same from the start to finish hehehe, the only &quot;semi-fancy&quot; is that I have very limited RAM (8GB), so all the feature engineering is done online with multiple passes.</p>\n\n<p>Features were divided in:</p>\n\n<p>Shallow: Counts of sections, e.g. instead of just count individual words/chars as features, I included specific things like what kind of metadata was included in the page, and so on.</p>\n\n<p>Deep(?): I tried to build some kick-ass features such as &quot;Does this webpage uses new JS standards?&quot;, very specific to this context.</p>\n\n<p>Feature number: Before the restart, 10M, After the restart, 8M, I got lended a machine with 16GB RAM, but never managed to do tfid, so they are just plain counts.</p>\n\n<p>Training:\nPretty much the winner was Xgboost, with deep=19, it took I think over 20 Hrs to train, I never managed to fit it memory, so I chunked the data or remove to as few as 250K features to fit it.</p>\n\n<p>Memory was always an issue, and after the restart, it was much worse for me, it was just too much data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96321,
      "author_name": "omgponies",
      "author_url": "",
      "post_date": "10/15/2015 17:30:47",
      "content": "<p>When I started running up against memory issues around the string manipulation (or rather mapping in RAM), I switched to Java and murmur3 hashing - there were still collisions (I don't have my notes in front of me, but I think it was around 1 collision in every 330k hashes, and some papers note that the collisions may even be useful - although targeted collisions like MinHash would obviously more useful).\nThat allowed both easy MP and also typed structures (using Trove) so that I could fit it all in memory on a 16GB machine (I don't recall how much, but I think 9GB).</p>\n\n<p>That didn't help if I wanted to output to CSV and load it into R of course.</p>\n\n<p>I need to go back through my code and see where I went wrong - I was excited to see what the winners were doing - but it was what I was doing... but obviously I wasn't if my score is so bad, so I wonder if my ngram code was buggy or the tf-idf side...</p>\n\n<p>Towards the end I even wrote a genetic algorithm layer to take all of the training output (40+ x their variations in RF/XGB/VW/etc) and that worked well - but I never had single models in the .98 range, so was clearly missing out.</p>\n\n<p>There is also a paper out there talking about the value in classification tasks using no alphanumeric chars at all and keeping only the &quot;punctuation&quot; (in HTML/JS that would be not strictly punctuation), but I never saw that get past .94ish.</p>\n\n<p>This was a fun one, but I am disappointed in my performance, in that my background should have helped here - oh well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96334,
      "author_name": "tobycheese",
      "author_url": "",
      "post_date": "10/15/2015 18:49:02",
      "content": "<p>[quote=NxGTR;96131]</p>\n\n<p>Good thing is you can always submit after competition ends, just keep on the good work :D</p>\n\n<p>[/quote]</p>\n\n<p>You too, I am amazed how well you consistently do, especially for multiple competitions in parallel! That masters badge isn't far off I'm sure of it.</p>\n\n<p>And yes, I am already planning to finish my solution and a write-up, before I have a peek at what everybody else did.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96391,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "10/16/2015 01:21:15",
      "content": "<p>At the start of the match, I only have a 4GB machine. Then I bought a 16GB machine, but still cannot fit the batch data after restart. I can't do SVD after tfidf, so I gave up xgboost and used LR to fit the sparse matrix, which may be a wrong choice.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "96058": "Before competition ends and everybody disappear, I just wanted to have the chance to thank all the incredible and unexpected support that for some reason I received in this competition. Several people contacted me by private message with the sole purpose to wish me luck, something I really appreciate.\r\n\r\nUnless something weird happen to the LB I feel that I let some people down, but do not worry, there is more from NxGTR, just wait for it :P\r\n\r\nI feel very sad, not for me or the competition performance, but for the poor guys at Deloitte hahaha :D, I am sorry guys, I guess you will have to miss my super skills :D\r\n\r\nGreat work from my team mate!, and you guys as always with outstanding performance.\r\n\r\nSee you around!, I always mess in everything :D",
    "96062": "Thank you NXGTR. And yeah let's have some weird thing happen!",
    "96065": "I thought I could get us a top 10 in physics, but that one is so weird, we droped from 7 to 13 during the last day even though we improved the score a lot... \r\n\r\nFor this competition, NxGTR really did great job showing how much a single model with simple feature engineering can do, some times simple methods are more valuable.  One disapointment is that GLC boost tree keep crushing because of memory error even though we have more than enough memory, it is basically the same as xgboost, not sure what happened.",
    "96129": "I feel your pain. I'm having spacy crash as well... which means my NLP-based features won't make it into the model. Of course it's my own fault that I have started so late and now run out of time, but still.",
    "96131": "[quote=Toby Cheese;96129]\r\nI feel your pain. I'm having spacy crash as well... which means my NLP-based features won't make it into the model. Of course it's my own fault that I have started so late and now run out of time, but still.\r\n[/quote]\r\n\r\nGood thing is you can always submit after competition ends, just keep on the good work :D",
    "96151": "[quote=NxGTR;96131]\r\n\r\n[quote=Toby Cheese;96129]\r\nI feel your pain. I'm having spacy crash as well... which means my NLP-based features won't make it into the model. Of course it's my own fault that I have started so late and now run out of time, but still.\r\n[/quote]\r\n\r\nGood thing is you can always submit after competition ends, just keep on the good work :D\r\n\r\n[/quote]\r\n\r\n@NxGTR, if our team had invited yours for a merge before the deadline, would you guys say yes?  \r\n\r\nIt's more and more difficult to end up in top 10 now without a large team (certainly not the case for mortehu). I am considering joining a big team of more than 5 people from now on in case we are competing for money..",
    "96153": "[quote=rcarson;96151]\r\n@NxGTR, if our team had invited yours for a merge before the deadline, would you guys say yes?  \r\n[/quote]\r\n\r\nI would have asked my teammate, but I would say yes, for sure.\r\nThis is my first Kaggle team up experience, and it was great, looking forward more teamups if chances arise.\r\n\r\nJust for fun, I can post a \"steal my entry\" (after competition finish, I think it breaks no rules), so we can find out how well we could have done.",
    "96175": "That was a really fun competition!  I still consider myself a complete novice and it was enjoyable to work with a more raw data set than some of the other competitions I have tried.  It was also semi refreshing to escape from the scripts from a bit.  As much as I enjoy seeing what others are doing, it was fun to just compete without worrying about people copying a well performing script at the end of the competition.\r\n\r\nI look forward to seeing some of the innovative approaches of the top performers!",
    "96250": "Hey, NxGTR, I remember You achieving .98-ish scores right at the start of the competition. Wouldn't You like to share some details of Your approach ;)?",
    "96258": "You get so close. Hope we can both make top 10 in next match. :)",
    "96288": "[quote=M.E.;96250]\r\nHey, NxGTR, I remember You achieving .98-ish scores right at the start of the competition. Wouldn't You like to share some details of Your approach ;)?\r\n[/quote]\r\n\r\nIt was nothing fancy really (pretty much what everybody did, count and count things), probably thats the reason my score pretty much remained the same from the start to finish hehehe, the only \"semi-fancy\" is that I have very limited RAM (8GB), so all the feature engineering is done online with multiple passes.\r\n\r\nFeatures were divided in:\r\n\r\nShallow: Counts of sections, e.g. instead of just count individual words/chars as features, I included specific things like what kind of metadata was included in the page, and so on.\r\n\r\nDeep(?): I tried to build some kick-ass features such as \"Does this webpage uses new JS standards?\", very specific to this context.\r\n\r\nFeature number: Before the restart, 10M, After the restart, 8M, I got lended a machine with 16GB RAM, but never managed to do tfid, so they are just plain counts.\r\n\r\nTraining:\r\nPretty much the winner was Xgboost, with deep=19, it took I think over 20 Hrs to train, I never managed to fit it memory, so I chunked the data or remove to as few as 250K features to fit it.\r\n\r\nMemory was always an issue, and after the restart, it was much worse for me, it was just too much data.",
    "96321": "When I started running up against memory issues around the string manipulation (or rather mapping in RAM), I switched to Java and murmur3 hashing - there were still collisions (I don't have my notes in front of me, but I think it was around 1 collision in every 330k hashes, and some papers note that the collisions may even be useful - although targeted collisions like MinHash would obviously more useful).\r\nThat allowed both easy MP and also typed structures (using Trove) so that I could fit it all in memory on a 16GB machine (I don't recall how much, but I think 9GB).\r\n\r\nThat didn't help if I wanted to output to CSV and load it into R of course.\r\n\r\nI need to go back through my code and see where I went wrong - I was excited to see what the winners were doing - but it was what I was doing... but obviously I wasn't if my score is so bad, so I wonder if my ngram code was buggy or the tf-idf side...\r\n\r\nTowards the end I even wrote a genetic algorithm layer to take all of the training output (40+ x their variations in RF/XGB/VW/etc) and that worked well - but I never had single models in the .98 range, so was clearly missing out.\r\n\r\n\r\nThere is also a paper out there talking about the value in classification tasks using no alphanumeric chars at all and keeping only the \"punctuation\" (in HTML/JS that would be not strictly punctuation), but I never saw that get past .94ish.\r\n\r\nThis was a fun one, but I am disappointed in my performance, in that my background should have helped here - oh well.",
    "96334": "[quote=NxGTR;96131]\r\n\r\nGood thing is you can always submit after competition ends, just keep on the good work :D\r\n\r\n[/quote]\r\n\r\nYou too, I am amazed how well you consistently do, especially for multiple competitions in parallel! That masters badge isn't far off I'm sure of it.\r\n\r\nAnd yes, I am already planning to finish my solution and a write-up, before I have a peek at what everybody else did.",
    "96391": "At the start of the match, I only have a 4GB machine. Then I bought a 16GB machine, but still cannot fit the batch data after restart. I can't do SVD after tfidf, so I gave up xgboost and used LR to fit the sparse matrix, which may be a wrong choice."
  },
  "source": "meta"
}