{
  "id": 16626,
  "title": "Beat the Benchmark 0.90388 with simple model",
  "url": "/competitions/dato-native/discussion/16626",
  "author_name": "",
  "post_date": "2015-09-23T03:32:10.007Z",
  "votes": 45,
  "comment_count": 42,
  "views": 9805,
  "content": "<p>In response <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16623/this-competition-is-not-popular-at-all\">to this post</a> and to drum up a little more interest in this competition, I'm posting some starter code that performs decently (0.90388), needs quite a bit of work, leaves room for a lot of improvement, and it is a really basic model (just 7 features with RandomForest), showing that many competitors are ignoring some really basic features.  It runs under 8 minutes from start to finish on a modern Mac Book Pro (uses multiprocessing, so runs 8 processes in parallel).  Just in case you're wondering, I'm not a proponent of high performance btb code this close to competition end, but really, just 7 basic count features?</p>\n\n<p>Enjoy!</p>\n\n<p>My feature set:</p>\n\n<pre><code>values['lines'] = text.count('\\n')\nvalues['spaces'] = text.count(' ')\nvalues['tabs'] = text.count('\\t')\nvalues['braces'] = text.count('{')\nvalues['brackets'] = text.count('[')\nvalues['words'] = len(re.split('\\s+', text))\nvalues['length'] = len(text)\n</code></pre>",
  "messages": [
    {
      "id": "93238",
      "postDate": "09/23/2015 03:32:10",
      "content": "<p>In response <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16623/this-competition-is-not-popular-at-all\">to this post</a> and to drum up a little more interest in this competition, I'm posting some starter code that performs decently (0.90388), needs quite a bit of work, leaves room for a lot of improvement, and it is a really basic model (just 7 features with RandomForest), showing that many competitors are ignoring some really basic features.  It runs under 8 minutes from start to finish on a modern Mac Book Pro (uses multiprocessing, so runs 8 processes in parallel).  Just in case you're wondering, I'm not a proponent of high performance btb code this close to competition end, but really, just 7 basic count features?</p>\n\n<p>Enjoy!</p>\n\n<p>My feature set:</p>\n\n<pre><code>values['lines'] = text.count('\\n')\nvalues['spaces'] = text.count(' ')\nvalues['tabs'] = text.count('\\t')\nvalues['braces'] = text.count('{')\nvalues['brackets'] = text.count('[')\nvalues['words'] = len(re.split('\\s+', text))\nvalues['length'] = len(text)\n</code></pre>",
      "rawMarkdown": "In response [to this post][1] and to drum up a little more interest in this competition, I'm posting some starter code that performs decently (0.90388), needs quite a bit of work, leaves room for a lot of improvement, and it is a really basic model (just 7 features with RandomForest), showing that many competitors are ignoring some really basic features.  It runs under 8 minutes from start to finish on a modern Mac Book Pro (uses multiprocessing, so runs 8 processes in parallel).  Just in case you're wondering, I'm not a proponent of high performance btb code this close to competition end, but really, just 7 basic count features?\r\n\r\nEnjoy!\r\n\r\nMy feature set:\r\n\r\n    values['lines'] = text.count('\\n')\r\n    values['spaces'] = text.count(' ')\r\n    values['tabs'] = text.count('\\t')\r\n    values['braces'] = text.count('{')\r\n    values['brackets'] = text.count('[')\r\n    values['words'] = len(re.split('\\s+', text))\r\n    values['length'] = len(text)\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16623/this-competition-is-not-popular-at-all",
      "votes": null
    },
    {
      "id": "93244",
      "postDate": "09/23/2015 04:19:09",
      "content": "<p>It is impressive how much more (I guess) complex the solution needs to be for less than 0.1 AUC gain :/</p>",
      "rawMarkdown": "It is impressive how much more (I guess) complex the solution needs to be for less than 0.1 AUC gain :/",
      "votes": null
    },
    {
      "id": "93249",
      "postDate": "09/23/2015 04:46:43",
      "content": "<p>0.90388 with this? </p>\n\n<p>David, the simplicity and performance of this code is impressive!</p>\n\n<p>This submission would automatically earn a new entrant 2920 Kaggle points if the competition were to end with current standings - a steal! :)</p>",
      "rawMarkdown": "0.90388 with this? \r\n\r\nDavid, the simplicity and performance of this code is impressive!\r\n\r\nThis submission would automatically earn a new entrant 2920 Kaggle points if the competition were to end with current standings - a steal! :)",
      "votes": null
    },
    {
      "id": "93251",
      "postDate": "09/23/2015 05:20:37",
      "content": "<p>Truly impressive!</p>",
      "rawMarkdown": "Truly impressive!",
      "votes": null
    },
    {
      "id": "93258",
      "postDate": "09/23/2015 07:42:46",
      "content": "<p>I'm really impressed about how well such basic feature selection can perform at detecting these native ads. I browsed a little about stumbleupon native ads system at the beginning of this competition and couldn't even imagine that pages you put through their ads system could be identified as sponsored since it doesn't mess up with the html code and it's just a URL that stumbleupon would show to targeted audiences.</p>",
      "rawMarkdown": "I'm really impressed about how well such basic feature selection can perform at detecting these native ads. I browsed a little about stumbleupon native ads system at the beginning of this competition and couldn't even imagine that pages you put through their ads system could be identified as sponsored since it doesn't mess up with the html code and it's just a URL that stumbleupon would show to targeted audiences.",
      "votes": null
    },
    {
      "id": "93288",
      "postDate": "09/23/2015 15:41:02",
      "content": "<p>Now 0.90388 is the new zero, Very impressive, David :)\n[quote=David Shinn;93238]</p>\n\n<p>In response <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16623/this-competition-is-not-popular-at-all\">to this post</a> and to drum up a little more interest in this competition, I'm posting some starter code that performs decently (0.90388), needs quite a bit of work, leaves room for a lot of improvement, and it is a really basic model (just 7 features with RandomForest), showing that many competitors are ignoring some really basic features.  It runs under 8 minutes from start to finish on a modern Mac Book Pro (uses multiprocessing, so runs 8 processes in parallel).  Just in case you're wondering, I'm not a proponent of high performance btb code this close to competition end, but really, just 7 basic count features?</p>\n\n<p>Enjoy!</p>\n\n<p>My feature set:</p>\n\n<pre><code>values['lines'] = text.count('\\n')\nvalues['spaces'] = text.count(' ')\nvalues['tabs'] = text.count('\\t')\nvalues['braces'] = text.count('{')\nvalues['brackets'] = text.count('[')\nvalues['words'] = len(re.split('\\s+', text))\nvalues['length'] = len(text)\n</code></pre>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Now 0.90388 is the new zero, Very impressive, David :)\r\n[quote=David Shinn;93238]\r\n\r\nIn response [to this post][1] and to drum up a little more interest in this competition, I'm posting some starter code that performs decently (0.90388), needs quite a bit of work, leaves room for a lot of improvement, and it is a really basic model (just 7 features with RandomForest), showing that many competitors are ignoring some really basic features.  It runs under 8 minutes from start to finish on a modern Mac Book Pro (uses multiprocessing, so runs 8 processes in parallel).  Just in case you're wondering, I'm not a proponent of high performance btb code this close to competition end, but really, just 7 basic count features?\r\n\r\nEnjoy!\r\n\r\nMy feature set:\r\n\r\n    values['lines'] = text.count('\\n')\r\n    values['spaces'] = text.count(' ')\r\n    values['tabs'] = text.count('\\t')\r\n    values['braces'] = text.count('{')\r\n    values['brackets'] = text.count('[')\r\n    values['words'] = len(re.split('\\s+', text))\r\n    values['length'] = len(text)\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16623/this-competition-is-not-popular-at-all\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "93326",
      "postDate": "09/23/2015 21:15:12",
      "content": "<p>Thank you for the example code.  I'm just curious, the choice of .imap() to do the threading versus other alternatives such as apply_async() and map_async()...  What made you pick .imap() and why does it seem to be running faster than some other analysis code that I wrote using apply_async()?  I found <a href=\"https://stackoverflow.com/questions/26520781/python-multiprocessing-poolwhats-the-difference-between-map-async-and-imap\">this on Stackoverflow</a> but I'm not sure I understand the details.</p>",
      "rawMarkdown": "Thank you for the example code.  I'm just curious, the choice of .imap() to do the threading versus other alternatives such as apply_async() and map_async()...  What made you pick .imap() and why does it seem to be running faster than some other analysis code that I wrote using apply_async()?  I found [this on Stackoverflow](https://stackoverflow.com/questions/26520781/python-multiprocessing-poolwhats-the-difference-between-map-async-and-imap) but I'm not sure I understand the details.",
      "votes": null
    },
    {
      "id": "93331",
      "postDate": "09/23/2015 22:46:01",
      "content": "<p>It's just what has worked.  I've always used <a href=\"http://stackoverflow.com/a/5666996/2626968\">this post</a> as a guide, something about the counter counting chunks, not actual elements.  I use imap so I can see progress.  Otherwise I just use map.  Still learning, of course.</p>",
      "rawMarkdown": "It's just what has worked.  I've always used [this post][1] as a guide, something about the counter counting chunks, not actual elements.  I use imap so I can see progress.  Otherwise I just use map.  Still learning, of course.\r\n\r\n\r\n  [1]: http://stackoverflow.com/a/5666996/2626968",
      "votes": null
    },
    {
      "id": "93347",
      "postDate": "09/24/2015 03:44:21",
      "content": "<p>Thanks a lot @David Shinn. This is very helpful. </p>",
      "rawMarkdown": "Thanks a lot @David Shinn. This is very helpful.",
      "votes": null
    },
    {
      "id": "93386",
      "postDate": "09/24/2015 21:23:33",
      "content": "<p>Beautiful!</p>",
      "rawMarkdown": "Beautiful!",
      "votes": null
    },
    {
      "id": "93826",
      "postDate": "10/01/2015 20:25:36",
      "content": "<p>Has anyone had luck modifying this to run on Windows?  Since you can't fork under Windows, I've had a tough time adjusting this so the processing/working function can use the train_keys list without having to pass the entire thing as a parameter.  Works great on Linux. :)</p>",
      "rawMarkdown": "Has anyone had luck modifying this to run on Windows?  Since you can't fork under Windows, I've had a tough time adjusting this so the processing/working function can use the train_keys list without having to pass the entire thing as a parameter.  Works great on Linux. :)",
      "votes": null
    },
    {
      "id": "93932",
      "postDate": "10/03/2015 02:10:17",
      "content": "<p>I only got an advanced windows PC. Looks like the script won't run through. Any suggestions on modifying this would be really helpful!</p>",
      "rawMarkdown": "I only got an advanced windows PC. Looks like the script won't run through. Any suggestions on modifying this would be really helpful!",
      "votes": null
    },
    {
      "id": "93933",
      "postDate": "10/03/2015 02:33:50",
      "content": "<p>[quote=paladin_o_newengland;93932]</p>\n\n<p>Any suggestions on modifying this would be really helpful!</p>\n\n<p>[/quote]</p>\n\n<p>Take this code:</p>\n\n<pre><code>p = multiprocessing.Pool()\nresults = p.imap(create_data, filepaths)\nwhile (True):\n    completed = results._index\n    print(&quot;\\r--- Completed {:,} out of {:,}&quot;.format(completed, num_tasks), end='')\n    sys.stdout.flush()\n    time.sleep(1)\n    if (completed == num_tasks): break\np.close()\np.join()\n</code></pre>\n\n<p>and replace it with:</p>\n\n<pre><code>results = map(create_data, filepaths)\n</code></pre>\n\n<p>It'll work, but you won't get parallel processing.</p>",
      "rawMarkdown": "[quote=paladin_o_newengland;93932]\r\n\r\nAny suggestions on modifying this would be really helpful!\r\n\r\n[/quote]\r\n\r\nTake this code:\r\n\r\n    p = multiprocessing.Pool()\r\n    results = p.imap(create_data, filepaths)\r\n    while (True):\r\n        completed = results._index\r\n        print(\"\\r--- Completed {:,} out of {:,}\".format(completed, num_tasks), end='')\r\n        sys.stdout.flush()\r\n        time.sleep(1)\r\n        if (completed == num_tasks): break\r\n    p.close()\r\n    p.join()\r\n\r\nand replace it with:\r\n\r\n    results = map(create_data, filepaths)\r\n\r\nIt'll work, but you won't get parallel processing.",
      "votes": null
    },
    {
      "id": "95052",
      "postDate": "10/04/2015 17:43:45",
      "content": "<p>I ran this code but for some reason all my test features are null. I am able to get the train features but not for the test. Any reason why?</p>",
      "rawMarkdown": "I ran this code but for some reason all my test features are null. I am able to get the train features but not for the test. Any reason why?",
      "votes": null
    },
    {
      "id": "95053",
      "postDate": "10/04/2015 17:56:48",
      "content": "<p>when i run this script i am getting this error : &quot;AttributeError: 'DataFrame' object has no attribute 'sponsored'&quot; , can you tell how to resolve this error.</p>",
      "rawMarkdown": "when i run this script i am getting this error : \"AttributeError: 'DataFrame' object has no attribute 'sponsored'\" , can you tell how to resolve this error.",
      "votes": null
    },
    {
      "id": "95066",
      "postDate": "10/04/2015 19:28:08",
      "content": "<p>[quote=ask788;95052]</p>\n\n<p>I ran this code but for some reason all my test features are null. I am able to get the train features but not for the test. Any reason why?</p>\n\n<p>[/quote]</p>\n\n<p>That's weird, I specifically set all test null values to 0, so you should at least have a data frame full of zeros.</p>\n\n<p>[quote=nknm;95053]</p>\n\n<p>when i run this script i am getting this error : &quot;AttributeError: 'DataFrame' object has no attribute 'sponsored'&quot; , can you tell how to resolve this error.</p>\n\n<p>[/quote]</p>\n\n<p>This probably means that the df_full dataframe does not have the column &quot;sponsored&quot;, which means either you took it out of the create_data function or something else is wrong.</p>\n\n<p>In general, I find it difficult to debug code that's been wrapped around multiprocessing, because it won't complain for code that breaks within the multiprocessing code until after it's all done, which is painful for this type of feature extraction.  The best thing you can do is</p>\n\n<ol>\n<li>Turn it into a non-multiprocessing code with the tip I added before, and</li>\n<li>Turn this line <code>filepaths = glob.glob('data/*/*.txt')</code> into <code>filepaths = glob.glob('data/0/*.txt')[:1000] + glob.glob('data/5/*.txt')[:1000]</code>for debugging purposes.  The code won't give you a proper submission file, but it should run through the feature extraction, training, and predicting sections quickly for you to experiment and figure out what's wrong with your code (or to add new features, etc).</li>\n</ol>\n\n<p>Isn't the learning process fun?</p>",
      "rawMarkdown": "[quote=ask788;95052]\r\n\r\nI ran this code but for some reason all my test features are null. I am able to get the train features but not for the test. Any reason why?\r\n\r\n[/quote]\r\n\r\nThat's weird, I specifically set all test null values to 0, so you should at least have a data frame full of zeros.\r\n\r\n[quote=nknm;95053]\r\n\r\nwhen i run this script i am getting this error : \"AttributeError: 'DataFrame' object has no attribute 'sponsored'\" , can you tell how to resolve this error.\r\n\r\n\r\n[/quote]\r\n\r\nThis probably means that the df_full dataframe does not have the column \"sponsored\", which means either you took it out of the create_data function or something else is wrong.\r\n\r\nIn general, I find it difficult to debug code that's been wrapped around multiprocessing, because it won't complain for code that breaks within the multiprocessing code until after it's all done, which is painful for this type of feature extraction.  The best thing you can do is\r\n\r\n 1. Turn it into a non-multiprocessing code with the tip I added before, and\r\n 2. Turn this line `filepaths = glob.glob('data/*/*.txt')` into `filepaths = glob.glob('data/0/*.txt')[:1000] + glob.glob('data/5/*.txt')[:1000]`for debugging purposes.  The code won't give you a proper submission file, but it should run through the feature extraction, training, and predicting sections quickly for you to experiment and figure out what's wrong with your code (or to add new features, etc).\r\n\r\nIsn't the learning process fun?",
      "votes": null
    },
    {
      "id": "95072",
      "postDate": "10/04/2015 20:21:50",
      "content": "<p>Hello Dave Shinn,</p>\n\n<p>Thank you for posting this, I hope to get it to run in the next few days. My error is num_tasks is staying at zero because the filepaths line (filepaths = glob.glob('data/<em>/</em>.txt')) isn't identifying the txt file. My data folder looks like the attached image below. Is there a way to extract the train TXT files from train_v2.csv that I am missing?</p>\n\n<p>Thank you,</p>",
      "rawMarkdown": "Hello Dave Shinn,\r\n\r\nThank you for posting this, I hope to get it to run in the next few days. My error is num_tasks is staying at zero because the filepaths line (filepaths = glob.glob('data/*/*.txt')) isn't identifying the txt file. My data folder looks like the attached image below. Is there a way to extract the train TXT files from train_v2.csv that I am missing?\r\n\r\nThank you,",
      "votes": null
    },
    {
      "id": "95074",
      "postDate": "10/04/2015 20:29:41",
      "content": "<p>[quote=psingman;95072]</p>\n\n<p>Hello Dave Shinn,</p>\n\n<p>Thank you for posting this, I hope to get it to run in the next few days. My error is num_tasks is staying at zero because the filepaths line (filepaths = glob.glob('data/<em>/</em>.txt')) isn't identifying the txt file. My data folder looks like the attached image below. Is there a way to extract the train TXT files from train_v2.csv that I am missing?</p>\n\n<p>Thank you,</p>\n\n<p>[/quote]</p>\n\n<p>I'm pretty sure that glob.glob needs a wild card, like: <code>filepaths = glob.glob('data/*/*.txt')</code>.</p>",
      "rawMarkdown": "[quote=psingman;95072]\r\n\r\nHello Dave Shinn,\r\n\r\nThank you for posting this, I hope to get it to run in the next few days. My error is num_tasks is staying at zero because the filepaths line (filepaths = glob.glob('data/*/*.txt')) isn't identifying the txt file. My data folder looks like the attached image below. Is there a way to extract the train TXT files from train_v2.csv that I am missing?\r\n\r\nThank you,\r\n\r\n[/quote]\r\n\r\nI'm pretty sure that glob.glob needs a wild card, like: `filepaths = glob.glob('data/*/*.txt')`.",
      "votes": null
    },
    {
      "id": "95075",
      "postDate": "10/04/2015 20:36:27",
      "content": "<p>[quote=psingman;95072]</p>\n\n<p>Hello Dave Shinn,</p>\n\n<p>Thank you for posting this, I hope to get it to run in the next few days. My error is num_tasks is staying at zero because the filepaths line (filepaths = glob.glob('data/<em>/</em>.txt')) isn't identifying the txt file. My data folder looks like the attached image below. Is there a way to extract the train TXT files from train_v2.csv that I am missing?</p>\n\n<p>Thank you,</p>\n\n<p>[/quote]</p>\n\n<p>Okay, just realized your problem.  You need to download all 6 zip files (0.zip, 1.zip, etc), then extract them so their contents reside under data/0/, data/1/, data/2/, etc.  There's over 400K files.</p>",
      "rawMarkdown": "[quote=psingman;95072]\r\n\r\nHello Dave Shinn,\r\n\r\nThank you for posting this, I hope to get it to run in the next few days. My error is num_tasks is staying at zero because the filepaths line (filepaths = glob.glob('data/*/*.txt')) isn't identifying the txt file. My data folder looks like the attached image below. Is there a way to extract the train TXT files from train_v2.csv that I am missing?\r\n\r\nThank you,\r\n\r\n[/quote]\r\n\r\nOkay, just realized your problem.  You need to download all 6 zip files (0.zip, 1.zip, etc), then extract them so their contents reside under data/0/, data/1/, data/2/, etc.  There's over 400K files.",
      "votes": null
    },
    {
      "id": "95091",
      "postDate": "10/05/2015 00:56:52",
      "content": "<p>I tried to run the benchmark code on window python. The below multiprocessing.Pool() step is seems to be running forever. This is 4th day it is still running. I have a pretty decent machine 32 GB RAM and I7 still not able to pass thru this step. Any idea how to make this code run fast?</p>\n\n<p>p = multiprocessing.Pool()\nresults = p.imap(create_data, filepaths)\nwhile (True):\n    completed = results._index\n    print(&quot;\\r--- Completed {:,} out of {:,}&quot;.format(completed, num_tasks), end='')\n    sys.stdout.flush()\n    time.sleep(1)\n    if (completed == num_tasks): break\np.close()\np.join()\ndf_full = pd.DataFrame(list(results))\nprint()</p>",
      "rawMarkdown": "I tried to run the benchmark code on window python. The below multiprocessing.Pool() step is seems to be running forever. This is 4th day it is still running. I have a pretty decent machine 32 GB RAM and I7 still not able to pass thru this step. Any idea how to make this code run fast?\r\n\r\np = multiprocessing.Pool()\r\nresults = p.imap(create_data, filepaths)\r\nwhile (True):\r\n    completed = results._index\r\n    print(\"\\r--- Completed {:,} out of {:,}\".format(completed, num_tasks), end='')\r\n    sys.stdout.flush()\r\n    time.sleep(1)\r\n    if (completed == num_tasks): break\r\np.close()\r\np.join()\r\ndf_full = pd.DataFrame(list(results))\r\nprint()",
      "votes": null
    },
    {
      "id": "95094",
      "postDate": "10/05/2015 01:19:15",
      "content": "<p>[quote=Vikrant Kumar;95091]</p>\n\n<p>I tried to run the benchmark code on window python. The below multiprocessing.Pool() step is seems to be running forever. This is 4th day it is still running. I have a pretty decent machine 32 GB RAM and I7 still not able to pass thru this step. Any idea how to make this code run fast?</p>\n\n<p>p = multiprocessing.Pool()\nresults = p.imap(create_data, filepaths)\nwhile (True):\n    completed = results._index\n    print(&quot;\\r--- Completed {:,} out of {:,}&quot;.format(completed, num_tasks), end='')\n    sys.stdout.flush()\n    time.sleep(1)\n    if (completed == num_tasks): break\np.close()\np.join()\ndf_full = pd.DataFrame(list(results))\nprint()</p>\n\n<p>[/quote]</p>\n\n<p>Unfortunately, it is likely a problem with the how multiprocessing works on Windows, where I can't help you.  However, take the tips from <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model/95066#post95066\">this part of the thread</a> and it'll set you up to debug the code in less than a couple of minutes.  My biggest advice here is first prototype something small that you expect to run in a reasonable amount of time (like a few minutes, not 4 days) and lose the multiprocessing.  Once you have the kinks worked out, try to get multiprocessing to work.  If that doesn't work, just forget the multiprocessing.</p>",
      "rawMarkdown": "[quote=Vikrant Kumar;95091]\r\n\r\nI tried to run the benchmark code on window python. The below multiprocessing.Pool() step is seems to be running forever. This is 4th day it is still running. I have a pretty decent machine 32 GB RAM and I7 still not able to pass thru this step. Any idea how to make this code run fast?\r\n\r\np = multiprocessing.Pool()\r\nresults = p.imap(create_data, filepaths)\r\nwhile (True):\r\n    completed = results._index\r\n    print(\"\\r--- Completed {:,} out of {:,}\".format(completed, num_tasks), end='')\r\n    sys.stdout.flush()\r\n    time.sleep(1)\r\n    if (completed == num_tasks): break\r\np.close()\r\np.join()\r\ndf_full = pd.DataFrame(list(results))\r\nprint()\r\n\r\n[/quote]\r\n\r\nUnfortunately, it is likely a problem with the how multiprocessing works on Windows, where I can't help you.  However, take the tips from [this part of the thread][1] and it'll set you up to debug the code in less than a couple of minutes.  My biggest advice here is first prototype something small that you expect to run in a reasonable amount of time (like a few minutes, not 4 days) and lose the multiprocessing.  Once you have the kinks worked out, try to get multiprocessing to work.  If that doesn't work, just forget the multiprocessing.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model/95066#post95066",
      "votes": null
    },
    {
      "id": "95096",
      "postDate": "10/05/2015 02:40:23",
      "content": "<p>Thanks David for response. Any idea where it will run fast? I can install something which can make it run fast like ubuntu or something.</p>",
      "rawMarkdown": "Thanks David for response. Any idea where it will run fast? I can install something which can make it run fast like ubuntu or something.",
      "votes": null
    },
    {
      "id": "95097",
      "postDate": "10/05/2015 02:41:16",
      "content": "<p>Thanks so much david! Replacing the multi processing with simply &quot;results = map(create_data, filepaths)&quot; works great for me. For those of you using windows too, change that would make it work. </p>",
      "rawMarkdown": "Thanks so much david! Replacing the multi processing with simply \"results = map(create_data, filepaths)\" works great for me. For those of you using windows too, change that would make it work.",
      "votes": null
    },
    {
      "id": "95098",
      "postDate": "10/05/2015 02:45:49",
      "content": "<p>Thank you, that definitely was my issue (facepalm). Just to make sure I'm not doing anything unnecessary, you have to unzip all the TXT files in 0.zip, 1.zip etc..? My 130GB laptop is running out of storage....</p>",
      "rawMarkdown": "Thank you, that definitely was my issue (facepalm). Just to make sure I'm not doing anything unnecessary, you have to unzip all the TXT files in 0.zip, 1.zip etc..? My 130GB laptop is running out of storage....",
      "votes": null
    },
    {
      "id": "95099",
      "postDate": "10/05/2015 02:46:05",
      "content": "<p>Hi Paladin \nDo I need to change the whole piece of code with &quot;results = map(create_data, filepaths)&quot; or only the first line &quot;p = multiprocessing.Pool()&quot;. I can see second line already is same as what you suggested. Can you be please precise. I will try it.</p>",
      "rawMarkdown": "Hi Paladin \r\nDo I need to change the whole piece of code with \"results = map(create_data, filepaths)\" or only the first line \"p = multiprocessing.Pool()\". I can see second line already is same as what you suggested. Can you be please precise. I will try it.",
      "votes": null
    },
    {
      "id": "95100",
      "postDate": "10/05/2015 03:05:58",
      "content": "<p>[quote=Vikrant Kumar;95099]</p>\n\n<p>Hi Paladin \nDo I need to change the whole piece of code with &quot;results = map(create_data, filepaths)&quot; or only the first line &quot;p = multiprocessing.Pool()&quot;. I can see second line already is same as what you suggested. Can you be please precise. I will try it.</p>\n\n<p>[/quote]</p>\n\n<p>I believe just change the results = map... line.  The other line sets up the processing pool which should be the same regardless of you using map, imap, etc.  See the <a href=\"https://docs.python.org/2/library/multiprocessing.html#module-multiprocessing.pool\">documentation here</a></p>",
      "rawMarkdown": "[quote=Vikrant Kumar;95099]\r\n\r\nHi Paladin \r\nDo I need to change the whole piece of code with \"results = map(create_data, filepaths)\" or only the first line \"p = multiprocessing.Pool()\". I can see second line already is same as what you suggested. Can you be please precise. I will try it.\r\n\r\n[/quote]\r\n\r\nI believe just change the results = map... line.  The other line sets up the processing pool which should be the same regardless of you using map, imap, etc.  See the [documentation here](https://docs.python.org/2/library/multiprocessing.html#module-multiprocessing.pool)",
      "votes": null
    },
    {
      "id": "95146",
      "postDate": "10/05/2015 14:16:21",
      "content": "<p>[quote=Vikrant Kumar;95099]</p>\n\n<p>Hi Paladin \nDo I need to change the whole piece of code with &quot;results = map(create_data, filepaths)&quot; or only the first line &quot;p = multiprocessing.Pool()&quot;. I can see second line already is same as what you suggested. Can you be please precise. I will try it.</p>\n\n<p>[/quote]</p>\n\n<p>I actually did just as David suggested in his response last page, use &quot;results = map(create_data, filepaths)&quot; to replace the whole thing below: (I didn't try other ways. Note running it will take time. Thanks. ) </p>\n\n<p>p = multiprocessing.Pool()\nresults = p.imap(create_data, filepaths)\nwhile (True):\n    completed = results._index\n    print(&quot;\\r--- Completed {:,} out of {:,}&quot;.format(completed, num_tasks), end='')\n    sys.stdout.flush()\n    time.sleep(1)\n    if (completed == num_tasks): break\np.close()\np.join()</p>",
      "rawMarkdown": "[quote=Vikrant Kumar;95099]\r\n\r\nHi Paladin \r\nDo I need to change the whole piece of code with \"results = map(create_data, filepaths)\" or only the first line \"p = multiprocessing.Pool()\". I can see second line already is same as what you suggested. Can you be please precise. I will try it.\r\n\r\n[/quote]\r\n\r\nI actually did just as David suggested in his response last page, use \"results = map(create_data, filepaths)\" to replace the whole thing below: (I didn't try other ways. Note running it will take time. Thanks. ) \r\n\r\np = multiprocessing.Pool()\r\nresults = p.imap(create_data, filepaths)\r\nwhile (True):\r\n    completed = results._index\r\n    print(\"\\r--- Completed {:,} out of {:,}\".format(completed, num_tasks), end='')\r\n    sys.stdout.flush()\r\n    time.sleep(1)\r\n    if (completed == num_tasks): break\r\np.close()\r\np.join()",
      "votes": null
    },
    {
      "id": "95230",
      "postDate": "10/06/2015 05:52:50",
      "content": "<p>[quote=David Shinn;93933]</p>\n\n<p>[quote=paladin_o_newengland;93932]</p>\n\n<p>Any suggestions on modifying this would be really helpful!</p>\n\n<p>[/quote]</p>\n\n<p>Take this code:</p>\n\n<pre><code>p = multiprocessing.Pool()\nresults = p.imap(create_data, filepaths)\nwhile (True):\n    completed = results._index\n    print(&quot;\\r--- Completed {:,} out of {:,}&quot;.format(completed, num_tasks), end='')\n    sys.stdout.flush()\n    time.sleep(1)\n    if (completed == num_tasks): break\np.close()\np.join()\n</code></pre>\n\n<p>and replace it with:</p>\n\n<pre><code>results = map(create_data, filepaths)\n</code></pre>\n\n<p>It'll work, but you won't get parallel processing.</p>\n\n<p>[/quote]\nHi\nI applied this but ended up with error</p>\n\n<hr>\n\n<p>AttributeError                            Traceback (most recent call last)\n in ()\n      3 #results = p.imap(create_data, filepaths)\n      4 while (True):\n----&gt; 5     completed = results._index\n      6     print(&quot;\\r--- Completed {:,} out of {:,}&quot;.format(completed, num_tasks), end='')\n      7     sys.stdout.flush()</p>\n\n<p>AttributeError: 'list' object has no attribute '_index'</p>\n\n<p>Any idea how to resolve this?</p>",
      "rawMarkdown": "[quote=David Shinn;93933]\r\n\r\n[quote=paladin_o_newengland;93932]\r\n\r\nAny suggestions on modifying this would be really helpful!\r\n\r\n[/quote]\r\n\r\nTake this code:\r\n\r\n    p = multiprocessing.Pool()\r\n    results = p.imap(create_data, filepaths)\r\n    while (True):\r\n        completed = results._index\r\n        print(\"\\r--- Completed {:,} out of {:,}\".format(completed, num_tasks), end='')\r\n        sys.stdout.flush()\r\n        time.sleep(1)\r\n        if (completed == num_tasks): break\r\n    p.close()\r\n    p.join()\r\n\r\nand replace it with:\r\n\r\n    results = map(create_data, filepaths)\r\n\r\nIt'll work, but you won't get parallel processing.\r\n\r\n[/quote]\r\nHi\r\nI applied this but ended up with error\r\n\r\n---------------------------------------------------------------------------\r\nAttributeError                            Traceback (most recent call last)\r\n<ipython-input-5-6a305321d825> in <module>()\r\n      3 #results = p.imap(create_data, filepaths)\r\n      4 while (True):\r\n----> 5     completed = results._index\r\n      6     print(\"\\r--- Completed {:,} out of {:,}\".format(completed, num_tasks), end='')\r\n      7     sys.stdout.flush()\r\n\r\nAttributeError: 'list' object has no attribute '_index'\r\n\r\nAny idea how to resolve this?",
      "votes": null
    },
    {
      "id": "95253",
      "postDate": "10/06/2015 13:31:36",
      "content": "<p>Hmm...I think the instructions were to replace <strong>all</strong> 10 lines with the 1 line.</p>\n\n<p>In the end, don't get caught up with the exact script's implementation, all it is doing is counting 7 things in the document using basic Python string manipulation and conveniently running the Random Forest and producing the submission file, all in one script.  If you can't get the script to work by hacking it, then rewrite it from scratch and only include the parts you fully understand, and your coding skills will be better for it.</p>",
      "rawMarkdown": "Hmm...I think the instructions were to replace **all** 10 lines with the 1 line.\r\n\r\nIn the end, don't get caught up with the exact script's implementation, all it is doing is counting 7 things in the document using basic Python string manipulation and conveniently running the Random Forest and producing the submission file, all in one script.  If you can't get the script to work by hacking it, then rewrite it from scratch and only include the parts you fully understand, and your coding skills will be better for it.",
      "votes": null
    },
    {
      "id": "95503",
      "postDate": "10/08/2015 14:32:21",
      "content": "<p>thanks for helping me to learn new things Mr.David Shinn.i ran successfully.</p>",
      "rawMarkdown": "thanks for helping me to learn new things Mr.David Shinn.i ran successfully.",
      "votes": null
    },
    {
      "id": "95636",
      "postDate": "10/10/2015 00:46:38",
      "content": "<p>David,</p>\n\n<p>Can you confirm your code gets you 0.90388? or did you do anything on top of this to get there? I ran this code and I'm only get a ~.07</p>\n\n<p>Thanks,\nJai</p>",
      "rawMarkdown": "David,\r\n\r\nCan you confirm your code gets you 0.90388? or did you do anything on top of this to get there? I ran this code and I'm only get a ~.07\r\n\r\nThanks,\r\nJai",
      "votes": null
    },
    {
      "id": "95638",
      "postDate": "10/10/2015 00:53:55",
      "content": "<p>Yep.  Proof is right off the leaderboard:</p>\n\n<pre><code>TeamId,TeamName,SubmissionDate,Score\n219616,&quot;David Shinn&quot;,&quot;2015-09-23 03:20:54&quot;,0.90388\n219815,&quot;Tim Kreienkamp&quot;,&quot;2015-09-23 17:53:10&quot;,0.90388\n214906,BigLeak,&quot;2015-09-25 17:39:49&quot;,0.90388\n220366,&quot;Jack B.&quot;,&quot;2015-09-26 00:38:58&quot;,0.90388\n220447,&quot;Rishab Gargeya&quot;,&quot;2015-09-26 09:33:16&quot;,0.90388\n213808,YS,&quot;2015-09-26 11:01:55&quot;,0.90388\n220669,&quot;surya venkat&quot;,&quot;2015-09-27 15:27:45&quot;,0.90388\n222164,&quot;Jeong-Yoon Lee&quot;,&quot;2015-10-02 22:18:38&quot;,0.90388\n222349,paladin_o_newengland,&quot;2015-10-03 17:07:10&quot;,0.90388\n222340,Mihawk,&quot;2015-10-04 03:13:29&quot;,0.90388\n207517,&quot;Ayman Khalafallah&quot;,&quot;2015-10-04 19:42:05&quot;,0.90388\n222826,MaksymPylypovych,&quot;2015-10-05 08:47:03&quot;,0.90388\n213434,&quot;Vikrant Kumar&quot;,&quot;2015-10-07 05:31:51&quot;,0.90388\n</code></pre>",
      "rawMarkdown": "Yep.  Proof is right off the leaderboard:\r\n\r\n    TeamId,TeamName,SubmissionDate,Score\r\n    219616,\"David Shinn\",\"2015-09-23 03:20:54\",0.90388\r\n    219815,\"Tim Kreienkamp\",\"2015-09-23 17:53:10\",0.90388\r\n    214906,BigLeak,\"2015-09-25 17:39:49\",0.90388\r\n    220366,\"Jack B.\",\"2015-09-26 00:38:58\",0.90388\r\n    220447,\"Rishab Gargeya\",\"2015-09-26 09:33:16\",0.90388\r\n    213808,YS,\"2015-09-26 11:01:55\",0.90388\r\n    220669,\"surya venkat\",\"2015-09-27 15:27:45\",0.90388\r\n    222164,\"Jeong-Yoon Lee\",\"2015-10-02 22:18:38\",0.90388\r\n    222349,paladin_o_newengland,\"2015-10-03 17:07:10\",0.90388\r\n    222340,Mihawk,\"2015-10-04 03:13:29\",0.90388\r\n    207517,\"Ayman Khalafallah\",\"2015-10-04 19:42:05\",0.90388\r\n    222826,MaksymPylypovych,\"2015-10-05 08:47:03\",0.90388\r\n    213434,\"Vikrant Kumar\",\"2015-10-07 05:31:51\",0.90388",
      "votes": null
    },
    {
      "id": "95639",
      "postDate": "10/10/2015 00:57:42",
      "content": "<p>David, Thanks for the quick response. How did you convert the output from clf.predict_proba into 1/0?</p>",
      "rawMarkdown": "David, Thanks for the quick response. How did you convert the output from clf.predict_proba into 1/0?",
      "votes": null
    },
    {
      "id": "95640",
      "postDate": "10/10/2015 01:03:54",
      "content": "<p>I didn't.  The evaluation metric is AUC, so you should be submitting a float probability between [0, 1].  Your submission file should look something like:</p>\n\n<pre><code>file,sponsored\n1000043_raw_html.txt,0.0153013876291\n1000097_raw_html.txt,0.01\n1000253_raw_html.txt,0.11\n100025_raw_html.txt,0.0\n1000313_raw_html.txt,0.18\n1000319_raw_html.txt,0.18\n1000325_raw_html.txt,0.03\n1000493_raw_html.txt,0.11\n1000553_raw_html.txt,0.64\n</code></pre>",
      "rawMarkdown": "I didn't.  The evaluation metric is AUC, so you should be submitting a float probability between [0, 1].  Your submission file should look something like:\r\n\r\n    file,sponsored\r\n    1000043_raw_html.txt,0.0153013876291\r\n    1000097_raw_html.txt,0.01\r\n    1000253_raw_html.txt,0.11\r\n    100025_raw_html.txt,0.0\r\n    1000313_raw_html.txt,0.18\r\n    1000319_raw_html.txt,0.18\r\n    1000325_raw_html.txt,0.03\r\n    1000493_raw_html.txt,0.11\r\n    1000553_raw_html.txt,0.64",
      "votes": null
    },
    {
      "id": "95641",
      "postDate": "10/10/2015 01:05:53",
      "content": "<p>David - thanks much. I appreciate your attitude towards teaching. </p>",
      "rawMarkdown": "David - thanks much. I appreciate your attitude towards teaching.",
      "votes": null
    },
    {
      "id": "95642",
      "postDate": "10/10/2015 01:13:34",
      "content": "<p>Jai - You're welcome.  The &quot;School of Kaggle&quot; has been good to me.  There's five whole days left and it is easy to improve this script to do much better than 0.90388, so go do some damage!</p>",
      "rawMarkdown": "Jai - You're welcome.  The \"School of Kaggle\" has been good to me.  There's five whole days left and it is easy to improve this script to do much better than 0.90388, so go do some damage!",
      "votes": null
    },
    {
      "id": "95643",
      "postDate": "10/10/2015 01:14:13",
      "content": "<p>Just FYI, in the example I converted the &quot;values&quot; result that is returned from a dictionary to using &quot;slots&quot; in order to <a href=\"https://stackoverflow.com/questions/472000/python-slots\">save memory</a>.  It seemed to help for me.</p>",
      "rawMarkdown": "Just FYI, in the example I converted the \"values\" result that is returned from a dictionary to using \"slots\" in order to [save memory](https://stackoverflow.com/questions/472000/python-slots).  It seemed to help for me.",
      "votes": null
    },
    {
      "id": "95646",
      "postDate": "10/10/2015 01:35:53",
      "content": "<p>Cool tip, I never knew that existed.  What line did you alter exactly?  Competitions with larger datasets definitely forces you to learn how to deal with memory and time efficiently.  I <strong>just</strong> learned how to pause a process using <code>kill -STOP</code> and <code>kill -CONT</code>, and in combination with <a href=\"http://stackoverflow.com/a/7485831/2626968\">this tip</a>, saved me from losing a 3 hour running process just now.</p>",
      "rawMarkdown": "Cool tip, I never knew that existed.  What line did you alter exactly?  Competitions with larger datasets definitely forces you to learn how to deal with memory and time efficiently.  I **just** learned how to pause a process using `kill -STOP` and `kill -CONT`, and in combination with [this tip][1], saved me from losing a 3 hour running process just now.\r\n\r\n\r\n  [1]: http://stackoverflow.com/a/7485831/2626968",
      "votes": null
    },
    {
      "id": "95648",
      "postDate": "10/10/2015 02:03:52",
      "content": "<p>I defined a class and then populated the values and returned it like this:</p>\n\n<pre><code>class ResultFileEntry(object):\n#https://utcc.utoronto.ca/~cks/space/blog/python/WhatSlotsAreGoodFor\n#http://tech.oyster.com/save-ram-with-python-slots/\n#https://stackoverflow.com/questions/1336791/dictionary-vs-object-which-is-more-efficient-and-why\n#https://stackoverflow.com/questions/472000/python-slots\n# We use slots here to save memory, a dynamic dictionary is not needed\n__slots__ = ['filename', 'sponsored', 'lines', 'spaces', 'tabs', 'braces', ...]\ndef __init__(self, filename):\n    self.file = filename\n    self.sponsored = None #make sure this is none because there is a check for this later during training\n    self.lines = 0\n    self.spaces = 0\n    self.tabs = 0\n    self.braces = 0\n    ...\n#https://docs.python.org/2.7/reference/datamodel.html#special-method-names\ndef __iter__(self):\n    yield self.file;\n    yield self.sponsored;\n    yield self.lines;\n    yield self.spaces;\n    yield self.tabs;\n    ...\n</code></pre>",
      "rawMarkdown": "I defined a class and then populated the values and returned it like this:\r\n\r\n    class ResultFileEntry(object):\r\n\t#https://utcc.utoronto.ca/~cks/space/blog/python/WhatSlotsAreGoodFor\r\n\t#http://tech.oyster.com/save-ram-with-python-slots/\r\n\t#https://stackoverflow.com/questions/1336791/dictionary-vs-object-which-is-more-efficient-and-why\r\n\t#https://stackoverflow.com/questions/472000/python-slots\r\n\t# We use slots here to save memory, a dynamic dictionary is not needed\r\n\t__slots__ = ['filename', 'sponsored', 'lines', 'spaces', 'tabs', 'braces', ...]\r\n\tdef __init__(self, filename):\r\n\t\tself.file = filename\r\n\t\tself.sponsored = None #make sure this is none because there is a check for this later during training\r\n\t\tself.lines = 0\r\n\t\tself.spaces = 0\r\n\t\tself.tabs = 0\r\n\t\tself.braces = 0\r\n\t\t...\r\n\t#https://docs.python.org/2.7/reference/datamodel.html#special-method-names\r\n\tdef __iter__(self):\r\n\t\tyield self.file;\r\n\t\tyield self.sponsored;\r\n\t\tyield self.lines;\r\n\t\tyield self.spaces;\r\n\t\tyield self.tabs;\r\n\t\t...",
      "votes": null
    },
    {
      "id": "95663",
      "postDate": "10/10/2015 08:28:09",
      "content": "<p>[quote=Jai;95639]</p>\n\n<p>David, Thanks for the quick response. How did you convert the output from clf.predict_proba into 1/0?</p>\n\n<p>[/quote]</p>\n\n<p>Don't do this. I did this by mistake (or rather, I didn't use predict_proba) for my first solution and with a cv of 0.79, I only got a lb score of 5.3. Took me some time to find this rather stupid mistake.</p>",
      "rawMarkdown": "[quote=Jai;95639]\r\n\r\nDavid, Thanks for the quick response. How did you convert the output from clf.predict_proba into 1/0?\r\n\r\n[/quote]\r\n\r\nDon't do this. I did this by mistake (or rather, I didn't use predict_proba) for my first solution and with a cv of 0.79, I only got a lb score of 5.3. Took me some time to find this rather stupid mistake.",
      "votes": null
    },
    {
      "id": "95814",
      "postDate": "10/12/2015 11:08:35",
      "content": "<p>Hi,</p>\n\n<p>First of all, thanks to <a href=\"https://www.kaggle.com/davidshinn\">@David Shinn</a> for the inspiration!</p>\n\n<p>Although it hasn't been tested on Windows, hope the attached script can deal with some asked technical issues here.</p>\n\n<p>Some quick note:</p>\n\n<ol>\n<li>No need to unzip [0-5].zip;</li>\n<li>pip3 install psutil, or comment out related code;</li>\n<li>Feel free to uncomment xgb related code to give it a shot;</li>\n<li>Some pool and data frame operations changed, if anything unclear please ask me;</li>\n<li>Wish it could let myself be focused on features, in the final couple of days.</li>\n<li>LB public score 0.90751</li>\n</ol>\n\n<p>Cheers!</p>",
      "rawMarkdown": "Hi,\r\n\r\nFirst of all, thanks to [@David Shinn][1] for the inspiration!\r\n\r\nAlthough it hasn't been tested on Windows, hope the attached script can deal with some asked technical issues here.\r\n\r\nSome quick note:\r\n\r\n 1. No need to unzip [0-5].zip;\r\n 2. pip3 install psutil, or comment out related code;\r\n 3. Feel free to uncomment xgb related code to give it a shot;\r\n 4. Some pool and data frame operations changed, if anything unclear please ask me;\r\n 5. Wish it could let myself be focused on features, in the final couple of days.\r\n 6. LB public score 0.90751\r\n\r\nCheers!\r\n\r\n\r\n  [1]: https://www.kaggle.com/davidshinn",
      "votes": null
    },
    {
      "id": "95818",
      "postDate": "10/12/2015 12:24:37",
      "content": "<p>[quote=psingman;95098]</p>\n\n<p>Thank you, that definitely was my issue (facepalm). Just to make sure I'm not doing anything unnecessary, you have to unzip all the TXT files in 0.zip, 1.zip etc..? My 130GB laptop is running out of storage....</p>\n\n<p>[/quote]</p>\n\n<p>Probably too late but try this to avoid unzipping:</p>\n\n<pre><code>import pandas as pd\nimport zipfile\n\n# read() gives you the content\nzipped = zipfile.ZipFile(zip_file_path, 'r')\nfor path in zipped.namelist():\n    if not str(path).endswith('_raw_html.txt'):\n        continue\n    zipped.read(path)\n\n# open() gives you the handle for pd\nzipped_csv = zipfile.ZipFile(zipped_csv_path, 'r')\npd.read_csv(zipped_csv.open(the_csv_file_name_inside_the_zip))\n</code></pre>",
      "rawMarkdown": "[quote=psingman;95098]\r\n\r\nThank you, that definitely was my issue (facepalm). Just to make sure I'm not doing anything unnecessary, you have to unzip all the TXT files in 0.zip, 1.zip etc..? My 130GB laptop is running out of storage....\r\n\r\n[/quote]\r\n\r\nProbably too late but try this to avoid unzipping:\r\n\r\n    import pandas as pd\r\n    import zipfile\r\n\r\n    # read() gives you the content\r\n    zipped = zipfile.ZipFile(zip_file_path, 'r')\r\n    for path in zipped.namelist():\r\n        if not str(path).endswith('_raw_html.txt'):\r\n            continue\r\n        zipped.read(path)\r\n\r\n    # open() gives you the handle for pd\r\n    zipped_csv = zipfile.ZipFile(zipped_csv_path, 'r')\r\n    pd.read_csv(zipped_csv.open(the_csv_file_name_inside_the_zip))",
      "votes": null
    },
    {
      "id": "96578",
      "postDate": "10/18/2015 19:00:27",
      "content": "<p>Thank you Barabbas. I'm not gonna run the script again but I'll use this format in the future to avoid unzipping.</p>",
      "rawMarkdown": "Thank you Barabbas. I'm not gonna run the script again but I'll use this format in the future to avoid unzipping.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 93244,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "09/23/2015 04:19:09",
      "content": "<p>It is impressive how much more (I guess) complex the solution needs to be for less than 0.1 AUC gain :/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93249,
      "author_name": "bekbolatov",
      "author_url": "",
      "post_date": "09/23/2015 04:46:43",
      "content": "<p>0.90388 with this? </p>\n\n<p>David, the simplicity and performance of this code is impressive!</p>\n\n<p>This submission would automatically earn a new entrant 2920 Kaggle points if the competition were to end with current standings - a steal! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93251,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "09/23/2015 05:20:37",
      "content": "<p>Truly impressive!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93258,
      "author_name": "fernandoprocy",
      "author_url": "",
      "post_date": "09/23/2015 07:42:46",
      "content": "<p>I'm really impressed about how well such basic feature selection can perform at detecting these native ads. I browsed a little about stumbleupon native ads system at the beginning of this competition and couldn't even imagine that pages you put through their ads system could be identified as sponsored since it doesn't mess up with the html code and it's just a URL that stumbleupon would show to targeted audiences.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93288,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "09/23/2015 15:41:02",
      "content": "<p>Now 0.90388 is the new zero, Very impressive, David :)\n[quote=David Shinn;93238]</p>\n\n<p>In response <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16623/this-competition-is-not-popular-at-all\">to this post</a> and to drum up a little more interest in this competition, I'm posting some starter code that performs decently (0.90388), needs quite a bit of work, leaves room for a lot of improvement, and it is a really basic model (just 7 features with RandomForest), showing that many competitors are ignoring some really basic features.  It runs under 8 minutes from start to finish on a modern Mac Book Pro (uses multiprocessing, so runs 8 processes in parallel).  Just in case you're wondering, I'm not a proponent of high performance btb code this close to competition end, but really, just 7 basic count features?</p>\n\n<p>Enjoy!</p>\n\n<p>My feature set:</p>\n\n<pre><code>values['lines'] = text.count('\\n')\nvalues['spaces'] = text.count(' ')\nvalues['tabs'] = text.count('\\t')\nvalues['braces'] = text.count('{')\nvalues['brackets'] = text.count('[')\nvalues['words'] = len(re.split('\\s+', text))\nvalues['length'] = len(text)\n</code></pre>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93326,
      "author_name": "firefly2442",
      "author_url": "",
      "post_date": "09/23/2015 21:15:12",
      "content": "<p>Thank you for the example code.  I'm just curious, the choice of .imap() to do the threading versus other alternatives such as apply_async() and map_async()...  What made you pick .imap() and why does it seem to be running faster than some other analysis code that I wrote using apply_async()?  I found <a href=\"https://stackoverflow.com/questions/26520781/python-multiprocessing-poolwhats-the-difference-between-map-async-and-imap\">this on Stackoverflow</a> but I'm not sure I understand the details.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93331,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "09/23/2015 22:46:01",
      "content": "<p>It's just what has worked.  I've always used <a href=\"http://stackoverflow.com/a/5666996/2626968\">this post</a> as a guide, something about the counter counting chunks, not actual elements.  I use imap so I can see progress.  Otherwise I just use map.  Still learning, of course.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93347,
      "author_name": "sudalairajkumar",
      "author_url": "",
      "post_date": "09/24/2015 03:44:21",
      "content": "<p>Thanks a lot @David Shinn. This is very helpful. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93386,
      "author_name": "tobycheese",
      "author_url": "",
      "post_date": "09/24/2015 21:23:33",
      "content": "<p>Beautiful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93826,
      "author_name": "firefly2442",
      "author_url": "",
      "post_date": "10/01/2015 20:25:36",
      "content": "<p>Has anyone had luck modifying this to run on Windows?  Since you can't fork under Windows, I've had a tough time adjusting this so the processing/working function can use the train_keys list without having to pass the entire thing as a parameter.  Works great on Linux. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93932,
      "author_name": "luyaoly",
      "author_url": "",
      "post_date": "10/03/2015 02:10:17",
      "content": "<p>I only got an advanced windows PC. Looks like the script won't run through. Any suggestions on modifying this would be really helpful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93933,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/03/2015 02:33:50",
      "content": "<p>[quote=paladin_o_newengland;93932]</p>\n\n<p>Any suggestions on modifying this would be really helpful!</p>\n\n<p>[/quote]</p>\n\n<p>Take this code:</p>\n\n<pre><code>p = multiprocessing.Pool()\nresults = p.imap(create_data, filepaths)\nwhile (True):\n    completed = results._index\n    print(&quot;\\r--- Completed {:,} out of {:,}&quot;.format(completed, num_tasks), end='')\n    sys.stdout.flush()\n    time.sleep(1)\n    if (completed == num_tasks): break\np.close()\np.join()\n</code></pre>\n\n<p>and replace it with:</p>\n\n<pre><code>results = map(create_data, filepaths)\n</code></pre>\n\n<p>It'll work, but you won't get parallel processing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95052,
      "author_name": "ask788",
      "author_url": "",
      "post_date": "10/04/2015 17:43:45",
      "content": "<p>I ran this code but for some reason all my test features are null. I am able to get the train features but not for the test. Any reason why?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95053,
      "author_name": "niranjanmudhiraj",
      "author_url": "",
      "post_date": "10/04/2015 17:56:48",
      "content": "<p>when i run this script i am getting this error : &quot;AttributeError: 'DataFrame' object has no attribute 'sponsored'&quot; , can you tell how to resolve this error.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95066,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/04/2015 19:28:08",
      "content": "<p>[quote=ask788;95052]</p>\n\n<p>I ran this code but for some reason all my test features are null. I am able to get the train features but not for the test. Any reason why?</p>\n\n<p>[/quote]</p>\n\n<p>That's weird, I specifically set all test null values to 0, so you should at least have a data frame full of zeros.</p>\n\n<p>[quote=nknm;95053]</p>\n\n<p>when i run this script i am getting this error : &quot;AttributeError: 'DataFrame' object has no attribute 'sponsored'&quot; , can you tell how to resolve this error.</p>\n\n<p>[/quote]</p>\n\n<p>This probably means that the df_full dataframe does not have the column &quot;sponsored&quot;, which means either you took it out of the create_data function or something else is wrong.</p>\n\n<p>In general, I find it difficult to debug code that's been wrapped around multiprocessing, because it won't complain for code that breaks within the multiprocessing code until after it's all done, which is painful for this type of feature extraction.  The best thing you can do is</p>\n\n<ol>\n<li>Turn it into a non-multiprocessing code with the tip I added before, and</li>\n<li>Turn this line <code>filepaths = glob.glob('data/*/*.txt')</code> into <code>filepaths = glob.glob('data/0/*.txt')[:1000] + glob.glob('data/5/*.txt')[:1000]</code>for debugging purposes.  The code won't give you a proper submission file, but it should run through the feature extraction, training, and predicting sections quickly for you to experiment and figure out what's wrong with your code (or to add new features, etc).</li>\n</ol>\n\n<p>Isn't the learning process fun?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95072,
      "author_name": "psingman",
      "author_url": "",
      "post_date": "10/04/2015 20:21:50",
      "content": "<p>Hello Dave Shinn,</p>\n\n<p>Thank you for posting this, I hope to get it to run in the next few days. My error is num_tasks is staying at zero because the filepaths line (filepaths = glob.glob('data/<em>/</em>.txt')) isn't identifying the txt file. My data folder looks like the attached image below. Is there a way to extract the train TXT files from train_v2.csv that I am missing?</p>\n\n<p>Thank you,</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95074,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/04/2015 20:29:41",
      "content": "<p>[quote=psingman;95072]</p>\n\n<p>Hello Dave Shinn,</p>\n\n<p>Thank you for posting this, I hope to get it to run in the next few days. My error is num_tasks is staying at zero because the filepaths line (filepaths = glob.glob('data/<em>/</em>.txt')) isn't identifying the txt file. My data folder looks like the attached image below. Is there a way to extract the train TXT files from train_v2.csv that I am missing?</p>\n\n<p>Thank you,</p>\n\n<p>[/quote]</p>\n\n<p>I'm pretty sure that glob.glob needs a wild card, like: <code>filepaths = glob.glob('data/*/*.txt')</code>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95075,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/04/2015 20:36:27",
      "content": "<p>[quote=psingman;95072]</p>\n\n<p>Hello Dave Shinn,</p>\n\n<p>Thank you for posting this, I hope to get it to run in the next few days. My error is num_tasks is staying at zero because the filepaths line (filepaths = glob.glob('data/<em>/</em>.txt')) isn't identifying the txt file. My data folder looks like the attached image below. Is there a way to extract the train TXT files from train_v2.csv that I am missing?</p>\n\n<p>Thank you,</p>\n\n<p>[/quote]</p>\n\n<p>Okay, just realized your problem.  You need to download all 6 zip files (0.zip, 1.zip, etc), then extract them so their contents reside under data/0/, data/1/, data/2/, etc.  There's over 400K files.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95091,
      "author_name": "vikrant24",
      "author_url": "",
      "post_date": "10/05/2015 00:56:52",
      "content": "<p>I tried to run the benchmark code on window python. The below multiprocessing.Pool() step is seems to be running forever. This is 4th day it is still running. I have a pretty decent machine 32 GB RAM and I7 still not able to pass thru this step. Any idea how to make this code run fast?</p>\n\n<p>p = multiprocessing.Pool()\nresults = p.imap(create_data, filepaths)\nwhile (True):\n    completed = results._index\n    print(&quot;\\r--- Completed {:,} out of {:,}&quot;.format(completed, num_tasks), end='')\n    sys.stdout.flush()\n    time.sleep(1)\n    if (completed == num_tasks): break\np.close()\np.join()\ndf_full = pd.DataFrame(list(results))\nprint()</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95094,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/05/2015 01:19:15",
      "content": "<p>[quote=Vikrant Kumar;95091]</p>\n\n<p>I tried to run the benchmark code on window python. The below multiprocessing.Pool() step is seems to be running forever. This is 4th day it is still running. I have a pretty decent machine 32 GB RAM and I7 still not able to pass thru this step. Any idea how to make this code run fast?</p>\n\n<p>p = multiprocessing.Pool()\nresults = p.imap(create_data, filepaths)\nwhile (True):\n    completed = results._index\n    print(&quot;\\r--- Completed {:,} out of {:,}&quot;.format(completed, num_tasks), end='')\n    sys.stdout.flush()\n    time.sleep(1)\n    if (completed == num_tasks): break\np.close()\np.join()\ndf_full = pd.DataFrame(list(results))\nprint()</p>\n\n<p>[/quote]</p>\n\n<p>Unfortunately, it is likely a problem with the how multiprocessing works on Windows, where I can't help you.  However, take the tips from <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model/95066#post95066\">this part of the thread</a> and it'll set you up to debug the code in less than a couple of minutes.  My biggest advice here is first prototype something small that you expect to run in a reasonable amount of time (like a few minutes, not 4 days) and lose the multiprocessing.  Once you have the kinks worked out, try to get multiprocessing to work.  If that doesn't work, just forget the multiprocessing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95096,
      "author_name": "vikrant24",
      "author_url": "",
      "post_date": "10/05/2015 02:40:23",
      "content": "<p>Thanks David for response. Any idea where it will run fast? I can install something which can make it run fast like ubuntu or something.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95097,
      "author_name": "luyaoly",
      "author_url": "",
      "post_date": "10/05/2015 02:41:16",
      "content": "<p>Thanks so much david! Replacing the multi processing with simply &quot;results = map(create_data, filepaths)&quot; works great for me. For those of you using windows too, change that would make it work. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95098,
      "author_name": "psingman",
      "author_url": "",
      "post_date": "10/05/2015 02:45:49",
      "content": "<p>Thank you, that definitely was my issue (facepalm). Just to make sure I'm not doing anything unnecessary, you have to unzip all the TXT files in 0.zip, 1.zip etc..? My 130GB laptop is running out of storage....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95099,
      "author_name": "vikrant24",
      "author_url": "",
      "post_date": "10/05/2015 02:46:05",
      "content": "<p>Hi Paladin \nDo I need to change the whole piece of code with &quot;results = map(create_data, filepaths)&quot; or only the first line &quot;p = multiprocessing.Pool()&quot;. I can see second line already is same as what you suggested. Can you be please precise. I will try it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95100,
      "author_name": "firefly2442",
      "author_url": "",
      "post_date": "10/05/2015 03:05:58",
      "content": "<p>[quote=Vikrant Kumar;95099]</p>\n\n<p>Hi Paladin \nDo I need to change the whole piece of code with &quot;results = map(create_data, filepaths)&quot; or only the first line &quot;p = multiprocessing.Pool()&quot;. I can see second line already is same as what you suggested. Can you be please precise. I will try it.</p>\n\n<p>[/quote]</p>\n\n<p>I believe just change the results = map... line.  The other line sets up the processing pool which should be the same regardless of you using map, imap, etc.  See the <a href=\"https://docs.python.org/2/library/multiprocessing.html#module-multiprocessing.pool\">documentation here</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95146,
      "author_name": "luyaoly",
      "author_url": "",
      "post_date": "10/05/2015 14:16:21",
      "content": "<p>[quote=Vikrant Kumar;95099]</p>\n\n<p>Hi Paladin \nDo I need to change the whole piece of code with &quot;results = map(create_data, filepaths)&quot; or only the first line &quot;p = multiprocessing.Pool()&quot;. I can see second line already is same as what you suggested. Can you be please precise. I will try it.</p>\n\n<p>[/quote]</p>\n\n<p>I actually did just as David suggested in his response last page, use &quot;results = map(create_data, filepaths)&quot; to replace the whole thing below: (I didn't try other ways. Note running it will take time. Thanks. ) </p>\n\n<p>p = multiprocessing.Pool()\nresults = p.imap(create_data, filepaths)\nwhile (True):\n    completed = results._index\n    print(&quot;\\r--- Completed {:,} out of {:,}&quot;.format(completed, num_tasks), end='')\n    sys.stdout.flush()\n    time.sleep(1)\n    if (completed == num_tasks): break\np.close()\np.join()</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95230,
      "author_name": "vikrant24",
      "author_url": "",
      "post_date": "10/06/2015 05:52:50",
      "content": "<p>[quote=David Shinn;93933]</p>\n\n<p>[quote=paladin_o_newengland;93932]</p>\n\n<p>Any suggestions on modifying this would be really helpful!</p>\n\n<p>[/quote]</p>\n\n<p>Take this code:</p>\n\n<pre><code>p = multiprocessing.Pool()\nresults = p.imap(create_data, filepaths)\nwhile (True):\n    completed = results._index\n    print(&quot;\\r--- Completed {:,} out of {:,}&quot;.format(completed, num_tasks), end='')\n    sys.stdout.flush()\n    time.sleep(1)\n    if (completed == num_tasks): break\np.close()\np.join()\n</code></pre>\n\n<p>and replace it with:</p>\n\n<pre><code>results = map(create_data, filepaths)\n</code></pre>\n\n<p>It'll work, but you won't get parallel processing.</p>\n\n<p>[/quote]\nHi\nI applied this but ended up with error</p>\n\n<hr>\n\n<p>AttributeError                            Traceback (most recent call last)\n in ()\n      3 #results = p.imap(create_data, filepaths)\n      4 while (True):\n----&gt; 5     completed = results._index\n      6     print(&quot;\\r--- Completed {:,} out of {:,}&quot;.format(completed, num_tasks), end='')\n      7     sys.stdout.flush()</p>\n\n<p>AttributeError: 'list' object has no attribute '_index'</p>\n\n<p>Any idea how to resolve this?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95253,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/06/2015 13:31:36",
      "content": "<p>Hmm...I think the instructions were to replace <strong>all</strong> 10 lines with the 1 line.</p>\n\n<p>In the end, don't get caught up with the exact script's implementation, all it is doing is counting 7 things in the document using basic Python string manipulation and conveniently running the Random Forest and producing the submission file, all in one script.  If you can't get the script to work by hacking it, then rewrite it from scratch and only include the parts you fully understand, and your coding skills will be better for it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95503,
      "author_name": "niranjanmudhiraj",
      "author_url": "",
      "post_date": "10/08/2015 14:32:21",
      "content": "<p>thanks for helping me to learn new things Mr.David Shinn.i ran successfully.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95636,
      "author_name": "dataparty",
      "author_url": "",
      "post_date": "10/10/2015 00:46:38",
      "content": "<p>David,</p>\n\n<p>Can you confirm your code gets you 0.90388? or did you do anything on top of this to get there? I ran this code and I'm only get a ~.07</p>\n\n<p>Thanks,\nJai</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95638,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/10/2015 00:53:55",
      "content": "<p>Yep.  Proof is right off the leaderboard:</p>\n\n<pre><code>TeamId,TeamName,SubmissionDate,Score\n219616,&quot;David Shinn&quot;,&quot;2015-09-23 03:20:54&quot;,0.90388\n219815,&quot;Tim Kreienkamp&quot;,&quot;2015-09-23 17:53:10&quot;,0.90388\n214906,BigLeak,&quot;2015-09-25 17:39:49&quot;,0.90388\n220366,&quot;Jack B.&quot;,&quot;2015-09-26 00:38:58&quot;,0.90388\n220447,&quot;Rishab Gargeya&quot;,&quot;2015-09-26 09:33:16&quot;,0.90388\n213808,YS,&quot;2015-09-26 11:01:55&quot;,0.90388\n220669,&quot;surya venkat&quot;,&quot;2015-09-27 15:27:45&quot;,0.90388\n222164,&quot;Jeong-Yoon Lee&quot;,&quot;2015-10-02 22:18:38&quot;,0.90388\n222349,paladin_o_newengland,&quot;2015-10-03 17:07:10&quot;,0.90388\n222340,Mihawk,&quot;2015-10-04 03:13:29&quot;,0.90388\n207517,&quot;Ayman Khalafallah&quot;,&quot;2015-10-04 19:42:05&quot;,0.90388\n222826,MaksymPylypovych,&quot;2015-10-05 08:47:03&quot;,0.90388\n213434,&quot;Vikrant Kumar&quot;,&quot;2015-10-07 05:31:51&quot;,0.90388\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95639,
      "author_name": "dataparty",
      "author_url": "",
      "post_date": "10/10/2015 00:57:42",
      "content": "<p>David, Thanks for the quick response. How did you convert the output from clf.predict_proba into 1/0?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95640,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/10/2015 01:03:54",
      "content": "<p>I didn't.  The evaluation metric is AUC, so you should be submitting a float probability between [0, 1].  Your submission file should look something like:</p>\n\n<pre><code>file,sponsored\n1000043_raw_html.txt,0.0153013876291\n1000097_raw_html.txt,0.01\n1000253_raw_html.txt,0.11\n100025_raw_html.txt,0.0\n1000313_raw_html.txt,0.18\n1000319_raw_html.txt,0.18\n1000325_raw_html.txt,0.03\n1000493_raw_html.txt,0.11\n1000553_raw_html.txt,0.64\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95641,
      "author_name": "dataparty",
      "author_url": "",
      "post_date": "10/10/2015 01:05:53",
      "content": "<p>David - thanks much. I appreciate your attitude towards teaching. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95642,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/10/2015 01:13:34",
      "content": "<p>Jai - You're welcome.  The &quot;School of Kaggle&quot; has been good to me.  There's five whole days left and it is easy to improve this script to do much better than 0.90388, so go do some damage!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95643,
      "author_name": "firefly2442",
      "author_url": "",
      "post_date": "10/10/2015 01:14:13",
      "content": "<p>Just FYI, in the example I converted the &quot;values&quot; result that is returned from a dictionary to using &quot;slots&quot; in order to <a href=\"https://stackoverflow.com/questions/472000/python-slots\">save memory</a>.  It seemed to help for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95646,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/10/2015 01:35:53",
      "content": "<p>Cool tip, I never knew that existed.  What line did you alter exactly?  Competitions with larger datasets definitely forces you to learn how to deal with memory and time efficiently.  I <strong>just</strong> learned how to pause a process using <code>kill -STOP</code> and <code>kill -CONT</code>, and in combination with <a href=\"http://stackoverflow.com/a/7485831/2626968\">this tip</a>, saved me from losing a 3 hour running process just now.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95648,
      "author_name": "firefly2442",
      "author_url": "",
      "post_date": "10/10/2015 02:03:52",
      "content": "<p>I defined a class and then populated the values and returned it like this:</p>\n\n<pre><code>class ResultFileEntry(object):\n#https://utcc.utoronto.ca/~cks/space/blog/python/WhatSlotsAreGoodFor\n#http://tech.oyster.com/save-ram-with-python-slots/\n#https://stackoverflow.com/questions/1336791/dictionary-vs-object-which-is-more-efficient-and-why\n#https://stackoverflow.com/questions/472000/python-slots\n# We use slots here to save memory, a dynamic dictionary is not needed\n__slots__ = ['filename', 'sponsored', 'lines', 'spaces', 'tabs', 'braces', ...]\ndef __init__(self, filename):\n    self.file = filename\n    self.sponsored = None #make sure this is none because there is a check for this later during training\n    self.lines = 0\n    self.spaces = 0\n    self.tabs = 0\n    self.braces = 0\n    ...\n#https://docs.python.org/2.7/reference/datamodel.html#special-method-names\ndef __iter__(self):\n    yield self.file;\n    yield self.sponsored;\n    yield self.lines;\n    yield self.spaces;\n    yield self.tabs;\n    ...\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95663,
      "author_name": "tobycheese",
      "author_url": "",
      "post_date": "10/10/2015 08:28:09",
      "content": "<p>[quote=Jai;95639]</p>\n\n<p>David, Thanks for the quick response. How did you convert the output from clf.predict_proba into 1/0?</p>\n\n<p>[/quote]</p>\n\n<p>Don't do this. I did this by mistake (or rather, I didn't use predict_proba) for my first solution and with a cv of 0.79, I only got a lb score of 5.3. Took me some time to find this rather stupid mistake.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95814,
      "author_name": "tmjiang",
      "author_url": "",
      "post_date": "10/12/2015 11:08:35",
      "content": "<p>Hi,</p>\n\n<p>First of all, thanks to <a href=\"https://www.kaggle.com/davidshinn\">@David Shinn</a> for the inspiration!</p>\n\n<p>Although it hasn't been tested on Windows, hope the attached script can deal with some asked technical issues here.</p>\n\n<p>Some quick note:</p>\n\n<ol>\n<li>No need to unzip [0-5].zip;</li>\n<li>pip3 install psutil, or comment out related code;</li>\n<li>Feel free to uncomment xgb related code to give it a shot;</li>\n<li>Some pool and data frame operations changed, if anything unclear please ask me;</li>\n<li>Wish it could let myself be focused on features, in the final couple of days.</li>\n<li>LB public score 0.90751</li>\n</ol>\n\n<p>Cheers!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95818,
      "author_name": "tmjiang",
      "author_url": "",
      "post_date": "10/12/2015 12:24:37",
      "content": "<p>[quote=psingman;95098]</p>\n\n<p>Thank you, that definitely was my issue (facepalm). Just to make sure I'm not doing anything unnecessary, you have to unzip all the TXT files in 0.zip, 1.zip etc..? My 130GB laptop is running out of storage....</p>\n\n<p>[/quote]</p>\n\n<p>Probably too late but try this to avoid unzipping:</p>\n\n<pre><code>import pandas as pd\nimport zipfile\n\n# read() gives you the content\nzipped = zipfile.ZipFile(zip_file_path, 'r')\nfor path in zipped.namelist():\n    if not str(path).endswith('_raw_html.txt'):\n        continue\n    zipped.read(path)\n\n# open() gives you the handle for pd\nzipped_csv = zipfile.ZipFile(zipped_csv_path, 'r')\npd.read_csv(zipped_csv.open(the_csv_file_name_inside_the_zip))\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96578,
      "author_name": "psingman",
      "author_url": "",
      "post_date": "10/18/2015 19:00:27",
      "content": "<p>Thank you Barabbas. I'm not gonna run the script again but I'll use this format in the future to avoid unzipping.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "93238": "In response [to this post][1] and to drum up a little more interest in this competition, I'm posting some starter code that performs decently (0.90388), needs quite a bit of work, leaves room for a lot of improvement, and it is a really basic model (just 7 features with RandomForest), showing that many competitors are ignoring some really basic features.  It runs under 8 minutes from start to finish on a modern Mac Book Pro (uses multiprocessing, so runs 8 processes in parallel).  Just in case you're wondering, I'm not a proponent of high performance btb code this close to competition end, but really, just 7 basic count features?\r\n\r\nEnjoy!\r\n\r\nMy feature set:\r\n\r\n    values['lines'] = text.count('\\n')\r\n    values['spaces'] = text.count(' ')\r\n    values['tabs'] = text.count('\\t')\r\n    values['braces'] = text.count('{')\r\n    values['brackets'] = text.count('[')\r\n    values['words'] = len(re.split('\\s+', text))\r\n    values['length'] = len(text)\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16623/this-competition-is-not-popular-at-all",
    "93244": "It is impressive how much more (I guess) complex the solution needs to be for less than 0.1 AUC gain :/",
    "93249": "0.90388 with this? \r\n\r\nDavid, the simplicity and performance of this code is impressive!\r\n\r\nThis submission would automatically earn a new entrant 2920 Kaggle points if the competition were to end with current standings - a steal! :)",
    "93251": "Truly impressive!",
    "93258": "I'm really impressed about how well such basic feature selection can perform at detecting these native ads. I browsed a little about stumbleupon native ads system at the beginning of this competition and couldn't even imagine that pages you put through their ads system could be identified as sponsored since it doesn't mess up with the html code and it's just a URL that stumbleupon would show to targeted audiences.",
    "93288": "Now 0.90388 is the new zero, Very impressive, David :)\r\n[quote=David Shinn;93238]\r\n\r\nIn response [to this post][1] and to drum up a little more interest in this competition, I'm posting some starter code that performs decently (0.90388), needs quite a bit of work, leaves room for a lot of improvement, and it is a really basic model (just 7 features with RandomForest), showing that many competitors are ignoring some really basic features.  It runs under 8 minutes from start to finish on a modern Mac Book Pro (uses multiprocessing, so runs 8 processes in parallel).  Just in case you're wondering, I'm not a proponent of high performance btb code this close to competition end, but really, just 7 basic count features?\r\n\r\nEnjoy!\r\n\r\nMy feature set:\r\n\r\n    values['lines'] = text.count('\\n')\r\n    values['spaces'] = text.count(' ')\r\n    values['tabs'] = text.count('\\t')\r\n    values['braces'] = text.count('{')\r\n    values['brackets'] = text.count('[')\r\n    values['words'] = len(re.split('\\s+', text))\r\n    values['length'] = len(text)\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16623/this-competition-is-not-popular-at-all\r\n\r\n[/quote]",
    "93326": "Thank you for the example code.  I'm just curious, the choice of .imap() to do the threading versus other alternatives such as apply_async() and map_async()...  What made you pick .imap() and why does it seem to be running faster than some other analysis code that I wrote using apply_async()?  I found [this on Stackoverflow](https://stackoverflow.com/questions/26520781/python-multiprocessing-poolwhats-the-difference-between-map-async-and-imap) but I'm not sure I understand the details.",
    "93331": "It's just what has worked.  I've always used [this post][1] as a guide, something about the counter counting chunks, not actual elements.  I use imap so I can see progress.  Otherwise I just use map.  Still learning, of course.\r\n\r\n\r\n  [1]: http://stackoverflow.com/a/5666996/2626968",
    "93347": "Thanks a lot @David Shinn. This is very helpful.",
    "93386": "Beautiful!",
    "93826": "Has anyone had luck modifying this to run on Windows?  Since you can't fork under Windows, I've had a tough time adjusting this so the processing/working function can use the train_keys list without having to pass the entire thing as a parameter.  Works great on Linux. :)",
    "93932": "I only got an advanced windows PC. Looks like the script won't run through. Any suggestions on modifying this would be really helpful!",
    "93933": "[quote=paladin_o_newengland;93932]\r\n\r\nAny suggestions on modifying this would be really helpful!\r\n\r\n[/quote]\r\n\r\nTake this code:\r\n\r\n    p = multiprocessing.Pool()\r\n    results = p.imap(create_data, filepaths)\r\n    while (True):\r\n        completed = results._index\r\n        print(\"\\r--- Completed {:,} out of {:,}\".format(completed, num_tasks), end='')\r\n        sys.stdout.flush()\r\n        time.sleep(1)\r\n        if (completed == num_tasks): break\r\n    p.close()\r\n    p.join()\r\n\r\nand replace it with:\r\n\r\n    results = map(create_data, filepaths)\r\n\r\nIt'll work, but you won't get parallel processing.",
    "95052": "I ran this code but for some reason all my test features are null. I am able to get the train features but not for the test. Any reason why?",
    "95053": "when i run this script i am getting this error : \"AttributeError: 'DataFrame' object has no attribute 'sponsored'\" , can you tell how to resolve this error.",
    "95066": "[quote=ask788;95052]\r\n\r\nI ran this code but for some reason all my test features are null. I am able to get the train features but not for the test. Any reason why?\r\n\r\n[/quote]\r\n\r\nThat's weird, I specifically set all test null values to 0, so you should at least have a data frame full of zeros.\r\n\r\n[quote=nknm;95053]\r\n\r\nwhen i run this script i am getting this error : \"AttributeError: 'DataFrame' object has no attribute 'sponsored'\" , can you tell how to resolve this error.\r\n\r\n\r\n[/quote]\r\n\r\nThis probably means that the df_full dataframe does not have the column \"sponsored\", which means either you took it out of the create_data function or something else is wrong.\r\n\r\nIn general, I find it difficult to debug code that's been wrapped around multiprocessing, because it won't complain for code that breaks within the multiprocessing code until after it's all done, which is painful for this type of feature extraction.  The best thing you can do is\r\n\r\n 1. Turn it into a non-multiprocessing code with the tip I added before, and\r\n 2. Turn this line `filepaths = glob.glob('data/*/*.txt')` into `filepaths = glob.glob('data/0/*.txt')[:1000] + glob.glob('data/5/*.txt')[:1000]`for debugging purposes.  The code won't give you a proper submission file, but it should run through the feature extraction, training, and predicting sections quickly for you to experiment and figure out what's wrong with your code (or to add new features, etc).\r\n\r\nIsn't the learning process fun?",
    "95072": "Hello Dave Shinn,\r\n\r\nThank you for posting this, I hope to get it to run in the next few days. My error is num_tasks is staying at zero because the filepaths line (filepaths = glob.glob('data/*/*.txt')) isn't identifying the txt file. My data folder looks like the attached image below. Is there a way to extract the train TXT files from train_v2.csv that I am missing?\r\n\r\nThank you,",
    "95074": "[quote=psingman;95072]\r\n\r\nHello Dave Shinn,\r\n\r\nThank you for posting this, I hope to get it to run in the next few days. My error is num_tasks is staying at zero because the filepaths line (filepaths = glob.glob('data/*/*.txt')) isn't identifying the txt file. My data folder looks like the attached image below. Is there a way to extract the train TXT files from train_v2.csv that I am missing?\r\n\r\nThank you,\r\n\r\n[/quote]\r\n\r\nI'm pretty sure that glob.glob needs a wild card, like: `filepaths = glob.glob('data/*/*.txt')`.",
    "95075": "[quote=psingman;95072]\r\n\r\nHello Dave Shinn,\r\n\r\nThank you for posting this, I hope to get it to run in the next few days. My error is num_tasks is staying at zero because the filepaths line (filepaths = glob.glob('data/*/*.txt')) isn't identifying the txt file. My data folder looks like the attached image below. Is there a way to extract the train TXT files from train_v2.csv that I am missing?\r\n\r\nThank you,\r\n\r\n[/quote]\r\n\r\nOkay, just realized your problem.  You need to download all 6 zip files (0.zip, 1.zip, etc), then extract them so their contents reside under data/0/, data/1/, data/2/, etc.  There's over 400K files.",
    "95091": "I tried to run the benchmark code on window python. The below multiprocessing.Pool() step is seems to be running forever. This is 4th day it is still running. I have a pretty decent machine 32 GB RAM and I7 still not able to pass thru this step. Any idea how to make this code run fast?\r\n\r\np = multiprocessing.Pool()\r\nresults = p.imap(create_data, filepaths)\r\nwhile (True):\r\n    completed = results._index\r\n    print(\"\\r--- Completed {:,} out of {:,}\".format(completed, num_tasks), end='')\r\n    sys.stdout.flush()\r\n    time.sleep(1)\r\n    if (completed == num_tasks): break\r\np.close()\r\np.join()\r\ndf_full = pd.DataFrame(list(results))\r\nprint()",
    "95094": "[quote=Vikrant Kumar;95091]\r\n\r\nI tried to run the benchmark code on window python. The below multiprocessing.Pool() step is seems to be running forever. This is 4th day it is still running. I have a pretty decent machine 32 GB RAM and I7 still not able to pass thru this step. Any idea how to make this code run fast?\r\n\r\np = multiprocessing.Pool()\r\nresults = p.imap(create_data, filepaths)\r\nwhile (True):\r\n    completed = results._index\r\n    print(\"\\r--- Completed {:,} out of {:,}\".format(completed, num_tasks), end='')\r\n    sys.stdout.flush()\r\n    time.sleep(1)\r\n    if (completed == num_tasks): break\r\np.close()\r\np.join()\r\ndf_full = pd.DataFrame(list(results))\r\nprint()\r\n\r\n[/quote]\r\n\r\nUnfortunately, it is likely a problem with the how multiprocessing works on Windows, where I can't help you.  However, take the tips from [this part of the thread][1] and it'll set you up to debug the code in less than a couple of minutes.  My biggest advice here is first prototype something small that you expect to run in a reasonable amount of time (like a few minutes, not 4 days) and lose the multiprocessing.  Once you have the kinks worked out, try to get multiprocessing to work.  If that doesn't work, just forget the multiprocessing.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model/95066#post95066",
    "95096": "Thanks David for response. Any idea where it will run fast? I can install something which can make it run fast like ubuntu or something.",
    "95097": "Thanks so much david! Replacing the multi processing with simply \"results = map(create_data, filepaths)\" works great for me. For those of you using windows too, change that would make it work.",
    "95098": "Thank you, that definitely was my issue (facepalm). Just to make sure I'm not doing anything unnecessary, you have to unzip all the TXT files in 0.zip, 1.zip etc..? My 130GB laptop is running out of storage....",
    "95099": "Hi Paladin \r\nDo I need to change the whole piece of code with \"results = map(create_data, filepaths)\" or only the first line \"p = multiprocessing.Pool()\". I can see second line already is same as what you suggested. Can you be please precise. I will try it.",
    "95100": "[quote=Vikrant Kumar;95099]\r\n\r\nHi Paladin \r\nDo I need to change the whole piece of code with \"results = map(create_data, filepaths)\" or only the first line \"p = multiprocessing.Pool()\". I can see second line already is same as what you suggested. Can you be please precise. I will try it.\r\n\r\n[/quote]\r\n\r\nI believe just change the results = map... line.  The other line sets up the processing pool which should be the same regardless of you using map, imap, etc.  See the [documentation here](https://docs.python.org/2/library/multiprocessing.html#module-multiprocessing.pool)",
    "95146": "[quote=Vikrant Kumar;95099]\r\n\r\nHi Paladin \r\nDo I need to change the whole piece of code with \"results = map(create_data, filepaths)\" or only the first line \"p = multiprocessing.Pool()\". I can see second line already is same as what you suggested. Can you be please precise. I will try it.\r\n\r\n[/quote]\r\n\r\nI actually did just as David suggested in his response last page, use \"results = map(create_data, filepaths)\" to replace the whole thing below: (I didn't try other ways. Note running it will take time. Thanks. ) \r\n\r\np = multiprocessing.Pool()\r\nresults = p.imap(create_data, filepaths)\r\nwhile (True):\r\n    completed = results._index\r\n    print(\"\\r--- Completed {:,} out of {:,}\".format(completed, num_tasks), end='')\r\n    sys.stdout.flush()\r\n    time.sleep(1)\r\n    if (completed == num_tasks): break\r\np.close()\r\np.join()",
    "95230": "[quote=David Shinn;93933]\r\n\r\n[quote=paladin_o_newengland;93932]\r\n\r\nAny suggestions on modifying this would be really helpful!\r\n\r\n[/quote]\r\n\r\nTake this code:\r\n\r\n    p = multiprocessing.Pool()\r\n    results = p.imap(create_data, filepaths)\r\n    while (True):\r\n        completed = results._index\r\n        print(\"\\r--- Completed {:,} out of {:,}\".format(completed, num_tasks), end='')\r\n        sys.stdout.flush()\r\n        time.sleep(1)\r\n        if (completed == num_tasks): break\r\n    p.close()\r\n    p.join()\r\n\r\nand replace it with:\r\n\r\n    results = map(create_data, filepaths)\r\n\r\nIt'll work, but you won't get parallel processing.\r\n\r\n[/quote]\r\nHi\r\nI applied this but ended up with error\r\n\r\n---------------------------------------------------------------------------\r\nAttributeError                            Traceback (most recent call last)\r\n<ipython-input-5-6a305321d825> in <module>()\r\n      3 #results = p.imap(create_data, filepaths)\r\n      4 while (True):\r\n----> 5     completed = results._index\r\n      6     print(\"\\r--- Completed {:,} out of {:,}\".format(completed, num_tasks), end='')\r\n      7     sys.stdout.flush()\r\n\r\nAttributeError: 'list' object has no attribute '_index'\r\n\r\nAny idea how to resolve this?",
    "95253": "Hmm...I think the instructions were to replace **all** 10 lines with the 1 line.\r\n\r\nIn the end, don't get caught up with the exact script's implementation, all it is doing is counting 7 things in the document using basic Python string manipulation and conveniently running the Random Forest and producing the submission file, all in one script.  If you can't get the script to work by hacking it, then rewrite it from scratch and only include the parts you fully understand, and your coding skills will be better for it.",
    "95503": "thanks for helping me to learn new things Mr.David Shinn.i ran successfully.",
    "95636": "David,\r\n\r\nCan you confirm your code gets you 0.90388? or did you do anything on top of this to get there? I ran this code and I'm only get a ~.07\r\n\r\nThanks,\r\nJai",
    "95638": "Yep.  Proof is right off the leaderboard:\r\n\r\n    TeamId,TeamName,SubmissionDate,Score\r\n    219616,\"David Shinn\",\"2015-09-23 03:20:54\",0.90388\r\n    219815,\"Tim Kreienkamp\",\"2015-09-23 17:53:10\",0.90388\r\n    214906,BigLeak,\"2015-09-25 17:39:49\",0.90388\r\n    220366,\"Jack B.\",\"2015-09-26 00:38:58\",0.90388\r\n    220447,\"Rishab Gargeya\",\"2015-09-26 09:33:16\",0.90388\r\n    213808,YS,\"2015-09-26 11:01:55\",0.90388\r\n    220669,\"surya venkat\",\"2015-09-27 15:27:45\",0.90388\r\n    222164,\"Jeong-Yoon Lee\",\"2015-10-02 22:18:38\",0.90388\r\n    222349,paladin_o_newengland,\"2015-10-03 17:07:10\",0.90388\r\n    222340,Mihawk,\"2015-10-04 03:13:29\",0.90388\r\n    207517,\"Ayman Khalafallah\",\"2015-10-04 19:42:05\",0.90388\r\n    222826,MaksymPylypovych,\"2015-10-05 08:47:03\",0.90388\r\n    213434,\"Vikrant Kumar\",\"2015-10-07 05:31:51\",0.90388",
    "95639": "David, Thanks for the quick response. How did you convert the output from clf.predict_proba into 1/0?",
    "95640": "I didn't.  The evaluation metric is AUC, so you should be submitting a float probability between [0, 1].  Your submission file should look something like:\r\n\r\n    file,sponsored\r\n    1000043_raw_html.txt,0.0153013876291\r\n    1000097_raw_html.txt,0.01\r\n    1000253_raw_html.txt,0.11\r\n    100025_raw_html.txt,0.0\r\n    1000313_raw_html.txt,0.18\r\n    1000319_raw_html.txt,0.18\r\n    1000325_raw_html.txt,0.03\r\n    1000493_raw_html.txt,0.11\r\n    1000553_raw_html.txt,0.64",
    "95641": "David - thanks much. I appreciate your attitude towards teaching.",
    "95642": "Jai - You're welcome.  The \"School of Kaggle\" has been good to me.  There's five whole days left and it is easy to improve this script to do much better than 0.90388, so go do some damage!",
    "95643": "Just FYI, in the example I converted the \"values\" result that is returned from a dictionary to using \"slots\" in order to [save memory](https://stackoverflow.com/questions/472000/python-slots).  It seemed to help for me.",
    "95646": "Cool tip, I never knew that existed.  What line did you alter exactly?  Competitions with larger datasets definitely forces you to learn how to deal with memory and time efficiently.  I **just** learned how to pause a process using `kill -STOP` and `kill -CONT`, and in combination with [this tip][1], saved me from losing a 3 hour running process just now.\r\n\r\n\r\n  [1]: http://stackoverflow.com/a/7485831/2626968",
    "95648": "I defined a class and then populated the values and returned it like this:\r\n\r\n    class ResultFileEntry(object):\r\n\t#https://utcc.utoronto.ca/~cks/space/blog/python/WhatSlotsAreGoodFor\r\n\t#http://tech.oyster.com/save-ram-with-python-slots/\r\n\t#https://stackoverflow.com/questions/1336791/dictionary-vs-object-which-is-more-efficient-and-why\r\n\t#https://stackoverflow.com/questions/472000/python-slots\r\n\t# We use slots here to save memory, a dynamic dictionary is not needed\r\n\t__slots__ = ['filename', 'sponsored', 'lines', 'spaces', 'tabs', 'braces', ...]\r\n\tdef __init__(self, filename):\r\n\t\tself.file = filename\r\n\t\tself.sponsored = None #make sure this is none because there is a check for this later during training\r\n\t\tself.lines = 0\r\n\t\tself.spaces = 0\r\n\t\tself.tabs = 0\r\n\t\tself.braces = 0\r\n\t\t...\r\n\t#https://docs.python.org/2.7/reference/datamodel.html#special-method-names\r\n\tdef __iter__(self):\r\n\t\tyield self.file;\r\n\t\tyield self.sponsored;\r\n\t\tyield self.lines;\r\n\t\tyield self.spaces;\r\n\t\tyield self.tabs;\r\n\t\t...",
    "95663": "[quote=Jai;95639]\r\n\r\nDavid, Thanks for the quick response. How did you convert the output from clf.predict_proba into 1/0?\r\n\r\n[/quote]\r\n\r\nDon't do this. I did this by mistake (or rather, I didn't use predict_proba) for my first solution and with a cv of 0.79, I only got a lb score of 5.3. Took me some time to find this rather stupid mistake.",
    "95814": "Hi,\r\n\r\nFirst of all, thanks to [@David Shinn][1] for the inspiration!\r\n\r\nAlthough it hasn't been tested on Windows, hope the attached script can deal with some asked technical issues here.\r\n\r\nSome quick note:\r\n\r\n 1. No need to unzip [0-5].zip;\r\n 2. pip3 install psutil, or comment out related code;\r\n 3. Feel free to uncomment xgb related code to give it a shot;\r\n 4. Some pool and data frame operations changed, if anything unclear please ask me;\r\n 5. Wish it could let myself be focused on features, in the final couple of days.\r\n 6. LB public score 0.90751\r\n\r\nCheers!\r\n\r\n\r\n  [1]: https://www.kaggle.com/davidshinn",
    "95818": "[quote=psingman;95098]\r\n\r\nThank you, that definitely was my issue (facepalm). Just to make sure I'm not doing anything unnecessary, you have to unzip all the TXT files in 0.zip, 1.zip etc..? My 130GB laptop is running out of storage....\r\n\r\n[/quote]\r\n\r\nProbably too late but try this to avoid unzipping:\r\n\r\n    import pandas as pd\r\n    import zipfile\r\n\r\n    # read() gives you the content\r\n    zipped = zipfile.ZipFile(zip_file_path, 'r')\r\n    for path in zipped.namelist():\r\n        if not str(path).endswith('_raw_html.txt'):\r\n            continue\r\n        zipped.read(path)\r\n\r\n    # open() gives you the handle for pd\r\n    zipped_csv = zipfile.ZipFile(zipped_csv_path, 'r')\r\n    pd.read_csv(zipped_csv.open(the_csv_file_name_inside_the_zip))",
    "96578": "Thank you Barabbas. I'm not gonna run the script again but I'll use this format in the future to avoid unzipping."
  },
  "source": "meta"
}