{
  "id": 15825,
  "title": "Welcome & Getting Started",
  "url": "/competitions/dato-native/discussion/15825",
  "author_name": "",
  "post_date": "2015-08-07T04:48:56.437Z",
  "votes": 3,
  "comment_count": 23,
  "views": 7422,
  "content": "<p>Welcome!</p>\n\n<p>To get rolling, we have two baseline scripts attached:</p>\n\n<ul>\n<li><p><strong>process_html.py:</strong> processes the html to extract the core content, links, and urls for images.</p></li>\n<li><p><strong>classifier.py:</strong> is an intentionally modest baseline script which extracts some basic features and changes.</p></li>\n</ul>\n\n<p>The baseline uses GraphLab Create, the core product our team, at Dato, built for other data geeks like us. </p>\n\n<p>You can use any tools for the competition, of course, but we'd love to see how far you can get with GLC.  GLC has been used for a number of competitions.</p>\n\n<p>Check out our <a href=\"http://blog.dato.com/kaggle-competition-stumbleupon-and-dato-ask-whats-truly-native\">post</a> for how to register and get started w/GLC and then, please</p>\n\n<p>beat-the-baseline!</p>\n\n<p>Game on!</p>",
  "messages": [
    {
      "id": "88773",
      "postDate": "08/07/2015 04:48:56",
      "content": "<p>Welcome!</p>\n\n<p>To get rolling, we have two baseline scripts attached:</p>\n\n<ul>\n<li><p><strong>process_html.py:</strong> processes the html to extract the core content, links, and urls for images.</p></li>\n<li><p><strong>classifier.py:</strong> is an intentionally modest baseline script which extracts some basic features and changes.</p></li>\n</ul>\n\n<p>The baseline uses GraphLab Create, the core product our team, at Dato, built for other data geeks like us. </p>\n\n<p>You can use any tools for the competition, of course, but we'd love to see how far you can get with GLC.  GLC has been used for a number of competitions.</p>\n\n<p>Check out our <a href=\"http://blog.dato.com/kaggle-competition-stumbleupon-and-dato-ask-whats-truly-native\">post</a> for how to register and get started w/GLC and then, please</p>\n\n<p>beat-the-baseline!</p>\n\n<p>Game on!</p>",
      "rawMarkdown": "Welcome!\r\n\r\nTo get rolling, we have two baseline scripts attached:\r\n\r\n - **process_html.py:** processes the html to extract the core content, links, and urls for images.\r\n   \r\n -  **classifier.py:** is an intentionally modest baseline script which extracts some basic features and changes.\r\n\r\nThe baseline uses GraphLab Create, the core product our team, at Dato, built for other data geeks like us. \r\n\r\nYou can use any tools for the competition, of course, but we'd love to see how far you can get with GLC.  GLC has been used for a number of competitions.\r\n\r\nCheck out our [post][1] for how to register and get started w/GLC and then, please\r\n\r\nbeat-the-baseline!\r\n\r\nGame on!\r\n\r\n\r\n  [1]: http://blog.dato.com/kaggle-competition-stumbleupon-and-dato-ask-whats-truly-native",
      "votes": null
    },
    {
      "id": "88784",
      "postDate": "08/07/2015 06:19:53",
      "content": "<p>So, are these files part of the contest as well?</p>\n\n<p>Just kidding.... because they are full of HTML code inside (i.e., there is an error in your files)</p>",
      "rawMarkdown": "So, are these files part of the contest as well?\r\n\r\nJust kidding.... because they are full of HTML code inside (i.e., there is an error in your files)",
      "votes": null
    },
    {
      "id": "88785",
      "postDate": "08/07/2015 06:32:03",
      "content": "<p>Hi NxGTR,</p>\n\n<p>The correct scripts are attached on here.  Sorry about the confusion.</p>\n\n<p>Thanks,</p>\n\n<p>Charlie</p>",
      "rawMarkdown": "Hi NxGTR,\r\n\r\nThe correct scripts are attached on here.  Sorry about the confusion.\r\n\r\nThanks,\r\n\r\nCharlie",
      "votes": null
    },
    {
      "id": "88882",
      "postDate": "08/08/2015 02:55:48",
      "content": "<p>Hello!  @Cloofa can you explain what the <code>last_bucket</code> means in the <code>process_html.py</code> file?  I don't understand why bucket 0 is being processed, but not being saved, as that process seems inefficient!</p>\n\n<p>EDIT: I only downloaded bucket 0, to just get started. I realized that the script will start downloading any bucket other than '0', but after finishing with that bucket, move on to the others like normal.</p>\n\n<p>TL;DR I would recommmend downloading more than one bucket, even if you are just wanting to play around with a subset of the data.</p>",
      "rawMarkdown": "Hello!  @Cloofa can you explain what the ```last_bucket``` means in the ```process_html.py``` file?  I don't understand why bucket 0 is being processed, but not being saved, as that process seems inefficient!\r\n\r\nEDIT: I only downloaded bucket 0, to just get started. I realized that the script will start downloading any bucket other than '0', but after finishing with that bucket, move on to the others like normal.\r\n\r\nTL;DR I would recommmend downloading more than one bucket, even if you are just wanting to play around with a subset of the data.",
      "votes": null
    },
    {
      "id": "88905",
      "postDate": "08/08/2015 15:27:57",
      "content": "<p>@Ryan Louie:  the script is naive and assumes you have all your files unzipped in the same folder.  Since the uncompressed files are large the JSON files may end up being close to 35GB depending on how much you extract.  last_bucket is used to track the bucket you are currently parsing, so when the bucket changes, you save the last JSON for that bucket.  I did this on purpose, in the case an I/O error happens, you can always continue from your last bucket by modifying a few lines.  Feel free to do edit this so that you can save for any given chunk of data.  </p>\n\n<p>Also, this is a very basic script to get started.  If you are familiar with regexp methods, I recommend using that instead of BeautifulSoup4, as it may be faster.  We went with this to make things more interpretable for users who are not as familiar with parsing HTML's.</p>",
      "rawMarkdown": "Ryan Louie:  the script is naive and assumes you have all your files unzipped in the same folder.  Since the uncompressed files are large the JSON files may end up being close to 35GB depending on how much you extract.  last_bucket is used to track the bucket you are currently parsing, so when the bucket changes, you save the last JSON for that bucket.  I did this on purpose, in the case an I/O error happens, you can always continue from your last bucket by modifying a few lines.  Feel free to do edit this so that you can save for any given chunk of data.  \r\n\r\nAlso, this is a very basic script to get started.  If you are familiar with regexp methods, I recommend using that instead of BeautifulSoup4, as it may be faster.  We went with this to make things more interpretable for users who are not as familiar with parsing HTML's.",
      "votes": null
    },
    {
      "id": "88917",
      "postDate": "08/08/2015 18:34:45",
      "content": "<p>So for example you process chunk 0,1,2,3,4\nThen seems you will not save json files for chunk 4. There's no chunk 5, so that no bucket != last_bucket meet for chunk 4?</p>",
      "rawMarkdown": "So for example you process chunk 0,1,2,3,4\r\nThen seems you will not save json files for chunk 4. There's no chunk 5, so that no bucket != last_bucket meet for chunk 4?",
      "votes": null
    },
    {
      "id": "88921",
      "postDate": "08/08/2015 18:52:34",
      "content": "<p>@Jiming Ye:  good catch, here is a quick simple fix attached.  The other thing you can do is persist to disk a large JSON blob after every X html's are processed, since the buckets themselves are random.</p>",
      "rawMarkdown": "Jiming Ye:  good catch, here is a quick simple fix attached.  The other thing you can do is persist to disk a large JSON blob after every X html's are processed, since the buckets themselves are random.",
      "votes": null
    },
    {
      "id": "88935",
      "postDate": "08/09/2015 01:49:31",
      "content": "<p>It trained well. But, when predicting, it failed with an error. It looked like the test set was not read right somehow?</p>",
      "rawMarkdown": "It trained well. But, when predicting, it failed with an error. It looked like the test set was not read right somehow?",
      "votes": null
    },
    {
      "id": "88953",
      "postDate": "08/09/2015 12:54:45",
      "content": "<p>@Deep</p>\n\n<p>remove line:</p>\n\n<pre><code>test = test.dropna()\n</code></pre>\n\n<p>and you should get proper submission with score: <code>0.69938</code></p>",
      "rawMarkdown": "Deep\r\n\r\nremove line:\r\n\r\n    test = test.dropna()\r\n\r\nand you should get proper submission with score: `0.69938`",
      "votes": null
    },
    {
      "id": "88983",
      "postDate": "08/09/2015 23:23:45",
      "content": "<p>@Dowakin, I removed that line, it still fails with same error as before: </p>\n\n<p>Traceback (most recent call last): <br>\n  File &quot;classifier.py&quot;, line 80, in  <br>\n    ypred = model.predict(test) <br>\n  File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/toolkits/classifier/logistic_classifier.py&quot;, line 630, in pred\nict <br>\n    missing_value_action = missing_value_action) <br>\n  File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/toolkits/_supervised_learning.py&quot;, line 117, in predict <br>\n    target = _graphlab.toolkits._main.run('supervised_learning_predict', options) <br>\n  File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/toolkits/_main.py&quot;, line 64, in run <br>\n    (success, message, params) = unity.run_toolkit(toolkit_name, options) <br>\n  File &quot;graphlab/cython/cy_unity.pyx&quot;, line 70, in graphlab.cython.cy_unity.UnityGlobalProxy.run_toolkit <br>\n  File &quot;graphlab/cython/cy_unity.pyx&quot;, line 74, in graphlab.cython.cy_unity.UnityGlobalProxy.run_toolkit <br>\nRuntimeError: Runtime Exception. basic_string::resize </p>",
      "rawMarkdown": "Dowakin, I removed that line, it still fails with same error as before: \r\n\r\nTraceback (most recent call last):                                                                                    \r\n  File \"classifier.py\", line 80, in <module>                                                                          \r\n    ypred = model.predict(test)                                                                                       \r\n  File \"/usr/local/lib/python2.7/dist-packages/graphlab/toolkits/classifier/logistic_classifier.py\", line 630, in pred\r\nict                                                                                                                   \r\n    missing_value_action = missing_value_action)                                                                    \r\n  File \"/usr/local/lib/python2.7/dist-packages/graphlab/toolkits/_supervised_learning.py\", line 117, in predict     \r\n    target = _graphlab.toolkits._main.run('supervised_learning_predict', options)                                    \r\n  File \"/usr/local/lib/python2.7/dist-packages/graphlab/toolkits/_main.py\", line 64, in run                           \r\n    (success, message, params) = unity.run_toolkit(toolkit_name, options)                                           \r\n  File \"graphlab/cython/cy_unity.pyx\", line 70, in graphlab.cython.cy_unity.UnityGlobalProxy.run_toolkit            \r\n  File \"graphlab/cython/cy_unity.pyx\", line 74, in graphlab.cython.cy_unity.UnityGlobalProxy.run_toolkit            \r\nRuntimeError: Runtime Exception. basic_string::resize",
      "votes": null
    },
    {
      "id": "88987",
      "postDate": "08/10/2015 00:04:34",
      "content": "<p>@Deep</p>\n\n<p>Oh, I just had error with submission, different number of rows.</p>\n\n<p>I installed everything in different conda env, just like from their instruction. So if you installed globally, you should try install it locally into new conda env.</p>",
      "rawMarkdown": "Deep\r\n\r\nOh, I just had error with submission, different number of rows.\r\n\r\nI installed everything in different conda env, just like from their instruction. So if you installed globally, you should try install it locally into new conda env.",
      "votes": null
    },
    {
      "id": "88998",
      "postDate": "08/10/2015 03:16:30",
      "content": "<p>@Dowakin, I was using virtualenv, but, I will try conda env. Thanks for the tip!</p>",
      "rawMarkdown": "Dowakin, I was using virtualenv, but, I will try conda env. Thanks for the tip!",
      "votes": null
    },
    {
      "id": "89840",
      "postDate": "08/19/2015 19:58:37",
      "content": "<p>I an new here ...i have this small issue when i ran the code -&gt; <strong>submission.save is creating an empty csv file although the sframe submission has data .Anyone else facing this issue ?</strong> </p>",
      "rawMarkdown": "I an new here ...i have this small issue when i ran the code -> **submission.save is creating an empty csv file although the sframe submission has data .Anyone else facing this issue ?**",
      "votes": null
    },
    {
      "id": "90014",
      "postDate": "08/21/2015 15:56:25",
      "content": "<p>Can somebody clarify which pass should bu used? \nPATH_TO_JSON = &quot;path/to/data/from/process_html.py&quot;</p>\n\n<p>I've used path to directory where I've downloaded all archives {0,1,2,3,4}\nbut script fails when parsing compleats.</p>\n\n<p>Probably, I'm missing something.\nI've got the following error.</p>\n\n<pre><code>Traceback (most recent call last):\n  File &quot;/Users/taras-sereda/PycharmProjects/dato/classifier.py&quot;, line 33, in &lt;module&gt;\n    sf = sf.unpack('X1',column_name_prefix='')\n  File &quot;/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sframe.py&quot;, line 4678, in unpack\n    new_sf = self[unpack_column].unpack(column_name_prefix, column_types, na_value, limit)\n  File &quot;/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sarray.py&quot;, line 2692, in unpack\n    raise TypeError(&quot;Only SArray of dict/list/array type supports unpack&quot;)\nTypeError: Only SArray of dict/list/array type supports unpack\n[INFO] Stopping the server connection.\n</code></pre>",
      "rawMarkdown": "Can somebody clarify which pass should bu used? \r\nPATH_TO_JSON = \"path/to/data/from/process_html.py\"\r\n\r\nI've used path to directory where I've downloaded all archives {0,1,2,3,4}\r\nbut script fails when parsing compleats.\r\n\r\nProbably, I'm missing something.\r\nI've got the following error.\r\n\r\n    Traceback (most recent call last):\r\n      File \"/Users/taras-sereda/PycharmProjects/dato/classifier.py\", line 33, in <module>\r\n        sf = sf.unpack('X1',column_name_prefix='')\r\n      File \"/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sframe.py\", line 4678, in unpack\r\n        new_sf = self[unpack_column].unpack(column_name_prefix, column_types, na_value, limit)\r\n      File \"/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sarray.py\", line 2692, in unpack\r\n        raise TypeError(\"Only SArray of dict/list/array type supports unpack\")\r\n    TypeError: Only SArray of dict/list/array type supports unpack\r\n    [INFO] Stopping the server connection.",
      "votes": null
    },
    {
      "id": "90058",
      "postDate": "08/21/2015 23:46:36",
      "content": "<p>[quote=taras sereda;90014]</p>\n\n<p>Can somebody clarify which pass should bu used? \nPATH_TO_JSON = &quot;path/to/data/from/process_html.py&quot;</p>\n\n<p>I've used path to directory where I've downloaded all archives {0,1,2,3,4}\nbut script fails when parsing compleats.</p>\n\n<p>Probably, I'm missing something.\nI've got the following error.</p>\n\n<pre><code>Traceback (most recent call last):\n  File &quot;/Users/taras-sereda/PycharmProjects/dato/classifier.py&quot;, line 33, in &lt;module&gt;\n    sf = sf.unpack('X1',column_name_prefix='')\n  File &quot;/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sframe.py&quot;, line 4678, in unpack\n    new_sf = self[unpack_column].unpack(column_name_prefix, column_types, na_value, limit)\n  File &quot;/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sarray.py&quot;, line 2692, in unpack\n    raise TypeError(&quot;Only SArray of dict/list/array type supports unpack&quot;)\nTypeError: Only SArray of dict/list/array type supports unpack\n[INFO] Stopping the server connection.\n</code></pre>\n\n<p>[/quote]</p>\n\n<p>My guess is that you need to run process_html.py first (takes about 8 hours for me) and you will get the json files at the output.\nHowever, I have no ideas what to do after that. In the classifier.py, the code tries to open something as csv file in  PATH_TO_JSON, making little sense. Any hints please?</p>",
      "rawMarkdown": "[quote=taras sereda;90014]\r\n\r\nCan somebody clarify which pass should bu used? \r\nPATH_TO_JSON = \"path/to/data/from/process_html.py\"\r\n\r\nI've used path to directory where I've downloaded all archives {0,1,2,3,4}\r\nbut script fails when parsing compleats.\r\n\r\nProbably, I'm missing something.\r\nI've got the following error.\r\n\r\n    Traceback (most recent call last):\r\n      File \"/Users/taras-sereda/PycharmProjects/dato/classifier.py\", line 33, in <module>\r\n        sf = sf.unpack('X1',column_name_prefix='')\r\n      File \"/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sframe.py\", line 4678, in unpack\r\n        new_sf = self[unpack_column].unpack(column_name_prefix, column_types, na_value, limit)\r\n      File \"/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sarray.py\", line 2692, in unpack\r\n        raise TypeError(\"Only SArray of dict/list/array type supports unpack\")\r\n    TypeError: Only SArray of dict/list/array type supports unpack\r\n    [INFO] Stopping the server connection.\r\n\r\n[/quote]\r\n\r\nMy guess is that you need to run process_html.py first (takes about 8 hours for me) and you will get the json files at the output.\r\nHowever, I have no ideas what to do after that. In the classifier.py, the code tries to open something as csv file in  PATH_TO_JSON, making little sense. Any hints please?",
      "votes": null
    },
    {
      "id": "90089",
      "postDate": "08/22/2015 07:33:12",
      "content": "<p>Thanks @FTD2016. I've solved an issue. \nInitially I was expecting a single run of script classifier.py.</p>\n\n<p>the problem was that I missunderstood which path to provide.\nPATH_TO_JSON = &quot;path/to/data/from/process_html.py&quot;</p>\n\n<p>In classifier.py. csv is used, because each parsed chunk, is not a True json file, actually it's a text file where each line is valid json. So this formulation is also a little bit misleading.</p>\n\n<p>the structure of each line in file is a parsed version of raw html in  each chunk.\nAs described in parsing script.</p>\n\n<pre><code>doc = {\n        &quot;id&quot;: urlid, \n        &quot;text&quot;:parse_text(soup),\n        &quot;title&quot;:parse_title(soup ),\n        &quot;links&quot;:parse_links(soup),\n        &quot;images&quot;:parse_images(soup),\n       }\n</code></pre>\n\n<p>Hope, that helped. </p>",
      "rawMarkdown": "Thanks @FTD2016. I've solved an issue. \r\nInitially I was expecting a single run of script classifier.py.\r\n\r\nthe problem was that I missunderstood which path to provide.\r\nPATH_TO_JSON = \"path/to/data/from/process_html.py\"\r\n\r\nIn classifier.py. csv is used, because each parsed chunk, is not a True json file, actually it's a text file where each line is valid json. So this formulation is also a little bit misleading.\r\n\r\nthe structure of each line in file is a parsed version of raw html in  each chunk.\r\nAs described in parsing script.\r\n\r\n    doc = {\r\n            \"id\": urlid, \r\n            \"text\":parse_text(soup),\r\n            \"title\":parse_title(soup ),\r\n            \"links\":parse_links(soup),\r\n            \"images\":parse_images(soup),\r\n           }\r\nHope, that helped.",
      "votes": null
    },
    {
      "id": "90167",
      "postDate": "08/23/2015 18:10:44",
      "content": "<p>Hi, how did you fix your issue? I ran the process_html.py script on a small subset of files which generated a chunk0.json. I then ran the classifier script. It runs fine till line 34</p>\n\n<pre><code>sf = gl.SFrame.read_csv(PATH_TO_JSON, header=False)\n</code></pre>\n\n<p>and I get &quot;Parsing completed&quot;.</p>\n\n<p>However it fails on the next line with a type mismatch.</p>\n\n<p>I would appreciate any help. Thanks!</p>",
      "rawMarkdown": "Hi, how did you fix your issue? I ran the process_html.py script on a small subset of files which generated a chunk0.json. I then ran the classifier script. It runs fine till line 34\r\n\r\n    sf = gl.SFrame.read_csv(PATH_TO_JSON, header=False)\r\n\r\nand I get \"Parsing completed\".\r\n\r\nHowever it fails on the next line with a type mismatch.\r\n\r\nI would appreciate any help. Thanks!",
      "votes": null
    },
    {
      "id": "91603",
      "postDate": "09/04/2015 18:53:57",
      "content": "<p>I want to ask if the code of process_html is bug free or if it needs debugging</p>",
      "rawMarkdown": "I want to ask if the code of process_html is bug free or if it needs debugging",
      "votes": null
    },
    {
      "id": "91778",
      "postDate": "09/07/2015 15:41:41",
      "content": "<p>[quote=cloofa;88905]\nAlso, this is a very basic script to get started.  If you are familiar with regexp methods, I recommend using that instead of BeautifulSoup4, as it may be faster.  We went with this to make things more interpretable for users who are not as familiar with parsing HTML's.\n[/quote]</p>\n\n<p>It's somewhat funny and refreshing to read this as the general opinion about regex being used for html parsing is &quot;you should know better and not do that&quot;.</p>",
      "rawMarkdown": "[quote=cloofa;88905]\r\nAlso, this is a very basic script to get started.  If you are familiar with regexp methods, I recommend using that instead of BeautifulSoup4, as it may be faster.  We went with this to make things more interpretable for users who are not as familiar with parsing HTML's.\r\n[/quote]\r\n\r\nIt's somewhat funny and refreshing to read this as the general opinion about regex being used for html parsing is \"you should know better and not do that\".",
      "votes": null
    },
    {
      "id": "92230",
      "postDate": "09/12/2015 04:45:32",
      "content": "<p>Hi!</p>\n\n<p>Did anybody get - TypeError:  is not JSON serializable when the process_html.py was run?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi!\r\n\r\nDid anybody get - TypeError: <filter object at 0x00000000163204E0> is not JSON serializable when the process_html.py was run?\r\n\r\nThanks",
      "votes": null
    },
    {
      "id": "93041",
      "postDate": "09/20/2015 19:24:04",
      "content": "<p>Hi,\nThanks for this starter code\nAnother issue\nbow_trn = gl.text_analytics.count_words(train['text_clean'])\ncomplains about count_words needs string SArrays ...\nAny idea ?\nTIA\nRgds\nBruno</p>",
      "rawMarkdown": "Hi,\r\nThanks for this starter code\r\nAnother issue\r\nbow_trn = gl.text_analytics.count_words(train['text_clean'])\r\ncomplains about count_words needs string SArrays ...\r\nAny idea ?\r\nTIA\r\nRgds\r\nBruno",
      "votes": null
    },
    {
      "id": "93068",
      "postDate": "09/21/2015 06:01:19",
      "content": "<p>Hi All,</p>\n\n<p>Anyone else encounter this error?  I am using Python 3.4.3 that comes with Anaconda.</p>\n\n<p>Traceback (most recent call last):\n  File &quot;process_html_new.py&quot;, line 158, in \n    main(sys.argv)\n  File &quot;process_html_new.py&quot;, line 139, in main\n    json.dump(entry, feedsjson)\n  File &quot;/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/<strong>init</strong>.py&quot;, line 178, in dump\n    for chunk in iterable:\n  File &quot;/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py&quot;, line 422, in _iterencode\n    yield from _iterencode_dict(o, _current_indent_level)\n  File &quot;/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py&quot;, line 396, in _iterencode_dict\n    yield from chunks\n  File &quot;/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py&quot;, line 429, in _iterencode\n    o = _default(o)\n  File &quot;/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py&quot;, line 173, in default\n    raise TypeError(repr(o) + &quot; is not JSON serializable&quot;)\nTypeError:  is not JSON serializable</p>\n\n<p>Regards,\nNaveen</p>",
      "rawMarkdown": "Hi All,\r\n\r\nAnyone else encounter this error?  I am using Python 3.4.3 that comes with Anaconda.\r\n\r\nTraceback (most recent call last):\r\n  File \"process_html_new.py\", line 158, in <module>\r\n    main(sys.argv)\r\n  File \"process_html_new.py\", line 139, in main\r\n    json.dump(entry, feedsjson)\r\n  File \"/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/__init__.py\", line 178, in dump\r\n    for chunk in iterable:\r\n  File \"/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py\", line 422, in _iterencode\r\n    yield from _iterencode_dict(o, _current_indent_level)\r\n  File \"/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py\", line 396, in _iterencode_dict\r\n    yield from chunks\r\n  File \"/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py\", line 429, in _iterencode\r\n    o = _default(o)\r\n  File \"/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py\", line 173, in default\r\n    raise TypeError(repr(o) + \" is not JSON serializable\")\r\nTypeError: <filter object at 0x102cde550> is not JSON serializable\r\n\r\nRegards,\r\nNaveen",
      "votes": null
    },
    {
      "id": "95533",
      "postDate": "10/08/2015 19:40:43",
      "content": "<p>Hi Competition Admin,\nI have had several submissions before the data leak. Recently, due to some personal commitments, I couldn't submit a new entry.</p>\n\n<p>Now, I just passed the deadline, which I overlooked. Would it be possible to allow previous submitters like myself to submit again?</p>\n\n<p>Very much appreciated. Thank you.</p>\n\n<p>Best,\nHarish</p>",
      "rawMarkdown": "Hi Competition Admin,\r\nI have had several submissions before the data leak. Recently, due to some personal commitments, I couldn't submit a new entry.\r\n\r\nNow, I just passed the deadline, which I overlooked. Would it be possible to allow previous submitters like myself to submit again?\r\n\r\nVery much appreciated. Thank you.\r\n\r\nBest,\r\nHarish",
      "votes": null
    },
    {
      "id": "96126",
      "postDate": "10/14/2015 17:01:44",
      "content": "<p>The attached script is an almost automatic workflow for GraphLab Create, except its input relies on the stored feature set from <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model/95814#post95814\">https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model/95814#post95814</a> , which is inspired by <a href=\"https://www.kaggle.com/davidshinn\">@David Shinn</a>'s nice work.</p>",
      "rawMarkdown": "The attached script is an almost automatic workflow for GraphLab Create, except its input relies on the stored feature set from https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model/95814#post95814 , which is inspired by [@David Shinn][1]'s nice work.\r\n\r\n\r\n  [1]: https://www.kaggle.com/davidshinn",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 88784,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "08/07/2015 06:19:53",
      "content": "<p>So, are these files part of the contest as well?</p>\n\n<p>Just kidding.... because they are full of HTML code inside (i.e., there is an error in your files)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88785,
      "author_name": "cloofa",
      "author_url": "",
      "post_date": "08/07/2015 06:32:03",
      "content": "<p>Hi NxGTR,</p>\n\n<p>The correct scripts are attached on here.  Sorry about the confusion.</p>\n\n<p>Thanks,</p>\n\n<p>Charlie</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88882,
      "author_name": "ryanlouie",
      "author_url": "",
      "post_date": "08/08/2015 02:55:48",
      "content": "<p>Hello!  @Cloofa can you explain what the <code>last_bucket</code> means in the <code>process_html.py</code> file?  I don't understand why bucket 0 is being processed, but not being saved, as that process seems inefficient!</p>\n\n<p>EDIT: I only downloaded bucket 0, to just get started. I realized that the script will start downloading any bucket other than '0', but after finishing with that bucket, move on to the others like normal.</p>\n\n<p>TL;DR I would recommmend downloading more than one bucket, even if you are just wanting to play around with a subset of the data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88905,
      "author_name": "cloofa",
      "author_url": "",
      "post_date": "08/08/2015 15:27:57",
      "content": "<p>@Ryan Louie:  the script is naive and assumes you have all your files unzipped in the same folder.  Since the uncompressed files are large the JSON files may end up being close to 35GB depending on how much you extract.  last_bucket is used to track the bucket you are currently parsing, so when the bucket changes, you save the last JSON for that bucket.  I did this on purpose, in the case an I/O error happens, you can always continue from your last bucket by modifying a few lines.  Feel free to do edit this so that you can save for any given chunk of data.  </p>\n\n<p>Also, this is a very basic script to get started.  If you are familiar with regexp methods, I recommend using that instead of BeautifulSoup4, as it may be faster.  We went with this to make things more interpretable for users who are not as familiar with parsing HTML's.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88917,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "08/08/2015 18:34:45",
      "content": "<p>So for example you process chunk 0,1,2,3,4\nThen seems you will not save json files for chunk 4. There's no chunk 5, so that no bucket != last_bucket meet for chunk 4?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88921,
      "author_name": "cloofa",
      "author_url": "",
      "post_date": "08/08/2015 18:52:34",
      "content": "<p>@Jiming Ye:  good catch, here is a quick simple fix attached.  The other thing you can do is persist to disk a large JSON blob after every X html's are processed, since the buckets themselves are random.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88935,
      "author_name": "deepcnn",
      "author_url": "",
      "post_date": "08/09/2015 01:49:31",
      "content": "<p>It trained well. But, when predicting, it failed with an error. It looked like the test set was not read right somehow?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88953,
      "author_name": "dowakin",
      "author_url": "",
      "post_date": "08/09/2015 12:54:45",
      "content": "<p>@Deep</p>\n\n<p>remove line:</p>\n\n<pre><code>test = test.dropna()\n</code></pre>\n\n<p>and you should get proper submission with score: <code>0.69938</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88983,
      "author_name": "deepcnn",
      "author_url": "",
      "post_date": "08/09/2015 23:23:45",
      "content": "<p>@Dowakin, I removed that line, it still fails with same error as before: </p>\n\n<p>Traceback (most recent call last): <br>\n  File &quot;classifier.py&quot;, line 80, in  <br>\n    ypred = model.predict(test) <br>\n  File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/toolkits/classifier/logistic_classifier.py&quot;, line 630, in pred\nict <br>\n    missing_value_action = missing_value_action) <br>\n  File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/toolkits/_supervised_learning.py&quot;, line 117, in predict <br>\n    target = _graphlab.toolkits._main.run('supervised_learning_predict', options) <br>\n  File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/toolkits/_main.py&quot;, line 64, in run <br>\n    (success, message, params) = unity.run_toolkit(toolkit_name, options) <br>\n  File &quot;graphlab/cython/cy_unity.pyx&quot;, line 70, in graphlab.cython.cy_unity.UnityGlobalProxy.run_toolkit <br>\n  File &quot;graphlab/cython/cy_unity.pyx&quot;, line 74, in graphlab.cython.cy_unity.UnityGlobalProxy.run_toolkit <br>\nRuntimeError: Runtime Exception. basic_string::resize </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88987,
      "author_name": "dowakin",
      "author_url": "",
      "post_date": "08/10/2015 00:04:34",
      "content": "<p>@Deep</p>\n\n<p>Oh, I just had error with submission, different number of rows.</p>\n\n<p>I installed everything in different conda env, just like from their instruction. So if you installed globally, you should try install it locally into new conda env.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88998,
      "author_name": "deepcnn",
      "author_url": "",
      "post_date": "08/10/2015 03:16:30",
      "content": "<p>@Dowakin, I was using virtualenv, but, I will try conda env. Thanks for the tip!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 89840,
      "author_name": "arnabsarkar",
      "author_url": "",
      "post_date": "08/19/2015 19:58:37",
      "content": "<p>I an new here ...i have this small issue when i ran the code -&gt; <strong>submission.save is creating an empty csv file although the sframe submission has data .Anyone else facing this issue ?</strong> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90014,
      "author_name": "tsereda",
      "author_url": "",
      "post_date": "08/21/2015 15:56:25",
      "content": "<p>Can somebody clarify which pass should bu used? \nPATH_TO_JSON = &quot;path/to/data/from/process_html.py&quot;</p>\n\n<p>I've used path to directory where I've downloaded all archives {0,1,2,3,4}\nbut script fails when parsing compleats.</p>\n\n<p>Probably, I'm missing something.\nI've got the following error.</p>\n\n<pre><code>Traceback (most recent call last):\n  File &quot;/Users/taras-sereda/PycharmProjects/dato/classifier.py&quot;, line 33, in &lt;module&gt;\n    sf = sf.unpack('X1',column_name_prefix='')\n  File &quot;/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sframe.py&quot;, line 4678, in unpack\n    new_sf = self[unpack_column].unpack(column_name_prefix, column_types, na_value, limit)\n  File &quot;/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sarray.py&quot;, line 2692, in unpack\n    raise TypeError(&quot;Only SArray of dict/list/array type supports unpack&quot;)\nTypeError: Only SArray of dict/list/array type supports unpack\n[INFO] Stopping the server connection.\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90058,
      "author_name": "therewasadream",
      "author_url": "",
      "post_date": "08/21/2015 23:46:36",
      "content": "<p>[quote=taras sereda;90014]</p>\n\n<p>Can somebody clarify which pass should bu used? \nPATH_TO_JSON = &quot;path/to/data/from/process_html.py&quot;</p>\n\n<p>I've used path to directory where I've downloaded all archives {0,1,2,3,4}\nbut script fails when parsing compleats.</p>\n\n<p>Probably, I'm missing something.\nI've got the following error.</p>\n\n<pre><code>Traceback (most recent call last):\n  File &quot;/Users/taras-sereda/PycharmProjects/dato/classifier.py&quot;, line 33, in &lt;module&gt;\n    sf = sf.unpack('X1',column_name_prefix='')\n  File &quot;/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sframe.py&quot;, line 4678, in unpack\n    new_sf = self[unpack_column].unpack(column_name_prefix, column_types, na_value, limit)\n  File &quot;/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sarray.py&quot;, line 2692, in unpack\n    raise TypeError(&quot;Only SArray of dict/list/array type supports unpack&quot;)\nTypeError: Only SArray of dict/list/array type supports unpack\n[INFO] Stopping the server connection.\n</code></pre>\n\n<p>[/quote]</p>\n\n<p>My guess is that you need to run process_html.py first (takes about 8 hours for me) and you will get the json files at the output.\nHowever, I have no ideas what to do after that. In the classifier.py, the code tries to open something as csv file in  PATH_TO_JSON, making little sense. Any hints please?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90089,
      "author_name": "tsereda",
      "author_url": "",
      "post_date": "08/22/2015 07:33:12",
      "content": "<p>Thanks @FTD2016. I've solved an issue. \nInitially I was expecting a single run of script classifier.py.</p>\n\n<p>the problem was that I missunderstood which path to provide.\nPATH_TO_JSON = &quot;path/to/data/from/process_html.py&quot;</p>\n\n<p>In classifier.py. csv is used, because each parsed chunk, is not a True json file, actually it's a text file where each line is valid json. So this formulation is also a little bit misleading.</p>\n\n<p>the structure of each line in file is a parsed version of raw html in  each chunk.\nAs described in parsing script.</p>\n\n<pre><code>doc = {\n        &quot;id&quot;: urlid, \n        &quot;text&quot;:parse_text(soup),\n        &quot;title&quot;:parse_title(soup ),\n        &quot;links&quot;:parse_links(soup),\n        &quot;images&quot;:parse_images(soup),\n       }\n</code></pre>\n\n<p>Hope, that helped. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90167,
      "author_name": "saurabhsaxena2",
      "author_url": "",
      "post_date": "08/23/2015 18:10:44",
      "content": "<p>Hi, how did you fix your issue? I ran the process_html.py script on a small subset of files which generated a chunk0.json. I then ran the classifier script. It runs fine till line 34</p>\n\n<pre><code>sf = gl.SFrame.read_csv(PATH_TO_JSON, header=False)\n</code></pre>\n\n<p>and I get &quot;Parsing completed&quot;.</p>\n\n<p>However it fails on the next line with a type mismatch.</p>\n\n<p>I would appreciate any help. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91603,
      "author_name": "ma2afify",
      "author_url": "",
      "post_date": "09/04/2015 18:53:57",
      "content": "<p>I want to ask if the code of process_html is bug free or if it needs debugging</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91778,
      "author_name": "wildwizard",
      "author_url": "",
      "post_date": "09/07/2015 15:41:41",
      "content": "<p>[quote=cloofa;88905]\nAlso, this is a very basic script to get started.  If you are familiar with regexp methods, I recommend using that instead of BeautifulSoup4, as it may be faster.  We went with this to make things more interpretable for users who are not as familiar with parsing HTML's.\n[/quote]</p>\n\n<p>It's somewhat funny and refreshing to read this as the general opinion about regex being used for html parsing is &quot;you should know better and not do that&quot;.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92230,
      "author_name": "dataparty",
      "author_url": "",
      "post_date": "09/12/2015 04:45:32",
      "content": "<p>Hi!</p>\n\n<p>Did anybody get - TypeError:  is not JSON serializable when the process_html.py was run?</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93041,
      "author_name": "bruno16",
      "author_url": "",
      "post_date": "09/20/2015 19:24:04",
      "content": "<p>Hi,\nThanks for this starter code\nAnother issue\nbow_trn = gl.text_analytics.count_words(train['text_clean'])\ncomplains about count_words needs string SArrays ...\nAny idea ?\nTIA\nRgds\nBruno</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93068,
      "author_name": "kumarnt",
      "author_url": "",
      "post_date": "09/21/2015 06:01:19",
      "content": "<p>Hi All,</p>\n\n<p>Anyone else encounter this error?  I am using Python 3.4.3 that comes with Anaconda.</p>\n\n<p>Traceback (most recent call last):\n  File &quot;process_html_new.py&quot;, line 158, in \n    main(sys.argv)\n  File &quot;process_html_new.py&quot;, line 139, in main\n    json.dump(entry, feedsjson)\n  File &quot;/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/<strong>init</strong>.py&quot;, line 178, in dump\n    for chunk in iterable:\n  File &quot;/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py&quot;, line 422, in _iterencode\n    yield from _iterencode_dict(o, _current_indent_level)\n  File &quot;/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py&quot;, line 396, in _iterencode_dict\n    yield from chunks\n  File &quot;/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py&quot;, line 429, in _iterencode\n    o = _default(o)\n  File &quot;/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py&quot;, line 173, in default\n    raise TypeError(repr(o) + &quot; is not JSON serializable&quot;)\nTypeError:  is not JSON serializable</p>\n\n<p>Regards,\nNaveen</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95533,
      "author_name": "hvsarma",
      "author_url": "",
      "post_date": "10/08/2015 19:40:43",
      "content": "<p>Hi Competition Admin,\nI have had several submissions before the data leak. Recently, due to some personal commitments, I couldn't submit a new entry.</p>\n\n<p>Now, I just passed the deadline, which I overlooked. Would it be possible to allow previous submitters like myself to submit again?</p>\n\n<p>Very much appreciated. Thank you.</p>\n\n<p>Best,\nHarish</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96126,
      "author_name": "tmjiang",
      "author_url": "",
      "post_date": "10/14/2015 17:01:44",
      "content": "<p>The attached script is an almost automatic workflow for GraphLab Create, except its input relies on the stored feature set from <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model/95814#post95814\">https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model/95814#post95814</a> , which is inspired by <a href=\"https://www.kaggle.com/davidshinn\">@David Shinn</a>'s nice work.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "88773": "Welcome!\r\n\r\nTo get rolling, we have two baseline scripts attached:\r\n\r\n - **process_html.py:** processes the html to extract the core content, links, and urls for images.\r\n   \r\n -  **classifier.py:** is an intentionally modest baseline script which extracts some basic features and changes.\r\n\r\nThe baseline uses GraphLab Create, the core product our team, at Dato, built for other data geeks like us. \r\n\r\nYou can use any tools for the competition, of course, but we'd love to see how far you can get with GLC.  GLC has been used for a number of competitions.\r\n\r\nCheck out our [post][1] for how to register and get started w/GLC and then, please\r\n\r\nbeat-the-baseline!\r\n\r\nGame on!\r\n\r\n\r\n  [1]: http://blog.dato.com/kaggle-competition-stumbleupon-and-dato-ask-whats-truly-native",
    "88784": "So, are these files part of the contest as well?\r\n\r\nJust kidding.... because they are full of HTML code inside (i.e., there is an error in your files)",
    "88785": "Hi NxGTR,\r\n\r\nThe correct scripts are attached on here.  Sorry about the confusion.\r\n\r\nThanks,\r\n\r\nCharlie",
    "88882": "Hello!  @Cloofa can you explain what the ```last_bucket``` means in the ```process_html.py``` file?  I don't understand why bucket 0 is being processed, but not being saved, as that process seems inefficient!\r\n\r\nEDIT: I only downloaded bucket 0, to just get started. I realized that the script will start downloading any bucket other than '0', but after finishing with that bucket, move on to the others like normal.\r\n\r\nTL;DR I would recommmend downloading more than one bucket, even if you are just wanting to play around with a subset of the data.",
    "88905": "Ryan Louie:  the script is naive and assumes you have all your files unzipped in the same folder.  Since the uncompressed files are large the JSON files may end up being close to 35GB depending on how much you extract.  last_bucket is used to track the bucket you are currently parsing, so when the bucket changes, you save the last JSON for that bucket.  I did this on purpose, in the case an I/O error happens, you can always continue from your last bucket by modifying a few lines.  Feel free to do edit this so that you can save for any given chunk of data.  \r\n\r\nAlso, this is a very basic script to get started.  If you are familiar with regexp methods, I recommend using that instead of BeautifulSoup4, as it may be faster.  We went with this to make things more interpretable for users who are not as familiar with parsing HTML's.",
    "88917": "So for example you process chunk 0,1,2,3,4\r\nThen seems you will not save json files for chunk 4. There's no chunk 5, so that no bucket != last_bucket meet for chunk 4?",
    "88921": "Jiming Ye:  good catch, here is a quick simple fix attached.  The other thing you can do is persist to disk a large JSON blob after every X html's are processed, since the buckets themselves are random.",
    "88935": "It trained well. But, when predicting, it failed with an error. It looked like the test set was not read right somehow?",
    "88953": "Deep\r\n\r\nremove line:\r\n\r\n    test = test.dropna()\r\n\r\nand you should get proper submission with score: `0.69938`",
    "88983": "Dowakin, I removed that line, it still fails with same error as before: \r\n\r\nTraceback (most recent call last):                                                                                    \r\n  File \"classifier.py\", line 80, in <module>                                                                          \r\n    ypred = model.predict(test)                                                                                       \r\n  File \"/usr/local/lib/python2.7/dist-packages/graphlab/toolkits/classifier/logistic_classifier.py\", line 630, in pred\r\nict                                                                                                                   \r\n    missing_value_action = missing_value_action)                                                                    \r\n  File \"/usr/local/lib/python2.7/dist-packages/graphlab/toolkits/_supervised_learning.py\", line 117, in predict     \r\n    target = _graphlab.toolkits._main.run('supervised_learning_predict', options)                                    \r\n  File \"/usr/local/lib/python2.7/dist-packages/graphlab/toolkits/_main.py\", line 64, in run                           \r\n    (success, message, params) = unity.run_toolkit(toolkit_name, options)                                           \r\n  File \"graphlab/cython/cy_unity.pyx\", line 70, in graphlab.cython.cy_unity.UnityGlobalProxy.run_toolkit            \r\n  File \"graphlab/cython/cy_unity.pyx\", line 74, in graphlab.cython.cy_unity.UnityGlobalProxy.run_toolkit            \r\nRuntimeError: Runtime Exception. basic_string::resize",
    "88987": "Deep\r\n\r\nOh, I just had error with submission, different number of rows.\r\n\r\nI installed everything in different conda env, just like from their instruction. So if you installed globally, you should try install it locally into new conda env.",
    "88998": "Dowakin, I was using virtualenv, but, I will try conda env. Thanks for the tip!",
    "89840": "I an new here ...i have this small issue when i ran the code -> **submission.save is creating an empty csv file although the sframe submission has data .Anyone else facing this issue ?**",
    "90014": "Can somebody clarify which pass should bu used? \r\nPATH_TO_JSON = \"path/to/data/from/process_html.py\"\r\n\r\nI've used path to directory where I've downloaded all archives {0,1,2,3,4}\r\nbut script fails when parsing compleats.\r\n\r\nProbably, I'm missing something.\r\nI've got the following error.\r\n\r\n    Traceback (most recent call last):\r\n      File \"/Users/taras-sereda/PycharmProjects/dato/classifier.py\", line 33, in <module>\r\n        sf = sf.unpack('X1',column_name_prefix='')\r\n      File \"/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sframe.py\", line 4678, in unpack\r\n        new_sf = self[unpack_column].unpack(column_name_prefix, column_types, na_value, limit)\r\n      File \"/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sarray.py\", line 2692, in unpack\r\n        raise TypeError(\"Only SArray of dict/list/array type supports unpack\")\r\n    TypeError: Only SArray of dict/list/array type supports unpack\r\n    [INFO] Stopping the server connection.",
    "90058": "[quote=taras sereda;90014]\r\n\r\nCan somebody clarify which pass should bu used? \r\nPATH_TO_JSON = \"path/to/data/from/process_html.py\"\r\n\r\nI've used path to directory where I've downloaded all archives {0,1,2,3,4}\r\nbut script fails when parsing compleats.\r\n\r\nProbably, I'm missing something.\r\nI've got the following error.\r\n\r\n    Traceback (most recent call last):\r\n      File \"/Users/taras-sereda/PycharmProjects/dato/classifier.py\", line 33, in <module>\r\n        sf = sf.unpack('X1',column_name_prefix='')\r\n      File \"/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sframe.py\", line 4678, in unpack\r\n        new_sf = self[unpack_column].unpack(column_name_prefix, column_types, na_value, limit)\r\n      File \"/usr/local/lib/python2.7/site-packages/graphlab/data_structures/sarray.py\", line 2692, in unpack\r\n        raise TypeError(\"Only SArray of dict/list/array type supports unpack\")\r\n    TypeError: Only SArray of dict/list/array type supports unpack\r\n    [INFO] Stopping the server connection.\r\n\r\n[/quote]\r\n\r\nMy guess is that you need to run process_html.py first (takes about 8 hours for me) and you will get the json files at the output.\r\nHowever, I have no ideas what to do after that. In the classifier.py, the code tries to open something as csv file in  PATH_TO_JSON, making little sense. Any hints please?",
    "90089": "Thanks @FTD2016. I've solved an issue. \r\nInitially I was expecting a single run of script classifier.py.\r\n\r\nthe problem was that I missunderstood which path to provide.\r\nPATH_TO_JSON = \"path/to/data/from/process_html.py\"\r\n\r\nIn classifier.py. csv is used, because each parsed chunk, is not a True json file, actually it's a text file where each line is valid json. So this formulation is also a little bit misleading.\r\n\r\nthe structure of each line in file is a parsed version of raw html in  each chunk.\r\nAs described in parsing script.\r\n\r\n    doc = {\r\n            \"id\": urlid, \r\n            \"text\":parse_text(soup),\r\n            \"title\":parse_title(soup ),\r\n            \"links\":parse_links(soup),\r\n            \"images\":parse_images(soup),\r\n           }\r\nHope, that helped.",
    "90167": "Hi, how did you fix your issue? I ran the process_html.py script on a small subset of files which generated a chunk0.json. I then ran the classifier script. It runs fine till line 34\r\n\r\n    sf = gl.SFrame.read_csv(PATH_TO_JSON, header=False)\r\n\r\nand I get \"Parsing completed\".\r\n\r\nHowever it fails on the next line with a type mismatch.\r\n\r\nI would appreciate any help. Thanks!",
    "91603": "I want to ask if the code of process_html is bug free or if it needs debugging",
    "91778": "[quote=cloofa;88905]\r\nAlso, this is a very basic script to get started.  If you are familiar with regexp methods, I recommend using that instead of BeautifulSoup4, as it may be faster.  We went with this to make things more interpretable for users who are not as familiar with parsing HTML's.\r\n[/quote]\r\n\r\nIt's somewhat funny and refreshing to read this as the general opinion about regex being used for html parsing is \"you should know better and not do that\".",
    "92230": "Hi!\r\n\r\nDid anybody get - TypeError: <filter object at 0x00000000163204E0> is not JSON serializable when the process_html.py was run?\r\n\r\nThanks",
    "93041": "Hi,\r\nThanks for this starter code\r\nAnother issue\r\nbow_trn = gl.text_analytics.count_words(train['text_clean'])\r\ncomplains about count_words needs string SArrays ...\r\nAny idea ?\r\nTIA\r\nRgds\r\nBruno",
    "93068": "Hi All,\r\n\r\nAnyone else encounter this error?  I am using Python 3.4.3 that comes with Anaconda.\r\n\r\nTraceback (most recent call last):\r\n  File \"process_html_new.py\", line 158, in <module>\r\n    main(sys.argv)\r\n  File \"process_html_new.py\", line 139, in main\r\n    json.dump(entry, feedsjson)\r\n  File \"/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/__init__.py\", line 178, in dump\r\n    for chunk in iterable:\r\n  File \"/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py\", line 422, in _iterencode\r\n    yield from _iterencode_dict(o, _current_indent_level)\r\n  File \"/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py\", line 396, in _iterencode_dict\r\n    yield from chunks\r\n  File \"/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py\", line 429, in _iterencode\r\n    o = _default(o)\r\n  File \"/Users/ntogar/anaconda/envs/kaggle/lib/python3.4/json/encoder.py\", line 173, in default\r\n    raise TypeError(repr(o) + \" is not JSON serializable\")\r\nTypeError: <filter object at 0x102cde550> is not JSON serializable\r\n\r\nRegards,\r\nNaveen",
    "95533": "Hi Competition Admin,\r\nI have had several submissions before the data leak. Recently, due to some personal commitments, I couldn't submit a new entry.\r\n\r\nNow, I just passed the deadline, which I overlooked. Would it be possible to allow previous submitters like myself to submit again?\r\n\r\nVery much appreciated. Thank you.\r\n\r\nBest,\r\nHarish",
    "96126": "The attached script is an almost automatic workflow for GraphLab Create, except its input relies on the stored feature set from https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model/95814#post95814 , which is inspired by [@David Shinn][1]'s nice work.\r\n\r\n\r\n  [1]: https://www.kaggle.com/davidshinn"
  },
  "source": "meta"
}