{
  "id": 16863,
  "title": "NLP vs. Structural Analysis?",
  "url": "/competitions/dato-native/discussion/16863",
  "author_name": "",
  "post_date": "2015-10-06T16:20:38.797Z",
  "votes": 9,
  "comment_count": 4,
  "views": 1290,
  "content": "<p>I hadn't intended to join the Dato competition, but after reading a few posts about how the competition was languishing (too few teams, data leakage), and taking the &quot;quick &amp; easy kaggle points&quot; bait (LOL), I figured I'd throw something together just to see how it fared.</p>\n\n<p>Since data prep for NLP (creating bags-of-words/ngrams, applying TF/IDF, word2vec, etc.) tends to be time-consuming I instead parsed all the files using a public domain HTML5 parser and created lots of simple features based on html/css/javascript analysis.  I also pulled out the domain names of any URLs I found.  From this pile of potential features I selected the 1000 most prevalent, produced an unoptimized model (with no CV or parameter tuning or anything, just seat-of-the-pants) and sent in a submission.</p>\n\n<p>Which put me in 36th with a score of 0.96664.  Since I had completely ignored the &quot;word&quot; content I was quite surprised.  And I imagine it might be possible to squeeze out another 0.01 through optimization and such (which would put me ~20th).</p>\n\n<p>So my question: Are other competitors using a similar approach, or are the prevailing models NLP-based?</p>",
  "messages": [
    {
      "id": "95273",
      "postDate": "10/06/2015 16:20:38",
      "content": "<p>I hadn't intended to join the Dato competition, but after reading a few posts about how the competition was languishing (too few teams, data leakage), and taking the &quot;quick &amp; easy kaggle points&quot; bait (LOL), I figured I'd throw something together just to see how it fared.</p>\n\n<p>Since data prep for NLP (creating bags-of-words/ngrams, applying TF/IDF, word2vec, etc.) tends to be time-consuming I instead parsed all the files using a public domain HTML5 parser and created lots of simple features based on html/css/javascript analysis.  I also pulled out the domain names of any URLs I found.  From this pile of potential features I selected the 1000 most prevalent, produced an unoptimized model (with no CV or parameter tuning or anything, just seat-of-the-pants) and sent in a submission.</p>\n\n<p>Which put me in 36th with a score of 0.96664.  Since I had completely ignored the &quot;word&quot; content I was quite surprised.  And I imagine it might be possible to squeeze out another 0.01 through optimization and such (which would put me ~20th).</p>\n\n<p>So my question: Are other competitors using a similar approach, or are the prevailing models NLP-based?</p>",
      "rawMarkdown": "I hadn't intended to join the Dato competition, but after reading a few posts about how the competition was languishing (too few teams, data leakage), and taking the \"quick & easy kaggle points\" bait (LOL), I figured I'd throw something together just to see how it fared.\r\n\r\nSince data prep for NLP (creating bags-of-words/ngrams, applying TF/IDF, word2vec, etc.) tends to be time-consuming I instead parsed all the files using a public domain HTML5 parser and created lots of simple features based on html/css/javascript analysis.  I also pulled out the domain names of any URLs I found.  From this pile of potential features I selected the 1000 most prevalent, produced an unoptimized model (with no CV or parameter tuning or anything, just seat-of-the-pants) and sent in a submission.\r\n\r\nWhich put me in 36th with a score of 0.96664.  Since I had completely ignored the \"word\" content I was quite surprised.  And I imagine it might be possible to squeeze out another 0.01 through optimization and such (which would put me ~20th).\r\n\r\nSo my question: Are other competitors using a similar approach, or are the prevailing models NLP-based?",
      "votes": null
    },
    {
      "id": "95282",
      "postDate": "10/06/2015 17:08:07",
      "content": "<p>I am using a primarily NLP-based approach, but apparently that will be changing starting this afternoon... :p</p>",
      "rawMarkdown": "I am using a primarily NLP-based approach, but apparently that will be changing starting this afternoon... :p",
      "votes": null
    },
    {
      "id": "95334",
      "postDate": "10/07/2015 05:19:32",
      "content": "<p>Interesting! I've so far concentrated on text extraction and I think I got it pretty well by now, having identified and filtered a lot of noise (using beautifulsoup and a bit of spacy, and a lot of manual inspection).</p>\n\n<p>But apart from that, I got not much else, and I fear of running out of time. A quick try in tf-idf and naive bayes on my data yielded me approx 0.8 on local cv but only 0.53 on the leaderboard, but I couldn't find out the reason for that discrepancy yet.</p>\n\n<p>The best thing of course would now be to join a team, but without a decent leaderboard score I doubt anyone would want to join with me. For me, this is one of the is the most interesting of all kaggle competitions I have joined so far, too bad I didn't realize this a bit earlier, as I have only really started working on this after the relaunch.</p>",
      "rawMarkdown": "Interesting! I've so far concentrated on text extraction and I think I got it pretty well by now, having identified and filtered a lot of noise (using beautifulsoup and a bit of spacy, and a lot of manual inspection).\r\n\r\nBut apart from that, I got not much else, and I fear of running out of time. A quick try in tf-idf and naive bayes on my data yielded me approx 0.8 on local cv but only 0.53 on the leaderboard, but I couldn't find out the reason for that discrepancy yet.\r\n\r\nThe best thing of course would now be to join a team, but without a decent leaderboard score I doubt anyone would want to join with me. For me, this is one of the is the most interesting of all kaggle competitions I have joined so far, too bad I didn't realize this a bit earlier, as I have only really started working on this after the relaunch.",
      "votes": null
    },
    {
      "id": "95370",
      "postDate": "10/07/2015 13:04:19",
      "content": "<p>I followed the same approach, but with a less careful selection of structural elements, and get a similar score (~0.95).</p>",
      "rawMarkdown": "I followed the same approach, but with a less careful selection of structural elements, and get a similar score (~0.95).",
      "votes": null
    },
    {
      "id": "95505",
      "postDate": "10/08/2015 14:34:44",
      "content": "<p>Same here-- grabbed some simple html features + domain names and got ~.95.</p>\n\n<p>The problem with the NLP-based approaches it there's so much html cruft that probably doesn't matter, but takes up a ton of time and space in your text vectorization.</p>",
      "rawMarkdown": "Same here-- grabbed some simple html features + domain names and got ~.95.\r\n\r\nThe problem with the NLP-based approaches it there's so much html cruft that probably doesn't matter, but takes up a ton of time and space in your text vectorization.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 95282,
      "author_name": "joshstone",
      "author_url": "",
      "post_date": "10/06/2015 17:08:07",
      "content": "<p>I am using a primarily NLP-based approach, but apparently that will be changing starting this afternoon... :p</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95334,
      "author_name": "tobycheese",
      "author_url": "",
      "post_date": "10/07/2015 05:19:32",
      "content": "<p>Interesting! I've so far concentrated on text extraction and I think I got it pretty well by now, having identified and filtered a lot of noise (using beautifulsoup and a bit of spacy, and a lot of manual inspection).</p>\n\n<p>But apart from that, I got not much else, and I fear of running out of time. A quick try in tf-idf and naive bayes on my data yielded me approx 0.8 on local cv but only 0.53 on the leaderboard, but I couldn't find out the reason for that discrepancy yet.</p>\n\n<p>The best thing of course would now be to join a team, but without a decent leaderboard score I doubt anyone would want to join with me. For me, this is one of the is the most interesting of all kaggle competitions I have joined so far, too bad I didn't realize this a bit earlier, as I have only really started working on this after the relaunch.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95370,
      "author_name": "mortezaa",
      "author_url": "",
      "post_date": "10/07/2015 13:04:19",
      "content": "<p>I followed the same approach, but with a less careful selection of structural elements, and get a similar score (~0.95).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95505,
      "author_name": "zachmayer",
      "author_url": "",
      "post_date": "10/08/2015 14:34:44",
      "content": "<p>Same here-- grabbed some simple html features + domain names and got ~.95.</p>\n\n<p>The problem with the NLP-based approaches it there's so much html cruft that probably doesn't matter, but takes up a ton of time and space in your text vectorization.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "95273": "I hadn't intended to join the Dato competition, but after reading a few posts about how the competition was languishing (too few teams, data leakage), and taking the \"quick & easy kaggle points\" bait (LOL), I figured I'd throw something together just to see how it fared.\r\n\r\nSince data prep for NLP (creating bags-of-words/ngrams, applying TF/IDF, word2vec, etc.) tends to be time-consuming I instead parsed all the files using a public domain HTML5 parser and created lots of simple features based on html/css/javascript analysis.  I also pulled out the domain names of any URLs I found.  From this pile of potential features I selected the 1000 most prevalent, produced an unoptimized model (with no CV or parameter tuning or anything, just seat-of-the-pants) and sent in a submission.\r\n\r\nWhich put me in 36th with a score of 0.96664.  Since I had completely ignored the \"word\" content I was quite surprised.  And I imagine it might be possible to squeeze out another 0.01 through optimization and such (which would put me ~20th).\r\n\r\nSo my question: Are other competitors using a similar approach, or are the prevailing models NLP-based?",
    "95282": "I am using a primarily NLP-based approach, but apparently that will be changing starting this afternoon... :p",
    "95334": "Interesting! I've so far concentrated on text extraction and I think I got it pretty well by now, having identified and filtered a lot of noise (using beautifulsoup and a bit of spacy, and a lot of manual inspection).\r\n\r\nBut apart from that, I got not much else, and I fear of running out of time. A quick try in tf-idf and naive bayes on my data yielded me approx 0.8 on local cv but only 0.53 on the leaderboard, but I couldn't find out the reason for that discrepancy yet.\r\n\r\nThe best thing of course would now be to join a team, but without a decent leaderboard score I doubt anyone would want to join with me. For me, this is one of the is the most interesting of all kaggle competitions I have joined so far, too bad I didn't realize this a bit earlier, as I have only really started working on this after the relaunch.",
    "95370": "I followed the same approach, but with a less careful selection of structural elements, and get a similar score (~0.95).",
    "95505": "Same here-- grabbed some simple html features + domain names and got ~.95.\r\n\r\nThe problem with the NLP-based approaches it there's so much html cruft that probably doesn't matter, but takes up a ton of time and space in your text vectorization."
  },
  "source": "meta"
}