{
  "id": 16025,
  "title": "Problem statement.",
  "url": "/competitions/dato-native/discussion/16025",
  "author_name": "",
  "post_date": "2015-08-19T12:31:18.463Z",
  "votes": 4,
  "comment_count": 6,
  "views": 1061,
  "content": "<p>Hi fellow Kagglers!</p>\n\n<p>I'm a bit confused about the problem statement and just wanted to be sure before proceeding further.</p>\n\n<p>So, we need to differentiate between the webpages that are actual stories and webpages that are just ads.\nE.g. webpage <em>213006_raw_html.txt</em> contains source code for this page: <a href=\"http://www.stumbleupon.com/su/AJfW9b/weightwatchers.tumblr.com/post/98341452736/fall-is-here-grab-your-hunter-green-coats-and/\" title=\"Weight Watchers  Fall is here. Grab your hunter green coats and...\">Weight Watchers  Fall is here. Grab your hunter green coats and...</a>, which is a sponsored page. And, then there are pages like <em>1186164_raw_html.txt</em>, that are actual stories.\nWe need to build a classifier that can differentiate between sponsored and non-sponsored pages.\nWe can extract features like:</p>\n\n<ul>\n<li>Presence of links redirecting to download something.</li>\n<li>Tags containing comments from readers etc.</li>\n</ul>\n\n<p>Am I right ? Am I missing something ?\nAnd, if I can ask what features did you use ? :)</p>\n\n<p>Any help would be appreciated!</p>",
  "messages": [
    {
      "id": "89786",
      "postDate": "08/19/2015 12:31:18",
      "content": "<p>Hi fellow Kagglers!</p>\n\n<p>I'm a bit confused about the problem statement and just wanted to be sure before proceeding further.</p>\n\n<p>So, we need to differentiate between the webpages that are actual stories and webpages that are just ads.\nE.g. webpage <em>213006_raw_html.txt</em> contains source code for this page: <a href=\"http://www.stumbleupon.com/su/AJfW9b/weightwatchers.tumblr.com/post/98341452736/fall-is-here-grab-your-hunter-green-coats-and/\" title=\"Weight Watchers  Fall is here. Grab your hunter green coats and...\">Weight Watchers  Fall is here. Grab your hunter green coats and...</a>, which is a sponsored page. And, then there are pages like <em>1186164_raw_html.txt</em>, that are actual stories.\nWe need to build a classifier that can differentiate between sponsored and non-sponsored pages.\nWe can extract features like:</p>\n\n<ul>\n<li>Presence of links redirecting to download something.</li>\n<li>Tags containing comments from readers etc.</li>\n</ul>\n\n<p>Am I right ? Am I missing something ?\nAnd, if I can ask what features did you use ? :)</p>\n\n<p>Any help would be appreciated!</p>",
      "rawMarkdown": "Hi fellow Kagglers!\r\n\r\nI'm a bit confused about the problem statement and just wanted to be sure before proceeding further.\r\n\r\nSo, we need to differentiate between the webpages that are actual stories and webpages that are just ads.\r\nE.g. webpage *213006_raw_html.txt* contains source code for this page: [Weight Watchers  Fall is here. Grab your hunter green coats and...][1], which is a sponsored page. And, then there are pages like *1186164_raw_html.txt*, that are actual stories.\r\nWe need to build a classifier that can differentiate between sponsored and non-sponsored pages.\r\nWe can extract features like:\r\n\r\n - Presence of links redirecting to download something.\r\n - Tags containing comments from readers etc.\r\n\r\nAm I right ? Am I missing something ?\r\nAnd, if I can ask what features did you use ? :)\r\n\r\nAny help would be appreciated!\r\n\r\n  [1]: http://www.stumbleupon.com/su/AJfW9b/weightwatchers.tumblr.com/post/98341452736/fall-is-here-grab-your-hunter-green-coats-and/ \"Weight Watchers  Fall is here. Grab your hunter green coats and...\"",
      "votes": null
    },
    {
      "id": "89806",
      "postDate": "08/19/2015 15:32:05",
      "content": "<p>You are correct.  There is a lot of room for creativity for feature extraction/engineering in this competition.  I would just mention that NLP techniques are a good candidate for feature generation as well (e.g. tfidf transform).</p>\n\n<p>It should be fun!</p>",
      "rawMarkdown": "You are correct.  There is a lot of room for creativity for feature extraction/engineering in this competition.  I would just mention that NLP techniques are a good candidate for feature generation as well (e.g. tfidf transform).\r\n\r\nIt should be fun!",
      "votes": null
    },
    {
      "id": "90562",
      "postDate": "08/27/2015 20:16:12",
      "content": "<p>OK, looking at that webpage, a human can tell that is a sponsored ad.</p>\n\n<p>But could someone tell me how you would consider this one a sponsored ad?\n1003244_raw_html.txt\nIt is listed as such in the training file.</p>\n\n<p>Same question for this one\n1003711_raw_html.txt\nAside from the headline which is somewhat pushy, and <em>maybe</em> that it all sounds a bit over-certain about a medical condition... but then both of those things are often true in actual news stories on The Huffington Post (and other websites) and those aren't sponsored content...we hope ;-) </p>",
      "rawMarkdown": "OK, looking at that webpage, a human can tell that is a sponsored ad.\r\n\r\nBut could someone tell me how you would consider this one a sponsored ad?\r\n1003244_raw_html.txt\r\nIt is listed as such in the training file.\r\n\r\nSame question for this one\r\n1003711_raw_html.txt\r\nAside from the headline which is somewhat pushy, and *maybe* that it all sounds a bit over-certain about a medical condition... but then both of those things are often true in actual news stories on The Huffington Post (and other websites) and those aren't sponsored content...we hope ;-)",
      "votes": null
    },
    {
      "id": "90652",
      "postDate": "08/28/2015 15:06:47",
      "content": "<p>So is target = 1 containing native advertising ? </p>\n\n<p>I have some doubts. for example : \n1532520_raw_html.txt is this : <a href=\"http://geek-news.mtv.com/2013/06/17/marvel-universe-infographic/\">http://geek-news.mtv.com/2013/06/17/marvel-universe-infographic/</a> which does not look like a native advertising to me. \nNeither does 1629138_raw_html.txt which is the old Xerox home page. </p>",
      "rawMarkdown": "So is target = 1 containing native advertising ? \r\n\r\nI have some doubts. for example : \r\n1532520_raw_html.txt is this : http://geek-news.mtv.com/2013/06/17/marvel-universe-infographic/ which does not look like a native advertising to me. \r\nNeither does 1629138_raw_html.txt which is the old Xerox home page.",
      "votes": null
    },
    {
      "id": "90657",
      "postDate": "08/28/2015 15:50:33",
      "content": "<p>That mtv link looks like a native ad for copypress imo.</p>",
      "rawMarkdown": "That mtv link looks like a native ad for copypress imo.",
      "votes": null
    },
    {
      "id": "90659",
      "postDate": "08/28/2015 15:55:31",
      "content": "<p>I thought that too but if you read the bottom you can see that MTV actually payed copypress to do this app, so it's not really native ad. Still, does not solve the Xerox one though. </p>",
      "rawMarkdown": "I thought that too but if you read the bottom you can see that MTV actually payed copypress to do this app, so it's not really native ad. Still, does not solve the Xerox one though.",
      "votes": null
    },
    {
      "id": "90672",
      "postDate": "08/28/2015 17:03:03",
      "content": "<p>1: is sponsored and 0: is native</p>",
      "rawMarkdown": "1: is sponsored and 0: is native",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 89806,
      "author_name": "",
      "author_url": "",
      "post_date": "08/19/2015 15:32:05",
      "content": "<p>You are correct.  There is a lot of room for creativity for feature extraction/engineering in this competition.  I would just mention that NLP techniques are a good candidate for feature generation as well (e.g. tfidf transform).</p>\n\n<p>It should be fun!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90562,
      "author_name": "jeffwinchell",
      "author_url": "",
      "post_date": "08/27/2015 20:16:12",
      "content": "<p>OK, looking at that webpage, a human can tell that is a sponsored ad.</p>\n\n<p>But could someone tell me how you would consider this one a sponsored ad?\n1003244_raw_html.txt\nIt is listed as such in the training file.</p>\n\n<p>Same question for this one\n1003711_raw_html.txt\nAside from the headline which is somewhat pushy, and <em>maybe</em> that it all sounds a bit over-certain about a medical condition... but then both of those things are often true in actual news stories on The Huffington Post (and other websites) and those aren't sponsored content...we hope ;-) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90652,
      "author_name": "pierregutierrez",
      "author_url": "",
      "post_date": "08/28/2015 15:06:47",
      "content": "<p>So is target = 1 containing native advertising ? </p>\n\n<p>I have some doubts. for example : \n1532520_raw_html.txt is this : <a href=\"http://geek-news.mtv.com/2013/06/17/marvel-universe-infographic/\">http://geek-news.mtv.com/2013/06/17/marvel-universe-infographic/</a> which does not look like a native advertising to me. \nNeither does 1629138_raw_html.txt which is the old Xerox home page. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90657,
      "author_name": "petecz",
      "author_url": "",
      "post_date": "08/28/2015 15:50:33",
      "content": "<p>That mtv link looks like a native ad for copypress imo.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90659,
      "author_name": "pierregutierrez",
      "author_url": "",
      "post_date": "08/28/2015 15:55:31",
      "content": "<p>I thought that too but if you read the bottom you can see that MTV actually payed copypress to do this app, so it's not really native ad. Still, does not solve the Xerox one though. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90672,
      "author_name": "cloofa",
      "author_url": "",
      "post_date": "08/28/2015 17:03:03",
      "content": "<p>1: is sponsored and 0: is native</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "89786": "Hi fellow Kagglers!\r\n\r\nI'm a bit confused about the problem statement and just wanted to be sure before proceeding further.\r\n\r\nSo, we need to differentiate between the webpages that are actual stories and webpages that are just ads.\r\nE.g. webpage *213006_raw_html.txt* contains source code for this page: [Weight Watchers  Fall is here. Grab your hunter green coats and...][1], which is a sponsored page. And, then there are pages like *1186164_raw_html.txt*, that are actual stories.\r\nWe need to build a classifier that can differentiate between sponsored and non-sponsored pages.\r\nWe can extract features like:\r\n\r\n - Presence of links redirecting to download something.\r\n - Tags containing comments from readers etc.\r\n\r\nAm I right ? Am I missing something ?\r\nAnd, if I can ask what features did you use ? :)\r\n\r\nAny help would be appreciated!\r\n\r\n  [1]: http://www.stumbleupon.com/su/AJfW9b/weightwatchers.tumblr.com/post/98341452736/fall-is-here-grab-your-hunter-green-coats-and/ \"Weight Watchers  Fall is here. Grab your hunter green coats and...\"",
    "89806": "You are correct.  There is a lot of room for creativity for feature extraction/engineering in this competition.  I would just mention that NLP techniques are a good candidate for feature generation as well (e.g. tfidf transform).\r\n\r\nIt should be fun!",
    "90562": "OK, looking at that webpage, a human can tell that is a sponsored ad.\r\n\r\nBut could someone tell me how you would consider this one a sponsored ad?\r\n1003244_raw_html.txt\r\nIt is listed as such in the training file.\r\n\r\nSame question for this one\r\n1003711_raw_html.txt\r\nAside from the headline which is somewhat pushy, and *maybe* that it all sounds a bit over-certain about a medical condition... but then both of those things are often true in actual news stories on The Huffington Post (and other websites) and those aren't sponsored content...we hope ;-)",
    "90652": "So is target = 1 containing native advertising ? \r\n\r\nI have some doubts. for example : \r\n1532520_raw_html.txt is this : http://geek-news.mtv.com/2013/06/17/marvel-universe-infographic/ which does not look like a native advertising to me. \r\nNeither does 1629138_raw_html.txt which is the old Xerox home page.",
    "90657": "That mtv link looks like a native ad for copypress imo.",
    "90659": "I thought that too but if you read the bottom you can see that MTV actually payed copypress to do this app, so it's not really native ad. Still, does not solve the Xerox one though.",
    "90672": "1: is sponsored and 0: is native"
  },
  "source": "meta"
}