{
  "id": 16607,
  "title": "Tfidf missing unable to detect native ads",
  "url": "/competitions/dato-native/discussion/16607",
  "author_name": "",
  "post_date": "2015-09-21T19:08:06.577Z",
  "votes": null,
  "comment_count": 5,
  "views": 1637,
  "content": "<p>I am using the NATO dataset to try some NLP techniques for the first time. </p>\n\n<p>I found that tfidf is very bad at predicting native ads when used alone. I have attached the confusion table of tdidf applied on a sample of appr. 12,000 ads.  </p>\n\n<p>My score of 91.2% only seems high because most ads are negative, but I do a very poor job at detecting positives. I wonder if anyone was able to achieve better results using tfidf alone?</p>\n\n<p>Thanks</p>\n\n<p>[confusion table also uploaded] <a href=\"http://www.tiikoni.com/tis/view/?id=3c8e169\">http://www.tiikoni.com/tis/view/?id=3c8e169</a></p>",
  "messages": [
    {
      "id": "93124",
      "postDate": "09/21/2015 19:08:06",
      "content": "<p>I am using the NATO dataset to try some NLP techniques for the first time. </p>\n\n<p>I found that tfidf is very bad at predicting native ads when used alone. I have attached the confusion table of tdidf applied on a sample of appr. 12,000 ads.  </p>\n\n<p>My score of 91.2% only seems high because most ads are negative, but I do a very poor job at detecting positives. I wonder if anyone was able to achieve better results using tfidf alone?</p>\n\n<p>Thanks</p>\n\n<p>[confusion table also uploaded] <a href=\"http://www.tiikoni.com/tis/view/?id=3c8e169\">http://www.tiikoni.com/tis/view/?id=3c8e169</a></p>",
      "rawMarkdown": "I am using the NATO dataset to try some NLP techniques for the first time. \r\n\r\nI found that tfidf is very bad at predicting native ads when used alone. I have attached the confusion table of tdidf applied on a sample of appr. 12,000 ads.  \r\n\r\nMy score of 91.2% only seems high because most ads are negative, but I do a very poor job at detecting positives. I wonder if anyone was able to achieve better results using tfidf alone?\r\n\r\nThanks\r\n\r\n\r\n[confusion table also uploaded] http://www.tiikoni.com/tis/view/?id=3c8e169",
      "votes": null
    },
    {
      "id": "93127",
      "postDate": "09/21/2015 19:58:54",
      "content": "<p>I didn't use tfidf. I am definitely a novice at NLP, but It seemed to me that, since the dataset is overwhelmingly dominated by background, the tfidf measure would only pick out words that have meaning to non-sponsored sites and sponsored sites would be completely underrepresented. </p>",
      "rawMarkdown": "I didn't use tfidf. I am definitely a novice at NLP, but It seemed to me that, since the dataset is overwhelmingly dominated by background, the tfidf measure would only pick out words that have meaning to non-sponsored sites and sponsored sites would be completely underrepresented.",
      "votes": null
    },
    {
      "id": "93153",
      "postDate": "09/22/2015 07:58:44",
      "content": "<p>I used tfidf no only because the DATO starter script uses this method to get a decent score, but also because using &quot;background&quot; related features such as number of words or number of links etc., gives even worse results than tfidf alone. I have attached the confusion table of a logistic regression that uses number of images, of links and of characters. No native ad is identified. </p>\n\n<p>Now I can try to combine the two methods. But I am afraid that combining a method that identifies a few native ads and a method that identifies zero native ads will not be helpful. </p>\n\n<p>That's why I wondered if and how anyone has been able to identify at least 50% of true positives. As per my previous mail, I can only identify about 17% of them (173/1010), which I think is mediocre. </p>\n\n<p>Thanks</p>",
      "rawMarkdown": "I used tfidf no only because the DATO starter script uses this method to get a decent score, but also because using \"background\" related features such as number of words or number of links etc., gives even worse results than tfidf alone. I have attached the confusion table of a logistic regression that uses number of images, of links and of characters. No native ad is identified. \r\n\r\nNow I can try to combine the two methods. But I am afraid that combining a method that identifies a few native ads and a method that identifies zero native ads will not be helpful. \r\n\r\nThat's why I wondered if and how anyone has been able to identify at least 50% of true positives. As per my previous mail, I can only identify about 17% of them (173/1010), which I think is mediocre. \r\n\r\nThanks",
      "votes": null
    },
    {
      "id": "95032",
      "postDate": "10/04/2015 13:46:24",
      "content": "<p>I think tfidf is a helpful transform here. Are you seeing a worse score with tfidf logistic regression over non-tfidf logistic regression? Could be a bug.</p>",
      "rawMarkdown": "I think tfidf is a helpful transform here. Are you seeing a worse score with tfidf logistic regression over non-tfidf logistic regression? Could be a bug.",
      "votes": null
    },
    {
      "id": "97370",
      "postDate": "10/26/2015 15:12:27",
      "content": "<p>I have redone tf idf using cleaner text and the result looks about the same. \nI have included the results of 3 estimation methods that use 3 different types of features: counting number of objects (as per DATO helper script), <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model\">David Shinn</a>'s features and tf idf. </p>\n\n<p>All but tf idf are unable to detect a single native ad. So a dumb algorithm that assigns each document to &quot;non native&quot; would perform just as well as them. TF IDF allocates about 10% of native ads correctly, as mentioned in my previous post. </p>\n\n<p>I am still very disappointed by this performance. Now that the competition is over, I hope Kaggle will release some of the best scripts to know what I should do to improve recall (or if mediocre recall is the norm).</p>",
      "rawMarkdown": "I have redone tf idf using cleaner text and the result looks about the same. \r\nI have included the results of 3 estimation methods that use 3 different types of features: counting number of objects (as per DATO helper script), [David Shinn][1]'s features and tf idf. \r\n\r\nAll but tf idf are unable to detect a single native ad. So a dumb algorithm that assigns each document to \"non native\" would perform just as well as them. TF IDF allocates about 10% of native ads correctly, as mentioned in my previous post. \r\n\r\nI am still very disappointed by this performance. Now that the competition is over, I hope Kaggle will release some of the best scripts to know what I should do to improve recall (or if mediocre recall is the norm).\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model",
      "votes": null
    },
    {
      "id": "97452",
      "postDate": "10/27/2015 17:29:06",
      "content": "<p>Have you considered evaluating your performance with the competition metric, ROC AUC?  You'll need to produce probability predictions (I don't think your code does that).  Alternatively, you can change the probability threshold to determine how you classify sponsored or not sponsored (e.g., &gt; 0.20 sponsored = 1) to produce better confusion matrices.  For imbalanced classes like this competition, the default 0.50 cutoff for classification prediction doesn't work very well.</p>",
      "rawMarkdown": "Have you considered evaluating your performance with the competition metric, ROC AUC?  You'll need to produce probability predictions (I don't think your code does that).  Alternatively, you can change the probability threshold to determine how you classify sponsored or not sponsored (e.g., > 0.20 sponsored = 1) to produce better confusion matrices.  For imbalanced classes like this competition, the default 0.50 cutoff for classification prediction doesn't work very well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 93127,
      "author_name": "gbrello",
      "author_url": "",
      "post_date": "09/21/2015 19:58:54",
      "content": "<p>I didn't use tfidf. I am definitely a novice at NLP, but It seemed to me that, since the dataset is overwhelmingly dominated by background, the tfidf measure would only pick out words that have meaning to non-sponsored sites and sponsored sites would be completely underrepresented. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93153,
      "author_name": "mchlkffl",
      "author_url": "",
      "post_date": "09/22/2015 07:58:44",
      "content": "<p>I used tfidf no only because the DATO starter script uses this method to get a decent score, but also because using &quot;background&quot; related features such as number of words or number of links etc., gives even worse results than tfidf alone. I have attached the confusion table of a logistic regression that uses number of images, of links and of characters. No native ad is identified. </p>\n\n<p>Now I can try to combine the two methods. But I am afraid that combining a method that identifies a few native ads and a method that identifies zero native ads will not be helpful. </p>\n\n<p>That's why I wondered if and how anyone has been able to identify at least 50% of true positives. As per my previous mail, I can only identify about 17% of them (173/1010), which I think is mediocre. </p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95032,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "10/04/2015 13:46:24",
      "content": "<p>I think tfidf is a helpful transform here. Are you seeing a worse score with tfidf logistic regression over non-tfidf logistic regression? Could be a bug.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 97370,
      "author_name": "mchlkffl",
      "author_url": "",
      "post_date": "10/26/2015 15:12:27",
      "content": "<p>I have redone tf idf using cleaner text and the result looks about the same. \nI have included the results of 3 estimation methods that use 3 different types of features: counting number of objects (as per DATO helper script), <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model\">David Shinn</a>'s features and tf idf. </p>\n\n<p>All but tf idf are unable to detect a single native ad. So a dumb algorithm that assigns each document to &quot;non native&quot; would perform just as well as them. TF IDF allocates about 10% of native ads correctly, as mentioned in my previous post. </p>\n\n<p>I am still very disappointed by this performance. Now that the competition is over, I hope Kaggle will release some of the best scripts to know what I should do to improve recall (or if mediocre recall is the norm).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 97452,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "10/27/2015 17:29:06",
      "content": "<p>Have you considered evaluating your performance with the competition metric, ROC AUC?  You'll need to produce probability predictions (I don't think your code does that).  Alternatively, you can change the probability threshold to determine how you classify sponsored or not sponsored (e.g., &gt; 0.20 sponsored = 1) to produce better confusion matrices.  For imbalanced classes like this competition, the default 0.50 cutoff for classification prediction doesn't work very well.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "93124": "I am using the NATO dataset to try some NLP techniques for the first time. \r\n\r\nI found that tfidf is very bad at predicting native ads when used alone. I have attached the confusion table of tdidf applied on a sample of appr. 12,000 ads.  \r\n\r\nMy score of 91.2% only seems high because most ads are negative, but I do a very poor job at detecting positives. I wonder if anyone was able to achieve better results using tfidf alone?\r\n\r\nThanks\r\n\r\n\r\n[confusion table also uploaded] http://www.tiikoni.com/tis/view/?id=3c8e169",
    "93127": "I didn't use tfidf. I am definitely a novice at NLP, but It seemed to me that, since the dataset is overwhelmingly dominated by background, the tfidf measure would only pick out words that have meaning to non-sponsored sites and sponsored sites would be completely underrepresented.",
    "93153": "I used tfidf no only because the DATO starter script uses this method to get a decent score, but also because using \"background\" related features such as number of words or number of links etc., gives even worse results than tfidf alone. I have attached the confusion table of a logistic regression that uses number of images, of links and of characters. No native ad is identified. \r\n\r\nNow I can try to combine the two methods. But I am afraid that combining a method that identifies a few native ads and a method that identifies zero native ads will not be helpful. \r\n\r\nThat's why I wondered if and how anyone has been able to identify at least 50% of true positives. As per my previous mail, I can only identify about 17% of them (173/1010), which I think is mediocre. \r\n\r\nThanks",
    "95032": "I think tfidf is a helpful transform here. Are you seeing a worse score with tfidf logistic regression over non-tfidf logistic regression? Could be a bug.",
    "97370": "I have redone tf idf using cleaner text and the result looks about the same. \r\nI have included the results of 3 estimation methods that use 3 different types of features: counting number of objects (as per DATO helper script), [David Shinn][1]'s features and tf idf. \r\n\r\nAll but tf idf are unable to detect a single native ad. So a dumb algorithm that assigns each document to \"non native\" would perform just as well as them. TF IDF allocates about 10% of native ads correctly, as mentioned in my previous post. \r\n\r\nI am still very disappointed by this performance. Now that the competition is over, I hope Kaggle will release some of the best scripts to know what I should do to improve recall (or if mediocre recall is the norm).\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model",
    "97452": "Have you considered evaluating your performance with the competition metric, ROC AUC?  You'll need to produce probability predictions (I don't think your code does that).  Alternatively, you can change the probability threshold to determine how you classify sponsored or not sponsored (e.g., > 0.20 sponsored = 1) to produce better confusion matrices.  For imbalanced classes like this competition, the default 0.50 cutoff for classification prediction doesn't work very well."
  },
  "source": "meta"
}