{
  "id": 78123,
  "title": "TfIdf works?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/78123",
  "author_name": "",
  "post_date": "2019-01-20T02:45:38.492883700Z",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Has anyone tried TfIdf?\nIt seems in some other text classification competitions TfIdf works well.</p>\n\n<p>I can think of the following two ways, but they might be time consuming.</p>\n\n<p>(A) Blend \"CNN Embedding\" model +\"Logistic Regression tfidf\" model.</p>\n\n<p>(B) Multiple input  CNN model, one input is embedding and the other is top N words from tfidf.</p>\n\n<p>Any thoughts?\nthanks!</p>",
  "messages": [
    {
      "id": "458595",
      "postDate": "01/20/2019 02:45:38",
      "content": "<p>Has anyone tried TfIdf?\nIt seems in some other text classification competitions TfIdf works well.</p>\n\n<p>I can think of the following two ways, but they might be time consuming.</p>\n\n<p>(A) Blend \"CNN Embedding\" model +\"Logistic Regression tfidf\" model.</p>\n\n<p>(B) Multiple input  CNN model, one input is embedding and the other is top N words from tfidf.</p>\n\n<p>Any thoughts?\nthanks!</p>",
      "rawMarkdown": "Has anyone tried TfIdf?\nIt seems in some other text classification competitions TfIdf works well.\n\nI can think of the following two ways, but they might be time consuming.\n\n(A) Blend \"CNN Embedding\" model +\"Logistic Regression tfidf\" model.\n\n(B) Multiple input  CNN model, one input is embedding and the other is top N words from tfidf.\n\nAny thoughts?\nthanks!",
      "votes": null
    },
    {
      "id": "458606",
      "postDate": "01/20/2019 03:23:55",
      "content": "<p>This one is great: <a href=\"https://www.kaggle.com/ryanzhang/tfidf-naivebayes-logreg-baseline\">https://www.kaggle.com/ryanzhang/tfidf-naivebayes-logreg-baseline</a></p>",
      "rawMarkdown": "This one is great: https://www.kaggle.com/ryanzhang/tfidf-naivebayes-logreg-baseline",
      "votes": null
    },
    {
      "id": "458899",
      "postDate": "01/20/2019 19:03:20",
      "content": "<p><a href=\"/higepon\">@higepon</a> have a look at a previous competition <a href=\"https://www.kaggle.com/abhishek/approaching-almost-any-nlp-problem-on-kaggle\">https://www.kaggle.com/abhishek/approaching-almost-any-nlp-problem-on-kaggle</a></p>",
      "rawMarkdown": "higepon have a look at a previous competition https://www.kaggle.com/abhishek/approaching-almost-any-nlp-problem-on-kaggle",
      "votes": null
    },
    {
      "id": "458900",
      "postDate": "01/20/2019 19:10:15",
      "content": "<p>Tried TFIDF &amp; FTRL without success...\nDue to kernel limits,  my NGRAM coverage was only [1-7] range :( </p>",
      "rawMarkdown": "Tried TFIDF &amp; FTRL without success...\nDue to kernel limits,  my NGRAM coverage was only [1-7] range :(",
      "votes": null
    },
    {
      "id": "459038",
      "postDate": "01/21/2019 04:15:29",
      "content": "<p>Wow thanks. This is exactly what I was looking for.\nI'll look into it.</p>\n\n<p>By the way I love their this trick.\n<code>\n   @contextmanager\n   def timer(task_name=\"timer\"):\n</code></p>",
      "rawMarkdown": "Wow thanks. This is exactly what I was looking for.\nI'll look into it.\n\nBy the way I love their this trick.\n```\n   @contextmanager\n   def timer(task_name=\"timer\"):\n```",
      "votes": null
    },
    {
      "id": "459044",
      "postDate": "01/21/2019 04:21:03",
      "content": "<p>Thank you so much.\nThe kernel is so useful. I bookmarked it in my kaggle folder.\nI like how the author made the kernel step by step.</p>\n\n<p>thanks!</p>",
      "rawMarkdown": "Thank you so much.\nThe kernel is so useful. I bookmarked it in my kaggle folder.\nI like how the author made the kernel step by step.\n\nthanks!",
      "votes": null
    },
    {
      "id": "459046",
      "postDate": "01/21/2019 04:22:06",
      "content": "<p>Thanks.\nThat's good to know.\nIn theory, it should work because it extracts important keywords.\nThe feature can be learned by LSTM or CNN, but it should take some time to get it.</p>",
      "rawMarkdown": "Thanks.\nThat's good to know.\nIn theory, it should work because it extracts important keywords.\nThe feature can be learned by LSTM or CNN, but it should take some time to get it.",
      "votes": null
    },
    {
      "id": "461054",
      "postDate": "01/25/2019 05:25:18",
      "content": "<p>Tried GRU plus TFIDF as in (B) -&gt; no improvement. my understanding is no new information for the model there. It can gather tfidf on its own.</p>",
      "rawMarkdown": "Tried GRU plus TFIDF as in (B) -&gt; no improvement. my understanding is no new information for the model there. It can gather tfidf on its own.",
      "votes": null
    },
    {
      "id": "461061",
      "postDate": "01/25/2019 06:10:38",
      "content": "<p>Thanks for trying it. It kinda makes sense that it didn't work. \nAs you pointed out GRU model can find it easily.</p>\n\n<p>I tried (A), but no luck so far.</p>",
      "rawMarkdown": "Thanks for trying it. It kinda makes sense that it didn't work. \nAs you pointed out GRU model can find it easily.\n\nI tried (A), but no luck so far.",
      "votes": null
    },
    {
      "id": "461069",
      "postDate": "01/25/2019 06:42:29",
      "content": "<p>I think TFIDF might not work in this context since the documents (which are questions in this case) are quite short. You'd ultimately end up with binary vectors. </p>",
      "rawMarkdown": "I think TFIDF might not work in this context since the documents (which are questions in this case) are quite short. You'd ultimately end up with binary vectors.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 458606,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "01/20/2019 03:23:55",
      "content": "<p>This one is great: <a href=\"https://www.kaggle.com/ryanzhang/tfidf-naivebayes-logreg-baseline\">https://www.kaggle.com/ryanzhang/tfidf-naivebayes-logreg-baseline</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 459038,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/21/2019 04:15:29",
          "content": "<p>Wow thanks. This is exactly what I was looking for.\nI'll look into it.</p>\n\n<p>By the way I love their this trick.\n<code>\n   @contextmanager\n   def timer(task_name=\"timer\"):\n</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 458899,
      "author_name": "cyberia",
      "author_url": "",
      "post_date": "01/20/2019 19:03:20",
      "content": "<p><a href=\"/higepon\">@higepon</a> have a look at a previous competition <a href=\"https://www.kaggle.com/abhishek/approaching-almost-any-nlp-problem-on-kaggle\">https://www.kaggle.com/abhishek/approaching-almost-any-nlp-problem-on-kaggle</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 459044,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/21/2019 04:21:03",
          "content": "<p>Thank you so much.\nThe kernel is so useful. I bookmarked it in my kaggle folder.\nI like how the author made the kernel step by step.</p>\n\n<p>thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 458900,
      "author_name": "a45632",
      "author_url": "",
      "post_date": "01/20/2019 19:10:15",
      "content": "<p>Tried TFIDF &amp; FTRL without success...\nDue to kernel limits,  my NGRAM coverage was only [1-7] range :( </p>",
      "votes": null,
      "replies": [
        {
          "id": 459046,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/21/2019 04:22:06",
          "content": "<p>Thanks.\nThat's good to know.\nIn theory, it should work because it extracts important keywords.\nThe feature can be learned by LSTM or CNN, but it should take some time to get it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 461054,
      "author_name": "isikkuntay",
      "author_url": "",
      "post_date": "01/25/2019 05:25:18",
      "content": "<p>Tried GRU plus TFIDF as in (B) -&gt; no improvement. my understanding is no new information for the model there. It can gather tfidf on its own.</p>",
      "votes": null,
      "replies": [
        {
          "id": 461061,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/25/2019 06:10:38",
          "content": "<p>Thanks for trying it. It kinda makes sense that it didn't work. \nAs you pointed out GRU model can find it easily.</p>\n\n<p>I tried (A), but no luck so far.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 461069,
      "author_name": "gamerx",
      "author_url": "",
      "post_date": "01/25/2019 06:42:29",
      "content": "<p>I think TFIDF might not work in this context since the documents (which are questions in this case) are quite short. You'd ultimately end up with binary vectors. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "458595": "Has anyone tried TfIdf?\nIt seems in some other text classification competitions TfIdf works well.\n\nI can think of the following two ways, but they might be time consuming.\n\n(A) Blend \"CNN Embedding\" model +\"Logistic Regression tfidf\" model.\n\n(B) Multiple input  CNN model, one input is embedding and the other is top N words from tfidf.\n\nAny thoughts?\nthanks!",
    "458606": "This one is great: https://www.kaggle.com/ryanzhang/tfidf-naivebayes-logreg-baseline",
    "458899": "higepon have a look at a previous competition https://www.kaggle.com/abhishek/approaching-almost-any-nlp-problem-on-kaggle",
    "458900": "Tried TFIDF &amp; FTRL without success...\nDue to kernel limits,  my NGRAM coverage was only [1-7] range :(",
    "459038": "Wow thanks. This is exactly what I was looking for.\nI'll look into it.\n\nBy the way I love their this trick.\n```\n   @contextmanager\n   def timer(task_name=\"timer\"):\n```",
    "459044": "Thank you so much.\nThe kernel is so useful. I bookmarked it in my kaggle folder.\nI like how the author made the kernel step by step.\n\nthanks!",
    "459046": "Thanks.\nThat's good to know.\nIn theory, it should work because it extracts important keywords.\nThe feature can be learned by LSTM or CNN, but it should take some time to get it.",
    "461054": "Tried GRU plus TFIDF as in (B) -&gt; no improvement. my understanding is no new information for the model there. It can gather tfidf on its own.",
    "461061": "Thanks for trying it. It kinda makes sense that it didn't work. \nAs you pointed out GRU model can find it easily.\n\nI tried (A), but no luck so far.",
    "461069": "I think TFIDF might not work in this context since the documents (which are questions in this case) are quite short. You'd ultimately end up with binary vectors."
  },
  "source": "meta"
}