{
  "id": 16823,
  "title": "What models might be worth trying?",
  "url": "/competitions/dato-native/discussion/16823",
  "author_name": "",
  "post_date": "2015-10-05T07:15:11.133Z",
  "votes": null,
  "comment_count": 2,
  "views": 558,
  "content": "<p>I only very recently joined this competition, and as I am rather new to text classification, I am wondering what models people have success with. Are the usual suspects such as XGboost worth a try?</p>\n\n<p>I only yesterday finished my first model, with some preprocesssing, tf-idf and Naive Bayes, I reached around 0.785 in local cv. I'm already getting into memory restrictions, so apart from improving my preprocessing and doing more feature engineering, I plan to try various online learners from scikit next.</p>",
  "messages": [
    {
      "id": "95108",
      "postDate": "10/05/2015 07:15:11",
      "content": "<p>I only very recently joined this competition, and as I am rather new to text classification, I am wondering what models people have success with. Are the usual suspects such as XGboost worth a try?</p>\n\n<p>I only yesterday finished my first model, with some preprocesssing, tf-idf and Naive Bayes, I reached around 0.785 in local cv. I'm already getting into memory restrictions, so apart from improving my preprocessing and doing more feature engineering, I plan to try various online learners from scikit next.</p>",
      "rawMarkdown": "I only very recently joined this competition, and as I am rather new to text classification, I am wondering what models people have success with. Are the usual suspects such as XGboost worth a try?\r\n\r\nI only yesterday finished my first model, with some preprocesssing, tf-idf and Naive Bayes, I reached around 0.785 in local cv. I'm already getting into memory restrictions, so apart from improving my preprocessing and doing more feature engineering, I plan to try various online learners from scikit next.",
      "votes": null
    },
    {
      "id": "95175",
      "postDate": "10/05/2015 17:30:06",
      "content": "<p>yup, the usual suspects , linear models and xgboost.  NLP is quite straightforward. The key is in feature engineering and less about the model (where a simple linear model can do almost as well as the more  complex ones).</p>",
      "rawMarkdown": "yup, the usual suspects , linear models and xgboost.  NLP is quite straightforward. The key is in feature engineering and less about the model (where a simple linear model can do almost as well as the more  complex ones).",
      "votes": null
    },
    {
      "id": "95181",
      "postDate": "10/05/2015 18:12:47",
      "content": "<p>Perfect, thanks!</p>\n\n<p>(as a side note, I found a bug in my preprocessing that caused some information to be lost, so that score I gave above might not mean a lot)</p>",
      "rawMarkdown": "Perfect, thanks!\r\n\r\n(as a side note, I found a bug in my preprocessing that caused some information to be lost, so that score I gave above might not mean a lot)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 95175,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "10/05/2015 17:30:06",
      "content": "<p>yup, the usual suspects , linear models and xgboost.  NLP is quite straightforward. The key is in feature engineering and less about the model (where a simple linear model can do almost as well as the more  complex ones).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95181,
      "author_name": "tobycheese",
      "author_url": "",
      "post_date": "10/05/2015 18:12:47",
      "content": "<p>Perfect, thanks!</p>\n\n<p>(as a side note, I found a bug in my preprocessing that caused some information to be lost, so that score I gave above might not mean a lot)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "95108": "I only very recently joined this competition, and as I am rather new to text classification, I am wondering what models people have success with. Are the usual suspects such as XGboost worth a try?\r\n\r\nI only yesterday finished my first model, with some preprocesssing, tf-idf and Naive Bayes, I reached around 0.785 in local cv. I'm already getting into memory restrictions, so apart from improving my preprocessing and doing more feature engineering, I plan to try various online learners from scikit next.",
    "95175": "yup, the usual suspects , linear models and xgboost.  NLP is quite straightforward. The key is in feature engineering and less about the model (where a simple linear model can do almost as well as the more  complex ones).",
    "95181": "Perfect, thanks!\r\n\r\n(as a side note, I found a bug in my preprocessing that caused some information to be lost, so that score I gave above might not mean a lot)"
  },
  "source": "meta"
}