{
  "id": 53249,
  "title": "Best single model?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53249",
  "author_name": "",
  "post_date": "2018-03-28T14:56:55.716120Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>I've been using all training data with feature engineering by counting on day/hour/os etc. here's what I got so far with catboost:</p>\n\n<ul>\n<li>sample rows: 90.60% </li>\n<li>1M rows: 94.14% </li>\n<li>all rows: 94.39%</li>\n</ul>\n\n<p>The gain doesn't look substantial with using a whole lot more data...</p>\n\n<p>I wonder what the current best single model there is and any pointers to breaking my current threshold would be much appreciated.</p>\n\n<p>Happy TalkingData :)</p>\n\n<p>-andrew</p>",
  "messages": [
    {
      "id": "305170",
      "postDate": "03/28/2018 14:56:55",
      "content": "<p>Hi all,</p>\n\n<p>I've been using all training data with feature engineering by counting on day/hour/os etc. here's what I got so far with catboost:</p>\n\n<ul>\n<li>sample rows: 90.60% </li>\n<li>1M rows: 94.14% </li>\n<li>all rows: 94.39%</li>\n</ul>\n\n<p>The gain doesn't look substantial with using a whole lot more data...</p>\n\n<p>I wonder what the current best single model there is and any pointers to breaking my current threshold would be much appreciated.</p>\n\n<p>Happy TalkingData :)</p>\n\n<p>-andrew</p>",
      "rawMarkdown": "Hi all,\n\nI've been using all training data with feature engineering by counting on day/hour/os etc. here's what I got so far with catboost:\n\n - sample rows: 90.60% \n - 1M rows: 94.14% \n - all rows: 94.39%\n\nThe gain doesn't look substantial with using a whole lot more data...\n\nI wonder what the current best single model there is and any pointers to breaking my current threshold would be much appreciated.\n\nHappy TalkingData :)\n\n-andrew",
      "votes": null
    },
    {
      "id": "307252",
      "postDate": "04/01/2018 07:24:06",
      "content": "<p>I would concentrate on good and fast feature engineering on small subsample, iterate and fail fast, learn and then go for the full data.</p>",
      "rawMarkdown": "I would concentrate on good and fast feature engineering on small subsample, iterate and fail fast, learn and then go for the full data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 307252,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "04/01/2018 07:24:06",
      "content": "<p>I would concentrate on good and fast feature engineering on small subsample, iterate and fail fast, learn and then go for the full data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "305170": "Hi all,\n\nI've been using all training data with feature engineering by counting on day/hour/os etc. here's what I got so far with catboost:\n\n - sample rows: 90.60% \n - 1M rows: 94.14% \n - all rows: 94.39%\n\nThe gain doesn't look substantial with using a whole lot more data...\n\nI wonder what the current best single model there is and any pointers to breaking my current threshold would be much appreciated.\n\nHappy TalkingData :)\n\n-andrew",
    "307252": "I would concentrate on good and fast feature engineering on small subsample, iterate and fail fast, learn and then go for the full data."
  },
  "source": "meta"
}