{
  "id": 59895,
  "title": "CV & Time-Series Split Target Encoding and some other codes",
  "url": "/competitions/avito-demand-prediction/discussion/59895",
  "author_name": "",
  "post_date": "2018-06-28T04:13:25.484243Z",
  "votes": 23,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, </p>\n\n<p>First, thanks kaggle and avito for organizing such a worderful competition and congrats to everyone! And thanks to teams sharing fantastic solutions!</p>\n\n<p>Although my whole solution is neither strong nor novel and not worth sharing, I want to share a script which generated a group of my strong features - target encoding (also called mean encoding). </p>\n\n<p><a href=\"https://www.kaggle.com/johnfarrell/adp-prepare-data-me-20-2-true\">code here</a></p>\n\n<p>I knew it long ago and found that the normal TE overfitting easily but <a href=\"https://www.coursera.org/learn/competitive-data-science\">this coursera course (week3)</a> got me know that it will only work when you take some generalizing tricks. In this competition I took cross-validated (CV) &amp; time-series split (TS) two kinds of generalization. </p>\n\n<p>Ideas are simple:</p>\n\n<ul>\n<li>Only use data from other CV folds to do target encoding</li>\n<li>Only use past data to do target encoding</li>\n<li>And combine them</li>\n</ul>\n\n<p>The experiment shows the lgb model using CV-only target encoding feature (and other common features) would lead to 0.0055(k=20, f=2) or 0.0063(k=50, f=5) gap between OOF valid score and LB, but with CV&amp;TS target encoding it gives consistent gaps of 0.0037(k={20, 50, 100}; f={2, 5, 10}), which seems much better!</p>\n\n<p>Will be happy if it helps~</p>\n\n<hr>\n\n<p>Note1: The target of regression is tested in this competition but I didn't test the classification target, maybe there're some bugs :p</p>\n\n<p>Note2: The code is arranged from <a href=\"https://zhuanlan.zhihu.com/p/26308272\">this article</a> (in Chinese). Credits to him~</p>",
  "messages": [
    {
      "id": "349393",
      "postDate": "06/28/2018 04:13:25",
      "content": "<p>Hi, </p>\n\n<p>First, thanks kaggle and avito for organizing such a worderful competition and congrats to everyone! And thanks to teams sharing fantastic solutions!</p>\n\n<p>Although my whole solution is neither strong nor novel and not worth sharing, I want to share a script which generated a group of my strong features - target encoding (also called mean encoding). </p>\n\n<p><a href=\"https://www.kaggle.com/johnfarrell/adp-prepare-data-me-20-2-true\">code here</a></p>\n\n<p>I knew it long ago and found that the normal TE overfitting easily but <a href=\"https://www.coursera.org/learn/competitive-data-science\">this coursera course (week3)</a> got me know that it will only work when you take some generalizing tricks. In this competition I took cross-validated (CV) &amp; time-series split (TS) two kinds of generalization. </p>\n\n<p>Ideas are simple:</p>\n\n<ul>\n<li>Only use data from other CV folds to do target encoding</li>\n<li>Only use past data to do target encoding</li>\n<li>And combine them</li>\n</ul>\n\n<p>The experiment shows the lgb model using CV-only target encoding feature (and other common features) would lead to 0.0055(k=20, f=2) or 0.0063(k=50, f=5) gap between OOF valid score and LB, but with CV&amp;TS target encoding it gives consistent gaps of 0.0037(k={20, 50, 100}; f={2, 5, 10}), which seems much better!</p>\n\n<p>Will be happy if it helps~</p>\n\n<hr>\n\n<p>Note1: The target of regression is tested in this competition but I didn't test the classification target, maybe there're some bugs :p</p>\n\n<p>Note2: The code is arranged from <a href=\"https://zhuanlan.zhihu.com/p/26308272\">this article</a> (in Chinese). Credits to him~</p>",
      "rawMarkdown": "Hi, \n\nFirst, thanks kaggle and avito for organizing such a worderful competition and congrats to everyone! And thanks to teams sharing fantastic solutions!\n\nAlthough my whole solution is neither strong nor novel and not worth sharing, I want to share a script which generated a group of my strong features - target encoding (also called mean encoding). \n\n[code here][1]\n\nI knew it long ago and found that the normal TE overfitting easily but [this coursera course (week3)][2] got me know that it will only work when you take some generalizing tricks. In this competition I took cross-validated (CV) &amp; time-series split (TS) two kinds of generalization. \n\nIdeas are simple:\n\n- Only use data from other CV folds to do target encoding\n- Only use past data to do target encoding\n- And combine them\n\nThe experiment shows the lgb model using CV-only target encoding feature (and other common features) would lead to 0.0055(k=20, f=2) or 0.0063(k=50, f=5) gap between OOF valid score and LB, but with CV&amp;TS target encoding it gives consistent gaps of 0.0037(k={20, 50, 100}; f={2, 5, 10}), which seems much better!\n\nWill be happy if it helps~\n\n\n\n-------\nNote1: The target of regression is tested in this competition but I didn't test the classification target, maybe there're some bugs :p\n\nNote2: The code is arranged from [this article][3] (in Chinese). Credits to him~\n\n\n  [1]: https://www.kaggle.com/johnfarrell/adp-prepare-data-me-20-2-true\n  [2]: https://www.coursera.org/learn/competitive-data-science\n  [3]: https://zhuanlan.zhihu.com/p/26308272",
      "votes": null
    },
    {
      "id": "349430",
      "postDate": "06/28/2018 05:08:06",
      "content": "<p>Other script:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/johnfarrell/adp-prepare-text-metadata\">Extract text metadata</a></li>\n<li><a href=\"https://www.kaggle.com/johnfarrell/adp-pytorch-fm-pretrain-dev\">Use pytorch to train FM and extract second last layer output as features available to other models like LGB</a></li>\n<li><a href=\"https://www.kaggle.com/johnfarrell/adp-pytorch-textcnn-context-only-dev\">Use the same pytorch template to train textcnn and extract second last layer feature</a></li>\n<li><a href=\"https://www.kaggle.com/johnfarrell/adp-blend-final\">Simple two-level stacking process</a></li>\n</ul>",
      "rawMarkdown": "Other script:\n\n- [Extract text metadata][1]\n- [Use pytorch to train FM and extract second last layer output as features available to other models like LGB][2]\n- [Use the same pytorch template to train textcnn and extract second last layer feature][3]\n- [Simple two-level stacking process][4]\n\n\n  [1]: https://www.kaggle.com/johnfarrell/adp-prepare-text-metadata\n  [2]: https://www.kaggle.com/johnfarrell/adp-pytorch-fm-pretrain-dev\n  [3]: https://www.kaggle.com/johnfarrell/adp-pytorch-textcnn-context-only-dev\n  [4]: https://www.kaggle.com/johnfarrell/adp-blend-final",
      "votes": null
    },
    {
      "id": "349660",
      "postDate": "06/28/2018 13:04:18",
      "content": "<p>Very helpful. Congrats~~ and happy kaggling!</p>",
      "rawMarkdown": "Very helpful. Congrats~~ and happy kaggling!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 349430,
      "author_name": "johnfarrell",
      "author_url": "",
      "post_date": "06/28/2018 05:08:06",
      "content": "<p>Other script:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/johnfarrell/adp-prepare-text-metadata\">Extract text metadata</a></li>\n<li><a href=\"https://www.kaggle.com/johnfarrell/adp-pytorch-fm-pretrain-dev\">Use pytorch to train FM and extract second last layer output as features available to other models like LGB</a></li>\n<li><a href=\"https://www.kaggle.com/johnfarrell/adp-pytorch-textcnn-context-only-dev\">Use the same pytorch template to train textcnn and extract second last layer feature</a></li>\n<li><a href=\"https://www.kaggle.com/johnfarrell/adp-blend-final\">Simple two-level stacking process</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349660,
      "author_name": "jyuan1986",
      "author_url": "",
      "post_date": "06/28/2018 13:04:18",
      "content": "<p>Very helpful. Congrats~~ and happy kaggling!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "349393": "Hi, \n\nFirst, thanks kaggle and avito for organizing such a worderful competition and congrats to everyone! And thanks to teams sharing fantastic solutions!\n\nAlthough my whole solution is neither strong nor novel and not worth sharing, I want to share a script which generated a group of my strong features - target encoding (also called mean encoding). \n\n[code here][1]\n\nI knew it long ago and found that the normal TE overfitting easily but [this coursera course (week3)][2] got me know that it will only work when you take some generalizing tricks. In this competition I took cross-validated (CV) &amp; time-series split (TS) two kinds of generalization. \n\nIdeas are simple:\n\n- Only use data from other CV folds to do target encoding\n- Only use past data to do target encoding\n- And combine them\n\nThe experiment shows the lgb model using CV-only target encoding feature (and other common features) would lead to 0.0055(k=20, f=2) or 0.0063(k=50, f=5) gap between OOF valid score and LB, but with CV&amp;TS target encoding it gives consistent gaps of 0.0037(k={20, 50, 100}; f={2, 5, 10}), which seems much better!\n\nWill be happy if it helps~\n\n\n\n-------\nNote1: The target of regression is tested in this competition but I didn't test the classification target, maybe there're some bugs :p\n\nNote2: The code is arranged from [this article][3] (in Chinese). Credits to him~\n\n\n  [1]: https://www.kaggle.com/johnfarrell/adp-prepare-data-me-20-2-true\n  [2]: https://www.coursera.org/learn/competitive-data-science\n  [3]: https://zhuanlan.zhihu.com/p/26308272",
    "349430": "Other script:\n\n- [Extract text metadata][1]\n- [Use pytorch to train FM and extract second last layer output as features available to other models like LGB][2]\n- [Use the same pytorch template to train textcnn and extract second last layer feature][3]\n- [Simple two-level stacking process][4]\n\n\n  [1]: https://www.kaggle.com/johnfarrell/adp-prepare-text-metadata\n  [2]: https://www.kaggle.com/johnfarrell/adp-pytorch-fm-pretrain-dev\n  [3]: https://www.kaggle.com/johnfarrell/adp-pytorch-textcnn-context-only-dev\n  [4]: https://www.kaggle.com/johnfarrell/adp-blend-final",
    "349660": "Very helpful. Congrats~~ and happy kaggling!"
  },
  "source": "meta"
}