{
  "id": 382886,
  "title": "17th place solution (T88 part)",
  "url": "/competitions/otto-recommender-system/discussion/382886",
  "author_name": "T88",
  "post_date": "2023-02-01T11:33:23.286000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p>kaggle &amp; otto, thanks for organizing a great competition.<br>\nHere I share my approach to my part of the team's solution.  <br>\n(public : 0.594 / private : 0.594) </p>\n<h1>Approach Summary</h1>\n<ul>\n<li>Candidate &amp; Re-rank 2stage model</li>\n<li>train : week3 / valid : week4  </li>\n<li>1 : 1 negative sampling (leave only sessions with labels)  </li>\n<li>LGB classification model (by type) </li>\n<li>Finally, ensemble the results of multiple experiments  </li>\n</ul>\n<h1>Candidate</h1>\n<ul>\n<li>re-visit  </li>\n<li>co-visit (normal)  </li>\n<li>co-visit (type_weighted)  </li>\n<li>co-visit (aid_x - aid_y ts diff weighted)  </li>\n<li>co-visit (clicks2carts, clicks2orders)  </li>\n<li>w2v aid similarity</li>\n<li>w2v session cluster most freq  <br>\n=&gt; Calculate the average of the embedding vectors of the aid's present in the history for each session, and perform clustering with k-means. The most frequent value for each cluster is used as a candidate.</li>\n</ul>\n<p>(avg. 180candidate / session )</p>\n<h1>Re-rank features</h1>\n<ul>\n<li>Various candidate ranks  <br>\n=&gt; Using all ranks lower than those extracted as candidates also increased scores significantly.  </li>\n<li>relative ts in session  </li>\n<li>aid count  </li>\n<li>session count  <br>\netc.   </li>\n</ul>\n<p>(use 91 features)  </p>",
  "messages": [
    {
      "id": 2125039,
      "postDate": "2023-02-01T11:33:23.287Z",
      "content": "<p>kaggle &amp; otto, thanks for organizing a great competition.<br>\nHere I share my approach to my part of the team's solution.  <br>\n(public : 0.594 / private : 0.594) </p>\n<h1>Approach Summary</h1>\n<ul>\n<li>Candidate &amp; Re-rank 2stage model</li>\n<li>train : week3 / valid : week4  </li>\n<li>1 : 1 negative sampling (leave only sessions with labels)  </li>\n<li>LGB classification model (by type) </li>\n<li>Finally, ensemble the results of multiple experiments  </li>\n</ul>\n<h1>Candidate</h1>\n<ul>\n<li>re-visit  </li>\n<li>co-visit (normal)  </li>\n<li>co-visit (type_weighted)  </li>\n<li>co-visit (aid_x - aid_y ts diff weighted)  </li>\n<li>co-visit (clicks2carts, clicks2orders)  </li>\n<li>w2v aid similarity</li>\n<li>w2v session cluster most freq  <br>\n=&gt; Calculate the average of the embedding vectors of the aid's present in the history for each session, and perform clustering with k-means. The most frequent value for each cluster is used as a candidate.</li>\n</ul>\n<p>(avg. 180candidate / session )</p>\n<h1>Re-rank features</h1>\n<ul>\n<li>Various candidate ranks  <br>\n=&gt; Using all ranks lower than those extracted as candidates also increased scores significantly.  </li>\n<li>relative ts in session  </li>\n<li>aid count  </li>\n<li>session count  <br>\netc.   </li>\n</ul>\n<p>(use 91 features)  </p>",
      "rawMarkdown": "kaggle & otto, thanks for organizing a great competition.\nHere I share my approach to my part of the team's solution.  \n(public : 0.594 / private : 0.594) \n\n# Approach Summary\n* Candidate & Re-rank 2stage model\n* train : week3 / valid : week4  \n* 1 : 1 negative sampling (leave only sessions with labels)  \n* LGB classification model (by type) \n* Finally, ensemble the results of multiple experiments  \n\n# Candidate  \n* re-visit  \n* co-visit (normal)  \n* co-visit (type_weighted)  \n* co-visit (aid_x - aid_y ts diff weighted)  \n* co-visit (clicks2carts, clicks2orders)  \n* w2v aid similarity\n* w2v session cluster most freq  \n=> Calculate the average of the embedding vectors of the aid's present in the history for each session, and perform clustering with k-means. The most frequent value for each cluster is used as a candidate.\n\n(avg. 180candidate / session )\n\n\n# Re-rank features  \n* Various candidate ranks  \n=> Using all ranks lower than those extracted as candidates also increased scores significantly.  \n* relative ts in session  \n* aid count  \n* session count  \n etc.   \n\n(use 91 features)  ",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2125039": "kaggle & otto, thanks for organizing a great competition.\nHere I share my approach to my part of the team's solution.  \n(public : 0.594 / private : 0.594) \n\n# Approach Summary\n* Candidate & Re-rank 2stage model\n* train : week3 / valid : week4  \n* 1 : 1 negative sampling (leave only sessions with labels)  \n* LGB classification model (by type) \n* Finally, ensemble the results of multiple experiments  \n\n# Candidate  \n* re-visit  \n* co-visit (normal)  \n* co-visit (type_weighted)  \n* co-visit (aid_x - aid_y ts diff weighted)  \n* co-visit (clicks2carts, clicks2orders)  \n* w2v aid similarity\n* w2v session cluster most freq  \n=> Calculate the average of the embedding vectors of the aid's present in the history for each session, and perform clustering with k-means. The most frequent value for each cluster is used as a candidate.\n\n(avg. 180candidate / session )\n\n\n# Re-rank features  \n* Various candidate ranks  \n=> Using all ranks lower than those extracted as candidates also increased scores significantly.  \n* relative ts in session  \n* aid count  \n* session count  \n etc.   \n\n(use 91 features)  "
  }
}