{
  "id": 383792,
  "title": "9th place solution🥇 [single model LB 0.602]",
  "url": "/competitions/otto-recommender-system/writeups/apolo-9th-place-solution-single-model-lb-0-602",
  "author_name": "",
  "post_date": "2023-02-13T06:41:20.807Z",
  "votes": 43,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, thank you to the competition organizers for a great competition.<br>\nI am very happy to win my first gold medal🥇</p>\n<p>Here is my solution!</p>\n<h1>Overview</h1>\n<p><strong>best single model</strong></p>\n<table>\n<thead>\n<tr>\n<th>orders CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.67087</td>\n<td>0.602</td>\n</tr>\n</tbody>\n</table>\n<p>(I have not calculated the CV for all orders/carts/clicks combined.)</p>\n<p>Probably the same as many other competitors, the candidates generation &amp; rerank method was used.</p>\n<h1>Candidates generation</h1>\n<p>On average, <strong>180 candidates</strong> were selected per session. (<strong>orders recall: 0.725</strong>)</p>\n<ul>\n<li>14 patterns of item-based CF</li>\n<li>2 patterns of user-based CF</li>\n<li>re-visit</li>\n</ul>\n<p>The candidates generated by the 14 patterns of item-based CF and 2 patterns of user-based CF were assembled to create the final candidates.</p>\n<p>Since <strong>sessions with a larger number of aids in the test data are more likely to have a larger number of events in the forecasting period</strong>, more candidates were created for sessions with a larger number of aids in the test data.</p>\n<pre><code>df = test_df[test_df[\"session\"].isin(target_session)].groupby(\"session\")[\"aid\"].count().reset_index()\ndf.columns = [\"session\", \"count\"]\ndf[\"count\"] = (df[\"count\"]**0.5*10).astype(\"int32\")\ndf[\"aid\"] = df.progress_apply(lambda row:list(dict(rec[row[\"session\"]].most_common(row[\"count\"])).keys()),axis=1)\ndf = df.explode([\"aid\"])\ndf = df[df[\"aid\"].notnull()].reset_index(drop=True)\n</code></pre>\n<h1>Feature Engineering</h1>\n<p>A total of <strong>226 features</strong> were created.</p>\n<ul>\n<li><strong>item-based CF features</strong></li>\n</ul>\n<p>First, 59 patterns of item-based CF were created.</p>\n<p>ex) pattern1</p>\n<pre><code>def get_aid_similarity1(df, topk=200):\n    session_info = df.drop_duplicates([\"session\", \"aid\"], keep=\"last\")\\\n                     .groupby(\"session\", as_index=False)[[\"aid\", \"type\", \"ts\"]].agg(list) \n\n    aid_similarity = {}\n    for session, aids, tps, tss in tqdm(zip(session_info[\"session\"],\n                                            session_info[\"aid\"],\n                                            session_info[\"type\"],\n                                            session_info[\"ts\"]),\n                                        total=len(session_info)):\n        for aid1, tp1, ts1 in zip(aids, tps, tss):\n            session_length = math.sqrt(len(aids))\n            aid_similarity.setdefault(aid1, Counter())\n            for aid2, tp2, ts2 in zip(aids, tps, tss):\n                if (aid1 == aid2):\n                    continue\n                aid_similarity[aid1][aid2] += (1/session_length)\n\n    # Exclude all but the topK to save time and memory\n    for aid1, aid2_dict in tqdm(aid_similarity.items()):\n        relations = dict(aid2_dict.most_common(topk))\n        # normalize\n        if len(relations) == 0:\n            continue\n        max_num = relations[max(relations, key=relations.get)]\n        if max_num == 0:\n            continue\n        aid_similarity[aid1] = {k: v / max_num for k, v in relations.items()}\n\n    del session_info; gc_clear()\n    return aid_similarity\n</code></pre>\n<p>ex) pattern2</p>\n<pre><code>def make_real_session(df, hours=2):\n    df[\"lag\"] = df[\"ts\"] - df.groupby(\"session\")[\"ts\"].shift(1)\n    df[\"real_session\"] = (df[\"lag\"] &gt; 1000*60*60*hours).astype('int8').fillna(0)\n    df[\"real_session\"] = df.groupby(\"session\")[\"real_session\"].cumsum()\n    del df[\"lag\"]; gc_clear()\n    return df\n\ndef get_aid_similarity2(df, topk=200):\n\n    df = make_real_session(df, hours=4)\n    session_info = df.groupby([\"session\", \"real_session\"], as_index=False)[[\"aid\", \"type\", \"ts\"]].agg(list) \n\n    aid_similarity = {}\n    aid_cnt = defaultdict(int)\n    for session, real_session, aids, tps, tss in tqdm(zip(session_info[\"session\"],\n                                                          session_info[\"real_session\"],\n                                                          session_info[\"aid\"],\n                                                          session_info[\"type\"],\n                                                          session_info[\"ts\"]),\n                                                      total=len(session_info)):\n        for aid1, tp1, ts1 in zip(aids, tps, tss):\n            aid_similarity.setdefault(aid1, Counter())\n            for aid2, tp2, ts2 in zip(aids, tps, tss):\n                if (abs(ts1-ts2)&gt;24*60*60*1000) or (aid1 == aid2):\n                    continue\n                aid_cnt[aid1] += 1\n                if min(tp1, tp2)==0:\n                    aid_similarity[aid1][aid2] += 1\n                elif min(tp1, tp2)==1:\n                    aid_similarity[aid1][aid2] += 3\n                elif min(tp1, tp2)==2:\n                    aid_similarity[aid1][aid2] += 6\n\n    # Exclude all but the topK to save time and memory\n    for aid1, aid2_dict in tqdm(aid_similarity.items()):  \n        for aid2, score in aid2_dict.items():  \n            aid_similarity[aid1][aid2] = score / math.sqrt(aid_cnt[aid1]*aid_cnt[aid2])\n        aid_similarity[aid1] = dict(aid2_dict.most_common(topk))\n\n    del session_info; gc_clear()\n    return aid_similarity\n</code></pre>\n<p>(Further increasing the <code>topk</code> parameter did not improve the score.)</p>\n<p>Then, feature creation was performed using the various <code>aid_similarity</code> created above.<br>\nUnlike <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">cdeotte's notebook</a>, features were created by summing the value of sim_item_dict.</p>\n<pre><code>def recommend_aid(df,\n                  _test_df,\n                  target_session,\n                  aid_similarity,\n                  feature_name,\n                  type_weighted=False, type_weights={0:1, 1:3, 2:6},\n                  time_weighted=False,\n                  only_last_aid=False,\n                  same_real_session=False, real_session_hours=2,\n                  drop_duplicates=False\n                  ):\n\n    # ==============================================================\n    #  test_df preprocess\n    # ==============================================================\n\n    if only_last_aid:\n        _test_df = _test_df.groupby(\"session\").last().reset_index()\n\n    if same_real_session:\n        _test_df = make_real_session(_test_df, hours=real_session_hours)\n        _test_df[\"real_session_max\"] = _test_df.groupby(\"session\")[\"real_session\"].transform(\"max\")\n        _test_df = _test_df[_test_df[\"real_session\"]==_test_df[\"real_session_max\"]].reset_index(drop=True)\n        del _test_df[\"lag\"], _test_df[\"real_session\"], _test_df[\"real_session_max\"]; gc_clear()\n\n    if drop_duplicates:\n        _test_df = _test_df.groupby([\"session\", \"aid\"], as_index=False)[[\"type\", \"ts\"]].max()\n\n    # ==============================================================\n    #  Create recommend dictionary\n    # ==============================================================\n\n    if type_weighted:\n\n        _test_df[\"type_weights\"] = _test_df[\"type\"].map(type_weights).astype(\"int8\")\n        session_info = _test_df[_test_df[\"session\"].isin(target_session)]\\\n                       .groupby(\"session\", as_index=False)[[\"aid\", \"type_weights\", \"ts\"]].agg(list) \n        del _test_df; gc_clear()\n\n        rec = {}\n        for session, aids, type_weights, tss in tqdm(zip(session_info[\"session\"],\n                                                         session_info[\"aid\"],\n                                                         session_info[\"type_weights\"],\n                                                         session_info[\"ts\"]),\n                                                     total=len(session_info)):\n            rec.setdefault(session, Counter())\n            if time_weighted:\n                time_weights = make_time_weights(tss)\n                for aid, time_weight, type_weight in zip(aids, time_weights, type_weights):\n                    rec[session] += {aid: v*time_weight*type_weight for (aid, v) in aid_similarity.get(aid, {}).items()}\n            else:\n                for aid, type_weight in zip(aids, type_weights):\n                    rec[session] += {aid: v*type_weight for (aid, v) in aid_similarity.get(aid, {}).items()}\n\n    else:\n        session_info = _test_df[_test_df[\"session\"].isin(target_session)]\\\n                       .groupby(\"session\", as_index=False)[[\"aid\", \"ts\"]].agg(list) \n        del _test_df; gc_clear()\n\n        rec = {}\n        for session, aids, tss in tqdm(zip(session_info[\"session\"],\n                                           session_info[\"aid\"],\n                                           session_info[\"ts\"]),\n                                       total=len(session_info)):\n            rec.setdefault(session, Counter())\n            if time_weighted:\n                time_weights = make_time_weights(tss)\n                for aid, time_weight in zip(aids, time_weights):\n                    rec[session] += {aid: v*time_weight for (aid, v) in aid_similarity.get(aid, {}).items()}\n            else:\n                for aid in aids:\n                    rec[session] += aid_similarity.get(aid, {})        \n    del session_info; gc_clear()\n\n    # ==============================================================\n    #  Create features\n    # ==============================================================\n    feature = []\n    for session, aid in tqdm(zip(df[\"session\"], df[\"aid\"]), total=len(df)):\n        feature.append( rec.get(session, {}).get(aid, 0) )\n    df[feature_name] = feature\n\n    del feature, rec; gc_clear()\n    return df\n</code></pre>\n<ul>\n<li><strong>user-based CF features</strong></li>\n</ul>\n<p>6 patterns of item-based CF were created.</p>\n<ul>\n<li><strong>other features</strong><ul>\n<li><strong>aid features</strong><ul>\n<li>number of visits per week (train)</li>\n<li>number of visits per day (test)</li>\n<li>last click/cart/order ts per aid</li></ul></li>\n<li><strong>session features</strong><ul>\n<li>number of visits per session (test)</li>\n<li>last click/cart/order ts per session</li></ul></li>\n<li><strong>aid x session features</strong><ul>\n<li>last click/cart/order ts per session and aid</li>\n<li>percentage of aids visited since the last ts of each session in the test data</li></ul></li></ul></li>\n</ul>\n<h1>Model training</h1>\n<p>Different pipelines were used depending on which of the orders/carts/clicks was the target.</p>\n<p>For orders and carts, models were trained in two stages (1st stage/2nd stage).</p>\n<h4>orders model</h4>\n<ul>\n<li><strong>1st stage</strong></li>\n</ul>\n<p>All candidates were trained on the model without negative sampling.<br>\nI didn't want to do negative sampling as much as possible because negative sampling lowers the score.</p>\n<ul>\n<li><strong>2nd stage</strong></li>\n</ul>\n<p>As a result of the 1st stage, only the top 50 candidates per session were selected for the 2nd stage.<br>\nThe predictions from the 1st stage were not used for the 2nd stage features.</p>\n<h4>carts model</h4>\n<ul>\n<li><strong>1st stage</strong></li>\n</ul>\n<p>Since I could not run the model on all the data without negative sampling, I trained the model with negative sampling (x 0.3) in the 1st stage.</p>\n<ul>\n<li><strong>2nd stage</strong></li>\n</ul>\n<p>The model of the 1st stage was used to create oof for all candidates, and then the top 50 candidates per session were selected for the 2nd stage.</p>\n<h4>clicks model</h4>\n<ul>\n<li><strong>1st stage</strong></li>\n</ul>\n<p>Only the 1st stage was run with negative sampling (x 0.1).</p>\n<p>Probably the score would be higher if the 2nd stage was performed as in the carts model, but since the clicks have less weight on the score, the 2nd stage for the clicks was not conducted for cost-effectiveness.</p>\n<h4>Algorithm and Parameters</h4>\n<p>I created one Catboost model each to predict clicks, carts, and orders.</p>\n<h5>Parameters</h5>\n<pre><code>scale_pos_weight = (y_trn==0).sum()/(y_trn==1).sum()\n\nCAT_PARAMS = {\n    'loss_function': 'Logloss',\n    'learning_rate': 0.02,\n    'max_depth': 5,\n    'task_type': 'GPU',\n    'scale_pos_weight': scale_pos_weight,\n}\n</code></pre>\n<h1>Ensemble</h1>\n<p>I created models with 5 seeds and ensembled them.<br>\n(I changed not only the seed of the model parameters, but also the seed in the test/ground truth split.)</p>\n<p>Unlike the ensemble method in <a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">@radek1's notebook</a>, the ensemble was performed in a way to apply weight to each of the predicted rankings.</p>\n<h1>Post process</h1>\n<p>For sessions with fewer than 20 aids to recommend, I recommended popular aids.</p>\n<h1>what didn't work</h1>\n<ul>\n<li>word2vec</li>\n<li>ALS</li>\n<li>BPR</li>\n</ul>\n<h1>Environment</h1>\n<p>only Google Colab Pro+ :)</p>",
  "messages": [
    {
      "id": "2130249",
      "postDate": "02/05/2023 09:46:27",
      "content": "<p>First of all, thank you to the competition organizers for a great competition.<br>\nI am very happy to win my first gold medal🥇</p>\n<p>Here is my solution!</p>\n<h1>Overview</h1>\n<p><strong>best single model</strong></p>\n<table>\n<thead>\n<tr>\n<th>orders CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.67087</td>\n<td>0.602</td>\n</tr>\n</tbody>\n</table>\n<p>(I have not calculated the CV for all orders/carts/clicks combined.)</p>\n<p>Probably the same as many other competitors, the candidates generation &amp; rerank method was used.</p>\n<h1>Candidates generation</h1>\n<p>On average, <strong>180 candidates</strong> were selected per session. (<strong>orders recall: 0.725</strong>)</p>\n<ul>\n<li>14 patterns of item-based CF</li>\n<li>2 patterns of user-based CF</li>\n<li>re-visit</li>\n</ul>\n<p>The candidates generated by the 14 patterns of item-based CF and 2 patterns of user-based CF were assembled to create the final candidates.</p>\n<p>Since <strong>sessions with a larger number of aids in the test data are more likely to have a larger number of events in the forecasting period</strong>, more candidates were created for sessions with a larger number of aids in the test data.</p>\n<pre><code>df = test_df[test_df[\"session\"].isin(target_session)].groupby(\"session\")[\"aid\"].count().reset_index()\ndf.columns = [\"session\", \"count\"]\ndf[\"count\"] = (df[\"count\"]**0.5*10).astype(\"int32\")\ndf[\"aid\"] = df.progress_apply(lambda row:list(dict(rec[row[\"session\"]].most_common(row[\"count\"])).keys()),axis=1)\ndf = df.explode([\"aid\"])\ndf = df[df[\"aid\"].notnull()].reset_index(drop=True)\n</code></pre>\n<h1>Feature Engineering</h1>\n<p>A total of <strong>226 features</strong> were created.</p>\n<ul>\n<li><strong>item-based CF features</strong></li>\n</ul>\n<p>First, 59 patterns of item-based CF were created.</p>\n<p>ex) pattern1</p>\n<pre><code>def get_aid_similarity1(df, topk=200):\n    session_info = df.drop_duplicates([\"session\", \"aid\"], keep=\"last\")\\\n                     .groupby(\"session\", as_index=False)[[\"aid\", \"type\", \"ts\"]].agg(list) \n\n    aid_similarity = {}\n    for session, aids, tps, tss in tqdm(zip(session_info[\"session\"],\n                                            session_info[\"aid\"],\n                                            session_info[\"type\"],\n                                            session_info[\"ts\"]),\n                                        total=len(session_info)):\n        for aid1, tp1, ts1 in zip(aids, tps, tss):\n            session_length = math.sqrt(len(aids))\n            aid_similarity.setdefault(aid1, Counter())\n            for aid2, tp2, ts2 in zip(aids, tps, tss):\n                if (aid1 == aid2):\n                    continue\n                aid_similarity[aid1][aid2] += (1/session_length)\n\n    # Exclude all but the topK to save time and memory\n    for aid1, aid2_dict in tqdm(aid_similarity.items()):\n        relations = dict(aid2_dict.most_common(topk))\n        # normalize\n        if len(relations) == 0:\n            continue\n        max_num = relations[max(relations, key=relations.get)]\n        if max_num == 0:\n            continue\n        aid_similarity[aid1] = {k: v / max_num for k, v in relations.items()}\n\n    del session_info; gc_clear()\n    return aid_similarity\n</code></pre>\n<p>ex) pattern2</p>\n<pre><code>def make_real_session(df, hours=2):\n    df[\"lag\"] = df[\"ts\"] - df.groupby(\"session\")[\"ts\"].shift(1)\n    df[\"real_session\"] = (df[\"lag\"] &gt; 1000*60*60*hours).astype('int8').fillna(0)\n    df[\"real_session\"] = df.groupby(\"session\")[\"real_session\"].cumsum()\n    del df[\"lag\"]; gc_clear()\n    return df\n\ndef get_aid_similarity2(df, topk=200):\n\n    df = make_real_session(df, hours=4)\n    session_info = df.groupby([\"session\", \"real_session\"], as_index=False)[[\"aid\", \"type\", \"ts\"]].agg(list) \n\n    aid_similarity = {}\n    aid_cnt = defaultdict(int)\n    for session, real_session, aids, tps, tss in tqdm(zip(session_info[\"session\"],\n                                                          session_info[\"real_session\"],\n                                                          session_info[\"aid\"],\n                                                          session_info[\"type\"],\n                                                          session_info[\"ts\"]),\n                                                      total=len(session_info)):\n        for aid1, tp1, ts1 in zip(aids, tps, tss):\n            aid_similarity.setdefault(aid1, Counter())\n            for aid2, tp2, ts2 in zip(aids, tps, tss):\n                if (abs(ts1-ts2)&gt;24*60*60*1000) or (aid1 == aid2):\n                    continue\n                aid_cnt[aid1] += 1\n                if min(tp1, tp2)==0:\n                    aid_similarity[aid1][aid2] += 1\n                elif min(tp1, tp2)==1:\n                    aid_similarity[aid1][aid2] += 3\n                elif min(tp1, tp2)==2:\n                    aid_similarity[aid1][aid2] += 6\n\n    # Exclude all but the topK to save time and memory\n    for aid1, aid2_dict in tqdm(aid_similarity.items()):  \n        for aid2, score in aid2_dict.items():  \n            aid_similarity[aid1][aid2] = score / math.sqrt(aid_cnt[aid1]*aid_cnt[aid2])\n        aid_similarity[aid1] = dict(aid2_dict.most_common(topk))\n\n    del session_info; gc_clear()\n    return aid_similarity\n</code></pre>\n<p>(Further increasing the <code>topk</code> parameter did not improve the score.)</p>\n<p>Then, feature creation was performed using the various <code>aid_similarity</code> created above.<br>\nUnlike <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">cdeotte's notebook</a>, features were created by summing the value of sim_item_dict.</p>\n<pre><code>def recommend_aid(df,\n                  _test_df,\n                  target_session,\n                  aid_similarity,\n                  feature_name,\n                  type_weighted=False, type_weights={0:1, 1:3, 2:6},\n                  time_weighted=False,\n                  only_last_aid=False,\n                  same_real_session=False, real_session_hours=2,\n                  drop_duplicates=False\n                  ):\n\n    # ==============================================================\n    #  test_df preprocess\n    # ==============================================================\n\n    if only_last_aid:\n        _test_df = _test_df.groupby(\"session\").last().reset_index()\n\n    if same_real_session:\n        _test_df = make_real_session(_test_df, hours=real_session_hours)\n        _test_df[\"real_session_max\"] = _test_df.groupby(\"session\")[\"real_session\"].transform(\"max\")\n        _test_df = _test_df[_test_df[\"real_session\"]==_test_df[\"real_session_max\"]].reset_index(drop=True)\n        del _test_df[\"lag\"], _test_df[\"real_session\"], _test_df[\"real_session_max\"]; gc_clear()\n\n    if drop_duplicates:\n        _test_df = _test_df.groupby([\"session\", \"aid\"], as_index=False)[[\"type\", \"ts\"]].max()\n\n    # ==============================================================\n    #  Create recommend dictionary\n    # ==============================================================\n\n    if type_weighted:\n\n        _test_df[\"type_weights\"] = _test_df[\"type\"].map(type_weights).astype(\"int8\")\n        session_info = _test_df[_test_df[\"session\"].isin(target_session)]\\\n                       .groupby(\"session\", as_index=False)[[\"aid\", \"type_weights\", \"ts\"]].agg(list) \n        del _test_df; gc_clear()\n\n        rec = {}\n        for session, aids, type_weights, tss in tqdm(zip(session_info[\"session\"],\n                                                         session_info[\"aid\"],\n                                                         session_info[\"type_weights\"],\n                                                         session_info[\"ts\"]),\n                                                     total=len(session_info)):\n            rec.setdefault(session, Counter())\n            if time_weighted:\n                time_weights = make_time_weights(tss)\n                for aid, time_weight, type_weight in zip(aids, time_weights, type_weights):\n                    rec[session] += {aid: v*time_weight*type_weight for (aid, v) in aid_similarity.get(aid, {}).items()}\n            else:\n                for aid, type_weight in zip(aids, type_weights):\n                    rec[session] += {aid: v*type_weight for (aid, v) in aid_similarity.get(aid, {}).items()}\n\n    else:\n        session_info = _test_df[_test_df[\"session\"].isin(target_session)]\\\n                       .groupby(\"session\", as_index=False)[[\"aid\", \"ts\"]].agg(list) \n        del _test_df; gc_clear()\n\n        rec = {}\n        for session, aids, tss in tqdm(zip(session_info[\"session\"],\n                                           session_info[\"aid\"],\n                                           session_info[\"ts\"]),\n                                       total=len(session_info)):\n            rec.setdefault(session, Counter())\n            if time_weighted:\n                time_weights = make_time_weights(tss)\n                for aid, time_weight in zip(aids, time_weights):\n                    rec[session] += {aid: v*time_weight for (aid, v) in aid_similarity.get(aid, {}).items()}\n            else:\n                for aid in aids:\n                    rec[session] += aid_similarity.get(aid, {})        \n    del session_info; gc_clear()\n\n    # ==============================================================\n    #  Create features\n    # ==============================================================\n    feature = []\n    for session, aid in tqdm(zip(df[\"session\"], df[\"aid\"]), total=len(df)):\n        feature.append( rec.get(session, {}).get(aid, 0) )\n    df[feature_name] = feature\n\n    del feature, rec; gc_clear()\n    return df\n</code></pre>\n<ul>\n<li><strong>user-based CF features</strong></li>\n</ul>\n<p>6 patterns of item-based CF were created.</p>\n<ul>\n<li><strong>other features</strong><ul>\n<li><strong>aid features</strong><ul>\n<li>number of visits per week (train)</li>\n<li>number of visits per day (test)</li>\n<li>last click/cart/order ts per aid</li></ul></li>\n<li><strong>session features</strong><ul>\n<li>number of visits per session (test)</li>\n<li>last click/cart/order ts per session</li></ul></li>\n<li><strong>aid x session features</strong><ul>\n<li>last click/cart/order ts per session and aid</li>\n<li>percentage of aids visited since the last ts of each session in the test data</li></ul></li></ul></li>\n</ul>\n<h1>Model training</h1>\n<p>Different pipelines were used depending on which of the orders/carts/clicks was the target.</p>\n<p>For orders and carts, models were trained in two stages (1st stage/2nd stage).</p>\n<h4>orders model</h4>\n<ul>\n<li><strong>1st stage</strong></li>\n</ul>\n<p>All candidates were trained on the model without negative sampling.<br>\nI didn't want to do negative sampling as much as possible because negative sampling lowers the score.</p>\n<ul>\n<li><strong>2nd stage</strong></li>\n</ul>\n<p>As a result of the 1st stage, only the top 50 candidates per session were selected for the 2nd stage.<br>\nThe predictions from the 1st stage were not used for the 2nd stage features.</p>\n<h4>carts model</h4>\n<ul>\n<li><strong>1st stage</strong></li>\n</ul>\n<p>Since I could not run the model on all the data without negative sampling, I trained the model with negative sampling (x 0.3) in the 1st stage.</p>\n<ul>\n<li><strong>2nd stage</strong></li>\n</ul>\n<p>The model of the 1st stage was used to create oof for all candidates, and then the top 50 candidates per session were selected for the 2nd stage.</p>\n<h4>clicks model</h4>\n<ul>\n<li><strong>1st stage</strong></li>\n</ul>\n<p>Only the 1st stage was run with negative sampling (x 0.1).</p>\n<p>Probably the score would be higher if the 2nd stage was performed as in the carts model, but since the clicks have less weight on the score, the 2nd stage for the clicks was not conducted for cost-effectiveness.</p>\n<h4>Algorithm and Parameters</h4>\n<p>I created one Catboost model each to predict clicks, carts, and orders.</p>\n<h5>Parameters</h5>\n<pre><code>scale_pos_weight = (y_trn==0).sum()/(y_trn==1).sum()\n\nCAT_PARAMS = {\n    'loss_function': 'Logloss',\n    'learning_rate': 0.02,\n    'max_depth': 5,\n    'task_type': 'GPU',\n    'scale_pos_weight': scale_pos_weight,\n}\n</code></pre>\n<h1>Ensemble</h1>\n<p>I created models with 5 seeds and ensembled them.<br>\n(I changed not only the seed of the model parameters, but also the seed in the test/ground truth split.)</p>\n<p>Unlike the ensemble method in <a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">@radek1's notebook</a>, the ensemble was performed in a way to apply weight to each of the predicted rankings.</p>\n<h1>Post process</h1>\n<p>For sessions with fewer than 20 aids to recommend, I recommended popular aids.</p>\n<h1>what didn't work</h1>\n<ul>\n<li>word2vec</li>\n<li>ALS</li>\n<li>BPR</li>\n</ul>\n<h1>Environment</h1>\n<p>only Google Colab Pro+ :)</p>",
      "rawMarkdown": "First of all, thank you to the competition organizers for a great competition.\nI am very happy to win my first gold medal🥇\n\nHere is my solution!\n\n# Overview\n\n**best single model**\n| orders CV | LB |\n| ---- | ---- |\n| 0.67087 | 0.602 |\n\n(I have not calculated the CV for all orders/carts/clicks combined.)\n\nProbably the same as many other competitors, the candidates generation & rerank method was used.\n\n# Candidates generation\n\nOn average, **180 candidates** were selected per session. (**orders recall: 0.725**)\n\n- 14 patterns of item-based CF\n- 2 patterns of user-based CF\n- re-visit\n\nThe candidates generated by the 14 patterns of item-based CF and 2 patterns of user-based CF were assembled to create the final candidates.\n\nSince **sessions with a larger number of aids in the test data are more likely to have a larger number of events in the forecasting period**, more candidates were created for sessions with a larger number of aids in the test data.\n\n```\ndf = test_df[test_df[\"session\"].isin(target_session)].groupby(\"session\")[\"aid\"].count().reset_index()\ndf.columns = [\"session\", \"count\"]\ndf[\"count\"] = (df[\"count\"]**0.5*10).astype(\"int32\")\ndf[\"aid\"] = df.progress_apply(lambda row:list(dict(rec[row[\"session\"]].most_common(row[\"count\"])).keys()),axis=1)\ndf = df.explode([\"aid\"])\ndf = df[df[\"aid\"].notnull()].reset_index(drop=True)\n```\n# Feature Engineering\n\nA total of **226 features** were created.\n\n- **item-based CF features**\n\nFirst, 59 patterns of item-based CF were created.\n\nex) pattern1\n```\ndef get_aid_similarity1(df, topk=200):\n    session_info = df.drop_duplicates([\"session\", \"aid\"], keep=\"last\")\\\n                     .groupby(\"session\", as_index=False)[[\"aid\", \"type\", \"ts\"]].agg(list) \n    \n    aid_similarity = {}\n    for session, aids, tps, tss in tqdm(zip(session_info[\"session\"],\n                                            session_info[\"aid\"],\n                                            session_info[\"type\"],\n                                            session_info[\"ts\"]),\n                                        total=len(session_info)):\n        for aid1, tp1, ts1 in zip(aids, tps, tss):\n            session_length = math.sqrt(len(aids))\n            aid_similarity.setdefault(aid1, Counter())\n            for aid2, tp2, ts2 in zip(aids, tps, tss):\n                if (aid1 == aid2):\n                    continue\n                aid_similarity[aid1][aid2] += (1/session_length)\n    \n    # Exclude all but the topK to save time and memory\n    for aid1, aid2_dict in tqdm(aid_similarity.items()):\n        relations = dict(aid2_dict.most_common(topk))\n        # normalize\n        if len(relations) == 0:\n            continue\n        max_num = relations[max(relations, key=relations.get)]\n        if max_num == 0:\n            continue\n        aid_similarity[aid1] = {k: v / max_num for k, v in relations.items()}\n\n    del session_info; gc_clear()\n    return aid_similarity\n```\nex) pattern2\n\n```\ndef make_real_session(df, hours=2):\n    df[\"lag\"] = df[\"ts\"] - df.groupby(\"session\")[\"ts\"].shift(1)\n    df[\"real_session\"] = (df[\"lag\"] > 1000*60*60*hours).astype('int8').fillna(0)\n    df[\"real_session\"] = df.groupby(\"session\")[\"real_session\"].cumsum()\n    del df[\"lag\"]; gc_clear()\n    return df\n\ndef get_aid_similarity2(df, topk=200):\n\n    df = make_real_session(df, hours=4)\n    session_info = df.groupby([\"session\", \"real_session\"], as_index=False)[[\"aid\", \"type\", \"ts\"]].agg(list) \n    \n    aid_similarity = {}\n    aid_cnt = defaultdict(int)\n    for session, real_session, aids, tps, tss in tqdm(zip(session_info[\"session\"],\n                                                          session_info[\"real_session\"],\n                                                          session_info[\"aid\"],\n                                                          session_info[\"type\"],\n                                                          session_info[\"ts\"]),\n                                                      total=len(session_info)):\n        for aid1, tp1, ts1 in zip(aids, tps, tss):\n            aid_similarity.setdefault(aid1, Counter())\n            for aid2, tp2, ts2 in zip(aids, tps, tss):\n                if (abs(ts1-ts2)>24*60*60*1000) or (aid1 == aid2):\n                    continue\n                aid_cnt[aid1] += 1\n                if min(tp1, tp2)==0:\n                    aid_similarity[aid1][aid2] += 1\n                elif min(tp1, tp2)==1:\n                    aid_similarity[aid1][aid2] += 3\n                elif min(tp1, tp2)==2:\n                    aid_similarity[aid1][aid2] += 6\n    \n    # Exclude all but the topK to save time and memory\n    for aid1, aid2_dict in tqdm(aid_similarity.items()):  \n        for aid2, score in aid2_dict.items():  \n            aid_similarity[aid1][aid2] = score / math.sqrt(aid_cnt[aid1]*aid_cnt[aid2])\n        aid_similarity[aid1] = dict(aid2_dict.most_common(topk))\n       \n    del session_info; gc_clear()\n    return aid_similarity\n```\n\n\n(Further increasing the `topk` parameter did not improve the score.)\n\nThen, feature creation was performed using the various `aid_similarity` created above.\nUnlike [cdeotte's notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575), features were created by summing the value of sim_item_dict.\n\n```\ndef recommend_aid(df,\n                  _test_df,\n                  target_session,\n                  aid_similarity,\n                  feature_name,\n                  type_weighted=False, type_weights={0:1, 1:3, 2:6},\n                  time_weighted=False,\n                  only_last_aid=False,\n                  same_real_session=False, real_session_hours=2,\n                  drop_duplicates=False\n                  ):\n    \n    # ==============================================================\n    #  test_df preprocess\n    # ==============================================================\n\n    if only_last_aid:\n        _test_df = _test_df.groupby(\"session\").last().reset_index()\n\n    if same_real_session:\n        _test_df = make_real_session(_test_df, hours=real_session_hours)\n        _test_df[\"real_session_max\"] = _test_df.groupby(\"session\")[\"real_session\"].transform(\"max\")\n        _test_df = _test_df[_test_df[\"real_session\"]==_test_df[\"real_session_max\"]].reset_index(drop=True)\n        del _test_df[\"lag\"], _test_df[\"real_session\"], _test_df[\"real_session_max\"]; gc_clear()\n    \n    if drop_duplicates:\n        _test_df = _test_df.groupby([\"session\", \"aid\"], as_index=False)[[\"type\", \"ts\"]].max()\n    \n    # ==============================================================\n    #  Create recommend dictionary\n    # ==============================================================\n        \n    if type_weighted:\n        \n        _test_df[\"type_weights\"] = _test_df[\"type\"].map(type_weights).astype(\"int8\")\n        session_info = _test_df[_test_df[\"session\"].isin(target_session)]\\\n                       .groupby(\"session\", as_index=False)[[\"aid\", \"type_weights\", \"ts\"]].agg(list) \n        del _test_df; gc_clear()\n\n        rec = {}\n        for session, aids, type_weights, tss in tqdm(zip(session_info[\"session\"],\n                                                         session_info[\"aid\"],\n                                                         session_info[\"type_weights\"],\n                                                         session_info[\"ts\"]),\n                                                     total=len(session_info)):\n            rec.setdefault(session, Counter())\n            if time_weighted:\n                time_weights = make_time_weights(tss)\n                for aid, time_weight, type_weight in zip(aids, time_weights, type_weights):\n                    rec[session] += {aid: v*time_weight*type_weight for (aid, v) in aid_similarity.get(aid, {}).items()}\n            else:\n                for aid, type_weight in zip(aids, type_weights):\n                    rec[session] += {aid: v*type_weight for (aid, v) in aid_similarity.get(aid, {}).items()}\n\n    else:\n        session_info = _test_df[_test_df[\"session\"].isin(target_session)]\\\n                       .groupby(\"session\", as_index=False)[[\"aid\", \"ts\"]].agg(list) \n        del _test_df; gc_clear()\n\n        rec = {}\n        for session, aids, tss in tqdm(zip(session_info[\"session\"],\n                                           session_info[\"aid\"],\n                                           session_info[\"ts\"]),\n                                       total=len(session_info)):\n            rec.setdefault(session, Counter())\n            if time_weighted:\n                time_weights = make_time_weights(tss)\n                for aid, time_weight in zip(aids, time_weights):\n                    rec[session] += {aid: v*time_weight for (aid, v) in aid_similarity.get(aid, {}).items()}\n            else:\n                for aid in aids:\n                    rec[session] += aid_similarity.get(aid, {})        \n    del session_info; gc_clear()\n    \n    # ==============================================================\n    #  Create features\n    # ==============================================================\n    feature = []\n    for session, aid in tqdm(zip(df[\"session\"], df[\"aid\"]), total=len(df)):\n        feature.append( rec.get(session, {}).get(aid, 0) )\n    df[feature_name] = feature\n\n    del feature, rec; gc_clear()\n    return df\n```\n\n- **user-based CF features**\n\n6 patterns of item-based CF were created.\n\n- **other features**\n   - **aid features**\n       - number of visits per week (train)\n       - number of visits per day (test)\n       - last click/cart/order ts per aid\n   - **session features**\n       - number of visits per session (test)\n       - last click/cart/order ts per session\n   - **aid x session features**\n       - last click/cart/order ts per session and aid\n       - percentage of aids visited since the last ts of each session in the test data\n\n# Model training\n\nDifferent pipelines were used depending on which of the orders/carts/clicks was the target.\n\nFor orders and carts, models were trained in two stages (1st stage/2nd stage).\n\n\n#### orders model\n- **1st stage**\n\nAll candidates were trained on the model without negative sampling.\nI didn't want to do negative sampling as much as possible because negative sampling lowers the score.\n\n- **2nd stage**\n\nAs a result of the 1st stage, only the top 50 candidates per session were selected for the 2nd stage.\nThe predictions from the 1st stage were not used for the 2nd stage features.\n\n#### carts model\n- **1st stage**\n\nSince I could not run the model on all the data without negative sampling, I trained the model with negative sampling (x 0.3) in the 1st stage.\n\n- **2nd stage**\n\nThe model of the 1st stage was used to create oof for all candidates, and then the top 50 candidates per session were selected for the 2nd stage.\n\n#### clicks model\n- **1st stage**\n\nOnly the 1st stage was run with negative sampling (x 0.1).\n\nProbably the score would be higher if the 2nd stage was performed as in the carts model, but since the clicks have less weight on the score, the 2nd stage for the clicks was not conducted for cost-effectiveness.\n\n#### Algorithm and Parameters\n\nI created one Catboost model each to predict clicks, carts, and orders.\n\n##### Parameters\n\n```\nscale_pos_weight = (y_trn==0).sum()/(y_trn==1).sum()\n\nCAT_PARAMS = {\n    'loss_function': 'Logloss',\n    'learning_rate': 0.02,\n    'max_depth': 5,\n    'task_type': 'GPU',\n    'scale_pos_weight': scale_pos_weight,\n}\n```\n\n# Ensemble\n\nI created models with 5 seeds and ensembled them.\n(I changed not only the seed of the model parameters, but also the seed in the test/ground truth split.)\n\nUnlike the ensemble method in [@radek1's notebook](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions), the ensemble was performed in a way to apply weight to each of the predicted rankings.\n\n# Post process\n\nFor sessions with fewer than 20 aids to recommend, I recommended popular aids.\n\n# what didn't work\n- word2vec\n- ALS\n- BPR\n\n# Environment\nonly Google Colab Pro+ :)",
      "votes": null
    },
    {
      "id": "2131156",
      "postDate": "02/05/2023 23:14:33",
      "content": "<p>Amazing job going solo and using single model. Congratulations on solo gold medal !!</p>",
      "rawMarkdown": "Amazing job going solo and using single model. Congratulations on solo gold medal !!",
      "votes": null
    },
    {
      "id": "2131782",
      "postDate": "02/06/2023 11:10:55",
      "content": "<p><a href=\"https://www.kaggle.com/shkanda\" target=\"_blank\">@shkanda</a> You can set the URL of your solution from Team tab. It will be linked from the leaderboard.<br>\n<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/team\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/team</a></p>",
      "rawMarkdown": "shkanda You can set the URL of your solution from Team tab. It will be linked from the leaderboard.\nhttps://www.kaggle.com/competitions/otto-recommender-system/team",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2131156,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/05/2023 23:14:33",
      "content": "<p>Amazing job going solo and using single model. Congratulations on solo gold medal !!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2131782,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "02/06/2023 11:10:55",
      "content": "<p><a href=\"https://www.kaggle.com/shkanda\" target=\"_blank\">@shkanda</a> You can set the URL of your solution from Team tab. It will be linked from the leaderboard.<br>\n<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/team\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/team</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2130249": "First of all, thank you to the competition organizers for a great competition.\nI am very happy to win my first gold medal🥇\n\nHere is my solution!\n\n# Overview\n\n**best single model**\n| orders CV | LB |\n| ---- | ---- |\n| 0.67087 | 0.602 |\n\n(I have not calculated the CV for all orders/carts/clicks combined.)\n\nProbably the same as many other competitors, the candidates generation & rerank method was used.\n\n# Candidates generation\n\nOn average, **180 candidates** were selected per session. (**orders recall: 0.725**)\n\n- 14 patterns of item-based CF\n- 2 patterns of user-based CF\n- re-visit\n\nThe candidates generated by the 14 patterns of item-based CF and 2 patterns of user-based CF were assembled to create the final candidates.\n\nSince **sessions with a larger number of aids in the test data are more likely to have a larger number of events in the forecasting period**, more candidates were created for sessions with a larger number of aids in the test data.\n\n```\ndf = test_df[test_df[\"session\"].isin(target_session)].groupby(\"session\")[\"aid\"].count().reset_index()\ndf.columns = [\"session\", \"count\"]\ndf[\"count\"] = (df[\"count\"]**0.5*10).astype(\"int32\")\ndf[\"aid\"] = df.progress_apply(lambda row:list(dict(rec[row[\"session\"]].most_common(row[\"count\"])).keys()),axis=1)\ndf = df.explode([\"aid\"])\ndf = df[df[\"aid\"].notnull()].reset_index(drop=True)\n```\n# Feature Engineering\n\nA total of **226 features** were created.\n\n- **item-based CF features**\n\nFirst, 59 patterns of item-based CF were created.\n\nex) pattern1\n```\ndef get_aid_similarity1(df, topk=200):\n    session_info = df.drop_duplicates([\"session\", \"aid\"], keep=\"last\")\\\n                     .groupby(\"session\", as_index=False)[[\"aid\", \"type\", \"ts\"]].agg(list) \n    \n    aid_similarity = {}\n    for session, aids, tps, tss in tqdm(zip(session_info[\"session\"],\n                                            session_info[\"aid\"],\n                                            session_info[\"type\"],\n                                            session_info[\"ts\"]),\n                                        total=len(session_info)):\n        for aid1, tp1, ts1 in zip(aids, tps, tss):\n            session_length = math.sqrt(len(aids))\n            aid_similarity.setdefault(aid1, Counter())\n            for aid2, tp2, ts2 in zip(aids, tps, tss):\n                if (aid1 == aid2):\n                    continue\n                aid_similarity[aid1][aid2] += (1/session_length)\n    \n    # Exclude all but the topK to save time and memory\n    for aid1, aid2_dict in tqdm(aid_similarity.items()):\n        relations = dict(aid2_dict.most_common(topk))\n        # normalize\n        if len(relations) == 0:\n            continue\n        max_num = relations[max(relations, key=relations.get)]\n        if max_num == 0:\n            continue\n        aid_similarity[aid1] = {k: v / max_num for k, v in relations.items()}\n\n    del session_info; gc_clear()\n    return aid_similarity\n```\nex) pattern2\n\n```\ndef make_real_session(df, hours=2):\n    df[\"lag\"] = df[\"ts\"] - df.groupby(\"session\")[\"ts\"].shift(1)\n    df[\"real_session\"] = (df[\"lag\"] > 1000*60*60*hours).astype('int8').fillna(0)\n    df[\"real_session\"] = df.groupby(\"session\")[\"real_session\"].cumsum()\n    del df[\"lag\"]; gc_clear()\n    return df\n\ndef get_aid_similarity2(df, topk=200):\n\n    df = make_real_session(df, hours=4)\n    session_info = df.groupby([\"session\", \"real_session\"], as_index=False)[[\"aid\", \"type\", \"ts\"]].agg(list) \n    \n    aid_similarity = {}\n    aid_cnt = defaultdict(int)\n    for session, real_session, aids, tps, tss in tqdm(zip(session_info[\"session\"],\n                                                          session_info[\"real_session\"],\n                                                          session_info[\"aid\"],\n                                                          session_info[\"type\"],\n                                                          session_info[\"ts\"]),\n                                                      total=len(session_info)):\n        for aid1, tp1, ts1 in zip(aids, tps, tss):\n            aid_similarity.setdefault(aid1, Counter())\n            for aid2, tp2, ts2 in zip(aids, tps, tss):\n                if (abs(ts1-ts2)>24*60*60*1000) or (aid1 == aid2):\n                    continue\n                aid_cnt[aid1] += 1\n                if min(tp1, tp2)==0:\n                    aid_similarity[aid1][aid2] += 1\n                elif min(tp1, tp2)==1:\n                    aid_similarity[aid1][aid2] += 3\n                elif min(tp1, tp2)==2:\n                    aid_similarity[aid1][aid2] += 6\n    \n    # Exclude all but the topK to save time and memory\n    for aid1, aid2_dict in tqdm(aid_similarity.items()):  \n        for aid2, score in aid2_dict.items():  \n            aid_similarity[aid1][aid2] = score / math.sqrt(aid_cnt[aid1]*aid_cnt[aid2])\n        aid_similarity[aid1] = dict(aid2_dict.most_common(topk))\n       \n    del session_info; gc_clear()\n    return aid_similarity\n```\n\n\n(Further increasing the `topk` parameter did not improve the score.)\n\nThen, feature creation was performed using the various `aid_similarity` created above.\nUnlike [cdeotte's notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575), features were created by summing the value of sim_item_dict.\n\n```\ndef recommend_aid(df,\n                  _test_df,\n                  target_session,\n                  aid_similarity,\n                  feature_name,\n                  type_weighted=False, type_weights={0:1, 1:3, 2:6},\n                  time_weighted=False,\n                  only_last_aid=False,\n                  same_real_session=False, real_session_hours=2,\n                  drop_duplicates=False\n                  ):\n    \n    # ==============================================================\n    #  test_df preprocess\n    # ==============================================================\n\n    if only_last_aid:\n        _test_df = _test_df.groupby(\"session\").last().reset_index()\n\n    if same_real_session:\n        _test_df = make_real_session(_test_df, hours=real_session_hours)\n        _test_df[\"real_session_max\"] = _test_df.groupby(\"session\")[\"real_session\"].transform(\"max\")\n        _test_df = _test_df[_test_df[\"real_session\"]==_test_df[\"real_session_max\"]].reset_index(drop=True)\n        del _test_df[\"lag\"], _test_df[\"real_session\"], _test_df[\"real_session_max\"]; gc_clear()\n    \n    if drop_duplicates:\n        _test_df = _test_df.groupby([\"session\", \"aid\"], as_index=False)[[\"type\", \"ts\"]].max()\n    \n    # ==============================================================\n    #  Create recommend dictionary\n    # ==============================================================\n        \n    if type_weighted:\n        \n        _test_df[\"type_weights\"] = _test_df[\"type\"].map(type_weights).astype(\"int8\")\n        session_info = _test_df[_test_df[\"session\"].isin(target_session)]\\\n                       .groupby(\"session\", as_index=False)[[\"aid\", \"type_weights\", \"ts\"]].agg(list) \n        del _test_df; gc_clear()\n\n        rec = {}\n        for session, aids, type_weights, tss in tqdm(zip(session_info[\"session\"],\n                                                         session_info[\"aid\"],\n                                                         session_info[\"type_weights\"],\n                                                         session_info[\"ts\"]),\n                                                     total=len(session_info)):\n            rec.setdefault(session, Counter())\n            if time_weighted:\n                time_weights = make_time_weights(tss)\n                for aid, time_weight, type_weight in zip(aids, time_weights, type_weights):\n                    rec[session] += {aid: v*time_weight*type_weight for (aid, v) in aid_similarity.get(aid, {}).items()}\n            else:\n                for aid, type_weight in zip(aids, type_weights):\n                    rec[session] += {aid: v*type_weight for (aid, v) in aid_similarity.get(aid, {}).items()}\n\n    else:\n        session_info = _test_df[_test_df[\"session\"].isin(target_session)]\\\n                       .groupby(\"session\", as_index=False)[[\"aid\", \"ts\"]].agg(list) \n        del _test_df; gc_clear()\n\n        rec = {}\n        for session, aids, tss in tqdm(zip(session_info[\"session\"],\n                                           session_info[\"aid\"],\n                                           session_info[\"ts\"]),\n                                       total=len(session_info)):\n            rec.setdefault(session, Counter())\n            if time_weighted:\n                time_weights = make_time_weights(tss)\n                for aid, time_weight in zip(aids, time_weights):\n                    rec[session] += {aid: v*time_weight for (aid, v) in aid_similarity.get(aid, {}).items()}\n            else:\n                for aid in aids:\n                    rec[session] += aid_similarity.get(aid, {})        \n    del session_info; gc_clear()\n    \n    # ==============================================================\n    #  Create features\n    # ==============================================================\n    feature = []\n    for session, aid in tqdm(zip(df[\"session\"], df[\"aid\"]), total=len(df)):\n        feature.append( rec.get(session, {}).get(aid, 0) )\n    df[feature_name] = feature\n\n    del feature, rec; gc_clear()\n    return df\n```\n\n- **user-based CF features**\n\n6 patterns of item-based CF were created.\n\n- **other features**\n   - **aid features**\n       - number of visits per week (train)\n       - number of visits per day (test)\n       - last click/cart/order ts per aid\n   - **session features**\n       - number of visits per session (test)\n       - last click/cart/order ts per session\n   - **aid x session features**\n       - last click/cart/order ts per session and aid\n       - percentage of aids visited since the last ts of each session in the test data\n\n# Model training\n\nDifferent pipelines were used depending on which of the orders/carts/clicks was the target.\n\nFor orders and carts, models were trained in two stages (1st stage/2nd stage).\n\n\n#### orders model\n- **1st stage**\n\nAll candidates were trained on the model without negative sampling.\nI didn't want to do negative sampling as much as possible because negative sampling lowers the score.\n\n- **2nd stage**\n\nAs a result of the 1st stage, only the top 50 candidates per session were selected for the 2nd stage.\nThe predictions from the 1st stage were not used for the 2nd stage features.\n\n#### carts model\n- **1st stage**\n\nSince I could not run the model on all the data without negative sampling, I trained the model with negative sampling (x 0.3) in the 1st stage.\n\n- **2nd stage**\n\nThe model of the 1st stage was used to create oof for all candidates, and then the top 50 candidates per session were selected for the 2nd stage.\n\n#### clicks model\n- **1st stage**\n\nOnly the 1st stage was run with negative sampling (x 0.1).\n\nProbably the score would be higher if the 2nd stage was performed as in the carts model, but since the clicks have less weight on the score, the 2nd stage for the clicks was not conducted for cost-effectiveness.\n\n#### Algorithm and Parameters\n\nI created one Catboost model each to predict clicks, carts, and orders.\n\n##### Parameters\n\n```\nscale_pos_weight = (y_trn==0).sum()/(y_trn==1).sum()\n\nCAT_PARAMS = {\n    'loss_function': 'Logloss',\n    'learning_rate': 0.02,\n    'max_depth': 5,\n    'task_type': 'GPU',\n    'scale_pos_weight': scale_pos_weight,\n}\n```\n\n# Ensemble\n\nI created models with 5 seeds and ensembled them.\n(I changed not only the seed of the model parameters, but also the seed in the test/ground truth split.)\n\nUnlike the ensemble method in [@radek1's notebook](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions), the ensemble was performed in a way to apply weight to each of the predicted rankings.\n\n# Post process\n\nFor sessions with fewer than 20 aids to recommend, I recommended popular aids.\n\n# what didn't work\n- word2vec\n- ALS\n- BPR\n\n# Environment\nonly Google Colab Pro+ :)",
    "2131156": "Amazing job going solo and using single model. Congratulations on solo gold medal !!",
    "2131782": "shkanda You can set the URL of your solution from Team tab. It will be linked from the leaderboard.\nhttps://www.kaggle.com/competitions/otto-recommender-system/team"
  },
  "source": "meta"
}