{
  "id": 383374,
  "title": "14th Place Solution(Ethan&qyxs part)",
  "url": "/competitions/otto-recommender-system/writeups/riotto-14th-place-solution-ethan-qyxs-part",
  "author_name": "",
  "post_date": "2023-02-03T12:45:37.499584900Z",
  "votes": 23,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks to OTTO and kaggle hosting such a good competition. Thanks my great teammates <a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> <a href=\"https://www.kaggle.com/rib\" target=\"_blank\">@rib</a> <a href=\"https://www.kaggle.com/ria\" target=\"_blank\">@ria</a>.</p>\n<h2>Validation</h2>\n<p>use the sessions in last week.</p>\n<h2>Retrieval</h2>\n<p>two traditional methods,</p>\n<ul>\n<li>u2i: the aids user action before</li>\n<li>i2i: itemcf, we do it seperately for each task. And we used some kinds of weight to optimize its performance, such as \"time_gap\", \"loc_gap\", \"type_difference\" between two aids.</li>\n</ul>\n<h2>Ranking</h2>\n<ul>\n<li>Features</li>\n</ul>\n<ol>\n<li>session based: count, last action time_gap(click/carts/orders), last action type </li>\n<li>aid based: count, ratio(ratio of click, ratio of carts to orders…), time</li>\n<li>session+aid based: count, last action time_gap(click/carts/orders), last action type</li>\n<li>w2v embeddings: difference and similarity of candiates between sessions' history aids, then use mean\\max\\min\\last\\std</li>\n<li>collaborative filtering score: collaborative filtering score of i2i, calculate the items between the candidates and user's historical aids.</li>\n<li>count and ratio features of aids after/before diffrent windows(1day, 3day) of each sessions' last time</li>\n</ol>\n<ul>\n<li>Data Augment</li>\n</ul>\n<ol>\n<li>use more weeks to train, we used 3 weeks for finnal submission.</li>\n<li>use different seeds to generate more traindata(different cutoff of sessions), we used 5 seeds at last. <br>\nthose bring 0.001 boost on local cv.</li>\n</ol>\n<ul>\n<li>Model</li>\n</ul>\n<ol>\n<li>lightgbm binary classifier, learning rate: 0.05, 1000 rounds</li>\n<li>catboost binary classifier, learning rate: 0.05, 6000 rounds<br>\nenemble these two results scored LB:0.600.</li>\n</ol>\n<h2>Interesting Findings</h2>\n<p>When optimize the task of orders, </p>\n<ol>\n<li>orders_prob = 0.7* orders_prob+0.2* carts_prob+0.1* click_prob, brings 0.0004 boost on orders' local cv(evaluation on the same session of three task).</li>\n<li>My teammate rib concats carts and orders' training data, and add one feature(0, 1) whether the sample comes from carts or not, can get 0.0003 boost on orders' score.</li>\n</ol>\n<h2>Enemble</h2>\n<p>For final submission, we used voting to enemble with Rib and ria's result(different candidats, LB: 0.601), LB goes to 0.602.</p>",
  "messages": [
    {
      "id": "2128056",
      "postDate": "02/03/2023 12:45:37",
      "content": "<p>Thanks to OTTO and kaggle hosting such a good competition. Thanks my great teammates <a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> <a href=\"https://www.kaggle.com/rib\" target=\"_blank\">@rib</a> <a href=\"https://www.kaggle.com/ria\" target=\"_blank\">@ria</a>.</p>\n<h2>Validation</h2>\n<p>use the sessions in last week.</p>\n<h2>Retrieval</h2>\n<p>two traditional methods,</p>\n<ul>\n<li>u2i: the aids user action before</li>\n<li>i2i: itemcf, we do it seperately for each task. And we used some kinds of weight to optimize its performance, such as \"time_gap\", \"loc_gap\", \"type_difference\" between two aids.</li>\n</ul>\n<h2>Ranking</h2>\n<ul>\n<li>Features</li>\n</ul>\n<ol>\n<li>session based: count, last action time_gap(click/carts/orders), last action type </li>\n<li>aid based: count, ratio(ratio of click, ratio of carts to orders…), time</li>\n<li>session+aid based: count, last action time_gap(click/carts/orders), last action type</li>\n<li>w2v embeddings: difference and similarity of candiates between sessions' history aids, then use mean\\max\\min\\last\\std</li>\n<li>collaborative filtering score: collaborative filtering score of i2i, calculate the items between the candidates and user's historical aids.</li>\n<li>count and ratio features of aids after/before diffrent windows(1day, 3day) of each sessions' last time</li>\n</ol>\n<ul>\n<li>Data Augment</li>\n</ul>\n<ol>\n<li>use more weeks to train, we used 3 weeks for finnal submission.</li>\n<li>use different seeds to generate more traindata(different cutoff of sessions), we used 5 seeds at last. <br>\nthose bring 0.001 boost on local cv.</li>\n</ol>\n<ul>\n<li>Model</li>\n</ul>\n<ol>\n<li>lightgbm binary classifier, learning rate: 0.05, 1000 rounds</li>\n<li>catboost binary classifier, learning rate: 0.05, 6000 rounds<br>\nenemble these two results scored LB:0.600.</li>\n</ol>\n<h2>Interesting Findings</h2>\n<p>When optimize the task of orders, </p>\n<ol>\n<li>orders_prob = 0.7* orders_prob+0.2* carts_prob+0.1* click_prob, brings 0.0004 boost on orders' local cv(evaluation on the same session of three task).</li>\n<li>My teammate rib concats carts and orders' training data, and add one feature(0, 1) whether the sample comes from carts or not, can get 0.0003 boost on orders' score.</li>\n</ol>\n<h2>Enemble</h2>\n<p>For final submission, we used voting to enemble with Rib and ria's result(different candidats, LB: 0.601), LB goes to 0.602.</p>",
      "rawMarkdown": "Thanks to OTTO and kaggle hosting such a good competition. Thanks my great teammates @juzqyxs @rib @ria.\n## Validation ##\nuse the sessions in last week.\n\n## Retrieval ##\ntwo traditional methods,\n- u2i: the aids user action before\n- i2i: itemcf, we do it seperately for each task. And we used some kinds of weight to optimize its performance, such as \"time_gap\", \"loc_gap\", \"type_difference\" between two aids.\n\n## Ranking ##\n- Features\n1. session based: count, last action time_gap(click/carts/orders), last action type \n2. aid based: count, ratio(ratio of click, ratio of carts to orders...), time\n3. session+aid based: count, last action time_gap(click/carts/orders), last action type\n4. w2v embeddings: difference and similarity of candiates between sessions' history aids, then use mean\\max\\min\\last\\std\n5. collaborative filtering score: collaborative filtering score of i2i, calculate the items between the candidates and user's historical aids.\n6. count and ratio features of aids after/before diffrent windows(1day, 3day) of each sessions' last time\n\n- Data Augment\n1. use more weeks to train, we used 3 weeks for finnal submission.\n2. use different seeds to generate more traindata(different cutoff of sessions), we used 5 seeds at last. \nthose bring 0.001 boost on local cv.\n\n- Model\n1. lightgbm binary classifier, learning rate: 0.05, 1000 rounds\n2. catboost binary classifier, learning rate: 0.05, 6000 rounds\nenemble these two results scored LB:0.600.\n\n## Interesting Findings ##\nWhen optimize the task of orders, \n1. orders_prob = 0.7* orders_prob+0.2* carts_prob+0.1* click_prob, brings 0.0004 boost on orders' local cv(evaluation on the same session of three task).\n2. My teammate rib concats carts and orders' training data, and add one feature(0, 1) whether the sample comes from carts or not, can get 0.0003 boost on orders' score.\n\n## Enemble ##\nFor final submission, we used voting to enemble with Rib and ria's result(different candidats, LB: 0.601), LB goes to 0.602.",
      "votes": null
    },
    {
      "id": "2128108",
      "postDate": "02/03/2023 13:20:21",
      "content": "<p>GM!GM!GM!😋</p>",
      "rawMarkdown": "GM!GM!GM!😋",
      "votes": null
    },
    {
      "id": "2128227",
      "postDate": "02/03/2023 14:37:35",
      "content": "<p>It's time!✌️✌️✌️</p>",
      "rawMarkdown": "It's time!✌️✌️✌️",
      "votes": null
    },
    {
      "id": "2128734",
      "postDate": "02/04/2023 01:58:32",
      "content": "<p>Congrat! GM finally! 👍 </p>",
      "rawMarkdown": "Congrat! GM finally! 👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2128108,
      "author_name": "hydantess",
      "author_url": "",
      "post_date": "02/03/2023 13:20:21",
      "content": "<p>GM!GM!GM!😋</p>",
      "votes": null,
      "replies": [
        {
          "id": 2128227,
          "author_name": "chenxin1991",
          "author_url": "",
          "post_date": "02/03/2023 14:37:35",
          "content": "<p>It's time!✌️✌️✌️</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2128734,
      "author_name": "hengzheng",
      "author_url": "",
      "post_date": "02/04/2023 01:58:32",
      "content": "<p>Congrat! GM finally! 👍 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2128056": "Thanks to OTTO and kaggle hosting such a good competition. Thanks my great teammates @juzqyxs @rib @ria.\n## Validation ##\nuse the sessions in last week.\n\n## Retrieval ##\ntwo traditional methods,\n- u2i: the aids user action before\n- i2i: itemcf, we do it seperately for each task. And we used some kinds of weight to optimize its performance, such as \"time_gap\", \"loc_gap\", \"type_difference\" between two aids.\n\n## Ranking ##\n- Features\n1. session based: count, last action time_gap(click/carts/orders), last action type \n2. aid based: count, ratio(ratio of click, ratio of carts to orders...), time\n3. session+aid based: count, last action time_gap(click/carts/orders), last action type\n4. w2v embeddings: difference and similarity of candiates between sessions' history aids, then use mean\\max\\min\\last\\std\n5. collaborative filtering score: collaborative filtering score of i2i, calculate the items between the candidates and user's historical aids.\n6. count and ratio features of aids after/before diffrent windows(1day, 3day) of each sessions' last time\n\n- Data Augment\n1. use more weeks to train, we used 3 weeks for finnal submission.\n2. use different seeds to generate more traindata(different cutoff of sessions), we used 5 seeds at last. \nthose bring 0.001 boost on local cv.\n\n- Model\n1. lightgbm binary classifier, learning rate: 0.05, 1000 rounds\n2. catboost binary classifier, learning rate: 0.05, 6000 rounds\nenemble these two results scored LB:0.600.\n\n## Interesting Findings ##\nWhen optimize the task of orders, \n1. orders_prob = 0.7* orders_prob+0.2* carts_prob+0.1* click_prob, brings 0.0004 boost on orders' local cv(evaluation on the same session of three task).\n2. My teammate rib concats carts and orders' training data, and add one feature(0, 1) whether the sample comes from carts or not, can get 0.0003 boost on orders' score.\n\n## Enemble ##\nFor final submission, we used voting to enemble with Rib and ria's result(different candidats, LB: 0.601), LB goes to 0.602.",
    "2128108": "GM!GM!GM!😋",
    "2128227": "It's time!✌️✌️✌️",
    "2128734": "Congrat! GM finally! 👍"
  },
  "source": "meta"
}