{
  "id": 382802,
  "title": "5th place Solution",
  "url": "/competitions/otto-recommender-system/writeups/nikhilmishra-6th-place-solution",
  "author_name": "",
  "post_date": "2023-02-01T04:24:53.143Z",
  "votes": 103,
  "comment_count": 32,
  "views": 0,
  "content": "<p>Firstly thanks to everyone for sharing so much on this problem. I learned a lot from all of you.</p>\n<p>Espsecially thanking (in alphabetical order)</p>\n<ol>\n<li><p>Carno: For sharing your numba pipeline. All of my candidate generation and many of my features were created using numba. I had almost never used numba before, so learnt a lot of numba in this comp.</p></li>\n<li><p>Chris: For sharing so much from comp start to end. I think every competitor owes you for that. I think you answered questions about the ranking model till the last day of the competition, so nothing but respect.</p></li>\n<li><p>Radek: For introducing us to polars, the speed of polars while joining tables really helped speed up my experiments</p></li>\n<li><p>Senkin: For your 1st place solution in H &amp; M. Many of my ideas in otto were inspired by this.</p></li>\n</ol>\n<p>My best performing solution on the private LB was almost single model (scoring almost the same as the ensemble), so will describe the single model (public LB: 0.604 and private 0.603)</p>\n<p><strong>Candiates Genreration(Numba) -&gt; Feature Creation (Numba , Polars) -&gt; Ranking  Model (Lightgbm) -&gt; Inference (Treelite)</strong></p>\n<h4>Candidates Generation</h4>\n<p>I think having a strong candidate generation method helped me a lot, so here's how I did it.</p>\n<p><strong>Number of Candidates Per Session:</strong> I generated 80 candidates most of the comp, and then jumped it up to 120 cands at the last week for some score boost (of around 0.001). I had a really decent max recall of 0.648 (on the validation set) for 80 candidates. I also tried 200 candidates in one experiment, but that did not help with the score.</p>\n<p>Also if I take the first 20 candidates from my candidates generation model, my score on LB would be 0.585. Th</p>\n<p>For candidates generation I used something similar to covisit matrices, I divided the user actions in a session for any given 2 aids, into various categories like</p>\n<p>a. Any action to Any action<br>\nb. Click to cart<br>\nc. Cart to order<br>\nd. Order to order,<br>\n… etc</p>\n<p>To keep memory usage low, I chose only the top (k * 100) candidates. K here is the number of candidates I wanted to generate</p>\n<p>Also, I normalized the weight by the frequency of the first item. So thinking its  basically like out of 100 times that milk was purchased, how many times were eggs purchased with it.</p>\n<p>The weight in the matrices was normalized by the number of items visited between the 2 aids we are talking about.</p>\n<p>Let's say if we have 5 aids, aids1, aids2,  aids3, aids4, aids5</p>\n<p>Then the weight of (aid1, aid5) will be (5-1)/(frequency of aid1).</p>\n<p>Also weight of (aid5, aid1) will be wt of ((aid1, aid5))/2 (just to add something like purchase of aid5 was driven by purchase of aid1 and not the other way round)</p>\n<p>For the inference time, to decide which top k candidates should I take, I used optuna, keeping the things like weight of each covisit matrix, weight of the recency of the item, normalized overall frequency of the item  etc as a parameter.</p>\n<h4>Feature Generation</h4>\n<ul>\n<li><p>Basic Features like frequency of the item, clicks to carts ratio, etc, recency of the item visited (this helps a lot if number of candidates in session is more than 20).</p></li>\n<li><p>Association of a generated candidate to any already seen aid in the session. This could be created by using the covisit matrix weights. Going really deep into such features helped me boost my score a lot. The idea is covisit matrix could be created in different ways to establish the relationship between 2 items, for example:</p>\n<p>a. Take only the average distance (number of aids between) between 2 aids.<br>\nb. Distance could also be measured in timestamp difference.<br>\nc. Consider only candidates in  the 1st neighbourhood (immediate candidates).<br>\nd.  Consider only relationships in the last week, etc.</p></li>\n</ul>\n<h4>Training a Ranking Model:</h4>\n<p>I used lightgbm, with 5% negative sampling and around 400 features, and data for last 2 weeks.<br>\nAdding data for second last week boosted the score by about 0.0005.</p>\n<p>Some things or tricks that worked for me:</p>\n<ol>\n<li><p>Training all clicks, carts, and orders with a single model (not 3 separate models)., this locally was easily seen to be getting around 0.001 to 0.002 better score (data is grouped into sessions not into session, type).</p></li>\n<li><p>Using separate labels for clicks, carts and orders whiel  ranking, with ranking label gain as Orders (6) -&gt; Carts(3) -&gt; Clicks(1), instead of just using 1 when the user performed an action(click, carts and orders) and 0 when they did not (this boosted the score by around 0.0005).</p></li>\n<li><p>Using the ranking of the stage 1 candidate generation model, as a feature of the stage 2 model. If you think, the stage 1 model can score 0.585 on the lb, so using this ranking was the important feature of my model.</p></li>\n</ol>\n<h4>Inference:</h4>\n<p>Nothing much to say here, except that I  used treelite for inferencing to reduce the inferencing time.</p>\n<p>And finally congratulations to all the winners ! It was really fun participating.</p>\n<p>P.S: I wrote this in a hurry before starting my office work, so let me know if I messed up some details</p>",
  "messages": [
    {
      "id": "2124549",
      "postDate": "02/01/2023 04:20:55",
      "content": "<p>Firstly thanks to everyone for sharing so much on this problem. I learned a lot from all of you.</p>\n<p>Espsecially thanking (in alphabetical order)</p>\n<ol>\n<li><p>Carno: For sharing your numba pipeline. All of my candidate generation and many of my features were created using numba. I had almost never used numba before, so learnt a lot of numba in this comp.</p></li>\n<li><p>Chris: For sharing so much from comp start to end. I think every competitor owes you for that. I think you answered questions about the ranking model till the last day of the competition, so nothing but respect.</p></li>\n<li><p>Radek: For introducing us to polars, the speed of polars while joining tables really helped speed up my experiments</p></li>\n<li><p>Senkin: For your 1st place solution in H &amp; M. Many of my ideas in otto were inspired by this.</p></li>\n</ol>\n<p>My best performing solution on the private LB was almost single model (scoring almost the same as the ensemble), so will describe the single model (public LB: 0.604 and private 0.603)</p>\n<p><strong>Candiates Genreration(Numba) -&gt; Feature Creation (Numba , Polars) -&gt; Ranking  Model (Lightgbm) -&gt; Inference (Treelite)</strong></p>\n<h4>Candidates Generation</h4>\n<p>I think having a strong candidate generation method helped me a lot, so here's how I did it.</p>\n<p><strong>Number of Candidates Per Session:</strong> I generated 80 candidates most of the comp, and then jumped it up to 120 cands at the last week for some score boost (of around 0.001). I had a really decent max recall of 0.648 (on the validation set) for 80 candidates. I also tried 200 candidates in one experiment, but that did not help with the score.</p>\n<p>Also if I take the first 20 candidates from my candidates generation model, my score on LB would be 0.585. Th</p>\n<p>For candidates generation I used something similar to covisit matrices, I divided the user actions in a session for any given 2 aids, into various categories like</p>\n<p>a. Any action to Any action<br>\nb. Click to cart<br>\nc. Cart to order<br>\nd. Order to order,<br>\n… etc</p>\n<p>To keep memory usage low, I chose only the top (k * 100) candidates. K here is the number of candidates I wanted to generate</p>\n<p>Also, I normalized the weight by the frequency of the first item. So thinking its  basically like out of 100 times that milk was purchased, how many times were eggs purchased with it.</p>\n<p>The weight in the matrices was normalized by the number of items visited between the 2 aids we are talking about.</p>\n<p>Let's say if we have 5 aids, aids1, aids2,  aids3, aids4, aids5</p>\n<p>Then the weight of (aid1, aid5) will be (5-1)/(frequency of aid1).</p>\n<p>Also weight of (aid5, aid1) will be wt of ((aid1, aid5))/2 (just to add something like purchase of aid5 was driven by purchase of aid1 and not the other way round)</p>\n<p>For the inference time, to decide which top k candidates should I take, I used optuna, keeping the things like weight of each covisit matrix, weight of the recency of the item, normalized overall frequency of the item  etc as a parameter.</p>\n<h4>Feature Generation</h4>\n<ul>\n<li><p>Basic Features like frequency of the item, clicks to carts ratio, etc, recency of the item visited (this helps a lot if number of candidates in session is more than 20).</p></li>\n<li><p>Association of a generated candidate to any already seen aid in the session. This could be created by using the covisit matrix weights. Going really deep into such features helped me boost my score a lot. The idea is covisit matrix could be created in different ways to establish the relationship between 2 items, for example:</p>\n<p>a. Take only the average distance (number of aids between) between 2 aids.<br>\nb. Distance could also be measured in timestamp difference.<br>\nc. Consider only candidates in  the 1st neighbourhood (immediate candidates).<br>\nd.  Consider only relationships in the last week, etc.</p></li>\n</ul>\n<h4>Training a Ranking Model:</h4>\n<p>I used lightgbm, with 5% negative sampling and around 400 features, and data for last 2 weeks.<br>\nAdding data for second last week boosted the score by about 0.0005.</p>\n<p>Some things or tricks that worked for me:</p>\n<ol>\n<li><p>Training all clicks, carts, and orders with a single model (not 3 separate models)., this locally was easily seen to be getting around 0.001 to 0.002 better score (data is grouped into sessions not into session, type).</p></li>\n<li><p>Using separate labels for clicks, carts and orders whiel  ranking, with ranking label gain as Orders (6) -&gt; Carts(3) -&gt; Clicks(1), instead of just using 1 when the user performed an action(click, carts and orders) and 0 when they did not (this boosted the score by around 0.0005).</p></li>\n<li><p>Using the ranking of the stage 1 candidate generation model, as a feature of the stage 2 model. If you think, the stage 1 model can score 0.585 on the lb, so using this ranking was the important feature of my model.</p></li>\n</ol>\n<h4>Inference:</h4>\n<p>Nothing much to say here, except that I  used treelite for inferencing to reduce the inferencing time.</p>\n<p>And finally congratulations to all the winners ! It was really fun participating.</p>\n<p>P.S: I wrote this in a hurry before starting my office work, so let me know if I messed up some details</p>",
      "rawMarkdown": "Firstly thanks to everyone for sharing so much on this problem. I learned a lot from all of you.\n\nEspsecially thanking (in alphabetical order)\n\n1. Carno: For sharing your numba pipeline. All of my candidate generation and many of my features were created using numba. I had almost never used numba before, so learnt a lot of numba in this comp.\n\n2. Chris: For sharing so much from comp start to end. I think every competitor owes you for that. I think you answered questions about the ranking model till the last day of the competition, so nothing but respect.\n\n3. Radek: For introducing us to polars, the speed of polars while joining tables really helped speed up my experiments\n\n4. Senkin: For your 1st place solution in H & M. Many of my ideas in otto were inspired by this.\n\nMy best performing solution on the private LB was almost single model (scoring almost the same as the ensemble), so will describe the single model (public LB: 0.604 and private 0.603)\n\n**Candiates Genreration(Numba) -> Feature Creation (Numba , Polars) -> Ranking  Model (Lightgbm) -> Inference (Treelite)**\n\n#### Candidates Generation\n\nI think having a strong candidate generation method helped me a lot, so here's how I did it.\n\n**Number of Candidates Per Session:** I generated 80 candidates most of the comp, and then jumped it up to 120 cands at the last week for some score boost (of around 0.001). I had a really decent max recall of 0.648 (on the validation set) for 80 candidates. I also tried 200 candidates in one experiment, but that did not help with the score.\n\nAlso if I take the first 20 candidates from my candidates generation model, my score on LB would be 0.585. Th\n\nFor candidates generation I used something similar to covisit matrices, I divided the user actions in a session for any given 2 aids, into various categories like\n\na. Any action to Any action\nb. Click to cart\nc. Cart to order\nd. Order to order,\n... etc\n\nTo keep memory usage low, I chose only the top (k * 100) candidates. K here is the number of candidates I wanted to generate\n\nAlso, I normalized the weight by the frequency of the first item. So thinking its  basically like out of 100 times that milk was purchased, how many times were eggs purchased with it.\n\nThe weight in the matrices was normalized by the number of items visited between the 2 aids we are talking about.\n\nLet's say if we have 5 aids, aids1, aids2,  aids3, aids4, aids5\n\nThen the weight of (aid1, aid5) will be (5-1)/(frequency of aid1).\n\nAlso weight of (aid5, aid1) will be wt of ((aid1, aid5))/2 (just to add something like purchase of aid5 was driven by purchase of aid1 and not the other way round)\n\nFor the inference time, to decide which top k candidates should I take, I used optuna, keeping the things like weight of each covisit matrix, weight of the recency of the item, normalized overall frequency of the item  etc as a parameter.\n\n#### Feature Generation\n\n* Basic Features like frequency of the item, clicks to carts ratio, etc, recency of the item visited (this helps a lot if number of candidates in session is more than 20).\n* Association of a generated candidate to any already seen aid in the session. This could be created by using the covisit matrix weights. Going really deep into such features helped me boost my score a lot. The idea is covisit matrix could be created in different ways to establish the relationship between 2 items, for example:\n\n  a. Take only the average distance (number of aids between) between 2 aids.\n  b. Distance could also be measured in timestamp difference.\n  c. Consider only candidates in  the 1st neighbourhood (immediate candidates).\n  d.  Consider only relationships in the last week, etc.\n\n\n#### Training a Ranking Model:\n\nI used lightgbm, with 5% negative sampling and around 400 features, and data for last 2 weeks.\nAdding data for second last week boosted the score by about 0.0005.\n\nSome things or tricks that worked for me:\n\n1. Training all clicks, carts, and orders with a single model (not 3 separate models)., this locally was easily seen to be getting around 0.001 to 0.002 better score (data is grouped into sessions not into session, type).\n\n2. Using separate labels for clicks, carts and orders whiel  ranking, with ranking label gain as Orders (6) -> Carts(3) -> Clicks(1), instead of just using 1 when the user performed an action(click, carts and orders) and 0 when they did not (this boosted the score by around 0.0005).\n\n3. Using the ranking of the stage 1 candidate generation model, as a feature of the stage 2 model. If you think, the stage 1 model can score 0.585 on the lb, so using this ranking was the important feature of my model.\n\n#### Inference:\nNothing much to say here, except that I  used treelite for inferencing to reduce the inferencing time.\n\nAnd finally congratulations to all the winners ! It was really fun participating.\n\nP.S: I wrote this in a hurry before starting my office work, so let me know if I messed up some details",
      "votes": null
    },
    {
      "id": "2124552",
      "postDate": "02/01/2023 04:23:41",
      "content": "<p>Congrats! 😄 Really liked the single model approach.</p>",
      "rawMarkdown": "Congrats! 😄 Really liked the single model approach.",
      "votes": null
    },
    {
      "id": "2124573",
      "postDate": "02/01/2023 04:45:15",
      "content": "<p>congratulations on a great finish too :D !</p>",
      "rawMarkdown": "congratulations on a great finish too :D !",
      "votes": null
    },
    {
      "id": "2124576",
      "postDate": "02/01/2023 04:49:01",
      "content": "<p>Congrats on the solo gold. </p>",
      "rawMarkdown": "Congrats on the solo gold.",
      "votes": null
    },
    {
      "id": "2124581",
      "postDate": "02/01/2023 04:53:31",
      "content": "<p>Thanks for sharing!</p>\n<p>I have a question.</p>\n<p><code>a. Any action to Any action\nb. Click to cart\nc. Cart to order\nd. Order to order ...</code></p>\n<p>In the above case then, do you have 4 or more types of covisitation matrix?</p>",
      "rawMarkdown": "Thanks for sharing!\n\nI have a question.\n\n```a. Any action to Any action\nb. Click to cart\nc. Cart to order\nd. Order to order ...```\n\nIn the above case then, do you have 4 or more types of covisitation matrix?",
      "votes": null
    },
    {
      "id": "2124590",
      "postDate": "02/01/2023 04:57:55",
      "content": "<p>firstly congratulations for the great result. Yes infact if you see total possible combinations will be  number of possible ordered pairs of clicks carts and orders plus one of (anyaction to any action).</p>\n<p>Also you can add something like clicks to (carts, orders), which I did not :D</p>",
      "rawMarkdown": "firstly congratulations for the great result. Yes infact if you see total possible combinations will be  number of possible ordered pairs of clicks carts and orders plus one of (anyaction to any action).\n\nAlso you can add something like clicks to (carts, orders), which I did not :D",
      "votes": null
    },
    {
      "id": "2124593",
      "postDate": "02/01/2023 05:01:34",
      "content": "<p>Thanks.</p>\n<p>Congrats on the results. I really like your detailed idea of the co-visiting matrix.</p>",
      "rawMarkdown": "Thanks.\n\nCongrats on the results. I really like your detailed idea of the co-visiting matrix.",
      "votes": null
    },
    {
      "id": "2124594",
      "postDate": "02/01/2023 05:02:01",
      "content": "<p>thanks a lot</p>",
      "rawMarkdown": "thanks a lot",
      "votes": null
    },
    {
      "id": "2124606",
      "postDate": "02/01/2023 05:09:04",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> for the great result. And thank you for sharing this great write-up! </p>",
      "rawMarkdown": "Congratulations @nikhilmishradev for the great result. And thank you for sharing this great write-up!",
      "votes": null
    },
    {
      "id": "2124638",
      "postDate": "02/01/2023 05:40:56",
      "content": "<p>Congrats! I also had 0.648 theoretical max recall with 100 candidates but my model wasn't learning anything. I guess your candidates were much better and positives were located in higher positions. High quality candidate generation was the key. </p>",
      "rawMarkdown": "Congrats! I also had 0.648 theoretical max recall with 100 candidates but my model wasn't learning anything. I guess your candidates were much better and positives were located in higher positions. High quality candidate generation was the key.",
      "votes": null
    },
    {
      "id": "2124667",
      "postDate": "02/01/2023 06:11:26",
      "content": "<p>I had 0.648 with 80 cands, with 120, I don't remember sorry :D. The most important way to boost the score was finding the association between the generated candidates and the already seen candidates in the sessions where num candidates &lt; 20. So the key for me was high quality feature generation for the ranker</p>",
      "rawMarkdown": "I had 0.648 with 80 cands, with 120, I don't remember sorry :D. The most important way to boost the score was finding the association between the generated candidates and the already seen candidates in the sessions where num candidates < 20. So the key for me was high quality feature generation for the ranker",
      "votes": null
    },
    {
      "id": "2124698",
      "postDate": "02/01/2023 06:47:20",
      "content": "<p>Congrats Nikhil! BTW are you planning to make the code open source?</p>",
      "rawMarkdown": "Congrats Nikhil! BTW are you planning to make the code open source?",
      "votes": null
    },
    {
      "id": "2124703",
      "postDate": "02/01/2023 06:52:47",
      "content": "<p>Thanks bro. Only if kaggle gives me a way to block you 😂</p>",
      "rawMarkdown": "Thanks bro. Only if kaggle gives me a way to block you 😂",
      "votes": null
    },
    {
      "id": "2124707",
      "postDate": "02/01/2023 06:56:47",
      "content": "<p>Congrats on the solo gold <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> . Inspiring.</p>",
      "rawMarkdown": "Congrats on the solo gold @nikhilmishradev . Inspiring.",
      "votes": null
    },
    {
      "id": "2124752",
      "postDate": "02/01/2023 07:46:56",
      "content": "<p>Nice explication. Well done. Congrats.</p>",
      "rawMarkdown": "Nice explication. Well done. Congrats.",
      "votes": null
    },
    {
      "id": "2125300",
      "postDate": "02/01/2023 15:25:08",
      "content": "<p>Congrats !!!  Interesting, you just train your model on last 2 weeks of data in test.csv ?? </p>",
      "rawMarkdown": "Congrats !!!  Interesting, you just train your model on last 2 weeks of data in test.csv ??",
      "votes": null
    },
    {
      "id": "2125317",
      "postDate": "02/01/2023 15:31:14",
      "content": "<p>good job, dude</p>",
      "rawMarkdown": "good job, dude",
      "votes": null
    },
    {
      "id": "2125429",
      "postDate": "02/01/2023 16:47:43",
      "content": "<p><a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> Congrats! Thanks for the tips on feature engineering!</p>",
      "rawMarkdown": "nikhilmishradev Congrats! Thanks for the tips on feature engineering!",
      "votes": null
    },
    {
      "id": "2125868",
      "postDate": "02/02/2023 00:33:28",
      "content": "<p>Congrats on your solo gold!<br>\n About tips 1&amp;2 of ranking model, I have similar strategy but only merging carts and orders. It seems useful to merge them 3 all. At least, you only need to train one model, saving much time.<br>\nThanks for your sharing!</p>",
      "rawMarkdown": "Congrats on your solo gold!\n About tips 1&2 of ranking model, I have similar strategy but only merging carts and orders. It seems useful to merge them 3 all. At least, you only need to train one model, saving much time.\nThanks for your sharing!",
      "votes": null
    },
    {
      "id": "2125909",
      "postDate": "02/02/2023 01:45:53",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> . It's amazing that you did so well solo. You are the 2nd best solo team. Very impressive!</p>",
      "rawMarkdown": "Congratulations @nikhilmishradev . It's amazing that you did so well solo. You are the 2nd best solo team. Very impressive!",
      "votes": null
    },
    {
      "id": "2126229",
      "postDate": "02/02/2023 06:17:36",
      "content": "<p>thanks a lot and congratulations <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for the wonderful result, as well as for your inspiring contributions to the community</p>",
      "rawMarkdown": "thanks a lot and congratulations @cdeotte for the wonderful result, as well as for your inspiring contributions to the community",
      "votes": null
    },
    {
      "id": "2126237",
      "postDate": "02/02/2023 06:20:30",
      "content": "<p>yes training a single model was a much lesser headache. Congratulations on the back to back great results on recommender systems</p>",
      "rawMarkdown": "yes training a single model was a much lesser headache. Congratulations on the back to back great results on recommender systems",
      "votes": null
    },
    {
      "id": "2126238",
      "postDate": "02/02/2023 06:21:44",
      "content": "<p>sorry for the confusion, when I say last 2 weeks  I mean last 2 weeks of the training data not the test data.</p>",
      "rawMarkdown": "sorry for the confusion, when I say last 2 weeks  I mean last 2 weeks of the training data not the test data.",
      "votes": null
    },
    {
      "id": "2126248",
      "postDate": "02/02/2023 06:30:24",
      "content": "<p>Good job! Nice!</p>",
      "rawMarkdown": "Good job! Nice!",
      "votes": null
    },
    {
      "id": "2126983",
      "postDate": "02/02/2023 15:48:32",
      "content": "<p>Congrats on the solo gold <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a>! Very interesting solution. Thanks for the write up!</p>",
      "rawMarkdown": "Congrats on the solo gold @nikhilmishradev! Very interesting solution. Thanks for the write up!",
      "votes": null
    },
    {
      "id": "2128013",
      "postDate": "02/03/2023 12:06:25",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> on solo gold!</p>",
      "rawMarkdown": "Congratulations @nikhilmishradev on solo gold!",
      "votes": null
    },
    {
      "id": "2128037",
      "postDate": "02/03/2023 12:19:54",
      "content": "<p>Congrats on solo gold on hard competition <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> </p>",
      "rawMarkdown": "Congrats on solo gold on hard competition @nikhilmishradev",
      "votes": null
    },
    {
      "id": "2128099",
      "postDate": "02/03/2023 13:16:32",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a>, always my inspiration. Waiting for you to get the final gold</p>",
      "rawMarkdown": "Thank you @tezdhar, always my inspiration. Waiting for you to get the final gold",
      "votes": null
    },
    {
      "id": "2128100",
      "postDate": "02/03/2023 13:16:44",
      "content": "<p>thank you <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> </p>",
      "rawMarkdown": "thank you @duykhanh99",
      "votes": null
    },
    {
      "id": "2130338",
      "postDate": "02/05/2023 11:42:04",
      "content": "<p>Amazing work <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> !</p>",
      "rawMarkdown": "Amazing work @nikhilmishradev !",
      "votes": null
    },
    {
      "id": "2130793",
      "postDate": "02/05/2023 17:32:03",
      "content": "<p>Thanks for the nice job!</p>",
      "rawMarkdown": "Thanks for the nice job!",
      "votes": null
    },
    {
      "id": "2133942",
      "postDate": "02/07/2023 17:15:45",
      "content": "<p>Thanks for sharing!Awesome work!</p>",
      "rawMarkdown": "Thanks for sharing!Awesome work!",
      "votes": null
    },
    {
      "id": "2136236",
      "postDate": "02/09/2023 07:24:44",
      "content": "<p>Hi Nikhil<br>\nthanks for write up. Are you planning to share the code as well?</p>",
      "rawMarkdown": "Hi Nikhil\nthanks for write up. Are you planning to share the code as well?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2124552,
      "author_name": "nlztrk",
      "author_url": "",
      "post_date": "02/01/2023 04:23:41",
      "content": "<p>Congrats! 😄 Really liked the single model approach.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2124573,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "02/01/2023 04:45:15",
          "content": "<p>congratulations on a great finish too :D !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2124576,
      "author_name": "kaggleqrdl",
      "author_url": "",
      "post_date": "02/01/2023 04:49:01",
      "content": "<p>Congrats on the solo gold. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2124581,
      "author_name": "songwonho",
      "author_url": "",
      "post_date": "02/01/2023 04:53:31",
      "content": "<p>Thanks for sharing!</p>\n<p>I have a question.</p>\n<p><code>a. Any action to Any action\nb. Click to cart\nc. Cart to order\nd. Order to order ...</code></p>\n<p>In the above case then, do you have 4 or more types of covisitation matrix?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2124590,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "02/01/2023 04:57:55",
          "content": "<p>firstly congratulations for the great result. Yes infact if you see total possible combinations will be  number of possible ordered pairs of clicks carts and orders plus one of (anyaction to any action).</p>\n<p>Also you can add something like clicks to (carts, orders), which I did not :D</p>",
          "votes": null,
          "replies": [
            {
              "id": 2124593,
              "author_name": "songwonho",
              "author_url": "",
              "post_date": "02/01/2023 05:01:34",
              "content": "<p>Thanks.</p>\n<p>Congrats on the results. I really like your detailed idea of the co-visiting matrix.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2124594,
                  "author_name": "nikhilmishradev",
                  "author_url": "",
                  "post_date": "02/01/2023 05:02:01",
                  "content": "<p>thanks a lot</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2124606,
      "author_name": "tahamhaider",
      "author_url": "",
      "post_date": "02/01/2023 05:09:04",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> for the great result. And thank you for sharing this great write-up! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2124638,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "02/01/2023 05:40:56",
      "content": "<p>Congrats! I also had 0.648 theoretical max recall with 100 candidates but my model wasn't learning anything. I guess your candidates were much better and positives were located in higher positions. High quality candidate generation was the key. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2124667,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "02/01/2023 06:11:26",
          "content": "<p>I had 0.648 with 80 cands, with 120, I don't remember sorry :D. The most important way to boost the score was finding the association between the generated candidates and the already seen candidates in the sessions where num candidates &lt; 20. So the key for me was high quality feature generation for the ranker</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2124698,
      "author_name": "himanshupoddar",
      "author_url": "",
      "post_date": "02/01/2023 06:47:20",
      "content": "<p>Congrats Nikhil! BTW are you planning to make the code open source?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2124703,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "02/01/2023 06:52:47",
          "content": "<p>Thanks bro. Only if kaggle gives me a way to block you 😂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2124707,
      "author_name": "manthanbhagat",
      "author_url": "",
      "post_date": "02/01/2023 06:56:47",
      "content": "<p>Congrats on the solo gold <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> . Inspiring.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2124752,
      "author_name": "santiagomota",
      "author_url": "",
      "post_date": "02/01/2023 07:46:56",
      "content": "<p>Nice explication. Well done. Congrats.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2125300,
      "author_name": "weiqiu",
      "author_url": "",
      "post_date": "02/01/2023 15:25:08",
      "content": "<p>Congrats !!!  Interesting, you just train your model on last 2 weeks of data in test.csv ?? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2126238,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "02/02/2023 06:21:44",
          "content": "<p>sorry for the confusion, when I say last 2 weeks  I mean last 2 weeks of the training data not the test data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2125317,
      "author_name": "artyomkhrennikov",
      "author_url": "",
      "post_date": "02/01/2023 15:31:14",
      "content": "<p>good job, dude</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2125429,
      "author_name": "minhbtnguyen",
      "author_url": "",
      "post_date": "02/01/2023 16:47:43",
      "content": "<p><a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> Congrats! Thanks for the tips on feature engineering!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2125868,
      "author_name": "sirius81",
      "author_url": "",
      "post_date": "02/02/2023 00:33:28",
      "content": "<p>Congrats on your solo gold!<br>\n About tips 1&amp;2 of ranking model, I have similar strategy but only merging carts and orders. It seems useful to merge them 3 all. At least, you only need to train one model, saving much time.<br>\nThanks for your sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2126237,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "02/02/2023 06:20:30",
          "content": "<p>yes training a single model was a much lesser headache. Congratulations on the back to back great results on recommender systems</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2125909,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/02/2023 01:45:53",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> . It's amazing that you did so well solo. You are the 2nd best solo team. Very impressive!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2126229,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "02/02/2023 06:17:36",
          "content": "<p>thanks a lot and congratulations <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for the wonderful result, as well as for your inspiring contributions to the community</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2126248,
      "author_name": "eldarazamatov",
      "author_url": "",
      "post_date": "02/02/2023 06:30:24",
      "content": "<p>Good job! Nice!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2126983,
      "author_name": "ravishah1",
      "author_url": "",
      "post_date": "02/02/2023 15:48:32",
      "content": "<p>Congrats on the solo gold <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a>! Very interesting solution. Thanks for the write up!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2128013,
      "author_name": "tezdhar",
      "author_url": "",
      "post_date": "02/03/2023 12:06:25",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> on solo gold!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2128099,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "02/03/2023 13:16:32",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a>, always my inspiration. Waiting for you to get the final gold</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2128037,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/03/2023 12:19:54",
      "content": "<p>Congrats on solo gold on hard competition <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2128100,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "02/03/2023 13:16:44",
          "content": "<p>thank you <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2130338,
      "author_name": "sohilsharma1996",
      "author_url": "",
      "post_date": "02/05/2023 11:42:04",
      "content": "<p>Amazing work <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2130793,
      "author_name": "qinzhida",
      "author_url": "",
      "post_date": "02/05/2023 17:32:03",
      "content": "<p>Thanks for the nice job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2133942,
      "author_name": "",
      "author_url": "",
      "post_date": "02/07/2023 17:15:45",
      "content": "<p>Thanks for sharing!Awesome work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2136236,
      "author_name": "purohit",
      "author_url": "",
      "post_date": "02/09/2023 07:24:44",
      "content": "<p>Hi Nikhil<br>\nthanks for write up. Are you planning to share the code as well?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2124549": "Firstly thanks to everyone for sharing so much on this problem. I learned a lot from all of you.\n\nEspsecially thanking (in alphabetical order)\n\n1. Carno: For sharing your numba pipeline. All of my candidate generation and many of my features were created using numba. I had almost never used numba before, so learnt a lot of numba in this comp.\n\n2. Chris: For sharing so much from comp start to end. I think every competitor owes you for that. I think you answered questions about the ranking model till the last day of the competition, so nothing but respect.\n\n3. Radek: For introducing us to polars, the speed of polars while joining tables really helped speed up my experiments\n\n4. Senkin: For your 1st place solution in H & M. Many of my ideas in otto were inspired by this.\n\nMy best performing solution on the private LB was almost single model (scoring almost the same as the ensemble), so will describe the single model (public LB: 0.604 and private 0.603)\n\n**Candiates Genreration(Numba) -> Feature Creation (Numba , Polars) -> Ranking  Model (Lightgbm) -> Inference (Treelite)**\n\n#### Candidates Generation\n\nI think having a strong candidate generation method helped me a lot, so here's how I did it.\n\n**Number of Candidates Per Session:** I generated 80 candidates most of the comp, and then jumped it up to 120 cands at the last week for some score boost (of around 0.001). I had a really decent max recall of 0.648 (on the validation set) for 80 candidates. I also tried 200 candidates in one experiment, but that did not help with the score.\n\nAlso if I take the first 20 candidates from my candidates generation model, my score on LB would be 0.585. Th\n\nFor candidates generation I used something similar to covisit matrices, I divided the user actions in a session for any given 2 aids, into various categories like\n\na. Any action to Any action\nb. Click to cart\nc. Cart to order\nd. Order to order,\n... etc\n\nTo keep memory usage low, I chose only the top (k * 100) candidates. K here is the number of candidates I wanted to generate\n\nAlso, I normalized the weight by the frequency of the first item. So thinking its  basically like out of 100 times that milk was purchased, how many times were eggs purchased with it.\n\nThe weight in the matrices was normalized by the number of items visited between the 2 aids we are talking about.\n\nLet's say if we have 5 aids, aids1, aids2,  aids3, aids4, aids5\n\nThen the weight of (aid1, aid5) will be (5-1)/(frequency of aid1).\n\nAlso weight of (aid5, aid1) will be wt of ((aid1, aid5))/2 (just to add something like purchase of aid5 was driven by purchase of aid1 and not the other way round)\n\nFor the inference time, to decide which top k candidates should I take, I used optuna, keeping the things like weight of each covisit matrix, weight of the recency of the item, normalized overall frequency of the item  etc as a parameter.\n\n#### Feature Generation\n\n* Basic Features like frequency of the item, clicks to carts ratio, etc, recency of the item visited (this helps a lot if number of candidates in session is more than 20).\n* Association of a generated candidate to any already seen aid in the session. This could be created by using the covisit matrix weights. Going really deep into such features helped me boost my score a lot. The idea is covisit matrix could be created in different ways to establish the relationship between 2 items, for example:\n\n  a. Take only the average distance (number of aids between) between 2 aids.\n  b. Distance could also be measured in timestamp difference.\n  c. Consider only candidates in  the 1st neighbourhood (immediate candidates).\n  d.  Consider only relationships in the last week, etc.\n\n\n#### Training a Ranking Model:\n\nI used lightgbm, with 5% negative sampling and around 400 features, and data for last 2 weeks.\nAdding data for second last week boosted the score by about 0.0005.\n\nSome things or tricks that worked for me:\n\n1. Training all clicks, carts, and orders with a single model (not 3 separate models)., this locally was easily seen to be getting around 0.001 to 0.002 better score (data is grouped into sessions not into session, type).\n\n2. Using separate labels for clicks, carts and orders whiel  ranking, with ranking label gain as Orders (6) -> Carts(3) -> Clicks(1), instead of just using 1 when the user performed an action(click, carts and orders) and 0 when they did not (this boosted the score by around 0.0005).\n\n3. Using the ranking of the stage 1 candidate generation model, as a feature of the stage 2 model. If you think, the stage 1 model can score 0.585 on the lb, so using this ranking was the important feature of my model.\n\n#### Inference:\nNothing much to say here, except that I  used treelite for inferencing to reduce the inferencing time.\n\nAnd finally congratulations to all the winners ! It was really fun participating.\n\nP.S: I wrote this in a hurry before starting my office work, so let me know if I messed up some details",
    "2124552": "Congrats! 😄 Really liked the single model approach.",
    "2124573": "congratulations on a great finish too :D !",
    "2124576": "Congrats on the solo gold.",
    "2124581": "Thanks for sharing!\n\nI have a question.\n\n```a. Any action to Any action\nb. Click to cart\nc. Cart to order\nd. Order to order ...```\n\nIn the above case then, do you have 4 or more types of covisitation matrix?",
    "2124590": "firstly congratulations for the great result. Yes infact if you see total possible combinations will be  number of possible ordered pairs of clicks carts and orders plus one of (anyaction to any action).\n\nAlso you can add something like clicks to (carts, orders), which I did not :D",
    "2124593": "Thanks.\n\nCongrats on the results. I really like your detailed idea of the co-visiting matrix.",
    "2124594": "thanks a lot",
    "2124606": "Congratulations @nikhilmishradev for the great result. And thank you for sharing this great write-up!",
    "2124638": "Congrats! I also had 0.648 theoretical max recall with 100 candidates but my model wasn't learning anything. I guess your candidates were much better and positives were located in higher positions. High quality candidate generation was the key.",
    "2124667": "I had 0.648 with 80 cands, with 120, I don't remember sorry :D. The most important way to boost the score was finding the association between the generated candidates and the already seen candidates in the sessions where num candidates < 20. So the key for me was high quality feature generation for the ranker",
    "2124698": "Congrats Nikhil! BTW are you planning to make the code open source?",
    "2124703": "Thanks bro. Only if kaggle gives me a way to block you 😂",
    "2124707": "Congrats on the solo gold @nikhilmishradev . Inspiring.",
    "2124752": "Nice explication. Well done. Congrats.",
    "2125300": "Congrats !!!  Interesting, you just train your model on last 2 weeks of data in test.csv ??",
    "2125317": "good job, dude",
    "2125429": "nikhilmishradev Congrats! Thanks for the tips on feature engineering!",
    "2125868": "Congrats on your solo gold!\n About tips 1&2 of ranking model, I have similar strategy but only merging carts and orders. It seems useful to merge them 3 all. At least, you only need to train one model, saving much time.\nThanks for your sharing!",
    "2125909": "Congratulations @nikhilmishradev . It's amazing that you did so well solo. You are the 2nd best solo team. Very impressive!",
    "2126229": "thanks a lot and congratulations @cdeotte for the wonderful result, as well as for your inspiring contributions to the community",
    "2126237": "yes training a single model was a much lesser headache. Congratulations on the back to back great results on recommender systems",
    "2126238": "sorry for the confusion, when I say last 2 weeks  I mean last 2 weeks of the training data not the test data.",
    "2126248": "Good job! Nice!",
    "2126983": "Congrats on the solo gold @nikhilmishradev! Very interesting solution. Thanks for the write up!",
    "2128013": "Congratulations @nikhilmishradev on solo gold!",
    "2128037": "Congrats on solo gold on hard competition @nikhilmishradev",
    "2128099": "Thank you @tezdhar, always my inspiration. Waiting for you to get the final gold",
    "2128100": "thank you @duykhanh99",
    "2130338": "Amazing work @nikhilmishradev !",
    "2130793": "Thanks for the nice job!",
    "2133942": "Thanks for sharing!Awesome work!",
    "2136236": "Hi Nikhil\nthanks for write up. Are you planning to share the code as well?"
  },
  "source": "meta"
}