{
  "id": 382897,
  "title": "69th story and solution",
  "url": "/competitions/otto-recommender-system/discussion/382897",
  "author_name": "",
  "post_date": "2023-02-01T12:28:37.159035200Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi there!<br>\nI worked solo and got to 69th position.</p>\n<p>My solution was similar in its structure to what most users here used. First - candidate generation, then - reranking using GBDT (LGBM and catboost tor orders, only LGBM for the rest). Everything was done on kaggle, had to use many notebooks to calculate everything, some of them ran for 4-6 hours so I had to wait overnight for the results. Had to re-write parts of my feature generation code for polars, as pandas is just too slow.<br>\nI guess it was a big mistake to start building w2vec features only during last week and a half. Probably, could have done much more if I had time to make more experiments with w2vec and probably also try doc2vec.</p>\n<p>I wanted to use average for results for diffents models. So, for orders I used 2 cross-validation sets. I generated 2 different sets of truncked sessions with different random seeds and built 2 different sets of features for them. Then I built LGBM model on 1 set of features and catboost model on another one, then used both models to make predictions on test data and averaged the prediction. The result was not much better than result of a single LGBM model, unfortunately.</p>\n<p>Used only co-visitation matrixes for candidate generation, and I see many people were able to get better recall results for top20 candidates. My numbers are:<br>\nclicks  52.72%<br>\ncarts:  40.63%<br>\norders: 64.71%<br>\nFor top 50 candidates it was:<br>\nclicks  60.43%<br>\ncarts:  45.21%<br>\norders: 67.86%<br>\nAt the last week I tried to switch to top75 candidates, and was suprised to see very close result to what I got for 50 candidates. The model was able to select almost no candidates from those additional coming from top75 candidates.<br>\nMy final cross-validation rates after re-ranking:<br>\nclicks: 54.54%<br>\ncarts: 42.08%<br>\norders: 65.80%</p>\n<h4>Clicks:</h4>\n<ul>\n<li><strong>n</strong> - 0 for the last viewed aid, 1 for aid last viewed before, e.t.c, 125 for aids never viewed</li>\n<li><strong>time_delta</strong> - time in seconds from a moment when aid was last viewed to the last action in session</li>\n<li><strong>type</strong> - 1 for aids ever put in cart, 2 for aids ever ordered, 0 for the rest</li>\n<li><strong>count_views</strong> - number of interactions with aid in a session</li>\n<li><strong>ts_diff</strong> - time in seconds between last event and event before last</li>\n<li><strong>time_viewed</strong> - time from user's click on aid untill next event clipped to 180 seconds and then summed for all views for the aid</li>\n<li><strong>daily_aid_count</strong> - normalized count of events with aid for the previous day</li>\n<li><strong>same_day_aid_count</strong> - normalized count of events with aid for the day</li>\n<li><strong>aid_count_weekly</strong> - normalized count of events with aid for the week</li>\n<li><strong>wgt_last</strong> - made a special co-visitation matrix to calculate only the very next aid after previous, applied it to the last aid session.</li>\n<li><strong>wgt_before_last</strong> - same matrix applied to aid before the last one</li>\n<li><strong>time_viewed_clipped</strong> - average time aid viewed in all interactions (before averaging values were clipped to 180)</li>\n<li><strong>aid_counts</strong> - total interactions with aid</li>\n<li><strong>is_promotion_day</strong> - checked test-period sessions for likely promoted aids. Should have probably removed that feature, bu finally didn't manage to find time to check results without it.</li>\n<li><strong>wgt_matrix</strong> - sum of co-visitaion matrix weights for the last 5 aids (same co-visitaion matrix that was used for candidate generation)</li>\n<li><strong>wgt_exp</strong> - sum of co-visitaion matrix weights for the last 10 aids normalized by n (divided weight by 1 for the last aid, by 2 for aid before it, then by 3 and so on)</li>\n<li><strong>similarity_first</strong> - w2vec similarity to last aid</li>\n<li><strong>similarity_second</strong> - w2vec similarity to aid before last</li>\n</ul>\n<h4>Cart/Order features</h4>\n<p>(they are mostly the same, will not explain again features also used in a model for clicks):</p>\n<ul>\n<li><strong>n</strong>                          </li>\n<li><strong>time_delta</strong></li>\n<li><strong>count_views</strong></li>\n<li><strong>ts_diff</strong></li>\n<li><strong>type_last</strong> - 0 if no buys for the aid, 1 if last buy is cart, 2 if last buy is order</li>\n<li><strong>time_viewed</strong></li>\n<li><strong>aid_counts_buys</strong> - total buys for aid</li>\n<li><strong>aid_counts_orders</strong> - total orders for aid</li>\n<li><strong>daily_aid_count</strong></li>\n<li><strong>session_time</strong> - time in seconds from first to last event</li>\n<li><strong>events_last_3hours</strong> - total number of events last 3 hours of session</li>\n<li><strong>buys_this_session</strong> - total number of cart/order events in session</li>\n<li><strong>buys_in_session</strong> - 0 if no buys, 1 if only carts, 2 if at least 1 order is present in session. Removed this feature for carts.</li>\n<li><strong>aid_count_weekly</strong> - number of carts/orders for the last week (normalized)</li>\n<li><strong>total_2order_conv</strong>  - feature depending on  <strong>type_last</strong>. If aid was has no buys, here is conversion rate from views to orders, else - conversion rate from carts to orders or from orders to orders. Similar feature was constructed for carts, with cart2cart and order2cart conversion rate. </li>\n<li><strong>clicks_before_buy</strong> - how many times on average item is viewed before first buy</li>\n<li><strong>time_viewed_clipped</strong> - for how long on average aid is viewed before first buy (before averaging values clipped to 180). This feature has low importance and I thought bout removing it, but experiment showed results go a bit down in that case.</li>\n<li><strong>w2v_20_mean</strong> - average w2vec similarity between candidate and last 20 aids (3 hours from last event)</li>\n<li><strong>w2v_20_min</strong> - minimal w2vec similarity between candidate and last 20 aids (3 hours from last event)</li>\n<li><strong>w2v_5_max</strong>  - maximal w2vec similarity between candidate and last 5 aids </li>\n<li><strong>w2v_5_min</strong>  -  minimal w2vec similarity between candidate and last 5 aids (this feature also has low importance, but its removal decreased result a bit)</li>\n<li><strong>history_mean</strong> - w2vec mean similarity between last aid and previous 4 aid before that.</li>\n<li><strong>wgt_b2order</strong> - co-visitation buy2order/buy2cart matrix feature</li>\n<li><strong>wgt_c2buy_short</strong> - co-visitation click2buy matrix feature (matrix counts only cases when there is 1 hour or less between click and buy event)</li>\n<li><strong>wgt_c2buy_full</strong> - co-visitation click2buy matrix feature for 30 last aids (if they are within 3 hours from last event)</li>\n<li><strong>wgt_c2buy_6_from_full</strong> - same as previous, but only 6 last aids.</li>\n</ul>",
  "messages": [
    {
      "id": "2125107",
      "postDate": "02/01/2023 12:28:37",
      "content": "<p>Hi there!<br>\nI worked solo and got to 69th position.</p>\n<p>My solution was similar in its structure to what most users here used. First - candidate generation, then - reranking using GBDT (LGBM and catboost tor orders, only LGBM for the rest). Everything was done on kaggle, had to use many notebooks to calculate everything, some of them ran for 4-6 hours so I had to wait overnight for the results. Had to re-write parts of my feature generation code for polars, as pandas is just too slow.<br>\nI guess it was a big mistake to start building w2vec features only during last week and a half. Probably, could have done much more if I had time to make more experiments with w2vec and probably also try doc2vec.</p>\n<p>I wanted to use average for results for diffents models. So, for orders I used 2 cross-validation sets. I generated 2 different sets of truncked sessions with different random seeds and built 2 different sets of features for them. Then I built LGBM model on 1 set of features and catboost model on another one, then used both models to make predictions on test data and averaged the prediction. The result was not much better than result of a single LGBM model, unfortunately.</p>\n<p>Used only co-visitation matrixes for candidate generation, and I see many people were able to get better recall results for top20 candidates. My numbers are:<br>\nclicks  52.72%<br>\ncarts:  40.63%<br>\norders: 64.71%<br>\nFor top 50 candidates it was:<br>\nclicks  60.43%<br>\ncarts:  45.21%<br>\norders: 67.86%<br>\nAt the last week I tried to switch to top75 candidates, and was suprised to see very close result to what I got for 50 candidates. The model was able to select almost no candidates from those additional coming from top75 candidates.<br>\nMy final cross-validation rates after re-ranking:<br>\nclicks: 54.54%<br>\ncarts: 42.08%<br>\norders: 65.80%</p>\n<h4>Clicks:</h4>\n<ul>\n<li><strong>n</strong> - 0 for the last viewed aid, 1 for aid last viewed before, e.t.c, 125 for aids never viewed</li>\n<li><strong>time_delta</strong> - time in seconds from a moment when aid was last viewed to the last action in session</li>\n<li><strong>type</strong> - 1 for aids ever put in cart, 2 for aids ever ordered, 0 for the rest</li>\n<li><strong>count_views</strong> - number of interactions with aid in a session</li>\n<li><strong>ts_diff</strong> - time in seconds between last event and event before last</li>\n<li><strong>time_viewed</strong> - time from user's click on aid untill next event clipped to 180 seconds and then summed for all views for the aid</li>\n<li><strong>daily_aid_count</strong> - normalized count of events with aid for the previous day</li>\n<li><strong>same_day_aid_count</strong> - normalized count of events with aid for the day</li>\n<li><strong>aid_count_weekly</strong> - normalized count of events with aid for the week</li>\n<li><strong>wgt_last</strong> - made a special co-visitation matrix to calculate only the very next aid after previous, applied it to the last aid session.</li>\n<li><strong>wgt_before_last</strong> - same matrix applied to aid before the last one</li>\n<li><strong>time_viewed_clipped</strong> - average time aid viewed in all interactions (before averaging values were clipped to 180)</li>\n<li><strong>aid_counts</strong> - total interactions with aid</li>\n<li><strong>is_promotion_day</strong> - checked test-period sessions for likely promoted aids. Should have probably removed that feature, bu finally didn't manage to find time to check results without it.</li>\n<li><strong>wgt_matrix</strong> - sum of co-visitaion matrix weights for the last 5 aids (same co-visitaion matrix that was used for candidate generation)</li>\n<li><strong>wgt_exp</strong> - sum of co-visitaion matrix weights for the last 10 aids normalized by n (divided weight by 1 for the last aid, by 2 for aid before it, then by 3 and so on)</li>\n<li><strong>similarity_first</strong> - w2vec similarity to last aid</li>\n<li><strong>similarity_second</strong> - w2vec similarity to aid before last</li>\n</ul>\n<h4>Cart/Order features</h4>\n<p>(they are mostly the same, will not explain again features also used in a model for clicks):</p>\n<ul>\n<li><strong>n</strong>                          </li>\n<li><strong>time_delta</strong></li>\n<li><strong>count_views</strong></li>\n<li><strong>ts_diff</strong></li>\n<li><strong>type_last</strong> - 0 if no buys for the aid, 1 if last buy is cart, 2 if last buy is order</li>\n<li><strong>time_viewed</strong></li>\n<li><strong>aid_counts_buys</strong> - total buys for aid</li>\n<li><strong>aid_counts_orders</strong> - total orders for aid</li>\n<li><strong>daily_aid_count</strong></li>\n<li><strong>session_time</strong> - time in seconds from first to last event</li>\n<li><strong>events_last_3hours</strong> - total number of events last 3 hours of session</li>\n<li><strong>buys_this_session</strong> - total number of cart/order events in session</li>\n<li><strong>buys_in_session</strong> - 0 if no buys, 1 if only carts, 2 if at least 1 order is present in session. Removed this feature for carts.</li>\n<li><strong>aid_count_weekly</strong> - number of carts/orders for the last week (normalized)</li>\n<li><strong>total_2order_conv</strong>  - feature depending on  <strong>type_last</strong>. If aid was has no buys, here is conversion rate from views to orders, else - conversion rate from carts to orders or from orders to orders. Similar feature was constructed for carts, with cart2cart and order2cart conversion rate. </li>\n<li><strong>clicks_before_buy</strong> - how many times on average item is viewed before first buy</li>\n<li><strong>time_viewed_clipped</strong> - for how long on average aid is viewed before first buy (before averaging values clipped to 180). This feature has low importance and I thought bout removing it, but experiment showed results go a bit down in that case.</li>\n<li><strong>w2v_20_mean</strong> - average w2vec similarity between candidate and last 20 aids (3 hours from last event)</li>\n<li><strong>w2v_20_min</strong> - minimal w2vec similarity between candidate and last 20 aids (3 hours from last event)</li>\n<li><strong>w2v_5_max</strong>  - maximal w2vec similarity between candidate and last 5 aids </li>\n<li><strong>w2v_5_min</strong>  -  minimal w2vec similarity between candidate and last 5 aids (this feature also has low importance, but its removal decreased result a bit)</li>\n<li><strong>history_mean</strong> - w2vec mean similarity between last aid and previous 4 aid before that.</li>\n<li><strong>wgt_b2order</strong> - co-visitation buy2order/buy2cart matrix feature</li>\n<li><strong>wgt_c2buy_short</strong> - co-visitation click2buy matrix feature (matrix counts only cases when there is 1 hour or less between click and buy event)</li>\n<li><strong>wgt_c2buy_full</strong> - co-visitation click2buy matrix feature for 30 last aids (if they are within 3 hours from last event)</li>\n<li><strong>wgt_c2buy_6_from_full</strong> - same as previous, but only 6 last aids.</li>\n</ul>",
      "rawMarkdown": "Hi there!\nI worked solo and got to 69th position.\n\nMy solution was similar in its structure to what most users here used. First - candidate generation, then - reranking using GBDT (LGBM and catboost tor orders, only LGBM for the rest). Everything was done on kaggle, had to use many notebooks to calculate everything, some of them ran for 4-6 hours so I had to wait overnight for the results. Had to re-write parts of my feature generation code for polars, as pandas is just too slow.\nI guess it was a big mistake to start building w2vec features only during last week and a half. Probably, could have done much more if I had time to make more experiments with w2vec and probably also try doc2vec.\n\nI wanted to use average for results for diffents models. So, for orders I used 2 cross-validation sets. I generated 2 different sets of truncked sessions with different random seeds and built 2 different sets of features for them. Then I built LGBM model on 1 set of features and catboost model on another one, then used both models to make predictions on test data and averaged the prediction. The result was not much better than result of a single LGBM model, unfortunately.\n\nUsed only co-visitation matrixes for candidate generation, and I see many people were able to get better recall results for top20 candidates. My numbers are:\nclicks  52.72%\ncarts:  40.63%\norders: 64.71%\nFor top 50 candidates it was:\nclicks  60.43%\ncarts:  45.21%\norders: 67.86%\nAt the last week I tried to switch to top75 candidates, and was suprised to see very close result to what I got for 50 candidates. The model was able to select almost no candidates from those additional coming from top75 candidates.\nMy final cross-validation rates after re-ranking:\nclicks: 54.54%\ncarts: 42.08%\norders: 65.80%\n\n#### Clicks:\n- **n** - 0 for the last viewed aid, 1 for aid last viewed before, e.t.c, 125 for aids never viewed\n- **time_delta** - time in seconds from a moment when aid was last viewed to the last action in session\n- **type** - 1 for aids ever put in cart, 2 for aids ever ordered, 0 for the rest\n- **count_views** - number of interactions with aid in a session\n- **ts_diff** - time in seconds between last event and event before last\n- **time_viewed** - time from user's click on aid untill next event clipped to 180 seconds and then summed for all views for the aid\n- **daily_aid_count** - normalized count of events with aid for the previous day\n- **same_day_aid_count** - normalized count of events with aid for the day\n- **aid_count_weekly** - normalized count of events with aid for the week\n- **wgt_last** - made a special co-visitation matrix to calculate only the very next aid after previous, applied it to the last aid session.\n- **wgt_before_last** - same matrix applied to aid before the last one\n- **time_viewed_clipped** - average time aid viewed in all interactions (before averaging values were clipped to 180)\n- **aid_counts** - total interactions with aid\n- **is_promotion_day** - checked test-period sessions for likely promoted aids. Should have probably removed that feature, bu finally didn't manage to find time to check results without it.\n- **wgt_matrix** - sum of co-visitaion matrix weights for the last 5 aids (same co-visitaion matrix that was used for candidate generation)\n- **wgt_exp** - sum of co-visitaion matrix weights for the last 10 aids normalized by n (divided weight by 1 for the last aid, by 2 for aid before it, then by 3 and so on)\n- **similarity_first** - w2vec similarity to last aid\n- **similarity_second** - w2vec similarity to aid before last\n\n#### Cart/Order features \n(they are mostly the same, will not explain again features also used in a model for clicks):\n\n- **n**                          \n- **time_delta**\n- **count_views**\n- **ts_diff**\n- **type_last** - 0 if no buys for the aid, 1 if last buy is cart, 2 if last buy is order\n- **time_viewed**\n- **aid_counts_buys** - total buys for aid\n- **aid_counts_orders** - total orders for aid\n- **daily_aid_count**\n- **session_time** - time in seconds from first to last event\n- **events_last_3hours** - total number of events last 3 hours of session\n- **buys_this_session** - total number of cart/order events in session\n- **buys_in_session** - 0 if no buys, 1 if only carts, 2 if at least 1 order is present in session. Removed this feature for carts.\n- **aid_count_weekly** - number of carts/orders for the last week (normalized)\n- **total_2order_conv**  - feature depending on  **type_last**. If aid was has no buys, here is conversion rate from views to orders, else - conversion rate from carts to orders or from orders to orders. Similar feature was constructed for carts, with cart2cart and order2cart conversion rate. \n- **clicks_before_buy** - how many times on average item is viewed before first buy\n- **time_viewed_clipped** - for how long on average aid is viewed before first buy (before averaging values clipped to 180). This feature has low importance and I thought bout removing it, but experiment showed results go a bit down in that case.\n- **w2v_20_mean** - average w2vec similarity between candidate and last 20 aids (3 hours from last event)\n- **w2v_20_min** - minimal w2vec similarity between candidate and last 20 aids (3 hours from last event)\n- **w2v_5_max**  - maximal w2vec similarity between candidate and last 5 aids \n- **w2v_5_min**  -  minimal w2vec similarity between candidate and last 5 aids (this feature also has low importance, but its removal decreased result a bit)\n- **history_mean** - w2vec mean similarity between last aid and previous 4 aid before that.\n- **wgt_b2order** - co-visitation buy2order/buy2cart matrix feature\n- **wgt_c2buy_short** - co-visitation click2buy matrix feature (matrix counts only cases when there is 1 hour or less between click and buy event)\n- **wgt_c2buy_full** - co-visitation click2buy matrix feature for 30 last aids (if they are within 3 hours from last event)\n- **wgt_c2buy_6_from_full** - same as previous, but only 6 last aids.",
      "votes": null
    },
    {
      "id": "2125177",
      "postDate": "02/01/2023 13:55:28",
      "content": "<p>Well done <a href=\"https://www.kaggle.com/artemfedorov\" target=\"_blank\">@artemfedorov</a> and thanks for sharing. You had some interesting features. </p>",
      "rawMarkdown": "Well done @artemfedorov and thanks for sharing. You had some interesting features.",
      "votes": null
    },
    {
      "id": "2125424",
      "postDate": "02/01/2023 16:43:53",
      "content": "<p>Thank you. Some of those original features didn't contribute much to final result, unfortunately.</p>",
      "rawMarkdown": "Thank you. Some of those original features didn't contribute much to final result, unfortunately.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2125177,
      "author_name": "gehallak",
      "author_url": "",
      "post_date": "02/01/2023 13:55:28",
      "content": "<p>Well done <a href=\"https://www.kaggle.com/artemfedorov\" target=\"_blank\">@artemfedorov</a> and thanks for sharing. You had some interesting features. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2125424,
          "author_name": "artemfedorov",
          "author_url": "",
          "post_date": "02/01/2023 16:43:53",
          "content": "<p>Thank you. Some of those original features didn't contribute much to final result, unfortunately.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2125107": "Hi there!\nI worked solo and got to 69th position.\n\nMy solution was similar in its structure to what most users here used. First - candidate generation, then - reranking using GBDT (LGBM and catboost tor orders, only LGBM for the rest). Everything was done on kaggle, had to use many notebooks to calculate everything, some of them ran for 4-6 hours so I had to wait overnight for the results. Had to re-write parts of my feature generation code for polars, as pandas is just too slow.\nI guess it was a big mistake to start building w2vec features only during last week and a half. Probably, could have done much more if I had time to make more experiments with w2vec and probably also try doc2vec.\n\nI wanted to use average for results for diffents models. So, for orders I used 2 cross-validation sets. I generated 2 different sets of truncked sessions with different random seeds and built 2 different sets of features for them. Then I built LGBM model on 1 set of features and catboost model on another one, then used both models to make predictions on test data and averaged the prediction. The result was not much better than result of a single LGBM model, unfortunately.\n\nUsed only co-visitation matrixes for candidate generation, and I see many people were able to get better recall results for top20 candidates. My numbers are:\nclicks  52.72%\ncarts:  40.63%\norders: 64.71%\nFor top 50 candidates it was:\nclicks  60.43%\ncarts:  45.21%\norders: 67.86%\nAt the last week I tried to switch to top75 candidates, and was suprised to see very close result to what I got for 50 candidates. The model was able to select almost no candidates from those additional coming from top75 candidates.\nMy final cross-validation rates after re-ranking:\nclicks: 54.54%\ncarts: 42.08%\norders: 65.80%\n\n#### Clicks:\n- **n** - 0 for the last viewed aid, 1 for aid last viewed before, e.t.c, 125 for aids never viewed\n- **time_delta** - time in seconds from a moment when aid was last viewed to the last action in session\n- **type** - 1 for aids ever put in cart, 2 for aids ever ordered, 0 for the rest\n- **count_views** - number of interactions with aid in a session\n- **ts_diff** - time in seconds between last event and event before last\n- **time_viewed** - time from user's click on aid untill next event clipped to 180 seconds and then summed for all views for the aid\n- **daily_aid_count** - normalized count of events with aid for the previous day\n- **same_day_aid_count** - normalized count of events with aid for the day\n- **aid_count_weekly** - normalized count of events with aid for the week\n- **wgt_last** - made a special co-visitation matrix to calculate only the very next aid after previous, applied it to the last aid session.\n- **wgt_before_last** - same matrix applied to aid before the last one\n- **time_viewed_clipped** - average time aid viewed in all interactions (before averaging values were clipped to 180)\n- **aid_counts** - total interactions with aid\n- **is_promotion_day** - checked test-period sessions for likely promoted aids. Should have probably removed that feature, bu finally didn't manage to find time to check results without it.\n- **wgt_matrix** - sum of co-visitaion matrix weights for the last 5 aids (same co-visitaion matrix that was used for candidate generation)\n- **wgt_exp** - sum of co-visitaion matrix weights for the last 10 aids normalized by n (divided weight by 1 for the last aid, by 2 for aid before it, then by 3 and so on)\n- **similarity_first** - w2vec similarity to last aid\n- **similarity_second** - w2vec similarity to aid before last\n\n#### Cart/Order features \n(they are mostly the same, will not explain again features also used in a model for clicks):\n\n- **n**                          \n- **time_delta**\n- **count_views**\n- **ts_diff**\n- **type_last** - 0 if no buys for the aid, 1 if last buy is cart, 2 if last buy is order\n- **time_viewed**\n- **aid_counts_buys** - total buys for aid\n- **aid_counts_orders** - total orders for aid\n- **daily_aid_count**\n- **session_time** - time in seconds from first to last event\n- **events_last_3hours** - total number of events last 3 hours of session\n- **buys_this_session** - total number of cart/order events in session\n- **buys_in_session** - 0 if no buys, 1 if only carts, 2 if at least 1 order is present in session. Removed this feature for carts.\n- **aid_count_weekly** - number of carts/orders for the last week (normalized)\n- **total_2order_conv**  - feature depending on  **type_last**. If aid was has no buys, here is conversion rate from views to orders, else - conversion rate from carts to orders or from orders to orders. Similar feature was constructed for carts, with cart2cart and order2cart conversion rate. \n- **clicks_before_buy** - how many times on average item is viewed before first buy\n- **time_viewed_clipped** - for how long on average aid is viewed before first buy (before averaging values clipped to 180). This feature has low importance and I thought bout removing it, but experiment showed results go a bit down in that case.\n- **w2v_20_mean** - average w2vec similarity between candidate and last 20 aids (3 hours from last event)\n- **w2v_20_min** - minimal w2vec similarity between candidate and last 20 aids (3 hours from last event)\n- **w2v_5_max**  - maximal w2vec similarity between candidate and last 5 aids \n- **w2v_5_min**  -  minimal w2vec similarity between candidate and last 5 aids (this feature also has low importance, but its removal decreased result a bit)\n- **history_mean** - w2vec mean similarity between last aid and previous 4 aid before that.\n- **wgt_b2order** - co-visitation buy2order/buy2cart matrix feature\n- **wgt_c2buy_short** - co-visitation click2buy matrix feature (matrix counts only cases when there is 1 hour or less between click and buy event)\n- **wgt_c2buy_full** - co-visitation click2buy matrix feature for 30 last aids (if they are within 3 hours from last event)\n- **wgt_c2buy_6_from_full** - same as previous, but only 6 last aids.",
    "2125177": "Well done @artemfedorov and thanks for sharing. You had some interesting features.",
    "2125424": "Thank you. Some of those original features didn't contribute much to final result, unfortunately."
  },
  "source": "meta"
}