{
  "id": 382103,
  "title": "Question about the interactions features",
  "url": "/competitions/otto-recommender-system/discussion/382103",
  "author_name": "",
  "post_date": "2023-01-29T15:54:18.288490700Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello Community, I have a question about the LEFT JOIN between the  tables, and the interactions features table.</p>\n<p>When doing the LEFT JOIN, you ends up with the same candidates multiple times when the USER ( session ) already interacted multiple times with the same item (aid).</p>\n<p>How did you deal with it ? knowing that you should have unique  pair for the Ranker.</p>\n<p>Thanks !</p>",
  "messages": [
    {
      "id": "2120446",
      "postDate": "01/29/2023 15:54:18",
      "content": "<p>Hello Community, I have a question about the LEFT JOIN between the  tables, and the interactions features table.</p>\n<p>When doing the LEFT JOIN, you ends up with the same candidates multiple times when the USER ( session ) already interacted multiple times with the same item (aid).</p>\n<p>How did you deal with it ? knowing that you should have unique  pair for the Ranker.</p>\n<p>Thanks !</p>",
      "rawMarkdown": "Hello Community, I have a question about the LEFT JOIN between the <session,candidates> tables, and the interactions features table.\n\nWhen doing the LEFT JOIN, you ends up with the same candidates multiple times when the USER ( session ) already interacted multiple times with the same item (aid).\n\nHow did you deal with it ? knowing that you should have unique <session,aid> pair for the Ranker.\n\nThanks !",
      "votes": null
    },
    {
      "id": "2120469",
      "postDate": "01/29/2023 16:09:09",
      "content": "<p>In my settings, only keep unique candidate item, assume that you have 2 strategies(revisit &amp; co-visit) to generate the candidate, then the train data of re-ranker will like, </p>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>aid</th>\n<th>s1_score</th>\n<th>s1_other_features</th>\n<th>s2_score</th>\n<th>s2_other features</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>1</td>\n<td>0.5</td>\n<td>…</td>\n<td>0.6</td>\n<td>…</td>\n</tr>\n<tr>\n<td>1</td>\n<td>2</td>\n<td>0.5</td>\n<td>…</td>\n<td>NaN</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "In my settings, only keep unique candidate item, assume that you have 2 strategies(revisit & co-visit) to generate the candidate, then the train data of re-ranker will like, \n| session |aid  |s1_score|s1_other_features|s2_score|s2_other features|\n| --- | --- |\n| 1 |1  |0.5|...|0.6|...|\n| 1 |2  |0.5|...|NaN|...|",
      "votes": null
    },
    {
      "id": "2120494",
      "postDate": "01/29/2023 16:35:25",
      "content": "<p>Yep, I agree. Personally, I'm using <br>\n<code>df.unique(subset=['session','aid'],keep=\"last\")</code></p>\n<p>But I think that this is not the right thing to do</p>",
      "rawMarkdown": "Yep, I agree. Personally, I'm using \n`df.unique(subset=['session','aid'],keep=\"last\")`\n\nBut I think that this is not the right thing to do",
      "votes": null
    },
    {
      "id": "2120884",
      "postDate": "01/29/2023 22:01:38",
      "content": "<p>Also, I have like &gt; 90% of NaNs using my interactions features, is it normal ? Here is an exemple of my feature, and the ratio of NaNs for each interaction feature:</p>\n<p>type                                           0.950608<br>\ninteraction_rank_per_session_clicks            0.953647<br>\ninteraction_rank_per_session_orders            0.982781<br>\ninteraction_rank_per_session_carts             0.999100<br>\ninteraction_rank_per_session                   0.950608<br>\ninteraction_rank_per_sessionUnique             0.950608<br>\ninteraction_rank_per_session_aid_type          0.950608<br>\ninteraction_rank_per_session_type              0.950608<br>\ninteraction_log_recency_score                  0.950608<br>\ninteraction_type_weighted_log_recency_score    0.950608<br>\ninteraction_nb_seen                            0.950608<br>\ninteraction_nb_clicks_interaction              0.950608<br>\ninteraction_nb_carts_interaction               0.950608<br>\ninteraction_nb_orders_interaction              0.950608<br>\nalready_clicked                                0.950608<br>\nalready_cartered                               0.950608<br>\nalready_ordered                                0.950608<br>\ninteraction_type_ratio                         0.950608<br>\ninteraction_type_median                        0.950608<br>\ninteraction_type_mode_user_aid                 0.950608<br>\nalready_see                                    0.950608</p>",
      "rawMarkdown": "Also, I have like > 90% of NaNs using my interactions features, is it normal ? Here is an exemple of my feature, and the ratio of NaNs for each interaction feature:\n\ntype                                           0.950608\ninteraction_rank_per_session_clicks            0.953647\ninteraction_rank_per_session_orders            0.982781\ninteraction_rank_per_session_carts             0.999100\ninteraction_rank_per_session                   0.950608\ninteraction_rank_per_sessionUnique             0.950608\ninteraction_rank_per_session_aid_type          0.950608\ninteraction_rank_per_session_type              0.950608\ninteraction_log_recency_score                  0.950608\ninteraction_type_weighted_log_recency_score    0.950608\ninteraction_nb_seen                            0.950608\ninteraction_nb_clicks_interaction              0.950608\ninteraction_nb_carts_interaction               0.950608\ninteraction_nb_orders_interaction              0.950608\nalready_clicked                                0.950608\nalready_cartered                               0.950608\nalready_ordered                                0.950608\ninteraction_type_ratio                         0.950608\ninteraction_type_median                        0.950608\ninteraction_type_mode_user_aid                 0.950608\nalready_see                                    0.950608",
      "votes": null
    },
    {
      "id": "2120906",
      "postDate": "01/29/2023 22:37:29",
      "content": "<p>Many of these are not NANs. For example</p>\n<pre><code>already_clicked 0.950608\nalready_cartered 0.950608\nalready_ordered 0.950608\n</code></pre>\n<p>There are no NANs here. Either they clicked or they did not click. Either it is <code>True</code> or <code>False</code>. There are no NANs. After you merge, just <code>fillna(False)</code>.</p>",
      "rawMarkdown": "Many of these are not NANs. For example\n\n    already_clicked 0.950608\n    already_cartered 0.950608\n    already_ordered 0.950608\n\nThere are no NANs here. Either they clicked or they did not click. Either it is `True` or `False`. There are no NANs. After you merge, just `fillna(False)`.",
      "votes": null
    },
    {
      "id": "2120916",
      "postDate": "01/29/2023 22:47:39",
      "content": "<p>I totally agree, but considering this feature selection <a href=\"https://www.uber.com/blog/research/maximum-relevance-and-minimum-redundancy-feature-selection-methods-for-a-marketing-machine-learning-platform\" target=\"_blank\">Maximum Relevance Minimum Redundance ( MRMR )</a>, it seems that all those features with NaNs are kind of redundants and introduced kind of noise in my model by adding complexity to my model which can explain my overfitting despite my good local validation score</p>\n<p>On the other hand, I tried some experiments by using onlythe [<strong>already_clicked</strong>, <strong>already_cartered</strong>, <strong>already_ordered</strong>] of the interactions features ( from the MRMR algorithm ), and   I got the same NDCG with a better validation score. </p>\n<p>I just started those experiments, I'll keep you informed if the LB agree with me or not 😂</p>",
      "rawMarkdown": "I totally agree, but considering this feature selection [Maximum Relevance Minimum Redundance ( MRMR )](https://www.uber.com/blog/research/maximum-relevance-and-minimum-redundancy-feature-selection-methods-for-a-marketing-machine-learning-platform), it seems that all those features with NaNs are kind of redundants and introduced kind of noise in my model by adding complexity to my model which can explain my overfitting despite my good local validation score\n\nOn the other hand, I tried some experiments by using onlythe [**already_clicked**, **already_cartered**, **already_ordered**] of the interactions features ( from the MRMR algorithm ), and   I got the same NDCG with a better validation score. \n\nI just started those experiments, I'll keep you informed if the LB agree with me or not 😂",
      "votes": null
    },
    {
      "id": "2121004",
      "postDate": "01/30/2023 00:57:17",
      "content": "<p>I just tried the submission, and it seems that my LB didn't increase nor decrease, which means that :</p>\n<p>-My features are redundants.</p>\n<p>-My potential overfitting is not due to my features interactions.</p>\n<p>I really don't know where I missed up, I need to investigate further,</p>\n<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "I just tried the submission, and it seems that my LB didn't increase nor decrease, which means that :\n\n-My features are redundants.\n\n-My potential overfitting is not due to my features interactions.\n\n I really don't know where I missed up, I need to investigate further,\n\nThanks @cdeotte",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2120469,
      "author_name": "gongbi",
      "author_url": "",
      "post_date": "01/29/2023 16:09:09",
      "content": "<p>In my settings, only keep unique candidate item, assume that you have 2 strategies(revisit &amp; co-visit) to generate the candidate, then the train data of re-ranker will like, </p>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>aid</th>\n<th>s1_score</th>\n<th>s1_other_features</th>\n<th>s2_score</th>\n<th>s2_other features</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>1</td>\n<td>0.5</td>\n<td>…</td>\n<td>0.6</td>\n<td>…</td>\n</tr>\n<tr>\n<td>1</td>\n<td>2</td>\n<td>0.5</td>\n<td>…</td>\n<td>NaN</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 2120494,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/29/2023 16:35:25",
          "content": "<p>Yep, I agree. Personally, I'm using <br>\n<code>df.unique(subset=['session','aid'],keep=\"last\")</code></p>\n<p>But I think that this is not the right thing to do</p>",
          "votes": null,
          "replies": [
            {
              "id": 2120884,
              "author_name": "rayanaay",
              "author_url": "",
              "post_date": "01/29/2023 22:01:38",
              "content": "<p>Also, I have like &gt; 90% of NaNs using my interactions features, is it normal ? Here is an exemple of my feature, and the ratio of NaNs for each interaction feature:</p>\n<p>type                                           0.950608<br>\ninteraction_rank_per_session_clicks            0.953647<br>\ninteraction_rank_per_session_orders            0.982781<br>\ninteraction_rank_per_session_carts             0.999100<br>\ninteraction_rank_per_session                   0.950608<br>\ninteraction_rank_per_sessionUnique             0.950608<br>\ninteraction_rank_per_session_aid_type          0.950608<br>\ninteraction_rank_per_session_type              0.950608<br>\ninteraction_log_recency_score                  0.950608<br>\ninteraction_type_weighted_log_recency_score    0.950608<br>\ninteraction_nb_seen                            0.950608<br>\ninteraction_nb_clicks_interaction              0.950608<br>\ninteraction_nb_carts_interaction               0.950608<br>\ninteraction_nb_orders_interaction              0.950608<br>\nalready_clicked                                0.950608<br>\nalready_cartered                               0.950608<br>\nalready_ordered                                0.950608<br>\ninteraction_type_ratio                         0.950608<br>\ninteraction_type_median                        0.950608<br>\ninteraction_type_mode_user_aid                 0.950608<br>\nalready_see                                    0.950608</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2120906,
                  "author_name": "cdeotte",
                  "author_url": "",
                  "post_date": "01/29/2023 22:37:29",
                  "content": "<p>Many of these are not NANs. For example</p>\n<pre><code>already_clicked 0.950608\nalready_cartered 0.950608\nalready_ordered 0.950608\n</code></pre>\n<p>There are no NANs here. Either they clicked or they did not click. Either it is <code>True</code> or <code>False</code>. There are no NANs. After you merge, just <code>fillna(False)</code>.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2120916,
                      "author_name": "rayanaay",
                      "author_url": "",
                      "post_date": "01/29/2023 22:47:39",
                      "content": "<p>I totally agree, but considering this feature selection <a href=\"https://www.uber.com/blog/research/maximum-relevance-and-minimum-redundancy-feature-selection-methods-for-a-marketing-machine-learning-platform\" target=\"_blank\">Maximum Relevance Minimum Redundance ( MRMR )</a>, it seems that all those features with NaNs are kind of redundants and introduced kind of noise in my model by adding complexity to my model which can explain my overfitting despite my good local validation score</p>\n<p>On the other hand, I tried some experiments by using onlythe [<strong>already_clicked</strong>, <strong>already_cartered</strong>, <strong>already_ordered</strong>] of the interactions features ( from the MRMR algorithm ), and   I got the same NDCG with a better validation score. </p>\n<p>I just started those experiments, I'll keep you informed if the LB agree with me or not 😂</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2121004,
                          "author_name": "rayanaay",
                          "author_url": "",
                          "post_date": "01/30/2023 00:57:17",
                          "content": "<p>I just tried the submission, and it seems that my LB didn't increase nor decrease, which means that :</p>\n<p>-My features are redundants.</p>\n<p>-My potential overfitting is not due to my features interactions.</p>\n<p>I really don't know where I missed up, I need to investigate further,</p>\n<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2120446": "Hello Community, I have a question about the LEFT JOIN between the <session,candidates> tables, and the interactions features table.\n\nWhen doing the LEFT JOIN, you ends up with the same candidates multiple times when the USER ( session ) already interacted multiple times with the same item (aid).\n\nHow did you deal with it ? knowing that you should have unique <session,aid> pair for the Ranker.\n\nThanks !",
    "2120469": "In my settings, only keep unique candidate item, assume that you have 2 strategies(revisit & co-visit) to generate the candidate, then the train data of re-ranker will like, \n| session |aid  |s1_score|s1_other_features|s2_score|s2_other features|\n| --- | --- |\n| 1 |1  |0.5|...|0.6|...|\n| 1 |2  |0.5|...|NaN|...|",
    "2120494": "Yep, I agree. Personally, I'm using \n`df.unique(subset=['session','aid'],keep=\"last\")`\n\nBut I think that this is not the right thing to do",
    "2120884": "Also, I have like > 90% of NaNs using my interactions features, is it normal ? Here is an exemple of my feature, and the ratio of NaNs for each interaction feature:\n\ntype                                           0.950608\ninteraction_rank_per_session_clicks            0.953647\ninteraction_rank_per_session_orders            0.982781\ninteraction_rank_per_session_carts             0.999100\ninteraction_rank_per_session                   0.950608\ninteraction_rank_per_sessionUnique             0.950608\ninteraction_rank_per_session_aid_type          0.950608\ninteraction_rank_per_session_type              0.950608\ninteraction_log_recency_score                  0.950608\ninteraction_type_weighted_log_recency_score    0.950608\ninteraction_nb_seen                            0.950608\ninteraction_nb_clicks_interaction              0.950608\ninteraction_nb_carts_interaction               0.950608\ninteraction_nb_orders_interaction              0.950608\nalready_clicked                                0.950608\nalready_cartered                               0.950608\nalready_ordered                                0.950608\ninteraction_type_ratio                         0.950608\ninteraction_type_median                        0.950608\ninteraction_type_mode_user_aid                 0.950608\nalready_see                                    0.950608",
    "2120906": "Many of these are not NANs. For example\n\n    already_clicked 0.950608\n    already_cartered 0.950608\n    already_ordered 0.950608\n\nThere are no NANs here. Either they clicked or they did not click. Either it is `True` or `False`. There are no NANs. After you merge, just `fillna(False)`.",
    "2120916": "I totally agree, but considering this feature selection [Maximum Relevance Minimum Redundance ( MRMR )](https://www.uber.com/blog/research/maximum-relevance-and-minimum-redundancy-feature-selection-methods-for-a-marketing-machine-learning-platform), it seems that all those features with NaNs are kind of redundants and introduced kind of noise in my model by adding complexity to my model which can explain my overfitting despite my good local validation score\n\nOn the other hand, I tried some experiments by using onlythe [**already_clicked**, **already_cartered**, **already_ordered**] of the interactions features ( from the MRMR algorithm ), and   I got the same NDCG with a better validation score. \n\nI just started those experiments, I'll keep you informed if the LB agree with me or not 😂",
    "2121004": "I just tried the submission, and it seems that my LB didn't increase nor decrease, which means that :\n\n-My features are redundants.\n\n-My potential overfitting is not due to my features interactions.\n\n I really don't know where I missed up, I need to investigate further,\n\nThanks @cdeotte"
  },
  "source": "meta"
}