{
  "id": 378893,
  "title": "Session (User) features are not effective in my GBM model.",
  "url": "/competitions/otto-recommender-system/discussion/378893",
  "author_name": "tetsuro731",
  "post_date": "2023-01-17T11:17:41.304000",
  "votes": 2,
  "comment_count": 18,
  "views": 0,
  "content": "<p>In this competition, I think most of the people use GBM model like a XGBoost or LightGBM.</p>\n<p>How to build a GBM model is explained in detail at the following discussion:<br>\n<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210</a></p>\n<p>I tried to use GBM model and I added \"session features\" and \"aid features\". <br>\nIn my model, aid features are effective but session features are not effective at all.<br>\nIn other words, feature importance of my session features are nearly zero whereas aid features have high importance.</p>\n<p>I'm not sure this situation is natural or not.</p>\n<p>For example, I generated session features as</p>\n<pre><code>user_features = train.groupby('session').agg({'session':'count','aid':'nunique','type':'mean'})\nuser_features.columns = ['user_user_count','user_item_count','user_buy_ratio']\n</code></pre>\n<p>,and I also generated aid features as</p>\n<pre><code>item_features = train.groupby('aid').agg({'aid':'count','session':'nunique','type':'mean'})\nitem_features.columns = ['item_item_count','item_user_count','item_buy_ratio']\n</code></pre>\n<p>.<br>\nIf possible, I want to know this situation is common for all users or only happened in my environment.</p>\n<h3>20220122</h3>\n<p>Thank you for sharing a lot of ideas and comments.<br>\nI found bug in my code and I finally found that my session feature can boost my model and LB.<br>\nIn my results, aid features are more effective than session features but both features have a potential to boost the score.</p>",
  "messages": [
    {
      "id": 2103770,
      "postDate": "2023-01-17T11:28:47.137Z",
      "content": "<p>Since we are predicting the preference of the items by the users, the relationship of the session with that item or similar items may be more valuable for your model instead of general statistics about that session. Therefore, extracting user-interaction features that can represent not only the user but also represent the relationship between (user and candidate) (<code>groupby([\"session\", \"aid\"])</code>) would affect your score more. But it would still be beneficial for you to generate extra features representing the session details.</p>",
      "rawMarkdown": "Since we are predicting the preference of the items by the users, the relationship of the session with that item or similar items may be more valuable for your model instead of general statistics about that session. Therefore, extracting user-interaction features that can represent not only the user but also represent the relationship between (user and candidate) (`groupby([\"session\", \"aid\"])`) would affect your score more. But it would still be beneficial for you to generate extra features representing the session details.",
      "votes": 7,
      "replies": [
        {
          "id": 2103789,
          "postDate": "2023-01-17T11:42:37.280Z",
          "content": "<p>Thank you for your comments!<br>\nI was surprised that the feature importance of the simple session features are almost zero when I tried to add the simple session features.<br>\nOn the other hand, several informative notebook say that we can generate session and aid features, which have a potential to boost <br>\nour model and LB score.<br>\nTherefore, I was confused by my results.</p>\n<p>Anyway, for now I guess it's not strange that the simple session information itself can't boost the score and we should consider the more effective interactive features. (if my understanding is correct)</p>",
          "rawMarkdown": "Thank you for your comments!\nI was surprised that the feature importance of the simple session features are almost zero when I tried to add the simple session features.\nOn the other hand, several informative notebook say that we can generate session and aid features, which have a potential to boost \nour model and LB score.\nTherefore, I was confused by my results.\n\nAnyway, for now I guess it's not strange that the simple session information itself can't boost the score and we should consider the more effective interactive features. (if my understanding is correct)\n"
        },
        {
          "id": 2105579,
          "postDate": "2023-01-18T16:06:54.713Z",
          "content": "<p>Sorry, I don't understand about the user-interaction features? Could you give some simple examples. Thanks very much!</p>",
          "rawMarkdown": "Sorry, I don't understand about the user-interaction features? Could you give some simple examples. Thanks very much!",
          "replies": [
            {
              "id": 2105716,
              "postDate": "2023-01-18T17:01:29.550Z",
              "content": "<p>for example:</p>\n<ol>\n<li>did user buy/cart/click the retrieved item?</li>\n<li>how many times the user buy/cart/click the retrieved item</li>\n</ol>",
              "rawMarkdown": "for example:\n1. did user buy/cart/click the retrieved item?\n2. how many times the user buy/cart/click the retrieved item"
            },
            {
              "id": 2106209,
              "postDate": "2023-01-19T02:44:40.473Z",
              "content": "<p>Thank you very much!!! I have another question. Is some mix-features necessary for prediction, maybe some features produced by groupby(['session', 'aid']), like:</p>\n<ol>\n<li>how many times item 'aid' does user 'session' buy/click/cart</li>\n<li>what's the ratio user 'session' buys/clicks/carts item 'aid'</li>\n</ol>\n<p>I found that many candidates will get NaN on these features. So I'm not sure whether the features like that are necessary. </p>",
              "rawMarkdown": "Thank you very much!!! I have another question. Is some mix-features necessary for prediction, maybe some features produced by groupby(['session', 'aid']), like:\n\n1. how many times item 'aid' does user 'session' buy/click/cart\n2. what's the ratio user 'session' buys/clicks/carts item 'aid'\n\nI found that many candidates will get NaN on these features. So I'm not sure whether the features like that are necessary. "
            },
            {
              "id": 2106275,
              "postDate": "2023-01-19T04:10:26.843Z",
              "content": "<p>We can do a simple A/B test to check the performance of these 'suspicious' features since the training for orders should be very fast. Just drop the features you doubt,and train, to see if the validation score improves. </p>",
              "rawMarkdown": "We can do a simple A/B test to check the performance of these 'suspicious' features since the training for orders should be very fast. Just drop the features you doubt,and train, to see if the validation score improves. "
            },
            {
              "id": 2114039,
              "postDate": "2023-01-24T17:07:27.827Z",
              "content": "<p>excuse me, i want to ask the questions of  training for orders, did you negative sample of the data, the 200 recall data is too slow to train even though for the orders, and the Memory is limited， but the CV of the negative sample of the data is very high and not equal to the lb, how to solve it? (maybe my user_item_feature is not userd and too waste time</p>",
              "rawMarkdown": "excuse me, i want to ask the questions of  training for orders, did you negative sample of the data, the 200 recall data is too slow to train even though for the orders, and the Memory is limited， but the CV of the negative sample of the data is very high and not equal to the lb, how to solve it? (maybe my user_item_feature is not userd and too waste time"
            }
          ]
        },
        {
          "id": 2105860,
          "postDate": "2023-01-18T19:34:36.203Z",
          "content": "<p><a href=\"https://www.kaggle.com/nlztrk\" target=\"_blank\">@nlztrk</a> Thanks for the clarification. A follow up question: when you create interaction features - how do you merge it back onto the candidate dataframe? Is it on <code>('session', 'aid')</code> and how are you dealing with nulls in this case? (which I believe can be significant) I might be missing something here. Looking forward to your response :) </p>",
          "rawMarkdown": "@nlztrk Thanks for the clarification. A follow up question: when you create interaction features - how do you merge it back onto the candidate dataframe? Is it on `('session', 'aid')` and how are you dealing with nulls in this case? (which I believe can be significant) I might be missing something here. Looking forward to your response :) ",
          "replies": [
            {
              "id": 2106772,
              "postDate": "2023-01-19T10:47:59.893Z",
              "content": "<p>If a candidate is not seen during the history of that specific session, the user-interaction stats derived from that candidate would be <strong>NaN</strong>. You can simply fill them with value that has no possibility to be seen, like -1, -999 etc. And it's actually a pretty good information IMHO that indicates user had any interactions with that item or item group.</p>",
              "rawMarkdown": "If a candidate is not seen during the history of that specific session, the user-interaction stats derived from that candidate would be **NaN**. You can simply fill them with value that has no possibility to be seen, like -1, -999 etc. And it's actually a pretty good information IMHO that indicates user had any interactions with that item or item group.",
              "votes": 3
            },
            {
              "id": 2107540,
              "postDate": "2023-01-19T21:00:50.723Z",
              "content": "<p><a href=\"https://www.kaggle.com/nlztrk\" target=\"_blank\">@nlztrk</a> Thanks a ton for the response!</p>",
              "rawMarkdown": "@nlztrk Thanks a ton for the response!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2103774,
      "postDate": "2023-01-17T11:33:38.963Z",
      "content": "<p>For me it is not the case. Such as the features you named 'user_user_count','user_item_count'---namely the session length and session unique item numbers ranked within top 50% of all my features.</p>",
      "rawMarkdown": "For me it is not the case. Such as the features you named 'user_user_count','user_item_count'---namely the session length and session unique item numbers ranked within top 50% of all my features.",
      "votes": 3,
      "replies": [
        {
          "id": 2103791,
          "postDate": "2023-01-17T11:45:27.640Z",
          "content": "<p>Thank you for your nice information.<br>\nIf you are correct, I guess my training code includes some bugs..</p>",
          "rawMarkdown": "Thank you for your nice information.\nIf you are correct, I guess my training code includes some bugs.."
        },
        {
          "id": 2106536,
          "postDate": "2023-01-19T08:05:47.773Z",
          "content": "<p>my observation aligned with <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a>, these 2 features (session length &amp; session unique item) are ranked top in my ranker model.</p>",
          "rawMarkdown": "my observation aligned with @buumoo, these 2 features (session length & session unique item) are ranked top in my ranker model.",
          "votes": 1,
          "replies": [
            {
              "id": 2110163,
              "postDate": "2023-01-22T02:46:07.747Z",
              "content": "<p>Thank you for your comment.<br>\nI fixed the bug in my code and my LB score had been improved.</p>",
              "rawMarkdown": "Thank you for your comment.\nI fixed the bug in my code and my LB score had been improved."
            }
          ]
        },
        {
          "id": 2110550,
          "postDate": "2023-01-22T09:12:16.873Z",
          "content": "<p>That's the part where models learn whether your candidates are generated with recency weighting or unique aids + covisitation.</p>",
          "rawMarkdown": "That's the part where models learn whether your candidates are generated with recency weighting or unique aids + covisitation.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2103761,
      "postDate": "2023-01-17T11:17:41.303Z",
      "content": "<p>In this competition, I think most of the people use GBM model like a XGBoost or LightGBM.</p>\n<p>How to build a GBM model is explained in detail at the following discussion:<br>\n<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210</a></p>\n<p>I tried to use GBM model and I added \"session features\" and \"aid features\". <br>\nIn my model, aid features are effective but session features are not effective at all.<br>\nIn other words, feature importance of my session features are nearly zero whereas aid features have high importance.</p>\n<p>I'm not sure this situation is natural or not.</p>\n<p>For example, I generated session features as</p>\n<pre><code>user_features = train.groupby('session').agg({'session':'count','aid':'nunique','type':'mean'})\nuser_features.columns = ['user_user_count','user_item_count','user_buy_ratio']\n</code></pre>\n<p>,and I also generated aid features as</p>\n<pre><code>item_features = train.groupby('aid').agg({'aid':'count','session':'nunique','type':'mean'})\nitem_features.columns = ['item_item_count','item_user_count','item_buy_ratio']\n</code></pre>\n<p>.<br>\nIf possible, I want to know this situation is common for all users or only happened in my environment.</p>\n<h3>20220122</h3>\n<p>Thank you for sharing a lot of ideas and comments.<br>\nI found bug in my code and I finally found that my session feature can boost my model and LB.<br>\nIn my results, aid features are more effective than session features but both features have a potential to boost the score.</p>",
      "rawMarkdown": "In this competition, I think most of the people use GBM model like a XGBoost or LightGBM.\n\nHow to build a GBM model is explained in detail at the following discussion:\nhttps://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\n\nI tried to use GBM model and I added \"session features\" and \"aid features\". \nIn my model, aid features are effective but session features are not effective at all.\nIn other words, feature importance of my session features are nearly zero whereas aid features have high importance.\n\nI'm not sure this situation is natural or not.\n\nFor example, I generated session features as\n\n```\nuser_features = train.groupby('session').agg({'session':'count','aid':'nunique','type':'mean'})\nuser_features.columns = ['user_user_count','user_item_count','user_buy_ratio']\n```\n\n,and I also generated aid features as\n\n```\nitem_features = train.groupby('aid').agg({'aid':'count','session':'nunique','type':'mean'})\nitem_features.columns = ['item_item_count','item_user_count','item_buy_ratio']\n```\n.\nIf possible, I want to know this situation is common for all users or only happened in my environment.\n\n### 20220122\nThank you for sharing a lot of ideas and comments.\nI found bug in my code and I finally found that my session feature can boost my model and LB.\nIn my results, aid features are more effective than session features but both features have a potential to boost the score.",
      "votes": 2
    },
    {
      "id": 2105016,
      "postDate": "2023-01-18T07:53:55.987Z",
      "content": "<p>For me, session features seems not to improve the scores but the feature importance is not near zero.</p>",
      "rawMarkdown": "For me, session features seems not to improve the scores but the feature importance is not near zero.",
      "replies": [
        {
          "id": 2110162,
          "postDate": "2023-01-22T02:44:53.547Z",
          "content": "<p>Thank you for your comment.<br>\nI fixed the bug in my code and I think my final results are similar to yours.</p>",
          "rawMarkdown": "Thank you for your comment.\nI fixed the bug in my code and I think my final results are similar to yours."
        }
      ]
    },
    {
      "id": 2105755,
      "postDate": "2023-01-18T17:42:26.407Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2103770,
      "author_name": "Anil Ozturk",
      "author_url": "",
      "post_date": "2023-01-17T11:28:47.137000",
      "content": "<p>Since we are predicting the preference of the items by the users, the relationship of the session with that item or similar items may be more valuable for your model instead of general statistics about that session. Therefore, extracting user-interaction features that can represent not only the user but also represent the relationship between (user and candidate) (<code>groupby([\"session\", \"aid\"])</code>) would affect your score more. But it would still be beneficial for you to generate extra features representing the session details.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2103789,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "2023-01-17T11:42:37.280000",
          "content": "<p>Thank you for your comments!<br>\nI was surprised that the feature importance of the simple session features are almost zero when I tried to add the simple session features.<br>\nOn the other hand, several informative notebook say that we can generate session and aid features, which have a potential to boost <br>\nour model and LB score.<br>\nTherefore, I was confused by my results.</p>\n<p>Anyway, for now I guess it's not strange that the simple session information itself can't boost the score and we should consider the more effective interactive features. (if my understanding is correct)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2105579,
          "author_name": "kimoyami",
          "author_url": "",
          "post_date": "2023-01-18T16:06:54.713000",
          "content": "<p>Sorry, I don't understand about the user-interaction features? Could you give some simple examples. Thanks very much!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2105716,
              "author_name": "BUUMOO",
              "author_url": "",
              "post_date": "2023-01-18T17:01:29.550000",
              "content": "<p>for example:</p>\n<ol>\n<li>did user buy/cart/click the retrieved item?</li>\n<li>how many times the user buy/cart/click the retrieved item</li>\n</ol>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2106209,
              "author_name": "kimoyami",
              "author_url": "",
              "post_date": "2023-01-19T02:44:40.473000",
              "content": "<p>Thank you very much!!! I have another question. Is some mix-features necessary for prediction, maybe some features produced by groupby(['session', 'aid']), like:</p>\n<ol>\n<li>how many times item 'aid' does user 'session' buy/click/cart</li>\n<li>what's the ratio user 'session' buys/clicks/carts item 'aid'</li>\n</ol>\n<p>I found that many candidates will get NaN on these features. So I'm not sure whether the features like that are necessary. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2106275,
              "author_name": "BUUMOO",
              "author_url": "",
              "post_date": "2023-01-19T04:10:26.843000",
              "content": "<p>We can do a simple A/B test to check the performance of these 'suspicious' features since the training for orders should be very fast. Just drop the features you doubt,and train, to see if the validation score improves. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2114039,
              "author_name": "wcq_glhf",
              "author_url": "",
              "post_date": "2023-01-24T17:07:27.827000",
              "content": "<p>excuse me, i want to ask the questions of  training for orders, did you negative sample of the data, the 200 recall data is too slow to train even though for the orders, and the Memory is limited， but the CV of the negative sample of the data is very high and not equal to the lb, how to solve it? (maybe my user_item_feature is not userd and too waste time</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2105860,
          "author_name": "Parth Tiwary",
          "author_url": "",
          "post_date": "2023-01-18T19:34:36.203000",
          "content": "<p><a href=\"https://www.kaggle.com/nlztrk\" target=\"_blank\">@nlztrk</a> Thanks for the clarification. A follow up question: when you create interaction features - how do you merge it back onto the candidate dataframe? Is it on <code>('session', 'aid')</code> and how are you dealing with nulls in this case? (which I believe can be significant) I might be missing something here. Looking forward to your response :) </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2106772,
              "author_name": "Anil Ozturk",
              "author_url": "",
              "post_date": "2023-01-19T10:47:59.893000",
              "content": "<p>If a candidate is not seen during the history of that specific session, the user-interaction stats derived from that candidate would be <strong>NaN</strong>. You can simply fill them with value that has no possibility to be seen, like -1, -999 etc. And it's actually a pretty good information IMHO that indicates user had any interactions with that item or item group.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2107540,
              "author_name": "Parth Tiwary",
              "author_url": "",
              "post_date": "2023-01-19T21:00:50.723000",
              "content": "<p><a href=\"https://www.kaggle.com/nlztrk\" target=\"_blank\">@nlztrk</a> Thanks a ton for the response!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2103774,
      "author_name": "BUUMOO",
      "author_url": "",
      "post_date": "2023-01-17T11:33:38.963000",
      "content": "<p>For me it is not the case. Such as the features you named 'user_user_count','user_item_count'---namely the session length and session unique item numbers ranked within top 50% of all my features.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2103791,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "2023-01-17T11:45:27.640000",
          "content": "<p>Thank you for your nice information.<br>\nIf you are correct, I guess my training code includes some bugs..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2106536,
          "author_name": "Aji Samudra",
          "author_url": "",
          "post_date": "2023-01-19T08:05:47.773000",
          "content": "<p>my observation aligned with <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a>, these 2 features (session length &amp; session unique item) are ranked top in my ranker model.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2110163,
              "author_name": "tetsuro731",
              "author_url": "",
              "post_date": "2023-01-22T02:46:07.747000",
              "content": "<p>Thank you for your comment.<br>\nI fixed the bug in my code and my LB score had been improved.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2110550,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2023-01-22T09:12:16.873000",
          "content": "<p>That's the part where models learn whether your candidates are generated with recency weighting or unique aids + covisitation.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2105016,
      "author_name": "ADAM.",
      "author_url": "",
      "post_date": "2023-01-18T07:53:55.987000",
      "content": "<p>For me, session features seems not to improve the scores but the feature importance is not near zero.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2110162,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "2023-01-22T02:44:53.547000",
          "content": "<p>Thank you for your comment.<br>\nI fixed the bug in my code and I think my final results are similar to yours.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2105755,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-18T17:42:26.407000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2103770": "Since we are predicting the preference of the items by the users, the relationship of the session with that item or similar items may be more valuable for your model instead of general statistics about that session. Therefore, extracting user-interaction features that can represent not only the user but also represent the relationship between (user and candidate) (`groupby([\"session\", \"aid\"])`) would affect your score more. But it would still be beneficial for you to generate extra features representing the session details.",
    "2103774": "For me it is not the case. Such as the features you named 'user_user_count','user_item_count'---namely the session length and session unique item numbers ranked within top 50% of all my features.",
    "2103761": "In this competition, I think most of the people use GBM model like a XGBoost or LightGBM.\n\nHow to build a GBM model is explained in detail at the following discussion:\nhttps://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\n\nI tried to use GBM model and I added \"session features\" and \"aid features\". \nIn my model, aid features are effective but session features are not effective at all.\nIn other words, feature importance of my session features are nearly zero whereas aid features have high importance.\n\nI'm not sure this situation is natural or not.\n\nFor example, I generated session features as\n\n```\nuser_features = train.groupby('session').agg({'session':'count','aid':'nunique','type':'mean'})\nuser_features.columns = ['user_user_count','user_item_count','user_buy_ratio']\n```\n\n,and I also generated aid features as\n\n```\nitem_features = train.groupby('aid').agg({'aid':'count','session':'nunique','type':'mean'})\nitem_features.columns = ['item_item_count','item_user_count','item_buy_ratio']\n```\n.\nIf possible, I want to know this situation is common for all users or only happened in my environment.\n\n### 20220122\nThank you for sharing a lot of ideas and comments.\nI found bug in my code and I finally found that my session feature can boost my model and LB.\nIn my results, aid features are more effective than session features but both features have a potential to boost the score.",
    "2105016": "For me, session features seems not to improve the scores but the feature importance is not near zero.",
    "2105755": ""
  }
}