{
  "id": 366138,
  "title": "Locate Real Users and Real Sessions EDA",
  "url": "/competitions/otto-recommender-system/discussion/366138",
  "author_name": "",
  "post_date": "2022-11-15T01:46:00.857832200Z",
  "votes": 80,
  "comment_count": 12,
  "views": 0,
  "content": "<h1>Locate Real Users and Sessions!</h1>\n<p>In Kaggle's Otto competition the word \"session\" actually means \"user\". We are given train data for <code>12,899,779</code> users (not \"sessions\") during a 4 week period and we are given test data for users during 1 week (in the future). We must predict what a \"user\" will do in the remainder of the 1 week that we do not have information about.</p>\n<p>I published a notebook <a href=\"https://www.kaggle.com/code/cdeotte/time-series-eda-users-and-real-sessions\" target=\"_blank\">here</a> which displays users and their real sessions time series EDA. We can identify real sessions by looking for gaps between user behavior. The following code locates real user sessions. (Note that some Kaggle datasets have divided <code>ts</code> by <code>1000</code> therefore you need to remove the <code>1000</code> below with those datasets):</p>\n<pre><code>train = train.sort_values(['session','ts'])\ntrain['d'] = train.groupby('session').ts.diff()\ntrain.d = (train.d &gt; 1000*60*60*2).astype('int8').fillna(0)\ntrain['d'] = train.groupby('session').d.cumsum()\n</code></pre>\n<p>After running the code above, the column <code>d</code> contains the real session number for each user. For example, we can analyze real sessions with code like <code>train.groupby(['session','d'])</code> or we can compute number of sessions with code like <code>train.groupby('session').d.transform('max')</code>.</p>\n<p>After plotting real user sessions, we observe that users exhibit regular patterns of session behavior. These observations can help us characterize and engineer features for users. These observations can also give us insight into predicting future <code>click</code>, <code>cart</code>, and <code>order</code> user behavior. </p>\n<h1>Morning User</h1>\n<p>Below is an example of a \"morning person\":<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/morning2.png\" alt=\"\"></p>\n<h1>Evening User</h1>\n<p>Below is an example of a \"night person\":<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/evening2.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2029857",
      "postDate": "11/15/2022 01:46:00",
      "content": "<h1>Locate Real Users and Sessions!</h1>\n<p>In Kaggle's Otto competition the word \"session\" actually means \"user\". We are given train data for <code>12,899,779</code> users (not \"sessions\") during a 4 week period and we are given test data for users during 1 week (in the future). We must predict what a \"user\" will do in the remainder of the 1 week that we do not have information about.</p>\n<p>I published a notebook <a href=\"https://www.kaggle.com/code/cdeotte/time-series-eda-users-and-real-sessions\" target=\"_blank\">here</a> which displays users and their real sessions time series EDA. We can identify real sessions by looking for gaps between user behavior. The following code locates real user sessions. (Note that some Kaggle datasets have divided <code>ts</code> by <code>1000</code> therefore you need to remove the <code>1000</code> below with those datasets):</p>\n<pre><code>train = train.sort_values(['session','ts'])\ntrain['d'] = train.groupby('session').ts.diff()\ntrain.d = (train.d &gt; 1000*60*60*2).astype('int8').fillna(0)\ntrain['d'] = train.groupby('session').d.cumsum()\n</code></pre>\n<p>After running the code above, the column <code>d</code> contains the real session number for each user. For example, we can analyze real sessions with code like <code>train.groupby(['session','d'])</code> or we can compute number of sessions with code like <code>train.groupby('session').d.transform('max')</code>.</p>\n<p>After plotting real user sessions, we observe that users exhibit regular patterns of session behavior. These observations can help us characterize and engineer features for users. These observations can also give us insight into predicting future <code>click</code>, <code>cart</code>, and <code>order</code> user behavior. </p>\n<h1>Morning User</h1>\n<p>Below is an example of a \"morning person\":<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/morning2.png\" alt=\"\"></p>\n<h1>Evening User</h1>\n<p>Below is an example of a \"night person\":<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/evening2.png\" alt=\"\"></p>",
      "rawMarkdown": "# Locate Real Users and Sessions!\nIn Kaggle's Otto competition the word \"session\" actually means \"user\". We are given train data for `12,899,779` users (not \"sessions\") during a 4 week period and we are given test data for users during 1 week (in the future). We must predict what a \"user\" will do in the remainder of the 1 week that we do not have information about.\n\nI published a notebook [here][1] which displays users and their real sessions time series EDA. We can identify real sessions by looking for gaps between user behavior. The following code locates real user sessions. (Note that some Kaggle datasets have divided `ts` by `1000` therefore you need to remove the `1000` below with those datasets):\n\n    train = train.sort_values(['session','ts'])\n    train['d'] = train.groupby('session').ts.diff()\n    train.d = (train.d > 1000*60*60*2).astype('int8').fillna(0)\n    train['d'] = train.groupby('session').d.cumsum()\n\nAfter running the code above, the column `d` contains the real session number for each user. For example, we can analyze real sessions with code like `train.groupby(['session','d'])` or we can compute number of sessions with code like `train.groupby('session').d.transform('max')`.\n\nAfter plotting real user sessions, we observe that users exhibit regular patterns of session behavior. These observations can help us characterize and engineer features for users. These observations can also give us insight into predicting future `click`, `cart`, and `order` user behavior. \n\n# Morning User\nBelow is an example of a \"morning person\":\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/morning2.png)\n\n# Evening User\nBelow is an example of a \"night person\":\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/evening2.png)\n\n[1]: https://www.kaggle.com/code/cdeotte/time-series-eda-users-and-real-sessions",
      "votes": null
    },
    {
      "id": "2029952",
      "postDate": "11/15/2022 05:04:55",
      "content": "<p>Great work! Maybe we can use  different active periods to recall candidates.</p>",
      "rawMarkdown": "Great work! Maybe we can use  different active periods to recall candidates.",
      "votes": null
    },
    {
      "id": "2030145",
      "postDate": "11/15/2022 08:17:42",
      "content": "<p>Thanks for posting it! It's really helpful. </p>\n<p>I see two possible issues (based on how data was generated, not your code):</p>\n<ol>\n<li><p>With train-test split some real sessions may be divided in half. I'm not sure what to do with it.</p></li>\n<li><p>With splitting the test set to get labels done randomly we don't know whether the next events belong to the new real session or are continuation of the previous one. I guess it's a place for various candidates generation processes (user-level and session-level).</p></li>\n</ol>\n<p>Thanks again for posting the topic and the code!  </p>",
      "rawMarkdown": "Thanks for posting it! It's really helpful. \n\nI see two possible issues (based on how data was generated, not your code):\n\n1. With train-test split some real sessions may be divided in half. I'm not sure what to do with it.\n\n2. With splitting the test set to get labels done randomly we don't know whether the next events belong to the new real session or are continuation of the previous one. I guess it's a place for various candidates generation processes (user-level and session-level).\n\nThanks again for posting the topic and the code!",
      "votes": null
    },
    {
      "id": "2030164",
      "postDate": "11/15/2022 08:36:15",
      "content": "<p>Great stuff! Such analyses often lead to great feature engineering ideas</p>",
      "rawMarkdown": "Great stuff! Such analyses often lead to great feature engineering ideas",
      "votes": null
    },
    {
      "id": "2030402",
      "postDate": "11/15/2022 11:44:59",
      "content": "<p>Nice visualization ! We can create user features like \"night person\" or \"daytime person\" ☀️</p>",
      "rawMarkdown": "Nice visualization ! We can create user features like \"night person\" or \"daytime person\" ☀️",
      "votes": null
    },
    {
      "id": "2054611",
      "postDate": "12/04/2022 09:29:17",
      "content": "<p>I found there's no intersection between the train and test sets in <code>session</code>. And we have 4 weeks data in train and 1 week data in test. How do you split the data to generate features and labels and validate the model performance?</p>",
      "rawMarkdown": "I found there's no intersection between the train and test sets in `session`. And we have 4 weeks data in train and 1 week data in test. How do you split the data to generate features and labels and validate the model performance?",
      "votes": null
    },
    {
      "id": "2054624",
      "postDate": "12/04/2022 09:40:53",
      "content": "<p>For validation, we make a new train and new validation dataset. New train is the first 3 weeks of Kaggle train. And new validation is the last 1 week of Kaggle train. See Radek's discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">here</a>.</p>\n<p>We do not need overlap with users because our models do not use <code>user id</code>. Our models only use <code>user features</code>. And the user features are common between train and validation. (And train and test).</p>",
      "rawMarkdown": "For validation, we make a new train and new validation dataset. New train is the first 3 weeks of Kaggle train. And new validation is the last 1 week of Kaggle train. See Radek's discussion [here][1].\n\nWe do not need overlap with users because our models do not use `user id`. Our models only use `user features`. And the user features are common between train and validation. (And train and test).\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991",
      "votes": null
    },
    {
      "id": "2054630",
      "postDate": "12/04/2022 09:51:47",
      "content": "<p>Make sense, thanks for the detailed explanation.</p>",
      "rawMarkdown": "Make sense, thanks for the detailed explanation.",
      "votes": null
    },
    {
      "id": "2055812",
      "postDate": "12/05/2022 13:08:09",
      "content": "<p>Chris, each user may have multiple sessions, right? so if we need to build user features, the only way we could do is like your mention before that finding which seesions are belong to which users, and then make user features based on the generated 'use_id'?  </p>",
      "rawMarkdown": "Chris, each user may have multiple sessions, right? so if we need to build user features, the only way we could do is like your mention before that finding which seesions are belong to which users, and then make user features based on the generated 'use_id'?",
      "votes": null
    },
    {
      "id": "2055832",
      "postDate": "12/05/2022 13:27:15",
      "content": "<p>We have a choice to either locate real sessions or ignore real sessions when building user features. For example, here is a user feature that ignores real sessions. It is  \"the average buy proportion for all time\": </p>\n<pre><code>df['buy'] = (df['type'] == 'orders').astype('int8')\ndf['mean_buy_ratio'] = df.groupby('session').transform.agg('mean') \n</code></pre>",
      "rawMarkdown": "We have a choice to either locate real sessions or ignore real sessions when building user features. For example, here is a user feature that ignores real sessions. It is  \"the average buy proportion for all time\": \n\n    df['buy'] = (df['type'] == 'orders').astype('int8')\n    df['mean_buy_ratio'] = df.groupby('session').transform.agg('mean')",
      "votes": null
    },
    {
      "id": "2055872",
      "postDate": "12/05/2022 14:02:28",
      "content": "<p>Thanks for explaining. It's helpful. Chris.</p>",
      "rawMarkdown": "Thanks for explaining. It's helpful. Chris.",
      "votes": null
    },
    {
      "id": "2055882",
      "postDate": "12/05/2022 14:09:55",
      "content": "<p><a href=\"https://www.kaggle.com/alvinai9603\" target=\"_blank\">@alvinai9603</a> Here is an example that requires real sessions. Let <code>d</code> indicate the real session number. The following user feature is the average time a user waits between clicks while they are actively browsing the website:</p>\n<pre><code>df = df.sort_values(['session','ts'])    \ndf['average_delta_time'] = df.groupby(['session','d']).ts.diff()\ndf['average_delta_time'] = df.groupby('session').average_delta_time.transform('mean')\n</code></pre>\n<p>This works because by using <code>groupby(['session','d'])</code> then the diff for the first item in each session is NAN. When we later compute the mean, it ignore NAN. Therefore the feature <code>average_delta_time</code> will only include the time between user clicks when a user is actively browsing the website.</p>",
      "rawMarkdown": "alvinai9603 Here is an example that requires real sessions. Let `d` indicate the real session number. The following user feature is the average time a user waits between clicks while they are actively browsing the website:\n\n    df = df.sort_values(['session','ts'])    \n    df['average_delta_time'] = df.groupby(['session','d']).ts.diff()\n    df['average_delta_time'] = df.groupby('session').average_delta_time.transform('mean')\n\nThis works because by using `groupby(['session','d'])` then the diff for the first item in each session is NAN. When we later compute the mean, it ignore NAN. Therefore the feature `average_delta_time` will only include the time between user clicks when a user is actively browsing the website.",
      "votes": null
    },
    {
      "id": "2055902",
      "postDate": "12/05/2022 14:32:11",
      "content": "<p>This is great. Look like we could do many feature works in the following days to improve our rank model. Thanks again.</p>",
      "rawMarkdown": "This is great. Look like we could do many feature works in the following days to improve our rank model. Thanks again.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2029952,
      "author_name": "jiahongxie",
      "author_url": "",
      "post_date": "11/15/2022 05:04:55",
      "content": "<p>Great work! Maybe we can use  different active periods to recall candidates.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2030145,
      "author_name": "piotrekga",
      "author_url": "",
      "post_date": "11/15/2022 08:17:42",
      "content": "<p>Thanks for posting it! It's really helpful. </p>\n<p>I see two possible issues (based on how data was generated, not your code):</p>\n<ol>\n<li><p>With train-test split some real sessions may be divided in half. I'm not sure what to do with it.</p></li>\n<li><p>With splitting the test set to get labels done randomly we don't know whether the next events belong to the new real session or are continuation of the previous one. I guess it's a place for various candidates generation processes (user-level and session-level).</p></li>\n</ol>\n<p>Thanks again for posting the topic and the code!  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2030164,
      "author_name": "narsil",
      "author_url": "",
      "post_date": "11/15/2022 08:36:15",
      "content": "<p>Great stuff! Such analyses often lead to great feature engineering ideas</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2030402,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "11/15/2022 11:44:59",
      "content": "<p>Nice visualization ! We can create user features like \"night person\" or \"daytime person\" ☀️</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2054611,
      "author_name": "wuwenmin",
      "author_url": "",
      "post_date": "12/04/2022 09:29:17",
      "content": "<p>I found there's no intersection between the train and test sets in <code>session</code>. And we have 4 weeks data in train and 1 week data in test. How do you split the data to generate features and labels and validate the model performance?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2054624,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "12/04/2022 09:40:53",
          "content": "<p>For validation, we make a new train and new validation dataset. New train is the first 3 weeks of Kaggle train. And new validation is the last 1 week of Kaggle train. See Radek's discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">here</a>.</p>\n<p>We do not need overlap with users because our models do not use <code>user id</code>. Our models only use <code>user features</code>. And the user features are common between train and validation. (And train and test).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2054630,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "12/04/2022 09:51:47",
          "content": "<p>Make sense, thanks for the detailed explanation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2055812,
          "author_name": "alvinai9603",
          "author_url": "",
          "post_date": "12/05/2022 13:08:09",
          "content": "<p>Chris, each user may have multiple sessions, right? so if we need to build user features, the only way we could do is like your mention before that finding which seesions are belong to which users, and then make user features based on the generated 'use_id'?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2055832,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "12/05/2022 13:27:15",
          "content": "<p>We have a choice to either locate real sessions or ignore real sessions when building user features. For example, here is a user feature that ignores real sessions. It is  \"the average buy proportion for all time\": </p>\n<pre><code>df['buy'] = (df['type'] == 'orders').astype('int8')\ndf['mean_buy_ratio'] = df.groupby('session').transform.agg('mean') \n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2055872,
          "author_name": "alvinai9603",
          "author_url": "",
          "post_date": "12/05/2022 14:02:28",
          "content": "<p>Thanks for explaining. It's helpful. Chris.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2055882,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "12/05/2022 14:09:55",
          "content": "<p><a href=\"https://www.kaggle.com/alvinai9603\" target=\"_blank\">@alvinai9603</a> Here is an example that requires real sessions. Let <code>d</code> indicate the real session number. The following user feature is the average time a user waits between clicks while they are actively browsing the website:</p>\n<pre><code>df = df.sort_values(['session','ts'])    \ndf['average_delta_time'] = df.groupby(['session','d']).ts.diff()\ndf['average_delta_time'] = df.groupby('session').average_delta_time.transform('mean')\n</code></pre>\n<p>This works because by using <code>groupby(['session','d'])</code> then the diff for the first item in each session is NAN. When we later compute the mean, it ignore NAN. Therefore the feature <code>average_delta_time</code> will only include the time between user clicks when a user is actively browsing the website.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2055902,
          "author_name": "alvinai9603",
          "author_url": "",
          "post_date": "12/05/2022 14:32:11",
          "content": "<p>This is great. Look like we could do many feature works in the following days to improve our rank model. Thanks again.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2029857": "# Locate Real Users and Sessions!\nIn Kaggle's Otto competition the word \"session\" actually means \"user\". We are given train data for `12,899,779` users (not \"sessions\") during a 4 week period and we are given test data for users during 1 week (in the future). We must predict what a \"user\" will do in the remainder of the 1 week that we do not have information about.\n\nI published a notebook [here][1] which displays users and their real sessions time series EDA. We can identify real sessions by looking for gaps between user behavior. The following code locates real user sessions. (Note that some Kaggle datasets have divided `ts` by `1000` therefore you need to remove the `1000` below with those datasets):\n\n    train = train.sort_values(['session','ts'])\n    train['d'] = train.groupby('session').ts.diff()\n    train.d = (train.d > 1000*60*60*2).astype('int8').fillna(0)\n    train['d'] = train.groupby('session').d.cumsum()\n\nAfter running the code above, the column `d` contains the real session number for each user. For example, we can analyze real sessions with code like `train.groupby(['session','d'])` or we can compute number of sessions with code like `train.groupby('session').d.transform('max')`.\n\nAfter plotting real user sessions, we observe that users exhibit regular patterns of session behavior. These observations can help us characterize and engineer features for users. These observations can also give us insight into predicting future `click`, `cart`, and `order` user behavior. \n\n# Morning User\nBelow is an example of a \"morning person\":\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/morning2.png)\n\n# Evening User\nBelow is an example of a \"night person\":\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/evening2.png)\n\n[1]: https://www.kaggle.com/code/cdeotte/time-series-eda-users-and-real-sessions",
    "2029952": "Great work! Maybe we can use  different active periods to recall candidates.",
    "2030145": "Thanks for posting it! It's really helpful. \n\nI see two possible issues (based on how data was generated, not your code):\n\n1. With train-test split some real sessions may be divided in half. I'm not sure what to do with it.\n\n2. With splitting the test set to get labels done randomly we don't know whether the next events belong to the new real session or are continuation of the previous one. I guess it's a place for various candidates generation processes (user-level and session-level).\n\nThanks again for posting the topic and the code!",
    "2030164": "Great stuff! Such analyses often lead to great feature engineering ideas",
    "2030402": "Nice visualization ! We can create user features like \"night person\" or \"daytime person\" ☀️",
    "2054611": "I found there's no intersection between the train and test sets in `session`. And we have 4 weeks data in train and 1 week data in test. How do you split the data to generate features and labels and validate the model performance?",
    "2054624": "For validation, we make a new train and new validation dataset. New train is the first 3 weeks of Kaggle train. And new validation is the last 1 week of Kaggle train. See Radek's discussion [here][1].\n\nWe do not need overlap with users because our models do not use `user id`. Our models only use `user features`. And the user features are common between train and validation. (And train and test).\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991",
    "2054630": "Make sense, thanks for the detailed explanation.",
    "2055812": "Chris, each user may have multiple sessions, right? so if we need to build user features, the only way we could do is like your mention before that finding which seesions are belong to which users, and then make user features based on the generated 'use_id'?",
    "2055832": "We have a choice to either locate real sessions or ignore real sessions when building user features. For example, here is a user feature that ignores real sessions. It is  \"the average buy proportion for all time\": \n\n    df['buy'] = (df['type'] == 'orders').astype('int8')\n    df['mean_buy_ratio'] = df.groupby('session').transform.agg('mean')",
    "2055872": "Thanks for explaining. It's helpful. Chris.",
    "2055882": "alvinai9603 Here is an example that requires real sessions. Let `d` indicate the real session number. The following user feature is the average time a user waits between clicks while they are actively browsing the website:\n\n    df = df.sort_values(['session','ts'])    \n    df['average_delta_time'] = df.groupby(['session','d']).ts.diff()\n    df['average_delta_time'] = df.groupby('session').average_delta_time.transform('mean')\n\nThis works because by using `groupby(['session','d'])` then the diff for the first item in each session is NAN. When we later compute the mean, it ignore NAN. Therefore the feature `average_delta_time` will only include the time between user clicks when a user is actively browsing the website.",
    "2055902": "This is great. Look like we could do many feature works in the following days to improve our rank model. Thanks again."
  },
  "source": "meta"
}