{
  "id": 367154,
  "title": "Dumb Question: so what do we actually need to predict?",
  "url": "/competitions/otto-recommender-system/discussion/367154",
  "author_name": "",
  "post_date": "2022-11-19T11:31:00.011512400Z",
  "votes": 34,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I am a bit confused and I cannot find my answer in the Discussion posts, which probably means my question is very dumb, but here I go.</p>\n<p>So in train we have 200,000 unique sessions, with varying event lengths.</p>\n<p>In the test set we have another 200,000 unique sessions (no session in the test set it also in the train set - right?)</p>\n<p>In the test set there are on average 5 events per session (that are truncated in time, so they were extracted from longer sessions I presume? like these 5 events are subsets of a longer set of events?).</p>\n<p>And the idea (after we looked at previous sessions in the train data and identified some kind of pattern) is to predict whatever these other new users in the test set will do AFTER the shown activity in the test set?</p>\n<p>In other words, if I have session 9999 in test set and I have 5 events for it (measured in clicks let's say), I can use my model + these 5 previous events as guidelines to predict for session 9999 up to 20 values for each type of event?</p>\n<p>Is this correct?</p>",
  "messages": [
    {
      "id": "2035991",
      "postDate": "11/19/2022 11:31:00",
      "content": "<p>I am a bit confused and I cannot find my answer in the Discussion posts, which probably means my question is very dumb, but here I go.</p>\n<p>So in train we have 200,000 unique sessions, with varying event lengths.</p>\n<p>In the test set we have another 200,000 unique sessions (no session in the test set it also in the train set - right?)</p>\n<p>In the test set there are on average 5 events per session (that are truncated in time, so they were extracted from longer sessions I presume? like these 5 events are subsets of a longer set of events?).</p>\n<p>And the idea (after we looked at previous sessions in the train data and identified some kind of pattern) is to predict whatever these other new users in the test set will do AFTER the shown activity in the test set?</p>\n<p>In other words, if I have session 9999 in test set and I have 5 events for it (measured in clicks let's say), I can use my model + these 5 previous events as guidelines to predict for session 9999 up to 20 values for each type of event?</p>\n<p>Is this correct?</p>",
      "rawMarkdown": "I am a bit confused and I cannot find my answer in the Discussion posts, which probably means my question is very dumb, but here I go.\n\nSo in train we have 200,000 unique sessions, with varying event lengths.\n\nIn the test set we have another 200,000 unique sessions (no session in the test set it also in the train set - right?)\n\nIn the test set there are on average 5 events per session (that are truncated in time, so they were extracted from longer sessions I presume? like these 5 events are subsets of a longer set of events?).\n\nAnd the idea (after we looked at previous sessions in the train data and identified some kind of pattern) is to predict whatever these other new users in the test set will do AFTER the shown activity in the test set?\n\nIn other words, if I have session 9999 in test set and I have 5 events for it (measured in clicks let's say), I can use my model + these 5 previous events as guidelines to predict for session 9999 up to 20 values for each type of event?\n\nIs this correct?",
      "votes": null
    },
    {
      "id": "2036145",
      "postDate": "11/19/2022 13:57:03",
      "content": "<p>Yes</p>\n<p>(Your message must have at least 10 characters.)</p>",
      "rawMarkdown": "Yes\n\n(Your message must have at least 10 characters.)",
      "votes": null
    },
    {
      "id": "2036172",
      "postDate": "11/19/2022 14:16:24",
      "content": "<p>Not a dumb question. I am sure that a bunch of people spent a while trying to figure out the same thing. This is all about learning! :) </p>",
      "rawMarkdown": "Not a dumb question. I am sure that a bunch of people spent a while trying to figure out the same thing. This is all about learning! :)",
      "votes": null
    },
    {
      "id": "2036201",
      "postDate": "11/19/2022 14:43:36",
      "content": "<p>Also, when you are doing a machine learning project for work, one of the first things you do is clarify the goal. I remember being in a meeting at the beginning of a project and I asked \"ok, so what is the definition of [goal] because to me it means…\" We discovered that none of us had the same definition of [goal]. It is better to find out at the beginning of the project than in the middle. </p>",
      "rawMarkdown": "Also, when you are doing a machine learning project for work, one of the first things you do is clarify the goal. I remember being in a meeting at the beginning of a project and I asked \"ok, so what is the definition of [goal] because to me it means...\" We discovered that none of us had the same definition of [goal]. It is better to find out at the beginning of the project than in the middle.",
      "votes": null
    },
    {
      "id": "2037767",
      "postDate": "11/20/2022 23:22:24",
      "content": "<p><a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> <br>\nAs I understand:</p>\n<ul>\n<li>the test sessions are different from the train sessions</li>\n<li>they may be part of a longer sessions; however, there is no official indication that test sessions are part of a longer session so we will have to address this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367495\" target=\"_blank\">\"cold start problem\"</a></li>\n<li>Ultimately, we are predicting up to 20 of the next items we think they will click/cart/order</li>\n<li>In your last example, I would feature engineer a matrix of correlations or clusters between users or items in the train set. Then I agree that you should use the 5 given events and using some sort of model and that matrix to make your predictions.</li>\n</ul>",
      "rawMarkdown": "andradaolteanu \nAs I understand:\n- the test sessions are different from the train sessions\n- they may be part of a longer sessions; however, there is no official indication that test sessions are part of a longer session so we will have to address this [\"cold start problem\"](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367495)\n- Ultimately, we are predicting up to 20 of the next items we think they will click/cart/order\n- In your last example, I would feature engineer a matrix of correlations or clusters between users or items in the train set. Then I agree that you should use the 5 given events and using some sort of model and that matrix to make your predictions.",
      "votes": null
    },
    {
      "id": "2037877",
      "postDate": "11/21/2022 02:34:56",
      "content": "<blockquote>\n  <p>they may be part of a longer sessions; however, there is no official indication that test sessions are part of a longer session so we will have to address this \"cold start problem\"</p>\n</blockquote>\n<p>The test data for each user (note \"session\" means \"user\" in this comp), is most likely <strong>half</strong> of that user's activity in the week of Aug 29th 2022. We can read the code from Otto's GitHub and we see that the test data was created by taking the user's full data (from the week of Aug 29th 2022) and randomly splitting it using <code>np.random.uniform</code>. Therefore the expected value is that we have <strong>half</strong> of the user's data (during the week of Aug 29th 2022).</p>\n<p>And our task is to predict the missing <strong>half</strong> of user data.</p>",
      "rawMarkdown": ">they may be part of a longer sessions; however, there is no official indication that test sessions are part of a longer session so we will have to address this \"cold start problem\"\n\nThe test data for each user (note \"session\" means \"user\" in this comp), is most likely **half** of that user's activity in the week of Aug 29th 2022. We can read the code from Otto's GitHub and we see that the test data was created by taking the user's full data (from the week of Aug 29th 2022) and randomly splitting it using `np.random.uniform`. Therefore the expected value is that we have **half** of the user's data (during the week of Aug 29th 2022).\n\nAnd our task is to predict the missing **half** of user data.",
      "votes": null
    },
    {
      "id": "2047739",
      "postDate": "11/29/2022 02:41:00",
      "content": "<p>Hi, I still don't get it. <br>\nHoping for discussions.</p>\n<ol>\n<li>If the sessions in test data are randomly split using<code>np.random.uniform</code>, I think there is data leaking problem.</li>\n<li>Why does the host of this competition make the 'userId' unavailable and use 'sessionId'?  Is this for privacy reasons?</li>\n</ol>",
      "rawMarkdown": "Hi, I still don't get it. \nHoping for discussions.\n1. If the sessions in test data are randomly split using` np.random.uniform`, I think there is data leaking problem.\n2. Why does the host of this competition make the 'userId' unavailable and use 'sessionId'?  Is this for privacy reasons?",
      "votes": null
    },
    {
      "id": "2048918",
      "postDate": "11/29/2022 19:23:06",
      "content": "<p>I agree with you <a href=\"https://www.kaggle.com/Megan\" target=\"_blank\">@Megan</a>, this is not a dumb question. I am one of these bunch of people</p>",
      "rawMarkdown": "I agree with you @Megan, this is not a dumb question. I am one of these bunch of people",
      "votes": null
    },
    {
      "id": "2051038",
      "postDate": "12/01/2022 06:58:37",
      "content": "<p>I have the same question</p>",
      "rawMarkdown": "I have the same question",
      "votes": null
    },
    {
      "id": "2051226",
      "postDate": "12/01/2022 09:38:13",
      "content": "<p>Hi luojinwei,</p>\n<p>one sessionId includes all interactions one user had with our shop during the whole period the trainset/testset covers. So it is basically a userId. </p>",
      "rawMarkdown": "Hi luojinwei,\n\none sessionId includes all interactions one user had with our shop during the whole period the trainset/testset covers. So it is basically a userId.",
      "votes": null
    },
    {
      "id": "2051813",
      "postDate": "12/01/2022 16:24:05",
      "content": "<p>Thanks for your question, i finally figure out what i need to do in this competition.😂</p>",
      "rawMarkdown": "Thanks for your question, i finally figure out what i need to do in this competition.😂",
      "votes": null
    },
    {
      "id": "2051862",
      "postDate": "12/01/2022 16:55:15",
      "content": "<p>okokok then maybe it's not that dumb 😂</p>",
      "rawMarkdown": "okokok then maybe it's not that dumb 😂",
      "votes": null
    },
    {
      "id": "2058265",
      "postDate": "12/07/2022 18:39:20",
      "content": "<p>Thank you for clarifying it.</p>",
      "rawMarkdown": "Thank you for clarifying it.",
      "votes": null
    },
    {
      "id": "2091947",
      "postDate": "01/08/2023 21:54:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>,</p>\n<p>I searched Otto's GitHub but didn't find any \"np.random.uniform\" and I couldn't find out how the test set was split in half.<br>\nI have two questions: </p>\n<ol>\n<li>I saw your reply <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939#2072959\" target=\"_blank\">here</a>, do you mean that each user's test data is split into a first half and a second half by time?</li>\n<li>From what point in time are we trying to predict events after? From the beginning of week 5, i.e. the whole week 5 event, or after the end of the test data we can see?</li>\n</ol>\n<p>I thought at first that the test data we got already specified half of the events, so all we needed to do was write those events correctly to the label of the CSV to get a recall of about 0.5, but in reality, I only got lb 0.061. so I found my understanding was incorrect.</p>",
      "rawMarkdown": "Hi @cdeotte,\n\nI searched Otto's GitHub but didn't find any \"np.random.uniform\" and I couldn't find out how the test set was split in half.\nI have two questions: \n\n1. I saw your reply [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939#2072959), do you mean that each user's test data is split into a first half and a second half by time?\n2. From what point in time are we trying to predict events after? From the beginning of week 5, i.e. the whole week 5 event, or after the end of the test data we can see?\n\nI thought at first that the test data we got already specified half of the events, so all we needed to do was write those events correctly to the label of the CSV to get a recall of about 0.5, but in reality, I only got lb 0.061. so I found my understanding was incorrect.",
      "votes": null
    },
    {
      "id": "2092003",
      "postDate": "01/09/2023 00:13:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chenlin1999\" target=\"_blank\">@chenlin1999</a>,<br>\nThe test data contains sessions truncated by timestamp, and you are to predict what occurs after the point of truncation. The timestamp that splits the session was chosen randomly. <br>\nSome basic information about the test sessions:<br>\nAll test sessions contain only data from week 5.<br>\nThe session ids are different from the session ids of the train data.<br>\nSessions with many events are rare compared with sessions with few events.</p>\n<p>Some examples:<br>\nA) A session has 100 events. These events happen on Monday, Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday of week 5. The session is truncated after Wednesday. So you have the events from Monday, Tuesday, and Wednesday to predict the events from Thursday, Friday, Sunday and Saturday.<br>\nB) A session has 100 events. These events happen on Monday, Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday of week 5. The session is truncated after Sunday. So you have the events from Monday, Tuesday, Wednesday, Thursday, Friday and Sunday to predict the events from Saturday.<br>\nC) A session has 100 events. These events happen on Monday, Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday of week 5. The session is truncated after Monday. So you have the events from Monday to predict the events from Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday.<br>\nD) A session has three events. Two of these events happen on Monday and the last one on Saturday of week 5. The session is truncated after Monday. So you have the two events from Monday to predict the events from Saturday.</p>\n<p>I hope this clears things up a little.</p>",
      "rawMarkdown": "Hi @chenlin1999,\nThe test data contains sessions truncated by timestamp, and you are to predict what occurs after the point of truncation. The timestamp that splits the session was chosen randomly. \nSome basic information about the test sessions:\nAll test sessions contain only data from week 5.\nThe session ids are different from the session ids of the train data.\nSessions with many events are rare compared with sessions with few events.\n\nSome examples:\nA) A session has 100 events. These events happen on Monday, Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday of week 5. The session is truncated after Wednesday. So you have the events from Monday, Tuesday, and Wednesday to predict the events from Thursday, Friday, Sunday and Saturday.\nB) A session has 100 events. These events happen on Monday, Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday of week 5. The session is truncated after Sunday. So you have the events from Monday, Tuesday, Wednesday, Thursday, Friday and Sunday to predict the events from Saturday.\nC) A session has 100 events. These events happen on Monday, Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday of week 5. The session is truncated after Monday. So you have the events from Monday to predict the events from Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday.\nD) A session has three events. Two of these events happen on Monday and the last one on Saturday of week 5. The session is truncated after Monday. So you have the two events from Monday to predict the events from Saturday.\n\nI hope this clears things up a little.",
      "votes": null
    },
    {
      "id": "2093279",
      "postDate": "01/09/2023 23:25:50",
      "content": "<p>Thank you very much, now I understand completely</p>",
      "rawMarkdown": "Thank you very much, now I understand completely",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2036145,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "11/19/2022 13:57:03",
      "content": "<p>Yes</p>\n<p>(Your message must have at least 10 characters.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2036172,
      "author_name": "megan3",
      "author_url": "",
      "post_date": "11/19/2022 14:16:24",
      "content": "<p>Not a dumb question. I am sure that a bunch of people spent a while trying to figure out the same thing. This is all about learning! :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 2048918,
          "author_name": "leiwong",
          "author_url": "",
          "post_date": "11/29/2022 19:23:06",
          "content": "<p>I agree with you <a href=\"https://www.kaggle.com/Megan\" target=\"_blank\">@Megan</a>, this is not a dumb question. I am one of these bunch of people</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2036201,
      "author_name": "megan3",
      "author_url": "",
      "post_date": "11/19/2022 14:43:36",
      "content": "<p>Also, when you are doing a machine learning project for work, one of the first things you do is clarify the goal. I remember being in a meeting at the beginning of a project and I asked \"ok, so what is the definition of [goal] because to me it means…\" We discovered that none of us had the same definition of [goal]. It is better to find out at the beginning of the project than in the middle. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2037767,
      "author_name": "ravishah1",
      "author_url": "",
      "post_date": "11/20/2022 23:22:24",
      "content": "<p><a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> <br>\nAs I understand:</p>\n<ul>\n<li>the test sessions are different from the train sessions</li>\n<li>they may be part of a longer sessions; however, there is no official indication that test sessions are part of a longer session so we will have to address this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367495\" target=\"_blank\">\"cold start problem\"</a></li>\n<li>Ultimately, we are predicting up to 20 of the next items we think they will click/cart/order</li>\n<li>In your last example, I would feature engineer a matrix of correlations or clusters between users or items in the train set. Then I agree that you should use the 5 given events and using some sort of model and that matrix to make your predictions.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2037877,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/21/2022 02:34:56",
          "content": "<blockquote>\n  <p>they may be part of a longer sessions; however, there is no official indication that test sessions are part of a longer session so we will have to address this \"cold start problem\"</p>\n</blockquote>\n<p>The test data for each user (note \"session\" means \"user\" in this comp), is most likely <strong>half</strong> of that user's activity in the week of Aug 29th 2022. We can read the code from Otto's GitHub and we see that the test data was created by taking the user's full data (from the week of Aug 29th 2022) and randomly splitting it using <code>np.random.uniform</code>. Therefore the expected value is that we have <strong>half</strong> of the user's data (during the week of Aug 29th 2022).</p>\n<p>And our task is to predict the missing <strong>half</strong> of user data.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2091947,
              "author_name": "chenlin1999",
              "author_url": "",
              "post_date": "01/08/2023 21:54:01",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>,</p>\n<p>I searched Otto's GitHub but didn't find any \"np.random.uniform\" and I couldn't find out how the test set was split in half.<br>\nI have two questions: </p>\n<ol>\n<li>I saw your reply <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939#2072959\" target=\"_blank\">here</a>, do you mean that each user's test data is split into a first half and a second half by time?</li>\n<li>From what point in time are we trying to predict events after? From the beginning of week 5, i.e. the whole week 5 event, or after the end of the test data we can see?</li>\n</ol>\n<p>I thought at first that the test data we got already specified half of the events, so all we needed to do was write those events correctly to the label of the CSV to get a recall of about 0.5, but in reality, I only got lb 0.061. so I found my understanding was incorrect.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2092003,
                  "author_name": "andreaswand",
                  "author_url": "",
                  "post_date": "01/09/2023 00:13:29",
                  "content": "<p>Hi <a href=\"https://www.kaggle.com/chenlin1999\" target=\"_blank\">@chenlin1999</a>,<br>\nThe test data contains sessions truncated by timestamp, and you are to predict what occurs after the point of truncation. The timestamp that splits the session was chosen randomly. <br>\nSome basic information about the test sessions:<br>\nAll test sessions contain only data from week 5.<br>\nThe session ids are different from the session ids of the train data.<br>\nSessions with many events are rare compared with sessions with few events.</p>\n<p>Some examples:<br>\nA) A session has 100 events. These events happen on Monday, Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday of week 5. The session is truncated after Wednesday. So you have the events from Monday, Tuesday, and Wednesday to predict the events from Thursday, Friday, Sunday and Saturday.<br>\nB) A session has 100 events. These events happen on Monday, Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday of week 5. The session is truncated after Sunday. So you have the events from Monday, Tuesday, Wednesday, Thursday, Friday and Sunday to predict the events from Saturday.<br>\nC) A session has 100 events. These events happen on Monday, Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday of week 5. The session is truncated after Monday. So you have the events from Monday to predict the events from Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday.<br>\nD) A session has three events. Two of these events happen on Monday and the last one on Saturday of week 5. The session is truncated after Monday. So you have the two events from Monday to predict the events from Saturday.</p>\n<p>I hope this clears things up a little.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2093279,
                      "author_name": "chenlin1999",
                      "author_url": "",
                      "post_date": "01/09/2023 23:25:50",
                      "content": "<p>Thank you very much, now I understand completely</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2047739,
          "author_name": "luojinwei",
          "author_url": "",
          "post_date": "11/29/2022 02:41:00",
          "content": "<p>Hi, I still don't get it. <br>\nHoping for discussions.</p>\n<ol>\n<li>If the sessions in test data are randomly split using<code>np.random.uniform</code>, I think there is data leaking problem.</li>\n<li>Why does the host of this competition make the 'userId' unavailable and use 'sessionId'?  Is this for privacy reasons?</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2051226,
          "author_name": "andreaswand",
          "author_url": "",
          "post_date": "12/01/2022 09:38:13",
          "content": "<p>Hi luojinwei,</p>\n<p>one sessionId includes all interactions one user had with our shop during the whole period the trainset/testset covers. So it is basically a userId. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2058265,
          "author_name": "leiwong",
          "author_url": "",
          "post_date": "12/07/2022 18:39:20",
          "content": "<p>Thank you for clarifying it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2051038,
      "author_name": "xhy2phd2016",
      "author_url": "",
      "post_date": "12/01/2022 06:58:37",
      "content": "<p>I have the same question</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2051813,
      "author_name": "rhine97",
      "author_url": "",
      "post_date": "12/01/2022 16:24:05",
      "content": "<p>Thanks for your question, i finally figure out what i need to do in this competition.😂</p>",
      "votes": null,
      "replies": [
        {
          "id": 2051862,
          "author_name": "andradaolteanu",
          "author_url": "",
          "post_date": "12/01/2022 16:55:15",
          "content": "<p>okokok then maybe it's not that dumb 😂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2035991": "I am a bit confused and I cannot find my answer in the Discussion posts, which probably means my question is very dumb, but here I go.\n\nSo in train we have 200,000 unique sessions, with varying event lengths.\n\nIn the test set we have another 200,000 unique sessions (no session in the test set it also in the train set - right?)\n\nIn the test set there are on average 5 events per session (that are truncated in time, so they were extracted from longer sessions I presume? like these 5 events are subsets of a longer set of events?).\n\nAnd the idea (after we looked at previous sessions in the train data and identified some kind of pattern) is to predict whatever these other new users in the test set will do AFTER the shown activity in the test set?\n\nIn other words, if I have session 9999 in test set and I have 5 events for it (measured in clicks let's say), I can use my model + these 5 previous events as guidelines to predict for session 9999 up to 20 values for each type of event?\n\nIs this correct?",
    "2036145": "Yes\n\n(Your message must have at least 10 characters.)",
    "2036172": "Not a dumb question. I am sure that a bunch of people spent a while trying to figure out the same thing. This is all about learning! :)",
    "2036201": "Also, when you are doing a machine learning project for work, one of the first things you do is clarify the goal. I remember being in a meeting at the beginning of a project and I asked \"ok, so what is the definition of [goal] because to me it means...\" We discovered that none of us had the same definition of [goal]. It is better to find out at the beginning of the project than in the middle.",
    "2037767": "andradaolteanu \nAs I understand:\n- the test sessions are different from the train sessions\n- they may be part of a longer sessions; however, there is no official indication that test sessions are part of a longer session so we will have to address this [\"cold start problem\"](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367495)\n- Ultimately, we are predicting up to 20 of the next items we think they will click/cart/order\n- In your last example, I would feature engineer a matrix of correlations or clusters between users or items in the train set. Then I agree that you should use the 5 given events and using some sort of model and that matrix to make your predictions.",
    "2037877": ">they may be part of a longer sessions; however, there is no official indication that test sessions are part of a longer session so we will have to address this \"cold start problem\"\n\nThe test data for each user (note \"session\" means \"user\" in this comp), is most likely **half** of that user's activity in the week of Aug 29th 2022. We can read the code from Otto's GitHub and we see that the test data was created by taking the user's full data (from the week of Aug 29th 2022) and randomly splitting it using `np.random.uniform`. Therefore the expected value is that we have **half** of the user's data (during the week of Aug 29th 2022).\n\nAnd our task is to predict the missing **half** of user data.",
    "2047739": "Hi, I still don't get it. \nHoping for discussions.\n1. If the sessions in test data are randomly split using` np.random.uniform`, I think there is data leaking problem.\n2. Why does the host of this competition make the 'userId' unavailable and use 'sessionId'?  Is this for privacy reasons?",
    "2048918": "I agree with you @Megan, this is not a dumb question. I am one of these bunch of people",
    "2051038": "I have the same question",
    "2051226": "Hi luojinwei,\n\none sessionId includes all interactions one user had with our shop during the whole period the trainset/testset covers. So it is basically a userId.",
    "2051813": "Thanks for your question, i finally figure out what i need to do in this competition.😂",
    "2051862": "okokok then maybe it's not that dumb 😂",
    "2058265": "Thank you for clarifying it.",
    "2091947": "Hi @cdeotte,\n\nI searched Otto's GitHub but didn't find any \"np.random.uniform\" and I couldn't find out how the test set was split in half.\nI have two questions: \n\n1. I saw your reply [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939#2072959), do you mean that each user's test data is split into a first half and a second half by time?\n2. From what point in time are we trying to predict events after? From the beginning of week 5, i.e. the whole week 5 event, or after the end of the test data we can see?\n\nI thought at first that the test data we got already specified half of the events, so all we needed to do was write those events correctly to the label of the CSV to get a recall of about 0.5, but in reality, I only got lb 0.061. so I found my understanding was incorrect.",
    "2092003": "Hi @chenlin1999,\nThe test data contains sessions truncated by timestamp, and you are to predict what occurs after the point of truncation. The timestamp that splits the session was chosen randomly. \nSome basic information about the test sessions:\nAll test sessions contain only data from week 5.\nThe session ids are different from the session ids of the train data.\nSessions with many events are rare compared with sessions with few events.\n\nSome examples:\nA) A session has 100 events. These events happen on Monday, Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday of week 5. The session is truncated after Wednesday. So you have the events from Monday, Tuesday, and Wednesday to predict the events from Thursday, Friday, Sunday and Saturday.\nB) A session has 100 events. These events happen on Monday, Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday of week 5. The session is truncated after Sunday. So you have the events from Monday, Tuesday, Wednesday, Thursday, Friday and Sunday to predict the events from Saturday.\nC) A session has 100 events. These events happen on Monday, Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday of week 5. The session is truncated after Monday. So you have the events from Monday to predict the events from Tuesday, Wednesday, Thursday, Friday, Sunday and Saturday.\nD) A session has three events. Two of these events happen on Monday and the last one on Saturday of week 5. The session is truncated after Monday. So you have the two events from Monday to predict the events from Saturday.\n\nI hope this clears things up a little.",
    "2093279": "Thank you very much, now I understand completely"
  },
  "source": "meta"
}