{
  "id": 190492,
  "title": "Sampling strategy?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190492",
  "author_name": "",
  "post_date": "2020-10-12T02:48:06.141632500Z",
  "votes": 11,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Hi, I saw many kernels that sample data using</p>\n<pre><code>train = train.iloc[90000000:,:]\n</code></pre>\n<p>after sorting using timestamp. What is the notion behind this? Or is there any better strategy as of now?</p>",
  "messages": [
    {
      "id": "1046802",
      "postDate": "10/12/2020 02:48:06",
      "content": "<p>Hi, I saw many kernels that sample data using</p>\n<pre><code>train = train.iloc[90000000:,:]\n</code></pre>\n<p>after sorting using timestamp. What is the notion behind this? Or is there any better strategy as of now?</p>",
      "rawMarkdown": "Hi, I saw many kernels that sample data using\n```\ntrain = train.iloc[90000000:,:]\n\n```\nafter sorting using timestamp. What is the notion behind this? Or is there any better strategy as of now?",
      "votes": null
    },
    {
      "id": "1046952",
      "postDate": "10/12/2020 06:12:00",
      "content": "<p>Some notebooks are using pandas directly to read the raw csv data which often results in memory errors when the entire dataset is used. Hence you see the sampling.</p>\n<p>Better strategy is to use the entire dataset 😛 (which I can confirm is possible).</p>",
      "rawMarkdown": "Some notebooks are using pandas directly to read the raw csv data which often results in memory errors when the entire dataset is used. Hence you see the sampling.\n\nBetter strategy is to use the entire dataset 😛 (which I can confirm is possible).",
      "votes": null
    },
    {
      "id": "1046976",
      "postDate": "10/12/2020 06:45:06",
      "content": "<p>haha, thanks.</p>",
      "rawMarkdown": "haha, thanks.",
      "votes": null
    },
    {
      "id": "1047001",
      "postDate": "10/12/2020 07:07:36",
      "content": "<p>As I understand it, in this  kernel -  is done to avoid leaks. The average values are taken from the first 90 million in time, and training is conducted on the last ~9 million + the average from the first part. This is the original and simplest way to divide data by time. Another option that came to mind was to divide the time separately for each user (train/val/test). The third method (which can be combined with the second) is to separate by users. What other options are possible?</p>",
      "rawMarkdown": "As I understand it, in this  kernel -  is done to avoid leaks. The average values are taken from the first 90 million in time, and training is conducted on the last ~9 million + the average from the first part. This is the original and simplest way to divide data by time. Another option that came to mind was to divide the time separately for each user (train/val/test). The third method (which can be combined with the second) is to separate by users. What other options are possible?",
      "votes": null
    },
    {
      "id": "1047026",
      "postDate": "10/12/2020 07:25:25",
      "content": "<p>Makes sense, thank you <a href=\"https://www.kaggle.com/sapr3s\" target=\"_blank\">@sapr3s</a> </p>",
      "rawMarkdown": "Makes sense, thank you @sapr3s",
      "votes": null
    },
    {
      "id": "1047380",
      "postDate": "10/12/2020 14:40:29",
      "content": "<p>Thanks for the clarification :)</p>",
      "rawMarkdown": "Thanks for the clarification :)",
      "votes": null
    },
    {
      "id": "1047918",
      "postDate": "10/13/2020 04:01:19",
      "content": "<p>Key things to keep in mind - Validation set should be of events occurring later than the data in train, some new user_ids and content_ids should be in the validation set.</p>",
      "rawMarkdown": "Key things to keep in mind - Validation set should be of events occurring later than the data in train, some new user_ids and content_ids should be in the validation set.",
      "votes": null
    },
    {
      "id": "1048832",
      "postDate": "10/13/2020 20:26:35",
      "content": "<p>I've wrote some code to create a simulated timeline. <br>\nBasically each user is assigned a random starting point at this timeline, and all their events timestamps are corrected by the chosen starting point (<code>corrected_timestamp = starting_point + timestamp</code>).<br>\nThen I can split my dataset chronologically, using the last few weeks/months as validation data.<br>\nThis insures me that I'll have both old and new users in my validation set (around 30-40% of new users, in my experiments)</p>",
      "rawMarkdown": "I've wrote some code to create a simulated timeline. \nBasically each user is assigned a random starting point at this timeline, and all their events timestamps are corrected by the chosen starting point (`corrected_timestamp = starting_point + timestamp`).\nThen I can split my dataset chronologically, using the last few weeks/months as validation data.\nThis insures me that I'll have both old and new users in my validation set (around 30-40% of new users, in my experiments)",
      "votes": null
    },
    {
      "id": "1048975",
      "postDate": "10/14/2020 02:29:46",
      "content": "<p>I have used the same strategy</p>",
      "rawMarkdown": "I have used the same strategy",
      "votes": null
    },
    {
      "id": "1049291",
      "postDate": "10/14/2020 09:46:29",
      "content": "<p>yes , I was able to load the full dataset(Thanks to <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>) , however I needed to use the famous reduce_mem_usage function to work with it .. especially if we are adding features using merge , or sorting stuff ..Not sure if there are other ways .</p>",
      "rawMarkdown": "yes , I was able to load the full dataset(Thanks to @rohanrao) , however I needed to use the famous reduce_mem_usage function to work with it .. especially if we are adding features using merge , or sorting stuff ..Not sure if there are other ways .",
      "votes": null
    },
    {
      "id": "1049455",
      "postDate": "10/14/2020 12:34:50",
      "content": "<p><a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a>  were you able to train with the full dataset?</p>",
      "rawMarkdown": "phoenix9032  were you able to train with the full dataset?",
      "votes": null
    },
    {
      "id": "1068157",
      "postDate": "11/03/2020 06:10:33",
      "content": "<p>Can you elaborate on this please ? </p>",
      "rawMarkdown": "Can you elaborate on this please ?",
      "votes": null
    },
    {
      "id": "1068187",
      "postDate": "11/03/2020 06:42:45",
      "content": "<ol>\n<li>Assign a random starting timestamp to all users(within reasonable bounds, say a month to an year). Lets call this t_start</li>\n<li>Create a new column called 'actual_timestamp'  = t_start + timestamp column</li>\n<li>Sort the entire train_set by this 'actual_timestamp' column</li>\n<li>Take the last 2.5M rows - lets call this the dev set</li>\n<li>Create groups of random sizes between 1 to 4000(taking 4 as the max num sessions in a single task_container_id and 1000 as the max num users in a group =&gt; 4 x 1000).</li>\n</ol>\n<p>This strategy can be used to debug and profile your inference pipeline.</p>",
      "rawMarkdown": "1. Assign a random starting timestamp to all users(within reasonable bounds, say a month to an year). Lets call this t_start\n2. Create a new column called 'actual_timestamp'  = t_start + timestamp column\n3. Sort the entire train_set by this 'actual_timestamp' column\n4. Take the last 2.5M rows - lets call this the dev set\n5. Create groups of random sizes between 1 to 4000(taking 4 as the max num sessions in a single task_container_id and 1000 as the max num users in a group => 4 x 1000).\n\nThis strategy can be used to debug and profile your inference pipeline.",
      "votes": null
    },
    {
      "id": "1068840",
      "postDate": "11/03/2020 19:39:54",
      "content": "<p>Thanks a lot!! And congrats for snatching the first rank.</p>",
      "rawMarkdown": "Thanks a lot!! And congrats for snatching the first rank.",
      "votes": null
    },
    {
      "id": "1069202",
      "postDate": "11/04/2020 07:40:56",
      "content": "<p>Rank is just ephemeral.  But thanks :)</p>",
      "rawMarkdown": "Rank is just ephemeral.  But thanks :)",
      "votes": null
    },
    {
      "id": "1069225",
      "postDate": "11/04/2020 08:17:18",
      "content": "<p>Even money won from this competition is ephemeral, enjoy it while you can hahaha </p>",
      "rawMarkdown": "Even money won from this competition is ephemeral, enjoy it while you can hahaha",
      "votes": null
    },
    {
      "id": "1069289",
      "postDate": "11/04/2020 09:34:25",
      "content": "<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a>  some how i still fail to visualize this.. could u give some example..<br>\nlike what timestamp can we assign to each user..</p>\n<p>Cant we actuall take last few time stamps for each user to build dev set ,post sorting with given timestamp?</p>",
      "rawMarkdown": "abhimanyud  some how i still fail to visualize this.. could u give some example..\nlike what timestamp can we assign to each user..\n\nCant we actuall take last few time stamps for each user to build dev set ,post sorting with given timestamp?",
      "votes": null
    },
    {
      "id": "1069293",
      "postDate": "11/04/2020 09:38:03",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>  check this notebook by <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> you might find what you're looking for here: <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a></p>",
      "rawMarkdown": "jaideepvalani  check this notebook by @its7171 you might find what you're looking for here: https://www.kaggle.com/its7171/cv-strategy",
      "votes": null
    },
    {
      "id": "1069308",
      "postDate": "11/04/2020 09:59:00",
      "content": "<p>Yes i just happened to check. Was wanting to ask.. are the files provided there are for full ds or partial ds</p>\n<p>Also i see some datasets uploaded have feather file for entire set,should that be good enough ?</p>",
      "rawMarkdown": "Yes i just happened to check. Was wanting to ask.. are the files provided there are for full ds or partial ds\n\nAlso i see some datasets uploaded have feather file for entire set,should that be good enough ?",
      "votes": null
    },
    {
      "id": "1069408",
      "postDate": "11/04/2020 12:22:07",
      "content": "<p>As mentioned in the notebook, it seems like complete dataset won't fit so you'll either have to do offline processing where you have more memory available or optimize the approach to fit in 16/13G. <br>\nYes, feather files can be stored/read very efficiently that's why most people working with pandas use it.</p>",
      "rawMarkdown": "As mentioned in the notebook, it seems like complete dataset won't fit so you'll either have to do offline processing where you have more memory available or optimize the approach to fit in 16/13G. \nYes, feather files can be stored/read very efficiently that's why most people working with pandas use it.",
      "votes": null
    },
    {
      "id": "1069417",
      "postDate": "11/04/2020 12:32:05",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> </p>\n<blockquote>\n  <p>Cant we actuall take last few time stamps for each user to build dev set ,post sorting with given timestamp?</p>\n</blockquote>\n<p>That's actually the right way to build a validation set.<br>\nWhat I was talking about was simulating the test environment. <br>\nSuppose we have only 3 users -</p>\n<p>We generate 3 random timestamps (making sure that they are within say 3 months of each other)<br>\nand assign them to each user. We assume that for each of these users this is the timestamp at which he first came to the app. Therefore if we add the original timestamp to it, we get a value that is representative of actual time of occurrence of the interaction .</p>\n<p>We can use this generated value to simulate groups of user data coming in just like they do in test_iter.</p>\n<p>Hope that helps.</p>",
      "rawMarkdown": "jaideepvalani \n> Cant we actuall take last few time stamps for each user to build dev set ,post sorting with given timestamp?\n\nThat's actually the right way to build a validation set.\nWhat I was talking about was simulating the test environment. \nSuppose we have only 3 users -\n\nWe generate 3 random timestamps (making sure that they are within say 3 months of each other)\nand assign them to each user. We assume that for each of these users this is the timestamp at which he first came to the app. Therefore if we add the original timestamp to it, we get a value that is representative of actual time of occurrence of the interaction .\n\nWe can use this generated value to simulate groups of user data coming in just like they do in test_iter.\n\nHope that helps.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1046952,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "10/12/2020 06:12:00",
      "content": "<p>Some notebooks are using pandas directly to read the raw csv data which often results in memory errors when the entire dataset is used. Hence you see the sampling.</p>\n<p>Better strategy is to use the entire dataset 😛 (which I can confirm is possible).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1046976,
          "author_name": "shahules",
          "author_url": "",
          "post_date": "10/12/2020 06:45:06",
          "content": "<p>haha, thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1049291,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "10/14/2020 09:46:29",
          "content": "<p>yes , I was able to load the full dataset(Thanks to <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>) , however I needed to use the famous reduce_mem_usage function to work with it .. especially if we are adding features using merge , or sorting stuff ..Not sure if there are other ways .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1049455,
          "author_name": "shahules",
          "author_url": "",
          "post_date": "10/14/2020 12:34:50",
          "content": "<p><a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a>  were you able to train with the full dataset?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1047001,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "10/12/2020 07:07:36",
      "content": "<p>As I understand it, in this  kernel -  is done to avoid leaks. The average values are taken from the first 90 million in time, and training is conducted on the last ~9 million + the average from the first part. This is the original and simplest way to divide data by time. Another option that came to mind was to divide the time separately for each user (train/val/test). The third method (which can be combined with the second) is to separate by users. What other options are possible?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1047026,
          "author_name": "shahules",
          "author_url": "",
          "post_date": "10/12/2020 07:25:25",
          "content": "<p>Makes sense, thank you <a href=\"https://www.kaggle.com/sapr3s\" target=\"_blank\">@sapr3s</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1047918,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "10/13/2020 04:01:19",
          "content": "<p>Key things to keep in mind - Validation set should be of events occurring later than the data in train, some new user_ids and content_ids should be in the validation set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1047380,
      "author_name": "domizianostingi",
      "author_url": "",
      "post_date": "10/12/2020 14:40:29",
      "content": "<p>Thanks for the clarification :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1048832,
      "author_name": "fredcaroli",
      "author_url": "",
      "post_date": "10/13/2020 20:26:35",
      "content": "<p>I've wrote some code to create a simulated timeline. <br>\nBasically each user is assigned a random starting point at this timeline, and all their events timestamps are corrected by the chosen starting point (<code>corrected_timestamp = starting_point + timestamp</code>).<br>\nThen I can split my dataset chronologically, using the last few weeks/months as validation data.<br>\nThis insures me that I'll have both old and new users in my validation set (around 30-40% of new users, in my experiments)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1048975,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "10/14/2020 02:29:46",
          "content": "<p>I have used the same strategy</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068157,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "11/03/2020 06:10:33",
          "content": "<p>Can you elaborate on this please ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068187,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "11/03/2020 06:42:45",
          "content": "<ol>\n<li>Assign a random starting timestamp to all users(within reasonable bounds, say a month to an year). Lets call this t_start</li>\n<li>Create a new column called 'actual_timestamp'  = t_start + timestamp column</li>\n<li>Sort the entire train_set by this 'actual_timestamp' column</li>\n<li>Take the last 2.5M rows - lets call this the dev set</li>\n<li>Create groups of random sizes between 1 to 4000(taking 4 as the max num sessions in a single task_container_id and 1000 as the max num users in a group =&gt; 4 x 1000).</li>\n</ol>\n<p>This strategy can be used to debug and profile your inference pipeline.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068840,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "11/03/2020 19:39:54",
          "content": "<p>Thanks a lot!! And congrats for snatching the first rank.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069202,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "11/04/2020 07:40:56",
          "content": "<p>Rank is just ephemeral.  But thanks :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069225,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "11/04/2020 08:17:18",
          "content": "<p>Even money won from this competition is ephemeral, enjoy it while you can hahaha </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069289,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "11/04/2020 09:34:25",
          "content": "<p><a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a>  some how i still fail to visualize this.. could u give some example..<br>\nlike what timestamp can we assign to each user..</p>\n<p>Cant we actuall take last few time stamps for each user to build dev set ,post sorting with given timestamp?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069293,
          "author_name": "dexarsal",
          "author_url": "",
          "post_date": "11/04/2020 09:38:03",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>  check this notebook by <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> you might find what you're looking for here: <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069308,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "11/04/2020 09:59:00",
          "content": "<p>Yes i just happened to check. Was wanting to ask.. are the files provided there are for full ds or partial ds</p>\n<p>Also i see some datasets uploaded have feather file for entire set,should that be good enough ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069408,
          "author_name": "dexarsal",
          "author_url": "",
          "post_date": "11/04/2020 12:22:07",
          "content": "<p>As mentioned in the notebook, it seems like complete dataset won't fit so you'll either have to do offline processing where you have more memory available or optimize the approach to fit in 16/13G. <br>\nYes, feather files can be stored/read very efficiently that's why most people working with pandas use it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069417,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "11/04/2020 12:32:05",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> </p>\n<blockquote>\n  <p>Cant we actuall take last few time stamps for each user to build dev set ,post sorting with given timestamp?</p>\n</blockquote>\n<p>That's actually the right way to build a validation set.<br>\nWhat I was talking about was simulating the test environment. <br>\nSuppose we have only 3 users -</p>\n<p>We generate 3 random timestamps (making sure that they are within say 3 months of each other)<br>\nand assign them to each user. We assume that for each of these users this is the timestamp at which he first came to the app. Therefore if we add the original timestamp to it, we get a value that is representative of actual time of occurrence of the interaction .</p>\n<p>We can use this generated value to simulate groups of user data coming in just like they do in test_iter.</p>\n<p>Hope that helps.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1046802": "Hi, I saw many kernels that sample data using\n```\ntrain = train.iloc[90000000:,:]\n\n```\nafter sorting using timestamp. What is the notion behind this? Or is there any better strategy as of now?",
    "1046952": "Some notebooks are using pandas directly to read the raw csv data which often results in memory errors when the entire dataset is used. Hence you see the sampling.\n\nBetter strategy is to use the entire dataset 😛 (which I can confirm is possible).",
    "1046976": "haha, thanks.",
    "1047001": "As I understand it, in this  kernel -  is done to avoid leaks. The average values are taken from the first 90 million in time, and training is conducted on the last ~9 million + the average from the first part. This is the original and simplest way to divide data by time. Another option that came to mind was to divide the time separately for each user (train/val/test). The third method (which can be combined with the second) is to separate by users. What other options are possible?",
    "1047026": "Makes sense, thank you @sapr3s",
    "1047380": "Thanks for the clarification :)",
    "1047918": "Key things to keep in mind - Validation set should be of events occurring later than the data in train, some new user_ids and content_ids should be in the validation set.",
    "1048832": "I've wrote some code to create a simulated timeline. \nBasically each user is assigned a random starting point at this timeline, and all their events timestamps are corrected by the chosen starting point (`corrected_timestamp = starting_point + timestamp`).\nThen I can split my dataset chronologically, using the last few weeks/months as validation data.\nThis insures me that I'll have both old and new users in my validation set (around 30-40% of new users, in my experiments)",
    "1048975": "I have used the same strategy",
    "1049291": "yes , I was able to load the full dataset(Thanks to @rohanrao) , however I needed to use the famous reduce_mem_usage function to work with it .. especially if we are adding features using merge , or sorting stuff ..Not sure if there are other ways .",
    "1049455": "phoenix9032  were you able to train with the full dataset?",
    "1068157": "Can you elaborate on this please ?",
    "1068187": "1. Assign a random starting timestamp to all users(within reasonable bounds, say a month to an year). Lets call this t_start\n2. Create a new column called 'actual_timestamp'  = t_start + timestamp column\n3. Sort the entire train_set by this 'actual_timestamp' column\n4. Take the last 2.5M rows - lets call this the dev set\n5. Create groups of random sizes between 1 to 4000(taking 4 as the max num sessions in a single task_container_id and 1000 as the max num users in a group => 4 x 1000).\n\nThis strategy can be used to debug and profile your inference pipeline.",
    "1068840": "Thanks a lot!! And congrats for snatching the first rank.",
    "1069202": "Rank is just ephemeral.  But thanks :)",
    "1069225": "Even money won from this competition is ephemeral, enjoy it while you can hahaha",
    "1069289": "abhimanyud  some how i still fail to visualize this.. could u give some example..\nlike what timestamp can we assign to each user..\n\nCant we actuall take last few time stamps for each user to build dev set ,post sorting with given timestamp?",
    "1069293": "jaideepvalani  check this notebook by @its7171 you might find what you're looking for here: https://www.kaggle.com/its7171/cv-strategy",
    "1069308": "Yes i just happened to check. Was wanting to ask.. are the files provided there are for full ds or partial ds\n\nAlso i see some datasets uploaded have feather file for entire set,should that be good enough ?",
    "1069408": "As mentioned in the notebook, it seems like complete dataset won't fit so you'll either have to do offline processing where you have more memory available or optimize the approach to fit in 16/13G. \nYes, feather files can be stored/read very efficiently that's why most people working with pandas use it.",
    "1069417": "jaideepvalani \n> Cant we actuall take last few time stamps for each user to build dev set ,post sorting with given timestamp?\n\nThat's actually the right way to build a validation set.\nWhat I was talking about was simulating the test environment. \nSuppose we have only 3 users -\n\nWe generate 3 random timestamps (making sure that they are within say 3 months of each other)\nand assign them to each user. We assume that for each of these users this is the timestamp at which he first came to the app. Therefore if we add the original timestamp to it, we get a value that is representative of actual time of occurrence of the interaction .\n\nWe can use this generated value to simulate groups of user data coming in just like they do in test_iter.\n\nHope that helps."
  },
  "source": "meta"
}