{
  "id": 308919,
  "title": "How To Setup Local CV",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919",
  "author_name": "Chris Deotte",
  "post_date": "2022-02-20T23:26:13.715000",
  "votes": 251,
  "comment_count": 51,
  "views": 0,
  "content": "<p>The first step in every Kaggle competition is to build a reliable local validation scheme. Then we use our local validation score to evaluate experiment ideas and/or tune hyperparameters.</p>\n<h1>Train Data</h1>\n<p>The last day in the transaction dataframe is <code>2020-09-22</code>. The public LB contains 1 week of transactions after this date. Therefore to create a local validation that mimics Kaggle's train test relationship, we can train on all transactions before <code>2020-9-15</code>. And validate on the last week in train data.</p>\n<pre><code>train = pd.read_csv('transactions_train.csv')\ntrain.t_dat = pd.to_datetime( train.t_dat )\ntrain = train.loc[ train.t_dat &lt;= pd.to_datetime('2020-09-15') ]\n</code></pre>\n<h1>Valid Data</h1>\n<p>The code below will create a dataframe with only the customers who made purchases during the last week of train (which are the only ones that affect competition metric). It formats the predictions as strings like sample_submission.csv</p>\n<pre><code>valid = pd.read_csv('transactions_train.csv')\nvalid.t_dat = pd.to_datetime( valid.t_dat )\nvalid = valid.loc[ valid.t_dat &gt;= pd.to_datetime('2020-09-16') ]\nvalid = valid.groupby('customer_id').article_id.apply(list).reset_index()\nvalid = valid.rename({'article_id':'prediction'},axis=1)\nvalid['prediction'] =\\\n    valid.prediction.apply(lambda x: ' '.join(['0'+str(k) for k in x]))\n</code></pre>\n<h1>Compute Validation Score mAP</h1>\n<p>To compute your validation score, use the <code>train</code> dataframe above (which does not contain the validation dates) and then make predictions for every customer in <code>sample_submission.csv</code>. Create a column of prediction strings just like you would when submitting to Kaggle and save as <code>submission.csv</code> (to be used below).</p>\n<h1>Important x2</h1>\n<p>Once you have your <code>submission.csv</code> file (made by using train dataframe above which excludes last of train data), run the code below. <strong>Important x1</strong>, we must execute the second line below. It guarantees that your submission and valid dataframes are in the same row order. Furthermore it removes all customers who do not make predictions during validation period because these customers do not affect metric score. <strong>Important x1</strong>, remember to write your <code>article_id</code>s with a prefix of zero in your dataframe predictionstrings otherwise your validation and LB score will be 0.</p>\n<pre><code>sub = pd.read_csv('submission.csv')\nsub = sub.set_index('customer_id').loc[valid.customer_id].reset_index()\nmapk( valid.prediction.str.split(), sub.prediction.str.split(), k=12)\n</code></pre>\n<h1>MapK Function</h1>\n<p>The MapK Function can be found on Kaggle's GitHub <a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\" target=\"_blank\">here</a> or in Kaerururu's notebook <a href=\"https://www.kaggle.com/kaerunantoka/h-m-how-to-calculate-map-12\" target=\"_blank\">here</a>.</p>\n<h1>CV LB Agreement</h1>\n<p>Using the code above, the best public notebook <a href=\"https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2\" target=\"_blank\">here</a> has validation score 0.023 and LB 0.020. If we combine my notebook <a href=\"https://www.kaggle.com/cdeotte/customers-who-bought-this-frequently-buy-this\" target=\"_blank\">here</a> with that notebook, then the validation score improves to 0.024 and LB improves to 0.021 <a href=\"https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">here</a>. So it appears when validation score improves then LB improves. </p>\n<h1>3 Folds</h1>\n<p>To have a more reliable validation, we can use 3 or more folds and average the results. When doing 3 folds, the best public notebook <a href=\"https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2\" target=\"_blank\">here</a> achieves 0.023, 0.021, 0.021 on folds 0,1,2 respectively and therefore achieves average 3-fold validation score of 0.022 which matches its LB score!</p>\n<ul>\n<li>Fold 0 - train &lt;= '2020-09-15', valid &gt;= '2020-09-16'</li>\n<li>Fold 1 - train &lt;= '2020-09-08', (valid &gt;= '2020-09-09')&amp;(valid &lt;= '2020-09-15')</li>\n<li>Fold 2 - train &lt;= '2020-09-01', (valid &gt;= '2020-09-02')&amp;(valid &lt;= '2020-09-08')</li>\n</ul>",
  "messages": [
    {
      "id": 1699104,
      "postDate": "2022-02-20T23:26:13.717Z",
      "content": "<p>The first step in every Kaggle competition is to build a reliable local validation scheme. Then we use our local validation score to evaluate experiment ideas and/or tune hyperparameters.</p>\n<h1>Train Data</h1>\n<p>The last day in the transaction dataframe is <code>2020-09-22</code>. The public LB contains 1 week of transactions after this date. Therefore to create a local validation that mimics Kaggle's train test relationship, we can train on all transactions before <code>2020-9-15</code>. And validate on the last week in train data.</p>\n<pre><code>train = pd.read_csv('transactions_train.csv')\ntrain.t_dat = pd.to_datetime( train.t_dat )\ntrain = train.loc[ train.t_dat &lt;= pd.to_datetime('2020-09-15') ]\n</code></pre>\n<h1>Valid Data</h1>\n<p>The code below will create a dataframe with only the customers who made purchases during the last week of train (which are the only ones that affect competition metric). It formats the predictions as strings like sample_submission.csv</p>\n<pre><code>valid = pd.read_csv('transactions_train.csv')\nvalid.t_dat = pd.to_datetime( valid.t_dat )\nvalid = valid.loc[ valid.t_dat &gt;= pd.to_datetime('2020-09-16') ]\nvalid = valid.groupby('customer_id').article_id.apply(list).reset_index()\nvalid = valid.rename({'article_id':'prediction'},axis=1)\nvalid['prediction'] =\\\n    valid.prediction.apply(lambda x: ' '.join(['0'+str(k) for k in x]))\n</code></pre>\n<h1>Compute Validation Score mAP</h1>\n<p>To compute your validation score, use the <code>train</code> dataframe above (which does not contain the validation dates) and then make predictions for every customer in <code>sample_submission.csv</code>. Create a column of prediction strings just like you would when submitting to Kaggle and save as <code>submission.csv</code> (to be used below).</p>\n<h1>Important x2</h1>\n<p>Once you have your <code>submission.csv</code> file (made by using train dataframe above which excludes last of train data), run the code below. <strong>Important x1</strong>, we must execute the second line below. It guarantees that your submission and valid dataframes are in the same row order. Furthermore it removes all customers who do not make predictions during validation period because these customers do not affect metric score. <strong>Important x1</strong>, remember to write your <code>article_id</code>s with a prefix of zero in your dataframe predictionstrings otherwise your validation and LB score will be 0.</p>\n<pre><code>sub = pd.read_csv('submission.csv')\nsub = sub.set_index('customer_id').loc[valid.customer_id].reset_index()\nmapk( valid.prediction.str.split(), sub.prediction.str.split(), k=12)\n</code></pre>\n<h1>MapK Function</h1>\n<p>The MapK Function can be found on Kaggle's GitHub <a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\" target=\"_blank\">here</a> or in Kaerururu's notebook <a href=\"https://www.kaggle.com/kaerunantoka/h-m-how-to-calculate-map-12\" target=\"_blank\">here</a>.</p>\n<h1>CV LB Agreement</h1>\n<p>Using the code above, the best public notebook <a href=\"https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2\" target=\"_blank\">here</a> has validation score 0.023 and LB 0.020. If we combine my notebook <a href=\"https://www.kaggle.com/cdeotte/customers-who-bought-this-frequently-buy-this\" target=\"_blank\">here</a> with that notebook, then the validation score improves to 0.024 and LB improves to 0.021 <a href=\"https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">here</a>. So it appears when validation score improves then LB improves. </p>\n<h1>3 Folds</h1>\n<p>To have a more reliable validation, we can use 3 or more folds and average the results. When doing 3 folds, the best public notebook <a href=\"https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2\" target=\"_blank\">here</a> achieves 0.023, 0.021, 0.021 on folds 0,1,2 respectively and therefore achieves average 3-fold validation score of 0.022 which matches its LB score!</p>\n<ul>\n<li>Fold 0 - train &lt;= '2020-09-15', valid &gt;= '2020-09-16'</li>\n<li>Fold 1 - train &lt;= '2020-09-08', (valid &gt;= '2020-09-09')&amp;(valid &lt;= '2020-09-15')</li>\n<li>Fold 2 - train &lt;= '2020-09-01', (valid &gt;= '2020-09-02')&amp;(valid &lt;= '2020-09-08')</li>\n</ul>",
      "rawMarkdown": "The first step in every Kaggle competition is to build a reliable local validation scheme. Then we use our local validation score to evaluate experiment ideas and/or tune hyperparameters.\n\n# Train Data\nThe last day in the transaction dataframe is `2020-09-22`. The public LB contains 1 week of transactions after this date. Therefore to create a local validation that mimics Kaggle's train test relationship, we can train on all transactions before `2020-9-15`. And validate on the last week in train data.\n\n    train = pd.read_csv('transactions_train.csv')\n    train.t_dat = pd.to_datetime( train.t_dat )\n    train = train.loc[ train.t_dat <= pd.to_datetime('2020-09-15') ]\n\n# Valid Data\nThe code below will create a dataframe with only the customers who made purchases during the last week of train (which are the only ones that affect competition metric). It formats the predictions as strings like sample_submission.csv\n\n    valid = pd.read_csv('transactions_train.csv')\n    valid.t_dat = pd.to_datetime( valid.t_dat )\n    valid = valid.loc[ valid.t_dat >= pd.to_datetime('2020-09-16') ]\n    valid = valid.groupby('customer_id').article_id.apply(list).reset_index()\n    valid = valid.rename({'article_id':'prediction'},axis=1)\n    valid['prediction'] =\\\n        valid.prediction.apply(lambda x: ' '.join(['0'+str(k) for k in x]))\n\n# Compute Validation Score mAP\nTo compute your validation score, use the `train` dataframe above (which does not contain the validation dates) and then make predictions for every customer in `sample_submission.csv`. Create a column of prediction strings just like you would when submitting to Kaggle and save as `submission.csv` (to be used below).\n\n# Important x2\nOnce you have your `submission.csv` file (made by using train dataframe above which excludes last of train data), run the code below. **Important x1**, we must execute the second line below. It guarantees that your submission and valid dataframes are in the same row order. Furthermore it removes all customers who do not make predictions during validation period because these customers do not affect metric score. **Important x1**, remember to write your `article_id`s with a prefix of zero in your dataframe predictionstrings otherwise your validation and LB score will be 0.\n\n    sub = pd.read_csv('submission.csv')\n    sub = sub.set_index('customer_id').loc[valid.customer_id].reset_index()\n    mapk( valid.prediction.str.split(), sub.prediction.str.split(), k=12)\n\n# MapK Function\nThe MapK Function can be found on Kaggle's GitHub [here][2] or in Kaerururu's notebook [here][1].\n\n# CV LB Agreement\nUsing the code above, the best public notebook [here][3] has validation score 0.023 and LB 0.020. If we combine my notebook [here][4] with that notebook, then the validation score improves to 0.024 and LB improves to 0.021 [here][5]. So it appears when validation score improves then LB improves. \n\n# 3 Folds\nTo have a more reliable validation, we can use 3 or more folds and average the results. When doing 3 folds, the best public notebook [here][3] achieves 0.023, 0.021, 0.021 on folds 0,1,2 respectively and therefore achieves average 3-fold validation score of 0.022 which matches its LB score!\n\n* Fold 0 - train <= '2020-09-15', valid >= '2020-09-16'\n* Fold 1 - train <= '2020-09-08', (valid >= '2020-09-09')&(valid <= '2020-09-15')\n* Fold 2 - train <= '2020-09-01', (valid >= '2020-09-02')&(valid <= '2020-09-08')\n\n[1]: https://www.kaggle.com/kaerunantoka/h-m-how-to-calculate-map-12\n[2]: https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\n[3]: https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2\n[4]: https://www.kaggle.com/cdeotte/customers-who-bought-this-frequently-buy-this\n[5]: https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021\n",
      "votes": 248
    },
    {
      "id": 1699322,
      "postDate": "2022-02-21T05:26:51.893Z",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>Thank you so much for your (as always) helpful and informative post. If I may be so bold as to ask a rather general (not related to this competition) and somewhat naive question (I am revising my understanding of cross-validation and <a href=\"https://www.kaggle.com/questions-and-answers/307921\" target=\"_blank\">I find my understanding is shaky</a>).</p>\n<p>My question is as follows; by comparing ones local CV to the LB is one not basically treating the LB as an ersatz hold-out test set, and if that is the case, would it not be better (if one has enough training data) to create ones own hold-out test set of a similar size, extracted from the training data, and where one can be sure that it is i.i.d, rather than assuming that the kaggle Public LB data is i.i.d? </p>\n<p>Away from kaggle there is no Public LB, so when people say <em>\"Trust your CV\"</em>, are they really not essentially  saying <em>\"Distrust the Public LB\"</em>?</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @cdeotte \n\nThank you so much for your (as always) helpful and informative post. If I may be so bold as to ask a rather general (not related to this competition) and somewhat naive question (I am revising my understanding of cross-validation and [I find my understanding is shaky](https://www.kaggle.com/questions-and-answers/307921)).\n\nMy question is as follows; by comparing ones local CV to the LB is one not basically treating the LB as an ersatz hold-out test set, and if that is the case, would it not be better (if one has enough training data) to create ones own hold-out test set of a similar size, extracted from the training data, and where one can be sure that it is i.i.d, rather than assuming that the kaggle Public LB data is i.i.d? \n\nAway from kaggle there is no Public LB, so when people say *\"Trust your CV\"*, are they really not essentially  saying *\"Distrust the Public LB\"*?\n\nAll the best,\ncarl",
      "votes": 5,
      "replies": [
        {
          "id": 1699816,
          "postDate": "2022-02-21T13:15:59.420Z",
          "content": "<p>There is much to say about this (and your other post). I will make a quick comment and update it in the days to come.</p>\n<blockquote>\n  <p>comparing ones local CV to the LB is one not basically treating the LB as an ersatz hold-out test</p>\n</blockquote>\n<p>Not exactly. The one difference between real life and Kaggle is that we are never sure about what's in Kaggle's test dataset. In some competitions, test dataset is very different than train data. The purpose of comparing CV to LB is a method of detective work to determine what is in Kaggle's test dataset. </p>\n<p>If CV scores and LB scores are related, then most likely the relationship between your local folds is the same relationship between Kaggle's train and test. If not, try other CV splitting techniques until it is. Also, we modify our modeling to generalize in the direction of test's shift from train. </p>",
          "rawMarkdown": "There is much to say about this (and your other post). I will make a quick comment and update it in the days to come.\n\n>comparing ones local CV to the LB is one not basically treating the LB as an ersatz hold-out test\n\nNot exactly. The one difference between real life and Kaggle is that we are never sure about what's in Kaggle's test dataset. In some competitions, test dataset is very different than train data. The purpose of comparing CV to LB is a method of detective work to determine what is in Kaggle's test dataset. \n\nIf CV scores and LB scores are related, then most likely the relationship between your local folds is the same relationship between Kaggle's train and test. If not, try other CV splitting techniques until it is. Also, we modify our modeling to generalize in the direction of test's shift from train. \n",
          "votes": 7
        },
        {
          "id": 1700283,
          "postDate": "2022-02-21T19:42:58.227Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<blockquote>\n  <p>\"<em>There is much to say about this…and update it in the days to come.</em>\"</p>\n</blockquote>\n<p>That would be wonderful; given that CV, or variants of, is an integral part of ML, and is ubiquitous in almost all kaggle competitions, your insights regarding general CV advice,  'hints and tips', or even bespoke approaches would be most appreciated by myself, and I somewhat suspect by many others too!</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @cdeotte \n\n> \"*There is much to say about this...and update it in the days to come.*\"\n\nThat would be wonderful; given that CV, or variants of, is an integral part of ML, and is ubiquitous in almost all kaggle competitions, your insights regarding general CV advice,  'hints and tips', or even bespoke approaches would be most appreciated by myself, and I somewhat suspect by many others too!\n\nAll the best,\ncarl",
          "votes": 1
        }
      ]
    },
    {
      "id": 1702052,
      "postDate": "2022-02-23T09:28:13.390Z",
      "content": "<p>Use can avoid this:</p>\n<pre><code>valid['prediction'] =\\\n    valid.prediction.apply(lambda x: ' '.join(['0'+str(k) for k in x]))\n</code></pre>\n<p>By this on line 1:</p>\n<pre><code>valid = pd.read_csv('transactions_train.csv', dtype={'article_id': 'str'})\n</code></pre>",
      "rawMarkdown": "Use can avoid this:\n```\nvalid['prediction'] =\\\n    valid.prediction.apply(lambda x: ' '.join(['0'+str(k) for k in x]))\n```\nBy this on line 1:\n```\nvalid = pd.read_csv('transactions_train.csv', dtype={'article_id': 'str'})\n```",
      "votes": 3,
      "replies": [
        {
          "id": 1702267,
          "postDate": "2022-02-23T13:25:06.253Z",
          "content": "<p>Note that the first line of code (in your comment) does two things (1) convert int to str, (2) combine multiple str into one prediction string. Your second line of code does only one thing (1) convert int to str.</p>\n<p>So if we use your second line, then we can change the first line to (i.e we can't remove the first line) </p>\n<pre><code>valid['prediction'] = valid.prediction.apply(lambda x: ' '.join(x))\n</code></pre>\n<p>However in general, IMO, it is best to load an integer as an integer because integers are smaller in memory than strings and computers can process integers faster than strings. Then we convert to string in the final line of code.</p>",
          "rawMarkdown": "Note that the first line of code (in your comment) does two things (1) convert int to str, (2) combine multiple str into one prediction string. Your second line of code does only one thing (1) convert int to str.\n\nSo if we use your second line, then we can change the first line to (i.e we can't remove the first line) \n\n    valid['prediction'] = valid.prediction.apply(lambda x: ' '.join(x))\n\nHowever in general, IMO, it is best to load an integer as an integer because integers are smaller in memory than strings and computers can process integers faster than strings. Then we convert to string in the final line of code.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1733968,
      "postDate": "2022-03-24T20:16:54.930Z",
      "content": "<p>This is very helpful, thank you so much!</p>",
      "rawMarkdown": "This is very helpful, thank you so much!",
      "votes": 1
    },
    {
      "id": 1731478,
      "postDate": "2022-03-22T12:06:04.980Z",
      "content": "<p>wonderful idea! Thank you for sharing this.</p>",
      "rawMarkdown": "wonderful idea! Thank you for sharing this.",
      "votes": 1
    },
    {
      "id": 1730261,
      "postDate": "2022-03-21T04:20:41.343Z",
      "content": "<p>Anyone having trouble installing ml_metrics for mapk package?</p>",
      "rawMarkdown": "Anyone having trouble installing ml_metrics for mapk package?",
      "votes": 1
    },
    {
      "id": 1714376,
      "postDate": "2022-03-06T22:41:18.773Z",
      "content": "<p>Why are you using combined weeks instead of just <br>\n2020-09-16 - 2020-09-22 <br>\n2020-09-09 - 2020-09-15 <br>\n2020-09-02 - 2020-09-08? </p>",
      "rawMarkdown": "Why are you using combined weeks instead of just \n2020-09-16 - 2020-09-22 \n2020-09-09 - 2020-09-15 \n2020-09-02 - 2020-09-08? ",
      "votes": 1,
      "replies": [
        {
          "id": 1714393,
          "postDate": "2022-03-07T00:30:18.553Z",
          "content": "<p>This is exactly what i'm saying. </p>\n<p>The extra dates are indicating that when we validation on <code>20-20-09-16 to 2020-09-22</code> then we must train on data <code>20-20-09-15</code> and before. (And when we validate on <code>20-20-09-09 to 2020-09-15</code> we must train on data <code>20-20--09-08</code> and before, etc etc)</p>",
          "rawMarkdown": "This is exactly what i'm saying. \n\nThe extra dates are indicating that when we validation on `20-20-09-16 to 2020-09-22` then we must train on data `20-20-09-15` and before. (And when we validate on `20-20-09-09 to 2020-09-15` we must train on data `20-20--09-08` and before, etc etc)",
          "votes": 1
        },
        {
          "id": 1714608,
          "postDate": "2022-03-07T06:43:12.067Z",
          "content": "<p>Yes, my bad, just saw &amp; symbol and my brain weirdly concluded that it's two weeks combined)</p>",
          "rawMarkdown": "Yes, my bad, just saw & symbol and my brain weirdly concluded that it's two weeks combined)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1713211,
      "postDate": "2022-03-05T18:52:48.277Z",
      "content": "<p>Thanks for sharing! Very useful notebook.</p>",
      "rawMarkdown": "Thanks for sharing! Very useful notebook.",
      "votes": 1
    },
    {
      "id": 1711936,
      "postDate": "2022-03-04T13:35:07.380Z",
      "content": "<p>Has anyone checked the CV results with sampling of the dataset?</p>",
      "rawMarkdown": "Has anyone checked the CV results with sampling of the dataset?",
      "votes": 1
    },
    {
      "id": 1710417,
      "postDate": "2022-03-03T03:05:57.317Z",
      "content": "<p>Nice, very useful sharing.</p>",
      "rawMarkdown": "Nice, very useful sharing.",
      "votes": 1
    },
    {
      "id": 1705706,
      "postDate": "2022-02-26T18:20:13.297Z",
      "content": "<p>Amazing Stuff .</p>",
      "rawMarkdown": "Amazing Stuff .",
      "votes": 1
    },
    {
      "id": 1702592,
      "postDate": "2022-02-23T18:22:01.430Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Great work @cdeotte ",
      "votes": 1
    },
    {
      "id": 1703473,
      "postDate": "2022-02-24T14:27:18.103Z",
      "content": "<p>Very useful notebook! Thanks for sharing.</p>",
      "rawMarkdown": "Very useful notebook! Thanks for sharing.",
      "votes": 2
    },
    {
      "id": 1702339,
      "postDate": "2022-02-23T14:28:17.037Z",
      "content": "<p>I think it's a smart idea. It will be helpful</p>",
      "rawMarkdown": "I think it's a smart idea. It will be helpful",
      "votes": 2
    },
    {
      "id": 1701235,
      "postDate": "2022-02-22T15:50:31.387Z",
      "content": "<p>Important information and high quality work 👏 Thank you very much <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Good luck for the competition!</p>",
      "rawMarkdown": "Important information and high quality work 👏 Thank you very much @cdeotte Good luck for the competition!",
      "votes": 2
    },
    {
      "id": 1701024,
      "postDate": "2022-02-22T13:16:00.657Z",
      "content": "<p>Thanks very much for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> great explanation on the importance of the time series element to the cross-validation.  I am just wondering if the season will have an impact on the recommendations. Something to investigate.</p>",
      "rawMarkdown": "Thanks very much for sharing @cdeotte great explanation on the importance of the time series element to the cross-validation.  I am just wondering if the season will have an impact on the recommendations. Something to investigate.",
      "votes": 2,
      "replies": [
        {
          "id": 1701154,
          "postDate": "2022-02-22T14:51:45.123Z",
          "content": "<p>Good suggestion. I suspect that time of year does affect what clothes customers purchase.</p>",
          "rawMarkdown": "Good suggestion. I suspect that time of year does affect what clothes customers purchase.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1699644,
      "postDate": "2022-02-21T10:28:11.827Z",
      "content": "<p>Thank's for this great explanation, one question however : How to deal with the very small ratio of public LB (it's only 1% of all private data, while the private LB will be the remaining 99%).</p>\n<p>Should we take this into account while validating ? like taking 1% of the validation week to compute the score that would probably be shown on the public LB and use the whole validation week to get the real CV score (which hopefully will be the same on private LB) ?</p>",
      "rawMarkdown": "Thank's for this great explanation, one question however : How to deal with the very small ratio of public LB (it's only 1% of all private data, while the private LB will be the remaining 99%).\n\nShould we take this into account while validating ? like taking 1% of the validation week to compute the score that would probably be shown on the public LB and use the whole validation week to get the real CV score (which hopefully will be the same on private LB) ?",
      "votes": 2,
      "replies": [
        {
          "id": 1699805,
          "postDate": "2022-02-21T13:02:58.253Z",
          "content": "<p>This is a common mistake in statistics and on Kaggle. Whether we can trust LB has nothing to do with the number 1%. Whether we can trust LB only relates to the number of rows in public LB (being scored). Since public LB is 1,000,000 million customers, then 1% is 10,000 customers being scored which is trustworthy.</p>\n<p>The common analogy is predicting elections. Consider 2 scenarios. Our school has 10,000 students and we poll 100 students (1%) to predict who will win the school election. Our state has 10,000,000 people and we poll 100 people (0.001%) to predict who will win the state election. In both cases, we trust our poll <strong>equally</strong> because the trustworthiness of our poll only depends on the number 100 and not 1% nor 0.001%.</p>\n<p>In both cases, if <code>p</code> is the proportion of people from our 100 sample who favor choice <code>A</code> for election. Then we are 95% confident that choice <code>A</code> will have result <code>p</code> with plus minus <code>(1.96 * sqrt(p * (1-p)) / 10)</code> in both cases.</p>",
          "rawMarkdown": "This is a common mistake in statistics and on Kaggle. Whether we can trust LB has nothing to do with the number 1%. Whether we can trust LB only relates to the number of rows in public LB (being scored). Since public LB is 1,000,000 million customers, then 1% is 10,000 customers being scored which is trustworthy.\n\nThe common analogy is predicting elections. Consider 2 scenarios. Our school has 10,000 students and we poll 100 students (1%) to predict who will win the school election. Our state has 10,000,000 people and we poll 100 people (0.001%) to predict who will win the state election. In both cases, we trust our poll **equally** because the trustworthiness of our poll only depends on the number 100 and not 1% nor 0.001%.\n\nIn both cases, if `p` is the proportion of people from our 100 sample who favor choice `A` for election. Then we are 95% confident that choice `A` will have result `p` with plus minus `(1.96 * sqrt(p * (1-p)) / 10)` in both cases.",
          "votes": 10
        },
        {
          "id": 1699893,
          "postDate": "2022-02-21T14:06:39.223Z",
          "content": "<p>Thank you for this very logical explanation, it makes much more sens now !</p>",
          "rawMarkdown": "Thank you for this very logical explanation, it makes much more sens now !",
          "votes": 1
        },
        {
          "id": 1701103,
          "postDate": "2022-02-22T14:22:17.677Z",
          "content": "<p>Please, ignore this comment.</p>",
          "rawMarkdown": "Please, ignore this comment."
        },
        {
          "id": 1701111,
          "postDate": "2022-02-22T14:26:07.893Z",
          "content": "<p>Please, ignore this comment as I have already found my own original answer, thanks. </p>",
          "rawMarkdown": "Please, ignore this comment as I have already found my own original answer, thanks. "
        }
      ]
    },
    {
      "id": 1699279,
      "postDate": "2022-02-21T04:43:05.333Z",
      "content": "<p>Hi, I understand that you truly recommend (not only in this post but also others too recently) to create a CV scheme that mimics the test data in order to get good results. </p>\n<p>Thank you for your recommendations.</p>",
      "rawMarkdown": "Hi, I understand that you truly recommend (not only in this post but also others too recently) to create a CV scheme that mimics the test data in order to get good results. \n\nThank you for your recommendations."
    },
    {
      "id": 1842859,
      "postDate": "2022-07-04T10:52:30.830Z",
      "content": "<p>Thanks for sharing. very useful to all concerned.</p>",
      "rawMarkdown": "Thanks for sharing. very useful to all concerned."
    },
    {
      "id": 1842355,
      "postDate": "2022-07-03T23:58:17.657Z",
      "content": "<p>Thank you so much, it's very  helpfull 👌.</p>",
      "rawMarkdown": "Thank you so much, it's very  helpfull 👌."
    },
    {
      "id": 1701884,
      "postDate": "2022-02-23T06:26:33.700Z",
      "content": "<p>Dear Chris, thank you for sharing relevant content and the clear explanation. I do have a question though; under <code>Valid Data</code> section you explain that for validation you select customers who made a purchse in the last week of training data. I think that one should select customers who purchased in the last week of training data through the <code>sales_channel_id = 2</code> (online channel) because i believe that for the test set we are supposed to recommend products for online customers. What do you think?</p>",
      "rawMarkdown": "Dear Chris, thank you for sharing relevant content and the clear explanation. I do have a question though; under `Valid Data` section you explain that for validation you select customers who made a purchse in the last week of training data. I think that one should select customers who purchased in the last week of training data through the `sales_channel_id = 2` (online channel) because i believe that for the test set we are supposed to recommend products for online customers. What do you think?",
      "replies": [
        {
          "id": 1702270,
          "postDate": "2022-02-23T13:26:49.190Z",
          "content": "<p>The evaluation page says</p>\n<blockquote>\n  <p>Customer that did not make any purchase during test period are excluded from the scoring.</p>\n</blockquote>\n<p>Where did you read that we should exclude <code>sales_channel_id=1</code> (in person sales) from prediction?</p>",
          "rawMarkdown": "The evaluation page says\n\n>Customer that did not make any purchase during test period are excluded from the scoring.\n\nWhere did you read that we should exclude `sales_channel_id=1` (in person sales) from prediction?",
          "votes": 1
        },
        {
          "id": 1702356,
          "postDate": "2022-02-23T14:41:49.897Z",
          "content": "<p>The description page says</p>\n<blockquote>\n  <p>Our online store offers shoppers an extensive selection of products to browse through. But with too many choices, customers might not quickly find what interests them or what they are looking for, and ultimately, they might not make a purchase. To enhance the shopping experience, product recommendations are key.</p>\n</blockquote>\n<p>It does not explicitly says that we should exclude <code>sales_channel_id=1</code> from the test set but I am guessing that it is.</p>",
          "rawMarkdown": "The description page says\n\n> Our online store offers shoppers an extensive selection of products to browse through. But with too many choices, customers might not quickly find what interests them or what they are looking for, and ultimately, they might not make a purchase. To enhance the shopping experience, product recommendations are key.\n\nIt does not explicitly says that we should exclude `sales_channel_id=1` from the test set but I am guessing that it is.",
          "votes": 2
        },
        {
          "id": 1702384,
          "postDate": "2022-02-23T15:06:43.137Z",
          "content": "<p>We should ask Kaggle. (This is important). It is my understanding that both online and in-person sales will be scored (for customers who made at least 1 purchase).</p>",
          "rawMarkdown": "We should ask Kaggle. (This is important). It is my understanding that both online and in-person sales will be scored (for customers who made at least 1 purchase).",
          "votes": 3
        },
        {
          "id": 1702408,
          "postDate": "2022-02-23T15:26:46.137Z",
          "content": "<p>Indeed it is a good idea to ask Kaggle for clarification. <br>\nOne more thing you mentioned <code>...(for customers who made at least 1 purchase)</code> I have checked the test set and there are approximately <code>10000</code> customers who are not seen in the train set so it is not only customers who made at least 1 purchase who are to be scored there are also new customers. So there is a user <code>cold start</code> side to this competition.</p>",
          "rawMarkdown": "Indeed it is a good idea to ask Kaggle for clarification. \nOne more thing you mentioned `...(for customers who made at least 1 purchase)` I have checked the test set and there are approximately `10000` customers who are not seen in the train set so it is not only customers who made at least 1 purchase who are to be scored there are also new customers. So there is a user `cold start` side to this competition.",
          "votes": 1
        },
        {
          "id": 1708473,
          "postDate": "2022-03-01T13:29:31.790Z",
          "content": "<p>Yes this is correct. \" All customers who made purchases during the test period are scored, regardless of whether they had purchase history in the training data\"</p>",
          "rawMarkdown": "Yes this is correct. \" All customers who made purchases during the test period are scored, regardless of whether they had purchase history in the training data\"",
          "votes": 6
        }
      ]
    },
    {
      "id": 1714288,
      "postDate": "2022-03-06T18:53:03.443Z",
      "content": "<p>Do you want<br>\n<code>valid = valid.loc[ train.t_dat &gt;= pd.to_datetime('2020-09-16') ]</code><br>\nto be<br>\n<code>valid = valid.loc[ valid.t_dat &gt;= pd.to_datetime('2020-09-16') ]</code></p>",
      "rawMarkdown": "Do you want\n`valid = valid.loc[ train.t_dat >= pd.to_datetime('2020-09-16') ]`\nto be\n`valid = valid.loc[ valid.t_dat >= pd.to_datetime('2020-09-16') ]`",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1714295,
          "postDate": "2022-03-06T18:58:51.913Z",
          "content": "<p>Yes, you are correct. Thank you. I updated my post.</p>",
          "rawMarkdown": "Yes, you are correct. Thank you. I updated my post."
        }
      ]
    },
    {
      "id": 1713210,
      "postDate": "2022-03-05T18:51:57.270Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1708206,
      "postDate": "2022-03-01T08:06:18.213Z",
      "content": "<p>You are right but i was thinking that can we use sklearn's timeseries split here or not instead of manually splitting the dataset.</p>",
      "rawMarkdown": "You are right but i was thinking that can we use sklearn's timeseries split here or not instead of manually splitting the dataset.",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 1715523,
          "postDate": "2022-03-08T05:35:49.370Z",
          "content": "<p>Nice pick, I am also curious about it.</p>",
          "rawMarkdown": "Nice pick, I am also curious about it."
        }
      ]
    },
    {
      "id": 1704960,
      "postDate": "2022-02-26T02:39:57.013Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1775948,
      "postDate": "2022-05-03T14:57:37.653Z",
      "content": "<p>Very nice sharing! Thanks!</p>",
      "rawMarkdown": "Very nice sharing! Thanks!",
      "votes": 1
    },
    {
      "id": 1747207,
      "postDate": "2022-04-06T13:34:49.633Z",
      "content": "<p>Thanks for the info</p>",
      "rawMarkdown": "Thanks for the info",
      "votes": 1
    },
    {
      "id": 1710181,
      "postDate": "2022-03-02T19:47:52.727Z",
      "content": "<p>Very helpful. Thanks for sharing</p>",
      "rawMarkdown": "Very helpful. Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1709105,
      "postDate": "2022-03-02T01:03:48.337Z",
      "content": "<p>good job, thanks</p>",
      "rawMarkdown": "good job, thanks",
      "votes": 1
    },
    {
      "id": 1705278,
      "postDate": "2022-02-26T10:24:29.463Z",
      "content": "<p>thanks man</p>",
      "rawMarkdown": "thanks man",
      "votes": 1
    },
    {
      "id": 1699954,
      "postDate": "2022-02-21T14:52:30.880Z",
      "content": "<p>Pretty useful! thanks.</p>",
      "rawMarkdown": "Pretty useful! thanks.",
      "votes": 1
    },
    {
      "id": 1711608,
      "postDate": "2022-03-04T06:34:09.903Z",
      "content": "<p>Thanks for sharing, it's useful.</p>",
      "rawMarkdown": "Thanks for sharing, it's useful.",
      "votes": 2
    },
    {
      "id": 1711255,
      "postDate": "2022-03-03T18:49:58.240Z",
      "content": "<p>Very useful. Thanks!</p>",
      "rawMarkdown": "Very useful. Thanks!",
      "votes": 2
    },
    {
      "id": 1704826,
      "postDate": "2022-02-25T21:29:30.230Z",
      "content": "<p>Very concise and direct.<br>\nThank you. </p>",
      "rawMarkdown": " Very concise and direct.\nThank you. \n",
      "votes": 2
    },
    {
      "id": 1704642,
      "postDate": "2022-02-25T17:30:04.640Z",
      "content": "<p>Amazing notebook!! Thanks for sharing</p>",
      "rawMarkdown": "Amazing notebook!! Thanks for sharing",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1699322,
      "author_name": "Carl McBride Ellis",
      "author_url": "",
      "post_date": "2022-02-21T05:26:51.893000",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>Thank you so much for your (as always) helpful and informative post. If I may be so bold as to ask a rather general (not related to this competition) and somewhat naive question (I am revising my understanding of cross-validation and <a href=\"https://www.kaggle.com/questions-and-answers/307921\" target=\"_blank\">I find my understanding is shaky</a>).</p>\n<p>My question is as follows; by comparing ones local CV to the LB is one not basically treating the LB as an ersatz hold-out test set, and if that is the case, would it not be better (if one has enough training data) to create ones own hold-out test set of a similar size, extracted from the training data, and where one can be sure that it is i.i.d, rather than assuming that the kaggle Public LB data is i.i.d? </p>\n<p>Away from kaggle there is no Public LB, so when people say <em>\"Trust your CV\"</em>, are they really not essentially  saying <em>\"Distrust the Public LB\"</em>?</p>\n<p>All the best,<br>\ncarl</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1699816,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-21T13:15:59.420000",
          "content": "<p>There is much to say about this (and your other post). I will make a quick comment and update it in the days to come.</p>\n<blockquote>\n  <p>comparing ones local CV to the LB is one not basically treating the LB as an ersatz hold-out test</p>\n</blockquote>\n<p>Not exactly. The one difference between real life and Kaggle is that we are never sure about what's in Kaggle's test dataset. In some competitions, test dataset is very different than train data. The purpose of comparing CV to LB is a method of detective work to determine what is in Kaggle's test dataset. </p>\n<p>If CV scores and LB scores are related, then most likely the relationship between your local folds is the same relationship between Kaggle's train and test. If not, try other CV splitting techniques until it is. Also, we modify our modeling to generalize in the direction of test's shift from train. </p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1700283,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-02-21T19:42:58.227000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<blockquote>\n  <p>\"<em>There is much to say about this…and update it in the days to come.</em>\"</p>\n</blockquote>\n<p>That would be wonderful; given that CV, or variants of, is an integral part of ML, and is ubiquitous in almost all kaggle competitions, your insights regarding general CV advice,  'hints and tips', or even bespoke approaches would be most appreciated by myself, and I somewhat suspect by many others too!</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1702052,
      "author_name": "Roman",
      "author_url": "",
      "post_date": "2022-02-23T09:28:13.390000",
      "content": "<p>Use can avoid this:</p>\n<pre><code>valid['prediction'] =\\\n    valid.prediction.apply(lambda x: ' '.join(['0'+str(k) for k in x]))\n</code></pre>\n<p>By this on line 1:</p>\n<pre><code>valid = pd.read_csv('transactions_train.csv', dtype={'article_id': 'str'})\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 1702267,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-23T13:25:06.253000",
          "content": "<p>Note that the first line of code (in your comment) does two things (1) convert int to str, (2) combine multiple str into one prediction string. Your second line of code does only one thing (1) convert int to str.</p>\n<p>So if we use your second line, then we can change the first line to (i.e we can't remove the first line) </p>\n<pre><code>valid['prediction'] = valid.prediction.apply(lambda x: ' '.join(x))\n</code></pre>\n<p>However in general, IMO, it is best to load an integer as an integer because integers are smaller in memory than strings and computers can process integers faster than strings. Then we convert to string in the final line of code.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1733968,
      "author_name": "jz",
      "author_url": "",
      "post_date": "2022-03-24T20:16:54.930000",
      "content": "<p>This is very helpful, thank you so much!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1731478,
      "author_name": "Pankaj Kumar",
      "author_url": "",
      "post_date": "2022-03-22T12:06:04.980000",
      "content": "<p>wonderful idea! Thank you for sharing this.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1730261,
      "author_name": "JasonBian",
      "author_url": "",
      "post_date": "2022-03-21T04:20:41.343000",
      "content": "<p>Anyone having trouble installing ml_metrics for mapk package?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1714376,
      "author_name": "Iurii Uspangaliev",
      "author_url": "",
      "post_date": "2022-03-06T22:41:18.773000",
      "content": "<p>Why are you using combined weeks instead of just <br>\n2020-09-16 - 2020-09-22 <br>\n2020-09-09 - 2020-09-15 <br>\n2020-09-02 - 2020-09-08? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1714393,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-03-07T00:30:18.553000",
          "content": "<p>This is exactly what i'm saying. </p>\n<p>The extra dates are indicating that when we validation on <code>20-20-09-16 to 2020-09-22</code> then we must train on data <code>20-20-09-15</code> and before. (And when we validate on <code>20-20-09-09 to 2020-09-15</code> we must train on data <code>20-20--09-08</code> and before, etc etc)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1714608,
          "author_name": "Iurii Uspangaliev",
          "author_url": "",
          "post_date": "2022-03-07T06:43:12.067000",
          "content": "<p>Yes, my bad, just saw &amp; symbol and my brain weirdly concluded that it's two weeks combined)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1713211,
      "author_name": "Kevan Rajaram",
      "author_url": "",
      "post_date": "2022-03-05T18:52:48.277000",
      "content": "<p>Thanks for sharing! Very useful notebook.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1711936,
      "author_name": "geekdad",
      "author_url": "",
      "post_date": "2022-03-04T13:35:07.380000",
      "content": "<p>Has anyone checked the CV results with sampling of the dataset?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1710417,
      "author_name": "r48n34",
      "author_url": "",
      "post_date": "2022-03-03T03:05:57.317000",
      "content": "<p>Nice, very useful sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1705706,
      "author_name": "Prabhav Sanga",
      "author_url": "",
      "post_date": "2022-02-26T18:20:13.297000",
      "content": "<p>Amazing Stuff .</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1702592,
      "author_name": "Asadullah AbdulJabbar",
      "author_url": "",
      "post_date": "2022-02-23T18:22:01.430000",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1703473,
      "author_name": "Pham Vu Dung",
      "author_url": "",
      "post_date": "2022-02-24T14:27:18.103000",
      "content": "<p>Very useful notebook! Thanks for sharing.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1702339,
      "author_name": "HIRO",
      "author_url": "",
      "post_date": "2022-02-23T14:28:17.037000",
      "content": "<p>I think it's a smart idea. It will be helpful</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1701235,
      "author_name": "FPiotro",
      "author_url": "",
      "post_date": "2022-02-22T15:50:31.387000",
      "content": "<p>Important information and high quality work 👏 Thank you very much <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Good luck for the competition!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1701024,
      "author_name": "James McNeill",
      "author_url": "",
      "post_date": "2022-02-22T13:16:00.657000",
      "content": "<p>Thanks very much for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> great explanation on the importance of the time series element to the cross-validation.  I am just wondering if the season will have an impact on the recommendations. Something to investigate.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1701154,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-22T14:51:45.123000",
          "content": "<p>Good suggestion. I suspect that time of year does affect what clothes customers purchase.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1699644,
      "author_name": "Mohamed Annis SOUAMES",
      "author_url": "",
      "post_date": "2022-02-21T10:28:11.827000",
      "content": "<p>Thank's for this great explanation, one question however : How to deal with the very small ratio of public LB (it's only 1% of all private data, while the private LB will be the remaining 99%).</p>\n<p>Should we take this into account while validating ? like taking 1% of the validation week to compute the score that would probably be shown on the public LB and use the whole validation week to get the real CV score (which hopefully will be the same on private LB) ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1699805,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-21T13:02:58.253000",
          "content": "<p>This is a common mistake in statistics and on Kaggle. Whether we can trust LB has nothing to do with the number 1%. Whether we can trust LB only relates to the number of rows in public LB (being scored). Since public LB is 1,000,000 million customers, then 1% is 10,000 customers being scored which is trustworthy.</p>\n<p>The common analogy is predicting elections. Consider 2 scenarios. Our school has 10,000 students and we poll 100 students (1%) to predict who will win the school election. Our state has 10,000,000 people and we poll 100 people (0.001%) to predict who will win the state election. In both cases, we trust our poll <strong>equally</strong> because the trustworthiness of our poll only depends on the number 100 and not 1% nor 0.001%.</p>\n<p>In both cases, if <code>p</code> is the proportion of people from our 100 sample who favor choice <code>A</code> for election. Then we are 95% confident that choice <code>A</code> will have result <code>p</code> with plus minus <code>(1.96 * sqrt(p * (1-p)) / 10)</code> in both cases.</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1699893,
          "author_name": "Mohamed Annis SOUAMES",
          "author_url": "",
          "post_date": "2022-02-21T14:06:39.223000",
          "content": "<p>Thank you for this very logical explanation, it makes much more sens now !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1701103,
          "author_name": "Mauricio Maroto",
          "author_url": "",
          "post_date": "2022-02-22T14:22:17.677000",
          "content": "<p>Please, ignore this comment.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1701111,
          "author_name": "Mauricio Maroto",
          "author_url": "",
          "post_date": "2022-02-22T14:26:07.893000",
          "content": "<p>Please, ignore this comment as I have already found my own original answer, thanks. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1699279,
      "author_name": "Mauricio Maroto",
      "author_url": "",
      "post_date": "2022-02-21T04:43:05.333000",
      "content": "<p>Hi, I understand that you truly recommend (not only in this post but also others too recently) to create a CV scheme that mimics the test data in order to get good results. </p>\n<p>Thank you for your recommendations.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1842859,
      "author_name": "Sohail Ahmed",
      "author_url": "",
      "post_date": "2022-07-04T10:52:30.830000",
      "content": "<p>Thanks for sharing. very useful to all concerned.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1842355,
      "author_name": "Sagouma Mohamed Sofiane",
      "author_url": "",
      "post_date": "2022-07-03T23:58:17.657000",
      "content": "<p>Thank you so much, it's very  helpfull 👌.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1701884,
      "author_name": "wti 200",
      "author_url": "",
      "post_date": "2022-02-23T06:26:33.700000",
      "content": "<p>Dear Chris, thank you for sharing relevant content and the clear explanation. I do have a question though; under <code>Valid Data</code> section you explain that for validation you select customers who made a purchse in the last week of training data. I think that one should select customers who purchased in the last week of training data through the <code>sales_channel_id = 2</code> (online channel) because i believe that for the test set we are supposed to recommend products for online customers. What do you think?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1702270,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-23T13:26:49.190000",
          "content": "<p>The evaluation page says</p>\n<blockquote>\n  <p>Customer that did not make any purchase during test period are excluded from the scoring.</p>\n</blockquote>\n<p>Where did you read that we should exclude <code>sales_channel_id=1</code> (in person sales) from prediction?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1702356,
          "author_name": "wti 200",
          "author_url": "",
          "post_date": "2022-02-23T14:41:49.897000",
          "content": "<p>The description page says</p>\n<blockquote>\n  <p>Our online store offers shoppers an extensive selection of products to browse through. But with too many choices, customers might not quickly find what interests them or what they are looking for, and ultimately, they might not make a purchase. To enhance the shopping experience, product recommendations are key.</p>\n</blockquote>\n<p>It does not explicitly says that we should exclude <code>sales_channel_id=1</code> from the test set but I am guessing that it is.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1702384,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-23T15:06:43.137000",
          "content": "<p>We should ask Kaggle. (This is important). It is my understanding that both online and in-person sales will be scored (for customers who made at least 1 purchase).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1702408,
          "author_name": "wti 200",
          "author_url": "",
          "post_date": "2022-02-23T15:26:46.137000",
          "content": "<p>Indeed it is a good idea to ask Kaggle for clarification. <br>\nOne more thing you mentioned <code>...(for customers who made at least 1 purchase)</code> I have checked the test set and there are approximately <code>10000</code> customers who are not seen in the train set so it is not only customers who made at least 1 purchase who are to be scored there are also new customers. So there is a user <code>cold start</code> side to this competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1708473,
          "author_name": "FridaRim",
          "author_url": "",
          "post_date": "2022-03-01T13:29:31.790000",
          "content": "<p>Yes this is correct. \" All customers who made purchases during the test period are scored, regardless of whether they had purchase history in the training data\"</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 1714288,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-06T18:53:03.443000",
      "content": "<p>Do you want<br>\n<code>valid = valid.loc[ train.t_dat &gt;= pd.to_datetime('2020-09-16') ]</code><br>\nto be<br>\n<code>valid = valid.loc[ valid.t_dat &gt;= pd.to_datetime('2020-09-16') ]</code></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1714295,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-03-06T18:58:51.913000",
          "content": "<p>Yes, you are correct. Thank you. I updated my post.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1713210,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-05T18:51:57.270000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1708206,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-01T08:06:18.213000",
      "content": "<p>You are right but i was thinking that can we use sklearn's timeseries split here or not instead of manually splitting the dataset.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1715523,
          "author_name": "Abdul Ghaffar Sid",
          "author_url": "",
          "post_date": "2022-03-08T05:35:49.370000",
          "content": "<p>Nice pick, I am also curious about it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1704960,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-26T02:39:57.013000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1775948,
      "author_name": "Zhicheng Wang",
      "author_url": "",
      "post_date": "2022-05-03T14:57:37.653000",
      "content": "<p>Very nice sharing! Thanks!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1747207,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-06T13:34:49.633000",
      "content": "<p>Thanks for the info</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1710181,
      "author_name": "Ishan Mehta115",
      "author_url": "",
      "post_date": "2022-03-02T19:47:52.727000",
      "content": "<p>Very helpful. Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1709105,
      "author_name": "DDDFG98",
      "author_url": "",
      "post_date": "2022-03-02T01:03:48.337000",
      "content": "<p>good job, thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1705278,
      "author_name": "Anak Agung Gede Kresna M",
      "author_url": "",
      "post_date": "2022-02-26T10:24:29.463000",
      "content": "<p>thanks man</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1699954,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-21T14:52:30.880000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1711608,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-04T06:34:09.903000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1711255,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-03T18:49:58.240000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1704826,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-25T21:29:30.230000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1704642,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-25T17:30:04.640000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1699104": "The first step in every Kaggle competition is to build a reliable local validation scheme. Then we use our local validation score to evaluate experiment ideas and/or tune hyperparameters.\n\n# Train Data\nThe last day in the transaction dataframe is `2020-09-22`. The public LB contains 1 week of transactions after this date. Therefore to create a local validation that mimics Kaggle's train test relationship, we can train on all transactions before `2020-9-15`. And validate on the last week in train data.\n\n    train = pd.read_csv('transactions_train.csv')\n    train.t_dat = pd.to_datetime( train.t_dat )\n    train = train.loc[ train.t_dat <= pd.to_datetime('2020-09-15') ]\n\n# Valid Data\nThe code below will create a dataframe with only the customers who made purchases during the last week of train (which are the only ones that affect competition metric). It formats the predictions as strings like sample_submission.csv\n\n    valid = pd.read_csv('transactions_train.csv')\n    valid.t_dat = pd.to_datetime( valid.t_dat )\n    valid = valid.loc[ valid.t_dat >= pd.to_datetime('2020-09-16') ]\n    valid = valid.groupby('customer_id').article_id.apply(list).reset_index()\n    valid = valid.rename({'article_id':'prediction'},axis=1)\n    valid['prediction'] =\\\n        valid.prediction.apply(lambda x: ' '.join(['0'+str(k) for k in x]))\n\n# Compute Validation Score mAP\nTo compute your validation score, use the `train` dataframe above (which does not contain the validation dates) and then make predictions for every customer in `sample_submission.csv`. Create a column of prediction strings just like you would when submitting to Kaggle and save as `submission.csv` (to be used below).\n\n# Important x2\nOnce you have your `submission.csv` file (made by using train dataframe above which excludes last of train data), run the code below. **Important x1**, we must execute the second line below. It guarantees that your submission and valid dataframes are in the same row order. Furthermore it removes all customers who do not make predictions during validation period because these customers do not affect metric score. **Important x1**, remember to write your `article_id`s with a prefix of zero in your dataframe predictionstrings otherwise your validation and LB score will be 0.\n\n    sub = pd.read_csv('submission.csv')\n    sub = sub.set_index('customer_id').loc[valid.customer_id].reset_index()\n    mapk( valid.prediction.str.split(), sub.prediction.str.split(), k=12)\n\n# MapK Function\nThe MapK Function can be found on Kaggle's GitHub [here][2] or in Kaerururu's notebook [here][1].\n\n# CV LB Agreement\nUsing the code above, the best public notebook [here][3] has validation score 0.023 and LB 0.020. If we combine my notebook [here][4] with that notebook, then the validation score improves to 0.024 and LB improves to 0.021 [here][5]. So it appears when validation score improves then LB improves. \n\n# 3 Folds\nTo have a more reliable validation, we can use 3 or more folds and average the results. When doing 3 folds, the best public notebook [here][3] achieves 0.023, 0.021, 0.021 on folds 0,1,2 respectively and therefore achieves average 3-fold validation score of 0.022 which matches its LB score!\n\n* Fold 0 - train <= '2020-09-15', valid >= '2020-09-16'\n* Fold 1 - train <= '2020-09-08', (valid >= '2020-09-09')&(valid <= '2020-09-15')\n* Fold 2 - train <= '2020-09-01', (valid >= '2020-09-02')&(valid <= '2020-09-08')\n\n[1]: https://www.kaggle.com/kaerunantoka/h-m-how-to-calculate-map-12\n[2]: https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\n[3]: https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2\n[4]: https://www.kaggle.com/cdeotte/customers-who-bought-this-frequently-buy-this\n[5]: https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021\n",
    "1699322": "Dear @cdeotte \n\nThank you so much for your (as always) helpful and informative post. If I may be so bold as to ask a rather general (not related to this competition) and somewhat naive question (I am revising my understanding of cross-validation and [I find my understanding is shaky](https://www.kaggle.com/questions-and-answers/307921)).\n\nMy question is as follows; by comparing ones local CV to the LB is one not basically treating the LB as an ersatz hold-out test set, and if that is the case, would it not be better (if one has enough training data) to create ones own hold-out test set of a similar size, extracted from the training data, and where one can be sure that it is i.i.d, rather than assuming that the kaggle Public LB data is i.i.d? \n\nAway from kaggle there is no Public LB, so when people say *\"Trust your CV\"*, are they really not essentially  saying *\"Distrust the Public LB\"*?\n\nAll the best,\ncarl",
    "1702052": "Use can avoid this:\n```\nvalid['prediction'] =\\\n    valid.prediction.apply(lambda x: ' '.join(['0'+str(k) for k in x]))\n```\nBy this on line 1:\n```\nvalid = pd.read_csv('transactions_train.csv', dtype={'article_id': 'str'})\n```",
    "1733968": "This is very helpful, thank you so much!",
    "1731478": "wonderful idea! Thank you for sharing this.",
    "1730261": "Anyone having trouble installing ml_metrics for mapk package?",
    "1714376": "Why are you using combined weeks instead of just \n2020-09-16 - 2020-09-22 \n2020-09-09 - 2020-09-15 \n2020-09-02 - 2020-09-08? ",
    "1713211": "Thanks for sharing! Very useful notebook.",
    "1711936": "Has anyone checked the CV results with sampling of the dataset?",
    "1710417": "Nice, very useful sharing.",
    "1705706": "Amazing Stuff .",
    "1702592": "Great work @cdeotte ",
    "1703473": "Very useful notebook! Thanks for sharing.",
    "1702339": "I think it's a smart idea. It will be helpful",
    "1701235": "Important information and high quality work 👏 Thank you very much @cdeotte Good luck for the competition!",
    "1701024": "Thanks very much for sharing @cdeotte great explanation on the importance of the time series element to the cross-validation.  I am just wondering if the season will have an impact on the recommendations. Something to investigate.",
    "1699644": "Thank's for this great explanation, one question however : How to deal with the very small ratio of public LB (it's only 1% of all private data, while the private LB will be the remaining 99%).\n\nShould we take this into account while validating ? like taking 1% of the validation week to compute the score that would probably be shown on the public LB and use the whole validation week to get the real CV score (which hopefully will be the same on private LB) ?",
    "1699279": "Hi, I understand that you truly recommend (not only in this post but also others too recently) to create a CV scheme that mimics the test data in order to get good results. \n\nThank you for your recommendations.",
    "1842859": "Thanks for sharing. very useful to all concerned.",
    "1842355": "Thank you so much, it's very  helpfull 👌.",
    "1701884": "Dear Chris, thank you for sharing relevant content and the clear explanation. I do have a question though; under `Valid Data` section you explain that for validation you select customers who made a purchse in the last week of training data. I think that one should select customers who purchased in the last week of training data through the `sales_channel_id = 2` (online channel) because i believe that for the test set we are supposed to recommend products for online customers. What do you think?",
    "1714288": "Do you want\n`valid = valid.loc[ train.t_dat >= pd.to_datetime('2020-09-16') ]`\nto be\n`valid = valid.loc[ valid.t_dat >= pd.to_datetime('2020-09-16') ]`",
    "1713210": "",
    "1708206": "You are right but i was thinking that can we use sklearn's timeseries split here or not instead of manually splitting the dataset.",
    "1704960": "",
    "1775948": "Very nice sharing! Thanks!",
    "1747207": "Thanks for the info",
    "1710181": "Very helpful. Thanks for sharing",
    "1709105": "good job, thanks",
    "1705278": "thanks man",
    "1699954": "Pretty useful! thanks.",
    "1711608": "Thanks for sharing, it's useful.",
    "1711255": "Very useful. Thanks!",
    "1704826": " Very concise and direct.\nThank you. \n",
    "1704642": "Amazing notebook!! Thanks for sharing"
  }
}