{
  "id": 307517,
  "title": "Why last items and the most popular items are better then ALS?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/307517",
  "author_name": "Mikhail Pchelintsev",
  "post_date": "2022-02-14T14:44:36.608000",
  "votes": 21,
  "comment_count": 7,
  "views": 0,
  "content": "<p>i'm checking public notebooks and ALS gives around 0.014 while heuristics around 0.02.<br>\nis there any intuition behind it?</p>",
  "messages": [
    {
      "id": 1689852,
      "postDate": "2022-02-14T14:44:36.610Z",
      "content": "<p>i'm checking public notebooks and ALS gives around 0.014 while heuristics around 0.02.<br>\nis there any intuition behind it?</p>",
      "rawMarkdown": "i'm checking public notebooks and ALS gives around 0.014 while heuristics around 0.02.\nis there any intuition behind it?",
      "votes": 20
    },
    {
      "id": 1690825,
      "postDate": "2022-02-15T05:50:30.530Z",
      "content": "<p>Since I read great notebooks of others, I've been really interested in that point too.</p>\n<p>The only thing what I found is that around 4,000 customers have purchased more than 10 same items. Furthermore, there are customers who purchased 650 similar black T-shirts, 199 white T-shirts, 188 black T-shirts…and so on.<br>\nI'm not a professional data scientist and can't estimate the impact of these features of data, but what I now feel is that some, and huge, transactions are made not for own sake but for some specific use, and probably have nothing to do with preference and recommendation.</p>\n<p>Below is an example of huge transactions:<br>\ntransactions_df.query(\"customer_id == 'd00063b94dcb1342869d4994844a2742b5d62927f36843164fb3f818f630bca9' and article_id == '0678342001'\")</p>\n<p>The above point might be off target, because there are still many everyday transactions.<br>\nSo I'm awaiting experts' answers…</p>",
      "rawMarkdown": "Since I read great notebooks of others, I've been really interested in that point too.\n\nThe only thing what I found is that around 4,000 customers have purchased more than 10 same items. Furthermore, there are customers who purchased 650 similar black T-shirts, 199 white T-shirts, 188 black T-shirts...and so on.\nI'm not a professional data scientist and can't estimate the impact of these features of data, but what I now feel is that some, and huge, transactions are made not for own sake but for some specific use, and probably have nothing to do with preference and recommendation.\n\nBelow is an example of huge transactions:\ntransactions_df.query(\"customer_id == 'd00063b94dcb1342869d4994844a2742b5d62927f36843164fb3f818f630bca9' and article_id == '0678342001'\")\n\nThe above point might be off target, because there are still many everyday transactions.\nSo I'm awaiting experts' answers...",
      "votes": 9,
      "replies": [
        {
          "id": 1694361,
          "postDate": "2022-02-17T11:21:47.467Z",
          "content": "<p>Yes, they might be buying those items for their own business in bulk, so I don't think that product recommendations will have any impact on them.</p>",
          "rawMarkdown": "Yes, they might be buying those items for their own business in bulk, so I don't think that product recommendations will have any impact on them.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1694982,
      "postDate": "2022-02-17T21:48:13.733Z",
      "content": "<p>Have you computed a local validation score for both methods? Are you sure heuristics will do better than ALS on private LB which is 99% of test data? When i compute validation score for the best public notebook (i.e. heuristics), . This is validating with the last week in train data.</p>\n<p>UPDATE1: I need to recompute val score for public notebook. My previous calculation had an error.<br>\nUPDATE2: validation score for best public notebook using last week of train data is 0.023!</p>",
      "rawMarkdown": "Have you computed a local validation score for both methods? Are you sure heuristics will do better than ALS on private LB which is 99% of test data? When i compute validation score for the best public notebook (i.e. heuristics), ~~i only get 0.0084~~. This is validating with the last week in train data.\n\nUPDATE1: I need to recompute val score for public notebook. My previous calculation had an error.\nUPDATE2: validation score for best public notebook using last week of train data is 0.023!",
      "votes": 7,
      "replies": [
        {
          "id": 1696887,
          "postDate": "2022-02-19T07:07:58.800Z",
          "content": "<p>Hey, I have a silly question about how you are setting up your validation pipeline,</p>\n<ol>\n<li>Are you only using the customer_ids that are in the transactions_train.csv for validation?</li>\n<li>Or are you(somehow) also using those customer_ids which are in sample_submission.csv for the predictions but not in transactions_train.csv for validation so that the validation properly reflect the LB?</li>\n</ol>",
          "rawMarkdown": "Hey, I have a silly question about how you are setting up your validation pipeline,\n1. Are you only using the customer_ids that are in the transactions_train.csv for validation?\n2. Or are you(somehow) also using those customer_ids which are in sample_submission.csv for the predictions but not in transactions_train.csv for validation so that the validation properly reflect the LB?"
        },
        {
          "id": 1697263,
          "postDate": "2022-02-19T13:47:56.013Z",
          "content": "<p>I use <strong>all</strong> customer_ids in sample submission. The metric will take care of everything. If a customer does not make a prediction during the validation period, then their predictions do not affect the metric score. And our models need to make predictions for customers not in train data because these customers may be in valid period. (At a minimum, we can just predict the 12 most popular items for customers not seen in train data).</p>\n<p>A more robust validation with \"folds\" will be to do this with 5 validation periods. For \"fold 1\", use the last week of train and then train model with weeks prior. For \"fold 2\", use the second to last week of train and then train model with weeks prior. For \"fold 3, fold4, fold5\", etc etc.</p>",
          "rawMarkdown": "I use **all** customer_ids in sample submission. The metric will take care of everything. If a customer does not make a prediction during the validation period, then their predictions do not affect the metric score. And our models need to make predictions for customers not in train data because these customers may be in valid period. (At a minimum, we can just predict the 12 most popular items for customers not seen in train data).\n\nA more robust validation with \"folds\" will be to do this with 5 validation periods. For \"fold 1\", use the last week of train and then train model with weeks prior. For \"fold 2\", use the second to last week of train and then train model with weeks prior. For \"fold 3, fold4, fold5\", etc etc.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1696400,
      "postDate": "2022-02-18T20:02:00.337Z",
      "content": "<p>ALS approaches as shown on public notebooks right now don't account for \"seasonal products\". The most commonly bought items in the past are not necessarily what people is buying at a specific week (e.g. winter vs summer clothes). </p>\n<p>The simple heuristics shown on public notebooks mostly make use of recent weeks or even the last day of data.</p>",
      "rawMarkdown": "ALS approaches as shown on public notebooks right now don't account for \"seasonal products\". The most commonly bought items in the past are not necessarily what people is buying at a specific week (e.g. winter vs summer clothes). \n\nThe simple heuristics shown on public notebooks mostly make use of recent weeks or even the last day of data.",
      "votes": 2
    },
    {
      "id": 1699760,
      "postDate": "2022-02-21T12:15:10.533Z",
      "content": "<p>This is because the LB is calculated with only 1% of the test data.<br>\nIf you will go and check this notebook <a href=\"https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2/notebook\" target=\"_blank\">https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2/notebook</a> based on heuristics and at the last if you check that how many customers were recommended the dummy predictions then you will find that approximately 85% of the customers were recommended that dummy predictions hence definetly this submission will almost fail on the private LB i.e is 99% of the test set.</p>",
      "rawMarkdown": "This is because the LB is calculated with only 1% of the test data.\nIf you will go and check this notebook https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2/notebook based on heuristics and at the last if you check that how many customers were recommended the dummy predictions then you will find that approximately 85% of the customers were recommended that dummy predictions hence definetly this submission will almost fail on the private LB i.e is 99% of the test set.",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1690825,
      "author_name": "Ryo Asashi",
      "author_url": "",
      "post_date": "2022-02-15T05:50:30.530000",
      "content": "<p>Since I read great notebooks of others, I've been really interested in that point too.</p>\n<p>The only thing what I found is that around 4,000 customers have purchased more than 10 same items. Furthermore, there are customers who purchased 650 similar black T-shirts, 199 white T-shirts, 188 black T-shirts…and so on.<br>\nI'm not a professional data scientist and can't estimate the impact of these features of data, but what I now feel is that some, and huge, transactions are made not for own sake but for some specific use, and probably have nothing to do with preference and recommendation.</p>\n<p>Below is an example of huge transactions:<br>\ntransactions_df.query(\"customer_id == 'd00063b94dcb1342869d4994844a2742b5d62927f36843164fb3f818f630bca9' and article_id == '0678342001'\")</p>\n<p>The above point might be off target, because there are still many everyday transactions.<br>\nSo I'm awaiting experts' answers…</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1694361,
          "author_name": "Susnato Dhar",
          "author_url": "",
          "post_date": "2022-02-17T11:21:47.467000",
          "content": "<p>Yes, they might be buying those items for their own business in bulk, so I don't think that product recommendations will have any impact on them.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1694982,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-02-17T21:48:13.733000",
      "content": "<p>Have you computed a local validation score for both methods? Are you sure heuristics will do better than ALS on private LB which is 99% of test data? When i compute validation score for the best public notebook (i.e. heuristics), . This is validating with the last week in train data.</p>\n<p>UPDATE1: I need to recompute val score for public notebook. My previous calculation had an error.<br>\nUPDATE2: validation score for best public notebook using last week of train data is 0.023!</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1696887,
          "author_name": "Susnato Dhar",
          "author_url": "",
          "post_date": "2022-02-19T07:07:58.800000",
          "content": "<p>Hey, I have a silly question about how you are setting up your validation pipeline,</p>\n<ol>\n<li>Are you only using the customer_ids that are in the transactions_train.csv for validation?</li>\n<li>Or are you(somehow) also using those customer_ids which are in sample_submission.csv for the predictions but not in transactions_train.csv for validation so that the validation properly reflect the LB?</li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1697263,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-19T13:47:56.013000",
          "content": "<p>I use <strong>all</strong> customer_ids in sample submission. The metric will take care of everything. If a customer does not make a prediction during the validation period, then their predictions do not affect the metric score. And our models need to make predictions for customers not in train data because these customers may be in valid period. (At a minimum, we can just predict the 12 most popular items for customers not seen in train data).</p>\n<p>A more robust validation with \"folds\" will be to do this with 5 validation periods. For \"fold 1\", use the last week of train and then train model with weeks prior. For \"fold 2\", use the second to last week of train and then train model with weeks prior. For \"fold 3, fold4, fold5\", etc etc.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1696400,
      "author_name": "Victor Gonzalez",
      "author_url": "",
      "post_date": "2022-02-18T20:02:00.337000",
      "content": "<p>ALS approaches as shown on public notebooks right now don't account for \"seasonal products\". The most commonly bought items in the past are not necessarily what people is buying at a specific week (e.g. winter vs summer clothes). </p>\n<p>The simple heuristics shown on public notebooks mostly make use of recent weeks or even the last day of data.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1699760,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-21T12:15:10.533000",
      "content": "<p>This is because the LB is calculated with only 1% of the test data.<br>\nIf you will go and check this notebook <a href=\"https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2/notebook\" target=\"_blank\">https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2/notebook</a> based on heuristics and at the last if you check that how many customers were recommended the dummy predictions then you will find that approximately 85% of the customers were recommended that dummy predictions hence definetly this submission will almost fail on the private LB i.e is 99% of the test set.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1689852": "i'm checking public notebooks and ALS gives around 0.014 while heuristics around 0.02.\nis there any intuition behind it?",
    "1690825": "Since I read great notebooks of others, I've been really interested in that point too.\n\nThe only thing what I found is that around 4,000 customers have purchased more than 10 same items. Furthermore, there are customers who purchased 650 similar black T-shirts, 199 white T-shirts, 188 black T-shirts...and so on.\nI'm not a professional data scientist and can't estimate the impact of these features of data, but what I now feel is that some, and huge, transactions are made not for own sake but for some specific use, and probably have nothing to do with preference and recommendation.\n\nBelow is an example of huge transactions:\ntransactions_df.query(\"customer_id == 'd00063b94dcb1342869d4994844a2742b5d62927f36843164fb3f818f630bca9' and article_id == '0678342001'\")\n\nThe above point might be off target, because there are still many everyday transactions.\nSo I'm awaiting experts' answers...",
    "1694982": "Have you computed a local validation score for both methods? Are you sure heuristics will do better than ALS on private LB which is 99% of test data? When i compute validation score for the best public notebook (i.e. heuristics), ~~i only get 0.0084~~. This is validating with the last week in train data.\n\nUPDATE1: I need to recompute val score for public notebook. My previous calculation had an error.\nUPDATE2: validation score for best public notebook using last week of train data is 0.023!",
    "1696400": "ALS approaches as shown on public notebooks right now don't account for \"seasonal products\". The most commonly bought items in the past are not necessarily what people is buying at a specific week (e.g. winter vs summer clothes). \n\nThe simple heuristics shown on public notebooks mostly make use of recent weeks or even the last day of data.",
    "1699760": "This is because the LB is calculated with only 1% of the test data.\nIf you will go and check this notebook https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2/notebook based on heuristics and at the last if you check that how many customers were recommended the dummy predictions then you will find that approximately 85% of the customers were recommended that dummy predictions hence definetly this submission will almost fail on the private LB i.e is 99% of the test set."
  }
}