{
  "id": 312653,
  "title": "Tradition approach outperforms deep learning model when data has cold-start/inactive users",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/312653",
  "author_name": "",
  "post_date": "2022-03-13T09:28:16.457935200Z",
  "votes": 88,
  "comment_count": 9,
  "views": 0,
  "content": "<h2>1. Score problem for deep model</h2>\n<ul>\n<li><p>In recent time, we are reported that deep model(deep learning, tree model,…) get better performance than traditional model in recommendation. But by viewing <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/code?competitionId=31254&amp;sortBy=scoreDescending\" target=\"_blank\">best score notebooks</a>, I am surprised that traditional approaches (trending, most purchased items for each group, …) outperform deep model in this data. I also tried to create some deep models in <a href=\"https://www.kaggle.com/astrung/recbole-lstm-sequential-for-recomendation-tutorial\" target=\"_blank\">this notebook</a> and <a href=\"https://www.kaggle.com/astrung/lstm-sequential-modelwith-item-features-tutorial\" target=\"_blank\">this notebook</a>, but it still get lower score than traditional model. Is it weird ? </p></li>\n<li><p>I think our data may has something which is different from others data in publications, so I start this notebook to investigate the problem:<br>\n<a href=\"https://www.kaggle.com/astrung/eda-extract-user-metadata-to-apply-deep-model/notebook\" target=\"_blank\">https://www.kaggle.com/astrung/eda-extract-user-metadata-to-apply-deep-model/notebook</a></p></li>\n<li><p>My notebook prove this hypothesis: our data has too many customers which theirs preference can not be predicted from past data, or data for each user is too small for deep model.</p></li>\n</ul>\n<h1>Customer problem</h1>\n<p>By EDA on customers, we can see that most of users in test data are cold start user, or user who is inactive in a long time, so their data isn't enough, or doesn't reflect their interest correctly:</p>\n<ul>\n<li><p>Inactive users: In our test data(1371980 users), <strong>509256 users(37% of all user) have been inactive for a year</strong>-they stop buying from before 2020. Beside, <strong>373171 users(27% of all user) have been inactive for 3 months</strong>(Sep, Aug, July) or more in 2020. In other words, <strong>882427(64% of all user) users have been inactive in our test data</strong>, they have gave up in a long time, and then they reappear in our test set. In most scenarios, customers give up on a system because their interest/priority factors have been changed, so their past data isn't enough to inference their desire now. Example: I gave up on H&amp;M a year ago because I have changed my fashion style, but now I have seen a sale-off/advertising campaign/hot trending in H&amp;M, so i comeback. With this type of user, my advice is avoiding use their past data correctly for recommendation. You should use items in sale-off/advertising campaign/hot trending, because they are likely the reason they come back. The more time they disappeared in past data, the more challenged correct prediction for their interest. </p></li>\n<li><p>Cold-start users: In recommendation, cold-start users are some users who don't have past data, or their past data is too small to inference correct recommendation. In this notebook, I demonstrate that most of users have very small transaction data. <strong>In detail, in 862724 users who have transactions in 2020, 547161 users(63% all users) have number of transactions &lt; 10. In other words, we have too many cold-start users in test data, and deep models don't work with this type of user</strong>. Deep model only works with users who have big number of transactions, not cold-start users or inactive users.</p></li>\n</ul>\n<h2>2.Solution</h2>\n<p>Deep model only works with active/non cold start users. But in our test data, there are only 9% users who satisfy this condition. So can we give up deep model ?</p>\n<p><strong>In this case, my advice is using hybrid approach. You should use sale-off/advertising campaign/hot trending for 92% users who is inactive/cold-start in test data, then you can use deep model with remaining users</strong>. We should use deep models for right use cases. By combining deep model into general approach, I got some higher score than origin general approach. After cleaning my notebook, I will publish it as soon as possible. </p>\n<h1>Dataset for anyone who want to use directly inactive/cold start user</h1>\n<p>I have already published inactive/cold-start user features in this customer metadata dataset, so you can use inactive/cold-start information directly for your hybrid approach:</p>\n<ul>\n<li>dataset: <a href=\"https://www.kaggle.com/astrung/hm-customer-metadata\" target=\"_blank\">https://www.kaggle.com/astrung/hm-customer-metadata</a></li>\n<li>notebook: <a href=\"https://www.kaggle.com/astrung/eda-extract-user-metadata-to-apply-deep-model/notebook\" target=\"_blank\">https://www.kaggle.com/astrung/eda-extract-user-metadata-to-apply-deep-model/notebook</a></li>\n</ul>\n<p>I also created a dataset for items, which extract sale-off/advertising campaign from transaction data. You can get sale-off information from this data for your general recommendation. If you want to find more information about article and campaign, please check my following notebook:</p>\n<ul>\n<li>sale-off item dataset: <a href=\"https://www.kaggle.com/astrung/hm-article-capaign\" target=\"_blank\">https://www.kaggle.com/astrung/hm-article-capaign</a></li>\n<li>sale-off item notebook: <a href=\"https://www.kaggle.com/astrung/eda-extract-campaign-from-transactions\" target=\"_blank\">https://www.kaggle.com/astrung/eda-extract-campaign-from-transactions</a></li>\n</ul>",
  "messages": [
    {
      "id": "1720944",
      "postDate": "03/13/2022 09:28:16",
      "content": "<h2>1. Score problem for deep model</h2>\n<ul>\n<li><p>In recent time, we are reported that deep model(deep learning, tree model,…) get better performance than traditional model in recommendation. But by viewing <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/code?competitionId=31254&amp;sortBy=scoreDescending\" target=\"_blank\">best score notebooks</a>, I am surprised that traditional approaches (trending, most purchased items for each group, …) outperform deep model in this data. I also tried to create some deep models in <a href=\"https://www.kaggle.com/astrung/recbole-lstm-sequential-for-recomendation-tutorial\" target=\"_blank\">this notebook</a> and <a href=\"https://www.kaggle.com/astrung/lstm-sequential-modelwith-item-features-tutorial\" target=\"_blank\">this notebook</a>, but it still get lower score than traditional model. Is it weird ? </p></li>\n<li><p>I think our data may has something which is different from others data in publications, so I start this notebook to investigate the problem:<br>\n<a href=\"https://www.kaggle.com/astrung/eda-extract-user-metadata-to-apply-deep-model/notebook\" target=\"_blank\">https://www.kaggle.com/astrung/eda-extract-user-metadata-to-apply-deep-model/notebook</a></p></li>\n<li><p>My notebook prove this hypothesis: our data has too many customers which theirs preference can not be predicted from past data, or data for each user is too small for deep model.</p></li>\n</ul>\n<h1>Customer problem</h1>\n<p>By EDA on customers, we can see that most of users in test data are cold start user, or user who is inactive in a long time, so their data isn't enough, or doesn't reflect their interest correctly:</p>\n<ul>\n<li><p>Inactive users: In our test data(1371980 users), <strong>509256 users(37% of all user) have been inactive for a year</strong>-they stop buying from before 2020. Beside, <strong>373171 users(27% of all user) have been inactive for 3 months</strong>(Sep, Aug, July) or more in 2020. In other words, <strong>882427(64% of all user) users have been inactive in our test data</strong>, they have gave up in a long time, and then they reappear in our test set. In most scenarios, customers give up on a system because their interest/priority factors have been changed, so their past data isn't enough to inference their desire now. Example: I gave up on H&amp;M a year ago because I have changed my fashion style, but now I have seen a sale-off/advertising campaign/hot trending in H&amp;M, so i comeback. With this type of user, my advice is avoiding use their past data correctly for recommendation. You should use items in sale-off/advertising campaign/hot trending, because they are likely the reason they come back. The more time they disappeared in past data, the more challenged correct prediction for their interest. </p></li>\n<li><p>Cold-start users: In recommendation, cold-start users are some users who don't have past data, or their past data is too small to inference correct recommendation. In this notebook, I demonstrate that most of users have very small transaction data. <strong>In detail, in 862724 users who have transactions in 2020, 547161 users(63% all users) have number of transactions &lt; 10. In other words, we have too many cold-start users in test data, and deep models don't work with this type of user</strong>. Deep model only works with users who have big number of transactions, not cold-start users or inactive users.</p></li>\n</ul>\n<h2>2.Solution</h2>\n<p>Deep model only works with active/non cold start users. But in our test data, there are only 9% users who satisfy this condition. So can we give up deep model ?</p>\n<p><strong>In this case, my advice is using hybrid approach. You should use sale-off/advertising campaign/hot trending for 92% users who is inactive/cold-start in test data, then you can use deep model with remaining users</strong>. We should use deep models for right use cases. By combining deep model into general approach, I got some higher score than origin general approach. After cleaning my notebook, I will publish it as soon as possible. </p>\n<h1>Dataset for anyone who want to use directly inactive/cold start user</h1>\n<p>I have already published inactive/cold-start user features in this customer metadata dataset, so you can use inactive/cold-start information directly for your hybrid approach:</p>\n<ul>\n<li>dataset: <a href=\"https://www.kaggle.com/astrung/hm-customer-metadata\" target=\"_blank\">https://www.kaggle.com/astrung/hm-customer-metadata</a></li>\n<li>notebook: <a href=\"https://www.kaggle.com/astrung/eda-extract-user-metadata-to-apply-deep-model/notebook\" target=\"_blank\">https://www.kaggle.com/astrung/eda-extract-user-metadata-to-apply-deep-model/notebook</a></li>\n</ul>\n<p>I also created a dataset for items, which extract sale-off/advertising campaign from transaction data. You can get sale-off information from this data for your general recommendation. If you want to find more information about article and campaign, please check my following notebook:</p>\n<ul>\n<li>sale-off item dataset: <a href=\"https://www.kaggle.com/astrung/hm-article-capaign\" target=\"_blank\">https://www.kaggle.com/astrung/hm-article-capaign</a></li>\n<li>sale-off item notebook: <a href=\"https://www.kaggle.com/astrung/eda-extract-campaign-from-transactions\" target=\"_blank\">https://www.kaggle.com/astrung/eda-extract-campaign-from-transactions</a></li>\n</ul>",
      "rawMarkdown": "## 1. Score problem for deep model\n* In recent time, we are reported that deep model(deep learning, tree model,...) get better performance than traditional model in recommendation. But by viewing [best score notebooks](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/code?competitionId=31254&sortBy=scoreDescending), I am surprised that traditional approaches (trending, most purchased items for each group, ...) outperform deep model in this data. I also tried to create some deep models in [this notebook](https://www.kaggle.com/astrung/recbole-lstm-sequential-for-recomendation-tutorial) and [this notebook](https://www.kaggle.com/astrung/lstm-sequential-modelwith-item-features-tutorial), but it still get lower score than traditional model. Is it weird ? \n\n* I think our data may has something which is different from others data in publications, so I start this notebook to investigate the problem:\nhttps://www.kaggle.com/astrung/eda-extract-user-metadata-to-apply-deep-model/notebook\n\n* My notebook prove this hypothesis: our data has too many customers which theirs preference can not be predicted from past data, or data for each user is too small for deep model.\n\n# Customer problem\nBy EDA on customers, we can see that most of users in test data are cold start user, or user who is inactive in a long time, so their data isn't enough, or doesn't reflect their interest correctly:\n\n* Inactive users: In our test data(1371980 users), **509256 users(37% of all user) have been inactive for a year**-they stop buying from before 2020. Beside, **373171 users(27% of all user) have been inactive for 3 months**(Sep, Aug, July) or more in 2020. In other words, **882427(64% of all user) users have been inactive in our test data**, they have gave up in a long time, and then they reappear in our test set. In most scenarios, customers give up on a system because their interest/priority factors have been changed, so their past data isn't enough to inference their desire now. Example: I gave up on H&M a year ago because I have changed my fashion style, but now I have seen a sale-off/advertising campaign/hot trending in H&M, so i comeback. With this type of user, my advice is avoiding use their past data correctly for recommendation. You should use items in sale-off/advertising campaign/hot trending, because they are likely the reason they come back. The more time they disappeared in past data, the more challenged correct prediction for their interest. \n\n* Cold-start users: In recommendation, cold-start users are some users who don't have past data, or their past data is too small to inference correct recommendation. In this notebook, I demonstrate that most of users have very small transaction data. **In detail, in 862724 users who have transactions in 2020, 547161 users(63% all users) have number of transactions < 10. In other words, we have too many cold-start users in test data, and deep models don't work with this type of user**. Deep model only works with users who have big number of transactions, not cold-start users or inactive users.\n\n## 2.Solution\nDeep model only works with active/non cold start users. But in our test data, there are only 9% users who satisfy this condition. So can we give up deep model ?\n\n**In this case, my advice is using hybrid approach. You should use sale-off/advertising campaign/hot trending for 92% users who is inactive/cold-start in test data, then you can use deep model with remaining users**. We should use deep models for right use cases. By combining deep model into general approach, I got some higher score than origin general approach. After cleaning my notebook, I will publish it as soon as possible. \n\n# Dataset for anyone who want to use directly inactive/cold start user\nI have already published inactive/cold-start user features in this customer metadata dataset, so you can use inactive/cold-start information directly for your hybrid approach:\n* dataset: https://www.kaggle.com/astrung/hm-customer-metadata\n* notebook: https://www.kaggle.com/astrung/eda-extract-user-metadata-to-apply-deep-model/notebook\n\nI also created a dataset for items, which extract sale-off/advertising campaign from transaction data. You can get sale-off information from this data for your general recommendation. If you want to find more information about article and campaign, please check my following notebook:\n* sale-off item dataset: https://www.kaggle.com/astrung/hm-article-capaign\n* sale-off item notebook: https://www.kaggle.com/astrung/eda-extract-campaign-from-transactions",
      "votes": null
    },
    {
      "id": "1721687",
      "postDate": "03/13/2022 23:57:24",
      "content": "<p>For whatever it's worth, I agree with a lot of what you have said. The only thing is that the scores on both public notebooks and leaderboard seem to be too low as yet to form an opinion.</p>",
      "rawMarkdown": "For whatever it's worth, I agree with a lot of what you have said. The only thing is that the scores on both public notebooks and leaderboard seem to be too low as yet to form an opinion.",
      "votes": null
    },
    {
      "id": "1721745",
      "postDate": "03/14/2022 01:37:17",
      "content": "<p>thank you for sharing your opinion. <br>\nAbout your ideal: \"both public notebooks and leaderboard seem to be too low as yet to form an opinion\", you mean that the gap score between deep model and traditional model is too small for concluding traditional model is better ? Or both score is too low when comparing with score of other recommendation datasets and publications ?</p>",
      "rawMarkdown": "thank you for sharing your opinion. \nAbout your ideal: \"both public notebooks and leaderboard seem to be too low as yet to form an opinion\", you mean that the gap score between deep model and traditional model is too small for concluding traditional model is better ? Or both score is too low when comparing with score of other recommendation datasets and publications ?",
      "votes": null
    },
    {
      "id": "1721756",
      "postDate": "03/14/2022 02:00:42",
      "content": "<p>IMO, both scores are too low.. in my back of napkins calculations, the scores represent only about 300-400 out of about 14k items..</p>",
      "rawMarkdown": "IMO, both scores are too low.. in my back of napkins calculations, the scores represent only about 300-400 out of about 14k items..",
      "votes": null
    },
    {
      "id": "1721786",
      "postDate": "03/14/2022 02:43:50",
      "content": "<p>Yes. I hope that we will have some notebooks with higher scores in this competition</p>",
      "rawMarkdown": "Yes. I hope that we will have some notebooks with higher scores in this competition",
      "votes": null
    },
    {
      "id": "1721812",
      "postDate": "03/14/2022 03:33:12",
      "content": "<p>Thanks for sharing the super detailed write-up! <br>\nI also think it's still a bit early in the competition-we might see more interesting DL + Hybrid based approaches appear in the next few weeks.</p>",
      "rawMarkdown": "Thanks for sharing the super detailed write-up! \nI also think it's still a bit early in the competition-we might see more interesting DL + Hybrid based approaches appear in the next few weeks.",
      "votes": null
    },
    {
      "id": "1721831",
      "postDate": "03/14/2022 03:51:05",
      "content": "<p>Thank you for your comment. I hope to see that, too.<br>\nIn early stage, it may be better to focus in 92% customers with general approach</p>",
      "rawMarkdown": "Thank you for your comment. I hope to see that, too.\nIn early stage, it may be better to focus in 92% customers with general approach",
      "votes": null
    },
    {
      "id": "1721972",
      "postDate": "03/14/2022 06:33:40",
      "content": "<p>I'm very new to the world of Recommender system problems, but I agree with you. <br>\nIt's one of the things I constantly struggle with and end up reminding myself, \"always start with the simple ideas first!\"</p>",
      "rawMarkdown": "I'm very new to the world of Recommender system problems, but I agree with you. \nIt's one of the things I constantly struggle with and end up reminding myself, \"always start with the simple ideas first!\"",
      "votes": null
    },
    {
      "id": "1738019",
      "postDate": "03/29/2022 00:56:38",
      "content": "<p>Thanks for great notebooks and discussions! I have read most of them.</p>\n<p>I validated your insight by hold-out method. Please see the discussion.<br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/315587\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/315587</a></p>",
      "rawMarkdown": "Thanks for great notebooks and discussions! I have read most of them.\n\nI validated your insight by hold-out method. Please see the discussion.\nhttps://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/315587",
      "votes": null
    },
    {
      "id": "1749656",
      "postDate": "04/08/2022 20:24:12",
      "content": "<p>Thanks, this is a good insight! It helps to explain some of my difficulties training a deep model.</p>\n<p>The customer dataset has a column labelled active.  Is it meaningful in relation to your datasets / features?</p>",
      "rawMarkdown": "Thanks, this is a good insight! It helps to explain some of my difficulties training a deep model.\n\nThe customer dataset has a column labelled active.  Is it meaningful in relation to your datasets / features?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1721687,
      "author_name": "atulverma",
      "author_url": "",
      "post_date": "03/13/2022 23:57:24",
      "content": "<p>For whatever it's worth, I agree with a lot of what you have said. The only thing is that the scores on both public notebooks and leaderboard seem to be too low as yet to form an opinion.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1721745,
          "author_name": "astrung",
          "author_url": "",
          "post_date": "03/14/2022 01:37:17",
          "content": "<p>thank you for sharing your opinion. <br>\nAbout your ideal: \"both public notebooks and leaderboard seem to be too low as yet to form an opinion\", you mean that the gap score between deep model and traditional model is too small for concluding traditional model is better ? Or both score is too low when comparing with score of other recommendation datasets and publications ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1721756,
          "author_name": "atulverma",
          "author_url": "",
          "post_date": "03/14/2022 02:00:42",
          "content": "<p>IMO, both scores are too low.. in my back of napkins calculations, the scores represent only about 300-400 out of about 14k items..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1721786,
          "author_name": "astrung",
          "author_url": "",
          "post_date": "03/14/2022 02:43:50",
          "content": "<p>Yes. I hope that we will have some notebooks with higher scores in this competition</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1721812,
      "author_name": "init27",
      "author_url": "",
      "post_date": "03/14/2022 03:33:12",
      "content": "<p>Thanks for sharing the super detailed write-up! <br>\nI also think it's still a bit early in the competition-we might see more interesting DL + Hybrid based approaches appear in the next few weeks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1721831,
          "author_name": "astrung",
          "author_url": "",
          "post_date": "03/14/2022 03:51:05",
          "content": "<p>Thank you for your comment. I hope to see that, too.<br>\nIn early stage, it may be better to focus in 92% customers with general approach</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1721972,
          "author_name": "init27",
          "author_url": "",
          "post_date": "03/14/2022 06:33:40",
          "content": "<p>I'm very new to the world of Recommender system problems, but I agree with you. <br>\nIt's one of the things I constantly struggle with and end up reminding myself, \"always start with the simple ideas first!\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1738019,
      "author_name": "hanejiyuto",
      "author_url": "",
      "post_date": "03/29/2022 00:56:38",
      "content": "<p>Thanks for great notebooks and discussions! I have read most of them.</p>\n<p>I validated your insight by hold-out method. Please see the discussion.<br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/315587\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/315587</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1749656,
      "author_name": "lane203j",
      "author_url": "",
      "post_date": "04/08/2022 20:24:12",
      "content": "<p>Thanks, this is a good insight! It helps to explain some of my difficulties training a deep model.</p>\n<p>The customer dataset has a column labelled active.  Is it meaningful in relation to your datasets / features?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1720944": "## 1. Score problem for deep model\n* In recent time, we are reported that deep model(deep learning, tree model,...) get better performance than traditional model in recommendation. But by viewing [best score notebooks](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/code?competitionId=31254&sortBy=scoreDescending), I am surprised that traditional approaches (trending, most purchased items for each group, ...) outperform deep model in this data. I also tried to create some deep models in [this notebook](https://www.kaggle.com/astrung/recbole-lstm-sequential-for-recomendation-tutorial) and [this notebook](https://www.kaggle.com/astrung/lstm-sequential-modelwith-item-features-tutorial), but it still get lower score than traditional model. Is it weird ? \n\n* I think our data may has something which is different from others data in publications, so I start this notebook to investigate the problem:\nhttps://www.kaggle.com/astrung/eda-extract-user-metadata-to-apply-deep-model/notebook\n\n* My notebook prove this hypothesis: our data has too many customers which theirs preference can not be predicted from past data, or data for each user is too small for deep model.\n\n# Customer problem\nBy EDA on customers, we can see that most of users in test data are cold start user, or user who is inactive in a long time, so their data isn't enough, or doesn't reflect their interest correctly:\n\n* Inactive users: In our test data(1371980 users), **509256 users(37% of all user) have been inactive for a year**-they stop buying from before 2020. Beside, **373171 users(27% of all user) have been inactive for 3 months**(Sep, Aug, July) or more in 2020. In other words, **882427(64% of all user) users have been inactive in our test data**, they have gave up in a long time, and then they reappear in our test set. In most scenarios, customers give up on a system because their interest/priority factors have been changed, so their past data isn't enough to inference their desire now. Example: I gave up on H&M a year ago because I have changed my fashion style, but now I have seen a sale-off/advertising campaign/hot trending in H&M, so i comeback. With this type of user, my advice is avoiding use their past data correctly for recommendation. You should use items in sale-off/advertising campaign/hot trending, because they are likely the reason they come back. The more time they disappeared in past data, the more challenged correct prediction for their interest. \n\n* Cold-start users: In recommendation, cold-start users are some users who don't have past data, or their past data is too small to inference correct recommendation. In this notebook, I demonstrate that most of users have very small transaction data. **In detail, in 862724 users who have transactions in 2020, 547161 users(63% all users) have number of transactions < 10. In other words, we have too many cold-start users in test data, and deep models don't work with this type of user**. Deep model only works with users who have big number of transactions, not cold-start users or inactive users.\n\n## 2.Solution\nDeep model only works with active/non cold start users. But in our test data, there are only 9% users who satisfy this condition. So can we give up deep model ?\n\n**In this case, my advice is using hybrid approach. You should use sale-off/advertising campaign/hot trending for 92% users who is inactive/cold-start in test data, then you can use deep model with remaining users**. We should use deep models for right use cases. By combining deep model into general approach, I got some higher score than origin general approach. After cleaning my notebook, I will publish it as soon as possible. \n\n# Dataset for anyone who want to use directly inactive/cold start user\nI have already published inactive/cold-start user features in this customer metadata dataset, so you can use inactive/cold-start information directly for your hybrid approach:\n* dataset: https://www.kaggle.com/astrung/hm-customer-metadata\n* notebook: https://www.kaggle.com/astrung/eda-extract-user-metadata-to-apply-deep-model/notebook\n\nI also created a dataset for items, which extract sale-off/advertising campaign from transaction data. You can get sale-off information from this data for your general recommendation. If you want to find more information about article and campaign, please check my following notebook:\n* sale-off item dataset: https://www.kaggle.com/astrung/hm-article-capaign\n* sale-off item notebook: https://www.kaggle.com/astrung/eda-extract-campaign-from-transactions",
    "1721687": "For whatever it's worth, I agree with a lot of what you have said. The only thing is that the scores on both public notebooks and leaderboard seem to be too low as yet to form an opinion.",
    "1721745": "thank you for sharing your opinion. \nAbout your ideal: \"both public notebooks and leaderboard seem to be too low as yet to form an opinion\", you mean that the gap score between deep model and traditional model is too small for concluding traditional model is better ? Or both score is too low when comparing with score of other recommendation datasets and publications ?",
    "1721756": "IMO, both scores are too low.. in my back of napkins calculations, the scores represent only about 300-400 out of about 14k items..",
    "1721786": "Yes. I hope that we will have some notebooks with higher scores in this competition",
    "1721812": "Thanks for sharing the super detailed write-up! \nI also think it's still a bit early in the competition-we might see more interesting DL + Hybrid based approaches appear in the next few weeks.",
    "1721831": "Thank you for your comment. I hope to see that, too.\nIn early stage, it may be better to focus in 92% customers with general approach",
    "1721972": "I'm very new to the world of Recommender system problems, but I agree with you. \nIt's one of the things I constantly struggle with and end up reminding myself, \"always start with the simple ideas first!\"",
    "1738019": "Thanks for great notebooks and discussions! I have read most of them.\n\nI validated your insight by hold-out method. Please see the discussion.\nhttps://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/315587",
    "1749656": "Thanks, this is a good insight! It helps to explain some of my difficulties training a deep model.\n\nThe customer dataset has a column labelled active.  Is it meaningful in relation to your datasets / features?"
  },
  "source": "meta"
}