{
  "id": 306946,
  "title": "Don't get yourself fooled. This is not similarity problem.",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/306946",
  "author_name": "",
  "post_date": "2022-02-11T16:07:25.524162500Z",
  "votes": 66,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I see a lot of posts saying about so-called \"similar\" competitions and most of them reference matching tasks whereas this is <strong>recommender system</strong> task.</p>\n<ul>\n<li>Does this mean that solutions based on similarities would not work? No.</li>\n<li>But then what is the difference? Glad you asked.<br>\nWhile finding items similar to those that customer has already bought might be a solution it still lacks a valuable piece of information - similarity between customers themselves. Thus if customer C1 bought items M1, M2 and M3 and customer C2 bought items M1, M2 and M4 we can say that this customers preferences are similar and we can recommend item M4 to customer C1 and item M3 to customer C2.</li>\n</ul>\n<p>So how do we find this kind of similarities between customers? Solution is called <strong>collaborative filtering</strong>.<br>\nOne (of the oldest) implementations is <strong>matrix factorization</strong>.</p>",
  "messages": [
    {
      "id": "1685915",
      "postDate": "02/11/2022 16:07:25",
      "content": "<p>I see a lot of posts saying about so-called \"similar\" competitions and most of them reference matching tasks whereas this is <strong>recommender system</strong> task.</p>\n<ul>\n<li>Does this mean that solutions based on similarities would not work? No.</li>\n<li>But then what is the difference? Glad you asked.<br>\nWhile finding items similar to those that customer has already bought might be a solution it still lacks a valuable piece of information - similarity between customers themselves. Thus if customer C1 bought items M1, M2 and M3 and customer C2 bought items M1, M2 and M4 we can say that this customers preferences are similar and we can recommend item M4 to customer C1 and item M3 to customer C2.</li>\n</ul>\n<p>So how do we find this kind of similarities between customers? Solution is called <strong>collaborative filtering</strong>.<br>\nOne (of the oldest) implementations is <strong>matrix factorization</strong>.</p>",
      "rawMarkdown": "I see a lot of posts saying about so-called \"similar\" competitions and most of them reference matching tasks whereas this is **recommender system** task.\n\n* Does this mean that solutions based on similarities would not work? No.\n* But then what is the difference? Glad you asked.\nWhile finding items similar to those that customer has already bought might be a solution it still lacks a valuable piece of information - similarity between customers themselves. Thus if customer C1 bought items M1, M2 and M3 and customer C2 bought items M1, M2 and M4 we can say that this customers preferences are similar and we can recommend item M4 to customer C1 and item M3 to customer C2.\n\nSo how do we find this kind of similarities between customers? Solution is called **collaborative filtering**.\nOne (of the oldest) implementations is **matrix factorization**.",
      "votes": null
    },
    {
      "id": "1686163",
      "postDate": "02/11/2022 20:06:51",
      "content": "<p>Thanks for sharing. Its a good heads up for us!</p>",
      "rawMarkdown": "Thanks for sharing. Its a good heads up for us!",
      "votes": null
    },
    {
      "id": "1686493",
      "postDate": "02/12/2022 06:05:19",
      "content": "<p>my suggested references from previous kaggle competition that tackle recommendation systems: <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations\" target=\"_blank\">https://www.kaggle.com/c/expedia-hotel-recommendations</a> &amp; <a href=\"https://www.kaggle.com/c/airbnb-recruiting-new-user-bookings\" target=\"_blank\">https://www.kaggle.com/c/airbnb-recruiting-new-user-bookings</a></p>",
      "rawMarkdown": "my suggested references from previous kaggle competition that tackle recommendation systems: https://www.kaggle.com/c/expedia-hotel-recommendations & https://www.kaggle.com/c/airbnb-recruiting-new-user-bookings",
      "votes": null
    },
    {
      "id": "1687670",
      "postDate": "02/13/2022 03:44:23",
      "content": "<p>Thanks for sharing , helps a lot ☺️</p>",
      "rawMarkdown": "Thanks for sharing , helps a lot ☺️",
      "votes": null
    },
    {
      "id": "1688292",
      "postDate": "02/13/2022 15:03:50",
      "content": "<p>It's been pretty funny so far to see people focusing on images, when there's a rich CF dataset right there.</p>",
      "rawMarkdown": "It's been pretty funny so far to see people focusing on images, when there's a rich CF dataset right there.",
      "votes": null
    },
    {
      "id": "1688297",
      "postDate": "02/13/2022 15:06:38",
      "content": "<p>That is exactly what I intuitively thought! I also think that in the world of fashion, it's not so much about what's similar, but also what suits each other.</p>",
      "rawMarkdown": "That is exactly what I intuitively thought! I also think that in the world of fashion, it's not so much about what's similar, but also what suits each other.",
      "votes": null
    },
    {
      "id": "1688329",
      "postDate": "02/13/2022 15:23:50",
      "content": "<p>Sorry for the noob Q: What does CF Stand for?</p>\n<p>I know we have metadata but curious about the acronym :)</p>",
      "rawMarkdown": "Sorry for the noob Q: What does CF Stand for?\n\nI know we have metadata but curious about the acronym :)",
      "votes": null
    },
    {
      "id": "1688344",
      "postDate": "02/13/2022 15:28:39",
      "content": "<p>it is Collaborative Filtering <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> </p>",
      "rawMarkdown": "it is Collaborative Filtering @init27",
      "votes": null
    },
    {
      "id": "1688451",
      "postDate": "02/13/2022 16:18:38",
      "content": "<p>I wouldn't say that similarities are not relevant here. It is a ranking problem where many strategies can be used to generate candidates to rank. One of the strategies can be based on visual similarity to your last transactions. Of course if it is any good time will tell.</p>",
      "rawMarkdown": "I wouldn't say that similarities are not relevant here. It is a ranking problem where many strategies can be used to generate candidates to rank. One of the strategies can be based on visual similarity to your last transactions. Of course if it is any good time will tell.",
      "votes": null
    },
    {
      "id": "1689551",
      "postDate": "02/14/2022 10:09:15",
      "content": "<p>I didn't say that similarities are not relevant. They are, but similarity between products is just one small piece of a puzzle. E.g. the fact that I bought jeans doesn't mean that recommending me similar jeans is a good idea. In fact it is a terrible idea - I have already bough jeans, I don't need another one.</p>\n<p>But our model should know what 'people like me' (1) buy after they have bought items 'similar' (2) to what I have bought.</p>\n<p>So there is two key points - people like me, thus similarity between people and similarity between items. But pay attention how similarity between items is used - not directly for recommendation.</p>",
      "rawMarkdown": "I didn't say that similarities are not relevant. They are, but similarity between products is just one small piece of a puzzle. E.g. the fact that I bought jeans doesn't mean that recommending me similar jeans is a good idea. In fact it is a terrible idea - I have already bough jeans, I don't need another one.\n\nBut our model should know what 'people like me' (1) buy after they have bought items 'similar' (2) to what I have bought.\n\nSo there is two key points - people like me, thus similarity between people and similarity between items. But pay attention how similarity between items is used - not directly for recommendation.",
      "votes": null
    },
    {
      "id": "1689728",
      "postDate": "02/14/2022 12:58:21",
      "content": "<p>I agree sometimes similarity is not the best approach. I was thinking about a strategy to recommend items between baskets that could be based on some visual aspect but not like you say that if you bought jeans trousers the next will be jeans trousers but maybe something that has the same texture or matching colours (matching means not exactly the same but maybe using colour palette characteristics).</p>",
      "rawMarkdown": "I agree sometimes similarity is not the best approach. I was thinking about a strategy to recommend items between baskets that could be based on some visual aspect but not like you say that if you bought jeans trousers the next will be jeans trousers but maybe something that has the same texture or matching colours (matching means not exactly the same but maybe using colour palette characteristics).",
      "votes": null
    },
    {
      "id": "1691300",
      "postDate": "02/15/2022 10:40:45",
      "content": "<p>I can not agree more with you. Those so called 'similar' competitions are not similar to this one, at least not so similar.</p>",
      "rawMarkdown": "I can not agree more with you. Those so called 'similar' competitions are not similar to this one, at least not so similar.",
      "votes": null
    },
    {
      "id": "1693367",
      "postDate": "02/16/2022 16:18:52",
      "content": "<p>You can use images to get similar products and recommend those. You can also use them for negative sampling. Even for narrowing down the products to be considered for a given customer. <br>\nMaking use of images at this point might seem a little bit esoteric or overcomplex. However, maybe some top solutions will make smart use of them. We will see!</p>",
      "rawMarkdown": "You can use images to get similar products and recommend those. You can also use them for negative sampling. Even for narrowing down the products to be considered for a given customer. \nMaking use of images at this point might seem a little bit esoteric or overcomplex. However, maybe some top solutions will make smart use of them. We will see!",
      "votes": null
    },
    {
      "id": "1694012",
      "postDate": "02/17/2022 05:42:09",
      "content": "<p>Is anyone trying as desperately as me to understand how to use \"implicit\" properly for this problem?  I'm new to data science but this did feel like implicit matrix factorization was a way to approach this.</p>",
      "rawMarkdown": "Is anyone trying as desperately as me to understand how to use \"implicit\" properly for this problem?  I'm new to data science but this did feel like implicit matrix factorization was a way to approach this.",
      "votes": null
    },
    {
      "id": "1694561",
      "postDate": "02/17/2022 14:20:06",
      "content": "<p><a href=\"https://www.kaggle.com/danofer\" target=\"_blank\">@danofer</a> You are absolutely right. Working on images without exploring just the csv files would be very time consuming and will take a long time to give us valuable insights vs just observing the columns.   </p>",
      "rawMarkdown": "danofer You are absolutely right. Working on images without exploring just the csv files would be very time consuming and will take a long time to give us valuable insights vs just observing the columns.",
      "votes": null
    },
    {
      "id": "1718605",
      "postDate": "03/11/2022 02:29:43",
      "content": "<p>Similar or not similar, This is a problem.</p>",
      "rawMarkdown": "Similar or not similar, This is a problem.",
      "votes": null
    },
    {
      "id": "1718895",
      "postDate": "03/11/2022 08:57:56",
      "content": "<p>Thanks for the reference. Looking at the scores of those earlier competitions, seems a long way to go.</p>",
      "rawMarkdown": "Thanks for the reference. Looking at the scores of those earlier competitions, seems a long way to go.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1686163,
      "author_name": "nebipeker",
      "author_url": "",
      "post_date": "02/11/2022 20:06:51",
      "content": "<p>Thanks for sharing. Its a good heads up for us!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1686493,
      "author_name": "sotaai",
      "author_url": "",
      "post_date": "02/12/2022 06:05:19",
      "content": "<p>my suggested references from previous kaggle competition that tackle recommendation systems: <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations\" target=\"_blank\">https://www.kaggle.com/c/expedia-hotel-recommendations</a> &amp; <a href=\"https://www.kaggle.com/c/airbnb-recruiting-new-user-bookings\" target=\"_blank\">https://www.kaggle.com/c/airbnb-recruiting-new-user-bookings</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1718895,
          "author_name": "atulverma",
          "author_url": "",
          "post_date": "03/11/2022 08:57:56",
          "content": "<p>Thanks for the reference. Looking at the scores of those earlier competitions, seems a long way to go.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1687670,
      "author_name": "arunasivapragasam",
      "author_url": "",
      "post_date": "02/13/2022 03:44:23",
      "content": "<p>Thanks for sharing , helps a lot ☺️</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1688292,
      "author_name": "danofer",
      "author_url": "",
      "post_date": "02/13/2022 15:03:50",
      "content": "<p>It's been pretty funny so far to see people focusing on images, when there's a rich CF dataset right there.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1688329,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/13/2022 15:23:50",
          "content": "<p>Sorry for the noob Q: What does CF Stand for?</p>\n<p>I know we have metadata but curious about the acronym :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688344,
          "author_name": "hervind",
          "author_url": "",
          "post_date": "02/13/2022 15:28:39",
          "content": "<p>it is Collaborative Filtering <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1693367,
          "author_name": "joelqv",
          "author_url": "",
          "post_date": "02/16/2022 16:18:52",
          "content": "<p>You can use images to get similar products and recommend those. You can also use them for negative sampling. Even for narrowing down the products to be considered for a given customer. <br>\nMaking use of images at this point might seem a little bit esoteric or overcomplex. However, maybe some top solutions will make smart use of them. We will see!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1688297,
      "author_name": "ismailbaris",
      "author_url": "",
      "post_date": "02/13/2022 15:06:38",
      "content": "<p>That is exactly what I intuitively thought! I also think that in the world of fashion, it's not so much about what's similar, but also what suits each other.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1688451,
      "author_name": "paweljankiewicz",
      "author_url": "",
      "post_date": "02/13/2022 16:18:38",
      "content": "<p>I wouldn't say that similarities are not relevant here. It is a ranking problem where many strategies can be used to generate candidates to rank. One of the strategies can be based on visual similarity to your last transactions. Of course if it is any good time will tell.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1689551,
          "author_name": "nroman",
          "author_url": "",
          "post_date": "02/14/2022 10:09:15",
          "content": "<p>I didn't say that similarities are not relevant. They are, but similarity between products is just one small piece of a puzzle. E.g. the fact that I bought jeans doesn't mean that recommending me similar jeans is a good idea. In fact it is a terrible idea - I have already bough jeans, I don't need another one.</p>\n<p>But our model should know what 'people like me' (1) buy after they have bought items 'similar' (2) to what I have bought.</p>\n<p>So there is two key points - people like me, thus similarity between people and similarity between items. But pay attention how similarity between items is used - not directly for recommendation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689728,
          "author_name": "paweljankiewicz",
          "author_url": "",
          "post_date": "02/14/2022 12:58:21",
          "content": "<p>I agree sometimes similarity is not the best approach. I was thinking about a strategy to recommend items between baskets that could be based on some visual aspect but not like you say that if you bought jeans trousers the next will be jeans trousers but maybe something that has the same texture or matching colours (matching means not exactly the same but maybe using colour palette characteristics).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1691300,
      "author_name": "lftuwujie",
      "author_url": "",
      "post_date": "02/15/2022 10:40:45",
      "content": "<p>I can not agree more with you. Those so called 'similar' competitions are not similar to this one, at least not so similar.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1694012,
      "author_name": "patgrady",
      "author_url": "",
      "post_date": "02/17/2022 05:42:09",
      "content": "<p>Is anyone trying as desperately as me to understand how to use \"implicit\" properly for this problem?  I'm new to data science but this did feel like implicit matrix factorization was a way to approach this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1694561,
      "author_name": "zuraiz",
      "author_url": "",
      "post_date": "02/17/2022 14:20:06",
      "content": "<p><a href=\"https://www.kaggle.com/danofer\" target=\"_blank\">@danofer</a> You are absolutely right. Working on images without exploring just the csv files would be very time consuming and will take a long time to give us valuable insights vs just observing the columns.   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1718605,
      "author_name": "shigengtian",
      "author_url": "",
      "post_date": "03/11/2022 02:29:43",
      "content": "<p>Similar or not similar, This is a problem.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1685915": "I see a lot of posts saying about so-called \"similar\" competitions and most of them reference matching tasks whereas this is **recommender system** task.\n\n* Does this mean that solutions based on similarities would not work? No.\n* But then what is the difference? Glad you asked.\nWhile finding items similar to those that customer has already bought might be a solution it still lacks a valuable piece of information - similarity between customers themselves. Thus if customer C1 bought items M1, M2 and M3 and customer C2 bought items M1, M2 and M4 we can say that this customers preferences are similar and we can recommend item M4 to customer C1 and item M3 to customer C2.\n\nSo how do we find this kind of similarities between customers? Solution is called **collaborative filtering**.\nOne (of the oldest) implementations is **matrix factorization**.",
    "1686163": "Thanks for sharing. Its a good heads up for us!",
    "1686493": "my suggested references from previous kaggle competition that tackle recommendation systems: https://www.kaggle.com/c/expedia-hotel-recommendations & https://www.kaggle.com/c/airbnb-recruiting-new-user-bookings",
    "1687670": "Thanks for sharing , helps a lot ☺️",
    "1688292": "It's been pretty funny so far to see people focusing on images, when there's a rich CF dataset right there.",
    "1688297": "That is exactly what I intuitively thought! I also think that in the world of fashion, it's not so much about what's similar, but also what suits each other.",
    "1688329": "Sorry for the noob Q: What does CF Stand for?\n\nI know we have metadata but curious about the acronym :)",
    "1688344": "it is Collaborative Filtering @init27",
    "1688451": "I wouldn't say that similarities are not relevant here. It is a ranking problem where many strategies can be used to generate candidates to rank. One of the strategies can be based on visual similarity to your last transactions. Of course if it is any good time will tell.",
    "1689551": "I didn't say that similarities are not relevant. They are, but similarity between products is just one small piece of a puzzle. E.g. the fact that I bought jeans doesn't mean that recommending me similar jeans is a good idea. In fact it is a terrible idea - I have already bough jeans, I don't need another one.\n\nBut our model should know what 'people like me' (1) buy after they have bought items 'similar' (2) to what I have bought.\n\nSo there is two key points - people like me, thus similarity between people and similarity between items. But pay attention how similarity between items is used - not directly for recommendation.",
    "1689728": "I agree sometimes similarity is not the best approach. I was thinking about a strategy to recommend items between baskets that could be based on some visual aspect but not like you say that if you bought jeans trousers the next will be jeans trousers but maybe something that has the same texture or matching colours (matching means not exactly the same but maybe using colour palette characteristics).",
    "1691300": "I can not agree more with you. Those so called 'similar' competitions are not similar to this one, at least not so similar.",
    "1693367": "You can use images to get similar products and recommend those. You can also use them for negative sampling. Even for narrowing down the products to be considered for a given customer. \nMaking use of images at this point might seem a little bit esoteric or overcomplex. However, maybe some top solutions will make smart use of them. We will see!",
    "1694012": "Is anyone trying as desperately as me to understand how to use \"implicit\" properly for this problem?  I'm new to data science but this did feel like implicit matrix factorization was a way to approach this.",
    "1694561": "danofer You are absolutely right. Working on images without exploring just the csv files would be very time consuming and will take a long time to give us valuable insights vs just observing the columns.",
    "1718605": "Similar or not similar, This is a problem.",
    "1718895": "Thanks for the reference. Looking at the scores of those earlier competitions, seems a long way to go."
  },
  "source": "meta"
}