{
  "id": 370756,
  "title": "Make NN Great Again!!",
  "url": "/competitions/otto-recommender-system/discussion/370756",
  "author_name": "KKY",
  "post_date": "2022-12-06T09:56:59.900000",
  "votes": 41,
  "comment_count": 19,
  "views": 0,
  "content": "<p>We know that in the h&amp;m recommendation competition, GBDT is the winning solution for most of the top teams. But in this competition,  we don't have any attribution info for users and products,  the action sequence is all we know.  Extract information from a pure sequence data, NN maybe more effective than GBDT!</p>\n<p>Let's share the model insight, idea, experiment result here,  start with me.</p>\n<hr>\n<p>model:  covisitation<br>\nvalidation on test:</p>\n<ul>\n<li>cv 0.6302</li>\n<li>lb 0.575</li>\n</ul>\n<p>model:  Transformer<br>\nvalidation on test:</p>\n<ul>\n<li>cv 0.6378</li>\n<li>lb 0.578</li>\n</ul>",
  "messages": [
    {
      "id": 2056623,
      "postDate": "2022-12-06T09:56:59.900Z",
      "content": "<p>We know that in the h&amp;m recommendation competition, GBDT is the winning solution for most of the top teams. But in this competition,  we don't have any attribution info for users and products,  the action sequence is all we know.  Extract information from a pure sequence data, NN maybe more effective than GBDT!</p>\n<p>Let's share the model insight, idea, experiment result here,  start with me.</p>\n<hr>\n<p>model:  covisitation<br>\nvalidation on test:</p>\n<ul>\n<li>cv 0.6302</li>\n<li>lb 0.575</li>\n</ul>\n<p>model:  Transformer<br>\nvalidation on test:</p>\n<ul>\n<li>cv 0.6378</li>\n<li>lb 0.578</li>\n</ul>",
      "rawMarkdown": "We know that in the h&m recommendation competition, GBDT is the winning solution for most of the top teams. But in this competition,  we don't have any attribution info for users and products,  the action sequence is all we know.  Extract information from a pure sequence data, NN maybe more effective than GBDT!\n\nLet's share the model insight, idea, experiment result here,  start with me.\n\n---\n\nmodel:  covisitation\nvalidation on test:\n- cv 0.6302\n- lb 0.575\n\nmodel:  Transformer\nvalidation on test:\n- cv 0.6378\n- lb 0.578",
      "votes": 41
    },
    {
      "id": 2062607,
      "postDate": "2022-12-12T08:52:38.053Z",
      "content": "<p>How do you  compute validation on test?</p>\n<p>For the co-visitation model with lb 0.575, this validation notebook from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> has a cv score of 0.5643.  But it does not use test it uses last week of train.</p>\n<p>I don't understand how you get 0.6302 for that model.</p>",
      "rawMarkdown": "How do you  compute validation on test?\n\nFor the co-visitation model with lb 0.575, this validation notebook from @cdeotte has a cv score of 0.5643.  But it does not use test it uses last week of train.\n\nI don't understand how you get 0.6302 for that model.",
      "votes": 1
    },
    {
      "id": 2057555,
      "postDate": "2022-12-07T07:54:15.623Z",
      "content": "<p>Would love to see a descriptive notebook on this!</p>",
      "rawMarkdown": "Would love to see a descriptive notebook on this!",
      "votes": 1
    },
    {
      "id": 2056739,
      "postDate": "2022-12-06T11:55:47.203Z",
      "content": "<p>Amazing! . I tried it before and gave up because it took too long to train. My idea is to take the type(clicks, carts, and orders) as type embedding like position embedding  and predict next even whether happen(0 or 1) so that this model both recall candidates and rerank for all types. </p>",
      "rawMarkdown": "Amazing! . I tried it before and gave up because it took too long to train. My idea is to take the type(clicks, carts, and orders) as type embedding like position embedding  and predict next even whether happen(0 or 1) so that this model both recall candidates and rerank for all types. ",
      "votes": 1,
      "replies": [
        {
          "id": 2056811,
          "postDate": "2022-12-06T13:02:43.453Z",
          "content": "<p>Yes, train nn is slow,  now it takes about 6 hours per epoch with full data. I am trying to figure out  how to speed up the data pipeline. </p>\n<p>How long does it take for your current (XGB/LGB) to train? </p>",
          "rawMarkdown": "Yes, train nn is slow,  now it takes about 6 hours per epoch with full data. I am trying to figure out  how to speed up the data pipeline. \n\nHow long does it take for your current (XGB/LGB) to train? "
        },
        {
          "id": 2056854,
          "postDate": "2022-12-06T14:03:53.737Z",
          "content": "<p>15s per epoch with downsample data except clicks data</p>",
          "rawMarkdown": "15s per epoch with downsample data except clicks data"
        }
      ]
    },
    {
      "id": 2056631,
      "postDate": "2022-12-06T10:03:28.797Z",
      "content": "<p>That is a really cool result! Thanks for sharing 🙂</p>\n<p>I have been thinking if rnns, transformers, maybe some weird CNN,could play a role here. Turns out that it might 🙂 Very inspiring 🙂</p>",
      "rawMarkdown": "That is a really cool result! Thanks for sharing 🙂\n\nI have been thinking if rnns, transformers, maybe some weird CNN,could play a role here. Turns out that it might 🙂 Very inspiring 🙂",
      "votes": 1,
      "replies": [
        {
          "id": 2056678,
          "postDate": "2022-12-06T11:03:57.457Z",
          "content": "<p>YES.  Let us do it!</p>",
          "rawMarkdown": "YES.  Let us do it!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2056936,
      "postDate": "2022-12-06T15:37:52.157Z",
      "content": "<p>Cool~ Given the setting of only ids, I also want to try a pure NN solution</p>",
      "rawMarkdown": "Cool~ Given the setting of only ids, I also want to try a pure NN solution",
      "votes": 2
    },
    {
      "id": 2056877,
      "postDate": "2022-12-06T14:13:23.943Z",
      "content": "<p>Nice result. What does your NN predict? Does it take user sequence as input and predict next item id? Or does it take pairs of \"user-item\" and predict likely of click, cart, order, etc?</p>",
      "rawMarkdown": "Nice result. What does your NN predict? Does it take user sequence as input and predict next item id? Or does it take pairs of \"user-item\" and predict likely of click, cart, order, etc?",
      "votes": 2,
      "replies": [
        {
          "id": 2056896,
          "postDate": "2022-12-06T14:41:10.130Z",
          "content": "<p>Take pairs of \"user-item\" and predict likely of click, cart, and order  😄   I tried the first method,  but too slow to train.</p>",
          "rawMarkdown": "Take pairs of \"user-item\" and predict likely of click, cart, and order  😄   I tried the first method,  but too slow to train.",
          "votes": 5
        },
        {
          "id": 2056900,
          "postDate": "2022-12-06T14:47:35.627Z",
          "content": "<p>Yes 1.8 million unique items is a lot. I wonder whether clustering the items into 1000 groups with something like Word2Vec and then feeding a time sequence of group numbers (i.e. only 1000 unique numbers) into your transformer can help the transformer learn what items the user history is click, cart, order.</p>",
          "rawMarkdown": "Yes 1.8 million unique items is a lot. I wonder whether clustering the items into 1000 groups with something like Word2Vec and then feeding a time sequence of group numbers (i.e. only 1000 unique numbers) into your transformer can help the transformer learn what items the user history is click, cart, order.",
          "votes": 4
        },
        {
          "id": 2057377,
          "postDate": "2022-12-07T03:48:58.803Z",
          "content": "<p>hey <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a>! How do you represent the user? I mean, we have no user ids in the dataset, and session ids are not overlapping between train and test!</p>\n<p>Do you represent a <code>user</code> in the user-item scenario somehow via the truncated portion of the session? If so, how do you do it? Is this where the transformer part comes in?</p>\n<p>Really inspiring, thanks for sharing! 🙂🙌  </p>",
          "rawMarkdown": "hey @evilpsycho42! How do you represent the user? I mean, we have no user ids in the dataset, and session ids are not overlapping between train and test!\n\nDo you represent a `user` in the user-item scenario somehow via the truncated portion of the session? If so, how do you do it? Is this where the transformer part comes in?\n\nReally inspiring, thanks for sharing! 🙂🙌  ",
          "votes": 1
        },
        {
          "id": 2060608,
          "postDate": "2022-12-10T06:52:36.707Z",
          "content": "<p>Use action sequence represent user</p>",
          "rawMarkdown": "Use action sequence represent user"
        },
        {
          "id": 2060609,
          "postDate": "2022-12-10T06:52:57.127Z",
          "content": "<p>Yes,  truncated action sequence</p>",
          "rawMarkdown": "Yes,  truncated action sequence"
        },
        {
          "id": 2060611,
          "postDate": "2022-12-10T06:53:21.113Z",
          "content": "<p>Pure Transformer </p>",
          "rawMarkdown": "Pure Transformer "
        }
      ]
    },
    {
      "id": 2121027,
      "postDate": "2023-01-30T02:06:24.523Z",
      "content": "<p>Could you please share your method once the competition ends? Thanks a lot!</p>",
      "rawMarkdown": "Could you please share your method once the competition ends? Thanks a lot!"
    },
    {
      "id": 2061712,
      "postDate": "2022-12-11T12:31:14.900Z",
      "content": "<p>Amazing result!! are you using the NN for candidates generation / ranking / both? </p>",
      "rawMarkdown": "Amazing result!! are you using the NN for candidates generation / ranking / both? "
    },
    {
      "id": 2087055,
      "postDate": "2023-01-05T09:24:23.580Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2062557,
      "postDate": "2022-12-12T08:07:40.167Z",
      "content": "<p>Cool! Did you model clicks/carts/orders sequence separately or put them together? BTW, do both the recall and ranking stages use the same architecture?</p>",
      "rawMarkdown": "Cool! Did you model clicks/carts/orders sequence separately or put them together? BTW, do both the recall and ranking stages use the same architecture?",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2062607,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2022-12-12T08:52:38.053000",
      "content": "<p>How do you  compute validation on test?</p>\n<p>For the co-visitation model with lb 0.575, this validation notebook from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> has a cv score of 0.5643.  But it does not use test it uses last week of train.</p>\n<p>I don't understand how you get 0.6302 for that model.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2057555,
      "author_name": "Jasleen Sondhi",
      "author_url": "",
      "post_date": "2022-12-07T07:54:15.623000",
      "content": "<p>Would love to see a descriptive notebook on this!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2056739,
      "author_name": "Try Harder",
      "author_url": "",
      "post_date": "2022-12-06T11:55:47.203000",
      "content": "<p>Amazing! . I tried it before and gave up because it took too long to train. My idea is to take the type(clicks, carts, and orders) as type embedding like position embedding  and predict next even whether happen(0 or 1) so that this model both recall candidates and rerank for all types. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2056811,
          "author_name": "KKY",
          "author_url": "",
          "post_date": "2022-12-06T13:02:43.453000",
          "content": "<p>Yes, train nn is slow,  now it takes about 6 hours per epoch with full data. I am trying to figure out  how to speed up the data pipeline. </p>\n<p>How long does it take for your current (XGB/LGB) to train? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2056854,
          "author_name": "Try Harder",
          "author_url": "",
          "post_date": "2022-12-06T14:03:53.737000",
          "content": "<p>15s per epoch with downsample data except clicks data</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2056631,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-12-06T10:03:28.797000",
      "content": "<p>That is a really cool result! Thanks for sharing 🙂</p>\n<p>I have been thinking if rnns, transformers, maybe some weird CNN,could play a role here. Turns out that it might 🙂 Very inspiring 🙂</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2056678,
          "author_name": "KKY",
          "author_url": "",
          "post_date": "2022-12-06T11:03:57.457000",
          "content": "<p>YES.  Let us do it!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2056936,
      "author_name": "sirius",
      "author_url": "",
      "post_date": "2022-12-06T15:37:52.157000",
      "content": "<p>Cool~ Given the setting of only ids, I also want to try a pure NN solution</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2056877,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-12-06T14:13:23.943000",
      "content": "<p>Nice result. What does your NN predict? Does it take user sequence as input and predict next item id? Or does it take pairs of \"user-item\" and predict likely of click, cart, order, etc?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2056896,
          "author_name": "KKY",
          "author_url": "",
          "post_date": "2022-12-06T14:41:10.130000",
          "content": "<p>Take pairs of \"user-item\" and predict likely of click, cart, and order  😄   I tried the first method,  but too slow to train.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 2056900,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-06T14:47:35.627000",
          "content": "<p>Yes 1.8 million unique items is a lot. I wonder whether clustering the items into 1000 groups with something like Word2Vec and then feeding a time sequence of group numbers (i.e. only 1000 unique numbers) into your transformer can help the transformer learn what items the user history is click, cart, order.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2057377,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-07T03:48:58.803000",
          "content": "<p>hey <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a>! How do you represent the user? I mean, we have no user ids in the dataset, and session ids are not overlapping between train and test!</p>\n<p>Do you represent a <code>user</code> in the user-item scenario somehow via the truncated portion of the session? If so, how do you do it? Is this where the transformer part comes in?</p>\n<p>Really inspiring, thanks for sharing! 🙂🙌  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2060608,
          "author_name": "KKY",
          "author_url": "",
          "post_date": "2022-12-10T06:52:36.707000",
          "content": "<p>Use action sequence represent user</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2060609,
          "author_name": "KKY",
          "author_url": "",
          "post_date": "2022-12-10T06:52:57.127000",
          "content": "<p>Yes,  truncated action sequence</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2060611,
          "author_name": "KKY",
          "author_url": "",
          "post_date": "2022-12-10T06:53:21.113000",
          "content": "<p>Pure Transformer </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2121027,
      "author_name": "TouTie",
      "author_url": "",
      "post_date": "2023-01-30T02:06:24.523000",
      "content": "<p>Could you please share your method once the competition ends? Thanks a lot!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2061712,
      "author_name": "EnricRovira",
      "author_url": "",
      "post_date": "2022-12-11T12:31:14.900000",
      "content": "<p>Amazing result!! are you using the NN for candidates generation / ranking / both? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2087055,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-05T09:24:23.580000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2062557,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-12T08:07:40.167000",
      "content": "<p>Cool! Did you model clicks/carts/orders sequence separately or put them together? BTW, do both the recall and ranking stages use the same architecture?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2056623": "We know that in the h&m recommendation competition, GBDT is the winning solution for most of the top teams. But in this competition,  we don't have any attribution info for users and products,  the action sequence is all we know.  Extract information from a pure sequence data, NN maybe more effective than GBDT!\n\nLet's share the model insight, idea, experiment result here,  start with me.\n\n---\n\nmodel:  covisitation\nvalidation on test:\n- cv 0.6302\n- lb 0.575\n\nmodel:  Transformer\nvalidation on test:\n- cv 0.6378\n- lb 0.578",
    "2062607": "How do you  compute validation on test?\n\nFor the co-visitation model with lb 0.575, this validation notebook from @cdeotte has a cv score of 0.5643.  But it does not use test it uses last week of train.\n\nI don't understand how you get 0.6302 for that model.",
    "2057555": "Would love to see a descriptive notebook on this!",
    "2056739": "Amazing! . I tried it before and gave up because it took too long to train. My idea is to take the type(clicks, carts, and orders) as type embedding like position embedding  and predict next even whether happen(0 or 1) so that this model both recall candidates and rerank for all types. ",
    "2056631": "That is a really cool result! Thanks for sharing 🙂\n\nI have been thinking if rnns, transformers, maybe some weird CNN,could play a role here. Turns out that it might 🙂 Very inspiring 🙂",
    "2056936": "Cool~ Given the setting of only ids, I also want to try a pure NN solution",
    "2056877": "Nice result. What does your NN predict? Does it take user sequence as input and predict next item id? Or does it take pairs of \"user-item\" and predict likely of click, cart, order, etc?",
    "2121027": "Could you please share your method once the competition ends? Thanks a lot!",
    "2061712": "Amazing result!! are you using the NN for candidates generation / ranking / both? ",
    "2087055": "",
    "2062557": "Cool! Did you model clicks/carts/orders sequence separately or put them together? BTW, do both the recall and ranking stages use the same architecture?"
  }
}