{
  "id": 382955,
  "title": "Idea about the \"Cross Target Stacking\" Method",
  "url": "/competitions/otto-recommender-system/discussion/382955",
  "author_name": "",
  "post_date": "2023-02-01T17:31:43.236467100Z",
  "votes": 22,
  "comment_count": 7,
  "views": 0,
  "content": "<h2>About this post</h2>\n<p>Our team's 34th solution and overview are summarized at the following link:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382812\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/382812</a></li>\n</ul>\n<p>In this thread, I'll explain about my stacking method which boosted both our local CV and LB score.<br>\nI'm not sure this method is commonly used or well known so I named this method \"Cross Target Stacking\" <br>\nbecause this is like a stacking using the different type of the target labels.</p>\n<h2>Ordinary Candidate-Rerank model</h2>\n<p>The following figure shows that the ordinary training/prediction strategies.<br>\nI think a lot of participants apply these kind of strategies.<br>\nI also try this strategies at first.</p>\n<p>At the training phase, the GBDT framework (like the LightGBM) is used for training and three independent models are generated for order/click/cart, respectively.</p>\n<p>The common candidates/features are used for the training but the target labels are separated for each type (order/cart/click).</p>\n<p>At the prediction phase, each models are used for the inference and we can generate the three independent score.<br>\nFinally, we can get the final submission results by concatenating there score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4624515%2Fed57783c4103c84f9c54148604c2e970%2F2023-02-02%202.15.20.png?generation=1675271741794866&amp;alt=media\" alt=\"\"></p>\n<h2>My Question</h2>\n<p>During the competition, I was wondering how to treat three different target labels.<br>\nIn the ordinary method, three different type of the labels (order/click/cart) are independent.<br>\nIn other words, we can generate \"order\" prediction score without click/cart target labels and we only need the <br>\n\"order\" target label and features when we train/predict \"order\" labels.</p>\n<p>However, I guessed click/carts target information might be made use of the order prediction.<br>\nFor the same reason, click/order might be made use of the \"cart\" prediction and cart/order might be made use of the \"click\" prediction as well.</p>\n<p>In order to check this hypothesis, I tried to add the cart/click target labels for \"order\" training on purpose.<br>\nObviously this features cause  leakage but I wanted to know cart/click target information are important to predict \"order\" or not.<br>\nBy adding there \"leakage\" features, my local CV score was drastically improved.<br>\nThis means that cart/click target label information is important to predict \"order\".</p>\n<h2>Cross Target Stacking Method</h2>\n<p>Of course, we can't use target label itself as a feature, so I tried to use the prediction score of cross target.<br>\nIn my definition, \"cross target\" represent the different type of target.<br>\nFor example, the cross target of \"order\" means \"cart/clicks\".<br>\nThis method is similar to the \"stacking\" method.<br>\nWhereas the ordinary stacking method uses the \"same\" type of prediction score, my stacking method uses the<br>\ndifferent type of target.<br>\nThat's why I named this method \"cross target stacking\".</p>\n<p>The following figure shows the overview of this method for training.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4624515%2Fdf4f95f5eae6d29d98cab90092ff7c39%2F2023-02-02%202.14.31.png?generation=1675271696594580&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4624515%2F275b803de2abd2a0d9d7cf372aa223a8%2F2023-02-02%202.17.03.png?generation=1675271844836461&amp;alt=media\" alt=\"\"></p>\n<p>There are two training/prediction phase.<br>\nThe first phase is similar to the ordinary training/prediction strategies.</p>\n<p>After the first training, we can get the OOF score for each type by using the cross validation.<br>\n(In my case number of fold=5)</p>\n<p>These prediction score can be added for the second training as features.<br>\nIn order to avoid overfitting, I only add cross target prediction score.</p>\n<p>For example, I added cart/click prediction scores to the \"order\" training but didn't add the \"order\"<br>\nprediction score itself.</p>\n<p>By using this procedures, our local CV score and LB are improved especially for the \"order\" training.</p>\n<h2>At the end</h2>\n<p>In fact, these strategies could boosted our score but I'm not sure this method is popular or not.<br>\nI'm also not sure this kind of method is adopted or tried by the Top-Kaggler in this competition.</p>\n<p>I would be happy if I could discuss this theme in this thread.</p>\n<p>At the end, I appreciate the collaboration with my team members and this competition supported by the organizers.<br>\nComments are welcome:)</p>",
  "messages": [
    {
      "id": "2125517",
      "postDate": "02/01/2023 17:31:43",
      "content": "<h2>About this post</h2>\n<p>Our team's 34th solution and overview are summarized at the following link:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382812\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/382812</a></li>\n</ul>\n<p>In this thread, I'll explain about my stacking method which boosted both our local CV and LB score.<br>\nI'm not sure this method is commonly used or well known so I named this method \"Cross Target Stacking\" <br>\nbecause this is like a stacking using the different type of the target labels.</p>\n<h2>Ordinary Candidate-Rerank model</h2>\n<p>The following figure shows that the ordinary training/prediction strategies.<br>\nI think a lot of participants apply these kind of strategies.<br>\nI also try this strategies at first.</p>\n<p>At the training phase, the GBDT framework (like the LightGBM) is used for training and three independent models are generated for order/click/cart, respectively.</p>\n<p>The common candidates/features are used for the training but the target labels are separated for each type (order/cart/click).</p>\n<p>At the prediction phase, each models are used for the inference and we can generate the three independent score.<br>\nFinally, we can get the final submission results by concatenating there score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4624515%2Fed57783c4103c84f9c54148604c2e970%2F2023-02-02%202.15.20.png?generation=1675271741794866&amp;alt=media\" alt=\"\"></p>\n<h2>My Question</h2>\n<p>During the competition, I was wondering how to treat three different target labels.<br>\nIn the ordinary method, three different type of the labels (order/click/cart) are independent.<br>\nIn other words, we can generate \"order\" prediction score without click/cart target labels and we only need the <br>\n\"order\" target label and features when we train/predict \"order\" labels.</p>\n<p>However, I guessed click/carts target information might be made use of the order prediction.<br>\nFor the same reason, click/order might be made use of the \"cart\" prediction and cart/order might be made use of the \"click\" prediction as well.</p>\n<p>In order to check this hypothesis, I tried to add the cart/click target labels for \"order\" training on purpose.<br>\nObviously this features cause  leakage but I wanted to know cart/click target information are important to predict \"order\" or not.<br>\nBy adding there \"leakage\" features, my local CV score was drastically improved.<br>\nThis means that cart/click target label information is important to predict \"order\".</p>\n<h2>Cross Target Stacking Method</h2>\n<p>Of course, we can't use target label itself as a feature, so I tried to use the prediction score of cross target.<br>\nIn my definition, \"cross target\" represent the different type of target.<br>\nFor example, the cross target of \"order\" means \"cart/clicks\".<br>\nThis method is similar to the \"stacking\" method.<br>\nWhereas the ordinary stacking method uses the \"same\" type of prediction score, my stacking method uses the<br>\ndifferent type of target.<br>\nThat's why I named this method \"cross target stacking\".</p>\n<p>The following figure shows the overview of this method for training.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4624515%2Fdf4f95f5eae6d29d98cab90092ff7c39%2F2023-02-02%202.14.31.png?generation=1675271696594580&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4624515%2F275b803de2abd2a0d9d7cf372aa223a8%2F2023-02-02%202.17.03.png?generation=1675271844836461&amp;alt=media\" alt=\"\"></p>\n<p>There are two training/prediction phase.<br>\nThe first phase is similar to the ordinary training/prediction strategies.</p>\n<p>After the first training, we can get the OOF score for each type by using the cross validation.<br>\n(In my case number of fold=5)</p>\n<p>These prediction score can be added for the second training as features.<br>\nIn order to avoid overfitting, I only add cross target prediction score.</p>\n<p>For example, I added cart/click prediction scores to the \"order\" training but didn't add the \"order\"<br>\nprediction score itself.</p>\n<p>By using this procedures, our local CV score and LB are improved especially for the \"order\" training.</p>\n<h2>At the end</h2>\n<p>In fact, these strategies could boosted our score but I'm not sure this method is popular or not.<br>\nI'm also not sure this kind of method is adopted or tried by the Top-Kaggler in this competition.</p>\n<p>I would be happy if I could discuss this theme in this thread.</p>\n<p>At the end, I appreciate the collaboration with my team members and this competition supported by the organizers.<br>\nComments are welcome:)</p>",
      "rawMarkdown": "## About this post\nOur team's 34th solution and overview are summarized at the following link:\n- https://www.kaggle.com/competitions/otto-recommender-system/discussion/382812\n\nIn this thread, I'll explain about my stacking method which boosted both our local CV and LB score.\nI'm not sure this method is commonly used or well known so I named this method \"Cross Target Stacking\" \nbecause this is like a stacking using the different type of the target labels.\n\n## Ordinary Candidate-Rerank model\n\nThe following figure shows that the ordinary training/prediction strategies.\nI think a lot of participants apply these kind of strategies.\nI also try this strategies at first.\n\nAt the training phase, the GBDT framework (like the LightGBM) is used for training and three independent models are generated for order/click/cart, respectively.\n\nThe common candidates/features are used for the training but the target labels are separated for each type (order/cart/click).\n\nAt the prediction phase, each models are used for the inference and we can generate the three independent score.\nFinally, we can get the final submission results by concatenating there score.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4624515%2Fed57783c4103c84f9c54148604c2e970%2F2023-02-02%202.15.20.png?generation=1675271741794866&alt=media)\n\n## My Question\n\nDuring the competition, I was wondering how to treat three different target labels.\nIn the ordinary method, three different type of the labels (order/click/cart) are independent.\nIn other words, we can generate \"order\" prediction score without click/cart target labels and we only need the \n\"order\" target label and features when we train/predict \"order\" labels.\n\nHowever, I guessed click/carts target information might be made use of the order prediction.\nFor the same reason, click/order might be made use of the \"cart\" prediction and cart/order might be made use of the \"click\" prediction as well.\n\nIn order to check this hypothesis, I tried to add the cart/click target labels for \"order\" training on purpose.\nObviously this features cause  leakage but I wanted to know cart/click target information are important to predict \"order\" or not.\nBy adding there \"leakage\" features, my local CV score was drastically improved.\nThis means that cart/click target label information is important to predict \"order\".\n\n## Cross Target Stacking Method\n\nOf course, we can't use target label itself as a feature, so I tried to use the prediction score of cross target.\nIn my definition, \"cross target\" represent the different type of target.\nFor example, the cross target of \"order\" means \"cart/clicks\".\nThis method is similar to the \"stacking\" method.\nWhereas the ordinary stacking method uses the \"same\" type of prediction score, my stacking method uses the\ndifferent type of target.\nThat's why I named this method \"cross target stacking\".\n\nThe following figure shows the overview of this method for training.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4624515%2Fdf4f95f5eae6d29d98cab90092ff7c39%2F2023-02-02%202.14.31.png?generation=1675271696594580&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4624515%2F275b803de2abd2a0d9d7cf372aa223a8%2F2023-02-02%202.17.03.png?generation=1675271844836461&alt=media)\n\nThere are two training/prediction phase.\nThe first phase is similar to the ordinary training/prediction strategies.\n\nAfter the first training, we can get the OOF score for each type by using the cross validation.\n(In my case number of fold=5)\n\nThese prediction score can be added for the second training as features.\nIn order to avoid overfitting, I only add cross target prediction score.\n\nFor example, I added cart/click prediction scores to the \"order\" training but didn't add the \"order\"\nprediction score itself.\n\nBy using this procedures, our local CV score and LB are improved especially for the \"order\" training.\n\n## At the end\nIn fact, these strategies could boosted our score but I'm not sure this method is popular or not.\nI'm also not sure this kind of method is adopted or tried by the Top-Kaggler in this competition.\n\nI would be happy if I could discuss this theme in this thread.\n\nAt the end, I appreciate the collaboration with my team members and this competition supported by the organizers.\nComments are welcome:)",
      "votes": null
    },
    {
      "id": "2125572",
      "postDate": "02/01/2023 18:09:31",
      "content": "<p>I used the same approach for \"orders\" model and it gave me ~0.0015.</p>",
      "rawMarkdown": "I used the same approach for \"orders\" model and it gave me ~0.0015.",
      "votes": null
    },
    {
      "id": "2125641",
      "postDate": "02/01/2023 19:02:03",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/tetsuro731\" target=\"_blank\">@tetsuro731</a> , I'm happy you shared  a detailed post of this approach. <br>\nI think I saw something similar in the RecSys Challenge 2021 (<a href=\"https://youtu.be/8ms6ZBXDx3E?t=5205\" target=\"_blank\">video presentation of recsys 2021 top teams</a>, the link should bring to the time 1:26:45 where a similar approach is shown, it is not in 2 phases like yours).</p>\n<p>This approach can bring many advantages. That year after the end of RecSys Challenge 2021 I tried training a classifier that predicted the absence of any kind of interaction and using it as an additional feature and saw a 0.5% improvement on the validation score, I would be really interested in knowing if this approach could also improve your solution in this challenge.</p>\n<p>My intuitive understanding of the usefulness of this approach is that the prediction of a model trained on another target type contains in itself manipulations of the features useful for that specific target, those manipulations may also be useful for other targets and so adding the score of the model as a feature can make easier to find those aspects that may be harder to find only seeing the current target type.</p>\n<p>The use of prediction from a model that predicts absence of interactions helps to more easily ignore candidates that are not interesting for any interaction type making their score lower and so less likely to end up in the top 20.</p>\n<p>Lastly, Congratulations for your position in the challenge!</p>",
      "rawMarkdown": "Hello @tetsuro731 , I'm happy you shared  a detailed post of this approach. \nI think I saw something similar in the RecSys Challenge 2021 ([video presentation of recsys 2021 top teams](https://youtu.be/8ms6ZBXDx3E?t=5205), the link should bring to the time 1:26:45 where a similar approach is shown, it is not in 2 phases like yours).\n\nThis approach can bring many advantages. That year after the end of RecSys Challenge 2021 I tried training a classifier that predicted the absence of any kind of interaction and using it as an additional feature and saw a 0.5% improvement on the validation score, I would be really interested in knowing if this approach could also improve your solution in this challenge.\n\nMy intuitive understanding of the usefulness of this approach is that the prediction of a model trained on another target type contains in itself manipulations of the features useful for that specific target, those manipulations may also be useful for other targets and so adding the score of the model as a feature can make easier to find those aspects that may be harder to find only seeing the current target type.\n\nThe use of prediction from a model that predicts absence of interactions helps to more easily ignore candidates that are not interesting for any interaction type making their score lower and so less likely to end up in the top 20.\n\nLastly, Congratulations for your position in the challenge!",
      "votes": null
    },
    {
      "id": "2125902",
      "postDate": "02/02/2023 01:33:11",
      "content": "<p>Thanks for sharing. I used the same idea here. In local cv, it boosted by almost 0.001. But didn't work on lb in my case.. confused..</p>",
      "rawMarkdown": "Thanks for sharing. I used the same idea here. In local cv, it boosted by almost 0.001. But didn't work on lb in my case.. confused..",
      "votes": null
    },
    {
      "id": "2127672",
      "postDate": "02/03/2023 03:28:21",
      "content": "<p>Thank you for your comment.<br>\nThat's consistent with my situation.<br>\nI applied to this method to \"order\" which could boost my lb from 0.593 to 0.594.</p>",
      "rawMarkdown": "Thank you for your comment.\nThat's consistent with my situation.\nI applied to this method to \"order\" which could boost my lb from 0.593 to 0.594.",
      "votes": null
    },
    {
      "id": "2127675",
      "postDate": "02/03/2023 03:31:51",
      "content": "<p>Thank you for your comment.<br>\nIn my case, these methods can boost our local CV.<br>\nHowever, as for LB, \"order\" stacking was the most effective to boost our LB but the other type stacking couldn't boost our LB effectively.<br>\nI think the effect of this method is depend on the \"type\".</p>",
      "rawMarkdown": "Thank you for your comment.\nIn my case, these methods can boost our local CV.\nHowever, as for LB, \"order\" stacking was the most effective to boost our LB but the other type stacking couldn't boost our LB effectively.\nI think the effect of this method is depend on the \"type\".",
      "votes": null
    },
    {
      "id": "2127676",
      "postDate": "02/03/2023 03:37:30",
      "content": "<p>Thank you for your informative comment/suggestion!<br>\nI think that's so interesting research.<br>\nIn this competition, these method can boost our LB ~0.001 which is not a large effect.<br>\nHowever, I think this kind of method is interesting and can be make use of the other situation as you said.</p>\n<p>I also congratulate you for your silver medal!</p>",
      "rawMarkdown": "Thank you for your informative comment/suggestion!\nI think that's so interesting research.\nIn this competition, these method can boost our LB ~0.001 which is not a large effect.\nHowever, I think this kind of method is interesting and can be make use of the other situation as you said.\n\nI also congratulate you for your silver medal!",
      "votes": null
    },
    {
      "id": "2127871",
      "postDate": "02/03/2023 09:33:23",
      "content": "<p>Stacking features generated in this way helped a bit also my model (not more 0.001 aniway).</p>\n<p>It is a sort of surrogate for a multioutput rank/classifier.</p>\n<p>If NN had been competitive to rank candidate with this dataset, we would certainly have used a schema like this to train:</p>\n<pre><code>X:session, candidate \ny:(candidate_is_a_session_order,candidate_is_a_session_cart,candidate_is_session_a_click)\n</code></pre>",
      "rawMarkdown": "Stacking features generated in this way helped a bit also my model (not more 0.001 aniway).\n\nIt is a sort of surrogate for a multioutput rank/classifier.\n\nIf NN had been competitive to rank candidate with this dataset, we would certainly have used a schema like this to train:\n```\nX:session, candidate \ny:(candidate_is_a_session_order,candidate_is_a_session_cart,candidate_is_session_a_click)\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2125572,
      "author_name": "sparakhin",
      "author_url": "",
      "post_date": "02/01/2023 18:09:31",
      "content": "<p>I used the same approach for \"orders\" model and it gave me ~0.0015.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2127672,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "02/03/2023 03:28:21",
          "content": "<p>Thank you for your comment.<br>\nThat's consistent with my situation.<br>\nI applied to this method to \"order\" which could boost my lb from 0.593 to 0.594.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2125641,
      "author_name": "pietromaldini1",
      "author_url": "",
      "post_date": "02/01/2023 19:02:03",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/tetsuro731\" target=\"_blank\">@tetsuro731</a> , I'm happy you shared  a detailed post of this approach. <br>\nI think I saw something similar in the RecSys Challenge 2021 (<a href=\"https://youtu.be/8ms6ZBXDx3E?t=5205\" target=\"_blank\">video presentation of recsys 2021 top teams</a>, the link should bring to the time 1:26:45 where a similar approach is shown, it is not in 2 phases like yours).</p>\n<p>This approach can bring many advantages. That year after the end of RecSys Challenge 2021 I tried training a classifier that predicted the absence of any kind of interaction and using it as an additional feature and saw a 0.5% improvement on the validation score, I would be really interested in knowing if this approach could also improve your solution in this challenge.</p>\n<p>My intuitive understanding of the usefulness of this approach is that the prediction of a model trained on another target type contains in itself manipulations of the features useful for that specific target, those manipulations may also be useful for other targets and so adding the score of the model as a feature can make easier to find those aspects that may be harder to find only seeing the current target type.</p>\n<p>The use of prediction from a model that predicts absence of interactions helps to more easily ignore candidates that are not interesting for any interaction type making their score lower and so less likely to end up in the top 20.</p>\n<p>Lastly, Congratulations for your position in the challenge!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2127676,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "02/03/2023 03:37:30",
          "content": "<p>Thank you for your informative comment/suggestion!<br>\nI think that's so interesting research.<br>\nIn this competition, these method can boost our LB ~0.001 which is not a large effect.<br>\nHowever, I think this kind of method is interesting and can be make use of the other situation as you said.</p>\n<p>I also congratulate you for your silver medal!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2125902,
      "author_name": "hookman",
      "author_url": "",
      "post_date": "02/02/2023 01:33:11",
      "content": "<p>Thanks for sharing. I used the same idea here. In local cv, it boosted by almost 0.001. But didn't work on lb in my case.. confused..</p>",
      "votes": null,
      "replies": [
        {
          "id": 2127675,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "02/03/2023 03:31:51",
          "content": "<p>Thank you for your comment.<br>\nIn my case, these methods can boost our local CV.<br>\nHowever, as for LB, \"order\" stacking was the most effective to boost our LB but the other type stacking couldn't boost our LB effectively.<br>\nI think the effect of this method is depend on the \"type\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2127871,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "02/03/2023 09:33:23",
      "content": "<p>Stacking features generated in this way helped a bit also my model (not more 0.001 aniway).</p>\n<p>It is a sort of surrogate for a multioutput rank/classifier.</p>\n<p>If NN had been competitive to rank candidate with this dataset, we would certainly have used a schema like this to train:</p>\n<pre><code>X:session, candidate \ny:(candidate_is_a_session_order,candidate_is_a_session_cart,candidate_is_session_a_click)\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2125517": "## About this post\nOur team's 34th solution and overview are summarized at the following link:\n- https://www.kaggle.com/competitions/otto-recommender-system/discussion/382812\n\nIn this thread, I'll explain about my stacking method which boosted both our local CV and LB score.\nI'm not sure this method is commonly used or well known so I named this method \"Cross Target Stacking\" \nbecause this is like a stacking using the different type of the target labels.\n\n## Ordinary Candidate-Rerank model\n\nThe following figure shows that the ordinary training/prediction strategies.\nI think a lot of participants apply these kind of strategies.\nI also try this strategies at first.\n\nAt the training phase, the GBDT framework (like the LightGBM) is used for training and three independent models are generated for order/click/cart, respectively.\n\nThe common candidates/features are used for the training but the target labels are separated for each type (order/cart/click).\n\nAt the prediction phase, each models are used for the inference and we can generate the three independent score.\nFinally, we can get the final submission results by concatenating there score.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4624515%2Fed57783c4103c84f9c54148604c2e970%2F2023-02-02%202.15.20.png?generation=1675271741794866&alt=media)\n\n## My Question\n\nDuring the competition, I was wondering how to treat three different target labels.\nIn the ordinary method, three different type of the labels (order/click/cart) are independent.\nIn other words, we can generate \"order\" prediction score without click/cart target labels and we only need the \n\"order\" target label and features when we train/predict \"order\" labels.\n\nHowever, I guessed click/carts target information might be made use of the order prediction.\nFor the same reason, click/order might be made use of the \"cart\" prediction and cart/order might be made use of the \"click\" prediction as well.\n\nIn order to check this hypothesis, I tried to add the cart/click target labels for \"order\" training on purpose.\nObviously this features cause  leakage but I wanted to know cart/click target information are important to predict \"order\" or not.\nBy adding there \"leakage\" features, my local CV score was drastically improved.\nThis means that cart/click target label information is important to predict \"order\".\n\n## Cross Target Stacking Method\n\nOf course, we can't use target label itself as a feature, so I tried to use the prediction score of cross target.\nIn my definition, \"cross target\" represent the different type of target.\nFor example, the cross target of \"order\" means \"cart/clicks\".\nThis method is similar to the \"stacking\" method.\nWhereas the ordinary stacking method uses the \"same\" type of prediction score, my stacking method uses the\ndifferent type of target.\nThat's why I named this method \"cross target stacking\".\n\nThe following figure shows the overview of this method for training.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4624515%2Fdf4f95f5eae6d29d98cab90092ff7c39%2F2023-02-02%202.14.31.png?generation=1675271696594580&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4624515%2F275b803de2abd2a0d9d7cf372aa223a8%2F2023-02-02%202.17.03.png?generation=1675271844836461&alt=media)\n\nThere are two training/prediction phase.\nThe first phase is similar to the ordinary training/prediction strategies.\n\nAfter the first training, we can get the OOF score for each type by using the cross validation.\n(In my case number of fold=5)\n\nThese prediction score can be added for the second training as features.\nIn order to avoid overfitting, I only add cross target prediction score.\n\nFor example, I added cart/click prediction scores to the \"order\" training but didn't add the \"order\"\nprediction score itself.\n\nBy using this procedures, our local CV score and LB are improved especially for the \"order\" training.\n\n## At the end\nIn fact, these strategies could boosted our score but I'm not sure this method is popular or not.\nI'm also not sure this kind of method is adopted or tried by the Top-Kaggler in this competition.\n\nI would be happy if I could discuss this theme in this thread.\n\nAt the end, I appreciate the collaboration with my team members and this competition supported by the organizers.\nComments are welcome:)",
    "2125572": "I used the same approach for \"orders\" model and it gave me ~0.0015.",
    "2125641": "Hello @tetsuro731 , I'm happy you shared  a detailed post of this approach. \nI think I saw something similar in the RecSys Challenge 2021 ([video presentation of recsys 2021 top teams](https://youtu.be/8ms6ZBXDx3E?t=5205), the link should bring to the time 1:26:45 where a similar approach is shown, it is not in 2 phases like yours).\n\nThis approach can bring many advantages. That year after the end of RecSys Challenge 2021 I tried training a classifier that predicted the absence of any kind of interaction and using it as an additional feature and saw a 0.5% improvement on the validation score, I would be really interested in knowing if this approach could also improve your solution in this challenge.\n\nMy intuitive understanding of the usefulness of this approach is that the prediction of a model trained on another target type contains in itself manipulations of the features useful for that specific target, those manipulations may also be useful for other targets and so adding the score of the model as a feature can make easier to find those aspects that may be harder to find only seeing the current target type.\n\nThe use of prediction from a model that predicts absence of interactions helps to more easily ignore candidates that are not interesting for any interaction type making their score lower and so less likely to end up in the top 20.\n\nLastly, Congratulations for your position in the challenge!",
    "2125902": "Thanks for sharing. I used the same idea here. In local cv, it boosted by almost 0.001. But didn't work on lb in my case.. confused..",
    "2127672": "Thank you for your comment.\nThat's consistent with my situation.\nI applied to this method to \"order\" which could boost my lb from 0.593 to 0.594.",
    "2127675": "Thank you for your comment.\nIn my case, these methods can boost our local CV.\nHowever, as for LB, \"order\" stacking was the most effective to boost our LB but the other type stacking couldn't boost our LB effectively.\nI think the effect of this method is depend on the \"type\".",
    "2127676": "Thank you for your informative comment/suggestion!\nI think that's so interesting research.\nIn this competition, these method can boost our LB ~0.001 which is not a large effect.\nHowever, I think this kind of method is interesting and can be make use of the other situation as you said.\n\nI also congratulate you for your silver medal!",
    "2127871": "Stacking features generated in this way helped a bit also my model (not more 0.001 aniway).\n\nIt is a sort of surrogate for a multioutput rank/classifier.\n\nIf NN had been competitive to rank candidate with this dataset, we would certainly have used a schema like this to train:\n```\nX:session, candidate \ny:(candidate_is_a_session_order,candidate_is_a_session_cart,candidate_is_session_a_click)\n```"
  },
  "source": "meta"
}