{
  "id": 371731,
  "title": "Very confused in understanding the candidates generation and feature-engineering",
  "url": "/competitions/otto-recommender-system/discussion/371731",
  "author_name": "",
  "post_date": "2022-12-12T00:30:14.003396Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello guys, thanks for the competition. I read some articles and I understood we need to use candidates to improve recall and construct user-item features to create the model. </p>\n<p>So, let's create an example to simplify my question.  Let <code>S = ['A', 'A', 'A', 'B', 'B', 'C', 'C',  'D' ]</code> a session where the letters are product ids. Also, let 'O = {B, C}' the purchased products from the same session  <code>S</code>. In this example 'A' was the first clicked in 'D` was the last. </p>\n<p>Ok, let's generate the candidates. I'm going to use two ideas from this <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070\" target=\"_blank\">article</a>:</p>\n<ul>\n<li>Repurchase products. (we will call <strong>GenerationIdea1</strong>)</li>\n<li>Item collaborative filtering. (we will call <strong>GenerationIdea2</strong>)</li>\n</ul>\n<p>So, from <strong>GenerationIdea1</strong> we will add to <code>Candidates</code> set the products <code>{B,C}</code> as they were previously purchased. <br>\nAnd, from  <strong>GenerationIdea2</strong> let's suppose the model outputted the products <code>{F}</code> (<em>yeah, I think the model can output a product that was not on the clicks, do you agree?</em>). Therefore, our final candidate set will be <code>Candidates = {B, C, F}</code>. </p>\n<p>Now we created the candidates, let's think about some user-item features. For simplification, let's just use:</p>\n<ul>\n<li>Session length</li>\n<li>Number of times the item was clicked. </li>\n<li>If it's the first item clicked</li>\n<li>if it's the last item clicked</li>\n</ul>\n<p>Finally, we will have a table like this:</p>\n<table>\n<thead>\n<tr>\n<th>ItemId</th>\n<th>session_length</th>\n<th>n_times_clicked</th>\n<th>is_first</th>\n<th>is_last</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B</td>\n<td>8</td>\n<td>2</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>C</td>\n<td>8</td>\n<td>2</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>F</td>\n<td>8</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>Ok, now I have some question:<br>\n1: Is my idea correct? <br>\n2: If my idea is correct, how the product item <code>F</code> could be useful for the model if it was not clicked? Probably the model will learn that <code>F</code> is not useful. <br>\n3: If we want to output 10 items, and our model can output just 3 items. How can we  predict the last 7 (10 - 3) items if we are input just 3 items (<code>F</code>, <code>B</code>, and <code>C</code> in this case)</p>\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "2062291",
      "postDate": "12/12/2022 00:30:14",
      "content": "<p>Hello guys, thanks for the competition. I read some articles and I understood we need to use candidates to improve recall and construct user-item features to create the model. </p>\n<p>So, let's create an example to simplify my question.  Let <code>S = ['A', 'A', 'A', 'B', 'B', 'C', 'C',  'D' ]</code> a session where the letters are product ids. Also, let 'O = {B, C}' the purchased products from the same session  <code>S</code>. In this example 'A' was the first clicked in 'D` was the last. </p>\n<p>Ok, let's generate the candidates. I'm going to use two ideas from this <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070\" target=\"_blank\">article</a>:</p>\n<ul>\n<li>Repurchase products. (we will call <strong>GenerationIdea1</strong>)</li>\n<li>Item collaborative filtering. (we will call <strong>GenerationIdea2</strong>)</li>\n</ul>\n<p>So, from <strong>GenerationIdea1</strong> we will add to <code>Candidates</code> set the products <code>{B,C}</code> as they were previously purchased. <br>\nAnd, from  <strong>GenerationIdea2</strong> let's suppose the model outputted the products <code>{F}</code> (<em>yeah, I think the model can output a product that was not on the clicks, do you agree?</em>). Therefore, our final candidate set will be <code>Candidates = {B, C, F}</code>. </p>\n<p>Now we created the candidates, let's think about some user-item features. For simplification, let's just use:</p>\n<ul>\n<li>Session length</li>\n<li>Number of times the item was clicked. </li>\n<li>If it's the first item clicked</li>\n<li>if it's the last item clicked</li>\n</ul>\n<p>Finally, we will have a table like this:</p>\n<table>\n<thead>\n<tr>\n<th>ItemId</th>\n<th>session_length</th>\n<th>n_times_clicked</th>\n<th>is_first</th>\n<th>is_last</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B</td>\n<td>8</td>\n<td>2</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>C</td>\n<td>8</td>\n<td>2</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>F</td>\n<td>8</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>Ok, now I have some question:<br>\n1: Is my idea correct? <br>\n2: If my idea is correct, how the product item <code>F</code> could be useful for the model if it was not clicked? Probably the model will learn that <code>F</code> is not useful. <br>\n3: If we want to output 10 items, and our model can output just 3 items. How can we  predict the last 7 (10 - 3) items if we are input just 3 items (<code>F</code>, <code>B</code>, and <code>C</code> in this case)</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Hello guys, thanks for the competition. I read some articles and I understood we need to use candidates to improve recall and construct user-item features to create the model. \n\nSo, let's create an example to simplify my question.  Let `S = ['A', 'A', 'A', 'B', 'B', 'C', 'C',  'D' ]` a session where the letters are product ids. Also, let 'O = {B, C}' the purchased products from the same session  `S`. In this example 'A' was the first clicked in 'D` was the last. \n\nOk, let's generate the candidates. I'm going to use two ideas from this [article](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070):\n- Repurchase products. (we will call **GenerationIdea1**)\n- Item collaborative filtering. (we will call **GenerationIdea2**)\n\nSo, from **GenerationIdea1** we will add to `Candidates` set the products `{B,C}` as they were previously purchased. \nAnd, from  **GenerationIdea2** let's suppose the model outputted the products `{F}` (*yeah, I think the model can output a product that was not on the clicks, do you agree?*). Therefore, our final candidate set will be `Candidates = {B, C, F}`. \n\nNow we created the candidates, let's think about some user-item features. For simplification, let's just use:\n- Session length\n- Number of times the item was clicked. \n- If it's the first item clicked\n- if it's the last item clicked\n\nFinally, we will have a table like this:\n\n| ItemId | session_length| n_times_clicked | is_first| is_last |\n| --- | --- |  --- | --- | --- |\n| B | 8 | 2 | 0 | 0 | \n| C | 8 | 2 | 0 | 0 |\n| F | 8 | 0 | 0 | 0 |\n\nOk, now I have some question:\n1: Is my idea correct? \n2: If my idea is correct, how the product item `F` could be useful for the model if it was not clicked? Probably the model will learn that `F` is not useful. \n3: If we want to output 10 items, and our model can output just 3 items. How can we  predict the last 7 (10 - 3) items if we are input just 3 items (`F`, `B`, and `C` in this case)\n\n\nThank you!",
      "votes": null
    },
    {
      "id": "2063123",
      "postDate": "12/12/2022 16:41:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kevintakano\" target=\"_blank\">@kevintakano</a> ! Yes, your general idea is correct. To summarize, we can split the procedure into a few steps:</p>\n<ol>\n<li>Candidate Generation: extract <code>K</code> items for each session <code>S</code>. In this phase, your goal is to restrict the search space to a limited number of relevant items (= <em>candidates</em>).</li>\n<li>Re-Ranking: use a re-ranking model to order the candidates from the most to the least relevant. Usually, this kind of model needs a set of features like a classical regression/classification problem, as you pointed out. </li>\n<li>Predictions: extract <code>N</code> recommendations from the re-ranking model.</li>\n</ol>\n<p>Obviously, <code>K</code>&gt;=<code>N</code>. To answer your questions:</p>\n<ol>\n<li>Yes, your idea is correct.</li>\n<li>The product <code>F</code> could be meaningful because we are not bound to recommend something inside the same session <code>S</code>, but instead the next item that we have to predict can be anything (even if <a href=\"https://www.kaggle.com/code/sbunzini/carts-and-orders-follow-clicks\" target=\"_blank\">this notebook</a> seems to show the opposite, especially for <em>carts/orders</em>)</li>\n<li>In general, you can consider multiple techniques to fill in the missing 7 items. For example, you can add the <em>most popular items</em> to your recommendations to reach the number of elements you need. This applies to both <em>candidate generation models</em> and the <em>re-ranker</em>.</li>\n</ol>\n<p>Hope I was clear enough😊</p>",
      "rawMarkdown": "Hi @kevintakano ! Yes, your general idea is correct. To summarize, we can split the procedure into a few steps:\n1. Candidate Generation: extract `K` items for each session `S`. In this phase, your goal is to restrict the search space to a limited number of relevant items (= *candidates*).\n2. Re-Ranking: use a re-ranking model to order the candidates from the most to the least relevant. Usually, this kind of model needs a set of features like a classical regression/classification problem, as you pointed out. \n3. Predictions: extract `N` recommendations from the re-ranking model.\n\nObviously, `K`>=`N`. To answer your questions:\n1. Yes, your idea is correct.\n2. The product `F` could be meaningful because we are not bound to recommend something inside the same session `S`, but instead the next item that we have to predict can be anything (even if [this notebook](https://www.kaggle.com/code/sbunzini/carts-and-orders-follow-clicks) seems to show the opposite, especially for *carts/orders*)\n3. In general, you can consider multiple techniques to fill in the missing 7 items. For example, you can add the *most popular items* to your recommendations to reach the number of elements you need. This applies to both *candidate generation models* and the *re-ranker*.\n\nHope I was clear enough😊",
      "votes": null
    },
    {
      "id": "2063144",
      "postDate": "12/12/2022 16:59:43",
      "content": "<p>Hi Kevin your idea is correct but you are missing the target column in your dataframe. The test data that we download from Kaggle is the first half of user activity and we need to predict the second half (i.e. the Kaggle leaderboard LB). Therefore your full session would be below and the last 5 activity is hidden from us.</p>\n<p>S = ['A', 'A', 'A', 'B', 'B', 'C', 'C', 'D'  || 'A', 'B', 'G', 'F', 'A']</p>\n<p>In total the user interacted with 13 items. Kaggle test data will include the first 8 and we need to predict the last 5. Therefore your data frame will include 3 columns of targets (using the ['A', 'B', 'G', 'F', 'A'])</p>\n<table>\n<thead>\n<tr>\n<th>ItemId</th>\n<th>session_length</th>\n<th>n_times_clicked</th>\n<th>is_first</th>\n<th>is_last</th>\n<th>click</th>\n<th>cart</th>\n<th>order</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B</td>\n<td>8</td>\n<td>2</td>\n<td>0</td>\n<td>0</td>\n<td>1</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>C</td>\n<td>8</td>\n<td>2</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>F</td>\n<td>8</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>The <code>X_train = df[:,1:5]</code> and <code>y_train = df[:,5:]</code> when you train your GBT ranker model</p>\n<p>In order to train a model, we need to split Kaggle's train data (with is 4 weeks) into 3 datasets. Read Radek's post <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">here</a> to learn how to split. The first 3 weeks are \"new train\". The first half of week 4 is \"new test\" and second half of week 4 is \"new LB (leaderboard)\". Then we train our model using the first two datasets as features and the third dataset as targets. Then we infer our model using Kaggle train, Kaggle test to predict Kaggle LB.</p>",
      "rawMarkdown": "Hi Kevin your idea is correct but you are missing the target column in your dataframe. The test data that we download from Kaggle is the first half of user activity and we need to predict the second half (i.e. the Kaggle leaderboard LB). Therefore your full session would be below and the last 5 activity is hidden from us.\n\nS = ['A', 'A', 'A', 'B', 'B', 'C', 'C', 'D'  || 'A', 'B', 'G', 'F', 'A']\n\nIn total the user interacted with 13 items. Kaggle test data will include the first 8 and we need to predict the last 5. Therefore your data frame will include 3 columns of targets (using the ['A', 'B', 'G', 'F', 'A'])\n\n| ItemId | session_length | n_times_clicked | is_first | is_last | click | cart | order |\n| --- | --- | --- | --- | --- | --- | --- | --- |\n| B | 8 | 2 | 0 | 0 | 1 | 0 | 0 |\n| C | 8 | 2 | 0 | 0 | 0 | 0 | 0 |\n| F | 8 | 0 | 0 | 0 | 0 | 0 | 1 |\n\nThe `X_train = df[:,1:5]` and `y_train = df[:,5:]` when you train your GBT ranker model\n\nIn order to train a model, we need to split Kaggle's train data (with is 4 weeks) into 3 datasets. Read Radek's post [here][1] to learn how to split. The first 3 weeks are \"new train\". The first half of week 4 is \"new test\" and second half of week 4 is \"new LB (leaderboard)\". Then we train our model using the first two datasets as features and the third dataset as targets. Then we infer our model using Kaggle train, Kaggle test to predict Kaggle LB.\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991",
      "votes": null
    },
    {
      "id": "2063473",
      "postDate": "12/13/2022 00:34:41",
      "content": "<p>If I understand correctly, we use the 3 datasets like follows:</p>\n<p>For Validation: use the 1st dataset as the feature and 2nd dataset as the target. Then we inference using the 1st 2 datasets as the feature and validation on the 3rd dataset to get the CV score.</p>\n<p>For Submission: use the 1st 2 datasets as the feature and the 3rd dataset as the target to train the model and inference using the kaggle test set.</p>\n<p>Correct me if I'm wrong. Thanks!<br>\n(I'm busy with work recently, didn't got time to start, just keep reading the discussions to keep updated)</p>",
      "rawMarkdown": "If I understand correctly, we use the 3 datasets like follows:\n\nFor Validation: use the 1st dataset as the feature and 2nd dataset as the target. Then we inference using the 1st 2 datasets as the feature and validation on the 3rd dataset to get the CV score.\n\nFor Submission: use the 1st 2 datasets as the feature and the 3rd dataset as the target to train the model and inference using the kaggle test set.\n\nCorrect me if I'm wrong. Thanks!\n(I'm busy with work recently, didn't got time to start, just keep reading the discussions to keep updated)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2063123,
      "author_name": "sbunzini",
      "author_url": "",
      "post_date": "12/12/2022 16:41:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kevintakano\" target=\"_blank\">@kevintakano</a> ! Yes, your general idea is correct. To summarize, we can split the procedure into a few steps:</p>\n<ol>\n<li>Candidate Generation: extract <code>K</code> items for each session <code>S</code>. In this phase, your goal is to restrict the search space to a limited number of relevant items (= <em>candidates</em>).</li>\n<li>Re-Ranking: use a re-ranking model to order the candidates from the most to the least relevant. Usually, this kind of model needs a set of features like a classical regression/classification problem, as you pointed out. </li>\n<li>Predictions: extract <code>N</code> recommendations from the re-ranking model.</li>\n</ol>\n<p>Obviously, <code>K</code>&gt;=<code>N</code>. To answer your questions:</p>\n<ol>\n<li>Yes, your idea is correct.</li>\n<li>The product <code>F</code> could be meaningful because we are not bound to recommend something inside the same session <code>S</code>, but instead the next item that we have to predict can be anything (even if <a href=\"https://www.kaggle.com/code/sbunzini/carts-and-orders-follow-clicks\" target=\"_blank\">this notebook</a> seems to show the opposite, especially for <em>carts/orders</em>)</li>\n<li>In general, you can consider multiple techniques to fill in the missing 7 items. For example, you can add the <em>most popular items</em> to your recommendations to reach the number of elements you need. This applies to both <em>candidate generation models</em> and the <em>re-ranker</em>.</li>\n</ol>\n<p>Hope I was clear enough😊</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2063144,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "12/12/2022 16:59:43",
      "content": "<p>Hi Kevin your idea is correct but you are missing the target column in your dataframe. The test data that we download from Kaggle is the first half of user activity and we need to predict the second half (i.e. the Kaggle leaderboard LB). Therefore your full session would be below and the last 5 activity is hidden from us.</p>\n<p>S = ['A', 'A', 'A', 'B', 'B', 'C', 'C', 'D'  || 'A', 'B', 'G', 'F', 'A']</p>\n<p>In total the user interacted with 13 items. Kaggle test data will include the first 8 and we need to predict the last 5. Therefore your data frame will include 3 columns of targets (using the ['A', 'B', 'G', 'F', 'A'])</p>\n<table>\n<thead>\n<tr>\n<th>ItemId</th>\n<th>session_length</th>\n<th>n_times_clicked</th>\n<th>is_first</th>\n<th>is_last</th>\n<th>click</th>\n<th>cart</th>\n<th>order</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B</td>\n<td>8</td>\n<td>2</td>\n<td>0</td>\n<td>0</td>\n<td>1</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>C</td>\n<td>8</td>\n<td>2</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>F</td>\n<td>8</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>The <code>X_train = df[:,1:5]</code> and <code>y_train = df[:,5:]</code> when you train your GBT ranker model</p>\n<p>In order to train a model, we need to split Kaggle's train data (with is 4 weeks) into 3 datasets. Read Radek's post <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">here</a> to learn how to split. The first 3 weeks are \"new train\". The first half of week 4 is \"new test\" and second half of week 4 is \"new LB (leaderboard)\". Then we train our model using the first two datasets as features and the third dataset as targets. Then we infer our model using Kaggle train, Kaggle test to predict Kaggle LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2063473,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "12/13/2022 00:34:41",
          "content": "<p>If I understand correctly, we use the 3 datasets like follows:</p>\n<p>For Validation: use the 1st dataset as the feature and 2nd dataset as the target. Then we inference using the 1st 2 datasets as the feature and validation on the 3rd dataset to get the CV score.</p>\n<p>For Submission: use the 1st 2 datasets as the feature and the 3rd dataset as the target to train the model and inference using the kaggle test set.</p>\n<p>Correct me if I'm wrong. Thanks!<br>\n(I'm busy with work recently, didn't got time to start, just keep reading the discussions to keep updated)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2062291": "Hello guys, thanks for the competition. I read some articles and I understood we need to use candidates to improve recall and construct user-item features to create the model. \n\nSo, let's create an example to simplify my question.  Let `S = ['A', 'A', 'A', 'B', 'B', 'C', 'C',  'D' ]` a session where the letters are product ids. Also, let 'O = {B, C}' the purchased products from the same session  `S`. In this example 'A' was the first clicked in 'D` was the last. \n\nOk, let's generate the candidates. I'm going to use two ideas from this [article](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070):\n- Repurchase products. (we will call **GenerationIdea1**)\n- Item collaborative filtering. (we will call **GenerationIdea2**)\n\nSo, from **GenerationIdea1** we will add to `Candidates` set the products `{B,C}` as they were previously purchased. \nAnd, from  **GenerationIdea2** let's suppose the model outputted the products `{F}` (*yeah, I think the model can output a product that was not on the clicks, do you agree?*). Therefore, our final candidate set will be `Candidates = {B, C, F}`. \n\nNow we created the candidates, let's think about some user-item features. For simplification, let's just use:\n- Session length\n- Number of times the item was clicked. \n- If it's the first item clicked\n- if it's the last item clicked\n\nFinally, we will have a table like this:\n\n| ItemId | session_length| n_times_clicked | is_first| is_last |\n| --- | --- |  --- | --- | --- |\n| B | 8 | 2 | 0 | 0 | \n| C | 8 | 2 | 0 | 0 |\n| F | 8 | 0 | 0 | 0 |\n\nOk, now I have some question:\n1: Is my idea correct? \n2: If my idea is correct, how the product item `F` could be useful for the model if it was not clicked? Probably the model will learn that `F` is not useful. \n3: If we want to output 10 items, and our model can output just 3 items. How can we  predict the last 7 (10 - 3) items if we are input just 3 items (`F`, `B`, and `C` in this case)\n\n\nThank you!",
    "2063123": "Hi @kevintakano ! Yes, your general idea is correct. To summarize, we can split the procedure into a few steps:\n1. Candidate Generation: extract `K` items for each session `S`. In this phase, your goal is to restrict the search space to a limited number of relevant items (= *candidates*).\n2. Re-Ranking: use a re-ranking model to order the candidates from the most to the least relevant. Usually, this kind of model needs a set of features like a classical regression/classification problem, as you pointed out. \n3. Predictions: extract `N` recommendations from the re-ranking model.\n\nObviously, `K`>=`N`. To answer your questions:\n1. Yes, your idea is correct.\n2. The product `F` could be meaningful because we are not bound to recommend something inside the same session `S`, but instead the next item that we have to predict can be anything (even if [this notebook](https://www.kaggle.com/code/sbunzini/carts-and-orders-follow-clicks) seems to show the opposite, especially for *carts/orders*)\n3. In general, you can consider multiple techniques to fill in the missing 7 items. For example, you can add the *most popular items* to your recommendations to reach the number of elements you need. This applies to both *candidate generation models* and the *re-ranker*.\n\nHope I was clear enough😊",
    "2063144": "Hi Kevin your idea is correct but you are missing the target column in your dataframe. The test data that we download from Kaggle is the first half of user activity and we need to predict the second half (i.e. the Kaggle leaderboard LB). Therefore your full session would be below and the last 5 activity is hidden from us.\n\nS = ['A', 'A', 'A', 'B', 'B', 'C', 'C', 'D'  || 'A', 'B', 'G', 'F', 'A']\n\nIn total the user interacted with 13 items. Kaggle test data will include the first 8 and we need to predict the last 5. Therefore your data frame will include 3 columns of targets (using the ['A', 'B', 'G', 'F', 'A'])\n\n| ItemId | session_length | n_times_clicked | is_first | is_last | click | cart | order |\n| --- | --- | --- | --- | --- | --- | --- | --- |\n| B | 8 | 2 | 0 | 0 | 1 | 0 | 0 |\n| C | 8 | 2 | 0 | 0 | 0 | 0 | 0 |\n| F | 8 | 0 | 0 | 0 | 0 | 0 | 1 |\n\nThe `X_train = df[:,1:5]` and `y_train = df[:,5:]` when you train your GBT ranker model\n\nIn order to train a model, we need to split Kaggle's train data (with is 4 weeks) into 3 datasets. Read Radek's post [here][1] to learn how to split. The first 3 weeks are \"new train\". The first half of week 4 is \"new test\" and second half of week 4 is \"new LB (leaderboard)\". Then we train our model using the first two datasets as features and the third dataset as targets. Then we infer our model using Kaggle train, Kaggle test to predict Kaggle LB.\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991",
    "2063473": "If I understand correctly, we use the 3 datasets like follows:\n\nFor Validation: use the 1st dataset as the feature and 2nd dataset as the target. Then we inference using the 1st 2 datasets as the feature and validation on the 3rd dataset to get the CV score.\n\nFor Submission: use the 1st 2 datasets as the feature and the 3rd dataset as the target to train the model and inference using the kaggle test set.\n\nCorrect me if I'm wrong. Thanks!\n(I'm busy with work recently, didn't got time to start, just keep reading the discussions to keep updated)"
  },
  "source": "meta"
}