{
  "id": 370210,
  "title": "How To Build a GBT Ranker Model",
  "url": "/competitions/otto-recommender-system/discussion/370210",
  "author_name": "Chris Deotte",
  "post_date": "2022-12-03T16:41:41.353000",
  "votes": 335,
  "comment_count": 224,
  "views": 0,
  "content": "<h1>Candidate Rerank Model</h1>\n<p>A strong and easy way to build a solution for this competition is a \"candidate rerank\" model explained <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">here</a>. Creating train data, training and inferring a GBT ranker model can be difficult to understand at first. I hope the following discussion post will help you on your journey.</p>\n<p>The general idea is; for each <code>session</code> (i.e. user), we find 50-200 candidates (that are likely to be correct predictions) then we train an GBT ranker model to select our final 20. To do this, we need to create train data for our ranker model. Below are some tips to help create train data and use GBT ranker models</p>\n<h1>Candidate DataFrame</h1>\n<p>The first step to creating train data for our ranker model is building a dataframe of candidates. One way to generate high quality candidates is using co-visitiation matrices, explained <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369\" target=\"_blank\">here</a>. Our candidate dataframe has one pair of <code>session</code> (i.e. user) and <code>aid</code> (i.e. item) per row. This dataframe has the following columns</p>\n<ul>\n<li>session (i.e. user)</li>\n<li>aid (i.e. item)</li>\n<li>user features</li>\n<li>item features</li>\n<li>user-item interaction features</li>\n<li>click target (i.e 0 or 1)</li>\n<li>cart target (i.e. 0 or 1)</li>\n<li>order target (i.e. 0 or 1)</li>\n</ul>\n<h1>Memory Management</h1>\n<p>In all of the following code, it is very important to continually change the <code>dtype</code> of each column to minimum data size. So after making a new column, make sure to <code>df[col] = df[col].astype('int32')</code> or the smallest possible dtype. If you do nothing then many libraries will use <code>int64</code> or <code>float64</code> which takes 8 bytes per value. This will use more RAM and disk space then needed. <strong>Always reduce dtypes!</strong></p>\n<p><strong>NOTE</strong> When using the latest version of cuDF locally (version 22.08 or later) , there is a new feature to set default bitwidth as 32. This will prevent the dataframe from using <code>int64</code> nor <code>float64</code> and makes it easier for the user. (Thanks Giba for pointing this out <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/371058\" target=\"_blank\">here</a>):</p>\n<pre><code>import cudf\ncudf.set_option(\"default_integer_bitwidth\", 32)\ncudf.set_option(\"default_float_bitwidth\", 32)\n</code></pre>\n<p>Even after reducing column dtypes, you may still have memory issues. The way to solve this is to process everything in chunks. If you are using <code>train.groupby('session')</code> to make user features, then perhaps first split the train data into X pieces (like 2, 4, 8 pieces). Then process user features for each piece separately and save to disk as <code>f'user_features_p{PIECE_NUMBER}.pqt'</code>. Later when you read them in, you can concatenate them together.</p>\n<p>Also consider using DASK which works with both CPU Pandas or GPU RAPIDS cuDF to use multiple CPU/GPU and/or disk when needed to avoid all memory errors! <a href=\"https://www.dask.org/\" target=\"_blank\">(info here)</a></p>\n<h1>Speed - Use GPU</h1>\n<p>All of the following code will take time to run since this data has around 13 million users and 2 million items. I suggest using an accelerated dataframe library for all the following processing like Nvidia's RAPIDS cuDF <a href=\"https://rapids.ai/\" target=\"_blank\">here</a> which uses GPU instead of CPU for accelerated speed !</p>\n<h1>Step 1</h1>\n<p>Our train data will be the first 3 weeks of Kaggle train data. And our validation data will be the last 1 week of Kaggle train. We then split our validation data into <strong>validation data A</strong> and <strong>validation data B</strong> where validation B are the ground truths. (Validation data A is like the Kaggle test data we download whereas validation data B is like the LB (leaderboard). Roughly speaking <code>A</code> is like first half of each user activity during week while <code>B</code> is second half).</p>\n<p>Consider using Radek's train and valid data <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">here</a> which already splits validation data into A and B. For every user (i.e. session) in our <strong>validation data A</strong> (i.e. last 1 week of Kaggle train), we generate X candidate aids. Let's say <code>X=50</code> for the rest of our discussion. We create a dataframe of shape <code>(</code>number_of_session x 50<code>, 2 )</code>:</p>\n<table>\n<thead>\n<tr>\n<th>user</th>\n<th>item</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0001</td>\n<td>6456</td>\n</tr>\n<tr>\n<td>0001</td>\n<td>4490</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>8486</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>7297</td>\n</tr>\n</tbody>\n</table>\n<p>Each session will appear 50 times. And there will be no duplicate <code>[session,aid]</code> pairs. We only use users from our validation data because we only have targets for users in validation data (in step 6 below).</p>\n<h1>Step 2</h1>\n<p>Create item features. Using our <strong>train data + valid data A</strong> (yes use test leak), we create item features in their own dataframe and save a parquet to disk. For example</p>\n<pre><code>item_features = train.groupby('aid').agg({'aid':'count','session':'nunique','type':'mean'})\nitem_features.columns = ['item_item_count','item_user_count','item_buy_ratio']\n# CONVERT COLUMNS TO INT32 and FLOAT32 HERE\nitem_features.to_parquet('item_features.pqt')\n</code></pre>\n<p><strong>NOTE</strong> if you have memory problems. Then break <code>train</code> into 10 dataframe parts. All rows pertained to a single item must be in the same dataframe part so that <code>train_part_1.groupby('aid')</code> will work correctly. After processing, save each part separately to disk. And later when you read them from disk, concatenate them together before merging them to candidate dataframe.</p>\n<h1>Step 3</h1>\n<p>Create user features. Using our <strong>validation data A</strong>, we create user features in their own dataframe and save a parquet to disk. For example</p>\n<pre><code>user_features = train.groupby('session').agg({'session':'count','aid':'nunique','type':'mean'})\nuser_features.columns = ['user_user_count','user_item_count','user_buy_ratio']\n# CONVERT COLUMNS TO INT32 and FLOAT32 HERE\nuser_features.to_parquet('user_features.pqt')\n</code></pre>\n<h1>Step 4</h1>\n<p>This step is optional. Step 4 will improve CV and LB, but your GBT ranker will work without step 4. Create user-item interaction features. Using our <strong>validation data A</strong>, we create <strong>multiple</strong> user-item feature dataframes and save them as parquets to disk. For each idea, we can make a new dataframe. One dataframe can contain all items that a user clicks. So make a dataframe with one column user, one column item, and a third column called <code>item_clicked</code>. Then for each unique item that a user clicked, we add a new row with <code>item_clicked = 1</code>. Note that our dataframe will have <strong>no duplicate rows</strong> of <code>['user','item']</code> pairs. Save this dataframe to disk. When we merge this to candidate dataframe we will <code>fillna(0)</code> to indicate the items not clicked.</p>\n<h1>Step 5</h1>\n<p>Add features to our candidate dataframe. To add features to our candidate dataframe, we read from disk and merge them on as follows</p>\n<pre><code>item_features = pd.read_parquet('item_features.pqt')\ncandidates = candidates.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\nuser_features = pd.read_parquet('user_features.pqt')\ncandidates = candidates.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\n</code></pre>\n<p>Now our candidate dataframe looks like</p>\n<table>\n<thead>\n<tr>\n<th>user</th>\n<th>item</th>\n<th>item_feat1</th>\n<th>item_feat2</th>\n<th>user_feat1</th>\n<th>user_feat2</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0001</td>\n<td>6456</td>\n<td>10</td>\n<td>12</td>\n<td>5.4</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>0001</td>\n<td>4490</td>\n<td>13</td>\n<td>15</td>\n<td>5.4</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>8486</td>\n<td>55</td>\n<td>8</td>\n<td>2</td>\n<td>1.2</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>7297</td>\n<td>70</td>\n<td>10</td>\n<td>2</td>\n<td>1.2</td>\n</tr>\n</tbody>\n</table>\n<p><strong>NOTE</strong> if we have memory problems merging all our features to our candidate dataframe, then we can do this in chunks</p>\n<pre><code>CHUNKS = 10\nchunk_size = np.ceil( len(candidates) / CHUNKS)\nfor k in range(CHUNKS):\n    df = candidates.iloc[k*chunk_size:(k+1)*chunk_size].copy()\n    df = df.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\n    df = df.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\n    df.to_parquet(f'candidate_with_features_p{k}.pqt')\n</code></pre>\n<h1>Step 6</h1>\n<p>We will now add targets to our candidate dataframe from step 1. The best way to add a column of targets is to use dataframe merge. First we make a dataframe of all the <code>target=1</code> as follows. Starting with a dataframe that contains the targets as a column of lists ( like Radek's ground truth labels <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation?select=test_labels.parquet\" target=\"_blank\">here</a>) such as:</p>\n<p><code>test_labels.parquet</code></p>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>type</th>\n<th>ground truth</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0001</td>\n<td>carts</td>\n<td>[3456, 4490, 5661, 7821, 9914 ]</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>carts</td>\n<td>[1222, 4656, 533, 8486]</td>\n</tr>\n</tbody>\n</table>\n<p>We use the following code to convert these lists into a dataframe of targets:</p>\n<pre><code>tar = pd.read_parquet('test_labels.parquet')\ntar = tar.loc[ tar['type']=='carts' ]\naids = tar.ground_truth.explode().astype('int32').rename('item')\ntar = tar[['session']].astype('int32').rename({'session':'user'},axis=1)\ntar = tar.merge(aids, left_index=True, right_index=True, how='left')\ntar['cart'] = 1\n</code></pre>\n<p>This produces a dataframe like</p>\n<table>\n<thead>\n<tr>\n<th>user</th>\n<th>item</th>\n<th>cart</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0001</td>\n<td>3456</td>\n<td>1</td>\n</tr>\n<tr>\n<td>0001</td>\n<td>4490</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>And we merge it to our candidate dataframe with the following line:</p>\n<pre><code>candidates = candidates.merge(cart_target,on=['user','item'],how='left').fillna(0)\n</code></pre>\n<table>\n<thead>\n<tr>\n<th>user</th>\n<th>item</th>\n<th>item_feat1</th>\n<th>item_feat2</th>\n<th>user_feat1</th>\n<th>user_feat2</th>\n<th>cart</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0001</td>\n<td>6456</td>\n<td>10</td>\n<td>12</td>\n<td>3</td>\n<td>0.5</td>\n<td>0</td>\n</tr>\n<tr>\n<td>0001</td>\n<td>4490</td>\n<td>13</td>\n<td>5</td>\n<td>5.4</td>\n<td>0.1</td>\n<td>1</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>8486</td>\n<td>55</td>\n<td>10</td>\n<td>5</td>\n<td>0.9</td>\n<td>1</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>7297</td>\n<td>70</td>\n<td>20</td>\n<td>2</td>\n<td>1.2</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<h1>Training</h1>\n<p>We now have train data for our GBT ranker model. We must train using <code>GroupKFold</code>. <strong>Important Note</strong>: when we train, we do not use the <code>user</code> and <code>item</code> columns as features, we only use the other columns. <code>FEATURES = candidates.columns[2 : -1*len(targets)]</code>. Note with XGB, we have 3 options for rankers by changing <code>objective</code> parameter, to either <code>rank:pairwise</code>, or <code>rank:ndcg</code>,  or <code>rank:map</code>. (We should also add and tune XGB parameters <code>'max_depth', 'subsample', 'colsample_bytree', 'learning_rate'</code>)</p>\n<pre><code>import xgboost as xgb\nfrom sklearn.model_selection import GroupKFold\n\nskf = GroupKFold(n_splits=5)\nfor fold,(train_idx, valid_idx) in enumerate(skf.split(candidates, candidates['click'], groups=candidates['user'] )):\n\n    X_train = candidates.loc[train_idx, FEATURES]\n    y_train = candidates.loc[train_idx, 'click']\n    X_valid = candidates.loc[valid_idx, FEATURES]\n    y_valid = candidates.loc[valid_idx, 'click']\n\n    # IF YOU HAVE 50 CANDIDATE WE USE 50 BELOW\n    dtrain = xgb.DMatrix(X_train, y_train, group=[50] * (len(train_idx)//50) ) \n    dvalid = xgb.DMatrix(X_valid, y_valid, group=[50] * (len(valid_idx)//50) ) \n\n    xgb_parms = {'objective':'rank:pairwise', 'tree_method':'gpu_hist'}\n    model = xgb.train(xgb_parms, \n        dtrain=dtrain,\n        evals=[(dtrain,'train'),(dvalid,'valid')],\n        num_boost_round=1000,\n        verbose_eval=100)\n    model.save_model(f'XGB_fold{fold}_click.xgb')\n</code></pre>\n<p><strong>NOTE</strong> If you have memory problems training XGB on GPU, consider downsampling negatives 2x, 4x, 10x, 20x with <code>frac = 0.5, 0.25, 0.1, or 0.05</code> (and then update group sizes in DMatrix). Or use DASK XGB with multiple GPUs. Here is example code:</p>\n<pre><code>positives = candidates.loc[candidates['click']==1]\nnegatives = candidates.loc[candidates['click']==0].sample(frac=0.5)\ncandidates = pd.concat([positives,negatives],axis=0,ignore_index=True)\n</code></pre>\n<h1>Inference</h1>\n<p>For inference, we create a new candidate dataframe (using our technique to generate candidates before) but this time from Kaggle's test data. Then we make item features from all 4 weeks of Kaggle train plus 1 week of Kaggle test. And we make user features from Kaggle test. We merge the features to our candidates. Then we use our saved models to infer predictions for clicks. Lastly we select 20 by sorting the predictions and choosing 20 with.</p>\n<pre><code>preds = np.zeros(len(test_candidates))\nfor fold in range(5):\n    model = xgb.Booster()\n    model.load_model(f'XGB_fold{fold}_click.xgb')\n    model.set_param({'predictor': 'gpu_predictor'})\n    dtest = xgb.DMatrix(data=test_candidates[FEATURES])\n    preds += model.predict(dtest)/5\npredictions = test_candidates[['user','item']].copy()\npredictions['pred'] = preds\n\npredictions = predictions.sort_values(['user','pred'], ascending=[True,False]).reset_index(drop=True)\npredictions['n'] = predictions.groupby('user').item.cumcount().astype('int8')\npredictions = predictions.loc[predictions.n&lt;20]\nsub = predictions.groupby('user').item.apply(list)\nsub = sub.to_frame().reset_index()\nsub.item = sub.item.apply(lambda x: \" \".join(map(str,x)))\nsub.columns = ['session_type','labels']\nsub.session_type = sub.session_type.astype('str')+ '_clicks'\n</code></pre>\n<p><strong>NOTE</strong> if you have memory errors. Consider loading 1/10th of the test data. Then merge features. Then infer. Next load the next 1/10th, merge features, infer. etc. etc. Lastly concatenate the predictions and make submission.csv</p>\n<h1>Enjoy</h1>\n<p>I hope this discussion post helps illustrate how to create data for a GBT ranking model and how to train and infer. Have fun!</p>",
  "messages": [
    {
      "id": 2053832,
      "postDate": "2022-12-03T16:41:41.353Z",
      "content": "<h1>Candidate Rerank Model</h1>\n<p>A strong and easy way to build a solution for this competition is a \"candidate rerank\" model explained <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">here</a>. Creating train data, training and inferring a GBT ranker model can be difficult to understand at first. I hope the following discussion post will help you on your journey.</p>\n<p>The general idea is; for each <code>session</code> (i.e. user), we find 50-200 candidates (that are likely to be correct predictions) then we train an GBT ranker model to select our final 20. To do this, we need to create train data for our ranker model. Below are some tips to help create train data and use GBT ranker models</p>\n<h1>Candidate DataFrame</h1>\n<p>The first step to creating train data for our ranker model is building a dataframe of candidates. One way to generate high quality candidates is using co-visitiation matrices, explained <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369\" target=\"_blank\">here</a>. Our candidate dataframe has one pair of <code>session</code> (i.e. user) and <code>aid</code> (i.e. item) per row. This dataframe has the following columns</p>\n<ul>\n<li>session (i.e. user)</li>\n<li>aid (i.e. item)</li>\n<li>user features</li>\n<li>item features</li>\n<li>user-item interaction features</li>\n<li>click target (i.e 0 or 1)</li>\n<li>cart target (i.e. 0 or 1)</li>\n<li>order target (i.e. 0 or 1)</li>\n</ul>\n<h1>Memory Management</h1>\n<p>In all of the following code, it is very important to continually change the <code>dtype</code> of each column to minimum data size. So after making a new column, make sure to <code>df[col] = df[col].astype('int32')</code> or the smallest possible dtype. If you do nothing then many libraries will use <code>int64</code> or <code>float64</code> which takes 8 bytes per value. This will use more RAM and disk space then needed. <strong>Always reduce dtypes!</strong></p>\n<p><strong>NOTE</strong> When using the latest version of cuDF locally (version 22.08 or later) , there is a new feature to set default bitwidth as 32. This will prevent the dataframe from using <code>int64</code> nor <code>float64</code> and makes it easier for the user. (Thanks Giba for pointing this out <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/371058\" target=\"_blank\">here</a>):</p>\n<pre><code>import cudf\ncudf.set_option(\"default_integer_bitwidth\", 32)\ncudf.set_option(\"default_float_bitwidth\", 32)\n</code></pre>\n<p>Even after reducing column dtypes, you may still have memory issues. The way to solve this is to process everything in chunks. If you are using <code>train.groupby('session')</code> to make user features, then perhaps first split the train data into X pieces (like 2, 4, 8 pieces). Then process user features for each piece separately and save to disk as <code>f'user_features_p{PIECE_NUMBER}.pqt'</code>. Later when you read them in, you can concatenate them together.</p>\n<p>Also consider using DASK which works with both CPU Pandas or GPU RAPIDS cuDF to use multiple CPU/GPU and/or disk when needed to avoid all memory errors! <a href=\"https://www.dask.org/\" target=\"_blank\">(info here)</a></p>\n<h1>Speed - Use GPU</h1>\n<p>All of the following code will take time to run since this data has around 13 million users and 2 million items. I suggest using an accelerated dataframe library for all the following processing like Nvidia's RAPIDS cuDF <a href=\"https://rapids.ai/\" target=\"_blank\">here</a> which uses GPU instead of CPU for accelerated speed !</p>\n<h1>Step 1</h1>\n<p>Our train data will be the first 3 weeks of Kaggle train data. And our validation data will be the last 1 week of Kaggle train. We then split our validation data into <strong>validation data A</strong> and <strong>validation data B</strong> where validation B are the ground truths. (Validation data A is like the Kaggle test data we download whereas validation data B is like the LB (leaderboard). Roughly speaking <code>A</code> is like first half of each user activity during week while <code>B</code> is second half).</p>\n<p>Consider using Radek's train and valid data <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">here</a> which already splits validation data into A and B. For every user (i.e. session) in our <strong>validation data A</strong> (i.e. last 1 week of Kaggle train), we generate X candidate aids. Let's say <code>X=50</code> for the rest of our discussion. We create a dataframe of shape <code>(</code>number_of_session x 50<code>, 2 )</code>:</p>\n<table>\n<thead>\n<tr>\n<th>user</th>\n<th>item</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0001</td>\n<td>6456</td>\n</tr>\n<tr>\n<td>0001</td>\n<td>4490</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>8486</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>7297</td>\n</tr>\n</tbody>\n</table>\n<p>Each session will appear 50 times. And there will be no duplicate <code>[session,aid]</code> pairs. We only use users from our validation data because we only have targets for users in validation data (in step 6 below).</p>\n<h1>Step 2</h1>\n<p>Create item features. Using our <strong>train data + valid data A</strong> (yes use test leak), we create item features in their own dataframe and save a parquet to disk. For example</p>\n<pre><code>item_features = train.groupby('aid').agg({'aid':'count','session':'nunique','type':'mean'})\nitem_features.columns = ['item_item_count','item_user_count','item_buy_ratio']\n# CONVERT COLUMNS TO INT32 and FLOAT32 HERE\nitem_features.to_parquet('item_features.pqt')\n</code></pre>\n<p><strong>NOTE</strong> if you have memory problems. Then break <code>train</code> into 10 dataframe parts. All rows pertained to a single item must be in the same dataframe part so that <code>train_part_1.groupby('aid')</code> will work correctly. After processing, save each part separately to disk. And later when you read them from disk, concatenate them together before merging them to candidate dataframe.</p>\n<h1>Step 3</h1>\n<p>Create user features. Using our <strong>validation data A</strong>, we create user features in their own dataframe and save a parquet to disk. For example</p>\n<pre><code>user_features = train.groupby('session').agg({'session':'count','aid':'nunique','type':'mean'})\nuser_features.columns = ['user_user_count','user_item_count','user_buy_ratio']\n# CONVERT COLUMNS TO INT32 and FLOAT32 HERE\nuser_features.to_parquet('user_features.pqt')\n</code></pre>\n<h1>Step 4</h1>\n<p>This step is optional. Step 4 will improve CV and LB, but your GBT ranker will work without step 4. Create user-item interaction features. Using our <strong>validation data A</strong>, we create <strong>multiple</strong> user-item feature dataframes and save them as parquets to disk. For each idea, we can make a new dataframe. One dataframe can contain all items that a user clicks. So make a dataframe with one column user, one column item, and a third column called <code>item_clicked</code>. Then for each unique item that a user clicked, we add a new row with <code>item_clicked = 1</code>. Note that our dataframe will have <strong>no duplicate rows</strong> of <code>['user','item']</code> pairs. Save this dataframe to disk. When we merge this to candidate dataframe we will <code>fillna(0)</code> to indicate the items not clicked.</p>\n<h1>Step 5</h1>\n<p>Add features to our candidate dataframe. To add features to our candidate dataframe, we read from disk and merge them on as follows</p>\n<pre><code>item_features = pd.read_parquet('item_features.pqt')\ncandidates = candidates.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\nuser_features = pd.read_parquet('user_features.pqt')\ncandidates = candidates.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\n</code></pre>\n<p>Now our candidate dataframe looks like</p>\n<table>\n<thead>\n<tr>\n<th>user</th>\n<th>item</th>\n<th>item_feat1</th>\n<th>item_feat2</th>\n<th>user_feat1</th>\n<th>user_feat2</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0001</td>\n<td>6456</td>\n<td>10</td>\n<td>12</td>\n<td>5.4</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>0001</td>\n<td>4490</td>\n<td>13</td>\n<td>15</td>\n<td>5.4</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>8486</td>\n<td>55</td>\n<td>8</td>\n<td>2</td>\n<td>1.2</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>7297</td>\n<td>70</td>\n<td>10</td>\n<td>2</td>\n<td>1.2</td>\n</tr>\n</tbody>\n</table>\n<p><strong>NOTE</strong> if we have memory problems merging all our features to our candidate dataframe, then we can do this in chunks</p>\n<pre><code>CHUNKS = 10\nchunk_size = np.ceil( len(candidates) / CHUNKS)\nfor k in range(CHUNKS):\n    df = candidates.iloc[k*chunk_size:(k+1)*chunk_size].copy()\n    df = df.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\n    df = df.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\n    df.to_parquet(f'candidate_with_features_p{k}.pqt')\n</code></pre>\n<h1>Step 6</h1>\n<p>We will now add targets to our candidate dataframe from step 1. The best way to add a column of targets is to use dataframe merge. First we make a dataframe of all the <code>target=1</code> as follows. Starting with a dataframe that contains the targets as a column of lists ( like Radek's ground truth labels <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation?select=test_labels.parquet\" target=\"_blank\">here</a>) such as:</p>\n<p><code>test_labels.parquet</code></p>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>type</th>\n<th>ground truth</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0001</td>\n<td>carts</td>\n<td>[3456, 4490, 5661, 7821, 9914 ]</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>carts</td>\n<td>[1222, 4656, 533, 8486]</td>\n</tr>\n</tbody>\n</table>\n<p>We use the following code to convert these lists into a dataframe of targets:</p>\n<pre><code>tar = pd.read_parquet('test_labels.parquet')\ntar = tar.loc[ tar['type']=='carts' ]\naids = tar.ground_truth.explode().astype('int32').rename('item')\ntar = tar[['session']].astype('int32').rename({'session':'user'},axis=1)\ntar = tar.merge(aids, left_index=True, right_index=True, how='left')\ntar['cart'] = 1\n</code></pre>\n<p>This produces a dataframe like</p>\n<table>\n<thead>\n<tr>\n<th>user</th>\n<th>item</th>\n<th>cart</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0001</td>\n<td>3456</td>\n<td>1</td>\n</tr>\n<tr>\n<td>0001</td>\n<td>4490</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>And we merge it to our candidate dataframe with the following line:</p>\n<pre><code>candidates = candidates.merge(cart_target,on=['user','item'],how='left').fillna(0)\n</code></pre>\n<table>\n<thead>\n<tr>\n<th>user</th>\n<th>item</th>\n<th>item_feat1</th>\n<th>item_feat2</th>\n<th>user_feat1</th>\n<th>user_feat2</th>\n<th>cart</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0001</td>\n<td>6456</td>\n<td>10</td>\n<td>12</td>\n<td>3</td>\n<td>0.5</td>\n<td>0</td>\n</tr>\n<tr>\n<td>0001</td>\n<td>4490</td>\n<td>13</td>\n<td>5</td>\n<td>5.4</td>\n<td>0.1</td>\n<td>1</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>8486</td>\n<td>55</td>\n<td>10</td>\n<td>5</td>\n<td>0.9</td>\n<td>1</td>\n</tr>\n<tr>\n<td>0002</td>\n<td>7297</td>\n<td>70</td>\n<td>20</td>\n<td>2</td>\n<td>1.2</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<h1>Training</h1>\n<p>We now have train data for our GBT ranker model. We must train using <code>GroupKFold</code>. <strong>Important Note</strong>: when we train, we do not use the <code>user</code> and <code>item</code> columns as features, we only use the other columns. <code>FEATURES = candidates.columns[2 : -1*len(targets)]</code>. Note with XGB, we have 3 options for rankers by changing <code>objective</code> parameter, to either <code>rank:pairwise</code>, or <code>rank:ndcg</code>,  or <code>rank:map</code>. (We should also add and tune XGB parameters <code>'max_depth', 'subsample', 'colsample_bytree', 'learning_rate'</code>)</p>\n<pre><code>import xgboost as xgb\nfrom sklearn.model_selection import GroupKFold\n\nskf = GroupKFold(n_splits=5)\nfor fold,(train_idx, valid_idx) in enumerate(skf.split(candidates, candidates['click'], groups=candidates['user'] )):\n\n    X_train = candidates.loc[train_idx, FEATURES]\n    y_train = candidates.loc[train_idx, 'click']\n    X_valid = candidates.loc[valid_idx, FEATURES]\n    y_valid = candidates.loc[valid_idx, 'click']\n\n    # IF YOU HAVE 50 CANDIDATE WE USE 50 BELOW\n    dtrain = xgb.DMatrix(X_train, y_train, group=[50] * (len(train_idx)//50) ) \n    dvalid = xgb.DMatrix(X_valid, y_valid, group=[50] * (len(valid_idx)//50) ) \n\n    xgb_parms = {'objective':'rank:pairwise', 'tree_method':'gpu_hist'}\n    model = xgb.train(xgb_parms, \n        dtrain=dtrain,\n        evals=[(dtrain,'train'),(dvalid,'valid')],\n        num_boost_round=1000,\n        verbose_eval=100)\n    model.save_model(f'XGB_fold{fold}_click.xgb')\n</code></pre>\n<p><strong>NOTE</strong> If you have memory problems training XGB on GPU, consider downsampling negatives 2x, 4x, 10x, 20x with <code>frac = 0.5, 0.25, 0.1, or 0.05</code> (and then update group sizes in DMatrix). Or use DASK XGB with multiple GPUs. Here is example code:</p>\n<pre><code>positives = candidates.loc[candidates['click']==1]\nnegatives = candidates.loc[candidates['click']==0].sample(frac=0.5)\ncandidates = pd.concat([positives,negatives],axis=0,ignore_index=True)\n</code></pre>\n<h1>Inference</h1>\n<p>For inference, we create a new candidate dataframe (using our technique to generate candidates before) but this time from Kaggle's test data. Then we make item features from all 4 weeks of Kaggle train plus 1 week of Kaggle test. And we make user features from Kaggle test. We merge the features to our candidates. Then we use our saved models to infer predictions for clicks. Lastly we select 20 by sorting the predictions and choosing 20 with.</p>\n<pre><code>preds = np.zeros(len(test_candidates))\nfor fold in range(5):\n    model = xgb.Booster()\n    model.load_model(f'XGB_fold{fold}_click.xgb')\n    model.set_param({'predictor': 'gpu_predictor'})\n    dtest = xgb.DMatrix(data=test_candidates[FEATURES])\n    preds += model.predict(dtest)/5\npredictions = test_candidates[['user','item']].copy()\npredictions['pred'] = preds\n\npredictions = predictions.sort_values(['user','pred'], ascending=[True,False]).reset_index(drop=True)\npredictions['n'] = predictions.groupby('user').item.cumcount().astype('int8')\npredictions = predictions.loc[predictions.n&lt;20]\nsub = predictions.groupby('user').item.apply(list)\nsub = sub.to_frame().reset_index()\nsub.item = sub.item.apply(lambda x: \" \".join(map(str,x)))\nsub.columns = ['session_type','labels']\nsub.session_type = sub.session_type.astype('str')+ '_clicks'\n</code></pre>\n<p><strong>NOTE</strong> if you have memory errors. Consider loading 1/10th of the test data. Then merge features. Then infer. Next load the next 1/10th, merge features, infer. etc. etc. Lastly concatenate the predictions and make submission.csv</p>\n<h1>Enjoy</h1>\n<p>I hope this discussion post helps illustrate how to create data for a GBT ranking model and how to train and infer. Have fun!</p>",
      "rawMarkdown": "# Candidate Rerank Model\nA strong and easy way to build a solution for this competition is a \"candidate rerank\" model explained [here][3]. Creating train data, training and inferring a GBT ranker model can be difficult to understand at first. I hope the following discussion post will help you on your journey.\n\nThe general idea is; for each `session` (i.e. user), we find 50-200 candidates (that are likely to be correct predictions) then we train an GBT ranker model to select our final 20. To do this, we need to create train data for our ranker model. Below are some tips to help create train data and use GBT ranker models\n\n# Candidate DataFrame\nThe first step to creating train data for our ranker model is building a dataframe of candidates. One way to generate high quality candidates is using co-visitiation matrices, explained [here][4]. Our candidate dataframe has one pair of `session` (i.e. user) and `aid` (i.e. item) per row. This dataframe has the following columns\n* session (i.e. user)\n* aid (i.e. item)\n* user features\n* item features\n* user-item interaction features\n* click target (i.e 0 or 1)\n* cart target (i.e. 0 or 1)\n* order target (i.e. 0 or 1)\n\n# Memory Management\nIn all of the following code, it is very important to continually change the `dtype` of each column to minimum data size. So after making a new column, make sure to `df[col] = df[col].astype('int32')` or the smallest possible dtype. If you do nothing then many libraries will use `int64` or `float64` which takes 8 bytes per value. This will use more RAM and disk space then needed. **Always reduce dtypes!**\n\n**NOTE** When using the latest version of cuDF locally (version 22.08 or later) , there is a new feature to set default bitwidth as 32. This will prevent the dataframe from using `int64` nor `float64` and makes it easier for the user. (Thanks Giba for pointing this out [here][7]):\n\n    import cudf\n    cudf.set_option(\"default_integer_bitwidth\", 32)\n    cudf.set_option(\"default_float_bitwidth\", 32)\n\nEven after reducing column dtypes, you may still have memory issues. The way to solve this is to process everything in chunks. If you are using `train.groupby('session')` to make user features, then perhaps first split the train data into X pieces (like 2, 4, 8 pieces). Then process user features for each piece separately and save to disk as `f'user_features_p{PIECE_NUMBER}.pqt'`. Later when you read them in, you can concatenate them together.\n\nAlso consider using DASK which works with both CPU Pandas or GPU RAPIDS cuDF to use multiple CPU/GPU and/or disk when needed to avoid all memory errors! [(info here)][6]\n\n# Speed - Use GPU\nAll of the following code will take time to run since this data has around 13 million users and 2 million items. I suggest using an accelerated dataframe library for all the following processing like Nvidia's RAPIDS cuDF [here][2] which uses GPU instead of CPU for accelerated speed !\n\n# Step 1\nOur train data will be the first 3 weeks of Kaggle train data. And our validation data will be the last 1 week of Kaggle train. We then split our validation data into **validation data A** and **validation data B** where validation B are the ground truths. (Validation data A is like the Kaggle test data we download whereas validation data B is like the LB (leaderboard). Roughly speaking `A` is like first half of each user activity during week while `B` is second half).\n\nConsider using Radek's train and valid data [here][5] which already splits validation data into A and B. For every user (i.e. session) in our **validation data A** (i.e. last 1 week of Kaggle train), we generate X candidate aids. Let's say `X=50` for the rest of our discussion. We create a dataframe of shape `( `number_of_session x 50`, 2 )`:\n\n| user | item | \n| --- | --- | \n| 0001 | 6456 | \n| 0001 | 4490 | \n| 0002 | 8486 | \n| 0002 | 7297 | \n\nEach session will appear 50 times. And there will be no duplicate `[session,aid]` pairs. We only use users from our validation data because we only have targets for users in validation data (in step 6 below).\n\n# Step 2\nCreate item features. Using our **train data + valid data A** (yes use test leak), we create item features in their own dataframe and save a parquet to disk. For example\n\n    item_features = train.groupby('aid').agg({'aid':'count','session':'nunique','type':'mean'})\n    item_features.columns = ['item_item_count','item_user_count','item_buy_ratio']\n    # CONVERT COLUMNS TO INT32 and FLOAT32 HERE\n    item_features.to_parquet('item_features.pqt')\n\n**NOTE** if you have memory problems. Then break `train` into 10 dataframe parts. All rows pertained to a single item must be in the same dataframe part so that `train_part_1.groupby('aid')` will work correctly. After processing, save each part separately to disk. And later when you read them from disk, concatenate them together before merging them to candidate dataframe.\n\n# Step 3\nCreate user features. Using our **validation data A**, we create user features in their own dataframe and save a parquet to disk. For example\n\n    user_features = train.groupby('session').agg({'session':'count','aid':'nunique','type':'mean'})\n    user_features.columns = ['user_user_count','user_item_count','user_buy_ratio']\n    # CONVERT COLUMNS TO INT32 and FLOAT32 HERE\n    user_features.to_parquet('user_features.pqt')\n\n# Step 4\nThis step is optional. Step 4 will improve CV and LB, but your GBT ranker will work without step 4. Create user-item interaction features. Using our **validation data A**, we create **multiple** user-item feature dataframes and save them as parquets to disk. For each idea, we can make a new dataframe. One dataframe can contain all items that a user clicks. So make a dataframe with one column user, one column item, and a third column called `item_clicked`. Then for each unique item that a user clicked, we add a new row with `item_clicked = 1`. Note that our dataframe will have **no duplicate rows** of `['user','item']` pairs. Save this dataframe to disk. When we merge this to candidate dataframe we will `fillna(0)` to indicate the items not clicked.\n\n# Step 5\nAdd features to our candidate dataframe. To add features to our candidate dataframe, we read from disk and merge them on as follows\n\n    item_features = pd.read_parquet('item_features.pqt')\n    candidates = candidates.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\n    user_features = pd.read_parquet('user_features.pqt')\n    candidates = candidates.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\n\nNow our candidate dataframe looks like\n\n| user | item | item_feat1 | item_feat2 | user_feat1 | user_feat2 |\n| --- | --- | --- | --- | --- | --- | \n| 0001 | 6456 | 10 | 12 | 5.4 | 0.5 |\n| 0001 | 4490 | 13 | 15 | 5.4 | 0.5 |\n| 0002 | 8486 | 55 | 8 | 2 | 1.2 |\n| 0002 | 7297 | 70 | 10 | 2 | 1.2 |\n\n**NOTE** if we have memory problems merging all our features to our candidate dataframe, then we can do this in chunks\n\n    CHUNKS = 10\n    chunk_size = np.ceil( len(candidates) / CHUNKS)\n    for k in range(CHUNKS):\n        df = candidates.iloc[k*chunk_size:(k+1)*chunk_size].copy()\n        df = df.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\n        df = df.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\n        df.to_parquet(f'candidate_with_features_p{k}.pqt')\n\n# Step 6\nWe will now add targets to our candidate dataframe from step 1. The best way to add a column of targets is to use dataframe merge. First we make a dataframe of all the `target=1` as follows. Starting with a dataframe that contains the targets as a column of lists ( like Radek's ground truth labels [here][1]) such as:\n\n`test_labels.parquet`\n| session | type | ground truth |\n | --- | --- | --- |\n| 0001 | carts | [3456, 4490, 5661, 7821, 9914 ] |\n| 0002 | carts | [1222, 4656, 533, 8486] |\n\nWe use the following code to convert these lists into a dataframe of targets:\n\n    tar = pd.read_parquet('test_labels.parquet')\n    tar = tar.loc[ tar['type']=='carts' ]\n    aids = tar.ground_truth.explode().astype('int32').rename('item')\n    tar = tar[['session']].astype('int32').rename({'session':'user'},axis=1)\n    tar = tar.merge(aids, left_index=True, right_index=True, how='left')\n    tar['cart'] = 1\n\nThis produces a dataframe like\n\n| user | item | cart | \n| --- | --- | --- | \n| 0001 | 3456 | 1 | \n| 0001 | 4490 | 1 | \n\nAnd we merge it to our candidate dataframe with the following line:\n\n    candidates = candidates.merge(cart_target,on=['user','item'],how='left').fillna(0)\n\n| user | item | item_feat1 | item_feat2 | user_feat1 | user_feat2 | cart |\n| --- | --- | --- | --- | --- | --- | --- |\n| 0001 | 6456 | 10 | 12 | 3 | 0.5 | 0 |\n| 0001 | 4490 | 13 | 5 | 5.4 | 0.1 | 1 |\n| 0002 | 8486 | 55 | 10 | 5 | 0.9 | 1 |\n| 0002 | 7297 | 70 | 20 | 2 | 1.2 | 0 |\n\n# Training\nWe now have train data for our GBT ranker model. We must train using `GroupKFold`. **Important Note**: when we train, we do not use the `user` and `item` columns as features, we only use the other columns. `FEATURES = candidates.columns[2 : -1*len(targets)]`. Note with XGB, we have 3 options for rankers by changing `objective` parameter, to either `rank:pairwise`, or `rank:ndcg`,  or `rank:map`. (We should also add and tune XGB parameters `'max_depth', 'subsample', 'colsample_bytree', 'learning_rate'`)\n\n    import xgboost as xgb\n    from sklearn.model_selection import GroupKFold\n\n    skf = GroupKFold(n_splits=5)\n    for fold,(train_idx, valid_idx) in enumerate(skf.split(candidates, candidates['click'], groups=candidates['user'] )):\n\n        X_train = candidates.loc[train_idx, FEATURES]\n        y_train = candidates.loc[train_idx, 'click']\n        X_valid = candidates.loc[valid_idx, FEATURES]\n        y_valid = candidates.loc[valid_idx, 'click']\n\n        # IF YOU HAVE 50 CANDIDATE WE USE 50 BELOW\n        dtrain = xgb.DMatrix(X_train, y_train, group=[50] * (len(train_idx)//50) ) \n        dvalid = xgb.DMatrix(X_valid, y_valid, group=[50] * (len(valid_idx)//50) ) \n\n        xgb_parms = {'objective':'rank:pairwise', 'tree_method':'gpu_hist'}\n        model = xgb.train(xgb_parms, \n            dtrain=dtrain,\n            evals=[(dtrain,'train'),(dvalid,'valid')],\n            num_boost_round=1000,\n            verbose_eval=100)\n        model.save_model(f'XGB_fold{fold}_click.xgb')\n\n**NOTE** If you have memory problems training XGB on GPU, consider downsampling negatives 2x, 4x, 10x, 20x with `frac = 0.5, 0.25, 0.1, or 0.05` (and then update group sizes in DMatrix). Or use DASK XGB with multiple GPUs. Here is example code:\n\n    positives = candidates.loc[candidates['click']==1]\n    negatives = candidates.loc[candidates['click']==0].sample(frac=0.5)\n    candidates = pd.concat([positives,negatives],axis=0,ignore_index=True)\n\n# Inference\nFor inference, we create a new candidate dataframe (using our technique to generate candidates before) but this time from Kaggle's test data. Then we make item features from all 4 weeks of Kaggle train plus 1 week of Kaggle test. And we make user features from Kaggle test. We merge the features to our candidates. Then we use our saved models to infer predictions for clicks. Lastly we select 20 by sorting the predictions and choosing 20 with.\n\n    preds = np.zeros(len(test_candidates))\n    for fold in range(5):\n        model = xgb.Booster()\n        model.load_model(f'XGB_fold{fold}_click.xgb')\n        model.set_param({'predictor': 'gpu_predictor'})\n        dtest = xgb.DMatrix(data=test_candidates[FEATURES])\n        preds += model.predict(dtest)/5\n    predictions = test_candidates[['user','item']].copy()\n    predictions['pred'] = preds\n\n    predictions = predictions.sort_values(['user','pred'], ascending=[True,False]).reset_index(drop=True)\n    predictions['n'] = predictions.groupby('user').item.cumcount().astype('int8')\n    predictions = predictions.loc[predictions.n<20]\n    sub = predictions.groupby('user').item.apply(list)\n    sub = sub.to_frame().reset_index()\n    sub.item = sub.item.apply(lambda x: \" \".join(map(str,x)))\n    sub.columns = ['session_type','labels']\n    sub.session_type = sub.session_type.astype('str')+ '_clicks'\n\n**NOTE** if you have memory errors. Consider loading 1/10th of the test data. Then merge features. Then infer. Next load the next 1/10th, merge features, infer. etc. etc. Lastly concatenate the predictions and make submission.csv\n\n# Enjoy\nI hope this discussion post helps illustrate how to create data for a GBT ranking model and how to train and infer. Have fun!\n\n[1]: https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation?select=test_labels.parquet\n[2]: https://rapids.ai/\n[3]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\n[4]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369\n[5]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991 \n[6]: https://www.dask.org/\n[7]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/371058",
      "votes": 335
    },
    {
      "id": 2059518,
      "postDate": "2022-12-09T00:47:14.520Z",
      "content": "<p>Will this work if our candidates for each session have mixed length? For one session we may have 25 candidates, for another one 30.</p>",
      "rawMarkdown": "Will this work if our candidates for each session have mixed length? For one session we may have 25 candidates, for another one 30.",
      "votes": 5,
      "replies": [
        {
          "id": 2059527,
          "postDate": "2022-12-09T01:20:53.497Z",
          "content": "<p>Yes, it will work. You just need to sort your dataframe by session and provide XGB with a list of group sizes in the order they appear in the dataframe.</p>\n<p>For example, sort your dataframe with <code>train = train.sort_values('session')</code>. Then provide a list of sizes of the sessions in order, for example <code>groups = train.groupby('session').aid.agg('count').values</code>. (Double check to make sure your dataframe library keeps the groups in the same order they appear in dataframe when doing groupby. I believe Pandas does). Then train with </p>\n<pre><code>dtrain = xgb.DMatrix(train[FEATS], train[TAR], group=groups )\n</code></pre>",
          "rawMarkdown": "Yes, it will work. You just need to sort your dataframe by session and provide XGB with a list of group sizes in the order they appear in the dataframe.\n\nFor example, sort your dataframe with `train = train.sort_values('session')`. Then provide a list of sizes of the sessions in order, for example `groups = train.groupby('session').aid.agg('count').values`. (Double check to make sure your dataframe library keeps the groups in the same order they appear in dataframe when doing groupby. I believe Pandas does). Then train with \n\n    dtrain = xgb.DMatrix(train[FEATS], train[TAR], group=groups )",
          "votes": 6
        },
        {
          "id": 2059937,
          "postDate": "2022-12-09T11:35:24.610Z",
          "content": "<p>I also have different number of candidates for each session.</p>\n<p>In my case I managed to make it work by using qid parameter instead of groups like this:</p>\n<p><code>\ndf=df.sort_values(by='session_id', ascending=True)\n</code><br>\n<code>dtrain = xgb.DMatrix(df[col_preds], df[col_target], qid=list(df['session_id']) ) \n</code></p>",
          "rawMarkdown": "I also have different number of candidates for each session.\n\nIn my case I managed to make it work by using qid parameter instead of groups like this:\n\n`\ndf=df.sort_values(by='session_id', ascending=True)\n`\n`dtrain = xgb.DMatrix(df[col_preds], df[col_target], qid=list(df['session_id']) ) \n`",
          "votes": 7
        },
        {
          "id": 2059954,
          "postDate": "2022-12-09T11:54:33.910Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2063532,
          "postDate": "2022-12-13T03:24:51.103Z",
          "content": "<p>For those working on google colab, xgb default version is 0.90, and does not have neither group or qid parameters.</p>\n<p>I upgraded it to version 1.7.2 and it worked.</p>\n<pre><code>! pip uninstall xgboost -qq -y\n! pip install xgboost==1.7.2 -qq\n</code></pre>",
          "rawMarkdown": "For those working on google colab, xgb default version is 0.90, and does not have neither group or qid parameters.\n\nI upgraded it to version 1.7.2 and it worked.\n\n```\n! pip uninstall xgboost -qq -y\n! pip install xgboost==1.7.2 -qq\n```",
          "votes": 4
        }
      ]
    },
    {
      "id": 2058286,
      "postDate": "2022-12-07T19:15:00.997Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> A new functionality of Cudf (since 22.08) is setting the default data type to be 32 or 64 bit. Using 32 bit can help save GPU memory when processing features.<br>\nThe code bellow sets 32 bit as the default:</p>\n<pre><code>import cudf\ncudf.set_option(\"default_integer_bitwidth\", 32)\ncudf.set_option(\"default_float_bitwidth\", 32)\n</code></pre>",
      "rawMarkdown": "@cdeotte A new functionality of Cudf (since 22.08) is setting the default data type to be 32 or 64 bit. Using 32 bit can help save GPU memory when processing features.\nThe code bellow sets 32 bit as the default:\n\n```\nimport cudf\ncudf.set_option(\"default_integer_bitwidth\", 32)\ncudf.set_option(\"default_float_bitwidth\", 32)\n```",
      "votes": 6,
      "replies": [
        {
          "id": 2058292,
          "postDate": "2022-12-07T19:21:22.653Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> I added this suggestion to the data management section of my post above.</p>",
          "rawMarkdown": "Thanks @titericz I added this suggestion to the data management section of my post above."
        }
      ]
    },
    {
      "id": 2078669,
      "postDate": "2022-12-28T14:12:18.810Z",
      "content": "<p>I can't understand what is happening, I'm keep getting this error: <strong><em>XGBoostError: [13:58:22] ../src/data/data.cc:694: Check failed: group_ptr_.back() == num_row_ (9624929 vs. 7699943) : Invalid group structure.  Number of rows obtained from groups doesn't equal to actual number of rows given by data.</em></strong> </p>\n<p>I read my dataframe from the disk:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F9f5888117f112e76546cd986da22172d%2FReading.JPG?generation=1672236414247488&amp;alt=media\" alt=\"\"></p>\n<p>Than I reduce the number of negatives like Chris proposed in his post:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F0431b5f064256829fe7638a54064f3c3%2FReducing.JPG?generation=1672236479917480&amp;alt=media\" alt=\"\"></p>\n<p>I check if my dataframe's length equals to the sum of the array with group counts:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F8692369e6eaeac04d2e2d5ca24b4d649%2FTrain.JPG?generation=1672236560514962&amp;alt=media\" alt=\"\"></p>\n<p>But I'm keep getting this error:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2Fca5323843c148064437fe919bf3a84b8%2FError_1.JPG?generation=1672236658532920&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2Fdc14b53811624616ca61d9db2ca45b27%2FError_2.JPG?generation=1672236667732271&amp;alt=media\" alt=\"\"></p>\n<p>Has anybody encountered anything similar to this? And if encountered what was the root cause?</p>",
      "rawMarkdown": "I can't understand what is happening, I'm keep getting this error: ***XGBoostError: [13:58:22] ../src/data/data.cc:694: Check failed: group_ptr_.back() == num_row_ (9624929 vs. 7699943) : Invalid group structure.  Number of rows obtained from groups doesn't equal to actual number of rows given by data.*** \n\nI read my dataframe from the disk:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F9f5888117f112e76546cd986da22172d%2FReading.JPG?generation=1672236414247488&alt=media)\n\nThan I reduce the number of negatives like Chris proposed in his post:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F0431b5f064256829fe7638a54064f3c3%2FReducing.JPG?generation=1672236479917480&alt=media)\n\nI check if my dataframe's length equals to the sum of the array with group counts:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F8692369e6eaeac04d2e2d5ca24b4d649%2FTrain.JPG?generation=1672236560514962&alt=media)\n\nBut I'm keep getting this error:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2Fca5323843c148064437fe919bf3a84b8%2FError_1.JPG?generation=1672236658532920&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2Fdc14b53811624616ca61d9db2ca45b27%2FError_2.JPG?generation=1672236667732271&alt=media)\n\nHas anybody encountered anything similar to this? And if encountered what was the root cause?",
      "votes": 3,
      "replies": [
        {
          "id": 2078777,
          "postDate": "2022-12-28T16:25:28.343Z",
          "content": "<p>The problem was with the <strong><em>train_groups</em></strong> and <strong><em>valid_groups</em></strong>. I'm now getting group sizes for train groups and validation groups in this way:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2Faeb5851bf0f0eb4ee8a5414dad4d1470%2FProblem%20Solved.JPG?generation=1672244683733045&amp;alt=media\" alt=\"\"></p>\n<p>This code solved the problem.</p>",
          "rawMarkdown": "The problem was with the ***train_groups*** and ***valid_groups***. I'm now getting group sizes for train groups and validation groups in this way:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2Faeb5851bf0f0eb4ee8a5414dad4d1470%2FProblem%20Solved.JPG?generation=1672244683733045&alt=media)\n\nThis code solved the problem.",
          "votes": 3,
          "replies": [
            {
              "id": 2078951,
              "postDate": "2022-12-28T19:37:25.203Z",
              "content": "<p>Looks good</p>",
              "rawMarkdown": "Looks good",
              "votes": 1
            },
            {
              "id": 2080659,
              "postDate": "2022-12-30T11:00:17.797Z",
              "content": "<p>hi,you have quite solved my problem about  the invalid number rows.but i have some new questions about you code.can you tell me your idea? the question 1 : where you get the dateset \"the_clicks_target_dataset.prt\" question 2 : your ' candidate ' is represent ' aids'? i am so confuse about it , if you have see it , please answer me , thanks!!</p>",
              "rawMarkdown": "hi,you have quite solved my problem about  the invalid number rows.but i have some new questions about you code.can you tell me your idea? the question 1 : where you get the dateset \"the_clicks_target_dataset.prt\" question 2 : your ' candidate ' is represent ' aids'? i am so confuse about it , if you have see it , please answer me , thanks!!\n",
              "votes": 1
            },
            {
              "id": 2080977,
              "postDate": "2022-12-30T17:21:14.793Z",
              "content": "<p>Hi, adairli00!</p>\n<ol>\n<li><em>clicks_target_dataset.parquet</em> - is a dataset after <strong>Step 6</strong> from Chris's post. It's a dataset of shape (#session x 50, #features) where one row is: session, aid, session features, aid features, session aid features and clciks target (0 or 1)</li>\n<li>You're right. Candiate is an aid, which I generated from co-visitation matrices for each session</li>\n</ol>",
              "rawMarkdown": "Hi, adairli00!\n1. *clicks_target_dataset.parquet* - is a dataset after **Step 6** from Chris's post. It's a dataset of shape (#session x 50, #features) where one row is: session, aid, session features, aid features, session aid features and clciks target (0 or 1)\n2. You're right. Candiate is an aid, which I generated from co-visitation matrices for each session"
            }
          ]
        }
      ]
    },
    {
      "id": 2061533,
      "postDate": "2022-12-11T08:03:33.400Z",
      "content": "<p>Hi Chris. Thanks a lot for this framework!   I was wondering that while merging candidate dataframe with the target,  should we do outer join instead of left, because on doing left join we might be missing out on some of the true (user, item) pairs which may not be in our candidate dataset<br>\n <code>candidates = candidates.merge(click_target,on=['user','item'],how='left').fillna(0)</code></p>",
      "rawMarkdown": "Hi Chris. Thanks a lot for this framework!   I was wondering that while merging candidate dataframe with the target,  should we do outer join instead of left, because on doing left join we might be missing out on some of the true (user, item) pairs which may not be in our candidate dataset\n ```candidates = candidates.merge(click_target,on=['user','item'],how='left').fillna(0)```",
      "votes": 3,
      "replies": [
        {
          "id": 2061875,
          "postDate": "2022-12-11T14:40:53.920Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kazama28\" target=\"_blank\">@kazama28</a> We want to use <code>how='left</code> because we want missing values and then convert them to zero with <code>fillna(0)</code>. These are the negative examples. To train our model we need both <code>target = 1</code> and <code>target = 0</code>. Without target = 0, our model cannot learn.</p>",
          "rawMarkdown": "Hi @kazama28 We want to use `how='left` because we want missing values and then convert them to zero with `fillna(0)`. These are the negative examples. To train our model we need both `target = 1` and `target = 0`. Without target = 0, our model cannot learn.",
          "votes": 1,
          "replies": [
            {
              "id": 2069772,
              "postDate": "2022-12-19T09:59:36.923Z",
              "content": "<p>But outer join (<code>how='outer'</code>) also preserve missing values, isn't it? As <a href=\"https://www.kaggle.com/kazama28\" target=\"_blank\">@kazama28</a> pointed out, by doing left join we would miss out groundtruths that are not in candidate set. For <code>type=click</code>, each <code>user</code> only has a single groundtruth, so if our candidate set miss that, we are left with negative examples only.</p>",
              "rawMarkdown": "But outer join (`how='outer'`) also preserve missing values, isn't it? As @kazama28 pointed out, by doing left join we would miss out groundtruths that are not in candidate set. For `type=click`, each `user` only has a single groundtruth, so if our candidate set miss that, we are left with negative examples only.",
              "votes": 1
            },
            {
              "id": 2074207,
              "postDate": "2022-12-23T20:46:17.927Z",
              "content": "<p>Good point. We can try both and see what produces better CV and better LB</p>",
              "rawMarkdown": "Good point. We can try both and see what produces better CV and better LB",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2059767,
      "postDate": "2022-12-09T07:47:35.763Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I have tried to implement the pipeline as you described above. I have managed to generate dataframes of candidates and features (items, users). However, when I tried to merge them, but I couldn't get the columns right in the merged dataframe as you described above. Could you have a look? Thanks a lot! 🙏</p>\n<p>This is your demo for merged dataframe</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fcf82ca1b474801c28527aabdc9d426e7%2FScreen%20Shot%202022-12-09%20at%2015.45.22.png?generation=1670571940780500&amp;alt=media\" alt=\"\"></p>\n<p>My candidates and item_features dataframes look like below</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Ff37405f6569eb49d42657643bfbabf12%2FScreen%20Shot%202022-12-09%20at%2015.37.35.png?generation=1670571720150971&amp;alt=media\" alt=\"\"></p>\n<p>This is the merged dataframe I got which is confusing to me.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fdf620b7d3ffb9545defd25ed136e6068%2FScreen%20Shot%202022-12-09%20at%2015.37.47.png?generation=1670571846965756&amp;alt=media\" alt=\"\"></p>\n<p>I have a feeling that maybe the problem lies on index, but tried to read the docs on <code>merge</code> and <code>set_index</code>, but have not figured out how to fix it.</p>",
      "rawMarkdown": "Hi @cdeotte, I have tried to implement the pipeline as you described above. I have managed to generate dataframes of candidates and features (items, users). However, when I tried to merge them, but I couldn't get the columns right in the merged dataframe as you described above. Could you have a look? Thanks a lot! 🙏\n\nThis is your demo for merged dataframe\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fcf82ca1b474801c28527aabdc9d426e7%2FScreen%20Shot%202022-12-09%20at%2015.45.22.png?generation=1670571940780500&alt=media)\n\nMy candidates and item_features dataframes look like below\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Ff37405f6569eb49d42657643bfbabf12%2FScreen%20Shot%202022-12-09%20at%2015.37.35.png?generation=1670571720150971&alt=media)\n\nThis is the merged dataframe I got which is confusing to me.\n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fdf620b7d3ffb9545defd25ed136e6068%2FScreen%20Shot%202022-12-09%20at%2015.37.47.png?generation=1670571846965756&alt=media)\n\nI have a feeling that maybe the problem lies on index, but tried to read the docs on `merge` and `set_index`, but have not figured out how to fix it.\n",
      "votes": 3
    },
    {
      "id": 2059625,
      "postDate": "2022-12-09T04:44:12.390Z",
      "content": "<p>Hi,Chris,your answer helped me a lot!!!<br>\nStep 4: I have a question. You use as an example whether an item is clicked as a feature of user-item interaction. But if I train a ranker to sort the candidates in click. Does this feature above leak my training label? That means that if I want to train a click reranker, I can't use both the user and item interaction features related to click? I am very confused about this and look forward to your answer.</p>",
      "rawMarkdown": "Hi,Chris,your answer helped me a lot!!!\nStep 4: I have a question. You use as an example whether an item is clicked as a feature of user-item interaction. But if I train a ranker to sort the candidates in click. Does this feature above leak my training label? That means that if I want to train a click reranker, I can't use both the user and item interaction features related to click? I am very confused about this and look forward to your answer.",
      "votes": 3,
      "replies": [
        {
          "id": 2059914,
          "postDate": "2022-12-09T10:57:40.477Z",
          "content": "<p><a href=\"https://www.kaggle.com/niejianfei\" target=\"_blank\">@niejianfei</a> there is no leak because we do not use ground truth to create interaction features. We begin with the 4 weeks of Kaggle train data. Then we build 3 datasets</p>\n<ul>\n<li>new train data (first 3 weeks)</li>\n<li>validation data A (first half of week 4)</li>\n<li>validation data B (second half of week 4)</li>\n</ul>\n<p>The ground truth is validation data B. We build interaction features from validation data A. Each of the test users activity is split in half. The first half of their activity is in Val A and the second half in Val B. For example, user A clicks 001 on monday, clicks 003 on Tuesday, clicks 002 on Wednesday and clicks 001 on Thursday. The Val A is <code>clicks = [001, 003]</code> from monday and tuesday. And Val B is <code>clicks = [002, 001]</code> from Wednesday and Thursday. Our interaction features are user A clicked 001 and 003. This is information from Monday and Tuesday. This will help us predict Wednesday and Thursday. </p>",
          "rawMarkdown": "@niejianfei there is no leak because we do not use ground truth to create interaction features. We begin with the 4 weeks of Kaggle train data. Then we build 3 datasets\n* new train data (first 3 weeks)\n* validation data A (first half of week 4)\n* validation data B (second half of week 4)\n\nThe ground truth is validation data B. We build interaction features from validation data A. Each of the test users activity is split in half. The first half of their activity is in Val A and the second half in Val B. For example, user A clicks 001 on monday, clicks 003 on Tuesday, clicks 002 on Wednesday and clicks 001 on Thursday. The Val A is `clicks = [001, 003]` from monday and tuesday. And Val B is `clicks = [002, 001]` from Wednesday and Thursday. Our interaction features are user A clicked 001 and 003. This is information from Monday and Tuesday. This will help us predict Wednesday and Thursday. ",
          "votes": 4
        },
        {
          "id": 2059915,
          "postDate": "2022-12-09T11:02:01.993Z",
          "content": "<p>Note these 3 datasets we make are comparable to when we make a submission to LB. For submissions, we have (1) kaggle train (first 4 weeks) (2) kaggle test (first half of week 5) (3) Leaderboard (second half of week 5)</p>",
          "rawMarkdown": "Note these 3 datasets we make are comparable to when we make a submission to LB. For submissions, we have (1) kaggle train (first 4 weeks) (2) kaggle test (first half of week 5) (3) Leaderboard (second half of week 5)",
          "votes": 2
        },
        {
          "id": 2059960,
          "postDate": "2022-12-09T12:06:52.090Z",
          "content": "<p>Chris, thank you very much. I get it now.I am a novice in data science competition, thank you so much for sharing!!</p>",
          "rawMarkdown": "Chris, thank you very much. I get it now.I am a novice in data science competition, thank you so much for sharing!!",
          "votes": 2
        }
      ]
    },
    {
      "id": 2069145,
      "postDate": "2022-12-18T16:44:24.107Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>,<br>\nFirst at all, thank you very much for this post and all the others in this complicated competition ! <br>\nI have a question : in \"Training\", you suggested to consider downsampling negatives ; but if we do so, we will not have 50 consecutive samples from the same user and groups won't be correct for XGB. Do you agree ?</p>\n<p>(I'm looking for an issue to solve memory problems, and I tell myself that I will not consider downsampling negatives, but features forward selection. Or maybe I should consider downsampling negatives, but take care to \"group=\" in xgb.DMatrix)</p>",
      "rawMarkdown": "Hi @cdeotte,\nFirst at all, thank you very much for this post and all the others in this complicated competition ! \nI have a question : in \"Training\", you suggested to consider downsampling negatives ; but if we do so, we will not have 50 consecutive samples from the same user and groups won't be correct for XGB. Do you agree ?\n\n(I'm looking for an issue to solve memory problems, and I tell myself that I will not consider downsampling negatives, but features forward selection. Or maybe I should consider downsampling negatives, but take care to \"group=\" in xgb.DMatrix)",
      "votes": 4,
      "replies": [
        {
          "id": 2069203,
          "postDate": "2022-12-18T18:05:56.417Z",
          "content": "<p>When training GBT rankers, we do not need to have the same number of groups per user. So after downsampling, we count how many candidates are in each user and input this updated list to \"group=\" in xgb.DMatrix. </p>\n<p>Note there are alternatives besides downsample to use less GPU RAM. Here are two more ideas:</p>\n<ul>\n<li>train with a subset of users (and 50 candidates each) each fold. And validate and infer with all users (like normal).</li>\n<li>use XGB dataloader together with <code>DeviceQuantileDMatrix</code>. This allows XGB to use less memory and train with larger datasets. Also consider training a fixed number of iterations and do not use a valid dataset nor early stopping during training (to maximize acceptable train data size). Sample code <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a></li>\n</ul>",
          "rawMarkdown": "When training GBT rankers, we do not need to have the same number of groups per user. So after downsampling, we count how many candidates are in each user and input this updated list to \"group=\" in xgb.DMatrix. \n\nNote there are alternatives besides downsample to use less GPU RAM. Here are two more ideas:\n* train with a subset of users (and 50 candidates each) each fold. And validate and infer with all users (like normal).\n* use XGB dataloader together with `DeviceQuantileDMatrix`. This allows XGB to use less memory and train with larger datasets. Also consider training a fixed number of iterations and do not use a valid dataset nor early stopping during training (to maximize acceptable train data size). Sample code [here][1]\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
          "votes": 7,
          "replies": [
            {
              "id": 2069282,
              "postDate": "2022-12-18T20:15:14.903Z",
              "content": "<p>Thank you for confirmation and alternatives !</p>",
              "rawMarkdown": "Thank you for confirmation and alternatives !",
              "votes": 1
            },
            {
              "id": 2069307,
              "postDate": "2022-12-18T20:50:43.717Z",
              "content": "<p>I added a comment to my original post reminding people to update group sizes in DMatrix if they downsample. Thanks for pointing this out.</p>",
              "rawMarkdown": "I added a comment to my original post reminding people to update group sizes in DMatrix if they downsample. Thanks for pointing this out.",
              "votes": 1
            },
            {
              "id": 2073840,
              "postDate": "2022-12-23T13:14:17.713Z",
              "content": "<p>Inspite of downsampling 20x, there is memory error. <br>\nI have 10 chunks of train data which I am concatenating into dask dataframe. </p>\n<p>I am also trying dask_cudf with multiple GPUs as shown <a href=\"https://developer.nvidia.com/blog/accelerating-xgboost-on-gpu-clusters-with-dask/\" target=\"_blank\">here</a> but still facing memory error. </p>\n<p>I am at a loss how to deal with this much scale. I am a novice so requesting for some tips.  </p>",
              "rawMarkdown": "Inspite of downsampling 20x, there is memory error. \nI have 10 chunks of train data which I am concatenating into dask dataframe. \n\n I am also trying dask_cudf with multiple GPUs as shown [here](https://developer.nvidia.com/blog/accelerating-xgboost-on-gpu-clusters-with-dask/) but still facing memory error. \n\nI am at a loss how to deal with this much scale. I am a novice so requesting for some tips.  ",
              "votes": 1
            },
            {
              "id": 2073865,
              "postDate": "2022-12-23T13:48:35.953Z",
              "content": "<p><a href=\"https://www.kaggle.com/sohamdats\" target=\"_blank\">@sohamdats</a> Do you have memory error when creating your dataframe, or do you have memory error when training XGB?</p>\n<p>I recommend creating your dataframe in one python script or jupyter notebook, then save dataframe to disk as parquet and make sure to reduce every column to least dtype (like float32). Next stop all code and clear all memory. Then start a new python script or jupyter notebook, read in the parquet and begin training XGB.</p>\n<p>When training XGB, use DeviceQuantileDMatrix . Train for a fixed number of iterations without validation data. So you will only give your XGB train data. Are you training with dask-xgb and using all your GPU during training or are you using 1 GPU with regular XGB?</p>",
              "rawMarkdown": "@sohamdats Do you have memory error when creating your dataframe, or do you have memory error when training XGB?\n\nI recommend creating your dataframe in one python script or jupyter notebook, then save dataframe to disk as parquet and make sure to reduce every column to least dtype (like float32). Next stop all code and clear all memory. Then start a new python script or jupyter notebook, read in the parquet and begin training XGB.\n\nWhen training XGB, use DeviceQuantileDMatrix . Train for a fixed number of iterations without validation data. So you will only give your XGB train data. Are you training with dask-xgb and using all your GPU during training or are you using 1 GPU with regular XGB?",
              "votes": 2
            },
            {
              "id": 2073924,
              "postDate": "2022-12-23T14:48:36.897Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> First of all thanks a lot for replying.</p>\n<p>I am facing memory error during training. I have saved 10 chunks in parquet format and loading them in a separate notebook for training. I am putting all of them together into a dask dataframe. I am reducing all data types.  </p>\n<p>First I tried to run the <a href=\"https://www.kaggle.com/code/radek1/training-an-xgboost-ranker-on-the-gpu/comments\" target=\"_blank\">Radek's notebook</a>  which threw a memory error during the ranker.fit stage. </p>\n<p>Then tried to run the <a href=\"https://developer.nvidia.com/blog/accelerating-xgboost-on-gpu-clusters-with-dask/\" target=\"_blank\">code</a> here. It is getting stuck in the load dataset method. This code uses DaskDeviceQuantileDMatrix but error is occuring before reaching there. </p>\n<p>I put n_workers=1 because &gt;1 was throwing error in kaggle notebook in GPU P100.  </p>",
              "rawMarkdown": "@cdeotte First of all thanks a lot for replying.\n\nI am facing memory error during training. I have saved 10 chunks in parquet format and loading them in a separate notebook for training. I am putting all of them together into a dask dataframe. I am reducing all data types.  \n\nFirst I tried to run the [Radek's notebook](https://www.kaggle.com/code/radek1/training-an-xgboost-ranker-on-the-gpu/comments)  which threw a memory error during the ranker.fit stage. \n\nThen tried to run the [code](https://developer.nvidia.com/blog/accelerating-xgboost-on-gpu-clusters-with-dask/) here. It is getting stuck in the load dataset method. This code uses DaskDeviceQuantileDMatrix but error is occuring before reaching there. \n\nI put n_workers=1 because >1 was throwing error in kaggle notebook in GPU P100.  ",
              "votes": 1
            },
            {
              "id": 2073930,
              "postDate": "2022-12-23T14:53:18.690Z",
              "content": "<p>If you are using dask-cudf and/or dask-xgb in Kaggle notebooks, then you should use 2xT4 GPU then you will have 2x16 = 32GB or GPU RAM. With 1xP100, you only have 16GB. </p>",
              "rawMarkdown": "If you are using dask-cudf and/or dask-xgb in Kaggle notebooks, then you should use 2xT4 GPU then you will have 2x16 = 32GB or GPU RAM. With 1xP100, you only have 16GB. ",
              "votes": 2
            },
            {
              "id": 2073932,
              "postDate": "2022-12-23T14:55:38.767Z",
              "content": "<p>Note if you already have your train data in multiple parquets on disk, then you do not need to use my <code>DeviceQuantileDMatrix</code> dataloader. If you search the internet for tutorials, you will see that <code>DeviceQuantileDMatrix</code> can load multiple parquets directly from disk.</p>\n<p>The purpose of my dataloader is to break a single dataframe into chunks because <code>DeviceQuantileDMatrix</code> needs chunks. However if you already have chunks on disk, then perhaps reading from disk is better.</p>",
              "rawMarkdown": "Note if you already have your train data in multiple parquets on disk, then you do not need to use my `DeviceQuantileDMatrix` dataloader. If you search the internet for tutorials, you will see that `DeviceQuantileDMatrix` can load multiple parquets directly from disk.\n\nThe purpose of my dataloader is to break a single dataframe into chunks because `DeviceQuantileDMatrix` needs chunks. However if you already have chunks on disk, then perhaps reading from disk is better.",
              "votes": 1
            },
            {
              "id": 2073935,
              "postDate": "2022-12-23T14:58:04.037Z",
              "content": "<p>Okay thanks a lot. I forgot that DeviceQuantileMatrix is a dataloader and I can read from disk itself. I will certainly try this. I can't thank you more. </p>",
              "rawMarkdown": "Okay thanks a lot. I forgot that DeviceQuantileMatrix is a dataloader and I can read from disk itself. I will certainly try this. I can't thank you more. ",
              "votes": 1
            },
            {
              "id": 2073952,
              "postDate": "2022-12-23T15:19:01.763Z",
              "content": "<p><a href=\"https://www.kaggle.com/sohamdats\" target=\"_blank\">@sohamdats</a> i just searched internet for info about reading from disk. I think the way to read from disk is to use my dataloader but change the \"next function\". My current dataloader from my code <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a> is like this</p>\n<pre><code>def next(self, input_data):\n    '''Yield next batch of data.'''\n    if self.it == self.batches:\n        return 0 # Return 0 when there's no more batch.\n\n    a = self.it * self.batch_size\n    b = min( (self.it + 1) * self.batch_size, len(self.df) )\n    dt = cudf.DataFrame(self.df.iloc[a:b])\n    input_data(data=dt[self.features], label=dt[self.target]) \n    self.it += 1\n    return 1\n</code></pre>\n<p>We notice that for each new request in the form of <code>self.it</code> we take a chunk from the dataframe in memory. Instead we can have it read a different parquet from disk for each new <code>self.it</code>. Something like the following</p>\n<pre><code>def next(self, input_data):\n    '''Yield next batch of data.'''\n    if self.it == self.batches:\n        return 0 # Return 0 when there's no more batch.\n\n    dt = cudf.read_parquet(f'train_{self.it}.parquet')\n    input_data(data=dt[self.features], label=dt[self.target]) \n    self.it += 1\n    return 1\n</code></pre>\n<p>I haven't tested this out, but perhaps reading from disk will use less memory than reading from dataframe in memory. Perhaps there is a way for <code>DeviceQuantileDMatrix</code> to read directly from disk without a dataloader, but i am not sure how to do that.</p>",
              "rawMarkdown": "@sohamdats i just searched internet for info about reading from disk. I think the way to read from disk is to use my dataloader but change the \"next function\". My current dataloader from my code [here][1] is like this\n\n    def next(self, input_data):\n        '''Yield next batch of data.'''\n        if self.it == self.batches:\n            return 0 # Return 0 when there's no more batch.\n        \n        a = self.it * self.batch_size\n        b = min( (self.it + 1) * self.batch_size, len(self.df) )\n        dt = cudf.DataFrame(self.df.iloc[a:b])\n        input_data(data=dt[self.features], label=dt[self.target]) \n        self.it += 1\n        return 1\n\nWe notice that for each new request in the form of `self.it` we take a chunk from the dataframe in memory. Instead we can have it read a different parquet from disk for each new `self.it`. Something like the following\n\n    def next(self, input_data):\n        '''Yield next batch of data.'''\n        if self.it == self.batches:\n            return 0 # Return 0 when there's no more batch.\n        \n        dt = cudf.read_parquet(f'train_{self.it}.parquet')\n        input_data(data=dt[self.features], label=dt[self.target]) \n        self.it += 1\n        return 1\n\nI haven't tested this out, but perhaps reading from disk will use less memory than reading from dataframe in memory. Perhaps there is a way for `DeviceQuantileDMatrix` to read directly from disk without a dataloader, but i am not sure how to do that.\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
              "votes": 1
            },
            {
              "id": 2075137,
              "postDate": "2022-12-25T05:55:05.583Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks a lot for taking out the time and replying. I am going to try this from your notebook as you suggested.</p>",
              "rawMarkdown": "@cdeotte Thanks a lot for taking out the time and replying. I am going to try this from your notebook as you suggested."
            },
            {
              "id": 2081885,
              "postDate": "2022-12-31T19:16:13.220Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I am confused about how we should specify the groups when using DeviceQuantileDMatrix, should it be specified in the \"next function\" (as a 'group' parameter to input_data)? </p>",
              "rawMarkdown": "Hi @cdeotte, I am confused about how we should specify the groups when using DeviceQuantileDMatrix, should it be specified in the \"next function\" (as a 'group' parameter to input_data)? \n"
            },
            {
              "id": 2081890,
              "postDate": "2022-12-31T19:24:15.130Z",
              "content": "<p><a href=\"https://www.kaggle.com/kazama28\" target=\"_blank\">@kazama28</a> The code is like this</p>\n<pre><code>    dtrain = xgb.DeviceQuantileDMatrix(Xy_train, max_bin=256)\n    dtrain.set_group( [50] * ( len(train_idx)//50) )\n</code></pre>\n<p>Note that <code>set_group</code> also works for normal DMatrix</p>",
              "rawMarkdown": "@kazama28 The code is like this\n\n        dtrain = xgb.DeviceQuantileDMatrix(Xy_train, max_bin=256)\n        dtrain.set_group( [50] * ( len(train_idx)//50) )\n\nNote that `set_group` also works for normal DMatrix",
              "votes": 1
            },
            {
              "id": 2082103,
              "postDate": "2023-01-01T05:33:55.047Z",
              "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
              "rawMarkdown": "Thanks a lot @cdeotte ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2054471,
      "postDate": "2022-12-04T07:03:44.707Z",
      "content": "<p>Thanks for sharing， I have a question that's why step 3 and step 4 only use the validation data  but not the train data and the validation data.  In the Inference step it seems to use similar set up which use 4week Kaggle train plus 1 week of Kaggle test to generate item feature but only use Kaggle test to generate user features. Could you explain it a bit ? Thanks!</p>",
      "rawMarkdown": "Thanks for sharing， I have a question that's why step 3 and step 4 only use the validation data  but not the train data and the validation data.  In the Inference step it seems to use similar set up which use 4week Kaggle train plus 1 week of Kaggle test to generate item feature but only use Kaggle test to generate user features. Could you explain it a bit ? Thanks!",
      "votes": 4,
      "replies": [
        {
          "id": 2054781,
          "postDate": "2022-12-04T13:04:18.230Z",
          "content": "<p>Great question. When making features we would like to use the most amount of data as possible. But…</p>\n<p>The reason that step 3 and 4 only use validation data is because there is no overlap between train users and valid users. We only have targets for validation data and we only generate candidates for validation users. If we created user features from train users, none of these features would get merged to our candidate dataframe and none would be used.</p>\n<p>If you use unsupervised learning (which is not explained in this discussion post) which does not require targets, then you can use information about train users in addition to information about validation users. For example, run PCA on all users (from train and valid). Then create features from the PCA data. Then when adding features to our candidate list, first use PCA on validation users, then merge PCA features built from all users.</p>\n<p>NOTE: I updated my discussion post to fix step 1. In step 1, we only generate candidates for the users in validation data. (Because we only have targets for validation users)</p>",
          "rawMarkdown": "Great question. When making features we would like to use the most amount of data as possible. But...\n\nThe reason that step 3 and 4 only use validation data is because there is no overlap between train users and valid users. We only have targets for validation data and we only generate candidates for validation users. If we created user features from train users, none of these features would get merged to our candidate dataframe and none would be used.\n\nIf you use unsupervised learning (which is not explained in this discussion post) which does not require targets, then you can use information about train users in addition to information about validation users. For example, run PCA on all users (from train and valid). Then create features from the PCA data. Then when adding features to our candidate list, first use PCA on validation users, then merge PCA features built from all users.\n\nNOTE: I updated my discussion post to fix step 1. In step 1, we only generate candidates for the users in validation data. (Because we only have targets for validation users)",
          "votes": 5
        },
        {
          "id": 2055536,
          "postDate": "2022-12-05T06:34:03.147Z",
          "content": "<p>I am also confused by this point… I understand why you have no need for user features in the train set during test set inference, but then what user features do you train on if they're all from the validation set? Or do you mean you're using the validation set for training?</p>\n<p>Perhaps my confusion (and maybe others') would be cleared up with a small, worked example. My initial confusion was when the user and item features use the <code>train</code> dataframe, but the text states to only use validation users. So then I wondered what that meant, for example. If you could import the data and run a very small example, end-to-end, it would be so helpful. Personally, I like to look at the data at each step and break things to understand the process. I completely understand if you'd rather not, however, and appreciate the work you've put in thus far 😊</p>\n<p>EDIT: I may have wrapped my head around it. You use the train and validation set in that the train is used for item features. You then exclusively use the validation set for user features. You then have user features, item features, and candidates for the validation set (which is now effectively the training set) and add on labels from the ground truth. The positive candidates are the correct candidates while the negative are the incorrect.</p>\n<p>In the spirit of sharing what I was trying as well: I was using the train/validation set as the ground truth and generating hard negative samples for each label, which, as you might imagine, can also be pretty memory intensive. I like your approach in that you're training using the candidates you specifically generated as opposed to random samples. I've tried this approach at work, after having the idea, and found that hard negative sampling performed better for my specific data. I hadn't read any literature on this method either, honestly, so I figured it was just an odd idea. I wonder what the winner will be in this case!</p>",
          "rawMarkdown": "I am also confused by this point... I understand why you have no need for user features in the train set during test set inference, but then what user features do you train on if they're all from the validation set? Or do you mean you're using the validation set for training?\n\nPerhaps my confusion (and maybe others') would be cleared up with a small, worked example. My initial confusion was when the user and item features use the `train` dataframe, but the text states to only use validation users. So then I wondered what that meant, for example. If you could import the data and run a very small example, end-to-end, it would be so helpful. Personally, I like to look at the data at each step and break things to understand the process. I completely understand if you'd rather not, however, and appreciate the work you've put in thus far 😊\n\n\nEDIT: I may have wrapped my head around it. You use the train and validation set in that the train is used for item features. You then exclusively use the validation set for user features. You then have user features, item features, and candidates for the validation set (which is now effectively the training set) and add on labels from the ground truth. The positive candidates are the correct candidates while the negative are the incorrect.\n\nIn the spirit of sharing what I was trying as well: I was using the train/validation set as the ground truth and generating hard negative samples for each label, which, as you might imagine, can also be pretty memory intensive. I like your approach in that you're training using the candidates you specifically generated as opposed to random samples. I've tried this approach at work, after having the idea, and found that hard negative sampling performed better for my specific data. I hadn't read any literature on this method either, honestly, so I figured it was just an odd idea. I wonder what the winner will be in this case!",
          "votes": 2
        },
        {
          "id": 2055851,
          "postDate": "2022-12-05T13:46:01.010Z",
          "content": "<p><a href=\"https://www.kaggle.com/dpalbrecht\" target=\"_blank\">@dpalbrecht</a> I agree where features come from is weird because there is <strong>no overlap</strong> between test users and train users. Note this is <strong>not the typical</strong> recommender system problem. Usually when building a recommender system model, we have lots of history on most users. In this competition we have <strong>no history</strong> for test users. Every test user is a \"cold start\" new user!</p>\n<p>The only long history we have is items. Therefore we can build item features from all our data. However when it comes to test users, we don't have their history so we cannot make features from early data. We can only make features from when these users engaged at Otto (i.e. the validation data and/or test data)</p>\n<p>Note that there are many ways to build unsupervised user features using PCA, auto encoder etc. For example, you can train an auto encoder on all user data (for all 5 weeks of data). Then give every user their embedding as a feature (during training and inference). This is a way to use all the user data if you want.</p>\n<p>You can also build a second fold (i.e \"time fold\"). Currently we are using one fold. (1) We train on weeks 1,2,3 and validate on week 4 of train. You could make a second fold where you (2) train on weeks 1,2 and validate on week 3. For the second fold, we would generate user features from the 3rd week of train data. This would be another way to use more users information during training.</p>",
          "rawMarkdown": "@dpalbrecht I agree where features come from is weird because there is **no overlap** between test users and train users. Note this is **not the typical** recommender system problem. Usually when building a recommender system model, we have lots of history on most users. In this competition we have **no history** for test users. Every test user is a \"cold start\" new user!\n\nThe only long history we have is items. Therefore we can build item features from all our data. However when it comes to test users, we don't have their history so we cannot make features from early data. We can only make features from when these users engaged at Otto (i.e. the validation data and/or test data)\n\nNote that there are many ways to build unsupervised user features using PCA, auto encoder etc. For example, you can train an auto encoder on all user data (for all 5 weeks of data). Then give every user their embedding as a feature (during training and inference). This is a way to use all the user data if you want.\n\nYou can also build a second fold (i.e \"time fold\"). Currently we are using one fold. (1) We train on weeks 1,2,3 and validate on week 4 of train. You could make a second fold where you (2) train on weeks 1,2 and validate on week 3. For the second fold, we would generate user features from the 3rd week of train data. This would be another way to use more users information during training.",
          "votes": 4
        },
        {
          "id": 2057381,
          "postDate": "2022-12-07T04:00:57.003Z",
          "content": "<p>Yeah, this formulation is feeling very odd to me for that reason. I'll have to try the PCA or autoencoder and time fold methods - those sound interesting. Thanks again for sharing!</p>",
          "rawMarkdown": "Yeah, this formulation is feeling very odd to me for that reason. I'll have to try the PCA or autoencoder and time fold methods - those sound interesting. Thanks again for sharing!",
          "votes": 1
        },
        {
          "id": 2058478,
          "postDate": "2022-12-08T01:06:27.830Z",
          "content": "<p>Hi, Chris <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , I got the similar question as <a href=\"https://www.kaggle.com/dpalbrecht\" target=\"_blank\">@dpalbrecht</a>. I saw you made:</p>\n<ol>\n<li>train data: week 1-3 .</li>\n<li>valid data: week 4.</li>\n<li>valid label: week 4.</li>\n<li>test data: week5.</li>\n<li>test unseem label: after week 5.</li>\n</ol>\n<p>You used valid data from week 4 to make features and then predicted the valid label from week 4. It there a data leakage? Because valid data has already contains the information of valid label, which likes using week 4 data to predict week 4 data. However, we need to use week 5 test data to predict the future interacted aids on the week6 or later week.</p>\n<p>I got you that \"there is no session overlap between train data and valid data\" so you use valid data to make item features, but I am worrying the data leakage which bring us overfitting. Please let me know if I am right. By the way, making session embeddings by PCA or AE is a good idea to follow.</p>",
          "rawMarkdown": "Hi, Chris @cdeotte , I got the similar question as @dpalbrecht. I saw you made:\n1. train data: week 1-3 .\n2. valid data: week 4.\n3. valid label: week 4.\n4. test data: week5.\n5. test unseem label: after week 5.\n\nYou used valid data from week 4 to make features and then predicted the valid label from week 4. It there a data leakage? Because valid data has already contains the information of valid label, which likes using week 4 data to predict week 4 data. However, we need to use week 5 test data to predict the future interacted aids on the week6 or later week.\n\nI got you that \"there is no session overlap between train data and valid data\" so you use valid data to make item features, but I am worrying the data leakage which bring us overfitting. Please let me know if I am right. By the way, making session embeddings by PCA or AE is a good idea to follow.",
          "votes": 2
        },
        {
          "id": 2058486,
          "postDate": "2022-12-08T01:17:02.917Z",
          "content": "<p>Thanks for asking this question. I need to clarify my post better. There are 5 datasets. Week 4 is split into valid data and ground truth. And public LB are ground truth from week 5. I need to find a clear way to explain this in my post. Maybe i'll add a  picture</p>\n<ol>\n<li>train data week 1-3</li>\n<li>valid data week 4 without ground truth (i.e. random first half of user activity week 4)</li>\n<li>valid data week 4 ground truth (i.e. random second half of user activity week 4)</li>\n<li>test data week 5 without ground truth (i.e. random first half of user activity week 5)</li>\n<li>test data week 5 ground truth (i.e. random second half of user activity week 5)</li>\n</ol>\n<p>All user data from Week 4 and week 5 is randomly split with <code>x = np.random.randint(1, 'user_number_of_events' )</code>. Then valid week 4 without ground truth is <code>user_events[:x]</code> and week  4 ground truth is <code>user_events[x:]</code> where <code>user_events</code> are in time order. I'll add a diagram to my post to clarify.</p>\n<p>We use 1+2 to make features during train. And we use 1+2+3+4 to make features during submission to LB. We do not use the valid ground truth #3 when building features to train our models.</p>",
          "rawMarkdown": "Thanks for asking this question. I need to clarify my post better. There are 5 datasets. Week 4 is split into valid data and ground truth. And public LB are ground truth from week 5. I need to find a clear way to explain this in my post. Maybe i'll add a  picture\n\n1. train data week 1-3\n2. valid data week 4 without ground truth (i.e. random first half of user activity week 4)\n3. valid data week 4 ground truth (i.e. random second half of user activity week 4)\n4. test data week 5 without ground truth (i.e. random first half of user activity week 5)\n5. test data week 5 ground truth (i.e. random second half of user activity week 5)\n\nAll user data from Week 4 and week 5 is randomly split with `x = np.random.randint(1, 'user_number_of_events' )`. Then valid week 4 without ground truth is `user_events[:x]` and week  4 ground truth is `user_events[x:]` where `user_events` are in time order. I'll add a diagram to my post to clarify.\n\nWe use 1+2 to make features during train. And we use 1+2+3+4 to make features during submission to LB. We do not use the valid ground truth #3 when building features to train our models.",
          "votes": 6
        },
        {
          "id": 2058488,
          "postDate": "2022-12-08T01:22:13.170Z",
          "content": "<p><a href=\"https://www.kaggle.com/alvinai9603\" target=\"_blank\">@alvinai9603</a> I updated my discussion to use <code>validation data A</code> and <code>validation data B</code>. Let me know if it is clearer now.</p>",
          "rawMarkdown": "@alvinai9603 I updated my discussion to use `validation data A` and `validation data B`. Let me know if it is clearer now.",
          "votes": 2
        },
        {
          "id": 2058515,
          "postDate": "2022-12-08T01:56:27.160Z",
          "content": "<p>Nice! It is much clear and there is not any data leakage as you said above. Thanks for sharing, Chris. 👍</p>",
          "rawMarkdown": "Nice! It is much clear and there is not any data leakage as you said above. Thanks for sharing, Chris. 👍",
          "votes": 1
        }
      ]
    },
    {
      "id": 2123349,
      "postDate": "2023-01-31T13:18:27.057Z",
      "content": "<p>It works, thx!</p>",
      "rawMarkdown": "It works, thx!",
      "votes": 1
    },
    {
      "id": 2114447,
      "postDate": "2023-01-25T01:53:11.867Z",
      "content": "<p>Hi Chris, Thanks for sharing. </p>\n<p>I am having trouble with memory management: I split the candidate dataframe into chunks as you suggested, but I still seem to run out of memory while loading the chunks.</p>\n<p>I used this code to split the dataframe</p>\n<pre><code>CHUNKS = 10\n\nchunk_size = int(np.ceil( len(candidates) / CHUNKS))\n\nfor k in range(CHUNKS):\n\n    df = candidates.iloc[k*chunk_size:(k+1)*chunk_size].copy()\n\n    df = df.merge(item_features, left_on='aid', right_index=True, how='left').fillna(0)\n\n    df = df.merge(user_features, left_on='session', right_index=True, how='left').fillna(0)\n\n    df.to_parquet(f'carts_candidate_with_features_p{k}_v1.pqt') \n</code></pre>\n<p>and this to read the dataframe, </p>\n<pre><code>df_container=[]\n\nfor k in range(CHUNKS):\n\n    df_container.append(cudf.read_parquet(f'carts_candidate_with_features_p{k}_v1.pqt'))\n\n\n candidates=cudf.concat(df_container)\n</code></pre>\n<p>But I am having memory problems. I tried other methods too but faced the same problem.</p>\n<p>I have generated 50 candidates per user, and then down sampled them by 20% and I have 49 features.<br>\nI also have converted the datatypes to <code>float32</code> and <code>int32</code>.</p>\n<p>Is there anything I can do?</p>",
      "rawMarkdown": "Hi Chris, Thanks for sharing. \n\nI am having trouble with memory management: I split the candidate dataframe into chunks as you suggested, but I still seem to run out of memory while loading the chunks.\n\nI used this code to split the dataframe\n\n\n```\nCHUNKS = 10\n\nchunk_size = int(np.ceil( len(candidates) / CHUNKS))\n\nfor k in range(CHUNKS):\n\n    df = candidates.iloc[k*chunk_size:(k+1)*chunk_size].copy()\n\n    df = df.merge(item_features, left_on='aid', right_index=True, how='left').fillna(0)\n\n    df = df.merge(user_features, left_on='session', right_index=True, how='left').fillna(0)\n\n    df.to_parquet(f'carts_candidate_with_features_p{k}_v1.pqt') \n```\n\n\nand this to read the dataframe, \n\n\n```\ndf_container=[]\n\nfor k in range(CHUNKS):\n\n    df_container.append(cudf.read_parquet(f'carts_candidate_with_features_p{k}_v1.pqt'))\n\n\n candidates=cudf.concat(df_container)\n\n```\n\n\n\n\nBut I am having memory problems. I tried other methods too but faced the same problem.\n\nI have generated 50 candidates per user, and then down sampled them by 20% and I have 49 features.\nI also have converted the datatypes to `float32` and `int32`.\n\nIs there anything I can do?\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 2114569,
          "postDate": "2023-01-25T04:51:06.043Z",
          "content": "<p>When do you get a memory error. Does it happen when you try to read all the chunks and concat into one dataframe? Or does it happen when you start training XGB?</p>\n<p>One thing you can do is read in half your chunks and train one XGB with half your chunks. Then clear memory. Then read in the other half of chunks and train a second XGB. Finally infer the test data with both XGB and average the predictions.</p>\n<p>A second idea is to use 2xT4 GPU and DASK which doubles your GPU RAM for cudf and xgb.</p>",
          "rawMarkdown": "When do you get a memory error. Does it happen when you try to read all the chunks and concat into one dataframe? Or does it happen when you start training XGB?\n\nOne thing you can do is read in half your chunks and train one XGB with half your chunks. Then clear memory. Then read in the other half of chunks and train a second XGB. Finally infer the test data with both XGB and average the predictions.\n\nA second idea is to use 2xT4 GPU and DASK which doubles your GPU RAM for cudf and xgb.",
          "votes": 1,
          "replies": [
            {
              "id": 2115788,
              "postDate": "2023-01-26T00:50:05.857Z",
              "content": "<p>Thanks for your response. It happens while reading the chunks and concating them into a single dataframe. I haven't even got to train a model on he complete dataset.</p>\n<blockquote>\n  <p>One thing you can do is read in half your chunks and train one XGB with half your chunks. Then clear memory. Then read in the other half of chunks and train a second XGB. Finally infer the test data with both XGB and average the predictions.</p>\n</blockquote>\n<p>Thanks for your suggestion. I will try this.</p>\n<blockquote>\n  <p>A second idea is to use 2xT4 GPU and DASK which doubles your GPU RAM for cudf and xgb.</p>\n</blockquote>\n<p>I will try this too. I am new to DASK and learning.</p>\n<p>Thank you so much for your time.</p>",
              "rawMarkdown": "Thanks for your response. It happens while reading the chunks and concating them into a single dataframe. I haven't even got to train a model on he complete dataset.\n\n>One thing you can do is read in half your chunks and train one XGB with half your chunks. Then clear memory. Then read in the other half of chunks and train a second XGB. Finally infer the test data with both XGB and average the predictions.\n\nThanks for your suggestion. I will try this.\n\n>A second idea is to use 2xT4 GPU and DASK which doubles your GPU RAM for cudf and xgb.\n\nI will try this too. I am new to DASK and learning.\n\nThank you so much for your time.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2113019,
      "postDate": "2023-01-24T03:04:29.950Z",
      "content": "<p>I have a problem step6, Is ground truth the same as valid data B?</p>",
      "rawMarkdown": "I have a problem step6, Is ground truth the same as valid data B?",
      "votes": 1,
      "replies": [
        {
          "id": 2114277,
          "postDate": "2023-01-24T20:27:59.957Z",
          "content": "<p><a href=\"https://www.kaggle.com/shun222\" target=\"_blank\">@shun222</a> Yes they are the same</p>",
          "rawMarkdown": "@shun222 Yes they are the same"
        }
      ]
    },
    {
      "id": 2109701,
      "postDate": "2023-01-21T15:40:20.133Z",
      "content": "<p>Wow, amazing.Cool</p>",
      "rawMarkdown": "Wow, amazing.Cool",
      "votes": 1
    },
    {
      "id": 2104955,
      "postDate": "2023-01-18T06:59:59.187Z",
      "content": "<p>Hi, I'm confused that we use 4 week data (3 weeks trian plus 1 week valid A) but use 5 weeks data (4 kaggle train plus 1 week kaggle test). Wouldn't this cause some feature distribution to be inconsistent ? like <code>item_item_count</code></p>",
      "rawMarkdown": "Hi, I'm confused that we use 4 week data (3 weeks trian plus 1 week valid A) but use 5 weeks data (4 kaggle train plus 1 week kaggle test). Wouldn't this cause some feature distribution to be inconsistent ? like `item_item_count`",
      "votes": 1,
      "replies": [
        {
          "id": 2105927,
          "postDate": "2023-01-18T20:37:50.570Z",
          "content": "<p>Great point <a href=\"https://www.kaggle.com/earlee0412\" target=\"_blank\">@earlee0412</a> For certain features, you do need to be careful such as counts. During inference, you can use the last 3 weeks of train plus 1 week kaggle test for total of 4 weeks to count.</p>",
          "rawMarkdown": "Great point @earlee0412 For certain features, you do need to be careful such as counts. During inference, you can use the last 3 weeks of train plus 1 week kaggle test for total of 4 weeks to count.",
          "votes": 2,
          "replies": [
            {
              "id": 2106185,
              "postDate": "2023-01-19T02:25:54.287Z",
              "content": "<p>agree so and thanks so much</p>",
              "rawMarkdown": "agree so and thanks so much",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2104876,
      "postDate": "2023-01-18T05:45:46.587Z",
      "content": "<p>Hi Chris, thank you for your sharing.</p>\n<p>I have a question about Infelence session.<br>\nI couldn't understand about \"we select 20 by sorting …\".<br>\nI couldn't sort predicts becouse My \"12899779_clicks\" predict only one aid(\"59625\")'s pred.</p>\n<p>Please tell me more about Infelence session.</p>",
      "rawMarkdown": "Hi Chris, thank you for your sharing.\n\nI have a question about Infelence session.\nI couldn't understand about \"we select 20 by sorting ...\".\nI couldn't sort predicts becouse My \"12899779_clicks\" predict only one aid(\"59625\")'s pred.\n\nPlease tell me more about Infelence session.",
      "votes": 1,
      "replies": [
        {
          "id": 2104894,
          "postDate": "2023-01-18T06:03:54.583Z",
          "content": "<p>In step 1, we generate X number of candidates where <code>X&gt;20</code>. Then in the \"Inference\" step, for each user we will have <code>X</code> predictions per user. And we sort those <code>X</code> predictions and take the best 20.</p>",
          "rawMarkdown": "In step 1, we generate X number of candidates where `X>20`. Then in the \"Inference\" step, for each user we will have `X` predictions per user. And we sort those `X` predictions and take the best 20.",
          "replies": [
            {
              "id": 2104909,
              "postDate": "2023-01-18T06:13:52.443Z",
              "content": "<p>I　unerstand.<br>\nThank you fou your help.</p>",
              "rawMarkdown": "I　unerstand.\nThank you fou your help.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2103732,
      "postDate": "2023-01-17T10:42:09.137Z",
      "content": "<p>For training, should we add the positive samples, i.e., the pairs of (session, item, target=1) that are not in the set candidates?</p>",
      "rawMarkdown": "For training, should we add the positive samples, i.e., the pairs of (session, item, target=1) that are not in the set candidates?",
      "votes": 1,
      "replies": [
        {
          "id": 2109728,
          "postDate": "2023-01-21T15:54:09.390Z",
          "content": "<p>I used it    </p>",
          "rawMarkdown": "I used it    ",
          "votes": 1
        },
        {
          "id": 2109796,
          "postDate": "2023-01-21T17:07:46.157Z",
          "content": "<p>The best way to find out for your model is to try it and compare to CV and LB versus not trying it.</p>",
          "rawMarkdown": "The best way to find out for your model is to try it and compare to CV and LB versus not trying it.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2100160,
      "postDate": "2023-01-15T00:50:02.993Z",
      "content": "<p>How to create user-item interaction features? This is improtant for rec sys.</p>",
      "rawMarkdown": "How to create user-item interaction features? This is improtant for rec sys.",
      "votes": 1
    },
    {
      "id": 2097920,
      "postDate": "2023-01-13T04:39:39.927Z",
      "content": "<p>Very practical guide！Thanks for sharing! </p>",
      "rawMarkdown": "Very practical guide！Thanks for sharing! ",
      "votes": 1
    },
    {
      "id": 2096386,
      "postDate": "2023-01-12T02:30:32.010Z",
      "content": "<p>Hi Chris, I appreciate your willingness to share valuable ideas and codes. It really helps me to try participation to this competition.<br>\nWe would generate 3 candidates for each action (clicks, carts, and orders) using co-visitation matrix.<br>\n(Below is the code snippets you had shared the 3 handcrafted reranked candidates using both history data and suggested data) </p>\n<p>pred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(<br>\n    lambda x: suggest_clicks(x)<br>\n)</p>\n<p>pred_df_buys = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(<br>\n    lambda x: suggest_buys(x)<br>\n)<br>\nI wonder if I have to use the separate candidate with each action (e.g. pred_df_clicks, pred_df_buys, pred_df_carts )  to make features for 3 separate ML ranker models?</p>",
      "rawMarkdown": "Hi Chris, I appreciate your willingness to share valuable ideas and codes. It really helps me to try participation to this competition.\nWe would generate 3 candidates for each action (clicks, carts, and orders) using co-visitation matrix.\n(Below is the code snippets you had shared the 3 handcrafted reranked candidates using both history data and suggested data) \n\npred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_clicks(x)\n)\n\npred_df_buys = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_buys(x)\n)\nI wonder if I have to use the separate candidate with each action (e.g. pred_df_clicks, pred_df_buys, pred_df_carts )  to make features for 3 separate ML ranker models?",
      "votes": 1,
      "replies": [
        {
          "id": 2096395,
          "postDate": "2023-01-12T02:40:52.497Z",
          "content": "<p>For each target (click,cart,order), we create one dataframe of candidates. There are 1.8 million test users. So each dataframe has <code>X</code> times 1.8 million where <code>X</code> is the number of candidates per user. We can try different <code>X</code>.</p>\n<p>Then for each dataframe, we merge on the targets of 0 or 1 whether the candidate is ground truth or not. Then we train GBT ranker for each target. So we have 3 trained GBT models.</p>",
          "rawMarkdown": "For each target (click,cart,order), we create one dataframe of candidates. There are 1.8 million test users. So each dataframe has `X` times 1.8 million where `X` is the number of candidates per user. We can try different `X`.\n\nThen for each dataframe, we merge on the targets of 0 or 1 whether the candidate is ground truth or not. Then we train GBT ranker for each target. So we have 3 trained GBT models.",
          "votes": 3,
          "replies": [
            {
              "id": 2096557,
              "postDate": "2023-01-12T06:15:07.947Z",
              "content": "<p>Ok, I see then I'll try it. Thank you very much.</p>",
              "rawMarkdown": "Ok, I see then I'll try it. Thank you very much.",
              "votes": 1
            },
            {
              "id": 2104649,
              "postDate": "2023-01-18T00:09:31.850Z",
              "content": "<p>Now, I was stuck in training phase. <br>\nI crated a dataframe (e.g. clicks_df_for_train (1 GB)) in TPU-VM kaggle-platform (enough for RAM size)<br>\nI tried LGBM in the platfrom, worked without any probelm, but didn't have good score, so I wanted to try XGBoost  before going into more feature engineering step to improve CV score.<br>\nFor xgb, I changed into GPU-supported platform but I failed to load the dataframe for training due to memory size.<br>\nI would appreciate some tips or advice for further proceeding. </p>",
              "rawMarkdown": "Now, I was stuck in training phase. \nI crated a dataframe (e.g. clicks_df_for_train (1 GB)) in TPU-VM kaggle-platform (enough for RAM size)\nI tried LGBM in the platfrom, worked without any probelm, but didn't have good score, so I wanted to try XGBoost  before going into more feature engineering step to improve CV score.\nFor xgb, I changed into GPU-supported platform but I failed to load the dataframe for training due to memory size.\nI would appreciate some tips or advice for further proceeding. "
            }
          ]
        }
      ]
    },
    {
      "id": 2095593,
      "postDate": "2023-01-11T13:45:31.300Z",
      "content": "<p>Thank you for your sharing!</p>\n<p>In the phase <strong>Training</strong>,  the size of candidates is 1801251 * 50?  1801251 is the number of sessions in the validation data A.</p>\n<p>My question is whether a group without a positive sample may affect the performance of XGBRanker.</p>\n<p>In addition, should I make sure each group is of the same size when downsampling negatives?</p>\n<p>Looking forward to your reply~</p>",
      "rawMarkdown": "Thank you for your sharing!\n\nIn the phase **Training**,  the size of candidates is 1801251 * 50?  1801251 is the number of sessions in the validation data A.\n\nMy question is whether a group without a positive sample may affect the performance of XGBRanker.\n\nIn addition, should I make sure each group is of the same size when downsampling negatives?\n\nLooking forward to your reply~",
      "votes": 1,
      "replies": [
        {
          "id": 2096306,
          "postDate": "2023-01-12T00:40:59.273Z",
          "content": "<blockquote>\n  <p>In the phase Training, the size of candidates is 1801251 * 50? 1801251 is the number of sessions in the validation data A.</p>\n</blockquote>\n<p>Yes if you use 50 candidates. This is a suggestion but you can try other numbers</p>\n<blockquote>\n  <p>My question is whether a group without a positive sample may affect the performance of XGBRanker.</p>\n</blockquote>\n<p>Maybe. The best way to check is to try both and compare CV score and LB score.</p>\n<blockquote>\n  <p>In addition, should I make sure each group is of the same size when downsampling negatives?</p>\n</blockquote>\n<p>No. This is not required by GBT ranker models. Just make sure to give XGB the correct sizes of each group. Also make sure that the dataframe you train with has each user together as consecutive rows. For example, first <code>df = df.sort_values('user')</code> then <code>group_size = df.groupby('user').user.agg('count').values</code>.</p>",
          "rawMarkdown": ">In the phase Training, the size of candidates is 1801251 * 50? 1801251 is the number of sessions in the validation data A.\n\nYes if you use 50 candidates. This is a suggestion but you can try other numbers\n\n>My question is whether a group without a positive sample may affect the performance of XGBRanker.\n\nMaybe. The best way to check is to try both and compare CV score and LB score.\n\n>In addition, should I make sure each group is of the same size when downsampling negatives?\n\nNo. This is not required by GBT ranker models. Just make sure to give XGB the correct sizes of each group. Also make sure that the dataframe you train with has each user together as consecutive rows. For example, first `df = df.sort_values('user')` then `group_size = df.groupby('user').user.agg('count').values`.",
          "votes": 4,
          "replies": [
            {
              "id": 2096351,
              "postDate": "2023-01-12T01:28:12.580Z",
              "content": "<p>Many thanks</p>",
              "rawMarkdown": "Many thanks",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2086622,
      "postDate": "2023-01-04T21:42:50.223Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. I have a question about the co-visitation matrix.</p>\n<p>Should we create respective co-visitation matrix to extract candidates for each of train/valid/test ? For exemple, for train I'll use the candidates extracted from the train dataset, for validation I'll use the candidates extracted from the validation set ect …</p>",
      "rawMarkdown": "Thanks for sharing @cdeotte. I have a question about the co-visitation matrix.\n\nShould we create respective co-visitation matrix to extract candidates for each of train/valid/test ? For exemple, for train I'll use the candidates extracted from the train dataset, for validation I'll use the candidates extracted from the validation set ect ...",
      "votes": 1,
      "replies": [
        {
          "id": 2096308,
          "postDate": "2023-01-12T00:42:20.280Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rayanaay\" target=\"_blank\">@rayanaay</a> . We make one set of co-visitation matrices for train/validate. And a second set of co-visitation matrices for inference (submission.csv).</p>",
          "rawMarkdown": "Hi @rayanaay . We make one set of co-visitation matrices for train/validate. And a second set of co-visitation matrices for inference (submission.csv).",
          "votes": 1
        }
      ]
    },
    {
      "id": 2085940,
      "postDate": "2023-01-04T13:32:06.997Z",
      "content": "<p>Thanks for sharing very informative artical.</p>",
      "rawMarkdown": "Thanks for sharing very informative artical.",
      "votes": 1
    },
    {
      "id": 2084084,
      "postDate": "2023-01-03T06:39:23.180Z",
      "content": "<p>How can we get training labels for training set? For example, if aid 100 was clicked in session 100 at some timestamp in training set, then the target should be 1 for the pair of session 100 and aid 100?</p>",
      "rawMarkdown": "How can we get training labels for training set? For example, if aid 100 was clicked in session 100 at some timestamp in training set, then the target should be 1 for the pair of session 100 and aid 100?",
      "votes": 1,
      "replies": [
        {
          "id": 2084450,
          "postDate": "2023-01-03T14:01:36.083Z",
          "content": "<p>I suggest using Radek's Kaggle dataset explained here. Radek has converted kaggle train data into \"new train data\", \"valid A data\" and \"valid B data\" (where valid B are train labels that you ask about). These 3 are just like \"kaggle's train.csv\", \"kaggle's test.csv\", and \"kaggle leaderboard\". Radek explains how he got them in his discussion post.</p>",
          "rawMarkdown": "I suggest using Radek's Kaggle dataset explained here. Radek has converted kaggle train data into \"new train data\", \"valid A data\" and \"valid B data\" (where valid B are train labels that you ask about). These 3 are just like \"kaggle's train.csv\", \"kaggle's test.csv\", and \"kaggle leaderboard\". Radek explains how he got them in his discussion post.",
          "votes": 1,
          "replies": [
            {
              "id": 2084539,
              "postDate": "2023-01-03T15:26:43.367Z",
              "content": "<p>Thanks for your reply! I am new to recommendation. My understandings are as follows. <br>\n1) \"new train data\": it is used for the generation of an \"aid2aid\" matrix.<br>\n2) \"valid A data\": we generate 50 candidates for each session in valid A data by the \"aid2aid\" matrix. Some candidates are from ground truths in valid B data, so the \"gt\" is set as 1. Some candidates are not in valid B data, so the \"gt\" is set as 0. With many rows of (session, aid, some features, gt), then we can train a GBT ranker.<br>\n3) \"valid B data\": it is the ground truth label for valid A data and is used to compute local CV. <br>\nAre the above correct? </p>",
              "rawMarkdown": "Thanks for your reply! I am new to recommendation. My understandings are as follows. \n1) \"new train data\": it is used for the generation of an \"aid2aid\" matrix.\n2) \"valid A data\": we generate 50 candidates for each session in valid A data by the \"aid2aid\" matrix. Some candidates are from ground truths in valid B data, so the \"gt\" is set as 1. Some candidates are not in valid B data, so the \"gt\" is set as 0. With many rows of (session, aid, some features, gt), then we can train a GBT ranker.\n3) \"valid B data\": it is the ground truth label for valid A data and is used to compute local CV. \nAre the above correct? ",
              "votes": 2
            },
            {
              "id": 2084552,
              "postDate": "2023-01-03T15:37:02.517Z",
              "content": "<p><a href=\"https://www.kaggle.com/nin7a1\" target=\"_blank\">@nin7a1</a> that is correct. However we can use \"test leakage\" in this competition, read discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939\" target=\"_blank\">here</a>. So in your \"#1\", you can use \"new train data\" plus \"valid A\" to generate \"aid2aid\" matrix. Yes, it will use leak from future data, (but we can replicate this when we make a submission to leaderboard because we can also use kaggle's train.csv and kaggle's test.csv to make \"aid2aid\" for submission)</p>",
              "rawMarkdown": "@nin7a1 that is correct. However we can use \"test leakage\" in this competition, read discussion [here][1]. So in your \"#1\", you can use \"new train data\" plus \"valid A\" to generate \"aid2aid\" matrix. Yes, it will use leak from future data, (but we can replicate this when we make a submission to leaderboard because we can also use kaggle's train.csv and kaggle's test.csv to make \"aid2aid\" for submission)\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939",
              "votes": 1
            },
            {
              "id": 2084558,
              "postDate": "2023-01-03T15:41:19.450Z",
              "content": "<p>Thank you!</p>",
              "rawMarkdown": "Thank you!",
              "votes": 1
            },
            {
              "id": 2086666,
              "postDate": "2023-01-04T23:16:41.723Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nI have a question about the training dataset ( the first 3 weeks ). Do you use this data only to generate features, or do you train your model on it too ?</p>\n<p>I'm confused because of this : <code>(where valid B are train labels that you ask about)</code> in your answer. If I well understood the pipeline:</p>\n<ul>\n<li>Valid A are the rows to consider ( input X for the ranker ).</li>\n<li>Valid B are the labels ( Y ).</li>\n<li>Whereas the first 3 weeks of data are used to extract features about the items.</li>\n</ul>\n<p>In summary, we are only using one single week to predict the labels for the test data</p>",
              "rawMarkdown": "@cdeotte \nI have a question about the training dataset ( the first 3 weeks ). Do you use this data only to generate features, or do you train your model on it too ?\n\nI'm confused because of this : `(where valid B are train labels that you ask about)` in your answer. If I well understood the pipeline:\n-  Valid A are the rows to consider ( input X for the ranker ).\n-  Valid B are the labels ( Y ).\n- Whereas the first 3 weeks of data are used to extract features about the items.\n\nIn summary, we are only using one single week to predict the labels for the test data",
              "votes": 1
            },
            {
              "id": 3076913,
              "postDate": "2024-12-20T11:48:56.903Z",
              "content": "<p>Good question <a href=\"https://www.kaggle.com/rayanaay\" target=\"_blank\">@rayanaay</a> , would like to know the answer to this if you figured it out since then :)<br>\nWould be great if this is clarified <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 🙏</p>",
              "rawMarkdown": "Good question @rayanaay , would like to know the answer to this if you figured it out since then :)\nWould be great if this is clarified @cdeotte 🙏"
            }
          ]
        }
      ]
    },
    {
      "id": 2083547,
      "postDate": "2023-01-02T15:48:00.560Z",
      "content": "<p>Hi Chris, its very impressive to me.i konw in the \"Inference\" stage you used kaggle's  train and test data to build a submission dataset.but i have one question to enquire you that how i built native CV to konw the score in my notebook like in your shared Code \"Compute Validation Score - [CV 565]\".  thank you about the sharing again!👍</p>",
      "rawMarkdown": "Hi Chris, its very impressive to me.i konw in the \"Inference\" stage you used kaggle's  train and test data to build a submission dataset.but i have one question to enquire you that how i built native CV to konw the score in my notebook like in your shared Code \"Compute Validation Score - [CV 565]\".  thank you about the sharing again!👍",
      "votes": 1,
      "replies": [
        {
          "id": 2083571,
          "postDate": "2023-01-02T16:12:42.703Z",
          "content": "<p>This is why we use 5fold GroupKFold. After building the train, valid A, and valid B datasets. Then we generate candidates and features. Then we train and infer our 5fold. This will give us OOF predictions for all the valid B data. We can then compute the local CV score comparing our OOF predictions with the valid B ground truth.</p>",
          "rawMarkdown": "This is why we use 5fold GroupKFold. After building the train, valid A, and valid B datasets. Then we generate candidates and features. Then we train and infer our 5fold. This will give us OOF predictions for all the valid B data. We can then compute the local CV score comparing our OOF predictions with the valid B ground truth.",
          "replies": [
            {
              "id": 2084063,
              "postDate": "2023-01-03T05:53:10.313Z",
              "content": "<p>thank you it's very helpful for me.but i have a new question : in your discussion a 'candidate' user contains 50 items, i want to ask that in the 50 items all present the 'clicks' not contains 'carts' and 'orders'? (the picture), if in 'clicks','carts' and 'orders' that one user will have 150 'items'. and other question : your discussion is rerank 'click' acton , but why you only mark 'carts' as type '1' rather than 'carts' and 'buys' as type '1' , i am curious about it , thank you again~✨<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12934577%2F253e4847908ff0372001d58afe023918%2FXnip2023-01-03_13-42-04.jpg?generation=1672725181023917&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "thank you it's very helpful for me.but i have a new question : in your discussion a 'candidate' user contains 50 items, i want to ask that in the 50 items all present the 'clicks' not contains 'carts' and 'orders'? (the picture), if in 'clicks','carts' and 'orders' that one user will have 150 'items'. and other question : your discussion is rerank 'click' acton , but why you only mark 'carts' as type '1' rather than 'carts' and 'buys' as type '1' , i am curious about it , thank you again~✨![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12934577%2F253e4847908ff0372001d58afe023918%2FXnip2023-01-03_13-42-04.jpg?generation=1672725181023917&alt=media)",
              "votes": 1
            },
            {
              "id": 2084452,
              "postDate": "2023-01-03T14:04:06.800Z",
              "content": "<p>We mark carts as type 1 because this is only one example. We build 3 candidate dataframes one for each clicks, carts, and orders. Each only has 50 candidates per user and each has candidates appropriate to the respective target. In step 6, we add the respective target to the respective dataframe. And we train and infer each separately by making 3 separate GBT ranker models. At last we concatenate all the predictions into a single submission.csv.</p>",
              "rawMarkdown": "We mark carts as type 1 because this is only one example. We build 3 candidate dataframes one for each clicks, carts, and orders. Each only has 50 candidates per user and each has candidates appropriate to the respective target. In step 6, we add the respective target to the respective dataframe. And we train and infer each separately by making 3 separate GBT ranker models. At last we concatenate all the predictions into a single submission.csv.",
              "votes": 2
            },
            {
              "id": 2084489,
              "postDate": "2023-01-03T14:47:12.900Z",
              "content": "<p>thank you!!!! I understand~</p>",
              "rawMarkdown": "thank you!!!! I understand~",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2082598,
      "postDate": "2023-01-01T17:58:26.140Z",
      "content": "<p>Why don't you use LightGBM ? is XGBoost more suited for this task ?<br>\nFrom my own experience, lightgbm uses less memory, isn't ? </p>",
      "rawMarkdown": "Why don't you use LightGBM ? is XGBoost more suited for this task ?\nFrom my own experience, lightgbm uses less memory, isn't ? ",
      "votes": 1,
      "replies": [
        {
          "id": 2082728,
          "postDate": "2023-01-01T21:07:20.913Z",
          "content": "<p>You can try XGBoost, LGBM, and CatBoost. They can all do ranker models. Its best to try all 3 and see which produces the best CV and LB.</p>",
          "rawMarkdown": "You can try XGBoost, LGBM, and CatBoost. They can all do ranker models. Its best to try all 3 and see which produces the best CV and LB.",
          "votes": 1,
          "replies": [
            {
              "id": 2082743,
              "postDate": "2023-01-01T21:36:47.127Z",
              "content": "<p>Great ! for the moment, I consider the problem as a binary problem, using  AUC as a metric. However, i'm getting 0.7 of AUC with I think, good features extracted from a graph neural network and aggregated features. Plus, when extracting the top 20 using the GBDT groupby each session, I'm getting a validation score on clicks of 0.400, which is very bad compared to the 0.600 of the top50 candidates extracted using the co-visitation-matrix. What do you think of that and what could be the source of the problem ? <br>\nI'm training 2 LGBM Classifier on the train dataset ( 2 chunks ) </p>",
              "rawMarkdown": "Great ! for the moment, I consider the problem as a binary problem, using  AUC as a metric. However, i'm getting 0.7 of AUC with I think, good features extracted from a graph neural network and aggregated features. Plus, when extracting the top 20 using the GBDT groupby each session, I'm getting a validation score on clicks of 0.400, which is very bad compared to the 0.600 of the top50 candidates extracted using the co-visitation-matrix. What do you think of that and what could be the source of the problem ? \nI'm training 2 LGBM Classifier on the train dataset ( 2 chunks ) ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2075850,
      "postDate": "2022-12-26T01:02:52.860Z",
      "content": "<p>Thanks for sharing very informative artical.</p>",
      "rawMarkdown": "Thanks for sharing very informative artical.",
      "votes": 1
    },
    {
      "id": 2075059,
      "postDate": "2022-12-25T02:14:50.993Z",
      "content": "<p>Where do you get user features and item features? It is not there in given dataset?</p>",
      "rawMarkdown": "Where do you get user features and item features? It is not there in given dataset?",
      "votes": 1,
      "replies": [
        {
          "id": 2075069,
          "postDate": "2022-12-25T03:00:48.480Z",
          "content": "<p>No it is not. The dataset from Kaggle only has <code>user, item, time stamp, event type</code> (they rename them differently but this is what they are). We must create features from these columns.</p>",
          "rawMarkdown": "No it is not. The dataset from Kaggle only has `user, item, time stamp, event type` (they rename them differently but this is what they are). We must create features from these columns."
        }
      ]
    },
    {
      "id": 2074187,
      "postDate": "2022-12-23T20:29:44.917Z",
      "content": "<p>Thanks for the post and all clarifications. </p>\n<p>After reading your post many times I conclude that validation A is more like a 'last_week_train' dataframe than validation for ml model. Because we use validation A to create the user session features and we don't use validation A as validation as we use in a holdout. Because you are using cross-validation kfold to train your data. <br>\nAlso, validation B is just used to create the ground truth, am I correct? </p>\n<p>EDIT: </p>\n<p>Also, we consider we don't have the same sessions in the test set from training. So, why do we break the validation in two parts (A and B) from the same user?</p>\n<p>Thank you,</p>",
      "rawMarkdown": "Thanks for the post and all clarifications. \n\nAfter reading your post many times I conclude that validation A is more like a 'last_week_train' dataframe than validation for ml model. Because we use validation A to create the user session features and we don't use validation A as validation as we use in a holdout. Because you are using cross-validation kfold to train your data. \nAlso, validation B is just used to create the ground truth, am I correct? \n\nEDIT: \n\nAlso, we consider we don't have the same sessions in the test set from training. So, why do we break the validation in two parts (A and B) from the same user?\n\nThank you,",
      "votes": 1,
      "replies": [
        {
          "id": 2074218,
          "postDate": "2022-12-23T20:53:20.180Z",
          "content": "<blockquote>\n  <p>After reading your post many times I conclude that validation A is more like a 'last_week_train' dataframe than validation for ml model.</p>\n</blockquote>\n<p>Perhaps, \"last_week_train\" would be a better name but note that it is more than <strong>normal</strong> train data. I used \"val A\" because it mimics \"test.csv\" that Kaggle provides. Using \"val A\" is actually a <strong>leak</strong> because some events in \"val A\" occur after events in \"val B\". Normally we can never use information from the future to predict the past.</p>\n<p>In normal competitions, we would not need <code>new train</code>, <code>val A</code> and <code>val B</code>. But in this comp there is a leak and Kaggle says we can use the leak. When we make a submission Kaggle gives us 3 dataframes. When we make a submission, Kaggle has <code>train.csv</code>, <code>test.csv</code> and <code>public LB</code>. Those 3 sets are different. The purpose of local validation is to mimic the leaderboard. So we must copy this setup (including the leak). Since Kaggle has 3 sets, we make 3 sets. Then we have the following comparison</p>\n<ul>\n<li>\"new train\" is like \"kaggle train\"</li>\n<li>\"val A\" is like \"kaggle test\"</li>\n<li>\"val B\" is like \"kaggle leaderboard\"</li>\n</ul>\n<p>We cannot use any information from \"val B\" because just like submission, we cannot access information from public leaderboard. In our local setup, we <strong>only</strong> use \"val B\" to compute our local validation score, we <strong>do not</strong> take any information from it. During submission, we can access kaggle train and kaggle test, so during validation we can also access \"new train\" and \"val A\".</p>",
          "rawMarkdown": ">After reading your post many times I conclude that validation A is more like a 'last_week_train' dataframe than validation for ml model.\n\nPerhaps, \"last_week_train\" would be a better name but note that it is more than **normal** train data. I used \"val A\" because it mimics \"test.csv\" that Kaggle provides. Using \"val A\" is actually a **leak** because some events in \"val A\" occur after events in \"val B\". Normally we can never use information from the future to predict the past.\n\nIn normal competitions, we would not need `new train`, `val A` and `val B`. But in this comp there is a leak and Kaggle says we can use the leak. When we make a submission Kaggle gives us 3 dataframes. When we make a submission, Kaggle has `train.csv`, `test.csv` and `public LB`. Those 3 sets are different. The purpose of local validation is to mimic the leaderboard. So we must copy this setup (including the leak). Since Kaggle has 3 sets, we make 3 sets. Then we have the following comparison\n* \"new train\" is like \"kaggle train\"\n* \"val A\" is like \"kaggle test\"\n* \"val B\" is like \"kaggle leaderboard\"\n\nWe cannot use any information from \"val B\" because just like submission, we cannot access information from public leaderboard. In our local setup, we **only** use \"val B\" to compute our local validation score, we **do not** take any information from it. During submission, we can access kaggle train and kaggle test, so during validation we can also access \"new train\" and \"val A\".",
          "votes": 2
        }
      ]
    },
    {
      "id": 2068857,
      "postDate": "2022-12-18T11:22:40.660Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, Thanks for the idea. I am new to this competition and i have some doubts, can you please explain them:<br>\n1) after splitting your training data into train +(val A,val B) you are generating 50 candidates only for val A right?<br>\n2) does that mean you are using your validation data A, val data B to train xgb ranker and training data only to create features for items?</p>",
      "rawMarkdown": "Hi @cdeotte, Thanks for the idea. I am new to this competition and i have some doubts, can you please explain them:\n1) after splitting your training data into train +(val A,val B) you are generating 50 candidates only for val A right?\n2) does that mean you are using your validation data A, val data B to train xgb ranker and training data only to create features for items?\n",
      "votes": 1,
      "replies": [
        {
          "id": 2069044,
          "postDate": "2022-12-18T14:21:09.447Z",
          "content": "<p>1) yes and note that val A and val B have the same users. So we are generating 50 candidates for both val A and val B).</p>\n<p>2) yes and we can use val A to create features for items too. (and for advanced Kagglers, we can use <strong>unsupervised</strong> methods to create user features from train).</p>",
          "rawMarkdown": "1) yes and note that val A and val B have the same users. So we are generating 50 candidates for both val A and val B).\n\n2) yes and we can use val A to create features for items too. (and for advanced Kagglers, we can use **unsupervised** methods to create user features from train).",
          "votes": 2,
          "replies": [
            {
              "id": 2072676,
              "postDate": "2022-12-22T10:39:26.787Z",
              "content": "<p>Hi Chris, I have one more doubt, are you training for each target('click', 'cart', 'buy') independently?<br>\nRegards<br>\nPriyanshu</p>",
              "rawMarkdown": "Hi Chris, I have one more doubt, are you training for each target('click', 'cart', 'buy') independently?\nRegards\nPriyanshu"
            },
            {
              "id": 2072894,
              "postDate": "2022-12-22T13:54:21.333Z",
              "content": "<p>Yes i train 3 independently. I make 3 dataframes. One dataframe is candidates for clicks, one for carts, one for orders. Then i train 3 XGB rankers. One takes click dataframe and predicts clicks. One takes carts and predicts carts. And the last takes orders and predicts orders.</p>",
              "rawMarkdown": "Yes i train 3 independently. I make 3 dataframes. One dataframe is candidates for clicks, one for carts, one for orders. Then i train 3 XGB rankers. One takes click dataframe and predicts clicks. One takes carts and predicts carts. And the last takes orders and predicts orders.",
              "votes": 2
            },
            {
              "id": 2076392,
              "postDate": "2022-12-26T13:07:49.110Z",
              "content": "<p>Thanks a lot for help <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. I have made a similar pipeline as you have described. My CV results are as follows:</p>\n<p>Clicks,    ndcg@20= .7xxxx          recall=  .59<br>\ncarts ,     ndcg@20= .9xxx           recall= .52<br>\nare you also getting same results?<br>\nregards<br>\nI am using  dividing val A data into 5 group folds training on 4 folds and testing on 1 fold<br>\nplease correct me if I  am wrong.<br>\nRegards<br>\nPriyanshu</p>",
              "rawMarkdown": "Thanks a lot for help @cdeotte. I have made a similar pipeline as you have described. My CV results are as follows:\n                 \nClicks,    ndcg@20= .7xxxx          recall=  .59\ncarts ,     ndcg@20= .9xxx           recall= .52\nare you also getting same results?\nregards\nI am using  dividing val A data into 5 group folds training on 4 folds and testing on 1 fold\nplease correct me if I  am wrong.\nRegards\nPriyanshu",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2066732,
      "postDate": "2022-12-16T02:50:26.957Z",
      "content": "<p>I love this</p>",
      "rawMarkdown": "I love this",
      "votes": 1
    },
    {
      "id": 2066098,
      "postDate": "2022-12-15T12:07:27.417Z",
      "content": "<p>Incredible, simple, and fast forward explanation !  that is what I was looking for :).<br>\nI have a question about the candidates generation. Let's say we use 3 methods for candidates generation, ( with the co-visitation matrix for example ), let's say 2 of them predict a good rec, while the last one predict a false negative, what do you think about weighting this true positive recommended item \"twice\", using the sample_weight parameter of LightGBM ?</p>",
      "rawMarkdown": "Incredible, simple, and fast forward explanation !  that is what I was looking for :).\nI have a question about the candidates generation. Let's say we use 3 methods for candidates generation, ( with the co-visitation matrix for example ), let's say 2 of them predict a good rec, while the last one predict a false negative, what do you think about weighting this true positive recommended item \"twice\", using the sample_weight parameter of LightGBM ?",
      "votes": 1
    },
    {
      "id": 2064603,
      "postDate": "2022-12-14T00:00:43.620Z",
      "content": "<p>Wow, this cleared up a lot! Thanks, Chris!</p>",
      "rawMarkdown": "Wow, this cleared up a lot! Thanks, Chris!",
      "votes": 1,
      "replies": [
        {
          "id": 2065595,
          "postDate": "2022-12-14T21:36:08.413Z",
          "content": "<p>Also, correct me if I'm wrong, but in step 2, is <code>train.groupby('session').agg({'session':'count','aid','nunique','type','mean'})</code> supposed to be <code>train.groupby('session').agg({'session':'count','aid':'nunique','type':'mean'})</code>?</p>",
          "rawMarkdown": "Also, correct me if I'm wrong, but in step 2, is `train.groupby('session').agg({'session':'count','aid','nunique','type','mean'})` supposed to be `train.groupby('session').agg({'session':'count','aid':'nunique','type':'mean'})`?",
          "votes": 1
        },
        {
          "id": 2065602,
          "postDate": "2022-12-14T21:50:12.463Z",
          "content": "<p>yes <a href=\"https://www.kaggle.com/alberteinsten\" target=\"_blank\">@alberteinsten</a> you are correct. Thank you, i updated my post</p>",
          "rawMarkdown": "yes @alberteinsten you are correct. Thank you, i updated my post",
          "votes": 1
        }
      ]
    },
    {
      "id": 2063769,
      "postDate": "2022-12-13T09:53:50.467Z",
      "content": "<p>Hi, I want to know whether the split chunk code is wrong. It could be written like this.</p>\n<pre><code>CHUNKS = 10\nchunk_size = ceil(len(candidates) / CHUNKS)\nfor k in range(CHUNKS):\n    df = candidates.iloc[k*chunk_size:(k+1)*chunk_size].copy()\n    df = df.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\n    df = df.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\n    df.to_parquet(f'candidate_with_features_p{k}.pqt')\n</code></pre>\n<p>Please correct me if I'm wrong, thanks.</p>",
      "rawMarkdown": "Hi, I want to know whether the split chunk code is wrong. It could be written like this.\n```\nCHUNKS = 10\nchunk_size = ceil(len(candidates) / CHUNKS)\nfor k in range(CHUNKS):\n    df = candidates.iloc[k*chunk_size:(k+1)*chunk_size].copy()\n    df = df.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\n    df = df.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\n    df.to_parquet(f'candidate_with_features_p{k}.pqt')\n```\nPlease correct me if I'm wrong, thanks.",
      "votes": 1,
      "replies": [
        {
          "id": 2064122,
          "postDate": "2022-12-13T14:30:44.733Z",
          "content": "<p>Yes, you are correct <a href=\"https://www.kaggle.com/jonneryr\" target=\"_blank\">@jonneryr</a> . Thanks for the correction. I updated my post</p>",
          "rawMarkdown": "Yes, you are correct @jonneryr . Thanks for the correction. I updated my post"
        }
      ]
    },
    {
      "id": 2063714,
      "postDate": "2022-12-13T08:25:19.977Z",
      "content": "<p>O, thats great!</p>",
      "rawMarkdown": "O, thats great!",
      "votes": 1
    },
    {
      "id": 2062366,
      "postDate": "2022-12-12T03:41:36.383Z",
      "content": "<p>Using Xgboost like gods using his grace. Holy f**k. </p>",
      "rawMarkdown": "Using Xgboost like gods using his grace. Holy f**k. ",
      "votes": 1
    },
    {
      "id": 2061563,
      "postDate": "2022-12-11T09:02:20.353Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, could you please elaborate <code>Step 4</code> like I'm five. I found it hard to understand to beginner like me. Thank you very much. </p>",
      "rawMarkdown": "Hi @cdeotte, could you please elaborate `Step 4` like I'm five. I found it hard to understand to beginner like me. Thank you very much. ",
      "votes": 1,
      "replies": [
        {
          "id": 2061877,
          "postDate": "2022-12-11T14:42:49Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/locbaop\" target=\"_blank\">@locbaop</a> Note that step 4 is optional. The GBT ranker model will work without step 4. So you should first build a successful GBT ranker without step 4, then add step 4. (and adding step 4 will improve CV and LB score)</p>",
          "rawMarkdown": "Hi @locbaop Note that step 4 is optional. The GBT ranker model will work without step 4. So you should first build a successful GBT ranker without step 4, then add step 4. (and adding step 4 will improve CV and LB score)\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2061103,
      "postDate": "2022-12-10T17:31:22.293Z",
      "content": "<p>great work bro</p>",
      "rawMarkdown": "great work bro",
      "votes": 1
    },
    {
      "id": 2060946,
      "postDate": "2022-12-10T15:25:13.117Z",
      "content": "<p>this is awesome! thanks for your explanation </p>",
      "rawMarkdown": "this is awesome! thanks for your explanation ",
      "votes": 1
    },
    {
      "id": 2060639,
      "postDate": "2022-12-10T07:51:52.007Z",
      "content": "<p>OMG!OMG! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> You are really a professional and meticulous person.👍</p>",
      "rawMarkdown": "OMG!OMG! @cdeotte You are really a professional and meticulous person.👍",
      "votes": 1
    },
    {
      "id": 2060470,
      "postDate": "2022-12-10T01:12:19.603Z",
      "content": "<p>This is awesome, thanks very much for the tips!</p>",
      "rawMarkdown": "This is awesome, thanks very much for the tips!",
      "votes": 1
    },
    {
      "id": 2055947,
      "postDate": "2022-12-05T15:20:14.797Z",
      "content": "<p>Thanks for sharing! <br>\nCurious about interaction features, <br>\nwould it make sense to add <code>item-context interaction</code> features as well? context here could be hour of the day, or day of week?<br>\nI have added the user features &amp; user-item interaction features to my GBT, but somehow still not performing better than heuristic</p>",
      "rawMarkdown": "Thanks for sharing! \nCurious about interaction features, \nwould it make sense to add `item-context interaction` features as well? context here could be hour of the day, or day of week?\nI have added the user features & user-item interaction features to my GBT, but somehow still not performing better than heuristic",
      "votes": 1,
      "replies": [
        {
          "id": 2055959,
          "postDate": "2022-12-05T15:33:30.147Z",
          "content": "<p>Yes we should add all types of features. Your \"item-context features\" are what i call \"item features\". I use the term \"item feature\" for any feature that is <strong>same</strong> for all users. Here is \"average hour\" that an item gets engaged with:</p>\n<pre><code>df['hour'] = (df.ts % (60*60*24*1000) ).astype('int32')\nitem_feature = df.groupby('aid').hour.agg('mean')\n</code></pre>\n<p>I call a feature \"item-user feature\" if the feature is different for each user. For example \"did user click this item?\" is a boolean feature that is different for each user given the same item.</p>",
          "rawMarkdown": "Yes we should add all types of features. Your \"item-context features\" are what i call \"item features\". I use the term \"item feature\" for any feature that is **same** for all users. Here is \"average hour\" that an item gets engaged with:\n\n    df['hour'] = (df.ts % (60*60*24*1000) ).astype('int32')\n    item_feature = df.groupby('aid').hour.agg('mean')\n\nI call a feature \"item-user feature\" if the feature is different for each user. For example \"did user click this item?\" is a boolean feature that is different for each user given the same item.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2054607,
      "postDate": "2022-12-04T09:27:26.677Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> your wonderful post has answered many of questions which are burning in my heart.  Thank you so much!</p>\n<p>In the sections about Memory management and Speed, I would like to know your opinion on pandas, cuDF, Dask vs Polars the image below. Also, does the image suggest Polars are faster than cuDF with GPU?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F155737f233c1ed896da44c6c5215a273%2FScreen%20Shot%202022-12-04%20at%2017.17.05.png?generation=1670145727436928&amp;alt=media\" alt=\"\"></p>\n<p>If using GPU, do you recommend it's better to use cuDF over other libraries? Thanks!</p>",
      "rawMarkdown": "Hi @cdeotte your wonderful post has answered many of questions which are burning in my heart.  Thank you so much!\n\nIn the sections about Memory management and Speed, I would like to know your opinion on pandas, cuDF, Dask vs Polars the image below. Also, does the image suggest Polars are faster than cuDF with GPU?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F155737f233c1ed896da44c6c5215a273%2FScreen%20Shot%202022-12-04%20at%2017.17.05.png?generation=1670145727436928&alt=media)\n\nIf using GPU, do you recommend it's better to use cuDF over other libraries? Thanks!",
      "votes": 1,
      "replies": [
        {
          "id": 2054777,
          "postDate": "2022-12-04T12:58:52.830Z",
          "content": "<p>Thanks for the list of libraries. I have not used all these libraries, so I can't comment which is better. Currently, I use cuDF for GPU and Pandas for CPU and DASK (and/or process in chunks) when i need to avoid memory error.</p>\n<p>I will explore more of these libraries. Thanks for the list.</p>",
          "rawMarkdown": "Thanks for the list of libraries. I have not used all these libraries, so I can't comment which is better. Currently, I use cuDF for GPU and Pandas for CPU and DASK (and/or process in chunks) when i need to avoid memory error.\n\nI will explore more of these libraries. Thanks for the list.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2054587,
      "postDate": "2022-12-04T09:06:29.497Z",
      "content": "<p>Another super amazing post! So much gold here! Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> so much for sharing 🙏❤️❤️❤️</p>",
      "rawMarkdown": "Another super amazing post! So much gold here! Thank you @cdeotte so much for sharing 🙏❤️❤️❤️",
      "votes": 1
    },
    {
      "id": 2053913,
      "postDate": "2022-12-03T18:01:00.507Z",
      "content": "<p>Many thanks for sharing !!! I learn so many things from your thread.  </p>",
      "rawMarkdown": "Many thanks for sharing !!! I learn so many things from your thread.  ",
      "votes": 1
    },
    {
      "id": 2053873,
      "postDate": "2022-12-03T17:12:56.060Z",
      "content": "<p>thanks for the wonderful thread, Chris. If you want to reveal, I am curious how much additional boost you got by ranking using GBDTs ?</p>",
      "rawMarkdown": "thanks for the wonderful thread, Chris. If you want to reveal, I am curious how much additional boost you got by ranking using GBDTs ?",
      "votes": 1
    },
    {
      "id": 2123828,
      "postDate": "2023-01-31T17:20:50.383Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I followed these steps to get my first version of all components necessary to produce a solution. Then I worked on improving separately:</p>\n<ol>\n<li>The features</li>\n<li>the candidate selection</li>\n<li>The ranker model (XGB + LGBM)</li>\n</ol>\n<p>I don't think I would have been able to build the whole pipeline without you.</p>",
      "rawMarkdown": "Thank you @cdeotte I followed these steps to get my first version of all components necessary to produce a solution. Then I worked on improving separately:\n\n1. The features\n2. the candidate selection\n3. The ranker model (XGB + LGBM)\n\nI don't think I would have been able to build the whole pipeline without you.",
      "votes": 2,
      "replies": [
        {
          "id": 2123858,
          "postDate": "2023-01-31T17:33:43.217Z",
          "content": "<p>Great job Elias</p>",
          "rawMarkdown": "Great job Elias",
          "votes": 2
        }
      ]
    },
    {
      "id": 2120892,
      "postDate": "2023-01-29T22:20:06.340Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  thank you for your valuable information.</p>\n<p>I'm struggling with the ranker for a long time, and I suspect my interactions features.</p>\n<p>Is it ok if my interactions features has &gt; 95% of NaNs ? from the feature importance of my GBDT, they're the most valuable variables, and I also beat the benchmark with a boost of 0.02, however my LB is still 0.565 with the ranker predictions.</p>\n<p>What do you think ?</p>",
      "rawMarkdown": "Hi @cdeotte  thank you for your valuable information.\n\nI'm struggling with the ranker for a long time, and I suspect my interactions features.\n\nIs it ok if my interactions features has > 95% of NaNs ? from the feature importance of my GBDT, they're the most valuable variables, and I also beat the benchmark with a boost of 0.02, however my LB is still 0.565 with the ranker predictions.\n\nWhat do you think ?",
      "votes": 2,
      "replies": [
        {
          "id": 2120903,
          "postDate": "2023-01-29T22:35:32.960Z",
          "content": "<p>If your NANs contain information then they are not really \"NANs\". For example, consider the interaction feature \"did user click/cart/order this item\". We can put a 1 when they do and a NAN when they don't. In this case, the NAN is actually a zero. And we really don't have any NANs. So many times NANs can actually contain information that your model will use.</p>",
          "rawMarkdown": "If your NANs contain information then they are not really \"NANs\". For example, consider the interaction feature \"did user click/cart/order this item\". We can put a 1 when they do and a NAN when they don't. In this case, the NAN is actually a zero. And we really don't have any NANs. So many times NANs can actually contain information that your model will use.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2111789,
      "postDate": "2023-01-23T07:19:24.397Z",
      "content": "<p>Hi Chris, I have a problem in setp1 and step 6, do we need to use different data for train and inference to create candidates ?</p>",
      "rawMarkdown": "Hi Chris, I have a problem in setp1 and step 6, do we need to use different data for train and inference to create candidates ?",
      "votes": 2,
      "replies": [
        {
          "id": 2111943,
          "postDate": "2023-01-23T10:08:31.377Z",
          "content": "<p>Yes. For train/validate we use \"new train\" plus \"valid data A\". And for inference we use kaggle's train.csv and test.csv.</p>",
          "rawMarkdown": "Yes. For train/validate we use \"new train\" plus \"valid data A\". And for inference we use kaggle's train.csv and test.csv.",
          "votes": 1,
          "replies": [
            {
              "id": 2111967,
              "postDate": "2023-01-23T10:21:03.097Z",
              "content": "<p>Thanks for your help !</p>",
              "rawMarkdown": "Thanks for your help !",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2085979,
      "postDate": "2023-01-04T13:44:19.230Z",
      "content": "<p>Hi Chris, its very impressive and helpful for us to learn. You are always the most value player in kaggle's competitions. Thank you very much!👍</p>",
      "rawMarkdown": "Hi Chris, its very impressive and helpful for us to learn. You are always the most value player in kaggle's competitions. Thank you very much!👍",
      "votes": 2,
      "replies": [
        {
          "id": 2096309,
          "postDate": "2023-01-12T00:42:35.477Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/songqizhou\" target=\"_blank\">@songqizhou</a> </p>",
          "rawMarkdown": "Thanks @songqizhou "
        }
      ]
    },
    {
      "id": 2084936,
      "postDate": "2023-01-03T21:01:46.460Z",
      "content": "<p>hi chris, your posts are always so helpful! And I wonder if I want to use co-visit matrix to get features, could I use the one generated with all data(like you shared) or I should only use the one using data from train + validation A. </p>",
      "rawMarkdown": "hi chris, your posts are always so helpful! And I wonder if I want to use co-visit matrix to get features, could I use the one generated with all data(like you shared) or I should only use the one using data from train + validation A. ",
      "votes": 2,
      "replies": [
        {
          "id": 2084983,
          "postDate": "2023-01-03T21:29:47.520Z",
          "content": "<p>Hi. You need to make both. You use the co-visitation matrix from \"new train\" + \"valid A\" to train your model. Then you use your second co-visitation matrix from \"kaggle train.csv\" + \"kaggle test.csv\" for infer.</p>\n<p>For example, my CV notebook <a href=\"https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565\" target=\"_blank\">here</a> uses \"new train\" + \"valid A\" to make co-visitation. And my submission notebook <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">here</a> uses \"kaggle train.csv\" + \"kaggle test.csv\" to make co-visitation. If this was a GBT reranker instead of heuristic rules. Then we would use my CV notebook to train the GBT reranker model. And then use my submission notebook to infer it.</p>",
          "rawMarkdown": "Hi. You need to make both. You use the co-visitation matrix from \"new train\" + \"valid A\" to train your model. Then you use your second co-visitation matrix from \"kaggle train.csv\" + \"kaggle test.csv\" for infer.\n\nFor example, my CV notebook [here][1] uses \"new train\" + \"valid A\" to make co-visitation. And my submission notebook [here][2] uses \"kaggle train.csv\" + \"kaggle test.csv\" to make co-visitation. If this was a GBT reranker instead of heuristic rules. Then we would use my CV notebook to train the GBT reranker model. And then use my submission notebook to infer it.\n\n[1]: https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565\n[2]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575",
          "votes": 3,
          "replies": [
            {
              "id": 2085200,
              "postDate": "2023-01-04T01:28:02.207Z",
              "content": "<p>very insightful, thank you!</p>",
              "rawMarkdown": "very insightful, thank you!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2084505,
      "postDate": "2023-01-03T15:02:41.387Z",
      "content": "<p>![Hi Chris you are truly teacher to me in this competition , and I quite have progress in the rerank part.. but I am facing a new problem about 'group'  in the modeling training part. when I use 'group' in your notebook , the error have appear that said \"num_row_ (4375050 vs. 4375063) : Invalid group structure\"(show at the picture blow)  ,but I cancel the 'group' it's OK to train my model. have you met this problem? how you fix up?and can you tell me the meaning of  'group' parameter and its important or not to the model training…thank you!! you answer make me grow up every day in this field ~~!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12934577%2F27c220232d0643fdf9eddd2f4ff5c291%2FXnip2023-01-03_22-46-09.jpg?generation=1672758123273487&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12934577%2F7cbed3e3a51c8f85d31b7b92b3870b65%2FXnip2023-01-03_22-45-24.jpg?generation=1672758146084252&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12934577%2F8575a8cc806a4af3de5e7769ffdd4ea2%2FXnip2023-01-03_22-41-10.jpg?generation=1672758159104699&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![Hi Chris you are truly teacher to me in this competition , and I quite have progress in the rerank part.. but I am facing a new problem about 'group'  in the modeling training part. when I use 'group' in your notebook , the error have appear that said \"num_row_ (4375050 vs. 4375063) : Invalid group structure\"(show at the picture blow)  ,but I cancel the 'group' it's OK to train my model. have you met this problem? how you fix up?and can you tell me the meaning of  'group' parameter and its important or not to the model training...thank you!! you answer make me grow up every day in this field ~~!\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12934577%2F27c220232d0643fdf9eddd2f4ff5c291%2FXnip2023-01-03_22-46-09.jpg?generation=1672758123273487&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12934577%2F7cbed3e3a51c8f85d31b7b92b3870b65%2FXnip2023-01-03_22-45-24.jpg?generation=1672758146084252&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12934577%2F8575a8cc806a4af3de5e7769ffdd4ea2%2FXnip2023-01-03_22-41-10.jpg?generation=1672758159104699&alt=media)",
      "votes": 2,
      "replies": [
        {
          "id": 2108272,
          "postDate": "2023-01-20T11:48:04.480Z",
          "content": "<p>I got the same problem about group size mismatch. After subsampling of negative sample and concatenate with positive sample. I only sorted my dataframe and  forgot to reindex it. Such broken index would interfere the behavior of groupkfolder.</p>",
          "rawMarkdown": "I got the same problem about group size mismatch. After subsampling of negative sample and concatenate with positive sample. I only sorted my dataframe and  forgot to reindex it. Such broken index would interfere the behavior of groupkfolder."
        }
      ]
    },
    {
      "id": 2078101,
      "postDate": "2022-12-28T04:45:14.573Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> the validation set you are using may be corrupted, for more details could you take a look at the discussion here? (kaggle spam system prevent me posting notebook links and other details, sorry for the trouble) Thanks a lot! </p>\n<p><a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation/discussion/374405\" target=\"_blank\">https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation/discussion/374405</a></p>",
      "rawMarkdown": "Hi @cdeotte the validation set you are using may be corrupted, for more details could you take a look at the discussion here? (kaggle spam system prevent me posting notebook links and other details, sorry for the trouble) Thanks a lot! \n\nhttps://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation/discussion/374405",
      "votes": 2,
      "replies": [
        {
          "id": 2078145,
          "postDate": "2022-12-28T05:07:04.023Z",
          "content": "<p>Thanks for the heads up. You are correct, Radek's validation dataset is not perfect. None-the-less, my current model has been tuned and discovered using Radek's original validation dataset with its slight errors.</p>",
          "rawMarkdown": "Thanks for the heads up. You are correct, Radek's validation dataset is not perfect. None-the-less, my current model has been tuned and discovered using Radek's original validation dataset with its slight errors.\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991#2032183"
        }
      ]
    },
    {
      "id": 2077952,
      "postDate": "2022-12-28T01:19:59.523Z",
      "content": "<p>Let's suppose I have this session S of tuples for the user JAMES, where the first element of the tuple is the product and the second is the type of event:</p>\n<p>S = [(\"notebook\", \"click\"), (\"notebook\", \"click\"), (\"notebook\", \"click\"), (\"mouse\", \"click\"), (\"mouse\", \"click\"), (\"mouse\", \"cart\"), (\"notebook\", \"cart\"), (\"notebook\", \"purchase\"), (\"mouse\", \"purchase\")]. </p>\n<p>And, we will create the user features:</p>\n<ul>\n<li>if is the first item (F1)</li>\n<li>if is the last item (F2)</li>\n<li>session length (F3)</li>\n</ul>\n<p>And for item features (<em>if you think it's not an item feature but actually a user feature pls tell me</em>):</p>\n<ul>\n<li>has this item already been clicked by user (F4)</li>\n<li>has this item already been added to cart by user  (F5)</li>\n<li>is this item popular? (F6)</li>\n<li>is this item bought in training data (at least once)?  (F7)</li>\n<li>times this item was clicked in all training data.  (F8)</li>\n</ul>\n<p>Also, we observed the co-visitation matrix for this data returned the items: [\"keyboard\"]. </p>\n<p>Our objective here is to create our candidate features to predict <strong>clicks</strong>. </p>\n<p>So, we will our candidate features this way:</p>\n<table>\n<thead>\n<tr>\n<th>USER</th>\n<th>ITEM</th>\n<th>F1.</th>\n<th>F2</th>\n<th>F3</th>\n<th>F4</th>\n<th>F5</th>\n<th>F6</th>\n<th>F7</th>\n<th>F8</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>JAMES</td>\n<td>NOTEBOOK</td>\n<td>1</td>\n<td>0</td>\n<td>9</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>381248</td>\n</tr>\n<tr>\n<td>JAMES</td>\n<td>MOUSE</td>\n<td>0</td>\n<td>1</td>\n<td>9</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1111232</td>\n</tr>\n<tr>\n<td>JAMES</td>\n<td>KEYBOARD</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>1</td>\n<td>1</td>\n<td>55555</td>\n</tr>\n</tbody>\n</table>\n<p>Observing this, my questions:</p>\n<ol>\n<li>Do my item features make sense? Maybe F4 and F5 is more like <strong>user features</strong> not <strong>item features</strong>.</li>\n<li>How the candidates from the co-visitation matrix (e.g the keyboard) could help the model if mostly of the columns/features are zeros? </li>\n<li>What's differ user-features, user-item-features from item-features? <br>\nThank you. </li>\n</ol>",
      "rawMarkdown": "Let's suppose I have this session S of tuples for the user JAMES, where the first element of the tuple is the product and the second is the type of event:\n\nS = [(\"notebook\", \"click\"), (\"notebook\", \"click\"), (\"notebook\", \"click\"), (\"mouse\", \"click\"), (\"mouse\", \"click\"), (\"mouse\", \"cart\"), (\"notebook\", \"cart\"), (\"notebook\", \"purchase\"), (\"mouse\", \"purchase\")]. \n\nAnd, we will create the user features:\n- if is the first item (F1)\n- if is the last item (F2)\n- session length (F3)\n\nAnd for item features (*if you think it's not an item feature but actually a user feature pls tell me*):\n- has this item already been clicked by user (F4)\n- has this item already been added to cart by user  (F5)\n- is this item popular? (F6)\n- is this item bought in training data (at least once)?  (F7)\n- times this item was clicked in all training data.  (F8)\n\nAlso, we observed the co-visitation matrix for this data returned the items: [\"keyboard\"]. \n\nOur objective here is to create our candidate features to predict **clicks**. \n\nSo, we will our candidate features this way:\n\n|USER | ITEM |  F1. |  F2 | F3  | F4 | F5 | F6 | F7 | F8 |\n| --- | --- | --- | --- | --- | --- | --- |--- |--- | --- |\n| JAMES|  NOTEBOOK |  1 |  0 | 9 | 1 | 1 | 1 | 1 | 381248 |\n| JAMES|  MOUSE |  0 |  1 | 9 | 1 | 1 | 1 | 1 | 1111232 |\n| JAMES|  KEYBOARD | 0 |  0 | 0 | 0 | 0 | 1 | 1 | 55555 |\n\nObserving this, my questions:\n\n1. Do my item features make sense? Maybe F4 and F5 is more like **user features** not **item features**.\n2. How the candidates from the co-visitation matrix (e.g the keyboard) could help the model if mostly of the columns/features are zeros? \n3. What's differ user-features, user-item-features from item-features? \nThank you. \n\n\n \n",
      "votes": 2,
      "replies": [
        {
          "id": 2077968,
          "postDate": "2022-12-28T02:01:24.537Z",
          "content": "<p>There are 3 types of features \"item\", \"user\", and \"user/item\". Your features are the following</p>\n<ul>\n<li>F1 - user/item</li>\n<li>F2 - user/item</li>\n<li>F3 - user</li>\n<li>F4 - user/item</li>\n<li>F5 - user/item</li>\n<li>F6 - item</li>\n<li>F7 - item</li>\n<li>F8 - item<br>\nAn \"item\" feature will always be the same for every row with this item regardless of user. A \"user\" feature will always be the same for every row with this user regardless of item. A \"user/item\" feature is everything else. </li>\n</ul>\n<p>Your feature F3 is a user feature, so your 3rd row should have a 9 in F3 column instead of 0. </p>",
          "rawMarkdown": "There are 3 types of features \"item\", \"user\", and \"user/item\". Your features are the following\n* F1 - user/item\n* F2 - user/item\n* F3 - user\n* F4 - user/item\n* F5 - user/item\n* F6 - item\n* F7 - item\n* F8 - item\nAn \"item\" feature will always be the same for every row with this item regardless of user. A \"user\" feature will always be the same for every row with this user regardless of item. A \"user/item\" feature is everything else. \n\nYour feature F3 is a user feature, so your 3rd row should have a 9 in F3 column instead of 0. ",
          "votes": 1,
          "replies": [
            {
              "id": 2078787,
              "postDate": "2022-12-28T16:38:30.063Z",
              "content": "<p>Thanks. </p>\n<p>And how items that were not clicked  (e.g came from co-visitation items) could help the algorithm if mostly of them have null/blank/zero columns ?</p>",
              "rawMarkdown": "Thanks. \n\nAnd how items that were not clicked  (e.g came from co-visitation items) could help the algorithm if mostly of them have null/blank/zero columns ?\n"
            }
          ]
        },
        {
          "id": 2080138,
          "postDate": "2022-12-29T22:30:05.853Z",
          "content": "<p>I believe at least F4 and F5 is user-item colaboration</p>\n<ol>\n<li>has this item already been clicked by user (F4)</li>\n<li>has this item already been added to cart by user (F5)</li>\n</ol>",
          "rawMarkdown": "I believe at least F4 and F5 is user-item colaboration\n1. has this item already been clicked by user (F4)\n2. has this item already been added to cart by user (F5)"
        }
      ]
    },
    {
      "id": 2077327,
      "postDate": "2022-12-27T13:24:56.363Z",
      "content": "<p>This is my first time participating in the competition. I will try this notebook!</p>",
      "rawMarkdown": "This is my first time participating in the competition. I will try this notebook!",
      "votes": 2
    },
    {
      "id": 2074808,
      "postDate": "2022-12-24T16:20:03.070Z",
      "content": "<p>If we use 5 fold for training, how to evaluate LOCAL CV . In particular, the entire validation set cannot be used directly.<br>\nThanks chris~</p>",
      "rawMarkdown": "If we use 5 fold for training, how to evaluate LOCAL CV . In particular, the entire validation set cannot be used directly.\nThanks chris~",
      "votes": 2,
      "replies": [
        {
          "id": 2074884,
          "postDate": "2022-12-24T18:14:52.293Z",
          "content": "<p>The targets are from <code>validation set B</code>. During training, we only use information from <code>new train</code> and <code>validation set A</code>. After making local predictions, we compute validation score using <code>validation set B</code>. </p>\n<p>Perhaps my use of names \"new train\", \"val set A\", and \"val set B\" are confusing. These are analogies of the datasets that Kaggle provides. Kaggle provides us with \"train.csv\", \"test.csv\", and \"public leaderboard\".</p>",
          "rawMarkdown": "The targets are from `validation set B`. During training, we only use information from `new train` and `validation set A`. After making local predictions, we compute validation score using `validation set B`. \n\nPerhaps my use of names \"new train\", \"val set A\", and \"val set B\" are confusing. These are analogies of the datasets that Kaggle provides. Kaggle provides us with \"train.csv\", \"test.csv\", and \"public leaderboard\"."
        }
      ]
    },
    {
      "id": 2072224,
      "postDate": "2022-12-21T21:53:26.533Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> again I come back to you with a question.<br>\nI'm kind of lost as I revisited the problem of creation of user and item features.<br>\nEven after adjusting the train kaggle data of kaggle by one week (to have three weeks in both training frames) the distribution of features for items and users seems vastly different.<br>\nIn my notation:<br>\ntrainA, testA -&gt; train and test taken from your otto validation<br>\ntrainB, testB -&gt; train and test taken from the parquets we use for LB inference.</p>\n<p>I adjusted time stamps as suggested by <a href=\"https://www.kaggle.com/aldparis\" target=\"_blank\">@aldparis</a> by two hours. So the trainB.ts.dt.day &gt; 7 in the code means we start from 8th of August (instead of 1st August) which then results in there weeks that we consider.<br>\ncode below illustrates my point:<br>\nAm I missing something? After improving my candidate generation process (co visitation) I want to be finally able to train my ranker. But i think as long as the feature distributions are so different I cant do that :-(.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11646918%2F47d953e7a2c949eeee56b460eaaaceb8%2FBildschirmfoto%20vom%202022-12-21%2023-04-59.png?generation=1671660337053128&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi @cdeotte again I come back to you with a question.\nI'm kind of lost as I revisited the problem of creation of user and item features.\nEven after adjusting the train kaggle data of kaggle by one week (to have three weeks in both training frames) the distribution of features for items and users seems vastly different.\nIn my notation:\ntrainA, testA -> train and test taken from your otto validation\ntrainB, testB -> train and test taken from the parquets we use for LB inference.\n\nI adjusted time stamps as suggested by @aldparis by two hours. So the trainB.ts.dt.day > 7 in the code means we start from 8th of August (instead of 1st August) which then results in there weeks that we consider.\ncode below illustrates my point:\nAm I missing something? After improving my candidate generation process (co visitation) I want to be finally able to train my ranker. But i think as long as the feature distributions are so different I cant do that :-(.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11646918%2F47d953e7a2c949eeee56b460eaaaceb8%2FBildschirmfoto%20vom%202022-12-21%2023-04-59.png?generation=1671660337053128&alt=media)",
      "votes": 2
    },
    {
      "id": 2067385,
      "postDate": "2022-12-16T16:52:41.373Z",
      "content": "<p>Thank you very much, Chris! I am trying your method with my own features, however, I got some trouble with setting the data for training the XGB Ranker. When I set <code>group = train_group</code> in the xgb.DMatrix wraper, the model seems to learn nothing (loss = 1.0, predictions are all the same for all instances), when I omit this, the model seems fine, but the performance is much worse than with manual choosing (Recall of clicks are ~0.42). I'd appreciate your help in inspecting this bug. Thank you!</p>",
      "rawMarkdown": "Thank you very much, Chris! I am trying your method with my own features, however, I got some trouble with setting the data for training the XGB Ranker. When I set `group = train_group` in the xgb.DMatrix wraper, the model seems to learn nothing (loss = 1.0, predictions are all the same for all instances), when I omit this, the model seems fine, but the performance is much worse than with manual choosing (Recall of clicks are ~0.42). I'd appreciate your help in inspecting this bug. Thank you!",
      "votes": 2,
      "replies": [
        {
          "id": 2069047,
          "postDate": "2022-12-18T14:25:05.197Z",
          "content": "<p><a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a> what loss are you using? If you are using a ranking metric like <code>MAP</code> then <code>metric=1</code> means that you have perfectly ranked the items therefore there is nothing else to learn.</p>",
          "rawMarkdown": "@shinomoriaoshi what loss are you using? If you are using a ranking metric like `MAP` then `metric=1` means that you have perfectly ranked the items therefore there is nothing else to learn.",
          "votes": 1,
          "replies": [
            {
              "id": 2069215,
              "postDate": "2022-12-18T18:16:26.097Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I used the objective \"rank:pairwise\". I suspected some groups only have the 0 class (no ground truth), but when I excluded them, nothing changed.</p>",
              "rawMarkdown": "@cdeotte I used the objective \"rank:pairwise\". I suspected some groups only have the 0 class (no ground truth), but when I excluded them, nothing changed."
            },
            {
              "id": 2069235,
              "postDate": "2022-12-18T18:47:45.683Z",
              "content": "<p>By default, <code>rank:pairwise</code> loss uses <code>map</code> metric. During training, XGB displays the current <code>map</code> metric. When it reaches <code>1.0</code>, there is nothing else to learn because you have achieved perfect ranking. For me, my map is 0.9XX for orders and carts. And 0.7XX for clicks during training.</p>\n<p>If you are achieving perfect ranking, then there must be something wrong with your ground truth. And/or you have a leaking feature.</p>",
              "rawMarkdown": "By default, `rank:pairwise` loss uses `map` metric. During training, XGB displays the current `map` metric. When it reaches `1.0`, there is nothing else to learn because you have achieved perfect ranking. For me, my map is 0.9XX for orders and carts. And 0.7XX for clicks during training.\n\nIf you are achieving perfect ranking, then there must be something wrong with your ground truth. And/or you have a leaking feature.",
              "votes": 4
            },
            {
              "id": 2069300,
              "postDate": "2022-12-18T20:40:35.333Z",
              "content": "<p>Thank you for pointing this out, I am using your code to generate candidates and ground truth so I believe there is nothing wrong with that. I will check my set of features, currently, I am using a customized Word2Vec model to derive the set of features, the leak could root from that.</p>",
              "rawMarkdown": "Thank you for pointing this out, I am using your code to generate candidates and ground truth so I believe there is nothing wrong with that. I will check my set of features, currently, I am using a customized Word2Vec model to derive the set of features, the leak could root from that."
            },
            {
              "id": 2070072,
              "postDate": "2022-12-19T15:21:26.230Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for sharing number for each type. My training MAP for clicks is comparable to what you have mentioned here, however, my validation MAP is 0.9XX. Any hunches why this might be the case? I am providing <code>qid</code> instead of <code>group</code> in DMatrix. I will try out with group and check if anything changes. </p>",
              "rawMarkdown": "@cdeotte thanks for sharing number for each type. My training MAP for clicks is comparable to what you have mentioned here, however, my validation MAP is 0.9XX. Any hunches why this might be the case? I am providing `qid` instead of `group` in DMatrix. I will try out with group and check if anything changes. "
            },
            {
              "id": 2070185,
              "postDate": "2022-12-19T17:21:56.057Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> circling back on the previous comment, it looks like on average we have around 11 aids per session in the train set and 3 aids per session in the validation set. In this case, should we consider a different split perhaps to better evaluate ranking on the validation set?</p>",
              "rawMarkdown": "@cdeotte circling back on the previous comment, it looks like on average we have around 11 aids per session in the train set and 3 aids per session in the validation set. In this case, should we consider a different split perhaps to better evaluate ranking on the validation set?"
            },
            {
              "id": 2070191,
              "postDate": "2022-12-19T17:27:57.910Z",
              "content": "<p>Regarding <code>clicks</code> there is only 1 ground truth for each user (i.e. the next click). For <code>carts</code> and <code>orders</code> the ground truth is all future <code>carts</code> and <code>orders</code> (during last week). But for <code>clicks</code> the ground truth is only the next <code>click</code>.</p>",
              "rawMarkdown": "Regarding `clicks` there is only 1 ground truth for each user (i.e. the next click). For `carts` and `orders` the ground truth is all future `carts` and `orders` (during last week). But for `clicks` the ground truth is only the next `click`.",
              "votes": 2
            },
            {
              "id": 2070290,
              "postDate": "2022-12-19T19:59:58.480Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for your response! After the merge laid out in your code in step 6, it looks like we drop nearly 800k ground truth aids, i.e. those sessions do not have any rows where <code>'click' = 1</code>. Perhaps you encountered it in your implementation. I am re-visiting my implementation again - I might have missed something in the process. </p>",
              "rawMarkdown": "@cdeotte thanks for your response! After the merge laid out in your code in step 6, it looks like we drop nearly 800k ground truth aids, i.e. those sessions do not have any rows where `'click' = 1`. Perhaps you encountered it in your implementation. I am re-visiting my implementation again - I might have missed something in the process. "
            },
            {
              "id": 2070312,
              "postDate": "2022-12-19T20:26:47.110Z",
              "content": "<p><a href=\"https://www.kaggle.com/parthpankajtiwary\" target=\"_blank\">@parthpankajtiwary</a> We use merge with <code>how='left'</code> so we do not drop any rows. We merge <code>'click'=1</code> and then we keep all rows and <code>fillna(0)</code>. Therefore afterward we have both positive and negative targets.</p>\n<p>Also note that we make 3 XGB ranker models. So we do step 6 for each target separately. We have one dataframe for clicks, carts, and orders. (Or a single dataframe with 3 target columns)</p>",
              "rawMarkdown": "@parthpankajtiwary We use merge with `how='left'` so we do not drop any rows. We merge `'click'=1` and then we keep all rows and `fillna(0)`. Therefore afterward we have both positive and negative targets.\n\nAlso note that we make 3 XGB ranker models. So we do step 6 for each target separately. We have one dataframe for clicks, carts, and orders. (Or a single dataframe with 3 target columns)",
              "votes": 1
            },
            {
              "id": 2070362,
              "postDate": "2022-12-19T22:48:28.710Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for the clarification! I have been able to resolve the issue on my end. All thanks to this amazing write-up and your responses :)</p>",
              "rawMarkdown": "@cdeotte Thanks for the clarification! I have been able to resolve the issue on my end. All thanks to this amazing write-up and your responses :)",
              "votes": 2
            },
            {
              "id": 2076724,
              "postDate": "2022-12-26T18:41:10.513Z",
              "content": "<p><a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a> Did you solve your <code>1.0</code> problem? I have just seen it with my own ranker. Before training, you need to sort your train data so that all users are on consecutive rows. And the <code>group counts</code> need to be in the same order as the unique users appear in row order.</p>",
              "rawMarkdown": "@shinomoriaoshi Did you solve your `1.0` problem? I have just seen it with my own ranker. Before training, you need to sort your train data so that all users are on consecutive rows. And the `group counts` need to be in the same order as the unique users appear in row order."
            },
            {
              "id": 2078770,
              "postDate": "2022-12-28T16:20:50.227Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I don't know if <a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a> solved this problem, but now I have it 😅. I will try to look at the order of train data and group counts. I hope this will help.</p>",
              "rawMarkdown": "@cdeotte I don't know if @shinomoriaoshi solved this problem, but now I have it 😅. I will try to look at the order of train data and group counts. I hope this will help."
            },
            {
              "id": 2078811,
              "postDate": "2022-12-28T17:04:10.420Z",
              "content": "<p>I solved my problem with train_map and valid_map equal 1.000000. I was simply leaking the actual target to the training process. When I corrected the leakage everything went back to normal, now for clicks my map equals to 0.9XXX</p>",
              "rawMarkdown": "I solved my problem with train_map and valid_map equal 1.000000. I was simply leaking the actual target to the training process. When I corrected the leakage everything went back to normal, now for clicks my map equals to 0.9XXX"
            },
            {
              "id": 2080950,
              "postDate": "2022-12-30T16:30:48.960Z",
              "content": "<p><a href=\"https://www.kaggle.com/parthpankajtiwary\" target=\"_blank\">@parthpankajtiwary</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nDid you find the reason for this? I am facing something similar.<br>\ntrain_map being too bad, valid-map is doing well? any guess on what could be the reason?</p>\n<pre><code>[]    train-:   valid-:\n[]    train-:   valid-:\n[]    train-:   valid-:\n[]    train-:   valid-:\n[]    train-:   valid-:\n[]    train-:   valid-:\n[]    train-:   valid-:\n</code></pre>",
              "rawMarkdown": "@parthpankajtiwary @cdeotte \nDid you find the reason for this? I am facing something similar.\ntrain_map being too bad, valid-map is doing well? any guess on what could be the reason?\n\n```python\n[0]\ttrain-map:0.19732\tvalid-map:0.92789\n[500]\ttrain-map:0.24034\tvalid-map:0.93368\n[1000]\ttrain-map:0.25308\tvalid-map:0.93437\n[1500]\ttrain-map:0.26383\tvalid-map:0.93486\n[2000]\ttrain-map:0.27324\tvalid-map:0.93514\n[2500]\ttrain-map:0.28226\tvalid-map:0.93527\n[2999]\ttrain-map:0.28984\tvalid-map:0.93530\n```"
            },
            {
              "id": 2082149,
              "postDate": "2023-01-01T06:52:47.987Z",
              "content": "<p>Awesome <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I was facing the 1.0 issue since a week. Sorting the data by users before getting the group count fixed it.</p>",
              "rawMarkdown": "Awesome @cdeotte, I was facing the 1.0 issue since a week. Sorting the data by users before getting the group count fixed it.",
              "votes": 1
            },
            {
              "id": 2082159,
              "postDate": "2023-01-01T07:08:36.150Z",
              "content": "<p>I have the same problem with you and my train-map and valid -map all weight 1.0, can you give some tips to me how you fix it?</p>",
              "rawMarkdown": "I have the same problem with you and my train-map and valid -map all weight 1.0, can you give some tips to me how you fix it?"
            },
            {
              "id": 2082484,
              "postDate": "2023-01-01T16:09:15.290Z",
              "content": "<p><a href=\"https://www.kaggle.com/adairli00\" target=\"_blank\">@adairli00</a> When using GBT ranker, your train data must have all rows for the same county together consectutive. So <code>train = train.sort_values('cfips')</code>. Then afterward, compute the size of each county group in the order they appear in your new sorted train dataframe. <code>groups = train.groupby('cfips').cfips.agg('count').values</code></p>",
              "rawMarkdown": "@adairli00 When using GBT ranker, your train data must have all rows for the same county together consectutive. So `train = train.sort_values('cfips')`. Then afterward, compute the size of each county group in the order they appear in your new sorted train dataframe. `groups = train.groupby('cfips').cfips.agg('count').values`",
              "votes": 2
            },
            {
              "id": 2082810,
              "postDate": "2023-01-02T01:12:32.587Z",
              "content": "<p>thank you Chris so much , it very helpful for me </p>",
              "rawMarkdown": "thank you Chris so much , it very helpful for me ",
              "votes": 1
            },
            {
              "id": 2104743,
              "postDate": "2023-01-18T02:21:14.397Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2104746,
              "postDate": "2023-01-18T02:22:07.050Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2104748,
              "postDate": "2023-01-18T02:23:45.377Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2067367,
      "postDate": "2022-12-16T16:29:16.853Z",
      "content": "<p>Thank you for sharing the idea! It's very useful! <br>\nThere's one more question. In your example, the ground truth for 'clicks' in a session contains several different aids, yet in my understanding there's only one ground truth for type 'clicks' in the real task, is that true?</p>",
      "rawMarkdown": "Thank you for sharing the idea! It's very useful! \nThere's one more question. In your example, the ground truth for 'clicks' in a session contains several different aids, yet in my understanding there's only one ground truth for type 'clicks' in the real task, is that true?",
      "votes": 2,
      "replies": [
        {
          "id": 2067403,
          "postDate": "2022-12-16T17:07:15.333Z",
          "content": "<p>Great point <a href=\"https://www.kaggle.com/samsonfha\" target=\"_blank\">@samsonfha</a> . You are correct. When adding ground truth label for click, we should only use the next one click. I will update my post. If you use Radek's ground truth file, then my code above still works because Radek only puts one click ground truth in each user's list.</p>",
          "rawMarkdown": "Great point @samsonfha . You are correct. When adding ground truth label for click, we should only use the next one click. I will update my post. If you use Radek's ground truth file, then my code above still works because Radek only puts one click ground truth in each user's list.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2065267,
      "postDate": "2022-12-14T13:14:14.120Z",
      "content": "<p>Thank you for this guide, it's very helpful.</p>\n<p>I have question about step 5. In your table aren't user_feat1 and user_feat2 should be the same for the same user regardless of any item? and we also need to merge user-item interaction feature on [user, item] right?</p>\n<p>And i think group sizes will always be different after merging ground truth label as it might add extra row</p>",
      "rawMarkdown": "Thank you for this guide, it's very helpful.\n\nI have question about step 5. In your table aren't user_feat1 and user_feat2 should be the same for the same user regardless of any item? and we also need to merge user-item interaction feature on [user, item] right?\n\nAnd i think group sizes will always be different after merging ground truth label as it might add extra row",
      "votes": 2,
      "replies": [
        {
          "id": 2065275,
          "postDate": "2022-12-14T13:29:36.680Z",
          "content": "<p>Yes, the user feats should be the same for same user. Thanks, i updated my post. Yes, we merge user-item on [user,item].</p>\n<p>When we merge ground truth, we use `how='left`` so it will not add an extra row. Furthermore the ground truth dataframe does not have any duplicate rows, so it will not add more rows.</p>",
          "rawMarkdown": "Yes, the user feats should be the same for same user. Thanks, i updated my post. Yes, we merge user-item on [user,item].\n\nWhen we merge ground truth, we use `how='left`` so it will not add an extra row. Furthermore the ground truth dataframe does not have any duplicate rows, so it will not add more rows.",
          "replies": [
            {
              "id": 2067064,
              "postDate": "2022-12-16T11:01:58.770Z",
              "content": "<p>Little confused about ground truth merge..</p>\n<p>If we use left merge with ground truth, wouldn't some positive label user item pair in ground truth that is not in candidate generated be left off? Unless we deliberately also put item from ground truth into the candidate aside from using covisitation matrix…</p>",
              "rawMarkdown": "Little confused about ground truth merge..\n\nIf we use left merge with ground truth, wouldn't some positive label user item pair in ground truth that is not in candidate generated be left off? Unless we deliberately also put item from ground truth into the candidate aside from using covisitation matrix..."
            },
            {
              "id": 2070032,
              "postDate": "2022-12-19T14:49:14.660Z",
              "content": "<p><a href=\"https://www.kaggle.com/bestsv6\" target=\"_blank\">@bestsv6</a> This is true, our candidates will not include all the ground truth. To improve our models and have more ground truth, we need to find techniques to find better candidates. But we cannot look at the list of ground truth to pick our candidates. That would be a leak because we cannot look at ground truth of Kaggle's leaderboard to pick candidates for our submission.</p>",
              "rawMarkdown": "@bestsv6 This is true, our candidates will not include all the ground truth. To improve our models and have more ground truth, we need to find techniques to find better candidates. But we cannot look at the list of ground truth to pick our candidates. That would be a leak because we cannot look at ground truth of Kaggle's leaderboard to pick candidates for our submission.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2062223,
      "postDate": "2022-12-11T21:41:50.410Z",
      "content": "<p><strong>UPDATE</strong> I fixed the code in <strong>Step 6</strong>. The correct code for creating target dataframe (that we can merge unto our candidate dataframe) is</p>\n<pre><code>tar = pd.read_parquet('test_labels.parquet')\ntar = tar.loc[ tar['type']=='clicks' ]\naids = tar.ground_truth.explode().astype('int32').rename('item')\ntar = tar[['session']].astype('int32').rename({'session':'user'},axis=1)\ntar = tar.merge(aids, left_index=True, right_index=True, how='left')\ntar['click'] = 1\n</code></pre>\n<p>The old wrong code was</p>\n<pre><code>tar = pd.read_parquet('test_labels.parquet')\ntar = tar.loc[ tar['type']=='clicks' ]\ntar = tar.labels.explode().astype('int32')\ntar.columns = ['user','item']\ntar['click'] = 1\n</code></pre>",
      "rawMarkdown": "**UPDATE** I fixed the code in **Step 6**. The correct code for creating target dataframe (that we can merge unto our candidate dataframe) is\n\n    tar = pd.read_parquet('test_labels.parquet')\n    tar = tar.loc[ tar['type']=='clicks' ]\n    aids = tar.ground_truth.explode().astype('int32').rename('item')\n    tar = tar[['session']].astype('int32').rename({'session':'user'},axis=1)\n    tar = tar.merge(aids, left_index=True, right_index=True, how='left')\n    tar['click'] = 1\n\nThe old wrong code was\n\n    tar = pd.read_parquet('test_labels.parquet')\n    tar = tar.loc[ tar['type']=='clicks' ]\n    tar = tar.labels.explode().astype('int32')\n    tar.columns = ['user','item']\n    tar['click'] = 1",
      "votes": 2
    },
    {
      "id": 2061240,
      "postDate": "2022-12-10T21:30:32.717Z",
      "content": "<p>tnx for this awesome post <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !<br>\ntoday i took the time to follow all the steps you describe and could improve CV<br>\n0.5655 (co visitation approach with some improvement)<br>\nto 0.5662.<br>\nWithout much changes/improvements in my code to the above described pipeline.<br>\nNow i should tune hyper parameters and create more and better candidates.</p>",
      "rawMarkdown": "tnx for this awesome post @cdeotte !\ntoday i took the time to follow all the steps you describe and could improve CV\n0.5655 (co visitation approach with some improvement)\nto 0.5662.\nWithout much changes/improvements in my code to the above described pipeline.\nNow i should tune hyper parameters and create more and better candidates.",
      "votes": 2,
      "replies": [
        {
          "id": 2061251,
          "postDate": "2022-12-10T21:51:41.873Z",
          "content": "<p>Great job Simon. Adding more high quality candidates (which may involve removing some of your existing candidates and replacing with better candidates because we don't want too many candidates for our GBT ranker), adding more (item, user, and interactions) features, and tuning your GBT ranker model will all improve CV and LB!</p>",
          "rawMarkdown": "Great job Simon. Adding more high quality candidates (which may involve removing some of your existing candidates and replacing with better candidates because we don't want too many candidates for our GBT ranker), adding more (item, user, and interactions) features, and tuning your GBT ranker model will all improve CV and LB!",
          "votes": 2
        },
        {
          "id": 2061971,
          "postDate": "2022-12-11T16:19:26.947Z",
          "content": "<p>Hi Chris. <br>\nI observed something strange.<br>\nWith the co visitation approach and the related CV given above I achieved 0.576 on LB. <br>\nBut when I generated candidates for real test data and generated features (with real train and real test data) I just got LB score 0.567.<br>\nIsn't this strange? Because the CV for the ranjer approach was higher. Or is there some other reason that the boost from CV to LB score us not as big for the ranker modell? <br>\nOtherwise I would look for bugs in the way I generated candidates and features for the real test data but to be honest I'm quiet sure I didn't make mistake (as the only thing I had to change in my code were the data sources…)</p>",
          "rawMarkdown": "Hi Chris. \nI observed something strange.\nWith the co visitation approach and the related CV given above I achieved 0.576 on LB. \nBut when I generated candidates for real test data and generated features (with real train and real test data) I just got LB score 0.567.\nIsn't this strange? Because the CV for the ranjer approach was higher. Or is there some other reason that the boost from CV to LB score us not as big for the ranker modell? \nOtherwise I would look for bugs in the way I generated candidates and features for the real test data but to be honest I'm quiet sure I didn't make mistake (as the only thing I had to change in my code were the data sources...)",
          "votes": 1
        },
        {
          "id": 2061998,
          "postDate": "2022-12-11T16:57:17.040Z",
          "content": "<p>For me, every improvement in CV is an improvement in LB, so I suspect something is happening in your pipeline. Here are two things to check out</p>\n<ul>\n<li>When training models and computing CV, we use our \"new train\" dataset (3 weeks of data) and our \"validation A\" dataset (roughly half week of data). Then when we build our submission we use \"Kaggle train\" (4 weeks of data) and \"Kaggle test\" (roughly half week of data).</li>\n<li>Next, make sure that your features <strong>generalize</strong>. For example, you <strong>cannot</strong> use a timestamp because the timestamps in the train datasets are different than timestamps in submission dataset. Instead of timestamps we use offsets. Like <code>ts minus dataset min ts</code>. Then the feature distribution in both CV and LB will be the same. Also if you count things. You cannot count things in 3 weeks of \"new train\" dataset for CV and then count things in 4 weeks of \"Kaggle train\" because the later counts will be 33% larger. You must be careful to make sure the distribution of CV features is the same as distribution of LB features.</li>\n</ul>",
          "rawMarkdown": "For me, every improvement in CV is an improvement in LB, so I suspect something is happening in your pipeline. Here are two things to check out\n* When training models and computing CV, we use our \"new train\" dataset (3 weeks of data) and our \"validation A\" dataset (roughly half week of data). Then when we build our submission we use \"Kaggle train\" (4 weeks of data) and \"Kaggle test\" (roughly half week of data).\n* Next, make sure that your features **generalize**. For example, you **cannot** use a timestamp because the timestamps in the train datasets are different than timestamps in submission dataset. Instead of timestamps we use offsets. Like `ts minus dataset min ts`. Then the feature distribution in both CV and LB will be the same. Also if you count things. You cannot count things in 3 weeks of \"new train\" dataset for CV and then count things in 4 weeks of \"Kaggle train\" because the later counts will be 33% larger. You must be careful to make sure the distribution of CV features is the same as distribution of LB features.",
          "votes": 7,
          "replies": [
            {
              "id": 2063582,
              "postDate": "2022-12-13T05:06:25.233Z",
              "content": "<blockquote>\n  <ul>\n  <li>Then when we build our submission we use \"Kaggle train\" (4 weeks of data) and \"Kaggle test\" (roughly half week of data).</li>\n  <li>Also if you count things. You cannot count things in 3 weeks of \"new train\" dataset for CV and then count things in 4 weeks of \"Kaggle train\" because the later counts will be 33% larger. You must be careful to make sure the distribution of CV features is the same as distribution of LB features.</li>\n  </ul>\n</blockquote>\n<p>Hi,Chris!<br>\nMy ranker model also has some aid count features in it. These aid count features are generated with new train data(3 weeks of kaggle train data). I'm using the Ranker model to infer kaggle test data. I need to calculate the characteristics of kaggle test users'  candidates. Should I use the last three weeks of data in kaggle train data plus half-week data from kaggle test data for the calculation of these aid count features in order to maintain a balanced distribution?</p>",
              "rawMarkdown": "> * Then when we build our submission we use \"Kaggle train\" (4 weeks of data) and \"Kaggle test\" (roughly half week of data).\n> *  Also if you count things. You cannot count things in 3 weeks of \"new train\" dataset for CV and then count things in 4 weeks of \"Kaggle train\" because the later counts will be 33% larger. You must be careful to make sure the distribution of CV features is the same as distribution of LB features.\n\nHi,Chris!\nMy ranker model also has some aid count features in it. These aid count features are generated with new train data(3 weeks of kaggle train data). I'm using the Ranker model to infer kaggle test data. I need to calculate the characteristics of kaggle test users'  candidates. Should I use the last three weeks of data in kaggle train data plus half-week data from kaggle test data for the calculation of these aid count features in order to maintain a balanced distribution?"
            }
          ]
        },
        {
          "id": 2062000,
          "postDate": "2022-12-11T16:59:46.073Z",
          "content": "<p>If you increased your CV score from 0.5655 to 0.5662 and you did not use \"validation dataset B\" (i.e. the ground truth) during training. Then you can certainly increase your LB by <code>+0.0007</code> also. You just need to carefully build the features using Kaggle train and Kaggle test so that they appear the same to your GBT model as the features that you built from \"new train\" and \"validation data A\".</p>",
          "rawMarkdown": "If you increased your CV score from 0.5655 to 0.5662 and you did not use \"validation dataset B\" (i.e. the ground truth) during training. Then you can certainly increase your LB by `+0.0007` also. You just need to carefully build the features using Kaggle train and Kaggle test so that they appear the same to your GBT model as the features that you built from \"new train\" and \"validation data A\".",
          "votes": 1
        },
        {
          "id": 2062073,
          "postDate": "2022-12-11T17:47:21.147Z",
          "content": "<p>One way to investigate is plot pairs of histograms for each of your features. For example if you have <code>feature_1</code>, then plot histogram of <code>feature_1</code> column from your train data (used to train model and compute CV), and plot histogram of <code>feature_1</code> column that you use during inference for LB sub. GBT models need to have the same histogram during train and infer. Most importantly check the max and min values, and the quantiles of each histogram. Looking at histograms with your eyeball should be enough to find the problem. A well designed generalized feature will have similar histograms for both train and infer.</p>",
          "rawMarkdown": "One way to investigate is plot pairs of histograms for each of your features. For example if you have `feature_1`, then plot histogram of `feature_1` column from your train data (used to train model and compute CV), and plot histogram of `feature_1` column that you use during inference for LB sub. GBT models need to have the same histogram during train and infer. Most importantly check the max and min values, and the quantiles of each histogram. Looking at histograms with your eyeball should be enough to find the problem. A well designed generalized feature will have similar histograms for both train and infer.",
          "votes": 4
        },
        {
          "id": 2062080,
          "postDate": "2022-12-11T17:56:44.947Z",
          "content": "<p>Hello Chris.<br>\nGreat for the detailed answer. I'm pretty sure you already named the bug! Indeed I have several \"count\" variables, which will certainly have different scale if I take real train and test Set. I will scale everything now by the time elapsed and report if that helped. Tnx!!</p>",
          "rawMarkdown": "Hello Chris.\nGreat for the detailed answer. I'm pretty sure you already named the bug! Indeed I have several \"count\" variables, which will certainly have different scale if I take real train and test Set. I will scale everything now by the time elapsed and report if that helped. Tnx!!",
          "votes": 1
        },
        {
          "id": 2064436,
          "postDate": "2022-12-13T18:49:16.723Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> &amp;  <a href=\"https://www.kaggle.com/niejianfei\" target=\"_blank\">@niejianfei</a>. Unfortunately I still have the same problem.<br>\nI described my pipeline in more detail <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/372030\" target=\"_blank\">here</a>. I would greatly appreciate if someone could point me to the mistake I make, I guess its something obvious i made wrong 🤒</p>",
          "rawMarkdown": "Hello @cdeotte &  @niejianfei. Unfortunately I still have the same problem.\nI described my pipeline in more detail [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/372030). I would greatly appreciate if someone could point me to the mistake I make, I guess its something obvious i made wrong 🤒"
        },
        {
          "id": 2064521,
          "postDate": "2022-12-13T20:43:22.973Z",
          "content": "<p><a href=\"https://www.kaggle.com/niejianfei\" target=\"_blank\">@niejianfei</a> Yes Jensen, that is what i do</p>\n<blockquote>\n  <p>Should I use the last three weeks of data in kaggle train data plus half-week data from kaggle test data for the calculation of these aid count features in order to maintain a balanced distribution?</p>\n</blockquote>",
          "rawMarkdown": "@niejianfei Yes Jensen, that is what i do\n> Should I use the last three weeks of data in kaggle train data plus half-week data from kaggle test data for the calculation of these aid count features in order to maintain a balanced distribution?",
          "replies": [
            {
              "id": 2064694,
              "postDate": "2022-12-14T03:48:06.380Z",
              "content": "<p>thank you for your answer</p>",
              "rawMarkdown": "thank you for your answer",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2057813,
      "postDate": "2022-12-07T11:09:21.423Z",
      "content": "<p>You are my angel.</p>",
      "rawMarkdown": "You are my angel.",
      "votes": 2
    },
    {
      "id": 2056545,
      "postDate": "2022-12-06T08:13:19.210Z",
      "content": "<p>Thank you so much. I have learned a lot from this thread about the ranker model !!</p>",
      "rawMarkdown": "Thank you so much. I have learned a lot from this thread about the ranker model !!",
      "votes": 2
    },
    {
      "id": 2054786,
      "postDate": "2022-12-04T13:06:42.103Z",
      "content": "<p><strong>UPDATE</strong>: I corrected step 1. In step 1, we generate candidates for every user in <strong>validation data</strong> (i.e. last 1 week of train). We do not generate candidates for users in <strong>train data</strong> (i.e. first 3 weeks of train). Because we only have targets for validation users in step 6.</p>",
      "rawMarkdown": "**UPDATE**: I corrected step 1. In step 1, we generate candidates for every user in **validation data** (i.e. last 1 week of train). We do not generate candidates for users in **train data** (i.e. first 3 weeks of train). Because we only have targets for validation users in step 6.",
      "votes": 2,
      "replies": [
        {
          "id": 2056266,
          "postDate": "2022-12-05T23:56:10.937Z",
          "content": "<p><strong>UPDATE</strong> In the inference step, we create a new candidate list for all users in <strong>Kaggle test data</strong>. (We do not create candidates for users in Kaggle train data).</p>",
          "rawMarkdown": "**UPDATE** In the inference step, we create a new candidate list for all users in **Kaggle test data**. (We do not create candidates for users in Kaggle train data).",
          "votes": 1
        }
      ]
    },
    {
      "id": 2054181,
      "postDate": "2022-12-03T23:22:11.353Z",
      "content": "<p>Hi Chris. Great project outline here! </p>\n<p>Am I right in saying the ranker training will be greatly impacted by the quality of our initial candidates? E.g if from our 50 candidates actually only 1 was correct, then for that session the ranker will only have positive sample to learn from? </p>",
      "rawMarkdown": "Hi Chris. Great project outline here! \n\nAm I right in saying the ranker training will be greatly impacted by the quality of our initial candidates? E.g if from our 50 candidates actually only 1 was correct, then for that session the ranker will only have positive sample to learn from? ",
      "votes": 2,
      "replies": [
        {
          "id": 2054259,
          "postDate": "2022-12-04T01:38:51.800Z",
          "content": "<p>Yes. It is important to provide good candidates. (The better the candidates the fewer you need. The worse the candidates the more you need). You can always compute the score <code>recall@50</code> or <code>recall@100</code> or <code>recall@200</code> on your 50, 100, or 200 candidate. The recall using all the candidates show you what the highest score you can achieve when you pick your final 20. </p>",
          "rawMarkdown": "Yes. It is important to provide good candidates. (The better the candidates the fewer you need. The worse the candidates the more you need). You can always compute the score `recall@50` or `recall@100` or `recall@200` on your 50, 100, or 200 candidate. The recall using all the candidates show you what the highest score you can achieve when you pick your final 20. ",
          "votes": 9
        }
      ]
    },
    {
      "id": 3021487,
      "postDate": "2024-10-18T14:56:28.010Z",
      "content": "<p>How are negative samples constructed?</p>",
      "rawMarkdown": "How are negative samples constructed?"
    },
    {
      "id": 2303763,
      "postDate": "2023-06-15T13:12:32.613Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . Viewing this thread in 2023, after the competition. The steps are clearly explained. I am replicating these for new project.</p>",
      "rawMarkdown": "Thanks @cdeotte . Viewing this thread in 2023, after the competition. The steps are clearly explained. I am replicating these for new project."
    },
    {
      "id": 2246867,
      "postDate": "2023-05-05T14:20:48.123Z",
      "content": "<p>Thx for your sharing, it did help a lots</p>",
      "rawMarkdown": "Thx for your sharing, it did help a lots\n"
    },
    {
      "id": 2124574,
      "postDate": "2023-02-01T04:48:53.600Z",
      "content": "<p>Thank you Chris, I learned a lot from your sharing！</p>\n<p>I have some doubts in the Inference step：<br>\nI would like to know why there is no need to specify a group in Inference?<br>\nHow does the model judge whether the data is in the same group during inference?<br>\nDoes this mean that the inference group must be the same as the training group?<br>\nI looked at the docs but didn't find what I want to know, am I missing something?</p>",
      "rawMarkdown": "Thank you Chris, I learned a lot from your sharing！\n\nI have some doubts in the Inference step：\nI would like to know why there is no need to specify a group in Inference?\nHow does the model judge whether the data is in the same group during inference?\nDoes this mean that the inference group must be the same as the training group?\nI looked at the docs but didn't find what I want to know, am I missing something?"
    },
    {
      "id": 2081143,
      "postDate": "2022-12-30T20:53:20.247Z",
      "content": "<p>My ranker performed really poorly. Although all 3 maps during training were &gt;= 0.9 my public score was 0.176. I concluded that the model was heavilly overfitted. Now I'm trying to reduce <em>subsample</em>, <em>colsample_bytree</em> and <em>max_depth</em> parameters. </p>\n<p>I also tried to give more weight to positive class by increasing <em>scale_pos_weight</em> parameter, but I got this error:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F26c16db43d75790ac7a16645dad18103%2FError.JPG?generation=1672433503878014&amp;alt=media\" alt=\"\"></p>\n<p>Has anybody tried to increase the weight of positive events and encountered this error? What was  the root cause?</p>",
      "rawMarkdown": "My ranker performed really poorly. Although all 3 maps during training were >= 0.9 my public score was 0.176. I concluded that the model was heavilly overfitted. Now I'm trying to reduce *subsample*, *colsample_bytree* and *max_depth* parameters. \n\nI also tried to give more weight to positive class by increasing *scale_pos_weight* parameter, but I got this error:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F26c16db43d75790ac7a16645dad18103%2FError.JPG?generation=1672433503878014&alt=media)\n\nHas anybody tried to increase the weight of positive events and encountered this error? What was  the root cause?"
    },
    {
      "id": 2078580,
      "postDate": "2022-12-28T12:48:15.617Z",
      "content": "<p>How should user features be used? None of the users in the test data have appeared in the training set?</p>",
      "rawMarkdown": "How should user features be used? None of the users in the test data have appeared in the training set?",
      "replies": [
        {
          "id": 2078583,
          "postDate": "2022-12-28T12:51:39.150Z",
          "content": "<p>We must create user features from the <code>test.csv</code> file. (In the 3 part validation scheme presented in my discussion post, this is called \"validation A\".) For each test user we have the first half of their activity and we must predict the second half. So a user feature could be </p>\n<ul>\n<li>how many items did they click</li>\n<li>what is the average hour of the day that they click</li>\n<li>what is the average day of the week that they click</li>\n<li>what proportion of their clicks do they order<br>\netc etc etc</li>\n</ul>",
          "rawMarkdown": "We must create user features from the `test.csv` file. (In the 3 part validation scheme presented in my discussion post, this is called \"validation A\".) For each test user we have the first half of their activity and we must predict the second half. So a user feature could be \n* how many items did they click\n* what is the average hour of the day that they click\n* what is the average day of the week that they click\n* what proportion of their clicks do they order\netc etc etc",
          "votes": 2,
          "replies": [
            {
              "id": 2117794,
              "postDate": "2023-01-27T14:51:39.820Z",
              "content": "<p>Thanks for sharing, may I ask how should we create targets from test.csv ? Because we don't have validation B dataset anymore in train and test.</p>",
              "rawMarkdown": "Thanks for sharing, may I ask how should we create targets from test.csv ? Because we don't have validation B dataset anymore in train and test."
            }
          ]
        }
      ]
    },
    {
      "id": 2063686,
      "postDate": "2022-12-13T07:47:16.933Z",
      "content": "<p>How the candidates generated? Since the number of history items of each user may less than you want (50 candidates for example).</p>",
      "rawMarkdown": "How the candidates generated? Since the number of history items of each user may less than you want (50 candidates for example).",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2064129,
          "postDate": "2022-12-13T14:37:53.497Z",
          "content": "<p>You can create candidates in many ways. The most common methods are co-visitation matrices and matrix factorization. For example the following code finds 20 candidates that the user has not clicked where <code>unique_aids</code> is a list of what the user has already clicked. And <code>top_20_clicks</code> is a co-visitation matrix:</p>\n<pre><code>from collections import Counter\nimport itertools\naids = list(itertools.chain(*[top_20_clicks[aid] for aid in UNIQUE_AIDS if aid in top_20_clicks]))\nCANDIDATES = [aid for aid, cnt in Counter(aids).most_common(20) if aid not in UNIQUE_AIDS]\nCANDIDATES += UNIQUE_AIDS[:20]\n</code></pre>\n<p>Covisitation matrix notebook is <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">here</a> and matrix factorization is <a href=\"https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "You can create candidates in many ways. The most common methods are co-visitation matrices and matrix factorization. For example the following code finds 20 candidates that the user has not clicked where `unique_aids` is a list of what the user has already clicked. And `top_20_clicks` is a co-visitation matrix:\n\n    from collections import Counter\n    import itertools\n    aids = list(itertools.chain(*[top_20_clicks[aid] for aid in UNIQUE_AIDS if aid in top_20_clicks]))\n    CANDIDATES = [aid for aid, cnt in Counter(aids).most_common(20) if aid not in UNIQUE_AIDS]\n    CANDIDATES += UNIQUE_AIDS[:20]\n\nCovisitation matrix notebook is [here][1] and matrix factorization is [here][2]\n\n[1]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\n[2]: https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu",
          "votes": 3
        },
        {
          "id": 2064567,
          "postDate": "2022-12-13T22:48:41.663Z",
          "content": "<p>So the candidates should include the history aid of a session, and adding the, for example, co-visitation matrix to it. Generate 20 candidates means (history aids + co-visitation matrix aids) is 20, or  history aids + (20 co-visitations aids), I notice you are using <code>most_common(20)</code> to get 20 most frequent aids from the co-visitation matrix, but <code>not in UNIQUE_AIDS</code> means the the history aids are not included, so the <code>CANDIDATES</code> list may always less than 20? Should I add the history aids to the <code>CANDIDATES</code>?</p>",
          "rawMarkdown": "So the candidates should include the history aid of a session, and adding the, for example, co-visitation matrix to it. Generate 20 candidates means (history aids + co-visitation matrix aids) is 20, or  history aids + (20 co-visitations aids), I notice you are using `most_common(20)` to get 20 most frequent aids from the co-visitation matrix, but `not in UNIQUE_AIDS` means the the history aids are not included, so the `CANDIDATES` list may always less than 20? Should I add the history aids to the `CANDIDATES`?",
          "isDeleted": true
        },
        {
          "id": 2064570,
          "postDate": "2022-12-13T22:50:40.860Z",
          "content": "<p><a href=\"https://www.kaggle.com/cocoshe\" target=\"_blank\">@cocoshe</a> yes, you should add the history too. Either the entire history or the last 20 items. I updated my post above to take 20 history and 20 co-visitation for a total of 40. Some users may have less than 40 candidates but that is fine. XGB does not need each group to be the same size.</p>\n<p>Of course there is no right or wrong way of making candidates. The goal is to first choose a number of how many candidates that your pipeline can handle (due to speed, memory, etc). Then for each user, generate the best X candidates you can in any way possible.</p>",
          "rawMarkdown": "@cocoshe yes, you should add the history too. Either the entire history or the last 20 items. I updated my post above to take 20 history and 20 co-visitation for a total of 40. Some users may have less than 40 candidates but that is fine. XGB does not need each group to be the same size.\n\nOf course there is no right or wrong way of making candidates. The goal is to first choose a number of how many candidates that your pipeline can handle (due to speed, memory, etc). Then for each user, generate the best X candidates you can in any way possible."
        },
        {
          "id": 2064578,
          "postDate": "2022-12-13T23:10:21.467Z",
          "content": "<p>\"there is no right or wrong way of making candidates\", I agree! I just try to learn the method as a basic way in the future:)</p>\n<p>BTW, I try to generate co-visitation matrix with something like <code>top_50_buy2buy</code>, and change the code here</p>\n<pre><code>  # SAVE TOP 50\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    # tmp = tmp.loc[tmp.n&lt;15].drop('n',axis=1)\n    tmp = tmp.loc[tmp.n&lt;50].drop('n',axis=1)  ################\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    # tmp.to_pandas().to_parquet(f'top_15_buy2buy_v{VER}_{PART}.pqt')  #####################\n    tmp.to_pandas().to_parquet(f'top_50_buy2buy_v{VER}_{PART}.pqt')\n</code></pre>\n<p>and I got the pqt like this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8787751%2F380e1aae98b03216594cf6999ed0d4f7%2F9EKB8PIVN46V40OE8W31.png?generation=1670972603642854&amp;alt=media\" alt=\"\"></p>\n<p>the co-visitation matrix seems can't fill the 20 candidates, I'm wondering if it is reasonable?</p>",
          "rawMarkdown": "\"there is no right or wrong way of making candidates\", I agree! I just try to learn the method as a basic way in the future:)\n\nBTW, I try to generate co-visitation matrix with something like `top_50_buy2buy`, and change the code here\n```  \n  # SAVE TOP 50\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    # tmp = tmp.loc[tmp.n<15].drop('n',axis=1)\n    tmp = tmp.loc[tmp.n<50].drop('n',axis=1)  ################\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    # tmp.to_pandas().to_parquet(f'top_15_buy2buy_v{VER}_{PART}.pqt')  #####################\n    tmp.to_pandas().to_parquet(f'top_50_buy2buy_v{VER}_{PART}.pqt')\n```\nand I got the pqt like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8787751%2F380e1aae98b03216594cf6999ed0d4f7%2F9EKB8PIVN46V40OE8W31.png?generation=1670972603642854&alt=media)\n\nthe co-visitation matrix seems can't fill the 20 candidates, I'm wondering if it is reasonable?",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 2064582,
          "postDate": "2022-12-13T23:18:26.423Z",
          "content": "<p>yeah, it's reasonable. The <code>buy2buy</code> won't have 50 for each item because there are not enough items bought together in train data. I suggest using both <code>top_20_buys</code> and <code>top_20_buy2buy</code> similar to my notebook <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">here</a> when generating candidates for <code>carts</code> and <code>orders</code>. And use <code>top_20_clicks</code> when generating candidates for <code>clicks</code>. </p>\n<p>Next brainstorm more co-visitation matrices to add more candidates. Lastly, you can add logic to determine how to weight and select candidates from all your co-visitation matrices (once you make more). Because we need to limit our candidates to maximum size and we want the highest quality candidates in that maximum size.</p>",
          "rawMarkdown": "yeah, it's reasonable. The `buy2buy` won't have 50 for each item because there are not enough items bought together in train data. I suggest using both `top_20_buys` and `top_20_buy2buy` similar to my notebook [here][1] when generating candidates for `carts` and `orders`. And use `top_20_clicks` when generating candidates for `clicks`. \n\nNext brainstorm more co-visitation matrices to add more candidates. Lastly, you can add logic to determine how to weight and select candidates from all your co-visitation matrices (once you make more). Because we need to limit our candidates to maximum size and we want the highest quality candidates in that maximum size.\n\n[1]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575",
          "votes": 5
        },
        {
          "id": 2065261,
          "postDate": "2022-12-14T13:02:04.850Z",
          "content": "<p>Appreciate so much! That really helps!👍</p>",
          "rawMarkdown": "Appreciate so much! That really helps!👍",
          "isDeleted": true
        },
        {
          "id": 2065276,
          "postDate": "2022-12-14T13:30:20.267Z",
          "content": "<p>One more question, the co-visitation matrix is used based on history aids?<br>\nWhen I try to generate session-aid pairs, I only have the aid history of a session. For example, <code>aid A</code> is a history aid of <code>session B</code>, so I need to find <code>{aid A: [aid X, aid Y, aid Z, ...]}</code> in co-visitation matrix.<br>\nThen I get something like this:</p>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>aid</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>session B</td>\n<td>aid A</td>\n</tr>\n<tr>\n<td>session B</td>\n<td>aid X</td>\n</tr>\n<tr>\n<td>session B</td>\n<td>aid Y</td>\n</tr>\n<tr>\n<td>session B</td>\n<td>aid Z</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>\n<p>I want to check if my thought is right when generating session-aid pairs. If so, any </p>",
          "rawMarkdown": "One more question, the co-visitation matrix is used based on history aids?\nWhen I try to generate session-aid pairs, I only have the aid history of a session. For example, `aid A` is a history aid of `session B`, so I need to find `{aid A: [aid X, aid Y, aid Z, ...]}` in co-visitation matrix.\nThen I get something like this:\n| session| aid  |\n| --- | --- |\n| session B| aid A |\n| session B| aid X |\n| session B| aid Y |\n| session B| aid Z |\n| ...| ...|\n\nI want to check if my thought is right when generating session-aid pairs. If so, any ",
          "isDeleted": true
        },
        {
          "id": 2065279,
          "postDate": "2022-12-14T13:34:34.617Z",
          "content": "<p>This is correct but you need to add logic to reduce the number of candidates. For example, if a user has 30 items in their history and your co-visitation matrix maps one aid to twenty aids, then when you apply the co-visitation matrix to every history item, you will have 600 candidates. That is too many. So you need to use</p>\n<pre><code>aids = list(itertools.chain(*[top_20_clicks[aid] for aid in UNIQUE_AIDS if aid in top_20_clicks]))\nCANDIDATES = [aid for aid, cnt in Counter(aids).most_common(20) if aid not in UNIQUE_AIDS]\n</code></pre>\n<p>This finds the top 20 most common among the 600</p>",
          "rawMarkdown": "This is correct but you need to add logic to reduce the number of candidates. For example, if a user has 30 items in their history and your co-visitation matrix maps one aid to twenty aids, then when you apply the co-visitation matrix to every history item, you will have 600 candidates. That is too many. So you need to use\n\n    aids = list(itertools.chain(*[top_20_clicks[aid] for aid in UNIQUE_AIDS if aid in top_20_clicks]))\n    CANDIDATES = [aid for aid, cnt in Counter(aids).most_common(20) if aid not in UNIQUE_AIDS]\n\nThis finds the top 20 most common among the 600",
          "votes": 2
        },
        {
          "id": 2065302,
          "postDate": "2022-12-14T14:12:52.160Z",
          "content": "<p>OK, I get it</p>",
          "rawMarkdown": "OK, I get it",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 2065463,
          "postDate": "2022-12-14T17:21:01.990Z",
          "content": "<p>I'm also confused about the dataset, where are the sessions from? <br>\nI check the <code>sample_submission.csv</code> of the competition, it is <code>5015409</code> which means 5015409/3=1671803, there are <code>1671803</code> unique sessions. But the dataset <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation?select=test_labels.parquet\" target=\"_blank\">here</a>, the dataframes' length like this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8787751%2Fea814ba3110b28e2deb2bc4223150880%2F4Q5ZO(FB0FK62HWVW.png?generation=1671038380797767&amp;alt=media\" alt=\"\"></p>\n<p>I don't know how to explain it, and the <code>test_label.parquet</code>…</p>",
          "rawMarkdown": "I'm also confused about the dataset, where are the sessions from? \nI check the `sample_submission.csv` of the competition, it is `5015409` which means 5015409/3=1671803, there are `1671803` unique sessions. But the dataset [here](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation?select=test_labels.parquet), the dataframes' length like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8787751%2Fea814ba3110b28e2deb2bc4223150880%2F4Q5ZO(FB0FK62HWVW.png?generation=1671038380797767&alt=media)\n\nI don't know how to explain it, and the `test_label.parquet`...",
          "isDeleted": true
        },
        {
          "id": 2065492,
          "postDate": "2022-12-14T18:01:13.563Z",
          "content": "<p>The test data provided by Kaggle has <code>1671803</code> unique sessions which matches the sample submission provided by Kaggle. In your plot above, you are exploring the \"new test\" which is the first half of week 4 of the Kaggle train data. That file was produced by Radek and does not have <code>1671803</code> unique sessions. It has <code>1801251</code> unique sessions and matches the number of unique sessions in Radek's ground truth file.</p>",
          "rawMarkdown": "The test data provided by Kaggle has `1671803` unique sessions which matches the sample submission provided by Kaggle. In your plot above, you are exploring the \"new test\" which is the first half of week 4 of the Kaggle train data. That file was produced by Radek and does not have `1671803` unique sessions. It has `1801251` unique sessions and matches the number of unique sessions in Radek's ground truth file."
        },
        {
          "id": 2065877,
          "postDate": "2022-12-15T06:41:49.670Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8787751%2F5931f98e76262068454b65e1e191785c%2FIMG_2459(20221215-143541).PNG?generation=1671086367447603&amp;alt=media\" alt=\"\"></p>\n<p>Anything wrong about the dataset split?</p>\n<p>And, can we use the same co-visitation matrix when we train and infer? Here's my thought:<br>\nI read the code <a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook\" target=\"_blank\">https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook</a>, and find that in Radek's prediction part, he just use the <code>test.parquet</code> of <code>Otto Full Optimized Memory Footprint</code>, I checked the shape, it is the same as <strong>test dataset of kaggle</strong>, so there are history aids of each session, which also means there are usually less than 20 aids for every sessions' history. So we need to fill the candidates until more than 20 aids for each session. Since he use the co-visitation matrix to generate some candidates when training on <strong>validation A</strong> and <strong>validation B</strong>, can we still use the same candidates when we predict on the test dataset?</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8787751%2F5931f98e76262068454b65e1e191785c%2FIMG_2459(20221215-143541).PNG?generation=1671086367447603&alt=media)\n\nAnything wrong about the dataset split?\n\nAnd, can we use the same co-visitation matrix when we train and infer? Here's my thought:\nI read the code [https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook), and find that in Radek's prediction part, he just use the `test.parquet` of `Otto Full Optimized Memory Footprint`, I checked the shape, it is the same as **test dataset of kaggle**, so there are history aids of each session, which also means there are usually less than 20 aids for every sessions' history. So we need to fill the candidates until more than 20 aids for each session. Since he use the co-visitation matrix to generate some candidates when training on **validation A** and **validation B**, can we still use the same candidates when we predict on the test dataset?",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 2066019,
          "postDate": "2022-12-15T09:56:57.027Z",
          "content": "<p><a href=\"https://www.kaggle.com/cocoshe\" target=\"_blank\">@cocoshe</a>, I'm also splitted data like you. But when come to train, what is the correct weeks to choose to create candidates dataframes?<br>\nwhen come to validate. I wonder how to choose which week to create item-features? user-features?</p>",
          "rawMarkdown": "@cocoshe, I'm also splitted data like you. But when come to train, what is the correct weeks to choose to create candidates dataframes?\nwhen come to validate. I wonder how to choose which week to create item-features? user-features?",
          "votes": 1
        },
        {
          "id": 2066063,
          "postDate": "2022-12-15T11:22:54.010Z",
          "content": "<p><code>Create item features. Using our train data + valid data A (yes use test leak), we create item features in their own dataframe and save a parquet to disk. For example</code> in <strong>Step 3</strong>, and <code>Create user features. Using our validation data A, we create user features in their own dataframe and save a parquet to disk. For example</code> in <strong>Step 4</strong>. I think that's what you want. <br>\nIt's seems reasonable, because we are preparing for the train part, so during training, our goal is to predict scores for session-aid pairs, and we use <strong>validation A</strong> as the \"test dataset\", so when we need the user feature, we only need to prepare the sessions we need, which means the sessions in <strong>validation A</strong>, for example: when we prepare the trian part,</p>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>aid</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>a</td>\n</tr>\n<tr>\n<td>A</td>\n<td>aa</td>\n</tr>\n<tr>\n<td>A</td>\n<td>aaa</td>\n</tr>\n<tr>\n<td>B</td>\n<td>b</td>\n</tr>\n<tr>\n<td>C</td>\n<td>c</td>\n</tr>\n<tr>\n<td>D</td>\n<td>d</td>\n</tr>\n</tbody>\n</table>\n<p>A,B,C,D are from <strong>validation A</strong>, also the <code>first part of Week 4</code>(I think we use the second part as label, also the <code>gt = 1</code>, and we can merge the label to <strong>validation A</strong> on <code>session</code> and <code>aid</code>), and the aids are your candidates from the co-visitation matrix maybe.<br>\nBTW, when we prepare the inference, we choose session in <strong>test dataset of kaggle</strong>, that's for submission! So we only need to prepare the session features in <strong>test dataset of kaggle</strong>, also the <code>Week 5</code><br>\nThen the features, just like <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> said in step 3 and step 4, we should get the features like this:</p>\n<ul>\n<li>user features</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>session_feature1</th>\n<th>session_feature2</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>a</td>\n<td>a</td>\n</tr>\n<tr>\n<td>A</td>\n<td>aa</td>\n<td>aa</td>\n</tr>\n<tr>\n<td>A</td>\n<td>aaa</td>\n<td>aaa</td>\n</tr>\n<tr>\n<td>B</td>\n<td>b</td>\n<td>b</td>\n</tr>\n<tr>\n<td>C</td>\n<td>c</td>\n<td>c</td>\n</tr>\n<tr>\n<td>D</td>\n<td>d</td>\n<td>d</td>\n</tr>\n</tbody>\n</table>\n<p>Then we use <code>merge</code> on <code>session</code> to add user features, so the item features are the same, just merge on <code>aid</code>.</p>\n<p>That's all I thought now, if anything wrong, please let me know!</p>",
          "rawMarkdown": "`Create item features. Using our train data + valid data A (yes use test leak), we create item features in their own dataframe and save a parquet to disk. For example` in **Step 3**, and `Create user features. Using our validation data A, we create user features in their own dataframe and save a parquet to disk. For example` in **Step 4**. I think that's what you want. \nIt's seems reasonable, because we are preparing for the train part, so during training, our goal is to predict scores for session-aid pairs, and we use **validation A** as the \"test dataset\", so when we need the user feature, we only need to prepare the sessions we need, which means the sessions in **validation A**, for example: when we prepare the trian part,\n| session | aid |\n| --- | --- |\n| A | a |\n| A | aa |\n| A | aaa |\n| B | b |\n| C | c |\n| D | d |\n\nA,B,C,D are from **validation A**, also the `first part of Week 4`(I think we use the second part as label, also the `gt = 1`, and we can merge the label to **validation A** on `session` and `aid`), and the aids are your candidates from the co-visitation matrix maybe.\nBTW, when we prepare the inference, we choose session in **test dataset of kaggle**, that's for submission! So we only need to prepare the session features in **test dataset of kaggle**, also the `Week 5`\nThen the features, just like @cdeotte said in step 3 and step 4, we should get the features like this:\n+ user features\n\n| session| session_feature1 | session_feature2 |\n| --- | --- | --- |\n| A | a | a |\n| A | aa | aa |\n| A | aaa | aaa |\n| B | b | b |\n| C | c | c |\n| D | d | d |\n\nThen we use `merge` on `session` to add user features, so the item features are the same, just merge on `aid`.\n\nThat's all I thought now, if anything wrong, please let me know!",
          "isDeleted": true
        },
        {
          "id": 2066248,
          "postDate": "2022-12-15T14:04:33.777Z",
          "content": "<p><a href=\"https://www.kaggle.com/cocoshe\" target=\"_blank\">@cocoshe</a> nice diagram. However your diagram is missing the leaderboard test data. Here is an analogy: \"validation A is to Kaggle test as validation B is to Kaggle leaderboard\". Both Kaggle test and Kaggle leaderboard come from week 5 data. For each user in week 5, their sequence of activity is randomly split (using <code>np.random.uniform(1, len(user_activitiy) )</code> ) and the first half of each user is in Kaggle test and the second half of each user is in Kaggle leaderboard. This is the same way that Radek made validation A and validation B from the last week of Kaggle train.</p>",
          "rawMarkdown": "@cocoshe nice diagram. However your diagram is missing the leaderboard test data. Here is an analogy: \"validation A is to Kaggle test as validation B is to Kaggle leaderboard\". Both Kaggle test and Kaggle leaderboard come from week 5 data. For each user in week 5, their sequence of activity is randomly split (using `np.random.uniform(1, len(user_activitiy) )` ) and the first half of each user is in Kaggle test and the second half of each user is in Kaggle leaderboard. This is the same way that Radek made validation A and validation B from the last week of Kaggle train.",
          "replies": [
            {
              "id": 2074225,
              "postDate": "2022-12-23T20:56:35.600Z",
              "content": "<blockquote>\n  <p>Here is an analogy: \"validation A is to Kaggle test as validation B is to Kaggle leaderboard\". </p>\n</blockquote>\n<p>Could you explain the difference between kaggle test and kaggle leaderboard? Kaggle test would be an f1-metric or something like this? </p>",
              "rawMarkdown": "> Here is an analogy: \"validation A is to Kaggle test as validation B is to Kaggle leaderboard\". \n\n\nCould you explain the difference between kaggle test and kaggle leaderboard? Kaggle test would be an f1-metric or something like this? "
            }
          ]
        },
        {
          "id": 2066253,
          "postDate": "2022-12-15T14:11:57.620Z",
          "content": "<p><a href=\"https://www.kaggle.com/locbaop\" target=\"_blank\">@locbaop</a> </p>\n<blockquote>\n  <p>I'm also splitted data like you. But when come to train, what is the correct weeks to choose to create candidates dataframes?<br>\n  when come to validate. I wonder how to choose which week to create item-features? user-features?</p>\n</blockquote>\n<p>During training we use \"new train\" + \"validation A\" to make features for items and users (and ground truth comes from \"validation B\"). And during inference we use \"kaggle train\" + \"kaggle test\" to make features for items and users (and ground truth comes from \"kaggle LB\"). Note that cocoshe's diagram is missing the \"kaggle LB\" which is the second half of week 5.</p>\n<p>(Here is an advanced comment ignore this if it is confusing: We also note that the users in \"new train\" and \"validation A\" do not overlap. And the users in \"new train\" and \"validation B\" do not overlap. So when making user features, we use \"validation A\" for supervised user features and we use \"new train\"+\"validation A\" for unsupervised user features. Our train dataframe of candidates will only have users from \"validation A\"+\"validation B\", so we do not make candidates for the users in \"new train\").</p>",
          "rawMarkdown": "@locbaop \n>I'm also splitted data like you. But when come to train, what is the correct weeks to choose to create candidates dataframes?\nwhen come to validate. I wonder how to choose which week to create item-features? user-features?\n\nDuring training we use \"new train\" + \"validation A\" to make features for items and users (and ground truth comes from \"validation B\"). And during inference we use \"kaggle train\" + \"kaggle test\" to make features for items and users (and ground truth comes from \"kaggle LB\"). Note that cocoshe's diagram is missing the \"kaggle LB\" which is the second half of week 5.\n\n(Here is an advanced comment ignore this if it is confusing: We also note that the users in \"new train\" and \"validation A\" do not overlap. And the users in \"new train\" and \"validation B\" do not overlap. So when making user features, we use \"validation A\" for supervised user features and we use \"new train\"+\"validation A\" for unsupervised user features. Our train dataframe of candidates will only have users from \"validation A\"+\"validation B\", so we do not make candidates for the users in \"new train\").",
          "votes": 1
        },
        {
          "id": 2066435,
          "postDate": "2022-12-15T17:34:17.113Z",
          "content": "<p>Strange about the <code>Both Kaggle test and Kaggle leaderboard come from week 5 data</code> you said. The <code>submission.csv</code> should contain <code>1671803</code> unique sessions, and they are all come from the test dataset on kaggle, also the Week 5 activities, if you split the Week 5 as you did to the Week 4, we can still get the same number of unique sessions?</p>\n<p>So in <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> notebook, I think he just read all the test data on kaggle when predicting, and do inference to select aid only among history, which also cause the aid number of each session is always less than 20.</p>",
          "rawMarkdown": "Strange about the `Both Kaggle test and Kaggle leaderboard come from week 5 data` you said. The `submission.csv` should contain `1671803` unique sessions, and they are all come from the test dataset on kaggle, also the Week 5 activities, if you split the Week 5 as you did to the Week 4, we can still get the same number of unique sessions?\n\nSo in @radek1 notebook, I think he just read all the test data on kaggle when predicting, and do inference to select aid only among history, which also cause the aid number of each session is always less than 20.",
          "isDeleted": true,
          "replies": [
            {
              "id": 2074224,
              "postDate": "2022-12-23T20:56:10.690Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2054040,
      "postDate": "2022-12-03T20:10:50.453Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2112772,
      "postDate": "2023-01-23T20:24:07.050Z",
      "content": "<p>Thanks, it really helps</p>",
      "rawMarkdown": "Thanks, it really helps",
      "votes": 1
    },
    {
      "id": 2066728,
      "postDate": "2022-12-16T02:43:56.837Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 2107408,
      "postDate": "2023-01-19T19:06:36.770Z",
      "content": "<p>Thank you for sharing this algorithm!</p>",
      "rawMarkdown": "Thank you for sharing this algorithm!"
    },
    {
      "id": 2097356,
      "postDate": "2023-01-12T16:00:52.487Z",
      "content": "<p>Thanks for sharing ~</p>",
      "rawMarkdown": "Thanks for sharing ~\n"
    },
    {
      "id": 2065506,
      "postDate": "2022-12-14T18:21:34.210Z",
      "content": "<p>Thanks for your explanation.</p>",
      "rawMarkdown": "Thanks for your explanation."
    }
  ],
  "comments": [
    {
      "id": 2059518,
      "author_name": "xarispan",
      "author_url": "",
      "post_date": "2022-12-09T00:47:14.520000",
      "content": "<p>Will this work if our candidates for each session have mixed length? For one session we may have 25 candidates, for another one 30.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2059527,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-09T01:20:53.497000",
          "content": "<p>Yes, it will work. You just need to sort your dataframe by session and provide XGB with a list of group sizes in the order they appear in the dataframe.</p>\n<p>For example, sort your dataframe with <code>train = train.sort_values('session')</code>. Then provide a list of sizes of the sessions in order, for example <code>groups = train.groupby('session').aid.agg('count').values</code>. (Double check to make sure your dataframe library keeps the groups in the same order they appear in dataframe when doing groupby. I believe Pandas does). Then train with </p>\n<pre><code>dtrain = xgb.DMatrix(train[FEATS], train[TAR], group=groups )\n</code></pre>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 2059937,
          "author_name": "Andrej Zubaľ",
          "author_url": "",
          "post_date": "2022-12-09T11:35:24.610000",
          "content": "<p>I also have different number of candidates for each session.</p>\n<p>In my case I managed to make it work by using qid parameter instead of groups like this:</p>\n<p><code>\ndf=df.sort_values(by='session_id', ascending=True)\n</code><br>\n<code>dtrain = xgb.DMatrix(df[col_preds], df[col_target], qid=list(df['session_id']) ) \n</code></p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 2059954,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-09T11:54:33.910000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2063532,
          "author_name": "HinePo",
          "author_url": "",
          "post_date": "2022-12-13T03:24:51.103000",
          "content": "<p>For those working on google colab, xgb default version is 0.90, and does not have neither group or qid parameters.</p>\n<p>I upgraded it to version 1.7.2 and it worked.</p>\n<pre><code>! pip uninstall xgboost -qq -y\n! pip install xgboost==1.7.2 -qq\n</code></pre>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2058286,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2022-12-07T19:15:00.997000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> A new functionality of Cudf (since 22.08) is setting the default data type to be 32 or 64 bit. Using 32 bit can help save GPU memory when processing features.<br>\nThe code bellow sets 32 bit as the default:</p>\n<pre><code>import cudf\ncudf.set_option(\"default_integer_bitwidth\", 32)\ncudf.set_option(\"default_float_bitwidth\", 32)\n</code></pre>",
      "votes": 6,
      "replies": [
        {
          "id": 2058292,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-07T19:21:22.653000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> I added this suggestion to the data management section of my post above.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2078669,
      "author_name": "Nick",
      "author_url": "",
      "post_date": "2022-12-28T14:12:18.810000",
      "content": "<p>I can't understand what is happening, I'm keep getting this error: <strong><em>XGBoostError: [13:58:22] ../src/data/data.cc:694: Check failed: group_ptr_.back() == num_row_ (9624929 vs. 7699943) : Invalid group structure.  Number of rows obtained from groups doesn't equal to actual number of rows given by data.</em></strong> </p>\n<p>I read my dataframe from the disk:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F9f5888117f112e76546cd986da22172d%2FReading.JPG?generation=1672236414247488&amp;alt=media\" alt=\"\"></p>\n<p>Than I reduce the number of negatives like Chris proposed in his post:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F0431b5f064256829fe7638a54064f3c3%2FReducing.JPG?generation=1672236479917480&amp;alt=media\" alt=\"\"></p>\n<p>I check if my dataframe's length equals to the sum of the array with group counts:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F8692369e6eaeac04d2e2d5ca24b4d649%2FTrain.JPG?generation=1672236560514962&amp;alt=media\" alt=\"\"></p>\n<p>But I'm keep getting this error:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2Fca5323843c148064437fe919bf3a84b8%2FError_1.JPG?generation=1672236658532920&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2Fdc14b53811624616ca61d9db2ca45b27%2FError_2.JPG?generation=1672236667732271&amp;alt=media\" alt=\"\"></p>\n<p>Has anybody encountered anything similar to this? And if encountered what was the root cause?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2078777,
          "author_name": "Nick",
          "author_url": "",
          "post_date": "2022-12-28T16:25:28.343000",
          "content": "<p>The problem was with the <strong><em>train_groups</em></strong> and <strong><em>valid_groups</em></strong>. I'm now getting group sizes for train groups and validation groups in this way:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2Faeb5851bf0f0eb4ee8a5414dad4d1470%2FProblem%20Solved.JPG?generation=1672244683733045&amp;alt=media\" alt=\"\"></p>\n<p>This code solved the problem.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2078951,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2022-12-28T19:37:25.203000",
              "content": "<p>Looks good</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2080659,
              "author_name": "adairli00",
              "author_url": "",
              "post_date": "2022-12-30T11:00:17.797000",
              "content": "<p>hi,you have quite solved my problem about  the invalid number rows.but i have some new questions about you code.can you tell me your idea? the question 1 : where you get the dateset \"the_clicks_target_dataset.prt\" question 2 : your ' candidate ' is represent ' aids'? i am so confuse about it , if you have see it , please answer me , thanks!!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2080977,
              "author_name": "Nick",
              "author_url": "",
              "post_date": "2022-12-30T17:21:14.793000",
              "content": "<p>Hi, adairli00!</p>\n<ol>\n<li><em>clicks_target_dataset.parquet</em> - is a dataset after <strong>Step 6</strong> from Chris's post. It's a dataset of shape (#session x 50, #features) where one row is: session, aid, session features, aid features, session aid features and clciks target (0 or 1)</li>\n<li>You're right. Candiate is an aid, which I generated from co-visitation matrices for each session</li>\n</ol>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2061533,
      "author_name": "Piyush Chauhan",
      "author_url": "",
      "post_date": "2022-12-11T08:03:33.400000",
      "content": "<p>Hi Chris. Thanks a lot for this framework!   I was wondering that while merging candidate dataframe with the target,  should we do outer join instead of left, because on doing left join we might be missing out on some of the true (user, item) pairs which may not be in our candidate dataset<br>\n <code>candidates = candidates.merge(click_target,on=['user','item'],how='left').fillna(0)</code></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2061875,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-11T14:40:53.920000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kazama28\" target=\"_blank\">@kazama28</a> We want to use <code>how='left</code> because we want missing values and then convert them to zero with <code>fillna(0)</code>. These are the negative examples. To train our model we need both <code>target = 1</code> and <code>target = 0</code>. Without target = 0, our model cannot learn.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2069772,
              "author_name": "NaN",
              "author_url": "",
              "post_date": "2022-12-19T09:59:36.923000",
              "content": "<p>But outer join (<code>how='outer'</code>) also preserve missing values, isn't it? As <a href=\"https://www.kaggle.com/kazama28\" target=\"_blank\">@kazama28</a> pointed out, by doing left join we would miss out groundtruths that are not in candidate set. For <code>type=click</code>, each <code>user</code> only has a single groundtruth, so if our candidate set miss that, we are left with negative examples only.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2074207,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2022-12-23T20:46:17.927000",
              "content": "<p>Good point. We can try both and see what produces better CV and better LB</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2059767,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "2022-12-09T07:47:35.763000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I have tried to implement the pipeline as you described above. I have managed to generate dataframes of candidates and features (items, users). However, when I tried to merge them, but I couldn't get the columns right in the merged dataframe as you described above. Could you have a look? Thanks a lot! 🙏</p>\n<p>This is your demo for merged dataframe</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fcf82ca1b474801c28527aabdc9d426e7%2FScreen%20Shot%202022-12-09%20at%2015.45.22.png?generation=1670571940780500&amp;alt=media\" alt=\"\"></p>\n<p>My candidates and item_features dataframes look like below</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Ff37405f6569eb49d42657643bfbabf12%2FScreen%20Shot%202022-12-09%20at%2015.37.35.png?generation=1670571720150971&amp;alt=media\" alt=\"\"></p>\n<p>This is the merged dataframe I got which is confusing to me.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fdf620b7d3ffb9545defd25ed136e6068%2FScreen%20Shot%202022-12-09%20at%2015.37.47.png?generation=1670571846965756&amp;alt=media\" alt=\"\"></p>\n<p>I have a feeling that maybe the problem lies on index, but tried to read the docs on <code>merge</code> and <code>set_index</code>, but have not figured out how to fix it.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2059625,
      "author_name": "Jensen",
      "author_url": "",
      "post_date": "2022-12-09T04:44:12.390000",
      "content": "<p>Hi,Chris,your answer helped me a lot!!!<br>\nStep 4: I have a question. You use as an example whether an item is clicked as a feature of user-item interaction. But if I train a ranker to sort the candidates in click. Does this feature above leak my training label? That means that if I want to train a click reranker, I can't use both the user and item interaction features related to click? I am very confused about this and look forward to your answer.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2059914,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-09T10:57:40.477000",
          "content": "<p><a href=\"https://www.kaggle.com/niejianfei\" target=\"_blank\">@niejianfei</a> there is no leak because we do not use ground truth to create interaction features. We begin with the 4 weeks of Kaggle train data. Then we build 3 datasets</p>\n<ul>\n<li>new train data (first 3 weeks)</li>\n<li>validation data A (first half of week 4)</li>\n<li>validation data B (second half of week 4)</li>\n</ul>\n<p>The ground truth is validation data B. We build interaction features from validation data A. Each of the test users activity is split in half. The first half of their activity is in Val A and the second half in Val B. For example, user A clicks 001 on monday, clicks 003 on Tuesday, clicks 002 on Wednesday and clicks 001 on Thursday. The Val A is <code>clicks = [001, 003]</code> from monday and tuesday. And Val B is <code>clicks = [002, 001]</code> from Wednesday and Thursday. Our interaction features are user A clicked 001 and 003. This is information from Monday and Tuesday. This will help us predict Wednesday and Thursday. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2059915,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-09T11:02:01.993000",
          "content": "<p>Note these 3 datasets we make are comparable to when we make a submission to LB. For submissions, we have (1) kaggle train (first 4 weeks) (2) kaggle test (first half of week 5) (3) Leaderboard (second half of week 5)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2059960,
          "author_name": "Jensen",
          "author_url": "",
          "post_date": "2022-12-09T12:06:52.090000",
          "content": "<p>Chris, thank you very much. I get it now.I am a novice in data science competition, thank you so much for sharing!!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2069145,
      "author_name": "aldparis",
      "author_url": "",
      "post_date": "2022-12-18T16:44:24.107000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>,<br>\nFirst at all, thank you very much for this post and all the others in this complicated competition ! <br>\nI have a question : in \"Training\", you suggested to consider downsampling negatives ; but if we do so, we will not have 50 consecutive samples from the same user and groups won't be correct for XGB. Do you agree ?</p>\n<p>(I'm looking for an issue to solve memory problems, and I tell myself that I will not consider downsampling negatives, but features forward selection. Or maybe I should consider downsampling negatives, but take care to \"group=\" in xgb.DMatrix)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2069203,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-18T18:05:56.417000",
          "content": "<p>When training GBT rankers, we do not need to have the same number of groups per user. So after downsampling, we count how many candidates are in each user and input this updated list to \"group=\" in xgb.DMatrix. </p>\n<p>Note there are alternatives besides downsample to use less GPU RAM. Here are two more ideas:</p>\n<ul>\n<li>train with a subset of users (and 50 candidates each) each fold. And validate and infer with all users (like normal).</li>\n<li>use XGB dataloader together with <code>DeviceQuantileDMatrix</code>. This allows XGB to use less memory and train with larger datasets. Also consider training a fixed number of iterations and do not use a valid dataset nor early stopping during training (to maximize acceptable train data size). Sample code <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a></li>\n</ul>",
          "votes": 7,
          "replies": [
            {
              "id": 2069282,
              "author_name": "aldparis",
              "author_url": "",
              "post_date": "2022-12-18T20:15:14.903000",
              "content": "<p>Thank you for confirmation and alternatives !</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2069307,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2022-12-18T20:50:43.717000",
              "content": "<p>I added a comment to my original post reminding people to update group sizes in DMatrix if they downsample. Thanks for pointing this out.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2073840,
              "author_name": "sapling",
              "author_url": "",
              "post_date": "2022-12-23T13:14:17.713000",
              "content": "<p>Inspite of downsampling 20x, there is memory error. <br>\nI have 10 chunks of train data which I am concatenating into dask dataframe. </p>\n<p>I am also trying dask_cudf with multiple GPUs as shown <a href=\"https://developer.nvidia.com/blog/accelerating-xgboost-on-gpu-clusters-with-dask/\" target=\"_blank\">here</a> but still facing memory error. </p>\n<p>I am at a loss how to deal with this much scale. I am a novice so requesting for some tips.  </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2073865,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2022-12-23T13:48:35.953000",
              "content": "<p><a href=\"https://www.kaggle.com/sohamdats\" target=\"_blank\">@sohamdats</a> Do you have memory error when creating your dataframe, or do you have memory error when training XGB?</p>\n<p>I recommend creating your dataframe in one python script or jupyter notebook, then save dataframe to disk as parquet and make sure to reduce every column to least dtype (like float32). Next stop all code and clear all memory. Then start a new python script or jupyter notebook, read in the parquet and begin training XGB.</p>\n<p>When training XGB, use DeviceQuantileDMatrix . Train for a fixed number of iterations without validation data. So you will only give your XGB train data. Are you training with dask-xgb and using all your GPU during training or are you using 1 GPU with regular XGB?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2073924,
              "author_name": "sapling",
              "author_url": "",
              "post_date": "2022-12-23T14:48:36.897000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> First of all thanks a lot for replying.</p>\n<p>I am facing memory error during training. I have saved 10 chunks in parquet format and loading them in a separate notebook for training. I am putting all of them together into a dask dataframe. I am reducing all data types.  </p>\n<p>First I tried to run the <a href=\"https://www.kaggle.com/code/radek1/training-an-xgboost-ranker-on-the-gpu/comments\" target=\"_blank\">Radek's notebook</a>  which threw a memory error during the ranker.fit stage. </p>\n<p>Then tried to run the <a href=\"https://developer.nvidia.com/blog/accelerating-xgboost-on-gpu-clusters-with-dask/\" target=\"_blank\">code</a> here. It is getting stuck in the load dataset method. This code uses DaskDeviceQuantileDMatrix but error is occuring before reaching there. </p>\n<p>I put n_workers=1 because &gt;1 was throwing error in kaggle notebook in GPU P100.  </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2073930,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2022-12-23T14:53:18.690000",
              "content": "<p>If you are using dask-cudf and/or dask-xgb in Kaggle notebooks, then you should use 2xT4 GPU then you will have 2x16 = 32GB or GPU RAM. With 1xP100, you only have 16GB. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2073932,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2022-12-23T14:55:38.767000",
              "content": "<p>Note if you already have your train data in multiple parquets on disk, then you do not need to use my <code>DeviceQuantileDMatrix</code> dataloader. If you search the internet for tutorials, you will see that <code>DeviceQuantileDMatrix</code> can load multiple parquets directly from disk.</p>\n<p>The purpose of my dataloader is to break a single dataframe into chunks because <code>DeviceQuantileDMatrix</code> needs chunks. However if you already have chunks on disk, then perhaps reading from disk is better.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2073935,
              "author_name": "sapling",
              "author_url": "",
              "post_date": "2022-12-23T14:58:04.037000",
              "content": "<p>Okay thanks a lot. I forgot that DeviceQuantileMatrix is a dataloader and I can read from disk itself. I will certainly try this. I can't thank you more. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2073952,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2022-12-23T15:19:01.763000",
              "content": "<p><a href=\"https://www.kaggle.com/sohamdats\" target=\"_blank\">@sohamdats</a> i just searched internet for info about reading from disk. I think the way to read from disk is to use my dataloader but change the \"next function\". My current dataloader from my code <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a> is like this</p>\n<pre><code>def next(self, input_data):\n    '''Yield next batch of data.'''\n    if self.it == self.batches:\n        return 0 # Return 0 when there's no more batch.\n\n    a = self.it * self.batch_size\n    b = min( (self.it + 1) * self.batch_size, len(self.df) )\n    dt = cudf.DataFrame(self.df.iloc[a:b])\n    input_data(data=dt[self.features], label=dt[self.target]) \n    self.it += 1\n    return 1\n</code></pre>\n<p>We notice that for each new request in the form of <code>self.it</code> we take a chunk from the dataframe in memory. Instead we can have it read a different parquet from disk for each new <code>self.it</code>. Something like the following</p>\n<pre><code>def next(self, input_data):\n    '''Yield next batch of data.'''\n    if self.it == self.batches:\n        return 0 # Return 0 when there's no more batch.\n\n    dt = cudf.read_parquet(f'train_{self.it}.parquet')\n    input_data(data=dt[self.features], label=dt[self.target]) \n    self.it += 1\n    return 1\n</code></pre>\n<p>I haven't tested this out, but perhaps reading from disk will use less memory than reading from dataframe in memory. Perhaps there is a way for <code>DeviceQuantileDMatrix</code> to read directly from disk without a dataloader, but i am not sure how to do that.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2075137,
              "author_name": "sapling",
              "author_url": "",
              "post_date": "2022-12-25T05:55:05.583000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks a lot for taking out the time and replying. I am going to try this from your notebook as you suggested.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2081885,
              "author_name": "Piyush Chauhan",
              "author_url": "",
              "post_date": "2022-12-31T19:16:13.220000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I am confused about how we should specify the groups when using DeviceQuantileDMatrix, should it be specified in the \"next function\" (as a 'group' parameter to input_data)? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2081890,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2022-12-31T19:24:15.130000",
              "content": "<p><a href=\"https://www.kaggle.com/kazama28\" target=\"_blank\">@kazama28</a> The code is like this</p>\n<pre><code>    dtrain = xgb.DeviceQuantileDMatrix(Xy_train, max_bin=256)\n    dtrain.set_group( [50] * ( len(train_idx)//50) )\n</code></pre>\n<p>Note that <code>set_group</code> also works for normal DMatrix</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2082103,
              "author_name": "Piyush Chauhan",
              "author_url": "",
              "post_date": "2023-01-01T05:33:55.047000",
              "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2054471,
      "author_name": "ningyuwhut",
      "author_url": "",
      "post_date": "2022-12-04T07:03:44.707000",
      "content": "<p>Thanks for sharing， I have a question that's why step 3 and step 4 only use the validation data  but not the train data and the validation data.  In the Inference step it seems to use similar set up which use 4week Kaggle train plus 1 week of Kaggle test to generate item feature but only use Kaggle test to generate user features. Could you explain it a bit ? Thanks!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2054781,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-04T13:04:18.230000",
          "content": "<p>Great question. When making features we would like to use the most amount of data as possible. But…</p>\n<p>The reason that step 3 and 4 only use validation data is because there is no overlap between train users and valid users. We only have targets for validation data and we only generate candidates for validation users. If we created user features from train users, none of these features would get merged to our candidate dataframe and none would be used.</p>\n<p>If you use unsupervised learning (which is not explained in this discussion post) which does not require targets, then you can use information about train users in addition to information about validation users. For example, run PCA on all users (from train and valid). Then create features from the PCA data. Then when adding features to our candidate list, first use PCA on validation users, then merge PCA features built from all users.</p>\n<p>NOTE: I updated my discussion post to fix step 1. In step 1, we only generate candidates for the users in validation data. (Because we only have targets for validation users)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 2055536,
          "author_name": "dpalbrecht",
          "author_url": "",
          "post_date": "2022-12-05T06:34:03.147000",
          "content": "<p>I am also confused by this point… I understand why you have no need for user features in the train set during test set inference, but then what user features do you train on if they're all from the validation set? Or do you mean you're using the validation set for training?</p>\n<p>Perhaps my confusion (and maybe others') would be cleared up with a small, worked example. My initial confusion was when the user and item features use the <code>train</code> dataframe, but the text states to only use validation users. So then I wondered what that meant, for example. If you could import the data and run a very small example, end-to-end, it would be so helpful. Personally, I like to look at the data at each step and break things to understand the process. I completely understand if you'd rather not, however, and appreciate the work you've put in thus far 😊</p>\n<p>EDIT: I may have wrapped my head around it. You use the train and validation set in that the train is used for item features. You then exclusively use the validation set for user features. You then have user features, item features, and candidates for the validation set (which is now effectively the training set) and add on labels from the ground truth. The positive candidates are the correct candidates while the negative are the incorrect.</p>\n<p>In the spirit of sharing what I was trying as well: I was using the train/validation set as the ground truth and generating hard negative samples for each label, which, as you might imagine, can also be pretty memory intensive. I like your approach in that you're training using the candidates you specifically generated as opposed to random samples. I've tried this approach at work, after having the idea, and found that hard negative sampling performed better for my specific data. I hadn't read any literature on this method either, honestly, so I figured it was just an odd idea. I wonder what the winner will be in this case!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2055851,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-05T13:46:01.010000",
          "content": "<p><a href=\"https://www.kaggle.com/dpalbrecht\" target=\"_blank\">@dpalbrecht</a> I agree where features come from is weird because there is <strong>no overlap</strong> between test users and train users. Note this is <strong>not the typical</strong> recommender system problem. Usually when building a recommender system model, we have lots of history on most users. In this competition we have <strong>no history</strong> for test users. Every test user is a \"cold start\" new user!</p>\n<p>The only long history we have is items. Therefore we can build item features from all our data. However when it comes to test users, we don't have their history so we cannot make features from early data. We can only make features from when these users engaged at Otto (i.e. the validation data and/or test data)</p>\n<p>Note that there are many ways to build unsupervised user features using PCA, auto encoder etc. For example, you can train an auto encoder on all user data (for all 5 weeks of data). Then give every user their embedding as a feature (during training and inference). This is a way to use all the user data if you want.</p>\n<p>You can also build a second fold (i.e \"time fold\"). Currently we are using one fold. (1) We train on weeks 1,2,3 and validate on week 4 of train. You could make a second fold where you (2) train on weeks 1,2 and validate on week 3. For the second fold, we would generate user features from the 3rd week of train data. This would be another way to use more users information during training.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2057381,
          "author_name": "dpalbrecht",
          "author_url": "",
          "post_date": "2022-12-07T04:00:57.003000",
          "content": "<p>Yeah, this formulation is feeling very odd to me for that reason. I'll have to try the PCA or autoencoder and time fold methods - those sound interesting. Thanks again for sharing!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2058478,
          "author_name": "Alvin.ai",
          "author_url": "",
          "post_date": "2022-12-08T01:06:27.830000",
          "content": "<p>Hi, Chris <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , I got the similar question as <a href=\"https://www.kaggle.com/dpalbrecht\" target=\"_blank\">@dpalbrecht</a>. I saw you made:</p>\n<ol>\n<li>train data: week 1-3 .</li>\n<li>valid data: week 4.</li>\n<li>valid label: week 4.</li>\n<li>test data: week5.</li>\n<li>test unseem label: after week 5.</li>\n</ol>\n<p>You used valid data from week 4 to make features and then predicted the valid label from week 4. It there a data leakage? Because valid data has already contains the information of valid label, which likes using week 4 data to predict week 4 data. However, we need to use week 5 test data to predict the future interacted aids on the week6 or later week.</p>\n<p>I got you that \"there is no session overlap between train data and valid data\" so you use valid data to make item features, but I am worrying the data leakage which bring us overfitting. Please let me know if I am right. By the way, making session embeddings by PCA or AE is a good idea to follow.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2058486,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-08T01:17:02.917000",
          "content": "<p>Thanks for asking this question. I need to clarify my post better. There are 5 datasets. Week 4 is split into valid data and ground truth. And public LB are ground truth from week 5. I need to find a clear way to explain this in my post. Maybe i'll add a  picture</p>\n<ol>\n<li>train data week 1-3</li>\n<li>valid data week 4 without ground truth (i.e. random first half of user activity week 4)</li>\n<li>valid data week 4 ground truth (i.e. random second half of user activity week 4)</li>\n<li>test data week 5 without ground truth (i.e. random first half of user activity week 5)</li>\n<li>test data week 5 ground truth (i.e. random second half of user activity week 5)</li>\n</ol>\n<p>All user data from Week 4 and week 5 is randomly split with <code>x = np.random.randint(1, 'user_number_of_events' )</code>. Then valid week 4 without ground truth is <code>user_events[:x]</code> and week  4 ground truth is <code>user_events[x:]</code> where <code>user_events</code> are in time order. I'll add a diagram to my post to clarify.</p>\n<p>We use 1+2 to make features during train. And we use 1+2+3+4 to make features during submission to LB. We do not use the valid ground truth #3 when building features to train our models.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 2058488,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-08T01:22:13.170000",
          "content": "<p><a href=\"https://www.kaggle.com/alvinai9603\" target=\"_blank\">@alvinai9603</a> I updated my discussion to use <code>validation data A</code> and <code>validation data B</code>. Let me know if it is clearer now.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2058515,
          "author_name": "Alvin.ai",
          "author_url": "",
          "post_date": "2022-12-08T01:56:27.160000",
          "content": "<p>Nice! It is much clear and there is not any data leakage as you said above. Thanks for sharing, Chris. 👍</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2123349,
      "author_name": "deanpp",
      "author_url": "",
      "post_date": "2023-01-31T13:18:27.057000",
      "content": "<p>It works, thx!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2114447,
      "author_name": "Abid Hasan",
      "author_url": "",
      "post_date": "2023-01-25T01:53:11.867000",
      "content": "<p>Hi Chris, Thanks for sharing. </p>\n<p>I am having trouble with memory management: I split the candidate dataframe into chunks as you suggested, but I still seem to run out of memory while loading the chunks.</p>\n<p>I used this code to split the dataframe</p>\n<pre><code>CHUNKS = 10\n\nchunk_size = int(np.ceil( len(candidates) / CHUNKS))\n\nfor k in range(CHUNKS):\n\n    df = candidates.iloc[k*chunk_size:(k+1)*chunk_size].copy()\n\n    df = df.merge(item_features, left_on='aid', right_index=True, how='left').fillna(0)\n\n    df = df.merge(user_features, left_on='session', right_index=True, how='left').fillna(0)\n\n    df.to_parquet(f'carts_candidate_with_features_p{k}_v1.pqt') \n</code></pre>\n<p>and this to read the dataframe, </p>\n<pre><code>df_container=[]\n\nfor k in range(CHUNKS):\n\n    df_container.append(cudf.read_parquet(f'carts_candidate_with_features_p{k}_v1.pqt'))\n\n\n candidates=cudf.concat(df_container)\n</code></pre>\n<p>But I am having memory problems. I tried other methods too but faced the same problem.</p>\n<p>I have generated 50 candidates per user, and then down sampled them by 20% and I have 49 features.<br>\nI also have converted the datatypes to <code>float32</code> and <code>int32</code>.</p>\n<p>Is there anything I can do?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2114569,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-01-25T04:51:06.043000",
          "content": "<p>When do you get a memory error. Does it happen when you try to read all the chunks and concat into one dataframe? Or does it happen when you start training XGB?</p>\n<p>One thing you can do is read in half your chunks and train one XGB with half your chunks. Then clear memory. Then read in the other half of chunks and train a second XGB. Finally infer the test data with both XGB and average the predictions.</p>\n<p>A second idea is to use 2xT4 GPU and DASK which doubles your GPU RAM for cudf and xgb.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2115788,
              "author_name": "Abid Hasan",
              "author_url": "",
              "post_date": "2023-01-26T00:50:05.857000",
              "content": "<p>Thanks for your response. It happens while reading the chunks and concating them into a single dataframe. I haven't even got to train a model on he complete dataset.</p>\n<blockquote>\n  <p>One thing you can do is read in half your chunks and train one XGB with half your chunks. Then clear memory. Then read in the other half of chunks and train a second XGB. Finally infer the test data with both XGB and average the predictions.</p>\n</blockquote>\n<p>Thanks for your suggestion. I will try this.</p>\n<blockquote>\n  <p>A second idea is to use 2xT4 GPU and DASK which doubles your GPU RAM for cudf and xgb.</p>\n</blockquote>\n<p>I will try this too. I am new to DASK and learning.</p>\n<p>Thank you so much for your time.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2113019,
      "author_name": "shun.222",
      "author_url": "",
      "post_date": "2023-01-24T03:04:29.950000",
      "content": "<p>I have a problem step6, Is ground truth the same as valid data B?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2114277,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-01-24T20:27:59.957000",
          "content": "<p><a href=\"https://www.kaggle.com/shun222\" target=\"_blank\">@shun222</a> Yes they are the same</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2109701,
      "author_name": "christeea",
      "author_url": "",
      "post_date": "2023-01-21T15:40:20.133000",
      "content": "<p>Wow, amazing.Cool</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2104955,
      "author_name": "EeyoreLee",
      "author_url": "",
      "post_date": "2023-01-18T06:59:59.187000",
      "content": "<p>Hi, I'm confused that we use 4 week data (3 weeks trian plus 1 week valid A) but use 5 weeks data (4 kaggle train plus 1 week kaggle test). Wouldn't this cause some feature distribution to be inconsistent ? like <code>item_item_count</code></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2105927,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-01-18T20:37:50.570000",
          "content": "<p>Great point <a href=\"https://www.kaggle.com/earlee0412\" target=\"_blank\">@earlee0412</a> For certain features, you do need to be careful such as counts. During inference, you can use the last 3 weeks of train plus 1 week kaggle test for total of 4 weeks to count.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2106185,
              "author_name": "EeyoreLee",
              "author_url": "",
              "post_date": "2023-01-19T02:25:54.287000",
              "content": "<p>agree so and thanks so much</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2104876,
      "author_name": "shun.222",
      "author_url": "",
      "post_date": "2023-01-18T05:45:46.587000",
      "content": "<p>Hi Chris, thank you for your sharing.</p>\n<p>I have a question about Infelence session.<br>\nI couldn't understand about \"we select 20 by sorting …\".<br>\nI couldn't sort predicts becouse My \"12899779_clicks\" predict only one aid(\"59625\")'s pred.</p>\n<p>Please tell me more about Infelence session.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2104894,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-01-18T06:03:54.583000",
          "content": "<p>In step 1, we generate X number of candidates where <code>X&gt;20</code>. Then in the \"Inference\" step, for each user we will have <code>X</code> predictions per user. And we sort those <code>X</code> predictions and take the best 20.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2104909,
              "author_name": "shun.222",
              "author_url": "",
              "post_date": "2023-01-18T06:13:52.443000",
              "content": "<p>I　unerstand.<br>\nThank you fou your help.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2103732,
      "author_name": "Nin7a1",
      "author_url": "",
      "post_date": "2023-01-17T10:42:09.137000",
      "content": "<p>For training, should we add the positive samples, i.e., the pairs of (session, item, target=1) that are not in the set candidates?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2109728,
          "author_name": "EeyoreLee",
          "author_url": "",
          "post_date": "2023-01-21T15:54:09.390000",
          "content": "<p>I used it    </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2109796,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-01-21T17:07:46.157000",
          "content": "<p>The best way to find out for your model is to try it and compare to CV and LB versus not trying it.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2100160,
      "author_name": "Pradeep Pujari",
      "author_url": "",
      "post_date": "2023-01-15T00:50:02.993000",
      "content": "<p>How to create user-item interaction features? This is improtant for rec sys.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2097920,
      "author_name": "SuperScience",
      "author_url": "",
      "post_date": "2023-01-13T04:39:39.927000",
      "content": "<p>Very practical guide！Thanks for sharing! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2096386,
      "author_name": "devstar3488",
      "author_url": "",
      "post_date": "2023-01-12T02:30:32.010000",
      "content": "<p>Hi Chris, I appreciate your willingness to share valuable ideas and codes. It really helps me to try participation to this competition.<br>\nWe would generate 3 candidates for each action (clicks, carts, and orders) using co-visitation matrix.<br>\n(Below is the code snippets you had shared the 3 handcrafted reranked candidates using both history data and suggested data) </p>\n<p>pred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(<br>\n    lambda x: suggest_clicks(x)<br>\n)</p>\n<p>pred_df_buys = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(<br>\n    lambda x: suggest_buys(x)<br>\n)<br>\nI wonder if I have to use the separate candidate with each action (e.g. pred_df_clicks, pred_df_buys, pred_df_carts )  to make features for 3 separate ML ranker models?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2096395,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-01-12T02:40:52.497000",
          "content": "<p>For each target (click,cart,order), we create one dataframe of candidates. There are 1.8 million test users. So each dataframe has <code>X</code> times 1.8 million where <code>X</code> is the number of candidates per user. We can try different <code>X</code>.</p>\n<p>Then for each dataframe, we merge on the targets of 0 or 1 whether the candidate is ground truth or not. Then we train GBT ranker for each target. So we have 3 trained GBT models.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2096557,
              "author_name": "devstar3488",
              "author_url": "",
              "post_date": "2023-01-12T06:15:07.947000",
              "content": "<p>Ok, I see then I'll try it. Thank you very much.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2104649,
              "author_name": "devstar3488",
              "author_url": "",
              "post_date": "2023-01-18T00:09:31.850000",
              "content": "<p>Now, I was stuck in training phase. <br>\nI crated a dataframe (e.g. clicks_df_for_train (1 GB)) in TPU-VM kaggle-platform (enough for RAM size)<br>\nI tried LGBM in the platfrom, worked without any probelm, but didn't have good score, so I wanted to try XGBoost  before going into more feature engineering step to improve CV score.<br>\nFor xgb, I changed into GPU-supported platform but I failed to load the dataframe for training due to memory size.<br>\nI would appreciate some tips or advice for further proceeding. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2095593,
      "author_name": "Bruce",
      "author_url": "",
      "post_date": "2023-01-11T13:45:31.300000",
      "content": "<p>Thank you for your sharing!</p>\n<p>In the phase <strong>Training</strong>,  the size of candidates is 1801251 * 50?  1801251 is the number of sessions in the validation data A.</p>\n<p>My question is whether a group without a positive sample may affect the performance of XGBRanker.</p>\n<p>In addition, should I make sure each group is of the same size when downsampling negatives?</p>\n<p>Looking forward to your reply~</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2096306,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-01-12T00:40:59.273000",
          "content": "<blockquote>\n  <p>In the phase Training, the size of candidates is 1801251 * 50? 1801251 is the number of sessions in the validation data A.</p>\n</blockquote>\n<p>Yes if you use 50 candidates. This is a suggestion but you can try other numbers</p>\n<blockquote>\n  <p>My question is whether a group without a positive sample may affect the performance of XGBRanker.</p>\n</blockquote>\n<p>Maybe. The best way to check is to try both and compare CV score and LB score.</p>\n<blockquote>\n  <p>In addition, should I make sure each group is of the same size when downsampling negatives?</p>\n</blockquote>\n<p>No. This is not required by GBT ranker models. Just make sure to give XGB the correct sizes of each group. Also make sure that the dataframe you train with has each user together as consecutive rows. For example, first <code>df = df.sort_values('user')</code> then <code>group_size = df.groupby('user').user.agg('count').values</code>.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2096351,
              "author_name": "Bruce",
              "author_url": "",
              "post_date": "2023-01-12T01:28:12.580000",
              "content": "<p>Many thanks</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2086622,
      "author_name": "Rayan-aay",
      "author_url": "",
      "post_date": "2023-01-04T21:42:50.223000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. I have a question about the co-visitation matrix.</p>\n<p>Should we create respective co-visitation matrix to extract candidates for each of train/valid/test ? For exemple, for train I'll use the candidates extracted from the train dataset, for validation I'll use the candidates extracted from the validation set ect …</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2096308,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-01-12T00:42:20.280000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rayanaay\" target=\"_blank\">@rayanaay</a> . We make one set of co-visitation matrices for train/validate. And a second set of co-visitation matrices for inference (submission.csv).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2085940,
      "author_name": "kunyuanS",
      "author_url": "",
      "post_date": "2023-01-04T13:32:06.997000",
      "content": "<p>Thanks for sharing very informative artical.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2084084,
      "author_name": "Nin7a1",
      "author_url": "",
      "post_date": "2023-01-03T06:39:23.180000",
      "content": "<p>How can we get training labels for training set? For example, if aid 100 was clicked in session 100 at some timestamp in training set, then the target should be 1 for the pair of session 100 and aid 100?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2084450,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-01-03T14:01:36.083000",
          "content": "<p>I suggest using Radek's Kaggle dataset explained here. Radek has converted kaggle train data into \"new train data\", \"valid A data\" and \"valid B data\" (where valid B are train labels that you ask about). These 3 are just like \"kaggle's train.csv\", \"kaggle's test.csv\", and \"kaggle leaderboard\". Radek explains how he got them in his discussion post.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2084539,
              "author_name": "Nin7a1",
              "author_url": "",
              "post_date": "2023-01-03T15:26:43.367000",
              "content": "<p>Thanks for your reply! I am new to recommendation. My understandings are as follows. <br>\n1) \"new train data\": it is used for the generation of an \"aid2aid\" matrix.<br>\n2) \"valid A data\": we generate 50 candidates for each session in valid A data by the \"aid2aid\" matrix. Some candidates are from ground truths in valid B data, so the \"gt\" is set as 1. Some candidates are not in valid B data, so the \"gt\" is set as 0. With many rows of (session, aid, some features, gt), then we can train a GBT ranker.<br>\n3) \"valid B data\": it is the ground truth label for valid A data and is used to compute local CV. <br>\nAre the above correct? </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2084552,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-01-03T15:37:02.517000",
              "content": "<p><a href=\"https://www.kaggle.com/nin7a1\" target=\"_blank\">@nin7a1</a> that is correct. However we can use \"test leakage\" in this competition, read discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939\" target=\"_blank\">here</a>. So in your \"#1\", you can use \"new train data\" plus \"valid A\" to generate \"aid2aid\" matrix. Yes, it will use leak from future data, (but we can replicate this when we make a submission to leaderboard because we can also use kaggle's train.csv and kaggle's test.csv to make \"aid2aid\" for submission)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2084558,
              "author_name": "Nin7a1",
              "author_url": "",
              "post_date": "2023-01-03T15:41:19.450000",
              "content": "<p>Thank you!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2086666,
              "author_name": "Rayan-aay",
              "author_url": "",
              "post_date": "2023-01-04T23:16:41.723000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nI have a question about the training dataset ( the first 3 weeks ). Do you use this data only to generate features, or do you train your model on it too ?</p>\n<p>I'm confused because of this : <code>(where valid B are train labels that you ask about)</code> in your answer. If I well understood the pipeline:</p>\n<ul>\n<li>Valid A are the rows to consider ( input X for the ranker ).</li>\n<li>Valid B are the labels ( Y ).</li>\n<li>Whereas the first 3 weeks of data are used to extract features about the items.</li>\n</ul>\n<p>In summary, we are only using one single week to predict the labels for the test data</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3076913,
              "author_name": "Jonathan Mallia",
              "author_url": "",
              "post_date": "2024-12-20T11:48:56.903000",
              "content": "<p>Good question <a href=\"https://www.kaggle.com/rayanaay\" target=\"_blank\">@rayanaay</a> , would like to know the answer to this if you figured it out since then :)<br>\nWould be great if this is clarified <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 🙏</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2083547,
      "author_name": "adairli00",
      "author_url": "",
      "post_date": "2023-01-02T15:48:00.560000",
      "content": "<p>Hi Chris, its very impressive to me.i konw in the \"Inference\" stage you used kaggle's  train and test data to build a submission dataset.but i have one question to enquire you that how i built native CV to konw the score in my notebook like in your shared Code \"Compute Validation Score - [CV 565]\".  thank you about the sharing again!👍</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2083571,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-01-02T16:12:42.703000",
          "content": "<p>This is why we use 5fold GroupKFold. After building the train, valid A, and valid B datasets. Then we generate candidates and features. Then we train and infer our 5fold. This will give us OOF predictions for all the valid B data. We can then compute the local CV score comparing our OOF predictions with the valid B ground truth.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2084063,
              "author_name": "adairli00",
              "author_url": "",
              "post_date": "2023-01-03T05:53:10.313000",
              "content": "<p>thank you it's very helpful for me.but i have a new question : in your discussion a 'candidate' user contains 50 items, i want to ask that in the 50 items all present the 'clicks' not contains 'carts' and 'orders'? (the picture), if in 'clicks','carts' and 'orders' that one user will have 150 'items'. and other question : your discussion is rerank 'click' acton , but why you only mark 'carts' as type '1' rather than 'carts' and 'buys' as type '1' , i am curious about it , thank you again~✨<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12934577%2F253e4847908ff0372001d58afe023918%2FXnip2023-01-03_13-42-04.jpg?generation=1672725181023917&amp;alt=media\" alt=\"\"></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2084452,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-01-03T14:04:06.800000",
              "content": "<p>We mark carts as type 1 because this is only one example. We build 3 candidate dataframes one for each clicks, carts, and orders. Each only has 50 candidates per user and each has candidates appropriate to the respective target. In step 6, we add the respective target to the respective dataframe. And we train and infer each separately by making 3 separate GBT ranker models. At last we concatenate all the predictions into a single submission.csv.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2084489,
              "author_name": "adairli00",
              "author_url": "",
              "post_date": "2023-01-03T14:47:12.900000",
              "content": "<p>thank you!!!! I understand~</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2082598,
      "author_name": "Rayan-aay",
      "author_url": "",
      "post_date": "2023-01-01T17:58:26.140000",
      "content": "<p>Why don't you use LightGBM ? is XGBoost more suited for this task ?<br>\nFrom my own experience, lightgbm uses less memory, isn't ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2082728,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-01-01T21:07:20.913000",
          "content": "<p>You can try XGBoost, LGBM, and CatBoost. They can all do ranker models. Its best to try all 3 and see which produces the best CV and LB.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2082743,
              "author_name": "Rayan-aay",
              "author_url": "",
              "post_date": "2023-01-01T21:36:47.127000",
              "content": "<p>Great ! for the moment, I consider the problem as a binary problem, using  AUC as a metric. However, i'm getting 0.7 of AUC with I think, good features extracted from a graph neural network and aggregated features. Plus, when extracting the top 20 using the GBDT groupby each session, I'm getting a validation score on clicks of 0.400, which is very bad compared to the 0.600 of the top50 candidates extracted using the co-visitation-matrix. What do you think of that and what could be the source of the problem ? <br>\nI'm training 2 LGBM Classifier on the train dataset ( 2 chunks ) </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2075850,
      "author_name": "Qucy Wei",
      "author_url": "",
      "post_date": "2022-12-26T01:02:52.860000",
      "content": "<p>Thanks for sharing very informative artical.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2075059,
      "author_name": "Pradeep Pujari",
      "author_url": "",
      "post_date": "2022-12-25T02:14:50.993000",
      "content": "<p>Where do you get user features and item features? It is not there in given dataset?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2075069,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-25T03:00:48.480000",
          "content": "<p>No it is not. The dataset from Kaggle only has <code>user, item, time stamp, event type</code> (they rename them differently but this is what they are). We must create features from these columns.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2074187,
      "author_name": "kevin takano",
      "author_url": "",
      "post_date": "2022-12-23T20:29:44.917000",
      "content": "<p>Thanks for the post and all clarifications. </p>\n<p>After reading your post many times I conclude that validation A is more like a 'last_week_train' dataframe than validation for ml model. Because we use validation A to create the user session features and we don't use validation A as validation as we use in a holdout. Because you are using cross-validation kfold to train your data. <br>\nAlso, validation B is just used to create the ground truth, am I correct? </p>\n<p>EDIT: </p>\n<p>Also, we consider we don't have the same sessions in the test set from training. So, why do we break the validation in two parts (A and B) from the same user?</p>\n<p>Thank you,</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2074218,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-23T20:53:20.180000",
          "content": "<blockquote>\n  <p>After reading your post many times I conclude that validation A is more like a 'last_week_train' dataframe than validation for ml model.</p>\n</blockquote>\n<p>Perhaps, \"last_week_train\" would be a better name but note that it is more than <strong>normal</strong> train data. I used \"val A\" because it mimics \"test.csv\" that Kaggle provides. Using \"val A\" is actually a <strong>leak</strong> because some events in \"val A\" occur after events in \"val B\". Normally we can never use information from the future to predict the past.</p>\n<p>In normal competitions, we would not need <code>new train</code>, <code>val A</code> and <code>val B</code>. But in this comp there is a leak and Kaggle says we can use the leak. When we make a submission Kaggle gives us 3 dataframes. When we make a submission, Kaggle has <code>train.csv</code>, <code>test.csv</code> and <code>public LB</code>. Those 3 sets are different. The purpose of local validation is to mimic the leaderboard. So we must copy this setup (including the leak). Since Kaggle has 3 sets, we make 3 sets. Then we have the following comparison</p>\n<ul>\n<li>\"new train\" is like \"kaggle train\"</li>\n<li>\"val A\" is like \"kaggle test\"</li>\n<li>\"val B\" is like \"kaggle leaderboard\"</li>\n</ul>\n<p>We cannot use any information from \"val B\" because just like submission, we cannot access information from public leaderboard. In our local setup, we <strong>only</strong> use \"val B\" to compute our local validation score, we <strong>do not</strong> take any information from it. During submission, we can access kaggle train and kaggle test, so during validation we can also access \"new train\" and \"val A\".</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2068857,
      "author_name": "Priyanshu Chaudhary",
      "author_url": "",
      "post_date": "2022-12-18T11:22:40.660000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, Thanks for the idea. I am new to this competition and i have some doubts, can you please explain them:<br>\n1) after splitting your training data into train +(val A,val B) you are generating 50 candidates only for val A right?<br>\n2) does that mean you are using your validation data A, val data B to train xgb ranker and training data only to create features for items?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2069044,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-18T14:21:09.447000",
          "content": "<p>1) yes and note that val A and val B have the same users. So we are generating 50 candidates for both val A and val B).</p>\n<p>2) yes and we can use val A to create features for items too. (and for advanced Kagglers, we can use <strong>unsupervised</strong> methods to create user features from train).</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2072676,
              "author_name": "Priyanshu Chaudhary",
              "author_url": "",
              "post_date": "2022-12-22T10:39:26.787000",
              "content": "<p>Hi Chris, I have one more doubt, are you training for each target('click', 'cart', 'buy') independently?<br>\nRegards<br>\nPriyanshu</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2072894,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2022-12-22T13:54:21.333000",
              "content": "<p>Yes i train 3 independently. I make 3 dataframes. One dataframe is candidates for clicks, one for carts, one for orders. Then i train 3 XGB rankers. One takes click dataframe and predicts clicks. One takes carts and predicts carts. And the last takes orders and predicts orders.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2076392,
              "author_name": "Priyanshu Chaudhary",
              "author_url": "",
              "post_date": "2022-12-26T13:07:49.110000",
              "content": "<p>Thanks a lot for help <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. I have made a similar pipeline as you have described. My CV results are as follows:</p>\n<p>Clicks,    ndcg@20= .7xxxx          recall=  .59<br>\ncarts ,     ndcg@20= .9xxx           recall= .52<br>\nare you also getting same results?<br>\nregards<br>\nI am using  dividing val A data into 5 group folds training on 4 folds and testing on 1 fold<br>\nplease correct me if I  am wrong.<br>\nRegards<br>\nPriyanshu</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2066732,
      "author_name": "apatheia",
      "author_url": "",
      "post_date": "2022-12-16T02:50:26.957000",
      "content": "<p>I love this</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2066098,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-15T12:07:27.417000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2064603,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-14T00:00:43.620000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2065595,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-14T21:36:08.413000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2065602,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-14T21:50:12.463000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2063769,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-13T09:53:50.467000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2064122,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-13T14:30:44.733000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2063714,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-13T08:25:19.977000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2062366,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-12T03:41:36.383000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2061563,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-11T09:02:20.353000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2061877,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-11T14:42:49",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2061103,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-10T17:31:22.293000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2060946,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-10T15:25:13.117000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2060639,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-10T07:51:52.007000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2060470,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-10T01:12:19.603000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2055947,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-05T15:20:14.797000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2055959,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-05T15:33:30.147000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2054607,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-04T09:27:26.677000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2054777,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-04T12:58:52.830000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2054587,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-04T09:06:29.497000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2053913,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-03T18:01:00.507000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2053873,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-03T17:12:56.060000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2123828,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-31T17:20:50.383000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2123858,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-31T17:33:43.217000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2120892,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-29T22:20:06.340000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2120903,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-29T22:35:32.960000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2111789,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-23T07:19:24.397000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2111943,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-23T10:08:31.377000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2111967,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-23T10:21:03.097000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2085979,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-04T13:44:19.230000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2096309,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-12T00:42:35.477000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2084936,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-03T21:01:46.460000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2084983,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-03T21:29:47.520000",
          "content": "",
          "votes": 3,
          "replies": [
            {
              "id": 2085200,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-04T01:28:02.207000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2084505,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-03T15:02:41.387000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2108272,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-20T11:48:04.480000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2078101,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-28T04:45:14.573000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2078145,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-28T05:07:04.023000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2077952,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-28T01:19:59.523000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2077968,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-28T02:01:24.537000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2078787,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-28T16:38:30.063000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2080138,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-29T22:30:05.853000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2077327,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-27T13:24:56.363000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2074808,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-24T16:20:03.070000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2074884,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-24T18:14:52.293000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2072224,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-21T21:53:26.533000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2067385,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-16T16:52:41.373000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2069047,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-18T14:25:05.197000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2069215,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-18T18:16:26.097000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2069235,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-18T18:47:45.683000",
              "content": "",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2069300,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-18T20:40:35.333000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2070072,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-19T15:21:26.230000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2070185,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-19T17:21:56.057000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2070191,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-19T17:27:57.910000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2070290,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-19T19:59:58.480000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2070312,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-19T20:26:47.110000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2070362,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-19T22:48:28.710000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2076724,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-26T18:41:10.513000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2078770,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-28T16:20:50.227000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2078811,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-28T17:04:10.420000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2080950,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-30T16:30:48.960000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082149,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-01T06:52:47.987000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2082159,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-01T07:08:36.150000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082484,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-01T16:09:15.290000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2082810,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-02T01:12:32.587000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2104743,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-18T02:21:14.397000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2104746,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-18T02:22:07.050000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2104748,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-18T02:23:45.377000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2067367,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-16T16:29:16.853000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2067403,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-16T17:07:15.333000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2065267,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-14T13:14:14.120000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2065275,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-14T13:29:36.680000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2067064,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-16T11:01:58.770000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2070032,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-19T14:49:14.660000",
              "content": "",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2062223,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-11T21:41:50.410000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2061240,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-10T21:30:32.717000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2061251,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-10T21:51:41.873000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2061971,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-11T16:19:26.947000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2061998,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-11T16:57:17.040000",
          "content": "",
          "votes": 7,
          "replies": [
            {
              "id": 2063582,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-13T05:06:25.233000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2062000,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-11T16:59:46.073000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2062073,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-11T17:47:21.147000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2062080,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-11T17:56:44.947000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2064436,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-13T18:49:16.723000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2064521,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-13T20:43:22.973000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2064694,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-14T03:48:06.380000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2057813,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-07T11:09:21.423000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2056545,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-06T08:13:19.210000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2054786,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-04T13:06:42.103000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2056266,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-05T23:56:10.937000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2054181,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-03T23:22:11.353000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2054259,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-04T01:38:51.800000",
          "content": "",
          "votes": 9,
          "replies": []
        }
      ]
    },
    {
      "id": 3021487,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-18T14:56:28.010000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2303763,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-15T13:12:32.613000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2246867,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-05T14:20:48.123000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2124574,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-01T04:48:53.600000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2081143,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-30T20:53:20.247000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2078580,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-28T12:48:15.617000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2078583,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-28T12:51:39.150000",
          "content": "",
          "votes": 2,
          "replies": [
            {
              "id": 2117794,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-27T14:51:39.820000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2063686,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-13T07:47:16.933000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2064129,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-13T14:37:53.497000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2064567,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-13T22:48:41.663000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2064570,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-13T22:50:40.860000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2064578,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-13T23:10:21.467000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2064582,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-13T23:18:26.423000",
          "content": "",
          "votes": 5,
          "replies": []
        },
        {
          "id": 2065261,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-14T13:02:04.850000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2065276,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-14T13:30:20.267000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2065279,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-14T13:34:34.617000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2065302,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-14T14:12:52.160000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2065463,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-14T17:21:01.990000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2065492,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-14T18:01:13.563000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2065877,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-15T06:41:49.670000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2066019,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-15T09:56:57.027000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2066063,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-15T11:22:54.010000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2066248,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-15T14:04:33.777000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2074225,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-23T20:56:35.600000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2066253,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-15T14:11:57.620000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2066435,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-15T17:34:17.113000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2074224,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-23T20:56:10.690000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2054040,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-03T20:10:50.453000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2112772,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-23T20:24:07.050000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2066728,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-16T02:43:56.837000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2107408,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-19T19:06:36.770000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2097356,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-12T16:00:52.487000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2065506,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-14T18:21:34.210000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2053832": "# Candidate Rerank Model\nA strong and easy way to build a solution for this competition is a \"candidate rerank\" model explained [here][3]. Creating train data, training and inferring a GBT ranker model can be difficult to understand at first. I hope the following discussion post will help you on your journey.\n\nThe general idea is; for each `session` (i.e. user), we find 50-200 candidates (that are likely to be correct predictions) then we train an GBT ranker model to select our final 20. To do this, we need to create train data for our ranker model. Below are some tips to help create train data and use GBT ranker models\n\n# Candidate DataFrame\nThe first step to creating train data for our ranker model is building a dataframe of candidates. One way to generate high quality candidates is using co-visitiation matrices, explained [here][4]. Our candidate dataframe has one pair of `session` (i.e. user) and `aid` (i.e. item) per row. This dataframe has the following columns\n* session (i.e. user)\n* aid (i.e. item)\n* user features\n* item features\n* user-item interaction features\n* click target (i.e 0 or 1)\n* cart target (i.e. 0 or 1)\n* order target (i.e. 0 or 1)\n\n# Memory Management\nIn all of the following code, it is very important to continually change the `dtype` of each column to minimum data size. So after making a new column, make sure to `df[col] = df[col].astype('int32')` or the smallest possible dtype. If you do nothing then many libraries will use `int64` or `float64` which takes 8 bytes per value. This will use more RAM and disk space then needed. **Always reduce dtypes!**\n\n**NOTE** When using the latest version of cuDF locally (version 22.08 or later) , there is a new feature to set default bitwidth as 32. This will prevent the dataframe from using `int64` nor `float64` and makes it easier for the user. (Thanks Giba for pointing this out [here][7]):\n\n    import cudf\n    cudf.set_option(\"default_integer_bitwidth\", 32)\n    cudf.set_option(\"default_float_bitwidth\", 32)\n\nEven after reducing column dtypes, you may still have memory issues. The way to solve this is to process everything in chunks. If you are using `train.groupby('session')` to make user features, then perhaps first split the train data into X pieces (like 2, 4, 8 pieces). Then process user features for each piece separately and save to disk as `f'user_features_p{PIECE_NUMBER}.pqt'`. Later when you read them in, you can concatenate them together.\n\nAlso consider using DASK which works with both CPU Pandas or GPU RAPIDS cuDF to use multiple CPU/GPU and/or disk when needed to avoid all memory errors! [(info here)][6]\n\n# Speed - Use GPU\nAll of the following code will take time to run since this data has around 13 million users and 2 million items. I suggest using an accelerated dataframe library for all the following processing like Nvidia's RAPIDS cuDF [here][2] which uses GPU instead of CPU for accelerated speed !\n\n# Step 1\nOur train data will be the first 3 weeks of Kaggle train data. And our validation data will be the last 1 week of Kaggle train. We then split our validation data into **validation data A** and **validation data B** where validation B are the ground truths. (Validation data A is like the Kaggle test data we download whereas validation data B is like the LB (leaderboard). Roughly speaking `A` is like first half of each user activity during week while `B` is second half).\n\nConsider using Radek's train and valid data [here][5] which already splits validation data into A and B. For every user (i.e. session) in our **validation data A** (i.e. last 1 week of Kaggle train), we generate X candidate aids. Let's say `X=50` for the rest of our discussion. We create a dataframe of shape `( `number_of_session x 50`, 2 )`:\n\n| user | item | \n| --- | --- | \n| 0001 | 6456 | \n| 0001 | 4490 | \n| 0002 | 8486 | \n| 0002 | 7297 | \n\nEach session will appear 50 times. And there will be no duplicate `[session,aid]` pairs. We only use users from our validation data because we only have targets for users in validation data (in step 6 below).\n\n# Step 2\nCreate item features. Using our **train data + valid data A** (yes use test leak), we create item features in their own dataframe and save a parquet to disk. For example\n\n    item_features = train.groupby('aid').agg({'aid':'count','session':'nunique','type':'mean'})\n    item_features.columns = ['item_item_count','item_user_count','item_buy_ratio']\n    # CONVERT COLUMNS TO INT32 and FLOAT32 HERE\n    item_features.to_parquet('item_features.pqt')\n\n**NOTE** if you have memory problems. Then break `train` into 10 dataframe parts. All rows pertained to a single item must be in the same dataframe part so that `train_part_1.groupby('aid')` will work correctly. After processing, save each part separately to disk. And later when you read them from disk, concatenate them together before merging them to candidate dataframe.\n\n# Step 3\nCreate user features. Using our **validation data A**, we create user features in their own dataframe and save a parquet to disk. For example\n\n    user_features = train.groupby('session').agg({'session':'count','aid':'nunique','type':'mean'})\n    user_features.columns = ['user_user_count','user_item_count','user_buy_ratio']\n    # CONVERT COLUMNS TO INT32 and FLOAT32 HERE\n    user_features.to_parquet('user_features.pqt')\n\n# Step 4\nThis step is optional. Step 4 will improve CV and LB, but your GBT ranker will work without step 4. Create user-item interaction features. Using our **validation data A**, we create **multiple** user-item feature dataframes and save them as parquets to disk. For each idea, we can make a new dataframe. One dataframe can contain all items that a user clicks. So make a dataframe with one column user, one column item, and a third column called `item_clicked`. Then for each unique item that a user clicked, we add a new row with `item_clicked = 1`. Note that our dataframe will have **no duplicate rows** of `['user','item']` pairs. Save this dataframe to disk. When we merge this to candidate dataframe we will `fillna(0)` to indicate the items not clicked.\n\n# Step 5\nAdd features to our candidate dataframe. To add features to our candidate dataframe, we read from disk and merge them on as follows\n\n    item_features = pd.read_parquet('item_features.pqt')\n    candidates = candidates.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\n    user_features = pd.read_parquet('user_features.pqt')\n    candidates = candidates.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\n\nNow our candidate dataframe looks like\n\n| user | item | item_feat1 | item_feat2 | user_feat1 | user_feat2 |\n| --- | --- | --- | --- | --- | --- | \n| 0001 | 6456 | 10 | 12 | 5.4 | 0.5 |\n| 0001 | 4490 | 13 | 15 | 5.4 | 0.5 |\n| 0002 | 8486 | 55 | 8 | 2 | 1.2 |\n| 0002 | 7297 | 70 | 10 | 2 | 1.2 |\n\n**NOTE** if we have memory problems merging all our features to our candidate dataframe, then we can do this in chunks\n\n    CHUNKS = 10\n    chunk_size = np.ceil( len(candidates) / CHUNKS)\n    for k in range(CHUNKS):\n        df = candidates.iloc[k*chunk_size:(k+1)*chunk_size].copy()\n        df = df.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\n        df = df.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\n        df.to_parquet(f'candidate_with_features_p{k}.pqt')\n\n# Step 6\nWe will now add targets to our candidate dataframe from step 1. The best way to add a column of targets is to use dataframe merge. First we make a dataframe of all the `target=1` as follows. Starting with a dataframe that contains the targets as a column of lists ( like Radek's ground truth labels [here][1]) such as:\n\n`test_labels.parquet`\n| session | type | ground truth |\n | --- | --- | --- |\n| 0001 | carts | [3456, 4490, 5661, 7821, 9914 ] |\n| 0002 | carts | [1222, 4656, 533, 8486] |\n\nWe use the following code to convert these lists into a dataframe of targets:\n\n    tar = pd.read_parquet('test_labels.parquet')\n    tar = tar.loc[ tar['type']=='carts' ]\n    aids = tar.ground_truth.explode().astype('int32').rename('item')\n    tar = tar[['session']].astype('int32').rename({'session':'user'},axis=1)\n    tar = tar.merge(aids, left_index=True, right_index=True, how='left')\n    tar['cart'] = 1\n\nThis produces a dataframe like\n\n| user | item | cart | \n| --- | --- | --- | \n| 0001 | 3456 | 1 | \n| 0001 | 4490 | 1 | \n\nAnd we merge it to our candidate dataframe with the following line:\n\n    candidates = candidates.merge(cart_target,on=['user','item'],how='left').fillna(0)\n\n| user | item | item_feat1 | item_feat2 | user_feat1 | user_feat2 | cart |\n| --- | --- | --- | --- | --- | --- | --- |\n| 0001 | 6456 | 10 | 12 | 3 | 0.5 | 0 |\n| 0001 | 4490 | 13 | 5 | 5.4 | 0.1 | 1 |\n| 0002 | 8486 | 55 | 10 | 5 | 0.9 | 1 |\n| 0002 | 7297 | 70 | 20 | 2 | 1.2 | 0 |\n\n# Training\nWe now have train data for our GBT ranker model. We must train using `GroupKFold`. **Important Note**: when we train, we do not use the `user` and `item` columns as features, we only use the other columns. `FEATURES = candidates.columns[2 : -1*len(targets)]`. Note with XGB, we have 3 options for rankers by changing `objective` parameter, to either `rank:pairwise`, or `rank:ndcg`,  or `rank:map`. (We should also add and tune XGB parameters `'max_depth', 'subsample', 'colsample_bytree', 'learning_rate'`)\n\n    import xgboost as xgb\n    from sklearn.model_selection import GroupKFold\n\n    skf = GroupKFold(n_splits=5)\n    for fold,(train_idx, valid_idx) in enumerate(skf.split(candidates, candidates['click'], groups=candidates['user'] )):\n\n        X_train = candidates.loc[train_idx, FEATURES]\n        y_train = candidates.loc[train_idx, 'click']\n        X_valid = candidates.loc[valid_idx, FEATURES]\n        y_valid = candidates.loc[valid_idx, 'click']\n\n        # IF YOU HAVE 50 CANDIDATE WE USE 50 BELOW\n        dtrain = xgb.DMatrix(X_train, y_train, group=[50] * (len(train_idx)//50) ) \n        dvalid = xgb.DMatrix(X_valid, y_valid, group=[50] * (len(valid_idx)//50) ) \n\n        xgb_parms = {'objective':'rank:pairwise', 'tree_method':'gpu_hist'}\n        model = xgb.train(xgb_parms, \n            dtrain=dtrain,\n            evals=[(dtrain,'train'),(dvalid,'valid')],\n            num_boost_round=1000,\n            verbose_eval=100)\n        model.save_model(f'XGB_fold{fold}_click.xgb')\n\n**NOTE** If you have memory problems training XGB on GPU, consider downsampling negatives 2x, 4x, 10x, 20x with `frac = 0.5, 0.25, 0.1, or 0.05` (and then update group sizes in DMatrix). Or use DASK XGB with multiple GPUs. Here is example code:\n\n    positives = candidates.loc[candidates['click']==1]\n    negatives = candidates.loc[candidates['click']==0].sample(frac=0.5)\n    candidates = pd.concat([positives,negatives],axis=0,ignore_index=True)\n\n# Inference\nFor inference, we create a new candidate dataframe (using our technique to generate candidates before) but this time from Kaggle's test data. Then we make item features from all 4 weeks of Kaggle train plus 1 week of Kaggle test. And we make user features from Kaggle test. We merge the features to our candidates. Then we use our saved models to infer predictions for clicks. Lastly we select 20 by sorting the predictions and choosing 20 with.\n\n    preds = np.zeros(len(test_candidates))\n    for fold in range(5):\n        model = xgb.Booster()\n        model.load_model(f'XGB_fold{fold}_click.xgb')\n        model.set_param({'predictor': 'gpu_predictor'})\n        dtest = xgb.DMatrix(data=test_candidates[FEATURES])\n        preds += model.predict(dtest)/5\n    predictions = test_candidates[['user','item']].copy()\n    predictions['pred'] = preds\n\n    predictions = predictions.sort_values(['user','pred'], ascending=[True,False]).reset_index(drop=True)\n    predictions['n'] = predictions.groupby('user').item.cumcount().astype('int8')\n    predictions = predictions.loc[predictions.n<20]\n    sub = predictions.groupby('user').item.apply(list)\n    sub = sub.to_frame().reset_index()\n    sub.item = sub.item.apply(lambda x: \" \".join(map(str,x)))\n    sub.columns = ['session_type','labels']\n    sub.session_type = sub.session_type.astype('str')+ '_clicks'\n\n**NOTE** if you have memory errors. Consider loading 1/10th of the test data. Then merge features. Then infer. Next load the next 1/10th, merge features, infer. etc. etc. Lastly concatenate the predictions and make submission.csv\n\n# Enjoy\nI hope this discussion post helps illustrate how to create data for a GBT ranking model and how to train and infer. Have fun!\n\n[1]: https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation?select=test_labels.parquet\n[2]: https://rapids.ai/\n[3]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\n[4]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369\n[5]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991 \n[6]: https://www.dask.org/\n[7]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/371058",
    "2059518": "Will this work if our candidates for each session have mixed length? For one session we may have 25 candidates, for another one 30.",
    "2058286": "@cdeotte A new functionality of Cudf (since 22.08) is setting the default data type to be 32 or 64 bit. Using 32 bit can help save GPU memory when processing features.\nThe code bellow sets 32 bit as the default:\n\n```\nimport cudf\ncudf.set_option(\"default_integer_bitwidth\", 32)\ncudf.set_option(\"default_float_bitwidth\", 32)\n```",
    "2078669": "I can't understand what is happening, I'm keep getting this error: ***XGBoostError: [13:58:22] ../src/data/data.cc:694: Check failed: group_ptr_.back() == num_row_ (9624929 vs. 7699943) : Invalid group structure.  Number of rows obtained from groups doesn't equal to actual number of rows given by data.*** \n\nI read my dataframe from the disk:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F9f5888117f112e76546cd986da22172d%2FReading.JPG?generation=1672236414247488&alt=media)\n\nThan I reduce the number of negatives like Chris proposed in his post:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F0431b5f064256829fe7638a54064f3c3%2FReducing.JPG?generation=1672236479917480&alt=media)\n\nI check if my dataframe's length equals to the sum of the array with group counts:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F8692369e6eaeac04d2e2d5ca24b4d649%2FTrain.JPG?generation=1672236560514962&alt=media)\n\nBut I'm keep getting this error:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2Fca5323843c148064437fe919bf3a84b8%2FError_1.JPG?generation=1672236658532920&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2Fdc14b53811624616ca61d9db2ca45b27%2FError_2.JPG?generation=1672236667732271&alt=media)\n\nHas anybody encountered anything similar to this? And if encountered what was the root cause?",
    "2061533": "Hi Chris. Thanks a lot for this framework!   I was wondering that while merging candidate dataframe with the target,  should we do outer join instead of left, because on doing left join we might be missing out on some of the true (user, item) pairs which may not be in our candidate dataset\n ```candidates = candidates.merge(click_target,on=['user','item'],how='left').fillna(0)```",
    "2059767": "Hi @cdeotte, I have tried to implement the pipeline as you described above. I have managed to generate dataframes of candidates and features (items, users). However, when I tried to merge them, but I couldn't get the columns right in the merged dataframe as you described above. Could you have a look? Thanks a lot! 🙏\n\nThis is your demo for merged dataframe\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fcf82ca1b474801c28527aabdc9d426e7%2FScreen%20Shot%202022-12-09%20at%2015.45.22.png?generation=1670571940780500&alt=media)\n\nMy candidates and item_features dataframes look like below\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Ff37405f6569eb49d42657643bfbabf12%2FScreen%20Shot%202022-12-09%20at%2015.37.35.png?generation=1670571720150971&alt=media)\n\nThis is the merged dataframe I got which is confusing to me.\n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fdf620b7d3ffb9545defd25ed136e6068%2FScreen%20Shot%202022-12-09%20at%2015.37.47.png?generation=1670571846965756&alt=media)\n\nI have a feeling that maybe the problem lies on index, but tried to read the docs on `merge` and `set_index`, but have not figured out how to fix it.\n",
    "2059625": "Hi,Chris,your answer helped me a lot!!!\nStep 4: I have a question. You use as an example whether an item is clicked as a feature of user-item interaction. But if I train a ranker to sort the candidates in click. Does this feature above leak my training label? That means that if I want to train a click reranker, I can't use both the user and item interaction features related to click? I am very confused about this and look forward to your answer.",
    "2069145": "Hi @cdeotte,\nFirst at all, thank you very much for this post and all the others in this complicated competition ! \nI have a question : in \"Training\", you suggested to consider downsampling negatives ; but if we do so, we will not have 50 consecutive samples from the same user and groups won't be correct for XGB. Do you agree ?\n\n(I'm looking for an issue to solve memory problems, and I tell myself that I will not consider downsampling negatives, but features forward selection. Or maybe I should consider downsampling negatives, but take care to \"group=\" in xgb.DMatrix)",
    "2054471": "Thanks for sharing， I have a question that's why step 3 and step 4 only use the validation data  but not the train data and the validation data.  In the Inference step it seems to use similar set up which use 4week Kaggle train plus 1 week of Kaggle test to generate item feature but only use Kaggle test to generate user features. Could you explain it a bit ? Thanks!",
    "2123349": "It works, thx!",
    "2114447": "Hi Chris, Thanks for sharing. \n\nI am having trouble with memory management: I split the candidate dataframe into chunks as you suggested, but I still seem to run out of memory while loading the chunks.\n\nI used this code to split the dataframe\n\n\n```\nCHUNKS = 10\n\nchunk_size = int(np.ceil( len(candidates) / CHUNKS))\n\nfor k in range(CHUNKS):\n\n    df = candidates.iloc[k*chunk_size:(k+1)*chunk_size].copy()\n\n    df = df.merge(item_features, left_on='aid', right_index=True, how='left').fillna(0)\n\n    df = df.merge(user_features, left_on='session', right_index=True, how='left').fillna(0)\n\n    df.to_parquet(f'carts_candidate_with_features_p{k}_v1.pqt') \n```\n\n\nand this to read the dataframe, \n\n\n```\ndf_container=[]\n\nfor k in range(CHUNKS):\n\n    df_container.append(cudf.read_parquet(f'carts_candidate_with_features_p{k}_v1.pqt'))\n\n\n candidates=cudf.concat(df_container)\n\n```\n\n\n\n\nBut I am having memory problems. I tried other methods too but faced the same problem.\n\nI have generated 50 candidates per user, and then down sampled them by 20% and I have 49 features.\nI also have converted the datatypes to `float32` and `int32`.\n\nIs there anything I can do?\n\n",
    "2113019": "I have a problem step6, Is ground truth the same as valid data B?",
    "2109701": "Wow, amazing.Cool",
    "2104955": "Hi, I'm confused that we use 4 week data (3 weeks trian plus 1 week valid A) but use 5 weeks data (4 kaggle train plus 1 week kaggle test). Wouldn't this cause some feature distribution to be inconsistent ? like `item_item_count`",
    "2104876": "Hi Chris, thank you for your sharing.\n\nI have a question about Infelence session.\nI couldn't understand about \"we select 20 by sorting ...\".\nI couldn't sort predicts becouse My \"12899779_clicks\" predict only one aid(\"59625\")'s pred.\n\nPlease tell me more about Infelence session.",
    "2103732": "For training, should we add the positive samples, i.e., the pairs of (session, item, target=1) that are not in the set candidates?",
    "2100160": "How to create user-item interaction features? This is improtant for rec sys.",
    "2097920": "Very practical guide！Thanks for sharing! ",
    "2096386": "Hi Chris, I appreciate your willingness to share valuable ideas and codes. It really helps me to try participation to this competition.\nWe would generate 3 candidates for each action (clicks, carts, and orders) using co-visitation matrix.\n(Below is the code snippets you had shared the 3 handcrafted reranked candidates using both history data and suggested data) \n\npred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_clicks(x)\n)\n\npred_df_buys = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_buys(x)\n)\nI wonder if I have to use the separate candidate with each action (e.g. pred_df_clicks, pred_df_buys, pred_df_carts )  to make features for 3 separate ML ranker models?",
    "2095593": "Thank you for your sharing!\n\nIn the phase **Training**,  the size of candidates is 1801251 * 50?  1801251 is the number of sessions in the validation data A.\n\nMy question is whether a group without a positive sample may affect the performance of XGBRanker.\n\nIn addition, should I make sure each group is of the same size when downsampling negatives?\n\nLooking forward to your reply~",
    "2086622": "Thanks for sharing @cdeotte. I have a question about the co-visitation matrix.\n\nShould we create respective co-visitation matrix to extract candidates for each of train/valid/test ? For exemple, for train I'll use the candidates extracted from the train dataset, for validation I'll use the candidates extracted from the validation set ect ...",
    "2085940": "Thanks for sharing very informative artical.",
    "2084084": "How can we get training labels for training set? For example, if aid 100 was clicked in session 100 at some timestamp in training set, then the target should be 1 for the pair of session 100 and aid 100?",
    "2083547": "Hi Chris, its very impressive to me.i konw in the \"Inference\" stage you used kaggle's  train and test data to build a submission dataset.but i have one question to enquire you that how i built native CV to konw the score in my notebook like in your shared Code \"Compute Validation Score - [CV 565]\".  thank you about the sharing again!👍",
    "2082598": "Why don't you use LightGBM ? is XGBoost more suited for this task ?\nFrom my own experience, lightgbm uses less memory, isn't ? ",
    "2075850": "Thanks for sharing very informative artical.",
    "2075059": "Where do you get user features and item features? It is not there in given dataset?",
    "2074187": "Thanks for the post and all clarifications. \n\nAfter reading your post many times I conclude that validation A is more like a 'last_week_train' dataframe than validation for ml model. Because we use validation A to create the user session features and we don't use validation A as validation as we use in a holdout. Because you are using cross-validation kfold to train your data. \nAlso, validation B is just used to create the ground truth, am I correct? \n\nEDIT: \n\nAlso, we consider we don't have the same sessions in the test set from training. So, why do we break the validation in two parts (A and B) from the same user?\n\nThank you,",
    "2068857": "Hi @cdeotte, Thanks for the idea. I am new to this competition and i have some doubts, can you please explain them:\n1) after splitting your training data into train +(val A,val B) you are generating 50 candidates only for val A right?\n2) does that mean you are using your validation data A, val data B to train xgb ranker and training data only to create features for items?\n",
    "2066732": "I love this",
    "2066098": "Incredible, simple, and fast forward explanation !  that is what I was looking for :).\nI have a question about the candidates generation. Let's say we use 3 methods for candidates generation, ( with the co-visitation matrix for example ), let's say 2 of them predict a good rec, while the last one predict a false negative, what do you think about weighting this true positive recommended item \"twice\", using the sample_weight parameter of LightGBM ?",
    "2064603": "Wow, this cleared up a lot! Thanks, Chris!",
    "2063769": "Hi, I want to know whether the split chunk code is wrong. It could be written like this.\n```\nCHUNKS = 10\nchunk_size = ceil(len(candidates) / CHUNKS)\nfor k in range(CHUNKS):\n    df = candidates.iloc[k*chunk_size:(k+1)*chunk_size].copy()\n    df = df.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\n    df = df.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\n    df.to_parquet(f'candidate_with_features_p{k}.pqt')\n```\nPlease correct me if I'm wrong, thanks.",
    "2063714": "O, thats great!",
    "2062366": "Using Xgboost like gods using his grace. Holy f**k. ",
    "2061563": "Hi @cdeotte, could you please elaborate `Step 4` like I'm five. I found it hard to understand to beginner like me. Thank you very much. ",
    "2061103": "great work bro",
    "2060946": "this is awesome! thanks for your explanation ",
    "2060639": "OMG!OMG! @cdeotte You are really a professional and meticulous person.👍",
    "2060470": "This is awesome, thanks very much for the tips!",
    "2055947": "Thanks for sharing! \nCurious about interaction features, \nwould it make sense to add `item-context interaction` features as well? context here could be hour of the day, or day of week?\nI have added the user features & user-item interaction features to my GBT, but somehow still not performing better than heuristic",
    "2054607": "Hi @cdeotte your wonderful post has answered many of questions which are burning in my heart.  Thank you so much!\n\nIn the sections about Memory management and Speed, I would like to know your opinion on pandas, cuDF, Dask vs Polars the image below. Also, does the image suggest Polars are faster than cuDF with GPU?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F155737f233c1ed896da44c6c5215a273%2FScreen%20Shot%202022-12-04%20at%2017.17.05.png?generation=1670145727436928&alt=media)\n\nIf using GPU, do you recommend it's better to use cuDF over other libraries? Thanks!",
    "2054587": "Another super amazing post! So much gold here! Thank you @cdeotte so much for sharing 🙏❤️❤️❤️",
    "2053913": "Many thanks for sharing !!! I learn so many things from your thread.  ",
    "2053873": "thanks for the wonderful thread, Chris. If you want to reveal, I am curious how much additional boost you got by ranking using GBDTs ?",
    "2123828": "Thank you @cdeotte I followed these steps to get my first version of all components necessary to produce a solution. Then I worked on improving separately:\n\n1. The features\n2. the candidate selection\n3. The ranker model (XGB + LGBM)\n\nI don't think I would have been able to build the whole pipeline without you.",
    "2120892": "Hi @cdeotte  thank you for your valuable information.\n\nI'm struggling with the ranker for a long time, and I suspect my interactions features.\n\nIs it ok if my interactions features has > 95% of NaNs ? from the feature importance of my GBDT, they're the most valuable variables, and I also beat the benchmark with a boost of 0.02, however my LB is still 0.565 with the ranker predictions.\n\nWhat do you think ?",
    "2111789": "Hi Chris, I have a problem in setp1 and step 6, do we need to use different data for train and inference to create candidates ?",
    "2085979": "Hi Chris, its very impressive and helpful for us to learn. You are always the most value player in kaggle's competitions. Thank you very much!👍",
    "2084936": "hi chris, your posts are always so helpful! And I wonder if I want to use co-visit matrix to get features, could I use the one generated with all data(like you shared) or I should only use the one using data from train + validation A. ",
    "2084505": "![Hi Chris you are truly teacher to me in this competition , and I quite have progress in the rerank part.. but I am facing a new problem about 'group'  in the modeling training part. when I use 'group' in your notebook , the error have appear that said \"num_row_ (4375050 vs. 4375063) : Invalid group structure\"(show at the picture blow)  ,but I cancel the 'group' it's OK to train my model. have you met this problem? how you fix up?and can you tell me the meaning of  'group' parameter and its important or not to the model training...thank you!! you answer make me grow up every day in this field ~~!\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12934577%2F27c220232d0643fdf9eddd2f4ff5c291%2FXnip2023-01-03_22-46-09.jpg?generation=1672758123273487&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12934577%2F7cbed3e3a51c8f85d31b7b92b3870b65%2FXnip2023-01-03_22-45-24.jpg?generation=1672758146084252&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12934577%2F8575a8cc806a4af3de5e7769ffdd4ea2%2FXnip2023-01-03_22-41-10.jpg?generation=1672758159104699&alt=media)",
    "2078101": "Hi @cdeotte the validation set you are using may be corrupted, for more details could you take a look at the discussion here? (kaggle spam system prevent me posting notebook links and other details, sorry for the trouble) Thanks a lot! \n\nhttps://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation/discussion/374405",
    "2077952": "Let's suppose I have this session S of tuples for the user JAMES, where the first element of the tuple is the product and the second is the type of event:\n\nS = [(\"notebook\", \"click\"), (\"notebook\", \"click\"), (\"notebook\", \"click\"), (\"mouse\", \"click\"), (\"mouse\", \"click\"), (\"mouse\", \"cart\"), (\"notebook\", \"cart\"), (\"notebook\", \"purchase\"), (\"mouse\", \"purchase\")]. \n\nAnd, we will create the user features:\n- if is the first item (F1)\n- if is the last item (F2)\n- session length (F3)\n\nAnd for item features (*if you think it's not an item feature but actually a user feature pls tell me*):\n- has this item already been clicked by user (F4)\n- has this item already been added to cart by user  (F5)\n- is this item popular? (F6)\n- is this item bought in training data (at least once)?  (F7)\n- times this item was clicked in all training data.  (F8)\n\nAlso, we observed the co-visitation matrix for this data returned the items: [\"keyboard\"]. \n\nOur objective here is to create our candidate features to predict **clicks**. \n\nSo, we will our candidate features this way:\n\n|USER | ITEM |  F1. |  F2 | F3  | F4 | F5 | F6 | F7 | F8 |\n| --- | --- | --- | --- | --- | --- | --- |--- |--- | --- |\n| JAMES|  NOTEBOOK |  1 |  0 | 9 | 1 | 1 | 1 | 1 | 381248 |\n| JAMES|  MOUSE |  0 |  1 | 9 | 1 | 1 | 1 | 1 | 1111232 |\n| JAMES|  KEYBOARD | 0 |  0 | 0 | 0 | 0 | 1 | 1 | 55555 |\n\nObserving this, my questions:\n\n1. Do my item features make sense? Maybe F4 and F5 is more like **user features** not **item features**.\n2. How the candidates from the co-visitation matrix (e.g the keyboard) could help the model if mostly of the columns/features are zeros? \n3. What's differ user-features, user-item-features from item-features? \nThank you. \n\n\n \n",
    "2077327": "This is my first time participating in the competition. I will try this notebook!",
    "2074808": "If we use 5 fold for training, how to evaluate LOCAL CV . In particular, the entire validation set cannot be used directly.\nThanks chris~",
    "2072224": "Hi @cdeotte again I come back to you with a question.\nI'm kind of lost as I revisited the problem of creation of user and item features.\nEven after adjusting the train kaggle data of kaggle by one week (to have three weeks in both training frames) the distribution of features for items and users seems vastly different.\nIn my notation:\ntrainA, testA -> train and test taken from your otto validation\ntrainB, testB -> train and test taken from the parquets we use for LB inference.\n\nI adjusted time stamps as suggested by @aldparis by two hours. So the trainB.ts.dt.day > 7 in the code means we start from 8th of August (instead of 1st August) which then results in there weeks that we consider.\ncode below illustrates my point:\nAm I missing something? After improving my candidate generation process (co visitation) I want to be finally able to train my ranker. But i think as long as the feature distributions are so different I cant do that :-(.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11646918%2F47d953e7a2c949eeee56b460eaaaceb8%2FBildschirmfoto%20vom%202022-12-21%2023-04-59.png?generation=1671660337053128&alt=media)",
    "2067385": "Thank you very much, Chris! I am trying your method with my own features, however, I got some trouble with setting the data for training the XGB Ranker. When I set `group = train_group` in the xgb.DMatrix wraper, the model seems to learn nothing (loss = 1.0, predictions are all the same for all instances), when I omit this, the model seems fine, but the performance is much worse than with manual choosing (Recall of clicks are ~0.42). I'd appreciate your help in inspecting this bug. Thank you!",
    "2067367": "Thank you for sharing the idea! It's very useful! \nThere's one more question. In your example, the ground truth for 'clicks' in a session contains several different aids, yet in my understanding there's only one ground truth for type 'clicks' in the real task, is that true?",
    "2065267": "Thank you for this guide, it's very helpful.\n\nI have question about step 5. In your table aren't user_feat1 and user_feat2 should be the same for the same user regardless of any item? and we also need to merge user-item interaction feature on [user, item] right?\n\nAnd i think group sizes will always be different after merging ground truth label as it might add extra row",
    "2062223": "**UPDATE** I fixed the code in **Step 6**. The correct code for creating target dataframe (that we can merge unto our candidate dataframe) is\n\n    tar = pd.read_parquet('test_labels.parquet')\n    tar = tar.loc[ tar['type']=='clicks' ]\n    aids = tar.ground_truth.explode().astype('int32').rename('item')\n    tar = tar[['session']].astype('int32').rename({'session':'user'},axis=1)\n    tar = tar.merge(aids, left_index=True, right_index=True, how='left')\n    tar['click'] = 1\n\nThe old wrong code was\n\n    tar = pd.read_parquet('test_labels.parquet')\n    tar = tar.loc[ tar['type']=='clicks' ]\n    tar = tar.labels.explode().astype('int32')\n    tar.columns = ['user','item']\n    tar['click'] = 1",
    "2061240": "tnx for this awesome post @cdeotte !\ntoday i took the time to follow all the steps you describe and could improve CV\n0.5655 (co visitation approach with some improvement)\nto 0.5662.\nWithout much changes/improvements in my code to the above described pipeline.\nNow i should tune hyper parameters and create more and better candidates.",
    "2057813": "You are my angel.",
    "2056545": "Thank you so much. I have learned a lot from this thread about the ranker model !!",
    "2054786": "**UPDATE**: I corrected step 1. In step 1, we generate candidates for every user in **validation data** (i.e. last 1 week of train). We do not generate candidates for users in **train data** (i.e. first 3 weeks of train). Because we only have targets for validation users in step 6.",
    "2054181": "Hi Chris. Great project outline here! \n\nAm I right in saying the ranker training will be greatly impacted by the quality of our initial candidates? E.g if from our 50 candidates actually only 1 was correct, then for that session the ranker will only have positive sample to learn from? ",
    "3021487": "How are negative samples constructed?",
    "2303763": "Thanks @cdeotte . Viewing this thread in 2023, after the competition. The steps are clearly explained. I am replicating these for new project.",
    "2246867": "Thx for your sharing, it did help a lots\n",
    "2124574": "Thank you Chris, I learned a lot from your sharing！\n\nI have some doubts in the Inference step：\nI would like to know why there is no need to specify a group in Inference?\nHow does the model judge whether the data is in the same group during inference?\nDoes this mean that the inference group must be the same as the training group?\nI looked at the docs but didn't find what I want to know, am I missing something?",
    "2081143": "My ranker performed really poorly. Although all 3 maps during training were >= 0.9 my public score was 0.176. I concluded that the model was heavilly overfitted. Now I'm trying to reduce *subsample*, *colsample_bytree* and *max_depth* parameters. \n\nI also tried to give more weight to positive class by increasing *scale_pos_weight* parameter, but I got this error:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2142668%2F26c16db43d75790ac7a16645dad18103%2FError.JPG?generation=1672433503878014&alt=media)\n\nHas anybody tried to increase the weight of positive events and encountered this error? What was  the root cause?",
    "2078580": "How should user features be used? None of the users in the test data have appeared in the training set?",
    "2063686": "How the candidates generated? Since the number of history items of each user may less than you want (50 candidates for example).",
    "2054040": "",
    "2112772": "Thanks, it really helps",
    "2066728": "Thanks for sharing",
    "2107408": "Thank you for sharing this algorithm!",
    "2097356": "Thanks for sharing ~\n",
    "2065506": "Thanks for your explanation."
  }
}