{
  "id": 382776,
  "title": "ex-27th place solution",
  "url": "/competitions/otto-recommender-system/writeups/dreamteam-ex-27th-place-solution",
  "author_name": "",
  "post_date": "2023-02-08T22:33:21.290Z",
  "votes": 34,
  "comment_count": 15,
  "views": 0,
  "content": "<h2>Appreciation</h2>\n<p>First, I would like to thank the organizers and those who shared knowledge. Especially we would like to thank <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> for sharing so many valuable notebooks and datasets, and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for explaining many questions in detail.</p>\n<p><strong>Will Release Some Reproducible Code When The LB is Finalized</strong></p>\n<h2>CV &amp; LB Flow</h2>\n<p>Since <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>'s CV setting tracks public LB perfectly, we directly use his datasets in this competition. The following graph shows how we leverage his datasets for feature engineering, local validation, and submission generation.</p>\n<ul>\n<li>Location Validation uses his validation dataset.</li>\n<li>Submission generation uses his full dataset and reranking models trained from local validation.</li>\n<li>To avoid data leakage, covist matrix and Item2Vec are trained based on validation's <code>train.parquet</code> and <code>test.parquet</code>, full's <code>train.parquet</code> and <code>test.parquet</code> separately.</li>\n<li>metrics that we tracked are recall@200 for recall strategy, wdcg@20 (type weighted average ndcg scores of click, cart and order rankers) for model training, validation recall@20, and PB score after submission. The wdcg@20, validation recall@20, and PB score are perfectly aligned. This allows faster iteration for different components in a parallel way.</li>\n<li><strong>It's a deep collaboration within the team, with each team member responsible for several parts, as you can see in the graph</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563032%2F93c4b8a615882a249af1509539a7c374%2Fotto-dream-team-solution.png?generation=1675291808419108&amp;alt=media\" alt=\"otto-dream-team-solution\"></li>\n</ul>\n<h2>Retrieving</h2>\n<p>Our retrieving strategy is quite simple, we do some hyperparameters tuning on <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">public notebook</a>. Mainly the <code>topN</code> of each kind of co-visitation matrix and the <code>N_REC</code> for each session. The optimal <code>topN</code> for each co-visitation matrix is <code>100</code>, and the optimal <code>N_REC</code> is <code>200</code>. Then we use <a href=\"https://www.kaggle.com/tuongkhang\" target=\"_blank\">@tuongkhang</a> 's <a href=\"https://www.kaggle.com/code/tuongkhang/otto-pipeline2-lb-0-576/notebook\" target=\"_blank\">public notebook</a> to generate the <code>recall200</code> candidates for both local validation and test submission. Then we were able to get recall@200 for each action type as follows:</p>\n<ul>\n<li>clicks recall = 0.68420</li>\n<li>carts recall = 0.54552</li>\n<li>orders recall = 0.72831</li>\n<li>overall recall = 0.66906</li>\n</ul>\n<p><strong>What didn't work</strong></p>\n<ul>\n<li>BPR-based candidates</li>\n<li>Graph embedding(LINE) - based candidates</li>\n</ul>\n<h2>Rerank</h2>\n<h3>Feature Enginneering</h3>\n<p><strong>We kept monitoring the feature importance in Google Sheets, this allowed us to discuss and figure out new features efficiently</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563032%2Ff465e0967c097c07277c07f51d3080d8%2Ffeature-importance.png?generation=1675296313006485&amp;alt=media\" alt=\"feature-importance\"></p>\n<p>The 1st version of our feature set only contains <strong>78</strong> features, and we got <code>0.592</code> on LB and <code>0.5802</code> on CV. Finally added up to ~500 features, and we achieved <code>0.596</code> on LB with a single LGB model.</p>\n<h4>User Features</h4>\n<ul>\n<li>count features: (event|click|cart|order|unique aids)</li>\n<li>type weighted aggregated score</li>\n<li>time-weighted aggregated score</li>\n<li>ratio features (click2cart, click2order, click2cart_or_order)</li>\n<li>time features (1st seen|click|cart|order, last seen|click|cart|order, click|cart|order hours' sin|cos mean and median)</li>\n</ul>\n<h4>Aid Features</h4>\n<ul>\n<li>count features (event|click|cart|order|unique users)</li>\n<li>type weighted aggregated score</li>\n<li>time-weighted aggregated score</li>\n<li>ratio features (click2cart, click2order, <code>click2cart_or_order</code>)</li>\n<li>time features (1st seen|click|cart|order, last seen|click|cart|order, click|cart|order hours' sin|cos mean and median)</li>\n<li>co-visitation features: <code>n_covisit_{click, cart, order}</code>, <code>n_incovisit_{click, cart, order}</code>, <code>in_cosivist_{rank, avg_rank}</code><br>\n<em>co-visitation features generation</em></li>\n</ul>\n<h4>User-Aid Features</h4>\n<ul>\n<li>count features (event|click|cart|order)</li>\n<li>type weighted aggregated score</li>\n<li>time-weighted aggregated score</li>\n<li>log recency aggregated score (thanks to <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> again)</li>\n<li>co-visitation features: <code>{time, type, buy}_covisit_occurs_{timeweighted, typeweighted, recency weighted score}</code> <code>{time, type, buy}_covisit_rank_{timeweighted, typeweighted, recency weighted score}</code></li>\n<li>jaccard similarity between target aid and user interacted aids</li>\n<li><code>{click, cart, order}_word2vec_similarity_{min, max, mean, sum}</code> between target aid and user interacted aids</li>\n<li>BPR: user2aid BPR score, aid2user BPR score</li>\n<li>affinity_timedecay_7 (exponential weighted time decay score of action types)</li>\n</ul>\n<h3>Rankers</h3>\n<ul>\n<li>LightGBM with <code>lambdarank</code> objective</li>\n<li>XGboost with <code>rank:pairwise</code> objective</li>\n<li>CatBoost with QueryCrossEntropy loss (the best)</li>\n</ul>\n<h3>Ensemble</h3>\n<p>We searched the weights of each ranker's predicted scores based on the recall@20 of the local validation candidates using optuna, which means the recall@20 is calculated by weighted summing the scores of Tens of Millions of samples and calculating the recall@20. By leveraging the <code>cudf</code> and my implemented <code>calc_recall_fast</code> function, each trial only takes 700ms for carts (60M samples). We set the weight range to be <code>(-1, 1)</code> to get a better score:</p>\n<pre><code> () -&gt; :\n    weights = [ / (pred_cols)] * (pred_cols)\n     i  ((weights)):\n        weights[i] = trial.suggest_float(, -, , step=)\n     calc_recall_fast(\n        model_preds,\n        ground_truths,\n        sess_len_cumsum,\n        weights,\n        gt_cnt\n    )\nstudy = optuna.create_study(direction=, study_name=)\nstudy.optimize(objective, n_trials=, show_progress_bar=)\n</code></pre>",
  "messages": [
    {
      "id": "2124361",
      "postDate": "02/01/2023 00:38:46",
      "content": "<h2>Appreciation</h2>\n<p>First, I would like to thank the organizers and those who shared knowledge. Especially we would like to thank <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> for sharing so many valuable notebooks and datasets, and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for explaining many questions in detail.</p>\n<p><strong>Will Release Some Reproducible Code When The LB is Finalized</strong></p>\n<h2>CV &amp; LB Flow</h2>\n<p>Since <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>'s CV setting tracks public LB perfectly, we directly use his datasets in this competition. The following graph shows how we leverage his datasets for feature engineering, local validation, and submission generation.</p>\n<ul>\n<li>Location Validation uses his validation dataset.</li>\n<li>Submission generation uses his full dataset and reranking models trained from local validation.</li>\n<li>To avoid data leakage, covist matrix and Item2Vec are trained based on validation's <code>train.parquet</code> and <code>test.parquet</code>, full's <code>train.parquet</code> and <code>test.parquet</code> separately.</li>\n<li>metrics that we tracked are recall@200 for recall strategy, wdcg@20 (type weighted average ndcg scores of click, cart and order rankers) for model training, validation recall@20, and PB score after submission. The wdcg@20, validation recall@20, and PB score are perfectly aligned. This allows faster iteration for different components in a parallel way.</li>\n<li><strong>It's a deep collaboration within the team, with each team member responsible for several parts, as you can see in the graph</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563032%2F93c4b8a615882a249af1509539a7c374%2Fotto-dream-team-solution.png?generation=1675291808419108&amp;alt=media\" alt=\"otto-dream-team-solution\"></li>\n</ul>\n<h2>Retrieving</h2>\n<p>Our retrieving strategy is quite simple, we do some hyperparameters tuning on <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">public notebook</a>. Mainly the <code>topN</code> of each kind of co-visitation matrix and the <code>N_REC</code> for each session. The optimal <code>topN</code> for each co-visitation matrix is <code>100</code>, and the optimal <code>N_REC</code> is <code>200</code>. Then we use <a href=\"https://www.kaggle.com/tuongkhang\" target=\"_blank\">@tuongkhang</a> 's <a href=\"https://www.kaggle.com/code/tuongkhang/otto-pipeline2-lb-0-576/notebook\" target=\"_blank\">public notebook</a> to generate the <code>recall200</code> candidates for both local validation and test submission. Then we were able to get recall@200 for each action type as follows:</p>\n<ul>\n<li>clicks recall = 0.68420</li>\n<li>carts recall = 0.54552</li>\n<li>orders recall = 0.72831</li>\n<li>overall recall = 0.66906</li>\n</ul>\n<p><strong>What didn't work</strong></p>\n<ul>\n<li>BPR-based candidates</li>\n<li>Graph embedding(LINE) - based candidates</li>\n</ul>\n<h2>Rerank</h2>\n<h3>Feature Enginneering</h3>\n<p><strong>We kept monitoring the feature importance in Google Sheets, this allowed us to discuss and figure out new features efficiently</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563032%2Ff465e0967c097c07277c07f51d3080d8%2Ffeature-importance.png?generation=1675296313006485&amp;alt=media\" alt=\"feature-importance\"></p>\n<p>The 1st version of our feature set only contains <strong>78</strong> features, and we got <code>0.592</code> on LB and <code>0.5802</code> on CV. Finally added up to ~500 features, and we achieved <code>0.596</code> on LB with a single LGB model.</p>\n<h4>User Features</h4>\n<ul>\n<li>count features: (event|click|cart|order|unique aids)</li>\n<li>type weighted aggregated score</li>\n<li>time-weighted aggregated score</li>\n<li>ratio features (click2cart, click2order, click2cart_or_order)</li>\n<li>time features (1st seen|click|cart|order, last seen|click|cart|order, click|cart|order hours' sin|cos mean and median)</li>\n</ul>\n<h4>Aid Features</h4>\n<ul>\n<li>count features (event|click|cart|order|unique users)</li>\n<li>type weighted aggregated score</li>\n<li>time-weighted aggregated score</li>\n<li>ratio features (click2cart, click2order, <code>click2cart_or_order</code>)</li>\n<li>time features (1st seen|click|cart|order, last seen|click|cart|order, click|cart|order hours' sin|cos mean and median)</li>\n<li>co-visitation features: <code>n_covisit_{click, cart, order}</code>, <code>n_incovisit_{click, cart, order}</code>, <code>in_cosivist_{rank, avg_rank}</code><br>\n<em>co-visitation features generation</em></li>\n</ul>\n<h4>User-Aid Features</h4>\n<ul>\n<li>count features (event|click|cart|order)</li>\n<li>type weighted aggregated score</li>\n<li>time-weighted aggregated score</li>\n<li>log recency aggregated score (thanks to <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> again)</li>\n<li>co-visitation features: <code>{time, type, buy}_covisit_occurs_{timeweighted, typeweighted, recency weighted score}</code> <code>{time, type, buy}_covisit_rank_{timeweighted, typeweighted, recency weighted score}</code></li>\n<li>jaccard similarity between target aid and user interacted aids</li>\n<li><code>{click, cart, order}_word2vec_similarity_{min, max, mean, sum}</code> between target aid and user interacted aids</li>\n<li>BPR: user2aid BPR score, aid2user BPR score</li>\n<li>affinity_timedecay_7 (exponential weighted time decay score of action types)</li>\n</ul>\n<h3>Rankers</h3>\n<ul>\n<li>LightGBM with <code>lambdarank</code> objective</li>\n<li>XGboost with <code>rank:pairwise</code> objective</li>\n<li>CatBoost with QueryCrossEntropy loss (the best)</li>\n</ul>\n<h3>Ensemble</h3>\n<p>We searched the weights of each ranker's predicted scores based on the recall@20 of the local validation candidates using optuna, which means the recall@20 is calculated by weighted summing the scores of Tens of Millions of samples and calculating the recall@20. By leveraging the <code>cudf</code> and my implemented <code>calc_recall_fast</code> function, each trial only takes 700ms for carts (60M samples). We set the weight range to be <code>(-1, 1)</code> to get a better score:</p>\n<pre><code> () -&gt; :\n    weights = [ / (pred_cols)] * (pred_cols)\n     i  ((weights)):\n        weights[i] = trial.suggest_float(, -, , step=)\n     calc_recall_fast(\n        model_preds,\n        ground_truths,\n        sess_len_cumsum,\n        weights,\n        gt_cnt\n    )\nstudy = optuna.create_study(direction=, study_name=)\nstudy.optimize(objective, n_trials=, show_progress_bar=)\n</code></pre>",
      "rawMarkdown": "## Appreciation\n\nFirst, I would like to thank the organizers and those who shared knowledge. Especially we would like to thank @radek1 for sharing so many valuable notebooks and datasets, and @cdeotte for explaining many questions in detail.\n\n**Will Release Some Reproducible Code When The LB is Finalized**\n\n## CV & LB Flow\nSince @radek1's CV setting tracks public LB perfectly, we directly use his datasets in this competition. The following graph shows how we leverage his datasets for feature engineering, local validation, and submission generation.\n\n* Location Validation uses his validation dataset.\n* Submission generation uses his full dataset and reranking models trained from local validation.\n* To avoid data leakage, covist matrix and Item2Vec are trained based on validation's `train.parquet` and `test.parquet`, full's `train.parquet` and `test.parquet` separately.\n* metrics that we tracked are recall@200 for recall strategy, wdcg@20 (type weighted average ndcg scores of click, cart and order rankers) for model training, validation recall@20, and PB score after submission. The wdcg@20, validation recall@20, and PB score are perfectly aligned. This allows faster iteration for different components in a parallel way.\n* **It's a deep collaboration within the team, with each team member responsible for several parts, as you can see in the graph**\n![otto-dream-team-solution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563032%2F93c4b8a615882a249af1509539a7c374%2Fotto-dream-team-solution.png?generation=1675291808419108&alt=media)\n\n\n## Retrieving\n\nOur retrieving strategy is quite simple, we do some hyperparameters tuning on @cdeotte 's [public notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575). Mainly the `topN` of each kind of co-visitation matrix and the `N_REC` for each session. The optimal `topN` for each co-visitation matrix is `100`, and the optimal `N_REC` is `200`. Then we use @tuongkhang 's [public notebook](https://www.kaggle.com/code/tuongkhang/otto-pipeline2-lb-0-576/notebook) to generate the `recall200` candidates for both local validation and test submission. Then we were able to get recall@200 for each action type as follows:\n\n* clicks recall = 0.68420\n* carts recall = 0.54552\n* orders recall = 0.72831\n* overall recall = 0.66906\n\n**What didn't work**\n* BPR-based candidates\n* Graph embedding(LINE) - based candidates\n\n## Rerank\n\n### Feature Enginneering\n**We kept monitoring the feature importance in Google Sheets, this allowed us to discuss and figure out new features efficiently**\n\n![feature-importance](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563032%2Ff465e0967c097c07277c07f51d3080d8%2Ffeature-importance.png?generation=1675296313006485&alt=media)\n\nThe 1st version of our feature set only contains **78** features, and we got `0.592` on LB and `0.5802` on CV. Finally added up to ~500 features, and we achieved `0.596` on LB with a single LGB model.\n\n#### User Features\n* count features: (event|click|cart|order|unique aids)\n* type weighted aggregated score\n* time-weighted aggregated score\n* ratio features (click2cart, click2order, click2cart_or_order)\n* time features (1st seen|click|cart|order, last seen|click|cart|order, click|cart|order hours' sin|cos mean and median)\n\n#### Aid Features\n* count features (event|click|cart|order|unique users)\n* type weighted aggregated score\n* time-weighted aggregated score\n* ratio features (click2cart, click2order, `click2cart_or_order`)\n* time features (1st seen|click|cart|order, last seen|click|cart|order, click|cart|order hours' sin|cos mean and median)\n* co-visitation features: `n_covisit_{click, cart, order}`, `n_incovisit_{click, cart, order}`, `in_cosivist_{rank, avg_rank}`\n*co-visitation features generation*\n\n\n#### User-Aid Features\n* count features (event|click|cart|order)\n* type weighted aggregated score\n* time-weighted aggregated score\n* log recency aggregated score (thanks to @radek1 again)\n* co-visitation features: `{time, type, buy}_covisit_occurs_{timeweighted, typeweighted, recency weighted score}` `{time, type, buy}_covisit_rank_{timeweighted, typeweighted, recency weighted score}`\n* jaccard similarity between target aid and user interacted aids\n* `{click, cart, order}_word2vec_similarity_{min, max, mean, sum}` between target aid and user interacted aids\n* BPR: user2aid BPR score, aid2user BPR score\n* affinity_timedecay_7 (exponential weighted time decay score of action types)\n\n### Rankers\n* LightGBM with `lambdarank` objective\n* XGboost with `rank:pairwise` objective\n* CatBoost with QueryCrossEntropy loss (the best)\n\n### Ensemble\nWe searched the weights of each ranker's predicted scores based on the recall@20 of the local validation candidates using optuna, which means the recall@20 is calculated by weighted summing the scores of Tens of Millions of samples and calculating the recall@20. By leveraging the `cudf` and my implemented `calc_recall_fast` function, each trial only takes 700ms for carts (60M samples). We set the weight range to be `(-1, 1)` to get a better score:\n```Python\ndef objective(trial: optuna.trial.Trial) -> float:\n    weights = [1 / len(pred_cols)] * len(pred_cols)\n    for i in range(len(weights)):\n        weights[i] = trial.suggest_float(f\"w{i}\", -1, 1, step=0.001)\n    return calc_recall_fast(\n        model_preds,\n        ground_truths,\n        sess_len_cumsum,\n        weights,\n        gt_cnt\n    )\nstudy = optuna.create_study(direction=\"maximize\", study_name=\"opt-model-weights\")\nstudy.optimize(objective, n_trials=1000, show_progress_bar=True)\n```",
      "votes": null
    },
    {
      "id": "2124523",
      "postDate": "02/01/2023 04:05:46",
      "content": "<p>Congrats. It's exactly same as our solution but I guess you have better features.</p>",
      "rawMarkdown": "Congrats. It's exactly same as our solution but I guess you have better features.",
      "votes": null
    },
    {
      "id": "2125828",
      "postDate": "02/01/2023 23:26:02",
      "content": "<p>Thanks, I'm updating the feature engineering and ranker parts. You can check later</p>",
      "rawMarkdown": "Thanks, I'm updating the feature engineering and ranker parts. You can check later",
      "votes": null
    },
    {
      "id": "2125917",
      "postDate": "02/02/2023 01:55:48",
      "content": "<p>Very clear pipeline plot! When you constructing the training samples for RANK model, do you use the train part (first three weeks) and random split it to input sequence and label?</p>",
      "rawMarkdown": "Very clear pipeline plot! When you constructing the training samples for RANK model, do you use the train part (first three weeks) and random split it to input sequence and label?",
      "votes": null
    },
    {
      "id": "2125926",
      "postDate": "02/02/2023 02:17:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a> , thanks!  In short, <code>(session, aid, label)</code> is generated based on the 4th week's 1st split and 2nd split, whereas the features are generated based on 3 weeks + 4th week's 1st split.</p>\n<p><code>(session, aid, label)</code> in the training samples are based on radek's <code>test.parquet</code> and <code>test_labels.parquet</code> in the his local validation dataset, where the features are generated based on the <code>train.parquet</code> and <code>test.parquet</code> in this dataset.</p>",
      "rawMarkdown": "Hi @evilpsycho42 , thanks!  In short, `(session, aid, label)` is generated based on the 4th week's 1st split and 2nd split, whereas the features are generated based on 3 weeks + 4th week's 1st split.\n\n`(session, aid, label)` in the training samples are based on radek's `test.parquet` and `test_labels.parquet` in the his local validation dataset, where the features are generated based on the `train.parquet` and `test.parquet` in this dataset.",
      "votes": null
    },
    {
      "id": "2126256",
      "postDate": "02/02/2023 06:35:05",
      "content": "<p>Congrats! and thanks for sharing</p>",
      "rawMarkdown": "Congrats! and thanks for sharing",
      "votes": null
    },
    {
      "id": "2126644",
      "postDate": "02/02/2023 11:01:01",
      "content": "<p>How much did you benefit from your ensemble method?</p>",
      "rawMarkdown": "How much did you benefit from your ensemble method?",
      "votes": null
    },
    {
      "id": "2126931",
      "postDate": "02/02/2023 15:14:22",
      "content": "<p>Not that much, cv from 0.584491 to 0.585179</p>",
      "rawMarkdown": "Not that much, cv from 0.584491 to 0.585179",
      "votes": null
    },
    {
      "id": "2132916",
      "postDate": "02/07/2023 05:41:02",
      "content": "<p>When our team also used optuna to adjust my model, we always found that the parameters searched by optuna and the final model effect could not be reproduced. Has your team encountered this?</p>",
      "rawMarkdown": "When our team also used optuna to adjust my model, we always found that the parameters searched by optuna and the final model effect could not be reproduced. Has your team encountered this?",
      "votes": null
    },
    {
      "id": "2133063",
      "postDate": "02/07/2023 08:28:51",
      "content": "<blockquote>\n  <p>the parameters searched by optuna and the final model effect could not be reproduced</p>\n</blockquote>\n<p>Did you compare this to the local CV score? If so, given the weights, the local score is determined.</p>",
      "rawMarkdown": "> the parameters searched by optuna and the final model effect could not be reproduced\n\nDid you compare this to the local CV score? If so, given the weights, the local score is determined.",
      "votes": null
    },
    {
      "id": "2133146",
      "postDate": "02/07/2023 09:15:33",
      "content": "<p>No, I use optuna to optimize the parameters of my tree model and not the model fusion parameters. However, the parameter substitution model searched by optuna cannot reproduce the corresponding results.</p>",
      "rawMarkdown": "No, I use optuna to optimize the parameters of my tree model and not the model fusion parameters. However, the parameter substitution model searched by optuna cannot reproduce the corresponding results.",
      "votes": null
    },
    {
      "id": "2133272",
      "postDate": "02/07/2023 10:38:31",
      "content": "<p>For catboost and xgboost GPU implementation, there is some randomness during training, even you set random_state, the result cannot be 100% reproduced.<br>\nBut for lgbm, if you set random_state, the result can be reproduced to be exactly the same. (gradient boosting mode, cpu training)</p>",
      "rawMarkdown": "For catboost and xgboost GPU implementation, there is some randomness during training, even you set random_state, the result cannot be 100% reproduced.\nBut for lgbm, if you set random_state, the result can be reproduced to be exactly the same. (gradient boosting mode, cpu training)",
      "votes": null
    },
    {
      "id": "2133317",
      "postDate": "02/07/2023 11:08:16",
      "content": "<p>I found that even after doing this; I still can't reproduce my results (I will substitute the model parameters in the picture); here is part of my code。</p>\n<pre><code> random\nrandom.seed()\n\n ():\n    scores = \n    reg_alpha = trial.suggest_float(, , ),\n    reg_lambda = trial.suggest_float(, , ),\n    num_leaves = trial.suggest_int(, , ),\n    min_child_samples =  trial.suggest_int(, , ),\n    colsample_bytree =  trial.suggest_float(, , ),\n    learning_rate = trial.suggest_float(, , ),\n     fold,(train_idx, valid_idx)  (skf.split(df_feature_sample, df_feature_sample[], groups=df_feature_sample[] )):\n        X_train = df_feature_sample.loc[train_idx, FEATURES]\n        y_train = df_feature_sample.loc[train_idx, ]\n        X_valid = df_feature_sample.loc[valid_idx, FEATURES]\n        y_valid = df_feature_sample.loc[valid_idx, ]\n        group_train=df_feature_sample.loc[train_idx].groupby().size()[df_feature_sample.loc[train_idx].session.drop_duplicates()]\n        group_vali=df_feature_sample.loc[valid_idx].groupby().size()[df_feature_sample.loc[valid_idx].session.drop_duplicates()]\n        ranker = lgb.LGBMRanker(\n                                objective=,\n                                metric=,\n                                boosting_type=,\n                                importance_type=,\n                                n_estimators=,\n                                eval_at= ,\n                                early_stopping_round=,\n                                n_jobs = -,\n                                random_state=,\n                                reg_alpha = reg_alpha,\n                                reg_lambda = reg_lambda,\n                                num_leaves = num_leaves,\n                                min_child_samples = min_child_samples,\n                                colsample_bytree = colsample_bytree,\n                                learning_rate=learning_rate,\n                                )        \n\n        ranker.fit(X_train,\n                   y_train, \n                   eval_set=[(X_train, y_train), (X_valid, y_valid)],\n                   group=group_train,\n                   eval_group=[group_train,group_vali])\n        score = gen_score(ranker)\n        scores += score\n         score&lt;:\n              * score\n     scores\n\n\nsampler = TPESampler(seed=) \nstudy = optuna.create_study(\n        direction=,\n        sampler=sampler,\n        study_name=)\nstudy.optimize(objective, n_trials=)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3562400%2Fbc7e0335bcba2de98cda856220ad99ae%2F111111.jpg?generation=1675767526395608&amp;alt=media\" alt=\"--\"></p>",
      "rawMarkdown": "I found that even after doing this; I still can't reproduce my results (I will substitute the model parameters in the picture); here is part of my code。\n```Python\nimport random\nrandom.seed(42)\n\ndef objective(trial):\n    scores = 0\n    reg_alpha = trial.suggest_float('reg_alpha', 1e-8, 10.0),\n    reg_lambda = trial.suggest_float('reg_lambda', 1e-8, 10.0),\n    num_leaves = trial.suggest_int('num_leaves', 2, 256),\n    min_child_samples =  trial.suggest_int('min_child_samples', 10, 128),\n    colsample_bytree =  trial.suggest_float('colsample_bytree', 0.1, 1),\n    learning_rate = trial.suggest_float('learning_rate', 0.1, 0.5),\n    for fold,(train_idx, valid_idx) in enumerate(skf.split(df_feature_sample, df_feature_sample['orders'], groups=df_feature_sample['session'] )):\n        X_train = df_feature_sample.loc[train_idx, FEATURES]\n        y_train = df_feature_sample.loc[train_idx, 'orders']\n        X_valid = df_feature_sample.loc[valid_idx, FEATURES]\n        y_valid = df_feature_sample.loc[valid_idx, 'orders']\n        group_train=df_feature_sample.loc[train_idx].groupby('session').size()[df_feature_sample.loc[train_idx].session.drop_duplicates()]\n        group_vali=df_feature_sample.loc[valid_idx].groupby('session').size()[df_feature_sample.loc[valid_idx].session.drop_duplicates()]\n        ranker = lgb.LGBMRanker(\n                                objective=\"lambdarank\",\n                                metric=\"ndcg\",\n                                boosting_type=\"dart\",\n                                importance_type='gain',\n                                n_estimators=700,\n                                eval_at= 20,\n                                early_stopping_round=80,\n                                n_jobs = -1,\n                                random_state=42,\n                                reg_alpha = reg_alpha,\n                                reg_lambda = reg_lambda,\n                                num_leaves = num_leaves,\n                                min_child_samples = min_child_samples,\n                                colsample_bytree = colsample_bytree,\n                                learning_rate=learning_rate,\n                                )        \n        \n        ranker.fit(X_train,\n                   y_train, \n                   eval_set=[(X_train, y_train), (X_valid, y_valid)],\n                   group=group_train,\n                   eval_group=[group_train,group_vali])\n        score = gen_score(ranker)\n        scores += score\n        if score<0.66:\n            return 5 * score\n    return scores\n\n\nsampler = TPESampler(seed=42) \nstudy = optuna.create_study(\n        direction='maximize',\n        sampler=sampler,\n        study_name='otto')\nstudy.optimize(objective, n_trials=100)\n```\n![--](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3562400%2Fbc7e0335bcba2de98cda856220ad99ae%2F111111.jpg?generation=1675767526395608&alt=media)",
      "votes": null
    },
    {
      "id": "2133399",
      "postDate": "02/07/2023 11:49:01",
      "content": "<p>First, I guess dropout used in DART mode induces some extent of randomness, make dart mode difficult to be 100% deterministic. Just my speculation.<br>\nSecond, the native early stop API provided by lgbm should not be used together with DART mode. It won't work. If you want to use early stop in DART mode, you need to write your own early stop callback. </p>",
      "rawMarkdown": "First, I guess dropout used in DART mode induces some extent of randomness, make dart mode difficult to be 100% deterministic. Just my speculation.\nSecond, the native early stop API provided by lgbm should not be used together with DART mode. It won't work. If you want to use early stop in DART mode, you need to write your own early stop callback.",
      "votes": null
    },
    {
      "id": "2133578",
      "postDate": "02/07/2023 13:48:17",
      "content": "<p>thanks, i think it's possible</p>",
      "rawMarkdown": "thanks, i think it's possible",
      "votes": null
    },
    {
      "id": "2135849",
      "postDate": "02/08/2023 22:30:35",
      "content": "<p>Both LGB and CatBoost's results are not reproducible. However, after many training experiments, I found that the gap between the tuning stage and the training stage can be decreased by increasing the <code>early_stopping_round</code>, I raised this param to <code>300</code> for this competition. The gap is relatively small. <br>\nIf you set a relatively small <code>early_stopping_round</code> e.g. <code>20</code> or <code>40</code>. Then most likely the early stopping is triggered before your model actually converges. Because of the randomness (even with a fixed seed), it could be triggered earlier during training than tuning. Then you'll see a big gap between training and tuning regarding the model performance.</p>",
      "rawMarkdown": "Both LGB and CatBoost's results are not reproducible. However, after many training experiments, I found that the gap between the tuning stage and the training stage can be decreased by increasing the `early_stopping_round`, I raised this param to `300` for this competition. The gap is relatively small. \nIf you set a relatively small `early_stopping_round` e.g. `20` or `40`. Then most likely the early stopping is triggered before your model actually converges. Because of the randomness (even with a fixed seed), it could be triggered earlier during training than tuning. Then you'll see a big gap between training and tuning regarding the model performance.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2124523,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "02/01/2023 04:05:46",
      "content": "<p>Congrats. It's exactly same as our solution but I guess you have better features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2125828,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "02/01/2023 23:26:02",
          "content": "<p>Thanks, I'm updating the feature engineering and ranker parts. You can check later</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2125917,
      "author_name": "evilpsycho42",
      "author_url": "",
      "post_date": "02/02/2023 01:55:48",
      "content": "<p>Very clear pipeline plot! When you constructing the training samples for RANK model, do you use the train part (first three weeks) and random split it to input sequence and label?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2125926,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "02/02/2023 02:17:40",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a> , thanks!  In short, <code>(session, aid, label)</code> is generated based on the 4th week's 1st split and 2nd split, whereas the features are generated based on 3 weeks + 4th week's 1st split.</p>\n<p><code>(session, aid, label)</code> in the training samples are based on radek's <code>test.parquet</code> and <code>test_labels.parquet</code> in the his local validation dataset, where the features are generated based on the <code>train.parquet</code> and <code>test.parquet</code> in this dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2126256,
      "author_name": "ajisamudra",
      "author_url": "",
      "post_date": "02/02/2023 06:35:05",
      "content": "<p>Congrats! and thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2126644,
      "author_name": "zachary666",
      "author_url": "",
      "post_date": "02/02/2023 11:01:01",
      "content": "<p>How much did you benefit from your ensemble method?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2126931,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "02/02/2023 15:14:22",
          "content": "<p>Not that much, cv from 0.584491 to 0.585179</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2132916,
      "author_name": "yasso1",
      "author_url": "",
      "post_date": "02/07/2023 05:41:02",
      "content": "<p>When our team also used optuna to adjust my model, we always found that the parameters searched by optuna and the final model effect could not be reproduced. Has your team encountered this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2133063,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "02/07/2023 08:28:51",
          "content": "<blockquote>\n  <p>the parameters searched by optuna and the final model effect could not be reproduced</p>\n</blockquote>\n<p>Did you compare this to the local CV score? If so, given the weights, the local score is determined.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2133146,
              "author_name": "yasso1",
              "author_url": "",
              "post_date": "02/07/2023 09:15:33",
              "content": "<p>No, I use optuna to optimize the parameters of my tree model and not the model fusion parameters. However, the parameter substitution model searched by optuna cannot reproduce the corresponding results.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2133272,
                  "author_name": "buumoo",
                  "author_url": "",
                  "post_date": "02/07/2023 10:38:31",
                  "content": "<p>For catboost and xgboost GPU implementation, there is some randomness during training, even you set random_state, the result cannot be 100% reproduced.<br>\nBut for lgbm, if you set random_state, the result can be reproduced to be exactly the same. (gradient boosting mode, cpu training)</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2133317,
                      "author_name": "yasso1",
                      "author_url": "",
                      "post_date": "02/07/2023 11:08:16",
                      "content": "<p>I found that even after doing this; I still can't reproduce my results (I will substitute the model parameters in the picture); here is part of my code。</p>\n<pre><code> random\nrandom.seed()\n\n ():\n    scores = \n    reg_alpha = trial.suggest_float(, , ),\n    reg_lambda = trial.suggest_float(, , ),\n    num_leaves = trial.suggest_int(, , ),\n    min_child_samples =  trial.suggest_int(, , ),\n    colsample_bytree =  trial.suggest_float(, , ),\n    learning_rate = trial.suggest_float(, , ),\n     fold,(train_idx, valid_idx)  (skf.split(df_feature_sample, df_feature_sample[], groups=df_feature_sample[] )):\n        X_train = df_feature_sample.loc[train_idx, FEATURES]\n        y_train = df_feature_sample.loc[train_idx, ]\n        X_valid = df_feature_sample.loc[valid_idx, FEATURES]\n        y_valid = df_feature_sample.loc[valid_idx, ]\n        group_train=df_feature_sample.loc[train_idx].groupby().size()[df_feature_sample.loc[train_idx].session.drop_duplicates()]\n        group_vali=df_feature_sample.loc[valid_idx].groupby().size()[df_feature_sample.loc[valid_idx].session.drop_duplicates()]\n        ranker = lgb.LGBMRanker(\n                                objective=,\n                                metric=,\n                                boosting_type=,\n                                importance_type=,\n                                n_estimators=,\n                                eval_at= ,\n                                early_stopping_round=,\n                                n_jobs = -,\n                                random_state=,\n                                reg_alpha = reg_alpha,\n                                reg_lambda = reg_lambda,\n                                num_leaves = num_leaves,\n                                min_child_samples = min_child_samples,\n                                colsample_bytree = colsample_bytree,\n                                learning_rate=learning_rate,\n                                )        \n\n        ranker.fit(X_train,\n                   y_train, \n                   eval_set=[(X_train, y_train), (X_valid, y_valid)],\n                   group=group_train,\n                   eval_group=[group_train,group_vali])\n        score = gen_score(ranker)\n        scores += score\n         score&lt;:\n              * score\n     scores\n\n\nsampler = TPESampler(seed=) \nstudy = optuna.create_study(\n        direction=,\n        sampler=sampler,\n        study_name=)\nstudy.optimize(objective, n_trials=)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3562400%2Fbc7e0335bcba2de98cda856220ad99ae%2F111111.jpg?generation=1675767526395608&amp;alt=media\" alt=\"--\"></p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2133399,
                          "author_name": "buumoo",
                          "author_url": "",
                          "post_date": "02/07/2023 11:49:01",
                          "content": "<p>First, I guess dropout used in DART mode induces some extent of randomness, make dart mode difficult to be 100% deterministic. Just my speculation.<br>\nSecond, the native early stop API provided by lgbm should not be used together with DART mode. It won't work. If you want to use early stop in DART mode, you need to write your own early stop callback. </p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2133578,
                              "author_name": "yasso1",
                              "author_url": "",
                              "post_date": "02/07/2023 13:48:17",
                              "content": "<p>thanks, i think it's possible</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2135849,
                                  "author_name": "wuwenmin",
                                  "author_url": "",
                                  "post_date": "02/08/2023 22:30:35",
                                  "content": "<p>Both LGB and CatBoost's results are not reproducible. However, after many training experiments, I found that the gap between the tuning stage and the training stage can be decreased by increasing the <code>early_stopping_round</code>, I raised this param to <code>300</code> for this competition. The gap is relatively small. <br>\nIf you set a relatively small <code>early_stopping_round</code> e.g. <code>20</code> or <code>40</code>. Then most likely the early stopping is triggered before your model actually converges. Because of the randomness (even with a fixed seed), it could be triggered earlier during training than tuning. Then you'll see a big gap between training and tuning regarding the model performance.</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2124361": "## Appreciation\n\nFirst, I would like to thank the organizers and those who shared knowledge. Especially we would like to thank @radek1 for sharing so many valuable notebooks and datasets, and @cdeotte for explaining many questions in detail.\n\n**Will Release Some Reproducible Code When The LB is Finalized**\n\n## CV & LB Flow\nSince @radek1's CV setting tracks public LB perfectly, we directly use his datasets in this competition. The following graph shows how we leverage his datasets for feature engineering, local validation, and submission generation.\n\n* Location Validation uses his validation dataset.\n* Submission generation uses his full dataset and reranking models trained from local validation.\n* To avoid data leakage, covist matrix and Item2Vec are trained based on validation's `train.parquet` and `test.parquet`, full's `train.parquet` and `test.parquet` separately.\n* metrics that we tracked are recall@200 for recall strategy, wdcg@20 (type weighted average ndcg scores of click, cart and order rankers) for model training, validation recall@20, and PB score after submission. The wdcg@20, validation recall@20, and PB score are perfectly aligned. This allows faster iteration for different components in a parallel way.\n* **It's a deep collaboration within the team, with each team member responsible for several parts, as you can see in the graph**\n![otto-dream-team-solution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563032%2F93c4b8a615882a249af1509539a7c374%2Fotto-dream-team-solution.png?generation=1675291808419108&alt=media)\n\n\n## Retrieving\n\nOur retrieving strategy is quite simple, we do some hyperparameters tuning on @cdeotte 's [public notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575). Mainly the `topN` of each kind of co-visitation matrix and the `N_REC` for each session. The optimal `topN` for each co-visitation matrix is `100`, and the optimal `N_REC` is `200`. Then we use @tuongkhang 's [public notebook](https://www.kaggle.com/code/tuongkhang/otto-pipeline2-lb-0-576/notebook) to generate the `recall200` candidates for both local validation and test submission. Then we were able to get recall@200 for each action type as follows:\n\n* clicks recall = 0.68420\n* carts recall = 0.54552\n* orders recall = 0.72831\n* overall recall = 0.66906\n\n**What didn't work**\n* BPR-based candidates\n* Graph embedding(LINE) - based candidates\n\n## Rerank\n\n### Feature Enginneering\n**We kept monitoring the feature importance in Google Sheets, this allowed us to discuss and figure out new features efficiently**\n\n![feature-importance](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563032%2Ff465e0967c097c07277c07f51d3080d8%2Ffeature-importance.png?generation=1675296313006485&alt=media)\n\nThe 1st version of our feature set only contains **78** features, and we got `0.592` on LB and `0.5802` on CV. Finally added up to ~500 features, and we achieved `0.596` on LB with a single LGB model.\n\n#### User Features\n* count features: (event|click|cart|order|unique aids)\n* type weighted aggregated score\n* time-weighted aggregated score\n* ratio features (click2cart, click2order, click2cart_or_order)\n* time features (1st seen|click|cart|order, last seen|click|cart|order, click|cart|order hours' sin|cos mean and median)\n\n#### Aid Features\n* count features (event|click|cart|order|unique users)\n* type weighted aggregated score\n* time-weighted aggregated score\n* ratio features (click2cart, click2order, `click2cart_or_order`)\n* time features (1st seen|click|cart|order, last seen|click|cart|order, click|cart|order hours' sin|cos mean and median)\n* co-visitation features: `n_covisit_{click, cart, order}`, `n_incovisit_{click, cart, order}`, `in_cosivist_{rank, avg_rank}`\n*co-visitation features generation*\n\n\n#### User-Aid Features\n* count features (event|click|cart|order)\n* type weighted aggregated score\n* time-weighted aggregated score\n* log recency aggregated score (thanks to @radek1 again)\n* co-visitation features: `{time, type, buy}_covisit_occurs_{timeweighted, typeweighted, recency weighted score}` `{time, type, buy}_covisit_rank_{timeweighted, typeweighted, recency weighted score}`\n* jaccard similarity between target aid and user interacted aids\n* `{click, cart, order}_word2vec_similarity_{min, max, mean, sum}` between target aid and user interacted aids\n* BPR: user2aid BPR score, aid2user BPR score\n* affinity_timedecay_7 (exponential weighted time decay score of action types)\n\n### Rankers\n* LightGBM with `lambdarank` objective\n* XGboost with `rank:pairwise` objective\n* CatBoost with QueryCrossEntropy loss (the best)\n\n### Ensemble\nWe searched the weights of each ranker's predicted scores based on the recall@20 of the local validation candidates using optuna, which means the recall@20 is calculated by weighted summing the scores of Tens of Millions of samples and calculating the recall@20. By leveraging the `cudf` and my implemented `calc_recall_fast` function, each trial only takes 700ms for carts (60M samples). We set the weight range to be `(-1, 1)` to get a better score:\n```Python\ndef objective(trial: optuna.trial.Trial) -> float:\n    weights = [1 / len(pred_cols)] * len(pred_cols)\n    for i in range(len(weights)):\n        weights[i] = trial.suggest_float(f\"w{i}\", -1, 1, step=0.001)\n    return calc_recall_fast(\n        model_preds,\n        ground_truths,\n        sess_len_cumsum,\n        weights,\n        gt_cnt\n    )\nstudy = optuna.create_study(direction=\"maximize\", study_name=\"opt-model-weights\")\nstudy.optimize(objective, n_trials=1000, show_progress_bar=True)\n```",
    "2124523": "Congrats. It's exactly same as our solution but I guess you have better features.",
    "2125828": "Thanks, I'm updating the feature engineering and ranker parts. You can check later",
    "2125917": "Very clear pipeline plot! When you constructing the training samples for RANK model, do you use the train part (first three weeks) and random split it to input sequence and label?",
    "2125926": "Hi @evilpsycho42 , thanks!  In short, `(session, aid, label)` is generated based on the 4th week's 1st split and 2nd split, whereas the features are generated based on 3 weeks + 4th week's 1st split.\n\n`(session, aid, label)` in the training samples are based on radek's `test.parquet` and `test_labels.parquet` in the his local validation dataset, where the features are generated based on the `train.parquet` and `test.parquet` in this dataset.",
    "2126256": "Congrats! and thanks for sharing",
    "2126644": "How much did you benefit from your ensemble method?",
    "2126931": "Not that much, cv from 0.584491 to 0.585179",
    "2132916": "When our team also used optuna to adjust my model, we always found that the parameters searched by optuna and the final model effect could not be reproduced. Has your team encountered this?",
    "2133063": "> the parameters searched by optuna and the final model effect could not be reproduced\n\nDid you compare this to the local CV score? If so, given the weights, the local score is determined.",
    "2133146": "No, I use optuna to optimize the parameters of my tree model and not the model fusion parameters. However, the parameter substitution model searched by optuna cannot reproduce the corresponding results.",
    "2133272": "For catboost and xgboost GPU implementation, there is some randomness during training, even you set random_state, the result cannot be 100% reproduced.\nBut for lgbm, if you set random_state, the result can be reproduced to be exactly the same. (gradient boosting mode, cpu training)",
    "2133317": "I found that even after doing this; I still can't reproduce my results (I will substitute the model parameters in the picture); here is part of my code。\n```Python\nimport random\nrandom.seed(42)\n\ndef objective(trial):\n    scores = 0\n    reg_alpha = trial.suggest_float('reg_alpha', 1e-8, 10.0),\n    reg_lambda = trial.suggest_float('reg_lambda', 1e-8, 10.0),\n    num_leaves = trial.suggest_int('num_leaves', 2, 256),\n    min_child_samples =  trial.suggest_int('min_child_samples', 10, 128),\n    colsample_bytree =  trial.suggest_float('colsample_bytree', 0.1, 1),\n    learning_rate = trial.suggest_float('learning_rate', 0.1, 0.5),\n    for fold,(train_idx, valid_idx) in enumerate(skf.split(df_feature_sample, df_feature_sample['orders'], groups=df_feature_sample['session'] )):\n        X_train = df_feature_sample.loc[train_idx, FEATURES]\n        y_train = df_feature_sample.loc[train_idx, 'orders']\n        X_valid = df_feature_sample.loc[valid_idx, FEATURES]\n        y_valid = df_feature_sample.loc[valid_idx, 'orders']\n        group_train=df_feature_sample.loc[train_idx].groupby('session').size()[df_feature_sample.loc[train_idx].session.drop_duplicates()]\n        group_vali=df_feature_sample.loc[valid_idx].groupby('session').size()[df_feature_sample.loc[valid_idx].session.drop_duplicates()]\n        ranker = lgb.LGBMRanker(\n                                objective=\"lambdarank\",\n                                metric=\"ndcg\",\n                                boosting_type=\"dart\",\n                                importance_type='gain',\n                                n_estimators=700,\n                                eval_at= 20,\n                                early_stopping_round=80,\n                                n_jobs = -1,\n                                random_state=42,\n                                reg_alpha = reg_alpha,\n                                reg_lambda = reg_lambda,\n                                num_leaves = num_leaves,\n                                min_child_samples = min_child_samples,\n                                colsample_bytree = colsample_bytree,\n                                learning_rate=learning_rate,\n                                )        \n        \n        ranker.fit(X_train,\n                   y_train, \n                   eval_set=[(X_train, y_train), (X_valid, y_valid)],\n                   group=group_train,\n                   eval_group=[group_train,group_vali])\n        score = gen_score(ranker)\n        scores += score\n        if score<0.66:\n            return 5 * score\n    return scores\n\n\nsampler = TPESampler(seed=42) \nstudy = optuna.create_study(\n        direction='maximize',\n        sampler=sampler,\n        study_name='otto')\nstudy.optimize(objective, n_trials=100)\n```\n![--](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3562400%2Fbc7e0335bcba2de98cda856220ad99ae%2F111111.jpg?generation=1675767526395608&alt=media)",
    "2133399": "First, I guess dropout used in DART mode induces some extent of randomness, make dart mode difficult to be 100% deterministic. Just my speculation.\nSecond, the native early stop API provided by lgbm should not be used together with DART mode. It won't work. If you want to use early stop in DART mode, you need to write your own early stop callback.",
    "2133578": "thanks, i think it's possible",
    "2135849": "Both LGB and CatBoost's results are not reproducible. However, after many training experiments, I found that the gap between the tuning stage and the training stage can be decreased by increasing the `early_stopping_round`, I raised this param to `300` for this competition. The gap is relatively small. \nIf you set a relatively small `early_stopping_round` e.g. `20` or `40`. Then most likely the early stopping is triggered before your model actually converges. Because of the randomness (even with a fixed seed), it could be triggered earlier during training than tuning. Then you'll see a big gap between training and tuning regarding the model performance."
  },
  "source": "meta"
}