{
  "id": 379605,
  "title": "Obvious Tips ",
  "url": "/competitions/otto-recommender-system/discussion/379605",
  "author_name": "",
  "post_date": "2023-01-20T09:05:30.714495200Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello community, I wanted to share something that, is clearly easy and obvious which can increase your local validation score. ( for my case, from 0.56 to 0.63 for orders )</p>\n<p>It may seems obvious, but when joining your candidates dataframe with the interactions features, you create duplicates rows. Then, if a right candidates is present multiple times, it will be duplicated in your predictions.</p>\n<p>In order to avoid this problem that took me around 2 days to find it, you can use the :</p>\n<ul>\n<li>polars : <code>df.unique(subset=['session','aid'], keep='last')</code>, In order to consider the last interaction. ( you can also average the prediction maybe ? )</li>\n<li>pandas : <code>df.drop_duplicates(subset=['session','aid'], keep='last')</code> </li>\n</ul>\n<p>I'm using polars :</p>\n<pre><code>candidates = pl.read_parquet(.(chunk))\n    scores = \n\n     model  tqdm(lgb_models):\n        scores += (model.predict(candidates[features_cols].to_pandas()))/(lgb_models)\n\n()\ntest = candidates.select([,])\ntest = test.with_columns(pl.Series(name=, values=scores))\ntest = test.sort(,reverse=).unique(subset=[,], keep=)\n\ntest_predictions = test.sort([, ], reverse=[,]).groupby([]).agg([\n        pl.col().limit().()\n    ])\n</code></pre>\n<p>Hope it helps.</p>",
  "messages": [
    {
      "id": "2108108",
      "postDate": "01/20/2023 09:05:30",
      "content": "<p>Hello community, I wanted to share something that, is clearly easy and obvious which can increase your local validation score. ( for my case, from 0.56 to 0.63 for orders )</p>\n<p>It may seems obvious, but when joining your candidates dataframe with the interactions features, you create duplicates rows. Then, if a right candidates is present multiple times, it will be duplicated in your predictions.</p>\n<p>In order to avoid this problem that took me around 2 days to find it, you can use the :</p>\n<ul>\n<li>polars : <code>df.unique(subset=['session','aid'], keep='last')</code>, In order to consider the last interaction. ( you can also average the prediction maybe ? )</li>\n<li>pandas : <code>df.drop_duplicates(subset=['session','aid'], keep='last')</code> </li>\n</ul>\n<p>I'm using polars :</p>\n<pre><code>candidates = pl.read_parquet(.(chunk))\n    scores = \n\n     model  tqdm(lgb_models):\n        scores += (model.predict(candidates[features_cols].to_pandas()))/(lgb_models)\n\n()\ntest = candidates.select([,])\ntest = test.with_columns(pl.Series(name=, values=scores))\ntest = test.sort(,reverse=).unique(subset=[,], keep=)\n\ntest_predictions = test.sort([, ], reverse=[,]).groupby([]).agg([\n        pl.col().limit().()\n    ])\n</code></pre>\n<p>Hope it helps.</p>",
      "rawMarkdown": "Hello community, I wanted to share something that, is clearly easy and obvious which can increase your local validation score. ( for my case, from 0.56 to 0.63 for orders )\n\nIt may seems obvious, but when joining your candidates dataframe with the interactions features, you create duplicates rows. Then, if a right candidates is present multiple times, it will be duplicated in your predictions.\n\nIn order to avoid this problem that took me around 2 days to find it, you can use the :\n- polars : `df.unique(subset=['session','aid'], keep='last')`, In order to consider the last interaction. ( you can also average the prediction maybe ? )\n- pandas : `df.drop_duplicates(subset=['session','aid'], keep='last')` \n\n\n\nI'm using polars :\n\n```python\ncandidates = pl.read_parquet(\"/kaggle/input/train-dataset-after-joining/train/train_chunk{}.parquet\".format(chunk))\n    scores = 0\n\n    for model in tqdm(lgb_models):\n        scores += (model.predict(candidates[features_cols].to_pandas()))/len(lgb_models)\n\nprint(\"done\")\ntest = candidates.select(['session','aid'])\ntest = test.with_columns(pl.Series(name='score', values=scores))\ntest = test.sort(\"score\",reverse=True).unique(subset=['session','aid'], keep='last')#.sort(\"session\")\n\ntest_predictions = test.sort(['session', 'score'], reverse=[False,True]).groupby(['session']).agg([\n        pl.col('aid').limit(20).list()\n    ])\n```\n\nHope it helps.",
      "votes": null
    },
    {
      "id": "2108241",
      "postDate": "01/20/2023 11:18:17",
      "content": "<p><a href=\"https://www.kaggle.com/rayanaay\" target=\"_blank\">@rayanaay</a> do you observe a difference b/w  <code>df.unique(subset=['session','aid'], keep='last')</code> and <code>df.unique(subset=['session','aid'])</code> in your overall local recall? </p>",
      "rawMarkdown": "rayanaay do you observe a difference b/w  `df.unique(subset=['session','aid'], keep='last')` and `df.unique(subset=['session','aid'])` in your overall local recall?",
      "votes": null
    },
    {
      "id": "2108247",
      "postDate": "01/20/2023 11:24:49",
      "content": "<p>Great question. Yes I did, still the \"last\" giving the best local validation compared to <code>keep=None</code>. </p>\n<p>In fact, I tried multiple strategies, and it seems that scoring all the interactions (with duplicates) then averaging the score of the ranker, give a little boost of + 0.0003</p>",
      "rawMarkdown": "Great question. Yes I did, still the \"last\" giving the best local validation compared to `keep=None`. \n\nIn fact, I tried multiple strategies, and it seems that scoring all the interactions (with duplicates) then averaging the score of the ranker, give a little boost of + 0.0003",
      "votes": null
    },
    {
      "id": "2108493",
      "postDate": "01/20/2023 15:13:45",
      "content": "<p><a href=\"https://www.kaggle.com/rayanaay\" target=\"_blank\">@rayanaay</a> Thanks for the response. Appreciate it :) </p>",
      "rawMarkdown": "rayanaay Thanks for the response. Appreciate it :)",
      "votes": null
    },
    {
      "id": "2115232",
      "postDate": "01/25/2023 15:51:47",
      "content": "<p>Great views, thanks for your precious sharing!</p>",
      "rawMarkdown": "Great views, thanks for your precious sharing!",
      "votes": null
    },
    {
      "id": "2115412",
      "postDate": "01/25/2023 18:40:01",
      "content": "<p>Glad to help ! :)</p>",
      "rawMarkdown": "Glad to help ! :)",
      "votes": null
    },
    {
      "id": "2123233",
      "postDate": "01/31/2023 11:53:36",
      "content": "<p>Appreciate your golden works share :) </p>",
      "rawMarkdown": "Appreciate your golden works share :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2108241,
      "author_name": "parthpankajtiwary",
      "author_url": "",
      "post_date": "01/20/2023 11:18:17",
      "content": "<p><a href=\"https://www.kaggle.com/rayanaay\" target=\"_blank\">@rayanaay</a> do you observe a difference b/w  <code>df.unique(subset=['session','aid'], keep='last')</code> and <code>df.unique(subset=['session','aid'])</code> in your overall local recall? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2108247,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/20/2023 11:24:49",
          "content": "<p>Great question. Yes I did, still the \"last\" giving the best local validation compared to <code>keep=None</code>. </p>\n<p>In fact, I tried multiple strategies, and it seems that scoring all the interactions (with duplicates) then averaging the score of the ranker, give a little boost of + 0.0003</p>",
          "votes": null,
          "replies": [
            {
              "id": 2108493,
              "author_name": "parthpankajtiwary",
              "author_url": "",
              "post_date": "01/20/2023 15:13:45",
              "content": "<p><a href=\"https://www.kaggle.com/rayanaay\" target=\"_blank\">@rayanaay</a> Thanks for the response. Appreciate it :) </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2115232,
      "author_name": "anrdaku",
      "author_url": "",
      "post_date": "01/25/2023 15:51:47",
      "content": "<p>Great views, thanks for your precious sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2115412,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/25/2023 18:40:01",
          "content": "<p>Glad to help ! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2123233,
      "author_name": "kingokind",
      "author_url": "",
      "post_date": "01/31/2023 11:53:36",
      "content": "<p>Appreciate your golden works share :) </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2108108": "Hello community, I wanted to share something that, is clearly easy and obvious which can increase your local validation score. ( for my case, from 0.56 to 0.63 for orders )\n\nIt may seems obvious, but when joining your candidates dataframe with the interactions features, you create duplicates rows. Then, if a right candidates is present multiple times, it will be duplicated in your predictions.\n\nIn order to avoid this problem that took me around 2 days to find it, you can use the :\n- polars : `df.unique(subset=['session','aid'], keep='last')`, In order to consider the last interaction. ( you can also average the prediction maybe ? )\n- pandas : `df.drop_duplicates(subset=['session','aid'], keep='last')` \n\n\n\nI'm using polars :\n\n```python\ncandidates = pl.read_parquet(\"/kaggle/input/train-dataset-after-joining/train/train_chunk{}.parquet\".format(chunk))\n    scores = 0\n\n    for model in tqdm(lgb_models):\n        scores += (model.predict(candidates[features_cols].to_pandas()))/len(lgb_models)\n\nprint(\"done\")\ntest = candidates.select(['session','aid'])\ntest = test.with_columns(pl.Series(name='score', values=scores))\ntest = test.sort(\"score\",reverse=True).unique(subset=['session','aid'], keep='last')#.sort(\"session\")\n\ntest_predictions = test.sort(['session', 'score'], reverse=[False,True]).groupby(['session']).agg([\n        pl.col('aid').limit(20).list()\n    ])\n```\n\nHope it helps.",
    "2108241": "rayanaay do you observe a difference b/w  `df.unique(subset=['session','aid'], keep='last')` and `df.unique(subset=['session','aid'])` in your overall local recall?",
    "2108247": "Great question. Yes I did, still the \"last\" giving the best local validation compared to `keep=None`. \n\nIn fact, I tried multiple strategies, and it seems that scoring all the interactions (with duplicates) then averaging the score of the ranker, give a little boost of + 0.0003",
    "2108493": "rayanaay Thanks for the response. Appreciate it :)",
    "2115232": "Great views, thanks for your precious sharing!",
    "2115412": "Glad to help ! :)",
    "2123233": "Appreciate your golden works share :)"
  },
  "source": "meta"
}