{
  "id": 599508,
  "title": "Features engineer",
  "url": "/competitions/aeroclub-recsys-2025/discussion/599508",
  "author_name": "",
  "post_date": "2025-08-17T00:44:28.762303500Z",
  "votes": 5,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Hello fellows,</p>\n<p>I'm really interested in what makes your LB scores have a huge improvement.</p>\n<p>For me, I the features that using companyID as a cluster could improve the score and local CV like 0.01-0.02.</p>\n<p>Other then that, I tried lots of different things like use searchRoute as cluster but didn't work out.</p>\n<p>I also tried the features in Json files. But I didn't get any useful features from it. Also, I noticed that the Json file has different data compared with the given data. </p>\n<p>BTW, just sharing some bad news from my side. My model training takes longer than I expect. Although I finally finiahed training my latest single model but the deadline is over like 4 mins 🥲 I wonder is it possible for me to select that one as my single model</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8174444%2Faee3c26802731b6296b48a8208352bfd%2F2025-08-17%202.35.44.png?generation=1755391356044590&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3270615",
      "postDate": "08/17/2025 00:44:28",
      "content": "<p>Hello fellows,</p>\n<p>I'm really interested in what makes your LB scores have a huge improvement.</p>\n<p>For me, I the features that using companyID as a cluster could improve the score and local CV like 0.01-0.02.</p>\n<p>Other then that, I tried lots of different things like use searchRoute as cluster but didn't work out.</p>\n<p>I also tried the features in Json files. But I didn't get any useful features from it. Also, I noticed that the Json file has different data compared with the given data. </p>\n<p>BTW, just sharing some bad news from my side. My model training takes longer than I expect. Although I finally finiahed training my latest single model but the deadline is over like 4 mins 🥲 I wonder is it possible for me to select that one as my single model</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8174444%2Faee3c26802731b6296b48a8208352bfd%2F2025-08-17%202.35.44.png?generation=1755391356044590&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hello fellows,\n\nI'm really interested in what makes your LB scores have a huge improvement.\n\nFor me, I the features that using companyID as a cluster could improve the score and local CV like 0.01-0.02.\n\nOther then that, I tried lots of different things like use searchRoute as cluster but didn't work out.\n\nI also tried the features in Json files. But I didn't get any useful features from it. Also, I noticed that the Json file has different data compared with the given data. \n\nBTW, just sharing some bad news from my side. My model training takes longer than I expect. Although I finally finiahed training my latest single model but the deadline is over like 4 mins 🥲 I wonder is it possible for me to select that one as my single model\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8174444%2Faee3c26802731b6296b48a8208352bfd%2F2025-08-17%202.35.44.png?generation=1755391356044590&alt=media)",
      "votes": null
    },
    {
      "id": "3270616",
      "postDate": "08/17/2025 00:45:51",
      "content": "<p>Using train/test for count stats, consider cluster stats of profileId, companyId and consider rank stats of ranker_id, (ranker_id, flight_hash) as public notebook postphone use.</p>",
      "rawMarkdown": "Using train/test for count stats, consider cluster stats of profileId, companyId and consider rank stats of ranker_id, (ranker_id, flight_hash) as public notebook postphone use.",
      "votes": null
    },
    {
      "id": "3270625",
      "postDate": "08/17/2025 01:25:38",
      "content": "<p>I also try to use count stats but it didn't work that will. It might because I only use the train data for it.<br>\nBut I'm not sure whether is using the train/test for making the features would be valided or against the rule (I know in most of the cases, it is considered invalid). Because in this way, your model has already seen the test data in the case.</p>",
      "rawMarkdown": "I also try to use count stats but it didn't work that will. It might because I only use the train data for it.\nBut I'm not sure whether is using the train/test for making the features would be valided or against the rule (I know in most of the cases, it is considered invalid). Because in this way, your model has already seen the test data in the case.",
      "votes": null
    },
    {
      "id": "3270626",
      "postDate": "08/17/2025 01:27:08",
      "content": "<p>I agree that use rank stats may be a good strategy in this case because lots of data is missing.</p>",
      "rawMarkdown": "I agree that use rank stats may be a good strategy in this case because lots of data is missing.",
      "votes": null
    },
    {
      "id": "3270633",
      "postDate": "08/17/2025 02:14:38",
      "content": "<p>there is a thread on this </p>",
      "rawMarkdown": "there is a thread on this",
      "votes": null
    },
    {
      "id": "3270669",
      "postDate": "08/17/2025 05:28:12",
      "content": "<p>Yes, using companyID and profileID as clusters.<br>\nFor profileID, just computing stats directly already gave me an improvement.<br>\nFor companyID (alone or in combination with profileID), the score improved only if the statistics for each ranker_id group were calculated using only those groups that occurred earlier in time than the group being evaluated.</p>\n<p>That is, I only added as features the statistics based on past actions of this client or this company.</p>\n<p>Additionally, I got some score improvement along with reduced data size and faster runtime by randomly truncating groups to a maximum of 50 options.</p>",
      "rawMarkdown": "Yes, using companyID and profileID as clusters.\nFor profileID, just computing stats directly already gave me an improvement.\nFor companyID (alone or in combination with profileID), the score improved only if the statistics for each ranker_id group were calculated using only those groups that occurred earlier in time than the group being evaluated.\n\nThat is, I only added as features the statistics based on past actions of this client or this company.\n\nAdditionally, I got some score improvement along with reduced data size and faster runtime by randomly truncating groups to a maximum of 50 options.",
      "votes": null
    },
    {
      "id": "3270674",
      "postDate": "08/17/2025 05:43:08",
      "content": "<p>It is not valid for real world data for you could not see future data, but for this game we have test dataset it's ok to use it, or even pseudo label though I think only use train data still could get good result.</p>",
      "rawMarkdown": "It is not valid for real world data for you could not see future data, but for this game we have test dataset it's ok to use it, or even pseudo label though I think only use train data still could get good result.",
      "votes": null
    },
    {
      "id": "3270688",
      "postDate": "08/17/2025 06:32:08",
      "content": "<p>Agree! Using only the train data for count-based and other group-based stats, my single XGB model can achieve 0.535+ on the private LB. By the way, how did you use the training set, and did you try any pseudo-labeling?</p>",
      "rawMarkdown": "Agree! Using only the train data for count-based and other group-based stats, my single XGB model can achieve 0.535+ on the private LB. By the way, how did you use the training set, and did you try any pseudo-labeling?",
      "votes": null
    },
    {
      "id": "3270699",
      "postDate": "08/17/2025 07:03:54",
      "content": "<p>I tried PL but most of the 'high confidence' samples were from low group size, i wasn't sure if i wanted to put in more samples of less group sizes, but it was working out fine at that time. </p>",
      "rawMarkdown": "I tried PL but most of the 'high confidence' samples were from low group size, i wasn't sure if i wanted to put in more samples of less group sizes, but it was working out fine at that time.",
      "votes": null
    },
    {
      "id": "3270701",
      "postDate": "08/17/2025 07:05:26",
      "content": "<p>For me a bunch of groupby stats based on companyID, profileId, booking route gave good boost. </p>",
      "rawMarkdown": "For me a bunch of groupby stats based on companyID, profileId, booking route gave good boost.",
      "votes": null
    },
    {
      "id": "3270705",
      "postDate": "08/17/2025 07:09:51",
      "content": "<p>I added some groupby features for profileId as well, but it didn’t help much. I’m wondering what you did differently?</p>",
      "rawMarkdown": "I added some groupby features for profileId as well, but it didn’t help much. I’m wondering what you did differently?",
      "votes": null
    },
    {
      "id": "3270708",
      "postDate": "08/17/2025 07:12:22",
      "content": "<p>I calculated these stats on a few engineered features, not the raw ones. I will do a detailed writeup today or tomorrow. </p>",
      "rawMarkdown": "I calculated these stats on a few engineered features, not the raw ones. I will do a detailed writeup today or tomorrow.",
      "votes": null
    },
    {
      "id": "3270722",
      "postDate": "08/17/2025 07:28:27",
      "content": "<p>pseudo could improve my single model local valid score but I did not submit them and did not use them as they seems could not help ensemble and pseudo model seems not match single model qulification. But I might late submit 1 pseudo model as interestingly I find my currrent best single model work best on LB also tough not PB, it's interesting to see pseudo single model performance online.</p>",
      "rawMarkdown": "pseudo could improve my single model local valid score but I did not submit them and did not use them as they seems could not help ensemble and pseudo model seems not match single model qulification. But I might late submit 1 pseudo model as interestingly I find my currrent best single model work best on LB also tough not PB, it's interesting to see pseudo single model performance online.",
      "votes": null
    },
    {
      "id": "3270806",
      "postDate": "08/17/2025 11:25:27",
      "content": "<p>I start with using train + test for counting from very early stage, so without adjustment I just late submit single model which only counting using train PB drop from 54156 to 52692.</p>",
      "rawMarkdown": "I start with using train + test for counting from very early stage, so without adjustment I just late submit single model which only counting using train PB drop from 54156 to 52692.",
      "votes": null
    },
    {
      "id": "3270837",
      "postDate": "08/17/2025 12:57:08",
      "content": "<p>Yeah, I tried to use train label stats, but local valid drop a lot seems for training each instance see future labels so it casue model overfit to these features. I do not use train label stats but your method would help, it is similar as pre-n-day-span stats ctr feature for recommendation system, only here the test is not prove in streaming mode so test has some feature gap but I would like try to add your feature to see the difference, thanks for sharing. </p>",
      "rawMarkdown": "Yeah, I tried to use train label stats, but local valid drop a lot seems for training each instance see future labels so it casue model overfit to these features. I do not use train label stats but your method would help, it is similar as pre-n-day-span stats ctr feature for recommendation system, only here the test is not prove in streaming mode so test has some feature gap but I would like try to add your feature to see the difference, thanks for sharing.",
      "votes": null
    },
    {
      "id": "3270867",
      "postDate": "08/17/2025 14:03:25",
      "content": "<p>Hi, could you please give an example of the count stats feature using train + test, and explain the meaning of test data leakage in this trick? I'm a bit confused. Since the test split doesn't have label \"selected = 0/1\", I assume the count stats using train + test must  be features that don't depend on the label.</p>\n<p>For example, \"marketingCarrier_code occurrence count\", how many times per marketingCarrier shows up, counting both selected = 1 and selected = 0, show how \"popular\" a marketingCarrier is.</p>\n<p>And the value would be the same for all occurrences in train + test, not time series , is this right? Like the code below:</p>\n<pre><code>df = pl.concat([train, test])\nmarketingCarrier_code_count = df.group_by().agg(pl.())\ndf = df.join(marketingCarrier_code_count, on=)\n</code></pre>\n<p>It means using the feature (the count computed on train + test together) both when training and when predicting. </p>",
      "rawMarkdown": "Hi, could you please give an example of the count stats feature using train + test, and explain the meaning of test data leakage in this trick? I'm a bit confused. Since the test split doesn't have label \"selected = 0/1\", I assume the count stats using train + test must  be features that don't depend on the label.\n\nFor example, \"marketingCarrier_code occurrence count\", how many times per marketingCarrier shows up, counting both selected = 1 and selected = 0, show how \"popular\" a marketingCarrier is.\n\nAnd the value would be the same for all occurrences in train + test, not time series , is this right? Like the code below:\n```python\ndf = pl.concat([train, test])\nmarketingCarrier_code_count = df.group_by(\"marketingCarrier_code\").agg(pl.len())\ndf = df.join(marketingCarrier_code_count, on=\"marketingCarrier_code\")\n```\n\nIt means using the feature (the count computed on train + test together) both when training and when predicting.",
      "votes": null
    },
    {
      "id": "3270921",
      "postDate": "08/17/2025 16:00:52",
      "content": "<p>I only count for show I did not used any labeled count yet</p>",
      "rawMarkdown": "I only count for show I did not used any labeled count yet",
      "votes": null
    },
    {
      "id": "3270977",
      "postDate": "08/17/2025 18:24:18",
      "content": "<p>Thank you for the companyID</p>",
      "rawMarkdown": "Thank you for the companyID",
      "votes": null
    },
    {
      "id": "3271095",
      "postDate": "08/18/2025 04:13:06",
      "content": "<p>Thank U! Great Feature Engineering!</p>",
      "rawMarkdown": "Thank U! Great Feature Engineering!",
      "votes": null
    },
    {
      "id": "3271289",
      "postDate": "08/18/2025 13:26:10",
      "content": "<p>Well adding your feature could boost online single model a lot PB from 54156 up to 54942, cool!</p>",
      "rawMarkdown": "Well adding your feature could boost online single model a lot PB from 54156 up to 54942, cool!",
      "votes": null
    },
    {
      "id": "3271304",
      "postDate": "08/18/2025 14:09:56",
      "content": "<p>Great! Thank you for checking it!<br>\nI really admired your approach to the task.</p>",
      "rawMarkdown": "Great! Thank you for checking it!\nI really admired your approach to the task.",
      "votes": null
    },
    {
      "id": "3271411",
      "postDate": "08/18/2025 20:04:36",
      "content": "<p>Thank you for sharing !</p>",
      "rawMarkdown": "Thank you for sharing !",
      "votes": null
    },
    {
      "id": "3271674",
      "postDate": "08/19/2025 11:48:25",
      "content": "<p>thanks for sharing your achievement dude.</p>",
      "rawMarkdown": "thanks for sharing your achievement dude.",
      "votes": null
    },
    {
      "id": "3271709",
      "postDate": "08/19/2025 13:34:29",
      "content": "<p>Thanks for sharing ur achievement…</p>",
      "rawMarkdown": "Thanks for sharing ur achievement...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3270616,
      "author_name": "goldenlock",
      "author_url": "",
      "post_date": "08/17/2025 00:45:51",
      "content": "<p>Using train/test for count stats, consider cluster stats of profileId, companyId and consider rank stats of ranker_id, (ranker_id, flight_hash) as public notebook postphone use.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3270625,
          "author_name": "dingyangwang",
          "author_url": "",
          "post_date": "08/17/2025 01:25:38",
          "content": "<p>I also try to use count stats but it didn't work that will. It might because I only use the train data for it.<br>\nBut I'm not sure whether is using the train/test for making the features would be valided or against the rule (I know in most of the cases, it is considered invalid). Because in this way, your model has already seen the test data in the case.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3270674,
              "author_name": "goldenlock",
              "author_url": "",
              "post_date": "08/17/2025 05:43:08",
              "content": "<p>It is not valid for real world data for you could not see future data, but for this game we have test dataset it's ok to use it, or even pseudo label though I think only use train data still could get good result.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3270688,
                  "author_name": "mango789",
                  "author_url": "",
                  "post_date": "08/17/2025 06:32:08",
                  "content": "<p>Agree! Using only the train data for count-based and other group-based stats, my single XGB model can achieve 0.535+ on the private LB. By the way, how did you use the training set, and did you try any pseudo-labeling?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3270699,
                      "author_name": "pheadrus",
                      "author_url": "",
                      "post_date": "08/17/2025 07:03:54",
                      "content": "<p>I tried PL but most of the 'high confidence' samples were from low group size, i wasn't sure if i wanted to put in more samples of less group sizes, but it was working out fine at that time. </p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 3270722,
                      "author_name": "goldenlock",
                      "author_url": "",
                      "post_date": "08/17/2025 07:28:27",
                      "content": "<p>pseudo could improve my single model local valid score but I did not submit them and did not use them as they seems could not help ensemble and pseudo model seems not match single model qulification. But I might late submit 1 pseudo model as interestingly I find my currrent best single model work best on LB also tough not PB, it's interesting to see pseudo single model performance online.</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 3270806,
                      "author_name": "goldenlock",
                      "author_url": "",
                      "post_date": "08/17/2025 11:25:27",
                      "content": "<p>I start with using train + test for counting from very early stage, so without adjustment I just late submit single model which only counting using train PB drop from 54156 to 52692.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3270867,
                          "author_name": "behrenskong",
                          "author_url": "",
                          "post_date": "08/17/2025 14:03:25",
                          "content": "<p>Hi, could you please give an example of the count stats feature using train + test, and explain the meaning of test data leakage in this trick? I'm a bit confused. Since the test split doesn't have label \"selected = 0/1\", I assume the count stats using train + test must  be features that don't depend on the label.</p>\n<p>For example, \"marketingCarrier_code occurrence count\", how many times per marketingCarrier shows up, counting both selected = 1 and selected = 0, show how \"popular\" a marketingCarrier is.</p>\n<p>And the value would be the same for all occurrences in train + test, not time series , is this right? Like the code below:</p>\n<pre><code>df = pl.concat([train, test])\nmarketingCarrier_code_count = df.group_by().agg(pl.())\ndf = df.join(marketingCarrier_code_count, on=)\n</code></pre>\n<p>It means using the feature (the count computed on train + test together) both when training and when predicting. </p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3270921,
                              "author_name": "goldenlock",
                              "author_url": "",
                              "post_date": "08/17/2025 16:00:52",
                              "content": "<p>I only count for show I did not used any labeled count yet</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3271095,
                                  "author_name": "behrenskong",
                                  "author_url": "",
                                  "post_date": "08/18/2025 04:13:06",
                                  "content": "<p>Thank U! Great Feature Engineering!</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3270626,
          "author_name": "dingyangwang",
          "author_url": "",
          "post_date": "08/17/2025 01:27:08",
          "content": "<p>I agree that use rank stats may be a good strategy in this case because lots of data is missing.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3270633,
              "author_name": "goldenlock",
              "author_url": "",
              "post_date": "08/17/2025 02:14:38",
              "content": "<p>there is a thread on this </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3270669,
      "author_name": "mikhailgolubchik",
      "author_url": "",
      "post_date": "08/17/2025 05:28:12",
      "content": "<p>Yes, using companyID and profileID as clusters.<br>\nFor profileID, just computing stats directly already gave me an improvement.<br>\nFor companyID (alone or in combination with profileID), the score improved only if the statistics for each ranker_id group were calculated using only those groups that occurred earlier in time than the group being evaluated.</p>\n<p>That is, I only added as features the statistics based on past actions of this client or this company.</p>\n<p>Additionally, I got some score improvement along with reduced data size and faster runtime by randomly truncating groups to a maximum of 50 options.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3270837,
          "author_name": "goldenlock",
          "author_url": "",
          "post_date": "08/17/2025 12:57:08",
          "content": "<p>Yeah, I tried to use train label stats, but local valid drop a lot seems for training each instance see future labels so it casue model overfit to these features. I do not use train label stats but your method would help, it is similar as pre-n-day-span stats ctr feature for recommendation system, only here the test is not prove in streaming mode so test has some feature gap but I would like try to add your feature to see the difference, thanks for sharing. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3271289,
          "author_name": "goldenlock",
          "author_url": "",
          "post_date": "08/18/2025 13:26:10",
          "content": "<p>Well adding your feature could boost online single model a lot PB from 54156 up to 54942, cool!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3271304,
              "author_name": "mikhailgolubchik",
              "author_url": "",
              "post_date": "08/18/2025 14:09:56",
              "content": "<p>Great! Thank you for checking it!<br>\nI really admired your approach to the task.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3270701,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "08/17/2025 07:05:26",
      "content": "<p>For me a bunch of groupby stats based on companyID, profileId, booking route gave good boost. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3270705,
          "author_name": "mango789",
          "author_url": "",
          "post_date": "08/17/2025 07:09:51",
          "content": "<p>I added some groupby features for profileId as well, but it didn’t help much. I’m wondering what you did differently?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3270708,
              "author_name": "pheadrus",
              "author_url": "",
              "post_date": "08/17/2025 07:12:22",
              "content": "<p>I calculated these stats on a few engineered features, not the raw ones. I will do a detailed writeup today or tomorrow. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3270977,
      "author_name": "navneetbende",
      "author_url": "",
      "post_date": "08/17/2025 18:24:18",
      "content": "<p>Thank you for the companyID</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3271411,
      "author_name": "abdelhakouanzougui2",
      "author_url": "",
      "post_date": "08/18/2025 20:04:36",
      "content": "<p>Thank you for sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3271674,
      "author_name": "muhammadtalha82",
      "author_url": "",
      "post_date": "08/19/2025 11:48:25",
      "content": "<p>thanks for sharing your achievement dude.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3271709,
      "author_name": "saraharshadbcs",
      "author_url": "",
      "post_date": "08/19/2025 13:34:29",
      "content": "<p>Thanks for sharing ur achievement…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3270615": "Hello fellows,\n\nI'm really interested in what makes your LB scores have a huge improvement.\n\nFor me, I the features that using companyID as a cluster could improve the score and local CV like 0.01-0.02.\n\nOther then that, I tried lots of different things like use searchRoute as cluster but didn't work out.\n\nI also tried the features in Json files. But I didn't get any useful features from it. Also, I noticed that the Json file has different data compared with the given data. \n\nBTW, just sharing some bad news from my side. My model training takes longer than I expect. Although I finally finiahed training my latest single model but the deadline is over like 4 mins 🥲 I wonder is it possible for me to select that one as my single model\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8174444%2Faee3c26802731b6296b48a8208352bfd%2F2025-08-17%202.35.44.png?generation=1755391356044590&alt=media)",
    "3270616": "Using train/test for count stats, consider cluster stats of profileId, companyId and consider rank stats of ranker_id, (ranker_id, flight_hash) as public notebook postphone use.",
    "3270625": "I also try to use count stats but it didn't work that will. It might because I only use the train data for it.\nBut I'm not sure whether is using the train/test for making the features would be valided or against the rule (I know in most of the cases, it is considered invalid). Because in this way, your model has already seen the test data in the case.",
    "3270626": "I agree that use rank stats may be a good strategy in this case because lots of data is missing.",
    "3270633": "there is a thread on this",
    "3270669": "Yes, using companyID and profileID as clusters.\nFor profileID, just computing stats directly already gave me an improvement.\nFor companyID (alone or in combination with profileID), the score improved only if the statistics for each ranker_id group were calculated using only those groups that occurred earlier in time than the group being evaluated.\n\nThat is, I only added as features the statistics based on past actions of this client or this company.\n\nAdditionally, I got some score improvement along with reduced data size and faster runtime by randomly truncating groups to a maximum of 50 options.",
    "3270674": "It is not valid for real world data for you could not see future data, but for this game we have test dataset it's ok to use it, or even pseudo label though I think only use train data still could get good result.",
    "3270688": "Agree! Using only the train data for count-based and other group-based stats, my single XGB model can achieve 0.535+ on the private LB. By the way, how did you use the training set, and did you try any pseudo-labeling?",
    "3270699": "I tried PL but most of the 'high confidence' samples were from low group size, i wasn't sure if i wanted to put in more samples of less group sizes, but it was working out fine at that time.",
    "3270701": "For me a bunch of groupby stats based on companyID, profileId, booking route gave good boost.",
    "3270705": "I added some groupby features for profileId as well, but it didn’t help much. I’m wondering what you did differently?",
    "3270708": "I calculated these stats on a few engineered features, not the raw ones. I will do a detailed writeup today or tomorrow.",
    "3270722": "pseudo could improve my single model local valid score but I did not submit them and did not use them as they seems could not help ensemble and pseudo model seems not match single model qulification. But I might late submit 1 pseudo model as interestingly I find my currrent best single model work best on LB also tough not PB, it's interesting to see pseudo single model performance online.",
    "3270806": "I start with using train + test for counting from very early stage, so without adjustment I just late submit single model which only counting using train PB drop from 54156 to 52692.",
    "3270837": "Yeah, I tried to use train label stats, but local valid drop a lot seems for training each instance see future labels so it casue model overfit to these features. I do not use train label stats but your method would help, it is similar as pre-n-day-span stats ctr feature for recommendation system, only here the test is not prove in streaming mode so test has some feature gap but I would like try to add your feature to see the difference, thanks for sharing.",
    "3270867": "Hi, could you please give an example of the count stats feature using train + test, and explain the meaning of test data leakage in this trick? I'm a bit confused. Since the test split doesn't have label \"selected = 0/1\", I assume the count stats using train + test must  be features that don't depend on the label.\n\nFor example, \"marketingCarrier_code occurrence count\", how many times per marketingCarrier shows up, counting both selected = 1 and selected = 0, show how \"popular\" a marketingCarrier is.\n\nAnd the value would be the same for all occurrences in train + test, not time series , is this right? Like the code below:\n```python\ndf = pl.concat([train, test])\nmarketingCarrier_code_count = df.group_by(\"marketingCarrier_code\").agg(pl.len())\ndf = df.join(marketingCarrier_code_count, on=\"marketingCarrier_code\")\n```\n\nIt means using the feature (the count computed on train + test together) both when training and when predicting.",
    "3270921": "I only count for show I did not used any labeled count yet",
    "3270977": "Thank you for the companyID",
    "3271095": "Thank U! Great Feature Engineering!",
    "3271289": "Well adding your feature could boost online single model a lot PB from 54156 up to 54942, cool!",
    "3271304": "Great! Thank you for checking it!\nI really admired your approach to the task.",
    "3271411": "Thank you for sharing !",
    "3271674": "thanks for sharing your achievement dude.",
    "3271709": "Thanks for sharing ur achievement..."
  },
  "source": "meta"
}