{
  "id": 347786,
  "title": "11th Place Solution (LightGBM with meta features)",
  "url": "/competitions/amex-default-prediction/discussion/347786",
  "author_name": "Shaorooon",
  "post_date": "2022-08-25T11:10:08.833000",
  "votes": 103,
  "comment_count": 44,
  "views": 0,
  "content": "<p>※Previously titled 12th Place Solution. It is now 11th due to the fixed ranking.</p>\n<p>Thank you to everyone who participated in the competition and to everyone involved in organizing it.<br>\nI learned a lot through this competition.</p>\n<h2>Score &amp; Result</h2>\n<ul>\n<li>My best submission<ul>\n<li>Local CV：0.79922</li>\n<li>Public: 0.80088</li>\n<li>Private: 0.80852</li></ul></li>\n<li>Results<ul>\n<li>Public:6th → Private 12th</li></ul></li>\n</ul>\n<h2>Feature Engineering</h2>\n<ul>\n<li>Base features are from public notebook<ul>\n<li><a href=\"https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds\" target=\"_blank\">https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds</a></li></ul></li>\n<li>Delete some features<ul>\n<li>Round up features</li>\n<li>Duplicate features (ex. groupby counts for all category features)</li></ul></li>\n<li>Add some features<ul>\n<li>Features aggregated by time period <ul>\n<li>min, max, mean, std for the last 3 and last 6 months</li></ul></li>\n<li>Rate and diff features in time series<ul>\n<li>ex. last - latest_3month_mean, last_3month_mean / last_6month_mean</li></ul></li>\n<li>Null count features</li>\n<li>Date features</li>\n<li><strong>Meta features (most important features!)</strong><ul>\n<li>how to make<ol>\n<li>Train_labels are assigned to train data (before aggregation by cid) and train model.</li>\n<li>Make oof prediction for train data.</li>\n<li>Aggregate oof prediction by time period.</li></ol></li>\n<li>Using this feature, I reached 0.800 PublicLB from 0.799 PublicLB in single model.</li>\n<li>Referring to the DSB2019's 2nd place solution method.<ul>\n<li><a href=\"https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388\" target=\"_blank\">https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388</a></li></ul></li></ul></li></ul></li>\n</ul>\n<h2>Validation strategy</h2>\n<ul>\n<li>Use stratfiedKfold(k=5).</li>\n<li>I think Public LB is more important than local CV to measure Private's performance.<ul>\n<li>Data size is about the same for train and public.</li>\n<li>In terms of time, public data is closer to private data than train data.</li>\n<li>Even after adversarial validation, train/private was farther away from the data than public/private.<ul>\n<li>train/private：AUC 0.99</li>\n<li>public/private：AUC 0.82</li></ul></li></ul></li>\n<li>While focusing on publicLB, we also looked at local CV to determine if there was any improvement.<ul>\n<li>Also checked local logloss because amex_metric was not stable.</li>\n<li>It was hard to find a few digits of publicLB.</li></ul></li>\n</ul>\n<h2>Model</h2>\n<ul>\n<li>LightGBM<ul>\n<li>Use dart.</li>\n<li>Hypyer_parameter is the same as base notebook.<ul>\n<li><a href=\"https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds\" target=\"_blank\">https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds</a></li></ul></li>\n<li>Get best amex metric model (use callback)<ul>\n<li><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332575\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/332575</a>    </li></ul></li></ul></li>\n</ul>\n<h2>Feature Selection</h2>\n<ul>\n<li>Adversarial validation<ul>\n<li>Delete features of high importance in train/private adversarial validation.<ul>\n<li>Drop_features: R_1, D59, S_11, B_29</li>\n<li>After the change, the AUC was 0.8.</li></ul></li></ul></li>\n<li>Null importance<ul>\n<li>Use features actual importance larger than mean importance with shuffled target.</li>\n<li>Before:4300 features → After:1300 features</li></ul></li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>Use 3 LightGBM models and rank ensemble (weighted average).</li>\n<li>Each model use different feature set.<ul>\n<li>Model_1: Not use meta features.</li>\n<li>Model_2: Use meta features &amp; large features (not use null importance feature selection).</li>\n<li>Model_3: Use meta features &amp; small features (use null importance feature selection).</li></ul></li>\n<li>Ensemble weight<ul>\n<li>model_1:model_2:model_3 = 4:4:2</li>\n<li>weight decided while looking at public LB</li></ul></li>\n</ul>\n<h2>Select Submission</h2>\n<ul>\n<li>I chose two sub's: <ol>\n<li>BestLB sub</li>\n<li>Sub with risk of time-series changes in features.<ul>\n<li>Features that were not important in adversarial validation are not used. </li></ul></li></ol></li>\n<li>The second model was the best in privateLB.</li>\n<li>Perhaps the trend of some features changed over time. adversarial validation was very helpful.</li>\n</ul>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": 1913556,
      "postDate": "2022-08-25T11:10:08.833Z",
      "content": "<p>※Previously titled 12th Place Solution. It is now 11th due to the fixed ranking.</p>\n<p>Thank you to everyone who participated in the competition and to everyone involved in organizing it.<br>\nI learned a lot through this competition.</p>\n<h2>Score &amp; Result</h2>\n<ul>\n<li>My best submission<ul>\n<li>Local CV：0.79922</li>\n<li>Public: 0.80088</li>\n<li>Private: 0.80852</li></ul></li>\n<li>Results<ul>\n<li>Public:6th → Private 12th</li></ul></li>\n</ul>\n<h2>Feature Engineering</h2>\n<ul>\n<li>Base features are from public notebook<ul>\n<li><a href=\"https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds\" target=\"_blank\">https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds</a></li></ul></li>\n<li>Delete some features<ul>\n<li>Round up features</li>\n<li>Duplicate features (ex. groupby counts for all category features)</li></ul></li>\n<li>Add some features<ul>\n<li>Features aggregated by time period <ul>\n<li>min, max, mean, std for the last 3 and last 6 months</li></ul></li>\n<li>Rate and diff features in time series<ul>\n<li>ex. last - latest_3month_mean, last_3month_mean / last_6month_mean</li></ul></li>\n<li>Null count features</li>\n<li>Date features</li>\n<li><strong>Meta features (most important features!)</strong><ul>\n<li>how to make<ol>\n<li>Train_labels are assigned to train data (before aggregation by cid) and train model.</li>\n<li>Make oof prediction for train data.</li>\n<li>Aggregate oof prediction by time period.</li></ol></li>\n<li>Using this feature, I reached 0.800 PublicLB from 0.799 PublicLB in single model.</li>\n<li>Referring to the DSB2019's 2nd place solution method.<ul>\n<li><a href=\"https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388\" target=\"_blank\">https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388</a></li></ul></li></ul></li></ul></li>\n</ul>\n<h2>Validation strategy</h2>\n<ul>\n<li>Use stratfiedKfold(k=5).</li>\n<li>I think Public LB is more important than local CV to measure Private's performance.<ul>\n<li>Data size is about the same for train and public.</li>\n<li>In terms of time, public data is closer to private data than train data.</li>\n<li>Even after adversarial validation, train/private was farther away from the data than public/private.<ul>\n<li>train/private：AUC 0.99</li>\n<li>public/private：AUC 0.82</li></ul></li></ul></li>\n<li>While focusing on publicLB, we also looked at local CV to determine if there was any improvement.<ul>\n<li>Also checked local logloss because amex_metric was not stable.</li>\n<li>It was hard to find a few digits of publicLB.</li></ul></li>\n</ul>\n<h2>Model</h2>\n<ul>\n<li>LightGBM<ul>\n<li>Use dart.</li>\n<li>Hypyer_parameter is the same as base notebook.<ul>\n<li><a href=\"https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds\" target=\"_blank\">https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds</a></li></ul></li>\n<li>Get best amex metric model (use callback)<ul>\n<li><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332575\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/332575</a>    </li></ul></li></ul></li>\n</ul>\n<h2>Feature Selection</h2>\n<ul>\n<li>Adversarial validation<ul>\n<li>Delete features of high importance in train/private adversarial validation.<ul>\n<li>Drop_features: R_1, D59, S_11, B_29</li>\n<li>After the change, the AUC was 0.8.</li></ul></li></ul></li>\n<li>Null importance<ul>\n<li>Use features actual importance larger than mean importance with shuffled target.</li>\n<li>Before:4300 features → After:1300 features</li></ul></li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>Use 3 LightGBM models and rank ensemble (weighted average).</li>\n<li>Each model use different feature set.<ul>\n<li>Model_1: Not use meta features.</li>\n<li>Model_2: Use meta features &amp; large features (not use null importance feature selection).</li>\n<li>Model_3: Use meta features &amp; small features (use null importance feature selection).</li></ul></li>\n<li>Ensemble weight<ul>\n<li>model_1:model_2:model_3 = 4:4:2</li>\n<li>weight decided while looking at public LB</li></ul></li>\n</ul>\n<h2>Select Submission</h2>\n<ul>\n<li>I chose two sub's: <ol>\n<li>BestLB sub</li>\n<li>Sub with risk of time-series changes in features.<ul>\n<li>Features that were not important in adversarial validation are not used. </li></ul></li></ol></li>\n<li>The second model was the best in privateLB.</li>\n<li>Perhaps the trend of some features changed over time. adversarial validation was very helpful.</li>\n</ul>\n<p>Thank you.</p>",
      "rawMarkdown": "※Previously titled 12th Place Solution. It is now 11th due to the fixed ranking.\n\nThank you to everyone who participated in the competition and to everyone involved in organizing it.\nI learned a lot through this competition.\n\n## Score & Result\n- My best submission\n    - Local CV：0.79922\n    - Public: 0.80088\n    - Private: 0.80852\n- Results\n    - Public:6th → Private 12th\n\n## Feature Engineering\n- Base features are from public notebook\n    - https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds\n- Delete some features\n    - Round up features\n    - Duplicate features (ex. groupby counts for all category features)\n- Add some features\n    - Features aggregated by time period \n        - min, max, mean, std for the last 3 and last 6 months\n    - Rate and diff features in time series\n        - ex. last - latest_3month_mean, last_3month_mean / last_6month_mean\n    - Null count features\n    - Date features\n    - **Meta features (most important features!)**\n        - how to make\n            1. Train_labels are assigned to train data (before aggregation by cid) and train model.\n            2. Make oof prediction for train data.\n            3. Aggregate oof prediction by time period.\n        - Using this feature, I reached 0.800 PublicLB from 0.799 PublicLB in single model.\n        - Referring to the DSB2019's 2nd place solution method.\n            - https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388\n\n## Validation strategy\n- Use stratfiedKfold(k=5).\n- I think Public LB is more important than local CV to measure Private's performance.\n    - Data size is about the same for train and public.\n    - In terms of time, public data is closer to private data than train data.\n    - Even after adversarial validation, train/private was farther away from the data than public/private.\n        - train/private：AUC 0.99\n        - public/private：AUC 0.82\n- While focusing on publicLB, we also looked at local CV to determine if there was any improvement.\n    - Also checked local logloss because amex_metric was not stable.\n    - It was hard to find a few digits of publicLB.\n\n## Model\n- LightGBM\n    - Use dart.\n    - Hypyer_parameter is the same as base notebook.\n        - https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds\n    - Get best amex metric model (use callback)\n        - https://www.kaggle.com/competitions/amex-default-prediction/discussion/332575    \n\n## Feature Selection\n- Adversarial validation\n    - Delete features of high importance in train/private adversarial validation.\n        - Drop_features: R_1, D59, S_11, B_29\n        - After the change, the AUC was 0.8.\n- Null importance\n    - Use features actual importance larger than mean importance with shuffled target.\n    - Before:4300 features → After:1300 features\n\n## Ensemble\n- Use 3 LightGBM models and rank ensemble (weighted average).\n- Each model use different feature set.\n    - Model_1: Not use meta features.\n    - Model_2: Use meta features & large features (not use null importance feature selection).\n    - Model_3: Use meta features & small features (use null importance feature selection).\n- Ensemble weight\n    - model_1:model_2:model_3 = 4:4:2\n    - weight decided while looking at public LB\n\n## Select Submission\n- I chose two sub's: \n    1. BestLB sub\n    2. Sub with risk of time-series changes in features.\n        - Features that were not important in adversarial validation are not used. \n- The second model was the best in privateLB.\n- Perhaps the trend of some features changed over time. adversarial validation was very helpful.\n\n\nThank you.",
      "votes": 103
    },
    {
      "id": 1913628,
      "postDate": "2022-08-25T12:20:54.807Z",
      "content": "<p>Congrats! <br>\nI almost did the same thing with what you did about meta features, though I didn't know it was called meta feature before. <br>\nThe only difference is that I used the oof prediction as 13 numeric features directly and you did some aggregation. <br>\nThese features also boosted my model a lot！</p>",
      "rawMarkdown": "Congrats! \nI almost did the same thing with what you did about meta features, though I didn't know it was called meta feature before. \nThe only difference is that I used the oof prediction as 13 numeric features directly and you did some aggregation. \nThese features also boosted my model a lot！",
      "votes": 3,
      "replies": [
        {
          "id": 1913820,
          "postDate": "2022-08-25T14:24:04.310Z",
          "content": "<p>Thank you for comment !</p>\n<blockquote>\n  <p>The only difference is that I used the oof prediction as 13 numeric features directly and you did some aggregation.</p>\n</blockquote>\n<p>It did not occur to me to use the 13 numeric features directly. That seems to work too.</p>",
          "rawMarkdown": "Thank you for comment !\n> The only difference is that I used the oof prediction as 13 numeric features directly and you did some aggregation.\n\nIt did not occur to me to use the 13 numeric features directly. That seems to work too.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1948795,
      "postDate": "2022-09-21T09:43:16.187Z",
      "content": "<p><a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a> very neat writeup and excellent solution. I learned a lot from it. </p>\n<p>I am very interested in your <code>adversarial validation</code> process, would you kindly elaborate a bit? </p>\n<p>thanks!</p>\n<p>xxxxyyyy80008</p>",
      "rawMarkdown": "@syaorn13 very neat writeup and excellent solution. I learned a lot from it. \n\nI am very interested in your `adversarial validation` process, would you kindly elaborate a bit? \n\nthanks!\n\nxxxxyyyy80008",
      "votes": 1,
      "replies": [
        {
          "id": 1950135,
          "postDate": "2022-09-22T06:42:37.367Z",
          "content": "<p><a href=\"https://www.kaggle.com/xxxxyyyy80008\" target=\"_blank\">@xxxxyyyy80008</a> <br>\nThe adversarial validation was done to see the difference between train and private distributions and for feature selection.<br>\nThe code is similar to the following notebook. (although the notebook is private and public)<br>\n<a href=\"https://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public\" target=\"_blank\">https://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public</a></p>",
          "rawMarkdown": "@xxxxyyyy80008 \nThe adversarial validation was done to see the difference between train and private distributions and for feature selection.\nThe code is similar to the following notebook. (although the notebook is private and public)\nhttps://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public"
        }
      ]
    },
    {
      "id": 1925829,
      "postDate": "2022-09-04T10:59:32.073Z",
      "content": "<p>Elegant and simple explanation, thanks! I would like to ask:</p>\n<p><strong>1.</strong> Why did you decide to remove the round-up features?<br>\n<strong>2.</strong> You first create all the potential features (4300) and performed feature selection based on null importances afterwards, correct? How much was the score boost after that?</p>",
      "rawMarkdown": "Elegant and simple explanation, thanks! I would like to ask:\n\n**1.** Why did you decide to remove the round-up features?\n**2.** You first create all the potential features (4300) and performed feature selection based on null importances afterwards, correct? How much was the score boost after that?",
      "votes": 1,
      "replies": [
        {
          "id": 1927923,
          "postDate": "2022-09-06T04:32:58.187Z",
          "content": "<p><a href=\"https://www.kaggle.com/delai50\" target=\"_blank\">@delai50</a> <br>\nThe round-up feature was removed because it was considered unimportant because it was very highly correlated with before round-up. Due to tight memory, we removed the less important features.</p>\n<p>As you said, I created 4300 features first and then narrowed them down with null importance. Local CV was about 0.0005 better, but private score was worse. I think it made sense because the score got better when I put them in the ensemble.</p>",
          "rawMarkdown": "@delai50 \nThe round-up feature was removed because it was considered unimportant because it was very highly correlated with before round-up. Due to tight memory, we removed the less important features.\n\nAs you said, I created 4300 features first and then narrowed them down with null importance. Local CV was about 0.0005 better, but private score was worse. I think it made sense because the score got better when I put them in the ensemble.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1923289,
      "postDate": "2022-09-02T05:42:11.737Z",
      "content": "<p>An interesting solution. Thank you so much for the detailed description! My congratulations!</p>",
      "rawMarkdown": "An interesting solution. Thank you so much for the detailed description! My congratulations!",
      "votes": 1
    },
    {
      "id": 1920207,
      "postDate": "2022-08-31T02:02:40.370Z",
      "content": "<p>Thanks for sharing your solution <a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a>! Very insightful your meta-features.</p>",
      "rawMarkdown": "Thanks for sharing your solution @syaorn13! Very insightful your meta-features.",
      "votes": 1
    },
    {
      "id": 1917672,
      "postDate": "2022-08-29T01:00:47.353Z",
      "content": "<p>Very nice solution. <br>\nHave you done meta feature oof prediction with lightgbm?</p>",
      "rawMarkdown": "Very nice solution. \nHave you done meta feature oof prediction with lightgbm?",
      "votes": 1,
      "replies": [
        {
          "id": 1917772,
          "postDate": "2022-08-29T04:02:36.537Z",
          "content": "<p><a href=\"https://www.kaggle.com/leewook\" target=\"_blank\">@leewook</a> <br>\nYes.<br>\nI used LightGBM for oof prediction to create meta features.</p>",
          "rawMarkdown": "@leewook \nYes.\nI used LightGBM for oof prediction to create meta features.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1915420,
      "postDate": "2022-08-27T00:51:13.177Z",
      "content": "<p>What a crazy meta feature!</p>",
      "rawMarkdown": "What a crazy meta feature!",
      "votes": 1
    },
    {
      "id": 1914364,
      "postDate": "2022-08-26T03:27:29.627Z",
      "content": "<p>Excellent idea, let me think a lot, this competition can use more data, the labels of default data are slightly difficult to define in the time series, but the labels of non-default data must be the same in the time series; we can use rolling time series to use more data.</p>",
      "rawMarkdown": "Excellent idea, let me think a lot, this competition can use more data, the labels of default data are slightly difficult to define in the time series, but the labels of non-default data must be the same in the time series; we can use rolling time series to use more data.",
      "votes": 1
    },
    {
      "id": 1913904,
      "postDate": "2022-08-25T15:13:11.477Z",
      "content": "<p>Congratulations！Amazing solutions!</p>",
      "rawMarkdown": "Congratulations！Amazing solutions!",
      "votes": 1
    },
    {
      "id": 1913822,
      "postDate": "2022-08-25T14:25:12.070Z",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a>, and thanks for sharing how to make meta-features, btw, you wrote-</p>\n<blockquote>\n  <p>Features that were not important in adversarial validation are not used.</p>\n</blockquote>\n<p>What Features did you remove from the train?</p>",
      "rawMarkdown": "Congratulations! @syaorn13, and thanks for sharing how to make meta-features, btw, you wrote-\n> Features that were not important in adversarial validation are not used.\n\nWhat Features did you remove from the train?",
      "votes": 1,
      "replies": [
        {
          "id": 1913836,
          "postDate": "2022-08-25T14:31:38.403Z",
          "content": "<p>Thank you comment !</p>\n<blockquote>\n  <p>What Features did you remove from the train?</p>\n</blockquote>\n<p>The deleted features are \"R_1, D59, S_11, B_29\".</p>",
          "rawMarkdown": "Thank you comment !\n> What Features did you remove from the train?\n\nThe deleted features are \"R_1, D59, S_11, B_29\".",
          "votes": 2
        },
        {
          "id": 1913863,
          "postDate": "2022-08-25T14:48:29.100Z",
          "content": "<p><a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a> Thanks for replying, I really liked your approach of checking the relationship between public/private using AV, for our case we only looked at train/private and we removed many features, (['R_1', 'D_59', 'S_11', 'S_9', 'S_15', 'S_24', 'R_27', 'S_13', 'D_45', 'D_62', 'D_77']) in order to make the train more like private test but that resulted into a huge data loss. Our best single models xgbs, lgbms, catboosts couldn't pass beyond 0.7945, and we only got close to 0.797 with ranked ensembling of 20+ models from these 3 model types, for removing a large number of features! We learnt a great lesson! </p>\n<p>Btw what was the AV score between public/private after you removed \"R_1, D59, S_11, B_29\"?</p>",
          "rawMarkdown": "@syaorn13 Thanks for replying, I really liked your approach of checking the relationship between public/private using AV, for our case we only looked at train/private and we removed many features, (['R_1', 'D_59', 'S_11', 'S_9', 'S_15', 'S_24', 'R_27', 'S_13', 'D_45', 'D_62', 'D_77']) in order to make the train more like private test but that resulted into a huge data loss. Our best single models xgbs, lgbms, catboosts couldn't pass beyond 0.7945, and we only got close to 0.797 with ranked ensembling of 20+ models from these 3 model types, for removing a large number of features! We learnt a great lesson! \n\nBtw what was the AV score between public/private after you removed \"R_1, D59, S_11, B_29\"?"
        },
        {
          "id": 1914335,
          "postDate": "2022-08-26T02:28:15.460Z",
          "content": "<p><a href=\"https://www.kaggle.com/susnato\" target=\"_blank\">@susnato</a> <br>\nThe adversarial validation I performed was train/private.<br>\nAfter removing \"R_1, D59, S_11, B_29\" the AUC dropped to about 0.8 with simple gbdt.<br>\nI thought that removing many features would have more disadvantages, so I removed only the 4 features that had the greatest impact.<br>\nThe optimal number of features to be removed may be different because we did not conduct multiple experiments.</p>",
          "rawMarkdown": "@susnato \nThe adversarial validation I performed was train/private.\nAfter removing \"R_1, D59, S_11, B_29\" the AUC dropped to about 0.8 with simple gbdt.\nI thought that removing many features would have more disadvantages, so I removed only the 4 features that had the greatest impact.\nThe optimal number of features to be removed may be different because we did not conduct multiple experiments.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1913762,
      "postDate": "2022-08-25T13:47:36.027Z",
      "content": "<p>This is a very unique approach, thanks for sharing! Your OOF and associated aggregation approach is indeed a great learning resource for me! Congratulations to you for the result, I believe it was very well deserved <a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a>!!</p>",
      "rawMarkdown": "This is a very unique approach, thanks for sharing! Your OOF and associated aggregation approach is indeed a great learning resource for me! Congratulations to you for the result, I believe it was very well deserved @syaorn13!!",
      "votes": 1
    },
    {
      "id": 1913679,
      "postDate": "2022-08-25T12:49:15.953Z",
      "content": "<p>Congratulations! I have nearly the same features but never used them because of the ram limits of kaggle notebooks. Can I ask if you ever did more experiments related to meta-features in this competition? For example predict the probability that the customer will default next month, and if they don't use that as a feature for the next transaction and aggregate those probabilties as the number of transactions increased?</p>",
      "rawMarkdown": "Congratulations! I have nearly the same features but never used them because of the ram limits of kaggle notebooks. Can I ask if you ever did more experiments related to meta-features in this competition? For example predict the probability that the customer will default next month, and if they don't use that as a feature for the next transaction and aggregate those probabilties as the number of transactions increased?",
      "votes": 1,
      "replies": [
        {
          "id": 1913847,
          "postDate": "2022-08-25T14:35:57.433Z",
          "content": "<p>Thank you !</p>\n<p>No, we have not experimented with such a meta feature.<br>\nThe only default information given is \"will they default within 120 days of the latest statement\", so I feel it is difficult to predict whether they will default in the following month.</p>",
          "rawMarkdown": "Thank you !\n\nNo, we have not experimented with such a meta feature.\nThe only default information given is \"will they default within 120 days of the latest statement\", so I feel it is difficult to predict whether they will default in the following month.",
          "votes": 2
        },
        {
          "id": 1913869,
          "postDate": "2022-08-25T14:52:20.850Z",
          "content": "<p>Alright, I assumed you could just use the current labels and change the meaning of it (keep the 1 and 0 ).  Anyway Congratulations again</p>",
          "rawMarkdown": "Alright, I assumed you could just use the current labels and change the meaning of it (keep the 1 and 0 ).  Anyway Congratulations again"
        }
      ]
    },
    {
      "id": 1913609,
      "postDate": "2022-08-25T11:55:26.187Z",
      "content": "<p>Congrats and thanks for sharing! Can you share your code to make Meta features? It's really fancy!</p>",
      "rawMarkdown": "Congrats and thanks for sharing! Can you share your code to make Meta features? It's really fancy!",
      "votes": 1,
      "replies": [
        {
          "id": 1913828,
          "postDate": "2022-08-25T14:27:26.283Z",
          "content": "<p>Thank you!<br>\nPlease refer to the code in the comment addressed to chris.</p>",
          "rawMarkdown": "Thank you!\nPlease refer to the code in the comment addressed to chris.",
          "votes": 1
        },
        {
          "id": 1914306,
          "postDate": "2022-08-26T01:28:03.287Z",
          "content": "<p>OK, by the way, can you describe how to make Date features in detail? I tried to make Date features from 'S_2' which depict holiday and weekend, but it has littlte improvement to my model. </p>",
          "rawMarkdown": "OK, by the way, can you describe how to make Date features in detail? I tried to make Date features from 'S_2' which depict holiday and weekend, but it has littlte improvement to my model. "
        }
      ]
    },
    {
      "id": 1913577,
      "postDate": "2022-08-25T11:33:40.530Z",
      "content": "<p>Elegent and great solution! The meta features idea is interesting. Thanks for sharing!</p>",
      "rawMarkdown": "Elegent and great solution! The meta features idea is interesting. Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1914667,
      "postDate": "2022-08-26T09:37:56.270Z",
      "content": "<p>Excellent writeup. One question re Meta features, how do you generate the features for the test set, since the target is obviously not available? </p>",
      "rawMarkdown": "Excellent writeup. One question re Meta features, how do you generate the features for the test set, since the target is obviously not available? \n",
      "votes": 2,
      "replies": [
        {
          "id": 1915103,
          "postDate": "2022-08-26T16:59:24.877Z",
          "content": "<p>The creation of meta features in the test data is done by predicting models trained on the train data.<br>\nA model is created by assigning train_label to the train data before aggregating by customer ID.</p>",
          "rawMarkdown": "The creation of meta features in the test data is done by predicting models trained on the train data.\nA model is created by assigning train_label to the train data before aggregating by customer ID.",
          "votes": 1
        },
        {
          "id": 1998007,
          "postDate": "2022-10-21T10:26:43.400Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1913596,
      "postDate": "2022-08-25T11:47:58.380Z",
      "content": "<p>Congratulations on fantastic solo gold achievement. </p>\n<p>I am fascinated by your meta features. After you create your OOF, what exactly are you aggregating? Is it like <code>train.groupby('month').agg({'oof':'mean'})</code>? And then add the result as a new column in train dataframe to be used when training another model?</p>",
      "rawMarkdown": "Congratulations on fantastic solo gold achievement. \n\nI am fascinated by your meta features. After you create your OOF, what exactly are you aggregating? Is it like `train.groupby('month').agg({'oof':'mean'})`? And then add the result as a new column in train dataframe to be used when training another model?",
      "votes": 2,
      "replies": [
        {
          "id": 1913630,
          "postDate": "2022-08-25T12:21:42.603Z",
          "content": "<p>I use a similar idea, and the link <a href=\"https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388\" target=\"_blank\">https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388</a> he mentioned describes this idea very well. In this competition, after you get the oof score of each statement (most costomer_ID has 13 statements), you create features like: <code>train.groupby(['customer_ID'])['oof_score'].agg(['mean', 'min', 'max', 'std'])</code> and then add those features in another model.</p>",
          "rawMarkdown": "I use a similar idea, and the link https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388 he mentioned describes this idea very well. In this competition, after you get the oof score of each statement (most costomer_ID has 13 statements), you create features like: `train.groupby(['customer_ID'])['oof_score'].agg(['mean', 'min', 'max', 'std'])` and then add those features in another model.",
          "votes": 5
        },
        {
          "id": 1913675,
          "postDate": "2022-08-25T12:46:52.903Z",
          "content": "<blockquote>\n  <p>train.groupby(['customer_ID'])['oof_score'].agg(['mean', 'min', 'max', 'std'])</p>\n</blockquote>\n<p>ah I did the same here, also used the predictions with the date to create some more features, saw some mild seasonality in defaults for customers that mostly done transactions end of month vs mid month etc.. The model I used with this scored 0.80752 on private LB, I've selected final submissions so poorly 😄</p>",
          "rawMarkdown": ">train.groupby(['customer_ID'])['oof_score'].agg(['mean', 'min', 'max', 'std'])\n\nah I did the same here, also used the predictions with the date to create some more features, saw some mild seasonality in defaults for customers that mostly done transactions end of month vs mid month etc.. The model I used with this scored 0.80752 on private LB, I've selected final submissions so poorly 😄"
        },
        {
          "id": 1913685,
          "postDate": "2022-08-25T12:52:19.737Z",
          "content": "<p>What confuses me is \"oof score of each statement\". Does the OOF include 13 predictions per customer or 1 prediction per customer?</p>",
          "rawMarkdown": "What confuses me is \"oof score of each statement\". Does the OOF include 13 predictions per customer or 1 prediction per customer?"
        },
        {
          "id": 1913703,
          "postDate": "2022-08-25T13:04:59.453Z",
          "content": "<p>It includes 13 predictions per customer. <br>\nFirst step is to assign labels of customer_IDs to each statement, for example, if one customer_ID, cid_A is defaulted customer, we assign the label 1 for all cid_A's statement. After the first step, if we have 1000 customer_IDs and every customer_ID has 13 statements, then we have 13000 labelled samples now. Then we trained a model using this prepared samples (13000 labelled samples) to get each sample a oof score. After this, we aggragated the oof scores (13 oof scores) to get the stat features (min, max, mean, std…) and assign those stat features to each of the 1000 customer_IDs for downstream models. </p>",
          "rawMarkdown": "It includes 13 predictions per customer. \nFirst step is to assign labels of customer_IDs to each statement, for example, if one customer_ID, cid_A is defaulted customer, we assign the label 1 for all cid_A's statement. After the first step, if we have 1000 customer_IDs and every customer_ID has 13 statements, then we have 13000 labelled samples now. Then we trained a model using this prepared samples (13000 labelled samples) to get each sample a oof score. After this, we aggragated the oof scores (13 oof scores) to get the stat features (min, max, mean, std...) and assign those stat features to each of the 1000 customer_IDs for downstream models. ",
          "votes": 5
        },
        {
          "id": 1913732,
          "postDate": "2022-08-25T13:25:39.507Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nThank you comment!</p>\n<p>OOF include 13 prediction. Created as follows<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930293%2Fff38c51bc4fe862e1f9ff919af07a536%2Foof_pred.png?generation=1661433784065595&amp;alt=media\" alt=\"\"></p>\n<p>OOF prediction is aggregated as follows.</p>\n<pre><code>def make_meta_features(df):\n    n_months_list = [3, 6, 13]\n    df_list = []\n\n    for n_month in n_months_list:\n        df_period = df[df[\"n_month_from_latest\"] &lt;= n_month-1]  # min n_month=0\n        df_period = df_period.groupby(\"customer_ID\")[\"prediction\"].agg(['mean', 'median', 'std', 'min', 'max'])\n        df_period.columns = [f\"meta_{col}_latest{n_month}\" for col in df_period.columns]\n        df_list.append(df_period)\n    df_periods = pd.concat(df_list, axis=1)\n    df_periods.reset_index(inplace = True)\n\n    df_periods[\"meta_last\"] = df.groupby(\"customer_ID\")[\"prediction\"].agg(['last']).values\n\n    # make slope (≒rate)\n    df_periods[\"meta_slope_last_latest3\"] = df_periods[\"meta_last\"] / df_periods[\"meta_mean_latest3\"]\n    df_periods[\"meta_slope_latest3_latest6\"] = df_periods[\"meta_mean_latest3\"] / df_periods[\"meta_mean_latest6\"]\n    df_periods[\"meta_slope_latest6_latest13\"] = df_periods[\"meta_mean_latest6\"] / df_periods[\"meta_mean_latest13\"]\n\n    # make meta pred\n    df_periods[\"meta_pred_latest3\"] = df_periods[\"meta_last\"] * df_periods[\"meta_slope_last_latest3\"]\n    df_periods[\"meta_pred_latest6\"] = df_periods[\"meta_last\"] * df_periods[\"meta_slope_latest3_latest6\"]\n    df_periods[\"meta_pred_latest13\"] = df_periods[\"meta_last\"] * df_periods[\"meta_slope_latest6_latest13\"]\n\n    return df_periods\n</code></pre>\n<blockquote>\n  <p>And then add the result as a new column in train dataframe to be used when training another model?</p>\n</blockquote>\n<p>Yes. A new model was created by adding meta features to the original features.</p>",
          "rawMarkdown": "@cdeotte \nThank you comment!\n\nOOF include 13 prediction. Created as follows\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930293%2Fff38c51bc4fe862e1f9ff919af07a536%2Foof_pred.png?generation=1661433784065595&alt=media)\n\nOOF prediction is aggregated as follows.\n\n```\ndef make_meta_features(df):\n    n_months_list = [3, 6, 13]\n    df_list = []\n\n    for n_month in n_months_list:\n        df_period = df[df[\"n_month_from_latest\"] <= n_month-1]  # min n_month=0\n        df_period = df_period.groupby(\"customer_ID\")[\"prediction\"].agg(['mean', 'median', 'std', 'min', 'max'])\n        df_period.columns = [f\"meta_{col}_latest{n_month}\" for col in df_period.columns]\n        df_list.append(df_period)\n    df_periods = pd.concat(df_list, axis=1)\n    df_periods.reset_index(inplace = True)\n\n    df_periods[\"meta_last\"] = df.groupby(\"customer_ID\")[\"prediction\"].agg(['last']).values\n    \n    # make slope (≒rate)\n    df_periods[\"meta_slope_last_latest3\"] = df_periods[\"meta_last\"] / df_periods[\"meta_mean_latest3\"]\n    df_periods[\"meta_slope_latest3_latest6\"] = df_periods[\"meta_mean_latest3\"] / df_periods[\"meta_mean_latest6\"]\n    df_periods[\"meta_slope_latest6_latest13\"] = df_periods[\"meta_mean_latest6\"] / df_periods[\"meta_mean_latest13\"]\n\n    # make meta pred\n    df_periods[\"meta_pred_latest3\"] = df_periods[\"meta_last\"] * df_periods[\"meta_slope_last_latest3\"]\n    df_periods[\"meta_pred_latest6\"] = df_periods[\"meta_last\"] * df_periods[\"meta_slope_latest3_latest6\"]\n    df_periods[\"meta_pred_latest13\"] = df_periods[\"meta_last\"] * df_periods[\"meta_slope_latest6_latest13\"]\n    \n    return df_periods\n\n```\n\n\n\n> And then add the result as a new column in train dataframe to be used when training another model?\n\nYes. A new model was created by adding meta features to the original features.",
          "votes": 11
        },
        {
          "id": 1913789,
          "postDate": "2022-08-25T14:05:09.347Z",
          "content": "<p>Wow, those are great features <a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a> <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> . Great idea, great job!</p>",
          "rawMarkdown": "Wow, those are great features @syaorn13 @lihaorocky . Great idea, great job!",
          "votes": 2
        },
        {
          "id": 1914512,
          "postDate": "2022-08-26T06:39:32.673Z",
          "content": "<p><a href=\"https://www.kaggle.com/sayorn13\" target=\"_blank\">@sayorn13</a> And you do this meta-features for every categorical feature? And I can't understand clear, how this meta-feature vary for different customer_ID's,  because prediction base on n_numbers_from_last?</p>",
          "rawMarkdown": "@sayorn13 And you do this meta-features for every categorical feature? And I can't understand clear, how this meta-feature vary for different customer_ID's,  because prediction base on n_numbers_from_last?"
        },
        {
          "id": 1920320,
          "postDate": "2022-08-31T04:26:35.233Z",
          "content": "<p><a href=\"https://www.kaggle.com/alexbruskin\" target=\"_blank\">@alexbruskin</a><br>\nI am not sure if this is the intent of the question, but I will answer.<br>\nThe meta feature is not created for all categorical features.<br>\nTo create the oofpred for the meta feature, we also use about 190 features that existed in train, in addition to n_month_from_latest.<br>\nPlease refer to the 5th solution for an easy-to-understand illustration.<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097</a></p>",
          "rawMarkdown": "@alexbruskin\nI am not sure if this is the intent of the question, but I will answer.\nThe meta feature is not created for all categorical features.\nTo create the oofpred for the meta feature, we also use about 190 features that existed in train, in addition to n_month_from_latest.\nPlease refer to the 5th solution for an easy-to-understand illustration.\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/348097",
          "votes": 1
        }
      ]
    },
    {
      "id": 1928208,
      "postDate": "2022-09-06T10:17:17.100Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a> , can you please share your strategies for feature creations using groupby and how to decide where to start from, for ex- in this competition mostly Kagglers aggregated on customer_id. Congratulations and Thank you.</p>",
      "rawMarkdown": "Hi @syaorn13 , can you please share your strategies for feature creations using groupby and how to decide where to start from, for ex- in this competition mostly Kagglers aggregated on customer_id. Congratulations and Thank you."
    },
    {
      "id": 1921604,
      "postDate": "2022-09-01T00:40:02.377Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a>, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi @syaorn13, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants"
    },
    {
      "id": 1920427,
      "postDate": "2022-08-31T06:50:33.380Z",
      "content": "<p>Thanksy for sharing !!!</p>",
      "rawMarkdown": "Thanksy for sharing !!!"
    },
    {
      "id": 1914404,
      "postDate": "2022-08-26T04:36:45.813Z",
      "content": "<p>hi shaorooon</p>\n<p>what do you mean \"train_labels are assigned to train data (before aggregation by cid) and train model.\", I am confuse won't that result info leakage，and how to get this info in test data</p>\n<p>I'm new and question might be silly, could you give me an explanation?</p>",
      "rawMarkdown": "hi shaorooon\n\nwhat do you mean \"train_labels are assigned to train data (before aggregation by cid) and train model.\", I am confuse won't that result info leakage，and how to get this info in test data\n\nI'm new and question might be silly, could you give me an explanation?",
      "replies": [
        {
          "id": 1915106,
          "postDate": "2022-08-26T17:03:24.400Z",
          "content": "<p>The assignment of labels to train data is performed with code like this<br>\n<code>\ntrain = pd.read_csv(\"train_data.csv\")  # original train data\ntrain_label = pd.read_csv(\"train_labels.csv\")\ntrain.merge(train_label, how=\"left\", on=[\"customer_ID\"])\n</code></p>\n<p>The creation of meta features in the test data is done by predicting models trained on the train data.</p>",
          "rawMarkdown": "The assignment of labels to train data is performed with code like this\n`\ntrain = pd.read_csv(\"train_data.csv\")  # original train data\ntrain_label = pd.read_csv(\"train_labels.csv\")\ntrain.merge(train_label, how=\"left\", on=[\"customer_ID\"])\n`\n\nThe creation of meta features in the test data is done by predicting models trained on the train data.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1913584,
      "postDate": "2022-08-25T11:38:12.963Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1913752,
          "postDate": "2022-08-25T13:40:10.177Z",
          "content": "<p>Thank you!</p>\n<p>The reasons why I think meta feature is effective are as follows.</p>\n<ul>\n<li>When aggregating by customer_ID, some information is inevitably lost.</li>\n<li>By assigning a target to the train data before aggregation, we can calculate the \"default probability of a given customer given only one month of data\".</li>\n<li>The meta feature calculated in this way has a new property of \"default probability at the time of each month,\" and by aggregating it, the part that could not be captured by the existing features can be compensated (e.g., how the default probability changed).</li>\n</ul>",
          "rawMarkdown": "Thank you!\n\nThe reasons why I think meta feature is effective are as follows.\n- When aggregating by customer_ID, some information is inevitably lost.\n- By assigning a target to the train data before aggregation, we can calculate the \"default probability of a given customer given only one month of data\".\n- The meta feature calculated in this way has a new property of \"default probability at the time of each month,\" and by aggregating it, the part that could not be captured by the existing features can be compensated (e.g., how the default probability changed).",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1913628,
      "author_name": "sirius",
      "author_url": "",
      "post_date": "2022-08-25T12:20:54.807000",
      "content": "<p>Congrats! <br>\nI almost did the same thing with what you did about meta features, though I didn't know it was called meta feature before. <br>\nThe only difference is that I used the oof prediction as 13 numeric features directly and you did some aggregation. <br>\nThese features also boosted my model a lot！</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1913820,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-08-25T14:24:04.310000",
          "content": "<p>Thank you for comment !</p>\n<blockquote>\n  <p>The only difference is that I used the oof prediction as 13 numeric features directly and you did some aggregation.</p>\n</blockquote>\n<p>It did not occur to me to use the 13 numeric features directly. That seems to work too.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1948795,
      "author_name": "xxxxyyyy80008",
      "author_url": "",
      "post_date": "2022-09-21T09:43:16.187000",
      "content": "<p><a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a> very neat writeup and excellent solution. I learned a lot from it. </p>\n<p>I am very interested in your <code>adversarial validation</code> process, would you kindly elaborate a bit? </p>\n<p>thanks!</p>\n<p>xxxxyyyy80008</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1950135,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-09-22T06:42:37.367000",
          "content": "<p><a href=\"https://www.kaggle.com/xxxxyyyy80008\" target=\"_blank\">@xxxxyyyy80008</a> <br>\nThe adversarial validation was done to see the difference between train and private distributions and for feature selection.<br>\nThe code is similar to the following notebook. (although the notebook is private and public)<br>\n<a href=\"https://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public\" target=\"_blank\">https://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1925829,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2022-09-04T10:59:32.073000",
      "content": "<p>Elegant and simple explanation, thanks! I would like to ask:</p>\n<p><strong>1.</strong> Why did you decide to remove the round-up features?<br>\n<strong>2.</strong> You first create all the potential features (4300) and performed feature selection based on null importances afterwards, correct? How much was the score boost after that?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1927923,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-09-06T04:32:58.187000",
          "content": "<p><a href=\"https://www.kaggle.com/delai50\" target=\"_blank\">@delai50</a> <br>\nThe round-up feature was removed because it was considered unimportant because it was very highly correlated with before round-up. Due to tight memory, we removed the less important features.</p>\n<p>As you said, I created 4300 features first and then narrowed them down with null importance. Local CV was about 0.0005 better, but private score was worse. I think it made sense because the score got better when I put them in the ensemble.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1923289,
      "author_name": "Nikita Mikheev",
      "author_url": "",
      "post_date": "2022-09-02T05:42:11.737000",
      "content": "<p>An interesting solution. Thank you so much for the detailed description! My congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1920207,
      "author_name": "Oscar Aguilar",
      "author_url": "",
      "post_date": "2022-08-31T02:02:40.370000",
      "content": "<p>Thanks for sharing your solution <a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a>! Very insightful your meta-features.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1917672,
      "author_name": "ds.wook",
      "author_url": "",
      "post_date": "2022-08-29T01:00:47.353000",
      "content": "<p>Very nice solution. <br>\nHave you done meta feature oof prediction with lightgbm?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1917772,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-08-29T04:02:36.537000",
          "content": "<p><a href=\"https://www.kaggle.com/leewook\" target=\"_blank\">@leewook</a> <br>\nYes.<br>\nI used LightGBM for oof prediction to create meta features.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1915420,
      "author_name": "Jackson You",
      "author_url": "",
      "post_date": "2022-08-27T00:51:13.177000",
      "content": "<p>What a crazy meta feature!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1914364,
      "author_name": "Fading_Vay",
      "author_url": "",
      "post_date": "2022-08-26T03:27:29.627000",
      "content": "<p>Excellent idea, let me think a lot, this competition can use more data, the labels of default data are slightly difficult to define in the time series, but the labels of non-default data must be the same in the time series; we can use rolling time series to use more data.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913904,
      "author_name": "SgangX",
      "author_url": "",
      "post_date": "2022-08-25T15:13:11.477000",
      "content": "<p>Congratulations！Amazing solutions!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913822,
      "author_name": "Susnato Dhar",
      "author_url": "",
      "post_date": "2022-08-25T14:25:12.070000",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a>, and thanks for sharing how to make meta-features, btw, you wrote-</p>\n<blockquote>\n  <p>Features that were not important in adversarial validation are not used.</p>\n</blockquote>\n<p>What Features did you remove from the train?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1913836,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-08-25T14:31:38.403000",
          "content": "<p>Thank you comment !</p>\n<blockquote>\n  <p>What Features did you remove from the train?</p>\n</blockquote>\n<p>The deleted features are \"R_1, D59, S_11, B_29\".</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1913863,
          "author_name": "Susnato Dhar",
          "author_url": "",
          "post_date": "2022-08-25T14:48:29.100000",
          "content": "<p><a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a> Thanks for replying, I really liked your approach of checking the relationship between public/private using AV, for our case we only looked at train/private and we removed many features, (['R_1', 'D_59', 'S_11', 'S_9', 'S_15', 'S_24', 'R_27', 'S_13', 'D_45', 'D_62', 'D_77']) in order to make the train more like private test but that resulted into a huge data loss. Our best single models xgbs, lgbms, catboosts couldn't pass beyond 0.7945, and we only got close to 0.797 with ranked ensembling of 20+ models from these 3 model types, for removing a large number of features! We learnt a great lesson! </p>\n<p>Btw what was the AV score between public/private after you removed \"R_1, D59, S_11, B_29\"?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1914335,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-08-26T02:28:15.460000",
          "content": "<p><a href=\"https://www.kaggle.com/susnato\" target=\"_blank\">@susnato</a> <br>\nThe adversarial validation I performed was train/private.<br>\nAfter removing \"R_1, D59, S_11, B_29\" the AUC dropped to about 0.8 with simple gbdt.<br>\nI thought that removing many features would have more disadvantages, so I removed only the 4 features that had the greatest impact.<br>\nThe optimal number of features to be removed may be different because we did not conduct multiple experiments.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1913762,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-08-25T13:47:36.027000",
      "content": "<p>This is a very unique approach, thanks for sharing! Your OOF and associated aggregation approach is indeed a great learning resource for me! Congratulations to you for the result, I believe it was very well deserved <a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a>!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913679,
      "author_name": "Tarrasque9",
      "author_url": "",
      "post_date": "2022-08-25T12:49:15.953000",
      "content": "<p>Congratulations! I have nearly the same features but never used them because of the ram limits of kaggle notebooks. Can I ask if you ever did more experiments related to meta-features in this competition? For example predict the probability that the customer will default next month, and if they don't use that as a feature for the next transaction and aggregate those probabilties as the number of transactions increased?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1913847,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-08-25T14:35:57.433000",
          "content": "<p>Thank you !</p>\n<p>No, we have not experimented with such a meta feature.<br>\nThe only default information given is \"will they default within 120 days of the latest statement\", so I feel it is difficult to predict whether they will default in the following month.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1913869,
          "author_name": "Tarrasque9",
          "author_url": "",
          "post_date": "2022-08-25T14:52:20.850000",
          "content": "<p>Alright, I assumed you could just use the current labels and change the meaning of it (keep the 1 and 0 ).  Anyway Congratulations again</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1913609,
      "author_name": "brilliw",
      "author_url": "",
      "post_date": "2022-08-25T11:55:26.187000",
      "content": "<p>Congrats and thanks for sharing! Can you share your code to make Meta features? It's really fancy!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1913828,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-08-25T14:27:26.283000",
          "content": "<p>Thank you!<br>\nPlease refer to the code in the comment addressed to chris.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1914306,
          "author_name": "brilliw",
          "author_url": "",
          "post_date": "2022-08-26T01:28:03.287000",
          "content": "<p>OK, by the way, can you describe how to make Date features in detail? I tried to make Date features from 'S_2' which depict holiday and weekend, but it has littlte improvement to my model. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1913577,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2022-08-25T11:33:40.530000",
      "content": "<p>Elegent and great solution! The meta features idea is interesting. Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1914667,
      "author_name": "Jonathan Mallia",
      "author_url": "",
      "post_date": "2022-08-26T09:37:56.270000",
      "content": "<p>Excellent writeup. One question re Meta features, how do you generate the features for the test set, since the target is obviously not available? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1915103,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-08-26T16:59:24.877000",
          "content": "<p>The creation of meta features in the test data is done by predicting models trained on the train data.<br>\nA model is created by assigning train_label to the train data before aggregating by customer ID.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1998007,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-10-21T10:26:43.400000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1913596,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-08-25T11:47:58.380000",
      "content": "<p>Congratulations on fantastic solo gold achievement. </p>\n<p>I am fascinated by your meta features. After you create your OOF, what exactly are you aggregating? Is it like <code>train.groupby('month').agg({'oof':'mean'})</code>? And then add the result as a new column in train dataframe to be used when training another model?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1913630,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-08-25T12:21:42.603000",
          "content": "<p>I use a similar idea, and the link <a href=\"https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388\" target=\"_blank\">https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388</a> he mentioned describes this idea very well. In this competition, after you get the oof score of each statement (most costomer_ID has 13 statements), you create features like: <code>train.groupby(['customer_ID'])['oof_score'].agg(['mean', 'min', 'max', 'std'])</code> and then add those features in another model.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1913675,
          "author_name": "JM",
          "author_url": "",
          "post_date": "2022-08-25T12:46:52.903000",
          "content": "<blockquote>\n  <p>train.groupby(['customer_ID'])['oof_score'].agg(['mean', 'min', 'max', 'std'])</p>\n</blockquote>\n<p>ah I did the same here, also used the predictions with the date to create some more features, saw some mild seasonality in defaults for customers that mostly done transactions end of month vs mid month etc.. The model I used with this scored 0.80752 on private LB, I've selected final submissions so poorly 😄</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1913685,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-25T12:52:19.737000",
          "content": "<p>What confuses me is \"oof score of each statement\". Does the OOF include 13 predictions per customer or 1 prediction per customer?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1913703,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-08-25T13:04:59.453000",
          "content": "<p>It includes 13 predictions per customer. <br>\nFirst step is to assign labels of customer_IDs to each statement, for example, if one customer_ID, cid_A is defaulted customer, we assign the label 1 for all cid_A's statement. After the first step, if we have 1000 customer_IDs and every customer_ID has 13 statements, then we have 13000 labelled samples now. Then we trained a model using this prepared samples (13000 labelled samples) to get each sample a oof score. After this, we aggragated the oof scores (13 oof scores) to get the stat features (min, max, mean, std…) and assign those stat features to each of the 1000 customer_IDs for downstream models. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1913732,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-08-25T13:25:39.507000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nThank you comment!</p>\n<p>OOF include 13 prediction. Created as follows<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930293%2Fff38c51bc4fe862e1f9ff919af07a536%2Foof_pred.png?generation=1661433784065595&amp;alt=media\" alt=\"\"></p>\n<p>OOF prediction is aggregated as follows.</p>\n<pre><code>def make_meta_features(df):\n    n_months_list = [3, 6, 13]\n    df_list = []\n\n    for n_month in n_months_list:\n        df_period = df[df[\"n_month_from_latest\"] &lt;= n_month-1]  # min n_month=0\n        df_period = df_period.groupby(\"customer_ID\")[\"prediction\"].agg(['mean', 'median', 'std', 'min', 'max'])\n        df_period.columns = [f\"meta_{col}_latest{n_month}\" for col in df_period.columns]\n        df_list.append(df_period)\n    df_periods = pd.concat(df_list, axis=1)\n    df_periods.reset_index(inplace = True)\n\n    df_periods[\"meta_last\"] = df.groupby(\"customer_ID\")[\"prediction\"].agg(['last']).values\n\n    # make slope (≒rate)\n    df_periods[\"meta_slope_last_latest3\"] = df_periods[\"meta_last\"] / df_periods[\"meta_mean_latest3\"]\n    df_periods[\"meta_slope_latest3_latest6\"] = df_periods[\"meta_mean_latest3\"] / df_periods[\"meta_mean_latest6\"]\n    df_periods[\"meta_slope_latest6_latest13\"] = df_periods[\"meta_mean_latest6\"] / df_periods[\"meta_mean_latest13\"]\n\n    # make meta pred\n    df_periods[\"meta_pred_latest3\"] = df_periods[\"meta_last\"] * df_periods[\"meta_slope_last_latest3\"]\n    df_periods[\"meta_pred_latest6\"] = df_periods[\"meta_last\"] * df_periods[\"meta_slope_latest3_latest6\"]\n    df_periods[\"meta_pred_latest13\"] = df_periods[\"meta_last\"] * df_periods[\"meta_slope_latest6_latest13\"]\n\n    return df_periods\n</code></pre>\n<blockquote>\n  <p>And then add the result as a new column in train dataframe to be used when training another model?</p>\n</blockquote>\n<p>Yes. A new model was created by adding meta features to the original features.</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 1913789,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-25T14:05:09.347000",
          "content": "<p>Wow, those are great features <a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a> <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> . Great idea, great job!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1914512,
          "author_name": "moritz",
          "author_url": "",
          "post_date": "2022-08-26T06:39:32.673000",
          "content": "<p><a href=\"https://www.kaggle.com/sayorn13\" target=\"_blank\">@sayorn13</a> And you do this meta-features for every categorical feature? And I can't understand clear, how this meta-feature vary for different customer_ID's,  because prediction base on n_numbers_from_last?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1920320,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-08-31T04:26:35.233000",
          "content": "<p><a href=\"https://www.kaggle.com/alexbruskin\" target=\"_blank\">@alexbruskin</a><br>\nI am not sure if this is the intent of the question, but I will answer.<br>\nThe meta feature is not created for all categorical features.<br>\nTo create the oofpred for the meta feature, we also use about 190 features that existed in train, in addition to n_month_from_latest.<br>\nPlease refer to the 5th solution for an easy-to-understand illustration.<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1928208,
      "author_name": "Gaurav Malik",
      "author_url": "",
      "post_date": "2022-09-06T10:17:17.100000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a> , can you please share your strategies for feature creations using groupby and how to decide where to start from, for ex- in this competition mostly Kagglers aggregated on customer_id. Congratulations and Thank you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1921604,
      "author_name": "Yang Liu",
      "author_url": "",
      "post_date": "2022-09-01T00:40:02.377000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/syaorn13\" target=\"_blank\">@syaorn13</a>, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1920427,
      "author_name": "LewsTherin",
      "author_url": "",
      "post_date": "2022-08-31T06:50:33.380000",
      "content": "<p>Thanksy for sharing !!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1914404,
      "author_name": "wenjun zhang323",
      "author_url": "",
      "post_date": "2022-08-26T04:36:45.813000",
      "content": "<p>hi shaorooon</p>\n<p>what do you mean \"train_labels are assigned to train data (before aggregation by cid) and train model.\", I am confuse won't that result info leakage，and how to get this info in test data</p>\n<p>I'm new and question might be silly, could you give me an explanation?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1915106,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-08-26T17:03:24.400000",
          "content": "<p>The assignment of labels to train data is performed with code like this<br>\n<code>\ntrain = pd.read_csv(\"train_data.csv\")  # original train data\ntrain_label = pd.read_csv(\"train_labels.csv\")\ntrain.merge(train_label, how=\"left\", on=[\"customer_ID\"])\n</code></p>\n<p>The creation of meta features in the test data is done by predicting models trained on the train data.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1913584,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T11:38:12.963000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1913752,
          "author_name": "Shaorooon",
          "author_url": "",
          "post_date": "2022-08-25T13:40:10.177000",
          "content": "<p>Thank you!</p>\n<p>The reasons why I think meta feature is effective are as follows.</p>\n<ul>\n<li>When aggregating by customer_ID, some information is inevitably lost.</li>\n<li>By assigning a target to the train data before aggregation, we can calculate the \"default probability of a given customer given only one month of data\".</li>\n<li>The meta feature calculated in this way has a new property of \"default probability at the time of each month,\" and by aggregating it, the part that could not be captured by the existing features can be compensated (e.g., how the default probability changed).</li>\n</ul>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1913556": "※Previously titled 12th Place Solution. It is now 11th due to the fixed ranking.\n\nThank you to everyone who participated in the competition and to everyone involved in organizing it.\nI learned a lot through this competition.\n\n## Score & Result\n- My best submission\n    - Local CV：0.79922\n    - Public: 0.80088\n    - Private: 0.80852\n- Results\n    - Public:6th → Private 12th\n\n## Feature Engineering\n- Base features are from public notebook\n    - https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds\n- Delete some features\n    - Round up features\n    - Duplicate features (ex. groupby counts for all category features)\n- Add some features\n    - Features aggregated by time period \n        - min, max, mean, std for the last 3 and last 6 months\n    - Rate and diff features in time series\n        - ex. last - latest_3month_mean, last_3month_mean / last_6month_mean\n    - Null count features\n    - Date features\n    - **Meta features (most important features!)**\n        - how to make\n            1. Train_labels are assigned to train data (before aggregation by cid) and train model.\n            2. Make oof prediction for train data.\n            3. Aggregate oof prediction by time period.\n        - Using this feature, I reached 0.800 PublicLB from 0.799 PublicLB in single model.\n        - Referring to the DSB2019's 2nd place solution method.\n            - https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388\n\n## Validation strategy\n- Use stratfiedKfold(k=5).\n- I think Public LB is more important than local CV to measure Private's performance.\n    - Data size is about the same for train and public.\n    - In terms of time, public data is closer to private data than train data.\n    - Even after adversarial validation, train/private was farther away from the data than public/private.\n        - train/private：AUC 0.99\n        - public/private：AUC 0.82\n- While focusing on publicLB, we also looked at local CV to determine if there was any improvement.\n    - Also checked local logloss because amex_metric was not stable.\n    - It was hard to find a few digits of publicLB.\n\n## Model\n- LightGBM\n    - Use dart.\n    - Hypyer_parameter is the same as base notebook.\n        - https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds\n    - Get best amex metric model (use callback)\n        - https://www.kaggle.com/competitions/amex-default-prediction/discussion/332575    \n\n## Feature Selection\n- Adversarial validation\n    - Delete features of high importance in train/private adversarial validation.\n        - Drop_features: R_1, D59, S_11, B_29\n        - After the change, the AUC was 0.8.\n- Null importance\n    - Use features actual importance larger than mean importance with shuffled target.\n    - Before:4300 features → After:1300 features\n\n## Ensemble\n- Use 3 LightGBM models and rank ensemble (weighted average).\n- Each model use different feature set.\n    - Model_1: Not use meta features.\n    - Model_2: Use meta features & large features (not use null importance feature selection).\n    - Model_3: Use meta features & small features (use null importance feature selection).\n- Ensemble weight\n    - model_1:model_2:model_3 = 4:4:2\n    - weight decided while looking at public LB\n\n## Select Submission\n- I chose two sub's: \n    1. BestLB sub\n    2. Sub with risk of time-series changes in features.\n        - Features that were not important in adversarial validation are not used. \n- The second model was the best in privateLB.\n- Perhaps the trend of some features changed over time. adversarial validation was very helpful.\n\n\nThank you.",
    "1913628": "Congrats! \nI almost did the same thing with what you did about meta features, though I didn't know it was called meta feature before. \nThe only difference is that I used the oof prediction as 13 numeric features directly and you did some aggregation. \nThese features also boosted my model a lot！",
    "1948795": "@syaorn13 very neat writeup and excellent solution. I learned a lot from it. \n\nI am very interested in your `adversarial validation` process, would you kindly elaborate a bit? \n\nthanks!\n\nxxxxyyyy80008",
    "1925829": "Elegant and simple explanation, thanks! I would like to ask:\n\n**1.** Why did you decide to remove the round-up features?\n**2.** You first create all the potential features (4300) and performed feature selection based on null importances afterwards, correct? How much was the score boost after that?",
    "1923289": "An interesting solution. Thank you so much for the detailed description! My congratulations!",
    "1920207": "Thanks for sharing your solution @syaorn13! Very insightful your meta-features.",
    "1917672": "Very nice solution. \nHave you done meta feature oof prediction with lightgbm?",
    "1915420": "What a crazy meta feature!",
    "1914364": "Excellent idea, let me think a lot, this competition can use more data, the labels of default data are slightly difficult to define in the time series, but the labels of non-default data must be the same in the time series; we can use rolling time series to use more data.",
    "1913904": "Congratulations！Amazing solutions!",
    "1913822": "Congratulations! @syaorn13, and thanks for sharing how to make meta-features, btw, you wrote-\n> Features that were not important in adversarial validation are not used.\n\nWhat Features did you remove from the train?",
    "1913762": "This is a very unique approach, thanks for sharing! Your OOF and associated aggregation approach is indeed a great learning resource for me! Congratulations to you for the result, I believe it was very well deserved @syaorn13!!",
    "1913679": "Congratulations! I have nearly the same features but never used them because of the ram limits of kaggle notebooks. Can I ask if you ever did more experiments related to meta-features in this competition? For example predict the probability that the customer will default next month, and if they don't use that as a feature for the next transaction and aggregate those probabilties as the number of transactions increased?",
    "1913609": "Congrats and thanks for sharing! Can you share your code to make Meta features? It's really fancy!",
    "1913577": "Elegent and great solution! The meta features idea is interesting. Thanks for sharing!",
    "1914667": "Excellent writeup. One question re Meta features, how do you generate the features for the test set, since the target is obviously not available? \n",
    "1913596": "Congratulations on fantastic solo gold achievement. \n\nI am fascinated by your meta features. After you create your OOF, what exactly are you aggregating? Is it like `train.groupby('month').agg({'oof':'mean'})`? And then add the result as a new column in train dataframe to be used when training another model?",
    "1928208": "Hi @syaorn13 , can you please share your strategies for feature creations using groupby and how to decide where to start from, for ex- in this competition mostly Kagglers aggregated on customer_id. Congratulations and Thank you.",
    "1921604": "Hi @syaorn13, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
    "1920427": "Thanksy for sharing !!!",
    "1914404": "hi shaorooon\n\nwhat do you mean \"train_labels are assigned to train data (before aggregation by cid) and train model.\", I am confuse won't that result info leakage，and how to get this info in test data\n\nI'm new and question might be silly, could you give me an explanation?",
    "1913584": ""
  }
}