{
  "id": 347688,
  "title": "Thank you all!",
  "url": "/competitions/amex-default-prediction/discussion/347688",
  "author_name": "raddar",
  "post_date": "2022-08-25T06:09:31.872000",
  "votes": 141,
  "comment_count": 34,
  "views": 0,
  "content": "<p>What a roller coaster competition! Personally, it was one of the few I enjoyed and I am happy I was part of. I hadn't done any serious kaggling for 4 years up to this point. I have worked on purely computer vision problems for the last 4 years, and it's really refreshing to remember all the tabular ML joys and frustrations :D</p>\n<p>My solution is kind of overcomplicated - ~100 models (mix of lgb, xgb, tabnets) in 1st layer (CV 0.787-0.798), 4 models (xgb pairwise rank, lgb gbdt, MLP, tabnet) on 2nd layer stacking layer (CV 0.8002-0.8003), and finally tabnet as a final stacking model in 3rd layer. My best private score was also my best CV model (CV 0.8006). In the end I think I missed out on some important features, but could not figure it out till the very end…</p>\n<p>Personally, tabnet models were my favorite. They were not good enough on their own, but when you take 20-30 of them, magic started to happen - by their own they score 0.787-0.790 on CV, however their ensemble reached 0.796 on CV, which was on par to many of my xgb/lgb models.</p>\n<p>I have put the new feature \"Team up\" to test how successful this new thing is. I was 95% sure I would not team up, but it was interesting to test out, how much attention a top20 can get with just this flag. I got 43 email requests, 8 linkedin requests and even 1 facebook request to team up. That's 5x more invites than in all my kaggle career. I think this feature is cool and long overdue. I am sorry for those who I did not respond to or disappointed.</p>\n<p>Finally, thank you for all your kind words about my dataset - that really means a lot to me!</p>",
  "messages": [
    {
      "id": 1913086,
      "postDate": "2022-08-25T06:09:31.873Z",
      "content": "<p>What a roller coaster competition! Personally, it was one of the few I enjoyed and I am happy I was part of. I hadn't done any serious kaggling for 4 years up to this point. I have worked on purely computer vision problems for the last 4 years, and it's really refreshing to remember all the tabular ML joys and frustrations :D</p>\n<p>My solution is kind of overcomplicated - ~100 models (mix of lgb, xgb, tabnets) in 1st layer (CV 0.787-0.798), 4 models (xgb pairwise rank, lgb gbdt, MLP, tabnet) on 2nd layer stacking layer (CV 0.8002-0.8003), and finally tabnet as a final stacking model in 3rd layer. My best private score was also my best CV model (CV 0.8006). In the end I think I missed out on some important features, but could not figure it out till the very end…</p>\n<p>Personally, tabnet models were my favorite. They were not good enough on their own, but when you take 20-30 of them, magic started to happen - by their own they score 0.787-0.790 on CV, however their ensemble reached 0.796 on CV, which was on par to many of my xgb/lgb models.</p>\n<p>I have put the new feature \"Team up\" to test how successful this new thing is. I was 95% sure I would not team up, but it was interesting to test out, how much attention a top20 can get with just this flag. I got 43 email requests, 8 linkedin requests and even 1 facebook request to team up. That's 5x more invites than in all my kaggle career. I think this feature is cool and long overdue. I am sorry for those who I did not respond to or disappointed.</p>\n<p>Finally, thank you for all your kind words about my dataset - that really means a lot to me!</p>",
      "rawMarkdown": "What a roller coaster competition! Personally, it was one of the few I enjoyed and I am happy I was part of. I hadn't done any serious kaggling for 4 years up to this point. I have worked on purely computer vision problems for the last 4 years, and it's really refreshing to remember all the tabular ML joys and frustrations :D\n\nMy solution is kind of overcomplicated - ~100 models (mix of lgb, xgb, tabnets) in 1st layer (CV 0.787-0.798), 4 models (xgb pairwise rank, lgb gbdt, MLP, tabnet) on 2nd layer stacking layer (CV 0.8002-0.8003), and finally tabnet as a final stacking model in 3rd layer. My best private score was also my best CV model (CV 0.8006). In the end I think I missed out on some important features, but could not figure it out till the very end...\n\nPersonally, tabnet models were my favorite. They were not good enough on their own, but when you take 20-30 of them, magic started to happen - by their own they score 0.787-0.790 on CV, however their ensemble reached 0.796 on CV, which was on par to many of my xgb/lgb models.\n\n\nI have put the new feature \"Team up\" to test how successful this new thing is. I was 95% sure I would not team up, but it was interesting to test out, how much attention a top20 can get with just this flag. I got 43 email requests, 8 linkedin requests and even 1 facebook request to team up. That's 5x more invites than in all my kaggle career. I think this feature is cool and long overdue. I am sorry for those who I did not respond to or disappointed.\n\nFinally, thank you for all your kind words about my dataset - that really means a lot to me!\n\n\n\n\n",
      "votes": 140
    },
    {
      "id": 1913156,
      "postDate": "2022-08-25T07:26:35.867Z",
      "content": "<p>Thank you so much for you impact!</p>\n<p>I personally think that this is the case when you got much more respect and recognition by helping others rather than getting a gold medal.</p>\n<p>Sad that you only tested the \"team up\" feature. If you was reading our request you would not be able to reject :) And we would get a chance to learn much more from you.</p>\n<p>Looking forward for new competitions to learn from you</p>\n<p>P.S. We got 0.797 with tabnet - I will share it later :)</p>",
      "rawMarkdown": "Thank you so much for you impact!\n\nI personally think that this is the case when you got much more respect and recognition by helping others rather than getting a gold medal.\n\nSad that you only tested the \"team up\" feature. If you was reading our request you would not be able to reject :) And we would get a chance to learn much more from you.\n\nLooking forward for new competitions to learn from you\n\nP.S. We got 0.797 with tabnet - I will share it later :)",
      "votes": 8,
      "replies": [
        {
          "id": 1913220,
          "postDate": "2022-08-25T08:32:27.513Z",
          "content": "<p>Just read it - looks compelling indeed :)</p>",
          "rawMarkdown": "Just read it - looks compelling indeed :)",
          "votes": 2
        },
        {
          "id": 1913386,
          "postDate": "2022-08-25T09:59:57.190Z",
          "content": "<p>0.797 tabnet is no joke! <br>\nHow did you do it? what is the trick? </p>",
          "rawMarkdown": "0.797 tabnet is no joke! \nHow did you do it? what is the trick? ",
          "votes": 2
        },
        {
          "id": 1915162,
          "postDate": "2022-08-26T18:03:10.520Z",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347880\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/347880</a></p>",
          "rawMarkdown": "https://www.kaggle.com/competitions/amex-default-prediction/discussion/347880",
          "votes": 2
        }
      ]
    },
    {
      "id": 1914233,
      "postDate": "2022-08-25T21:21:41.800Z",
      "content": "<p>Thanks Radar for your contributions in this competition. Your many discussions, notebooks, and datasets were helpful. </p>\n<p>Your participation and enthusiasm made the competition more enjoyable. You facilitated lots of interesting discussions and encouraged more people to enter this competition with your memory friendly dataset.</p>",
      "rawMarkdown": "Thanks Radar for your contributions in this competition. Your many discussions, notebooks, and datasets were helpful. \n\nYour participation and enthusiasm made the competition more enjoyable. You facilitated lots of interesting discussions and encouraged more people to enter this competition with your memory friendly dataset.",
      "votes": 5
    },
    {
      "id": 1914235,
      "postDate": "2022-08-25T21:25:36.580Z",
      "content": "<p>I played around with TabNet too. I tried to use its unsupervised functionality to learn the test data. And I tried using test pseudo labels and other knowledge distillation tricks to boost the CV. But i couldn't boost the CV very much. In contrast, i got my Transformer CV 0.788/LB 0.790 up to CV 0.798/LB 0.799. I thought I could do something similar with TabNet but never figured out how to get it to work with TabNet.</p>\n<p>(TabNet also gave an error if i tried to train classification with soft targets. And i didn't dig into the source code enough to update the loss function)</p>",
      "rawMarkdown": "I played around with TabNet too. I tried to use its unsupervised functionality to learn the test data. And I tried using test pseudo labels and other knowledge distillation tricks to boost the CV. But i couldn't boost the CV very much. In contrast, i got my Transformer CV 0.788/LB 0.790 up to CV 0.798/LB 0.799. I thought I could do something similar with TabNet but never figured out how to get it to work with TabNet.\n\n(TabNet also gave an error if i tried to train classification with soft targets. And i didn't dig into the source code enough to update the loss function)",
      "votes": 3
    },
    {
      "id": 1913391,
      "postDate": "2022-08-25T10:00:49.830Z",
      "content": "<p>You have come back after 4 years’ absence! I would like to ask you about score and rank of the public leaderboard in Kaggle. You're a top of top person as a Kaggler. You made many kinds of overcomplicated approximately 100 models and then reached 0.8006 in CV. However actually there might be no difference between +-0.05 statistically comparison to others. What does it suggested in difference between them in the real world? I like Kaggle and I really respect for you, but the difference is significantly small. Is this competition special?</p>",
      "rawMarkdown": "You have come back after 4 years’ absence! I would like to ask you about score and rank of the public leaderboard in Kaggle. You're a top of top person as a Kaggler. You made many kinds of overcomplicated approximately 100 models and then reached 0.8006 in CV. However actually there might be no difference between +-0.05 statistically comparison to others. What does it suggested in difference between them in the real world? I like Kaggle and I really respect for you, but the difference is significantly small. Is this competition special?",
      "votes": 1,
      "replies": [
        {
          "id": 1913538,
          "postDate": "2022-08-25T11:00:35.417Z",
          "content": "<p>I would think that kaggle as a whole is like being a cherry on top of the cake (in terms of all the modelling aspects). In real life scenario you would settle with the solution which balances between complexity, inference time, explainability, etc. This very well may mean that taking a single model from that 100 model ensemble would suffice. However, kaggle allows curious minded data scientists to push the limits of what ML can achieve. And this is why kaggle is unique in its way - in real environment you would not do this, and if you did - your colleagues would think you are insane :)</p>",
          "rawMarkdown": "I would think that kaggle as a whole is like being a cherry on top of the cake (in terms of all the modelling aspects). In real life scenario you would settle with the solution which balances between complexity, inference time, explainability, etc. This very well may mean that taking a single model from that 100 model ensemble would suffice. However, kaggle allows curious minded data scientists to push the limits of what ML can achieve. And this is why kaggle is unique in its way - in real environment you would not do this, and if you did - your colleagues would think you are insane :)\n\n",
          "votes": 7
        },
        {
          "id": 1913571,
          "postDate": "2022-08-25T11:26:54.270Z",
          "content": "<p>I really appreciate your convincing answer. I admire that you also think about balances which you explained. At first I would like to reach the level which can be insane! and then I'll pretend an average data scientist!</p>",
          "rawMarkdown": "I really appreciate your convincing answer. I admire that you also think about balances which you explained. At first I would like to reach the level which can be insane! and then I'll pretend an average data scientist!"
        }
      ]
    },
    {
      "id": 1913172,
      "postDate": "2022-08-25T07:40:04.163Z",
      "content": "<p>If it had not been said enough, your dataset rocks ! Thanks you for sharing it</p>",
      "rawMarkdown": "If it had not been said enough, your dataset rocks ! Thanks you for sharing it",
      "votes": 1,
      "replies": [
        {
          "id": 1914083,
          "postDate": "2022-08-25T17:59:50.903Z",
          "content": "<p>Without his dataset many would have probably not hung in there in the competition! A great contribution to the community.</p>",
          "rawMarkdown": "Without his dataset many would have probably not hung in there in the competition! A great contribution to the community.\n"
        }
      ]
    },
    {
      "id": 1913493,
      "postDate": "2022-08-25T10:41:34.070Z",
      "content": "<p>Thanks for sharing your hindsights &amp; data set. If you want more features feel free to reach out to me next time :-)</p>",
      "rawMarkdown": "Thanks for sharing your hindsights & data set. If you want more features feel free to reach out to me next time :-)",
      "votes": 2
    },
    {
      "id": 1930220,
      "postDate": "2022-09-07T16:51:01.880Z",
      "content": "<p>Thanks for the dataset and everything that you shared, hope to see you active on Kaggle again!</p>",
      "rawMarkdown": "Thanks for the dataset and everything that you shared, hope to see you active on Kaggle again!"
    },
    {
      "id": 1923619,
      "postDate": "2022-09-02T11:51:33.807Z",
      "content": "<p>I mark.good</p>",
      "rawMarkdown": "I mark.good"
    },
    {
      "id": 1921632,
      "postDate": "2022-09-01T01:07:02.223Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi @raddar May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "replies": [
        {
          "id": 1926484,
          "postDate": "2022-09-04T21:42:16.510Z",
          "content": "<p>no problem - done</p>",
          "rawMarkdown": "no problem - done"
        }
      ]
    },
    {
      "id": 1917423,
      "postDate": "2022-08-28T18:14:17.057Z",
      "content": "<p>Thanks for the sharing of your data set, meant people with very limited time (like myself :)) could compete!</p>\n<p>I was wondering if you could share your validation scheme for your stacking? </p>\n<p>Did you reserve a hold-out test set and then use this for the multiple layer stacking or used nested CV?</p>",
      "rawMarkdown": "Thanks for the sharing of your data set, meant people with very limited time (like myself :)) could compete!\n\nI was wondering if you could share your validation scheme for your stacking? \n\nDid you reserve a hold-out test set and then use this for the multiple layer stacking or used nested CV?"
    },
    {
      "id": 1915169,
      "postDate": "2022-08-26T18:07:06.260Z",
      "content": "<p>Can you please elaborate on your tabnet strategy? 20-30 models of seed-averaging or any additional changes?</p>",
      "rawMarkdown": "Can you please elaborate on your tabnet strategy? 20-30 models of seed-averaging or any additional changes?",
      "replies": [
        {
          "id": 1915322,
          "postDate": "2022-08-26T21:32:04.323Z",
          "content": "<p>each trained with different set of params. linear blend on top to reach 0.796</p>",
          "rawMarkdown": "each trained with different set of params. linear blend on top to reach 0.796",
          "votes": 2
        }
      ]
    },
    {
      "id": 1914681,
      "postDate": "2022-08-26T09:56:49.847Z",
      "content": "<p>Thanks Raddar for your dataset sharing and lots of insights from discussion forum!</p>",
      "rawMarkdown": "Thanks Raddar for your dataset sharing and lots of insights from discussion forum!"
    },
    {
      "id": 1914356,
      "postDate": "2022-08-26T03:10:25.023Z",
      "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> Like for everyone, your dataset saved my life. But for me it was your Days Overdue notebook that gave me the most benefit and inspiration. I mean, you even directly suggested the possible next steps at the end of the notebook. </p>\n<p>I really appreciate the work you did, and how openly you shared with the community!</p>",
      "rawMarkdown": "@raddar Like for everyone, your dataset saved my life. But for me it was your Days Overdue notebook that gave me the most benefit and inspiration. I mean, you even directly suggested the possible next steps at the end of the notebook. \n\nI really appreciate the work you did, and how openly you shared with the community!"
    },
    {
      "id": 1914025,
      "postDate": "2022-08-25T16:59:11.717Z",
      "content": "<p>Nice to see you could make TabNet part of the solution. I trained TabNet models with maximum of ~1300 features, it did not improve my ensemble or stacking. Did you use more than 1300 features?</p>",
      "rawMarkdown": "Nice to see you could make TabNet part of the solution. I trained TabNet models with maximum of ~1300 features, it did not improve my ensemble or stacking. Did you use more than 1300 features?",
      "replies": [
        {
          "id": 1914040,
          "postDate": "2022-08-25T17:16:05.250Z",
          "content": "<p>1000 features~</p>",
          "rawMarkdown": "1000 features~",
          "votes": 2
        }
      ]
    },
    {
      "id": 1913727,
      "postDate": "2022-08-25T13:21:26.430Z",
      "content": "<p>Many thanks to you for your dataset and your insights <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>! Congratulations to you for the result as well!</p>",
      "rawMarkdown": "Many thanks to you for your dataset and your insights @raddar! Congratulations to you for the result as well!"
    },
    {
      "id": 1913649,
      "postDate": "2022-08-25T12:33:28.533Z",
      "content": "<p>👀 I clicked your flag with no response… 👀, thanks for your awesome dataset anyway. </p>",
      "rawMarkdown": "👀 I clicked your flag with no response... 👀, thanks for your awesome dataset anyway. "
    },
    {
      "id": 1913247,
      "postDate": "2022-08-25T08:56:03.047Z",
      "content": "<p>Thank you very much for your notebooks, discussion, and the dataset.</p>",
      "rawMarkdown": "Thank you very much for your notebooks, discussion, and the dataset."
    },
    {
      "id": 1913188,
      "postDate": "2022-08-25T07:55:44.710Z",
      "content": "<p>Thanks for sharing! Newbie in competitions here. Do you have any suggestions on source material for these multi-level kind of models you worked on?</p>",
      "rawMarkdown": "Thanks for sharing! Newbie in competitions here. Do you have any suggestions on source material for these multi-level kind of models you worked on?",
      "replies": [
        {
          "id": 1913214,
          "postDate": "2022-08-25T08:26:56.987Z",
          "content": "<p>It's a complex task. I do not have any resources on that, I learned that on my own. But i think there is for sure some good material on that. Maybe someone else can point it out</p>",
          "rawMarkdown": "It's a complex task. I do not have any resources on that, I learned that on my own. But i think there is for sure some good material on that. Maybe someone else can point it out",
          "votes": 1
        },
        {
          "id": 1913444,
          "postDate": "2022-08-25T10:25:46.873Z",
          "content": "<p>Thank you anyway and congrats!</p>",
          "rawMarkdown": "Thank you anyway and congrats!"
        },
        {
          "id": 1914163,
          "postDate": "2022-08-25T19:41:01.573Z",
          "content": "<p><a href=\"https://towardsdatascience.com/simple-model-stacking-explained-and-automated-1b54e4357916\" target=\"_blank\">https://towardsdatascience.com/simple-model-stacking-explained-and-automated-1b54e4357916</a></p>",
          "rawMarkdown": "https://towardsdatascience.com/simple-model-stacking-explained-and-automated-1b54e4357916"
        }
      ]
    },
    {
      "id": 1913206,
      "postDate": "2022-08-25T08:13:29.953Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1913218,
          "postDate": "2022-08-25T08:31:24.500Z",
          "content": "<p>I used all the standard features most people have used (min/max/average/mean/median/stdev). I used numpy <code>np.nanmean</code> equivalents - this means ignoring NA's when calculating aggregates. And then simply number of NA's as features.</p>\n<p>What I did more is I build 188 xgboost models for each variable to predict the target (using 13 data points + some stats). So each xgboost prediction then could be considered as an aggregator function. </p>\n<p>My best single one-seed LGB was 0.7979 on CV - so you have me beat :)</p>",
          "rawMarkdown": "I used all the standard features most people have used (min/max/average/mean/median/stdev). I used numpy `np.nanmean` equivalents - this means ignoring NA's when calculating aggregates. And then simply number of NA's as features.\n\nWhat I did more is I build 188 xgboost models for each variable to predict the target (using 13 data points + some stats). So each xgboost prediction then could be considered as an aggregator function. \n\nMy best single one-seed LGB was 0.7979 on CV - so you have me beat :)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1917001,
      "postDate": "2022-08-28T10:28:08.850Z",
      "content": "<p>Thank you for the insights. </p>",
      "rawMarkdown": "Thank you for the insights. "
    },
    {
      "id": 1913839,
      "postDate": "2022-08-25T14:33:09.793Z",
      "content": "<p>thank you for your impressive work!</p>",
      "rawMarkdown": "thank you for your impressive work!"
    }
  ],
  "comments": [
    {
      "id": 1913156,
      "author_name": "Pavel Vodolazov",
      "author_url": "",
      "post_date": "2022-08-25T07:26:35.867000",
      "content": "<p>Thank you so much for you impact!</p>\n<p>I personally think that this is the case when you got much more respect and recognition by helping others rather than getting a gold medal.</p>\n<p>Sad that you only tested the \"team up\" feature. If you was reading our request you would not be able to reject :) And we would get a chance to learn much more from you.</p>\n<p>Looking forward for new competitions to learn from you</p>\n<p>P.S. We got 0.797 with tabnet - I will share it later :)</p>",
      "votes": 8,
      "replies": [
        {
          "id": 1913220,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-08-25T08:32:27.513000",
          "content": "<p>Just read it - looks compelling indeed :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1913386,
          "author_name": "The Devastator",
          "author_url": "",
          "post_date": "2022-08-25T09:59:57.190000",
          "content": "<p>0.797 tabnet is no joke! <br>\nHow did you do it? what is the trick? </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1915162,
          "author_name": "Pavel Vodolazov",
          "author_url": "",
          "post_date": "2022-08-26T18:03:10.520000",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347880\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/347880</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1914233,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-08-25T21:21:41.800000",
      "content": "<p>Thanks Radar for your contributions in this competition. Your many discussions, notebooks, and datasets were helpful. </p>\n<p>Your participation and enthusiasm made the competition more enjoyable. You facilitated lots of interesting discussions and encouraged more people to enter this competition with your memory friendly dataset.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1914235,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-08-25T21:25:36.580000",
      "content": "<p>I played around with TabNet too. I tried to use its unsupervised functionality to learn the test data. And I tried using test pseudo labels and other knowledge distillation tricks to boost the CV. But i couldn't boost the CV very much. In contrast, i got my Transformer CV 0.788/LB 0.790 up to CV 0.798/LB 0.799. I thought I could do something similar with TabNet but never figured out how to get it to work with TabNet.</p>\n<p>(TabNet also gave an error if i tried to train classification with soft targets. And i didn't dig into the source code enough to update the loss function)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1913391,
      "author_name": "Daisy",
      "author_url": "",
      "post_date": "2022-08-25T10:00:49.830000",
      "content": "<p>You have come back after 4 years’ absence! I would like to ask you about score and rank of the public leaderboard in Kaggle. You're a top of top person as a Kaggler. You made many kinds of overcomplicated approximately 100 models and then reached 0.8006 in CV. However actually there might be no difference between +-0.05 statistically comparison to others. What does it suggested in difference between them in the real world? I like Kaggle and I really respect for you, but the difference is significantly small. Is this competition special?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1913538,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-08-25T11:00:35.417000",
          "content": "<p>I would think that kaggle as a whole is like being a cherry on top of the cake (in terms of all the modelling aspects). In real life scenario you would settle with the solution which balances between complexity, inference time, explainability, etc. This very well may mean that taking a single model from that 100 model ensemble would suffice. However, kaggle allows curious minded data scientists to push the limits of what ML can achieve. And this is why kaggle is unique in its way - in real environment you would not do this, and if you did - your colleagues would think you are insane :)</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1913571,
          "author_name": "Daisy",
          "author_url": "",
          "post_date": "2022-08-25T11:26:54.270000",
          "content": "<p>I really appreciate your convincing answer. I admire that you also think about balances which you explained. At first I would like to reach the level which can be insane! and then I'll pretend an average data scientist!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1913172,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2022-08-25T07:40:04.163000",
      "content": "<p>If it had not been said enough, your dataset rocks ! Thanks you for sharing it</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1914083,
          "author_name": "S Charlesworth",
          "author_url": "",
          "post_date": "2022-08-25T17:59:50.903000",
          "content": "<p>Without his dataset many would have probably not hung in there in the competition! A great contribution to the community.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1913493,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2022-08-25T10:41:34.070000",
      "content": "<p>Thanks for sharing your hindsights &amp; data set. If you want more features feel free to reach out to me next time :-)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1930220,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2022-09-07T16:51:01.880000",
      "content": "<p>Thanks for the dataset and everything that you shared, hope to see you active on Kaggle again!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1923619,
      "author_name": "Mark",
      "author_url": "",
      "post_date": "2022-09-02T11:51:33.807000",
      "content": "<p>I mark.good</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1921632,
      "author_name": "Yang Liu",
      "author_url": "",
      "post_date": "2022-09-01T01:07:02.223000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1926484,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-09-04T21:42:16.510000",
          "content": "<p>no problem - done</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1917423,
      "author_name": "FChmiel",
      "author_url": "",
      "post_date": "2022-08-28T18:14:17.057000",
      "content": "<p>Thanks for the sharing of your data set, meant people with very limited time (like myself :)) could compete!</p>\n<p>I was wondering if you could share your validation scheme for your stacking? </p>\n<p>Did you reserve a hold-out test set and then use this for the multiple layer stacking or used nested CV?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1915169,
      "author_name": "Pavel Vodolazov",
      "author_url": "",
      "post_date": "2022-08-26T18:07:06.260000",
      "content": "<p>Can you please elaborate on your tabnet strategy? 20-30 models of seed-averaging or any additional changes?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1915322,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-08-26T21:32:04.323000",
          "content": "<p>each trained with different set of params. linear blend on top to reach 0.796</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1914681,
      "author_name": "Dr. Alvinleenh",
      "author_url": "",
      "post_date": "2022-08-26T09:56:49.847000",
      "content": "<p>Thanks Raddar for your dataset sharing and lots of insights from discussion forum!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1914356,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2022-08-26T03:10:25.023000",
      "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> Like for everyone, your dataset saved my life. But for me it was your Days Overdue notebook that gave me the most benefit and inspiration. I mean, you even directly suggested the possible next steps at the end of the notebook. </p>\n<p>I really appreciate the work you did, and how openly you shared with the community!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1914025,
      "author_name": "HinePo",
      "author_url": "",
      "post_date": "2022-08-25T16:59:11.717000",
      "content": "<p>Nice to see you could make TabNet part of the solution. I trained TabNet models with maximum of ~1300 features, it did not improve my ensemble or stacking. Did you use more than 1300 features?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1914040,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-08-25T17:16:05.250000",
          "content": "<p>1000 features~</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1913727,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-08-25T13:21:26.430000",
      "content": "<p>Many thanks to you for your dataset and your insights <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>! Congratulations to you for the result as well!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1913649,
      "author_name": "MichaelG",
      "author_url": "",
      "post_date": "2022-08-25T12:33:28.533000",
      "content": "<p>👀 I clicked your flag with no response… 👀, thanks for your awesome dataset anyway. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1913247,
      "author_name": "Gaju Ahmed",
      "author_url": "",
      "post_date": "2022-08-25T08:56:03.047000",
      "content": "<p>Thank you very much for your notebooks, discussion, and the dataset.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1913188,
      "author_name": "Antonio Intini",
      "author_url": "",
      "post_date": "2022-08-25T07:55:44.710000",
      "content": "<p>Thanks for sharing! Newbie in competitions here. Do you have any suggestions on source material for these multi-level kind of models you worked on?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1913214,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-08-25T08:26:56.987000",
          "content": "<p>It's a complex task. I do not have any resources on that, I learned that on my own. But i think there is for sure some good material on that. Maybe someone else can point it out</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1913444,
          "author_name": "Antonio Intini",
          "author_url": "",
          "post_date": "2022-08-25T10:25:46.873000",
          "content": "<p>Thank you anyway and congrats!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1914163,
          "author_name": "Chris Miles",
          "author_url": "",
          "post_date": "2022-08-25T19:41:01.573000",
          "content": "<p><a href=\"https://towardsdatascience.com/simple-model-stacking-explained-and-automated-1b54e4357916\" target=\"_blank\">https://towardsdatascience.com/simple-model-stacking-explained-and-automated-1b54e4357916</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1913206,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T08:13:29.953000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1913218,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-08-25T08:31:24.500000",
          "content": "<p>I used all the standard features most people have used (min/max/average/mean/median/stdev). I used numpy <code>np.nanmean</code> equivalents - this means ignoring NA's when calculating aggregates. And then simply number of NA's as features.</p>\n<p>What I did more is I build 188 xgboost models for each variable to predict the target (using 13 data points + some stats). So each xgboost prediction then could be considered as an aggregator function. </p>\n<p>My best single one-seed LGB was 0.7979 on CV - so you have me beat :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1917001,
      "author_name": "Harry Mapodile",
      "author_url": "",
      "post_date": "2022-08-28T10:28:08.850000",
      "content": "<p>Thank you for the insights. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1913839,
      "author_name": "ffloyd",
      "author_url": "",
      "post_date": "2022-08-25T14:33:09.793000",
      "content": "<p>thank you for your impressive work!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1913086": "What a roller coaster competition! Personally, it was one of the few I enjoyed and I am happy I was part of. I hadn't done any serious kaggling for 4 years up to this point. I have worked on purely computer vision problems for the last 4 years, and it's really refreshing to remember all the tabular ML joys and frustrations :D\n\nMy solution is kind of overcomplicated - ~100 models (mix of lgb, xgb, tabnets) in 1st layer (CV 0.787-0.798), 4 models (xgb pairwise rank, lgb gbdt, MLP, tabnet) on 2nd layer stacking layer (CV 0.8002-0.8003), and finally tabnet as a final stacking model in 3rd layer. My best private score was also my best CV model (CV 0.8006). In the end I think I missed out on some important features, but could not figure it out till the very end...\n\nPersonally, tabnet models were my favorite. They were not good enough on their own, but when you take 20-30 of them, magic started to happen - by their own they score 0.787-0.790 on CV, however their ensemble reached 0.796 on CV, which was on par to many of my xgb/lgb models.\n\n\nI have put the new feature \"Team up\" to test how successful this new thing is. I was 95% sure I would not team up, but it was interesting to test out, how much attention a top20 can get with just this flag. I got 43 email requests, 8 linkedin requests and even 1 facebook request to team up. That's 5x more invites than in all my kaggle career. I think this feature is cool and long overdue. I am sorry for those who I did not respond to or disappointed.\n\nFinally, thank you for all your kind words about my dataset - that really means a lot to me!\n\n\n\n\n",
    "1913156": "Thank you so much for you impact!\n\nI personally think that this is the case when you got much more respect and recognition by helping others rather than getting a gold medal.\n\nSad that you only tested the \"team up\" feature. If you was reading our request you would not be able to reject :) And we would get a chance to learn much more from you.\n\nLooking forward for new competitions to learn from you\n\nP.S. We got 0.797 with tabnet - I will share it later :)",
    "1914233": "Thanks Radar for your contributions in this competition. Your many discussions, notebooks, and datasets were helpful. \n\nYour participation and enthusiasm made the competition more enjoyable. You facilitated lots of interesting discussions and encouraged more people to enter this competition with your memory friendly dataset.",
    "1914235": "I played around with TabNet too. I tried to use its unsupervised functionality to learn the test data. And I tried using test pseudo labels and other knowledge distillation tricks to boost the CV. But i couldn't boost the CV very much. In contrast, i got my Transformer CV 0.788/LB 0.790 up to CV 0.798/LB 0.799. I thought I could do something similar with TabNet but never figured out how to get it to work with TabNet.\n\n(TabNet also gave an error if i tried to train classification with soft targets. And i didn't dig into the source code enough to update the loss function)",
    "1913391": "You have come back after 4 years’ absence! I would like to ask you about score and rank of the public leaderboard in Kaggle. You're a top of top person as a Kaggler. You made many kinds of overcomplicated approximately 100 models and then reached 0.8006 in CV. However actually there might be no difference between +-0.05 statistically comparison to others. What does it suggested in difference between them in the real world? I like Kaggle and I really respect for you, but the difference is significantly small. Is this competition special?",
    "1913172": "If it had not been said enough, your dataset rocks ! Thanks you for sharing it",
    "1913493": "Thanks for sharing your hindsights & data set. If you want more features feel free to reach out to me next time :-)",
    "1930220": "Thanks for the dataset and everything that you shared, hope to see you active on Kaggle again!",
    "1923619": "I mark.good",
    "1921632": "Hi @raddar May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
    "1917423": "Thanks for the sharing of your data set, meant people with very limited time (like myself :)) could compete!\n\nI was wondering if you could share your validation scheme for your stacking? \n\nDid you reserve a hold-out test set and then use this for the multiple layer stacking or used nested CV?",
    "1915169": "Can you please elaborate on your tabnet strategy? 20-30 models of seed-averaging or any additional changes?",
    "1914681": "Thanks Raddar for your dataset sharing and lots of insights from discussion forum!",
    "1914356": "@raddar Like for everyone, your dataset saved my life. But for me it was your Days Overdue notebook that gave me the most benefit and inspiration. I mean, you even directly suggested the possible next steps at the end of the notebook. \n\nI really appreciate the work you did, and how openly you shared with the community!",
    "1914025": "Nice to see you could make TabNet part of the solution. I trained TabNet models with maximum of ~1300 features, it did not improve my ensemble or stacking. Did you use more than 1300 features?",
    "1913727": "Many thanks to you for your dataset and your insights @raddar! Congratulations to you for the result as well!",
    "1913649": "👀 I clicked your flag with no response... 👀, thanks for your awesome dataset anyway. ",
    "1913247": "Thank you very much for your notebooks, discussion, and the dataset.",
    "1913188": "Thanks for sharing! Newbie in competitions here. Do you have any suggestions on source material for these multi-level kind of models you worked on?",
    "1913206": "",
    "1917001": "Thank you for the insights. ",
    "1913839": "thank you for your impressive work!"
  }
}