{
  "id": 366409,
  "title": "Private 5th Solution (A Beginner part)",
  "url": "/competitions/open-problems-multimodal/writeups/lucky-shake-private-5th-solution-a-beginner-part",
  "author_name": "",
  "post_date": "2022-11-16T05:06:56.250Z",
  "votes": 36,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>Intro</h1>\n<p>This is a scheme from a beginner. Public notebooks, the top scheme of last year, and ensemble have helped me a lot.</p>\n<h2>Citeseq</h2>\n<h3>Data preprocessing and feature engineering</h3>\n<p>I used two different methods<br>\n①Preprocessing method of public notebook from <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> . <br>\n②Using PCA to reduce dimensions to 128 + Direct features based on absolute correlation to targets. <br>\nI have normalized the row after both methods.<br>\nThe first method is more effective, and the second method is only used for ensemble. </p>\n<h3>Model</h3>\n<p>①MLP without BN and drop (Adam as the optimizer). I tried different activation functions and ensemble them, and this greatly improved CV. <br>\n②LGBM. I trained two LGBM models (different data preprocessing).</p>\n<h3>CV</h3>\n<p>①groupkfold on donor<br>\n②groupkfold on donor and day<br>\nThe first one scored higher on public LB, but the second one performed slightly better on private LB.</p>\n<h1>ensemble</h1>\n<p>Through oof prediction, I selected 4 NN models to mix with 2 LGBM models, and determined the weight. The best CV score was 0.896(donor).</p>\n<h2>Multiome</h2>\n<h3>Data preprocessing and feature engineering</h3>\n<p>I used the top method of last year and made some adjustments.<br>\nThis method has the following steps:<br>\n①tf-idf<br>\n②log1p<br>\n③sklearn.preprocessing.Normalizer(norm=\"l2\") or sklearn.preprocessing.Normalizer(norm=\"max\")<br>\n④PCA(512)<br>\n⑤row normalization<br>\n⑥Select the first 64 items in 512 and the first 100 items in 512 generated by direct dimension reduction as all features.</p>\n<h3>Model</h3>\n<p>MLP with drop(AdamW as the optimizer). I used different activation functions and ensemble them.</p>\n<h3>CV</h3>\n<p>①groupkfold on donor<br>\n②groupkfold on donor and day</p>\n<h3>ensemble</h3>\n<p>I used 4 NN and ensembled them. The best CV score was 0.670(donor).</p>\n<p>In this way, I got the submission of 0.814 public LB and 0.772 private LB. I think this is probably the easiest way to win the gold medal.</p>\n<p>Finally, I would like to thank my two teammates <a href=\"https://www.kaggle.com/jcerpentier\" target=\"_blank\">@jcerpentier</a> <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> . I have learned a lot from them. </p>\n<p>Looking forward to the next progress!</p>",
  "messages": [
    {
      "id": "2031456",
      "postDate": "11/16/2022 05:05:11",
      "content": "<h1>Intro</h1>\n<p>This is a scheme from a beginner. Public notebooks, the top scheme of last year, and ensemble have helped me a lot.</p>\n<h2>Citeseq</h2>\n<h3>Data preprocessing and feature engineering</h3>\n<p>I used two different methods<br>\n①Preprocessing method of public notebook from <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> . <br>\n②Using PCA to reduce dimensions to 128 + Direct features based on absolute correlation to targets. <br>\nI have normalized the row after both methods.<br>\nThe first method is more effective, and the second method is only used for ensemble. </p>\n<h3>Model</h3>\n<p>①MLP without BN and drop (Adam as the optimizer). I tried different activation functions and ensemble them, and this greatly improved CV. <br>\n②LGBM. I trained two LGBM models (different data preprocessing).</p>\n<h3>CV</h3>\n<p>①groupkfold on donor<br>\n②groupkfold on donor and day<br>\nThe first one scored higher on public LB, but the second one performed slightly better on private LB.</p>\n<h1>ensemble</h1>\n<p>Through oof prediction, I selected 4 NN models to mix with 2 LGBM models, and determined the weight. The best CV score was 0.896(donor).</p>\n<h2>Multiome</h2>\n<h3>Data preprocessing and feature engineering</h3>\n<p>I used the top method of last year and made some adjustments.<br>\nThis method has the following steps:<br>\n①tf-idf<br>\n②log1p<br>\n③sklearn.preprocessing.Normalizer(norm=\"l2\") or sklearn.preprocessing.Normalizer(norm=\"max\")<br>\n④PCA(512)<br>\n⑤row normalization<br>\n⑥Select the first 64 items in 512 and the first 100 items in 512 generated by direct dimension reduction as all features.</p>\n<h3>Model</h3>\n<p>MLP with drop(AdamW as the optimizer). I used different activation functions and ensemble them.</p>\n<h3>CV</h3>\n<p>①groupkfold on donor<br>\n②groupkfold on donor and day</p>\n<h3>ensemble</h3>\n<p>I used 4 NN and ensembled them. The best CV score was 0.670(donor).</p>\n<p>In this way, I got the submission of 0.814 public LB and 0.772 private LB. I think this is probably the easiest way to win the gold medal.</p>\n<p>Finally, I would like to thank my two teammates <a href=\"https://www.kaggle.com/jcerpentier\" target=\"_blank\">@jcerpentier</a> <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> . I have learned a lot from them. </p>\n<p>Looking forward to the next progress!</p>",
      "rawMarkdown": "# Intro\nThis is a scheme from a beginner. Public notebooks, the top scheme of last year, and ensemble have helped me a lot.\n## Citeseq\n###  Data preprocessing and feature engineering\nI used two different methods\n①Preprocessing method of public notebook from @pourchot . \n②Using PCA to reduce dimensions to 128 + Direct features based on absolute correlation to targets. \nI have normalized the row after both methods.\nThe first method is more effective, and the second method is only used for ensemble. \n###  Model\n①MLP without BN and drop (Adam as the optimizer). I tried different activation functions and ensemble them, and this greatly improved CV. \n②LGBM. I trained two LGBM models (different data preprocessing).\n###  CV\n①groupkfold on donor\n②groupkfold on donor and day\nThe first one scored higher on public LB, but the second one performed slightly better on private LB.\n# ensemble\nThrough oof prediction, I selected 4 NN models to mix with 2 LGBM models, and determined the weight. The best CV score was 0.896(donor).\n## Multiome\n### Data preprocessing and feature engineering\nI used the top method of last year and made some adjustments.\nThis method has the following steps:\n①tf-idf\n②log1p\n③sklearn.preprocessing.Normalizer(norm=\"l2\") or sklearn.preprocessing.Normalizer(norm=\"max\")\n④PCA(512)\n⑤row normalization\n⑥Select the first 64 items in 512 and the first 100 items in 512 generated by direct dimension reduction as all features.\n###  Model\nMLP with drop(AdamW as the optimizer). I used different activation functions and ensemble them.\n###  CV\n①groupkfold on donor\n②groupkfold on donor and day\n###  ensemble\nI used 4 NN and ensembled them. The best CV score was 0.670(donor).\n\nIn this way, I got the submission of 0.814 public LB and 0.772 private LB. I think this is probably the easiest way to win the gold medal.\n\nFinally, I would like to thank my two teammates @jcerpentier @ahmedelfazouan . I have learned a lot from them. \n\nLooking forward to the next progress!",
      "votes": null
    },
    {
      "id": "2031672",
      "postDate": "11/16/2022 07:56:09",
      "content": "<p>Thanks to the organizers and my teammates for this well-organized and interesting competition!</p>\n<p>Aside from <a href=\"https://www.kaggle.com/qqzzxxdd\" target=\"_blank\">@qqzzxxdd</a> 's great writeup, I'd like to add a brief summary of things we did that <strong>did not</strong> end up working.</p>\n<ul>\n<li><p>Initially, we were working on CV splits by donor. This resulted in our highest public solution, and was 1 of the 2 submissions that we picked as the final (1 CV, 1 LB). As some may have noticed, it is extremely challenging to push the CV score above 0.9 with single model in this context. We tried using model soups ( <a href=\"https://arxiv.org/abs/2203.05482\" target=\"_blank\">https://arxiv.org/abs/2203.05482</a> ), which did not help much. Alternatively, I manually implemented a similar method that reached 0.8997 CV score. But we noticed that CV and LB did not align as well as expected for donor.</p></li>\n<li><p>A lot of keras tuner/manual experiments for the best MLP. Turned out the best model was just always the same 4 layer MLP.</p></li>\n<li><p>Ridge ensemble (different alphas) also reached good CV results in donor context, but did not perform well for us on LB.</p></li>\n</ul>\n<p>Lastly, I think it is mandatory to thank <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> , <a href=\"https://www.kaggle.com/jsmithperera\" target=\"_blank\">@jsmithperera</a> and <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a> (sorry if I forgot any) for their great notebooks! They were very helpful to us, and probably many others, throughout this competition.</p>",
      "rawMarkdown": "Thanks to the organizers and my teammates for this well-organized and interesting competition!\n\nAside from @qqzzxxdd 's great writeup, I'd like to add a brief summary of things we did that **did not** end up working.\n\n- Initially, we were working on CV splits by donor. This resulted in our highest public solution, and was 1 of the 2 submissions that we picked as the final (1 CV, 1 LB). As some may have noticed, it is extremely challenging to push the CV score above 0.9 with single model in this context. We tried using model soups ( https://arxiv.org/abs/2203.05482 ), which did not help much. Alternatively, I manually implemented a similar method that reached 0.8997 CV score. But we noticed that CV and LB did not align as well as expected for donor.\n\n- A lot of keras tuner/manual experiments for the best MLP. Turned out the best model was just always the same 4 layer MLP.\n\n- Ridge ensemble (different alphas) also reached good CV results in donor context, but did not perform well for us on LB.\n\nLastly, I think it is mandatory to thank @pourchot , @jsmithperera and @vslaykovsky (sorry if I forgot any) for their great notebooks! They were very helpful to us, and probably many others, throughout this competition.",
      "votes": null
    },
    {
      "id": "2031764",
      "postDate": "11/16/2022 08:35:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jcerpentier\" target=\"_blank\">@jcerpentier</a> , Congrats 🎉🍾 ! </p>",
      "rawMarkdown": "Hi @jcerpentier , Congrats 🎉🍾 !",
      "votes": null
    },
    {
      "id": "2031782",
      "postDate": "11/16/2022 08:40:01",
      "content": "<p>Congrats the team ! You were better than some grand masters and it is a big challenge 😉</p>",
      "rawMarkdown": "Congrats the team ! You were better than some grand masters and it is a big challenge 😉",
      "votes": null
    },
    {
      "id": "2032716",
      "postDate": "11/16/2022 19:40:04",
      "content": "<p>Very useful for  beginners like me thanks <a href=\"https://www.kaggle.com/qqzzxxdd\" target=\"_blank\">@qqzzxxdd</a> for sharing</p>",
      "rawMarkdown": "Very useful for  beginners like me thanks @qqzzxxdd for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2031672,
      "author_name": "jcerpentier",
      "author_url": "",
      "post_date": "11/16/2022 07:56:09",
      "content": "<p>Thanks to the organizers and my teammates for this well-organized and interesting competition!</p>\n<p>Aside from <a href=\"https://www.kaggle.com/qqzzxxdd\" target=\"_blank\">@qqzzxxdd</a> 's great writeup, I'd like to add a brief summary of things we did that <strong>did not</strong> end up working.</p>\n<ul>\n<li><p>Initially, we were working on CV splits by donor. This resulted in our highest public solution, and was 1 of the 2 submissions that we picked as the final (1 CV, 1 LB). As some may have noticed, it is extremely challenging to push the CV score above 0.9 with single model in this context. We tried using model soups ( <a href=\"https://arxiv.org/abs/2203.05482\" target=\"_blank\">https://arxiv.org/abs/2203.05482</a> ), which did not help much. Alternatively, I manually implemented a similar method that reached 0.8997 CV score. But we noticed that CV and LB did not align as well as expected for donor.</p></li>\n<li><p>A lot of keras tuner/manual experiments for the best MLP. Turned out the best model was just always the same 4 layer MLP.</p></li>\n<li><p>Ridge ensemble (different alphas) also reached good CV results in donor context, but did not perform well for us on LB.</p></li>\n</ul>\n<p>Lastly, I think it is mandatory to thank <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> , <a href=\"https://www.kaggle.com/jsmithperera\" target=\"_blank\">@jsmithperera</a> and <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a> (sorry if I forgot any) for their great notebooks! They were very helpful to us, and probably many others, throughout this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2031764,
          "author_name": "pourchot",
          "author_url": "",
          "post_date": "11/16/2022 08:35:26",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jcerpentier\" target=\"_blank\">@jcerpentier</a> , Congrats 🎉🍾 ! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2031782,
          "author_name": "pourchot",
          "author_url": "",
          "post_date": "11/16/2022 08:40:01",
          "content": "<p>Congrats the team ! You were better than some grand masters and it is a big challenge 😉</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2032716,
      "author_name": "shreyamishra0307",
      "author_url": "",
      "post_date": "11/16/2022 19:40:04",
      "content": "<p>Very useful for  beginners like me thanks <a href=\"https://www.kaggle.com/qqzzxxdd\" target=\"_blank\">@qqzzxxdd</a> for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2031456": "# Intro\nThis is a scheme from a beginner. Public notebooks, the top scheme of last year, and ensemble have helped me a lot.\n## Citeseq\n###  Data preprocessing and feature engineering\nI used two different methods\n①Preprocessing method of public notebook from @pourchot . \n②Using PCA to reduce dimensions to 128 + Direct features based on absolute correlation to targets. \nI have normalized the row after both methods.\nThe first method is more effective, and the second method is only used for ensemble. \n###  Model\n①MLP without BN and drop (Adam as the optimizer). I tried different activation functions and ensemble them, and this greatly improved CV. \n②LGBM. I trained two LGBM models (different data preprocessing).\n###  CV\n①groupkfold on donor\n②groupkfold on donor and day\nThe first one scored higher on public LB, but the second one performed slightly better on private LB.\n# ensemble\nThrough oof prediction, I selected 4 NN models to mix with 2 LGBM models, and determined the weight. The best CV score was 0.896(donor).\n## Multiome\n### Data preprocessing and feature engineering\nI used the top method of last year and made some adjustments.\nThis method has the following steps:\n①tf-idf\n②log1p\n③sklearn.preprocessing.Normalizer(norm=\"l2\") or sklearn.preprocessing.Normalizer(norm=\"max\")\n④PCA(512)\n⑤row normalization\n⑥Select the first 64 items in 512 and the first 100 items in 512 generated by direct dimension reduction as all features.\n###  Model\nMLP with drop(AdamW as the optimizer). I used different activation functions and ensemble them.\n###  CV\n①groupkfold on donor\n②groupkfold on donor and day\n###  ensemble\nI used 4 NN and ensembled them. The best CV score was 0.670(donor).\n\nIn this way, I got the submission of 0.814 public LB and 0.772 private LB. I think this is probably the easiest way to win the gold medal.\n\nFinally, I would like to thank my two teammates @jcerpentier @ahmedelfazouan . I have learned a lot from them. \n\nLooking forward to the next progress!",
    "2031672": "Thanks to the organizers and my teammates for this well-organized and interesting competition!\n\nAside from @qqzzxxdd 's great writeup, I'd like to add a brief summary of things we did that **did not** end up working.\n\n- Initially, we were working on CV splits by donor. This resulted in our highest public solution, and was 1 of the 2 submissions that we picked as the final (1 CV, 1 LB). As some may have noticed, it is extremely challenging to push the CV score above 0.9 with single model in this context. We tried using model soups ( https://arxiv.org/abs/2203.05482 ), which did not help much. Alternatively, I manually implemented a similar method that reached 0.8997 CV score. But we noticed that CV and LB did not align as well as expected for donor.\n\n- A lot of keras tuner/manual experiments for the best MLP. Turned out the best model was just always the same 4 layer MLP.\n\n- Ridge ensemble (different alphas) also reached good CV results in donor context, but did not perform well for us on LB.\n\nLastly, I think it is mandatory to thank @pourchot , @jsmithperera and @vslaykovsky (sorry if I forgot any) for their great notebooks! They were very helpful to us, and probably many others, throughout this competition.",
    "2031764": "Hi @jcerpentier , Congrats 🎉🍾 !",
    "2031782": "Congrats the team ! You were better than some grand masters and it is a big challenge 😉",
    "2032716": "Very useful for  beginners like me thanks @qqzzxxdd for sharing"
  },
  "source": "meta"
}