{
  "id": 209339,
  "title": "7th rank solution (Public 5th)",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/writeups/lucky-star-7th-rank-solution-public-5th",
  "author_name": "",
  "post_date": "2021-01-07T08:31:23.162744Z",
  "votes": 23,
  "comment_count": 13,
  "views": 0,
  "content": "<p>First of all, Thanks to Kaggle and INGV (Istituto nazionale di geofisica e vulcanologia) National Institute of Geophysics and Volcanology for organizing this competition.</p>\n<p>Secondly, Shout out to people who posted great notebooks which helped me and my teammate to get started. I believe lot of participants have found those notebooks useful too.</p>\n<p>Finally, Thanks to my teammate <a href=\"https://www.kaggle.com/laplaceplanet\" target=\"_blank\">@laplaceplanet</a> who provided valuable efforts in this competition.</p>\n<p><strong>SOLUTION:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F681389%2F322c10dd8f57b98111de59f939be1cfe%2Fsolution_kaggle.jpg?generation=1610007608705304&amp;alt=media\" alt=\"\"></p>\n<p>a) <strong>Feature Engineering:</strong></p>\n<p>We mostly used the public features. Credits to <a href=\"https://www.kaggle.com/amanooo\" target=\"_blank\">@amanooo</a> <a href=\"https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft\" target=\"_blank\">https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft</a> and <a href=\"https://www.kaggle.com/carpediemamigo\" target=\"_blank\">@carpediemamigo</a> <a href=\"https://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh\" target=\"_blank\">https://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh</a> (for creating 7730 features).<br>\nApart from aforementioned notebooks used Librosa <a href=\"https://librosa.org/doc/latest/index.html\" target=\"_blank\">https://librosa.org/doc/latest/index.html</a> for creating mfcc (Mel-Frequency Cepstral Coefficients) features. <br>\nEnded with ~8000 (public 7830 and 200 different features)</p>\n<p>b) <strong>Feature selection:</strong></p>\n<p>We spent a lot of time in finding the important features, as what we have seen in the past, that less features have helped in similar competitions. One important insight in this <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390\" target=\"_blank\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390</a> is </p>\n<blockquote>\n  <p>One of our best final LGB model only used four features: (i) number of peaks of at least support 2 on the denoised signal, (ii) 20% percentile on std of rolling window of size 50, (iii) 4th and (iv) 18th Mel-frequency cepstral coefficients mean. </p>\n</blockquote>\n<p>We used this insights to keep less features. For feature selection we incorporated couple of methods. <br>\n1) Used tsfresh <a href=\"https://tsfresh.readthedocs.io/en/latest/api/tsfresh.feature_selection.html\" target=\"_blank\">https://tsfresh.readthedocs.io/en/latest/api/tsfresh.feature_selection.html</a> to reduce 7730 features to 2000.<br>\n2) Used Permutation importance <a href=\"https://www.kaggle.com/dansbecker/permutation-importance\" target=\"_blank\">https://www.kaggle.com/dansbecker/permutation-importance</a> on 2300 features and selected top 189 features. </p>\n<p>c) <strong>Model Training:</strong></p>\n<p>1) Trained NN model ( we found it is better than catboost and lightgbm). Used 5 fold cross validation strategy. NN architecture is:-<br>\n`<br>\ndef get_model():<br>\ntf.keras.backend.clear_session()<br>\ninp = Input(shape=(len(features + added_features)))<br>\nx = BatchNormalization()(inp)<br>\nx = Dropout(0.1)(x)</p>\n<p>for n in [512, 1024, 2048, 3072, 2048, 1024, 512, 256]:<br>\n    x = Dense(n)(x)<br>\n    x = Activation(tf.keras.activations.swish)(x)<br>\n    x = Dropout(rate=0.1)(x)</p>\n<p>out = Dense(1, activation=\"linear\")(x)<br>\nmodel = Model(inp, out)<br>\nmodel.compile(loss=\"mean_absolute_error\", optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3))<br>\nmodel.summary()<br>\nreturn model<br>\n`<br>\ncredit to <a href=\"https://www.kaggle.com/laplaceplanet\" target=\"_blank\">@laplaceplanet</a> for NN architecture. <strong>Swish activation</strong> provided better results over other activation functions.</p>\n<p>2) Extracted last layer of NN architecture and used that as input to catboost Model. This improved lb  quite a bit.</p>\n<p>This pipeline helped us achieve 3822095 on LB.</p>\n<p><strong>Model Robustness:</strong></p>\n<p>Used the above pipeline and trained model with different seed ( 5 different seed), then average results over 5 seeds, that achieved Public LB 3730808 and  Private LB  3830808</p>",
  "messages": [
    {
      "id": "1142211",
      "postDate": "01/07/2021 08:31:23",
      "content": "<p>First of all, Thanks to Kaggle and INGV (Istituto nazionale di geofisica e vulcanologia) National Institute of Geophysics and Volcanology for organizing this competition.</p>\n<p>Secondly, Shout out to people who posted great notebooks which helped me and my teammate to get started. I believe lot of participants have found those notebooks useful too.</p>\n<p>Finally, Thanks to my teammate <a href=\"https://www.kaggle.com/laplaceplanet\" target=\"_blank\">@laplaceplanet</a> who provided valuable efforts in this competition.</p>\n<p><strong>SOLUTION:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F681389%2F322c10dd8f57b98111de59f939be1cfe%2Fsolution_kaggle.jpg?generation=1610007608705304&amp;alt=media\" alt=\"\"></p>\n<p>a) <strong>Feature Engineering:</strong></p>\n<p>We mostly used the public features. Credits to <a href=\"https://www.kaggle.com/amanooo\" target=\"_blank\">@amanooo</a> <a href=\"https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft\" target=\"_blank\">https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft</a> and <a href=\"https://www.kaggle.com/carpediemamigo\" target=\"_blank\">@carpediemamigo</a> <a href=\"https://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh\" target=\"_blank\">https://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh</a> (for creating 7730 features).<br>\nApart from aforementioned notebooks used Librosa <a href=\"https://librosa.org/doc/latest/index.html\" target=\"_blank\">https://librosa.org/doc/latest/index.html</a> for creating mfcc (Mel-Frequency Cepstral Coefficients) features. <br>\nEnded with ~8000 (public 7830 and 200 different features)</p>\n<p>b) <strong>Feature selection:</strong></p>\n<p>We spent a lot of time in finding the important features, as what we have seen in the past, that less features have helped in similar competitions. One important insight in this <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390\" target=\"_blank\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390</a> is </p>\n<blockquote>\n  <p>One of our best final LGB model only used four features: (i) number of peaks of at least support 2 on the denoised signal, (ii) 20% percentile on std of rolling window of size 50, (iii) 4th and (iv) 18th Mel-frequency cepstral coefficients mean. </p>\n</blockquote>\n<p>We used this insights to keep less features. For feature selection we incorporated couple of methods. <br>\n1) Used tsfresh <a href=\"https://tsfresh.readthedocs.io/en/latest/api/tsfresh.feature_selection.html\" target=\"_blank\">https://tsfresh.readthedocs.io/en/latest/api/tsfresh.feature_selection.html</a> to reduce 7730 features to 2000.<br>\n2) Used Permutation importance <a href=\"https://www.kaggle.com/dansbecker/permutation-importance\" target=\"_blank\">https://www.kaggle.com/dansbecker/permutation-importance</a> on 2300 features and selected top 189 features. </p>\n<p>c) <strong>Model Training:</strong></p>\n<p>1) Trained NN model ( we found it is better than catboost and lightgbm). Used 5 fold cross validation strategy. NN architecture is:-<br>\n`<br>\ndef get_model():<br>\ntf.keras.backend.clear_session()<br>\ninp = Input(shape=(len(features + added_features)))<br>\nx = BatchNormalization()(inp)<br>\nx = Dropout(0.1)(x)</p>\n<p>for n in [512, 1024, 2048, 3072, 2048, 1024, 512, 256]:<br>\n    x = Dense(n)(x)<br>\n    x = Activation(tf.keras.activations.swish)(x)<br>\n    x = Dropout(rate=0.1)(x)</p>\n<p>out = Dense(1, activation=\"linear\")(x)<br>\nmodel = Model(inp, out)<br>\nmodel.compile(loss=\"mean_absolute_error\", optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3))<br>\nmodel.summary()<br>\nreturn model<br>\n`<br>\ncredit to <a href=\"https://www.kaggle.com/laplaceplanet\" target=\"_blank\">@laplaceplanet</a> for NN architecture. <strong>Swish activation</strong> provided better results over other activation functions.</p>\n<p>2) Extracted last layer of NN architecture and used that as input to catboost Model. This improved lb  quite a bit.</p>\n<p>This pipeline helped us achieve 3822095 on LB.</p>\n<p><strong>Model Robustness:</strong></p>\n<p>Used the above pipeline and trained model with different seed ( 5 different seed), then average results over 5 seeds, that achieved Public LB 3730808 and  Private LB  3830808</p>",
      "rawMarkdown": "First of all, Thanks to Kaggle and INGV (Istituto nazionale di geofisica e vulcanologia) National Institute of Geophysics and Volcanology for organizing this competition.\n\nSecondly, Shout out to people who posted great notebooks which helped me and my teammate to get started. I believe lot of participants have found those notebooks useful too.\n\nFinally, Thanks to my teammate @laplaceplanet who provided valuable efforts in this competition.\n\n**SOLUTION:**\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F681389%2F322c10dd8f57b98111de59f939be1cfe%2Fsolution_kaggle.jpg?generation=1610007608705304&alt=media)\n\na) **Feature Engineering:**\n\nWe mostly used the public features. Credits to @amanooo https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft and @carpediemamigo https://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh (for creating 7730 features).\nApart from aforementioned notebooks used Librosa https://librosa.org/doc/latest/index.html for creating mfcc (Mel-Frequency Cepstral Coefficients) features. \nEnded with ~8000 (public 7830 and 200 different features)\n\nb) **Feature selection:**\n\nWe spent a lot of time in finding the important features, as what we have seen in the past, that less features have helped in similar competitions. One important insight in this https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390 is \n> One of our best final LGB model only used four features: (i) number of peaks of at least support 2 on the denoised signal, (ii) 20% percentile on std of rolling window of size 50, (iii) 4th and (iv) 18th Mel-frequency cepstral coefficients mean. \n\nWe used this insights to keep less features. For feature selection we incorporated couple of methods. \n1) Used tsfresh https://tsfresh.readthedocs.io/en/latest/api/tsfresh.feature_selection.html to reduce 7730 features to 2000.\n2) Used Permutation importance https://www.kaggle.com/dansbecker/permutation-importance on 2300 features and selected top 189 features. \n\nc) **Model Training:**\n\n1) Trained NN model ( we found it is better than catboost and lightgbm). Used 5 fold cross validation strategy. NN architecture is:-\n`\ndef get_model():\ntf.keras.backend.clear_session()\ninp = Input(shape=(len(features + added_features)))\nx = BatchNormalization()(inp)\nx = Dropout(0.1)(x)\n\nfor n in [512, 1024, 2048, 3072, 2048, 1024, 512, 256]:\n    x = Dense(n)(x)\n    x = Activation(tf.keras.activations.swish)(x)\n    x = Dropout(rate=0.1)(x)\n\nout = Dense(1, activation=\"linear\")(x)\nmodel = Model(inp, out)\nmodel.compile(loss=\"mean_absolute_error\", optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3))\nmodel.summary()\nreturn model\n`\ncredit to @laplaceplanet for NN architecture. **Swish activation** provided better results over other activation functions.\n\n2) Extracted last layer of NN architecture and used that as input to catboost Model. This improved lb  quite a bit.\n\nThis pipeline helped us achieve 3822095 on LB.\n\n**Model Robustness:**\n\nUsed the above pipeline and trained model with different seed ( 5 different seed), then average results over 5 seeds, that achieved Public LB 3730808 and  Private LB  3830808",
      "votes": null
    },
    {
      "id": "1142272",
      "postDate": "01/07/2021 09:13:05",
      "content": "<p>Great Job. Thanks for sharing your work. </p>",
      "rawMarkdown": "Great Job. Thanks for sharing your work.",
      "votes": null
    },
    {
      "id": "1142278",
      "postDate": "01/07/2021 09:21:52",
      "content": "<p>Good solution. Thanks for sharing</p>",
      "rawMarkdown": "Good solution. Thanks for sharing",
      "votes": null
    },
    {
      "id": "1142280",
      "postDate": "01/07/2021 09:22:25",
      "content": "<p>Great job, congrats!. Btw, how much do you think the swish activation helped?</p>",
      "rawMarkdown": "Great job, congrats!. Btw, how much do you think the swish activation helped?",
      "votes": null
    },
    {
      "id": "1142376",
      "postDate": "01/07/2021 10:43:48",
      "content": "<p>Congrats, nice solution! How did you come up with the NN architecture, i.e. the number of hidden layers and the number of neurons in the layers?</p>",
      "rawMarkdown": "Congrats, nice solution! How did you come up with the NN architecture, i.e. the number of hidden layers and the number of neurons in the layers?",
      "votes": null
    },
    {
      "id": "1142452",
      "postDate": "01/07/2021 11:48:11",
      "content": "<p>Congratulations ! Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations ! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "1142768",
      "postDate": "01/07/2021 15:26:55",
      "content": "<p>Great job, thank you! </p>",
      "rawMarkdown": "Great job, thank you!",
      "votes": null
    },
    {
      "id": "1142929",
      "postDate": "01/07/2021 17:11:41",
      "content": "<p>Great job and thanks for sharing! <br>\nWhich other activation functions you used and how was their performance?</p>",
      "rawMarkdown": "Great job and thanks for sharing! \nWhich other activation functions you used and how was their performance?",
      "votes": null
    },
    {
      "id": "1143659",
      "postDate": "01/08/2021 01:48:49",
      "content": "<p><a href=\"https://www.kaggle.com/manjjimnav\" target=\"_blank\">@manjjimnav</a>  used relu, leaky relu and tanh. Later found that people have mentioned about using swish and getting better results, so tried that.</p>",
      "rawMarkdown": "manjjimnav  used relu, leaky relu and tanh. Later found that people have mentioned about using swish and getting better results, so tried that.",
      "votes": null
    },
    {
      "id": "1143661",
      "postDate": "01/08/2021 01:50:09",
      "content": "<p><a href=\"https://www.kaggle.com/leventelippenszky\" target=\"_blank\">@leventelippenszky</a> Tried different combinations and checked local MAE loss. It is preferable to use neurons in 2^n .</p>",
      "rawMarkdown": "leventelippenszky Tried different combinations and checked local MAE loss. It is preferable to use neurons in 2^n .",
      "votes": null
    },
    {
      "id": "1143664",
      "postDate": "01/08/2021 01:51:57",
      "content": "<p><a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> , when used activations other than swash, results were worse compared to decision trees results, so i would say that swash helped a lot.</p>",
      "rawMarkdown": "enric1296 , when used activations other than swash, results were worse compared to decision trees results, so i would say that swash helped a lot.",
      "votes": null
    },
    {
      "id": "1147095",
      "postDate": "01/10/2021 09:33:43",
      "content": "<p>thx for sharing. look forward to understanding and learning from your solution when i have a moment.</p>",
      "rawMarkdown": "thx for sharing. look forward to understanding and learning from your solution when i have a moment.",
      "votes": null
    },
    {
      "id": "1260915",
      "postDate": "04/02/2021 14:23:44",
      "content": "<p>Congrats!! Just for my curiosity, how much time did it take all the work? (I'm trying to understand it and I just wanted to have an estimative time :)) ) Thank you!!</p>",
      "rawMarkdown": "Congrats!! Just for my curiosity, how much time did it take all the work? (I'm trying to understand it and I just wanted to have an estimative time :)) ) Thank you!!",
      "votes": null
    },
    {
      "id": "1262114",
      "postDate": "04/03/2021 19:50:25",
      "content": "<p>2-3 weeks should be sufficient, with 3-4 hours each day. Lot of help was taken from public kernel too</p>",
      "rawMarkdown": "2-3 weeks should be sufficient, with 3-4 hours each day. Lot of help was taken from public kernel too",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1142272,
      "author_name": "obougacha",
      "author_url": "",
      "post_date": "01/07/2021 09:13:05",
      "content": "<p>Great Job. Thanks for sharing your work. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1142278,
      "author_name": "byungsunbae",
      "author_url": "",
      "post_date": "01/07/2021 09:21:52",
      "content": "<p>Good solution. Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1142280,
      "author_name": "enric1296",
      "author_url": "",
      "post_date": "01/07/2021 09:22:25",
      "content": "<p>Great job, congrats!. Btw, how much do you think the swish activation helped?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143664,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "01/08/2021 01:51:57",
          "content": "<p><a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> , when used activations other than swash, results were worse compared to decision trees results, so i would say that swash helped a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1142376,
      "author_name": "leventelippenszky",
      "author_url": "",
      "post_date": "01/07/2021 10:43:48",
      "content": "<p>Congrats, nice solution! How did you come up with the NN architecture, i.e. the number of hidden layers and the number of neurons in the layers?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143661,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "01/08/2021 01:50:09",
          "content": "<p><a href=\"https://www.kaggle.com/leventelippenszky\" target=\"_blank\">@leventelippenszky</a> Tried different combinations and checked local MAE loss. It is preferable to use neurons in 2^n .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1142452,
      "author_name": "adaubas",
      "author_url": "",
      "post_date": "01/07/2021 11:48:11",
      "content": "<p>Congratulations ! Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1142768,
      "author_name": "sergey7",
      "author_url": "",
      "post_date": "01/07/2021 15:26:55",
      "content": "<p>Great job, thank you! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1142929,
      "author_name": "manjjimnav",
      "author_url": "",
      "post_date": "01/07/2021 17:11:41",
      "content": "<p>Great job and thanks for sharing! <br>\nWhich other activation functions you used and how was their performance?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143659,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "01/08/2021 01:48:49",
          "content": "<p><a href=\"https://www.kaggle.com/manjjimnav\" target=\"_blank\">@manjjimnav</a>  used relu, leaky relu and tanh. Later found that people have mentioned about using swish and getting better results, so tried that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1147095,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "01/10/2021 09:33:43",
      "content": "<p>thx for sharing. look forward to understanding and learning from your solution when i have a moment.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1260915,
      "author_name": "andreeatodor",
      "author_url": "",
      "post_date": "04/02/2021 14:23:44",
      "content": "<p>Congrats!! Just for my curiosity, how much time did it take all the work? (I'm trying to understand it and I just wanted to have an estimative time :)) ) Thank you!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1262114,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "04/03/2021 19:50:25",
          "content": "<p>2-3 weeks should be sufficient, with 3-4 hours each day. Lot of help was taken from public kernel too</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1142211": "First of all, Thanks to Kaggle and INGV (Istituto nazionale di geofisica e vulcanologia) National Institute of Geophysics and Volcanology for organizing this competition.\n\nSecondly, Shout out to people who posted great notebooks which helped me and my teammate to get started. I believe lot of participants have found those notebooks useful too.\n\nFinally, Thanks to my teammate @laplaceplanet who provided valuable efforts in this competition.\n\n**SOLUTION:**\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F681389%2F322c10dd8f57b98111de59f939be1cfe%2Fsolution_kaggle.jpg?generation=1610007608705304&alt=media)\n\na) **Feature Engineering:**\n\nWe mostly used the public features. Credits to @amanooo https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft and @carpediemamigo https://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh (for creating 7730 features).\nApart from aforementioned notebooks used Librosa https://librosa.org/doc/latest/index.html for creating mfcc (Mel-Frequency Cepstral Coefficients) features. \nEnded with ~8000 (public 7830 and 200 different features)\n\nb) **Feature selection:**\n\nWe spent a lot of time in finding the important features, as what we have seen in the past, that less features have helped in similar competitions. One important insight in this https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390 is \n> One of our best final LGB model only used four features: (i) number of peaks of at least support 2 on the denoised signal, (ii) 20% percentile on std of rolling window of size 50, (iii) 4th and (iv) 18th Mel-frequency cepstral coefficients mean. \n\nWe used this insights to keep less features. For feature selection we incorporated couple of methods. \n1) Used tsfresh https://tsfresh.readthedocs.io/en/latest/api/tsfresh.feature_selection.html to reduce 7730 features to 2000.\n2) Used Permutation importance https://www.kaggle.com/dansbecker/permutation-importance on 2300 features and selected top 189 features. \n\nc) **Model Training:**\n\n1) Trained NN model ( we found it is better than catboost and lightgbm). Used 5 fold cross validation strategy. NN architecture is:-\n`\ndef get_model():\ntf.keras.backend.clear_session()\ninp = Input(shape=(len(features + added_features)))\nx = BatchNormalization()(inp)\nx = Dropout(0.1)(x)\n\nfor n in [512, 1024, 2048, 3072, 2048, 1024, 512, 256]:\n    x = Dense(n)(x)\n    x = Activation(tf.keras.activations.swish)(x)\n    x = Dropout(rate=0.1)(x)\n\nout = Dense(1, activation=\"linear\")(x)\nmodel = Model(inp, out)\nmodel.compile(loss=\"mean_absolute_error\", optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3))\nmodel.summary()\nreturn model\n`\ncredit to @laplaceplanet for NN architecture. **Swish activation** provided better results over other activation functions.\n\n2) Extracted last layer of NN architecture and used that as input to catboost Model. This improved lb  quite a bit.\n\nThis pipeline helped us achieve 3822095 on LB.\n\n**Model Robustness:**\n\nUsed the above pipeline and trained model with different seed ( 5 different seed), then average results over 5 seeds, that achieved Public LB 3730808 and  Private LB  3830808",
    "1142272": "Great Job. Thanks for sharing your work.",
    "1142278": "Good solution. Thanks for sharing",
    "1142280": "Great job, congrats!. Btw, how much do you think the swish activation helped?",
    "1142376": "Congrats, nice solution! How did you come up with the NN architecture, i.e. the number of hidden layers and the number of neurons in the layers?",
    "1142452": "Congratulations ! Thanks for sharing.",
    "1142768": "Great job, thank you!",
    "1142929": "Great job and thanks for sharing! \nWhich other activation functions you used and how was their performance?",
    "1143659": "manjjimnav  used relu, leaky relu and tanh. Later found that people have mentioned about using swish and getting better results, so tried that.",
    "1143661": "leventelippenszky Tried different combinations and checked local MAE loss. It is preferable to use neurons in 2^n .",
    "1143664": "enric1296 , when used activations other than swash, results were worse compared to decision trees results, so i would say that swash helped a lot.",
    "1147095": "thx for sharing. look forward to understanding and learning from your solution when i have a moment.",
    "1260915": "Congrats!! Just for my curiosity, how much time did it take all the work? (I'm trying to understand it and I just wanted to have an estimative time :)) ) Thank you!!",
    "1262114": "2-3 weeks should be sufficient, with 3-4 hours each day. Lot of help was taken from public kernel too"
  },
  "source": "meta"
}