{
  "id": 349789,
  "title": "260th place solution",
  "url": "/competitions/amex-default-prediction/writeups/ce-ea-260th-place-solution",
  "author_name": "",
  "post_date": "2022-09-02T17:59:51.871770800Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thank to AMEX for this competition and thank to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for his wonedful data compression notebook.<br>\nMy model is the enesemble of DART and Neural networks.</p>\n<p>Neural Network models: Tabular data(same used for Tree learners), Gru, Transformer and Stack Model(Gru+Tabular Model).</p>\n<p>Following are the scenarios where neural nets struggled compared to the tree based learners.<br>\n1 . Outlier numbers in some features.(lightgbm process it as quantiles, hence roboust)<br>\n2 . Understanding if the value is missing; even if we fix it as -1, still will be difficult to process.<br>\n3 . Dominant easy examples.(Gradients are not getting propogated to hard/varied examples as most examples are easy to classify).</p>\n<p><strong>1. Preprocessing:</strong><br>\nClip the values of the features to 95th or 99th percentile of the feature. <a href=\"https://www.kaggle.com/code/narendra/amex-tabular-nn-data-clip-map\" target=\"_blank\">Tabular Data Clip values</a>, <a href=\"https://www.kaggle.com/code/narendra/amex-sequence-data-clip-map\" target=\"_blank\">Sequential Data Clip values</a> </p>\n<p><strong>2. Add Embeddings to the Missing values</strong><br>\nFor both FFN (tabular) and sequential data(GRU, transformer), for each feature given more information that the values are missing or not. Adding this information helps the model to converge faster and better generalization.</p>\n<p>Tabular Model: <a href=\"https://www.kaggle.com/code/narendra/amex-nn-ranking-train-1024#model\" target=\"_blank\">here</a><br>\nGru Model: <a href=\"https://www.kaggle.com/code/narendra/amex-gru-with-tab-embedds-train-v3#Sequence-Model\" target=\"_blank\">here</a><br>\nStack Model: <a href=\"https://www.kaggle.com/code/narendra/amex-model-stack\" target=\"_blank\">here</a><br>\nTransformer Model: <a href=\"https://www.kaggle.com/code/narendra/amex-transformer-train-v3\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": "1924100",
      "postDate": "09/02/2022 17:59:51",
      "content": "<p>Thank to AMEX for this competition and thank to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for his wonedful data compression notebook.<br>\nMy model is the enesemble of DART and Neural networks.</p>\n<p>Neural Network models: Tabular data(same used for Tree learners), Gru, Transformer and Stack Model(Gru+Tabular Model).</p>\n<p>Following are the scenarios where neural nets struggled compared to the tree based learners.<br>\n1 . Outlier numbers in some features.(lightgbm process it as quantiles, hence roboust)<br>\n2 . Understanding if the value is missing; even if we fix it as -1, still will be difficult to process.<br>\n3 . Dominant easy examples.(Gradients are not getting propogated to hard/varied examples as most examples are easy to classify).</p>\n<p><strong>1. Preprocessing:</strong><br>\nClip the values of the features to 95th or 99th percentile of the feature. <a href=\"https://www.kaggle.com/code/narendra/amex-tabular-nn-data-clip-map\" target=\"_blank\">Tabular Data Clip values</a>, <a href=\"https://www.kaggle.com/code/narendra/amex-sequence-data-clip-map\" target=\"_blank\">Sequential Data Clip values</a> </p>\n<p><strong>2. Add Embeddings to the Missing values</strong><br>\nFor both FFN (tabular) and sequential data(GRU, transformer), for each feature given more information that the values are missing or not. Adding this information helps the model to converge faster and better generalization.</p>\n<p>Tabular Model: <a href=\"https://www.kaggle.com/code/narendra/amex-nn-ranking-train-1024#model\" target=\"_blank\">here</a><br>\nGru Model: <a href=\"https://www.kaggle.com/code/narendra/amex-gru-with-tab-embedds-train-v3#Sequence-Model\" target=\"_blank\">here</a><br>\nStack Model: <a href=\"https://www.kaggle.com/code/narendra/amex-model-stack\" target=\"_blank\">here</a><br>\nTransformer Model: <a href=\"https://www.kaggle.com/code/narendra/amex-transformer-train-v3\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Thank to AMEX for this competition and thank to @raddar for his wonedful data compression notebook.\nMy model is the enesemble of DART and Neural networks.\n\nNeural Network models: Tabular data(same used for Tree learners), Gru, Transformer and Stack Model(Gru+Tabular Model).\n\nFollowing are the scenarios where neural nets struggled compared to the tree based learners.\n1 . Outlier numbers in some features.(lightgbm process it as quantiles, hence roboust)\n2 . Understanding if the value is missing; even if we fix it as -1, still will be difficult to process.\n3 . Dominant easy examples.(Gradients are not getting propogated to hard/varied examples as most examples are easy to classify).\n\n**1. Preprocessing:**\nClip the values of the features to 95th or 99th percentile of the feature. [Tabular Data Clip values](https://www.kaggle.com/code/narendra/amex-tabular-nn-data-clip-map), [Sequential Data Clip values](https://www.kaggle.com/code/narendra/amex-sequence-data-clip-map) \n\n\n**2. Add Embeddings to the Missing values**\nFor both FFN (tabular) and sequential data(GRU, transformer), for each feature given more information that the values are missing or not. Adding this information helps the model to converge faster and better generalization.\n\nTabular Model: [here](https://www.kaggle.com/code/narendra/amex-nn-ranking-train-1024#model)\nGru Model: [here](https://www.kaggle.com/code/narendra/amex-gru-with-tab-embedds-train-v3#Sequence-Model)\nStack Model: [here](https://www.kaggle.com/code/narendra/amex-model-stack)\nTransformer Model: [here](https://www.kaggle.com/code/narendra/amex-transformer-train-v3)",
      "votes": null
    },
    {
      "id": "1929563",
      "postDate": "09/07/2022 07:10:25",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/narendra\" target=\"_blank\">@narendra</a> . You're right, the meaningful features is crucial for neural nets. So far, I still use pretrain technology to fillna or deal with outline value. </p>",
      "rawMarkdown": "Congratulations @narendra . You're right, the meaningful features is crucial for neural nets. So far, I still use pretrain technology to fillna or deal with outline value.",
      "votes": null
    },
    {
      "id": "1936753",
      "postDate": "09/13/2022 03:26:20",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/jacksonyou\" target=\"_blank\">@jacksonyou</a> </p>",
      "rawMarkdown": "thanks @jacksonyou",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1929563,
      "author_name": "jacksonyou",
      "author_url": "",
      "post_date": "09/07/2022 07:10:25",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/narendra\" target=\"_blank\">@narendra</a> . You're right, the meaningful features is crucial for neural nets. So far, I still use pretrain technology to fillna or deal with outline value. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1936753,
          "author_name": "narendra",
          "author_url": "",
          "post_date": "09/13/2022 03:26:20",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/jacksonyou\" target=\"_blank\">@jacksonyou</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1924100": "Thank to AMEX for this competition and thank to @raddar for his wonedful data compression notebook.\nMy model is the enesemble of DART and Neural networks.\n\nNeural Network models: Tabular data(same used for Tree learners), Gru, Transformer and Stack Model(Gru+Tabular Model).\n\nFollowing are the scenarios where neural nets struggled compared to the tree based learners.\n1 . Outlier numbers in some features.(lightgbm process it as quantiles, hence roboust)\n2 . Understanding if the value is missing; even if we fix it as -1, still will be difficult to process.\n3 . Dominant easy examples.(Gradients are not getting propogated to hard/varied examples as most examples are easy to classify).\n\n**1. Preprocessing:**\nClip the values of the features to 95th or 99th percentile of the feature. [Tabular Data Clip values](https://www.kaggle.com/code/narendra/amex-tabular-nn-data-clip-map), [Sequential Data Clip values](https://www.kaggle.com/code/narendra/amex-sequence-data-clip-map) \n\n\n**2. Add Embeddings to the Missing values**\nFor both FFN (tabular) and sequential data(GRU, transformer), for each feature given more information that the values are missing or not. Adding this information helps the model to converge faster and better generalization.\n\nTabular Model: [here](https://www.kaggle.com/code/narendra/amex-nn-ranking-train-1024#model)\nGru Model: [here](https://www.kaggle.com/code/narendra/amex-gru-with-tab-embedds-train-v3#Sequence-Model)\nStack Model: [here](https://www.kaggle.com/code/narendra/amex-model-stack)\nTransformer Model: [here](https://www.kaggle.com/code/narendra/amex-transformer-train-v3)",
    "1929563": "Congratulations @narendra . You're right, the meaningful features is crucial for neural nets. So far, I still use pretrain technology to fillna or deal with outline value.",
    "1936753": "thanks @jacksonyou"
  },
  "source": "meta"
}