{
  "id": 331214,
  "title": "What do we know so far? - Insights from the top discussions ",
  "url": "/competitions/amex-default-prediction/discussion/331214",
  "author_name": "",
  "post_date": "2022-06-16T07:38:39.150604600Z",
  "votes": 19,
  "comment_count": 1,
  "views": 0,
  "content": "<h4>What do we know so far? - Insights from the top discussions</h4>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">Graphical explanation of the competition metric</a><br>\nby <a href=\"ambrosm\" target=\"_blank\">ambrosm</a></p>\n<ul>\n<li>Showed us some intuitive explanation of the competition's metric</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327094\" target=\"_blank\">Last month per customer</a><br>\nby <a href=\"inversion\" target=\"_blank\">inversion</a></p>\n<p>Showes us how to use pandas to just keep the last statement month per customer:</p>\n<pre><code>X_train =  (train_data\n            .groupby('customer_ID')\n            .tail(1)\n            .set_index('customer_ID', drop=True)\n            .sort_index()\n            .fillna(-999)\n            .drop(['S_2'], axis='columns'))\n</code></pre>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327759\" target=\"_blank\">Good Correlation!</a><br>\n<a href=\"librauee\" target=\"_blank\">librauee</a></p>\n<ul>\n<li>Found that LB and local CV has a good correlation,</li>\n</ul>\n<p>From <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328020\" target=\"_blank\">10x fast metric (numpy)</a> <br>\nby <a href=\"yunchonggan\" target=\"_blank\">yunchonggan</a></p>\n<ul>\n<li>Shared an optimized version of the competition's metric that is ~10x faster (it is simple numpy)</li>\n</ul>\n<p>From <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327148\" target=\"_blank\">Insights from a previous default prediction competition</a> <br>\nby <a href=\"datahobbit\" target=\"_blank\">datahobbit</a></p>\n<ul>\n<li>A list links to resources default prediction competitions</li>\n</ul>\n<p>From <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116\" target=\"_blank\">Normalized Gini Coefficient (G). Default Rate (D).</a> <br>\nby <a href=\"mpwolke\" target=\"_blank\">mpwolke</a></p>\n<ul>\n<li>Shared an intuitive explanation to the competition's metric: Gini</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327765\" target=\"_blank\">GBDT or NN,which is the winner of this competition</a><br>\nby <a href=\"senkin13\" target=\"_blank\">senkin13</a></p>\n<ul>\n<li>Debated if NN or GBM will win this competition.</li>\n<li>According to twitter: 70% Thinks: GBM</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327135\" target=\"_blank\">Articles, Research Papers and Methodologies </a><br>\nby <a href=\"ruchi798\" target=\"_blank\">ruchi798</a></p>\n<ul>\n<li>Shared Articles, Research Papers and Methodologies related to this competition</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327400\" target=\"_blank\">⚡[FAST LOADING] only 1.4GB Training Data using Feather</a><br>\nby <a href=\"seefun\" target=\"_blank\">seefun</a></p>\n<ul>\n<li>Optimized the dataset loading time even further using feather</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651\" target=\"_blank\">Strange Histograms</a><br>\nby <a href=\"cdeotte\" target=\"_blank\">cdeotte</a></p>\n<ul>\n<li>Found out that features have strange histograms: For example, at first glance, it appears that column B_8 is a binary feature that has values 0 and 1.</li>\n<li>But if we zoom in, we see that there are many values between 0 and 0.01. And many values between 1.0 and 1.01</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234\" target=\"_blank\">What does it mean?</a><br>\nby <a href=\"yashlab\" target=\"_blank\">yashlab</a></p>\n<ul>\n<li>Discussed a section from the rules saying that negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.</li>\n<li>This is because the dataset is unbalnaced</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143\" target=\"_blank\">⚡ 9x Data Compression achieved with Feather🕊️</a><br>\nby <a href=\"ruchi798\" target=\"_blank\">ruchi798</a></p>\n<ul>\n<li>Compared multiple methods for compressing the data and showed the resulting compute time</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">How To Reduce Data Size</a><br>\nby <a href=\"cdeotte\" target=\"_blank\">cdeotte</a></p>\n<ul>\n<li>Shared a long post about reducing the data size the right way</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327649\" target=\"_blank\">The data has uniform random noise injected</a><br>\nby <a href=\"raddar\" target=\"_blank\">raddar</a></p>\n<ul>\n<li>Discussing if it would be possible to de-normalize to hlep them</li>\n</ul>\n<p>From <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327142\" target=\"_blank\">Understanding the Input data using Pandas</a> <br>\nby <a href=\"balabaskar\" target=\"_blank\">balabaskar</a></p>\n<ul>\n<li>Discussed some general features of the data, such as:</li>\n<li>There were a total of 5531451 rows in the training dataset and 11363762 rows in the testing dataset.</li>\n<li>A total of 190 columns were found in the datasets.</li>\n<li>Memory of the datasets is very large, to the point that the kernel would crash trying to load the datasets.</li>\n</ul>\n<p>From <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926\" target=\"_blank\">How to identify Public and Private</a> <br>\nby <a href=\"ryotak12\" target=\"_blank\">ryotak12</a></p>\n<ul>\n<li>Talked about a possibility to identify Public and Private and did some LB probing to make sure about those results.</li>\n</ul>",
  "messages": [
    {
      "id": "1822271",
      "postDate": "06/16/2022 07:38:39",
      "content": "<h4>What do we know so far? - Insights from the top discussions</h4>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">Graphical explanation of the competition metric</a><br>\nby <a href=\"ambrosm\" target=\"_blank\">ambrosm</a></p>\n<ul>\n<li>Showed us some intuitive explanation of the competition's metric</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327094\" target=\"_blank\">Last month per customer</a><br>\nby <a href=\"inversion\" target=\"_blank\">inversion</a></p>\n<p>Showes us how to use pandas to just keep the last statement month per customer:</p>\n<pre><code>X_train =  (train_data\n            .groupby('customer_ID')\n            .tail(1)\n            .set_index('customer_ID', drop=True)\n            .sort_index()\n            .fillna(-999)\n            .drop(['S_2'], axis='columns'))\n</code></pre>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327759\" target=\"_blank\">Good Correlation!</a><br>\n<a href=\"librauee\" target=\"_blank\">librauee</a></p>\n<ul>\n<li>Found that LB and local CV has a good correlation,</li>\n</ul>\n<p>From <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328020\" target=\"_blank\">10x fast metric (numpy)</a> <br>\nby <a href=\"yunchonggan\" target=\"_blank\">yunchonggan</a></p>\n<ul>\n<li>Shared an optimized version of the competition's metric that is ~10x faster (it is simple numpy)</li>\n</ul>\n<p>From <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327148\" target=\"_blank\">Insights from a previous default prediction competition</a> <br>\nby <a href=\"datahobbit\" target=\"_blank\">datahobbit</a></p>\n<ul>\n<li>A list links to resources default prediction competitions</li>\n</ul>\n<p>From <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116\" target=\"_blank\">Normalized Gini Coefficient (G). Default Rate (D).</a> <br>\nby <a href=\"mpwolke\" target=\"_blank\">mpwolke</a></p>\n<ul>\n<li>Shared an intuitive explanation to the competition's metric: Gini</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327765\" target=\"_blank\">GBDT or NN,which is the winner of this competition</a><br>\nby <a href=\"senkin13\" target=\"_blank\">senkin13</a></p>\n<ul>\n<li>Debated if NN or GBM will win this competition.</li>\n<li>According to twitter: 70% Thinks: GBM</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327135\" target=\"_blank\">Articles, Research Papers and Methodologies </a><br>\nby <a href=\"ruchi798\" target=\"_blank\">ruchi798</a></p>\n<ul>\n<li>Shared Articles, Research Papers and Methodologies related to this competition</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327400\" target=\"_blank\">⚡[FAST LOADING] only 1.4GB Training Data using Feather</a><br>\nby <a href=\"seefun\" target=\"_blank\">seefun</a></p>\n<ul>\n<li>Optimized the dataset loading time even further using feather</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651\" target=\"_blank\">Strange Histograms</a><br>\nby <a href=\"cdeotte\" target=\"_blank\">cdeotte</a></p>\n<ul>\n<li>Found out that features have strange histograms: For example, at first glance, it appears that column B_8 is a binary feature that has values 0 and 1.</li>\n<li>But if we zoom in, we see that there are many values between 0 and 0.01. And many values between 1.0 and 1.01</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234\" target=\"_blank\">What does it mean?</a><br>\nby <a href=\"yashlab\" target=\"_blank\">yashlab</a></p>\n<ul>\n<li>Discussed a section from the rules saying that negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.</li>\n<li>This is because the dataset is unbalnaced</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143\" target=\"_blank\">⚡ 9x Data Compression achieved with Feather🕊️</a><br>\nby <a href=\"ruchi798\" target=\"_blank\">ruchi798</a></p>\n<ul>\n<li>Compared multiple methods for compressing the data and showed the resulting compute time</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">How To Reduce Data Size</a><br>\nby <a href=\"cdeotte\" target=\"_blank\">cdeotte</a></p>\n<ul>\n<li>Shared a long post about reducing the data size the right way</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327649\" target=\"_blank\">The data has uniform random noise injected</a><br>\nby <a href=\"raddar\" target=\"_blank\">raddar</a></p>\n<ul>\n<li>Discussing if it would be possible to de-normalize to hlep them</li>\n</ul>\n<p>From <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327142\" target=\"_blank\">Understanding the Input data using Pandas</a> <br>\nby <a href=\"balabaskar\" target=\"_blank\">balabaskar</a></p>\n<ul>\n<li>Discussed some general features of the data, such as:</li>\n<li>There were a total of 5531451 rows in the training dataset and 11363762 rows in the testing dataset.</li>\n<li>A total of 190 columns were found in the datasets.</li>\n<li>Memory of the datasets is very large, to the point that the kernel would crash trying to load the datasets.</li>\n</ul>\n<p>From <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926\" target=\"_blank\">How to identify Public and Private</a> <br>\nby <a href=\"ryotak12\" target=\"_blank\">ryotak12</a></p>\n<ul>\n<li>Talked about a possibility to identify Public and Private and did some LB probing to make sure about those results.</li>\n</ul>",
      "rawMarkdown": "#### What do we know so far? - Insights from the top discussions \n\n[Graphical explanation of the competition metric](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464)\nby [ambrosm](ambrosm)\n\n- Showed us some intuitive explanation of the competition's metric\n\n[Last month per customer](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327094)\nby [inversion](inversion)\n\nShowes us how to use pandas to just keep the last statement month per customer:\n\n```python\nX_train =  (train_data\n            .groupby('customer_ID')\n            .tail(1)\n            .set_index('customer_ID', drop=True)\n            .sort_index()\n            .fillna(-999)\n            .drop(['S_2'], axis='columns'))\n```\n\n[Good Correlation!](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327759)\n[librauee](librauee)\n\n- Found that LB and local CV has a good correlation,\n\nFrom [10x fast metric (numpy)](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328020) \nby [yunchonggan](yunchonggan)\n\n- Shared an optimized version of the competition's metric that is ~10x faster (it is simple numpy)\n\n\nFrom [Insights from a previous default prediction competition](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327148) \nby [datahobbit](datahobbit)\n\n- A list links to resources default prediction competitions\n\nFrom [Normalized Gini Coefficient (G). Default Rate (D).](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116) \nby [mpwolke](mpwolke)\n\n- Shared an intuitive explanation to the competition's metric: Gini\n\n\n[GBDT or NN,which is the winner of this competition](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327765)\nby [senkin13](senkin13)\n\n- Debated if NN or GBM will win this competition.\n- According to twitter: 70% Thinks: GBM\n\n[Articles, Research Papers and Methodologies ](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327135)\nby [ruchi798](ruchi798)\n\n- Shared Articles, Research Papers and Methodologies related to this competition\n\n[⚡[FAST LOADING] only 1.4GB Training Data using Feather](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327400)\nby [seefun](seefun)\n\n- Optimized the dataset loading time even further using feather\n\n[Strange Histograms](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651)\nby [cdeotte](cdeotte)\n\n- Found out that features have strange histograms: For example, at first glance, it appears that column B_8 is a binary feature that has values 0 and 1.\n- But if we zoom in, we see that there are many values between 0 and 0.01. And many values between 1.0 and 1.01\n\n\n[What does it mean?](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234)\nby [yashlab](yashlab)\n\n- Discussed a section from the rules saying that negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.\n- This is because the dataset is unbalnaced\n\n[⚡ 9x Data Compression achieved with Feather🕊️](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143)\nby [ruchi798](ruchi798)\n\n- Compared multiple methods for compressing the data and showed the resulting compute time\n\n[How To Reduce Data Size](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054)\nby [cdeotte](cdeotte)\n\n- Shared a long post about reducing the data size the right way\n\n\n[The data has uniform random noise injected](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327649)\nby [raddar](raddar)\n\n- Discussing if it would be possible to de-normalize to hlep them\n\nFrom [Understanding the Input data using Pandas](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327142) \nby [balabaskar](balabaskar)\n\n- Discussed some general features of the data, such as:\n- There were a total of 5531451 rows in the training dataset and 11363762 rows in the testing dataset.\n- A total of 190 columns were found in the datasets.\n- Memory of the datasets is very large, to the point that the kernel would crash trying to load the datasets.\n\nFrom [How to identify Public and Private](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926) \nby [ryotak12](ryotak12)\n\n- Talked about a possibility to identify Public and Private and did some LB probing to make sure about those results.",
      "votes": null
    },
    {
      "id": "1895744",
      "postDate": "08/12/2022 11:14:28",
      "content": "<p>Very useful, thanks!</p>",
      "rawMarkdown": "Very useful, thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1895744,
      "author_name": "kimberlynie",
      "author_url": "",
      "post_date": "08/12/2022 11:14:28",
      "content": "<p>Very useful, thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1822271": "#### What do we know so far? - Insights from the top discussions \n\n[Graphical explanation of the competition metric](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464)\nby [ambrosm](ambrosm)\n\n- Showed us some intuitive explanation of the competition's metric\n\n[Last month per customer](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327094)\nby [inversion](inversion)\n\nShowes us how to use pandas to just keep the last statement month per customer:\n\n```python\nX_train =  (train_data\n            .groupby('customer_ID')\n            .tail(1)\n            .set_index('customer_ID', drop=True)\n            .sort_index()\n            .fillna(-999)\n            .drop(['S_2'], axis='columns'))\n```\n\n[Good Correlation!](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327759)\n[librauee](librauee)\n\n- Found that LB and local CV has a good correlation,\n\nFrom [10x fast metric (numpy)](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328020) \nby [yunchonggan](yunchonggan)\n\n- Shared an optimized version of the competition's metric that is ~10x faster (it is simple numpy)\n\n\nFrom [Insights from a previous default prediction competition](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327148) \nby [datahobbit](datahobbit)\n\n- A list links to resources default prediction competitions\n\nFrom [Normalized Gini Coefficient (G). Default Rate (D).](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116) \nby [mpwolke](mpwolke)\n\n- Shared an intuitive explanation to the competition's metric: Gini\n\n\n[GBDT or NN,which is the winner of this competition](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327765)\nby [senkin13](senkin13)\n\n- Debated if NN or GBM will win this competition.\n- According to twitter: 70% Thinks: GBM\n\n[Articles, Research Papers and Methodologies ](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327135)\nby [ruchi798](ruchi798)\n\n- Shared Articles, Research Papers and Methodologies related to this competition\n\n[⚡[FAST LOADING] only 1.4GB Training Data using Feather](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327400)\nby [seefun](seefun)\n\n- Optimized the dataset loading time even further using feather\n\n[Strange Histograms](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651)\nby [cdeotte](cdeotte)\n\n- Found out that features have strange histograms: For example, at first glance, it appears that column B_8 is a binary feature that has values 0 and 1.\n- But if we zoom in, we see that there are many values between 0 and 0.01. And many values between 1.0 and 1.01\n\n\n[What does it mean?](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327234)\nby [yashlab](yashlab)\n\n- Discussed a section from the rules saying that negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.\n- This is because the dataset is unbalnaced\n\n[⚡ 9x Data Compression achieved with Feather🕊️](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143)\nby [ruchi798](ruchi798)\n\n- Compared multiple methods for compressing the data and showed the resulting compute time\n\n[How To Reduce Data Size](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054)\nby [cdeotte](cdeotte)\n\n- Shared a long post about reducing the data size the right way\n\n\n[The data has uniform random noise injected](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327649)\nby [raddar](raddar)\n\n- Discussing if it would be possible to de-normalize to hlep them\n\nFrom [Understanding the Input data using Pandas](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327142) \nby [balabaskar](balabaskar)\n\n- Discussed some general features of the data, such as:\n- There were a total of 5531451 rows in the training dataset and 11363762 rows in the testing dataset.\n- A total of 190 columns were found in the datasets.\n- Memory of the datasets is very large, to the point that the kernel would crash trying to load the datasets.\n\nFrom [How to identify Public and Private](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926) \nby [ryotak12](ryotak12)\n\n- Talked about a possibility to identify Public and Private and did some LB probing to make sure about those results.",
    "1895744": "Very useful, thanks!"
  },
  "source": "meta"
}