{
  "id": 332218,
  "title": "FYI: Nvidia blog - How American Express Uses Deep Learning for Better Decision Making",
  "url": "/competitions/amex-default-prediction/discussion/332218",
  "author_name": "",
  "post_date": "2022-06-20T21:08:11.254271200Z",
  "votes": 42,
  "comment_count": 5,
  "views": 0,
  "content": "<p><img src=\"https://i.ibb.co/ySqhK5d/Selection-059.png\" alt=\"https://i.ibb.co/ySqhK5d/Selection-059.png\"></p>\n<p><a href=\"https://blogs.nvidia.com/blog/2019/12/10/american-express-deep-learning/\" target=\"_blank\">https://blogs.nvidia.com/blog/2019/12/10/american-express-deep-learning/</a><br>\n<a href=\"https://info.nvidia.com/dl-powers-better-decisions-finance-reg-page.html?ondemandrgt=yes\" target=\"_blank\">https://info.nvidia.com/dl-powers-better-decisions-finance-reg-page.html?ondemandrgt=yes</a></p>\n<p>register and watch the replay</p>\n<p>related: <a href=\"https://www.youtube.com/watch?v=fiSfB74yvlk&amp;t=2806s\" target=\"_blank\">https://www.youtube.com/watch?v=fiSfB74yvlk&amp;t=2806s</a></p>",
  "messages": [
    {
      "id": "1827129",
      "postDate": "06/20/2022 21:08:11",
      "content": "<p><img src=\"https://i.ibb.co/ySqhK5d/Selection-059.png\" alt=\"https://i.ibb.co/ySqhK5d/Selection-059.png\"></p>\n<p><a href=\"https://blogs.nvidia.com/blog/2019/12/10/american-express-deep-learning/\" target=\"_blank\">https://blogs.nvidia.com/blog/2019/12/10/american-express-deep-learning/</a><br>\n<a href=\"https://info.nvidia.com/dl-powers-better-decisions-finance-reg-page.html?ondemandrgt=yes\" target=\"_blank\">https://info.nvidia.com/dl-powers-better-decisions-finance-reg-page.html?ondemandrgt=yes</a></p>\n<p>register and watch the replay</p>\n<p>related: <a href=\"https://www.youtube.com/watch?v=fiSfB74yvlk&amp;t=2806s\" target=\"_blank\">https://www.youtube.com/watch?v=fiSfB74yvlk&amp;t=2806s</a></p>",
      "rawMarkdown": "![https://i.ibb.co/ySqhK5d/Selection-059.png](https://i.ibb.co/ySqhK5d/Selection-059.png)\n\nhttps://blogs.nvidia.com/blog/2019/12/10/american-express-deep-learning/\nhttps://info.nvidia.com/dl-powers-better-decisions-finance-reg-page.html?ondemandrgt=yes\n\nregister and watch the replay\n\nrelated: https://www.youtube.com/watch?v=fiSfB74yvlk&t=2806s",
      "votes": null
    },
    {
      "id": "1829744",
      "postDate": "06/22/2022 22:23:26",
      "content": "<p>amex paper:</p>\n<p>Sequential Deep Learning for Credit Risk Monitoring with Tabular Financial Data<br>\n<a href=\"https://arxiv.org/pdf/2012.15330.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.15330.pdf</a></p>\n<p>Using Generative Adversarial Networks to SynthesizeArtificial Financial Datasets<br>\n<a href=\"https://arxiv.org/pdf/2002.02271.pdf\" target=\"_blank\">https://arxiv.org/pdf/2002.02271.pdf</a></p>\n<hr>\n<p>An Automated System for Data Attribute Anomaly Detection<br>\n<a href=\"http://proceedings.mlr.press/v71/love18a/love18a.pdf\" target=\"_blank\">http://proceedings.mlr.press/v71/love18a/love18a.pdf</a></p>\n<p>DataQC [21] is an internal automated tool developed at American Express for data quality assessment.It allows users to evaluate similarities and differences between two provided datasets.  The toolperforms a comprehensive set of data quality tests to quickly highlight how one dataset is differentfrom another.  These tests include comparison of feature means, rates of missing values, uni- andmultivariate distributions, and extreme values. The tool produces detailed findings and quantitivescores for all the tests that it runs</p>\n<hr>\n<p>others (related):</p>\n<p><a href=\"https://www.synthesized.io/post/data-rebalancing-for-banking-and-insurance\" target=\"_blank\">https://www.synthesized.io/post/data-rebalancing-for-banking-and-insurance</a></p>",
      "rawMarkdown": "amex paper:\n\nSequential Deep Learning for Credit Risk Monitoring with Tabular Financial Data\nhttps://arxiv.org/pdf/2012.15330.pdf\n\nUsing Generative Adversarial Networks to SynthesizeArtificial Financial Datasets\nhttps://arxiv.org/pdf/2002.02271.pdf\n\n---\n\nAn Automated System for Data Attribute Anomaly Detection\nhttp://proceedings.mlr.press/v71/love18a/love18a.pdf\n\nDataQC [21] is an internal automated tool developed at American Express for data quality assessment.It allows users to evaluate similarities and differences between two provided datasets.  The toolperforms a comprehensive set of data quality tests to quickly highlight how one dataset is differentfrom another.  These tests include comparison of feature means, rates of missing values, uni- andmultivariate distributions, and extreme values. The tool produces detailed findings and quantitivescores for all the tests that it runs\n\n\n---\n\nothers (related):\n\nhttps://www.synthesized.io/post/data-rebalancing-for-banking-and-insurance",
      "votes": null
    },
    {
      "id": "1829807",
      "postDate": "06/23/2022 00:04:30",
      "content": "<p>Insightful! Thank you for sharing <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
      "rawMarkdown": "Insightful! Thank you for sharing @hengck23",
      "votes": null
    },
    {
      "id": "1830178",
      "postDate": "06/23/2022 08:45:05",
      "content": "<p>Thanks for sharing this!</p>",
      "rawMarkdown": "Thanks for sharing this!",
      "votes": null
    },
    {
      "id": "1834838",
      "postDate": "06/27/2022 08:18:50",
      "content": "<p>Looks great. Thanks for bringing all of this material together. Really interesting to see how data quality checking can be improved soo much by using deep learning. I have regularly seen many conditional searches through each feature on a time series step (e.g. one month). However, the big challenge always appeared to be reviewing the entire dataset to understand the inner workings of differences. Will have to review the articles a bit more to gain more insight.</p>",
      "rawMarkdown": "Looks great. Thanks for bringing all of this material together. Really interesting to see how data quality checking can be improved soo much by using deep learning. I have regularly seen many conditional searches through each feature on a time series step (e.g. one month). However, the big challenge always appeared to be reviewing the entire dataset to understand the inner workings of differences. Will have to review the articles a bit more to gain more insight.",
      "votes": null
    },
    {
      "id": "1836493",
      "postDate": "06/28/2022 17:35:35",
      "content": "<p>Looking at their Preprocessing step, they rely heavily on target values….. to transform the feature</p>\n<ul>\n<li>Missing values for each feature were imputed <strong>based on the target label</strong> (default/non-default) rates within 10 bins defined<br>\nby the feature’s percentiles.</li>\n<li>For dealing with numerical outliers, we developed a novel<br>\ncapping procedure to extract the most significant part of the<br>\nfeature distribution by leveraging splits obtained from <strong>training a decision tree model</strong></li>\n<li>Categorical data were transformed to numerical features<br>\nusing a procedure known as Laplace smoothing [18], which<br>\ncontains two main steps:<br>\n(1) <strong>Calculate the average of the target variable within each\ncategory.</strong></li>\n</ul>",
      "rawMarkdown": "Looking at their Preprocessing step, they rely heavily on target values..... to transform the feature\n- Missing values for each feature were imputed **based on the target label** (default/non-default) rates within 10 bins defined\nby the feature’s percentiles.\n- For dealing with numerical outliers, we developed a novel\ncapping procedure to extract the most significant part of the\nfeature distribution by leveraging splits obtained from **training a decision tree model**\n- Categorical data were transformed to numerical features\nusing a procedure known as Laplace smoothing [18], which\ncontains two main steps:\n(1) **Calculate the average of the target variable within each\ncategory.**",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1829744,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/22/2022 22:23:26",
      "content": "<p>amex paper:</p>\n<p>Sequential Deep Learning for Credit Risk Monitoring with Tabular Financial Data<br>\n<a href=\"https://arxiv.org/pdf/2012.15330.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.15330.pdf</a></p>\n<p>Using Generative Adversarial Networks to SynthesizeArtificial Financial Datasets<br>\n<a href=\"https://arxiv.org/pdf/2002.02271.pdf\" target=\"_blank\">https://arxiv.org/pdf/2002.02271.pdf</a></p>\n<hr>\n<p>An Automated System for Data Attribute Anomaly Detection<br>\n<a href=\"http://proceedings.mlr.press/v71/love18a/love18a.pdf\" target=\"_blank\">http://proceedings.mlr.press/v71/love18a/love18a.pdf</a></p>\n<p>DataQC [21] is an internal automated tool developed at American Express for data quality assessment.It allows users to evaluate similarities and differences between two provided datasets.  The toolperforms a comprehensive set of data quality tests to quickly highlight how one dataset is differentfrom another.  These tests include comparison of feature means, rates of missing values, uni- andmultivariate distributions, and extreme values. The tool produces detailed findings and quantitivescores for all the tests that it runs</p>\n<hr>\n<p>others (related):</p>\n<p><a href=\"https://www.synthesized.io/post/data-rebalancing-for-banking-and-insurance\" target=\"_blank\">https://www.synthesized.io/post/data-rebalancing-for-banking-and-insurance</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1834838,
          "author_name": "datajmcn",
          "author_url": "",
          "post_date": "06/27/2022 08:18:50",
          "content": "<p>Looks great. Thanks for bringing all of this material together. Really interesting to see how data quality checking can be improved soo much by using deep learning. I have regularly seen many conditional searches through each feature on a time series step (e.g. one month). However, the big challenge always appeared to be reviewing the entire dataset to understand the inner workings of differences. Will have to review the articles a bit more to gain more insight.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1829807,
      "author_name": "arti1117",
      "author_url": "",
      "post_date": "06/23/2022 00:04:30",
      "content": "<p>Insightful! Thank you for sharing <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1830178,
      "author_name": "angelicacassandra",
      "author_url": "",
      "post_date": "06/23/2022 08:45:05",
      "content": "<p>Thanks for sharing this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1836493,
      "author_name": "kingychiu",
      "author_url": "",
      "post_date": "06/28/2022 17:35:35",
      "content": "<p>Looking at their Preprocessing step, they rely heavily on target values….. to transform the feature</p>\n<ul>\n<li>Missing values for each feature were imputed <strong>based on the target label</strong> (default/non-default) rates within 10 bins defined<br>\nby the feature’s percentiles.</li>\n<li>For dealing with numerical outliers, we developed a novel<br>\ncapping procedure to extract the most significant part of the<br>\nfeature distribution by leveraging splits obtained from <strong>training a decision tree model</strong></li>\n<li>Categorical data were transformed to numerical features<br>\nusing a procedure known as Laplace smoothing [18], which<br>\ncontains two main steps:<br>\n(1) <strong>Calculate the average of the target variable within each\ncategory.</strong></li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1827129": "![https://i.ibb.co/ySqhK5d/Selection-059.png](https://i.ibb.co/ySqhK5d/Selection-059.png)\n\nhttps://blogs.nvidia.com/blog/2019/12/10/american-express-deep-learning/\nhttps://info.nvidia.com/dl-powers-better-decisions-finance-reg-page.html?ondemandrgt=yes\n\nregister and watch the replay\n\nrelated: https://www.youtube.com/watch?v=fiSfB74yvlk&t=2806s",
    "1829744": "amex paper:\n\nSequential Deep Learning for Credit Risk Monitoring with Tabular Financial Data\nhttps://arxiv.org/pdf/2012.15330.pdf\n\nUsing Generative Adversarial Networks to SynthesizeArtificial Financial Datasets\nhttps://arxiv.org/pdf/2002.02271.pdf\n\n---\n\nAn Automated System for Data Attribute Anomaly Detection\nhttp://proceedings.mlr.press/v71/love18a/love18a.pdf\n\nDataQC [21] is an internal automated tool developed at American Express for data quality assessment.It allows users to evaluate similarities and differences between two provided datasets.  The toolperforms a comprehensive set of data quality tests to quickly highlight how one dataset is differentfrom another.  These tests include comparison of feature means, rates of missing values, uni- andmultivariate distributions, and extreme values. The tool produces detailed findings and quantitivescores for all the tests that it runs\n\n\n---\n\nothers (related):\n\nhttps://www.synthesized.io/post/data-rebalancing-for-banking-and-insurance",
    "1829807": "Insightful! Thank you for sharing @hengck23",
    "1830178": "Thanks for sharing this!",
    "1834838": "Looks great. Thanks for bringing all of this material together. Really interesting to see how data quality checking can be improved soo much by using deep learning. I have regularly seen many conditional searches through each feature on a time series step (e.g. one month). However, the big challenge always appeared to be reviewing the entire dataset to understand the inner workings of differences. Will have to review the articles a bit more to gain more insight.",
    "1836493": "Looking at their Preprocessing step, they rely heavily on target values..... to transform the feature\n- Missing values for each feature were imputed **based on the target label** (default/non-default) rates within 10 bins defined\nby the feature’s percentiles.\n- For dealing with numerical outliers, we developed a novel\ncapping procedure to extract the most significant part of the\nfeature distribution by leveraging splits obtained from **training a decision tree model**\n- Categorical data were transformed to numerical features\nusing a procedure known as Laplace smoothing [18], which\ncontains two main steps:\n(1) **Calculate the average of the target variable within each\ncategory.**"
  },
  "source": "meta"
}