{
  "id": 327189,
  "title": "Welcome to Amex modeling challenge",
  "url": "/competitions/amex-default-prediction/discussion/327189",
  "author_name": "Di Xu",
  "post_date": "2022-05-26T03:23:03.608000",
  "votes": 105,
  "comment_count": 77,
  "views": 0,
  "content": "<p>Hello everyone, </p>\n<p>We're excited to launch this competition and publish an industrial scale data for the community to stress test your modeling ingenuity and mine the insights we may not have discovered before. A salient feature about this dataset is the way how the records are organized across different time periods so you can explore both traditional machine learning algorithms designed for tabular data as well as emerging algorithms that can handle longitudinal data effectively.</p>\n<p>I hope all of you will enjoy this competition and look forward to the insights and learnings from all the great minds in this community.</p>\n<p>Di Xu - AI labs and AI governance, American Express</p>",
  "messages": [
    {
      "id": 1801687,
      "postDate": "2022-05-26T03:23:03.610Z",
      "content": "<p>Hello everyone, </p>\n<p>We're excited to launch this competition and publish an industrial scale data for the community to stress test your modeling ingenuity and mine the insights we may not have discovered before. A salient feature about this dataset is the way how the records are organized across different time periods so you can explore both traditional machine learning algorithms designed for tabular data as well as emerging algorithms that can handle longitudinal data effectively.</p>\n<p>I hope all of you will enjoy this competition and look forward to the insights and learnings from all the great minds in this community.</p>\n<p>Di Xu - AI labs and AI governance, American Express</p>",
      "rawMarkdown": "Hello everyone, \n\nWe're excited to launch this competition and publish an industrial scale data for the community to stress test your modeling ingenuity and mine the insights we may not have discovered before. A salient feature about this dataset is the way how the records are organized across different time periods so you can explore both traditional machine learning algorithms designed for tabular data as well as emerging algorithms that can handle longitudinal data effectively.\n\nI hope all of you will enjoy this competition and look forward to the insights and learnings from all the great minds in this community.\n\nDi Xu - AI labs and AI governance, American Express",
      "votes": 104
    },
    {
      "id": 1853157,
      "postDate": "2022-07-12T16:45:27.797Z",
      "content": "<p>Hi Amex Team,</p>\n<p>I was wondering if there is a data dictionary on what these variables mean? I know there are descriptions about the data under the data tab of the competition but I was wondering if there is more details on that. For instance P_2 is a paid variable but what does the two represent? If there is a data dictionary of sorts that would be helpful as well.</p>\n<p>Thank you</p>\n<p>-Mudassir Ali</p>",
      "rawMarkdown": "Hi Amex Team,\n\nI was wondering if there is a data dictionary on what these variables mean? I know there are descriptions about the data under the data tab of the competition but I was wondering if there is more details on that. For instance P_2 is a paid variable but what does the two represent? If there is a data dictionary of sorts that would be helpful as well.\n\nThank you\n\n-Mudassir Ali",
      "votes": 11,
      "replies": [
        {
          "id": 1868143,
          "postDate": "2022-07-23T18:26:50.127Z",
          "content": "<p>Hi Mudassir - often these competitions keep the meanings of the variables secret just to keep the work more purely data focused.  I was a bit surprised even to see the grouping of data into types, e.g. categorical.  I would not hold much hope for a real answer to your question!</p>",
          "rawMarkdown": "Hi Mudassir - often these competitions keep the meanings of the variables secret just to keep the work more purely data focused.  I was a bit surprised even to see the grouping of data into types, e.g. categorical.  I would not hold much hope for a real answer to your question!",
          "votes": 3
        }
      ]
    },
    {
      "id": 1803661,
      "postDate": "2022-05-28T04:07:08.317Z",
      "content": "<p>Hi Amex Team, Thank you for all your efforts on hosting this competition,</p>\n<p>Can you please elaborate on the following,</p>\n<blockquote>\n  <p>Note that the negative class has been subsampled for this dataset at 5%.</p>\n</blockquote>\n<p>In train dataset distribution of target 25/75, how does this relate to sub-sample rate of 5% mentioned in the data description page.</p>\n<ul>\n<li>Time period for train dataset is from March 17 to March 18, For customers who do not have all 13 statements, is it due to them being inactive or There can be customers in the training data who were acquired after March 17.</li>\n</ul>",
      "rawMarkdown": "Hi Amex Team, Thank you for all your efforts on hosting this competition,\n\nCan you please elaborate on the following,\n> Note that the negative class has been subsampled for this dataset at 5%.\n\nIn train dataset distribution of target 25/75, how does this relate to sub-sample rate of 5% mentioned in the data description page.\n\n\n-  Time period for train dataset is from March 17 to March 18, For customers who do not have all 13 statements, is it due to them being inactive or There can be customers in the training data who were acquired after March 17.",
      "votes": 8,
      "replies": [
        {
          "id": 1806207,
          "postDate": "2022-05-30T22:12:29.040Z",
          "content": "<p>My guess is that the true negative samples would be 20x more. So instead of seeing 1 part positive and 3 part negative, we would actually see 1 part positive and 60 parts negative. So in real life the proportion of positive targets is <code>0.016 = 1/61</code> instead of <code>0.25 = 1/4</code>.</p>\n<p>If the host didn't do this, then the dataset would be 1TB instead of 50GB in order to provide the same number of positive samples for our models to learn the pattern of positive samples.</p>",
          "rawMarkdown": "My guess is that the true negative samples would be 20x more. So instead of seeing 1 part positive and 3 part negative, we would actually see 1 part positive and 60 parts negative. So in real life the proportion of positive targets is `0.016 = 1/61` instead of `0.25 = 1/4`.\n\nIf the host didn't do this, then the dataset would be 1TB instead of 50GB in order to provide the same number of positive samples for our models to learn the pattern of positive samples.",
          "votes": 24
        },
        {
          "id": 1806411,
          "postDate": "2022-05-31T06:21:58.320Z",
          "content": "<p>Makes a lot of sense now, Thank you</p>",
          "rawMarkdown": "Makes a lot of sense now, Thank you"
        },
        {
          "id": 1809697,
          "postDate": "2022-06-03T00:49:58.673Z",
          "content": "<p>the downsample is only for the negative sample </p>",
          "rawMarkdown": "the downsample is only for the negative sample "
        },
        {
          "id": 1816535,
          "postDate": "2022-06-10T09:55:20.717Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for the clarification. Since our guess is that the only 5% of the negative class is taken, will it impact our final classification cause we are now disturbing the original distribution? As in other imbalanced data cases we use like SMOTE/SMOTEEN to create extra samples for the minority classes which also disturbs the original distribution. </p>\n<p>Considering the host has already done that by downsampling only the majority class in this case, will it impact the classification ? Is it advisable to maintain the same ratio of 1/61 and downsample the positive targets in the present 50GB data from 25/(25+75) i.e. 1/4 to 1/61 ratio? Please correct me if I am wrong anywhere and let me know if you require any clarifications.</p>",
          "rawMarkdown": "@cdeotte Thanks for the clarification. Since our guess is that the only 5% of the negative class is taken, will it impact our final classification cause we are now disturbing the original distribution? As in other imbalanced data cases we use like SMOTE/SMOTEEN to create extra samples for the minority classes which also disturbs the original distribution. \n\nConsidering the host has already done that by downsampling only the majority class in this case, will it impact the classification ? Is it advisable to maintain the same ratio of 1/61 and downsample the positive targets in the present 50GB data from 25/(25+75) i.e. 1/4 to 1/61 ratio? Please correct me if I am wrong anywhere and let me know if you require any clarifications."
        }
      ]
    },
    {
      "id": 1811785,
      "postDate": "2022-06-05T06:11:08.043Z",
      "content": "<p>If the challenge for this competition is an \"industrial-scale dataset\", then presumably the winning architecture is one that can stream through the <br>\ndata, like a <code>reduce()</code> function so to speak, rather than one that attempts to load the entire dataset into memory in one go.</p>",
      "rawMarkdown": "If the challenge for this competition is an \"industrial-scale dataset\", then presumably the winning architecture is one that can stream through the \ndata, like a `reduce()` function so to speak, rather than one that attempts to load the entire dataset into memory in one go.",
      "votes": 5
    },
    {
      "id": 1802220,
      "postDate": "2022-05-26T14:53:34.167Z",
      "content": "<p>Thank you for hosting this awesome Competition. My mind isn't great though I did my best to learn and explore new Data.</p>",
      "rawMarkdown": "Thank you for hosting this awesome Competition. My mind isn't great though I did my best to learn and explore new Data.",
      "votes": 5,
      "replies": [
        {
          "id": 2972210,
          "postDate": "2024-08-28T05:42:21.390Z",
          "content": "<p>Hi Marilia, I am facing some issue in loading this dataset since the size of the dataset is too large. Can you suggest some easy and simple solution to load and work on this dataset.</p>",
          "rawMarkdown": "Hi Marilia, I am facing some issue in loading this dataset since the size of the dataset is too large. Can you suggest some easy and simple solution to load and work on this dataset."
        }
      ]
    },
    {
      "id": 1809261,
      "postDate": "2022-06-02T14:57:09.787Z",
      "content": "<p>Could you elaborate on the relationship between labels and training set?<br>\nI see that they are not in the same order and that there are several duplicat customerID's.<br>\nHow are they paired?</p>",
      "rawMarkdown": "Could you elaborate on the relationship between labels and training set?\nI see that they are not in the same order and that there are several duplicat customerID's.\nHow are they paired?",
      "votes": 3,
      "replies": [
        {
          "id": 1811752,
          "postDate": "2022-06-05T05:20:08.187Z",
          "content": "<p>Duplicate ID's are due to this being \"panel\" data A.K.A. \"longitudinal\" data. Read <a href=\"https://en.wikipedia.org/wiki/Panel_data\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "Duplicate ID's are due to this being \"panel\" data A.K.A. \"longitudinal\" data. Read [here](https://en.wikipedia.org/wiki/Panel_data).",
          "votes": 3
        }
      ]
    },
    {
      "id": 1859748,
      "postDate": "2022-07-17T23:23:16.357Z",
      "content": "<p>Awesome, I can't wait to get started!</p>",
      "rawMarkdown": "Awesome, I can't wait to get started!\n",
      "votes": 1
    },
    {
      "id": 1853475,
      "postDate": "2022-07-12T22:59:06.660Z",
      "content": "<p>Thank you very much for hosting this challenging competition !</p>",
      "rawMarkdown": "Thank you very much for hosting this challenging competition !",
      "votes": 1
    },
    {
      "id": 1844767,
      "postDate": "2022-07-05T19:27:31.403Z",
      "content": "<p>Interesting.</p>",
      "rawMarkdown": "Interesting.",
      "votes": 1
    },
    {
      "id": 1842620,
      "postDate": "2022-07-04T06:42:28.800Z",
      "content": "<p>Maybe we can open more decimals?</p>",
      "rawMarkdown": "Maybe we can open more decimals?",
      "votes": 1
    },
    {
      "id": 1841393,
      "postDate": "2022-07-03T04:38:36.953Z",
      "content": "<p>Awesome competition, looking forward to putting in a solid submission. This will be my first Kaggle competition, and although I've applied ML professionally, the pure focus on performance without having to deal with other client restraints is refreshing!</p>",
      "rawMarkdown": "Awesome competition, looking forward to putting in a solid submission. This will be my first Kaggle competition, and although I've applied ML professionally, the pure focus on performance without having to deal with other client restraints is refreshing!",
      "votes": 1
    },
    {
      "id": 1838776,
      "postDate": "2022-06-30T19:43:27.430Z",
      "content": "<p>Thank you for hosting this fantastic competition ! </p>",
      "rawMarkdown": "Thank you for hosting this fantastic competition ! ",
      "votes": 1
    },
    {
      "id": 1833381,
      "postDate": "2022-06-26T00:04:05.417Z",
      "content": "<p>Can you elaborate data set having extremely large number of NAN values even core features as well</p>",
      "rawMarkdown": "Can you elaborate data set having extremely large number of NAN values even core features as well",
      "votes": 1
    },
    {
      "id": 1820036,
      "postDate": "2022-06-14T10:15:25.133Z",
      "content": "<p>Hi Amex Team, Thank you for all your efforts on hosting this competition.</p>",
      "rawMarkdown": "Hi Amex Team, Thank you for all your efforts on hosting this competition.",
      "votes": 1,
      "replies": [
        {
          "id": 1820794,
          "postDate": "2022-06-15T02:14:10.270Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1819835,
      "postDate": "2022-06-14T06:59:52.043Z",
      "content": "<p>Thank you very much for this competition</p>",
      "rawMarkdown": "Thank you very much for this competition",
      "votes": 1
    },
    {
      "id": 1818518,
      "postDate": "2022-06-12T19:19:49.720Z",
      "content": "<p>Hello, this is my first kaggle competition!  If anyone wants to collaborate please let me know!</p>",
      "rawMarkdown": "Hello, this is my first kaggle competition!  If anyone wants to collaborate please let me know!",
      "votes": 1,
      "replies": [
        {
          "id": 1857518,
          "postDate": "2022-07-16T07:07:23.103Z",
          "content": "<p>Yes I can join</p>",
          "rawMarkdown": "Yes I can join"
        }
      ]
    },
    {
      "id": 1816497,
      "postDate": "2022-06-10T09:04:26.187Z",
      "content": "<p>Thanks very much for setting up this great competition. It will be fascinating to see the insights that can be gained by the community. Outside of the winning solutions, I am sure there will be a wide range of plausible ML ideas that can be taken forward by AMEX for future model reviews. The ML analysis produced could help to act as model validation as well</p>",
      "rawMarkdown": "Thanks very much for setting up this great competition. It will be fascinating to see the insights that can be gained by the community. Outside of the winning solutions, I am sure there will be a wide range of plausible ML ideas that can be taken forward by AMEX for future model reviews. The ML analysis produced could help to act as model validation as well",
      "votes": 1
    },
    {
      "id": 1801697,
      "postDate": "2022-05-26T04:05:29.580Z",
      "content": "<p>This will be a great fun 🙌</p>",
      "rawMarkdown": "This will be a great fun 🙌",
      "votes": 1
    },
    {
      "id": 1866831,
      "postDate": "2022-07-22T19:07:42.707Z",
      "content": "<p><a href=\"https://www.kaggle.com/dixu178087\" target=\"_blank\">@dixu178087</a> I would like some clarification on the rules:</p>\n<p><code>To the extent your Submission makes use of generally commercially available software not owned by you that you used to generate your Submission, but that can be procured by the Competition Sponsor without undue expense, you do not grant the license in the preceding sentence to that software.</code></p>\n<p>I am confused what is <code>undue expense</code> here. If I were to use commercial software with free licence for individuals (same licence non-free for businesses) - my submission would comply with this rule or not?</p>\n<p>Thank you.</p>",
      "rawMarkdown": "@dixu178087 I would like some clarification on the rules:\n\n`To the extent your Submission makes use of generally commercially available software not owned by you that you used to generate your Submission, but that can be procured by the Competition Sponsor without undue expense, you do not grant the license in the preceding sentence to that software.`\n\nI am confused what is `undue expense` here. If I were to use commercial software with free licence for individuals (same licence non-free for businesses) - my submission would comply with this rule or not?\n\nThank you.\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 1890302,
          "postDate": "2022-08-08T16:45:06.707Z",
          "content": "<p>Undue expense in the context of a major bank offering $100,000 in prize money for code.</p>\n<p>So open-source is free. You may also be using some paid cloud API, or some retail software that requires a software license. This is all fine.</p>\n<p>I think <code>undue expense</code> here is really to protect against a major vendor solving it with the proprietary data-science platform, then asking millions in licencing fees on top of the prize money.</p>",
          "rawMarkdown": "Undue expense in the context of a major bank offering $100,000 in prize money for code.\n\nSo open-source is free. You may also be using some paid cloud API, or some retail software that requires a software license. This is all fine.\n\nI think `undue expense` here is really to protect against a major vendor solving it with the proprietary data-science platform, then asking millions in licencing fees on top of the prize money.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1855525,
      "postDate": "2022-07-14T16:56:23.577Z",
      "content": "<p>If the challenge for this competition is an \"industrial-scale dataset\", then presumably the winning architecture is one that can stream through the<br>\ndata, like a reduce() function so to speak, rather than one that attempts to load the entire dataset into memory in one go.</p>",
      "rawMarkdown": "If the challenge for this competition is an \"industrial-scale dataset\", then presumably the winning architecture is one that can stream through the\ndata, like a reduce() function so to speak, rather than one that attempts to load the entire dataset into memory in one go.",
      "votes": 2
    },
    {
      "id": 1822354,
      "postDate": "2022-06-16T09:30:58.347Z",
      "content": "<p>How do we interpret the date variable \"S_2\"? Which date is this exactly? e.g. payment date?<br>\nLabels are mapped to the customer and not customer date combination? So if a customer is default, do we interpret the data as default for entire year?</p>",
      "rawMarkdown": "How do we interpret the date variable \"S_2\"? Which date is this exactly? e.g. payment date?\nLabels are mapped to the customer and not customer date combination? So if a customer is default, do we interpret the data as default for entire year?",
      "votes": 2,
      "replies": [
        {
          "id": 1839949,
          "postDate": "2022-07-01T20:20:32.933Z",
          "content": "<p>S_ variables are spend variables but it appears to be the only date in each row so the assumption is that is the statement date.</p>",
          "rawMarkdown": "S_ variables are spend variables but it appears to be the only date in each row so the assumption is that is the statement date.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1808374,
      "postDate": "2022-06-01T18:37:32.847Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/dixu178087\" target=\"_blank\">@dixu178087</a> oculd we know if those are US data or if they are more global ?</p>",
      "rawMarkdown": "Hey @dixu178087 oculd we know if those are US data or if they are more global ?",
      "votes": 2
    },
    {
      "id": 1804978,
      "postDate": "2022-05-29T16:11:20.367Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dixu178087\" target=\"_blank\">@dixu178087</a> ,<br>\nAs an Ex #TeamAmex member, this is such a good initiative.<br>\nLooking forward to reading the results.</p>\n<p>Thanks</p>",
      "rawMarkdown": "Hi @dixu178087 ,\nAs an Ex #TeamAmex member, this is such a good initiative.\nLooking forward to reading the results.\n\nThanks",
      "votes": 2
    },
    {
      "id": 1810033,
      "postDate": "2022-06-03T07:50:30.103Z",
      "content": "<p>I'm an amex customer and I thought it would be fun to participate in this contest.</p>",
      "rawMarkdown": "I'm an amex customer and I thought it would be fun to participate in this contest.",
      "votes": 1
    },
    {
      "id": 1864543,
      "postDate": "2022-07-21T06:09:19.930Z",
      "content": "<p>We have started working on this. After extracting the data we have seen that the data had been given with proper manner. The count is exactly matching with the labels data. We will provide two solutions 1. With the given label data 2. We have generated our label data with our logic.</p>",
      "rawMarkdown": "We have started working on this. After extracting the data we have seen that the data had been given with proper manner. The count is exactly matching with the labels data. We will provide two solutions 1. With the given label data 2. We have generated our label data with our logic.",
      "votes": -2
    },
    {
      "id": 2972213,
      "postDate": "2024-08-28T05:42:54.917Z",
      "content": "<p>Hi everyone, I am facing some issue in loading this dataset since the size of the dataset is too large. Can you suggest some easy and simple solution to load and work on this dataset.</p>",
      "rawMarkdown": "Hi everyone, I am facing some issue in loading this dataset since the size of the dataset is too large. Can you suggest some easy and simple solution to load and work on this dataset."
    },
    {
      "id": 2447901,
      "postDate": "2023-09-20T09:55:37.733Z",
      "content": "<p>Great dataset but it would be even better if you could explain the variables for a novice starter. </p>",
      "rawMarkdown": "Great dataset but it would be even better if you could explain the variables for a novice starter. "
    },
    {
      "id": 1912508,
      "postDate": "2022-08-24T19:15:46.190Z",
      "content": "<p>Interesting</p>",
      "rawMarkdown": "Interesting"
    },
    {
      "id": 1907545,
      "postDate": "2022-08-20T21:49:49.950Z",
      "content": "<p>Thanks for your sharing,I have a clear mind after learning yours!!</p>",
      "rawMarkdown": "Thanks for your sharing,I have a clear mind after learning yours!!"
    },
    {
      "id": 1907021,
      "postDate": "2022-08-20T11:59:38.397Z",
      "content": "<p>AMAZING HOSTING</p>",
      "rawMarkdown": "AMAZING HOSTING"
    },
    {
      "id": 1900674,
      "postDate": "2022-08-16T07:27:50.920Z",
      "content": "<p>I'm looking forward to getting a good result.</p>",
      "rawMarkdown": "I'm looking forward to getting a good result."
    },
    {
      "id": 1899873,
      "postDate": "2022-08-15T15:02:58.163Z",
      "content": "<p>I'm going to do my best to finish this game, it's very exciting</p>",
      "rawMarkdown": "I'm going to do my best to finish this game, it's very exciting"
    },
    {
      "id": 1898617,
      "postDate": "2022-08-14T17:13:48.600Z",
      "content": "<p>Two questions:</p>\n<ol>\n<li>I don't see any dates in the labels dataset, does that mean that the  labels represent balances as of last date?</li>\n<li>How do you define \"default in the future\" ? Any default within the next 4 months (120 days)? </li>\n</ol>",
      "rawMarkdown": "Two questions:\n1. I don't see any dates in the labels dataset, does that mean that the  labels represent balances as of last date?\n2. How do you define \"default in the future\" ? Any default within the next 4 months (120 days)? "
    },
    {
      "id": 1896348,
      "postDate": "2022-08-12T19:07:36.633Z",
      "content": "<p>Thank you for hosting this amazing competition</p>",
      "rawMarkdown": "Thank you for hosting this amazing competition"
    },
    {
      "id": 1895618,
      "postDate": "2022-08-12T09:19:23.413Z",
      "content": "<p>What is the meaning of 'Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.'</p>",
      "rawMarkdown": "What is the meaning of 'Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.'"
    },
    {
      "id": 1885971,
      "postDate": "2022-08-05T13:19:38.043Z",
      "content": "<p>here we go！</p>",
      "rawMarkdown": "here we go！"
    },
    {
      "id": 1885221,
      "postDate": "2022-08-05T02:49:30.380Z",
      "content": "<p>i'll join now!</p>",
      "rawMarkdown": "i'll join now!"
    },
    {
      "id": 1859784,
      "postDate": "2022-07-18T01:03:08.457Z",
      "content": "<p>We will work out!</p>",
      "rawMarkdown": "We will work out!"
    },
    {
      "id": 1855941,
      "postDate": "2022-07-15T03:05:50.133Z",
      "content": "<p>It's a great competition similar to my recently job. It help a lot, thank you!</p>",
      "rawMarkdown": "It's a great competition similar to my recently job. It help a lot, thank you!"
    },
    {
      "id": 1903682,
      "postDate": "2022-08-17T15:14:17.737Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1899577,
      "postDate": "2022-08-15T11:14:46.607Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1855970,
      "postDate": "2022-07-15T03:34:28.983Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 1853726,
      "postDate": "2022-07-13T05:22:15.063Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1842954,
      "postDate": "2022-07-04T12:23:36.657Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1846522,
          "postDate": "2022-07-07T06:07:29.990Z",
          "content": "<p>From my understanding, 51% of test data is only used  to make the models generalize better. Since there is no upper cap on number of submissions, with each submission we might end up tuning the parameters to make the model perform better on test data as well. Hence to prevent the data leakage, algorithm might be picking 51% of test data points randomly and computing the scores. The entire test data evaluation might be opened up during the final evaluation</p>",
          "rawMarkdown": "From my understanding, 51% of test data is only used  to make the models generalize better. Since there is no upper cap on number of submissions, with each submission we might end up tuning the parameters to make the model perform better on test data as well. Hence to prevent the data leakage, algorithm might be picking 51% of test data points randomly and computing the scores. The entire test data evaluation might be opened up during the final evaluation\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 1835266,
      "postDate": "2022-06-27T16:05:43.133Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1821949,
      "postDate": "2022-06-16T00:40:19.197Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1822972,
          "postDate": "2022-06-16T22:23:29.053Z",
          "content": "<blockquote>\n  <p>Hi Di Xu, please can you help me with this:</p>\n<pre><code>D_* = Delinquency variables\nS_* = Spend variables\nP_* = Payment variables\nB_* = Balance variables\nR_* = Risk variables\n</code></pre>\n  <p>What does it mean '*' in each variable ?. maybe days from the latest credit card statement ? </p>\n  <p>So, D_39 is customer Delinquency at 39 days after latest credit card statement ?</p>\n  <p>Thank's</p>\n</blockquote>",
          "rawMarkdown": "> Hi Di Xu, please can you help me with this:\n> \n>     D_* = Delinquency variables\n>     S_* = Spend variables\n>     P_* = Payment variables\n>     B_* = Balance variables\n>     R_* = Risk variables\n> \n> What does it mean '*' in each variable ?. maybe days from the latest credit card statement ? \n> \n> So, D_39 is customer Delinquency at 39 days after latest credit card statement ?\n> \n> Thank's\n\n",
          "votes": 4
        }
      ]
    },
    {
      "id": 1818397,
      "postDate": "2022-06-12T16:30:35.167Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1809713,
      "postDate": "2022-06-03T01:25:12.357Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1807808,
      "postDate": "2022-06-01T10:43:35.227Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 1871781,
      "postDate": "2022-07-26T13:29:15.257Z",
      "content": "<p>Thank you for hosting!</p>",
      "rawMarkdown": "Thank you for hosting!",
      "votes": 1
    },
    {
      "id": 1912112,
      "postDate": "2022-08-24T14:13:32.570Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    },
    {
      "id": 1908340,
      "postDate": "2022-08-21T14:59:59.090Z",
      "content": "<p>Thanks for you hosting!</p>",
      "rawMarkdown": "Thanks for you hosting!"
    },
    {
      "id": 1906667,
      "postDate": "2022-08-20T05:37:34.737Z",
      "content": "<p>Thanks！ Good for starter!</p>",
      "rawMarkdown": "Thanks！ Good for starter!"
    },
    {
      "id": 1903327,
      "postDate": "2022-08-17T10:13:36.003Z",
      "content": "<p>I am a starter, thanks</p>",
      "rawMarkdown": "I am a starter, thanks"
    },
    {
      "id": 1900404,
      "postDate": "2022-08-16T02:16:31.963Z",
      "content": "<p>Thank you for hosting!!!</p>",
      "rawMarkdown": "Thank you for hosting!!!"
    },
    {
      "id": 1899162,
      "postDate": "2022-08-15T05:16:54.547Z",
      "content": "<p>Thanks for hosting</p>",
      "rawMarkdown": "Thanks for hosting"
    },
    {
      "id": 1897805,
      "postDate": "2022-08-14T05:22:29.403Z",
      "content": "<p>Thank you for hosting!</p>",
      "rawMarkdown": "Thank you for hosting!"
    },
    {
      "id": 1896894,
      "postDate": "2022-08-13T08:52:32.360Z",
      "content": "<p>Thank you for hosting!</p>",
      "rawMarkdown": "Thank you for hosting!"
    },
    {
      "id": 1896892,
      "postDate": "2022-08-13T08:52:04.357Z",
      "content": "<p>Thank you. great</p>",
      "rawMarkdown": "Thank you. great"
    },
    {
      "id": 1894950,
      "postDate": "2022-08-11T20:22:05.553Z",
      "content": "<p>great thanks so much</p>",
      "rawMarkdown": "great thanks so much"
    },
    {
      "id": 1889767,
      "postDate": "2022-08-08T11:19:35.707Z",
      "content": "<p>Thank you for hosting!</p>",
      "rawMarkdown": "Thank you for hosting!"
    },
    {
      "id": 1886038,
      "postDate": "2022-08-05T14:24:20.727Z",
      "content": "<p>Thanks for the hospitality!</p>",
      "rawMarkdown": "Thanks for the hospitality!"
    },
    {
      "id": 1874891,
      "postDate": "2022-07-28T15:07:29.110Z",
      "content": "<p>Thank you for hosting!</p>",
      "rawMarkdown": "Thank you for hosting!"
    },
    {
      "id": 1865554,
      "postDate": "2022-07-22T00:15:56.780Z",
      "content": "<p>Thank you for hosting!</p>",
      "rawMarkdown": "Thank you for hosting!"
    },
    {
      "id": 1863975,
      "postDate": "2022-07-20T16:32:52.170Z",
      "content": "<p>Thanks for Hosting!!</p>",
      "rawMarkdown": "Thanks for Hosting!!"
    }
  ],
  "comments": [
    {
      "id": 1853157,
      "author_name": "Mudassir",
      "author_url": "",
      "post_date": "2022-07-12T16:45:27.797000",
      "content": "<p>Hi Amex Team,</p>\n<p>I was wondering if there is a data dictionary on what these variables mean? I know there are descriptions about the data under the data tab of the competition but I was wondering if there is more details on that. For instance P_2 is a paid variable but what does the two represent? If there is a data dictionary of sorts that would be helpful as well.</p>\n<p>Thank you</p>\n<p>-Mudassir Ali</p>",
      "votes": 11,
      "replies": [
        {
          "id": 1868143,
          "author_name": "Bri",
          "author_url": "",
          "post_date": "2022-07-23T18:26:50.127000",
          "content": "<p>Hi Mudassir - often these competitions keep the meanings of the variables secret just to keep the work more purely data focused.  I was a bit surprised even to see the grouping of data into types, e.g. categorical.  I would not hold much hope for a real answer to your question!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1803661,
      "author_name": "Seeker",
      "author_url": "",
      "post_date": "2022-05-28T04:07:08.317000",
      "content": "<p>Hi Amex Team, Thank you for all your efforts on hosting this competition,</p>\n<p>Can you please elaborate on the following,</p>\n<blockquote>\n  <p>Note that the negative class has been subsampled for this dataset at 5%.</p>\n</blockquote>\n<p>In train dataset distribution of target 25/75, how does this relate to sub-sample rate of 5% mentioned in the data description page.</p>\n<ul>\n<li>Time period for train dataset is from March 17 to March 18, For customers who do not have all 13 statements, is it due to them being inactive or There can be customers in the training data who were acquired after March 17.</li>\n</ul>",
      "votes": 8,
      "replies": [
        {
          "id": 1806207,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-05-30T22:12:29.040000",
          "content": "<p>My guess is that the true negative samples would be 20x more. So instead of seeing 1 part positive and 3 part negative, we would actually see 1 part positive and 60 parts negative. So in real life the proportion of positive targets is <code>0.016 = 1/61</code> instead of <code>0.25 = 1/4</code>.</p>\n<p>If the host didn't do this, then the dataset would be 1TB instead of 50GB in order to provide the same number of positive samples for our models to learn the pattern of positive samples.</p>",
          "votes": 24,
          "replies": []
        },
        {
          "id": 1806411,
          "author_name": "Seeker",
          "author_url": "",
          "post_date": "2022-05-31T06:21:58.320000",
          "content": "<p>Makes a lot of sense now, Thank you</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1809697,
          "author_name": "jxlijunhao",
          "author_url": "",
          "post_date": "2022-06-03T00:49:58.673000",
          "content": "<p>the downsample is only for the negative sample </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1816535,
          "author_name": "KH_ARJUN",
          "author_url": "",
          "post_date": "2022-06-10T09:55:20.717000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for the clarification. Since our guess is that the only 5% of the negative class is taken, will it impact our final classification cause we are now disturbing the original distribution? As in other imbalanced data cases we use like SMOTE/SMOTEEN to create extra samples for the minority classes which also disturbs the original distribution. </p>\n<p>Considering the host has already done that by downsampling only the majority class in this case, will it impact the classification ? Is it advisable to maintain the same ratio of 1/61 and downsample the positive targets in the present 50GB data from 25/(25+75) i.e. 1/4 to 1/61 ratio? Please correct me if I am wrong anywhere and let me know if you require any clarifications.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1811785,
      "author_name": "James McGuigan",
      "author_url": "",
      "post_date": "2022-06-05T06:11:08.043000",
      "content": "<p>If the challenge for this competition is an \"industrial-scale dataset\", then presumably the winning architecture is one that can stream through the <br>\ndata, like a <code>reduce()</code> function so to speak, rather than one that attempts to load the entire dataset into memory in one go.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1802220,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2022-05-26T14:53:34.167000",
      "content": "<p>Thank you for hosting this awesome Competition. My mind isn't great though I did my best to learn and explore new Data.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2972210,
          "author_name": "Akashbnsl89",
          "author_url": "",
          "post_date": "2024-08-28T05:42:21.390000",
          "content": "<p>Hi Marilia, I am facing some issue in loading this dataset since the size of the dataset is too large. Can you suggest some easy and simple solution to load and work on this dataset.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1809261,
      "author_name": "Ongstad Tormod",
      "author_url": "",
      "post_date": "2022-06-02T14:57:09.787000",
      "content": "<p>Could you elaborate on the relationship between labels and training set?<br>\nI see that they are not in the same order and that there are several duplicat customerID's.<br>\nHow are they paired?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1811752,
          "author_name": "Kris Smith",
          "author_url": "",
          "post_date": "2022-06-05T05:20:08.187000",
          "content": "<p>Duplicate ID's are due to this being \"panel\" data A.K.A. \"longitudinal\" data. Read <a href=\"https://en.wikipedia.org/wiki/Panel_data\" target=\"_blank\">here</a>.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1859748,
      "author_name": "Vladimir Gritsichine",
      "author_url": "",
      "post_date": "2022-07-17T23:23:16.357000",
      "content": "<p>Awesome, I can't wait to get started!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1853475,
      "author_name": "le_petit_bois",
      "author_url": "",
      "post_date": "2022-07-12T22:59:06.660000",
      "content": "<p>Thank you very much for hosting this challenging competition !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1844767,
      "author_name": "Utkarsh Arya",
      "author_url": "",
      "post_date": "2022-07-05T19:27:31.403000",
      "content": "<p>Interesting.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1842620,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2022-07-04T06:42:28.800000",
      "content": "<p>Maybe we can open more decimals?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1841393,
      "author_name": "Xavier R Nogueira",
      "author_url": "",
      "post_date": "2022-07-03T04:38:36.953000",
      "content": "<p>Awesome competition, looking forward to putting in a solid submission. This will be my first Kaggle competition, and although I've applied ML professionally, the pure focus on performance without having to deal with other client restraints is refreshing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1838776,
      "author_name": "Hou zhangyi ",
      "author_url": "",
      "post_date": "2022-06-30T19:43:27.430000",
      "content": "<p>Thank you for hosting this fantastic competition ! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1833381,
      "author_name": "Umer Dadabhoy",
      "author_url": "",
      "post_date": "2022-06-26T00:04:05.417000",
      "content": "<p>Can you elaborate data set having extremely large number of NAN values even core features as well</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1820036,
      "author_name": "Aman Gupta",
      "author_url": "",
      "post_date": "2022-06-14T10:15:25.133000",
      "content": "<p>Hi Amex Team, Thank you for all your efforts on hosting this competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1820794,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-15T02:14:10.270000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1819835,
      "author_name": "super_bmw",
      "author_url": "",
      "post_date": "2022-06-14T06:59:52.043000",
      "content": "<p>Thank you very much for this competition</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1818518,
      "author_name": "Billy",
      "author_url": "",
      "post_date": "2022-06-12T19:19:49.720000",
      "content": "<p>Hello, this is my first kaggle competition!  If anyone wants to collaborate please let me know!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1857518,
          "author_name": "sameer shaikh",
          "author_url": "",
          "post_date": "2022-07-16T07:07:23.103000",
          "content": "<p>Yes I can join</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1816497,
      "author_name": "James McNeill",
      "author_url": "",
      "post_date": "2022-06-10T09:04:26.187000",
      "content": "<p>Thanks very much for setting up this great competition. It will be fascinating to see the insights that can be gained by the community. Outside of the winning solutions, I am sure there will be a wide range of plausible ML ideas that can be taken forward by AMEX for future model reviews. The ML analysis produced could help to act as model validation as well</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1801697,
      "author_name": "tarick.morty",
      "author_url": "",
      "post_date": "2022-05-26T04:05:29.580000",
      "content": "<p>This will be a great fun 🙌</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1866831,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2022-07-22T19:07:42.707000",
      "content": "<p><a href=\"https://www.kaggle.com/dixu178087\" target=\"_blank\">@dixu178087</a> I would like some clarification on the rules:</p>\n<p><code>To the extent your Submission makes use of generally commercially available software not owned by you that you used to generate your Submission, but that can be procured by the Competition Sponsor without undue expense, you do not grant the license in the preceding sentence to that software.</code></p>\n<p>I am confused what is <code>undue expense</code> here. If I were to use commercial software with free licence for individuals (same licence non-free for businesses) - my submission would comply with this rule or not?</p>\n<p>Thank you.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1890302,
          "author_name": "James McGuigan",
          "author_url": "",
          "post_date": "2022-08-08T16:45:06.707000",
          "content": "<p>Undue expense in the context of a major bank offering $100,000 in prize money for code.</p>\n<p>So open-source is free. You may also be using some paid cloud API, or some retail software that requires a software license. This is all fine.</p>\n<p>I think <code>undue expense</code> here is really to protect against a major vendor solving it with the proprietary data-science platform, then asking millions in licencing fees on top of the prize money.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1855525,
      "author_name": "Bilal Suppal",
      "author_url": "",
      "post_date": "2022-07-14T16:56:23.577000",
      "content": "<p>If the challenge for this competition is an \"industrial-scale dataset\", then presumably the winning architecture is one that can stream through the<br>\ndata, like a reduce() function so to speak, rather than one that attempts to load the entire dataset into memory in one go.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1822354,
      "author_name": "Kundan Kumar",
      "author_url": "",
      "post_date": "2022-06-16T09:30:58.347000",
      "content": "<p>How do we interpret the date variable \"S_2\"? Which date is this exactly? e.g. payment date?<br>\nLabels are mapped to the customer and not customer date combination? So if a customer is default, do we interpret the data as default for entire year?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1839949,
          "author_name": "Denis Perracchio",
          "author_url": "",
          "post_date": "2022-07-01T20:20:32.933000",
          "content": "<p>S_ variables are spend variables but it appears to be the only date in each row so the assumption is that is the statement date.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1808374,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2022-06-01T18:37:32.847000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/dixu178087\" target=\"_blank\">@dixu178087</a> oculd we know if those are US data or if they are more global ?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1804978,
      "author_name": "KritiDoneria",
      "author_url": "",
      "post_date": "2022-05-29T16:11:20.367000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dixu178087\" target=\"_blank\">@dixu178087</a> ,<br>\nAs an Ex #TeamAmex member, this is such a good initiative.<br>\nLooking forward to reading the results.</p>\n<p>Thanks</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1810033,
      "author_name": "AORAN SHEN",
      "author_url": "",
      "post_date": "2022-06-03T07:50:30.103000",
      "content": "<p>I'm an amex customer and I thought it would be fun to participate in this contest.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1864543,
      "author_name": "Sudip Ray",
      "author_url": "",
      "post_date": "2022-07-21T06:09:19.930000",
      "content": "<p>We have started working on this. After extracting the data we have seen that the data had been given with proper manner. The count is exactly matching with the labels data. We will provide two solutions 1. With the given label data 2. We have generated our label data with our logic.</p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 2972213,
      "author_name": "Akashbnsl89",
      "author_url": "",
      "post_date": "2024-08-28T05:42:54.917000",
      "content": "<p>Hi everyone, I am facing some issue in loading this dataset since the size of the dataset is too large. Can you suggest some easy and simple solution to load and work on this dataset.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2447901,
      "author_name": "Mikasa",
      "author_url": "",
      "post_date": "2023-09-20T09:55:37.733000",
      "content": "<p>Great dataset but it would be even better if you could explain the variables for a novice starter. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1912508,
      "author_name": "BenBlueBlue",
      "author_url": "",
      "post_date": "2022-08-24T19:15:46.190000",
      "content": "<p>Interesting</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1907545,
      "author_name": "PudgeRain",
      "author_url": "",
      "post_date": "2022-08-20T21:49:49.950000",
      "content": "<p>Thanks for your sharing,I have a clear mind after learning yours!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1907021,
      "author_name": "SiyuHuang2022",
      "author_url": "",
      "post_date": "2022-08-20T11:59:38.397000",
      "content": "<p>AMAZING HOSTING</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1900674,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-16T07:27:50.920000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1899873,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-15T15:02:58.163000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1898617,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-14T17:13:48.600000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1896348,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-12T19:07:36.633000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1895618,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-12T09:19:23.413000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1885971,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-05T13:19:38.043000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1885221,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-05T02:49:30.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1859784,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-18T01:03:08.457000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1855941,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-15T03:05:50.133000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1903682,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-17T15:14:17.737000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1899577,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-15T11:14:46.607000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1855970,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-15T03:34:28.983000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1853726,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-13T05:22:15.063000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1842954,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-04T12:23:36.657000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1846522,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-07T06:07:29.990000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1835266,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-27T16:05:43.133000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1821949,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-16T00:40:19.197000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1822972,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-16T22:23:29.053000",
          "content": "",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1818397,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-12T16:30:35.167000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1809713,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-03T01:25:12.357000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1807808,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-01T10:43:35.227000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1871781,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-26T13:29:15.257000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912112,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-24T14:13:32.570000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1908340,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-21T14:59:59.090000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1906667,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-20T05:37:34.737000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1903327,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-17T10:13:36.003000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1900404,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-16T02:16:31.963000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1899162,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-15T05:16:54.547000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1897805,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-14T05:22:29.403000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1896894,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-13T08:52:32.360000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1896892,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-13T08:52:04.357000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1894950,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-11T20:22:05.553000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1889767,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-08T11:19:35.707000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1886038,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-05T14:24:20.727000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1874891,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-28T15:07:29.110000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1865554,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-22T00:15:56.780000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1863975,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-20T16:32:52.170000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1801687": "Hello everyone, \n\nWe're excited to launch this competition and publish an industrial scale data for the community to stress test your modeling ingenuity and mine the insights we may not have discovered before. A salient feature about this dataset is the way how the records are organized across different time periods so you can explore both traditional machine learning algorithms designed for tabular data as well as emerging algorithms that can handle longitudinal data effectively.\n\nI hope all of you will enjoy this competition and look forward to the insights and learnings from all the great minds in this community.\n\nDi Xu - AI labs and AI governance, American Express",
    "1853157": "Hi Amex Team,\n\nI was wondering if there is a data dictionary on what these variables mean? I know there are descriptions about the data under the data tab of the competition but I was wondering if there is more details on that. For instance P_2 is a paid variable but what does the two represent? If there is a data dictionary of sorts that would be helpful as well.\n\nThank you\n\n-Mudassir Ali",
    "1803661": "Hi Amex Team, Thank you for all your efforts on hosting this competition,\n\nCan you please elaborate on the following,\n> Note that the negative class has been subsampled for this dataset at 5%.\n\nIn train dataset distribution of target 25/75, how does this relate to sub-sample rate of 5% mentioned in the data description page.\n\n\n-  Time period for train dataset is from March 17 to March 18, For customers who do not have all 13 statements, is it due to them being inactive or There can be customers in the training data who were acquired after March 17.",
    "1811785": "If the challenge for this competition is an \"industrial-scale dataset\", then presumably the winning architecture is one that can stream through the \ndata, like a `reduce()` function so to speak, rather than one that attempts to load the entire dataset into memory in one go.",
    "1802220": "Thank you for hosting this awesome Competition. My mind isn't great though I did my best to learn and explore new Data.",
    "1809261": "Could you elaborate on the relationship between labels and training set?\nI see that they are not in the same order and that there are several duplicat customerID's.\nHow are they paired?",
    "1859748": "Awesome, I can't wait to get started!\n",
    "1853475": "Thank you very much for hosting this challenging competition !",
    "1844767": "Interesting.",
    "1842620": "Maybe we can open more decimals?",
    "1841393": "Awesome competition, looking forward to putting in a solid submission. This will be my first Kaggle competition, and although I've applied ML professionally, the pure focus on performance without having to deal with other client restraints is refreshing!",
    "1838776": "Thank you for hosting this fantastic competition ! ",
    "1833381": "Can you elaborate data set having extremely large number of NAN values even core features as well",
    "1820036": "Hi Amex Team, Thank you for all your efforts on hosting this competition.",
    "1819835": "Thank you very much for this competition",
    "1818518": "Hello, this is my first kaggle competition!  If anyone wants to collaborate please let me know!",
    "1816497": "Thanks very much for setting up this great competition. It will be fascinating to see the insights that can be gained by the community. Outside of the winning solutions, I am sure there will be a wide range of plausible ML ideas that can be taken forward by AMEX for future model reviews. The ML analysis produced could help to act as model validation as well",
    "1801697": "This will be a great fun 🙌",
    "1866831": "@dixu178087 I would like some clarification on the rules:\n\n`To the extent your Submission makes use of generally commercially available software not owned by you that you used to generate your Submission, but that can be procured by the Competition Sponsor without undue expense, you do not grant the license in the preceding sentence to that software.`\n\nI am confused what is `undue expense` here. If I were to use commercial software with free licence for individuals (same licence non-free for businesses) - my submission would comply with this rule or not?\n\nThank you.\n\n",
    "1855525": "If the challenge for this competition is an \"industrial-scale dataset\", then presumably the winning architecture is one that can stream through the\ndata, like a reduce() function so to speak, rather than one that attempts to load the entire dataset into memory in one go.",
    "1822354": "How do we interpret the date variable \"S_2\"? Which date is this exactly? e.g. payment date?\nLabels are mapped to the customer and not customer date combination? So if a customer is default, do we interpret the data as default for entire year?",
    "1808374": "Hey @dixu178087 oculd we know if those are US data or if they are more global ?",
    "1804978": "Hi @dixu178087 ,\nAs an Ex #TeamAmex member, this is such a good initiative.\nLooking forward to reading the results.\n\nThanks",
    "1810033": "I'm an amex customer and I thought it would be fun to participate in this contest.",
    "1864543": "We have started working on this. After extracting the data we have seen that the data had been given with proper manner. The count is exactly matching with the labels data. We will provide two solutions 1. With the given label data 2. We have generated our label data with our logic.",
    "2972213": "Hi everyone, I am facing some issue in loading this dataset since the size of the dataset is too large. Can you suggest some easy and simple solution to load and work on this dataset.",
    "2447901": "Great dataset but it would be even better if you could explain the variables for a novice starter. ",
    "1912508": "Interesting",
    "1907545": "Thanks for your sharing,I have a clear mind after learning yours!!",
    "1907021": "AMAZING HOSTING",
    "1900674": "I'm looking forward to getting a good result.",
    "1899873": "I'm going to do my best to finish this game, it's very exciting",
    "1898617": "Two questions:\n1. I don't see any dates in the labels dataset, does that mean that the  labels represent balances as of last date?\n2. How do you define \"default in the future\" ? Any default within the next 4 months (120 days)? ",
    "1896348": "Thank you for hosting this amazing competition",
    "1895618": "What is the meaning of 'Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.'",
    "1885971": "here we go！",
    "1885221": "i'll join now!",
    "1859784": "We will work out!",
    "1855941": "It's a great competition similar to my recently job. It help a lot, thank you!",
    "1903682": "",
    "1899577": "",
    "1855970": "",
    "1853726": "",
    "1842954": "",
    "1835266": "",
    "1821949": "",
    "1818397": "",
    "1809713": "",
    "1807808": "",
    "1871781": "Thank you for hosting!",
    "1912112": "Thank you!",
    "1908340": "Thanks for you hosting!",
    "1906667": "Thanks！ Good for starter!",
    "1903327": "I am a starter, thanks",
    "1900404": "Thank you for hosting!!!",
    "1899162": "Thanks for hosting",
    "1897805": "Thank you for hosting!",
    "1896894": "Thank you for hosting!",
    "1896892": "Thank you. great",
    "1894950": "great thanks so much",
    "1889767": "Thank you for hosting!",
    "1886038": "Thanks for the hospitality!",
    "1874891": "Thank you for hosting!",
    "1865554": "Thank you for hosting!",
    "1863975": "Thanks for Hosting!!"
  }
}