{
  "id": 326904,
  "title": "New to Kaggle or Machine Learning? Come Say Hi!",
  "url": "/competitions/amex-default-prediction/discussion/326904",
  "author_name": "Addison Howard",
  "post_date": "2022-05-24T21:15:26.790000",
  "votes": 51,
  "comment_count": 52,
  "views": 0,
  "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! </p>\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n<p>New to Kaggle? Take a look at a few videos Dr. Rachael Tatman has put together to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\" target=\"_blank\">site etiquette</a>, or <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\" target=\"_blank\">Kaggle lingo</a>.</p>\n<p><strong>Remember</strong>: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our <a href=\"https://www.kaggle.com/community-guidelines\" target=\"_blank\">Kaggle community guidelines</a>.</p>\n<p>Happy Modeling!</p>",
  "messages": [
    {
      "id": 1800285,
      "postDate": "2022-05-24T21:15:26.790Z",
      "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! </p>\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n<p>New to Kaggle? Take a look at a few videos Dr. Rachael Tatman has put together to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\" target=\"_blank\">site etiquette</a>, or <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\" target=\"_blank\">Kaggle lingo</a>.</p>\n<p><strong>Remember</strong>: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our <a href=\"https://www.kaggle.com/community-guidelines\" target=\"_blank\">Kaggle community guidelines</a>.</p>\n<p>Happy Modeling!</p>",
      "rawMarkdown": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! \n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos Dr. Rachael Tatman has put together to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), or [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s).\n\n**Remember**: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our [Kaggle community guidelines](https://www.kaggle.com/community-guidelines).\n\nHappy Modeling!",
      "votes": 50
    },
    {
      "id": 1832500,
      "postDate": "2022-06-25T05:27:52.430Z",
      "content": "<p>Hi! <br>\nI'm a begginer of Kaggle, but want to attend aggressively for my career making!</p>\n<p>I want to ask the way to deal with such the big data.<br>\nOnly downloading take long time…<br>\nHow to deal with this problem? </p>",
      "rawMarkdown": "Hi! \nI'm a begginer of Kaggle, but want to attend aggressively for my career making!\n\nI want to ask the way to deal with such the big data.\nOnly downloading take long time...\nHow to deal with this problem? ",
      "votes": 3,
      "replies": [
        {
          "id": 1832774,
          "postDate": "2022-06-25T11:35:24.733Z",
          "content": "<p>I also have the same problem. I have only 4GB RAM space. </p>",
          "rawMarkdown": "I also have the same problem. I have only 4GB RAM space. "
        },
        {
          "id": 1833754,
          "postDate": "2022-06-26T09:55:54.507Z",
          "content": "<p>One of the ways to deal with such a big amounts of data is using parquet data sets. Parquet data set is a compressed version of .csv.  You can find  notebooks in the Internet that already contain parquet datasets (it's size is about 5-10 times less than size of original datasets). You can work with parquet datasets this way:</p>\n<blockquote>\n  <p>import pandas as pd<br>\n  import numpy as np<br>\n  import gc<br>\n  import matplotlib.pyplot as plt<br>\n  import seaborn as sns<br>\n  train = pd.read_parquet(r'C:\\Users\\aljutov\\Desktop\\American_Express_competition\\train_data.parquet\\train_data.parquet')<br>\n  test = pd.read_parquet(r'C:\\Users\\aljutov\\Desktop\\American_Express_competition\\test_data.parquet\\test_data.parquet')</p>\n</blockquote>\n<p>after that begin EDA as usual…</p>",
          "rawMarkdown": "One of the ways to deal with such a big amounts of data is using parquet data sets. Parquet data set is a compressed version of .csv.  You can find  notebooks in the Internet that already contain parquet datasets (it's size is about 5-10 times less than size of original datasets). You can work with parquet datasets this way:\n\n> \nimport pandas as pd\nimport numpy as np\nimport gc\nimport matplotlib.pyplot as plt\nimport seaborn as sns\ntrain = pd.read_parquet(r'C:\\Users\\aljutov\\Desktop\\American_Express_competition\\train_data.parquet\\train_data.parquet')\ntest = pd.read_parquet(r'C:\\Users\\aljutov\\Desktop\\American_Express_competition\\test_data.parquet\\test_data.parquet')\n\nafter that begin EDA as usual...\n",
          "votes": 5
        },
        {
          "id": 1882782,
          "postDate": "2022-08-03T12:27:48.457Z",
          "content": "<p>Hi How did you create parquet files please can you share entire procedure</p>",
          "rawMarkdown": "Hi How did you create parquet files please can you share entire procedure",
          "votes": 1
        },
        {
          "id": 1897711,
          "postDate": "2022-08-14T03:09:48.320Z",
          "content": "<p>Do you know how to create parquet datasets now?</p>",
          "rawMarkdown": "Do you know how to create parquet datasets now?"
        }
      ]
    },
    {
      "id": 1888828,
      "postDate": "2022-08-07T19:47:09.917Z",
      "content": "<p>Hello Guys, I am still fairly new to machine learning and I am joining this competition not necessarily to win (although I want to) but to progress in my career as a successful Data Scientist. I am looking to make lifelong friends in this community with whom I can learn and add value to. So please, feel free to reach out to me. My LinkedIn is embedded in my profile. Wishing everyone success and good luck in this competition.</p>",
      "rawMarkdown": "Hello Guys, I am still fairly new to machine learning and I am joining this competition not necessarily to win (although I want to) but to progress in my career as a successful Data Scientist. I am looking to make lifelong friends in this community with whom I can learn and add value to. So please, feel free to reach out to me. My LinkedIn is embedded in my profile. Wishing everyone success and good luck in this competition.",
      "votes": 1,
      "replies": [
        {
          "id": 1895045,
          "postDate": "2022-08-11T22:48:19.963Z",
          "content": "<p>Welcome Olusola!</p>",
          "rawMarkdown": "Welcome Olusola!"
        }
      ]
    },
    {
      "id": 1850019,
      "postDate": "2022-07-10T02:29:53.950Z",
      "content": "<p>I want to improve my data analysis skill.</p>",
      "rawMarkdown": "I want to improve my data analysis skill.",
      "votes": 1
    },
    {
      "id": 1818947,
      "postDate": "2022-06-13T10:03:49.680Z",
      "content": "<p>Hey! I'm know some theory in DS and want to start practice! But, I faced with memory problem (for this dataset -- Store Sales - Time Series Forecasting, 124.76 MiB), in this competition 50.31 GiB of data - how people can to do something with that batch? Сan I do something without adding video cards, etc. only with kaggle environment power?</p>",
      "rawMarkdown": "Hey! I'm know some theory in DS and want to start practice! But, I faced with memory problem (for this dataset -- Store Sales - Time Series Forecasting, 124.76 MiB), in this competition 50.31 GiB of data - how people can to do something with that batch? Сan I do something without adding video cards, etc. only with kaggle environment power?",
      "votes": 1,
      "replies": [
        {
          "id": 1820356,
          "postDate": "2022-06-14T14:59:42.467Z",
          "content": "<p>Hi, </p>\n<p>Try to use parametrs <em>iterator=True, chunksize=1000000</em> when reading data with pd.read_csv</p>",
          "rawMarkdown": "Hi, \n\nTry to use parametrs *iterator=True, chunksize=1000000* when reading data with pd.read_csv"
        },
        {
          "id": 1821522,
          "postDate": "2022-06-15T16:13:45.630Z",
          "content": "<p>Check below for more understanding:</p>\n<ul>\n<li><a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/abhianalytic/amex-read-full-data-without-any-transformation</a></li>\n<li><a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328565</a></li>\n</ul>",
          "rawMarkdown": "Check below for more understanding:\n- [https://www.kaggle.com/code/abhianalytic/amex-read-full-data-without-any-transformation](url)\n- [https://www.kaggle.com/competitions/amex-default-prediction/discussion/328565](url)\n"
        }
      ]
    },
    {
      "id": 1853297,
      "postDate": "2022-07-12T19:36:41.447Z",
      "content": "<p>Hi! This is my first Kaggle competition, however I have worked on a few ML projects in my job as an environmental data scientist ( I would say I am probably at a low-intermediate skill level.). That said, in environmental contexts we often use geospatial data which due to it's large memory footprint and complexity often limits my ability to go \"all out\" on predictions (also our clients like to understand things simply, so too much feature engineering is avoided). I'm looking forward to being able to do exactly what I want with the data without worrying about clients or other more mundane things.</p>\n<p>To start, since I am running my analysis on a simple laptop, I decided to use dask for the entire workflow. To start, imported the train/test data, formatted their dtypes, and re-saved to partitioned .parquet files. This takes 12 minutes for all the provided data! If you want to do the same thing to kick start your work flow check out my notebook here:  <a href=\"https://www.kaggle.com/code/xaviernogueira/dask-pre-processing-csv-parquet-in-parallel\" target=\"_blank\">https://www.kaggle.com/code/xaviernogueira/dask-pre-processing-csv-parquet-in-parallel</a>. I would appreciate an upvote if you find it useful :) </p>",
      "rawMarkdown": "Hi! This is my first Kaggle competition, however I have worked on a few ML projects in my job as an environmental data scientist ( I would say I am probably at a low-intermediate skill level.). That said, in environmental contexts we often use geospatial data which due to it's large memory footprint and complexity often limits my ability to go \"all out\" on predictions (also our clients like to understand things simply, so too much feature engineering is avoided). I'm looking forward to being able to do exactly what I want with the data without worrying about clients or other more mundane things.\n\nTo start, since I am running my analysis on a simple laptop, I decided to use dask for the entire workflow. To start, imported the train/test data, formatted their dtypes, and re-saved to partitioned .parquet files. This takes 12 minutes for all the provided data! If you want to do the same thing to kick start your work flow check out my notebook here:  https://www.kaggle.com/code/xaviernogueira/dask-pre-processing-csv-parquet-in-parallel. I would appreciate an upvote if you find it useful :) ",
      "votes": 2
    },
    {
      "id": 1911530,
      "postDate": "2022-08-24T06:49:53.517Z",
      "content": "<p>Hello Everyone! I am new to Kaggle and currently unaware of how exactly the private leaderboard works. My doubt is:</p>\n<p>According to the competition rules, the public leaderboard is decided based on 51% of test data. And private leaderboard will be decided based on the other 49%. So that means the test set that we have used in our codes is only 51% of the actual test set, and after the competition deadline, Kaggle uses the other 41% (which is not currently provided to us), right? </p>\n<p>So how exactly does Kaggle run our code with a different test set? I am worried about this because I have not used the dataset provided by AmEx (as the file size is large). Instead, I have used <a href=\"https://www.kaggle.com/datasets/munumbutt/amexfeather\" target=\"_blank\">this</a> feather dataset. Now in my case, how will Kaggle run my code for the private leaderboard?</p>",
      "rawMarkdown": "Hello Everyone! I am new to Kaggle and currently unaware of how exactly the private leaderboard works. My doubt is:\n\nAccording to the competition rules, the public leaderboard is decided based on 51% of test data. And private leaderboard will be decided based on the other 49%. So that means the test set that we have used in our codes is only 51% of the actual test set, and after the competition deadline, Kaggle uses the other 41% (which is not currently provided to us), right? \n\nSo how exactly does Kaggle run our code with a different test set? I am worried about this because I have not used the dataset provided by AmEx (as the file size is large). Instead, I have used [this](https://www.kaggle.com/datasets/munumbutt/amexfeather) feather dataset. Now in my case, how will Kaggle run my code for the private leaderboard?"
    },
    {
      "id": 1900544,
      "postDate": "2022-08-16T05:20:26.613Z",
      "content": "<p>Hi,<br>\nI hope everyone will do well and build a good team.</p>",
      "rawMarkdown": "Hi,\nI hope everyone will do well and build a good team."
    },
    {
      "id": 1899470,
      "postDate": "2022-08-15T09:23:57.790Z",
      "content": "<p>Happy Modeling!Happy Kaggling!</p>",
      "rawMarkdown": "Happy Modeling!Happy Kaggling!"
    },
    {
      "id": 1880285,
      "postDate": "2022-08-01T15:29:30.803Z",
      "content": "<p>hi <br>\nthis is my first participation.<br>\nI am pursuing PGP in Data Science, I have built basic models during my project work.<br>\nI used the train_data for model building and got recall &amp; f1 score of 0.96(for 1) on the training dataset.<br>\nBut  evaluating my model on \"amex_metric\" is giving me a score of 0.0209.<br>\nI want to understand the possible reasons for the same.<br>\nHow can I improve my Score</p>",
      "rawMarkdown": "hi \nthis is my first participation.\nI am pursuing PGP in Data Science, I have built basic models during my project work.\nI used the train_data for model building and got recall & f1 score of 0.96(for 1) on the training dataset.\nBut  evaluating my model on \"amex_metric\" is giving me a score of 0.0209.\nI want to understand the possible reasons for the same.\nHow can I improve my Score"
    },
    {
      "id": 1867745,
      "postDate": "2022-07-23T13:28:31.310Z",
      "content": "<p>Hello, this will be my first Kaggle submission !!</p>",
      "rawMarkdown": "Hello, this will be my first Kaggle submission !!"
    },
    {
      "id": 1865363,
      "postDate": "2022-07-21T18:50:15.377Z",
      "content": "<p>Hi, <br>\nThis is my first competition and I am a bit confused regarding the daily submission. It says that it will be automatically graded? How does this work? Will importing nrow=500, instead of the entire dataset, affect the grade? Also, it says I have to match the submission CSV but I don't understand how I am supposed to be creating a csv.<br>\nThanks!</p>",
      "rawMarkdown": "Hi, \nThis is my first competition and I am a bit confused regarding the daily submission. It says that it will be automatically graded? How does this work? Will importing nrow=500, instead of the entire dataset, affect the grade? Also, it says I have to match the submission CSV but I don't understand how I am supposed to be creating a csv.\nThanks!"
    },
    {
      "id": 1855128,
      "postDate": "2022-07-14T09:52:39.497Z",
      "content": "<p>Hello everyone! I'm somewhat of a new joiner on Kaggle and to machine learning as well. This is going to be my first official competition - besides the tutorial one so wish me luck! I'm certainly excited about what I can unravel with my newly acquired skills.</p>",
      "rawMarkdown": "Hello everyone! I'm somewhat of a new joiner on Kaggle and to machine learning as well. This is going to be my first official competition - besides the tutorial one so wish me luck! I'm certainly excited about what I can unravel with my newly acquired skills."
    },
    {
      "id": 1832772,
      "postDate": "2022-06-25T11:34:03.863Z",
      "content": "<p>I have a laptop with only 4GB RAM. Could anyone please write down the best way to proceed with this system requirement? <br>\nIs there any cloud storage we can opt temporarily? </p>",
      "rawMarkdown": "I have a laptop with only 4GB RAM. Could anyone please write down the best way to proceed with this system requirement? \nIs there any cloud storage we can opt temporarily? "
    },
    {
      "id": 1831965,
      "postDate": "2022-06-24T14:53:02.037Z",
      "content": "<p>So the loading process is done. But now Im stuck with what to do with the data. I have mostly done randomforest model. But i dont think RandomForest will perform well on this huge dataset. Any suggestions?</p>",
      "rawMarkdown": "So the loading process is done. But now Im stuck with what to do with the data. I have mostly done randomforest model. But i dont think RandomForest will perform well on this huge dataset. Any suggestions?\n"
    },
    {
      "id": 1828400,
      "postDate": "2022-06-21T16:54:35.573Z",
      "content": "<p>Hi<br>\nThe data is still downloading…any solution </p>",
      "rawMarkdown": "Hi\nThe data is still downloading...any solution "
    },
    {
      "id": 1822974,
      "postDate": "2022-06-16T22:26:10.170Z",
      "content": "<p>Hi Adisson, Thank you for sharing.</p>\n<p>Please can you help me with this:</p>\n<p>D_* = Delinquency variables<br>\nS_* = Spend variables<br>\nP_* = Payment variables<br>\nB_* = Balance variables<br>\nR_* = Risk variables</p>\n<p>What does it mean '*' in each variable ?. maybe days from the latest credit card statement ?</p>\n<p>So, D_39 is customer Delinquency at 39 days after latest credit card statement ?</p>\n<p>Thank's</p>",
      "rawMarkdown": "Hi Adisson, Thank you for sharing.\n\nPlease can you help me with this:\n\nD_* = Delinquency variables\nS_* = Spend variables\nP_* = Payment variables\nB_* = Balance variables\nR_* = Risk variables\n\nWhat does it mean '*' in each variable ?. maybe days from the latest credit card statement ?\n\nSo, D_39 is customer Delinquency at 39 days after latest credit card statement ?\n\nThank's",
      "replies": [
        {
          "id": 1879268,
          "postDate": "2022-08-01T00:54:12.867Z",
          "content": "<p>The * after the letter just refers to that number in each column name, meaning D_* means that any column that leads with \"D\" is a delinquency variable regardless of what number comes next, meaning D_39 and D_42 are both delinquency variables. I also wouldn't infer too much from the variable names, given the way that Amex is masking/anonymizing the data we're not meant to be able to easily infer something from the letter for variable type and the number after it. Hope this helps </p>",
          "rawMarkdown": "The * after the letter just refers to that number in each column name, meaning D_* means that any column that leads with \"D\" is a delinquency variable regardless of what number comes next, meaning D_39 and D_42 are both delinquency variables. I also wouldn't infer too much from the variable names, given the way that Amex is masking/anonymizing the data we're not meant to be able to easily infer something from the letter for variable type and the number after it. Hope this helps \n\n\n"
        }
      ]
    },
    {
      "id": 1821214,
      "postDate": "2022-06-15T10:22:21.617Z",
      "content": "<p>This is very nice, thank you for the video!</p>",
      "rawMarkdown": "This is very nice, thank you for the video!\n"
    },
    {
      "id": 1820967,
      "postDate": "2022-06-15T06:38:56.920Z",
      "content": "<p>This is our first competition on Kaggle, we are working with the dataset but are not able to understand all the features of the dataset completely and how all the features are aggregated as<br>\nD_* = Delinquency variables<br>\nS_* = Spend variables<br>\nP_* = Payment variables<br>\nB_* = Balance variables<br>\nR_* = Risk variables<br>\nIt would be a great help if anyone can share the data dictionary and resources to understand the dataset.</p>",
      "rawMarkdown": "This is our first competition on Kaggle, we are working with the dataset but are not able to understand all the features of the dataset completely and how all the features are aggregated as\nD_* = Delinquency variables\nS_* = Spend variables\nP_* = Payment variables\nB_* = Balance variables\nR_* = Risk variables\nIt would be a great help if anyone can share the data dictionary and resources to understand the dataset.",
      "replies": [
        {
          "id": 1821460,
          "postDate": "2022-06-15T14:50:14.027Z",
          "content": "<p>I think what the aggregated features are like the main categories whereas given feature in the tran_data.csv are subcategories of those main categories. correct me if I am wrong I am also new here. </p>",
          "rawMarkdown": "I think what the aggregated features are like the main categories whereas given feature in the tran_data.csv are subcategories of those main categories. correct me if I am wrong I am also new here. "
        }
      ]
    },
    {
      "id": 1817018,
      "postDate": "2022-06-10T18:58:10.427Z",
      "content": "<p>Thanks for sharing! I will watch the videos as soon as possible.</p>",
      "rawMarkdown": "Thanks for sharing! I will watch the videos as soon as possible."
    },
    {
      "id": 1816556,
      "postDate": "2022-06-10T10:11:41.170Z",
      "content": "<p>Thanks Kaggle for the Site etiquette video and Kaggle lingo video.😊😊</p>",
      "rawMarkdown": "Thanks Kaggle for the Site etiquette video and Kaggle lingo video.😊😊"
    },
    {
      "id": 1812447,
      "postDate": "2022-06-05T20:24:45.397Z",
      "content": "<p>Thankyou for sharing looking forward to this experience and modelling.🙌🙌</p>",
      "rawMarkdown": "Thankyou for sharing looking forward to this experience and modelling.🙌🙌"
    },
    {
      "id": 1809116,
      "postDate": "2022-06-02T12:06:36.963Z",
      "content": "<p>thanx , Rachel</p>",
      "rawMarkdown": "thanx , Rachel"
    },
    {
      "id": 1805469,
      "postDate": "2022-05-30T08:06:43.103Z",
      "content": "<p>Thanks, Rachael for the very informative video on Kaggle lingo.</p>",
      "rawMarkdown": "Thanks, Rachael for the very informative video on Kaggle lingo."
    },
    {
      "id": 1864305,
      "postDate": "2022-07-20T23:33:30.620Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1831525,
      "postDate": "2022-06-24T08:13:49.823Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1817980,
      "postDate": "2022-06-12T05:04:33.987Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1816762,
      "postDate": "2022-06-10T14:26:15.087Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1913644,
      "postDate": "2022-08-25T12:30:20.213Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 1911119,
      "postDate": "2022-08-24T00:14:05.387Z",
      "content": "<p>thank for sharing!</p>",
      "rawMarkdown": "thank for sharing!"
    },
    {
      "id": 1897902,
      "postDate": "2022-08-14T06:21:45.877Z",
      "content": "<p>Thanks for u sharing</p>",
      "rawMarkdown": "Thanks for u sharing"
    },
    {
      "id": 1897620,
      "postDate": "2022-08-13T23:47:58.517Z",
      "content": "<p>Thanks for the videos!</p>",
      "rawMarkdown": "Thanks for the videos!"
    },
    {
      "id": 1888029,
      "postDate": "2022-08-07T09:59:16.417Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 1883719,
      "postDate": "2022-08-04T03:39:00.117Z",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!"
    },
    {
      "id": 1867499,
      "postDate": "2022-07-23T09:12:49.753Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 1859546,
      "postDate": "2022-07-17T17:49:28.647Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 1843255,
      "postDate": "2022-07-04T16:45:24.323Z",
      "content": "<p>Thanks for being so informative </p>",
      "rawMarkdown": "Thanks for being so informative "
    },
    {
      "id": 1840364,
      "postDate": "2022-07-02T07:38:39.823Z",
      "content": "<p>Thanks for sharing！</p>",
      "rawMarkdown": "Thanks for sharing！"
    },
    {
      "id": 1827652,
      "postDate": "2022-06-21T07:52:22.493Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 1816364,
      "postDate": "2022-06-10T06:38:20.970Z",
      "content": "<p>Thanks for the details  👍</p>",
      "rawMarkdown": "Thanks for the details  👍"
    },
    {
      "id": 1814691,
      "postDate": "2022-06-08T07:04:22.697Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing.\n"
    },
    {
      "id": 1811405,
      "postDate": "2022-06-04T16:44:52.197Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing."
    },
    {
      "id": 1811087,
      "postDate": "2022-06-04T09:51:47.673Z",
      "content": "<p>Thanks for the interesting videos</p>",
      "rawMarkdown": "Thanks for the interesting videos"
    },
    {
      "id": 1805485,
      "postDate": "2022-05-30T08:24:33.927Z",
      "content": "<p>Thanks for the video!</p>",
      "rawMarkdown": "Thanks for the video!"
    }
  ],
  "comments": [
    {
      "id": 1832500,
      "author_name": "Taro Kuroda",
      "author_url": "",
      "post_date": "2022-06-25T05:27:52.430000",
      "content": "<p>Hi! <br>\nI'm a begginer of Kaggle, but want to attend aggressively for my career making!</p>\n<p>I want to ask the way to deal with such the big data.<br>\nOnly downloading take long time…<br>\nHow to deal with this problem? </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1832774,
          "author_name": "Abin Singh",
          "author_url": "",
          "post_date": "2022-06-25T11:35:24.733000",
          "content": "<p>I also have the same problem. I have only 4GB RAM space. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1833754,
          "author_name": "Aleksey Bannikov",
          "author_url": "",
          "post_date": "2022-06-26T09:55:54.507000",
          "content": "<p>One of the ways to deal with such a big amounts of data is using parquet data sets. Parquet data set is a compressed version of .csv.  You can find  notebooks in the Internet that already contain parquet datasets (it's size is about 5-10 times less than size of original datasets). You can work with parquet datasets this way:</p>\n<blockquote>\n  <p>import pandas as pd<br>\n  import numpy as np<br>\n  import gc<br>\n  import matplotlib.pyplot as plt<br>\n  import seaborn as sns<br>\n  train = pd.read_parquet(r'C:\\Users\\aljutov\\Desktop\\American_Express_competition\\train_data.parquet\\train_data.parquet')<br>\n  test = pd.read_parquet(r'C:\\Users\\aljutov\\Desktop\\American_Express_competition\\test_data.parquet\\test_data.parquet')</p>\n</blockquote>\n<p>after that begin EDA as usual…</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1882782,
          "author_name": "KULKARNI ISHWARI NITIN",
          "author_url": "",
          "post_date": "2022-08-03T12:27:48.457000",
          "content": "<p>Hi How did you create parquet files please can you share entire procedure</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1897711,
          "author_name": "WD",
          "author_url": "",
          "post_date": "2022-08-14T03:09:48.320000",
          "content": "<p>Do you know how to create parquet datasets now?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1888828,
      "author_name": "Olusola Awodiya",
      "author_url": "",
      "post_date": "2022-08-07T19:47:09.917000",
      "content": "<p>Hello Guys, I am still fairly new to machine learning and I am joining this competition not necessarily to win (although I want to) but to progress in my career as a successful Data Scientist. I am looking to make lifelong friends in this community with whom I can learn and add value to. So please, feel free to reach out to me. My LinkedIn is embedded in my profile. Wishing everyone success and good luck in this competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1895045,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-11T22:48:19.963000",
          "content": "<p>Welcome Olusola!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1850019,
      "author_name": "kusashi",
      "author_url": "",
      "post_date": "2022-07-10T02:29:53.950000",
      "content": "<p>I want to improve my data analysis skill.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1818947,
      "author_name": "reallsanya",
      "author_url": "",
      "post_date": "2022-06-13T10:03:49.680000",
      "content": "<p>Hey! I'm know some theory in DS and want to start practice! But, I faced with memory problem (for this dataset -- Store Sales - Time Series Forecasting, 124.76 MiB), in this competition 50.31 GiB of data - how people can to do something with that batch? Сan I do something without adding video cards, etc. only with kaggle environment power?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1820356,
          "author_name": "tom",
          "author_url": "",
          "post_date": "2022-06-14T14:59:42.467000",
          "content": "<p>Hi, </p>\n<p>Try to use parametrs <em>iterator=True, chunksize=1000000</em> when reading data with pd.read_csv</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1821522,
          "author_name": "AbhishekKumar",
          "author_url": "",
          "post_date": "2022-06-15T16:13:45.630000",
          "content": "<p>Check below for more understanding:</p>\n<ul>\n<li><a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/abhianalytic/amex-read-full-data-without-any-transformation</a></li>\n<li><a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328565</a></li>\n</ul>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1853297,
      "author_name": "Xavier R Nogueira",
      "author_url": "",
      "post_date": "2022-07-12T19:36:41.447000",
      "content": "<p>Hi! This is my first Kaggle competition, however I have worked on a few ML projects in my job as an environmental data scientist ( I would say I am probably at a low-intermediate skill level.). That said, in environmental contexts we often use geospatial data which due to it's large memory footprint and complexity often limits my ability to go \"all out\" on predictions (also our clients like to understand things simply, so too much feature engineering is avoided). I'm looking forward to being able to do exactly what I want with the data without worrying about clients or other more mundane things.</p>\n<p>To start, since I am running my analysis on a simple laptop, I decided to use dask for the entire workflow. To start, imported the train/test data, formatted their dtypes, and re-saved to partitioned .parquet files. This takes 12 minutes for all the provided data! If you want to do the same thing to kick start your work flow check out my notebook here:  <a href=\"https://www.kaggle.com/code/xaviernogueira/dask-pre-processing-csv-parquet-in-parallel\" target=\"_blank\">https://www.kaggle.com/code/xaviernogueira/dask-pre-processing-csv-parquet-in-parallel</a>. I would appreciate an upvote if you find it useful :) </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1911530,
      "author_name": "Sumit Adwani",
      "author_url": "",
      "post_date": "2022-08-24T06:49:53.517000",
      "content": "<p>Hello Everyone! I am new to Kaggle and currently unaware of how exactly the private leaderboard works. My doubt is:</p>\n<p>According to the competition rules, the public leaderboard is decided based on 51% of test data. And private leaderboard will be decided based on the other 49%. So that means the test set that we have used in our codes is only 51% of the actual test set, and after the competition deadline, Kaggle uses the other 41% (which is not currently provided to us), right? </p>\n<p>So how exactly does Kaggle run our code with a different test set? I am worried about this because I have not used the dataset provided by AmEx (as the file size is large). Instead, I have used <a href=\"https://www.kaggle.com/datasets/munumbutt/amexfeather\" target=\"_blank\">this</a> feather dataset. Now in my case, how will Kaggle run my code for the private leaderboard?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1900544,
      "author_name": "Ajay Patsariya",
      "author_url": "",
      "post_date": "2022-08-16T05:20:26.613000",
      "content": "<p>Hi,<br>\nI hope everyone will do well and build a good team.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1899470,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-15T09:23:57.790000",
      "content": "<p>Happy Modeling!Happy Kaggling!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1880285,
      "author_name": "ATULKAUL2304",
      "author_url": "",
      "post_date": "2022-08-01T15:29:30.803000",
      "content": "<p>hi <br>\nthis is my first participation.<br>\nI am pursuing PGP in Data Science, I have built basic models during my project work.<br>\nI used the train_data for model building and got recall &amp; f1 score of 0.96(for 1) on the training dataset.<br>\nBut  evaluating my model on \"amex_metric\" is giving me a score of 0.0209.<br>\nI want to understand the possible reasons for the same.<br>\nHow can I improve my Score</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1867745,
      "author_name": "Mahendra Singh Rajpoot",
      "author_url": "",
      "post_date": "2022-07-23T13:28:31.310000",
      "content": "<p>Hello, this will be my first Kaggle submission !!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1865363,
      "author_name": "Clara Sastra",
      "author_url": "",
      "post_date": "2022-07-21T18:50:15.377000",
      "content": "<p>Hi, <br>\nThis is my first competition and I am a bit confused regarding the daily submission. It says that it will be automatically graded? How does this work? Will importing nrow=500, instead of the entire dataset, affect the grade? Also, it says I have to match the submission CSV but I don't understand how I am supposed to be creating a csv.<br>\nThanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1855128,
      "author_name": "ChrisRP",
      "author_url": "",
      "post_date": "2022-07-14T09:52:39.497000",
      "content": "<p>Hello everyone! I'm somewhat of a new joiner on Kaggle and to machine learning as well. This is going to be my first official competition - besides the tutorial one so wish me luck! I'm certainly excited about what I can unravel with my newly acquired skills.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1832772,
      "author_name": "Abin Singh",
      "author_url": "",
      "post_date": "2022-06-25T11:34:03.863000",
      "content": "<p>I have a laptop with only 4GB RAM. Could anyone please write down the best way to proceed with this system requirement? <br>\nIs there any cloud storage we can opt temporarily? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1831965,
      "author_name": "Humanoid007",
      "author_url": "",
      "post_date": "2022-06-24T14:53:02.037000",
      "content": "<p>So the loading process is done. But now Im stuck with what to do with the data. I have mostly done randomforest model. But i dont think RandomForest will perform well on this huge dataset. Any suggestions?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1828400,
      "author_name": "Firas Hamid Malik",
      "author_url": "",
      "post_date": "2022-06-21T16:54:35.573000",
      "content": "<p>Hi<br>\nThe data is still downloading…any solution </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1822974,
      "author_name": "Gustav C",
      "author_url": "",
      "post_date": "2022-06-16T22:26:10.170000",
      "content": "<p>Hi Adisson, Thank you for sharing.</p>\n<p>Please can you help me with this:</p>\n<p>D_* = Delinquency variables<br>\nS_* = Spend variables<br>\nP_* = Payment variables<br>\nB_* = Balance variables<br>\nR_* = Risk variables</p>\n<p>What does it mean '*' in each variable ?. maybe days from the latest credit card statement ?</p>\n<p>So, D_39 is customer Delinquency at 39 days after latest credit card statement ?</p>\n<p>Thank's</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1879268,
          "author_name": "Markham Lee",
          "author_url": "",
          "post_date": "2022-08-01T00:54:12.867000",
          "content": "<p>The * after the letter just refers to that number in each column name, meaning D_* means that any column that leads with \"D\" is a delinquency variable regardless of what number comes next, meaning D_39 and D_42 are both delinquency variables. I also wouldn't infer too much from the variable names, given the way that Amex is masking/anonymizing the data we're not meant to be able to easily infer something from the letter for variable type and the number after it. Hope this helps </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1821214,
      "author_name": "menno",
      "author_url": "",
      "post_date": "2022-06-15T10:22:21.617000",
      "content": "<p>This is very nice, thank you for the video!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1820967,
      "author_name": "Shivendra",
      "author_url": "",
      "post_date": "2022-06-15T06:38:56.920000",
      "content": "<p>This is our first competition on Kaggle, we are working with the dataset but are not able to understand all the features of the dataset completely and how all the features are aggregated as<br>\nD_* = Delinquency variables<br>\nS_* = Spend variables<br>\nP_* = Payment variables<br>\nB_* = Balance variables<br>\nR_* = Risk variables<br>\nIt would be a great help if anyone can share the data dictionary and resources to understand the dataset.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1821460,
          "author_name": "Dewanik Koirala",
          "author_url": "",
          "post_date": "2022-06-15T14:50:14.027000",
          "content": "<p>I think what the aggregated features are like the main categories whereas given feature in the tran_data.csv are subcategories of those main categories. correct me if I am wrong I am also new here. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1817018,
      "author_name": "Felipe Renck",
      "author_url": "",
      "post_date": "2022-06-10T18:58:10.427000",
      "content": "<p>Thanks for sharing! I will watch the videos as soon as possible.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1816556,
      "author_name": "KH_ARJUN",
      "author_url": "",
      "post_date": "2022-06-10T10:11:41.170000",
      "content": "<p>Thanks Kaggle for the Site etiquette video and Kaggle lingo video.😊😊</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1812447,
      "author_name": "Nishad Ingale",
      "author_url": "",
      "post_date": "2022-06-05T20:24:45.397000",
      "content": "<p>Thankyou for sharing looking forward to this experience and modelling.🙌🙌</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1809116,
      "author_name": "Harshit Gaur",
      "author_url": "",
      "post_date": "2022-06-02T12:06:36.963000",
      "content": "<p>thanx , Rachel</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1805469,
      "author_name": "TianmaXingkong(TX)",
      "author_url": "",
      "post_date": "2022-05-30T08:06:43.103000",
      "content": "<p>Thanks, Rachael for the very informative video on Kaggle lingo.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1864305,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-20T23:33:30.620000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1831525,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-24T08:13:49.823000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1817980,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-12T05:04:33.987000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1816762,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-10T14:26:15.087000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1913644,
      "author_name": "Peter Doohyun Choi",
      "author_url": "",
      "post_date": "2022-08-25T12:30:20.213000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1911119,
      "author_name": "Rinat Akhmetov",
      "author_url": "",
      "post_date": "2022-08-24T00:14:05.387000",
      "content": "<p>thank for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1897902,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-14T06:21:45.877000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1897620,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-13T23:47:58.517000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1888029,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-07T09:59:16.417000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1883719,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-04T03:39:00.117000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1867499,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-23T09:12:49.753000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1859546,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-17T17:49:28.647000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1843255,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-04T16:45:24.323000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1840364,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-02T07:38:39.823000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1827652,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-21T07:52:22.493000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1816364,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-10T06:38:20.970000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1814691,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-08T07:04:22.697000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1811405,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-04T16:44:52.197000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1811087,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-04T09:51:47.673000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1805485,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-30T08:24:33.927000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1800285": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! \n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos Dr. Rachael Tatman has put together to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), or [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s).\n\n**Remember**: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our [Kaggle community guidelines](https://www.kaggle.com/community-guidelines).\n\nHappy Modeling!",
    "1832500": "Hi! \nI'm a begginer of Kaggle, but want to attend aggressively for my career making!\n\nI want to ask the way to deal with such the big data.\nOnly downloading take long time...\nHow to deal with this problem? ",
    "1888828": "Hello Guys, I am still fairly new to machine learning and I am joining this competition not necessarily to win (although I want to) but to progress in my career as a successful Data Scientist. I am looking to make lifelong friends in this community with whom I can learn and add value to. So please, feel free to reach out to me. My LinkedIn is embedded in my profile. Wishing everyone success and good luck in this competition.",
    "1850019": "I want to improve my data analysis skill.",
    "1818947": "Hey! I'm know some theory in DS and want to start practice! But, I faced with memory problem (for this dataset -- Store Sales - Time Series Forecasting, 124.76 MiB), in this competition 50.31 GiB of data - how people can to do something with that batch? Сan I do something without adding video cards, etc. only with kaggle environment power?",
    "1853297": "Hi! This is my first Kaggle competition, however I have worked on a few ML projects in my job as an environmental data scientist ( I would say I am probably at a low-intermediate skill level.). That said, in environmental contexts we often use geospatial data which due to it's large memory footprint and complexity often limits my ability to go \"all out\" on predictions (also our clients like to understand things simply, so too much feature engineering is avoided). I'm looking forward to being able to do exactly what I want with the data without worrying about clients or other more mundane things.\n\nTo start, since I am running my analysis on a simple laptop, I decided to use dask for the entire workflow. To start, imported the train/test data, formatted their dtypes, and re-saved to partitioned .parquet files. This takes 12 minutes for all the provided data! If you want to do the same thing to kick start your work flow check out my notebook here:  https://www.kaggle.com/code/xaviernogueira/dask-pre-processing-csv-parquet-in-parallel. I would appreciate an upvote if you find it useful :) ",
    "1911530": "Hello Everyone! I am new to Kaggle and currently unaware of how exactly the private leaderboard works. My doubt is:\n\nAccording to the competition rules, the public leaderboard is decided based on 51% of test data. And private leaderboard will be decided based on the other 49%. So that means the test set that we have used in our codes is only 51% of the actual test set, and after the competition deadline, Kaggle uses the other 41% (which is not currently provided to us), right? \n\nSo how exactly does Kaggle run our code with a different test set? I am worried about this because I have not used the dataset provided by AmEx (as the file size is large). Instead, I have used [this](https://www.kaggle.com/datasets/munumbutt/amexfeather) feather dataset. Now in my case, how will Kaggle run my code for the private leaderboard?",
    "1900544": "Hi,\nI hope everyone will do well and build a good team.",
    "1899470": "Happy Modeling!Happy Kaggling!",
    "1880285": "hi \nthis is my first participation.\nI am pursuing PGP in Data Science, I have built basic models during my project work.\nI used the train_data for model building and got recall & f1 score of 0.96(for 1) on the training dataset.\nBut  evaluating my model on \"amex_metric\" is giving me a score of 0.0209.\nI want to understand the possible reasons for the same.\nHow can I improve my Score",
    "1867745": "Hello, this will be my first Kaggle submission !!",
    "1865363": "Hi, \nThis is my first competition and I am a bit confused regarding the daily submission. It says that it will be automatically graded? How does this work? Will importing nrow=500, instead of the entire dataset, affect the grade? Also, it says I have to match the submission CSV but I don't understand how I am supposed to be creating a csv.\nThanks!",
    "1855128": "Hello everyone! I'm somewhat of a new joiner on Kaggle and to machine learning as well. This is going to be my first official competition - besides the tutorial one so wish me luck! I'm certainly excited about what I can unravel with my newly acquired skills.",
    "1832772": "I have a laptop with only 4GB RAM. Could anyone please write down the best way to proceed with this system requirement? \nIs there any cloud storage we can opt temporarily? ",
    "1831965": "So the loading process is done. But now Im stuck with what to do with the data. I have mostly done randomforest model. But i dont think RandomForest will perform well on this huge dataset. Any suggestions?\n",
    "1828400": "Hi\nThe data is still downloading...any solution ",
    "1822974": "Hi Adisson, Thank you for sharing.\n\nPlease can you help me with this:\n\nD_* = Delinquency variables\nS_* = Spend variables\nP_* = Payment variables\nB_* = Balance variables\nR_* = Risk variables\n\nWhat does it mean '*' in each variable ?. maybe days from the latest credit card statement ?\n\nSo, D_39 is customer Delinquency at 39 days after latest credit card statement ?\n\nThank's",
    "1821214": "This is very nice, thank you for the video!\n",
    "1820967": "This is our first competition on Kaggle, we are working with the dataset but are not able to understand all the features of the dataset completely and how all the features are aggregated as\nD_* = Delinquency variables\nS_* = Spend variables\nP_* = Payment variables\nB_* = Balance variables\nR_* = Risk variables\nIt would be a great help if anyone can share the data dictionary and resources to understand the dataset.",
    "1817018": "Thanks for sharing! I will watch the videos as soon as possible.",
    "1816556": "Thanks Kaggle for the Site etiquette video and Kaggle lingo video.😊😊",
    "1812447": "Thankyou for sharing looking forward to this experience and modelling.🙌🙌",
    "1809116": "thanx , Rachel",
    "1805469": "Thanks, Rachael for the very informative video on Kaggle lingo.",
    "1864305": "",
    "1831525": "",
    "1817980": "",
    "1816762": "",
    "1913644": "Thanks for sharing",
    "1911119": "thank for sharing!",
    "1897902": "Thanks for u sharing",
    "1897620": "Thanks for the videos!",
    "1888029": "Thanks for sharing",
    "1883719": "thanks for sharing!",
    "1867499": "Thanks for sharing!",
    "1859546": "Thanks for sharing",
    "1843255": "Thanks for being so informative ",
    "1840364": "Thanks for sharing！",
    "1827652": "Thanks for sharing",
    "1816364": "Thanks for the details  👍",
    "1814691": "Thank you for sharing.\n",
    "1811405": "Thank you for sharing.",
    "1811087": "Thanks for the interesting videos",
    "1805485": "Thanks for the video!"
  }
}