{
  "id": 341561,
  "title": "Are we guessing blindly here ??",
  "url": "/competitions/amex-default-prediction/discussion/341561",
  "author_name": "",
  "post_date": "2022-08-03T11:19:36.521059600Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I tried to dissect the data in various way just so that it make some sense, but fails at each attempt.  As an armature its hard for me to do any analysis without knowing what the columns refers to and what values I am looking at. I know companies want to hide the meaning of data but how can one perform any kind of analysis without knowing what the columns refers to and what value you are looking at.</p>\n<p>Like in column B_36 one of the value is '0.009968', so what is the this value. B denotes balance variable but there are 42 balance variable.</p>\n<p>So are we just guessing blinding and happy if model score 0.7+ or is there a method to understand this madness.</p>",
  "messages": [
    {
      "id": "1882666",
      "postDate": "08/03/2022 11:19:36",
      "content": "<p>I tried to dissect the data in various way just so that it make some sense, but fails at each attempt.  As an armature its hard for me to do any analysis without knowing what the columns refers to and what values I am looking at. I know companies want to hide the meaning of data but how can one perform any kind of analysis without knowing what the columns refers to and what value you are looking at.</p>\n<p>Like in column B_36 one of the value is '0.009968', so what is the this value. B denotes balance variable but there are 42 balance variable.</p>\n<p>So are we just guessing blinding and happy if model score 0.7+ or is there a method to understand this madness.</p>",
      "rawMarkdown": "I tried to dissect the data in various way just so that it make some sense, but fails at each attempt.  As an armature its hard for me to do any analysis without knowing what the columns refers to and what values I am looking at. I know companies want to hide the meaning of data but how can one perform any kind of analysis without knowing what the columns refers to and what value you are looking at.\n\nLike in column B_36 one of the value is '0.009968', so what is the this value. B denotes balance variable but there are 42 balance variable.\n\nSo are we just guessing blinding and happy if model score 0.7+ or is there a method to understand this madness.",
      "votes": null
    },
    {
      "id": "1882882",
      "postDate": "08/03/2022 13:15:37",
      "content": "<p>I beg to differ, we don't really need know every detail regarding variables to proceed, of course it is good to know though. All we need to know is how these variable affect the target. This being said there is some astute analysis on the data performed by the grand masters, I would implore you to read these discussions. I believe that we will you gather perspective.</p>",
      "rawMarkdown": "I beg to differ, we don't really need know every detail regarding variables to proceed, of course it is good to know though. All we need to know is how these variable affect the target. This being said there is some astute analysis on the data performed by the grand masters, I would implore you to read these discussions. I believe that we will you gather perspective.",
      "votes": null
    },
    {
      "id": "1882970",
      "postDate": "08/03/2022 14:00:14",
      "content": "<p>It's part of the challenge and fun!</p>",
      "rawMarkdown": "It's part of the challenge and fun!",
      "votes": null
    },
    {
      "id": "1883534",
      "postDate": "08/03/2022 22:32:30",
      "content": "<p>It is not necessary at all to know the source of data or the meaning of columns, and Kaggle datasets are often fully or partially anonymized. At least in this competition we know the general purpose of BDPRS columns. It would be helpful to know the meaning of columns for high-level feature engineering, but even that part often requires domain knowledge that most of us don't have.</p>\n<p>I think the bigger problem in this competition is that so many data columns have a uniform noise injected, yet even that has been solved for the most part.</p>",
      "rawMarkdown": "It is not necessary at all to know the source of data or the meaning of columns, and Kaggle datasets are often fully or partially anonymized. At least in this competition we know the general purpose of BDPRS columns. It would be helpful to know the meaning of columns for high-level feature engineering, but even that part often requires domain knowledge that most of us don't have.\n\nI think the bigger problem in this competition is that so many data columns have a uniform noise injected, yet even that has been solved for the most part.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1882882,
      "author_name": "sarthmirashi07",
      "author_url": "",
      "post_date": "08/03/2022 13:15:37",
      "content": "<p>I beg to differ, we don't really need know every detail regarding variables to proceed, of course it is good to know though. All we need to know is how these variable affect the target. This being said there is some astute analysis on the data performed by the grand masters, I would implore you to read these discussions. I believe that we will you gather perspective.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1882970,
      "author_name": "scharlesworth",
      "author_url": "",
      "post_date": "08/03/2022 14:00:14",
      "content": "<p>It's part of the challenge and fun!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1883534,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "08/03/2022 22:32:30",
      "content": "<p>It is not necessary at all to know the source of data or the meaning of columns, and Kaggle datasets are often fully or partially anonymized. At least in this competition we know the general purpose of BDPRS columns. It would be helpful to know the meaning of columns for high-level feature engineering, but even that part often requires domain knowledge that most of us don't have.</p>\n<p>I think the bigger problem in this competition is that so many data columns have a uniform noise injected, yet even that has been solved for the most part.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1882666": "I tried to dissect the data in various way just so that it make some sense, but fails at each attempt.  As an armature its hard for me to do any analysis without knowing what the columns refers to and what values I am looking at. I know companies want to hide the meaning of data but how can one perform any kind of analysis without knowing what the columns refers to and what value you are looking at.\n\nLike in column B_36 one of the value is '0.009968', so what is the this value. B denotes balance variable but there are 42 balance variable.\n\nSo are we just guessing blinding and happy if model score 0.7+ or is there a method to understand this madness.",
    "1882882": "I beg to differ, we don't really need know every detail regarding variables to proceed, of course it is good to know though. All we need to know is how these variable affect the target. This being said there is some astute analysis on the data performed by the grand masters, I would implore you to read these discussions. I believe that we will you gather perspective.",
    "1882970": "It's part of the challenge and fun!",
    "1883534": "It is not necessary at all to know the source of data or the meaning of columns, and Kaggle datasets are often fully or partially anonymized. At least in this competition we know the general purpose of BDPRS columns. It would be helpful to know the meaning of columns for high-level feature engineering, but even that part often requires domain knowledge that most of us don't have.\n\nI think the bigger problem in this competition is that so many data columns have a uniform noise injected, yet even that has been solved for the most part."
  },
  "source": "meta"
}