{
  "id": 331180,
  "title": "difference between 'empty', '0.0', '-1.0', '-1' && category vs label?",
  "url": "/competitions/amex-default-prediction/discussion/331180",
  "author_name": "datashovel",
  "post_date": "2022-06-16T04:22:08.623000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Sorry of this is already covered.  If so please just refer me to the applicable thread.</p>\n<p>First thing I wanted to do was analyze what kind of data is stored for each data point.  Most appear to be some value on a continuous spectrum of possible values.  The others I've labeled in the following categories.  As is mentioned from the documentation, 11 are categorical [B30, B38, D114, D116, D117, D120, D126, D63, D64, D66, D68].  But there is a distinct difference in that D_63 and D_64 contain \"label\" values and not \"1.0, 2.0, etc\" values:</p>\n<p>\"category\" - B_30, B_38, D_66, D_68, D_114, D_116, D_117, D_120, D_126<br>\n\"boolean\" - B_31<br>\n\"checkbox?\" - D_87<br>\n\"label\" - D_63, D_64</p>\n<h1>1</h1>\n<p>Would it be accurate to describe B_31 as a boolean value, and D_87 some equivalent of \"if the checkbox is checked then you see 1.0, otherwise you see empty\"?  Just curious if someone could elaborate on what the difference is between B_31 and D_87 and / or whether it could matter when it comes to how you model that data?  I'm imagining the scenario could be that B_31 might've been some radio button group on an application form where it was a required field when the person applied, so everyone had to choose either \"true\" or \"false\".  Whereas maybe D_87 could've been some checkbox on an application form where the person either did check it \"1.0\" or didn't check it \"empty\"?  I'm assuming here there was a deliberate reason B_31 and D_87 have different possible values though they could both have been represented to us by a boolean data type.</p>\n<h1>2</h1>\n<p>Next, the \"category\" and \"label\" data types have several values which may or may not mean what I think they mean.  And that may or may not matter.</p>\n<p>\"empty\", \"0.0\", \"-1\", \"-1.0\".  I wanted to see if someone knows enough to be able to differentiate between what each of these values mean?  My assumption otherwise is that any record, with a \"1.0, 2.0, 3.0, CL, CO, O, R, U, etc\" value, means that the record is \"part of that category\".</p>\n<p>This is just a guess to get the conversation started, but my current assumption is maybe this is what they each mean:</p>\n<p>\"empty\" - there is no information about this category, for this record<br>\n\"0.0\" - this record is maybe part of some \"other\" category besides what is listed as possible categories<br>\n\"-1 or -1.0\" - this record is maybe \"n/a\" when it comes to this category?</p>\n<h1>3</h1>\n<p>Third, I'd like to know if there was a reason D_63 and D_64 have \"label\" values instead of the more ambiguous \"1.0, 2.0, 3.0, etc\" values.  And if there was, whether it could matter in terms of how I model this data?</p>",
  "messages": [
    {
      "id": 1822096,
      "postDate": "2022-06-16T04:22:08.623Z",
      "content": "<p>Sorry of this is already covered.  If so please just refer me to the applicable thread.</p>\n<p>First thing I wanted to do was analyze what kind of data is stored for each data point.  Most appear to be some value on a continuous spectrum of possible values.  The others I've labeled in the following categories.  As is mentioned from the documentation, 11 are categorical [B30, B38, D114, D116, D117, D120, D126, D63, D64, D66, D68].  But there is a distinct difference in that D_63 and D_64 contain \"label\" values and not \"1.0, 2.0, etc\" values:</p>\n<p>\"category\" - B_30, B_38, D_66, D_68, D_114, D_116, D_117, D_120, D_126<br>\n\"boolean\" - B_31<br>\n\"checkbox?\" - D_87<br>\n\"label\" - D_63, D_64</p>\n<h1>1</h1>\n<p>Would it be accurate to describe B_31 as a boolean value, and D_87 some equivalent of \"if the checkbox is checked then you see 1.0, otherwise you see empty\"?  Just curious if someone could elaborate on what the difference is between B_31 and D_87 and / or whether it could matter when it comes to how you model that data?  I'm imagining the scenario could be that B_31 might've been some radio button group on an application form where it was a required field when the person applied, so everyone had to choose either \"true\" or \"false\".  Whereas maybe D_87 could've been some checkbox on an application form where the person either did check it \"1.0\" or didn't check it \"empty\"?  I'm assuming here there was a deliberate reason B_31 and D_87 have different possible values though they could both have been represented to us by a boolean data type.</p>\n<h1>2</h1>\n<p>Next, the \"category\" and \"label\" data types have several values which may or may not mean what I think they mean.  And that may or may not matter.</p>\n<p>\"empty\", \"0.0\", \"-1\", \"-1.0\".  I wanted to see if someone knows enough to be able to differentiate between what each of these values mean?  My assumption otherwise is that any record, with a \"1.0, 2.0, 3.0, CL, CO, O, R, U, etc\" value, means that the record is \"part of that category\".</p>\n<p>This is just a guess to get the conversation started, but my current assumption is maybe this is what they each mean:</p>\n<p>\"empty\" - there is no information about this category, for this record<br>\n\"0.0\" - this record is maybe part of some \"other\" category besides what is listed as possible categories<br>\n\"-1 or -1.0\" - this record is maybe \"n/a\" when it comes to this category?</p>\n<h1>3</h1>\n<p>Third, I'd like to know if there was a reason D_63 and D_64 have \"label\" values instead of the more ambiguous \"1.0, 2.0, 3.0, etc\" values.  And if there was, whether it could matter in terms of how I model this data?</p>",
      "rawMarkdown": "Sorry of this is already covered.  If so please just refer me to the applicable thread.\n\nFirst thing I wanted to do was analyze what kind of data is stored for each data point.  Most appear to be some value on a continuous spectrum of possible values.  The others I've labeled in the following categories.  As is mentioned from the documentation, 11 are categorical [B30, B38, D114, D116, D117, D120, D126, D63, D64, D66, D68].  But there is a distinct difference in that D_63 and D_64 contain \"label\" values and not \"1.0, 2.0, etc\" values:\n\n\"category\" - B_30, B_38, D_66, D_68, D_114, D_116, D_117, D_120, D_126\n\"boolean\" - B_31\n\"checkbox?\" - D_87\n\"label\" - D_63, D_64\n\n#1\nWould it be accurate to describe B_31 as a boolean value, and D_87 some equivalent of \"if the checkbox is checked then you see 1.0, otherwise you see empty\"?  Just curious if someone could elaborate on what the difference is between B_31 and D_87 and / or whether it could matter when it comes to how you model that data?  I'm imagining the scenario could be that B_31 might've been some radio button group on an application form where it was a required field when the person applied, so everyone had to choose either \"true\" or \"false\".  Whereas maybe D_87 could've been some checkbox on an application form where the person either did check it \"1.0\" or didn't check it \"empty\"?  I'm assuming here there was a deliberate reason B_31 and D_87 have different possible values though they could both have been represented to us by a boolean data type.\n\n#2\nNext, the \"category\" and \"label\" data types have several values which may or may not mean what I think they mean.  And that may or may not matter.\n\n\"empty\", \"0.0\", \"-1\", \"-1.0\".  I wanted to see if someone knows enough to be able to differentiate between what each of these values mean?  My assumption otherwise is that any record, with a \"1.0, 2.0, 3.0, CL, CO, O, R, U, etc\" value, means that the record is \"part of that category\".\n\nThis is just a guess to get the conversation started, but my current assumption is maybe this is what they each mean:\n\n\"empty\" - there is no information about this category, for this record\n\"0.0\" - this record is maybe part of some \"other\" category besides what is listed as possible categories\n\"-1 or -1.0\" - this record is maybe \"n/a\" when it comes to this category?\n\n#3\nThird, I'd like to know if there was a reason D_63 and D_64 have \"label\" values instead of the more ambiguous \"1.0, 2.0, 3.0, etc\" values.  And if there was, whether it could matter in terms of how I model this data?\n\n",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1822096": "Sorry of this is already covered.  If so please just refer me to the applicable thread.\n\nFirst thing I wanted to do was analyze what kind of data is stored for each data point.  Most appear to be some value on a continuous spectrum of possible values.  The others I've labeled in the following categories.  As is mentioned from the documentation, 11 are categorical [B30, B38, D114, D116, D117, D120, D126, D63, D64, D66, D68].  But there is a distinct difference in that D_63 and D_64 contain \"label\" values and not \"1.0, 2.0, etc\" values:\n\n\"category\" - B_30, B_38, D_66, D_68, D_114, D_116, D_117, D_120, D_126\n\"boolean\" - B_31\n\"checkbox?\" - D_87\n\"label\" - D_63, D_64\n\n#1\nWould it be accurate to describe B_31 as a boolean value, and D_87 some equivalent of \"if the checkbox is checked then you see 1.0, otherwise you see empty\"?  Just curious if someone could elaborate on what the difference is between B_31 and D_87 and / or whether it could matter when it comes to how you model that data?  I'm imagining the scenario could be that B_31 might've been some radio button group on an application form where it was a required field when the person applied, so everyone had to choose either \"true\" or \"false\".  Whereas maybe D_87 could've been some checkbox on an application form where the person either did check it \"1.0\" or didn't check it \"empty\"?  I'm assuming here there was a deliberate reason B_31 and D_87 have different possible values though they could both have been represented to us by a boolean data type.\n\n#2\nNext, the \"category\" and \"label\" data types have several values which may or may not mean what I think they mean.  And that may or may not matter.\n\n\"empty\", \"0.0\", \"-1\", \"-1.0\".  I wanted to see if someone knows enough to be able to differentiate between what each of these values mean?  My assumption otherwise is that any record, with a \"1.0, 2.0, 3.0, CL, CO, O, R, U, etc\" value, means that the record is \"part of that category\".\n\nThis is just a guess to get the conversation started, but my current assumption is maybe this is what they each mean:\n\n\"empty\" - there is no information about this category, for this record\n\"0.0\" - this record is maybe part of some \"other\" category besides what is listed as possible categories\n\"-1 or -1.0\" - this record is maybe \"n/a\" when it comes to this category?\n\n#3\nThird, I'd like to know if there was a reason D_63 and D_64 have \"label\" values instead of the more ambiguous \"1.0, 2.0, 3.0, etc\" values.  And if there was, whether it could matter in terms of how I model this data?\n\n"
  }
}