{
  "id": 493689,
  "title": "How to set dtype for unspecified transforms?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/493689",
  "author_name": "",
  "post_date": "2024-04-14T14:36:37.650610700Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The competition dataset categorizes data using transformation groups:</p>\n<ol>\n<li><code>P - Transform DPD</code> (Days past due) and <code>A - Transform amount</code> (numerical columns)</li>\n<li><code>M - Masking categories</code> (categorical columns)</li>\n<li><code>D - Transform date</code> (temporal columns)</li>\n<li><code>T - Unspecified Transform</code> and <code>L - Unspecified Transform</code></li>\n</ol>\n<p>When applying aggregation, I'm uncertain about which methods to use for the unspecified columns. Without fishing for your secret sauce to data preprocessing, how should I approach my issue?</p>",
  "messages": [
    {
      "id": "2751798",
      "postDate": "04/14/2024 14:36:37",
      "content": "<p>The competition dataset categorizes data using transformation groups:</p>\n<ol>\n<li><code>P - Transform DPD</code> (Days past due) and <code>A - Transform amount</code> (numerical columns)</li>\n<li><code>M - Masking categories</code> (categorical columns)</li>\n<li><code>D - Transform date</code> (temporal columns)</li>\n<li><code>T - Unspecified Transform</code> and <code>L - Unspecified Transform</code></li>\n</ol>\n<p>When applying aggregation, I'm uncertain about which methods to use for the unspecified columns. Without fishing for your secret sauce to data preprocessing, how should I approach my issue?</p>",
      "rawMarkdown": "The competition dataset categorizes data using transformation groups:\n\n1. `P - Transform DPD` (Days past due) and `A - Transform amount` (numerical columns)\n2. `M - Masking categories` (categorical columns)\n3. `D - Transform date` (temporal columns)\n4. `T - Unspecified Transform` and `L - Unspecified Transform `\n\nWhen applying aggregation, I'm uncertain about which methods to use for the unspecified columns. Without fishing for your secret sauce to data preprocessing, how should I approach my issue?",
      "votes": null
    },
    {
      "id": "2752054",
      "postDate": "04/14/2024 17:19:04",
      "content": "<p>one way is to divide the features further as numeric and categorical type, For example, we can run col[-1] == 'T' to check it is dtype along with it we also check col.dtype == Int (just for example), so we have two types or more types of calls as type T float or type T category. then we treat them separately, just the same way as we treat P and M separately.</p>",
      "rawMarkdown": "one way is to divide the features further as numeric and categorical type, For example, we can run col[-1] == 'T' to check it is dtype along with it we also check col.dtype == Int (just for example), so we have two types or more types of calls as type T float or type T category. then we treat them separately, just the same way as we treat P and M separately.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2752054,
      "author_name": "shreyas9181",
      "author_url": "",
      "post_date": "04/14/2024 17:19:04",
      "content": "<p>one way is to divide the features further as numeric and categorical type, For example, we can run col[-1] == 'T' to check it is dtype along with it we also check col.dtype == Int (just for example), so we have two types or more types of calls as type T float or type T category. then we treat them separately, just the same way as we treat P and M separately.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2751798": "The competition dataset categorizes data using transformation groups:\n\n1. `P - Transform DPD` (Days past due) and `A - Transform amount` (numerical columns)\n2. `M - Masking categories` (categorical columns)\n3. `D - Transform date` (temporal columns)\n4. `T - Unspecified Transform` and `L - Unspecified Transform `\n\nWhen applying aggregation, I'm uncertain about which methods to use for the unspecified columns. Without fishing for your secret sauce to data preprocessing, how should I approach my issue?",
    "2752054": "one way is to divide the features further as numeric and categorical type, For example, we can run col[-1] == 'T' to check it is dtype along with it we also check col.dtype == Int (just for example), so we have two types or more types of calls as type T float or type T category. then we treat them separately, just the same way as we treat P and M separately."
  },
  "source": "meta"
}