{
  "id": 477365,
  "title": "Aggregation Function Ideas?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/477365",
  "author_name": "",
  "post_date": "2024-02-15T19:20:43.959255800Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm trying to figure out the best way to aggregate the data from depth 1 and 2. </p>\n<p>I've seen some notebooks using max() on every feature. While max() makes sense for some features, there are so many different features that it seems hard to have one function for everything.</p>\n<p>Wondering if anyone had any insight?</p>",
  "messages": [
    {
      "id": "2654018",
      "postDate": "02/15/2024 19:20:43",
      "content": "<p>I'm trying to figure out the best way to aggregate the data from depth 1 and 2. </p>\n<p>I've seen some notebooks using max() on every feature. While max() makes sense for some features, there are so many different features that it seems hard to have one function for everything.</p>\n<p>Wondering if anyone had any insight?</p>",
      "rawMarkdown": "I'm trying to figure out the best way to aggregate the data from depth 1 and 2. \n\nI've seen some notebooks using max() on every feature. While max() makes sense for some features, there are so many different features that it seems hard to have one function for everything.\n\nWondering if anyone had any insight?",
      "votes": null
    },
    {
      "id": "2707183",
      "postDate": "03/20/2024 10:47:03",
      "content": "<p>You can try to add any functions you want, nobody can know in advance which will work better. But there are a few problems to keep in mind:</p>\n<ol>\n<li>Handle the memory use, as any additional calculations will need extra RAM. </li>\n<li>Dimensionality - you will need to filter out unuseful columns.</li>\n<li>Model timing and results, as higher dimensionality will need extra time and it can add noise if columns are not useful.</li>\n</ol>\n<p>It's a time consuming try-test loop to findout the golden formula</p>",
      "rawMarkdown": "You can try to add any functions you want, nobody can know in advance which will work better. But there are a few problems to keep in mind:\n1. Handle the memory use, as any additional calculations will need extra RAM. \n2. Dimensionality - you will need to filter out unuseful columns.\n3. Model timing and results, as higher dimensionality will need extra time and it can add noise if columns are not useful.\n\nIt's a time consuming try-test loop to findout the golden formula",
      "votes": null
    },
    {
      "id": "2707269",
      "postDate": "03/20/2024 12:11:31",
      "content": "<p>my advice would be to look at particular case_ids.  start by filtering for 1 random good guy (0) and 1 random bad guy (1).  try to look at all the fields from each table.</p>\n<p>try to understand the person and their application details …how much money they make, how much the loan is worth, their family situation … was their any difference?  maybe the bad guy have not been approved in the first place?  why? and create features that help answer that why.</p>",
      "rawMarkdown": "my advice would be to look at particular case_ids.  start by filtering for 1 random good guy (0) and 1 random bad guy (1).  try to look at all the fields from each table.\n\ntry to understand the person and their application details ...how much money they make, how much the loan is worth, their family situation ... was their any difference?  maybe the bad guy have not been approved in the first place?  why? and create features that help answer that why.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2707183,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "03/20/2024 10:47:03",
      "content": "<p>You can try to add any functions you want, nobody can know in advance which will work better. But there are a few problems to keep in mind:</p>\n<ol>\n<li>Handle the memory use, as any additional calculations will need extra RAM. </li>\n<li>Dimensionality - you will need to filter out unuseful columns.</li>\n<li>Model timing and results, as higher dimensionality will need extra time and it can add noise if columns are not useful.</li>\n</ol>\n<p>It's a time consuming try-test loop to findout the golden formula</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2707269,
      "author_name": "romandovega",
      "author_url": "",
      "post_date": "03/20/2024 12:11:31",
      "content": "<p>my advice would be to look at particular case_ids.  start by filtering for 1 random good guy (0) and 1 random bad guy (1).  try to look at all the fields from each table.</p>\n<p>try to understand the person and their application details …how much money they make, how much the loan is worth, their family situation … was their any difference?  maybe the bad guy have not been approved in the first place?  why? and create features that help answer that why.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2654018": "I'm trying to figure out the best way to aggregate the data from depth 1 and 2. \n\nI've seen some notebooks using max() on every feature. While max() makes sense for some features, there are so many different features that it seems hard to have one function for everything.\n\nWondering if anyone had any insight?",
    "2707183": "You can try to add any functions you want, nobody can know in advance which will work better. But there are a few problems to keep in mind:\n1. Handle the memory use, as any additional calculations will need extra RAM. \n2. Dimensionality - you will need to filter out unuseful columns.\n3. Model timing and results, as higher dimensionality will need extra time and it can add noise if columns are not useful.\n\nIt's a time consuming try-test loop to findout the golden formula",
    "2707269": "my advice would be to look at particular case_ids.  start by filtering for 1 random good guy (0) and 1 random bad guy (1).  try to look at all the fields from each table.\n\ntry to understand the person and their application details ...how much money they make, how much the loan is worth, their family situation ... was their any difference?  maybe the bad guy have not been approved in the first place?  why? and create features that help answer that why."
  },
  "source": "meta"
}