{
  "id": 330449,
  "title": "Could someone explain me what is magic feature?",
  "url": "/competitions/amex-default-prediction/discussion/330449",
  "author_name": "",
  "post_date": "2022-06-12T12:09:43.245457800Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I was going through previous tabular competition solutions and the authors always refer to 'Magic feature'.  I want to know whether this feature already available in dataset or engineered feature?<br>\nCould someone explain me pls? Thanks in advance </p>",
  "messages": [
    {
      "id": "1818223",
      "postDate": "06/12/2022 12:09:43",
      "content": "<p>I was going through previous tabular competition solutions and the authors always refer to 'Magic feature'.  I want to know whether this feature already available in dataset or engineered feature?<br>\nCould someone explain me pls? Thanks in advance </p>",
      "rawMarkdown": "I was going through previous tabular competition solutions and the authors always refer to 'Magic feature'.  I want to know whether this feature already available in dataset or engineered feature?\nCould someone explain me pls? Thanks in advance",
      "votes": null
    },
    {
      "id": "1818239",
      "postDate": "06/12/2022 12:21:46",
      "content": "<p>When one feature has an unusual predictive power in comparison to others in the data table, it is referred to as a magic feature. Some tabular competitions do have such features, using them effectively can boost the scores drastically.</p>",
      "rawMarkdown": "When one feature has an unusual predictive power in comparison to others in the data table, it is referred to as a magic feature. Some tabular competitions do have such features, using them effectively can boost the scores drastically.",
      "votes": null
    },
    {
      "id": "1818248",
      "postDate": "06/12/2022 12:26:29",
      "content": "<p>Thank you for your explanation <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "rawMarkdown": "Thank you for your explanation @ravi20076",
      "votes": null
    },
    {
      "id": "1818358",
      "postDate": "06/12/2022 15:06:20",
      "content": "<p>For example in this competition, <code>P_2_last</code> has greater predictive power than others, so we may consider this a magic feature.</p>",
      "rawMarkdown": "For example in this competition, `P_2_last` has greater predictive power than others, so we may consider this a magic feature.",
      "votes": null
    },
    {
      "id": "1818404",
      "postDate": "06/12/2022 16:44:50",
      "content": "<p>In my experience, \"magic\" usually refers to non-obvious engineered features which give an edge over previous benchmark results that use only the raw data or simple feature engineering. \"Magic\" also often involves exploiting data artifacts or leaks that are unlikely to generalize in real-world modeling. Examples include frequency counts of numerical feature values in an <a href=\"https://www.kaggle.com/code/cdeotte/200-magical-models-santander-0-920\" target=\"_blank\">old Santander competition</a> and the Instant Gratification competition that cobbled together 100s of distinct synthetic datasets into a combined file that included a categorical feature used to re-separate them. </p>\n<p>The \"magic\" of these features is not that they are necessarily highly predictive, but that discovering and fully utilizing them is often the key differentiator for which teams take top competition spots. The usage of the term is messy and broad, but I do think there's usually a tie-in to data preparation issues that leave behind artificial traces of predictive signal. </p>\n<p>Edit: for emphasis, it's \"magic\" because these features are conjured from competition data quirks and wouldn't be found in a proper industry model.</p>",
      "rawMarkdown": "In my experience, \"magic\" usually refers to non-obvious engineered features which give an edge over previous benchmark results that use only the raw data or simple feature engineering. \"Magic\" also often involves exploiting data artifacts or leaks that are unlikely to generalize in real-world modeling. Examples include frequency counts of numerical feature values in an [old Santander competition](https://www.kaggle.com/code/cdeotte/200-magical-models-santander-0-920) and the Instant Gratification competition that cobbled together 100s of distinct synthetic datasets into a combined file that included a categorical feature used to re-separate them. \n\nThe \"magic\" of these features is not that they are necessarily highly predictive, but that discovering and fully utilizing them is often the key differentiator for which teams take top competition spots. The usage of the term is messy and broad, but I do think there's usually a tie-in to data preparation issues that leave behind artificial traces of predictive signal. \n\nEdit: for emphasis, it's \"magic\" because these features are conjured from competition data quirks and wouldn't be found in a proper industry model.",
      "votes": null
    },
    {
      "id": "1818652",
      "postDate": "06/13/2022 03:16:50",
      "content": "<p>Thank you for the detailed explanation <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> </p>",
      "rawMarkdown": "Thank you for the detailed explanation @aquatic",
      "votes": null
    },
    {
      "id": "1818653",
      "postDate": "06/13/2022 03:17:34",
      "content": "<p>Thanks you <a href=\"https://www.kaggle.com/susnato\" target=\"_blank\">@susnato</a> </p>",
      "rawMarkdown": "Thanks you @susnato",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1818239,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "06/12/2022 12:21:46",
      "content": "<p>When one feature has an unusual predictive power in comparison to others in the data table, it is referred to as a magic feature. Some tabular competitions do have such features, using them effectively can boost the scores drastically.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1818248,
          "author_name": "thamotharan",
          "author_url": "",
          "post_date": "06/12/2022 12:26:29",
          "content": "<p>Thank you for your explanation <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1818358,
          "author_name": "susnato",
          "author_url": "",
          "post_date": "06/12/2022 15:06:20",
          "content": "<p>For example in this competition, <code>P_2_last</code> has greater predictive power than others, so we may consider this a magic feature.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1818653,
          "author_name": "thamotharan",
          "author_url": "",
          "post_date": "06/13/2022 03:17:34",
          "content": "<p>Thanks you <a href=\"https://www.kaggle.com/susnato\" target=\"_blank\">@susnato</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1818404,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "06/12/2022 16:44:50",
      "content": "<p>In my experience, \"magic\" usually refers to non-obvious engineered features which give an edge over previous benchmark results that use only the raw data or simple feature engineering. \"Magic\" also often involves exploiting data artifacts or leaks that are unlikely to generalize in real-world modeling. Examples include frequency counts of numerical feature values in an <a href=\"https://www.kaggle.com/code/cdeotte/200-magical-models-santander-0-920\" target=\"_blank\">old Santander competition</a> and the Instant Gratification competition that cobbled together 100s of distinct synthetic datasets into a combined file that included a categorical feature used to re-separate them. </p>\n<p>The \"magic\" of these features is not that they are necessarily highly predictive, but that discovering and fully utilizing them is often the key differentiator for which teams take top competition spots. The usage of the term is messy and broad, but I do think there's usually a tie-in to data preparation issues that leave behind artificial traces of predictive signal. </p>\n<p>Edit: for emphasis, it's \"magic\" because these features are conjured from competition data quirks and wouldn't be found in a proper industry model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1818652,
          "author_name": "thamotharan",
          "author_url": "",
          "post_date": "06/13/2022 03:16:50",
          "content": "<p>Thank you for the detailed explanation <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1818223": "I was going through previous tabular competition solutions and the authors always refer to 'Magic feature'.  I want to know whether this feature already available in dataset or engineered feature?\nCould someone explain me pls? Thanks in advance",
    "1818239": "When one feature has an unusual predictive power in comparison to others in the data table, it is referred to as a magic feature. Some tabular competitions do have such features, using them effectively can boost the scores drastically.",
    "1818248": "Thank you for your explanation @ravi20076",
    "1818358": "For example in this competition, `P_2_last` has greater predictive power than others, so we may consider this a magic feature.",
    "1818404": "In my experience, \"magic\" usually refers to non-obvious engineered features which give an edge over previous benchmark results that use only the raw data or simple feature engineering. \"Magic\" also often involves exploiting data artifacts or leaks that are unlikely to generalize in real-world modeling. Examples include frequency counts of numerical feature values in an [old Santander competition](https://www.kaggle.com/code/cdeotte/200-magical-models-santander-0-920) and the Instant Gratification competition that cobbled together 100s of distinct synthetic datasets into a combined file that included a categorical feature used to re-separate them. \n\nThe \"magic\" of these features is not that they are necessarily highly predictive, but that discovering and fully utilizing them is often the key differentiator for which teams take top competition spots. The usage of the term is messy and broad, but I do think there's usually a tie-in to data preparation issues that leave behind artificial traces of predictive signal. \n\nEdit: for emphasis, it's \"magic\" because these features are conjured from competition data quirks and wouldn't be found in a proper industry model.",
    "1818652": "Thank you for the detailed explanation @aquatic",
    "1818653": "Thanks you @susnato"
  },
  "source": "meta"
}