{
  "id": 329436,
  "title": "What can we still do when we have no information about the feaures",
  "url": "/competitions/amex-default-prediction/discussion/329436",
  "author_name": "",
  "post_date": "2022-06-06T16:39:44.902249800Z",
  "votes": 61,
  "comment_count": 7,
  "views": 0,
  "content": "<h3>What can we still do when we have no information about the features</h3>\n<p>Hi Everyone,</p>\n<p>The dataset for this competition is anonymized, Feature engineering won't be easy. <br>\nThere already had been <a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/overview\" target=\"_blank\">some</a> <a href=\"https://www.kaggle.com/competitions/santander-customer-transaction-prediction/overview\" target=\"_blank\">past</a> <a href=\"https://www.kaggle.com/c/lish-moa\" target=\"_blank\">competitions</a> <a href=\"https://www.kaggle.com/c/jane-street-market-prediction\" target=\"_blank\">on</a> kaggle that also had anonymized features, so we can learn from them what should and shouldn't we try for this competition. </p>\n<p>So making things simple, there are three different approaches that had proven to work. </p>\n<h4>Heavy EDA and \"reverse engineering\" the data</h4>\n<p>This is obviously the preferred approach and mostly will yield the top performance overall when done correctly. This was, by the way, the approach used in the famous <a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111284\" target=\"_blank\">IEEE top solution</a> and On <a href=\"https://www.kaggle.com/competitions/santander-customer-transaction-prediction/discussion/88909\" target=\"_blank\">Santander Customer Transaction Prediction\n</a>. Both are monumental famous competitions on Kaggle and the solutions developed in those competitions had ripple effects throughout the entire data-science field worldwide (More on that in the next section about feature engineering).</p>\n<p>So, in this approach: the main workhorse we use is EDA. EDA EDA and EDA.</p>\n<p><strong>What are we searching for?</strong></p>\n<p>It depends, for example: on IEEE the main goal was to identify unique users so a UID had to be developed. To do this, some had used adversarial validation and checked the top features and some others simply performed calculations based on the columns and found a way to extract one of the other columns. [I searched for this notebook now, but I couldn't find it. I remember that it was incredible and I learned a lot from it so if anyone has it: share it with us please!]</p>\n<p>It comes down to looking at the data and coming up with ideas and theories.</p>\n<h4>\"Blind\" Feature Engineering</h4>\n<p>OK. Reverse engineering the data is nice and all. But at the end of the day: asking \"to have the right ideas\" is like asking to win the lottery. We can keep on trying forever to find the \"magic\" and end up simply not lucky enough.</p>\n<p>There are multiple methods we can apply for creating \"blind\" features. </p>\n<p><strong>Simple Feature Engineering</strong></p>\n<p>I highly advise anyone on Kaggle to check out this [<a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575\" target=\"_blank\">https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575</a>] post by Chris. <br>\nThis is one of the most upvoted posts all over kaggle. [And is a famous post out of kaggle as well: Data scientists refer to it all the time]</p>\n<p>It has information about </p>\n<ul>\n<li>Combining / Transforming / Interaction</li>\n<li>Categorical Features</li>\n<li>Frequency Encoding</li>\n<li>Aggregations / Group Statistics</li>\n</ul>\n<p><strong>More about Aggregations / Group Statistics</strong></p>\n<p>Again from Chris we got <a href=\"https://www.kaggle.com/code/cdeotte/xgb-fraud-with-magic-0-9600/notebook\" target=\"_blank\">this</a> notebook.<br>\nIt has important information about aggregation features you might want to try \"blind\" and then perform some feature selection on the outcome you get. <br>\nThis way you will be able to get meaningful features without knowing the meaning of the columns themselves. </p>\n<p><strong>Bruteforce Feature Engineering</strong></p>\n<p><a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111245\" target=\"_blank\">original post</a> - Credit to the original author! </p>\n<p>If you just want to create all features interactions and then feature select some of them, here is a code to do so:</p>\n<pre><code>for i in range(len(cols)):\n    for j in range(i + 1, len(cols)):\n        X['{}_div_{}'.format(cols[i], cols[j])] = X[cols[i]] * X[cols[j]]\n        X['{}_mul_{}'.format(cols[i], cols[j])] = X[cols[i]] * X[cols[j]]\n        X['{}_sub1_{}'.format(cols[i], cols[j])] = X[cols[i]] - X[cols[j]]\n        X['{}_sub2_{}'.format(cols[j], cols[i])] = X[cols[j]] - X[cols[i]]\n</code></pre>\n<p>You can them compute a correlation matrix for all new generated features<br>\nThen removing all the features with a high correlation. Keeping those that correlate with target value better.</p>\n<pre><code>bruteforce_cols = [col for col in X.columns if '_mul_' in col or '_sub_' in col of '_div_' in col]\ncorr_matrix = X[bruteforce_cols].corr()\n\nto_drop = list()\n\nfor i in range(1, len(corr_matrix)):\n    for j in range(i):\n        if corr_matrix.iloc[i, j] &gt;= 0.98:\n            if abs(pd.concat([X[corr_matrix.index[i]], y], axis=1).corr().iloc[0][1]) &gt; abs(pd.concat([X[corr_matrix.columns[j]], y], axis=1).corr().iloc[0][1]):\n                to_drop.append(corr_matrix.columns[j])\n            else:\n                to_drop.append(corr_matrix.index[i])\n\nto_drop = list(set(to_drop))\n</code></pre>\n<p><strong>Stacking</strong></p>\n<p>One more trick we can apply that we don't have to know the meaning of the columns for is to simply use \"heavy stacking\". <br>\nThis was done so many times on Kaggle so we won't need to dive into that again. </p>\n<p><strong>Pseudolabeling</strong></p>\n<p>The main idea is to use a model for labeling an unlabeled dataset and then use this dataset in addition to the training set. <br>\nThere isn't much to say about Pseudolabeling as well. This also had been used so many times on Kaggle and when done properly it most probably improves performance.</p>\n<h4>Representation Learning</h4>\n<p><strong>DEA and Ladder Networks</strong></p>\n<ul>\n<li><p>DEA (Denoising auto encoders) were used in the MOA competition and on TPS competitions in the past. </p></li>\n<li><p>Ladder Networks (<a href=\"https://arxiv.org/abs/1507.02672):\" target=\"_blank\">https://arxiv.org/abs/1507.02672):</a> Using denoising autoencoder in a unique fashion. This might be an interesting approach to try as well as simply vanilla DEA.</p></li>\n</ul>\n<p><strong>Unique Neural network architectures</strong></p>\n<ul>\n<li>If you don't know about this: Check out the \"MOA Net\" that was introduced on the Mechanisms of action competition's 2nd place winner. <br>\nThis might sound crazy at first but this network uses convolutions for tabular data. <br>\nNeedless to say, by its 2nd place we can safely assume that it works well.</li>\n</ul>\n<p>What about you? <br>\nDo you know any other approach or trick we can use when we have no information about the features?</p>",
  "messages": [
    {
      "id": "1813247",
      "postDate": "06/06/2022 16:39:44",
      "content": "<h3>What can we still do when we have no information about the features</h3>\n<p>Hi Everyone,</p>\n<p>The dataset for this competition is anonymized, Feature engineering won't be easy. <br>\nThere already had been <a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/overview\" target=\"_blank\">some</a> <a href=\"https://www.kaggle.com/competitions/santander-customer-transaction-prediction/overview\" target=\"_blank\">past</a> <a href=\"https://www.kaggle.com/c/lish-moa\" target=\"_blank\">competitions</a> <a href=\"https://www.kaggle.com/c/jane-street-market-prediction\" target=\"_blank\">on</a> kaggle that also had anonymized features, so we can learn from them what should and shouldn't we try for this competition. </p>\n<p>So making things simple, there are three different approaches that had proven to work. </p>\n<h4>Heavy EDA and \"reverse engineering\" the data</h4>\n<p>This is obviously the preferred approach and mostly will yield the top performance overall when done correctly. This was, by the way, the approach used in the famous <a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111284\" target=\"_blank\">IEEE top solution</a> and On <a href=\"https://www.kaggle.com/competitions/santander-customer-transaction-prediction/discussion/88909\" target=\"_blank\">Santander Customer Transaction Prediction\n</a>. Both are monumental famous competitions on Kaggle and the solutions developed in those competitions had ripple effects throughout the entire data-science field worldwide (More on that in the next section about feature engineering).</p>\n<p>So, in this approach: the main workhorse we use is EDA. EDA EDA and EDA.</p>\n<p><strong>What are we searching for?</strong></p>\n<p>It depends, for example: on IEEE the main goal was to identify unique users so a UID had to be developed. To do this, some had used adversarial validation and checked the top features and some others simply performed calculations based on the columns and found a way to extract one of the other columns. [I searched for this notebook now, but I couldn't find it. I remember that it was incredible and I learned a lot from it so if anyone has it: share it with us please!]</p>\n<p>It comes down to looking at the data and coming up with ideas and theories.</p>\n<h4>\"Blind\" Feature Engineering</h4>\n<p>OK. Reverse engineering the data is nice and all. But at the end of the day: asking \"to have the right ideas\" is like asking to win the lottery. We can keep on trying forever to find the \"magic\" and end up simply not lucky enough.</p>\n<p>There are multiple methods we can apply for creating \"blind\" features. </p>\n<p><strong>Simple Feature Engineering</strong></p>\n<p>I highly advise anyone on Kaggle to check out this [<a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575\" target=\"_blank\">https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575</a>] post by Chris. <br>\nThis is one of the most upvoted posts all over kaggle. [And is a famous post out of kaggle as well: Data scientists refer to it all the time]</p>\n<p>It has information about </p>\n<ul>\n<li>Combining / Transforming / Interaction</li>\n<li>Categorical Features</li>\n<li>Frequency Encoding</li>\n<li>Aggregations / Group Statistics</li>\n</ul>\n<p><strong>More about Aggregations / Group Statistics</strong></p>\n<p>Again from Chris we got <a href=\"https://www.kaggle.com/code/cdeotte/xgb-fraud-with-magic-0-9600/notebook\" target=\"_blank\">this</a> notebook.<br>\nIt has important information about aggregation features you might want to try \"blind\" and then perform some feature selection on the outcome you get. <br>\nThis way you will be able to get meaningful features without knowing the meaning of the columns themselves. </p>\n<p><strong>Bruteforce Feature Engineering</strong></p>\n<p><a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111245\" target=\"_blank\">original post</a> - Credit to the original author! </p>\n<p>If you just want to create all features interactions and then feature select some of them, here is a code to do so:</p>\n<pre><code>for i in range(len(cols)):\n    for j in range(i + 1, len(cols)):\n        X['{}_div_{}'.format(cols[i], cols[j])] = X[cols[i]] * X[cols[j]]\n        X['{}_mul_{}'.format(cols[i], cols[j])] = X[cols[i]] * X[cols[j]]\n        X['{}_sub1_{}'.format(cols[i], cols[j])] = X[cols[i]] - X[cols[j]]\n        X['{}_sub2_{}'.format(cols[j], cols[i])] = X[cols[j]] - X[cols[i]]\n</code></pre>\n<p>You can them compute a correlation matrix for all new generated features<br>\nThen removing all the features with a high correlation. Keeping those that correlate with target value better.</p>\n<pre><code>bruteforce_cols = [col for col in X.columns if '_mul_' in col or '_sub_' in col of '_div_' in col]\ncorr_matrix = X[bruteforce_cols].corr()\n\nto_drop = list()\n\nfor i in range(1, len(corr_matrix)):\n    for j in range(i):\n        if corr_matrix.iloc[i, j] &gt;= 0.98:\n            if abs(pd.concat([X[corr_matrix.index[i]], y], axis=1).corr().iloc[0][1]) &gt; abs(pd.concat([X[corr_matrix.columns[j]], y], axis=1).corr().iloc[0][1]):\n                to_drop.append(corr_matrix.columns[j])\n            else:\n                to_drop.append(corr_matrix.index[i])\n\nto_drop = list(set(to_drop))\n</code></pre>\n<p><strong>Stacking</strong></p>\n<p>One more trick we can apply that we don't have to know the meaning of the columns for is to simply use \"heavy stacking\". <br>\nThis was done so many times on Kaggle so we won't need to dive into that again. </p>\n<p><strong>Pseudolabeling</strong></p>\n<p>The main idea is to use a model for labeling an unlabeled dataset and then use this dataset in addition to the training set. <br>\nThere isn't much to say about Pseudolabeling as well. This also had been used so many times on Kaggle and when done properly it most probably improves performance.</p>\n<h4>Representation Learning</h4>\n<p><strong>DEA and Ladder Networks</strong></p>\n<ul>\n<li><p>DEA (Denoising auto encoders) were used in the MOA competition and on TPS competitions in the past. </p></li>\n<li><p>Ladder Networks (<a href=\"https://arxiv.org/abs/1507.02672):\" target=\"_blank\">https://arxiv.org/abs/1507.02672):</a> Using denoising autoencoder in a unique fashion. This might be an interesting approach to try as well as simply vanilla DEA.</p></li>\n</ul>\n<p><strong>Unique Neural network architectures</strong></p>\n<ul>\n<li>If you don't know about this: Check out the \"MOA Net\" that was introduced on the Mechanisms of action competition's 2nd place winner. <br>\nThis might sound crazy at first but this network uses convolutions for tabular data. <br>\nNeedless to say, by its 2nd place we can safely assume that it works well.</li>\n</ul>\n<p>What about you? <br>\nDo you know any other approach or trick we can use when we have no information about the features?</p>",
      "rawMarkdown": "### What can we still do when we have no information about the features\n\nHi Everyone,\n\n\nThe dataset for this competition is anonymized, Feature engineering won't be easy. \nThere already had been [some](https://www.kaggle.com/competitions/ieee-fraud-detection/overview) [past](https://www.kaggle.com/competitions/santander-customer-transaction-prediction/overview) [competitions](https://www.kaggle.com/c/lish-moa) [on](https://www.kaggle.com/c/jane-street-market-prediction) kaggle that also had anonymized features, so we can learn from them what should and shouldn't we try for this competition. \n\n\nSo making things simple, there are three different approaches that had proven to work. \n\n\n#### Heavy EDA and \"reverse engineering\" the data\n\nThis is obviously the preferred approach and mostly will yield the top performance overall when done correctly. This was, by the way, the approach used in the famous [IEEE top solution](https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111284) and On [Santander Customer Transaction Prediction\n](https://www.kaggle.com/competitions/santander-customer-transaction-prediction/discussion/88909). Both are monumental famous competitions on Kaggle and the solutions developed in those competitions had ripple effects throughout the entire data-science field worldwide (More on that in the next section about feature engineering).\n\nSo, in this approach: the main workhorse we use is EDA. EDA EDA and EDA.\n\n**What are we searching for?**\n\nIt depends, for example: on IEEE the main goal was to identify unique users so a UID had to be developed. To do this, some had used adversarial validation and checked the top features and some others simply performed calculations based on the columns and found a way to extract one of the other columns. [I searched for this notebook now, but I couldn't find it. I remember that it was incredible and I learned a lot from it so if anyone has it: share it with us please!]\n\nIt comes down to looking at the data and coming up with ideas and theories.\n\n\n#### \"Blind\" Feature Engineering\n\nOK. Reverse engineering the data is nice and all. But at the end of the day: asking \"to have the right ideas\" is like asking to win the lottery. We can keep on trying forever to find the \"magic\" and end up simply not lucky enough.\n\nThere are multiple methods we can apply for creating \"blind\" features. \n\n**Simple Feature Engineering**\n\nI highly advise anyone on Kaggle to check out this [https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575] post by Chris. \nThis is one of the most upvoted posts all over kaggle. [And is a famous post out of kaggle as well: Data scientists refer to it all the time]\n\nIt has information about \n- Combining / Transforming / Interaction\n- Categorical Features\n- Frequency Encoding\n- Aggregations / Group Statistics\n\n\n**More about Aggregations / Group Statistics**\n\nAgain from Chris we got [this](https://www.kaggle.com/code/cdeotte/xgb-fraud-with-magic-0-9600/notebook) notebook.\nIt has important information about aggregation features you might want to try \"blind\" and then perform some feature selection on the outcome you get. \nThis way you will be able to get meaningful features without knowing the meaning of the columns themselves. \n\n\n**Bruteforce Feature Engineering**\n\n[original post](https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111245) - Credit to the original author! \n\nIf you just want to create all features interactions and then feature select some of them, here is a code to do so:\n\n```\nfor i in range(len(cols)):\n    for j in range(i + 1, len(cols)):\n        X['{}_div_{}'.format(cols[i], cols[j])] = X[cols[i]] * X[cols[j]]\n        X['{}_mul_{}'.format(cols[i], cols[j])] = X[cols[i]] * X[cols[j]]\n        X['{}_sub1_{}'.format(cols[i], cols[j])] = X[cols[i]] - X[cols[j]]\n        X['{}_sub2_{}'.format(cols[j], cols[i])] = X[cols[j]] - X[cols[i]]\n```\n        \nYou can them compute a correlation matrix for all new generated features\nThen removing all the features with a high correlation. Keeping those that correlate with target value better.\n\n```\nbruteforce_cols = [col for col in X.columns if '_mul_' in col or '_sub_' in col of '_div_' in col]\ncorr_matrix = X[bruteforce_cols].corr()\n\nto_drop = list()\n\nfor i in range(1, len(corr_matrix)):\n    for j in range(i):\n        if corr_matrix.iloc[i, j] >= 0.98:\n            if abs(pd.concat([X[corr_matrix.index[i]], y], axis=1).corr().iloc[0][1]) > abs(pd.concat([X[corr_matrix.columns[j]], y], axis=1).corr().iloc[0][1]):\n                to_drop.append(corr_matrix.columns[j])\n            else:\n                to_drop.append(corr_matrix.index[i])\n\nto_drop = list(set(to_drop))\n```\n\n\n**Stacking**\n\nOne more trick we can apply that we don't have to know the meaning of the columns for is to simply use \"heavy stacking\". \nThis was done so many times on Kaggle so we won't need to dive into that again. \n\n\n**Pseudolabeling**\n\nThe main idea is to use a model for labeling an unlabeled dataset and then use this dataset in addition to the training set. \nThere isn't much to say about Pseudolabeling as well. This also had been used so many times on Kaggle and when done properly it most probably improves performance.\n\n\n#### Representation Learning\n\n\n**DEA and Ladder Networks**\n\n- DEA (Denoising auto encoders) were used in the MOA competition and on TPS competitions in the past. \n\n- Ladder Networks (https://arxiv.org/abs/1507.02672): Using denoising autoencoder in a unique fashion. This might be an interesting approach to try as well as simply vanilla DEA.\n\n\n**Unique Neural network architectures**\n\n- If you don't know about this: Check out the \"MOA Net\" that was introduced on the Mechanisms of action competition's 2nd place winner. \nThis might sound crazy at first but this network uses convolutions for tabular data. \nNeedless to say, by its 2nd place we can safely assume that it works well.\n\n\nWhat about you? \nDo you know any other approach or trick we can use when we have no information about the features?",
      "votes": null
    },
    {
      "id": "1813361",
      "postDate": "06/06/2022 18:48:00",
      "content": "<p>Nice compile of techniques. I would first combine different variables or use some dimensionally reduction method<br>\nCheers!</p>",
      "rawMarkdown": "Nice compile of techniques. I would first combine different variables or use some dimensionally reduction method\nCheers!",
      "votes": null
    },
    {
      "id": "1813736",
      "postDate": "06/07/2022 07:14:01",
      "content": "<p>Great information. Thanks for sharing . Specially notebook from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . Very helpful…🙏</p>",
      "rawMarkdown": "Great information. Thanks for sharing . Specially notebook from @cdeotte . Very helpful...🙏",
      "votes": null
    },
    {
      "id": "1816022",
      "postDate": "06/09/2022 18:12:00",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1820703",
      "postDate": "06/14/2022 21:41:53",
      "content": "<p>Great topic!</p>",
      "rawMarkdown": "Great topic!",
      "votes": null
    },
    {
      "id": "1821572",
      "postDate": "06/15/2022 17:33:18",
      "content": "<p>Good summary. Autoencoders are extremly helpful especially working with high dimensional data like we have here.</p>",
      "rawMarkdown": "Good summary. Autoencoders are extremly helpful especially working with high dimensional data like we have here.",
      "votes": null
    },
    {
      "id": "1873119",
      "postDate": "07/27/2022 13:01:08",
      "content": "<p><a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> <code>X['{}_div_{}'.format(cols[i], cols[j])] = X[cols[i]] * X[cols[j]]</code> perhaps in this line there's a typo and instead of multiplication <code>*</code> there should be a division <code>/</code>?</p>",
      "rawMarkdown": "thedevastator `X['{}_div_{}'.format(cols[i], cols[j])] = X[cols[i]] * X[cols[j]]` perhaps in this line there's a typo and instead of multiplication `*` there should be a division `/`?",
      "votes": null
    },
    {
      "id": "1873183",
      "postDate": "07/27/2022 13:41:22",
      "content": "<p>Great topic! I'm new and would never have thought of this without reading. 🙂</p>",
      "rawMarkdown": "Great topic! I'm new and would never have thought of this without reading. 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1813361,
      "author_name": "paulojunqueira",
      "author_url": "",
      "post_date": "06/06/2022 18:48:00",
      "content": "<p>Nice compile of techniques. I would first combine different variables or use some dimensionally reduction method<br>\nCheers!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1813736,
      "author_name": "cid007",
      "author_url": "",
      "post_date": "06/07/2022 07:14:01",
      "content": "<p>Great information. Thanks for sharing . Specially notebook from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . Very helpful…🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1816022,
      "author_name": "duikei",
      "author_url": "",
      "post_date": "06/09/2022 18:12:00",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1820703,
      "author_name": "rodrigostallsikora",
      "author_url": "",
      "post_date": "06/14/2022 21:41:53",
      "content": "<p>Great topic!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1821572,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "06/15/2022 17:33:18",
      "content": "<p>Good summary. Autoencoders are extremly helpful especially working with high dimensional data like we have here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1873119,
      "author_name": "lisatirskikh",
      "author_url": "",
      "post_date": "07/27/2022 13:01:08",
      "content": "<p><a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> <code>X['{}_div_{}'.format(cols[i], cols[j])] = X[cols[i]] * X[cols[j]]</code> perhaps in this line there's a typo and instead of multiplication <code>*</code> there should be a division <code>/</code>?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1873183,
      "author_name": "bellapisani",
      "author_url": "",
      "post_date": "07/27/2022 13:41:22",
      "content": "<p>Great topic! I'm new and would never have thought of this without reading. 🙂</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1813247": "### What can we still do when we have no information about the features\n\nHi Everyone,\n\n\nThe dataset for this competition is anonymized, Feature engineering won't be easy. \nThere already had been [some](https://www.kaggle.com/competitions/ieee-fraud-detection/overview) [past](https://www.kaggle.com/competitions/santander-customer-transaction-prediction/overview) [competitions](https://www.kaggle.com/c/lish-moa) [on](https://www.kaggle.com/c/jane-street-market-prediction) kaggle that also had anonymized features, so we can learn from them what should and shouldn't we try for this competition. \n\n\nSo making things simple, there are three different approaches that had proven to work. \n\n\n#### Heavy EDA and \"reverse engineering\" the data\n\nThis is obviously the preferred approach and mostly will yield the top performance overall when done correctly. This was, by the way, the approach used in the famous [IEEE top solution](https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111284) and On [Santander Customer Transaction Prediction\n](https://www.kaggle.com/competitions/santander-customer-transaction-prediction/discussion/88909). Both are monumental famous competitions on Kaggle and the solutions developed in those competitions had ripple effects throughout the entire data-science field worldwide (More on that in the next section about feature engineering).\n\nSo, in this approach: the main workhorse we use is EDA. EDA EDA and EDA.\n\n**What are we searching for?**\n\nIt depends, for example: on IEEE the main goal was to identify unique users so a UID had to be developed. To do this, some had used adversarial validation and checked the top features and some others simply performed calculations based on the columns and found a way to extract one of the other columns. [I searched for this notebook now, but I couldn't find it. I remember that it was incredible and I learned a lot from it so if anyone has it: share it with us please!]\n\nIt comes down to looking at the data and coming up with ideas and theories.\n\n\n#### \"Blind\" Feature Engineering\n\nOK. Reverse engineering the data is nice and all. But at the end of the day: asking \"to have the right ideas\" is like asking to win the lottery. We can keep on trying forever to find the \"magic\" and end up simply not lucky enough.\n\nThere are multiple methods we can apply for creating \"blind\" features. \n\n**Simple Feature Engineering**\n\nI highly advise anyone on Kaggle to check out this [https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575] post by Chris. \nThis is one of the most upvoted posts all over kaggle. [And is a famous post out of kaggle as well: Data scientists refer to it all the time]\n\nIt has information about \n- Combining / Transforming / Interaction\n- Categorical Features\n- Frequency Encoding\n- Aggregations / Group Statistics\n\n\n**More about Aggregations / Group Statistics**\n\nAgain from Chris we got [this](https://www.kaggle.com/code/cdeotte/xgb-fraud-with-magic-0-9600/notebook) notebook.\nIt has important information about aggregation features you might want to try \"blind\" and then perform some feature selection on the outcome you get. \nThis way you will be able to get meaningful features without knowing the meaning of the columns themselves. \n\n\n**Bruteforce Feature Engineering**\n\n[original post](https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111245) - Credit to the original author! \n\nIf you just want to create all features interactions and then feature select some of them, here is a code to do so:\n\n```\nfor i in range(len(cols)):\n    for j in range(i + 1, len(cols)):\n        X['{}_div_{}'.format(cols[i], cols[j])] = X[cols[i]] * X[cols[j]]\n        X['{}_mul_{}'.format(cols[i], cols[j])] = X[cols[i]] * X[cols[j]]\n        X['{}_sub1_{}'.format(cols[i], cols[j])] = X[cols[i]] - X[cols[j]]\n        X['{}_sub2_{}'.format(cols[j], cols[i])] = X[cols[j]] - X[cols[i]]\n```\n        \nYou can them compute a correlation matrix for all new generated features\nThen removing all the features with a high correlation. Keeping those that correlate with target value better.\n\n```\nbruteforce_cols = [col for col in X.columns if '_mul_' in col or '_sub_' in col of '_div_' in col]\ncorr_matrix = X[bruteforce_cols].corr()\n\nto_drop = list()\n\nfor i in range(1, len(corr_matrix)):\n    for j in range(i):\n        if corr_matrix.iloc[i, j] >= 0.98:\n            if abs(pd.concat([X[corr_matrix.index[i]], y], axis=1).corr().iloc[0][1]) > abs(pd.concat([X[corr_matrix.columns[j]], y], axis=1).corr().iloc[0][1]):\n                to_drop.append(corr_matrix.columns[j])\n            else:\n                to_drop.append(corr_matrix.index[i])\n\nto_drop = list(set(to_drop))\n```\n\n\n**Stacking**\n\nOne more trick we can apply that we don't have to know the meaning of the columns for is to simply use \"heavy stacking\". \nThis was done so many times on Kaggle so we won't need to dive into that again. \n\n\n**Pseudolabeling**\n\nThe main idea is to use a model for labeling an unlabeled dataset and then use this dataset in addition to the training set. \nThere isn't much to say about Pseudolabeling as well. This also had been used so many times on Kaggle and when done properly it most probably improves performance.\n\n\n#### Representation Learning\n\n\n**DEA and Ladder Networks**\n\n- DEA (Denoising auto encoders) were used in the MOA competition and on TPS competitions in the past. \n\n- Ladder Networks (https://arxiv.org/abs/1507.02672): Using denoising autoencoder in a unique fashion. This might be an interesting approach to try as well as simply vanilla DEA.\n\n\n**Unique Neural network architectures**\n\n- If you don't know about this: Check out the \"MOA Net\" that was introduced on the Mechanisms of action competition's 2nd place winner. \nThis might sound crazy at first but this network uses convolutions for tabular data. \nNeedless to say, by its 2nd place we can safely assume that it works well.\n\n\nWhat about you? \nDo you know any other approach or trick we can use when we have no information about the features?",
    "1813361": "Nice compile of techniques. I would first combine different variables or use some dimensionally reduction method\nCheers!",
    "1813736": "Great information. Thanks for sharing . Specially notebook from @cdeotte . Very helpful...🙏",
    "1816022": "Thanks for sharing!",
    "1820703": "Great topic!",
    "1821572": "Good summary. Autoencoders are extremly helpful especially working with high dimensional data like we have here.",
    "1873119": "thedevastator `X['{}_div_{}'.format(cols[i], cols[j])] = X[cols[i]] * X[cols[j]]` perhaps in this line there's a typo and instead of multiplication `*` there should be a division `/`?",
    "1873183": "Great topic! I'm new and would never have thought of this without reading. 🙂"
  },
  "source": "meta"
}