{
  "id": 320756,
  "title": "💥 Create 391 features in a few minutes! Multiple, customizable and scalable feature stores with kedro & featuretools.",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/320756",
  "author_name": "",
  "post_date": "2022-04-23T09:45:21.963477600Z",
  "votes": 14,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Have you ever had a problem that you used Jupyter Notebook for data cleaning, feature engineering and model training, was modifying some cells, running the model again, back and forth a few times and finally you were a bit lost which version of objects was used by the final model? 🥴</p>\n<p>When I joined <a href=\"https://getindata.com\" target=\"_blank\">GetInData</a> a few months ago, I started getting familiarised with <strong>MLops</strong>, especially especially <a href=\"https://kedro.readthedocs.io/en/stable/introduction/introduction.html\" target=\"_blank\">Kedro</a>. At first it was new, frightening and irritating, but after I got to know it more and more, I started seeing huge value in the structured approach.</p>\n<blockquote>\n  <p>We decided to take part in the Kaggle H&amp;M competition and see if <strong>Kedro</strong> is a good fit for such projects. After some time, we have concluded that <strong>benefits outweigh struggles</strong> when:</p>\n  <ul>\n  <li>the project can be structured into <strong>independent modules</strong></li>\n  <li>multiple Data Scientists and Machine Learning Engineers <strong>work simultaneously</strong> on the project and can develop independent modules which will be connected later using pipelines</li>\n  <li>you want to <strong>experiment</strong> with various approaches applied to each part of the processing and modelling part (e.g. data cleaning, feature engineering, ML algorithms, hyper-parameter tuning, validation, ensembling etc.) and see which configuration of them works best</li>\n  <li>you need <strong>reproducibility and scalability</strong></li>\n  </ul>\n</blockquote>\n<p>👉 I decided to <strong>share a toy example with you</strong> - please find it <a href=\"https://github.com/adrian-dembek/kaggle-hm-kedro\" target=\"_blank\">HERE</a> on GitHub. </p>\n<p>The solution can create <strong>multiple feature stores</strong> using <a href=\"https://featuretools.alteryx.com\" target=\"_blank\">featuretools</a>. You can play with configuration on your own and decide to go with even deeper feature engineering. It uses already <em>sampled data</em> so when you clone the repo and run a few commands, it should be working smoothly. The configuration in the repo created <strong>391 features</strong>, but in production we created over… 18'000 🤯. We are still developing our solution, so stay tuned! 🤓</p>\n<p>I am looking forward to your comments - especially on <strong>how Data Scientists see Kedro</strong>, do you think it would be beneficial for you to get to know the tool more? Also, feel free to ask questions!</p>",
  "messages": [
    {
      "id": "1765266",
      "postDate": "04/23/2022 09:45:21",
      "content": "<p>Have you ever had a problem that you used Jupyter Notebook for data cleaning, feature engineering and model training, was modifying some cells, running the model again, back and forth a few times and finally you were a bit lost which version of objects was used by the final model? 🥴</p>\n<p>When I joined <a href=\"https://getindata.com\" target=\"_blank\">GetInData</a> a few months ago, I started getting familiarised with <strong>MLops</strong>, especially especially <a href=\"https://kedro.readthedocs.io/en/stable/introduction/introduction.html\" target=\"_blank\">Kedro</a>. At first it was new, frightening and irritating, but after I got to know it more and more, I started seeing huge value in the structured approach.</p>\n<blockquote>\n  <p>We decided to take part in the Kaggle H&amp;M competition and see if <strong>Kedro</strong> is a good fit for such projects. After some time, we have concluded that <strong>benefits outweigh struggles</strong> when:</p>\n  <ul>\n  <li>the project can be structured into <strong>independent modules</strong></li>\n  <li>multiple Data Scientists and Machine Learning Engineers <strong>work simultaneously</strong> on the project and can develop independent modules which will be connected later using pipelines</li>\n  <li>you want to <strong>experiment</strong> with various approaches applied to each part of the processing and modelling part (e.g. data cleaning, feature engineering, ML algorithms, hyper-parameter tuning, validation, ensembling etc.) and see which configuration of them works best</li>\n  <li>you need <strong>reproducibility and scalability</strong></li>\n  </ul>\n</blockquote>\n<p>👉 I decided to <strong>share a toy example with you</strong> - please find it <a href=\"https://github.com/adrian-dembek/kaggle-hm-kedro\" target=\"_blank\">HERE</a> on GitHub. </p>\n<p>The solution can create <strong>multiple feature stores</strong> using <a href=\"https://featuretools.alteryx.com\" target=\"_blank\">featuretools</a>. You can play with configuration on your own and decide to go with even deeper feature engineering. It uses already <em>sampled data</em> so when you clone the repo and run a few commands, it should be working smoothly. The configuration in the repo created <strong>391 features</strong>, but in production we created over… 18'000 🤯. We are still developing our solution, so stay tuned! 🤓</p>\n<p>I am looking forward to your comments - especially on <strong>how Data Scientists see Kedro</strong>, do you think it would be beneficial for you to get to know the tool more? Also, feel free to ask questions!</p>",
      "rawMarkdown": "Have you ever had a problem that you used Jupyter Notebook for data cleaning, feature engineering and model training, was modifying some cells, running the model again, back and forth a few times and finally you were a bit lost which version of objects was used by the final model? 🥴\n\nWhen I joined [GetInData](https://getindata.com) a few months ago, I started getting familiarised with **MLops**, especially especially [Kedro](https://kedro.readthedocs.io/en/stable/introduction/introduction.html). At first it was new, frightening and irritating, but after I got to know it more and more, I started seeing huge value in the structured approach.\n\n> We decided to take part in the Kaggle H&M competition and see if **Kedro** is a good fit for such projects. After some time, we have concluded that **benefits outweigh struggles** when:\n- the project can be structured into **independent modules**\n- multiple Data Scientists and Machine Learning Engineers **work simultaneously** on the project and can develop independent modules which will be connected later using pipelines\n- you want to **experiment** with various approaches applied to each part of the processing and modelling part (e.g. data cleaning, feature engineering, ML algorithms, hyper-parameter tuning, validation, ensembling etc.) and see which configuration of them works best\n- you need **reproducibility and scalability**\n\n👉 I decided to **share a toy example with you** - please find it [HERE](https://github.com/adrian-dembek/kaggle-hm-kedro) on GitHub. \n\nThe solution can create **multiple feature stores** using [featuretools](https://featuretools.alteryx.com). You can play with configuration on your own and decide to go with even deeper feature engineering. It uses already *sampled data* so when you clone the repo and run a few commands, it should be working smoothly. The configuration in the repo created **391 features**, but in production we created over... 18'000 🤯. We are still developing our solution, so stay tuned! 🤓\n\n\nI am looking forward to your comments - especially on **how Data Scientists see Kedro**, do you think it would be beneficial for you to get to know the tool more? Also, feel free to ask questions!",
      "votes": null
    },
    {
      "id": "1765810",
      "postDate": "04/23/2022 22:12:30",
      "content": "<p><a href=\"https://www.kaggle.com/adriande\" target=\"_blank\">@adriande</a> Great idea!</p>\n<p>The next best step is to find the best features among the 391 newly created ones. You can use Featurewiz to find the best features using the MRMR algorithm.</p>\n<p>You can take a look here:<br>\n<a href=\"https://github.com/AutoViML/featurewiz\"><img src=\"https://i.ibb.co/ZLdZMZg/featurewiz-logos.png\" alt=\"featurewiz-logos\"></a></p>\n<p>You can reduce features and then build a simpler model than a bloated model with numerous features.<br>\nHope this helps,<br>\nRam</p>",
      "rawMarkdown": "adriande Great idea!\n\nThe next best step is to find the best features among the 391 newly created ones. You can use Featurewiz to find the best features using the MRMR algorithm.\n\nYou can take a look here:\n<a href=\"https://github.com/AutoViML/featurewiz\"><img src=\"https://i.ibb.co/ZLdZMZg/featurewiz-logos.png\" alt=\"featurewiz-logos\" border=\"0\"></a>\n\nYou can reduce features and then build a simpler model than a bloated model with numerous features.\nHope this helps,\nRam",
      "votes": null
    },
    {
      "id": "1766002",
      "postDate": "04/24/2022 05:27:58",
      "content": "<p><a href=\"https://www.kaggle.com/rsesha\" target=\"_blank\">@rsesha</a> Thank you! </p>\n<p>I will take a look definitely as the dimensionality reduction is the next step for sure. What we did so far is used feature preselection tools provided by featuretools. These 391 features would be great if all of them were informative… sometimes they end up being just a single-value feature and would be time-wasting to include such features in the model.</p>\n<p>We have a node for such cases and can easily add it to the pipeline :) </p>\n<pre><code>def automatically_preselect_features(fs,\n                                      remove_single_value,\n                                      remove_low_information,\n                                      remove_highly_null,\n                                      pct_null_threshold,\n                                      remove_highly_correlated,\n                                      pct_corr_threshold\n                                     ):\n\n    import featuretools as ft\n\n    curr_n_cols = fs.shape[1]\n    print('input df number of features: ' + str(curr_n_cols))\n\n    if remove_single_value==True:\n        fs = ft.selection.remove_single_value_features(fs, count_nan_as_value=True) # to prevent removing \"if had_sth==1 else null\" features\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' features with single value')\n        curr_n_cols = fs.shape[1]    \n        print('curr_n_cols:', curr_n_cols)\n\n    if remove_low_information==True:\n        fs = ft.selection.remove_low_information_features(fs)\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' features with low information')\n        curr_n_cols = fs.shape[1]\n        print('curr_n_cols:', curr_n_cols)\n\n    if remove_highly_null==True:\n        fs = ft.selection.remove_highly_null_features(fs, pct_null_threshold=pct_null_threshold)\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' highly null features')\n        curr_n_cols = fs.shape[1]  \n        print('curr_n_cols:', curr_n_cols)\n\n    if remove_highly_correlated==True:\n        fs = ft.selection.remove_highly_correlated_features(fs, pct_corr_threshold=pct_corr_threshold)\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' highly correlated features')\n        curr_n_cols = fs.shape[1]  \n        print('curr_n_cols:', curr_n_cols)\n\n    # not working\n    # ft.selection.remove_highly_correlated_features(feature_matrix_customers)\n\n    print('Remaining number of features after automatic selection: ' + str(curr_n_cols))\n\n\n    return fs\n</code></pre>",
      "rawMarkdown": "rsesha Thank you! \n\nI will take a look definitely as the dimensionality reduction is the next step for sure. What we did so far is used feature preselection tools provided by featuretools. These 391 features would be great if all of them were informative... sometimes they end up being just a single-value feature and would be time-wasting to include such features in the model.\n\nWe have a node for such cases and can easily add it to the pipeline :) \n\n```\ndef automatically_preselect_features(fs,\n                                      remove_single_value,\n                                      remove_low_information,\n                                      remove_highly_null,\n                                      pct_null_threshold,\n                                      remove_highly_correlated,\n                                      pct_corr_threshold\n                                     ):\n\n    import featuretools as ft\n    \n    curr_n_cols = fs.shape[1]\n    print('input df number of features: ' + str(curr_n_cols))\n    \n    if remove_single_value==True:\n        fs = ft.selection.remove_single_value_features(fs, count_nan_as_value=True) # to prevent removing \"if had_sth==1 else null\" features\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' features with single value')\n        curr_n_cols = fs.shape[1]    \n        print('curr_n_cols:', curr_n_cols)\n        \n    if remove_low_information==True:\n        fs = ft.selection.remove_low_information_features(fs)\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' features with low information')\n        curr_n_cols = fs.shape[1]\n        print('curr_n_cols:', curr_n_cols)\n    \n    if remove_highly_null==True:\n        fs = ft.selection.remove_highly_null_features(fs, pct_null_threshold=pct_null_threshold)\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' highly null features')\n        curr_n_cols = fs.shape[1]  \n        print('curr_n_cols:', curr_n_cols)\n        \n    if remove_highly_correlated==True:\n        fs = ft.selection.remove_highly_correlated_features(fs, pct_corr_threshold=pct_corr_threshold)\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' highly correlated features')\n        curr_n_cols = fs.shape[1]  \n        print('curr_n_cols:', curr_n_cols)\n\n    # not working\n    # ft.selection.remove_highly_correlated_features(feature_matrix_customers)\n        \n    print('Remaining number of features after automatic selection: ' + str(curr_n_cols))\n    \n    \n    return fs\n```",
      "votes": null
    },
    {
      "id": "1767446",
      "postDate": "04/25/2022 11:11:55",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/adriande\" target=\"_blank\">@adriande</a> :</p>\n<p>It seems like a good idea to do feature selection using correlation thresholds like below:<br>\n<code>ft.selection.remove_highly_correlated_features(fs, pct_corr_threshold=pct_corr_threshold)</code></p>\n<p>However, you will hit a major roadblock as almost everyone who has ventured in this path has done:<br>\n<code>1. If two variables are highly correlated, which of the two should I remove and which one should I keep?</code></p>\n<p>That's the question <a href=\"https://github.com/AutoViML/featurewiz\" target=\"_blank\">featurewiz</a> solves using the SULOV method. See how it eliminates one of the two correlated variables here.</p>\n<p><a href=\"https://github.com/AutoViML/featurewiz#2--feature-selection\"><img src=\"https://i.ibb.co/86LCg05/SULOV.jpg\" alt=\"SULOV\"></a></p>\n<p>All you have to do feature selection in your pipeline is do the following:<br>\n<code>\nfrom featurewiz import FeatureWiz\n</code><br>\n<code>\nfeatures = FeatureWiz(corr_limit=0.70, feature_engg='', category_encoders='', \n                  dask_xgboost_flag=False, nrows=None, verbose=2)\n</code><br>\n<code>X_train_selected = features.fit_transform(X_train, y_train)\n</code><br>\n<code>X_test_selected = features.transform(X_test)\n</code><br>\n<code>features.features  ### provides the list of selected features ###\n</code><br>\nYou can contact me via DM if you have any interest in exploring this further.<br>\nThanks</p>",
      "rawMarkdown": "Hi @adriande :\n\nIt seems like a good idea to do feature selection using correlation thresholds like below:\n`ft.selection.remove_highly_correlated_features(fs, pct_corr_threshold=pct_corr_threshold)`\n\nHowever, you will hit a major roadblock as almost everyone who has ventured in this path has done:\n`1. If two variables are highly correlated, which of the two should I remove and which one should I keep?`\n\nThat's the question [featurewiz](https://github.com/AutoViML/featurewiz) solves using the SULOV method. See how it eliminates one of the two correlated variables here.\n\n<a href=\"https://github.com/AutoViML/featurewiz#2--feature-selection\"><img src=\"https://i.ibb.co/86LCg05/SULOV.jpg\" alt=\"SULOV\" border=\"0\"></a>\n\nAll you have to do feature selection in your pipeline is do the following:\n`\nfrom featurewiz import FeatureWiz\n`\n`\nfeatures = FeatureWiz(corr_limit=0.70, feature_engg='', category_encoders='', \n                  dask_xgboost_flag=False, nrows=None, verbose=2)\n`\n`X_train_selected = features.fit_transform(X_train, y_train)\n`\n`X_test_selected = features.transform(X_test)\n`\n`features.features  ### provides the list of selected features ###\n`\nYou can contact me via DM if you have any interest in exploring this further.\nThanks",
      "votes": null
    },
    {
      "id": "1768546",
      "postDate": "04/26/2022 11:44:56",
      "content": "<p>Thank you for such a brilliant idea! I do think that this comepition is mostly about finding some underlying useful features and trying to use dimensionality reduction techniques to find the most useful features,  since everyone is using Lgbmranker, I don't really think the model will make a huge difference.</p>\n<p>I had some very noob questions for you sir: may I ask how could I use your github toy sample repository in google colab? I could only clone them but after that colab could not apply those conda commands.</p>\n<p>Also, may I ask could I use already sampled data for this new tech?  like 1% or 5% percent data from the original dataset? Or must I use the kedro sample pipeline for data sampling? </p>\n<p>Thank you for your patience! I have really learned a lot from your discussion!</p>",
      "rawMarkdown": "Thank you for such a brilliant idea! I do think that this comepition is mostly about finding some underlying useful features and trying to use dimensionality reduction techniques to find the most useful features,  since everyone is using Lgbmranker, I don't really think the model will make a huge difference.\n\nI had some very noob questions for you sir: may I ask how could I use your github toy sample repository in google colab? I could only clone them but after that colab could not apply those conda commands.\n\nAlso, may I ask could I use already sampled data for this new tech?  like 1% or 5% percent data from the original dataset? Or must I use the kedro sample pipeline for data sampling? \n\nThank you for your patience! I have really learned a lot from your discussion!",
      "votes": null
    },
    {
      "id": "1769207",
      "postDate": "04/27/2022 04:23:42",
      "content": "<p>Thanks for your share. It's amazing!!</p>",
      "rawMarkdown": "Thanks for your share. It's amazing!!",
      "votes": null
    },
    {
      "id": "1770162",
      "postDate": "04/28/2022 03:45:49",
      "content": "<p>Thanks for sharing. It's impressive!</p>",
      "rawMarkdown": "Thanks for sharing. It's impressive!",
      "votes": null
    },
    {
      "id": "1770503",
      "postDate": "04/28/2022 10:14:44",
      "content": "<p>Nice work 👍</p>",
      "rawMarkdown": "Nice work 👍",
      "votes": null
    },
    {
      "id": "1778637",
      "postDate": "05/05/2022 14:36:33",
      "content": "<p>Thanks for your share🙌</p>",
      "rawMarkdown": "Thanks for your share🙌",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1765810,
      "author_name": "rsesha",
      "author_url": "",
      "post_date": "04/23/2022 22:12:30",
      "content": "<p><a href=\"https://www.kaggle.com/adriande\" target=\"_blank\">@adriande</a> Great idea!</p>\n<p>The next best step is to find the best features among the 391 newly created ones. You can use Featurewiz to find the best features using the MRMR algorithm.</p>\n<p>You can take a look here:<br>\n<a href=\"https://github.com/AutoViML/featurewiz\"><img src=\"https://i.ibb.co/ZLdZMZg/featurewiz-logos.png\" alt=\"featurewiz-logos\"></a></p>\n<p>You can reduce features and then build a simpler model than a bloated model with numerous features.<br>\nHope this helps,<br>\nRam</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1766002,
      "author_name": "adriande",
      "author_url": "",
      "post_date": "04/24/2022 05:27:58",
      "content": "<p><a href=\"https://www.kaggle.com/rsesha\" target=\"_blank\">@rsesha</a> Thank you! </p>\n<p>I will take a look definitely as the dimensionality reduction is the next step for sure. What we did so far is used feature preselection tools provided by featuretools. These 391 features would be great if all of them were informative… sometimes they end up being just a single-value feature and would be time-wasting to include such features in the model.</p>\n<p>We have a node for such cases and can easily add it to the pipeline :) </p>\n<pre><code>def automatically_preselect_features(fs,\n                                      remove_single_value,\n                                      remove_low_information,\n                                      remove_highly_null,\n                                      pct_null_threshold,\n                                      remove_highly_correlated,\n                                      pct_corr_threshold\n                                     ):\n\n    import featuretools as ft\n\n    curr_n_cols = fs.shape[1]\n    print('input df number of features: ' + str(curr_n_cols))\n\n    if remove_single_value==True:\n        fs = ft.selection.remove_single_value_features(fs, count_nan_as_value=True) # to prevent removing \"if had_sth==1 else null\" features\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' features with single value')\n        curr_n_cols = fs.shape[1]    \n        print('curr_n_cols:', curr_n_cols)\n\n    if remove_low_information==True:\n        fs = ft.selection.remove_low_information_features(fs)\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' features with low information')\n        curr_n_cols = fs.shape[1]\n        print('curr_n_cols:', curr_n_cols)\n\n    if remove_highly_null==True:\n        fs = ft.selection.remove_highly_null_features(fs, pct_null_threshold=pct_null_threshold)\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' highly null features')\n        curr_n_cols = fs.shape[1]  \n        print('curr_n_cols:', curr_n_cols)\n\n    if remove_highly_correlated==True:\n        fs = ft.selection.remove_highly_correlated_features(fs, pct_corr_threshold=pct_corr_threshold)\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' highly correlated features')\n        curr_n_cols = fs.shape[1]  \n        print('curr_n_cols:', curr_n_cols)\n\n    # not working\n    # ft.selection.remove_highly_correlated_features(feature_matrix_customers)\n\n    print('Remaining number of features after automatic selection: ' + str(curr_n_cols))\n\n\n    return fs\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1767446,
          "author_name": "rsesha",
          "author_url": "",
          "post_date": "04/25/2022 11:11:55",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/adriande\" target=\"_blank\">@adriande</a> :</p>\n<p>It seems like a good idea to do feature selection using correlation thresholds like below:<br>\n<code>ft.selection.remove_highly_correlated_features(fs, pct_corr_threshold=pct_corr_threshold)</code></p>\n<p>However, you will hit a major roadblock as almost everyone who has ventured in this path has done:<br>\n<code>1. If two variables are highly correlated, which of the two should I remove and which one should I keep?</code></p>\n<p>That's the question <a href=\"https://github.com/AutoViML/featurewiz\" target=\"_blank\">featurewiz</a> solves using the SULOV method. See how it eliminates one of the two correlated variables here.</p>\n<p><a href=\"https://github.com/AutoViML/featurewiz#2--feature-selection\"><img src=\"https://i.ibb.co/86LCg05/SULOV.jpg\" alt=\"SULOV\"></a></p>\n<p>All you have to do feature selection in your pipeline is do the following:<br>\n<code>\nfrom featurewiz import FeatureWiz\n</code><br>\n<code>\nfeatures = FeatureWiz(corr_limit=0.70, feature_engg='', category_encoders='', \n                  dask_xgboost_flag=False, nrows=None, verbose=2)\n</code><br>\n<code>X_train_selected = features.fit_transform(X_train, y_train)\n</code><br>\n<code>X_test_selected = features.transform(X_test)\n</code><br>\n<code>features.features  ### provides the list of selected features ###\n</code><br>\nYou can contact me via DM if you have any interest in exploring this further.<br>\nThanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1768546,
      "author_name": "skybookreader",
      "author_url": "",
      "post_date": "04/26/2022 11:44:56",
      "content": "<p>Thank you for such a brilliant idea! I do think that this comepition is mostly about finding some underlying useful features and trying to use dimensionality reduction techniques to find the most useful features,  since everyone is using Lgbmranker, I don't really think the model will make a huge difference.</p>\n<p>I had some very noob questions for you sir: may I ask how could I use your github toy sample repository in google colab? I could only clone them but after that colab could not apply those conda commands.</p>\n<p>Also, may I ask could I use already sampled data for this new tech?  like 1% or 5% percent data from the original dataset? Or must I use the kedro sample pipeline for data sampling? </p>\n<p>Thank you for your patience! I have really learned a lot from your discussion!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1769207,
      "author_name": "blackbur",
      "author_url": "",
      "post_date": "04/27/2022 04:23:42",
      "content": "<p>Thanks for your share. It's amazing!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1770162,
      "author_name": "",
      "author_url": "",
      "post_date": "04/28/2022 03:45:49",
      "content": "<p>Thanks for sharing. It's impressive!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1770503,
      "author_name": "eagerbird",
      "author_url": "",
      "post_date": "04/28/2022 10:14:44",
      "content": "<p>Nice work 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1778637,
      "author_name": "tiger0",
      "author_url": "",
      "post_date": "05/05/2022 14:36:33",
      "content": "<p>Thanks for your share🙌</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1765266": "Have you ever had a problem that you used Jupyter Notebook for data cleaning, feature engineering and model training, was modifying some cells, running the model again, back and forth a few times and finally you were a bit lost which version of objects was used by the final model? 🥴\n\nWhen I joined [GetInData](https://getindata.com) a few months ago, I started getting familiarised with **MLops**, especially especially [Kedro](https://kedro.readthedocs.io/en/stable/introduction/introduction.html). At first it was new, frightening and irritating, but after I got to know it more and more, I started seeing huge value in the structured approach.\n\n> We decided to take part in the Kaggle H&M competition and see if **Kedro** is a good fit for such projects. After some time, we have concluded that **benefits outweigh struggles** when:\n- the project can be structured into **independent modules**\n- multiple Data Scientists and Machine Learning Engineers **work simultaneously** on the project and can develop independent modules which will be connected later using pipelines\n- you want to **experiment** with various approaches applied to each part of the processing and modelling part (e.g. data cleaning, feature engineering, ML algorithms, hyper-parameter tuning, validation, ensembling etc.) and see which configuration of them works best\n- you need **reproducibility and scalability**\n\n👉 I decided to **share a toy example with you** - please find it [HERE](https://github.com/adrian-dembek/kaggle-hm-kedro) on GitHub. \n\nThe solution can create **multiple feature stores** using [featuretools](https://featuretools.alteryx.com). You can play with configuration on your own and decide to go with even deeper feature engineering. It uses already *sampled data* so when you clone the repo and run a few commands, it should be working smoothly. The configuration in the repo created **391 features**, but in production we created over... 18'000 🤯. We are still developing our solution, so stay tuned! 🤓\n\n\nI am looking forward to your comments - especially on **how Data Scientists see Kedro**, do you think it would be beneficial for you to get to know the tool more? Also, feel free to ask questions!",
    "1765810": "adriande Great idea!\n\nThe next best step is to find the best features among the 391 newly created ones. You can use Featurewiz to find the best features using the MRMR algorithm.\n\nYou can take a look here:\n<a href=\"https://github.com/AutoViML/featurewiz\"><img src=\"https://i.ibb.co/ZLdZMZg/featurewiz-logos.png\" alt=\"featurewiz-logos\" border=\"0\"></a>\n\nYou can reduce features and then build a simpler model than a bloated model with numerous features.\nHope this helps,\nRam",
    "1766002": "rsesha Thank you! \n\nI will take a look definitely as the dimensionality reduction is the next step for sure. What we did so far is used feature preselection tools provided by featuretools. These 391 features would be great if all of them were informative... sometimes they end up being just a single-value feature and would be time-wasting to include such features in the model.\n\nWe have a node for such cases and can easily add it to the pipeline :) \n\n```\ndef automatically_preselect_features(fs,\n                                      remove_single_value,\n                                      remove_low_information,\n                                      remove_highly_null,\n                                      pct_null_threshold,\n                                      remove_highly_correlated,\n                                      pct_corr_threshold\n                                     ):\n\n    import featuretools as ft\n    \n    curr_n_cols = fs.shape[1]\n    print('input df number of features: ' + str(curr_n_cols))\n    \n    if remove_single_value==True:\n        fs = ft.selection.remove_single_value_features(fs, count_nan_as_value=True) # to prevent removing \"if had_sth==1 else null\" features\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' features with single value')\n        curr_n_cols = fs.shape[1]    \n        print('curr_n_cols:', curr_n_cols)\n        \n    if remove_low_information==True:\n        fs = ft.selection.remove_low_information_features(fs)\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' features with low information')\n        curr_n_cols = fs.shape[1]\n        print('curr_n_cols:', curr_n_cols)\n    \n    if remove_highly_null==True:\n        fs = ft.selection.remove_highly_null_features(fs, pct_null_threshold=pct_null_threshold)\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' highly null features')\n        curr_n_cols = fs.shape[1]  \n        print('curr_n_cols:', curr_n_cols)\n        \n    if remove_highly_correlated==True:\n        fs = ft.selection.remove_highly_correlated_features(fs, pct_corr_threshold=pct_corr_threshold)\n        removed_n_cols = curr_n_cols - fs.shape[1]\n        print('removed ' + str(removed_n_cols) + ' highly correlated features')\n        curr_n_cols = fs.shape[1]  \n        print('curr_n_cols:', curr_n_cols)\n\n    # not working\n    # ft.selection.remove_highly_correlated_features(feature_matrix_customers)\n        \n    print('Remaining number of features after automatic selection: ' + str(curr_n_cols))\n    \n    \n    return fs\n```",
    "1767446": "Hi @adriande :\n\nIt seems like a good idea to do feature selection using correlation thresholds like below:\n`ft.selection.remove_highly_correlated_features(fs, pct_corr_threshold=pct_corr_threshold)`\n\nHowever, you will hit a major roadblock as almost everyone who has ventured in this path has done:\n`1. If two variables are highly correlated, which of the two should I remove and which one should I keep?`\n\nThat's the question [featurewiz](https://github.com/AutoViML/featurewiz) solves using the SULOV method. See how it eliminates one of the two correlated variables here.\n\n<a href=\"https://github.com/AutoViML/featurewiz#2--feature-selection\"><img src=\"https://i.ibb.co/86LCg05/SULOV.jpg\" alt=\"SULOV\" border=\"0\"></a>\n\nAll you have to do feature selection in your pipeline is do the following:\n`\nfrom featurewiz import FeatureWiz\n`\n`\nfeatures = FeatureWiz(corr_limit=0.70, feature_engg='', category_encoders='', \n                  dask_xgboost_flag=False, nrows=None, verbose=2)\n`\n`X_train_selected = features.fit_transform(X_train, y_train)\n`\n`X_test_selected = features.transform(X_test)\n`\n`features.features  ### provides the list of selected features ###\n`\nYou can contact me via DM if you have any interest in exploring this further.\nThanks",
    "1768546": "Thank you for such a brilliant idea! I do think that this comepition is mostly about finding some underlying useful features and trying to use dimensionality reduction techniques to find the most useful features,  since everyone is using Lgbmranker, I don't really think the model will make a huge difference.\n\nI had some very noob questions for you sir: may I ask how could I use your github toy sample repository in google colab? I could only clone them but after that colab could not apply those conda commands.\n\nAlso, may I ask could I use already sampled data for this new tech?  like 1% or 5% percent data from the original dataset? Or must I use the kedro sample pipeline for data sampling? \n\nThank you for your patience! I have really learned a lot from your discussion!",
    "1769207": "Thanks for your share. It's amazing!!",
    "1770162": "Thanks for sharing. It's impressive!",
    "1770503": "Nice work 👍",
    "1778637": "Thanks for your share🙌"
  },
  "source": "meta"
}