{
  "id": 179033,
  "title": "Optimization of the data preparation",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/179033",
  "author_name": "Dmitrij Kozachuk",
  "post_date": "2020-09-01T07:21:16.774000",
  "votes": 33,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I've noticed that most of popular notebooks with Quantile Regression have the same data preparation. It's at first glance hard to understand in some parts. In this topic I will write some improvements of these parts without any time improvement and only for better understanding. You can use them if you want (for instance, if you are a perfectionist in the code style or just for ideas for the future). I will extend this topic with each improvement that I can find in the next days.</p>\n<p><strong>Improvements:</strong></p>\n<p><strong>1.</strong> This one:</p>\n<pre><code>base = data.loc[data.Weeks == data.min_week]\nbase = base[['Patient','FVC']].copy()\nbase.columns = ['Patient','min_FVC']\nbase['nb'] = 1\nbase['nb'] = base.groupby('Patient')['nb'].transform('cumsum')\nbase = base[base.nb==1]\nbase.drop('nb', axis=1, inplace=True)\n</code></pre>\n<p>into this one:</p>\n<pre><code>base = (\n    data\n    .loc[data.Weeks == data.min_week][['Patient','FVC']]\n    .rename({'FVC': 'min_FVC'}, axis=1)\n    .groupby('Patient')\n    .first()\n    .reset_index()\n)\n</code></pre>\n<p><strong>2.</strong> This one:</p>\n<pre><code>COLS = ['Sex','SmokingStatus'] #,'Age'\nFE = []\nfor col in COLS:\n    for mod in sorted(data[col].unique()):\n        FE.append(mod)\n        data[mod] = (data[col] == mod).astype(int)\n</code></pre>\n<p>into this one:</p>\n<pre><code>FE = list(data.Sex.unique()) + list(data.SmokingStatus.unique())\ndata = pd.concat([\n    data,\n    pd.get_dummies(data.Sex),\n    pd.get_dummies(data.SmokingStatus)\n], axis=1)\n</code></pre>\n<p><strong>3.</strong>[optional] This one:</p>\n<pre><code>data['age'] = (data['Age'] - data['Age'].min()) / (data['Age'].max() - data['Age'].min())\ndata['BASE'] = (data['min_FVC'] - data['min_FVC'].min() ) / ( data['min_FVC'].max() - data['min_FVC'].min() )\ndata['week'] = (data['base_week'] - data['base_week'].min() ) / ( data['base_week'].max() - data['base_week'].min() )\ndata['percent'] = (data['Percent'] - data['Percent'].min() ) / ( data['Percent'].max() - data['Percent'].min() )\nFE += ['age','percent','week','BASE']\n</code></pre>\n<p>into this one:</p>\n<pre><code>def get_fillness(series):\n    return (series - series.min()) / (series.max() - series.min())\n\ndata['age'] = get_fillness(data['Age'])\ndata['BASE'] = get_fillness(data['min_FVC'])\ndata['week'] = get_fillness(data['base_week'])\ndata['percent'] = get_fillness(data['Percent'])\n\nFE += ['age','percent','week','BASE']\n</code></pre>\n<p><strong>Proofs:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F593e310c99d10311e44d575bed88484b%2FProof.png?generation=1598903196314916&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F007734f1ed6097bd990fe7a7bca70dea%2FProof2.png?generation=1598903680222466&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 993847,
      "postDate": "2020-09-01T07:21:16.773Z",
      "content": "<p>I've noticed that most of popular notebooks with Quantile Regression have the same data preparation. It's at first glance hard to understand in some parts. In this topic I will write some improvements of these parts without any time improvement and only for better understanding. You can use them if you want (for instance, if you are a perfectionist in the code style or just for ideas for the future). I will extend this topic with each improvement that I can find in the next days.</p>\n<p><strong>Improvements:</strong></p>\n<p><strong>1.</strong> This one:</p>\n<pre><code>base = data.loc[data.Weeks == data.min_week]\nbase = base[['Patient','FVC']].copy()\nbase.columns = ['Patient','min_FVC']\nbase['nb'] = 1\nbase['nb'] = base.groupby('Patient')['nb'].transform('cumsum')\nbase = base[base.nb==1]\nbase.drop('nb', axis=1, inplace=True)\n</code></pre>\n<p>into this one:</p>\n<pre><code>base = (\n    data\n    .loc[data.Weeks == data.min_week][['Patient','FVC']]\n    .rename({'FVC': 'min_FVC'}, axis=1)\n    .groupby('Patient')\n    .first()\n    .reset_index()\n)\n</code></pre>\n<p><strong>2.</strong> This one:</p>\n<pre><code>COLS = ['Sex','SmokingStatus'] #,'Age'\nFE = []\nfor col in COLS:\n    for mod in sorted(data[col].unique()):\n        FE.append(mod)\n        data[mod] = (data[col] == mod).astype(int)\n</code></pre>\n<p>into this one:</p>\n<pre><code>FE = list(data.Sex.unique()) + list(data.SmokingStatus.unique())\ndata = pd.concat([\n    data,\n    pd.get_dummies(data.Sex),\n    pd.get_dummies(data.SmokingStatus)\n], axis=1)\n</code></pre>\n<p><strong>3.</strong>[optional] This one:</p>\n<pre><code>data['age'] = (data['Age'] - data['Age'].min()) / (data['Age'].max() - data['Age'].min())\ndata['BASE'] = (data['min_FVC'] - data['min_FVC'].min() ) / ( data['min_FVC'].max() - data['min_FVC'].min() )\ndata['week'] = (data['base_week'] - data['base_week'].min() ) / ( data['base_week'].max() - data['base_week'].min() )\ndata['percent'] = (data['Percent'] - data['Percent'].min() ) / ( data['Percent'].max() - data['Percent'].min() )\nFE += ['age','percent','week','BASE']\n</code></pre>\n<p>into this one:</p>\n<pre><code>def get_fillness(series):\n    return (series - series.min()) / (series.max() - series.min())\n\ndata['age'] = get_fillness(data['Age'])\ndata['BASE'] = get_fillness(data['min_FVC'])\ndata['week'] = get_fillness(data['base_week'])\ndata['percent'] = get_fillness(data['Percent'])\n\nFE += ['age','percent','week','BASE']\n</code></pre>\n<p><strong>Proofs:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F593e310c99d10311e44d575bed88484b%2FProof.png?generation=1598903196314916&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F007734f1ed6097bd990fe7a7bca70dea%2FProof2.png?generation=1598903680222466&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I've noticed that most of popular notebooks with Quantile Regression have the same data preparation. It's at first glance hard to understand in some parts. In this topic I will write some improvements of these parts without any time improvement and only for better understanding. You can use them if you want (for instance, if you are a perfectionist in the code style or just for ideas for the future). I will extend this topic with each improvement that I can find in the next days.\n\n**Improvements:**\n\n**1.** This one:\n```\nbase = data.loc[data.Weeks == data.min_week]\nbase = base[['Patient','FVC']].copy()\nbase.columns = ['Patient','min_FVC']\nbase['nb'] = 1\nbase['nb'] = base.groupby('Patient')['nb'].transform('cumsum')\nbase = base[base.nb==1]\nbase.drop('nb', axis=1, inplace=True)\n```\ninto this one:\n```\nbase = (\n    data\n    .loc[data.Weeks == data.min_week][['Patient','FVC']]\n    .rename({'FVC': 'min_FVC'}, axis=1)\n    .groupby('Patient')\n    .first()\n    .reset_index()\n)\n```\n**2.** This one:\n```\nCOLS = ['Sex','SmokingStatus'] #,'Age'\nFE = []\nfor col in COLS:\n    for mod in sorted(data[col].unique()):\n        FE.append(mod)\n        data[mod] = (data[col] == mod).astype(int)\n```\ninto this one:\n```\nFE = list(data.Sex.unique()) + list(data.SmokingStatus.unique())\ndata = pd.concat([\n    data,\n    pd.get_dummies(data.Sex),\n    pd.get_dummies(data.SmokingStatus)\n], axis=1)\n```\n**3.**[optional] This one:\n```\ndata['age'] = (data['Age'] - data['Age'].min()) / (data['Age'].max() - data['Age'].min())\ndata['BASE'] = (data['min_FVC'] - data['min_FVC'].min() ) / ( data['min_FVC'].max() - data['min_FVC'].min() )\ndata['week'] = (data['base_week'] - data['base_week'].min() ) / ( data['base_week'].max() - data['base_week'].min() )\ndata['percent'] = (data['Percent'] - data['Percent'].min() ) / ( data['Percent'].max() - data['Percent'].min() )\nFE += ['age','percent','week','BASE']\n```\ninto this one:\n```\ndef get_fillness(series):\n    return (series - series.min()) / (series.max() - series.min())\n\ndata['age'] = get_fillness(data['Age'])\ndata['BASE'] = get_fillness(data['min_FVC'])\ndata['week'] = get_fillness(data['base_week'])\ndata['percent'] = get_fillness(data['Percent'])\n\nFE += ['age','percent','week','BASE']\n```\n\n**Proofs:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F593e310c99d10311e44d575bed88484b%2FProof.png?generation=1598903196314916&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F007734f1ed6097bd990fe7a7bca70dea%2FProof2.png?generation=1598903680222466&alt=media)",
      "votes": 32
    },
    {
      "id": 994317,
      "postDate": "2020-09-01T14:13:57.027Z",
      "content": "<p>Thanks for sharing!<br>\nAbout normalization: most of kernels normalize the tabular data with the train AND the test sets (which is actually hidden and made during submission because the training portion is included in the inference kernel). However it's a good practice to normalize train data with only train data and test data with only test data because of potential leaks. In this challenge we only want that our solution performs the best on the competition set so should we continue to normalize with (train+test) ? so if we voluntary introduces a leak during training, would the performance on the competition test set be better ?</p>",
      "rawMarkdown": "Thanks for sharing!\nAbout normalization: most of kernels normalize the tabular data with the train AND the test sets (which is actually hidden and made during submission because the training portion is included in the inference kernel). However it's a good practice to normalize train data with only train data and test data with only test data because of potential leaks. In this challenge we only want that our solution performs the best on the competition set so should we continue to normalize with (train+test) ? so if we voluntary introduces a leak during training, would the performance on the competition test set be better ?",
      "votes": 5,
      "replies": [
        {
          "id": 994404,
          "postDate": "2020-09-01T14:59:49.437Z",
          "content": "<p>Agree, in my opinion that's a data leakage. Thank you for pay attention.</p>",
          "rawMarkdown": "Agree, in my opinion that's a data leakage. Thank you for pay attention.",
          "votes": 1
        },
        {
          "id": 998624,
          "postDate": "2020-09-04T22:13:45.077Z",
          "content": "<p>It actually goes further than this. In a cross-validation we should be normalising based only on the training fold not the complete training data.</p>",
          "rawMarkdown": "It actually goes further than this. In a cross-validation we should be normalising based only on the training fold not the complete training data.",
          "votes": 3
        }
      ]
    },
    {
      "id": 994028,
      "postDate": "2020-09-01T10:10:11.937Z",
      "content": "<p>Thanks for sharing. I always use my own code because of this reason. Some of them are still too long though. You can use this one-liner for baseline features. (You have to make sure the data is sorted by <code>[Patient, Weeks]</code>)</p>\n<p><code>df['FVC_Baseline'] = df.groupby('Patient').transform('first')['FVC']</code></p>",
      "rawMarkdown": "Thanks for sharing. I always use my own code because of this reason. Some of them are still too long though. You can use this one-liner for baseline features. (You have to make sure the data is sorted by `[Patient, Weeks]`)\n\n`df['FVC_Baseline'] = df.groupby('Patient').transform('first')['FVC']`",
      "votes": 3,
      "replies": [
        {
          "id": 994060,
          "postDate": "2020-09-01T10:37:56.080Z",
          "content": "<p>Agree, good optimization! Before grouping we can apply <code>.sort(by=(Patient, Weeks))</code> and it will work in every case. In the evening I will update the topic with corresponding check, that results are the same.</p>",
          "rawMarkdown": "Agree, good optimization! Before grouping we can apply `.sort(by=(Patient, Weeks))` and it will work in every case. In the evening I will update the topic with corresponding check, that results are the same."
        }
      ]
    },
    {
      "id": 998433,
      "postDate": "2020-09-04T18:12:08.413Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/koza4ukdmitrij\" target=\"_blank\">@koza4ukdmitrij</a>, very clean approach to optimize some code which I used..thanks for the hints! 👍👍</p>",
      "rawMarkdown": "Hey @koza4ukdmitrij, very clean approach to optimize some code which I used..thanks for the hints! 👍👍",
      "votes": 1,
      "replies": [
        {
          "id": 998441,
          "postDate": "2020-09-04T18:16:26.700Z",
          "content": "<p>You're welcome! </p>",
          "rawMarkdown": "You're welcome! "
        }
      ]
    },
    {
      "id": 995388,
      "postDate": "2020-09-02T12:02:14.140Z",
      "content": "<p>Thanks for sharing , helps to simplify the data preparation script 😃</p>",
      "rawMarkdown": "Thanks for sharing , helps to simplify the data preparation script 😃",
      "votes": 1,
      "replies": [
        {
          "id": 995402,
          "postDate": "2020-09-02T12:10:34.427Z",
          "content": "<p>Yeah, that's one of the aims)</p>",
          "rawMarkdown": "Yeah, that's one of the aims)",
          "votes": 1
        }
      ]
    },
    {
      "id": 994241,
      "postDate": "2020-09-01T13:34:15.193Z",
      "content": "<p>This is pretty good. I actually used some of this myself and you might not have any problems with this data set but when you do the pd.get_dummies() it's better to make the column into categorical first because if for e.g., only males were present on the train set then you would not get the right one-hot-encoding.</p>",
      "rawMarkdown": "This is pretty good. I actually used some of this myself and you might not have any problems with this data set but when you do the pd.get_dummies() it's better to make the column into categorical first because if for e.g., only males were present on the train set then you would not get the right one-hot-encoding.",
      "votes": 1,
      "replies": [
        {
          "id": 994252,
          "postDate": "2020-09-01T13:41:24.537Z",
          "content": "<p>Thanks! Looks like there is no such problems in the old version too, cause it was <code>sorted(data['Sex'].unique())</code> instead of <code>['male', \"female']</code>.</p>",
          "rawMarkdown": "Thanks! Looks like there is no such problems in the old version too, cause it was `sorted(data['Sex'].unique())` instead of `['male', \"female']`."
        }
      ]
    },
    {
      "id": 998377,
      "postDate": "2020-09-04T17:34:31.380Z",
      "content": "<p>Thanks for making this clean. I did use this code first, but then I was having a hard time to understand the script. I decided to rewrite almost the whole notebook with the addition of comments, and now navigating to a section and understanding scripts feels easy. <a href=\"https://www.kaggle.com/chrisden\" target=\"_blank\">@chrisden</a> did a good job with his <a href=\"https://www.kaggle.com/chrisden/6-82-quantile-reg-lr-schedulers-checkpoints\" target=\"_blank\">well-documented notebook</a>.</p>",
      "rawMarkdown": "Thanks for making this clean. I did use this code first, but then I was having a hard time to understand the script. I decided to rewrite almost the whole notebook with the addition of comments, and now navigating to a section and understanding scripts feels easy. @chrisden did a good job with his [well-documented notebook](https://www.kaggle.com/chrisden/6-82-quantile-reg-lr-schedulers-checkpoints).",
      "votes": 2,
      "replies": [
        {
          "id": 998397,
          "postDate": "2020-09-04T17:47:40.123Z",
          "content": "<p>Thanks a lot for your notebook anyway: writing a code with an attractive result is more usefull than the one with optimization and lower metric on my side. You're really welcome to use code from the topic. </p>",
          "rawMarkdown": "Thanks a lot for your notebook anyway: writing a code with an attractive result is more usefull than the one with optimization and lower metric on my side. You're really welcome to use code from the topic. "
        },
        {
          "id": 998436,
          "postDate": "2020-09-04T18:13:39.703Z",
          "content": "<p>Dear  <a href=\"https://www.kaggle.com/aadhavvignesh\" target=\"_blank\">@aadhavvignesh</a>, thanks a lot for the Kudos for my Notebook! 😃</p>",
          "rawMarkdown": "Dear  @aadhavvignesh, thanks a lot for the Kudos for my Notebook! 😃",
          "votes": 1
        }
      ]
    },
    {
      "id": 994222,
      "postDate": "2020-09-01T13:26:41.427Z",
      "content": "<p>Thanks for sharing. </p>",
      "rawMarkdown": "Thanks for sharing. ",
      "replies": [
        {
          "id": 994239,
          "postDate": "2020-09-01T13:33:49.267Z",
          "content": "<p>You're welcome!</p>",
          "rawMarkdown": "You're welcome!",
          "votes": 2
        }
      ]
    },
    {
      "id": 1000457,
      "postDate": "2020-09-06T14:39:50.263Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1000480,
          "postDate": "2020-09-06T14:55:29.687Z",
          "content": "<p>Thanks for your response ;)</p>",
          "rawMarkdown": "Thanks for your response ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1001145,
      "postDate": "2020-09-07T05:11:08.307Z",
      "content": "<p>Thanks for making this clean.</p>",
      "rawMarkdown": "Thanks for making this clean."
    }
  ],
  "comments": [
    {
      "id": 994317,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-09-01T14:13:57.027000",
      "content": "<p>Thanks for sharing!<br>\nAbout normalization: most of kernels normalize the tabular data with the train AND the test sets (which is actually hidden and made during submission because the training portion is included in the inference kernel). However it's a good practice to normalize train data with only train data and test data with only test data because of potential leaks. In this challenge we only want that our solution performs the best on the competition set so should we continue to normalize with (train+test) ? so if we voluntary introduces a leak during training, would the performance on the competition test set be better ?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 994404,
          "author_name": "Dmitrij Kozachuk",
          "author_url": "",
          "post_date": "2020-09-01T14:59:49.437000",
          "content": "<p>Agree, in my opinion that's a data leakage. Thank you for pay attention.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 998624,
          "author_name": "jameschapman19",
          "author_url": "",
          "post_date": "2020-09-04T22:13:45.077000",
          "content": "<p>It actually goes further than this. In a cross-validation we should be normalising based only on the training fold not the complete training data.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 994028,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2020-09-01T10:10:11.937000",
      "content": "<p>Thanks for sharing. I always use my own code because of this reason. Some of them are still too long though. You can use this one-liner for baseline features. (You have to make sure the data is sorted by <code>[Patient, Weeks]</code>)</p>\n<p><code>df['FVC_Baseline'] = df.groupby('Patient').transform('first')['FVC']</code></p>",
      "votes": 3,
      "replies": [
        {
          "id": 994060,
          "author_name": "Dmitrij Kozachuk",
          "author_url": "",
          "post_date": "2020-09-01T10:37:56.080000",
          "content": "<p>Agree, good optimization! Before grouping we can apply <code>.sort(by=(Patient, Weeks))</code> and it will work in every case. In the evening I will update the topic with corresponding check, that results are the same.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 998433,
      "author_name": "from coffee import *",
      "author_url": "",
      "post_date": "2020-09-04T18:12:08.413000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/koza4ukdmitrij\" target=\"_blank\">@koza4ukdmitrij</a>, very clean approach to optimize some code which I used..thanks for the hints! 👍👍</p>",
      "votes": 1,
      "replies": [
        {
          "id": 998441,
          "author_name": "Dmitrij Kozachuk",
          "author_url": "",
          "post_date": "2020-09-04T18:16:26.700000",
          "content": "<p>You're welcome! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 995388,
      "author_name": "D1nall",
      "author_url": "",
      "post_date": "2020-09-02T12:02:14.140000",
      "content": "<p>Thanks for sharing , helps to simplify the data preparation script 😃</p>",
      "votes": 1,
      "replies": [
        {
          "id": 995402,
          "author_name": "Dmitrij Kozachuk",
          "author_url": "",
          "post_date": "2020-09-02T12:10:34.427000",
          "content": "<p>Yeah, that's one of the aims)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 994241,
      "author_name": "johnny",
      "author_url": "",
      "post_date": "2020-09-01T13:34:15.193000",
      "content": "<p>This is pretty good. I actually used some of this myself and you might not have any problems with this data set but when you do the pd.get_dummies() it's better to make the column into categorical first because if for e.g., only males were present on the train set then you would not get the right one-hot-encoding.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 994252,
          "author_name": "Dmitrij Kozachuk",
          "author_url": "",
          "post_date": "2020-09-01T13:41:24.537000",
          "content": "<p>Thanks! Looks like there is no such problems in the old version too, cause it was <code>sorted(data['Sex'].unique())</code> instead of <code>['male', \"female']</code>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 998377,
      "author_name": "Aadhav Vignesh",
      "author_url": "",
      "post_date": "2020-09-04T17:34:31.380000",
      "content": "<p>Thanks for making this clean. I did use this code first, but then I was having a hard time to understand the script. I decided to rewrite almost the whole notebook with the addition of comments, and now navigating to a section and understanding scripts feels easy. <a href=\"https://www.kaggle.com/chrisden\" target=\"_blank\">@chrisden</a> did a good job with his <a href=\"https://www.kaggle.com/chrisden/6-82-quantile-reg-lr-schedulers-checkpoints\" target=\"_blank\">well-documented notebook</a>.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 998397,
          "author_name": "Dmitrij Kozachuk",
          "author_url": "",
          "post_date": "2020-09-04T17:47:40.123000",
          "content": "<p>Thanks a lot for your notebook anyway: writing a code with an attractive result is more usefull than the one with optimization and lower metric on my side. You're really welcome to use code from the topic. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 998436,
          "author_name": "from coffee import *",
          "author_url": "",
          "post_date": "2020-09-04T18:13:39.703000",
          "content": "<p>Dear  <a href=\"https://www.kaggle.com/aadhavvignesh\" target=\"_blank\">@aadhavvignesh</a>, thanks a lot for the Kudos for my Notebook! 😃</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 994222,
      "author_name": "Smuch",
      "author_url": "",
      "post_date": "2020-09-01T13:26:41.427000",
      "content": "<p>Thanks for sharing. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 994239,
          "author_name": "Dmitrij Kozachuk",
          "author_url": "",
          "post_date": "2020-09-01T13:33:49.267000",
          "content": "<p>You're welcome!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1000457,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-06T14:39:50.263000",
      "content": "",
      "votes": -1,
      "replies": [
        {
          "id": 1000480,
          "author_name": "Dmitrij Kozachuk",
          "author_url": "",
          "post_date": "2020-09-06T14:55:29.687000",
          "content": "<p>Thanks for your response ;)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1001145,
      "author_name": "Ranejeb",
      "author_url": "",
      "post_date": "2020-09-07T05:11:08.307000",
      "content": "<p>Thanks for making this clean.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "993847": "I've noticed that most of popular notebooks with Quantile Regression have the same data preparation. It's at first glance hard to understand in some parts. In this topic I will write some improvements of these parts without any time improvement and only for better understanding. You can use them if you want (for instance, if you are a perfectionist in the code style or just for ideas for the future). I will extend this topic with each improvement that I can find in the next days.\n\n**Improvements:**\n\n**1.** This one:\n```\nbase = data.loc[data.Weeks == data.min_week]\nbase = base[['Patient','FVC']].copy()\nbase.columns = ['Patient','min_FVC']\nbase['nb'] = 1\nbase['nb'] = base.groupby('Patient')['nb'].transform('cumsum')\nbase = base[base.nb==1]\nbase.drop('nb', axis=1, inplace=True)\n```\ninto this one:\n```\nbase = (\n    data\n    .loc[data.Weeks == data.min_week][['Patient','FVC']]\n    .rename({'FVC': 'min_FVC'}, axis=1)\n    .groupby('Patient')\n    .first()\n    .reset_index()\n)\n```\n**2.** This one:\n```\nCOLS = ['Sex','SmokingStatus'] #,'Age'\nFE = []\nfor col in COLS:\n    for mod in sorted(data[col].unique()):\n        FE.append(mod)\n        data[mod] = (data[col] == mod).astype(int)\n```\ninto this one:\n```\nFE = list(data.Sex.unique()) + list(data.SmokingStatus.unique())\ndata = pd.concat([\n    data,\n    pd.get_dummies(data.Sex),\n    pd.get_dummies(data.SmokingStatus)\n], axis=1)\n```\n**3.**[optional] This one:\n```\ndata['age'] = (data['Age'] - data['Age'].min()) / (data['Age'].max() - data['Age'].min())\ndata['BASE'] = (data['min_FVC'] - data['min_FVC'].min() ) / ( data['min_FVC'].max() - data['min_FVC'].min() )\ndata['week'] = (data['base_week'] - data['base_week'].min() ) / ( data['base_week'].max() - data['base_week'].min() )\ndata['percent'] = (data['Percent'] - data['Percent'].min() ) / ( data['Percent'].max() - data['Percent'].min() )\nFE += ['age','percent','week','BASE']\n```\ninto this one:\n```\ndef get_fillness(series):\n    return (series - series.min()) / (series.max() - series.min())\n\ndata['age'] = get_fillness(data['Age'])\ndata['BASE'] = get_fillness(data['min_FVC'])\ndata['week'] = get_fillness(data['base_week'])\ndata['percent'] = get_fillness(data['Percent'])\n\nFE += ['age','percent','week','BASE']\n```\n\n**Proofs:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F593e310c99d10311e44d575bed88484b%2FProof.png?generation=1598903196314916&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F007734f1ed6097bd990fe7a7bca70dea%2FProof2.png?generation=1598903680222466&alt=media)",
    "994317": "Thanks for sharing!\nAbout normalization: most of kernels normalize the tabular data with the train AND the test sets (which is actually hidden and made during submission because the training portion is included in the inference kernel). However it's a good practice to normalize train data with only train data and test data with only test data because of potential leaks. In this challenge we only want that our solution performs the best on the competition set so should we continue to normalize with (train+test) ? so if we voluntary introduces a leak during training, would the performance on the competition test set be better ?",
    "994028": "Thanks for sharing. I always use my own code because of this reason. Some of them are still too long though. You can use this one-liner for baseline features. (You have to make sure the data is sorted by `[Patient, Weeks]`)\n\n`df['FVC_Baseline'] = df.groupby('Patient').transform('first')['FVC']`",
    "998433": "Hey @koza4ukdmitrij, very clean approach to optimize some code which I used..thanks for the hints! 👍👍",
    "995388": "Thanks for sharing , helps to simplify the data preparation script 😃",
    "994241": "This is pretty good. I actually used some of this myself and you might not have any problems with this data set but when you do the pd.get_dummies() it's better to make the column into categorical first because if for e.g., only males were present on the train set then you would not get the right one-hot-encoding.",
    "998377": "Thanks for making this clean. I did use this code first, but then I was having a hard time to understand the script. I decided to rewrite almost the whole notebook with the addition of comments, and now navigating to a section and understanding scripts feels easy. @chrisden did a good job with his [well-documented notebook](https://www.kaggle.com/chrisden/6-82-quantile-reg-lr-schedulers-checkpoints).",
    "994222": "Thanks for sharing. ",
    "1000457": "",
    "1001145": "Thanks for making this clean."
  }
}