{
  "id": 80250,
  "title": "A framework for fast feature extraction",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/80250",
  "author_name": "",
  "post_date": "2019-02-12T05:48:42.416539Z",
  "votes": 116,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I present to you the framework i am using for extracting features from this dataset.</p>\n\n<p>Please let me know what you think:</p>\n\n<pre><code>import numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nfrom joblib import Parallel, delayed\n\n\nclass FeatureGenerator(object):\n    def __init__(self, dtype, n_jobs=1, chunk_size=None):\n        self.chunk_size = chunk_size\n        self.dtype = dtype\n        self.filename = None\n        self.n_jobs = n_jobs\n        self.test_files = []\n        if self.dtype == 'train':\n            self.filename = '../input/train.csv'\n            self.total_data = int(629145481 / self.chunk_size)\n        else:\n            submission = pd.read_csv('../input/sample_submission.csv')\n            for seg_id in submission.seg_id.values:\n                self.test_files.append((seg_id, '../input/test/' + seg_id + '.csv'))\n            self.total_data = int(len(submission))\n\n    def read_chunks(self):\n        if self.dtype == 'train':\n            iter_df = pd.read_csv(self.filename, iterator=True, chunksize=self.chunk_size,\n                                  dtype={'acoustic_data': np.float64, 'time_to_failure': np.float64})\n            for counter, df in enumerate(iter_df):\n                x = df.acoustic_data.values\n                y = df.time_to_failure.values[-1]\n                seg_id = 'train_' + str(counter)\n                yield seg_id, x, y\n        else:\n            for seg_id, f in self.test_files:\n                df = pd.read_csv(f, dtype={'acoustic_data': np.float64})\n                x = df.acoustic_data.values\n                yield seg_id, x, -999\n\n    def features(self, x, y, seg_id):\n        feature_dict = dict()\n        feature_dict['target'] = y\n        feature_dict['seg_id'] = seg_id\n\n        # create features here\n        # for example:\n        # feature_dict['mean'] = np.mean(x)\n\n        return feature_dict\n\n    def generate(self):\n        feature_list = []\n        res = Parallel(n_jobs=self.n_jobs,\n                       backend='threading')(delayed(self.features)(x, y, s)\n                                            for s, x, y in tqdm(self.read_chunks(), total=self.total_data))\n        for r in res:\n            feature_list.append(r)\n        return pd.DataFrame(feature_list)\n\n\ntraining_fg = FeatureGenerator(dtype='train', n_jobs=4, chunk_size=150000)\ntraining_data = training_fg.generate()\n\ntest_fg = FeatureGenerator(dtype='test', n_jobs=4, chunk_size=None)\ntest_data = test_fg.generate()\n\ntraining_data.to_csv(\"../input/train_features.csv\", index=False)\ntest_data.to_csv(\"../input/test_features.csv\", index=False)\n</code></pre>",
  "messages": [
    {
      "id": "469947",
      "postDate": "02/12/2019 05:48:42",
      "content": "<p>I present to you the framework i am using for extracting features from this dataset.</p>\n\n<p>Please let me know what you think:</p>\n\n<pre><code>import numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nfrom joblib import Parallel, delayed\n\n\nclass FeatureGenerator(object):\n    def __init__(self, dtype, n_jobs=1, chunk_size=None):\n        self.chunk_size = chunk_size\n        self.dtype = dtype\n        self.filename = None\n        self.n_jobs = n_jobs\n        self.test_files = []\n        if self.dtype == 'train':\n            self.filename = '../input/train.csv'\n            self.total_data = int(629145481 / self.chunk_size)\n        else:\n            submission = pd.read_csv('../input/sample_submission.csv')\n            for seg_id in submission.seg_id.values:\n                self.test_files.append((seg_id, '../input/test/' + seg_id + '.csv'))\n            self.total_data = int(len(submission))\n\n    def read_chunks(self):\n        if self.dtype == 'train':\n            iter_df = pd.read_csv(self.filename, iterator=True, chunksize=self.chunk_size,\n                                  dtype={'acoustic_data': np.float64, 'time_to_failure': np.float64})\n            for counter, df in enumerate(iter_df):\n                x = df.acoustic_data.values\n                y = df.time_to_failure.values[-1]\n                seg_id = 'train_' + str(counter)\n                yield seg_id, x, y\n        else:\n            for seg_id, f in self.test_files:\n                df = pd.read_csv(f, dtype={'acoustic_data': np.float64})\n                x = df.acoustic_data.values\n                yield seg_id, x, -999\n\n    def features(self, x, y, seg_id):\n        feature_dict = dict()\n        feature_dict['target'] = y\n        feature_dict['seg_id'] = seg_id\n\n        # create features here\n        # for example:\n        # feature_dict['mean'] = np.mean(x)\n\n        return feature_dict\n\n    def generate(self):\n        feature_list = []\n        res = Parallel(n_jobs=self.n_jobs,\n                       backend='threading')(delayed(self.features)(x, y, s)\n                                            for s, x, y in tqdm(self.read_chunks(), total=self.total_data))\n        for r in res:\n            feature_list.append(r)\n        return pd.DataFrame(feature_list)\n\n\ntraining_fg = FeatureGenerator(dtype='train', n_jobs=4, chunk_size=150000)\ntraining_data = training_fg.generate()\n\ntest_fg = FeatureGenerator(dtype='test', n_jobs=4, chunk_size=None)\ntest_data = test_fg.generate()\n\ntraining_data.to_csv(\"../input/train_features.csv\", index=False)\ntest_data.to_csv(\"../input/test_features.csv\", index=False)\n</code></pre>",
      "rawMarkdown": "I present to you the framework i am using for extracting features from this dataset.\n\nPlease let me know what you think:\n\n\n    import numpy as np\n    import pandas as pd\n    from tqdm import tqdm\n    from joblib import Parallel, delayed\n\n\n    class FeatureGenerator(object):\n        def __init__(self, dtype, n_jobs=1, chunk_size=None):\n            self.chunk_size = chunk_size\n            self.dtype = dtype\n            self.filename = None\n            self.n_jobs = n_jobs\n            self.test_files = []\n            if self.dtype == 'train':\n                self.filename = '../input/train.csv'\n                self.total_data = int(629145481 / self.chunk_size)\n            else:\n                submission = pd.read_csv('../input/sample_submission.csv')\n                for seg_id in submission.seg_id.values:\n                    self.test_files.append((seg_id, '../input/test/' + seg_id + '.csv'))\n                self.total_data = int(len(submission))\n\n        def read_chunks(self):\n            if self.dtype == 'train':\n                iter_df = pd.read_csv(self.filename, iterator=True, chunksize=self.chunk_size,\n                                      dtype={'acoustic_data': np.float64, 'time_to_failure': np.float64})\n                for counter, df in enumerate(iter_df):\n                    x = df.acoustic_data.values\n                    y = df.time_to_failure.values[-1]\n                    seg_id = 'train_' + str(counter)\n                    yield seg_id, x, y\n            else:\n                for seg_id, f in self.test_files:\n                    df = pd.read_csv(f, dtype={'acoustic_data': np.float64})\n                    x = df.acoustic_data.values\n                    yield seg_id, x, -999\n\n        def features(self, x, y, seg_id):\n            feature_dict = dict()\n            feature_dict['target'] = y\n            feature_dict['seg_id'] = seg_id\n\n            # create features here\n            # for example:\n            # feature_dict['mean'] = np.mean(x)\n\n            return feature_dict\n\n        def generate(self):\n            feature_list = []\n            res = Parallel(n_jobs=self.n_jobs,\n                           backend='threading')(delayed(self.features)(x, y, s)\n                                                for s, x, y in tqdm(self.read_chunks(), total=self.total_data))\n            for r in res:\n                feature_list.append(r)\n            return pd.DataFrame(feature_list)\n\n\n    training_fg = FeatureGenerator(dtype='train', n_jobs=4, chunk_size=150000)\n    training_data = training_fg.generate()\n\n    test_fg = FeatureGenerator(dtype='test', n_jobs=4, chunk_size=None)\n    test_data = test_fg.generate()\n\n    training_data.to_csv(\"../input/train_features.csv\", index=False)\n    test_data.to_csv(\"../input/test_features.csv\", index=False)",
      "votes": null
    },
    {
      "id": "469958",
      "postDate": "02/12/2019 06:07:08",
      "content": "<p>Thanks for the share!\n<a href=\"https://joblib.readthedocs.io/en/latest/\">Joblib docs</a> link for the unaware (like me).</p>",
      "rawMarkdown": "Thanks for the share!\n[Joblib docs](https://joblib.readthedocs.io/en/latest/) link for the unaware (like me).",
      "votes": null
    },
    {
      "id": "469960",
      "postDate": "02/12/2019 06:12:06",
      "content": "<p>Don't delete it.</p>",
      "rawMarkdown": "Don't delete it.",
      "votes": null
    },
    {
      "id": "469964",
      "postDate": "02/12/2019 06:16:34",
      "content": "<p>delete the post? :D</p>",
      "rawMarkdown": "delete the post? :D",
      "votes": null
    },
    {
      "id": "470148",
      "postDate": "02/12/2019 13:24:04",
      "content": "<p>This is fantastic, Abhiskeh, thank you! I was having trouble getting over memory error problems when running locally, so your code is incredibly useful.</p>\n\n<p>I wasn't aware of the joblib library, so this is also a great learning experience for me.</p>\n\n<p>I'm really grateful to you for sharing this valuable lesson.</p>",
      "rawMarkdown": "This is fantastic, Abhiskeh, thank you! I was having trouble getting over memory error problems when running locally, so your code is incredibly useful.\n\nI wasn't aware of the joblib library, so this is also a great learning experience for me.\n\nI'm really grateful to you for sharing this valuable lesson.",
      "votes": null
    },
    {
      "id": "470224",
      "postDate": "02/12/2019 16:04:38",
      "content": "<p>Thanks it helped me alot</p>",
      "rawMarkdown": "Thanks it helped me alot",
      "votes": null
    },
    {
      "id": "470250",
      "postDate": "02/12/2019 16:41:50",
      "content": "<p>Thanks Abhi. Only one small suggestion, if I may: in the function FeatureGenerator.features(…), you really \"should\" remove the parameter y; the features should be a function only of x.</p>",
      "rawMarkdown": "Thanks Abhi. Only one small suggestion, if I may: in the function FeatureGenerator.features(…), you really \"should\" remove the parameter y; the features should be a function only of x.",
      "votes": null
    },
    {
      "id": "470252",
      "postDate": "02/12/2019 16:43:46",
      "content": "<p>We beg you: <strong>please</strong> don't delete the post!!!</p>",
      "rawMarkdown": "We beg you: **please** don't delete the post!!!",
      "votes": null
    },
    {
      "id": "470295",
      "postDate": "02/12/2019 18:00:27",
      "content": "<p>y is only stored so that you dont lose the target variable. i dont use it for creating features, neither should anyone else. :)</p>",
      "rawMarkdown": "y is only stored so that you dont lose the target variable. i dont use it for creating features, neither should anyone else. :)",
      "votes": null
    },
    {
      "id": "471318",
      "postDate": "02/14/2019 09:52:31",
      "content": "<p>Very Useful indeed thanks for sharing</p>",
      "rawMarkdown": "Very Useful indeed thanks for sharing",
      "votes": null
    },
    {
      "id": "471523",
      "postDate": "02/14/2019 15:17:48",
      "content": "<p>This is a great share. Thanks</p>",
      "rawMarkdown": "This is a great share. Thanks",
      "votes": null
    },
    {
      "id": "471864",
      "postDate": "02/15/2019 02:12:17",
      "content": "<p>hu</p>",
      "rawMarkdown": "hu",
      "votes": null
    },
    {
      "id": "475313",
      "postDate": "02/20/2019 15:52:22",
      "content": "<p>thanks for this, it is faster than mine</p>",
      "rawMarkdown": "thanks for this, it is faster than mine",
      "votes": null
    },
    {
      "id": "476962",
      "postDate": "02/23/2019 14:44:12",
      "content": "<p>Thanks Abhishek...concepts in the post helped me to achieve slightly better performance.</p>",
      "rawMarkdown": "Thanks Abhishek...concepts in the post helped me to achieve slightly better performance.",
      "votes": null
    },
    {
      "id": "478968",
      "postDate": "02/26/2019 21:18:18",
      "content": "<p>I would love to see this in R :) - But anyways, job well done.</p>",
      "rawMarkdown": "I would love to see this in R :) - But anyways, job well done.",
      "votes": null
    },
    {
      "id": "482140",
      "postDate": "03/02/2019 11:49:54",
      "content": "<p>thx</p>",
      "rawMarkdown": "thx",
      "votes": null
    },
    {
      "id": "482557",
      "postDate": "03/03/2019 08:11:54",
      "content": "<p>thanks</p>",
      "rawMarkdown": "thanks",
      "votes": null
    },
    {
      "id": "486302",
      "postDate": "03/08/2019 15:16:23",
      "content": "<p>Good code, thank you. The only unknown thing for me here is RAM utilization: despite reading file in chunks and using joblib, this code utilizes ~10 GiB of RAM (but output dataframes - <code>training_data</code> and <code>test_data</code> - both occupies less than 1 MiB of space), and I was not able to release this space using gc. Moreover, this piece of code:</p>\n\n<p>```\nreader = pd.read_csv('../input/train.csv',\n                    dtype={'acoustic_data': np.int16,\n                           'time_to_failure': np.float64},\n                    chunksize=10_000_000,\n                    iterator=True)</p>\n\n<p>for chunk in reader:\n    print(chunk.shape)\n```</p>\n\n<p>will also consume ~10 GiB of RAM, which could not be released by deleting variables and using gc.collect(). Does anybody knows, is it a bug or feature? Is there any way to release memory after reading csv file in pandas?</p>\n\n<p>P.S. I am monitoring RAM utilization using RAM indicator in kernel. Hoping that this is correct approach</p>",
      "rawMarkdown": "Good code, thank you. The only unknown thing for me here is RAM utilization: despite reading file in chunks and using joblib, this code utilizes ~10 GiB of RAM (but output dataframes - `training_data` and `test_data` - both occupies less than 1 MiB of space), and I was not able to release this space using gc. Moreover, this piece of code:\n\n```\nreader = pd.read_csv('../input/train.csv',\n                    dtype={'acoustic_data': np.int16,\n                           'time_to_failure': np.float64},\n                    chunksize=10_000_000,\n                    iterator=True)\n\nfor chunk in reader:\n    print(chunk.shape)\n```\n\nwill also consume ~10 GiB of RAM, which could not be released by deleting variables and using gc.collect(). Does anybody knows, is it a bug or feature? Is there any way to release memory after reading csv file in pandas?\n\nP.S. I am monitoring RAM utilization using RAM indicator in kernel. Hoping that this is correct approach",
      "votes": null
    },
    {
      "id": "506896",
      "postDate": "04/04/2019 02:22:32",
      "content": "<p>Thanks it is very helpful </p>",
      "rawMarkdown": "Thanks it is very helpful",
      "votes": null
    },
    {
      "id": "532497",
      "postDate": "05/17/2019 05:50:21",
      "content": "<p>That's really nice</p>",
      "rawMarkdown": "That's really nice",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 469958,
      "author_name": "init27",
      "author_url": "",
      "post_date": "02/12/2019 06:07:08",
      "content": "<p>Thanks for the share!\n<a href=\"https://joblib.readthedocs.io/en/latest/\">Joblib docs</a> link for the unaware (like me).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 469960,
      "author_name": "apartmentguru",
      "author_url": "",
      "post_date": "02/12/2019 06:12:06",
      "content": "<p>Don't delete it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 469964,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "02/12/2019 06:16:34",
          "content": "<p>delete the post? :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 470252,
          "author_name": "danjel",
          "author_url": "",
          "post_date": "02/12/2019 16:43:46",
          "content": "<p>We beg you: <strong>please</strong> don't delete the post!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 470148,
      "author_name": "marcogorelli",
      "author_url": "",
      "post_date": "02/12/2019 13:24:04",
      "content": "<p>This is fantastic, Abhiskeh, thank you! I was having trouble getting over memory error problems when running locally, so your code is incredibly useful.</p>\n\n<p>I wasn't aware of the joblib library, so this is also a great learning experience for me.</p>\n\n<p>I'm really grateful to you for sharing this valuable lesson.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 470224,
      "author_name": "kaushikruth123",
      "author_url": "",
      "post_date": "02/12/2019 16:04:38",
      "content": "<p>Thanks it helped me alot</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 470250,
      "author_name": "danjel",
      "author_url": "",
      "post_date": "02/12/2019 16:41:50",
      "content": "<p>Thanks Abhi. Only one small suggestion, if I may: in the function FeatureGenerator.features(…), you really \"should\" remove the parameter y; the features should be a function only of x.</p>",
      "votes": null,
      "replies": [
        {
          "id": 470295,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "02/12/2019 18:00:27",
          "content": "<p>y is only stored so that you dont lose the target variable. i dont use it for creating features, neither should anyone else. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 471318,
      "author_name": "pavansanagapati",
      "author_url": "",
      "post_date": "02/14/2019 09:52:31",
      "content": "<p>Very Useful indeed thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 471523,
      "author_name": "itsmesunil",
      "author_url": "",
      "post_date": "02/14/2019 15:17:48",
      "content": "<p>This is a great share. Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 471864,
      "author_name": "pandagon",
      "author_url": "",
      "post_date": "02/15/2019 02:12:17",
      "content": "<p>hu</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 475313,
      "author_name": "ikanoo",
      "author_url": "",
      "post_date": "02/20/2019 15:52:22",
      "content": "<p>thanks for this, it is faster than mine</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 476962,
      "author_name": "ybshankar010",
      "author_url": "",
      "post_date": "02/23/2019 14:44:12",
      "content": "<p>Thanks Abhishek...concepts in the post helped me to achieve slightly better performance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 478968,
      "author_name": "theupgrade",
      "author_url": "",
      "post_date": "02/26/2019 21:18:18",
      "content": "<p>I would love to see this in R :) - But anyways, job well done.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 482140,
      "author_name": "ralphy",
      "author_url": "",
      "post_date": "03/02/2019 11:49:54",
      "content": "<p>thx</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 482557,
      "author_name": "chenhao2334",
      "author_url": "",
      "post_date": "03/03/2019 08:11:54",
      "content": "<p>thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 486302,
      "author_name": "paulvonbragin",
      "author_url": "",
      "post_date": "03/08/2019 15:16:23",
      "content": "<p>Good code, thank you. The only unknown thing for me here is RAM utilization: despite reading file in chunks and using joblib, this code utilizes ~10 GiB of RAM (but output dataframes - <code>training_data</code> and <code>test_data</code> - both occupies less than 1 MiB of space), and I was not able to release this space using gc. Moreover, this piece of code:</p>\n\n<p>```\nreader = pd.read_csv('../input/train.csv',\n                    dtype={'acoustic_data': np.int16,\n                           'time_to_failure': np.float64},\n                    chunksize=10_000_000,\n                    iterator=True)</p>\n\n<p>for chunk in reader:\n    print(chunk.shape)\n```</p>\n\n<p>will also consume ~10 GiB of RAM, which could not be released by deleting variables and using gc.collect(). Does anybody knows, is it a bug or feature? Is there any way to release memory after reading csv file in pandas?</p>\n\n<p>P.S. I am monitoring RAM utilization using RAM indicator in kernel. Hoping that this is correct approach</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 506896,
      "author_name": "chenlongwang",
      "author_url": "",
      "post_date": "04/04/2019 02:22:32",
      "content": "<p>Thanks it is very helpful </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 532497,
      "author_name": "shubhamsingh060",
      "author_url": "",
      "post_date": "05/17/2019 05:50:21",
      "content": "<p>That's really nice</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "469947": "I present to you the framework i am using for extracting features from this dataset.\n\nPlease let me know what you think:\n\n\n    import numpy as np\n    import pandas as pd\n    from tqdm import tqdm\n    from joblib import Parallel, delayed\n\n\n    class FeatureGenerator(object):\n        def __init__(self, dtype, n_jobs=1, chunk_size=None):\n            self.chunk_size = chunk_size\n            self.dtype = dtype\n            self.filename = None\n            self.n_jobs = n_jobs\n            self.test_files = []\n            if self.dtype == 'train':\n                self.filename = '../input/train.csv'\n                self.total_data = int(629145481 / self.chunk_size)\n            else:\n                submission = pd.read_csv('../input/sample_submission.csv')\n                for seg_id in submission.seg_id.values:\n                    self.test_files.append((seg_id, '../input/test/' + seg_id + '.csv'))\n                self.total_data = int(len(submission))\n\n        def read_chunks(self):\n            if self.dtype == 'train':\n                iter_df = pd.read_csv(self.filename, iterator=True, chunksize=self.chunk_size,\n                                      dtype={'acoustic_data': np.float64, 'time_to_failure': np.float64})\n                for counter, df in enumerate(iter_df):\n                    x = df.acoustic_data.values\n                    y = df.time_to_failure.values[-1]\n                    seg_id = 'train_' + str(counter)\n                    yield seg_id, x, y\n            else:\n                for seg_id, f in self.test_files:\n                    df = pd.read_csv(f, dtype={'acoustic_data': np.float64})\n                    x = df.acoustic_data.values\n                    yield seg_id, x, -999\n\n        def features(self, x, y, seg_id):\n            feature_dict = dict()\n            feature_dict['target'] = y\n            feature_dict['seg_id'] = seg_id\n\n            # create features here\n            # for example:\n            # feature_dict['mean'] = np.mean(x)\n\n            return feature_dict\n\n        def generate(self):\n            feature_list = []\n            res = Parallel(n_jobs=self.n_jobs,\n                           backend='threading')(delayed(self.features)(x, y, s)\n                                                for s, x, y in tqdm(self.read_chunks(), total=self.total_data))\n            for r in res:\n                feature_list.append(r)\n            return pd.DataFrame(feature_list)\n\n\n    training_fg = FeatureGenerator(dtype='train', n_jobs=4, chunk_size=150000)\n    training_data = training_fg.generate()\n\n    test_fg = FeatureGenerator(dtype='test', n_jobs=4, chunk_size=None)\n    test_data = test_fg.generate()\n\n    training_data.to_csv(\"../input/train_features.csv\", index=False)\n    test_data.to_csv(\"../input/test_features.csv\", index=False)",
    "469958": "Thanks for the share!\n[Joblib docs](https://joblib.readthedocs.io/en/latest/) link for the unaware (like me).",
    "469960": "Don't delete it.",
    "469964": "delete the post? :D",
    "470148": "This is fantastic, Abhiskeh, thank you! I was having trouble getting over memory error problems when running locally, so your code is incredibly useful.\n\nI wasn't aware of the joblib library, so this is also a great learning experience for me.\n\nI'm really grateful to you for sharing this valuable lesson.",
    "470224": "Thanks it helped me alot",
    "470250": "Thanks Abhi. Only one small suggestion, if I may: in the function FeatureGenerator.features(…), you really \"should\" remove the parameter y; the features should be a function only of x.",
    "470252": "We beg you: **please** don't delete the post!!!",
    "470295": "y is only stored so that you dont lose the target variable. i dont use it for creating features, neither should anyone else. :)",
    "471318": "Very Useful indeed thanks for sharing",
    "471523": "This is a great share. Thanks",
    "471864": "hu",
    "475313": "thanks for this, it is faster than mine",
    "476962": "Thanks Abhishek...concepts in the post helped me to achieve slightly better performance.",
    "478968": "I would love to see this in R :) - But anyways, job well done.",
    "482140": "thx",
    "482557": "thanks",
    "486302": "Good code, thank you. The only unknown thing for me here is RAM utilization: despite reading file in chunks and using joblib, this code utilizes ~10 GiB of RAM (but output dataframes - `training_data` and `test_data` - both occupies less than 1 MiB of space), and I was not able to release this space using gc. Moreover, this piece of code:\n\n```\nreader = pd.read_csv('../input/train.csv',\n                    dtype={'acoustic_data': np.int16,\n                           'time_to_failure': np.float64},\n                    chunksize=10_000_000,\n                    iterator=True)\n\nfor chunk in reader:\n    print(chunk.shape)\n```\n\nwill also consume ~10 GiB of RAM, which could not be released by deleting variables and using gc.collect(). Does anybody knows, is it a bug or feature? Is there any way to release memory after reading csv file in pandas?\n\nP.S. I am monitoring RAM utilization using RAM indicator in kernel. Hoping that this is correct approach",
    "506896": "Thanks it is very helpful",
    "532497": "That's really nice"
  },
  "source": "meta"
}