{
  "id": 70837,
  "title": "How do you handle iterative features adding?",
  "url": "/competitions/PLAsTiCC-2018/discussion/70837",
  "author_name": "",
  "post_date": "2018-11-07T17:16:05.521334700Z",
  "votes": 1,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I am quite a novice to the IT/programming. So, I do the following\n1. calculate basic features set for test \n2. save them to csv file \n3. load that csv file as a dataframe, add new features  df['new_feature'] = new_feature_function\n4. save newcsv file with added features. Do all the same for train\nThe test features file is huge.\nAlternative way is to save every new feature as a numpy array of [[object_id, feature]], it has no label, and I save numpy array feature in a file named as a feature, not as good for manipulating as pandas header, and this yields to many files. </p>\n\n<p>Both solutions are bulky and not elegant. I hope there are better ways to handle feature adding for such large sets of objects, that I am not just aware of. Any good elegant solutions on this?  </p>",
  "messages": [
    {
      "id": "417063",
      "postDate": "11/07/2018 17:16:05",
      "content": "<p>I am quite a novice to the IT/programming. So, I do the following\n1. calculate basic features set for test \n2. save them to csv file \n3. load that csv file as a dataframe, add new features  df['new_feature'] = new_feature_function\n4. save newcsv file with added features. Do all the same for train\nThe test features file is huge.\nAlternative way is to save every new feature as a numpy array of [[object_id, feature]], it has no label, and I save numpy array feature in a file named as a feature, not as good for manipulating as pandas header, and this yields to many files. </p>\n\n<p>Both solutions are bulky and not elegant. I hope there are better ways to handle feature adding for such large sets of objects, that I am not just aware of. Any good elegant solutions on this?  </p>",
      "rawMarkdown": "I am quite a novice to the IT/programming. So, I do the following\n1. calculate basic features set for test \n2. save them to csv file \n3. load that csv file as a dataframe, add new features  df['new_feature'] = new_feature_function\n4. save newcsv file with added features. Do all the same for train\nThe test features file is huge.\nAlternative way is to save every new feature as a numpy array of [[object_id, feature]], it has no label, and I save numpy array feature in a file named as a feature, not as good for manipulating as pandas header, and this yields to many files. \n\nBoth solutions are bulky and not elegant. I hope there are better ways to handle feature adding for such large sets of objects, that I am not just aware of. Any good elegant solutions on this?",
      "votes": null
    },
    {
      "id": "417065",
      "postDate": "11/07/2018 17:20:04",
      "content": "<p>It was suggested to split test by object_id, and work with chunks of it.</p>",
      "rawMarkdown": "It was suggested to split test by object_id, and work with chunks of it.",
      "votes": null
    },
    {
      "id": "417067",
      "postDate": "11/07/2018 17:29:53",
      "content": "<p>I compute multiple test sets with aggregated features only on multiple kernels and then combine them as below:\n<code>\nfrom functools import reduce\nn=500000\nfor i, dfs in enumerate(zip(\n   pd.read_csv( '../input/path-to-csv/test_set1.csv' , iterator=True, chunksize= n),\n   pd.read_csv('../input/path-to-csv/test_set2.csv', iterator=True, chunksize= n),\n   pd.read_csv('../input/path-to-csv/test_set3.csv', iterator=True, chunksize= n))):\n    df = reduce(lambda left,right: pd.merge(left,right,on='object_id'), dfs)\n</code></p>",
      "rawMarkdown": "I compute multiple test sets with aggregated features only on multiple kernels and then combine them as below:\n```\nfrom functools import reduce\nn=500000\nfor i, dfs in enumerate(zip(\n   pd.read_csv( '../input/path-to-csv/test_set1.csv' , iterator=True, chunksize= n),\n   pd.read_csv('../input/path-to-csv/test_set2.csv', iterator=True, chunksize= n),\n   pd.read_csv('../input/path-to-csv/test_set3.csv', iterator=True, chunksize= n))):\n    df = reduce(lambda left,right: pd.merge(left,right,on='object_id'), dfs)\n```",
      "votes": null
    },
    {
      "id": "417078",
      "postDate": "11/07/2018 17:51:16",
      "content": "<p>My processing occurs in stages / code blocks. Up top I have something like:</p>\n\n<blockquote>\n  <p>DSET_PREP_1 = False\n  DSET_PREP_2 = DSET_PREP_1 OR False\n  DSET_PREP_3 = DSET_PREP_2 OR False\n  DSET_PREP_4 = DSET_PREP_3 OR False\n  DSET_PREP_5 = DSET_PREP_4 OR False</p>\n</blockquote>\n\n<p>Then in the body of the code</p>\n\n<blockquote>\n  <p>if not DSET_PREP1:\n     prep1 = pd.read_feather('./prep1.ftr')\n  else:\n     # .... compute prep1, then save it to prep1.ftr</p>\n  \n  <p>if not DSET_PREP2:\n     prep2 = pd.read_feather('./prep2.ftr')\n  else:\n     # .... compute prep2 using prep1, then save it to prep2.ftr</p>\n  \n  <p>if not DSET_PREP3:\n     prep3 = pd.read_feather('./prep3.ftr')\n  else:\n     # .... compute prep3 using prep2, then save it to prep3.ftr</p>\n</blockquote>\n\n<p>etc.</p>\n\n<p>So if I do something such as add a new stage of features in the pipeline, I'll simply set that stages DSET_PREP to True, and it'll recompute everything downstream.</p>\n\n<p>In other competitions with large amounts of data, sometimes I'll write <code>.ftr's</code> on a per-feature basis and then merge them.</p>",
      "rawMarkdown": "My processing occurs in stages / code blocks. Up top I have something like:\n\n\n&gt; DSET_PREP_1 = False\n&gt; DSET_PREP_2 = DSET_PREP_1 OR False\n&gt; DSET_PREP_3 = DSET_PREP_2 OR False\n&gt; DSET_PREP_4 = DSET_PREP_3 OR False\n&gt; DSET_PREP_5 = DSET_PREP_4 OR False\n\nThen in the body of the code\n\n&gt; if not DSET_PREP1:\n&gt;    prep1 = pd.read_feather('./prep1.ftr')\n&gt; else:\n&gt;    # .... compute prep1, then save it to prep1.ftr\n&gt; \n&gt; if not DSET_PREP2:\n&gt;    prep2 = pd.read_feather('./prep2.ftr')\n&gt; else:\n&gt;    # .... compute prep2 using prep1, then save it to prep2.ftr\n&gt; \n&gt; if not DSET_PREP3:\n&gt;    prep3 = pd.read_feather('./prep3.ftr')\n&gt; else:\n&gt;    # .... compute prep3 using prep2, then save it to prep3.ftr\n\netc.\n\nSo if I do something such as add a new stage of features in the pipeline, I'll simply set that stages DSET_PREP to True, and it'll recompute everything downstream.\n\nIn other competitions with large amounts of data, sometimes I'll write `.ftr's` on a per-feature basis and then merge them.",
      "votes": null
    },
    {
      "id": "417162",
      "postDate": "11/07/2018 21:45:31",
      "content": "<p>You can store all your data files in an efficient database (Spark, BigQuery, Athena etc. Needs to be a columnar database. Most of the online services have free tiers), do your basic feature engineering there, add advanced features (the ones you can't do in SQL) as new tables with object_id as key, and whenever you need a dataset, use a simple SQL script to pull all relevant columns together and then export as CSV or directly load into Python/R. This way you can avoid ever loading the full time series dataset into memory when training, and adding / removing features can be done quite fast and separately from your training script.</p>",
      "rawMarkdown": "You can store all your data files in an efficient database (Spark, BigQuery, Athena etc. Needs to be a columnar database. Most of the online services have free tiers), do your basic feature engineering there, add advanced features (the ones you can't do in SQL) as new tables with object_id as key, and whenever you need a dataset, use a simple SQL script to pull all relevant columns together and then export as CSV or directly load into Python/R. This way you can avoid ever loading the full time series dataset into memory when training, and adding / removing features can be done quite fast and separately from your training script.",
      "votes": null
    },
    {
      "id": "417231",
      "postDate": "11/08/2018 00:59:07",
      "content": "<p>I did in chunks, works, but I have a problem of passing object id as I chunk by object ids... I do it later <code>df_test['object_id'] = meta_test['object_id']</code> as they are ascending in both np.unique and meta file, but it's not not optimal (they may change order in general case)... a bit stuck how to make it work properly</p>",
      "rawMarkdown": "I did in chunks, works, but I have a problem of passing object id as I chunk by object ids... I do it later ```df_test['object_id'] = meta_test['object_id'] ``` as they are ascending in both np.unique and meta file, but it's not not optimal (they may change order in general case)... a bit stuck how to make it work properly",
      "votes": null
    },
    {
      "id": "417378",
      "postDate": "11/08/2018 06:56:56",
      "content": "<p>You can transform your numpy arrays into pandas data frames then join on object_id.</p>",
      "rawMarkdown": "You can transform your numpy arrays into pandas data frames then join on object_id.",
      "votes": null
    },
    {
      "id": "417497",
      "postDate": "11/08/2018 11:23:25",
      "content": "<p>this is probably better as npy files are small and csv is huge... but then i save each feature in npy array named as a feature like 'skew', and it's not as easy to manipulate as df['skew']... Can I load numpy file names as headers in that new dataframe? I can do it manually, but it's nice to make just a string list of features to join, load those numpy arrays and join. </p>",
      "rawMarkdown": "this is probably better as npy files are small and csv is huge... but then i save each feature in npy array named as a feature like 'skew', and it's not as easy to manipulate as df['skew']... Can I load numpy file names as headers in that new dataframe? I can do it manually, but it's nice to make just a string list of features to join, load those numpy arrays and join.",
      "votes": null
    },
    {
      "id": "417511",
      "postDate": "11/08/2018 11:58:35",
      "content": "<p>thank you <a href=\"/authman\">@authman</a> I have not met ftr type, will read about it now. Why did you choose ftr and not npy or something else, by the way, is it better for the size?</p>",
      "rawMarkdown": "thank you @authman I have not met ftr type, will read about it now. Why did you choose ftr and not npy or something else, by the way, is it better for the size?",
      "votes": null
    },
    {
      "id": "417558",
      "postDate": "11/08/2018 13:00:15",
      "content": "<p>Here is a code I used in the past.  Data is a pandas dataframe, and features are stored as numpy arrays.</p>\n\n<pre><code>for feat in features:\n    if feat in data.columns:\n        continue\n    print(feat)\n    data[feat] = np.load('../data/' +  feat + '.npy')\n</code></pre>\n\n<p>This assumes that the rows are ordered consistently across feature files.  But why wouldn't they if you construct all of the in a consistent way?</p>",
      "rawMarkdown": "Here is a code I used in the past.  Data is a pandas dataframe, and features are stored as numpy arrays.\n\n    for feat in features:\n        if feat in data.columns:\n            continue\n        print(feat)\n        data[feat] = np.load('../data/' +  feat + '.npy')\n\nThis assumes that the rows are ordered consistently across feature files.  But why wouldn't they if you construct all of the in a consistent way?",
      "votes": null
    },
    {
      "id": "417654",
      "postDate": "11/08/2018 15:33:29",
      "content": "<p>Thank you ! </p>",
      "rawMarkdown": "Thank you !",
      "votes": null
    },
    {
      "id": "417677",
      "postDate": "11/08/2018 16:17:59",
      "content": "<p><code>This way you can avoid ever loading the full time series dataset into memory when training</code> -- the best solution so far, proper professional way, but too advanced for me at this stage. I'll try it in the next competitions when my skills improve</p>",
      "rawMarkdown": "```This way you can avoid ever loading the full time series dataset into memory when training``` -- the best solution so far, proper professional way, but too advanced for me at this stage. I'll try it in the next competitions when my skills improve",
      "votes": null
    },
    {
      "id": "417699",
      "postDate": "11/08/2018 16:51:11",
      "content": "<p>I'd say if you know a little bit of SQL and don't mind paying less than one dollar per month for some cloud storage, it is probably one of the easier methods. There are a few datasets on Kaggle that are hosted on BigQuery, which is a good place to practice. I can share some of my scripts if you would like to know more.</p>",
      "rawMarkdown": "I'd say if you know a little bit of SQL and don't mind paying less than one dollar per month for some cloud storage, it is probably one of the easier methods. There are a few datasets on Kaggle that are hosted on BigQuery, which is a good place to practice. I can share some of my scripts if you would like to know more.",
      "votes": null
    },
    {
      "id": "417758",
      "postDate": "11/08/2018 18:36:02",
      "content": "<p>It's the test, not train, that gives me pain... At the moment I am fighting with my program bugs trying to load calculated features into classifier by chunks... trying to put it all together and make it work... If I manage this, I'll come back to database. </p>\n\n<p>What do you use (Spark, BigQuery, Hadoop) ? \nI know SELECT and WHERE for SQL :), could be enough, but I need to see what to install and how and where...did not work with databases yet</p>",
      "rawMarkdown": "It's the test, not train, that gives me pain... At the moment I am fighting with my program bugs trying to load calculated features into classifier by chunks... trying to put it all together and make it work... If I manage this, I'll come back to database. \n\nWhat do you use (Spark, BigQuery, Hadoop) ? \nI know SELECT and WHERE for SQL :), could be enough, but I need to see what to install and how and where...did not work with databases yet",
      "votes": null
    },
    {
      "id": "417835",
      "postDate": "11/08/2018 21:06:44",
      "content": "<p>I just submitted a script for a utility I'm using to read in the training/test files in chunks, separating observations into pandas.DataFrames by object_id. I'm sure there are better ways to store and retrieve the data, but I've settled on this approach for now. I've got to focus my efforts on improved training features!</p>\n\n<p>Have a look here: <a href=\"https://www.kaggle.com/gregbehm/plasticc-data-reader-get-objects-by-id\">https://www.kaggle.com/gregbehm/plasticc-data-reader-get-objects-by-id</a></p>\n\n<p>I'm using a modified version with <code>pandas.read_sql_query</code> to read SQLite database files I've created, but for demo purposes in this kernel I show reading CSV files.</p>\n\n<p>There's demo code in main() at the bottom of the script.</p>",
      "rawMarkdown": "I just submitted a script for a utility I'm using to read in the training/test files in chunks, separating observations into pandas.DataFrames by object_id. I'm sure there are better ways to store and retrieve the data, but I've settled on this approach for now. I've got to focus my efforts on improved training features!\n\nHave a look here: https://www.kaggle.com/gregbehm/plasticc-data-reader-get-objects-by-id\n\nI'm using a modified version with ```pandas.read_sql_query``` to read SQLite database files I've created, but for demo purposes in this kernel I show reading CSV files.\n\nThere's demo code in main() at the bottom of the script.",
      "votes": null
    },
    {
      "id": "418066",
      "postDate": "11/09/2018 08:32:58",
      "content": "<p>Here is a basic idea of how to do it:</p>\n\n<p>Suppose I have tables called <code>test_meta</code> and <code>test_series</code> that I uploaded to BigQuery (not a quick process, but the easiest way to do it is to spin up a cheapest compute engine, download Kaggle dataset there, use <code>gsutil -m cp</code> to copy everything to a cloud storage bucket then import to BigQuery).</p>\n\n<ol>\n<li>First I want to add some basic features like mean and standard deviation. Instead of directly adding new columns to my dataset, I can run a query that can calculate these values and store it as a view. e.g. my <code>additional stats</code> view:</li>\n</ol>\n\n<p><code>\nSELECT <br>\n    object_id, AVG(flux_err) AS flux_err_mean, <br>\n    STDDEV(flux_err) AS flux_err_std, <br>\n    AVG(detected) AS detected_mean <br>\nFROM `project.astro.test_series` <br>\nGROUP BY object_id\n</code></p>\n\n<ol>\n<li><p>Then I have to append some new feature columns that I calculated from the time series to the final feature set. I store the features in the table <code>test_features</code>.</p></li>\n<li><p>Finally, I want to combine the simple features together with the new features. I can do the following:</p></li>\n</ol>\n\n<p><code>\nSELECT\n  a.*, <br>\n  b.flux_err_mean, <br>\n  b.flux_err_std, <br>\n  b.detected_mean, <br>\n  c.hostgal_photoz, <br>\n  c.ddf <br>\nFROM\n  `project.astro.test_features` a, <br>\n  `project.astro.additional_stats` b, <br>\n  `project.astro.test_meta` c <br>\nWHERE\n  a.object_id = b.object_id <br>\n  AND a.object_id = c.object_id\n</code></p>\n\n<p>You can mix in all the features you want and exclude the ones you don't need.\nThis script should be able to finish under one minute for the training set and at most 10 minutes for test (c.f. maybe close to one hour of Pandas preprocessing in Python). Now you can export the resulting table to Cloud Storage then download from there, or directly pull into your Python environment using the BigQuery API (not recommended because of the size). Because you did most of the full-dataset operations in the cloud, you don't have to deal with split time series files and finding solutions to not split time series of one object into two files. You receive a file (or files) that is ready to be ingested by a training algorithm.</p>\n\n<p>Google BigQuery offers 1TB free processing per month, and the storage cost of the data is tiny, so you normally do not have to worry about billing. It is a great way to perform fast data preprocessing on a less powerful computer.</p>",
      "rawMarkdown": "Here is a basic idea of how to do it:\n\nSuppose I have tables called `test_meta` and `test_series` that I uploaded to BigQuery (not a quick process, but the easiest way to do it is to spin up a cheapest compute engine, download Kaggle dataset there, use `gsutil -m cp` to copy everything to a cloud storage bucket then import to BigQuery).\n\n1. First I want to add some basic features like mean and standard deviation. Instead of directly adding new columns to my dataset, I can run a query that can calculate these values and store it as a view. e.g. my `additional stats` view:\n\n```\nSELECT  \n    object_id, AVG(flux_err) AS flux_err_mean,  \n    STDDEV(flux_err) AS flux_err_std,  \n    AVG(detected) AS detected_mean  \nFROM `project.astro.test_series`  \nGROUP BY object_id\n```\n\n2. Then I have to append some new feature columns that I calculated from the time series to the final feature set. I store the features in the table `test_features`.\n\n3. Finally, I want to combine the simple features together with the new features. I can do the following:\n\n```\nSELECT\n  a.*,  \n  b.flux_err_mean,  \n  b.flux_err_std,  \n  b.detected_mean,  \n  c.hostgal_photoz,  \n  c.ddf  \nFROM\n  `project.astro.test_features` a,  \n  `project.astro.additional_stats` b,  \n  `project.astro.test_meta` c   \nWHERE\n  a.object_id = b.object_id  \n  AND a.object_id = c.object_id\n```\n\nYou can mix in all the features you want and exclude the ones you don't need.\nThis script should be able to finish under one minute for the training set and at most 10 minutes for test (c.f. maybe close to one hour of Pandas preprocessing in Python). Now you can export the resulting table to Cloud Storage then download from there, or directly pull into your Python environment using the BigQuery API (not recommended because of the size). Because you did most of the full-dataset operations in the cloud, you don't have to deal with split time series files and finding solutions to not split time series of one object into two files. You receive a file (or files) that is ready to be ingested by a training algorithm.\n\nGoogle BigQuery offers 1TB free processing per month, and the storage cost of the data is tiny, so you normally do not have to worry about billing. It is a great way to perform fast data preprocessing on a less powerful computer.",
      "votes": null
    },
    {
      "id": "418286",
      "postDate": "11/09/2018 15:40:02",
      "content": "<p>This script should be able to finish under one minute for the training set and at most 10 minutes for test (c.f. maybe close to one hour of Pandas preprocessing in Python). -- this is impressive!</p>\n\n<ol>\n<li>I have GCP already and copied kaggle datasets there. Do i still need to copy everything to a cloud storage bucket? (I just have them in my instance main directory, instance has ssd drive)</li>\n<li>Some features are calculated via python packages, so then to run if from sql query I need first to load that flux row to a numpy, than process, than save. I can load via panda sql, I think, do i need those sql then, will it speed up the process for advanced features? Basic features that can be done via sql query I already have...</li>\n</ol>\n\n<p>thank you for your commands!</p>",
      "rawMarkdown": "This script should be able to finish under one minute for the training set and at most 10 minutes for test (c.f. maybe close to one hour of Pandas preprocessing in Python). -- this is impressive!\n\n1. I have GCP already and copied kaggle datasets there. Do i still need to copy everything to a cloud storage bucket? (I just have them in my instance main directory, instance has ssd drive)\n2. Some features are calculated via python packages, so then to run if from sql query I need first to load that flux row to a numpy, than process, than save. I can load via panda sql, I think, do i need those sql then, will it speed up the process for advanced features? Basic features that can be done via sql query I already have...\n\nthank you for your commands!",
      "votes": null
    },
    {
      "id": "418305",
      "postDate": "11/09/2018 16:12:55",
      "content": "<ol>\n<li>The point of GCS bucket is to transfer files between your PC, BigQuery and maybe a compute engine. BigQuery cannot import or export CSV larger than a certain value directly through the web interface.</li>\n<li>You may still use your local copies of the database for calculating additional features, or use the BigQuery API to pull from the database if the query result is not too large. If you need to do this often, then Internet can indeed be a bottleneck (unless you don't mind running some processing on Compute Engine), which is the main drawback of this method. If you need to process many features locally and are not too memory-constrained, there is another way. You can use Dask to split your test set data into partitions while guranteeing no object_id's will be split in two files.</li>\n</ol>\n\n<p>```\nimport dask.dataframe as dd  </p>\n\n<h1>by setting the index below, we ensure it will not be split between files</h1>\n\n<p>dd.read_csv('test_dataset.csv').set_index('object_id').repartition(150).to_parquet('partitioned_test_set/', write_index=True)\n```</p>\n\n<p>(You can use to_csv as it is easier, but parquet takes less space and reads faster)\nNext time, you can just do:</p>\n\n<p><code>\ntest = dd.read_parquet('partitioned_test_set/') <br>\nfor part in test.partitions: <br>\n    part_content = part.compute() <br>\n    # do your processing per partition <br>\n    # save your results for this partition\n</code></p>\n\n<p>This does partitioning easily and avoids the problem of having to check the next file chunk to see if the last object in this chunk is partly there.\n(Also Dask can do a lot more)</p>",
      "rawMarkdown": "1. The point of GCS bucket is to transfer files between your PC, BigQuery and maybe a compute engine. BigQuery cannot import or export CSV larger than a certain value directly through the web interface.\n2. You may still use your local copies of the database for calculating additional features, or use the BigQuery API to pull from the database if the query result is not too large. If you need to do this often, then Internet can indeed be a bottleneck (unless you don't mind running some processing on Compute Engine), which is the main drawback of this method. If you need to process many features locally and are not too memory-constrained, there is another way. You can use Dask to split your test set data into partitions while guranteeing no object_id's will be split in two files.\n\n```\nimport dask.dataframe as dd  \n#by setting the index below, we ensure it will not be split between files  \ndd.read_csv('test_dataset.csv').set_index('object_id').repartition(150).to_parquet('partitioned_test_set/', write_index=True)\n```\n\n(You can use to_csv as it is easier, but parquet takes less space and reads faster)\nNext time, you can just do:\n\n```\ntest = dd.read_parquet('partitioned_test_set/')  \nfor part in test.partitions:  \n    part_content = part.compute()  \n    # do your processing per partition  \n    # save your results for this partition\n```\n\nThis does partitioning easily and avoids the problem of having to check the next file chunk to see if the last object in this chunk is partly there.\n(Also Dask can do a lot more)",
      "votes": null
    },
    {
      "id": "418345",
      "postDate": "11/09/2018 17:40:37",
      "content": "<p><em>raw speeeeeeeed</em></p>",
      "rawMarkdown": "_raw speeeeeeeed_",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 417065,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/07/2018 17:20:04",
      "content": "<p>It was suggested to split test by object_id, and work with chunks of it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 417231,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/08/2018 00:59:07",
          "content": "<p>I did in chunks, works, but I have a problem of passing object id as I chunk by object ids... I do it later <code>df_test['object_id'] = meta_test['object_id']</code> as they are ascending in both np.unique and meta file, but it's not not optimal (they may change order in general case)... a bit stuck how to make it work properly</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417378,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 06:56:56",
          "content": "<p>You can transform your numpy arrays into pandas data frames then join on object_id.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417497,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/08/2018 11:23:25",
          "content": "<p>this is probably better as npy files are small and csv is huge... but then i save each feature in npy array named as a feature like 'skew', and it's not as easy to manipulate as df['skew']... Can I load numpy file names as headers in that new dataframe? I can do it manually, but it's nice to make just a string list of features to join, load those numpy arrays and join. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417558,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 13:00:15",
          "content": "<p>Here is a code I used in the past.  Data is a pandas dataframe, and features are stored as numpy arrays.</p>\n\n<pre><code>for feat in features:\n    if feat in data.columns:\n        continue\n    print(feat)\n    data[feat] = np.load('../data/' +  feat + '.npy')\n</code></pre>\n\n<p>This assumes that the rows are ordered consistently across feature files.  But why wouldn't they if you construct all of the in a consistent way?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417654,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/08/2018 15:33:29",
          "content": "<p>Thank you ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417067,
      "author_name": "iprapas",
      "author_url": "",
      "post_date": "11/07/2018 17:29:53",
      "content": "<p>I compute multiple test sets with aggregated features only on multiple kernels and then combine them as below:\n<code>\nfrom functools import reduce\nn=500000\nfor i, dfs in enumerate(zip(\n   pd.read_csv( '../input/path-to-csv/test_set1.csv' , iterator=True, chunksize= n),\n   pd.read_csv('../input/path-to-csv/test_set2.csv', iterator=True, chunksize= n),\n   pd.read_csv('../input/path-to-csv/test_set3.csv', iterator=True, chunksize= n))):\n    df = reduce(lambda left,right: pd.merge(left,right,on='object_id'), dfs)\n</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417078,
      "author_name": "authman",
      "author_url": "",
      "post_date": "11/07/2018 17:51:16",
      "content": "<p>My processing occurs in stages / code blocks. Up top I have something like:</p>\n\n<blockquote>\n  <p>DSET_PREP_1 = False\n  DSET_PREP_2 = DSET_PREP_1 OR False\n  DSET_PREP_3 = DSET_PREP_2 OR False\n  DSET_PREP_4 = DSET_PREP_3 OR False\n  DSET_PREP_5 = DSET_PREP_4 OR False</p>\n</blockquote>\n\n<p>Then in the body of the code</p>\n\n<blockquote>\n  <p>if not DSET_PREP1:\n     prep1 = pd.read_feather('./prep1.ftr')\n  else:\n     # .... compute prep1, then save it to prep1.ftr</p>\n  \n  <p>if not DSET_PREP2:\n     prep2 = pd.read_feather('./prep2.ftr')\n  else:\n     # .... compute prep2 using prep1, then save it to prep2.ftr</p>\n  \n  <p>if not DSET_PREP3:\n     prep3 = pd.read_feather('./prep3.ftr')\n  else:\n     # .... compute prep3 using prep2, then save it to prep3.ftr</p>\n</blockquote>\n\n<p>etc.</p>\n\n<p>So if I do something such as add a new stage of features in the pipeline, I'll simply set that stages DSET_PREP to True, and it'll recompute everything downstream.</p>\n\n<p>In other competitions with large amounts of data, sometimes I'll write <code>.ftr's</code> on a per-feature basis and then merge them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 417511,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/08/2018 11:58:35",
          "content": "<p>thank you <a href=\"/authman\">@authman</a> I have not met ftr type, will read about it now. Why did you choose ftr and not npy or something else, by the way, is it better for the size?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418345,
          "author_name": "authman",
          "author_url": "",
          "post_date": "11/09/2018 17:40:37",
          "content": "<p><em>raw speeeeeeeed</em></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417162,
      "author_name": "mithrillion",
      "author_url": "",
      "post_date": "11/07/2018 21:45:31",
      "content": "<p>You can store all your data files in an efficient database (Spark, BigQuery, Athena etc. Needs to be a columnar database. Most of the online services have free tiers), do your basic feature engineering there, add advanced features (the ones you can't do in SQL) as new tables with object_id as key, and whenever you need a dataset, use a simple SQL script to pull all relevant columns together and then export as CSV or directly load into Python/R. This way you can avoid ever loading the full time series dataset into memory when training, and adding / removing features can be done quite fast and separately from your training script.</p>",
      "votes": null,
      "replies": [
        {
          "id": 417677,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/08/2018 16:17:59",
          "content": "<p><code>This way you can avoid ever loading the full time series dataset into memory when training</code> -- the best solution so far, proper professional way, but too advanced for me at this stage. I'll try it in the next competitions when my skills improve</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417699,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "11/08/2018 16:51:11",
          "content": "<p>I'd say if you know a little bit of SQL and don't mind paying less than one dollar per month for some cloud storage, it is probably one of the easier methods. There are a few datasets on Kaggle that are hosted on BigQuery, which is a good place to practice. I can share some of my scripts if you would like to know more.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417758,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/08/2018 18:36:02",
          "content": "<p>It's the test, not train, that gives me pain... At the moment I am fighting with my program bugs trying to load calculated features into classifier by chunks... trying to put it all together and make it work... If I manage this, I'll come back to database. </p>\n\n<p>What do you use (Spark, BigQuery, Hadoop) ? \nI know SELECT and WHERE for SQL :), could be enough, but I need to see what to install and how and where...did not work with databases yet</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417835,
          "author_name": "gregbehm",
          "author_url": "",
          "post_date": "11/08/2018 21:06:44",
          "content": "<p>I just submitted a script for a utility I'm using to read in the training/test files in chunks, separating observations into pandas.DataFrames by object_id. I'm sure there are better ways to store and retrieve the data, but I've settled on this approach for now. I've got to focus my efforts on improved training features!</p>\n\n<p>Have a look here: <a href=\"https://www.kaggle.com/gregbehm/plasticc-data-reader-get-objects-by-id\">https://www.kaggle.com/gregbehm/plasticc-data-reader-get-objects-by-id</a></p>\n\n<p>I'm using a modified version with <code>pandas.read_sql_query</code> to read SQLite database files I've created, but for demo purposes in this kernel I show reading CSV files.</p>\n\n<p>There's demo code in main() at the bottom of the script.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418066,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "11/09/2018 08:32:58",
          "content": "<p>Here is a basic idea of how to do it:</p>\n\n<p>Suppose I have tables called <code>test_meta</code> and <code>test_series</code> that I uploaded to BigQuery (not a quick process, but the easiest way to do it is to spin up a cheapest compute engine, download Kaggle dataset there, use <code>gsutil -m cp</code> to copy everything to a cloud storage bucket then import to BigQuery).</p>\n\n<ol>\n<li>First I want to add some basic features like mean and standard deviation. Instead of directly adding new columns to my dataset, I can run a query that can calculate these values and store it as a view. e.g. my <code>additional stats</code> view:</li>\n</ol>\n\n<p><code>\nSELECT <br>\n    object_id, AVG(flux_err) AS flux_err_mean, <br>\n    STDDEV(flux_err) AS flux_err_std, <br>\n    AVG(detected) AS detected_mean <br>\nFROM `project.astro.test_series` <br>\nGROUP BY object_id\n</code></p>\n\n<ol>\n<li><p>Then I have to append some new feature columns that I calculated from the time series to the final feature set. I store the features in the table <code>test_features</code>.</p></li>\n<li><p>Finally, I want to combine the simple features together with the new features. I can do the following:</p></li>\n</ol>\n\n<p><code>\nSELECT\n  a.*, <br>\n  b.flux_err_mean, <br>\n  b.flux_err_std, <br>\n  b.detected_mean, <br>\n  c.hostgal_photoz, <br>\n  c.ddf <br>\nFROM\n  `project.astro.test_features` a, <br>\n  `project.astro.additional_stats` b, <br>\n  `project.astro.test_meta` c <br>\nWHERE\n  a.object_id = b.object_id <br>\n  AND a.object_id = c.object_id\n</code></p>\n\n<p>You can mix in all the features you want and exclude the ones you don't need.\nThis script should be able to finish under one minute for the training set and at most 10 minutes for test (c.f. maybe close to one hour of Pandas preprocessing in Python). Now you can export the resulting table to Cloud Storage then download from there, or directly pull into your Python environment using the BigQuery API (not recommended because of the size). Because you did most of the full-dataset operations in the cloud, you don't have to deal with split time series files and finding solutions to not split time series of one object into two files. You receive a file (or files) that is ready to be ingested by a training algorithm.</p>\n\n<p>Google BigQuery offers 1TB free processing per month, and the storage cost of the data is tiny, so you normally do not have to worry about billing. It is a great way to perform fast data preprocessing on a less powerful computer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418286,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/09/2018 15:40:02",
          "content": "<p>This script should be able to finish under one minute for the training set and at most 10 minutes for test (c.f. maybe close to one hour of Pandas preprocessing in Python). -- this is impressive!</p>\n\n<ol>\n<li>I have GCP already and copied kaggle datasets there. Do i still need to copy everything to a cloud storage bucket? (I just have them in my instance main directory, instance has ssd drive)</li>\n<li>Some features are calculated via python packages, so then to run if from sql query I need first to load that flux row to a numpy, than process, than save. I can load via panda sql, I think, do i need those sql then, will it speed up the process for advanced features? Basic features that can be done via sql query I already have...</li>\n</ol>\n\n<p>thank you for your commands!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418305,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "11/09/2018 16:12:55",
          "content": "<ol>\n<li>The point of GCS bucket is to transfer files between your PC, BigQuery and maybe a compute engine. BigQuery cannot import or export CSV larger than a certain value directly through the web interface.</li>\n<li>You may still use your local copies of the database for calculating additional features, or use the BigQuery API to pull from the database if the query result is not too large. If you need to do this often, then Internet can indeed be a bottleneck (unless you don't mind running some processing on Compute Engine), which is the main drawback of this method. If you need to process many features locally and are not too memory-constrained, there is another way. You can use Dask to split your test set data into partitions while guranteeing no object_id's will be split in two files.</li>\n</ol>\n\n<p>```\nimport dask.dataframe as dd  </p>\n\n<h1>by setting the index below, we ensure it will not be split between files</h1>\n\n<p>dd.read_csv('test_dataset.csv').set_index('object_id').repartition(150).to_parquet('partitioned_test_set/', write_index=True)\n```</p>\n\n<p>(You can use to_csv as it is easier, but parquet takes less space and reads faster)\nNext time, you can just do:</p>\n\n<p><code>\ntest = dd.read_parquet('partitioned_test_set/') <br>\nfor part in test.partitions: <br>\n    part_content = part.compute() <br>\n    # do your processing per partition <br>\n    # save your results for this partition\n</code></p>\n\n<p>This does partitioning easily and avoids the problem of having to check the next file chunk to see if the last object in this chunk is partly there.\n(Also Dask can do a lot more)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "417063": "I am quite a novice to the IT/programming. So, I do the following\n1. calculate basic features set for test \n2. save them to csv file \n3. load that csv file as a dataframe, add new features  df['new_feature'] = new_feature_function\n4. save newcsv file with added features. Do all the same for train\nThe test features file is huge.\nAlternative way is to save every new feature as a numpy array of [[object_id, feature]], it has no label, and I save numpy array feature in a file named as a feature, not as good for manipulating as pandas header, and this yields to many files. \n\nBoth solutions are bulky and not elegant. I hope there are better ways to handle feature adding for such large sets of objects, that I am not just aware of. Any good elegant solutions on this?",
    "417065": "It was suggested to split test by object_id, and work with chunks of it.",
    "417067": "I compute multiple test sets with aggregated features only on multiple kernels and then combine them as below:\n```\nfrom functools import reduce\nn=500000\nfor i, dfs in enumerate(zip(\n   pd.read_csv( '../input/path-to-csv/test_set1.csv' , iterator=True, chunksize= n),\n   pd.read_csv('../input/path-to-csv/test_set2.csv', iterator=True, chunksize= n),\n   pd.read_csv('../input/path-to-csv/test_set3.csv', iterator=True, chunksize= n))):\n    df = reduce(lambda left,right: pd.merge(left,right,on='object_id'), dfs)\n```",
    "417078": "My processing occurs in stages / code blocks. Up top I have something like:\n\n\n&gt; DSET_PREP_1 = False\n&gt; DSET_PREP_2 = DSET_PREP_1 OR False\n&gt; DSET_PREP_3 = DSET_PREP_2 OR False\n&gt; DSET_PREP_4 = DSET_PREP_3 OR False\n&gt; DSET_PREP_5 = DSET_PREP_4 OR False\n\nThen in the body of the code\n\n&gt; if not DSET_PREP1:\n&gt;    prep1 = pd.read_feather('./prep1.ftr')\n&gt; else:\n&gt;    # .... compute prep1, then save it to prep1.ftr\n&gt; \n&gt; if not DSET_PREP2:\n&gt;    prep2 = pd.read_feather('./prep2.ftr')\n&gt; else:\n&gt;    # .... compute prep2 using prep1, then save it to prep2.ftr\n&gt; \n&gt; if not DSET_PREP3:\n&gt;    prep3 = pd.read_feather('./prep3.ftr')\n&gt; else:\n&gt;    # .... compute prep3 using prep2, then save it to prep3.ftr\n\netc.\n\nSo if I do something such as add a new stage of features in the pipeline, I'll simply set that stages DSET_PREP to True, and it'll recompute everything downstream.\n\nIn other competitions with large amounts of data, sometimes I'll write `.ftr's` on a per-feature basis and then merge them.",
    "417162": "You can store all your data files in an efficient database (Spark, BigQuery, Athena etc. Needs to be a columnar database. Most of the online services have free tiers), do your basic feature engineering there, add advanced features (the ones you can't do in SQL) as new tables with object_id as key, and whenever you need a dataset, use a simple SQL script to pull all relevant columns together and then export as CSV or directly load into Python/R. This way you can avoid ever loading the full time series dataset into memory when training, and adding / removing features can be done quite fast and separately from your training script.",
    "417231": "I did in chunks, works, but I have a problem of passing object id as I chunk by object ids... I do it later ```df_test['object_id'] = meta_test['object_id'] ``` as they are ascending in both np.unique and meta file, but it's not not optimal (they may change order in general case)... a bit stuck how to make it work properly",
    "417378": "You can transform your numpy arrays into pandas data frames then join on object_id.",
    "417497": "this is probably better as npy files are small and csv is huge... but then i save each feature in npy array named as a feature like 'skew', and it's not as easy to manipulate as df['skew']... Can I load numpy file names as headers in that new dataframe? I can do it manually, but it's nice to make just a string list of features to join, load those numpy arrays and join.",
    "417511": "thank you @authman I have not met ftr type, will read about it now. Why did you choose ftr and not npy or something else, by the way, is it better for the size?",
    "417558": "Here is a code I used in the past.  Data is a pandas dataframe, and features are stored as numpy arrays.\n\n    for feat in features:\n        if feat in data.columns:\n            continue\n        print(feat)\n        data[feat] = np.load('../data/' +  feat + '.npy')\n\nThis assumes that the rows are ordered consistently across feature files.  But why wouldn't they if you construct all of the in a consistent way?",
    "417654": "Thank you !",
    "417677": "```This way you can avoid ever loading the full time series dataset into memory when training``` -- the best solution so far, proper professional way, but too advanced for me at this stage. I'll try it in the next competitions when my skills improve",
    "417699": "I'd say if you know a little bit of SQL and don't mind paying less than one dollar per month for some cloud storage, it is probably one of the easier methods. There are a few datasets on Kaggle that are hosted on BigQuery, which is a good place to practice. I can share some of my scripts if you would like to know more.",
    "417758": "It's the test, not train, that gives me pain... At the moment I am fighting with my program bugs trying to load calculated features into classifier by chunks... trying to put it all together and make it work... If I manage this, I'll come back to database. \n\nWhat do you use (Spark, BigQuery, Hadoop) ? \nI know SELECT and WHERE for SQL :), could be enough, but I need to see what to install and how and where...did not work with databases yet",
    "417835": "I just submitted a script for a utility I'm using to read in the training/test files in chunks, separating observations into pandas.DataFrames by object_id. I'm sure there are better ways to store and retrieve the data, but I've settled on this approach for now. I've got to focus my efforts on improved training features!\n\nHave a look here: https://www.kaggle.com/gregbehm/plasticc-data-reader-get-objects-by-id\n\nI'm using a modified version with ```pandas.read_sql_query``` to read SQLite database files I've created, but for demo purposes in this kernel I show reading CSV files.\n\nThere's demo code in main() at the bottom of the script.",
    "418066": "Here is a basic idea of how to do it:\n\nSuppose I have tables called `test_meta` and `test_series` that I uploaded to BigQuery (not a quick process, but the easiest way to do it is to spin up a cheapest compute engine, download Kaggle dataset there, use `gsutil -m cp` to copy everything to a cloud storage bucket then import to BigQuery).\n\n1. First I want to add some basic features like mean and standard deviation. Instead of directly adding new columns to my dataset, I can run a query that can calculate these values and store it as a view. e.g. my `additional stats` view:\n\n```\nSELECT  \n    object_id, AVG(flux_err) AS flux_err_mean,  \n    STDDEV(flux_err) AS flux_err_std,  \n    AVG(detected) AS detected_mean  \nFROM `project.astro.test_series`  \nGROUP BY object_id\n```\n\n2. Then I have to append some new feature columns that I calculated from the time series to the final feature set. I store the features in the table `test_features`.\n\n3. Finally, I want to combine the simple features together with the new features. I can do the following:\n\n```\nSELECT\n  a.*,  \n  b.flux_err_mean,  \n  b.flux_err_std,  \n  b.detected_mean,  \n  c.hostgal_photoz,  \n  c.ddf  \nFROM\n  `project.astro.test_features` a,  \n  `project.astro.additional_stats` b,  \n  `project.astro.test_meta` c   \nWHERE\n  a.object_id = b.object_id  \n  AND a.object_id = c.object_id\n```\n\nYou can mix in all the features you want and exclude the ones you don't need.\nThis script should be able to finish under one minute for the training set and at most 10 minutes for test (c.f. maybe close to one hour of Pandas preprocessing in Python). Now you can export the resulting table to Cloud Storage then download from there, or directly pull into your Python environment using the BigQuery API (not recommended because of the size). Because you did most of the full-dataset operations in the cloud, you don't have to deal with split time series files and finding solutions to not split time series of one object into two files. You receive a file (or files) that is ready to be ingested by a training algorithm.\n\nGoogle BigQuery offers 1TB free processing per month, and the storage cost of the data is tiny, so you normally do not have to worry about billing. It is a great way to perform fast data preprocessing on a less powerful computer.",
    "418286": "This script should be able to finish under one minute for the training set and at most 10 minutes for test (c.f. maybe close to one hour of Pandas preprocessing in Python). -- this is impressive!\n\n1. I have GCP already and copied kaggle datasets there. Do i still need to copy everything to a cloud storage bucket? (I just have them in my instance main directory, instance has ssd drive)\n2. Some features are calculated via python packages, so then to run if from sql query I need first to load that flux row to a numpy, than process, than save. I can load via panda sql, I think, do i need those sql then, will it speed up the process for advanced features? Basic features that can be done via sql query I already have...\n\nthank you for your commands!",
    "418305": "1. The point of GCS bucket is to transfer files between your PC, BigQuery and maybe a compute engine. BigQuery cannot import or export CSV larger than a certain value directly through the web interface.\n2. You may still use your local copies of the database for calculating additional features, or use the BigQuery API to pull from the database if the query result is not too large. If you need to do this often, then Internet can indeed be a bottleneck (unless you don't mind running some processing on Compute Engine), which is the main drawback of this method. If you need to process many features locally and are not too memory-constrained, there is another way. You can use Dask to split your test set data into partitions while guranteeing no object_id's will be split in two files.\n\n```\nimport dask.dataframe as dd  \n#by setting the index below, we ensure it will not be split between files  \ndd.read_csv('test_dataset.csv').set_index('object_id').repartition(150).to_parquet('partitioned_test_set/', write_index=True)\n```\n\n(You can use to_csv as it is easier, but parquet takes less space and reads faster)\nNext time, you can just do:\n\n```\ntest = dd.read_parquet('partitioned_test_set/')  \nfor part in test.partitions:  \n    part_content = part.compute()  \n    # do your processing per partition  \n    # save your results for this partition\n```\n\nThis does partitioning easily and avoids the problem of having to check the next file chunk to see if the last object in this chunk is partly there.\n(Also Dask can do a lot more)",
    "418345": "_raw speeeeeeeed_"
  },
  "source": "meta"
}