{
  "id": 193682,
  "title": "Reading and Processing data take too much time!",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/193682",
  "author_name": "mooneral",
  "post_date": "2020-10-28T08:58:47.904000",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi Kagglers,<br>\nI've been working on this dataset for a couple of days and I have written functions to read training data and do some transformation like fft/cwt then PCA but it takes too much time. Currently, it takes about 10 hours for this process to complete (approx 8 sec/iter for a single training file).<br>\nIf you are working on this data set what are your approach/suggestions to make this process much faster? Or Am I taking the wrong avenue?</p>",
  "messages": [
    {
      "id": 1062857,
      "postDate": "2020-10-28T08:58:47.903Z",
      "content": "<p>Hi Kagglers,<br>\nI've been working on this dataset for a couple of days and I have written functions to read training data and do some transformation like fft/cwt then PCA but it takes too much time. Currently, it takes about 10 hours for this process to complete (approx 8 sec/iter for a single training file).<br>\nIf you are working on this data set what are your approach/suggestions to make this process much faster? Or Am I taking the wrong avenue?</p>",
      "rawMarkdown": "Hi Kagglers,\nI've been working on this dataset for a couple of days and I have written functions to read training data and do some transformation like fft/cwt then PCA but it takes too much time. Currently, it takes about 10 hours for this process to complete (approx 8 sec/iter for a single training file).\nIf you are working on this data set what are your approach/suggestions to make this process much faster? Or Am I taking the wrong avenue?",
      "votes": 2
    },
    {
      "id": 1074875,
      "postDate": "2020-11-11T07:13:08.657Z",
      "content": "<p>import dataset into database and run the feature enginining at database. It could speed up the data preparing process, especially the featuring engineering repetitively. </p>\n<p>In database, the feature engineering could be done in many mode, many time and save the results in structured tempoaray place(materialized view)</p>\n<p>My approach:<br>\nStage A( ETL mode by manual or airflow)</p>\n<ol>\n<li>Import data into data base (around 2 hour)Database: Postgres 12</li>\n<li>install database development Plugin: Py/Python(install plpython3u extension). It can let data do science computation at database with ETL data to jupyternootbook. </li>\n<li>create function to working necessary feature engineering work.</li>\n<li>create view of x_train, x_test as view<br>\nStage B( machine learning part)</li>\n<li>using pandas to the prepared data from database.</li>\n<li>feed into lightgbm or optuna for train and hyperparameter optimization</li>\n<li>submit it if you are OK with the local CV results.</li>\n</ol>\n<p>Considering the Stage A need plenty of time and repetitive work, I think it deserved to let it run at database.</p>\n<p>For example, create an function to process the sensors parameters like mean, std, skew, kurt, missing_count can be get in two steps lik following mode:</p>\n<ol>\n<li>Create an function in database(PG12): </li>\n</ol>\n<p>CREATE OR REPLACE FUNCTION public.py_thdf_stat_min_max(hour_id integer, shift integer)<br>\n RETURNS real[]<br>\n LANGUAGE plpython3u<br>\n COST 500<br>\nAS $function$<br>\n    import pandas as pd</p>\n<pre><code>cols = ['sensor_1','sensor_2','sensor_3', 'sensor_4','sensor_5',\n         'sensor_6','sensor_7','sensor_8','sensor_9','sensor_10'\n        ]\n\nagg_cols = []\nresult_sql = plpy.execute(f\"\"\"\n                          select tm.* \n                          from \n                            t_measures tm\n                          inner join \n                            v_train_per_hour vtph \n                          on \n                            tm.segment_id = vtph.segment_id \n                          and \n                            vtph.hour_to_eruption={hour_id} \n                          order by \n                            ts_id\"\"\")\n\ndf = pd.DataFrame.from_records(result_sql)[cols]\n\nif shift == 0:\n    agg_cols = ['size','count','mean','median','std','skew','kurt','min','max']\nelse:\n    agg_cols = ['mean','median','std','skew','kurt','min','max']\n    df = df.fillna(0).diff(periods=shift)\n\ndf_stat = df.agg(agg_cols)\ndf_stat['hour_id'] = hour_id\ndf_stat['shift'] = shift\n\nreturn df_stat.values.tolist()\n</code></pre>\n<p>$function$<br>\n;</p>\n<ol>\n<li><p>create an materialized view to get data. <br>\nNote: In order to speed up the process, the data has been stored in array data type. <br>\nCREATE MATERIALIZED VIEW public.mv_train_tsdf_stat<br>\nTABLESPACE pg_default<br>\nAS SELECT t.segment_id,<br>\nt.time_to_eruption,<br>\nt.time_to_eruption / 100 / 3600 AS hour_to_eruption,<br>\npy_tdf_stat(t.segment_id, 0) AS df_stat,<br>\npy_tdf_stat(t.segment_id, 1) AS df_stat_diff_1,<br>\npy_tdf_stat(t.segment_id, 100) AS df_stat_diff_100<br>\nFROM train t<br>\nWITH DATA;</p></li>\n<li><p>create view to generate the x_train for next step;<br>\nCREATE MATERIALIZED VIEW public.x_train<br>\nTABLESPACE pg_default<br>\nAS WITH seg2hour AS (<br>\n     SELECT vsl.segment_id,<br>\n        vsl.time_to_eruption,<br>\n        vhl.hour_to_eruption,<br>\n        vsl.nlag,<br>\n        unnest(((((((((vsl.sensor1_stat[2:] || vsl.sensor2_stat[2:]) || vsl.sensor3_stat[2:]) || vsl.sensor4_stat[2:]) || vsl.sensor5_stat[2:]) || vsl.sensor6_stat[2:]) || vsl.sensor7_stat[2:]) || vsl.sensor8_stat[2:]) || vsl.sensor9_stat[2:]) || vsl.sensor10_stat[2:]) AS seg_stat,<br>\n        unnest(((((((((vhl.sensor1_stat[2:] || vhl.sensor2_stat[2:]) || vhl.sensor3_stat[2:]) || vhl.sensor4_stat[2:]) || vhl.sensor5_stat[2:]) || vhl.sensor6_stat[2:]) || vhl.sensor7_stat[2:]) || vhl.sensor8_stat[2:]) || vhl.sensor9_stat[2:]) || vhl.sensor10_stat[2:]) AS hour_stat<br>\n       FROM v_train_stat_per_sensor_lags vsl<br>\n         JOIN v_train_stat_per_hour_lags vhl ON vhl.nlag = vsl.nlag<br>\n    )<br>\nSELECT seg2hour.segment_id,<br>\nseg2hour.time_to_eruption,<br>\nseg2hour.hour_to_eruption,<br>\ncount(*) AS count,<br>\ncorr(COALESCE(NULLIF(seg2hour.seg_stat, 'NaN'::real), 0::real)::double precision, COALESCE(NULLIF(seg2hour.hour_stat, 'NaN'::real), 0::real)::double precision) AS corr<br>\nFROM seg2hour<br>\nGROUP BY seg2hour.segment_id, seg2hour.time_to_eruption, seg2hour.hour_to_eruption<br>\nORDER BY seg2hour.segment_id, seg2hour.hour_to_eruption<br>\nWITH DATA;</p></li>\n</ol>\n<p>To sumup,  one step to import dataset in to postgresql database, and do the feature engineering repetitively work in datbase. It can significantly save featuring engineering time.  And eay, because it is just SQL language. </p>\n<p>Only let small piece of data runing into training/predicting. </p>\n<p>That all of my sharing.</p>\n<p>If you find it useful, pls vote it up.</p>\n<p>WangYong</p>",
      "rawMarkdown": "import dataset into database and run the feature enginining at database. It could speed up the data preparing process, especially the featuring engineering repetitively. \n\nIn database, the feature engineering could be done in many mode, many time and save the results in structured tempoaray place(materialized view)\n\nMy approach:\nStage A( ETL mode by manual or airflow)\n1. Import data into data base (around 2 hour)Database: Postgres 12\n2. install database development Plugin: Py/Python(install plpython3u extension). It can let data do science computation at database with ETL data to jupyternootbook. \n3. create function to working necessary feature engineering work.\n4. create view of x_train, x_test as view\nStage B( machine learning part)\n5. using pandas to the prepared data from database.\n6. feed into lightgbm or optuna for train and hyperparameter optimization\n7. submit it if you are OK with the local CV results.\n\nConsidering the Stage A need plenty of time and repetitive work, I think it deserved to let it run at database.\n\nFor example, create an function to process the sensors parameters like mean, std, skew, kurt, missing_count can be get in two steps lik following mode:\n\n1. Create an function in database(PG12): \n\nCREATE OR REPLACE FUNCTION public.py_thdf_stat_min_max(hour_id integer, shift integer)\n RETURNS real[]\n LANGUAGE plpython3u\n COST 500\nAS $function$\n    import pandas as pd\n    \n    cols = ['sensor_1','sensor_2','sensor_3', 'sensor_4','sensor_5',\n    \t\t 'sensor_6','sensor_7','sensor_8','sensor_9','sensor_10'\n    \t\t]\n        \n    agg_cols = []\n    result_sql = plpy.execute(f\"\"\"\n\t\t\t\t\t\t\t  select tm.* \n\t\t\t\t\t\t\t  from \n\t\t  \t\t\t\t\t    t_measures tm\n\t\t\t\t\t\t\t  inner join \n\t\t\t\t\t\t\t    v_train_per_hour vtph \n\t\t\t\t\t\t\t  on \n\t\t\t\t\t\t\t    tm.segment_id = vtph.segment_id \n\t\t\t\t\t\t\t  and \n\t\t\t\t\t\t\t    vtph.hour_to_eruption={hour_id} \n\t\t\t\t\t\t\t  order by \n\t\t\t\t\t\t\t    ts_id\"\"\")\n    \n    df = pd.DataFrame.from_records(result_sql)[cols]\n    \n    if shift == 0:\n    \tagg_cols = ['size','count','mean','median','std','skew','kurt','min','max']\n    else:\n        agg_cols = ['mean','median','std','skew','kurt','min','max']\n        df = df.fillna(0).diff(periods=shift)\n    \n    df_stat = df.agg(agg_cols)\n    df_stat['hour_id'] = hour_id\n    df_stat['shift'] = shift\n    \n    return df_stat.values.tolist()\n\n$function$\n;\n\n2. create an materialized view to get data. \nNote: In order to speed up the process, the data has been stored in array data type. \nCREATE MATERIALIZED VIEW public.mv_train_tsdf_stat\nTABLESPACE pg_default\nAS SELECT t.segment_id,\n    t.time_to_eruption,\n    t.time_to_eruption / 100 / 3600 AS hour_to_eruption,\n    py_tdf_stat(t.segment_id, 0) AS df_stat,\n    py_tdf_stat(t.segment_id, 1) AS df_stat_diff_1,\n    py_tdf_stat(t.segment_id, 100) AS df_stat_diff_100\n   FROM train t\nWITH DATA;\n\n3. create view to generate the x_train for next step;\nCREATE MATERIALIZED VIEW public.x_train\nTABLESPACE pg_default\nAS WITH seg2hour AS (\n         SELECT vsl.segment_id,\n            vsl.time_to_eruption,\n            vhl.hour_to_eruption,\n            vsl.nlag,\n            unnest(((((((((vsl.sensor1_stat[2:] || vsl.sensor2_stat[2:]) || vsl.sensor3_stat[2:]) || vsl.sensor4_stat[2:]) || vsl.sensor5_stat[2:]) || vsl.sensor6_stat[2:]) || vsl.sensor7_stat[2:]) || vsl.sensor8_stat[2:]) || vsl.sensor9_stat[2:]) || vsl.sensor10_stat[2:]) AS seg_stat,\n            unnest(((((((((vhl.sensor1_stat[2:] || vhl.sensor2_stat[2:]) || vhl.sensor3_stat[2:]) || vhl.sensor4_stat[2:]) || vhl.sensor5_stat[2:]) || vhl.sensor6_stat[2:]) || vhl.sensor7_stat[2:]) || vhl.sensor8_stat[2:]) || vhl.sensor9_stat[2:]) || vhl.sensor10_stat[2:]) AS hour_stat\n           FROM v_train_stat_per_sensor_lags vsl\n             JOIN v_train_stat_per_hour_lags vhl ON vhl.nlag = vsl.nlag\n        )\n SELECT seg2hour.segment_id,\n    seg2hour.time_to_eruption,\n    seg2hour.hour_to_eruption,\n    count(*) AS count,\n    corr(COALESCE(NULLIF(seg2hour.seg_stat, 'NaN'::real), 0::real)::double precision, COALESCE(NULLIF(seg2hour.hour_stat, 'NaN'::real), 0::real)::double precision) AS corr\n   FROM seg2hour\n  GROUP BY seg2hour.segment_id, seg2hour.time_to_eruption, seg2hour.hour_to_eruption\n  ORDER BY seg2hour.segment_id, seg2hour.hour_to_eruption\nWITH DATA;\n\nTo sumup,  one step to import dataset in to postgresql database, and do the feature engineering repetitively work in datbase. It can significantly save featuring engineering time.  And eay, because it is just SQL language. \n\nOnly let small piece of data runing into training/predicting. \n\nThat all of my sharing.\n\nIf you find it useful, pls vote it up.\n\nWangYong",
      "votes": 1
    },
    {
      "id": 1064096,
      "postDate": "2020-10-29T17:33:44.300Z",
      "content": "<p>Hi, i will  just say, accordingly to my intuition regarding the spectral analysis, to work only on one sensor at the beginning, especially if you're doing fft/cwt computations. There are plenty of notebook in this competition that can help you to choose the sensor with less NA values and more 'events'. The postulate i have, after reading some articles on the seismic activity regarding volcano, is to chase for a phenomena visible by spectral analysis, and that this phenomena is available on all sensors. May be i'm wrong about this hypothesis, to be tested with correlation matrix for instance, but  at least it simplify your job at the beginning of your spectral study, reduce the time need to processed the data and help you to improve your hypothesis.. You can also split the train set to generate a new train(80 % ) &amp;test set(20 % ) for having a rather good idea of the model capacity to fill the objective. Once you get some interesting results, apply the model to others sensors.</p>\n<p>Just my 2 cents feeling.</p>",
      "rawMarkdown": "Hi, i will  just say, accordingly to my intuition regarding the spectral analysis, to work only on one sensor at the beginning, especially if you're doing fft/cwt computations. There are plenty of notebook in this competition that can help you to choose the sensor with less NA values and more 'events'. The postulate i have, after reading some articles on the seismic activity regarding volcano, is to chase for a phenomena visible by spectral analysis, and that this phenomena is available on all sensors. May be i'm wrong about this hypothesis, to be tested with correlation matrix for instance, but  at least it simplify your job at the beginning of your spectral study, reduce the time need to processed the data and help you to improve your hypothesis.. You can also split the train set to generate a new train(80 % ) &test set(20 % ) for having a rather good idea of the model capacity to fill the objective. Once you get some interesting results, apply the model to others sensors.\n\nJust my 2 cents feeling.",
      "votes": 1
    },
    {
      "id": 1063091,
      "postDate": "2020-10-28T13:51:50.533Z",
      "content": "<p>What I did is having a separate pre-processing stage and saving the result in one or several csv files. Then I upload the csv files in a database I created and run training and inference using them. This strategy is allowed by the competition rules, if I am not mistaken.</p>",
      "rawMarkdown": "What I did is having a separate pre-processing stage and saving the result in one or several csv files. Then I upload the csv files in a database I created and run training and inference using them. This strategy is allowed by the competition rules, if I am not mistaken.",
      "votes": 1
    },
    {
      "id": 1075992,
      "postDate": "2020-11-12T06:57:23.663Z",
      "content": "<p>You can use my dataset. It contains all the challenge data in a way that it is fast to read and process.<br>\n<a href=\"https://www.kaggle.com/itamargr/volcaniceruptionsdata\" target=\"_blank\">https://www.kaggle.com/itamargr/volcaniceruptionsdata</a></p>\n<p>Itamar</p>",
      "rawMarkdown": "You can use my dataset. It contains all the challenge data in a way that it is fast to read and process.\nhttps://www.kaggle.com/itamargr/volcaniceruptionsdata\n\nItamar",
      "votes": 2
    },
    {
      "id": 1064693,
      "postDate": "2020-10-30T12:26:38.343Z",
      "content": "<p>In my processing, the following two points have made the speed considerably faster during preprocessing.<br>\nIt may overlap with the contents described by other people, but I will describe it just in case.</p>\n<ul>\n<li>Convert the CSV file to a parquet file in advance, and use the parquet file during preprocessing. -&gt; (1)</li>\n<li>The new feature will be kept as a list and finally combined using the concat function. -&gt; (2)</li>\n</ul>\n<p>(Excuse me my English is  very poor.)</p>\n<p>(1)</p>\n<pre><code>for idx in tqdm.tqdm_notebook(range(len(train_df))):\n    segment_id = train_df.iloc[idx,0]\n    detail_df = pd.read_csv(f'{TRAIN_FILE_PATH}/{segment_id}.csv')\n    detail_df.to_parquet(f'{segment_id}.parquet')\n</code></pre>\n<p>(2)</p>\n<pre><code>fe_types = ['min','max','range','mean',]\n\ndef preprocessing(segments, df_len, path, interpolate=False):\n    maps = {}\n\n    for feat in sensors:\n        for fe_type in fe_types:\n            maps[f'{feat}_{fe_type}'] = []\n\n    for idx in tqdm.tqdm_notebook(range(df_len)):\n        segment_id = segments[idx]\n        detail_df = pd.read_parquet(f'{path}/{segment_id}.parquet').astype(np.float32)\n        sensor_data_len = len(detail_df)\n\n        for feat in sensors:\n            maps[feat+'_min'].append(detail_df[feat].min())\n            maps[feat+'_max'].append(detail_df[feat].max())\n            maps[feat+'_range'].append(maps[feat+'_max'][idx] - maps[feat+'_min'][idx])\n            maps[feat+'_mean'].append(detail_df[feat].mean())\n    return maps\n\nmaps = preprocessing(train_df.segment_id.tolist(), len(train_df), TRAIN_PARQUET_PATH)\nfor feat in maps.keys():\n    train_df = pd.concat([train_df, pd.DataFrame(maps[feat], columns=[feat])], axis=1)    \n</code></pre>",
      "rawMarkdown": "In my processing, the following two points have made the speed considerably faster during preprocessing.\nIt may overlap with the contents described by other people, but I will describe it just in case.\n- Convert the CSV file to a parquet file in advance, and use the parquet file during preprocessing. -> (1)\n- The new feature will be kept as a list and finally combined using the concat function. -> (2)\n\n(Excuse me my English is  very poor.)\n\n(1)\n```Python\nfor idx in tqdm.tqdm_notebook(range(len(train_df))):\n    segment_id = train_df.iloc[idx,0]\n    detail_df = pd.read_csv(f'{TRAIN_FILE_PATH}/{segment_id}.csv')\n    detail_df.to_parquet(f'{segment_id}.parquet')\n```\n\n(2)\n```Python\nfe_types = ['min','max','range','mean',]\n\ndef preprocessing(segments, df_len, path, interpolate=False):\n    maps = {}\n\n    for feat in sensors:\n        for fe_type in fe_types:\n            maps[f'{feat}_{fe_type}'] = []\n    \n    for idx in tqdm.tqdm_notebook(range(df_len)):\n        segment_id = segments[idx]\n        detail_df = pd.read_parquet(f'{path}/{segment_id}.parquet').astype(np.float32)\n        sensor_data_len = len(detail_df)\n        \n        for feat in sensors:\n            maps[feat+'_min'].append(detail_df[feat].min())\n            maps[feat+'_max'].append(detail_df[feat].max())\n            maps[feat+'_range'].append(maps[feat+'_max'][idx] - maps[feat+'_min'][idx])\n            maps[feat+'_mean'].append(detail_df[feat].mean())\n    return maps\n\nmaps = preprocessing(train_df.segment_id.tolist(), len(train_df), TRAIN_PARQUET_PATH)\nfor feat in maps.keys():\n    train_df = pd.concat([train_df, pd.DataFrame(maps[feat], columns=[feat])], axis=1)    \n```",
      "votes": 2
    },
    {
      "id": 1063981,
      "postDate": "2020-10-29T15:04:17.610Z",
      "content": "<p>The code below use pure Python to consolidate csv files into one (20gb)</p>\n<p></p><pre><br>\nimport glob<br>\nimport csv<br>\ndef consolidate_train():<br>\n    files = list(glob.glob('train*.csv'))<br>\n    with open('consolidated-train.csv', 'w') as csvoutput:<br>\n        for i in files:<br>\n            with open(i,'r') as csvinput:<br>\n                writer = csv.writer(csvoutput, lineterminator='\\n')<br>\n                reader = csv.reader(csvinput)<br>\n                all = []                <br>\n                row = next(reader)<br>\n                row.append('segment_id')<br>\n                row.append('period')<br>\n                all.append(row)<br>\n                c = 0                <br>\n                for row in reader:<br>\n                    row.append(str(i[6:-4]))    <br>\n                    row.append(c)<br>\n                    all.append(row)<br>\n                    c = c+1<br>\n                print(str(i[6:-4]))<br>\n                writer.writerows(all)</pre><p></p>\n<p>consolidate_train()<br>\n</p>\n<p></p><pre><br>\ndef consolidate_train_drop_headers():<br>\n    with open(\"consolidated-train.csv\",\"r\") as inputfile, open(\"output.csv\",\"w\",newline=\"\") as outputfile:<br>\n       csv_in = csv.reader(inputfile)<br>\n       csv_out = csv.writer(outputfile)<br>\n       title = next(csv_in)<br>\n       csv_out.writerow(title)<br>\n       for row in csv_in:<br>\n            if row != title:<br>\n                 csv_out.writerow(row)</pre><p></p>\n<p>consolidate_train_drop_headers()<br>\n</p>",
      "rawMarkdown": "The code below use pure Python to consolidate csv files into one (20gb)\n\n<pre>\nimport glob\nimport csv\ndef consolidate_train():\n    files = list(glob.glob('train\\*.csv'))\n    with open('consolidated-train.csv', 'w') as csvoutput:\n        for i in files:\n            with open(i,'r') as csvinput:\n                writer = csv.writer(csvoutput, lineterminator='\\n')\n                reader = csv.reader(csvinput)\n                all = []                \n                row = next(reader)\n                row.append('segment_id')\n                row.append('period')\n                all.append(row)\n                c = 0                \n                for row in reader:\n                    row.append(str(i[6:-4]))    \n                    row.append(c)\n                    all.append(row)\n                    c = c+1\n                print(str(i[6:-4]))\n                writer.writerows(all)\n\nconsolidate_train()\n<\\pre>\n\n\n\n<pre>\ndef consolidate_train_drop_headers():\n    with open(\"consolidated-train.csv\",\"r\") as inputfile, open(\"output.csv\",\"w\",newline=\"\") as outputfile:\n       csv_in = csv.reader(inputfile)\n       csv_out = csv.writer(outputfile)\n       title = next(csv_in)\n       csv_out.writerow(title)\n       for row in csv_in:\n            if row != title:\n                 csv_out.writerow(row)\n\nconsolidate_train_drop_headers()\n<\\pre>",
      "votes": 1
    },
    {
      "id": 1063312,
      "postDate": "2020-10-28T18:17:47.687Z",
      "content": "<p><a href=\"https://www.kaggle.com/javagarm\" target=\"_blank\">@javagarm</a> the problem is preparing the training set, as other libraries than <code>deep learning</code> do not support <code>gpu</code>. </p>",
      "rawMarkdown": "@javagarm the problem is preparing the training set, as other libraries than `deep learning` do not support `gpu`. "
    },
    {
      "id": 1063306,
      "postDate": "2020-10-28T18:14:03.293Z",
      "content": "<p><a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a> Indeed I took the same approach but somehow thought there may be a better way to speed up the preprocessing. I've tried <code>dask</code> and <code>numba</code> but no help. </p>",
      "rawMarkdown": "@catadanna Indeed I took the same approach but somehow thought there may be a better way to speed up the preprocessing. I've tried `dask` and `numba` but no help. ",
      "replies": [
        {
          "id": 1063310,
          "postDate": "2020-10-28T18:15:59.753Z",
          "content": "<p>Good luck! I keep my option, though. I aleready have my database with csv files to train, and I add more pre-processed files to it, over and over.</p>",
          "rawMarkdown": "Good luck! I keep my option, though. I aleready have my database with csv files to train, and I add more pre-processed files to it, over and over."
        }
      ]
    },
    {
      "id": 1062877,
      "postDate": "2020-10-28T09:24:09.570Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1074875,
      "author_name": "WangYong",
      "author_url": "",
      "post_date": "2020-11-11T07:13:08.657000",
      "content": "<p>import dataset into database and run the feature enginining at database. It could speed up the data preparing process, especially the featuring engineering repetitively. </p>\n<p>In database, the feature engineering could be done in many mode, many time and save the results in structured tempoaray place(materialized view)</p>\n<p>My approach:<br>\nStage A( ETL mode by manual or airflow)</p>\n<ol>\n<li>Import data into data base (around 2 hour)Database: Postgres 12</li>\n<li>install database development Plugin: Py/Python(install plpython3u extension). It can let data do science computation at database with ETL data to jupyternootbook. </li>\n<li>create function to working necessary feature engineering work.</li>\n<li>create view of x_train, x_test as view<br>\nStage B( machine learning part)</li>\n<li>using pandas to the prepared data from database.</li>\n<li>feed into lightgbm or optuna for train and hyperparameter optimization</li>\n<li>submit it if you are OK with the local CV results.</li>\n</ol>\n<p>Considering the Stage A need plenty of time and repetitive work, I think it deserved to let it run at database.</p>\n<p>For example, create an function to process the sensors parameters like mean, std, skew, kurt, missing_count can be get in two steps lik following mode:</p>\n<ol>\n<li>Create an function in database(PG12): </li>\n</ol>\n<p>CREATE OR REPLACE FUNCTION public.py_thdf_stat_min_max(hour_id integer, shift integer)<br>\n RETURNS real[]<br>\n LANGUAGE plpython3u<br>\n COST 500<br>\nAS $function$<br>\n    import pandas as pd</p>\n<pre><code>cols = ['sensor_1','sensor_2','sensor_3', 'sensor_4','sensor_5',\n         'sensor_6','sensor_7','sensor_8','sensor_9','sensor_10'\n        ]\n\nagg_cols = []\nresult_sql = plpy.execute(f\"\"\"\n                          select tm.* \n                          from \n                            t_measures tm\n                          inner join \n                            v_train_per_hour vtph \n                          on \n                            tm.segment_id = vtph.segment_id \n                          and \n                            vtph.hour_to_eruption={hour_id} \n                          order by \n                            ts_id\"\"\")\n\ndf = pd.DataFrame.from_records(result_sql)[cols]\n\nif shift == 0:\n    agg_cols = ['size','count','mean','median','std','skew','kurt','min','max']\nelse:\n    agg_cols = ['mean','median','std','skew','kurt','min','max']\n    df = df.fillna(0).diff(periods=shift)\n\ndf_stat = df.agg(agg_cols)\ndf_stat['hour_id'] = hour_id\ndf_stat['shift'] = shift\n\nreturn df_stat.values.tolist()\n</code></pre>\n<p>$function$<br>\n;</p>\n<ol>\n<li><p>create an materialized view to get data. <br>\nNote: In order to speed up the process, the data has been stored in array data type. <br>\nCREATE MATERIALIZED VIEW public.mv_train_tsdf_stat<br>\nTABLESPACE pg_default<br>\nAS SELECT t.segment_id,<br>\nt.time_to_eruption,<br>\nt.time_to_eruption / 100 / 3600 AS hour_to_eruption,<br>\npy_tdf_stat(t.segment_id, 0) AS df_stat,<br>\npy_tdf_stat(t.segment_id, 1) AS df_stat_diff_1,<br>\npy_tdf_stat(t.segment_id, 100) AS df_stat_diff_100<br>\nFROM train t<br>\nWITH DATA;</p></li>\n<li><p>create view to generate the x_train for next step;<br>\nCREATE MATERIALIZED VIEW public.x_train<br>\nTABLESPACE pg_default<br>\nAS WITH seg2hour AS (<br>\n     SELECT vsl.segment_id,<br>\n        vsl.time_to_eruption,<br>\n        vhl.hour_to_eruption,<br>\n        vsl.nlag,<br>\n        unnest(((((((((vsl.sensor1_stat[2:] || vsl.sensor2_stat[2:]) || vsl.sensor3_stat[2:]) || vsl.sensor4_stat[2:]) || vsl.sensor5_stat[2:]) || vsl.sensor6_stat[2:]) || vsl.sensor7_stat[2:]) || vsl.sensor8_stat[2:]) || vsl.sensor9_stat[2:]) || vsl.sensor10_stat[2:]) AS seg_stat,<br>\n        unnest(((((((((vhl.sensor1_stat[2:] || vhl.sensor2_stat[2:]) || vhl.sensor3_stat[2:]) || vhl.sensor4_stat[2:]) || vhl.sensor5_stat[2:]) || vhl.sensor6_stat[2:]) || vhl.sensor7_stat[2:]) || vhl.sensor8_stat[2:]) || vhl.sensor9_stat[2:]) || vhl.sensor10_stat[2:]) AS hour_stat<br>\n       FROM v_train_stat_per_sensor_lags vsl<br>\n         JOIN v_train_stat_per_hour_lags vhl ON vhl.nlag = vsl.nlag<br>\n    )<br>\nSELECT seg2hour.segment_id,<br>\nseg2hour.time_to_eruption,<br>\nseg2hour.hour_to_eruption,<br>\ncount(*) AS count,<br>\ncorr(COALESCE(NULLIF(seg2hour.seg_stat, 'NaN'::real), 0::real)::double precision, COALESCE(NULLIF(seg2hour.hour_stat, 'NaN'::real), 0::real)::double precision) AS corr<br>\nFROM seg2hour<br>\nGROUP BY seg2hour.segment_id, seg2hour.time_to_eruption, seg2hour.hour_to_eruption<br>\nORDER BY seg2hour.segment_id, seg2hour.hour_to_eruption<br>\nWITH DATA;</p></li>\n</ol>\n<p>To sumup,  one step to import dataset in to postgresql database, and do the feature engineering repetitively work in datbase. It can significantly save featuring engineering time.  And eay, because it is just SQL language. </p>\n<p>Only let small piece of data runing into training/predicting. </p>\n<p>That all of my sharing.</p>\n<p>If you find it useful, pls vote it up.</p>\n<p>WangYong</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1064096,
      "author_name": "JcLambert",
      "author_url": "",
      "post_date": "2020-10-29T17:33:44.300000",
      "content": "<p>Hi, i will  just say, accordingly to my intuition regarding the spectral analysis, to work only on one sensor at the beginning, especially if you're doing fft/cwt computations. There are plenty of notebook in this competition that can help you to choose the sensor with less NA values and more 'events'. The postulate i have, after reading some articles on the seismic activity regarding volcano, is to chase for a phenomena visible by spectral analysis, and that this phenomena is available on all sensors. May be i'm wrong about this hypothesis, to be tested with correlation matrix for instance, but  at least it simplify your job at the beginning of your spectral study, reduce the time need to processed the data and help you to improve your hypothesis.. You can also split the train set to generate a new train(80 % ) &amp;test set(20 % ) for having a rather good idea of the model capacity to fill the objective. Once you get some interesting results, apply the model to others sensors.</p>\n<p>Just my 2 cents feeling.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1063091,
      "author_name": "Catadanna",
      "author_url": "",
      "post_date": "2020-10-28T13:51:50.533000",
      "content": "<p>What I did is having a separate pre-processing stage and saving the result in one or several csv files. Then I upload the csv files in a database I created and run training and inference using them. This strategy is allowed by the competition rules, if I am not mistaken.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1075992,
      "author_name": "Itamar Gilad",
      "author_url": "",
      "post_date": "2020-11-12T06:57:23.663000",
      "content": "<p>You can use my dataset. It contains all the challenge data in a way that it is fast to read and process.<br>\n<a href=\"https://www.kaggle.com/itamargr/volcaniceruptionsdata\" target=\"_blank\">https://www.kaggle.com/itamargr/volcaniceruptionsdata</a></p>\n<p>Itamar</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1064693,
      "author_name": "cocoa",
      "author_url": "",
      "post_date": "2020-10-30T12:26:38.343000",
      "content": "<p>In my processing, the following two points have made the speed considerably faster during preprocessing.<br>\nIt may overlap with the contents described by other people, but I will describe it just in case.</p>\n<ul>\n<li>Convert the CSV file to a parquet file in advance, and use the parquet file during preprocessing. -&gt; (1)</li>\n<li>The new feature will be kept as a list and finally combined using the concat function. -&gt; (2)</li>\n</ul>\n<p>(Excuse me my English is  very poor.)</p>\n<p>(1)</p>\n<pre><code>for idx in tqdm.tqdm_notebook(range(len(train_df))):\n    segment_id = train_df.iloc[idx,0]\n    detail_df = pd.read_csv(f'{TRAIN_FILE_PATH}/{segment_id}.csv')\n    detail_df.to_parquet(f'{segment_id}.parquet')\n</code></pre>\n<p>(2)</p>\n<pre><code>fe_types = ['min','max','range','mean',]\n\ndef preprocessing(segments, df_len, path, interpolate=False):\n    maps = {}\n\n    for feat in sensors:\n        for fe_type in fe_types:\n            maps[f'{feat}_{fe_type}'] = []\n\n    for idx in tqdm.tqdm_notebook(range(df_len)):\n        segment_id = segments[idx]\n        detail_df = pd.read_parquet(f'{path}/{segment_id}.parquet').astype(np.float32)\n        sensor_data_len = len(detail_df)\n\n        for feat in sensors:\n            maps[feat+'_min'].append(detail_df[feat].min())\n            maps[feat+'_max'].append(detail_df[feat].max())\n            maps[feat+'_range'].append(maps[feat+'_max'][idx] - maps[feat+'_min'][idx])\n            maps[feat+'_mean'].append(detail_df[feat].mean())\n    return maps\n\nmaps = preprocessing(train_df.segment_id.tolist(), len(train_df), TRAIN_PARQUET_PATH)\nfor feat in maps.keys():\n    train_df = pd.concat([train_df, pd.DataFrame(maps[feat], columns=[feat])], axis=1)    \n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1063981,
      "author_name": "JP",
      "author_url": "",
      "post_date": "2020-10-29T15:04:17.610000",
      "content": "<p>The code below use pure Python to consolidate csv files into one (20gb)</p>\n<p></p><pre><br>\nimport glob<br>\nimport csv<br>\ndef consolidate_train():<br>\n    files = list(glob.glob('train*.csv'))<br>\n    with open('consolidated-train.csv', 'w') as csvoutput:<br>\n        for i in files:<br>\n            with open(i,'r') as csvinput:<br>\n                writer = csv.writer(csvoutput, lineterminator='\\n')<br>\n                reader = csv.reader(csvinput)<br>\n                all = []                <br>\n                row = next(reader)<br>\n                row.append('segment_id')<br>\n                row.append('period')<br>\n                all.append(row)<br>\n                c = 0                <br>\n                for row in reader:<br>\n                    row.append(str(i[6:-4]))    <br>\n                    row.append(c)<br>\n                    all.append(row)<br>\n                    c = c+1<br>\n                print(str(i[6:-4]))<br>\n                writer.writerows(all)</pre><p></p>\n<p>consolidate_train()<br>\n</p>\n<p></p><pre><br>\ndef consolidate_train_drop_headers():<br>\n    with open(\"consolidated-train.csv\",\"r\") as inputfile, open(\"output.csv\",\"w\",newline=\"\") as outputfile:<br>\n       csv_in = csv.reader(inputfile)<br>\n       csv_out = csv.writer(outputfile)<br>\n       title = next(csv_in)<br>\n       csv_out.writerow(title)<br>\n       for row in csv_in:<br>\n            if row != title:<br>\n                 csv_out.writerow(row)</pre><p></p>\n<p>consolidate_train_drop_headers()<br>\n</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1063312,
      "author_name": "mooneral",
      "author_url": "",
      "post_date": "2020-10-28T18:17:47.687000",
      "content": "<p><a href=\"https://www.kaggle.com/javagarm\" target=\"_blank\">@javagarm</a> the problem is preparing the training set, as other libraries than <code>deep learning</code> do not support <code>gpu</code>. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1063306,
      "author_name": "mooneral",
      "author_url": "",
      "post_date": "2020-10-28T18:14:03.293000",
      "content": "<p><a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a> Indeed I took the same approach but somehow thought there may be a better way to speed up the preprocessing. I've tried <code>dask</code> and <code>numba</code> but no help. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1063310,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2020-10-28T18:15:59.753000",
          "content": "<p>Good luck! I keep my option, though. I aleready have my database with csv files to train, and I add more pre-processed files to it, over and over.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1062877,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-28T09:24:09.570000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1062857": "Hi Kagglers,\nI've been working on this dataset for a couple of days and I have written functions to read training data and do some transformation like fft/cwt then PCA but it takes too much time. Currently, it takes about 10 hours for this process to complete (approx 8 sec/iter for a single training file).\nIf you are working on this data set what are your approach/suggestions to make this process much faster? Or Am I taking the wrong avenue?",
    "1074875": "import dataset into database and run the feature enginining at database. It could speed up the data preparing process, especially the featuring engineering repetitively. \n\nIn database, the feature engineering could be done in many mode, many time and save the results in structured tempoaray place(materialized view)\n\nMy approach:\nStage A( ETL mode by manual or airflow)\n1. Import data into data base (around 2 hour)Database: Postgres 12\n2. install database development Plugin: Py/Python(install plpython3u extension). It can let data do science computation at database with ETL data to jupyternootbook. \n3. create function to working necessary feature engineering work.\n4. create view of x_train, x_test as view\nStage B( machine learning part)\n5. using pandas to the prepared data from database.\n6. feed into lightgbm or optuna for train and hyperparameter optimization\n7. submit it if you are OK with the local CV results.\n\nConsidering the Stage A need plenty of time and repetitive work, I think it deserved to let it run at database.\n\nFor example, create an function to process the sensors parameters like mean, std, skew, kurt, missing_count can be get in two steps lik following mode:\n\n1. Create an function in database(PG12): \n\nCREATE OR REPLACE FUNCTION public.py_thdf_stat_min_max(hour_id integer, shift integer)\n RETURNS real[]\n LANGUAGE plpython3u\n COST 500\nAS $function$\n    import pandas as pd\n    \n    cols = ['sensor_1','sensor_2','sensor_3', 'sensor_4','sensor_5',\n    \t\t 'sensor_6','sensor_7','sensor_8','sensor_9','sensor_10'\n    \t\t]\n        \n    agg_cols = []\n    result_sql = plpy.execute(f\"\"\"\n\t\t\t\t\t\t\t  select tm.* \n\t\t\t\t\t\t\t  from \n\t\t  \t\t\t\t\t    t_measures tm\n\t\t\t\t\t\t\t  inner join \n\t\t\t\t\t\t\t    v_train_per_hour vtph \n\t\t\t\t\t\t\t  on \n\t\t\t\t\t\t\t    tm.segment_id = vtph.segment_id \n\t\t\t\t\t\t\t  and \n\t\t\t\t\t\t\t    vtph.hour_to_eruption={hour_id} \n\t\t\t\t\t\t\t  order by \n\t\t\t\t\t\t\t    ts_id\"\"\")\n    \n    df = pd.DataFrame.from_records(result_sql)[cols]\n    \n    if shift == 0:\n    \tagg_cols = ['size','count','mean','median','std','skew','kurt','min','max']\n    else:\n        agg_cols = ['mean','median','std','skew','kurt','min','max']\n        df = df.fillna(0).diff(periods=shift)\n    \n    df_stat = df.agg(agg_cols)\n    df_stat['hour_id'] = hour_id\n    df_stat['shift'] = shift\n    \n    return df_stat.values.tolist()\n\n$function$\n;\n\n2. create an materialized view to get data. \nNote: In order to speed up the process, the data has been stored in array data type. \nCREATE MATERIALIZED VIEW public.mv_train_tsdf_stat\nTABLESPACE pg_default\nAS SELECT t.segment_id,\n    t.time_to_eruption,\n    t.time_to_eruption / 100 / 3600 AS hour_to_eruption,\n    py_tdf_stat(t.segment_id, 0) AS df_stat,\n    py_tdf_stat(t.segment_id, 1) AS df_stat_diff_1,\n    py_tdf_stat(t.segment_id, 100) AS df_stat_diff_100\n   FROM train t\nWITH DATA;\n\n3. create view to generate the x_train for next step;\nCREATE MATERIALIZED VIEW public.x_train\nTABLESPACE pg_default\nAS WITH seg2hour AS (\n         SELECT vsl.segment_id,\n            vsl.time_to_eruption,\n            vhl.hour_to_eruption,\n            vsl.nlag,\n            unnest(((((((((vsl.sensor1_stat[2:] || vsl.sensor2_stat[2:]) || vsl.sensor3_stat[2:]) || vsl.sensor4_stat[2:]) || vsl.sensor5_stat[2:]) || vsl.sensor6_stat[2:]) || vsl.sensor7_stat[2:]) || vsl.sensor8_stat[2:]) || vsl.sensor9_stat[2:]) || vsl.sensor10_stat[2:]) AS seg_stat,\n            unnest(((((((((vhl.sensor1_stat[2:] || vhl.sensor2_stat[2:]) || vhl.sensor3_stat[2:]) || vhl.sensor4_stat[2:]) || vhl.sensor5_stat[2:]) || vhl.sensor6_stat[2:]) || vhl.sensor7_stat[2:]) || vhl.sensor8_stat[2:]) || vhl.sensor9_stat[2:]) || vhl.sensor10_stat[2:]) AS hour_stat\n           FROM v_train_stat_per_sensor_lags vsl\n             JOIN v_train_stat_per_hour_lags vhl ON vhl.nlag = vsl.nlag\n        )\n SELECT seg2hour.segment_id,\n    seg2hour.time_to_eruption,\n    seg2hour.hour_to_eruption,\n    count(*) AS count,\n    corr(COALESCE(NULLIF(seg2hour.seg_stat, 'NaN'::real), 0::real)::double precision, COALESCE(NULLIF(seg2hour.hour_stat, 'NaN'::real), 0::real)::double precision) AS corr\n   FROM seg2hour\n  GROUP BY seg2hour.segment_id, seg2hour.time_to_eruption, seg2hour.hour_to_eruption\n  ORDER BY seg2hour.segment_id, seg2hour.hour_to_eruption\nWITH DATA;\n\nTo sumup,  one step to import dataset in to postgresql database, and do the feature engineering repetitively work in datbase. It can significantly save featuring engineering time.  And eay, because it is just SQL language. \n\nOnly let small piece of data runing into training/predicting. \n\nThat all of my sharing.\n\nIf you find it useful, pls vote it up.\n\nWangYong",
    "1064096": "Hi, i will  just say, accordingly to my intuition regarding the spectral analysis, to work only on one sensor at the beginning, especially if you're doing fft/cwt computations. There are plenty of notebook in this competition that can help you to choose the sensor with less NA values and more 'events'. The postulate i have, after reading some articles on the seismic activity regarding volcano, is to chase for a phenomena visible by spectral analysis, and that this phenomena is available on all sensors. May be i'm wrong about this hypothesis, to be tested with correlation matrix for instance, but  at least it simplify your job at the beginning of your spectral study, reduce the time need to processed the data and help you to improve your hypothesis.. You can also split the train set to generate a new train(80 % ) &test set(20 % ) for having a rather good idea of the model capacity to fill the objective. Once you get some interesting results, apply the model to others sensors.\n\nJust my 2 cents feeling.",
    "1063091": "What I did is having a separate pre-processing stage and saving the result in one or several csv files. Then I upload the csv files in a database I created and run training and inference using them. This strategy is allowed by the competition rules, if I am not mistaken.",
    "1075992": "You can use my dataset. It contains all the challenge data in a way that it is fast to read and process.\nhttps://www.kaggle.com/itamargr/volcaniceruptionsdata\n\nItamar",
    "1064693": "In my processing, the following two points have made the speed considerably faster during preprocessing.\nIt may overlap with the contents described by other people, but I will describe it just in case.\n- Convert the CSV file to a parquet file in advance, and use the parquet file during preprocessing. -> (1)\n- The new feature will be kept as a list and finally combined using the concat function. -> (2)\n\n(Excuse me my English is  very poor.)\n\n(1)\n```Python\nfor idx in tqdm.tqdm_notebook(range(len(train_df))):\n    segment_id = train_df.iloc[idx,0]\n    detail_df = pd.read_csv(f'{TRAIN_FILE_PATH}/{segment_id}.csv')\n    detail_df.to_parquet(f'{segment_id}.parquet')\n```\n\n(2)\n```Python\nfe_types = ['min','max','range','mean',]\n\ndef preprocessing(segments, df_len, path, interpolate=False):\n    maps = {}\n\n    for feat in sensors:\n        for fe_type in fe_types:\n            maps[f'{feat}_{fe_type}'] = []\n    \n    for idx in tqdm.tqdm_notebook(range(df_len)):\n        segment_id = segments[idx]\n        detail_df = pd.read_parquet(f'{path}/{segment_id}.parquet').astype(np.float32)\n        sensor_data_len = len(detail_df)\n        \n        for feat in sensors:\n            maps[feat+'_min'].append(detail_df[feat].min())\n            maps[feat+'_max'].append(detail_df[feat].max())\n            maps[feat+'_range'].append(maps[feat+'_max'][idx] - maps[feat+'_min'][idx])\n            maps[feat+'_mean'].append(detail_df[feat].mean())\n    return maps\n\nmaps = preprocessing(train_df.segment_id.tolist(), len(train_df), TRAIN_PARQUET_PATH)\nfor feat in maps.keys():\n    train_df = pd.concat([train_df, pd.DataFrame(maps[feat], columns=[feat])], axis=1)    \n```",
    "1063981": "The code below use pure Python to consolidate csv files into one (20gb)\n\n<pre>\nimport glob\nimport csv\ndef consolidate_train():\n    files = list(glob.glob('train\\*.csv'))\n    with open('consolidated-train.csv', 'w') as csvoutput:\n        for i in files:\n            with open(i,'r') as csvinput:\n                writer = csv.writer(csvoutput, lineterminator='\\n')\n                reader = csv.reader(csvinput)\n                all = []                \n                row = next(reader)\n                row.append('segment_id')\n                row.append('period')\n                all.append(row)\n                c = 0                \n                for row in reader:\n                    row.append(str(i[6:-4]))    \n                    row.append(c)\n                    all.append(row)\n                    c = c+1\n                print(str(i[6:-4]))\n                writer.writerows(all)\n\nconsolidate_train()\n<\\pre>\n\n\n\n<pre>\ndef consolidate_train_drop_headers():\n    with open(\"consolidated-train.csv\",\"r\") as inputfile, open(\"output.csv\",\"w\",newline=\"\") as outputfile:\n       csv_in = csv.reader(inputfile)\n       csv_out = csv.writer(outputfile)\n       title = next(csv_in)\n       csv_out.writerow(title)\n       for row in csv_in:\n            if row != title:\n                 csv_out.writerow(row)\n\nconsolidate_train_drop_headers()\n<\\pre>",
    "1063312": "@javagarm the problem is preparing the training set, as other libraries than `deep learning` do not support `gpu`. ",
    "1063306": "@catadanna Indeed I took the same approach but somehow thought there may be a better way to speed up the preprocessing. I've tried `dask` and `numba` but no help. ",
    "1062877": ""
  }
}