{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Feature Engineering: Aggregation Functions\n\nFollow up of my baseline agregation notebook https://www.kaggle.com/code/lucasmorin/amex-feature-engineering\n\nThis notebooks provides aggregation functions idea. The functions that currently do not appears in other notebooks mostly come from the time series manipulation package tsfresh (https://tsfresh.readthedocs.io/en/latest/), work on past competitions (Optiver / Google Brain) and personnal finance education / experience (drawdowns, max over min). \n\n\nNote the goal is to use pandas feature agregation functions. A GPU approach would probably be faster. However as we are only allowed 2 GPU sessions at once I usually prefer 10 standard CPU approach.\n\n**This Notebook is part of a serie built for the AMEX competition:**\n- [Base Feature engineering](https://www.kaggle.com/code/lucasmorin/amex-feature-engineering-base)\n- [Baseline lgbm](https://www.kaggle.com/code/lucasmorin/amex-lgbm-features-eng)\n- [Feature Engineering 2: aggregation function](https://www.kaggle.com/code/lucasmorin/amex-feature-engineering-2-aggreg-functions)\n- [Feature Engineering 3: transformation function](https://www.kaggle.com/code/lucasmorin/amex-feature-engineering-3-transform-functions)\n\n**With associated Data Sets:**\n- [Base Feature engineering](https://www.kaggle.com/datasets/lucasmorin/amex-base-fe)\n- [Feature Engineering 2 - aggregation function](https://www.kaggle.com/datasets/lucasmorin/amex-fe2)\n- [Feature Engineering 3 - transform function](https://www.kaggle.com/datasets/lucasmorin/amex-fe3)\n\n**Please make sure to upvote everything you use / find interesting / usefull**","metadata":{"papermill":{"duration":0.016481,"end_time":"2021-09-05T16:32:48.105499","exception":false,"start_time":"2021-09-05T16:32:48.089018","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport warnings\nimport gc","metadata":{"papermill":{"duration":1.11019,"end_time":"2021-09-05T16:32:49.233491","exception":false,"start_time":"2021-09-05T16:32:48.123301","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-10T16:21:52.362381Z","iopub.execute_input":"2022-07-10T16:21:52.36281Z","iopub.status.idle":"2022-07-10T16:21:52.368692Z","shell.execute_reply.started":"2022-07-10T16:21:52.362777Z","shell.execute_reply":"2022-07-10T16:21:52.367522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Pandas groupby - agg\n\nThe general idea is to use the groupby pandas approach. We can group by customer then apply aggregation functions.\n\nThere is different way to then perform the aggregation:\n\n    - Using base functions\n    \n    - Using list of functions with agg\n    \n    - Using dictionnary with agg","metadata":{}},{"cell_type":"markdown","source":"# Pandas base function\n\nPandas allows 13 agregation functions to be called by name. Their implementation is optimized, so you might prefer using their name.\n\n-    mean(): Compute mean of groups\n-    sum(): Compute sum of group values\n-    size(): Compute group sizes\n-    count(): Compute count of group\n-    std(): Standard deviation of groups\n-    var(): Compute variance of groups\n-    sem(): Standard error of the mean of groups\n-    describe(): Generates descriptive statistics\n-    first(): Compute first of group values\n-    last(): Compute last of group values\n-    nth() : Take nth value, or a subset if n is a list\n-    min(): Compute min of group values\n-    max(): Compute max of group values\n\nsum, mean, std, and sem even have a cython optimised implementation.\n\nHowever, the handling of edge case is not always ideal and you might want specific implementation to deal with them.\n\nFor exemple, depending on the usecase, you could want different result for a mean including a missing value.","metadata":{}},{"cell_type":"markdown","source":"general usage as functions:","metadata":{}},{"cell_type":"code","source":"df = pd.DataFrame({'group':[1,1,2,2],'values':[4,1,1,2],'values2':[0,1,1,2]})\ndf","metadata":{"execution":{"iopub.status.busy":"2022-07-10T16:21:52.432487Z","iopub.execute_input":"2022-07-10T16:21:52.432877Z","iopub.status.idle":"2022-07-10T16:21:52.443784Z","shell.execute_reply.started":"2022-07-10T16:21:52.432845Z","shell.execute_reply":"2022-07-10T16:21:52.4429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.groupby('group').mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-10T16:21:52.493207Z","iopub.execute_input":"2022-07-10T16:21:52.493576Z","iopub.status.idle":"2022-07-10T16:21:52.504872Z","shell.execute_reply.started":"2022-07-10T16:21:52.493545Z","shell.execute_reply":"2022-07-10T16:21:52.504055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Aggregation with list of function\n\nIt is possible to aggregate with multiple functions at a time with the agg function.\nThe agg function takke functions as inputs (notice we use np functions here).","metadata":{}},{"cell_type":"code","source":"df.groupby('group').agg([np.mean,np.std])","metadata":{"execution":{"iopub.status.busy":"2022-07-10T16:21:52.546236Z","iopub.execute_input":"2022-07-10T16:21:52.546985Z","iopub.status.idle":"2022-07-10T16:21:52.569361Z","shell.execute_reply.started":"2022-07-10T16:21:52.546932Z","shell.execute_reply":"2022-07-10T16:21:52.568566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can get back to pandas function by calling the function by their name (it only works for the 13 functions mentionned above).","metadata":{}},{"cell_type":"code","source":"df.groupby('group').agg(['mean','std'])","metadata":{"execution":{"iopub.status.busy":"2022-07-10T16:21:52.602269Z","iopub.execute_input":"2022-07-10T16:21:52.602683Z","iopub.status.idle":"2022-07-10T16:21:52.622014Z","shell.execute_reply.started":"2022-07-10T16:21:52.602647Z","shell.execute_reply":"2022-07-10T16:21:52.621084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Aggregation with dictionnary of functions\n\nApplying all fuctions to all features might be a bit over the top. Often you want to apply specific functions to specific columns.\nAgg allows this by passing a dictionnary:\n\n\n\n","metadata":{}},{"cell_type":"code","source":"dict_FE = {'values':['mean','median'],\n'values2':['mean','std']}\n\ndf.groupby('group').agg(dict_FE)","metadata":{"execution":{"iopub.status.busy":"2022-07-10T16:21:52.660017Z","iopub.execute_input":"2022-07-10T16:21:52.660742Z","iopub.status.idle":"2022-07-10T16:21:52.682879Z","shell.execute_reply.started":"2022-07-10T16:21:52.660691Z","shell.execute_reply":"2022-07-10T16:21:52.681549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, agg usually create a multindexed dataframe... it is often usefull to rename the columns:","metadata":{}},{"cell_type":"code","source":"df_agg = df.groupby('group').agg(dict_FE)\ndf_agg.columns = [c[0]+'-'+c[1] for c in df_agg.columns]\ndf_agg","metadata":{"execution":{"iopub.status.busy":"2022-07-10T16:21:52.714642Z","iopub.execute_input":"2022-07-10T16:21:52.715312Z","iopub.status.idle":"2022-07-10T16:21:52.733164Z","shell.execute_reply.started":"2022-07-10T16:21:52.715265Z","shell.execute_reply":"2022-07-10T16:21:52.73196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Tools","metadata":{"papermill":{"duration":0.015314,"end_time":"2021-09-05T16:32:52.390177","exception":false,"start_time":"2021-09-05T16:32:52.374863","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def _roll(a, shift):\n    \"\"\" Roll 1D array elements. Improves the performance of numpy.roll()\"\"\"\n\n\n    if not isinstance(a, np.ndarray):\n        a = np.asarray(a)\n    idx = shift % len(a)\n    return np.concatenate([a[-idx:], a[:-idx]])\n\n\ndef _get_length_sequences_where(x):\n    \"\"\" This method calculates the length of all sub-sequences where the array x is either True or 1. \"\"\"\n    if len(x) == 0:\n        return [0]\n    else:\n        res = [len(list(group)) for value, group in itertools.groupby(x) if value == 1]\n        return res if len(res) > 0 else [0]\n\ndef _aggregate_on_chunks(x, f_agg, chunk_len):\n    \"\"\"Takes the time series x and constructs a lower sampled version of it by applying the aggregation function f_agg on\n    consecutive chunks of length chunk_len\"\"\"\n    \n    return [\n        getattr(x[i * chunk_len : (i + 1) * chunk_len], f_agg)()\n        for i in range(int(np.ceil(len(x) / chunk_len)))\n    ]\n\ndef _into_subchunks(x, subchunk_length, every_n=1):\n    \"\"\"Split the time series x into subwindows of length \"subchunk_length\", starting every \"every_n\".\"\"\"\n    len_x = len(x)\n\n    assert subchunk_length > 1\n    assert every_n > 0\n\n    # how often can we shift a window of size subchunk_length over the input?\n    num_shifts = (len_x - subchunk_length) // every_n + 1\n    shift_starts = every_n * np.arange(num_shifts)\n    indices = np.arange(subchunk_length)\n\n    indexer = np.expand_dims(indices, axis=0) + np.expand_dims(shift_starts, axis=1)\n    return np.asarray(x)[indexer]\n\n\ndef set_property(key, value):\n    \"\"\"\n    This method returns a decorator that sets the property key of the function to value\n    \"\"\"\n\n    def decorate_func(func):\n        setattr(func, key, value)\n        if func.__doc__ and key == \"fctype\":\n            func.__doc__ = (\n                func.__doc__ + \"\\n\\n    *This function is of type: \" + value + \"*\\n\"\n            )\n        return func\n\n    return decorate_func","metadata":{"papermill":{"duration":0.032465,"end_time":"2021-09-05T16:32:52.438138","exception":false,"start_time":"2021-09-05T16:32:52.405673","status":"completed"},"tags":[],"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T16:21:52.747482Z","iopub.execute_input":"2022-07-10T16:21:52.747856Z","iopub.status.idle":"2022-07-10T16:21:52.761777Z","shell.execute_reply.started":"2022-07-10T16:21:52.747826Z","shell.execute_reply":"2022-07-10T16:21:52.760601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Aggregation functions\n\nSo feature engineering becomes having and choosing features to build from groups. Below is a set of functions that are usefull for time series aggregations. Feel free to comment if you think one is missing.","metadata":{}},{"cell_type":"code","source":"def median(x):\n    return np.median(x)\n\ndef variation_coefficient(x):\n    mean = np.mean(x)\n    if mean != 0:\n        return np.std(x) / mean\n    else:\n        return np.nan\n\ndef variance(x):\n    return np.var(x)\n\ndef skewness(x):\n    if not isinstance(x, pd.Series):\n        x = pd.Series(x)\n    return pd.Series.skew(x)\n\ndef kurtosis(x):\n    if not isinstance(x, pd.Series):\n        x = pd.Series(x)\n    return pd.Series.kurtosis(x)\n\ndef standard_deviation(x):\n    return np.std(x)\n\ndef large_standard_deviation(x):\n    if (np.max(x)-np.min(x)) == 0:\n        return np.nan\n    else:\n        return np.std(x)/(np.max(x)-np.min(x))\n\ndef variation_coefficient(x):\n    mean = np.mean(x)\n    if mean != 0:\n        return np.std(x) / mean\n    else:\n        return np.nan\n\ndef variance_std_ratio(x):\n    y = np.var(x)\n    if y != 0:\n        return y/np.sqrt(y)\n    else:\n        return np.nan\n\ndef ratio_beyond_r_sigma(x, r):\n    if x.size == 0:\n        return np.nan\n    else:\n        return np.sum(np.abs(x - np.mean(x)) > r * np.asarray(np.std(x))) / x.size\n\ndef range_ratio(x):\n    mean_median_difference = np.abs(np.mean(x) - np.median(x))\n    max_min_difference = np.max(x) - np.min(x)\n    if max_min_difference == 0:\n        return np.nan\n    else:\n        return mean_median_difference / max_min_difference\n    \ndef has_duplicate_max(x):\n    return np.sum(x == np.max(x)) >= 2\n\ndef has_duplicate_min(x):\n    return np.sum(x == np.min(x)) >= 2\n\ndef has_duplicate(x):\n    return x.size != np.unique(x).size\n\ndef count_duplicate_max(x):\n    return np.sum(x == np.max(x))\n\ndef count_duplicate_min(x):\n    return np.sum(x == np.min(x))\n\ndef count_duplicate(x):\n    return x.size - np.unique(x).size\n\ndef sum_values(x):\n    if len(x) == 0:\n        return 0\n    return np.sum(x)\n\ndef log_return(list_stock_prices):\n    return np.log(list_stock_prices).diff() \n\ndef realized_volatility(series):\n    return np.sqrt(np.sum(series**2))\n\ndef realized_abs_skew(series):\n    return np.power(np.abs(np.sum(series**3)),1/3)\n\ndef realized_skew(series):\n    return np.sign(np.sum(series**3))*np.power(np.abs(np.sum(series**3)),1/3)\n\ndef realized_vol_skew(series):\n    return np.power(np.abs(np.sum(series**6)),1/6)\n\ndef realized_quarticity(series):\n    return np.power(np.sum(series**4),1/4)\n\ndef count_unique(series):\n    return len(np.unique(series))\n\ndef count(series):\n    return series.size\n\n#drawdons functions are mine\ndef maximum_drawdown(series):\n    series = np.asarray(series)\n    if len(series)<2:\n        return 0\n    k = series[np.argmax(np.maximum.accumulate(series) - series)]\n    i = np.argmax(np.maximum.accumulate(series) - series)\n    if len(series[:i])<1:\n        return np.NaN\n    else:\n        j = np.max(series[:i])\n    return j-k\n\ndef maximum_drawup(series):\n    series = np.asarray(series)\n    if len(series)<2:\n        return 0\n    \n\n    series = - series\n    k = series[np.argmax(np.maximum.accumulate(series) - series)]\n    i = np.argmax(np.maximum.accumulate(series) - series)\n    if len(series[:i])<1:\n        return np.NaN\n    else:\n        j = np.max(series[:i])\n    return j-k\n\ndef drawdown_duration(series):\n    series = np.asarray(series)\n    if len(series)<2:\n        return 0\n\n    k = np.argmax(np.maximum.accumulate(series) - series)\n    i = np.argmax(np.maximum.accumulate(series) - series)\n    if len(series[:i]) == 0:\n        j=k\n    else:\n        j = np.argmax(series[:i])\n    return k-j\n\ndef drawup_duration(series):\n    series = np.asarray(series)\n    if len(series)<2:\n        return 0\n\n    series=-series\n    k = np.argmax(np.maximum.accumulate(series) - series)\n    i = np.argmax(np.maximum.accumulate(series) - series)\n    if len(series[:i]) == 0:\n        j=k\n    else:\n        j = np.argmax(series[:i])\n    return k-j\n\ndef max_over_min(series):\n    if len(series)<2:\n        return 0\n    if np.min(series) == 0:\n        return np.nan\n    return np.max(series)/np.min(series)\n\ndef mean_n_absolute_max(x, number_of_maxima = 1):\n    \"\"\" Calculates the arithmetic mean of the n absolute maximum values of the time series.\"\"\"\n    assert (\n        number_of_maxima > 0\n    ), f\" number_of_maxima={number_of_maxima} which is not greater than 1\"\n\n    n_absolute_maximum_values = np.sort(np.absolute(x))[-number_of_maxima:]\n\n    return np.mean(n_absolute_maximum_values) if len(x) > number_of_maxima else np.NaN\n\n\ndef count_above(x, t):\n    if len(x)==0:\n        return np.nan\n    else:\n        return np.sum(x >= t) / len(x)\n\ndef count_below(x, t):\n    if len(x)==0:\n        return np.nan\n    else:\n        return np.sum(x <= t) / len(x)\n\n#number of valleys = number_peaks(-x, n)\ndef number_peaks(x, n):\n    \"\"\"\n    Calculates the number of peaks of at least support n in the time series x. A peak of support n is defined as a\n    subsequence of x where a value occurs, which is bigger than its n neighbours to the left and to the right.\n    \"\"\"\n    x_reduced = x[n:-n]\n\n    res = None\n    for i in range(1, n + 1):\n        result_first = x_reduced > _roll(x, i)[n:-n]\n\n        if res is None:\n            res = result_first\n        else:\n            res &= result_first\n\n        res &= x_reduced > _roll(x, -i)[n:-n]\n    return np.sum(res)\n\ndef mean_abs_change(x):\n    return np.mean(np.abs(np.diff(x)))\n\ndef mean_change(x):\n    x = np.asarray(x)\n    return (x[-1] - x[0]) / (len(x) - 1) if len(x) > 1 else np.NaN\n\ndef mean_second_derivative_central(x):\n    x = np.asarray(x)\n    return (x[-1] - x[-2] - x[1] + x[0]) / (2 * (len(x) - 2)) if len(x) > 2 else np.NaN\n\n\ndef root_mean_square(x):\n    return np.sqrt(np.mean(np.square(x))) if len(x) > 0 else np.NaN\n\ndef absolute_sum_of_changes(x):\n    return np.sum(np.abs(np.diff(x)))\n\ndef longest_strike_below_mean(x):\n    if not isinstance(x, (np.ndarray, pd.Series)):\n        x = np.asarray(x)\n    return np.max(_get_length_sequences_where(x < np.mean(x))) if x.size > 0 else 0\n\ndef longest_strike_above_mean(x):\n    if not isinstance(x, (np.ndarray, pd.Series)):\n        x = np.asarray(x)\n    return np.max(_get_length_sequences_where(x > np.mean(x))) if x.size > 0 else 0\n\ndef count_above_mean(x):\n    m = np.mean(x)\n    return np.where(x > m)[0].size\n\ndef count_below_mean(x):\n    m = np.mean(x)\n    return np.where(x < m)[0].size\n\ndef last_location_of_maximum(x):\n    x = np.asarray(x)\n    return 1.0 - np.argmax(x[::-1]) / len(x) if len(x) > 0 else np.NaN\n\ndef first_location_of_maximum(x):\n    if not isinstance(x, (np.ndarray, pd.Series)):\n        x = np.asarray(x)\n    return np.argmax(x) / len(x) if len(x) > 0 else np.NaN\n\ndef last_location_of_minimum(x):\n    x = np.asarray(x)\n    return 1.0 - np.argmin(x[::-1]) / len(x) if len(x) > 0 else np.NaN\n\ndef first_location_of_minimum(x):\n    if not isinstance(x, (np.ndarray, pd.Series)):\n        x = np.asarray(x)\n    return np.argmin(x) / len(x) if len(x) > 0 else np.NaN\n\n# Test non-consecutive non-reoccuring values ?\ndef percentage_of_reoccurring_values_to_all_values(x):\n    if len(x) == 0:\n        return np.nan\n    unique, counts = np.unique(x, return_counts=True)\n    if counts.shape[0] == 0:\n        return 0\n    return np.sum(counts > 1) / float(counts.shape[0])\n\ndef percentage_of_reoccurring_datapoints_to_all_datapoints(x):\n    if len(x) == 0:\n        return np.nan\n    if not isinstance(x, pd.Series):\n        x = pd.Series(x)\n    value_counts = x.value_counts()\n    reoccuring_values = value_counts[value_counts > 1].sum()\n    if np.isnan(reoccuring_values):\n        return 0\n\n    return reoccuring_values / x.size\n\n\ndef sum_of_reoccurring_values(x):\n    unique, counts = np.unique(x, return_counts=True)\n    counts[counts < 2] = 0\n    counts[counts > 1] = 1\n    return np.sum(counts * unique)\n\ndef sum_of_reoccurring_data_points(x):\n    unique, counts = np.unique(x, return_counts=True)\n    counts[counts < 2] = 0\n    return np.sum(counts * unique)\n\ndef ratio_value_number_to_time_series_length(x):\n    if not isinstance(x, (np.ndarray, pd.Series)):\n        x = np.asarray(x)\n    if x.size == 0:\n        return np.nan\n\n    return np.unique(x).size / x.size\n\ndef abs_energy(x):\n    if not isinstance(x, (np.ndarray, pd.Series)):\n        x = np.asarray(x)\n    return np.dot(x, x)\n\ndef quantile(x, q):\n    if len(x) == 0:\n        return np.NaN\n    return np.quantile(x, q)\n\n# crossing the mean ? other levels ? \ndef number_crossing_m(x, m):\n    if not isinstance(x, (np.ndarray, pd.Series)):\n        x = np.asarray(x)\n    # From https://stackoverflow.com/questions/3843017/efficiently-detect-sign-changes-in-python\n    positive = x > m\n    return np.where(np.diff(positive))[0].size\n\ndef absolute_maximum(x):\n    return np.max(np.absolute(x)) if len(x) > 0 else np.NaN\n\ndef value_count(x, value):\n    if not isinstance(x, (np.ndarray, pd.Series)):\n        x = np.asarray(x)\n    if np.isnan(value):\n        return np.isnan(x).sum()\n    else:\n        return x[x == value].size\n\ndef range_count(x, min, max):\n    return np.sum((x >= min) & (x < max))\n\ndef mean_diff(x):\n    return np.nanmean(np.diff(x.values))","metadata":{"papermill":{"duration":0.093138,"end_time":"2021-09-05T16:32:52.546848","exception":false,"start_time":"2021-09-05T16:32:52.45371","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-10T16:21:52.839416Z","iopub.execute_input":"2022-07-10T16:21:52.839879Z","iopub.status.idle":"2022-07-10T16:21:52.91458Z","shell.execute_reply.started":"2022-07-10T16:21:52.839839Z","shell.execute_reply":"2022-07-10T16:21:52.913127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Lambda functions to facilitate application","metadata":{"papermill":{"duration":0.015317,"end_time":"2021-09-05T16:32:52.577833","exception":false,"start_time":"2021-09-05T16:32:52.562516","status":"completed"},"tags":[]}},{"cell_type":"code","source":"count_above_0 = lambda x: count_above(x,0)\ncount_above_0.__name__ = 'count_above_0'\n\ncount_below_0 = lambda x: count_below(x,0)\ncount_below_0.__name__ = 'count_below_0'\n\nvalue_count_0 = lambda x: value_count(x,0)\nvalue_count_0.__name__ = 'value_count_0'\n\ncount_near_0 = lambda x: range_count(x,-0.00001,0.00001)\ncount_near_0.__name__ = 'count_near_0_0'\n\nratio_beyond_01_sigma = lambda x: ratio_beyond_r_sigma(x,0.1)\nratio_beyond_01_sigma.__name__ = 'ratio_beyond_01_sigma'\n\nratio_beyond_02_sigma = lambda x: ratio_beyond_r_sigma(x,0.2)\nratio_beyond_02_sigma.__name__ = 'ratio_beyond_02_sigma'\n\nratio_beyond_03_sigma = lambda x: ratio_beyond_r_sigma(x,0.3)\nratio_beyond_03_sigma.__name__ = 'ratio_beyond_03_sigma'\n\nnumber_crossing_0 = lambda x: number_crossing_m(x,0)\nnumber_crossing_0.__name__ = 'number_crossing_0'\n\nquantile_01 = lambda x: quantile(x,0.1)\nquantile_01.__name__ = 'quantile_01'\n\nquantile_025 = lambda x: quantile(x,0.25)\nquantile_025.__name__ = 'quantile_025'\n\nquantile_075 = lambda x: quantile(x,0.75)\nquantile_075.__name__ = 'quantile_075'\n\nquantile_09 = lambda x: quantile(x,0.9)\nquantile_09.__name__ = 'quantile_09'\n\nnumber_peaks_2 = lambda x: number_peaks(x,2)\nnumber_peaks_2.__name__ = 'number_peaks_2'\n\nmean_n_absolute_max_2 = lambda x: mean_n_absolute_max(x,2)\nmean_n_absolute_max_2.__name__ = 'mean_n_absolute_max_2'\n\nnumber_peaks_5 = lambda x: number_peaks(x,5)\nnumber_peaks_5.__name__ = 'number_peaks_5'\n\nmean_n_absolute_max_5 = lambda x: mean_n_absolute_max(x,5)\nmean_n_absolute_max_5.__name__ = 'mean_n_absolute_max_5'\n\nnumber_peaks_10 = lambda x: number_peaks(x,10)\nnumber_peaks_10.__name__ = 'number_peaks_10'\n\nmean_n_absolute_max_10 = lambda x: mean_n_absolute_max(x,10)\nmean_n_absolute_max_10.__name__ = 'mean_n_absolute_max_10'","metadata":{"papermill":{"duration":0.032107,"end_time":"2021-09-05T16:32:52.625491","exception":false,"start_time":"2021-09-05T16:32:52.593384","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-10T16:21:52.91723Z","iopub.execute_input":"2022-07-10T16:21:52.917743Z","iopub.status.idle":"2022-07-10T16:21:52.932545Z","shell.execute_reply.started":"2022-07-10T16:21:52.917706Z","shell.execute_reply":"2022-07-10T16:21:52.931407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base_stats = ['mean','sum','size','count','std','first','last','min','max',median,skewness,kurtosis]\nhigher_order_stats = [abs_energy,root_mean_square,sum_values,realized_volatility,realized_abs_skew,realized_skew,realized_vol_skew,realized_quarticity]\nadditional_quantiles = [quantile_01,quantile_025,quantile_075,quantile_09]\nother_min_max = [absolute_maximum,max_over_min]\nmin_max_positions = [last_location_of_maximum,first_location_of_maximum,last_location_of_minimum,first_location_of_minimum]\npeaks = [number_peaks_2, mean_n_absolute_max_2, number_peaks_5, mean_n_absolute_max_5, number_peaks_10, mean_n_absolute_max_10]\ncounts = [count_unique,count,count_above_0,count_below_0,value_count_0,count_near_0]\nreoccuring_values = [count_above_mean,count_below_mean,percentage_of_reoccurring_values_to_all_values,percentage_of_reoccurring_datapoints_to_all_datapoints,sum_of_reoccurring_values,sum_of_reoccurring_data_points,ratio_value_number_to_time_series_length]\ncount_duplicate = [count_duplicate,count_duplicate_min,count_duplicate_max]\nvariations = [mean_diff,mean_abs_change,mean_change,mean_second_derivative_central,absolute_sum_of_changes,number_crossing_0]\nranges = [variance_std_ratio,ratio_beyond_01_sigma,ratio_beyond_02_sigma,ratio_beyond_03_sigma,large_standard_deviation,range_ratio]\n\nall_functions = base_stats + higher_order_stats + additional_quantiles + other_min_max + min_max_positions + peaks + counts + variations + ranges \n\n#+ reoccuring_values + count_duplicate : usually very slow.\n","metadata":{"papermill":{"duration":0.026553,"end_time":"2021-09-05T16:32:52.667871","exception":false,"start_time":"2021-09-05T16:32:52.641318","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-10T16:21:52.934224Z","iopub.execute_input":"2022-07-10T16:21:52.935394Z","iopub.status.idle":"2022-07-10T16:21:52.950623Z","shell.execute_reply.started":"2022-07-10T16:21:52.935348Z","shell.execute_reply":"2022-07-10T16:21:52.949701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After that you can put this features in the first notebook for aggregation :-)","metadata":{}},{"cell_type":"code","source":"DEBUG = False\n\n# https://stackoverflow.com/questions/2130016/splitting-a-list-into-n-parts-of-approximately-equal-length\ndef split(a, n):\n    k, m = divmod(len(a), n)\n    return (a[i*k+min(i, m):(i+1)*k+min(i+1, m)] for i in range(n))\n\n#main_features = [f'B_{i}' for i in [11,14,17]]+['D_39','D_131']+[f'S_{i}' for i in [16,23]]+['P_2','P_3']\nmain_features = ['P_2']\n\ndef prepare_dataset(train_test = 'train'):\n    \n    data = pd.read_parquet('../input/amex-data-integer-dtypes-parquet-format/'+train_test+'.parquet')\n    \n    if DEBUG:\n        data = data.iloc[:int((len(data)/60))]\n    \n    split_ids = split(data.customer_ID.unique(),10)\n\n    df_list = []\n\n    for (i,ids) in enumerate(split_ids):\n        print(i)\n        data_ids = data[data.customer_ID.isin(ids)]\n\n        data_agg = data_ids[main_features].groupby(data_ids.customer_ID).agg(all_functions)\n       \n        data_agg.columns = [c[0]+'-'+c[1] for c in data_agg.columns]\n\n        \n        df_list.append(data_agg)\n        gc.collect()\n\n    pd.concat(df_list,axis=0).astype('float16').to_pickle(train_test+'_data_agg2.pkl')","metadata":{"execution":{"iopub.status.busy":"2022-07-10T16:21:52.954134Z","iopub.execute_input":"2022-07-10T16:21:52.954754Z","iopub.status.idle":"2022-07-10T16:21:52.970338Z","shell.execute_reply.started":"2022-07-10T16:21:52.954694Z","shell.execute_reply":"2022-07-10T16:21:52.969168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nprepare_dataset(train_test = 'train')","metadata":{"execution":{"iopub.status.busy":"2022-07-10T16:21:52.973048Z","iopub.execute_input":"2022-07-10T16:21:52.973719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nprepare_dataset(train_test = 'test')","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}