{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# ParquetAnalyzer Class\n\nParquetAnalyzer is an abstraction of how to retrieve parquet features.\n\n## How to use\nYou can interact with it via its public methods and properties, which are listed in the class docstring.\n\nYou can ignore all the private methods if you are not developing new features.\n\nThere are examples of how to use ParquetAnalyzer after this block.\n\nYou can copy this class definition into your notebook to use it.\n\n## How to develop new parquet processing features\nTo add new preprocessing features, go to the end of the class definition where `_engineer_features_for_subject` is defined.\n\nYou will notice an area in that method denoted \"NEW PARQUET PROCESSING CODE\". This is where you can define new code to transform data from one subject's `parquet` DataFrame to create a new feature. This new feature should be added to the `features` dictionary.\n\n`features` must remain as a \"flat\" dictionary - i.e. every value must be a float. If you add a data that is a form of time series, for example, please add each entry of that series as a new key-value pair.\n\nSee `_create_hourly_features` for an example of parquet processing code that is called in this area.\n\nThe rest of the code will ensure that the features you add to the `features` dictionary will be incorporated in the final DataFrame created when you call a `ParquetAnalyzer`'s `engineer_features_for_subjects` method.","metadata":{}},{"cell_type":"code","source":"# ParquetAnalyzer Class Definition\n\nfrom concurrent.futures import ThreadPoolExecutor\nfrom typing import Optional\nimport os\n\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport plotly.express as px\nimport re\nfrom tqdm import tqdm\n\nclass ParquetAnalyzer:\n    \"\"\"\n    A ParquetAnalyzer object gives access to raw or engineered features from available parquet data.\n\n    Public methods and properties for an instance of ParquetAnalyzer named `pa` are as follows:\n        - pa.parquet_ids\n        - pa.get_raw_features(subject_id)\n        - pa.engineer_features_for_subject(subject_id)\n        - pa.engineer_features_for_subjects()\n    Refer to each public method or property's docstring for more information.\n    \"\"\"\n    \n    # --- PUBLIC METHODS --- (these provide no-fuss access to parquet data)\n    \n    def __init__(self, partition: str = \"train\"):\n        \"\"\"\n        Initialize an ParquetAnalyzer object.\n\n        Arguments:\n            - partition - Either \"train\" or \"test\", corresponding to series_train.parquet and series_test.parquet.\n        \"\"\"\n        # Take inventory of which subject ID's have parquet data, and the paths to that parquet data.\n        self._parquet_paths = {}\n        self._root = f\"/kaggle/input/child-mind-institute-problematic-internet-use/series_{partition}.parquet/\"\n        for id_subdir in os.listdir(self._root):\n            parquet_path = os.path.join(self._root, id_subdir, 'part-0.parquet')\n            parquet_id = self._path_to_subject_id(parquet_path)\n            self._parquet_paths[parquet_id] = parquet_path\n    \n    @property\n    def parquet_ids(self) -> list[str]:\n        \"\"\"\n        A list of subject ID's which have parquet data available.\n        An example of an ID for a subject is \"00115b9f\".\n        \"\"\"\n        return self._parquet_ids\n    \n    def get_raw_features(self, subject_id: str) -> pd.DataFrame:\n        \"\"\"\n        Return a pd.DataFrame of raw parquet data for a single subject.\n        \n        The following raw time signals are its columns (\"+\" indicates signals that are most likely to be useful):\n\n        Arguments:\n            - subject_id - The ID of a subject who has parquet data (such as \"00115b9f\")\n            \n        Returns:\n            - A table whose rows correspond to timesteps identified by the index \"step\", and whose columns are the following values at each timestep:\n                - X is a raw accelerometer signal captured in the X-axis. \n                    There may be direction-specific motion of interest, but for now assume ENMO summarizes these three, so don't necessarily need to use yet.\n                - Y is similar to X, but for the Y-axis.\n                - Z is similar to X, but for the Z-axis.\n                - enmo is approximately motion magnitude (Euclidean Norm Minus One). \n                    Could hypothesize this is low but possibly jittery (high frequency) for the chronically online, for example.\n                - anglez is angle from horizontal plane. \n                - non-wear_flag\n                - light is the amount of light sensed by the device. \n                    Probably corresponds to being indoors or outdoors - so a lower value might indicate being indoors. \n                    Even lower values might also suggest being indoors in the dark.\n                - battery_voltage \n                    Shouldn't matter, but lower battery might cause noise or attenuate signals?\n                - time_of_day could be used to partition data. \n                    If there's any motion late at night, it might imply the user is a night owl.\n                - weekday (i.e. which day of the week it is - 6 or 7 would be the weekend)\n                    Could also be interesting to partition on.\n                - quarter of the year, valued as one of 1, 2, 3, or 4.\n                    Also interesting for partitioning.\n                - relative_date_PCIAT is how many days since PCIAT was administered\n                    Feels like it wouldn't matter?\n        \"\"\"\n        return self._get_raw_features(subject_id)\n    \n    def engineer_features_for_subject(self, subject_id: str) -> dict[str, float]:\n        \"\"\"Calculate a dictionary of engineered features for a single subject.\n        \n        Arguments:\n            - subject_id: The ID of a subject who has parquet data (such as \"00115b9f\")\n            - plot (Optional): Plot features for this subject for helper functions which support this ability.\n            \n        Returns:\n            - A dictionary of relevant features from this parquet. Keys are feature names, values are feature values. \n        \"\"\"\n        return self._engineer_features_for_subject(subject_id)\n    \n    def engineer_features_for_subjects(self, subject_ids: Optional[list[str]] = None, parallel: bool = True) -> pd.DataFrame:\n        \"\"\"Calculate a pd.DataFrame of engineered features for ALL available subjects.\n        \n        Arguments:\n            - subject_ids (Optional) - \n                List of parquets to engineer features for. \n                Default behavior is to engineer features for all subjects.\n            - parallel (Optional) - \n                Set to True (default behavior) to process parquets in parallel.\n                Set to False to process parquets in series via a loop.\n    \n        Returns:\n            - A table whose columns are parquet-derived features, and whose rows are subjects indexed by their ID.\n        \"\"\"\n        return self._engineer_features_for_subjects(subject_ids, parallel)\n    \n    # --- END OF PUBLIC METHODS ---\n    \n    # --- PRIVATE METHODS --- (only need to interact with these if you want to modify how parquets are read)\n    \n    @property\n    def _parquet_ids(self) -> list[str]:\n        \"\"\"\n        A list of subject ID's which have parquet data available.\n        \n        See `parquet_ids` property for more information.\n        \"\"\"\n        return list(self._parquet_paths.keys())\n    \n    def _get_raw_features(self, subject_id: str) -> pd.DataFrame:\n        \"\"\"\n        Return a DataFrame of the given subject's raw parquet data.\n        \n        See `get_raw_features` method for more information.\n        \"\"\"\n        parquet_path = self._subject_id_to_path(subject_id)\n        raw_parquet = pd.read_parquet(parquet_path)\n        raw_parquet.set_index('step', inplace=True)\n        return raw_parquet\n        \n    def _engineer_features_for_subjects(self, subject_ids: Optional[list[str]] = None, parallel: bool = True, enable_hdcza: bool = False) -> pd.DataFrame:\n        \"\"\"\n        Retrieve relevant features from each parquet, and make a DataFrame out of these features.\n\n        See `engineer_features_for_subjects` method for more information.\n        \"\"\"\n        \n        if subject_ids is None:\n            subject_ids = self._parquet_ids\n        \n        if parallel:\n            # Extract parquet data from their files in parallel.\n            with ThreadPoolExecutor() as executor:\n                all_subjects_parquet_features: list[dict[str,float]] = list(tqdm(\n                    executor.map(lambda subject_id: self._engineer_features_for_subject(subject_id, enable_hdcza), subject_ids), total=len(subject_ids), desc=\"Engineering features from parquets\"\n                ))\n        else:\n            # Extract parquet data from their files in series. This may be easier to interpret.\n            all_subjects_parquet_features: list[dict[str,float]] = []\n            for subject_id in tqdm(subject_ids, total=len(subject_ids), desc=\"Engineering features from parquets\"):\n                parquet_features = self._engineer_features_for_subject(subject_id)\n                all_subjects_parquet_features.append(parquet_features)\n                \n        # Convert the list of  to a DataFrame\n        pqdf = pd.DataFrame(all_subjects_parquet_features).set_index(\"id\")\n\n        return pqdf\n    \n    @staticmethod\n    def _path_to_subject_id(path: str) -> str:\n        \"\"\"\n        Extract a subject's ID from the path to their parquet data.\n        This subject ID is in the same format as the subject ID's in `train.csv` or `test.csv`.\n        For example, \"00115b9f\" is a valid subject ID that this function might return.\n        \"\"\"\n        return os.path.basename(os.path.dirname(path)).split(\"=\")[1]\n    \n    def _subject_id_to_path(self, subject_id: str) -> str:\n        \"\"\"\n        Get the path to the given subject's parquet data.\n        \"\"\"\n        return self._parquet_paths[subject_id]\n\n    @staticmethod\n    def _mask_times(times: pd.Series, start: str, end: Optional[str] = None) -> pd.Series:\n        \"\"\"\n        Convenience function for producing a mask on a Series of timestamps between the provided `start` and `end` times.\n        This is done in the range: [`start`, `end`). i.e. the `start` time is included, but the `end` time is not.\n\n        The `start` and `end` times should be defined in 24-hour time in this exact format: \"HH:MM:SS\"\n        For example, \"18:59:00\" and \"06:01:30\" are valid.\n        \n        Arguments:\n            - times - A Series of pandas datetimes.\n            - start - A start time whose format is specified above.\n            - end (Optional) - An end time whose format is specified above. \n                If `end` is omitted, then the mask is true for all entries after `start`.\n        \n        Returns:\n            - A Series of boolean values which is True-valued for times in [`start`, `end`).\n        \"\"\"\n        # Ensure the time format is right.\n        for cutoff in [start, end]:\n            if cutoff is not None:\n                assert bool(re.match(r'^(?:[01]\\d|2[0-3]):[0-5]\\d:[0-5]\\d$', cutoff)), Exception(f\"Provide `{cutoff}` in 24-hour HH:MM:SS format.\")\n\n        # Return the mask.\n        after_start = times >= pd.Timestamp(\"1970-01-01 \" + start)\n        if end is None:\n            return after_start\n        \n        after_end = times < pd.Timestamp(\"1970-01-01 \" + end)\n        return after_start & after_end\n    \n    # --- \n    # The remaining methods here are used to derive features from one subject's `parquet` DataFrame.\n    # ---\n    \n    @staticmethod\n    def _create_hourly_features(parquet: pd.DataFrame, signals: list[str]) -> dict[str, float]:\n        \"\"\"\n        Create features related to average signal values within each hour of the parquet data.\n        \n        Arguments:\n            - parquet, a subject's parquet data.\n                Precondition: The `hour` column has been defined.\n            - signals, a list of column names in `parquet`.\n            \n        Returns:\n            - A dictionary containing the following key-value pairs:\n                - For each of the 24 hours and for each signal in signals, the \n        \n        Modifications to `parquet`:\n            None.\n        \"\"\"\n        hourly_features = {}\n        \n        # Take the mean within each hour group.\n        hourly_average_df = parquet.groupby('hour')[signals].mean()\n        hours = hourly_average_df.index\n        \n        # Approximate derivative as first-order difference.\n        hourly_average_derivative_df = hourly_average_df.diff()\n        \n        for signal in signals:\n            # Add the hourly averages as features.\n            for hour in hours:\n                hourly_features[f'{signal}_hr{hour}_avg'] = hourly_average_df.at[hour, signal]\n                \n            # Record which hours the maximum & minimum average values occur at.\n            hourly_features[f'{signal}_min_hr'] = hourly_average_df[signal].idxmin()\n            hourly_features[f'{signal}_max_hr'] = hourly_average_df[signal].idxmax()\n            \n            # Record which hours the maximum & minimum changes in average values occur at.\n            hourly_features[f'{signal}_diffmin_hr'] = hourly_average_derivative_df[signal].idxmin()\n            hourly_features[f'{signal}_diffmax_hr'] = hourly_average_derivative_df[signal].idxmax()\n\n        return hourly_features\n\n\n    @staticmethod\n    def _hdcza(parquet: pd.DataFrame, include_gaps = False, verbose = False):\n        \"\"\"\n        Implementation of HDCZA algorithm.\n        \n        HDCZA is the Heuristic algorithm looking at Distribution of Change in Z-Angle.\n        This uses z-angle only to predict the SPT-window (sleep period time window).\n\n        Returns 3 features:\n            - spt_start_time: Timestamp (i.e. an integer) representing start time of the sleep period window in SECONDS.\n            - spt_end_time: Timestamp (i.e. an integer) representing start time of the sleep period window in SECONDS.\n            - spt_duration: Integer representing duration of sleep in MINUTES.\n        \n        Version 2. Modified 11/17.\n\n        TODO: \n        - resolve issue of WHY these unreasonably long stretches appear, even when the 60 minute gaps are not used to merge blocks.\n        - resolve issue where some parquets still have non-monotonically increasing datetimes. For now, just forcing them to be sorted.\n        \"\"\"\n        df = parquet.copy()\n\n        # ---Prep for the HDCZA algorithm---\n\n        # Find where new weeks begin.\n        weekday_diff = df['weekday'].diff() < 0\n        week_start_idxs = [0] + np.where(weekday_diff.to_numpy())[0].tolist() + [len(df)]\n        # Create monotonically increasing enumerations of week and day.\n        df['absolute_week'] = pd.cut(pd.Series(range(len(df))), bins=week_start_idxs, labels=range(len(week_start_idxs)-1), include_lowest=True, right=False).astype(int)\n        df['absolute_day'] = df['weekday'] + (df['absolute_week']*7)\n\n        # Define the \"noon-to-noon\" day for grouping purposes.\n        df['after_noon'] = ParquetAnalyzer._mask_times(df['time_of_day'], \"12:00:00\")\n        df['noon2noon_day'] = df['absolute_day']\n        df.loc[df['after_noon'], 'noon2noon_day'] += 1\n\n        # The continuous monotonically increasing datetime.\n        df['datetime'] = pd.to_datetime(\"1970-01-01\") + pd.to_timedelta(df['absolute_day'], unit='days') + (df['time_of_day']-pd.Timestamp(\"1970-01-01\"))\n\n        # Force monotonic series.\n        df = df.drop_duplicates(subset='datetime', keep=\"last\").sort_values(by='datetime').reset_index()\n        \n        # Assert monotonically increasing datetime.\n        break_mask = (df['datetime'].diff() < pd.Timedelta(0))\n        if break_mask.sum() != 0:\n        #     # Below is what *should* be done instead of sorting datetimes. Need to check true reason why datetimes are out of order.\n            break_idxs = np.where(break_mask.to_numpy())[0].tolist()\n            display(df.iloc[break_idxs[0]-2:break_idxs[0]+3])\n            raise Exception(\"Found datetime decreases; first instance displayed above.\")\n\n        # ---Start the HDCZA algorithm---\n\n        # Steps 1 and 2 are assumed to be precalculated in anglez column.\n        \n        # Step 3: Calculate the absolute differences between successive averages\n        df['abs_diff'] = df['anglez'].diff().abs()\n        \n        # Step 4: Calculate a 5-minute rolling median of these differences\n        df.index = df['datetime'].copy()\n        df['rolling_median'] = df['abs_diff'].rolling('5min').median()\n        # df.reset_index(inplace=True)\n\n        # Step 5: Drop NaN values resulting from rolling calculation\n        df = df.dropna(subset=['rolling_median'])\n        \n        # Step 6: Calculate 10th percentile of the rolling median for each day\n        threshold = df.groupby('noon2noon_day')['rolling_median'].quantile(0.1).reset_index()\n        threshold['critical_threshold'] = threshold['rolling_median'] * 15\n        # Merge the critical threshold back to the main DataFrame\n        df = df.merge(threshold[['noon2noon_day', 'critical_threshold']], on='noon2noon_day', how='left')\n\n        # Step 7. Group stretches of low anglez activity into \"blocks\" and filter the data down to the blocks of low activity which last longer than 30 minutes.\n        df['below_threshold'] = df['rolling_median'] < df['critical_threshold']\n        # First, get blocks of continuous stretches of being below the activity critical threshold.\n        df['block_id'] = (df['below_threshold'].diff() == 1).cumsum()\n        # Filter the data down to just the blocks which are below the threshold.\n        df = df[df['below_threshold']]\n        \n        # Get the list of block IDs whose corresponding blocks are longer than 30 minutes\n        df['time_diff'] = df['datetime'].diff()\n        block_durations = df.groupby('block_id')['time_diff'].sum()\n        long_block_ids = list(block_durations[block_durations >= pd.Timedelta('30 minutes')].index)\n        # For the remainder of the algorithm, only consider blocks which are longer than 30 minutes.\n        df = df[df['block_id'].isin(long_block_ids)]\n\n\n        # Step 8: Include short time gaps between blocks as part of the blocks (gaps of less than 60 minutes)\n        if include_gaps:\n            time_diffs = df['datetime'].diff()\n            interblock_gaps = (time_diffs.where(df['block_id'].diff() != 0)).ffill().bfill()\n            df.loc[(interblock_gaps < pd.Timedelta('60 minutes')), 'block_id'] = None\n            df.loc[:,'block_id'] = df['block_id'].ffill().bfill()\n\n        # Calculate new block durations\n        block_durations = df.groupby('block_id')['time_diff'].sum()\n        block_durations = block_durations.reset_index().rename(columns={'time_diff': 'duration'})\n\n        # Filter out blocks which last unreasonably long for now. (1 day)\n        # TODO resolve issue of WHY these unreasonably long stretches appear,\n        # even when the 60 minute gaps are not used to merge blocks.\n        block_durations = block_durations[block_durations['duration'] < pd.Timedelta('1 day')]\n        df = df.merge(block_durations[['block_id', 'duration']], on='block_id', how='left')\n        \n        # Step 9: Get longest valid block in the day - this is the SPT-window.\n        longest_block_duration = block_durations.loc[block_durations['duration'].idxmax()]\n        spt_window = longest_block_duration[['duration']].loc['duration']\n\n        # Get the start and end times of the longest block.\n        longest_block_id = longest_block_duration[['block_id']].loc['block_id']\n        spt_start_time = df[df['block_id']==longest_block_id].iloc[0]['datetime']\n        spt_end_time = df[df['block_id']==longest_block_id].iloc[-1]['datetime']\n        \n        return {\n            'spt_start_timestamp': spt_start_time.timestamp(),\n            'spt_end_timestamp': spt_end_time.timestamp(),\n            'spt_duration_seconds': spt_window.total_seconds(),\n            'spt_start_hour': spt_start_time.hour,\n            'spt_end_hour': spt_end_time.hour,\n        }\n\n\n    def _add_common_features(self, parquet: pd.DataFrame) -> pd.DataFrame:\n        \"\"\"\n        Prepare columns that will be used across multiple preprocessing techniques.\n        \"\"\"\n        # Prepare columns that will be used across multiple preprocessing techniques.\n        \n        # Interpret time of day as nanoseconds and cast to pandas datetime type.\n        # Ignore the 1970-01-01 component - this column only represents the time of day.\n        parquet['time_of_day'] = pd.to_datetime(parquet['time_of_day'], unit='ns')\n\n        # Create hour bin to group data by.\n        parquet['hour'] = parquet['time_of_day'].dt.hour\n        \n        # Create smoothed signals using a small window to reduce noise.\n        signals = ['enmo', 'anglez', 'light']\n        for signal in signals:\n            parquet[\"smooth_\"+signal] = parquet[signal].rolling(window=5).mean()\n\n        return parquet\n    \n    def _engineer_features_for_subject(self, subject_id: str, enable_hdcza: bool = False, verbose: bool = True) -> dict[str, float]:\n        \"\"\"Calculate relevant features based on one subject's parquet file.\n\n        See `engineer_features_for_subject` method for more information.\n        \"\"\"\n        parquet: pd.DataFrame = self._get_raw_features(subject_id)\n        parquet = self._add_common_features(parquet) # NEW - moved common features to different function\n        \n        # **********************************************\n        # ***  NEW PARQUET PROCESSING CODE GOES HERE ***\n        # **********************************************\n\n        # This `features` dictionary is where you add new features for a given subject.\n        features = {\"id\": subject_id}\n        \n        # Use parquet to bin a few signals on an hourly basis, take the average within each hour, and derive some features from that.\n        hourly_averages = self._create_hourly_features(parquet, ['smooth_enmo', 'smooth_anglez', 'smooth_light'])\n        features.update(hourly_averages)\n\n        # Create sleep features from van Hees paper\n        spt_features = self._hdcza(parquet, verbose=verbose)\n        features.update(spt_features)\n        \n        # (more features to be calculated...)\n        \n        # ***********************************************\n        # ******* END NEW PARQUET PROCESSING CODE *******\n        # ***********************************************\n        return features","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-18T04:28:32.433312Z","iopub.execute_input":"2024-11-18T04:28:32.433753Z","iopub.status.idle":"2024-11-18T04:28:32.486156Z","shell.execute_reply.started":"2024-11-18T04:28:32.433715Z","shell.execute_reply":"2024-11-18T04:28:32.484528Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pa_train = ParquetAnalyzer(\"train\")\nids = pa_train.parquet_ids\n\npq_df_train = pa_train.engineer_features_for_subject(ids[0])\ndisplay(pq_df_train)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-18T04:28:32.488748Z","iopub.execute_input":"2024-11-18T04:28:32.489388Z","iopub.status.idle":"2024-11-18T04:28:32.658473Z","shell.execute_reply.started":"2024-11-18T04:28:32.489337Z","shell.execute_reply":"2024-11-18T04:28:32.657146Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ParquetAnalyzer Usage Example","metadata":{}},{"cell_type":"code","source":"# Sample Code for Extracting Features from Parquet Data using ParquetAnalyzer\n\n# Prepare a ParquetAnalyzer object to read parquet data.\npa_train = ParquetAnalyzer(\"train\")\n\n# Extract IDs of subjects who have parquet data.\nids = pa_train.parquet_ids\nprint(\"First 10 IDs of subjects with parquet data:\", ids[:10])\n\n# Get the raw parquet data of subject 3ab539f0.\nexample_raw_features = pa_train.get_raw_features('3ab539f0')\nprint(\"Raw features for 3ab539f0:\")\ndisplay(example_raw_features)\n\n# Engineer features for the parquet data of subject 3ab539f0.\nexample_engineered_features = pa_train.engineer_features_for_subject('3ab539f0')\nprint(\"Engineered features for 3ab539f0:\")\nprint(example_engineered_features)\n\n# Engineer features for a few subjects at once.\n# Write results to `pqdf`.\nexample_pq_df_train = pa_train.engineer_features_for_subjects(ids[:10], parallel=False)\nprint(\"Engineered features for 10 subjects:\")\ndisplay(example_pq_df_train)\n\n# Let's say you have the rest of the training data the CSV:\ndf_train = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv', index_col='id')\n\n# Parquet engineered features from multiple subjects can be combined with features from that CSV like this:\ndf_train_with_parquet = pd.concat(\n    [example_pq_df_train, df_train], \n    axis=1, join='outer')\nprint(\"Full training DataFrame with parquet data added:\")\ndisplay(df_train_with_parquet)\n\n# You can do something similar with the test data. In that case, use ParquetAnalyzer(\"test\").","metadata":{"execution":{"iopub.status.busy":"2024-11-18T04:28:32.737364Z","iopub.execute_input":"2024-11-18T04:28:32.737786Z","iopub.status.idle":"2024-11-18T04:28:37.434893Z","shell.execute_reply.started":"2024-11-18T04:28:32.737748Z","shell.execute_reply":"2024-11-18T04:28:37.433716Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Engineering Features for All Parquets\nDepending on which preprocessing steps are used, this may take a few minutes when using parallel processing.","metadata":{}},{"cell_type":"code","source":"%%time\npa_train = ParquetAnalyzer(\"train\")\npq_df_train = pa_train.engineer_features_for_subjects(parallel=True)\ndisplay(pq_df_train)","metadata":{"execution":{"iopub.status.busy":"2024-11-18T04:28:37.437392Z","iopub.execute_input":"2024-11-18T04:28:37.437808Z","iopub.status.idle":"2024-11-18T04:33:55.795794Z","shell.execute_reply.started":"2024-11-18T04:28:37.437763Z","shell.execute_reply":"2024-11-18T04:33:55.794470Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Plotting Some Parquet Engineered Features","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport plotly.express as px\n\npq_df_train_subset = pq_df_train.iloc[:5]\n\n# Filter columns that match the pattern \"something_<number>\"\nsignal_columns = [col for col in pq_df_train_subset.columns if col.startswith('smooth_enmo_hr')]\n\n# Melt the DataFrame to long format\npq_df_smooth_enmo_hr = pq_df_train_subset[signal_columns].copy()\npq_df_smooth_enmo_hr.columns = pq_df_smooth_enmo_hr.columns.str.replace(r'smooth_enmo_hr(\\d+)_avg', r'\\1', regex=True)\nlong_df = pq_df_smooth_enmo_hr.reset_index().melt(id_vars='id', var_name='signal', value_name='value')\n\n# Create an interactive plot using Plotly\nfig = px.line(long_df, x='signal', y='value', color='id', \n              title='Interactive Signal Plot', \n              labels={'value': 'Signal Value', 'signal': 'Signal Type'})\n\n# Show the plot\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-20T22:20:55.647470Z","iopub.execute_input":"2024-10-20T22:20:55.647895Z","iopub.status.idle":"2024-10-20T22:20:55.746565Z","shell.execute_reply.started":"2024-10-20T22:20:55.647842Z","shell.execute_reply":"2024-10-20T22:20:55.745506Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}