{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":96164,"databundleVersionId":12993472,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id = Section1></a>\n**1. INTRODUCTION**","metadata":{}},{"cell_type":"markdown","source":"The cryptocurrency market stands as a pulsating, digital frontier, a landscape of unprecedented dynamism that continuously generates vast streams of data. Within this evolving realm lie unparalleled opportunities for those capable of extracting meaningful insights from the cacophony of information.\n\nYet, this lucrative territory presents a formidable challenge: an inherently low signal-to-noise ratio that makes the identification of genuinely predictive patterns exceptionally difficult. Price movements are not random; they are orchestrated by an intricate dance of liquidity shifts, real-time order flow dynamics, subtle sentiment changes, and underlying structural inefficiencies.\n\nTo truly unravel these complex forces and unlock market foresight demands the precision and power of sophisticated quantitative techniques.","metadata":{}},{"cell_type":"markdown","source":"<a id = Section2></a>\n**2. PROBLEM STATEMENT**","metadata":{}},{"cell_type":"markdown","source":"Core Mission: To construct a robust machine learning model for predicting short-term crypto future price movements. Precisely the following objectives:\n\n1.**Information Fusion:** Synthesize these diverse, high-dimensional data sources into a single, highly effective directional price movement signal.          \n2.**Real-World Replication:** Tackle the raw, noisy complexities of live market data, mirroring the analytical intensity of DRW's daily operations.              \n3.**Unveiling Hidden Structures:** Design a learning model that not only identifies explicit patterns but also efficiently captures implicit interactions between all features to sharpen predictive accuracy.","metadata":{}},{"cell_type":"markdown","source":"<a id = Section4></a>\n**3.ENVIRONMENT SETUP & LIBRARIES IMPORT**","metadata":{}},{"cell_type":"code","source":"import sys\nsys.path.append('/kaggle/working/')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!rm -rf /kaggle/working/scipy\n!rm -rf /kaggle/working/statsmodels","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!rm -rf scipy\n!rm -rf statsmodels","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!pip install numpy==1.26.4 scipy==1.11.4 statsmodels==0.14.0 --no-cache-dir --force-reinstall","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#-------------------------------------------------------------------------------------------\nimport pandas as pd\nimport numpy as np\nfrom matplotlib import pyplot as plt                  # To perform data visualisation\nimport seaborn as sns                                 # To perform data visualisation\nimport plotly.express as px                           # To perform data visualisation\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\nimport statsmodels.api as sm\nfrom scipy.stats import skew, kurtosis\n%matplotlib inline\n                                             \nfrom sklearn.linear_model import LinearRegression     # To perform prediction\nfrom sklearn.tree import DecisionTreeRegressor\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.linear_model import Lasso, LassoCV, RidgeCV, ElasticNetCV\nimport lightgbm as lgb\nimport xgboost as xgb\nfrom sklearn.svm import SVR\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import LSTM, Dense, Dropout, BatchNormalization\nfrom tensorflow.keras.callbacks import EarlyStopping, ReduceLROnPlateau\nimport tensorflow as tf\nfrom tensorflow.keras import backend as K\n\nfrom sklearn.feature_selection import mutual_info_regression\nfrom sklearn.preprocessing import StandardScaler\nimport math\n\nfrom scipy.stats import chi2_contingency              # To perform chi sqr test\nfrom scipy import stats\nfrom scipy.stats import skewtest, norm\nfrom sklearn.preprocessing import PowerTransformer\n\nfrom sklearn.preprocessing import StandardScaler      # To perform feature scaling\nfrom sklearn.model_selection import train_test_split  # To perform train test split\n\nfrom sklearn.model_selection import GridSearchCV      # To perform hyperparameter tuning\nfrom sklearn.model_selection import RandomizedSearchCV, TimeSeriesSplit\nfrom scipy.stats import pearsonr                       # To perform pearson correlation\nfrom scipy.stats import loguniform\nfrom scipy.stats import uniform\nfrom random import randint\nimport copy\n\nfrom sklearn.metrics import mean_squared_error\nfrom sklearn.metrics import make_scorer\nfrom sklearn.metrics import r2_score\nfrom sklearn.inspection import permutation_importance\nimport gc\n#import mlflow\n\nsns.set_theme(style=\"white\", palette=\"muted\")\n%matplotlib inline\n\n#-------------------------------------------------------------------------------------------\nimport warnings                                       # Importing warning to disable runtime warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-30T23:33:07.089497Z","iopub.execute_input":"2025-11-30T23:33:07.089799Z","iopub.status.idle":"2025-11-30T23:33:33.317596Z","shell.execute_reply.started":"2025-11-30T23:33:07.089776Z","shell.execute_reply":"2025-11-30T23:33:33.317046Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id = Section5></a>\n**5.IMPORTING DATASETS**","metadata":{}},{"cell_type":"code","source":"#Class for importing datasets from the kaggle competition\n\nfilepath={'train_path':'/kaggle/input/drw-crypto-market-prediction/train.parquet',\n         'test_path':'/kaggle/input/drw-crypto-market-prediction/test.parquet'}\nclass data_import():\n    #Constructor\n    def __init__(self, filepath):\n        self.filepath = filepath\n        self.y_train_DRW = None\n        self.x_train_DRW = None\n        self.y_test_DRW = None\n        self.x_test_DRW = None\n    \n    #Function for importing train dataset\n    def file_import_train(self,num_records=250000):\n        train_path = self.filepath['train_path']\n        df_train = pd.read_parquet(train_path)\n\n        # Get the tail using df_train.tail() and then .compute()\n        df_train = df_train.tail(num_records)\n        \n        self.y_train_DRW=df_train[['label']]\n        self.x_train_DRW=df_train.drop(['label'],axis=1)\n       \n        return self.x_train_DRW,self.y_train_DRW\n\n    #Function for importing test dataset\n    def file_import_test(self,num_records=250000):\n        test_path = self.filepath['test_path']\n        df_test = pd.read_parquet(test_path)\n       \n        df_test = df_test.tail(num_records)\n        \n        self.y_test_DRW=df_test[['label']]\n        self.x_test_DRW=df_test.drop(['label'],axis=1)\n        \n        return self.x_test_DRW,self.y_test_DRW\n\n    def memory_optimize(self):\n        self.file_import_train()\n        self.file_import_test()\n\n        for df in self.x_train_DRW, self.y_train_DRW, self.x_test_DRW, self.y_test_DRW:\n            column_dtypes = df.dtypes\n            for col in df.columns:\n                current_dtype = column_dtypes[col] # Access the dtype from the Series of dtypes\n                if str(current_dtype) == 'float64':\n                    df[col]=df[col].astype('float32') # Converting to float 32 to optimize RAM\n                elif str(current_dtype) == 'int64':\n                    df[col]=df[col].astype('float32') # Converting to float 32 to optimize RAM\n        print('All dtypes converted to float32')\n        return self.x_train_DRW, self.y_train_DRW, self.x_test_DRW, self.y_test_DRW\n        \n#Calling the data_import class to populate train and test datasets from kaggle competition\ndata_import=data_import(filepath)\n#x_train_DRW,y_train_DRW=data_import.file_import_train()  #Train split on train.parquet dataset\n#x_test_DRW,y_test_DRW=data_import.file_import_test()     #Test split on test.parquet dataset\nx_train_DRW, y_train_DRW, x_test_DRW, y_test_DRW=data_import.memory_optimize()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-30T23:33:52.678151Z","iopub.execute_input":"2025-11-30T23:33:52.679044Z","iopub.status.idle":"2025-11-30T23:34:39.234028Z","shell.execute_reply.started":"2025-11-30T23:33:52.679015Z","shell.execute_reply":"2025-11-30T23:34:39.233323Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id = Section6></a>\n**6. DATA PREPROCESSING**\n\n1.HANDLING MISSING VALUES,INCONSISTENCIES/DUPLICATES/REDUNDANCIES                           \n2.HANDLING OUTLIERS","metadata":{}},{"cell_type":"markdown","source":"<a id = Section6.1></a>\n**6.1 HANDLING MISSING VALUES,INCONSISTENCIES/DUPLICATES/REDUNDANCIES**","metadata":{}},{"cell_type":"code","source":"dataframes = {'x_train_DRW': x_train_DRW, 'x_test_DRW': x_test_DRW}\n\n#Class to examine basic data characteristics - rows, columns, descriptive statistics, duplicates and missing values\nclass data_properties():\n    \n    # Constructor\n    def __init__(self, dataframes_dict):\n        # The __init__ should use the argument passed to it\n        self.dataframes = dataframes_dict\n\n    # Function to reflect shape, info, and descriptive statistics of each dataframe\n    def dataset_description(self):\n        # Loop through the key (name) and value (the DataFrame object)\n        # in the dictionary.\n        for name, df_object in self.dataframes.items():\n            print(f\"\\n==========================================\")\n            print(f\"Properties for DataFrame: '{name}'\")\n            print(f\"==========================================\")\n            \n            # Now, 'df_object' is an actual DataFrame, so you can use .shape, etc.\n            print(f\"Shape: {df_object.shape}\\n\")\n            \n            print(\"Info:\")\n            df_object.info() # .info() prints directly, so no need for a print() call\n            \n            print(f\"\\nDescriptive Statistics:\")\n            # .describe() returns a DataFrame, so we print it\n            print(df_object.describe())\n\n            print(f\"Total missing values in DataFrame: {df_object.isna().sum().sum()}\")\n            \n            #Treating duplicates if any\n            if df_object.duplicated().any():\n                print(f\"Duplicates found in '{name}'. Removing duplicates.\")\n                df_object.drop_duplicates(inplace=True)\n                print(\"Duplicates removed.\")\n            else:\n                print(f\"No duplicates found in {name}.\")\n                        \n            print(\"\\n\")\n\n            # Check for infinite values in dataset\n            has_infinite = df_object.isin([np.inf, -np.inf]).any().any()\n            \n            if has_infinite:\n                print(f\"The DataFrame '{name}' contains infinite values.\")\n            \n                # Finding which columns contain infinite values\n                cols_with_inf = df_object.columns[df_object.isin([np.inf, -np.inf]).any()]\n                print(\"Columns with infinite values:\", list(cols_with_inf))\n\n                # Dropping infinite values rows and columns\n                for col in list(cols_with_inf):\n                    if df_object[col].value_counts().values[0]==250000:   #if all values in a column are infinite\n                        df_object.drop(col, axis=1, inplace=True)\n            else:\n                print(f\"The DataFrame '{name}' does not contain infinite values.\")\n\n    #Function to test skewness - left or right \n    def skewness_test(self):\n        left_skewed_col=[]\n        right_skewed_col=[]\n        for df_object,df_content in self.dataframes.items():\n            for col in df_content.columns:\n                if df_content[col].mean()<df_content[col].median():\n                    left_skewed_col.append(col)\n                elif df_content[col].mean()>df_content[col].median():\n                    right_skewed_col.append(col)\n            \n            #Examining skewness for one feature through histogram and violin plot\n            print(f\"The feature 'X247' in {df_object} - Skewness depiction through Histogram, Violon Plot and Q-Q plot\")\n            # Create subplots: 1 row, 2 columns\n            fig = make_subplots(rows=1, cols=2,\n                                subplot_titles=('Histogram Distribution of X247', 'Violin plot of X247'))\n            \n            # Add Histogram\n            fig.add_trace(go.Histogram(x=df_content['X247'], name='Histogram', nbinsx=500), \n                          row=1, col=1)\n            \n            # Add Violin plot\n            fig.add_trace(go.Violin(y=df_content['X247'], name='Violin plot', box_visible=True, meanline_visible=True, points='outliers'), row=1, col=2)\n            \n            # Update layout for better appearance\n            fig.update_layout(\n                title_text=\"Distribution of X247\", \n                showlegend=False)\n            \n            # Update x-axis label for histogram\n            fig.update_xaxes(title_text=\"X247\", row=1, col=1)\n            \n            # Update y-axis label for boxplot (optional, as it's the same as the x-axis for histogram)\n            fig.update_yaxes(title_text=\"X247\", row=1, col=2)\n            fig.show()\n\n            #Examining skewness through Q-Q plot\n            sm.qqplot(df_content['X247'], line='s')\n            plt.title('Q-Q plot for X247')\n            plt.show()\n\n        print(f\"The columns with left skewness: {left_skewed_col}\")\n        print(f\"The columns with right skewness: {right_skewed_col}\")\n\n        \n# Create an instance of the class\ndata_prop_instance = data_properties(dataframes)\n\n# Run the dataset_description method\ndataset_prop=data_prop_instance.dataset_description()\ndataset_prop\n\n# Run the left skewness test method\ndataset_skewness=data_prop_instance.skewness_test()\ndataset_skewness\n\n#Inference:\n#1.Found no missing values in any column of train and test datasets.\n#2.There is a presence of right skewness in some columns eg:- bid_qty and ask_qty (as listed in output) as mean is greater than the median.\n#3.There is a presence of left skewness in some columns (as listed in output) eg:- X9, X17 etc.) as median is greater than the mean.\n#4.Two kind of outliers found - one is the natural variation in values and other is unusual extremely high value (eg:- bid qty of 1114.93 when 75th percentile value is 13.08)\n#5.Hence, these extreme values need to be clipped with IQR clipping technique.","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-30T23:34:47.413221Z","iopub.execute_input":"2025-11-30T23:34:47.413731Z","iopub.status.idle":"2025-11-30T23:35:44.610979Z","shell.execute_reply.started":"2025-11-30T23:34:47.413695Z","shell.execute_reply":"2025-11-30T23:35:44.609970Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**6.2 TRAIN TEST SPLIT ON ORIGINAL DATASET**","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\n#Class for train test splitting\n\nclass DataSplitter():\n    #constructor\n    def __init__(self, x_dataset, y_dataset, test_size=None, random_state=None, shuffle=None):\n        # The __init__ should use the argument passed to it\n        self.x_dataset=x_dataset\n        self.y_dataset=y_dataset\n        self.test_size=test_size\n        self.random_state=random_state\n        self.shuffle=shuffle\n        \n    # Function to create train test set\n    def train_test_set(self):\n        x_train, x_test, y_train, y_test = train_test_split(self.x_dataset, self.y_dataset, test_size=self.test_size, random_state=self.random_state, shuffle=self.shuffle)\n        print('Train Test Split Created')\n        return x_train, x_test, y_train, y_test\n\nDataSplitter_instance=DataSplitter(x_train_DRW,y_train_DRW,test_size=0.3, random_state=42, shuffle=False)\n\n#Train Test split on train.parquet dataset\nx_train_original, x_test_original, y_train_original, y_test_original=DataSplitter_instance.train_test_set()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-30T23:36:54.377347Z","iopub.execute_input":"2025-11-30T23:36:54.377643Z","iopub.status.idle":"2025-11-30T23:36:55.457922Z","shell.execute_reply.started":"2025-11-30T23:36:54.377622Z","shell.execute_reply":"2025-11-30T23:36:55.457049Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id = Section6.2></a>\n**6.2 HANDLING OUTLIERS**","metadata":{}},{"cell_type":"code","source":"#Class for 1)Treating extreme outliers with IQR - upper and lower clipping technique and applying power transformation, 2)Feature Scaling and \n## 3)Memory Optimization\n\nclass outlier_treatement():\n    #constructor\n    def __init__(self, dataframe, scaling_fit_dataset, threshold_for_std_dev):\n        # The __init__ should use the argument passed to it\n        self.dataframe=dataframe\n        self.scaling_fit_dataset=scaling_fit_dataset\n        self.scaler = StandardScaler()\n        self.pt = PowerTransformer(method='yeo-johnson', standardize=False)  # Initialize PowerTransformer with Yeo-Johnson method for positive/negative skewness post extreme values removal\n        self.data_dict={'self.dataframe':self.dataframe,'self.scaling_fit_dataset':self.scaling_fit_dataset}\n        self.skewness_threshold = 1.0\n        self.kurtosis_threshold = 1.0\n        self.skewed_cols=[]\n        self.col_unique_one = []\n        self.skewed_col_final = []\n        self.dataframe_transformed = None\n        self.scaling_fit_dataset_transformed = None\n        self.zero_std_dev_col=[]\n        self.threshold_for_std_dev = threshold_for_std_dev\n        self.final_cols_scaling=[]\n        self.final_cols_scaling_df=pd.DataFrame()\n        \n    # Function to handle extreme outliers (erroneous values, inputs etc) through IQR - Upper and Lower bound technique\n    def outlier_handling(self):\n        for df_name, df_object in self.data_dict.items():\n            for col in df_object.columns:\n                upper_bound=self.scaling_fit_dataset[col].quantile(0.95)\n                lower_bound=self.scaling_fit_dataset[col].quantile(0.05)\n                df_object[col] = df_object[col].clip(lower=lower_bound, upper=upper_bound)  \n        \n        #Check if the skewness still persists post outlier treatment and check for distribution normality\n        for col in self.scaling_fit_dataset.columns:\n            col_skewness= self.scaling_fit_dataset[col].skew()\n            col_kurtosis=kurtosis(self.scaling_fit_dataset[col], fisher=True)\n            \n            if (abs(col_skewness)>self.skewness_threshold) or (abs(col_kurtosis)>self.kurtosis_threshold): #if the column still skewed (+ve) and kurtosis (+ve) post outlier treatment\n                self.skewed_cols.append(col)\n\n        # Creating copies of passed datasets\n        self.dataframe_transformed = self.dataframe.copy()\n        self.scaling_fit_dataset_transformed = self.scaling_fit_dataset.copy()\n\n        # Transformation step for Train and Test dataset:\n        # Only apply power transformation if skewed columns are found\n        if self.skewed_cols:\n            # Step 1: Identify skewed columns with ≤ 1 unique value (zero or near-zero variance)\n            for col in self.skewed_cols:\n                if self.scaling_fit_dataset[col].nunique() <= 1 and self.scaling_fit_dataset[col].std() < 1e-6:\n                    self.col_unique_one.append(col)\n            print(f'Columns with ≤1 unique value (excluded from power transform): {self.col_unique_one}')\n\n            # Step 2: Filter final skewed columns with enough variance\n            self.skewed_col_final = [col for col in self.skewed_cols if col not in self.col_unique_one]\n            print(f'Columns with >1 unique value (included in power transform): {self.skewed_col_final}')\n            \n            if self.skewed_col_final:\n                # Fit transformer only on valid columns\n                self.pt.fit(self.scaling_fit_dataset[self.skewed_col_final])\n\n                # Apply transform\n                self.dataframe_transformed[self.skewed_col_final] = self.pt.transform(self.dataframe_transformed[self.skewed_col_final])\n                self.scaling_fit_dataset_transformed[self.skewed_col_final] = self.pt.transform(self.scaling_fit_dataset_transformed[self.skewed_col_final])\n\n                # Cap inf/very large values after power transformation \n                for df_t in [self.dataframe_transformed, self.scaling_fit_dataset_transformed]:\n                    for col in self.skewed_col_final: # Only apply to columns that were power-transformed\n                        max_finite_val = np.finfo(np.float32).max / 2\n                        min_finite_val = np.finfo(np.float32).min / 2 \n                        # Replace actual infinities\n                        df_t[col] = df_t[col].replace([np.inf, -np.inf], [max_finite_val, min_finite_val])\n                        # Clip any remaining values that are still too large/small,\n                        # even if not strictly 'inf' (e.g., 1.7e308)\n                        df_t[col] = df_t[col].clip(lower=min_finite_val, upper=max_finite_val)\n\n                        # Clipping within 85th/15th percentile for max and min finite values as x_test_original has extreme value at 90th/10th percentile too\n                        upper_bound= self.scaling_fit_dataset_transformed[col].quantile(0.90)\n                        lower_bound= self.scaling_fit_dataset_transformed[col].quantile(0.10)\n                        df_t[col] = df_t[col].clip(lower=lower_bound, upper=upper_bound)\n            \n            else:\n                print(\"No valid skewed columns to transform after excluding low-variance ones.\")\n\n        else:\n            print(\"No skewed columns found for power transformation.\")\n\n        return self.dataframe_transformed, self.scaling_fit_dataset_transformed\n\n    # Function to scale features as per standard scaler\n    def feature_standard_scaling(self):\n        self.outlier_handling()\n        \n        # Creating copies of passed datasets\n        self.dataframe_scaled = self.dataframe_transformed.copy() \n        self.scaling_fit_dataset_scaled = self.scaling_fit_dataset_transformed.copy() \n\n        # Check for columns standard deviation in train dataset\n        for col in self.scaling_fit_dataset_scaled.columns:\n            col_std_dev = self.scaling_fit_dataset_scaled[col].std()\n            if col_std_dev < self.threshold_for_std_dev:  #if std deviation falls below threshold, they would possibly be close to zero and \n                                                          #thus create close to infinite values when applied standard scaling\n                self.zero_std_dev_col.append(col)\n            else:\n                self.final_cols_scaling.append(col)       #Final cols with non-zero std deviation\n        \n        print(f'The features with close to zero std. devaiation: {self.zero_std_dev_col}')\n        print(f'The final features to be scaled: {self.final_cols_scaling}')\n            \n\n        if self.final_cols_scaling:\n            #Fitting standard scaler on train data\n            self.scaler.fit(self.scaling_fit_dataset_transformed[self.final_cols_scaling])\n\n            #Transforming Train and Test dataset through scaler\n            self.scaling_fit_dataset_scaled[self.final_cols_scaling] = self.scaler.transform(self.scaling_fit_dataset_scaled[self.final_cols_scaling])\n            self.dataframe_scaled[self.final_cols_scaling] = self.scaler.transform(self.dataframe_scaled[self.final_cols_scaling])\n\n            #Retaining only above zero std dev features\n            self.scaling_fit_dataset_scaled = self.scaling_fit_dataset_scaled[self.final_cols_scaling]\n            self.dataframe_scaled = self.dataframe_scaled[self.final_cols_scaling]\n            \n        return self.dataframe_scaled\n\n    # Function for memory optimization - converting 64 bits to 32 bits\n    def memory_optimize(self):\n        self.feature_standard_scaling()\n        column_dtypes = self.dataframe_scaled.dtypes\n        for col in self.dataframe_scaled.columns:\n            current_dtype = column_dtypes[col] # Access the dtype from the Series of dtypes\n    \n            if str(current_dtype) != 'float32':\n                self.dataframe_scaled[col]=self.dataframe_scaled[col].astype('float32') #Converting to float32 for RAM optimization\n\n            gc.collect()\n\n        #Treating duplicates columns if any\n        self.dataframe_scaled = self.dataframe_scaled.loc[:, ~self.dataframe_scaled.columns.duplicated()]\n        \n        # Clipping again if ~infinite values exist\n        self.dataframe_scaled.replace([np.inf, -np.inf], [np.finfo(np.float32).max, np.finfo(np.float32).min], inplace=True)\n\n        return self.dataframe_scaled\n\nprocessed_dataframes_final = {}\n\n# outlier_treatement class methods call on x_train_original\nx_processor = outlier_treatement(dataframe=x_train_original, scaling_fit_dataset=x_train_original, threshold_for_std_dev=1e-9)\nx_train_original_processed = x_processor.memory_optimize()\nprocessed_dataframes_final['x_train_original_processed'] = x_train_original_processed\n\n# outlier_treatement class methods call on x_test_original\nx_processor.dataframe = x_test_original.copy()\nx_test_original_processed = x_processor.memory_optimize()\nprocessed_dataframes_final['x_test_original_processed'] = x_test_original_processed\n\n# outlier_treatement class methods call on y_train_original\ny_processor = outlier_treatement(dataframe=y_train_original, scaling_fit_dataset=y_train_original, threshold_for_std_dev=1e-9)\ny_train_original_processed = y_processor.memory_optimize()\nprocessed_dataframes_final['y_train_original_processed'] = y_train_original_processed\n\n# outlier_treatement class methods call on y_test_original\ny_processor.dataframe = y_test_original.copy()\ny_test_original_processed = y_processor.memory_optimize()\nprocessed_dataframes_final['y_test_original_processed'] = y_test_original_processed\n\n# Storing results in new variables\nx_train_original_processed=processed_dataframes_final['x_train_original_processed']\nx_test_original_processed=processed_dataframes_final['x_test_original_processed']\ny_train_original_processed=processed_dataframes_final['y_train_original_processed']\ny_test_original_processed=processed_dataframes_final['y_test_original_processed']\n\nprint('All dtypes converted to float32')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-30T23:37:06.971670Z","iopub.execute_input":"2025-11-30T23:37:06.972258Z","iopub.status.idle":"2025-11-30T23:49:42.487728Z","shell.execute_reply.started":"2025-11-30T23:37:06.972232Z","shell.execute_reply":"2025-11-30T23:49:42.487023Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#ANALYSING THE ACTUAL DRIFT BETWEEN X_TRAIN AND X_TEST POST TRAIN TEST SPLIT DONE WITH SHUFFLE=FALSE\ndef analyze_drift(x_train, x_test):\n    train_stats = x_train.describe().T[['mean', 'std']]\n    test_stats = x_test.describe().T[['mean', 'std']]\n    drift_df = train_stats.join(test_stats, lsuffix='_train', rsuffix='_test')\n    drift_df['mean_diff'] = (drift_df['mean_test'] - drift_df['mean_train']).abs()\n    drift_df['std_ratio'] = drift_df['std_test'] / drift_df['std_train']\n    return drift_df.sort_values(by='mean_diff', ascending=False)\n\ndrift_report = analyze_drift(x_train_original_processed, x_test_original_processed)\ndrift_report","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-30T23:51:19.318878Z","iopub.execute_input":"2025-11-30T23:51:19.319781Z","iopub.status.idle":"2025-11-30T23:51:24.999322Z","shell.execute_reply.started":"2025-11-30T23:51:19.319744Z","shell.execute_reply":"2025-11-30T23:51:24.998681Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id = Section7></a>\n**7. INITIAL FEATURE EXPLORATION - REDUCING DIMENSIONALITY**\n\n1.RELATIVE IMPORTANCE OF FEATURES - PEARSON'S CORRELATION AND SPEARMAN'S CORRELATION                 \n2.RELATIVE IMPORTANCE OF FEATURES - MUTUAL INFORMATION SCORE                          \n3.RELATIVE IMPORTANCE OF FEATURES - DECISION TREE AND RANDOM FOREST APPROACH ","metadata":{}},{"cell_type":"markdown","source":"<a id = Section7.1></a>\n7.1 RELATIVE IMPORTANCE OF FEATURES - PEARSON'S CORRELATION AND SPEARMAN'S CORRELATION","metadata":{}},{"cell_type":"code","source":"#Class for initial data exploration through pair plots, pearson/spearman's correlation through heatmaps\n\nclass initial_data_exploration():\n    #constructor\n    def __init__(self,x_train_df,y_train_df,label_col=None,num_features_for_plot=None, sample_rows_for_plot=None):\n        # The __init__ should use the argument passed to it\n        self.x_train_df = x_train_df\n        self.y_train_df = y_train_df\n        self.label_col = label_col\n        self.num_features_for_plot = num_features_for_plot\n        self.sample_rows_for_plot = sample_rows_for_plot\n        features_to_plot = self.x_train_df.iloc[:, :self.num_features_for_plot] # Assuming 0-indexed features\n        self.merged_df = pd.concat([features_to_plot, self.y_train_df[[self.label_col]]], axis=1)\n        print(f\"Merged DataFrame for plots created with shape: {self.merged_df.shape}\")\n            \n    # Function for quick visual analysis of linear/non-linear relationships through pairplots\n    def plot_initial_relationship(self):\n        \n        # Generate the pair plot for a sample of 50 records for quick observation\n        sns.pairplot(self.merged_df.iloc[0:self.sample_rows_for_plot,:])\n        \n        # Display the plot\n        plt.tight_layout()\n        plt.show()\n        \n        \n    # Function for pearson correlation test\n    def pearson_corr(self):\n        #Pearson Correlation Examination\n        corr_matrix_pearson=self.merged_df.corr(method='pearson')\n        # Create the heatmap\n        plt.figure(figsize=(12, 10)) # Adjust figure size as needed\n        sns.heatmap(corr_matrix_pearson, annot=True, cmap='coolwarm', fmt=\".2f\")\n        \n        # Set the title of the heatmap\n        plt.title('Pearson Correlation Heatmap')\n        \n        # Display the plot\n        plt.tight_layout()\n        plt.show()\n        \n\n    # Function for spearman's correlation test\n    def spearman_corr(self):\n        #Spearman's Correlation Examination\n        corr_matrix_spearman=self.merged_df.corr(method='spearman')\n        # Create the heatmap\n        plt.figure(figsize=(12, 10)) # Adjust figure size as needed\n        sns.heatmap(corr_matrix_spearman, annot=True, cmap='coolwarm', fmt=\".2f\")\n        \n        # Set the title of the heatmap\n        plt.title(\"Spearman's Correlation Heatmap\")\n        \n        # Display the plot\n        plt.tight_layout()\n        plt.show()\n        \n    \ninitial_data_explor_instance=initial_data_exploration(x_train_original_processed,y_train_original_processed,label_col='label',num_features_for_plot=10,sample_rows_for_plot=50)\ninitial_data_explor_instance.plot_initial_relationship()\ninitial_data_explor_instance.pearson_corr()\ninitial_data_explor_instance.spearman_corr()\n\n# Inference:\n#1.Some predictors like bid_qty,as_qty,buy_qty etc. have non-linear and complex relationship with label variable.\n#2.Some predictors like X1,X2,X3 etc. have possible linearity in their relationship with label (can be tested by pearson's correlation)\n#3.Also, there's high linearity amonst some predictors like between x1 and x2,x3,x4 each indicating multicollinear features (can be tested by pearson's correlation)\n#4.Even though features like X1,X2 etc. percieved to have a linear relationships with label, their correlation percentage is way less - 3 to 4 percent.\n#5.This indicates complex relationships of these features with label. Hence, there's need to run non-linear test to examine feature importance with label.\n#6.But there's presence of high collinearity between predictors like X2 and X3, X3 and X4 (corr. percent atleast 90) indicating one feature of these\n##multicollinear feature combinations should be retained and others should be discarded on the basis of pearson's correlation threshold.","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-30T23:51:28.901228Z","iopub.execute_input":"2025-11-30T23:51:28.901501Z","iopub.status.idle":"2025-11-30T23:51:54.769579Z","shell.execute_reply.started":"2025-11-30T23:51:28.901481Z","shell.execute_reply":"2025-11-30T23:51:54.768748Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 7.1 INITIAL PRUNING ON ORIGINAL FEATURES - MODEL AGNOSTIC TECHNIQUES","metadata":{}},{"cell_type":"markdown","source":"7.1.1 INITIAL PRUNING - REMOVING MULTI COLLINEARITY \n\nUsing pearson's correlation among predictors, those predictors having correlation coefficient (0.80) with each other can be the potential \ncandidates for discarding. This threshold can be decided basis crypto SME discussion. In our case, we have made a judgment of threshold as 80 percent.","metadata":{}},{"cell_type":"code","source":"#Class for removing multicollinearity\nclass remove_multicollinearity():\n    #constructor\n    def __init__(self,x_train_df,y_train_df,threshold=None,rows_to_include=None,label=None,method=None):\n        self.x_train_df = x_train_df\n        self.y_train_df = y_train_df\n        self.threshold = threshold\n        self.rows_to_include = rows_to_include\n        self.label = label\n        self.method = method\n        self.correlation_matrix_predictors = self.x_train_df.tail(rows_to_include).corr(method=self.method)\n\n        self.mi_scores = mutual_info_regression(self.x_train_df.tail(self.rows_to_include), self.y_train_df[self.label].tail(self.rows_to_include))\n        \n        # Create a pandas dataframe to store feature names and their MI scores\n        self.mi_df = pd.DataFrame({'MI Score':self.mi_scores, 'Feature':self.x_train_df.columns}).sort_values(by='MI Score',ascending=False)\n\n        # Create a dictionary for faster MI score lookups\n        self.mi_scores_dict = self.mi_df.set_index('Feature')['MI Score'].to_dict()\n\n\n    #Function for removing multicollinearity\n    def multicollinear_treatment(self):\n        highly_correlated_pairs = []\n\n        # Iterate through the matrix\n        for i in range(len(self.correlation_matrix_predictors.columns)):\n            for j in range(i + 1, len(self.correlation_matrix_predictors.columns)):\n                col1 = self.correlation_matrix_predictors.columns[i]\n                col2 = self.correlation_matrix_predictors.columns[j]\n                correlation_value = self.correlation_matrix_predictors.iloc[i, j]\n        \n                # Check if the absolute correlation is above the threshold\n                if abs(correlation_value) > self.threshold:\n                    highly_correlated_pairs.append((col1, col2, correlation_value))\n        \n        # Print the highly correlated pairs\n        print(f\"Highly correlated predictor pairs (absolute correlation > {self.threshold}):\")\n        if highly_correlated_pairs:\n            for col1, col2, corr_value in highly_correlated_pairs:\n                print(f\"  - {col1} and {col2}: {corr_value:.4f}\")\n        else:\n            print(\"No highly correlated predictor pairs found above the specified threshold.\")\n        \n        # Identify features to drop (On the basis of MI scores)\n        features_to_drop = set()\n        \n        # mi_scores_dict for faster lookups\n        for col1, col2, _ in highly_correlated_pairs:\n            # Use .get() with a default of -np.inf for safety in case a feature is somehow missing\n            mi_col1 = self.mi_scores_dict.get(col1, -np.inf)\n            mi_col2 = self.mi_scores_dict.get(col2, -np.inf)\n\n            if mi_col1 > mi_col2:\n                features_to_drop.add(col2)\n            else:\n                features_to_drop.add(col1)\n        \n        # Drop the identified features from x_train_DRW\n        x_train_reduced = self.x_train_df.drop(columns=list(features_to_drop), errors='ignore')\n        \n        # Print the count of unique dropped features\n        print(f\"\\nNumber of unique features dropped due to high collinearity: {len(set(features_to_drop))}\")\n        print(f\"Shape of x_train reduced post Multicollinearity Treatment: {x_train_reduced.shape}\")   #Dataframe with removed multicollinearity\n        \n        return x_train_reduced\n\nremove_multicollinearity_instance=remove_multicollinearity(x_train_original_processed,y_train_original_processed,threshold=0.80,rows_to_include=50000,label='label',method='pearson')\nx_train_reduced_post_pearson=remove_multicollinearity_instance.multicollinear_treatment()\nx_test_reduced_post_pearson=x_test_original_processed[x_train_reduced_post_pearson.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-30T23:52:24.764859Z","iopub.execute_input":"2025-11-30T23:52:24.765135Z","iopub.status.idle":"2025-11-30T23:57:21.228794Z","shell.execute_reply.started":"2025-11-30T23:52:24.765116Z","shell.execute_reply":"2025-11-30T23:57:21.227936Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Model Agnositic technique like Mutual Information Score can be leveraged to further reduce dimensionality.","metadata":{}},{"cell_type":"markdown","source":"<a id = Section7.2></a>\n**7.1.2 RELATIVE IMPORTANCE OF FEATURES - MUTUAL INFORMATION SCORE**","metadata":{}},{"cell_type":"code","source":"# Class for analyzing and visualizing Mutual Information scores on x_train_reduced_post_pearson\nclass MIFeatureAnalyzer:\n    # Constructor\n    def __init__(self, x_train_df, y_train_df, label_col=None, rows_to_include=None, threshold=None):\n        self.x_train_df = x_train_df\n        self.y_train_df = y_train_df\n        self.rows_to_include = rows_to_include\n        self.label = label_col\n        self.threshold=threshold\n        self.mi_scores = mutual_info_regression(self.x_train_df.tail(self.rows_to_include), self.y_train_df[self.label].tail(self.rows_to_include))\n\n        # Create a pandas dataframe to store feature names and their MI scores\n        self.mi_df = pd.DataFrame({'MI Score':self.mi_scores, 'Feature':self.x_train_df.columns}).sort_values(by='MI Score',ascending=False)\n        \n        # Below Threshold features list\n        self.below_threshold_features=list(self.mi_df[self.mi_df['MI Score']<self.threshold]['Feature'].values)\n\n        # Above Threshold features df\n        self.above_threshold_features_mi=self.mi_df[self.mi_df['MI Score']>=self.threshold]\n\n    def MI_plotting(self):\n        # --- Visualization ---\n        # Top 50 Features basis MI score\n        fig, my_ax = plt.subplots(nrows=1, ncols=2,figsize=(16, 5))\n        sns.barplot(y='MI Score', x='Feature', data=self.mi_df.head(50),ax=my_ax[0])\n        my_ax[0].set_title('Top 50 Features - Mutual Information Score')\n        my_ax[0].set_xlabel('Importance - Mutual Information Score')\n        my_ax[0].set_ylabel('Feature')\n        my_ax[0].tick_params(axis='x', rotation=90)\n        \n        # Create a line plot of the ranked mutual information scores\n        sns.lineplot(x=range(len(self.mi_df)), y=self.mi_df['MI Score'],ax=my_ax[1])\n        my_ax[1].set_title('Mutual Information Score in Descending Order')\n        my_ax[1].set_ylabel('Importance - Mutual Information Score')\n        my_ax[1].set_xlabel('Feature Rank')\n        \n        plt.grid(linestyle=':')\n        plt.tight_layout()\n        plt.show()\n\n    def above_threshold_features_MI(self):\n        # Retaining features with importance above threshold - MI Scores\n        print(f\"Features Importances List - MI score: {self.mi_df.shape}\")\n        print(f\"Above Threshold Features Importances  List - MI score: {self.above_threshold_features_mi.shape}\")\n        above_threshold_features_MI=self.above_threshold_features_mi\n        return above_threshold_features_MI\n\n    def MI_feature_drop(self):\n        x_train_reduced_post_pearson_MI=self.x_train_df.copy()\n        for col in self.below_threshold_features:\n            x_train_reduced_post_pearson_MI.drop([col],axis=1,inplace=True)\n        \n        print(f\"Shape of dataset reduced through MI score: {x_train_reduced_post_pearson_MI.shape}\")\n\n        return x_train_reduced_post_pearson_MI\n\nremove_post_MI_instance=MIFeatureAnalyzer(x_train_reduced_post_pearson, y_train_original_processed, label_col='label', rows_to_include=50000, threshold=0.05)\nremove_post_MI_instance.MI_plotting()\n#Inference:\n#1.The line chart depicts the elbow point at somewhere around 0.05 MI score, hence we can keep the threshold at 0.05.\nx_train_reduced_post_pearson_MI=remove_post_MI_instance.MI_feature_drop()\nx_test_reduced_post_pearson_MI=x_test_reduced_post_pearson[x_train_reduced_post_pearson_MI.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-30T23:57:48.030916Z","iopub.execute_input":"2025-11-30T23:57:48.031533Z","iopub.status.idle":"2025-11-30T23:58:29.685252Z","shell.execute_reply.started":"2025-11-30T23:57:48.031507Z","shell.execute_reply":"2025-11-30T23:58:29.684561Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Post this initial pruning step, two workflows are created:\n\n**Workflow 1** - Get the most important original features post feature selection pipeline from the 51 non-redundant original features.       \n**Workflow 2** - Engineer the top features from 51 non-redundant original features and then get the most important features post feature selection pipeline. (As engineering the 120+ features will be computationally expensive)","metadata":{}},{"cell_type":"markdown","source":"**Workflow 1:**\nPost leveraging model agnostic feature selection techniques (Multicollinaerity treatment, MI score feature selection), two kinds of feature selection strategies can be used: \n\n**1. Linear Based - Lasso/Ridge/Elastic Regression, Random Forest/XG/LGBM/SVR**   \n**2. Non-Linear Based - ANN/RNN/LSTM**","metadata":{}},{"cell_type":"markdown","source":"# 7.2 WORKFLOW 1 (POST INITIAL PRUNING) - LINEAR FEATURE SELECTION PIPELINE - ORIGINAL NON-REDUNDANT FEATURES. ","metadata":{}},{"cell_type":"markdown","source":"**7.2.1 RELATIVE IMPORTANCE OF FEATURES - LASSO - L1 REGRESSION APPROACH**","metadata":{}},{"cell_type":"code","source":"# Class for analyzing feature importance through Lasso - L1 Regression\n\nclass lasso_FeatureAnalyzer:\n    # Constructor\n    def __init__(self, x_train_df, y_train_df, cv=None, random_state=None, max_iter=None):  \n        self.x_train_df = x_train_df\n        self.y_train_df = y_train_df\n        self.cv=cv\n        self.random_state=random_state\n        self.max_iter=max_iter\n        self.custom_alphas = np.logspace(-3, 0, 100)\n        self.lasso_cv_model = LassoCV(cv=self.cv,            \n                                      random_state=self.random_state,\n                                      max_iter=self.max_iter,\n                                      alphas=self.custom_alphas\n                                     ).fit(self.x_train_df,self.y_train_df)\n\n        print(f\"Optimal alpha found by LassoCV: {self.lasso_cv_model.alpha_}\")\n\n        # Get coefficients from the best Lasso model\n        self.lasso_coefficients_df = pd.DataFrame({'Feature Coefficient':self.lasso_cv_model.coef_, 'Feature':self.x_train_df.columns})\n\n        # Identify selected features (non-zero coefficients)\n        self.selected_features_lasso_df = self.lasso_coefficients_df[self.lasso_coefficients_df['Feature Coefficient'] != 0].sort_values(by='Feature Coefficient', key=abs, ascending=False)\n\n    def lasso_plotting(self):\n        # --- Visualization ---\n        # Top 50 Important Features\n        fig, my_ax = plt.subplots(nrows=1, ncols=2,figsize=(16, 7))\n        sns.barplot(y='Feature Coefficient', x='Feature', data=self.selected_features_lasso_df,ax=my_ax[0])\n        my_ax[0].set_title('Top Features Basis Lasso - L1 Regression Feature Importances')\n        my_ax[0].set_ylabel('Feature Coefficient Value')\n        my_ax[0].set_xlabel('Feature')\n        my_ax[0].tick_params(axis='x', rotation=90)\n        \n        # Create a line plot of the ranked feature importances\n        sns.lineplot(x=range(len(self.selected_features_lasso_df)), y=self.selected_features_lasso_df['Feature Coefficient'],ax=my_ax[1])\n        my_ax[1].set_title('Feature Importance Basis Lasso - L1 Regression in Descending Order')\n        my_ax[1].set_ylabel('Feature Coefficient Value')\n        my_ax[1].set_xlabel('Feature Rank')\n        \n        plt.tight_layout()\n        plt.show()\n\n    def lasso_feature_drop(self):\n        self.x_train_reduced_post_pearson_MI_lasso=self.x_train_df.copy()\n        self.x_train_reduced_post_pearson_MI_lasso= self.x_train_reduced_post_pearson_MI_lasso[self.selected_features_lasso_df['Feature']]\n\n        print(f\"Shape of dataset reduced through Pearson, MI score & Lasso: {self.x_train_reduced_post_pearson_MI_lasso.shape}\")\n\n        return self.x_train_reduced_post_pearson_MI_lasso,self.selected_features_lasso_df,self.lasso_cv_model\n\nlasso_FeatureAnalyzer_instance=lasso_FeatureAnalyzer(x_train_reduced_post_pearson_MI, y_train_original_processed, cv=5, random_state=42, max_iter=10000)\nlasso_FeatureAnalyzer_instance.lasso_plotting()\nx_train_reduced_post_pearson_MI_lasso, Imp_features_lasso,model_lasso=lasso_FeatureAnalyzer_instance.lasso_feature_drop()\nx_test_reduced_post_pearson_MI_lasso=x_test_reduced_post_pearson_MI[x_train_reduced_post_pearson_MI_lasso.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-30T23:58:47.654737Z","iopub.execute_input":"2025-11-30T23:58:47.655016Z","iopub.status.idle":"2025-11-30T23:58:49.909477Z","shell.execute_reply.started":"2025-11-30T23:58:47.654998Z","shell.execute_reply":"2025-11-30T23:58:49.908754Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"These 9 features are considered important by linear (Lasso L-1) model. This suggests they have a fundamental, linear relationship (though very less influence as pearson of 0.20) with the target variable.","metadata":{}},{"cell_type":"code","source":"#Training the lasso regressor on top features as per Lasso\nlasso_FeatureAnalyzer_instance_two=lasso_FeatureAnalyzer(x_train_reduced_post_pearson_MI, y_train_original_processed, cv=5, random_state=42, max_iter=10000)\nx_train_reduced_post_pearson_MI_lasso_two, Imp_features_lasso_two, model_lasso_two=lasso_FeatureAnalyzer_instance_two.lasso_feature_drop()\n\n#Predictions from lasso \nlasso_pred=model_lasso_two.predict(x_test_reduced_post_pearson_MI)\nlasso_pred_df=pd.DataFrame(lasso_pred,columns=['pred_label from Lasso top features'])\n\nm1 = lasso_pred_df['pred_label from Lasso top features'].round(3)\nm2 = y_test_original_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:\ncorr_pandas = m1.corr(m2)\nprint(f\"Pearson Correlation for Lasso from Lasso's best features - original: {corr_pandas}\")\n\n#Accuracy as per RMSE\nrmse_lasso = np.sqrt(mean_squared_error(m1,m2))\nprint(f\"Root Mean Squared Error for Lasso from Lasso's best features - original: {rmse_lasso:.4f}\")\n\n#Accuracy as per R2\nr2_lasso = r2_score(m1,m2)\nprint(f\"R2 score for Lasso from Lasso's best features - original: {r2_lasso:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-30T23:58:57.991036Z","iopub.execute_input":"2025-11-30T23:58:57.991317Z","iopub.status.idle":"2025-11-30T23:58:59.623133Z","shell.execute_reply.started":"2025-11-30T23:58:57.991296Z","shell.execute_reply":"2025-11-30T23:58:59.621632Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.2.2 RELATIVE IMPORTANCE OF FEATURES - RIDGE - L2 REGRESSION APPROACH**","metadata":{}},{"cell_type":"code","source":"# Class for analyzing feature importance through Ridge - L2 Regression\n\nclass ridge_FeatureAnalyzer:\n    # Constructor\n    def __init__(self, x_train_df, y_train_df, cv=None, random_state=None, max_iter=None):  \n        self.x_train_df = x_train_df\n        self.y_train_df = y_train_df\n        self.cv=cv\n        self.random_state=random_state\n        self.max_iter=max_iter\n        self.custom_alphas = np.logspace(-3, 0, 100)\n        self.ridge_cv_model = RidgeCV(cv=self.cv,            \n                                      alphas=self.custom_alphas\n                                     ).fit(self.x_train_df,self.y_train_df)\n\n        print(f\"Optimal alpha found by RidgeCV: {self.ridge_cv_model.alpha_}\")\n        #print(self.ridge_cv_model.coef_)\n        #print(self.x_train_df.columns)\n        # Get coefficients from the best Lasso model\n        self.ridge_coefficients_df = pd.DataFrame({'Feature Coefficient':self.ridge_cv_model.coef_.reshape(-1), 'Feature':self.x_train_df.columns})\n\n        # Identify selected features (non-zero coefficients)\n        self.selected_features_ridge_df = self.ridge_coefficients_df[self.ridge_coefficients_df['Feature Coefficient'] != 0].sort_values(by='Feature Coefficient', key=abs, ascending=False)\n\n    def ridge_plotting(self):\n        # --- Visualization ---\n        # Top 50 Important Features\n        fig, my_ax = plt.subplots(nrows=1, ncols=2,figsize=(16, 7))\n        sns.barplot(y='Feature Coefficient', x='Feature', data=self.selected_features_ridge_df,ax=my_ax[0])\n        my_ax[0].set_title('Top Features Basis Ridge - L2 Regression Feature Importances')\n        my_ax[0].set_ylabel('Feature Coefficient Value')\n        my_ax[0].set_xlabel('Feature')\n        my_ax[0].tick_params(axis='x', rotation=90)\n        \n        # Create a line plot of the ranked feature importances\n        sns.lineplot(x=range(len(self.selected_features_ridge_df)), y=self.selected_features_ridge_df['Feature Coefficient'],ax=my_ax[1])\n        my_ax[1].set_title('Feature Importance Basis Ridge - L2 Regression in Descending Order')\n        my_ax[1].set_ylabel('Feature Coefficient Value')\n        my_ax[1].set_xlabel('Feature Rank')\n        \n        plt.tight_layout()\n        plt.show()\n\n    def ridge_feature_drop(self):\n        self.x_train_reduced_post_pearson_MI_ridge=self.x_train_df.copy()\n        self.x_train_reduced_post_pearson_MI_ridge= self.x_train_reduced_post_pearson_MI_ridge[self.selected_features_ridge_df['Feature']]\n\n        print(f\"Shape of dataset reduced through Pearson, MI score & Ridge: {self.x_train_reduced_post_pearson_MI_ridge.shape}\")\n\n        return self.x_train_reduced_post_pearson_MI_ridge,self.selected_features_ridge_df,self.ridge_cv_model\n\nridge_FeatureAnalyzer_instance=ridge_FeatureAnalyzer(x_train_reduced_post_pearson_MI, y_train_original_processed, cv=5)\nridge_FeatureAnalyzer_instance.ridge_plotting()\nx_train_reduced_post_pearson_MI_ridge, Imp_features_ridge,model_ridge=ridge_FeatureAnalyzer_instance.ridge_feature_drop()\nx_test_reduced_post_pearson_MI_ridge=x_test_reduced_post_pearson_MI[x_train_reduced_post_pearson_MI_ridge.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-30T23:59:25.505852Z","iopub.execute_input":"2025-11-30T23:59:25.506146Z","iopub.status.idle":"2025-12-01T00:00:14.265014Z","shell.execute_reply.started":"2025-11-30T23:59:25.506124Z","shell.execute_reply":"2025-12-01T00:00:14.264231Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Ridge model trained on Best Features basis Ridge coefficients\nridge_FeatureAnalyzer_instance_best=ridge_FeatureAnalyzer(x_train_reduced_post_pearson_MI_ridge, y_train_original_processed, cv=5)\nx_train_reduced_post_pearson_MI_ridge_best, Imp_features_ridge_best,model_ridge_best=ridge_FeatureAnalyzer_instance_best.ridge_feature_drop()\n\n#Predictions from Ridge \nridge_pred=model_ridge_best.predict(x_test_reduced_post_pearson_MI_ridge)\nridge_pred_df=pd.DataFrame(ridge_pred,columns=['pred_label from Ridge top features'])\n\nm1_ridge = ridge_pred_df['pred_label from Ridge top features'].round(3)\nm2_ridge = y_test_original_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:\ncorr_pandas_ridge = m1_ridge.corr(m2_ridge)\nprint(f\"Pearson Correlation for Ridge from Ridge's best features - original: {corr_pandas_ridge}\")\n\n#Accuracy as per RMSE\nrmse_ridge = np.sqrt(mean_squared_error(m1_ridge,m2_ridge))\nprint(f\"Root Mean Squared Error for Ridge from Ridge's best features - original: {rmse_ridge:.4f}\")\n\n#Accuracy as per R2\nr2_ridge = r2_score(m1_ridge,m2_ridge)\nprint(f\"R2 score for Ridge from Ridge's best features - original: {r2_ridge:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:00:28.649122Z","iopub.execute_input":"2025-12-01T00:00:28.649373Z","iopub.status.idle":"2025-12-01T00:01:15.358262Z","shell.execute_reply.started":"2025-12-01T00:00:28.649356Z","shell.execute_reply":"2025-12-01T00:01:15.357142Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.2.3 RELATIVE IMPORTANCE OF FEATURES - ELASTIC NET REGRESSION APPROACH**","metadata":{}},{"cell_type":"code","source":"# Class for analyzing feature importance through Elastic Net Regression\n\nclass elastic_FeatureAnalyzer:\n    # Constructor\n    def __init__(self, x_train_df, y_train_df, cv=None, max_iter=None, l1_ratio=None):  \n        self.x_train_df = x_train_df\n        self.y_train_df = y_train_df\n        self.cv=cv\n        self.max_iter=max_iter\n        self.l1_ratio=l1_ratio\n        self.custom_alphas = np.logspace(-3, 0, 100)\n        self.elastic_cv_model = ElasticNetCV(cv=self.cv,            \n                                      alphas=self.custom_alphas,\n                                      max_iter=self.max_iter,\n                                      l1_ratio=self.l1_ratio\n                                     ).fit(self.x_train_df,self.y_train_df)\n\n        print(f\"Optimal alpha found by ElasticCV: {self.elastic_cv_model.alpha_}\")\n        \n        # Get coefficients from the best Elastic model\n        self.elastic_coefficients_df = pd.DataFrame({'Feature Coefficient':self.elastic_cv_model.coef_.reshape(-1), 'Feature':self.x_train_df.columns})\n\n        # Identify selected features (non-zero coefficients)\n        self.selected_features_elastic_df = self.elastic_coefficients_df[self.elastic_coefficients_df['Feature Coefficient'] != 0].sort_values(by='Feature Coefficient', key=abs, ascending=False)\n\n    def elastic_plotting(self):\n        # --- Visualization ---\n        # Top 50 Important Features\n        fig, my_ax = plt.subplots(nrows=1, ncols=2,figsize=(16, 7))\n        sns.barplot(y='Feature Coefficient', x='Feature', data=self.selected_features_elastic_df,ax=my_ax[0])\n        my_ax[0].set_title('Top Features Basis Elastic Regression Feature Importances')\n        my_ax[0].set_ylabel('Feature Coefficient Value')\n        my_ax[0].set_xlabel('Feature')\n        my_ax[0].tick_params(axis='x', rotation=90)\n        \n        # Create a line plot of the ranked feature importances\n        sns.lineplot(x=range(len(self.selected_features_elastic_df)), y=self.selected_features_elastic_df['Feature Coefficient'],ax=my_ax[1])\n        my_ax[1].set_title('Feature Importance Basis Elastic Regression in Descending Order')\n        my_ax[1].set_ylabel('Feature Coefficient Value')\n        my_ax[1].set_xlabel('Feature Rank')\n        \n        plt.tight_layout()\n        plt.show()\n\n    def elastic_feature_drop(self):\n        self.x_train_reduced_post_pearson_MI_elastic=self.x_train_df.copy()\n        self.x_train_reduced_post_pearson_MI_elastic= self.x_train_reduced_post_pearson_MI_elastic[self.selected_features_elastic_df['Feature']]\n\n        print(f\"Shape of dataset reduced through Pearson, MI score & Elastic: {self.x_train_reduced_post_pearson_MI_elastic.shape}\")\n\n        return self.x_train_reduced_post_pearson_MI_elastic,self.selected_features_elastic_df,self.elastic_cv_model\n\nelastic_FeatureAnalyzer_instance=elastic_FeatureAnalyzer(x_train_reduced_post_pearson_MI, y_train_original_processed, cv=5, max_iter=1000, l1_ratio=0.5)\nelastic_FeatureAnalyzer_instance.elastic_plotting()\nx_train_reduced_post_pearson_MI_elastic, Imp_features_elastic,model_elastic=elastic_FeatureAnalyzer_instance.elastic_feature_drop()\nx_test_reduced_post_pearson_MI_elastic=x_test_reduced_post_pearson_MI[x_train_reduced_post_pearson_MI_elastic.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:01:34.173971Z","iopub.execute_input":"2025-12-01T00:01:34.174239Z","iopub.status.idle":"2025-12-01T00:01:36.464145Z","shell.execute_reply.started":"2025-12-01T00:01:34.174223Z","shell.execute_reply":"2025-12-01T00:01:36.463534Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Elastic Net model trained on Best Features basis Ridge coefficients\nelastic_FeatureAnalyzer_instance_best=elastic_FeatureAnalyzer(x_train_reduced_post_pearson_MI_elastic, y_train_original_processed, cv=5, max_iter=1000, l1_ratio=0.5)\nx_train_reduced_post_pearson_MI_elastic_best, Imp_features_elastic_best,model_elastic_best=elastic_FeatureAnalyzer_instance_best.elastic_feature_drop()\n\n#Predictions from Ridge \nelastic_pred=model_elastic_best.predict(x_test_reduced_post_pearson_MI_elastic)\nelastic_pred_df=pd.DataFrame(elastic_pred,columns=['pred_label from Elastic top features'])\n\nm1_elastic = elastic_pred_df['pred_label from Elastic top features'].round(3)\nm2_elastic = y_test_original_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:\ncorr_pandas_elastic = m1_elastic.corr(m2_elastic)\nprint(f\"Pearson Correlation for Elastic from Elastic's best features - original: {corr_pandas_elastic}\")\n\n#Accuracy as per RMSE\nrmse_elastic = np.sqrt(mean_squared_error(m1_elastic,m2_elastic))\nprint(f\"Root Mean Squared Error for elastic from elastic's best features - original: {rmse_elastic:.4f}\")\n\n#Accuracy as per R2\nr2_elastic = r2_score(m1_elastic, m2_elastic)\nprint(f\"R2 score for elastic from elastic's best features - original: {r2_elastic:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:01:45.870862Z","iopub.execute_input":"2025-12-01T00:01:45.871139Z","iopub.status.idle":"2025-12-01T00:01:46.787456Z","shell.execute_reply.started":"2025-12-01T00:01:45.871121Z","shell.execute_reply":"2025-12-01T00:01:46.785066Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.2.4 RELATIVE IMPORTANCE OF FEATURES - RANDOM FOREST**","metadata":{}},{"cell_type":"code","source":"# Class for analyzing feature importance through Random Forest Regressor\nclass RF_FeatureAnalyzer:\n    # Constructor\n    def __init__(self, x_train_df, y_train_df, rows_to_include=None,label_col=None, threshold=None, **rf_params):\n        self.x_train_df = x_train_df\n        self.y_train_df = y_train_df\n        self.rows_to_include = rows_to_include\n        self.label = label_col\n        self.threshold=threshold\n        self.rf_params = rf_params\n        \n        # Initialize the Random Forest Regressor model\n        self.rf_regressor = RandomForestRegressor(**self.rf_params)\n        \n        # Train the RF regressor\n        self.rf_regressor.fit(self.x_train_df.tail(self.rows_to_include), self.y_train_df.tail(self.rows_to_include)['label'])\n\n        # Get feature importances\n        self.importances_reg_rf = self.rf_regressor.feature_importances_\n        \n        # Create a DataFrame for better readability and sorting\n        self.feature_names_reg_rf = self.x_train_df.columns\n        self.importance_reg_rf = pd.DataFrame({\n            'Feature': self.feature_names_reg_rf,\n            'Importance': self.importances_reg_rf\n        }).sort_values(by='Importance',ascending=False)\n\n        # Above Threshold features df\n        self.above_threshold_features_rf=self.importance_reg_rf[self.importance_reg_rf['Importance']>=self.threshold]\n\n    def RF_plotting(self):\n        # --- Visualization ---\n        # Top 50 Important Features\n        fig, my_ax = plt.subplots(nrows=1, ncols=2,figsize=(16, 7))\n        sns.barplot(y='Importance', x='Feature', data=self.importance_reg_rf.head(50),ax=my_ax[0])\n        my_ax[0].set_title('Top 50 Features Basis Random Forest Regressor Feature Importances')\n        my_ax[0].set_ylabel('Importance')\n        my_ax[0].set_xlabel('Feature')\n        my_ax[0].tick_params(axis='x', rotation=90)\n        \n        # Create a line plot of the ranked feature importances\n        sns.lineplot(x=range(len(self.importance_reg_rf)), y=self.importance_reg_rf['Importance'],ax=my_ax[1])\n        my_ax[1].set_title('Feature Importance Basis RF Regressor in Descending Order')\n        my_ax[1].set_ylabel('Feature Importance')\n        my_ax[1].set_xlabel('Feature Rank')\n        \n        plt.tight_layout()\n        plt.show()\n\n    def above_threshold_features(self):\n        # Retaining features with importance above threshold - RF Regressor\n        print(f\"Features Importances List - RF Regressor: {self.importance_reg_rf.shape}\")\n        print(f\"Above Threshold Features Importances  List - RF Regressor: {self.above_threshold_features_rf.shape}\")\n        above_threshold_features_rf=self.above_threshold_features_rf\n        self.x_train_reduced_post_pearson_MI=self.x_train_df.copy()\n        self.above_threshold_features_rf=self.above_threshold_features_rf\n        self.x_train_reduced_post_pearson_MI_RF=self.x_train_reduced_post_pearson_MI[self.above_threshold_features_rf['Feature']]\n        return self.above_threshold_features_rf, self.x_train_reduced_post_pearson_MI_RF, self.rf_regressor\n\ntuned_hyperparameters = {'n_estimators': 150,         # Number of trees in the forest. \n                        'max_depth': 15,              # Maximum depth of the tree. Prevents overfitting and speeds up.\n                        'min_samples_split': 10,      # Minimum number of samples required to split an internal node.\n                        'min_samples_leaf': 5,        # Minimum number of samples required to be at a leaf node.\n                        'max_features': 0.7,          # The number of features to consider when looking for the best split.\n                        'n_jobs': 2,                  # Number of jobs kept as 2 to optimize RAM\n                        'random_state': 42            # Seed for reproducibility of results.\n                        }\n\nRF_FeatureAnalyzer_instance=RF_FeatureAnalyzer(x_train_reduced_post_pearson_MI, y_train_original_processed, rows_to_include=50000,label_col='label', threshold=0.015, **tuned_hyperparameters)\nRF_FeatureAnalyzer_instance.RF_plotting()\n#Inference:\n#1.The line chart depicts the elbow point at somewhere around 0.015 RF importance score, hence we can keep the threshold at 0.015.\nabove_threshold_features_RF,x_train_reduced_post_pearson_MI_RF, model_RF=RF_FeatureAnalyzer_instance.above_threshold_features()\nx_test_reduced_post_pearson_MI_RF=x_test_reduced_post_pearson_MI[x_train_reduced_post_pearson_MI_RF.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:02:17.643386Z","iopub.execute_input":"2025-12-01T00:02:17.643740Z","iopub.status.idle":"2025-12-01T00:03:56.158633Z","shell.execute_reply.started":"2025-12-01T00:02:17.643694Z","shell.execute_reply":"2025-12-01T00:03:56.158049Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Training the RF regressor on top features as per RF\nRF_FeatureAnalyzer_instance_two=RF_FeatureAnalyzer(x_train_reduced_post_pearson_MI_RF, y_train_original_processed, rows_to_include=50000,label_col='label', threshold=0.015, **tuned_hyperparameters)\nx_train_reduced_post_pearson_MI_RF_two, Imp_features_RF_two, model_RF_two=RF_FeatureAnalyzer_instance_two.above_threshold_features()\n\n#Predictions from lasso \nRF_pred=model_RF_two.predict(x_test_reduced_post_pearson_MI_RF)\nRF_pred_df=pd.DataFrame(RF_pred,columns=['pred_label from RF top features - Original'])\n\nm1_RF = RF_pred_df['pred_label from RF top features - Original'].round(3)\nm2_RF = y_test_original_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:\ncorr_pandas_RF = m1_RF.corr(m2_RF)\nprint(f\"Pearson Correlation for Random Forest from RF's best features - original: {corr_pandas_RF}\")\n\n#Accuracy as per RMSE\nrmse_RF = np.sqrt(mean_squared_error(m1_RF,m2_RF))\nprint(f\"Root Mean Squared Error for Random Forest from RF's best features - original: {rmse_RF:.4f}\")\n\n#Accuracy as per R2\nr2_RF = r2_score(m1_RF,m2_RF)\nprint(f\"R2 score for RF from RF's best features - original: {r2_RF:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:04:07.751957Z","iopub.execute_input":"2025-12-01T00:04:07.752247Z","iopub.status.idle":"2025-12-01T00:05:13.702145Z","shell.execute_reply.started":"2025-12-01T00:04:07.752228Z","shell.execute_reply":"2025-12-01T00:05:13.701339Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.2.5 RELATIVE IMPORTANCE OF FEATURES - LIGTWEIGHT GBM APPROACH**","metadata":{}},{"cell_type":"code","source":"# Class for analyzing feature importance through Lightweight GBM Regressor\nclass LGBM_FeatureAnalyzer:\n    # Constructor\n    def __init__(self, x_train_df, y_train_df, label_col=None,threshold=None, random_state=None):\n        self.x_train_df = x_train_df\n        self.y_train_df = y_train_df\n        self.label = label_col\n        self.threshold=threshold\n        self.random_state = random_state\n        \n        # Initialize the LGBM Regressor model\n        self.lgbm_regressor = lgb.LGBMRegressor(random_state=self.random_state)\n        \n        # Train the LGBM regressor\n        self.lgbm_regressor.fit(self.x_train_df, self.y_train_df['label'])\n\n        # Get feature importances\n        self.importances_reg_lgbm = self.lgbm_regressor.feature_importances_\n        \n        # Create a DataFrame for better readability and sorting\n        self.feature_names_reg_lgbm = self.x_train_df.columns\n        self.importance_reg_lgbm = pd.DataFrame({\n            'Feature': self.feature_names_reg_lgbm,\n            'Importance': self.importances_reg_lgbm\n        }).sort_values(by='Importance',ascending=False)\n\n        # Above Threshold features df\n        self.above_threshold_features_lgbm=self.importance_reg_lgbm[self.importance_reg_lgbm['Importance']>=self.threshold]\n\n\n    def LGBM_plotting(self):\n        # --- Visualization ---\n        # Top 50 Important Features\n        fig, my_ax = plt.subplots(nrows=1, ncols=2,figsize=(16, 7))\n        sns.barplot(y='Importance', x='Feature', data=self.importance_reg_lgbm.head(50),ax=my_ax[0])\n        my_ax[0].set_title('Top 50 Features Basis LGBM Regressor Feature Importances')\n        my_ax[0].set_ylabel('Importance')\n        my_ax[0].set_xlabel('Feature')\n        my_ax[0].tick_params(axis='x', rotation=90)\n        \n        # Create a line plot of the ranked feature importances\n        sns.lineplot(x=range(len(self.importance_reg_lgbm)), y=self.importance_reg_lgbm['Importance'],ax=my_ax[1])\n        my_ax[1].set_title('Feature Importance Basis LGBM Regressor in Descending Order')\n        my_ax[1].set_ylabel('Feature Importance')\n        my_ax[1].set_xlabel('Feature Rank')\n        \n        plt.tight_layout()\n        plt.show()\n\n\n    def above_threshold_features_LGBM(self):\n        # Retaining features with importance above threshold - RF Regressor\n        print(f\"Features Importances List - LGBM Regressor: {self.importance_reg_lgbm.shape}\")\n        print(f\"Above Threshold Features Importances  List - LGBM Regressor: {self.above_threshold_features_lgbm.shape}\")\n\n        self.x_train_reduced_post_pearson_MI=self.x_train_df.copy()\n        self.above_threshold_features_lgbm=self.above_threshold_features_lgbm\n        self.x_train_reduced_post_pearson_MI_LGBM=self.x_train_reduced_post_pearson_MI[self.above_threshold_features_lgbm['Feature']]\n        return self.above_threshold_features_lgbm, self.x_train_reduced_post_pearson_MI_LGBM, self.lgbm_regressor\n\nremove_post_LGBM_instance=LGBM_FeatureAnalyzer(x_train_reduced_post_pearson_MI, y_train_original_processed, label_col='label', threshold=30, random_state=42)\nremove_post_LGBM_instance.LGBM_plotting()\n#Inference:\n#1.The line chart depicts the elbow point at somewhere around LGBM importance score of 30, hence we can keep the threshold at 30.\nabove_threshold_features_lgbm, x_train_reduced_post_pearson_MI_LGBM, model_LGBM = remove_post_LGBM_instance.above_threshold_features_LGBM()\nx_test_reduced_post_pearson_MI_LGBM=x_test_reduced_post_pearson_MI[x_train_reduced_post_pearson_MI_LGBM.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:05:30.793827Z","iopub.execute_input":"2025-12-01T00:05:30.794385Z","iopub.status.idle":"2025-12-01T00:05:34.279548Z","shell.execute_reply.started":"2025-12-01T00:05:30.794364Z","shell.execute_reply":"2025-12-01T00:05:34.278917Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Training the LGBM regressor on top features as per LGBM\nLGBM_FeatureAnalyzer_instance_two=LGBM_FeatureAnalyzer(x_train_reduced_post_pearson_MI_LGBM, y_train_original_processed, label_col='label', threshold=30, random_state=42)\nx_train_reduced_post_pearson_MI_LGBM_two, Imp_features_LGBM_two, model_LGBM_two=LGBM_FeatureAnalyzer_instance_two.above_threshold_features_LGBM()\n\n#Predictions from LGBM \nLGBM_pred=model_LGBM_two.predict(x_test_reduced_post_pearson_MI_LGBM)\nLGBM_pred_df=pd.DataFrame(LGBM_pred,columns=['pred_label from LGBM top features - Original'])\n\nm1_LGBM = LGBM_pred_df['pred_label from LGBM top features - Original'].round(3)\nm2_LGBM = y_test_original_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:\ncorr_pandas_LGBM = m1_LGBM.corr(m2_LGBM)\nprint(f\"Pearson Correlation for LGBM Regressor from LGBM's best features - original: {corr_pandas_LGBM}\")\n\n#Accuracy as per RMSE\nrmse_LGBM = np.sqrt(mean_squared_error(m1_LGBM,m2_LGBM))\nprint(f\"Root Mean Squared Error for LGBM Regressor from LGBM's best features - original: {rmse_RF:.4f}\")\n\n#Accuracy as per R2\nr2_LGBM = r2_score(m1_LGBM,m2_LGBM)\nprint(f\"R2 score for LGBM from LGBM's best features - original: {r2_LGBM:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:05:50.026996Z","iopub.execute_input":"2025-12-01T00:05:50.027295Z","iopub.status.idle":"2025-12-01T00:05:52.567685Z","shell.execute_reply.started":"2025-12-01T00:05:50.027275Z","shell.execute_reply":"2025-12-01T00:05:52.566876Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.2.6 RELATIVE IMPORTANCE OF FEATURES - XG Boost**","metadata":{}},{"cell_type":"code","source":"# Class for analyzing feature importance through XGB Regressor\nclass XG_FeatureAnalyzer:\n    # Constructor\n    def __init__(self, x_train_df, y_train_df, label_col=None,threshold=None, random_state=None, n_estimators=None, learning_rate=None, max_depth=None, gamma=None, subsample=None, colsample_bytree=None, n_jobs=None):\n        self.x_train_df = x_train_df\n        self.y_train_df = y_train_df\n        self.label = label_col\n        self.threshold=threshold\n        self.random_state = random_state\n        self.n_estimators = n_estimators\n        self.learning_rate = learning_rate      \n        self.max_depth = max_depth               \n        self.gamma = gamma                \n        self.subsample = subsample             \n        self.colsample_bytree = colsample_bytree      \n        self.n_jobs = n_jobs\n        \n        # Initialize the XG Regressor model\n        self.xg_regressor = xgb.XGBRegressor(random_state=self.random_state, \n                                             n_estimators=self.n_estimators,\n                                             learning_rate=self.learning_rate,\n                                             max_depth=self.max_depth,\n                                             gamma=self.gamma, \n                                             subsample=self.subsample, \n                                             colsample_bytree=self.colsample_bytree,\n                                             n_jobs=self.n_jobs)\n        \n        # Train the XG regressor\n        self.xg_regressor.fit(self.x_train_df, self.y_train_df['label'])\n\n        # Get feature importances\n        self.importances_reg_xg = self.xg_regressor.feature_importances_\n        \n        # Create a DataFrame for better readability and sorting\n        self.feature_names_reg_xg = self.x_train_df.columns\n        self.importance_reg_xg = pd.DataFrame({\n            'Feature': self.feature_names_reg_xg,\n            'Importance': self.importances_reg_xg\n        }).sort_values(by='Importance',ascending=False)\n\n        # Above Threshold features df\n        self.above_threshold_features_xg=self.importance_reg_xg[self.importance_reg_xg['Importance']>=self.threshold]\n\n    def XG_plotting(self):\n        # --- Visualization ---\n        # Top 50 Important Features\n        fig, my_ax = plt.subplots(nrows=1, ncols=2,figsize=(16, 7))\n        sns.barplot(y='Importance', x='Feature', data=self.importance_reg_xg.head(50),ax=my_ax[0])\n        my_ax[0].set_title('Top 50 Features Basis XG Regressor Feature Importances')\n        my_ax[0].set_ylabel('Importance')\n        my_ax[0].set_xlabel('Feature')\n        my_ax[0].tick_params(axis='x', rotation=90)\n        \n        # Create a line plot of the ranked feature importances\n        sns.lineplot(x=range(len(self.importance_reg_xg)), y=self.importance_reg_xg['Importance'],ax=my_ax[1])\n        my_ax[1].set_title('Feature Importance Basis XG Regressor in Descending Order')\n        my_ax[1].set_ylabel('Feature Importance')\n        my_ax[1].set_xlabel('Feature Rank')\n        \n        plt.tight_layout()\n        plt.show()\n\n\n    def above_threshold_features_XG(self):\n        # Retaining features with importance above threshold - RF Regressor\n        print(f\"Features Importances List - XG Regressor: {self.importance_reg_xg.shape}\")\n        print(f\"Above Threshold Features Importances  List - XG Regressor: {self.above_threshold_features_xg.shape}\")\n\n        self.x_train_reduced_post_pearson_MI=self.x_train_df.copy()\n        self.above_threshold_features_xg=self.above_threshold_features_xg\n        self.x_train_reduced_post_pearson_MI_XG=self.x_train_reduced_post_pearson_MI[self.above_threshold_features_xg['Feature']]\n        return self.above_threshold_features_xg, self.x_train_reduced_post_pearson_MI_XG, self.xg_regressor\n\nremove_post_XG_instance=XG_FeatureAnalyzer(x_train_reduced_post_pearson_MI, y_train_original_processed, label_col='label', threshold=0.015, random_state=42, n_estimators=2000, learning_rate=0.05, max_depth=5, gamma=0.1, subsample=0.8, colsample_bytree=0.7, n_jobs=-1)\nremove_post_XG_instance.XG_plotting()\n#Inference:\n#1.The line chart depicts the elbow point at somewhere around XG importance score of 30, hence we can keep the threshold at 0.015.\nabove_threshold_features_xg, x_train_reduced_post_pearson_MI_XG, model_xg =remove_post_XG_instance.above_threshold_features_XG()\nx_test_reduced_post_pearson_MI_XG=x_test_reduced_post_pearson_MI[x_train_reduced_post_pearson_MI_XG.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:06:10.117051Z","iopub.execute_input":"2025-12-01T00:06:10.117785Z","iopub.status.idle":"2025-12-01T00:06:34.151140Z","shell.execute_reply.started":"2025-12-01T00:06:10.117752Z","shell.execute_reply":"2025-12-01T00:06:34.150515Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Training the XG regressor on top features as per XG\nXG_FeatureAnalyzer_instance_two=XG_FeatureAnalyzer(x_train_reduced_post_pearson_MI_XG, y_train_original_processed, label_col='label', threshold=30, random_state=42)\nx_train_reduced_post_pearson_MI_XG_two, Imp_features_XG_two, model_XG_two=XG_FeatureAnalyzer_instance_two.above_threshold_features_XG()\n\n#Predictions from XG \nXG_pred=model_XG_two.predict(x_test_reduced_post_pearson_MI_XG)\nXG_pred_df=pd.DataFrame(XG_pred,columns=['pred_label from XG top features - Original'])\n\nm1_XG = XG_pred_df['pred_label from XG top features - Original'].round(3)\nm2_XG = y_test_original_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:\ncorr_pandas_XG = m1_XG.corr(m2_XG)\nprint(f\"Pearson Correlation for XG Regressor from XG's best features - original: {corr_pandas_RF}\")\n\n#Accuracy as per RMSE\nrmse_XG = np.sqrt(mean_squared_error(m1_XG,m2_XG))\nprint(f\"Root Mean Squared Error for XG Regressor from XG's best features - original: {rmse_RF:.4f}\")\n\n#Accuracy as per R2\nr2_XG = r2_score(m1_XG,m2_XG)\nprint(f\"R2 score for XG from XG's best features - original: {r2_XG:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:06:47.645280Z","iopub.execute_input":"2025-12-01T00:06:47.645876Z","iopub.status.idle":"2025-12-01T00:06:49.666321Z","shell.execute_reply.started":"2025-12-01T00:06:47.645852Z","shell.execute_reply":"2025-12-01T00:06:49.665646Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.2.7 RELATIVE IMPORTANCE OF FEATURES - SVR (Support Vector Regression)**","metadata":{}},{"cell_type":"code","source":"# Class for analyzing feature importance through SVM Regressor\nclass SVR_FeatureAnalyzer:\n    # Constructor\n    def __init__(self, x_train_df, y_train_df, x_test_df, y_test_df, label_col=None, timestamps_num=None, threshold=None, kernel=None, C=None, epsilon=None, cache_size=None, n_repeats=None, random_state=None, n_jobs=None):\n        self.timestamps_num = timestamps_num\n        self.x_train_df = x_train_df.tail(self.timestamps_num)\n        self.y_train_df = y_train_df.tail(self.timestamps_num)\n        self.x_test_df = x_test_df\n        self.y_test_df = y_test_df\n        self.label = label_col\n        self.threshold = threshold\n        self.kernel = kernel\n        self.C = C\n        self.epsilon = epsilon\n        self.cache_size = cache_size\n        self.n_repeats = n_repeats\n        self.random_state = random_state\n        self.n_jobs = n_jobs\n        \n        # Initialize the SVM Regressor model\n        self.svr_model = SVR(kernel=self.kernel, C=self.C, epsilon=self.epsilon, cache_size=self.cache_size)\n        \n        # Train the SVM regressor\n        self.svr_model.fit(self.x_train_df, self.y_train_df['label'])\n\n        # Predict from SVM regressor\n        self.y_pred_array = self.svr_model.predict(x_test_df)\n        self.y_pred = pd.DataFrame(self.y_pred_array, columns=['Prediction from SVR'])\n\n        # Evaluation metrics on SVM regressor - Pearson/RMSE/R2\n        print(self.y_test_df)\n        print(self.y_pred)\n        \n        self.SVR_pearson = self.y_pred['Prediction from SVR'].corr(self.y_test_df['label'])\n        self.SVR_RMSE = np.sqrt(mean_squared_error(self.y_pred['Prediction from SVR'],self.y_test_df['label']))\n        self.SVR_R2 = r2_score(self.y_pred['Prediction from SVR'],self.y_test_df['label'])\n\n        print(\"\\n--- SVR Model Results (RBF Kernel) ---\")\n        print(f\"Test Set R-squared (R²): {self.SVR_R2:.4f}\")\n        print(f\"Test Set RMSE: {self.SVR_RMSE:.4f}\")\n        print(f\"Test Set Pearson Correlation (Actual vs. Predicted): {self.SVR_pearson:.4f}\")\n        \n        # Feature Importance using Permutation Importance (Model Agnostic)\n        print(\"\\n--- 5. Feature Importance: Permutation Importance ---\")\n        \n        # Use permutation importance on the held-out test set (X_test)\n        # The scoring metric used here (default is r2) measures the decrease in score \n        # when a feature is randomly shuffled.\n        self.result = permutation_importance(self.svr_model, self.x_test_df.tail(self.timestamps_num), \n                                        self.y_test_df.tail(self.timestamps_num), n_repeats=self.n_repeats, random_state=self.random_state, n_jobs=self.n_jobs)\n\n        # Get the indices that sort the importance scores in descending order\n        self.sorted_idx = self.result.importances_mean.argsort()[::-1]\n        print(self.sorted_idx)\n\n        # Create a DataFrame for better readability and sorting\n        self.feature_names_reg_svr = self.x_train_df.columns\n        self.importance_reg_svr = pd.DataFrame({\n            'Feature': self.feature_names_reg_svr,\n            'Importance': self.sorted_idx\n        }).sort_values(by='Importance',ascending=False)\n\n        # Above Threshold features df\n        self.above_threshold_features_svr=self.importance_reg_svr[self.importance_reg_svr['Importance']>=self.threshold]\n\n    def SVR_plotting(self):\n        \n        # Top 50 Important Features as per SVR\n        fig, my_ax = plt.subplots(nrows=1, ncols=2,figsize=(16, 7))\n        sns.barplot(y='Importance', x='Feature', data=self.importance_reg_svr.head(50),ax=my_ax[0])\n        my_ax[0].set_title('Top 50 Features Basis SVM Regressor Feature Importances')\n        my_ax[0].set_ylabel('Importance')\n        my_ax[0].set_xlabel('Feature')\n        my_ax[0].tick_params(axis='x', rotation=90)\n        \n        # Create a line plot of the ranked feature importances\n        sns.lineplot(x=range(len(self.importance_reg_svr)), y=self.importance_reg_svr['Importance'],ax=my_ax[1])\n        my_ax[1].set_title('Feature Importance Basis SVM Regressor in Descending Order')\n        my_ax[1].set_ylabel('Feature Importance')\n        my_ax[1].set_xlabel('Feature Rank')\n        \n        plt.tight_layout()\n        plt.show()\n\n    def above_threshold_features_SVR(self):\n        # Retaining features with importance above threshold - RF Regressor\n        print(f\"Features Importances List - SVM Regressor: {self.importance_reg_svr.shape}\")\n        print(f\"Above Threshold Features Importances  List - SVM Regressor: {self.above_threshold_features_svr.shape}\")\n\n        self.x_train_reduced_post_pearson_MI=self.x_train_df.copy()\n        self.above_threshold_features_svr=self.above_threshold_features_svr\n        self.x_train_reduced_post_pearson_MI_SVR=self.x_train_reduced_post_pearson_MI[self.above_threshold_features_svr['Feature']]\n        return self.above_threshold_features_svr, self.x_train_reduced_post_pearson_MI_SVR, self.svr_model\n\n\nSVR_FeatureAnalyzer_instance=SVR_FeatureAnalyzer(x_train_reduced_post_pearson_MI, y_train_original_processed, x_test_reduced_post_pearson_MI, y_test_original_processed, label_col='label', threshold=0, timestamps_num=10000, kernel='rbf', C=1, epsilon=0.1, cache_size=500, n_repeats=5, random_state=42, n_jobs=-1)\nSVR_FeatureAnalyzer_instance.SVR_plotting()\n#Inference:\n#1.The line chart depicts the elbow point at somewhere around LGBM importance score of 30, hence we can keep the threshold at 30.\n\nabove_threshold_features_svr, x_train_reduced_post_pearson_MI_SVR, model_svr = SVR_FeatureAnalyzer_instance.above_threshold_features_SVR()\nx_test_reduced_post_pearson_MI_SVR=x_test_reduced_post_pearson_MI[x_train_reduced_post_pearson_MI_SVR.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:06:58.609424Z","iopub.execute_input":"2025-12-01T00:06:58.610093Z","iopub.status.idle":"2025-12-01T00:17:03.085085Z","shell.execute_reply.started":"2025-12-01T00:06:58.610068Z","shell.execute_reply":"2025-12-01T00:17:03.084323Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Training the SVR regressor on top features as per SVR                                             \nSVR_FeatureAnalyzer_instance_two=SVR_FeatureAnalyzer(x_train_reduced_post_pearson_MI_SVR, y_train_original_processed, x_test_reduced_post_pearson_MI_SVR, y_test_original_processed, label_col='label', threshold=0, timestamps_num=10000, kernel='rbf', C=1, epsilon=0.1, cache_size=500, n_repeats=5, random_state=42, n_jobs=-1)\nabove_threshold_features_svr_two, x_train_reduced_post_pearson_MI_SVR_two, model_SVR_two=SVR_FeatureAnalyzer_instance_two.above_threshold_features_SVR()\n\n#Predictions from SVR \nSVR_pred=model_SVR_two.predict(x_test_reduced_post_pearson_MI_SVR)\nSVR_pred_df=pd.DataFrame(SVR_pred,columns=['pred_label from SVR top features - Original'])\n\nm1_SVR = SVR_pred_df['pred_label from SVR top features - Original'].round(3)\nm2_SVR = y_test_original_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:\ncorr_pandas_SVR = m1_SVR.corr(m2_SVR)\nprint(f\"Pearson Correlation for SVR Regressor from SVR's best features - original: {corr_pandas_SVR}\")\n\n#Accuracy as per RMSE\nrmse_SVR = np.sqrt(mean_squared_error(m1_SVR,m2_SVR))\nprint(f\"Root Mean Squared Error for SVR Regressor from SVR's best features - original: {rmse_SVR:.4f}\")\n\n#Accuracy as per R2\nr2_SVR = r2_score(m1_SVR,m2_SVR)\nprint(f\"R2 score for SVR from SVR's best features - original: {r2_SVR:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:18:52.583129Z","iopub.execute_input":"2025-12-01T00:18:52.583730Z","iopub.status.idle":"2025-12-01T00:28:53.598374Z","shell.execute_reply.started":"2025-12-01T00:18:52.583686Z","shell.execute_reply":"2025-12-01T00:28:53.597523Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 7.2 WORKFLOW 1 (POST INITIAL PRUNING) - NON-LINEAR FEATURE SELECTION PIPELINE - ORIGINAL NON-REDUNDANT FEATURES. ","metadata":{}},{"cell_type":"markdown","source":"**7.2.1 RELATIVE IMPORTANCE OF FEATURES - Deep Neural Networks (ANNs)**","metadata":{}},{"cell_type":"code","source":"class ANN_FeatureAnalyzer:\n    # --- Existing __init__ method here (omitted for brevity) ---\n    def __init__(self, x_train_df, y_train_df, x_test_df, y_test_df, label_col=None, threshold=None, timestamps_num=None, activation=None, output_activation=None, optimizer=None, loss=None, metrics=None, monitor=None, patience=None, restore_best_weights=None, verbose=None, validation_split=None, epochs=None, batch_size=None, squared=None, n_repeats=None, random_state=None, n_jobs=None, scoring=None):\n        self.timestamps_num = timestamps_num\n        self.x_train_df = x_train_df.tail(self.timestamps_num)\n        self.y_train_df = y_train_df.tail(self.timestamps_num)\n        self.x_test_df = x_test_df\n        self.y_test_df = y_test_df\n        self.label = label_col\n        self.input_dim = self.x_train_df.shape[1]\n        self.threshold = threshold\n        self.activation = activation\n        self.output_activation = output_activation\n        self.optimizer = optimizer\n        self.loss = loss\n        self.metrics = metrics\n        self.monitor = monitor\n        self.patience = patience\n        self.restore_best_weights = restore_best_weights\n        self.verbose = verbose\n        self.validation_split = validation_split\n        self.epochs = epochs\n        self.batch_size = batch_size\n        self.squared = squared\n        self.n_repeats = n_repeats\n        self.random_state = random_state\n        self.n_jobs = n_jobs\n        self.early_stopping = None # Will be instantiated in model_prep\n        self.scoring = scoring\n        \n    # --- Your Keras Model Builder (with corrected return) ---\n    def build_ann_model(self):\n        self.ann_model = Sequential([\n            Dense(64, activation=self.activation, input_shape=(self.input_dim,)),\n            Dropout(0.2),\n            Dense(32, activation=self.activation),\n            Dropout(0.2),\n            Dense(16, activation=self.activation),\n            Dense(1, activation=self.output_activation)\n        ])\n        # CORRECTED: Return the Keras model object\n        return self.ann_model \n\n    # --- Keras Compilation and Setup (Your provided method) ---\n    def model_prep(self):\n        # 1. Build the model\n        self.build_ann_model()\n        \n        # 2. Compile the model\n        self.ann_model.compile(optimizer=self.optimizer, loss=self.loss, metrics=self.metrics)\n\n        # 3. Define Early Stopping (instantiated here)\n        self.early_stopping = EarlyStopping(\n            monitor=self.monitor, \n            patience=self.patience, \n            restore_best_weights=self.restore_best_weights, \n            verbose=self.verbose\n        )\n        return self.ann_model # Return the compiled model\n\n    # -----------------------------------------------------------------\n    # SCALAR METHODS (MANDATORY FOR SCIKIT-LEARN COMPLIANCE)\n    # These are used internally by permutation_importance\n    # -----------------------------------------------------------------\n\n    def fit(self, X_train_data=None, y_train_data=None):\n        \"\"\"\n        [MANDATORY SCORING METHOD] Satisfies the Scikit-learn API check.\n        Performs the actual Keras training by calling model_prep().\n        \"\"\"\n        # The Canvas code correctly determines which data to use:\n        X_data = X_train_data if X_train_data is not None else self.x_train_df\n        y_data = y_train_data if y_train_data is not None else self.y_train_df\n\n        self.input_dim = X_data.shape[1] \n        \n        # Call model_prep to build, compile, and define callbacks\n        self.model_prep() \n        \n        # Perform the actual training (using user's variable name self.ann_fit)\n        # NOTE: We use self.x_train_df and self.y_train_df for Keras training\n        self.ann_fit = self.ann_model.fit(\n            X_data, y_data, \n            validation_split=self.validation_split, \n            epochs=self.epochs, \n            batch_size=self.batch_size,\n            callbacks=[self.early_stopping], # Ensure callbacks are passed as a list\n            verbose=self.verbose\n        )\n        # MUST return the estimator instance itself\n        return self \n\n    def predict(self, X):\n        \"\"\"\n        [MANDATORY SCORING METHOD] Uses the trained Keras model for prediction.\n        \"\"\"\n        if not hasattr(self, 'ann_model'):\n            # Note: We rely on permutation_importance calling .fit() first.\n            raise AttributeError(\"ANN model has not been built/trained. Call .fit() first.\")\n            \n        # X is passed from permutation_importance as the data slice\n        # Keras predict returns 2D, flatten to 1D for scikit-learn metrics\n        return self.ann_model.predict(X, verbose=0).flatten()\n\n    def score(self, X, y):\n        \"\"\"\n        [MANDATORY SCORING METHOD] Calculates the R^2 score for the Scikit-learn API.\n        \"\"\"\n        y_pred = self.predict(X)\n        \n        # Ensure target 'y' is also a flat array for r2_score comparison\n        if isinstance(y, pd.DataFrame) or isinstance(y, pd.Series):\n             y_flat = y.values.flatten()\n        else:\n             y_flat = y.flatten()\n\n        return r2_score(y_flat, y_pred)\n    \n    # -----------------------------------------------------------------\n    # COMPREHENSIVE EVALUATION METHOD\n    # -----------------------------------------------------------------\n    \n    def model_score(self):\n        \"\"\"\n        Trains the model, runs predictions on the full test set, \n        and calculates all primary evaluation metrics (RMSE, R2, Pearson r).\n        \"\"\"\n        # Call fit() with required arguments to train the model\n        self.fit(self.x_train_df, self.y_train_df)\n        \n        # Evaluate ANN on the full X_test\n        # Predictions from ANN - Use the full test set data stored in self.\n        self.y_pred_ann = pd.DataFrame(\n            self.ann_model.predict(self.x_test_df, verbose=self.verbose).flatten(), \n            columns=['Predictions from ANN']\n        )\n        \n        # Calculate metrics (using self.label for robustness)\n        target_col = self.label or 'label'\n        \n        self.ann_rmse = mean_squared_error(\n            self.y_test_df[target_col], \n            self.y_pred_ann['Predictions from ANN'], \n            squared=self.squared\n        )\n        self.ann_r2 = r2_score(\n            self.y_test_df[target_col], \n            self.y_pred_ann['Predictions from ANN']\n        )\n        self.ann_pearson, _ = pearsonr(\n            self.y_test_df[target_col], \n            self.y_pred_ann['Predictions from ANN']\n        )\n        \n        # Print results for quick check\n        print(f\"ANN Model Results: R2={self.ann_r2:.4f}, RMSE={self.ann_rmse:.4f}, Pearson={self.ann_pearson:.4f}\")\n        \n        return self.ann_r2, self.ann_model\n    \n    # -----------------------------------------------------------------\n    # Feature Analysis Method \n    # -----------------------------------------------------------------\n    \n    def model_permutation(self):\n        \"\"\"\n        Calculates permutation importance on the fitted estimator.\n        \"\"\"\n        # Ensure the model is trained before running feature importance\n        if not hasattr(self, 'ann_model') or not hasattr(self, 'ann_fit'):\n            self.fit(self.x_train_df, self.y_train_df)\n            \n        # 1. Calculate Permutation Importance\n        self.result = permutation_importance(\n            self, # The estimator (now Scikit-learn compliant)\n            # Using the last portion of the test set for relevance\n            self.x_test_df.tail(self.timestamps_num), \n            self.y_test_df.tail(self.timestamps_num), \n            n_repeats=self.n_repeats, \n            random_state=self.random_state, \n            n_jobs=self.n_jobs, \n            scoring=self.scoring\n        )\n        \n        # 2. Get the indices that sort the importance scores in descending order\n        self.sorted_idx = self.result.importances_mean.argsort()[::-1]\n        \n        # 3. Create correctly sorted arrays using the index array (self.sorted_idx)\n        sorted_importance_means = self.result.importances_mean[self.sorted_idx]\n        sorted_feature_names = self.x_train_df.columns[self.sorted_idx]\n\n        # 4. Create the final DataFrame using the actual sorted scores\n        self.importance_reg_ANN = pd.DataFrame({\n            'Feature': sorted_feature_names,\n            # Use the actual sorted scores, NOT the index array\n            'Importance': sorted_importance_means \n        })\n        \n        # 5. Filter features above the threshold\n        self.above_threshold_features_ANN_df = self.importance_reg_ANN[\n            self.importance_reg_ANN['Importance'] >= self.threshold\n        ]\n        \n        print(\"Feature Importances (ANN Regressor):\\n\", self.importance_reg_ANN)\n        \n        return self.importance_reg_ANN, self.above_threshold_features_ANN_df, self.ann_model\n    \n    # --- Placeholder for ANN_plotting (Ensuring it calls model_permutation) ---\n    def ANN_plotting(self):\n        # We need to import matplotlib and seaborn at the top of the file for this to work\n        import matplotlib.pyplot as plt\n        import seaborn as sns\n        \n        # The plotting method MUST call the permutation method first\n        self.importance_reg_ANN, self.above_threshold_features_ANN_df, self.ann_model = self.model_permutation()\n        \n        # Now plot the results\n        fig, my_ax = plt.subplots(nrows=1, ncols=2,figsize=(16, 7))\n        sns.barplot(y='Importance', x='Feature', data=self.importance_reg_ANN.head(50),ax=my_ax[0])\n        my_ax[0].set_title('Top 50 Features Basis ANN Regressor Feature Importances')\n        my_ax[0].set_ylabel('Importance')\n        my_ax[0].set_xlabel('Feature')\n        my_ax[0].tick_params(axis='x', rotation=90)\n        \n        # Create a line plot of the ranked feature importances\n        sns.lineplot(x=range(len(self.importance_reg_ANN)), y=self.importance_reg_ANN['Importance'],ax=my_ax[1])\n        my_ax[1].set_title('Feature Importance Basis ANN Regressor in Descending Order')\n        my_ax[1].set_ylabel('Feature Importance')\n        my_ax[1].set_xlabel('Feature Rank')\n        \n        plt.tight_layout()\n        plt.show()\n\n    # --- Placeholder for above_threshold_features_ANN (Ensuring it calls ANN_plotting) ---\n    def above_threshold_features_ANN(self):\n        # Calling ANN_plotting ensures all calculations are run and saved to self.\n        # This will indirectly ensure self.model_permutation() is run.\n        self.ANN_plotting() \n        \n        # Retaining features with importance above threshold\n        print(f\"Features Importances List - ANN Regressor: {self.importance_reg_ANN.shape}\")\n        \n        # The following variables were created and saved inside model_permutation()\n        x_train_reduced = self.x_train_df[self.above_threshold_features_ANN_df['Feature']]\n        \n        return self.above_threshold_features_ANN_df, x_train_reduced, self.ann_model\n\nANN_FeatureAnalyzer_instance=ANN_FeatureAnalyzer(x_train_reduced_post_pearson_MI, y_train_original_processed, x_test_reduced_post_pearson_MI, y_test_original_processed, label_col='label', threshold=-0.025, timestamps_num=10000, activation='relu', output_activation='linear', optimizer='adam', loss='mse', metrics=['mae'], monitor='val_loss', patience=15, restore_best_weights=True, verbose=0, validation_split=0.2, epochs=100, batch_size=256, squared=None, n_repeats=3, random_state=42, n_jobs=-1, scoring='r2')\nabove_threshold_features_ANN, x_train_reduced_post_pearson_MI_ANN, model_ANN = ANN_FeatureAnalyzer_instance.above_threshold_features_ANN()\n\nx_test_reduced_post_pearson_MI_ANN=x_test_reduced_post_pearson_MI[x_train_reduced_post_pearson_MI_ANN.columns]\n#Inference:\n#1.The line chart depicts the elbow point at somewhere around ANN Permutation importance score of -0.025, hence we can keep the threshold at -0.025","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:05:30.426367Z","iopub.execute_input":"2025-12-01T01:05:30.426932Z","iopub.status.idle":"2025-12-01T01:06:37.019984Z","shell.execute_reply.started":"2025-12-01T01:05:30.426909Z","shell.execute_reply":"2025-12-01T01:06:37.019129Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Training the ANN regressor on top features as per ANN\nANN_FeatureAnalyzer_instance_two=ANN_FeatureAnalyzer(x_train_reduced_post_pearson_MI_ANN, y_train_original_processed, x_test_reduced_post_pearson_MI_ANN, y_test_original_processed, label_col='label', threshold=-0.025, timestamps_num=10000, activation='relu', output_activation='linear', optimizer='adam', loss='mse', metrics=['mae'], monitor='val_loss', patience=15, restore_best_weights=True, verbose=0, validation_split=0.2, epochs=100, batch_size=256, squared=None, n_repeats=3, random_state=42, n_jobs=-1, scoring='r2')\nabove_threshold_features_ANN_two, x_train_reduced_post_pearson_MI_ANN_two, model_ANN_two = ANN_FeatureAnalyzer_instance_two.above_threshold_features_ANN()\n\n#Training the ANN regressor on top features as per ANN \n#ANN_FeatureAnalyzer_instance.fit(X_train_data=x_train_reduced_post_pearson_MI_ANN, y_train_data=ANN_FeatureAnalyzer_instance.y_train_df)\n\n# Extract the ANN model from the ANN feature analyzer class\n#model_ANN_reduced = ANN_FeatureAnalyzer_instance.ann_model\n\n#Predictions from ANN model\nANN_pred=model_ANN_two.predict(x_test_reduced_post_pearson_MI_ANN)\nANN_pred_df=pd.DataFrame(ANN_pred,columns=['Predictions from ANN top features - Original'])\n\nm1_ANN = ANN_pred_df['Predictions from ANN top features - Original'].round(3)\nm2_ANN = y_test_original_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation\ncorr_pandas_ANN = m1_ANN.corr(m2_ANN)\nprint(f\"Pearson Correlation from ANN's best features - Original: {corr_pandas_ANN}\")\n\n#Accuracy as per RMSE\nrmse_ANN = np.sqrt(mean_squared_error(m1_ANN, m2_ANN))\nprint(f\"Root Mean Squared Error from ANN's best features - Original: {rmse_ANN:.4f}\")\n\n#Accuracy as per R2\nr2_ANN = r2_score(m1_ANN,m2_ANN)\nprint(f\"R2 score for ANN from ANN's best features - original: {r2_ANN:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:08:30.641210Z","iopub.execute_input":"2025-12-01T01:08:30.641954Z","iopub.status.idle":"2025-12-01T01:09:30.346349Z","shell.execute_reply.started":"2025-12-01T01:08:30.641927Z","shell.execute_reply":"2025-12-01T01:09:30.345417Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 7.4 WORKFLOW 2 (POST INITIAL PRUNING) - FEATURE ENGINEERING ON ORIGINAL NON-REDUNDANT FEATURES.\n\n1. CREATING ROLLING AVERAGES FEATURES - SMOOTHING OUT THE NOISE\n2. CREATING EWM AVERAGE FEATURES\n3. CREATING LAGGING FEATURES","metadata":{}},{"cell_type":"code","source":"# Class for Combining the sets - x_train_reduced_post_pearson_MI/x_test_reduced_post_pearson_MI and y_train_original_processed/y_test_original_processed\n# as feature engineering should be done on entire x_train_DRW and then further split to train test sets\n\nclass combine_df():\n    # constructor\n    def __init__(self, x_train_df, x_test_df, y_train_df, y_test_df):\n        self.x_train_df = x_train_df.copy()\n        self.x_test_df = x_test_df.copy()\n        self.y_train_df = y_train_df.copy()\n        self.y_test_df = y_test_df.copy()\n\n    def combine(self):\n        self.combined_df_x=pd.concat([self.x_train_df,self.x_test_df],axis=0).reset_index(drop=True)\n        self.combined_df_y=pd.concat([self.y_train_df,self.y_test_df],axis=0).reset_index(drop=True)\n        return self.combined_df_x, self.combined_df_y\n        \n# Combining the sets - x_train_reduced_post_pearson_MI/x_test_reduced_post_pearson_MI and y_train_original_processed/y_test_original_processed\ncombine_df_instance=combine_df(x_train_reduced_post_pearson_MI,x_test_reduced_post_pearson_MI,y_train_original_processed, y_test_original_processed)\ncombined_df_x, combined_df_y=combine_df_instance.combine()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:52:50.640136Z","iopub.execute_input":"2025-12-01T00:52:50.640875Z","iopub.status.idle":"2025-12-01T00:52:50.679370Z","shell.execute_reply.started":"2025-12-01T00:52:50.640851Z","shell.execute_reply":"2025-12-01T00:52:50.678638Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# # Class to engineer new features - Rolling Averages, EWMA and Lagging features for all Original non-redundant data post initial pruning- x_train_reduced_post_pearson) ---\nimport gc\n\ndef batch_columns(col_list, batch_size):\n    for i in range(0, len(col_list), batch_size):\n        yield col_list[i:i+batch_size]\n\nclass feature_engg():\n    def __init__(self, combined_df_x, combined_df_y):\n        self.combined_df_x = combined_df_x.copy()\n        self.combined_df_y = combined_df_y.copy()\n        \n    def rolling_avg(self, batch_size=10):\n        all_batches = []\n        for col_batch in batch_columns(list(self.combined_df_x.columns), batch_size):\n            rolling_avg_1_wk_df=pd.DataFrame(index=self.combined_df_x.index)\n            rolling_avg_1_month_df=pd.DataFrame(index=self.combined_df_x.index)\n            rolling_avg_1_qtr_df=pd.DataFrame(index=self.combined_df_x.index)\n\n            for col in col_batch:\n                rolling_avg_1_wk_df[f'Rolling Avg - 15 mins Period {col}'] = self.combined_df_x[col].rolling(15).mean()\n                rolling_avg_1_month_df[f'Rolling Avg - 30 mins Period {col}'] = self.combined_df_x[col].rolling(30).mean()\n                rolling_avg_1_qtr_df[f'Rolling Avg - 60 mins Period {col}'] = self.combined_df_x[col].rolling(60).mean()\n\n            batch_df = pd.concat([rolling_avg_1_wk_df, rolling_avg_1_month_df, rolling_avg_1_qtr_df], axis=1)\n            all_batches.append(batch_df)\n            del rolling_avg_1_wk_df, rolling_avg_1_month_df, rolling_avg_1_qtr_df, batch_df\n            gc.collect()\n\n        return pd.concat(all_batches, axis=1)\n\n    def EWMA_features(self, batch_size=10):\n        all_batches = []\n        for col_batch in batch_columns(list(self.combined_df_x.columns), batch_size):\n            EWM_avg_1_wk_df=pd.DataFrame(index=self.combined_df_x.index)\n            EWM_avg_1_month_df=pd.DataFrame(index=self.combined_df_x.index)\n            EWM_avg_1_qtr_df=pd.DataFrame(index=self.combined_df_x.index)\n            \n            for col in col_batch:\n                EWM_avg_1_wk_df[f'EWM Avg - 15 mins Period {col}'] = self.combined_df_x[col].ewm(span=15,min_periods=15).mean()\n                EWM_avg_1_month_df[f'EWM Avg - 30 mins Period {col}'] = self.combined_df_x[col].ewm(span=30,min_periods=30).mean()\n                EWM_avg_1_qtr_df[f'Rolling Avg - 60 mins Period {col}'] = self.combined_df_x[col].ewm(span=60,min_periods=60).mean()\n            \n            batch_df = pd.concat([EWM_avg_1_wk_df, EWM_avg_1_month_df, EWM_avg_1_qtr_df], axis=1)\n            all_batches.append(batch_df)\n            del EWM_avg_1_wk_df, EWM_avg_1_month_df, EWM_avg_1_qtr_df, batch_df\n            gc.collect()\n\n        return pd.concat(all_batches, axis=1)\n\n    def lagging_features(self, batch_size=10):\n        all_batches = []\n        for col_batch in batch_columns(list(self.combined_df_x.columns), batch_size):\n            Lag_10mins_df=pd.DataFrame(index=self.combined_df_x.index)\n            Lag_30mins_df=pd.DataFrame(index=self.combined_df_x.index)\n            Lag_60mins_df=pd.DataFrame(index=self.combined_df_x.index)\n            Lag_720mins_df=pd.DataFrame(index=self.combined_df_x.index)\n            Lag_1440mins_df=pd.DataFrame(index=self.combined_df_x.index)\n            \n            for col in col_batch:\n                Lag_10mins_df[f'Lag - 15 Minutes {col}'] = self.combined_df_x[col].shift(periods=15)\n                Lag_30mins_df[f'Lag - 30 Minutes {col}'] = self.combined_df_x[col].shift(periods=30)\n                Lag_60mins_df[f'Lag - 60 Minutes {col}'] = self.combined_df_x[col].shift(periods=60)\n                # Longer lags commented out in original\n                # Lag_720mins_df[f'Lag - 720 Minutes {col}'] = self.combined_df_x[col].shift(periods=720)\n                # Lag_1440mins_df[f'Lag - 1440 Minutes {col}'] = self.combined_df_x[col].shift(periods=1440)\n\n            batch_df = pd.concat([Lag_10mins_df, Lag_30mins_df, Lag_60mins_df, Lag_720mins_df, Lag_1440mins_df], axis=1)\n            all_batches.append(batch_df)\n            del Lag_10mins_df, Lag_30mins_df, Lag_60mins_df, Lag_720mins_df, Lag_1440mins_df, batch_df\n            gc.collect()\n\n        return pd.concat(all_batches, axis=1)\n\n    def final_engg_features(self):\n        rolling_features_df = self.rolling_avg()\n        ewm_features_df = self.EWMA_features()\n        lag_features_df = self.lagging_features()\n\n        self.Final_Engg_df_x = pd.concat([self.combined_df_x, rolling_features_df, ewm_features_df, lag_features_df], axis=1)\n        self.Final_Engg_df_x = self.Final_Engg_df_x.loc[:, ~self.Final_Engg_df_x.columns.duplicated()]\n        self.Final_Engg_df_x = self.Final_Engg_df_x.dropna().reset_index(drop=True)\n\n        self.y_engg = self.combined_df_y.loc[self.Final_Engg_df_x.index]\n\n        # Convert to float32\n        for df in (self.Final_Engg_df_x, self.y_engg):\n            for col in df.columns:\n                if df[col].dtype != 'float32':\n                    df[col] = df[col].astype('float32')\n\n        print(f\"Shape of Final Engineered Dataset: {self.Final_Engg_df_x.shape}\")\n        print(f\"Shape of aligned y_train_engg: {self.y_engg.shape}\")\n        \n        return self.Final_Engg_df_x, self.y_engg\n\nfeature_engg_instance = feature_engg(combined_df_x, combined_df_y)\nx_engg, y_engg = feature_engg_instance.final_engg_features()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:53:33.055650Z","iopub.execute_input":"2025-12-01T00:53:33.056263Z","iopub.status.idle":"2025-12-01T00:53:45.785304Z","shell.execute_reply.started":"2025-12-01T00:53:33.056240Z","shell.execute_reply":"2025-12-01T00:53:45.784576Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.3.2 TRAIN TEST SPLIT ON ENGINEERED DATASET**","metadata":{}},{"cell_type":"code","source":"DataSplitter_instance_engg=DataSplitter(x_engg, y_engg,test_size=0.3, random_state=42, shuffle=False)\nx_train_engg, x_test_engg, y_train_engg, y_test_engg=DataSplitter_instance_engg.train_test_set()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:54:11.134408Z","iopub.execute_input":"2025-12-01T00:54:11.134690Z","iopub.status.idle":"2025-12-01T00:54:11.569344Z","shell.execute_reply.started":"2025-12-01T00:54:11.134670Z","shell.execute_reply":"2025-12-01T00:54:11.568541Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.3.3 OUTLIER TREATMENT PIPELINE ON ENGINEERED DATASET**","metadata":{}},{"cell_type":"code","source":"processed_dataframes_final_engg = {}\n\n# outlier_treatement class methods call on x_train_engg exclusively on engineered features\nx_processor_engg = outlier_treatement(dataframe=x_train_engg.iloc[:,50:], scaling_fit_dataset=x_train_engg.iloc[:,50:], threshold_for_std_dev=1e-9)\nx_train_engg_processed = x_processor_engg.memory_optimize()\nprocessed_dataframes_final_engg['x_train_engg_processed'] = x_train_engg_processed\n\n# outlier_treatement class methods call on x_test_engg exclusively on engineered features\nx_processor_engg.dataframe = x_test_engg.iloc[:,50:].copy()\nx_test_engg_processed = x_processor_engg.memory_optimize()\nprocessed_dataframes_final_engg['x_test_engg_processed'] = x_test_engg_processed\n\n# outlier_treatement class methods call on y_train_engg exclusively on engineered features\ny_processor_engg = outlier_treatement(dataframe=y_train_engg, scaling_fit_dataset=y_train_engg, threshold_for_std_dev=1e-9)\ny_train_engg_processed = y_processor_engg.memory_optimize()\nprocessed_dataframes_final_engg['y_train_engg_processed'] = y_train_engg_processed\n\n# outlier_treatement class methods call on y_test_engg exclusively on engineered features\ny_processor_engg.dataframe = y_test_engg.copy()\ny_test_engg_processed = y_processor_engg.memory_optimize()\nprocessed_dataframes_final_engg['y_test_engg_processed'] = y_test_engg_processed\n\n# Storing results in new variables\nx_train_engg_processed=processed_dataframes_final_engg['x_train_engg_processed']\nx_test_engg_processed=processed_dataframes_final_engg['x_test_engg_processed']\ny_train_engg_processed=processed_dataframes_final_engg['y_train_engg_processed']\ny_test_engg_processed=processed_dataframes_final_engg['y_test_engg_processed']\n\n#Combining the original features back with transformed engineered features\nx_train_engg_processed=pd.concat([x_train_engg.iloc[:,0:50],x_train_engg_processed],axis=1)\nx_test_engg_processed=pd.concat([x_test_engg.iloc[:,0:50],x_test_engg_processed],axis=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T00:54:20.087916Z","iopub.execute_input":"2025-12-01T00:54:20.088521Z","iopub.status.idle":"2025-12-01T01:02:08.593815Z","shell.execute_reply.started":"2025-12-01T00:54:20.088498Z","shell.execute_reply":"2025-12-01T01:02:08.592980Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"###### free the dictionary not needed anymore\ndel processed_dataframes_final_engg","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:10:17.444990Z","iopub.execute_input":"2025-12-01T01:10:17.445791Z","iopub.status.idle":"2025-12-01T01:10:17.449263Z","shell.execute_reply.started":"2025-12-01T01:10:17.445757Z","shell.execute_reply":"2025-12-01T01:10:17.448720Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Final dropping of features with extreme std deviation features in x_test_engg_processed to increase signal to noise\nx_test_engg_processed_reduced = x_test_engg_processed.drop(columns=x_test_engg_processed.columns[x_test_engg_processed.std() > 3])\nx_train_engg_processed_reduced=x_train_engg_processed[x_test_engg_processed_reduced.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:10:23.806386Z","iopub.execute_input":"2025-12-01T01:10:23.806949Z","iopub.status.idle":"2025-12-01T01:10:23.992127Z","shell.execute_reply.started":"2025-12-01T01:10:23.806925Z","shell.execute_reply":"2025-12-01T01:10:23.991507Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **7.3.4 FEATURE IMPORTANCE ON ENGINEERED DATASET**","metadata":{}},{"cell_type":"markdown","source":"**7.3.4.1 INITIAL VARIANCE THRESHOLD - FEATURE SELECTION**","metadata":{}},{"cell_type":"code","source":"# Class to analyse variances for all features\nclass feature_variance_analyzer():\n    #Constructor\n    def __init__(self, x_train_engg_df):\n        self.x_train_engg_df = x_train_engg_df\n\n    # Function to plot feature variances\n    def feature_variance_plotting(self):\n        # Calculate variances\n        self.variances = self.x_train_engg_df.var().sort_values(ascending=True) \n        print(\"Top 10 features with lowest variance:\\n\", self.variances.head(10))\n        print(\"\\nTop 10 features with highest variance:\\n\", self.variances.tail(10))\n        \n        # Visualize the distribution of variances\n        plt.figure(figsize=(12, 6))\n        sns.histplot(self.variances, bins=50, kde=True)\n        plt.title('Distribution of Feature Variances (Scaled Data)')\n        plt.xlabel('Variance')\n        plt.ylabel('Number of Features')\n        plt.grid(True)\n        plt.show()\n        \n        # Plot cumulative sum of variances to see how many features contribute to variance\n        # This can help decide a cutoff if you're looking for a sharp drop-off\n        plt.figure(figsize=(12, 6))\n        plt.plot(np.arange(len(self.variances)), np.cumsum(self.variances.values))\n        plt.title('Cumulative Sum of Variances')\n        plt.xlabel('Number of Features (sorted by variance)')\n        plt.ylabel('Cumulative Variance')\n        plt.grid(True)\n        plt.show()\n        return self.variances\n\n    # Function to drop features zero threshold\n    def dropping_below_variance_threshold(self):\n        self.feature_variance_plotting()\n        more_than_zero_variance_cols = list(self.variances[self.variances > 0].index)\n        self.x_train_engg_processed_post_variance=self.x_train_engg_df[more_than_zero_variance_cols]\n        print(f\"The shape of engineered dataset post dropping zero variances features: {self.x_train_engg_processed_post_variance.shape}\")\n        return self.x_train_engg_processed_post_variance\n        \nfeature_variance_analyzer_instance_engg=feature_variance_analyzer(x_train_engg_processed_reduced)\nx_train_engg_processed_dropped_zero_variance=feature_variance_analyzer_instance_engg.dropping_below_variance_threshold()\nx_test_engg_processed_dropped_zero_variance=x_test_engg_processed_reduced[x_train_engg_processed_dropped_zero_variance.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:10:29.670453Z","iopub.execute_input":"2025-12-01T01:10:29.671316Z","iopub.status.idle":"2025-12-01T01:10:30.720642Z","shell.execute_reply.started":"2025-12-01T01:10:29.671291Z","shell.execute_reply":"2025-12-01T01:10:30.719887Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.3.4.2 PRUNING BY MUTICOLLINEARITY CHECK ON ENGINEERED FEATURES**","metadata":{}},{"cell_type":"code","source":"remove_multicollinearity_instance_engg=remove_multicollinearity(x_train_engg_processed_dropped_zero_variance, y_train_engg_processed,threshold=0.80,rows_to_include=40000,label='label',method='pearson')\nx_train_engg_dropped_variance_pearson=remove_multicollinearity_instance_engg.multicollinear_treatment()\nx_test_engg_dropped_variance_pearson=x_test_engg_processed_dropped_zero_variance[x_train_engg_dropped_variance_pearson.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:10:41.242664Z","iopub.execute_input":"2025-12-01T01:10:41.242994Z","iopub.status.idle":"2025-12-01T01:12:37.567300Z","shell.execute_reply.started":"2025-12-01T01:10:41.242974Z","shell.execute_reply":"2025-12-01T01:12:37.566467Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.3.4.3 EXAMINING FEATURE IMPORTANCE THROUGH MUTUAL INFORMANTION ON REMAINING FEATURES**","metadata":{}},{"cell_type":"code","source":"MIFeatureAnalyzer_instance_engg=MIFeatureAnalyzer(x_train_engg_dropped_variance_pearson, y_train_engg_processed, label_col='label', rows_to_include=50000, threshold=0.1)\nMIFeatureAnalyzer_instance_engg.MI_plotting()\n#Inference:\n#1.The line chart depicts the elbow point at somewhere around 0.1 MI score, hence we can keep the threshold at 0.1.\nx_train_engg_dropped_variance_pearson_MI=MIFeatureAnalyzer_instance_engg.MI_feature_drop()\nx_test_engg_dropped_variance_pearson_MI=x_test_engg_dropped_variance_pearson[x_train_engg_dropped_variance_pearson_MI.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:13:38.285676Z","iopub.execute_input":"2025-12-01T01:13:38.286015Z","iopub.status.idle":"2025-12-01T01:13:53.363173Z","shell.execute_reply.started":"2025-12-01T01:13:38.285994Z","shell.execute_reply":"2025-12-01T01:13:53.362205Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.3.4.4 EXAMINING FEATURE IMPORTANCE THROUGH LASSO L-1 ON ENGINEERED DATASET**","metadata":{}},{"cell_type":"code","source":"lasso_FeatureAnalyzer_instance_engg=lasso_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI, y_train_engg_processed, cv=5, random_state=42, max_iter=10000)\nlasso_FeatureAnalyzer_instance_engg.lasso_plotting()\nx_train_engg_dropped_variance_pearson_MI_lasso, Imp_features_lasso_engg,model_lasso_engg=lasso_FeatureAnalyzer_instance_engg.lasso_feature_drop()\nx_test_engg_dropped_variance_pearson_MI_lasso=x_test_engg_dropped_variance_pearson_MI[x_train_engg_dropped_variance_pearson_MI_lasso.columns]\n\n#Training the lasso regressor on top engg features as per Lasso\nlasso_FeatureAnalyzer_instance_two_engg=lasso_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI_lasso, y_train_engg_processed, cv=5, random_state=42, max_iter=10000)\nx_train_engg_dropped_variance_pearson_MI_lasso_two, Imp_features_lasso_two_engg, model_lasso_two_engg=lasso_FeatureAnalyzer_instance_two_engg.lasso_feature_drop()\n\n#Predictions from lasso best features - engineered\nlasso_pred_engg=model_lasso_two_engg.predict(x_test_engg_dropped_variance_pearson_MI_lasso)\nlasso_pred_df_engg=pd.DataFrame(lasso_pred_engg,columns=['Pred_label from Lasso top features - Engineered'])\n\nm1_lasso_engg = lasso_pred_df_engg['Pred_label from Lasso top features - Engineered'].round(3)\nm2_lasso_engg = y_test_engg_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:\ncorr_pandas_lasso_engg = m1_lasso_engg.corr(m2_lasso_engg)\nprint(f\"Pearson Correlation for Lasso from Lasso's best features - Engineered: {corr_pandas_lasso_engg}\")\n\n#Accuracy as per RMSE\nrmse_pandas_lasso_engg = np.sqrt(mean_squared_error(m1_lasso_engg,m2_lasso_engg))\nprint(f\"Root Mean Squared Error for Lasso from Lasso's best features - Engineered: {rmse_pandas_lasso_engg:.4f}\")\n\n#Accuracy as per R2 Score\nr2_lasso_engg = r2_score(m1_lasso_engg,m2_lasso_engg)\nprint(f\"R2 for Lasso from Lasso's best features - Engineered: {r2_lasso_engg:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:15:51.718326Z","iopub.execute_input":"2025-12-01T01:15:51.718654Z","iopub.status.idle":"2025-12-01T01:15:55.281420Z","shell.execute_reply.started":"2025-12-01T01:15:51.718633Z","shell.execute_reply":"2025-12-01T01:15:55.278741Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.3.4.5 EXAMINING FEATURE IMPORTANCES THROUGH RIDGE L-2 ON ENGINEERED DATASET**","metadata":{}},{"cell_type":"code","source":"ridge_FeatureAnalyzer_instance_engg=ridge_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI, y_train_engg_processed, cv=5)\nridge_FeatureAnalyzer_instance_engg.ridge_plotting()\nx_train_engg_dropped_variance_pearson_MI_ridge, Imp_features_ridge_engg,model_ridge_engg=ridge_FeatureAnalyzer_instance_engg.ridge_feature_drop()\nx_test_engg_dropped_variance_pearson_MI_ridge=x_test_engg_dropped_variance_pearson_MI[x_train_engg_dropped_variance_pearson_MI_ridge.columns]\n\n#Ridge model trained on Best Features basis Ridge coefficients\nridge_FeatureAnalyzer_instance_two_engg=ridge_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI_ridge, y_train_engg_processed, cv=5)\nx_train_engg_dropped_variance_pearson_MI_ridge_two, Imp_features_ridge_engg_two, model_ridge_engg_two=ridge_FeatureAnalyzer_instance_two_engg.ridge_feature_drop()\n\n#Predictions from Ridge from Best features - engineered\nridge_pred_engg=model_ridge_engg_two.predict(x_test_engg_dropped_variance_pearson_MI_ridge)\nridge_pred_df_engg=pd.DataFrame(ridge_pred_engg,columns=['Pred_label from Ridge Top features - Engineered'])\n\nm1_ridge_engg = ridge_pred_df_engg['Pred_label from Ridge Top features - Engineered'].round(3)\nm2_ridge_engg = y_test_engg_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:\ncorr_pandas_ridge_engg = m1_ridge_engg.corr(m2_ridge_engg)\nprint(f\"Pearson Correlation for Ridge from Ridge's best features - Engineered: {corr_pandas_ridge_engg}\")\n\n#Accuracy as per RMSE\nrmse_ridge_engg = np.sqrt(mean_squared_error(m1_ridge_engg,m2_ridge_engg))\nprint(f\"Root Mean Squared Error for Ridge from Ridge's best features - Engineered: {rmse_ridge_engg:.4f}\")\n\n#Accuracy as per R2 Score\nr2_ridge_engg = r2_score(m1_ridge_engg,m2_ridge_engg)\nprint(f\"R2 for Ridge from Ridge's best features - Engineered: {r2_ridge_engg:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:18:26.141397Z","iopub.execute_input":"2025-12-01T01:18:26.141879Z","iopub.status.idle":"2025-12-01T01:19:44.777116Z","shell.execute_reply.started":"2025-12-01T01:18:26.141856Z","shell.execute_reply":"2025-12-01T01:19:44.776076Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.3.4.6 EXAMINING FEATURE IMPORTANCE THROUGH ELASTIC REGRESSION ON ENGINEERED DATASET**","metadata":{}},{"cell_type":"code","source":"elastic_FeatureAnalyzer_instance_engg=elastic_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI, y_train_engg_processed, cv=5, max_iter=1000, l1_ratio=0.5)\nelastic_FeatureAnalyzer_instance_engg.elastic_plotting()\nx_train_engg_dropped_variance_pearson_MI_elastic, Imp_features_elastic_engg,model_elastic_engg=elastic_FeatureAnalyzer_instance_engg.elastic_feature_drop()\nx_test_engg_dropped_variance_pearson_MI_elastic=x_test_engg_dropped_variance_pearson_MI[x_train_engg_dropped_variance_pearson_MI_elastic.columns]\n\n#Elastic Net model trained on Best Features - engineered basis Ridge coefficients \nelastic_FeatureAnalyzer_instance_two_engg=elastic_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI_elastic, y_train_engg_processed, cv=5, max_iter=1000, l1_ratio=0.5)\nx_train_engg_dropped_variance_pearson_MI_elastic_two, Imp_features_elastic_engg_two, model_elastic_engg_two=elastic_FeatureAnalyzer_instance_two_engg.elastic_feature_drop()\n\n#Predictions from Elastic Net - Best features engineered \nelastic_pred_engg=model_elastic_engg_two.predict(x_test_engg_dropped_variance_pearson_MI_elastic)\nelastic_pred_df_engg=pd.DataFrame(elastic_pred_engg,columns=['Pred_label from Elastic top features - Engineered'])\n\nm1_elastic_engg = elastic_pred_df_engg['Pred_label from Elastic top features - Engineered'].round(3)\nm2_elastic_engg = y_test_engg_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:\ncorr_pandas_elastic_engg = m1_elastic_engg.corr(m2_elastic_engg)\nprint(f\"Pearson Correlation for Elastic from Elastic's best features - Engineered: {corr_pandas_elastic_engg}\")\n\n#Accuracy as per RMSE\nrmse_elastic_engg = np.sqrt(mean_squared_error(m1_elastic_engg,m2_elastic_engg))\nprint(f\"Root Mean Squared Error for elastic from elastic's best features - Engineered: {rmse_elastic_engg:.4f}\")\n\n#Accuracy as per R2 Score\nr2_elastic_engg = r2_score(m1_elastic_engg,m2_elastic_engg)\nprint(f\"R2 for Elastic from Elastic's best features - Engineered: {r2_elastic_engg:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:21:38.204990Z","iopub.execute_input":"2025-12-01T01:21:38.205551Z","iopub.status.idle":"2025-12-01T01:21:41.710122Z","shell.execute_reply.started":"2025-12-01T01:21:38.205528Z","shell.execute_reply":"2025-12-01T01:21:41.709260Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.3.4.7 EXAMINING FEATURE IMPORTANCE THROUGH LGBM REGRESSOR ON ENGINEERED FEATURES**","metadata":{}},{"cell_type":"code","source":"remove_post_LGBM_instance_engg=LGBM_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI, y_train_engg_processed, label_col='label', threshold=30, random_state=42)\nremove_post_LGBM_instance_engg.LGBM_plotting()\n#Inference:\n#1.The line chart depicts the elbow point at somewhere around LGBM importance score of 30, hence we can keep the threshold at 30.\nabove_threshold_features_LGBM_engg,x_train_engg_dropped_variance_pearson_MI_LGBM, model_LGBM_engg = remove_post_LGBM_instance_engg.above_threshold_features_LGBM()\nx_test_engg_dropped_variance_pearson_MI_LGBM=x_test_engg_dropped_variance_pearson_MI[x_train_engg_dropped_variance_pearson_MI_LGBM.columns]\n\n#Training the LGBM regressor on top features as per LGBM\nLGBM_FeatureAnalyzer_instance_engg_two=LGBM_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI_LGBM, y_train_engg_processed, label_col='label', threshold=30, random_state=42)\nx_train_engg_dropped_variance_pearson_MI_LGBM_two, Imp_features_LGBM_two, model_LGBM_engg_two=LGBM_FeatureAnalyzer_instance_engg_two.above_threshold_features_LGBM()\n\n#Predictions from LGBM \nLGBM_pred_engg=model_LGBM_engg_two.predict(x_test_engg_dropped_variance_pearson_MI_LGBM)\nLGBM_pred_df_engg=pd.DataFrame(LGBM_pred_engg,columns=['Pred_label from LGBM top features - Engineered'])\n\nm1_LGBM_engg = LGBM_pred_df_engg['Pred_label from LGBM top features - Engineered'].round(3)\nm2_LGBM_engg = y_test_engg_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation on Best Features - LGBM:\ncorr_pandas_LGBM_engg = m1_LGBM_engg.corr(m2_LGBM_engg)\nprint(f\"Pearson Correlation for LGBM Regressor from LGBM's best features - Engineered: {corr_pandas_LGBM_engg}\")\n\n#Accuracy as per RMSE\nrmse_LGBM_engg = np.sqrt(mean_squared_error(m1_LGBM_engg,m2_LGBM_engg))\nprint(f\"Root Mean Squared Error for LGBM Regressor from LGBM's best features - Engineered: {rmse_LGBM_engg:.4f}\")\n\n#Accuracy as per R2 score\nr2_LGBM_engg = r2_score(m1_LGBM_engg,m2_LGBM_engg)\nprint(f\"R2 Score for LGBM Regressor from LGBM's best features - Engineered: {r2_LGBM_engg:.4f}\")\n\n#Accuracy as per R2 Score\nr2_LGBM_engg = r2_score(m1_LGBM_engg,m2_LGBM_engg)\nprint(f\"R2 for LGBM from LGBM's best features - Engineered: {r2_LGBM_engg:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:21:52.951993Z","iopub.execute_input":"2025-12-01T01:21:52.952499Z","iopub.status.idle":"2025-12-01T01:21:58.451058Z","shell.execute_reply.started":"2025-12-01T01:21:52.952478Z","shell.execute_reply":"2025-12-01T01:21:58.450316Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.3.4.8 EXAMINING FEATURE IMPORTANCE THROUGH RANDOM FOREST REGRESSOR ON ENGINEERED FEATURES**","metadata":{}},{"cell_type":"code","source":"RF_FeatureAnalyzer_instance_engg=RF_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI, y_train_engg_processed, rows_to_include=50000,label_col='label', threshold=0.015, **tuned_hyperparameters)\nRF_FeatureAnalyzer_instance_engg.RF_plotting()\n#Inference:\n#1.The line chart depicts the elbow point at somewhere around 0.015 RF importance score, hence we can keep the threshold at 0.015.\nabove_threshold_features_RF_engg,x_train_engg_dropped_variance_pearson_MI_RF, model_RF_engg=RF_FeatureAnalyzer_instance_engg.above_threshold_features()\nx_test_engg_dropped_variance_pearson_MI_RF=x_test_engg_dropped_variance_pearson_MI[x_train_engg_dropped_variance_pearson_MI_RF.columns]\n\n#Training the RF regressor on top features - engineered as per RF\nRF_FeatureAnalyzer_instance_engg_two=RF_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI_RF, y_train_engg_processed, rows_to_include=50000,label_col='label', threshold=0.015, **tuned_hyperparameters)\nx_train_engg_dropped_variance_pearson_MI_RF_two, Imp_features_RF_engg_two, model_RF_engg_two=RF_FeatureAnalyzer_instance_engg_two.above_threshold_features()\n\n#Predictions from lasso \nRF_pred_engg=model_RF_engg_two.predict(x_test_engg_dropped_variance_pearson_MI_RF)\nRF_pred_engg_df=pd.DataFrame(RF_pred_engg,columns=['Pred_label from RF top features - Engineered'])\n\nm1_RF_engg = RF_pred_engg_df['Pred_label from RF top features - Engineered'].round(3)\nm2_RF_engg = y_test_engg_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:\ncorr_pandas_RF_engg = m1_RF_engg.corr(m2_RF_engg)\nprint(f\"Pearson Correlation for Random Forest from RF's best features - Engineered: {corr_pandas_RF_engg}\")\n\n\n#Accuracy as per RMSE\nrmse_RF_engg = np.sqrt(mean_squared_error(m1_RF_engg,m2_RF_engg))\nprint(f\"Root Mean Squared Error for Random Forest from RF's best features - Engineered: {rmse_RF_engg:.4f}\")\n\n#Accuracy as per R2 Score\nr2_RF_engg = r2_score(m1_RF_engg,m2_RF_engg)\nprint(f\"R2 Score for Random Forest from RF's best features - Engineered: {r2_RF_engg:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:22:33.668169Z","iopub.execute_input":"2025-12-01T01:22:33.668770Z","iopub.status.idle":"2025-12-01T01:24:08.906576Z","shell.execute_reply.started":"2025-12-01T01:22:33.668739Z","shell.execute_reply":"2025-12-01T01:24:08.905774Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.3.4.9 EXAMINING FEATURE IMPORTANCE THROUGH XG BOOST REGRESSOR ON ENGINEERED FEATURES**","metadata":{}},{"cell_type":"code","source":"XG_FeatureAnalyzer_instance_engg=XG_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI, y_train_engg_processed, label_col='label', threshold=0.016, random_state=42, n_estimators=2000, learning_rate=0.05, max_depth=5, gamma=0.1, subsample=0.8, colsample_bytree=0.7, n_jobs=-1)\nXG_FeatureAnalyzer_instance_engg.XG_plotting()\n#Inference:\n#1.The line chart depicts the elbow point at somewhere around LGBM importance score of 30, hence we can keep the threshold at 0.015.\nabove_threshold_features_XG_engg,x_train_engg_dropped_variance_pearson_MI_XG, model_XG_engg =XG_FeatureAnalyzer_instance_engg.above_threshold_features_XG()\nx_test_engg_dropped_variance_pearson_MI_XG=x_test_engg_dropped_variance_pearson_MI[x_train_engg_dropped_variance_pearson_MI_XG.columns]\n\n#Training the XG regressor on top features as per XG - Engineered\nXG_FeatureAnalyzer_instance_engg_two=XG_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI_XG, y_train_engg_processed, label_col='label', threshold=0.015, random_state=42)\nx_train_engg_dropped_variance_pearson_MI_XG_two, Imp_features_XG_engg_two, model_XG_engg_two=XG_FeatureAnalyzer_instance_engg_two.above_threshold_features_XG()\n\n#Predictions from XG top features - Engineered\nXG_pred_engg=model_XG_engg_two.predict(x_test_engg_dropped_variance_pearson_MI_XG)\nXG_pred_engg_df=pd.DataFrame(XG_pred_engg,columns=['Pred_label from XG top features - Engineered'])\n\nm1_XG_engg = XG_pred_engg_df['Pred_label from XG top features - Engineered'].round(3)\nm2_XG_engg = y_test_engg_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:\ncorr_pandas_XG_engg = m1_XG_engg.corr(m2_XG_engg)\nprint(f\"Pearson Correlation for XG Regressor from XG's best features - Engineered: {corr_pandas_XG_engg}\")\n\n#Accuracy as per RMSE\nrmse_XG_engg = np.sqrt(mean_squared_error(m1_XG_engg,m2_XG_engg))\nprint(f\"Root Mean Squared Error for XG Regressor from XG's best features - original: {rmse_XG_engg:.4f}\")\n\n#Accuracy as per R2\nr2_XG_engg = r2_score(m1_XG_engg,m2_XG_engg)\nprint(f\"R2 Score for XG Regressor from XG's best features - original: {r2_XG_engg:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:24:28.029472Z","iopub.execute_input":"2025-12-01T01:24:28.029796Z","iopub.status.idle":"2025-12-01T01:24:49.940359Z","shell.execute_reply.started":"2025-12-01T01:24:28.029773Z","shell.execute_reply":"2025-12-01T01:24:49.939645Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7.3.4.5 EXAMINING FEATURE IMPORTANCE THROUGH ANN REGRESSOR ON ENGINEERED FEATURES**","metadata":{}},{"cell_type":"code","source":"ANN_FeatureAnalyzer_instance_engg=ANN_FeatureAnalyzer(x_train_engg_dropped_variance_pearson_MI, y_train_engg_processed, x_test_engg_dropped_variance_pearson_MI, y_test_engg_processed, label_col='label', threshold=-0.025, timestamps_num=10000, activation='relu', output_activation='linear', optimizer='adam', loss='mse', metrics=['mae'], monitor='val_loss', patience=15, restore_best_weights=True, verbose=0, validation_split=0.2, epochs=100, batch_size=256, squared=None, n_repeats=3, random_state=42, n_jobs=-1, scoring='r2')\nabove_threshold_features_ANN_engg, x_train_reduced_post_pearson_MI_ANN_engg, model_ANN_engg = ANN_FeatureAnalyzer_instance_engg.above_threshold_features_ANN() ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:26:01.656981Z","iopub.execute_input":"2025-12-01T01:26:01.657265Z","iopub.status.idle":"2025-12-01T01:26:58.787496Z","shell.execute_reply.started":"2025-12-01T01:26:01.657249Z","shell.execute_reply":"2025-12-01T01:26:58.786721Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"x_test_reduced_post_pearson_MI_ANN_engg=x_test_engg_dropped_variance_pearson_MI[x_train_reduced_post_pearson_MI_ANN_engg.columns]\n\nANN_FeatureAnalyzer_instance_engg.fit(X_train_data=x_train_reduced_post_pearson_MI_ANN_engg, y_train_data=ANN_FeatureAnalyzer_instance_engg.y_train_df)\n\n# Get the new, correctly-sized model object. \n# This variable (model_ANN_reduced_engg) is guaranteed to be the 39-feature model.\nmodel_ANN_reduced_engg = ANN_FeatureAnalyzer_instance_engg.ann_model\n\n#Predictions from ANN\nANN_pred_engg=model_ANN_reduced_engg.predict(x_test_reduced_post_pearson_MI_ANN_engg)\nANN_pred_engg_df=pd.DataFrame(ANN_pred_engg,columns=['Predictions from ANN top features - Engineered'])\n\nm1_ANN_engg = ANN_pred_engg_df['Predictions from ANN top features - Engineered'].round(3)\nm2_ANN_engg = y_test_engg_processed['label'].reset_index(drop=True).round(3)\n\n#Accuracy as per Pearson Correlation:ed\ncorr_pandas_ANN_engg = m1_ANN_engg.corr(m2_ANN_engg)\nprint(f\"Pearson Correlation from ANN's best features - Engineered: {corr_pandas_ANN_engg}\")\n\n#Accuracy as per RMSE\nrmse_ANN_engg = np.sqrt(mean_squared_error(m1_ANN_engg, m2_ANN_engg))\nprint(f\"Root Mean Squared Error from ANN's best features - Engineered: {rmse_ANN_engg:.4f}\")\n\n#Accuracy as per R2\nr2_ANN_engg = r2_score(m1_ANN_engg, m2_ANN_engg)\nprint(f\"R2 Score from ANN's best features - Engineered: {r2_ANN_engg:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:27:11.263406Z","iopub.execute_input":"2025-12-01T01:27:11.263689Z","iopub.status.idle":"2025-12-01T01:27:22.792901Z","shell.execute_reply.started":"2025-12-01T01:27:11.263668Z","shell.execute_reply":"2025-12-01T01:27:22.792148Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"# Final Best Features \nfinal_best_df = pd.concat([x_train_engg_dropped_variance_pearson_MI_LGBM, y_train_engg_processed],axis=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:29:58.663381Z","iopub.execute_input":"2025-12-01T01:29:58.664108Z","iopub.status.idle":"2025-12-01T01:29:58.679937Z","shell.execute_reply.started":"2025-12-01T01:29:58.664083Z","shell.execute_reply":"2025-12-01T01:29:58.679177Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"final_engg_df.to_csv('submission.csv', index=False)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"x_test_DRW, y_test_DRW","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id = Section8></a>\n**8. PREPROCESSING PIPELINE ON TEST DATASET (FROM COMPETITION)**","metadata":{}},{"cell_type":"markdown","source":"The test dataset from competition - x_test_parquet needs to undergo the same transformation pipeline as on the x_train_parquet. Precisely the below key steps need to be performed on x_test_parquet:\n\n1. x_test_DRW already deduced from x_test_parquet (during our import datasets class).\n2. Transformation/preprocessing pipeline on x_test_DRW (Pipeline already trained on x_train_original). \n3. Aligning transformed x_test_DRW (x_test_DRW_processed) with reduced features of x_train_original post transformation, pearson and MI feature dropping (As features are engineered post this step) yielding x_test_DRW_processed_reduced\n4. Engineering the features in x_test_DRW_processed_reduced aligned with feature engineering on x_train_eng (rolling avg/EWM/Lagging) yielding x_test_DRW_engineered. \n5. Transformation/preprocessing pipeline on x_test_DRW_engineered (Pipeline already trained on x_train_original).\n6. Aligning transformed x_test_DRW_engineered with final best engineered features - 'x_train_non-linear_engineered'.","metadata":{}},{"cell_type":"code","source":"# Class for preprocessing of unseen test - x_test_DRW\nclass test_preprocessing_pipeline():\n    # Constructor\n    def __init__(self, x_test_df, y_test_df, x_processor, x_train_reduced, feature_engg_instance, x_processor_engg, x_train_eng_final):\n        self.x_test_df = x_test_df\n        self.y_test_df = y_test_df\n        self.x_processor = x_processor\n        self.x_train_reduced = x_train_reduced\n        self.feature_engg_instance = feature_engg_instance\n        self.x_processor_engg = x_processor_engg\n        self.x_train_eng_final = x_train_eng_final\n\n   # Function for transformation pipeline on final test data - 'x_test_DRW'\n    def transformation(self):\n        self.x_processor.dataframe = self.x_test_df.copy()\n        self.x_test_DRW_processed = self.x_processor.memory_optimize()\n\n        # Convert to float32 for RAM optimization\n        #for col in self.x_test_DRW_processed.select_dtypes(include=['float64']).columns:\n             #self.x_test_DRW_processed[col] = self.x_test_DRW_processed[col].astype('float32')\n        \n        del self.x_test_df  # free memory\n        return self.x_test_DRW_processed\n\n     # Function for feature dropping pipeline on final processed test data - 'x_test_DRW_processed'\n    def feature_drop(self):\n        self.transformation()\n        self.x_test_DRW_processed_reduced = self.x_test_DRW_processed[self.x_train_reduced.columns]\n         \n        del self.x_test_DRW_processed  # free memory\n        return self.x_test_DRW_processed_reduced\n\n     # Function for engineering pipeline on final reduced test data - 'x_test_DRW_processed_reduced'\n    def engineering(self):  \n        self.feature_drop()\n        self.feature_engg_instance.combined_df_x = self.x_test_DRW_processed_reduced\n        self.feature_engg_instance.combined_df_y = self.y_test_df\n\n        self.x_test_DRW_engineered, self.y_test_DRW_engineered = self.feature_engg_instance.final_engg_features()\n\n        del self.x_test_DRW_processed_reduced, self.y_test_df  # free memory\n        return self.x_test_DRW_engineered, self.y_test_DRW_engineered\n    \n    # Function for transformation pipeline on final engineered test data - 'x_test_DRW_engineered'\n    def transformation_engg(self): \n        self.engineering()\n        self.x_processor_engg.dataframe = self.x_test_DRW_engineered.copy()\n        self.x_test_DRW_engg_processed = self.x_processor_engg.memory_optimize()\n\n        # Convert engineered features to float32\n        for col in self.x_test_DRW_engg_processed.select_dtypes(include=['float64']).columns:\n            self.x_test_DRW_engg_processed[col] = self.x_test_DRW_engg_processed[col].astype('float32')\n\n        # Combine first 51 original features back\n        self.x_test_DRW_engg_processed = pd.concat([self.x_test_DRW_engineered.iloc[:, 0:51], self.x_test_DRW_engg_processed],axis=1)\n\n        del self.x_test_DRW_engineered  # free memory\n        return self.x_test_DRW_engg_processed\n    \n    # Function for feature dropping pipeline on final engineered and processed test data - 'x_test_DRW_engg_processed'\n    def feature_drop_engg(self):\n        self.transformation_engg()\n        self.x_test_final = self.x_test_DRW_engg_processed[self.x_train_eng_final.columns]\n\n        del self.x_test_DRW_engg_processed  # free memory\n        return self.x_test_final","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:44:54.953589Z","iopub.execute_input":"2025-12-01T01:44:54.954201Z","iopub.status.idle":"2025-12-01T01:44:54.962358Z","shell.execute_reply.started":"2025-12-01T01:44:54.954177Z","shell.execute_reply":"2025-12-01T01:44:54.961695Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_preprocessing_pipeline_instance=test_preprocessing_pipeline(x_test_DRW, y_test_DRW, x_processor, x_train_reduced_post_pearson_MI, feature_engg_instance, x_processor_engg, x_train_engg_dropped_variance_pearson_MI_LGBM)\nx_test_DRW_final = test_preprocessing_pipeline_instance.feature_drop_engg()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T02:30:26.197155Z","iopub.execute_input":"2025-12-01T02:30:26.197804Z","execution_failed":"2025-12-01T02:31:38.563Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**9. Train Test Split on Final Best Features**","metadata":{}},{"cell_type":"code","source":"DataSplitter_instance_final=DataSplitter(x_train_engg_dropped_variance_pearson_MI_LGBM,y_train_engg_processed,test_size=0.3, random_state=42, shuffle=False)\n\n#Train Test split on train.parquet dataset\nx_train_final, x_test_final, y_train_final, y_test_final=DataSplitter_instance_final.train_test_set()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T01:54:56.444291Z","iopub.execute_input":"2025-12-01T01:54:56.445089Z","iopub.status.idle":"2025-12-01T01:54:56.468147Z","shell.execute_reply.started":"2025-12-01T01:54:56.445065Z","shell.execute_reply":"2025-12-01T01:54:56.467079Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id = Section9></a>\n**10. FINAL PREDICTION**\n","metadata":{}},{"cell_type":"code","source":"#Predictions from Final Best Estimator - LGBM (trained on best top performing engineered features)\npred_final=model_LGBM_engg_two.predict(x_test_DRW_final)\npred_final_df=pd.DataFrame(pred_final,columns=['Pred_label from Top Best features'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-01T02:02:24.668519Z","iopub.execute_input":"2025-12-01T02:02:24.668933Z","iopub.status.idle":"2025-12-01T02:02:24.967124Z","shell.execute_reply.started":"2025-12-01T02:02:24.668908Z","shell.execute_reply":"2025-12-01T02:02:24.965900Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}