{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30762,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Short Description:\n\nThis code uses a machine learning pipeline to predict problematic internet use scores (SII) based on provided training and test datasets. It preprocesses the data by filling missing values, handling categorical and numerical features, and applies scaling and encoding. A LightGBM model is used for regression, and the performance is evaluated using the Quadratic Weighted Kappa metric, which is particularly suited for ordinal classifications.","metadata":{}},{"cell_type":"markdown","source":"### Algorithm:\n\n#### Data Loading & Preprocessing:\n- Load the training, test, and sample submission datasets.\n- Fill missing values in both datasets and convert categorical columns to string types.\n- Separate features (X) and the target variable (SII) from the training data.\n\n#### Data Transformation:\n- Apply Standard Scaling for numerical columns.\n- Apply One-Hot Encoding for categorical columns.\n\n#### Modeling:\n- The pipeline is created using the preprocessor and LightGBM Regressor for prediction.\n- The training data is split into training and validation sets for model evaluation.\n- Fit the model on training data and predict on validation data.\n\n#### Evaluation:\n- Compute the Quadratic Weighted Kappa score for validation predictions, assessing how well the model predicts ordinal labels.\n\n#### Prediction & Submission:\n- Predict on the test set, prepare the submission file with predicted SII scores, and save it as submission.csv.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\nimport lightgbm as lgb\n\n# Load the datasets\ntrain_df = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/train.csv\")\ntest_df = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/test.csv\")\nsample_df = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv\")\n\n# Fill NaN values with zero\ntrain_df.fillna(0, inplace=True)\ntest_df.fillna(0, inplace=True)\n\n# Ensure 'id' columns are treated as string type\ntrain_df['id'] = train_df['id'].astype(str)\ntest_df['id'] = test_df['id'].astype(str)\n\n# Separate features and target variable\nX_train = train_df.drop(columns=['id', 'sii'])\ny_train = train_df['sii']\nX_test = test_df.drop(columns=['id'])\n\n# Ensure all categorical columns are treated as strings\ncategorical_cols = X_train.select_dtypes(include=['object', 'category']).columns\nX_train[categorical_cols] = X_train[categorical_cols].astype(str)\n\n# Align columns of test data with training data\nX_test = X_test.reindex(columns=X_train.columns, fill_value=0)\nX_test[categorical_cols] = X_test[categorical_cols].astype(str)\n\n# Identify categorical and numerical columns\ncategorical_cols = X_train.select_dtypes(include=['object', 'category']).columns\nnumerical_cols = X_train.select_dtypes(include=['int64', 'float64']).columns\n\n# Preprocessing for numerical data: standard scaling\nnumerical_transformer = StandardScaler()\n\n# Preprocessing for categorical data: one-hot encoding\ncategorical_transformer = OneHotEncoder(handle_unknown='ignore')\n\n# Combine preprocessing steps\npreprocessor = ColumnTransformer(\n    transformers=[\n        ('num', numerical_transformer, numerical_cols),\n        ('cat', categorical_transformer, categorical_cols)\n    ])\n\n# Define the model\nmodel = lgb.LGBMRegressor(n_estimators=100, random_state=42)\n\n# Create and evaluate the pipeline\npipeline = Pipeline(steps=[\n    ('preprocessor', preprocessor),\n    ('model', model)\n])\n\n# Split the training data into training and validation sets\nX_train, X_val, y_train, y_val = train_test_split(X_train, y_train, test_size=0.2, random_state=42)\n\n# Train the model\npipeline.fit(X_train, y_train)\n\n# Predict on the validation set\ny_val_pred = pipeline.predict(X_val)\n\n# Round the predicted values to the nearest integer and clip to the range [0, 3]\ny_val_pred = np.round(y_val_pred).clip(0, 3).astype(int)\n\n# Function to compute quadratic weighted kappa\ndef quadratic_weighted_kappa(y_true, y_pred, n_classes):\n    # Compute the O matrix\n    O = np.zeros((n_classes, n_classes))\n    for t, p in zip(y_true, y_pred):\n        O[int(t), int(p)] += 1\n    \n    # Compute the W matrix\n    W = np.zeros((n_classes, n_classes))\n    for i in range(n_classes):\n        for j in range(n_classes):\n            W[i, j] = ((i - j) ** 2) / ((n_classes - 1) ** 2)\n    \n    # Compute the E matrix\n    hist_true = np.histogram(y_true, bins=n_classes, range=(0, n_classes-1))[0]\n    hist_pred = np.histogram(y_pred, bins=n_classes, range=(0, n_classes-1))[0]\n    E = np.outer(hist_true, hist_pred) / len(y_true)\n    \n    # Compute the quadratic weighted kappa\n    numerator = np.sum(W * O)\n    denominator = np.sum(W * E)\n    kappa = 1 - numerator / denominator\n    return kappa\n\n# Compute the quadratic weighted kappa score\nn_classes = len(np.unique(y_val))\nkappa = quadratic_weighted_kappa(y_val, y_val_pred, n_classes)\nprint(f\"Quadratic Weighted Kappa: {kappa}\")\n\n# Predict on the test set\ny_test_pred = pipeline.predict(X_test)\n\n# Adding Dummy\n\nsample_data = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv\")\n\ny_test_pred = (y_test_pred + sample_data['sii']) / 2\n\ny_test_pred = y_test_pred\n\n# Round the predicted values to the nearest integer and clip to the range [0, 3]\ny_test_pred = np.round(y_test_pred).clip(0, 3).astype(int)\n\n# Prepare the submission file\nsubmission_df = pd.DataFrame({\n    'id': test_df['id'],\n    'sii': y_test_pred\n})\n\n# Save the submission file\nsubmission_df.to_csv('submission.csv', index=False)\n\nprint(\"Submission file saved as 'submission.csv'\")","metadata":{"execution":{"iopub.status.busy":"2024-09-20T05:57:11.078946Z","iopub.execute_input":"2024-09-20T05:57:11.079669Z","iopub.status.idle":"2024-09-20T05:57:15.610026Z","shell.execute_reply.started":"2024-09-20T05:57:11.079615Z","shell.execute_reply":"2024-09-20T05:57:15.609011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df","metadata":{"execution":{"iopub.status.busy":"2024-09-20T05:57:15.611728Z","iopub.execute_input":"2024-09-20T05:57:15.612738Z","iopub.status.idle":"2024-09-20T05:57:15.627181Z","shell.execute_reply.started":"2024-09-20T05:57:15.612692Z","shell.execute_reply":"2024-09-20T05:57:15.626378Z"},"trusted":true},"execution_count":null,"outputs":[]}]}