{
  "id": 535118,
  "title": "🚀 Starter Notebook: Predicting Problematic Internet Use Scores Using LightGBM",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/535118",
  "author_name": "",
  "post_date": "2024-09-20T10:49:29.474733900Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>📔 <strong>Notebook URL:</strong> <a href=\"https://www.kaggle.com/code/mirzamilanfarabi/predicting-problematic-internet-use-with-lightgbm\" target=\"_blank\">Click Here</a></p>\n<p><strong>Description:</strong> This code uses a machine learning pipeline to predict problematic internet use scores (SII) based on provided training and test datasets. It preprocesses the data by filling missing values, handling categorical and numerical features, and applies scaling and encoding. A LightGBM model is used for regression, and the performance is evaluated using the Quadratic Weighted Kappa metric, which is particularly suited for ordinal classifications.</p>\n<p><strong>Algorithm:</strong></p>\n<p><strong>1. Data Loading:</strong></p>\n<ul>\n<li>Load the training (train.csv), test (test.csv), and sample submission (sample_submission.csv) datasets using pandas.</li>\n</ul>\n<p><strong>2. Data Preprocessing:</strong></p>\n<ul>\n<li>Fill Missing Values: Replace any NaN values in both the training and test datasets with zeros.</li>\n<li>Ensure String Type for IDs: Convert the id column in both the training and test datasets to strings.</li>\n<li>Separate Features and Target:<br>\n    - Features (X_train) are all columns in the training set except id and sii (the target).<br>\n    - Target (y_train) is the sii column in the training data.<br>\n    - For the test set (X_test), drop the id column.</li>\n</ul>\n<p><strong>3. Handle Categorical Variables:</strong></p>\n<ul>\n<li>Identify categorical columns in the training set and ensure that they are treated as string type.</li>\n<li>Align the test dataset columns to match the training set by reindexing and filling missing columns with zeros.</li>\n</ul>\n<p><strong>4. Column Identification:</strong></p>\n<ul>\n<li>Categorical Columns: Identify all columns in the training set that are categorical (string type).</li>\n<li>Numerical Columns: Identify all numerical columns in the training set.</li>\n</ul>\n<p><strong>5. Preprocessing Pipelines:</strong></p>\n<ul>\n<li>Numerical Data Transformation: Apply StandardScaler for numerical features.</li>\n<li>Categorical Data Transformation: Apply OneHotEncoder for categorical features, ignoring unknown categories during transformation.</li>\n<li>Combine Transformers: Use ColumnTransformer to combine both numerical and categorical preprocessing.</li>\n</ul>\n<p><strong>6. Model Definition:</strong></p>\n<ul>\n<li>Define the model using LightGBM Regressor with 100 estimators and a fixed random state for reproducibility.</li>\n</ul>\n<p><strong>7. Pipeline Creation:</strong></p>\n<ul>\n<li>Create a Pipeline consisting of the preprocessor (transformations for both numerical and categorical data) and the LightGBM model.</li>\n</ul>\n<p><strong>8. Train-Test Split:</strong></p>\n<ul>\n<li>Split the training data into training and validation sets using an 80-20 split for model evaluation.</li>\n</ul>\n<p><strong>9. Model Training:</strong></p>\n<ul>\n<li>Train the pipeline (preprocessing + model) on the training portion of the split data (X_train and y_train).</li>\n</ul>\n<p><strong>10. Validation Prediction:</strong></p>\n<ul>\n<li>Predict on the validation set and round the predicted values to the nearest integer.</li>\n<li>Clip the predicted values to the range [0, 3], as the target variable has integer values within this range.</li>\n</ul>\n<p><strong>11. Quadratic Weighted Kappa Calculation:</strong></p>\n<ul>\n<li>Define a function to compute the Quadratic Weighted Kappa score, which is suitable for ordinal predictions.</li>\n<li>Calculate the observed (O), expected (E), and weight (W) matrices based on the true and predicted values.<br>\n    Compute the kappa score using the formula:<br>\n    kappa=1−∑W⋅O∑W⋅E<br>\n    kappa=1−∑W⋅E∑W⋅O​</li>\n</ul>\n<p><strong>12. Evaluation:</strong></p>\n<ul>\n<li>Calculate the Quadratic Weighted Kappa score for the validation set predictions and print it.</li>\n</ul>\n<p><strong>13. Test Set Predictions:</strong></p>\n<ul>\n<li>Predict on the test set using the trained pipeline.</li>\n<li>Round and clip the test set predictions to integer values within [0, 3].</li>\n</ul>\n<p><strong>14. Submission Preparation:</strong></p>\n<ul>\n<li>Prepare a submission file containing the 'id' and predicted 'sii' values for the test data.</li>\n<li>Save the submission file as submission.csv.</li>\n</ul>\n<p><strong>Code:</strong></p>\n<pre><code> pandas  pd\n numpy  np\n sklearn.model_selection  train_test_split\n sklearn.preprocessing  StandardScaler, OneHotEncoder\n sklearn.compose  ColumnTransformer\n sklearn.pipeline  Pipeline\n lightgbm  lgb\n\n\ntrain_df = pd.read_csv()\ntest_df = pd.read_csv()\nsample_df = pd.read_csv()\n\n\ntrain_df.fillna(, inplace=)\ntest_df.fillna(, inplace=)\n\n\ntrain_df[] = train_df[].astype()\ntest_df[] = test_df[].astype()\n\n\nX_train = train_df.drop(columns=[, ])\ny_train = train_df[]\nX_test = test_df.drop(columns=[])\n\n\ncategorical_cols = X_train.select_dtypes(include=[, ]).columns\nX_train[categorical_cols] = X_train[categorical_cols].astype()\n\n\nX_test = X_test.reindex(columns=X_train.columns, fill_value=)\nX_test[categorical_cols] = X_test[categorical_cols].astype()\n\n\ncategorical_cols = X_train.select_dtypes(include=[, ]).columns\nnumerical_cols = X_train.select_dtypes(include=[, ]).columns\n\n\nnumerical_transformer = StandardScaler()\n\n\ncategorical_transformer = OneHotEncoder(handle_unknown=)\n\n\npreprocessor = ColumnTransformer(\n    transformers=[\n        (, numerical_transformer, numerical_cols),\n        (, categorical_transformer, categorical_cols)\n    ])\n\n\nmodel = lgb.LGBMRegressor(n_estimators=, random_state=)\n\n\npipeline = Pipeline(steps=[\n    (, preprocessor),\n    (, model)\n])\n\n\nX_train, X_val, y_train, y_val = train_test_split(X_train, y_train, test_size=, random_state=)\n\n\npipeline.fit(X_train, y_train)\n\n\ny_val_pred = pipeline.predict(X_val)\n\n\ny_val_pred = np.(y_val_pred).clip(, ).astype()\n\n\n ():\n    \n    O = np.zeros((n_classes, n_classes))\n     t, p  (y_true, y_pred):\n        O[(t), (p)] += \n\n    \n    W = np.zeros((n_classes, n_classes))\n     i  (n_classes):\n         j  (n_classes):\n            W[i, j] = ((i - j) ** ) / ((n_classes - ) ** )\n\n    \n    hist_true = np.histogram(y_true, bins=n_classes, =(, n_classes-))[]\n    hist_pred = np.histogram(y_pred, bins=n_classes, =(, n_classes-))[]\n    E = np.outer(hist_true, hist_pred) / (y_true)\n\n    \n    numerator = np.(W * O)\n    denominator = np.(W * E)\n    kappa =  - numerator / denominator\n     kappa\n\n\nn_classes = (np.unique(y_val))\nkappa = quadratic_weighted_kappa(y_val, y_val_pred, n_classes)\n()\n\n\ny_test_pred = pipeline.predict(X_test)\n\n\ny_test_pred = np.(y_test_pred).clip(, ).astype()\n\n\nsubmission_df = pd.DataFrame({\n    : test_df[],\n    : y_test_pred\n})\n\n\nsubmission_df.to_csv(, index=)\n\n()\n</code></pre>\n<pre><code>submission_df\n</code></pre>",
  "messages": [
    {
      "id": "2993910",
      "postDate": "09/20/2024 10:49:29",
      "content": "<p>📔 <strong>Notebook URL:</strong> <a href=\"https://www.kaggle.com/code/mirzamilanfarabi/predicting-problematic-internet-use-with-lightgbm\" target=\"_blank\">Click Here</a></p>\n<p><strong>Description:</strong> This code uses a machine learning pipeline to predict problematic internet use scores (SII) based on provided training and test datasets. It preprocesses the data by filling missing values, handling categorical and numerical features, and applies scaling and encoding. A LightGBM model is used for regression, and the performance is evaluated using the Quadratic Weighted Kappa metric, which is particularly suited for ordinal classifications.</p>\n<p><strong>Algorithm:</strong></p>\n<p><strong>1. Data Loading:</strong></p>\n<ul>\n<li>Load the training (train.csv), test (test.csv), and sample submission (sample_submission.csv) datasets using pandas.</li>\n</ul>\n<p><strong>2. Data Preprocessing:</strong></p>\n<ul>\n<li>Fill Missing Values: Replace any NaN values in both the training and test datasets with zeros.</li>\n<li>Ensure String Type for IDs: Convert the id column in both the training and test datasets to strings.</li>\n<li>Separate Features and Target:<br>\n    - Features (X_train) are all columns in the training set except id and sii (the target).<br>\n    - Target (y_train) is the sii column in the training data.<br>\n    - For the test set (X_test), drop the id column.</li>\n</ul>\n<p><strong>3. Handle Categorical Variables:</strong></p>\n<ul>\n<li>Identify categorical columns in the training set and ensure that they are treated as string type.</li>\n<li>Align the test dataset columns to match the training set by reindexing and filling missing columns with zeros.</li>\n</ul>\n<p><strong>4. Column Identification:</strong></p>\n<ul>\n<li>Categorical Columns: Identify all columns in the training set that are categorical (string type).</li>\n<li>Numerical Columns: Identify all numerical columns in the training set.</li>\n</ul>\n<p><strong>5. Preprocessing Pipelines:</strong></p>\n<ul>\n<li>Numerical Data Transformation: Apply StandardScaler for numerical features.</li>\n<li>Categorical Data Transformation: Apply OneHotEncoder for categorical features, ignoring unknown categories during transformation.</li>\n<li>Combine Transformers: Use ColumnTransformer to combine both numerical and categorical preprocessing.</li>\n</ul>\n<p><strong>6. Model Definition:</strong></p>\n<ul>\n<li>Define the model using LightGBM Regressor with 100 estimators and a fixed random state for reproducibility.</li>\n</ul>\n<p><strong>7. Pipeline Creation:</strong></p>\n<ul>\n<li>Create a Pipeline consisting of the preprocessor (transformations for both numerical and categorical data) and the LightGBM model.</li>\n</ul>\n<p><strong>8. Train-Test Split:</strong></p>\n<ul>\n<li>Split the training data into training and validation sets using an 80-20 split for model evaluation.</li>\n</ul>\n<p><strong>9. Model Training:</strong></p>\n<ul>\n<li>Train the pipeline (preprocessing + model) on the training portion of the split data (X_train and y_train).</li>\n</ul>\n<p><strong>10. Validation Prediction:</strong></p>\n<ul>\n<li>Predict on the validation set and round the predicted values to the nearest integer.</li>\n<li>Clip the predicted values to the range [0, 3], as the target variable has integer values within this range.</li>\n</ul>\n<p><strong>11. Quadratic Weighted Kappa Calculation:</strong></p>\n<ul>\n<li>Define a function to compute the Quadratic Weighted Kappa score, which is suitable for ordinal predictions.</li>\n<li>Calculate the observed (O), expected (E), and weight (W) matrices based on the true and predicted values.<br>\n    Compute the kappa score using the formula:<br>\n    kappa=1−∑W⋅O∑W⋅E<br>\n    kappa=1−∑W⋅E∑W⋅O​</li>\n</ul>\n<p><strong>12. Evaluation:</strong></p>\n<ul>\n<li>Calculate the Quadratic Weighted Kappa score for the validation set predictions and print it.</li>\n</ul>\n<p><strong>13. Test Set Predictions:</strong></p>\n<ul>\n<li>Predict on the test set using the trained pipeline.</li>\n<li>Round and clip the test set predictions to integer values within [0, 3].</li>\n</ul>\n<p><strong>14. Submission Preparation:</strong></p>\n<ul>\n<li>Prepare a submission file containing the 'id' and predicted 'sii' values for the test data.</li>\n<li>Save the submission file as submission.csv.</li>\n</ul>\n<p><strong>Code:</strong></p>\n<pre><code> pandas  pd\n numpy  np\n sklearn.model_selection  train_test_split\n sklearn.preprocessing  StandardScaler, OneHotEncoder\n sklearn.compose  ColumnTransformer\n sklearn.pipeline  Pipeline\n lightgbm  lgb\n\n\ntrain_df = pd.read_csv()\ntest_df = pd.read_csv()\nsample_df = pd.read_csv()\n\n\ntrain_df.fillna(, inplace=)\ntest_df.fillna(, inplace=)\n\n\ntrain_df[] = train_df[].astype()\ntest_df[] = test_df[].astype()\n\n\nX_train = train_df.drop(columns=[, ])\ny_train = train_df[]\nX_test = test_df.drop(columns=[])\n\n\ncategorical_cols = X_train.select_dtypes(include=[, ]).columns\nX_train[categorical_cols] = X_train[categorical_cols].astype()\n\n\nX_test = X_test.reindex(columns=X_train.columns, fill_value=)\nX_test[categorical_cols] = X_test[categorical_cols].astype()\n\n\ncategorical_cols = X_train.select_dtypes(include=[, ]).columns\nnumerical_cols = X_train.select_dtypes(include=[, ]).columns\n\n\nnumerical_transformer = StandardScaler()\n\n\ncategorical_transformer = OneHotEncoder(handle_unknown=)\n\n\npreprocessor = ColumnTransformer(\n    transformers=[\n        (, numerical_transformer, numerical_cols),\n        (, categorical_transformer, categorical_cols)\n    ])\n\n\nmodel = lgb.LGBMRegressor(n_estimators=, random_state=)\n\n\npipeline = Pipeline(steps=[\n    (, preprocessor),\n    (, model)\n])\n\n\nX_train, X_val, y_train, y_val = train_test_split(X_train, y_train, test_size=, random_state=)\n\n\npipeline.fit(X_train, y_train)\n\n\ny_val_pred = pipeline.predict(X_val)\n\n\ny_val_pred = np.(y_val_pred).clip(, ).astype()\n\n\n ():\n    \n    O = np.zeros((n_classes, n_classes))\n     t, p  (y_true, y_pred):\n        O[(t), (p)] += \n\n    \n    W = np.zeros((n_classes, n_classes))\n     i  (n_classes):\n         j  (n_classes):\n            W[i, j] = ((i - j) ** ) / ((n_classes - ) ** )\n\n    \n    hist_true = np.histogram(y_true, bins=n_classes, =(, n_classes-))[]\n    hist_pred = np.histogram(y_pred, bins=n_classes, =(, n_classes-))[]\n    E = np.outer(hist_true, hist_pred) / (y_true)\n\n    \n    numerator = np.(W * O)\n    denominator = np.(W * E)\n    kappa =  - numerator / denominator\n     kappa\n\n\nn_classes = (np.unique(y_val))\nkappa = quadratic_weighted_kappa(y_val, y_val_pred, n_classes)\n()\n\n\ny_test_pred = pipeline.predict(X_test)\n\n\ny_test_pred = np.(y_test_pred).clip(, ).astype()\n\n\nsubmission_df = pd.DataFrame({\n    : test_df[],\n    : y_test_pred\n})\n\n\nsubmission_df.to_csv(, index=)\n\n()\n</code></pre>\n<pre><code>submission_df\n</code></pre>",
      "rawMarkdown": "📔 **Notebook URL:** [Click Here](https://www.kaggle.com/code/mirzamilanfarabi/predicting-problematic-internet-use-with-lightgbm)\n\n**Description:** This code uses a machine learning pipeline to predict problematic internet use scores (SII) based on provided training and test datasets. It preprocesses the data by filling missing values, handling categorical and numerical features, and applies scaling and encoding. A LightGBM model is used for regression, and the performance is evaluated using the Quadratic Weighted Kappa metric, which is particularly suited for ordinal classifications.\n\n**Algorithm:**\n\n**1. Data Loading:**\n- Load the training (train.csv), test (test.csv), and sample submission (sample_submission.csv) datasets using pandas.\n\n**2. Data Preprocessing:**\n- Fill Missing Values: Replace any NaN values in both the training and test datasets with zeros.\n- Ensure String Type for IDs: Convert the id column in both the training and test datasets to strings.\n- Separate Features and Target:\n        - Features (X_train) are all columns in the training set except id and sii (the target).\n        - Target (y_train) is the sii column in the training data.\n        - For the test set (X_test), drop the id column.\n\n**3. Handle Categorical Variables:**\n- Identify categorical columns in the training set and ensure that they are treated as string type.\n- Align the test dataset columns to match the training set by reindexing and filling missing columns with zeros.\n\n**4. Column Identification:**\n- Categorical Columns: Identify all columns in the training set that are categorical (string type).\n- Numerical Columns: Identify all numerical columns in the training set.\n\n**5. Preprocessing Pipelines:**\n- Numerical Data Transformation: Apply StandardScaler for numerical features.\n- Categorical Data Transformation: Apply OneHotEncoder for categorical features, ignoring unknown categories during transformation.\n- Combine Transformers: Use ColumnTransformer to combine both numerical and categorical preprocessing.\n\n**6. Model Definition:**\n- Define the model using LightGBM Regressor with 100 estimators and a fixed random state for reproducibility.\n\n**7. Pipeline Creation:**\n- Create a Pipeline consisting of the preprocessor (transformations for both numerical and categorical data) and the LightGBM model.\n\n**8. Train-Test Split:**\n- Split the training data into training and validation sets using an 80-20 split for model evaluation.\n\n**9. Model Training:**\n- Train the pipeline (preprocessing + model) on the training portion of the split data (X_train and y_train).\n\n**10. Validation Prediction:**\n- Predict on the validation set and round the predicted values to the nearest integer.\n- Clip the predicted values to the range [0, 3], as the target variable has integer values within this range.\n\n**11. Quadratic Weighted Kappa Calculation:**\n- Define a function to compute the Quadratic Weighted Kappa score, which is suitable for ordinal predictions.\n- Calculate the observed (O), expected (E), and weight (W) matrices based on the true and predicted values.\n        Compute the kappa score using the formula:\n        kappa=1−∑W⋅O∑W⋅E\n        kappa=1−∑W⋅E∑W⋅O​\n\n**12. Evaluation:**\n- Calculate the Quadratic Weighted Kappa score for the validation set predictions and print it.\n\n**13. Test Set Predictions:**\n- Predict on the test set using the trained pipeline.\n- Round and clip the test set predictions to integer values within [0, 3].\n\n**14. Submission Preparation:**\n- Prepare a submission file containing the 'id' and predicted 'sii' values for the test data.\n- Save the submission file as submission.csv.\n\n**Code:**\n```python\nimport pandas as pd\nimport numpy as np\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\nimport lightgbm as lgb\n\n# Load the datasets\ntrain_df = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/train.csv\")\ntest_df = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/test.csv\")\nsample_df = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv\")\n\n# Fill NaN values with zero\ntrain_df.fillna(0, inplace=True)\ntest_df.fillna(0, inplace=True)\n\n# Ensure 'id' columns are treated as string type\ntrain_df['id'] = train_df['id'].astype(str)\ntest_df['id'] = test_df['id'].astype(str)\n\n# Separate features and target variable\nX_train = train_df.drop(columns=['id', 'sii'])\ny_train = train_df['sii']\nX_test = test_df.drop(columns=['id'])\n\n# Ensure all categorical columns are treated as strings\ncategorical_cols = X_train.select_dtypes(include=['object', 'category']).columns\nX_train[categorical_cols] = X_train[categorical_cols].astype(str)\n\n# Align columns of test data with training data\nX_test = X_test.reindex(columns=X_train.columns, fill_value=0)\nX_test[categorical_cols] = X_test[categorical_cols].astype(str)\n\n# Identify categorical and numerical columns\ncategorical_cols = X_train.select_dtypes(include=['object', 'category']).columns\nnumerical_cols = X_train.select_dtypes(include=['int64', 'float64']).columns\n\n# Preprocessing for numerical data: standard scaling\nnumerical_transformer = StandardScaler()\n\n# Preprocessing for categorical data: one-hot encoding\ncategorical_transformer = OneHotEncoder(handle_unknown='ignore')\n\n# Combine preprocessing steps\npreprocessor = ColumnTransformer(\n    transformers=[\n        ('num', numerical_transformer, numerical_cols),\n        ('cat', categorical_transformer, categorical_cols)\n    ])\n\n# Define the model\nmodel = lgb.LGBMRegressor(n_estimators=100, random_state=42)\n\n# Create and evaluate the pipeline\npipeline = Pipeline(steps=[\n    ('preprocessor', preprocessor),\n    ('model', model)\n])\n\n# Split the training data into training and validation sets\nX_train, X_val, y_train, y_val = train_test_split(X_train, y_train, test_size=0.2, random_state=42)\n\n# Train the model\npipeline.fit(X_train, y_train)\n\n# Predict on the validation set\ny_val_pred = pipeline.predict(X_val)\n\n# Round the predicted values to the nearest integer and clip to the range [0, 3]\ny_val_pred = np.round(y_val_pred).clip(0, 3).astype(int)\n\n# Function to compute quadratic weighted kappa\ndef quadratic_weighted_kappa(y_true, y_pred, n_classes):\n    # Compute the O matrix\n    O = np.zeros((n_classes, n_classes))\n    for t, p in zip(y_true, y_pred):\n        O[int(t), int(p)] += 1\n    \n    # Compute the W matrix\n    W = np.zeros((n_classes, n_classes))\n    for i in range(n_classes):\n        for j in range(n_classes):\n            W[i, j] = ((i - j) ** 2) / ((n_classes - 1) ** 2)\n    \n    # Compute the E matrix\n    hist_true = np.histogram(y_true, bins=n_classes, range=(0, n_classes-1))[0]\n    hist_pred = np.histogram(y_pred, bins=n_classes, range=(0, n_classes-1))[0]\n    E = np.outer(hist_true, hist_pred) / len(y_true)\n    \n    # Compute the quadratic weighted kappa\n    numerator = np.sum(W * O)\n    denominator = np.sum(W * E)\n    kappa = 1 - numerator / denominator\n    return kappa\n\n# Compute the quadratic weighted kappa score\nn_classes = len(np.unique(y_val))\nkappa = quadratic_weighted_kappa(y_val, y_val_pred, n_classes)\nprint(f\"Quadratic Weighted Kappa: {kappa}\")\n\n# Predict on the test set\ny_test_pred = pipeline.predict(X_test)\n\n# Round the predicted values to the nearest integer and clip to the range [0, 3]\ny_test_pred = np.round(y_test_pred).clip(0, 3).astype(int)\n\n# Prepare the submission file\nsubmission_df = pd.DataFrame({\n    'id': test_df['id'],\n    'sii': y_test_pred\n})\n\n# Save the submission file\nsubmission_df.to_csv('submission.csv', index=False)\n\nprint(\"Submission file saved as 'submission.csv'\")\n```\n\n```python\nsubmission_df\n```",
      "votes": null
    },
    {
      "id": "3008978",
      "postDate": "10/07/2024 11:09:41",
      "content": "<p>very useful algorithm, thanks!</p>",
      "rawMarkdown": "very useful algorithm, thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3008978,
      "author_name": "oksanatemruk",
      "author_url": "",
      "post_date": "10/07/2024 11:09:41",
      "content": "<p>very useful algorithm, thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2993910": "📔 **Notebook URL:** [Click Here](https://www.kaggle.com/code/mirzamilanfarabi/predicting-problematic-internet-use-with-lightgbm)\n\n**Description:** This code uses a machine learning pipeline to predict problematic internet use scores (SII) based on provided training and test datasets. It preprocesses the data by filling missing values, handling categorical and numerical features, and applies scaling and encoding. A LightGBM model is used for regression, and the performance is evaluated using the Quadratic Weighted Kappa metric, which is particularly suited for ordinal classifications.\n\n**Algorithm:**\n\n**1. Data Loading:**\n- Load the training (train.csv), test (test.csv), and sample submission (sample_submission.csv) datasets using pandas.\n\n**2. Data Preprocessing:**\n- Fill Missing Values: Replace any NaN values in both the training and test datasets with zeros.\n- Ensure String Type for IDs: Convert the id column in both the training and test datasets to strings.\n- Separate Features and Target:\n        - Features (X_train) are all columns in the training set except id and sii (the target).\n        - Target (y_train) is the sii column in the training data.\n        - For the test set (X_test), drop the id column.\n\n**3. Handle Categorical Variables:**\n- Identify categorical columns in the training set and ensure that they are treated as string type.\n- Align the test dataset columns to match the training set by reindexing and filling missing columns with zeros.\n\n**4. Column Identification:**\n- Categorical Columns: Identify all columns in the training set that are categorical (string type).\n- Numerical Columns: Identify all numerical columns in the training set.\n\n**5. Preprocessing Pipelines:**\n- Numerical Data Transformation: Apply StandardScaler for numerical features.\n- Categorical Data Transformation: Apply OneHotEncoder for categorical features, ignoring unknown categories during transformation.\n- Combine Transformers: Use ColumnTransformer to combine both numerical and categorical preprocessing.\n\n**6. Model Definition:**\n- Define the model using LightGBM Regressor with 100 estimators and a fixed random state for reproducibility.\n\n**7. Pipeline Creation:**\n- Create a Pipeline consisting of the preprocessor (transformations for both numerical and categorical data) and the LightGBM model.\n\n**8. Train-Test Split:**\n- Split the training data into training and validation sets using an 80-20 split for model evaluation.\n\n**9. Model Training:**\n- Train the pipeline (preprocessing + model) on the training portion of the split data (X_train and y_train).\n\n**10. Validation Prediction:**\n- Predict on the validation set and round the predicted values to the nearest integer.\n- Clip the predicted values to the range [0, 3], as the target variable has integer values within this range.\n\n**11. Quadratic Weighted Kappa Calculation:**\n- Define a function to compute the Quadratic Weighted Kappa score, which is suitable for ordinal predictions.\n- Calculate the observed (O), expected (E), and weight (W) matrices based on the true and predicted values.\n        Compute the kappa score using the formula:\n        kappa=1−∑W⋅O∑W⋅E\n        kappa=1−∑W⋅E∑W⋅O​\n\n**12. Evaluation:**\n- Calculate the Quadratic Weighted Kappa score for the validation set predictions and print it.\n\n**13. Test Set Predictions:**\n- Predict on the test set using the trained pipeline.\n- Round and clip the test set predictions to integer values within [0, 3].\n\n**14. Submission Preparation:**\n- Prepare a submission file containing the 'id' and predicted 'sii' values for the test data.\n- Save the submission file as submission.csv.\n\n**Code:**\n```python\nimport pandas as pd\nimport numpy as np\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\nimport lightgbm as lgb\n\n# Load the datasets\ntrain_df = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/train.csv\")\ntest_df = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/test.csv\")\nsample_df = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv\")\n\n# Fill NaN values with zero\ntrain_df.fillna(0, inplace=True)\ntest_df.fillna(0, inplace=True)\n\n# Ensure 'id' columns are treated as string type\ntrain_df['id'] = train_df['id'].astype(str)\ntest_df['id'] = test_df['id'].astype(str)\n\n# Separate features and target variable\nX_train = train_df.drop(columns=['id', 'sii'])\ny_train = train_df['sii']\nX_test = test_df.drop(columns=['id'])\n\n# Ensure all categorical columns are treated as strings\ncategorical_cols = X_train.select_dtypes(include=['object', 'category']).columns\nX_train[categorical_cols] = X_train[categorical_cols].astype(str)\n\n# Align columns of test data with training data\nX_test = X_test.reindex(columns=X_train.columns, fill_value=0)\nX_test[categorical_cols] = X_test[categorical_cols].astype(str)\n\n# Identify categorical and numerical columns\ncategorical_cols = X_train.select_dtypes(include=['object', 'category']).columns\nnumerical_cols = X_train.select_dtypes(include=['int64', 'float64']).columns\n\n# Preprocessing for numerical data: standard scaling\nnumerical_transformer = StandardScaler()\n\n# Preprocessing for categorical data: one-hot encoding\ncategorical_transformer = OneHotEncoder(handle_unknown='ignore')\n\n# Combine preprocessing steps\npreprocessor = ColumnTransformer(\n    transformers=[\n        ('num', numerical_transformer, numerical_cols),\n        ('cat', categorical_transformer, categorical_cols)\n    ])\n\n# Define the model\nmodel = lgb.LGBMRegressor(n_estimators=100, random_state=42)\n\n# Create and evaluate the pipeline\npipeline = Pipeline(steps=[\n    ('preprocessor', preprocessor),\n    ('model', model)\n])\n\n# Split the training data into training and validation sets\nX_train, X_val, y_train, y_val = train_test_split(X_train, y_train, test_size=0.2, random_state=42)\n\n# Train the model\npipeline.fit(X_train, y_train)\n\n# Predict on the validation set\ny_val_pred = pipeline.predict(X_val)\n\n# Round the predicted values to the nearest integer and clip to the range [0, 3]\ny_val_pred = np.round(y_val_pred).clip(0, 3).astype(int)\n\n# Function to compute quadratic weighted kappa\ndef quadratic_weighted_kappa(y_true, y_pred, n_classes):\n    # Compute the O matrix\n    O = np.zeros((n_classes, n_classes))\n    for t, p in zip(y_true, y_pred):\n        O[int(t), int(p)] += 1\n    \n    # Compute the W matrix\n    W = np.zeros((n_classes, n_classes))\n    for i in range(n_classes):\n        for j in range(n_classes):\n            W[i, j] = ((i - j) ** 2) / ((n_classes - 1) ** 2)\n    \n    # Compute the E matrix\n    hist_true = np.histogram(y_true, bins=n_classes, range=(0, n_classes-1))[0]\n    hist_pred = np.histogram(y_pred, bins=n_classes, range=(0, n_classes-1))[0]\n    E = np.outer(hist_true, hist_pred) / len(y_true)\n    \n    # Compute the quadratic weighted kappa\n    numerator = np.sum(W * O)\n    denominator = np.sum(W * E)\n    kappa = 1 - numerator / denominator\n    return kappa\n\n# Compute the quadratic weighted kappa score\nn_classes = len(np.unique(y_val))\nkappa = quadratic_weighted_kappa(y_val, y_val_pred, n_classes)\nprint(f\"Quadratic Weighted Kappa: {kappa}\")\n\n# Predict on the test set\ny_test_pred = pipeline.predict(X_test)\n\n# Round the predicted values to the nearest integer and clip to the range [0, 3]\ny_test_pred = np.round(y_test_pred).clip(0, 3).astype(int)\n\n# Prepare the submission file\nsubmission_df = pd.DataFrame({\n    'id': test_df['id'],\n    'sii': y_test_pred\n})\n\n# Save the submission file\nsubmission_df.to_csv('submission.csv', index=False)\n\nprint(\"Submission file saved as 'submission.csv'\")\n```\n\n```python\nsubmission_df\n```",
    "3008978": "very useful algorithm, thanks!"
  },
  "source": "meta"
}