{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30839,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\nI'll prepare a set of guidelines to help you compete effectively in this competition. These will cover data exploration, feature engineering, model selection, evaluation, and submission strategies.\n\nCompetition Guidelines: Predicting Problematic Internet Usage\n1. Understand the Problem Statement\nThe goal is to predict the level of problematic internet use (sii) in children and adolescents based on physical activity and fitness data.\nThe evaluation metric is Quadratic Weighted Kappa (QWK), which measures agreement between predicted and actual values.\nThe target variable (sii) is ordinal (ordered categories), so we must treat it appropriately.\n2. Explore the Data\nStep 1: Load and Inspect the Data\nRead the dataset and check for missing values, distributions, and correlations.\nIdentify categorical vs numerical features.\nCheck if any features are highly imbalanced or contain outliers.\nStep 2: Understand Feature Relevance\nLook at which physical fitness indicators (e.g., posture, BMI, heart rate, endurance, flexibility, etc.) are most correlated with sii.\nExplore relationships between different fitness measures and internet usage.\nUse domain knowledge (e.g., kids with poor posture or high sedentary time might have problematic internet use).\n3. Data Preprocessing\nStep 1: Handle Missing Data\nImpute missing values based on mean/median/mode for numerical data or most frequent category for categorical data.\nStep 2: Normalize & Encode Features\nOrdinal Encoding: Use ordinal encoding for ordered categorical variables.\nOne-Hot Encoding: Use one-hot encoding for non-ordered categorical features.\nScaling: Normalize numerical features (e.g., StandardScaler, MinMaxScaler) to make them comparable.\nStep 3: Feature Engineering\nDerive new features based on domain knowledge (e.g., activity-to-sedentary ratio).\nApply Principal Component Analysis (PCA) or other dimensionality reduction if needed.\n4. Model Selection & Training\nStep 1: Choose an Appropriate Model\nSince the target (sii) is ordinal, consider:\nOrdinal Regression Models: Logistic Regression with ordered classes.\nTree-Based Models: Random Forest, XGBoost, LightGBM, or CatBoost.\nDeep Learning: Neural Networks with ordinal loss functions (if sufficient data).\nStep 2: Handle Imbalanced Data\nIf sii classes are imbalanced, apply:\nClass Weighting (in models like Logistic Regression or Decision Trees).\nSMOTE (Synthetic Minority Over-sampling Technique).\nStep 3: Cross-Validation\nUse Stratified K-Fold Cross-Validation (because sii is ordinal).\nTune hyperparameters using GridSearchCV or Optuna.\n5. Optimize for Quadratic Weighted Kappa\nSince QWK is the competition’s evaluation metric, optimize directly for it.\nConvert raw model outputs to ordinal class predictions to maximize QWK.\nTry Regression + Round Strategy:\nTrain a regression model and round the predictions to the nearest class.\n6. Submission Strategy\nStep 1: Format Predictions Correctly\nEnsure the submission file is structured as:\nbash\nCopy code\nid,sii\n000046df,0\n000089ff,1\n00012558,2\n00017ccd,3\nUse integer values only for sii.\nStep 2: Post-Processing\nIf the model outputs probabilities, use argmax to pick the most likely sii class.\nConsider using calibration techniques (Platt Scaling, Isotonic Regression) to refine predictions.\nStep 3: Ensemble Models\nBlend multiple models (e.g., XGBoost + Neural Network + Random Forest) for better performance.\nUse weighted averaging or stacking.\n7. Monitor and Improve\nAnalyze which features contribute the most (e.g., SHAP values, feature importance).\nLook for data leakage (i.e., features that give away the target).\nPerform error analysis to understand where the model is misclassifying.\nNext Steps\nI’ll now inspect the data dictionary from the CSV you uploaded to better understand the available features. Let me analyze it and display relevant information.\nThe data dictionary provides the following key details:\n\nColumns include Instrument, Field, Description, Type, and possible Values.\nFeatures are categorized into groups such as:\nDemographics (e.g., age, sex, enrollment season)\nChildren's Global Assessment Scale (CGAS)\nPhysical Activity & Fitness Metrics (likely critical for predictions)\nNext Steps\nI will now:\n\nExtract relevant physical fitness features from the dictionary.\nIdentify the target variable (sii).\nSummarize the types of data available.\nLet me process the data dictionary further. ​​\n\nI have extracted and displayed the physical activity & fitness features along with any details on the target variable (sii) for your review. Let me know if you need further analysis or if you want me to suggest specific feature engineering strategies based on these variables. 🚀 ​","metadata":{}},{"cell_type":"code","source":"#installing all the necessary packages \nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns \nfrom sklearn.model_selection import train_test_split \nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import classification_report, confusion_matrix\nfrom sklearn.preprocessing import StandardScaler  \nfrom sklearn.decomposition import PCA #Principle component analysis to improve model performance ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:35:59.498967Z","iopub.execute_input":"2025-01-31T14:35:59.499351Z","iopub.status.idle":"2025-01-31T14:36:03.012411Z","shell.execute_reply.started":"2025-01-31T14:35:59.499311Z","shell.execute_reply":"2025-01-31T14:36:03.011307Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load datasets\n\ntrain_data = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest_data = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\ntrain_data.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:36:06.093286Z","iopub.execute_input":"2025-01-31T14:36:06.093625Z","iopub.status.idle":"2025-01-31T14:36:06.227323Z","shell.execute_reply.started":"2025-01-31T14:36:06.093600Z","shell.execute_reply":"2025-01-31T14:36:06.225973Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data.describe().transpose()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:36:11.813599Z","iopub.execute_input":"2025-01-31T14:36:11.814058Z","iopub.status.idle":"2025-01-31T14:36:11.955567Z","shell.execute_reply.started":"2025-01-31T14:36:11.814027Z","shell.execute_reply":"2025-01-31T14:36:11.954558Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:36:16.152313Z","iopub.execute_input":"2025-01-31T14:36:16.152636Z","iopub.status.idle":"2025-01-31T14:36:16.181895Z","shell.execute_reply.started":"2025-01-31T14:36:16.152613Z","shell.execute_reply":"2025-01-31T14:36:16.180823Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data['sii'].value_counts()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:36:28.557525Z","iopub.execute_input":"2025-01-31T14:36:28.557997Z","iopub.status.idle":"2025-01-31T14:36:28.567580Z","shell.execute_reply.started":"2025-01-31T14:36:28.557962Z","shell.execute_reply":"2025-01-31T14:36:28.566591Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#select columns with more than 50% non null values and filling the missing values to ensure that the selected columns are the most accurate \nthreshold = 0.5 * len(train_data)\ncolumns_with_data = train_data.columns[train_data.isnull().sum() < threshold]\ntrain_data = train_data[columns_with_data]\n#replace all missing values with 0 \ntrain_data = train_data.fillna(0)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:37:40.880237Z","iopub.execute_input":"2025-01-31T14:37:40.880569Z","iopub.status.idle":"2025-01-31T14:37:40.899399Z","shell.execute_reply.started":"2025-01-31T14:37:40.880545Z","shell.execute_reply":"2025-01-31T14:37:40.898223Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#define the target column \ntarget_column = 'sii'\ntrain_data_cleaned = train_data.dropna(subset=[target_column])\n#check the results \ntrain_data_cleaned.head()\ntrain_data_cleaned.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:37:42.648104Z","iopub.execute_input":"2025-01-31T14:37:42.648426Z","iopub.status.idle":"2025-01-31T14:37:42.669731Z","shell.execute_reply.started":"2025-01-31T14:37:42.648402Z","shell.execute_reply":"2025-01-31T14:37:42.668534Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#categorical columns in the dataset \ncategorical_columns = ['Basic_Demos-Enroll_Season', 'CGAS-Season', 'Physical-Season', 'FGC-Season', 'BIA-Season', 'PCIAT-Season', 'SDS-Season', 'PreInt_EduHx-Season']\n#plotting boxplots for 'sii' against each categorical column \nplt.figure(figsize=(16,24))\nfor i, col in enumerate(categorical_columns, 1):\n    plt.subplot(4,2,i)\n    sns.boxplot(x=col, y='sii', data=train_data_cleaned)\n    plt.xticks(rotation=45)\n    plt.title(f\"'sii' vs {col}\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:38:02.500863Z","iopub.execute_input":"2025-01-31T14:38:02.501230Z","iopub.status.idle":"2025-01-31T14:38:04.448755Z","shell.execute_reply.started":"2025-01-31T14:38:02.501202Z","shell.execute_reply":"2025-01-31T14:38:04.447478Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#plot target column 'sii' with numerical columns \n#numerical against sii \nnumerical_cols = train_data_cleaned.select_dtypes(include=['float', 'int64']).columns\n#set the number of plots per row \nplots_per_row = 5\nn_rows = (len(numerical_cols) + plots_per_row -1) // plots_per_row\nplt.figure(figsize=(20, 4 * n_rows))\nfor i, col in enumerate(numerical_cols): \n    plt.subplot(n_rows, plots_per_row, i + 1)\n    sns.boxplot(x='sii', y=col, data=train_data_cleaned)\n    plt.title(col)\n    plt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:38:41.528343Z","iopub.execute_input":"2025-01-31T14:38:41.528840Z","iopub.status.idle":"2025-01-31T14:39:33.484073Z","shell.execute_reply.started":"2025-01-31T14:38:41.528795Z","shell.execute_reply":"2025-01-31T14:39:33.483040Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#indentify categorical columns for seasons\nseason_cols = [\n    'basic_Demos-Enroll_Season', \n    'CGAS-Season', \n    'Physical-Season', \n    'FGC-Season', \n    'BIA-Season', \n    'PCIAT-Season', \n    'SDS-Season', \n    'PreInt_EduHx-Season'  \n]\n#create a mapping dict for seasons \nseason_mapping = {\n    'Spring': 0, \n    'Summer': 1, \n    'Fall': 2,\n    'Winter': 3\n}\n#Apply manual encoding to the categorical columns\nfor col in season_cols:\n    if col in train_data_cleaned.columns:\n        train_data_cleaned[col] = train_data_cleaned[col].replace(season_mapping)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:41:01.739991Z","iopub.execute_input":"2025-01-31T14:41:01.740347Z","iopub.status.idle":"2025-01-31T14:41:01.747473Z","shell.execute_reply.started":"2025-01-31T14:41:01.740318Z","shell.execute_reply":"2025-01-31T14:41:01.746309Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#drop the id column if present \ntrain_data_no_id = train_data_cleaned.drop(columns=['id'], errors='ignore')\ntrain_data_numeric = train_data_no_id.select_dtypes(include=['number'])\n#calculate the correlation matrix \ncorrelation_matrix = train_data_numeric.corr()\n#plot the heatmap \nplt.figure(figsize=(30,30))\nsns.heatmap(correlation_matrix, annot=True, fmt='1f', cmap='coolwarm', square=True)\nplt.title('correlation heatmap')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:43:26.508189Z","iopub.execute_input":"2025-01-31T14:43:26.508527Z","iopub.status.idle":"2025-01-31T14:43:38.309921Z","shell.execute_reply.started":"2025-01-31T14:43:26.508503Z","shell.execute_reply":"2025-01-31T14:43:38.308281Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Get the intersection of columns in teh train_data_cleaned and test_data \n# ✅ Get the intersection of column names\ncommon_columns = train_data_cleaned.columns.intersection(test_data.columns)\n# ✅ Filter both DataFrames to keep only the common columns\nX = train_data_cleaned[common_columns].drop(columns=['id'])\ny = train_data_cleaned['sii']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:45:27.943041Z","iopub.execute_input":"2025-01-31T14:45:27.943413Z","iopub.status.idle":"2025-01-31T14:45:27.952267Z","shell.execute_reply.started":"2025-01-31T14:45:27.943384Z","shell.execute_reply":"2025-01-31T14:45:27.951185Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ✅ Step 2: Split Data into Training & Validation Sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=2) # 20% test split","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:45:30.903036Z","iopub.execute_input":"2025-01-31T14:45:30.903387Z","iopub.status.idle":"2025-01-31T14:45:30.912428Z","shell.execute_reply.started":"2025-01-31T14:45:30.903361Z","shell.execute_reply":"2025-01-31T14:45:30.911547Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\n\n# ✅ Step 1: Ensure all categorical features are converted to numeric\nX_train_encoded = pd.get_dummies(X_train)  # Converts categorical to numeric\nX_test_encoded = pd.get_dummies(X_test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:47:35.673073Z","iopub.execute_input":"2025-01-31T14:47:35.673585Z","iopub.status.idle":"2025-01-31T14:47:35.686927Z","shell.execute_reply.started":"2025-01-31T14:47:35.673523Z","shell.execute_reply":"2025-01-31T14:47:35.685851Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ✅ Step 2: Align columns (ensure train & test have the same features)\nX_train_encoded, X_test_encoded = X_train_encoded.align(X_test_encoded, join='left', axis=1, fill_value=0)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:47:44.465776Z","iopub.execute_input":"2025-01-31T14:47:44.466150Z","iopub.status.idle":"2025-01-31T14:47:44.472029Z","shell.execute_reply.started":"2025-01-31T14:47:44.466121Z","shell.execute_reply":"2025-01-31T14:47:44.470823Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ✅ Step 3: Scale the numeric data\nscaler = StandardScaler()\nX_train_scaled = scaler.fit_transform(X_train_encoded)\nX_test_scaled = scaler.transform(X_test_encoded)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:48:01.493143Z","iopub.execute_input":"2025-01-31T14:48:01.493639Z","iopub.status.idle":"2025-01-31T14:48:01.531777Z","shell.execute_reply.started":"2025-01-31T14:48:01.493599Z","shell.execute_reply":"2025-01-31T14:48:01.530697Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ✅ Print shape confirmation\nprint(\"X_train_scaled shape:\", X_train_scaled.shape)\nprint(\"X_test_scaled shape:\", X_test_scaled.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:48:10.278323Z","iopub.execute_input":"2025-01-31T14:48:10.278694Z","iopub.status.idle":"2025-01-31T14:48:10.285037Z","shell.execute_reply.started":"2025-01-31T14:48:10.278638Z","shell.execute_reply":"2025-01-31T14:48:10.283824Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Apply PCA to reduce dimensionality and simplify the dataset PCA reduces dimensionality while preserving important patterns in the data.\n# ✅ Step 2: Apply PCA\n\npca = PCA(n_components=0.95) #Keep 95% of the variance \nX_train_pca = pca.fit_transform(X_train_scaled)\nX_test_pca = pca.transform(X_test_scaled)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:48:32.005006Z","iopub.execute_input":"2025-01-31T14:48:32.005370Z","iopub.status.idle":"2025-01-31T14:48:32.066934Z","shell.execute_reply.started":"2025-01-31T14:48:32.005343Z","shell.execute_reply":"2025-01-31T14:48:32.065251Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Create instance of Random Forest Classifier \nfrom sklearn.ensemble import RandomForestClassifier\n\n# ✅ Step 1: Create Random Forest Classifier Instance\nrf_classifier = RandomForestClassifier(\n    n_estimators=100,  # Number of trees in the forest\n    max_depth=10,  # No depth limit (fully grown trees)\n    min_samples_split=10,  # Minimum samples needed to split a node\n    min_samples_leaf=4,  # Minimum samples per leaf\n    random_state=42,  # Ensures reproducibility\n    n_jobs=-1  # Uses all available CPU cores for faster training\n)\n\n# ✅ Print Model Parameters\nprint(rf_classifier)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:48:34.926993Z","iopub.execute_input":"2025-01-31T14:48:34.927338Z","iopub.status.idle":"2025-01-31T14:48:34.938952Z","shell.execute_reply.started":"2025-01-31T14:48:34.927312Z","shell.execute_reply":"2025-01-31T14:48:34.937861Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"rf_classifier.fit(X_train_pca, y_train)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:48:39.831674Z","iopub.execute_input":"2025-01-31T14:48:39.832096Z","iopub.status.idle":"2025-01-31T14:48:40.423418Z","shell.execute_reply.started":"2025-01-31T14:48:39.832064Z","shell.execute_reply":"2025-01-31T14:48:40.422235Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Make predictions on the test set\ny_pred = rf_classifier.predict(X_test_pca)\n# Generate a classification report\nprint(\"Classification Report:\")\nprint(classification_report(y_test, y_pred))\n# Generate a confusion matrix\nprint(\"Confusion Matrix:\")\nprint(confusion_matrix(y_test, y_pred))\n# Calculate the accuracy on the test set\naccuracy = rf_classifier.score(X_test_pca, y_test)\nprint(f\"Model Accuracy: {accuracy}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:49:48.750758Z","iopub.execute_input":"2025-01-31T14:49:48.751242Z","iopub.status.idle":"2025-01-31T14:49:48.891217Z","shell.execute_reply.started":"2025-01-31T14:49:48.751198Z","shell.execute_reply":"2025-01-31T14:49:48.889750Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#indentify categorical columns for seasons\nseason_cols = [\n    'basic_Demos-Enroll_Season', \n    'CGAS-Season', \n    'Physical-Season', \n    'FGC-Season', \n    'BIA-Season', \n    'PCIAT-Season', \n    'SDS-Season', \n    'PreInt_EduHx-Season'  \n]\n#create a mapping dict for seasons \nseason_mapping = {\n    'Spring': 0, \n    'Summer': 1, \n    'Fall': 2,\n    'Winter': 3\n}\n#Apply manual encoding to the categorical columns\nfor col in season_cols:\n    if col in test_data.columns:\n        test_data[col] = test_data[col].map(season_mapping)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:50:03.927518Z","iopub.execute_input":"2025-01-31T14:50:03.927947Z","iopub.status.idle":"2025-01-31T14:50:03.939398Z","shell.execute_reply.started":"2025-01-31T14:50:03.927916Z","shell.execute_reply":"2025-01-31T14:50:03.937965Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ✅ Step 1: Handle Missing Values\ntest_data.fillna(0, inplace=True)\n# ✅ Step 2: Find Common Columns\ncommon_columns = train_data_cleaned.columns.intersection(test_data.columns)\n# ✅ Step 3: Prepare the Test Data\nX_test_data = test_data[common_columns].drop(columns=['id'])  #Drop the id col \n# ✅ Step 4: One-Hot Encode Categorical Columns\nX_test_encoded = pd.get_dummies(X_test_data)\n# ✅ Step 5: Align Columns to Match Training Data\nX_test_encoded = X_test_encoded.reindex(columns=X_train_encoded.columns, fill_value=0)\n# ✅ Step 6: Scale the Test Data Using the Fitted Scaler\nX_test_scaled = scaler.transform(X_test_encoded)\n# ✅ Step 7: Apply PCA\nX_test_pca = pca.transform(X_test_scaled)\n# ✅ Step 8: Make Predictions\npredictions = rf_classifier.predict(X_test_pca)\n# ✅ Step 9: Create Submission File\nsubmission = pd.DataFrame({\n    'id': test_data['id'], #include the id from test_data \n    'sii': predictions  #predictions from the model \n})\nsubmission.to_csv('submission.csv', index=False)\n#Check submission data frame \nprint(submission.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-31T14:53:16.031547Z","iopub.execute_input":"2025-01-31T14:53:16.032003Z","iopub.status.idle":"2025-01-31T14:53:16.095819Z","shell.execute_reply.started":"2025-01-31T14:53:16.031969Z","shell.execute_reply":"2025-01-31T14:53:16.094631Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}