{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30761,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":" <a id='top'></a>\n <p style=\"color:white;background-color:blue; font-family:newtimeroman; line-height:250%; text-align:center; border-radius: 5px 25px;\">\n CHILD MIND INSTITUTE : PROBLEMATIC INTERNET USE <img src =\"https://storage.googleapis.com/kaggle-organizations/3842/thumbnail.jpg\" style=\"width:2.5rem\" align =\"left\">  </p>\n","metadata":{}},{"cell_type":"markdown","source":"\n<strong style='background-color:#B4E5F0'> Important Note: </strong> <br></br> \n&emsp; This Notebook is solely made for educational purpose and as an author of the notebook I allow all the readers to reuse the code as per their requirement, this notebook is fully open source\n---","metadata":{}},{"cell_type":"markdown","source":"\n<div class=\"alert alert-info\"/>  Table Of Content </div>\n\n* 1 [Import libraries](#1)\n* 2 [Import Data Paths](#2)\n* 3 [Data processing](#3)\n* 4 [Feature Engineering](#4)\n* 5 [Visualizations](#5)\n* 6 [Status Decoding](#6)\n\n<hr size='3' color='grey'/> \n","metadata":{}},{"cell_type":"markdown","source":"## Understanding problem statement:\n","metadata":{}},{"cell_type":"markdown","source":"before we deep dive into the data let's first give some minutes to understand the problem statement, what exactly the domain people are expecting from us, what is the proposed problem and what we should be doing, and also let's understand what impact it can make by our efforts that we are putting here.\n\n**Problem Statement**\n> Can you predict the level of problematic internet usage exhibited by children and adolescents, based on their physical activity? The goal of this competition is to develop a predictive model that analyzes children's physical activity and fitness data to identify early signs of problematic internet use. Identifying these patterns can help trigger interventions to encourage healthier digital habits.\n\n**What are we trying to achieve here?**\n>In today’s digital age, problematic internet use among children and adolescents is a growing concern. Better understanding this issue is crucial for addressing mental health problems such as depression and anxiety.\n\n>Current methods for measuring problematic internet use in children and adolescents are often complex and require professional assessments. This creates access, cultural, and linguistic barriers for many families. Due to these limitations, problematic internet use is often not measured directly, but is instead associated with issues such as depression and anxiety in youth.\n\n>Conversely, physical & fitness measures are extremely accessible and widely available with minimal intervention or clinical expertise. Changes in physical habits, such as poorer posture, irregular diet, and reduced physical activity, are common in excessive technology users. We propose using these easily obtainable physical fitness indicators as proxies for identifying problematic internet use, especially in contexts lacking clinical expertise or suitable assessment tools.\n\n**Why we should be doing it?**\n>Your work will contribute to a healthier, happier future where children are better equipped to navigate the digital landscape responsibly.\n\n**What we can do?**\n<br>\nas expected by compitition hosts, we can create a predictive model which will help us to predict the problemactic internet use\n</br>\n**How?**\n<br></br> To understand how, we need to understand the data,we need to understand what we are going to predict, what is our actual challenge in it and how can we solve that challenge. \n---","metadata":{}},{"cell_type":"markdown","source":"## Understanding the Dataset\n\nNow I suppose we got a fair understanding of what exactly we want to do it, but we still don't know what our dataset is having, what things are there and what thing's are not there?\nso before going to dataset, since we have already seen the problem statement,</br>\n**what thing's you think might be useful for preparing a good dataset?**\n> may be some grouping factors like:  age, gender,location, genetics, to group the people and understand the cases seperately\n> some health data like fitness measures BMI, BMR, height, weight, musclemass, avg water content, Blood Pressure, glucose levels, any specific disorder symptoms, any vitamin and mineral difficiencies, any past trauma and injuries, some school conditions,\n> some bad habit's data : daily screen time, avg useage of adult content, avg social media presence, online interactions, \n\nok we are not doctor's but something like above features about the usecase might help us to categorise the patient into a healthy and unhealthy internet user.\n\nNow let's see what domain experts have suggested us to consider as features for our study, may be these all are not correct, not all might be that much useful yet we have to decide which is good which is bad metric to study, but let's see what we have in our plate.\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ndisplay(df.shape,df.columns)\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-09-22T12:45:18.062769Z","iopub.execute_input":"2024-09-22T12:45:18.063985Z","iopub.status.idle":"2024-09-22T12:45:18.527769Z","shell.execute_reply.started":"2024-09-22T12:45:18.063921Z","shell.execute_reply":"2024-09-22T12:45:18.526866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df2 = pd.read_parquet(\"/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet/id=00115b9f/part-0.parquet\")\ndf2.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-09-22T12:45:18.530072Z","iopub.execute_input":"2024-09-22T12:45:18.530592Z","iopub.status.idle":"2024-09-22T12:45:18.680296Z","shell.execute_reply.started":"2024-09-22T12:45:18.530525Z","shell.execute_reply":"2024-09-22T12:45:18.679298Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"ok, as expected, doctor's have also collected similar data with more granularity in it :)","metadata":{}},{"cell_type":"code","source":"dict_df = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/data_dictionary.csv\")\ndisplay(dict_df.head(40),dict_df.tail(41))","metadata":{"execution":{"iopub.status.busy":"2024-09-22T12:45:18.681754Z","iopub.execute_input":"2024-09-22T12:45:18.682199Z","iopub.status.idle":"2024-09-22T12:45:18.724708Z","shell.execute_reply.started":"2024-09-22T12:45:18.682152Z","shell.execute_reply":"2024-09-22T12:45:18.723706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(dict_df['Instrument'].unique(),dict_df['Instrument'].nunique())","metadata":{"execution":{"iopub.status.busy":"2024-09-22T12:45:18.725846Z","iopub.execute_input":"2024-09-22T12:45:18.726168Z","iopub.status.idle":"2024-09-22T12:45:18.736434Z","shell.execute_reply.started":"2024-09-22T12:45:18.726133Z","shell.execute_reply":"2024-09-22T12:45:18.735373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"so here we are using 11 unique groups of measurements, each group will have certain measurements and each measurement in the study group might be affecting to Sii, ok now quickely let's do some basic exploration of this data and understand \n* how data is varying, based on which we will decide\n* what parameters to cosider for further study, \n* what features are more relevant, \n* do we need any data imputations, and \n* check some baseline models","metadata":{}},{"cell_type":"markdown","source":"<a id ='1'></a><br><div class=\"alert alert-info\">\n  <strong> 1. Importing Libraries</strong>\n</div>\n\n[🏠](#top)","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n\nimport os\nimport seaborn as sns\nimport plotly.express as px","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-09-22T12:45:18.739103Z","iopub.execute_input":"2024-09-22T12:45:18.739419Z","iopub.status.idle":"2024-09-22T12:45:20.842900Z","shell.execute_reply.started":"2024-09-22T12:45:18.739386Z","shell.execute_reply":"2024-09-22T12:45:20.842025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id ='2'></a><br><div class=\"alert alert-info\">\n  <strong>2. Import Data Paths </strong>\n</div>\n\n[🏠](#top)","metadata":{}},{"cell_type":"markdown","source":"## Set data paths","metadata":{}},{"cell_type":"code","source":"INPUT = '/kaggle/input/child-mind-institute-problematic-internet-use'\n\nTRAIN_SET = INPUT + '/series_train.parquet'\nTEST_SET = INPUT + '/series_test.parquet'\n\ntrain_df = pd.read_csv(INPUT + '/train.csv')\ntest_df = pd.read_csv(INPUT + '/test.csv')\nmeta_df = pd.read_csv(INPUT + '/data_dictionary.csv')\nsample_df = pd.read_csv(INPUT + '/sample_submission.csv')\n\n","metadata":{"execution":{"iopub.status.busy":"2024-09-22T12:45:20.844088Z","iopub.execute_input":"2024-09-22T12:45:20.844560Z","iopub.status.idle":"2024-09-22T12:45:20.905422Z","shell.execute_reply.started":"2024-09-22T12:45:20.844527Z","shell.execute_reply":"2024-09-22T12:45:20.904435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_df.head(),train_df.info(),train_df.describe())","metadata":{"execution":{"iopub.status.busy":"2024-09-22T12:45:20.906766Z","iopub.execute_input":"2024-09-22T12:45:20.907184Z","iopub.status.idle":"2024-09-22T12:45:21.117933Z","shell.execute_reply.started":"2024-09-22T12:45:20.907138Z","shell.execute_reply":"2024-09-22T12:45:21.116998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(test_df.head(),test_df.info(),test_df.describe())","metadata":{"execution":{"iopub.status.busy":"2024-09-22T12:45:21.119059Z","iopub.execute_input":"2024-09-22T12:45:21.119931Z","iopub.status.idle":"2024-09-22T12:45:21.249551Z","shell.execute_reply.started":"2024-09-22T12:45:21.119856Z","shell.execute_reply":"2024-09-22T12:45:21.248612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's bad, in test dataset we only have 59 columns but in train dataset we have 80 columns, so our features in train and test dataset don't match... which is not a good sign of study","metadata":{}},{"cell_type":"markdown","source":"there are total 3960 records but only 2736 records have SII value which means other data is missing we might need imputation for sii in future for more fine tunnig,\n\nso to consider as our inner most data recidency we have 59 columns and 2736 records as our training set, a true knowledge corpus, out of which we need to either predict model accurately or do some imputations in it depending on the future results","metadata":{}},{"cell_type":"markdown","source":"<a id ='3'></a><br><div class=\"alert alert-info\">\n  <strong>3. Data Processing </strong>\n</div>\n\n[🏠](#top)","metadata":{}},{"cell_type":"markdown","source":"## Training Dataset","metadata":{}},{"cell_type":"code","source":"feature_cols = list(test_df.columns)+['sii']\ntrain_df = train_df[feature_cols].copy()\ndisplay(train_df.head(),train_df.shape)","metadata":{"execution":{"iopub.status.busy":"2024-09-22T12:45:21.250715Z","iopub.execute_input":"2024-09-22T12:45:21.251102Z","iopub.status.idle":"2024-09-22T12:45:21.281139Z","shell.execute_reply.started":"2024-09-22T12:45:21.251065Z","shell.execute_reply":"2024-09-22T12:45:21.280222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"df = train_df.copy()","metadata":{"execution":{"iopub.status.busy":"2024-09-22T12:45:21.282428Z","iopub.execute_input":"2024-09-22T12:45:21.282793Z","iopub.status.idle":"2024-09-22T12:45:21.289429Z","shell.execute_reply.started":"2024-09-22T12:45:21.282756Z","shell.execute_reply":"2024-09-22T12:45:21.288544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## T-SNE Cluster Analysis of the Dataset","metadata":{}},{"cell_type":"markdown","source":"this is just to see if there exists any exclusive clusters in the dataset, if yes then boom we can achieve our target very early than expected timeline,","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.cluster import KMeans\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.manifold import TSNE\nimport plotly.express as px\n\n# Step 1: Automatically detect categorical columns (dtype 'object' or 'category')\ncategorical_columns = df.select_dtypes(include=['object', 'category']).columns.tolist()\ndf['sii']= df['sii'].fillna(4)\n# Step 2: Handle missing values in categorical columns by filling with 'Unknown'\ndf[categorical_columns] = df[categorical_columns].fillna('Unknown')\n# df = df[df['sii'] !='Unknown']\n# Step 3: Convert categorical columns to numerical using One-Hot Encoding\ndf_encoded = pd.get_dummies(df, columns=categorical_columns, drop_first=True)\n\n# Step 4: Impute missing values in the numerical columns\nnumerical_columns = df_encoded.select_dtypes(include=['float64', 'int64']).columns  # Detect numerical columns\nimputer = SimpleImputer(strategy='mean')  # You can also use 'median' or other strategies\ndf_encoded[numerical_columns] = imputer.fit_transform(df_encoded[numerical_columns])\ndf_encoded\n# # Step 5: Preprocess the data (drop the target column 'sii_Unknown' if necessary)\nX = df_encoded.drop('sii_Unknown', axis=1, errors='ignore')  # If 'sii_Unknown' is created by encoding\n\n# Step 6: Standardize the numerical features\nscaler = StandardScaler()\nX_scaled = scaler.fit_transform(X)\n\n# Step 7: Apply KMeans clustering\nn_clusters = 5  # Number of clusters based on unique values in 'sii'\nkmeans = KMeans(n_clusters=n_clusters, random_state=42)\nclusters = kmeans.fit_predict(X_scaled)\n\n# # Step 8: Use t-SNE for dimensionality reduction to 2D for visualization\ntsne = TSNE(n_components=2, random_state=42)\nX_tsne = tsne.fit_transform(X_scaled)\nprint(X_tsne)\n# # Step 9: Create a DataFrame for plotting\ndf_plot = pd.DataFrame(X_tsne, columns=['Dim1', 'Dim2'])\ndf_plot['Cluster'] = clusters\ndf_plot['sii'] = df['sii']  # Original 'sii' column for labeling in the plot\n\n# Step 10: Visualize the clusters using Plotly\nfig = px.scatter(df_plot, x='Dim1', y='Dim2', color='sii',\n                  title=\"t-SNE Visualization of Clusters with Automatically Detected Categorical Columns\",\n                )\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-09-22T13:18:39.892752Z","iopub.execute_input":"2024-09-22T13:18:39.893197Z","iopub.status.idle":"2024-09-22T13:19:07.341859Z","shell.execute_reply.started":"2024-09-22T13:18:39.893152Z","shell.execute_reply":"2024-09-22T13:19:07.340843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"in 2 dimensional space it's not that clear about the cluster, let's do a 3d analysis","metadata":{}},{"cell_type":"code","source":"from sklearn.manifold import TSNE\nimport plotly.express as px\n\n# Perform t-SNE to reduce the dataset to 3 components\ntsne = TSNE(n_components=3, random_state=42, perplexity=30, n_iter=3000)\n\n# Fit and transform the encoded data\ndf_tsne_3d = tsne.fit_transform(df_encoded)\n\n# Create a DataFrame for the t-SNE results and include the cluster labels and sii column\ndf_tsne_plot = pd.DataFrame(df_tsne_3d, columns=['Dim1', 'Dim2', 'Dim3'])\ndf_tsne_plot['Cluster'] = clusters\ndf_tsne_plot['sii'] = df['sii']\n\n# Plot the 3D scatter plot\nfig = px.scatter_3d(df_tsne_plot, x='Dim1', y='Dim2', z='Dim3', color='sii', symbol='Cluster',\n                    title=\"3D t-SNE Visualization of Clusters\",\n                    labels={'Dim1': 'Component 1', 'Dim2': 'Component 2', 'Dim3': 'Component 3'},\n                    color_discrete_sequence=px.colors.qualitative.Plotly)\n\n# Show the plot\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-09-22T13:19:26.221074Z","iopub.execute_input":"2024-09-22T13:19:26.221730Z","iopub.status.idle":"2024-09-22T13:21:24.816144Z","shell.execute_reply.started":"2024-09-22T13:19:26.221684Z","shell.execute_reply":"2024-09-22T13:21:24.815078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"the major yellow spot at centre is showing us the missing values, it's like a big piece of cake seperated, I think we need to deep dive here to remove the dimensionality, let's try UMAP","metadata":{}},{"cell_type":"code","source":"import umap\nimport plotly.express as px\nimport pandas as pd\n\n# Step 1: Initialize UMAP with 3 components\numap_3d = umap.UMAP(n_components=3)\n\n# Step 2: Fit and transform the encoded dataset (df_encoded)\ndf_umap_3d = umap_3d.fit_transform(df_encoded)\n\n# Step 3: Create a DataFrame for the UMAP results and include cluster labels and sii column\ndf_umap_plot = pd.DataFrame(df_umap_3d, columns=['Dim1', 'Dim2', 'Dim3'])\ndf_umap_plot['Cluster'] = clusters\ndf_umap_plot['sii'] = df['sii']\n\n# Step 4: Plot the 3D scatter plot\nfig = px.scatter_3d(df_umap_plot, x='Dim1', y='Dim2', z='Dim3', color='sii',\n                    title=\"3D UMAP Visualization of Clusters\",\n                    labels={'Dim1': 'Component 1', 'Dim2': 'Component 2', 'Dim3': 'Component 3'},\n                    color_discrete_sequence=px.colors.qualitative.Plotly)\n\n# Step 5: Show the plot\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-09-22T13:23:58.862409Z","iopub.execute_input":"2024-09-22T13:23:58.863478Z","iopub.status.idle":"2024-09-22T13:24:34.049586Z","shell.execute_reply.started":"2024-09-22T13:23:58.863432Z","shell.execute_reply":"2024-09-22T13:24:34.048509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"hmm, so based on the clusters that we are getting we can conlude that there is some relative behaviour in sii and the data that we have, but it's not that much strong enough as if a direct model is applicable on it, we might need to do more indepth analysis of features and check the relation of each feature group with sii...\n\nbut one conclusion, we can not directly apply a pre-build model on the dataset, it needs a good data churning out there.\n","metadata":{}},{"cell_type":"markdown","source":"well well, we couldn't get a pattern between sii and umap clusters in the dataset here. but let's see what these cluster's are ... may be they will give us path to our main analysis","metadata":{}},{"cell_type":"code","source":"umap_result = df_umap_3d\n# Assuming `umap_result` is your UMAP-transformed data (with 2 or 3 components)\nkmeans = KMeans(n_clusters=4)  # Set n_clusters to the number of clusters you want\ncluster_labels = kmeans.fit_predict(df_umap_3d)\n\nfig = px.scatter_3d(\n    x=umap_result[:, 0],  # UMAP component 1\n    y=umap_result[:, 1],  # UMAP component 2\n    z=umap_result[:, 2],  # UMAP component 3\n    color=cluster_labels,  # Color by cluster labels\n    symbol=df['sii'],  # Optionally, shape by original categories (e.g., 'sii')\n    title=\"UMAP Clustering Results\"\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-22T13:26:17.821548Z","iopub.execute_input":"2024-09-22T13:26:17.822288Z","iopub.status.idle":"2024-09-22T13:26:17.950299Z","shell.execute_reply.started":"2024-09-22T13:26:17.822245Z","shell.execute_reply":"2024-09-22T13:26:17.949393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"aahaa... so now we have a umap based dimensionality reduction of our dataset and we got few clusters as well of our dataset, what we will do is we will keep these clusters as it is in our main dataset to see how they are related to their parent features and how these clusters are dependant on the sii values, ofcourse they are not relevant at the current moment but in future studies we will utilize their relation someway or other,","metadata":{}},{"cell_type":"markdown","source":"## Conclusion of Dimensionality Analysis:\n* Umap is able to categorise high dimensional data into 4 clusters, \n* which indicates there is high chance of data dependancy and data formations\n* there is need of analysing relation between high dimensional umap clusters with sii values, may be they are related or may not be.\n* check the relation in lower space (normal dataset) features with sii values and umap cluster values.\n* try to find relation between grouped dimensions and sii","metadata":{}},{"cell_type":"markdown","source":"### Work in progress!!","metadata":{}},{"cell_type":"code","source":"","metadata":{"execution":{"iopub.status.busy":"2024-09-22T13:17:14.642526Z","iopub.execute_input":"2024-09-22T13:17:14.642942Z","iopub.status.idle":"2024-09-22T13:17:14.649487Z","shell.execute_reply.started":"2024-09-22T13:17:14.642898Z","shell.execute_reply":"2024-09-22T13:17:14.648412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}