{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"from sklearn.cluster import KMeans, AgglomerativeClustering\nfrom sklearn.decomposition import PCA\nimport numpy as np\nimport pandas as pd \nimport os\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom sklearn.preprocessing import StandardScaler\n\n#Loading the CMI train dataset. \ntrain_data = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\n\n#Filling missing values in the data ith 0s. \ntrain_data.fillna(value=0, inplace=True)\n\n#Choosing two features for the clusters. I chose Physical-Waist_Circumference and Physical-HeartRate because I was curious to see if the plot would demonstrate that a higher waist circumference would indicate a higher heartrate. \nX = train_data[['Physical-Waist_Circumference']]\ny = train_data[['Physical-HeartRate']]\n\n#Running the standard scaler on both features. \nscaler = StandardScaler()\n\nX = scaler.fit_transform(X)\ny = scaler.fit_transform(y)\nplt.scatter(X, y)\nplt.xlabel (\"Physical-Waist_Circumference\")\nplt.ylabel (\"Physical-HeartRate\")\nplt.show()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-11-22T00:48:27.250088Z","iopub.execute_input":"2024-11-22T00:48:27.251265Z","iopub.status.idle":"2024-11-22T00:48:27.455495Z","shell.execute_reply.started":"2024-11-22T00:48:27.251202Z","shell.execute_reply":"2024-11-22T00:48:27.454467Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# K-Means Clustering","metadata":{}},{"cell_type":"code","source":"#Creating the model instance with 2 clusters. \nkmeans = KMeans(n_clusters=2, random_state = 0)\n#Running a prediction on the model. \nkmeans = kmeans.fit_predict(X, y)\n\n#Creating a scatterplot of the two features with the Kmeans prediction set as the color to see the clusters. \nplt.scatter(X, y, c=kmeans)\n\nplt.xlabel (\"Physical-Waist_Circumference\")\nplt.ylabel (\"Physical-HeartRate\")\nplt.show()\n#I can see in the plot that the model has differentiated two large clusters in the data. One of them being in yellow on the right side, and the other being the purple verticle line of data on the left. ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-22T00:48:27.457733Z","iopub.execute_input":"2024-11-22T00:48:27.458163Z","iopub.status.idle":"2024-11-22T00:48:27.692700Z","shell.execute_reply.started":"2024-11-22T00:48:27.458121Z","shell.execute_reply":"2024-11-22T00:48:27.691664Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Agglomerative Clustering","metadata":{}},{"cell_type":"code","source":"#Initiating the model with 2 clusters. \nagg = AgglomerativeClustering(n_clusters=2)\n\n#Making a prediction. \nagg_clusters = agg.fit_predict(X, y)\n\nplt.scatter(X, y, c=agg_clusters)\n\nplt.xlabel (\"Physical-Waist_Circumference\")\nplt.ylabel (\"Physical-HeartRate\")\nplt.show()\n#Looking at the plot, the Agglomerative Clustering model has created two similar models to the K-Means, except that in this one, the colors of each cluster seems to be switched. ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-22T00:48:27.694371Z","iopub.execute_input":"2024-11-22T00:48:27.694784Z","iopub.status.idle":"2024-11-22T00:48:28.537690Z","shell.execute_reply.started":"2024-11-22T00:48:27.694741Z","shell.execute_reply":"2024-11-22T00:48:28.536683Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Cite: https://www.geeksforgeeks.org/implementing-pca-in-python-with-scikit-learn/ I used this resource to know what argument to give the PCA model which is n_components. \nPCA=PCA(n_components=1)\nPCA_transform = PCA.fit_transform(X,y)\n\nplt.scatter(X, y, c=PCA_transform)\nplt.xlabel (\"Physical-Waist_Circumference\")\nplt.ylabel (\"Physical-HeartRate\")\nplt.show()\n\n#When I scatterplot the features with the color set to display the PCA model, it looks like the PCA has created a gradient on the right cluster, although I'm not sure how to interpret it. ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-22T00:48:28.538851Z","iopub.execute_input":"2024-11-22T00:48:28.539136Z","iopub.status.idle":"2024-11-22T00:48:28.764815Z","shell.execute_reply.started":"2024-11-22T00:48:28.539109Z","shell.execute_reply":"2024-11-22T00:48:28.763765Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# When comparing the 2 clustering algorithms and the PCA, and assuming that I've correctly modeled and plotted the cluster predictions, I can see that all three have separated very similar clusters. These clusters seem to be two large circles side by side on the plot. I believe this is due to the large distance of space between the line of data on the left, and the cluster of data on the right. The PCA produced a gradient effect on the rightmost cluster. I'm interested to learn what it means and how to analyze it. ","metadata":{}}]}