{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 1. What is Clustering?","metadata":{}},{"cell_type":"markdown","source":"Many things around us can be categorized as “this and that” or to be less vague and more specific, we have groupings that could be binary or groups that can be more than two, like a type of pizza base or type of car that you might want to purchase. The choices are always clear – or, how the technical lingo wants to put it – predefined groups and the process predicting that is an important process in the Data Science stack called Classification.\n\nBut what if we bring into play a quest where we don’t have pre-defined choices initially, rather, we derive those choices! Choices that are based out of hidden patterns, underlying similarities between the constituent variables, salient features from the data etc. This process is known as Clustering in Machine Learning or Cluster Analysis, where we group the data together into an unknown number of groups and later use that information for further business processes.\n\n\nSo, to put it in simple words, in machine learning clustering is the process by which we create groups in a data, like customers, products, employees, text documents, in such a way that objects falling into one group exhibit many similar properties with each other and are different from objects that fall in the other groups that got created during the process.","metadata":{}},{"cell_type":"markdown","source":"## 1.1. Data Exploration","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf=pd.read_csv(\"/kaggle/input/tabular-playground-series-jul-2022/data.csv\")\nss=pd.read_csv(\"/kaggle/input/tabular-playground-series-jul-2022/sample_submission.csv\")","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-01T18:22:22.024560Z","iopub.execute_input":"2022-07-01T18:22:22.024970Z","iopub.status.idle":"2022-07-01T18:22:22.897888Z","shell.execute_reply.started":"2022-07-01T18:22:22.024936Z","shell.execute_reply":"2022-07-01T18:22:22.896451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\nsns.set(rc={'figure.figsize':(24,20)})\nsns.heatmap(df.corr(),annot=True,fmt='.2f')","metadata":{"execution":{"iopub.status.busy":"2022-07-01T18:22:22.900618Z","iopub.execute_input":"2022-07-01T18:22:22.901309Z","iopub.status.idle":"2022-07-01T18:22:26.709727Z","shell.execute_reply.started":"2022-07-01T18:22:22.901252Z","shell.execute_reply":"2022-07-01T18:22:26.708915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set(rc={'figure.figsize':(15,15)})\nfor i, column in enumerate(list(df.columns), 1):\n    plt.subplot(5,6,i)\n    p=sns.histplot(x=column,data=df.sample(1000),stat='count',kde=True,color='green')","metadata":{"execution":{"iopub.status.busy":"2022-07-01T18:22:26.710981Z","iopub.execute_input":"2022-07-01T18:22:26.711765Z","iopub.status.idle":"2022-07-01T18:22:32.654620Z","shell.execute_reply.started":"2022-07-01T18:22:26.711732Z","shell.execute_reply":"2022-07-01T18:22:32.653221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. What are clustering algorithms?","metadata":{}},{"cell_type":"markdown","source":"Clustering is an unsupervised machine learning task. You might also hear this referred to as cluster analysis because of the way this method works. Using a clustering algorithm means you're going to give the algorithm a lot of input data with no labels and let it find any groupings in the data it can.\n\nThose groupings are called clusters. A cluster is a group of data points that are similar to each other based on their relation to surrounding data points. Clustering is used for things like feature engineering or pattern discovery.\n\n**When you're starting with data you know nothing about, clustering might be a good place to get some insight.**","metadata":{}},{"cell_type":"markdown","source":"# 3. Types of clustering algorithms\nThere are different types of clustering algorithms that handle all kinds of unique data.","metadata":{}},{"cell_type":"markdown","source":"## 3.1. Density-based\nIn density-based clustering, data is grouped by areas of high concentrations of data points surrounded by areas of low concentrations of data points. Basically the algorithm finds the places that are dense with data points and calls those clusters.\nThe great thing about this is that the clusters can be any shape. You aren't constrained to expected conditions.\nThe clustering algorithms under this type don't try to assign outliers to clusters, so they get ignored.\n\nDensity-based clustering connects areas of high example density into clusters. This allows for arbitrary-shaped distributions as long as dense areas can be connected. These algorithms have difficulty with data of varying densities and high dimensions. Further, by design, these algorithms do not assign outliers to clusters.\n\n![](https://www.researchgate.net/publication/334279038/figure/fig5/AS:960330599514121@1605972065828/Density-based-clustering-techniques-DBSCAN_Q320.jpg)","metadata":{}},{"cell_type":"markdown","source":"## 3.2. Distribution-based\nWith a distribution-based clustering approach, all of the data points are considered parts of a cluster based on the probability that they belong to a given cluster.\nIt works like this: there is a center-point, and as the distance of a data point from the center increases, the probability of it being a part of that cluster decreases.\nIf you aren't sure of how the distribution in your data might be, you should consider a different type of algorithm.\n\nThis clustering approach assumes data is composed of distributions, such as Gaussian distributions. In Figure 3, the distribution-based algorithm clusters data into three Gaussian distributions. As distance from the distribution's center increases, the probability that a point belongs to the distribution decreases. The bands show that decrease in probability. When you do not know the type of distribution in your data, you should use a different algorithm.\n\n![](https://www.researchgate.net/publication/332053160/figure/fig2/AS:741417534107649@1553779123709/Distribution-model-of-clustering.png)","metadata":{}},{"cell_type":"markdown","source":"## 3.3. Centroid-based\nCentroid-based clustering is the one you probably hear about the most. It's a little sensitive to the initial parameters you give it, but it's fast and efficient.\n\nThese types of algorithms separate data points based on multiple centroids in the data. Each data point is assigned to a cluster based on its squared distance from the centroid. This is the most commonly used type of clustering.\n\n![](https://www.researchgate.net/publication/334279038/figure/fig6/AS:960330603712512@1605972066015/Centroid-based-clustering-algorithm.png)","metadata":{}},{"cell_type":"markdown","source":"## 3.4. Hierarchical-based\nHierarchical-based clustering is typically used on hierarchical data, like you would get from a company database or taxonomies. It builds a tree of clusters so everything is organized from the top-down.\nThis is more restrictive than the other clustering types, but it's perfect for specific kinds of data sets.\n\n**Hierarchical clustering creates a tree of clusters. Hierarchical clustering, not surprisingly, is well suited to hierarchical data, such as taxonomies.**\n\n![](https://www.analytixlabs.co.in/blog/wp-content/uploads/2020/07/image-3-28-1-600x400.jpg)","metadata":{}},{"cell_type":"markdown","source":"### 3.4.1. Divisive Approach\nThis approach of hierarchical clustering follows a top-down approach where we consider that all the data points belong to one large cluster and try to divide the data into smaller groups based on a termination logic or, a point beyond which there will be no further division of data points. This termination logic can be based on the minimum sum of squares of error inside a cluster or for categorical data, the metric can be the GINI coefficient inside a cluster.\n\n![](https://www.analytixlabs.co.in/blog/wp-content/uploads/2020/07/image-4-17-1-600x323.jpg)","metadata":{}},{"cell_type":"markdown","source":"### 3.4.2. Agglomerative Approach\nAgglomerative is quite the contrary to Divisive, where all the “N” data points are considered to be a single member of “N” clusters that the data is comprised into. We iteratively combine these numerous “N” clusters to fewer number of clusters, let’s say “k” clusters and hence assign the data points to each of these clusters accordingly. This approach is a bottom-up one, and also uses a termination logic in combining the clusters. This logic can be a number based criterion (no more clusters beyond this point) or a distance criterion (clusters should not be too far apart to be merged) or variance criterion (increase in the variance of the cluster being merged should not exceed a threshold, Ward Method)","metadata":{}},{"cell_type":"markdown","source":"# 4. Types of Clustering Algorithms","metadata":{}},{"cell_type":"markdown","source":"# >> 4.1. k-Means Clustering","metadata":{}},{"cell_type":"markdown","source":"k-means clustering is a method of vector quantization, originally from signal processing, that aims to partition n observations into k clusters in which each observation belongs to the cluster with the nearest mean, serving as a prototype of the cluster. \n\nK-means clustering uses “centroids”, K different randomly-initiated points in the data, and assigns every data point to the nearest centroid. After every point has been assigned, the centroid is moved to the average of all of the points assigned to it.\n\n![](https://media.geeksforgeeks.org/wp-content/uploads/20190812011831/Screenshot-2019-08-12-at-1.09.42-AM.png)\n\n#### Where is K-means clustering used?\nkmeans algorithm is very popular and used in a variety of applications such as market segmentation, document clustering, image segmentation and image compression, etc. The goal usually when we undergo a cluster analysis is either: Get a meaningful intuition of the structure of the data we're dealing with.","metadata":{}},{"cell_type":"markdown","source":"## 4.1.1. Advantages of k-means\n- Relatively simple to implement\n- Scales to large data sets.\n- Guarantees convergence.\n- Can warm-start the positions of centroids.\n- Easily adapts to new examples.\n- Generalizes to clusters of different shapes and sizes, such as elliptical clusters.\n\n## 4.1.2. Disadvantages of k-means\n- Choosing K manually.\n   - Use the “Loss vs. Clusters” plot to find the optimal (k), as discussed in Interpret Results.\n- Being dependent on initial values.\n   - For a low , you can mitigate this dependence by running k-means several times with different initial values and picking the best result. As  increases, you need advanced versions of k-means to pick better values of the initial centroids (called k-means seeding). For a full discussion of k- means seeding see, A Comparative Study of Efficient Initialization Methods for the K-Means Clustering Algorithm by M. Emre Celebi, Hassan A. Kingravi, Patricio A. Vela.\n- Clustering data of varying sizes and density.\n   - k-means has trouble clustering data where clusters are of varying sizes and density. To cluster such data, you need to generalize k-means as described in the Advantages section.\n- Clustering outliers.\n   - Centroids can be dragged by outliers, or outliers might get their own cluster instead of being ignored. Consider removing or clipping outliers before clustering.\n- Scaling with number of dimensions.\n   - As the number of dimensions increases, a distance-based similarity measure converges to a constant value between any given examples. Reduce dimensionality either by using PCA on the feature data, or by using “spectral clustering” to modify the clustering algorithm as explained below.","metadata":{}},{"cell_type":"markdown","source":"## 4.1.3. Code Example:","metadata":{}},{"cell_type":"code","source":"from numpy import unique\nfrom numpy import where\nfrom matplotlib import pyplot\nfrom sklearn.datasets import make_classification\nfrom sklearn.cluster import KMeans\n\n# initialize the data set we'll work with\ntraining_data, _ = make_classification(\n    n_samples=1000,\n    n_features=2,\n    n_informative=2,\n    n_redundant=0,\n    n_clusters_per_class=1,\n    random_state=4\n)\n\n# define the model\nkmeans_model = KMeans(n_clusters=2)\n\n# assign each data point to a cluster\nkmeans_result = kmeans_model.fit_predict(training_data)\n\n# get all of the unique clusters\nkmeans_clusters = unique(kmeans_result)\n\n# plot the DBSCAN clusters\nfor dbscan_cluster in kmeans_clusters:\n    # get data points that fall in this cluster\n    index = where(kmeans_result == kmeans_clusters)\n    # make the plot\n    pyplot.scatter(training_data[index, 0], training_data[index, 1])\n\n# show the DBSCAN plot\npyplot.show()","metadata":{"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# >> 4.2. DBSCAN clustering algorithm","metadata":{}},{"cell_type":"markdown","source":"DBSCAN is a density-based clustering algorithm that works on the assumption that clusters are dense regions in space separated by regions of lower density. It groups 'densely grouped' data points into a single cluster.\n\n#### Is DBSCAN faster than KMeans?\nDBSCAN produces a varying number of clusters, based on the input data. Here's a list of advantages of KMeans and DBScan: KMeans is much faster than DBScan. DBScan doesn't need number of clusters.\n\n![](https://www.analytixlabs.co.in/blog/wp-content/uploads/2020/07/image-12-1-600x314.jpg)\n\n\n#### What is the basic principle of DBSCAN clustering?\nThe principle of DBSCAN is to find the neighborhoods of data points exceeds certain density threshold. The density threshold is defined by two parameters: the radius of the neighborhood (eps) and the minimum number of neighbors/data points (minPts) within the radius of the neighborhood.\n\n#### What is the difference between KMeans and DBSCAN?\nK-means needs a prototype-based concept of a cluster. DBSCAN needs a density-based concept. K-means has difficulty with non-globular clusters and clusters of multiple sizes. DBSCAN is used to handle clusters of multiple sizes and structures and is not powerfully influenced by noise or outliers.","metadata":{}},{"cell_type":"markdown","source":"## 4.2.1. Advantages of DBSCAN clustering\n- Can easily deal with noise, not affected by outliers.\n- Doesn’t require prior specification of clusters.\n- It has no strict shapes, it can correctly accommodate many data points.\n\n## 4.2.2. Disadvantages of DBSCAN clustering\n- Sensitive to the clustering hyper-parameters – the eps and the min_points.\n- Cannot work with datasets of varying densities.\n- Fails if the data is too sparse.\n- The density measures (Reachability and Connectivity) can be affected by sampling.","metadata":{}},{"cell_type":"markdown","source":"## 4.2.3. Code Example:","metadata":{}},{"cell_type":"code","source":"from numpy import unique\nfrom numpy import where\nfrom matplotlib import pyplot\nfrom sklearn.datasets import make_classification\nfrom sklearn.cluster import DBSCAN\n\n# initialize the data set we'll work with\ntraining_data, _ = make_classification(\n    n_samples=1000,\n    n_features=2,\n    n_informative=2,\n    n_redundant=0,\n    n_clusters_per_class=1,\n    random_state=4\n)\n\n# define the model\ndbscan_model = DBSCAN(eps=0.25, min_samples=9)\n\n# train the model\ndbscan_model.fit(training_data)\n\n# assign each data point to a cluster\ndbscan_result = dbscan_model.predict(training_data)\n\n# get all of the unique clusters\ndbscan_cluster = unique(dbscan_result)\n\n# plot the DBSCAN clusters\nfor dbscan_cluster in dbscan_clusters:\n    # get data points that fall in this cluster\n    index = where(dbscan_result == dbscan_clusters)\n    # make the plot\n    pyplot.scatter(training_data[index, 0], training_data[index, 1])\n\n# show the DBSCAN plot\npyplot.show()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-01T18:22:33.550083Z","iopub.status.idle":"2022-07-01T18:22:33.550960Z","shell.execute_reply.started":"2022-07-01T18:22:33.550709Z","shell.execute_reply":"2022-07-01T18:22:33.550736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# >> 4.3. Gaussian Mixture Model algorithm","metadata":{}},{"cell_type":"markdown","source":"Gaussian mixture models (GMMs) are often used for data clustering. You can use GMMs to perform either hard clustering or soft clustering on query data. To perform hard clustering, the GMM assigns query data points to the multivariate normal components that maximize the component posterior probability, given the data.\n\nGaussian Mixture Models (GMMs) assume that there are a certain number of Gaussian distributions, and each of these distributions represent a cluster. Hence, a Gaussian Mixture Model tends to group the data points belonging to a single distribution together.\n\n![](https://www.analytixlabs.co.in/blog/wp-content/uploads/2020/07/image-14-1-600x450.jpg)\n\nAt its simplest, GMM is also a type of clustering algorithm. As its name implies, each cluster is modelled according to a different Gaussian distribution. This flexible and probabilistic approach to modelling the data means that rather than having hard assignments into clusters like k-means, we have soft assignments.","metadata":{}},{"cell_type":"markdown","source":"## 4.3.1. Advantages of Gaussian Mixture Model algorithm\n- The associativity of a data point to a cluster is quantified using probability metrics – which can be easily interpreted.\n- Proven to be accurate for real-time data sets.\n- Some versions of GMM allows for mixed membership of data points, hence it can be a good alternative to Fuzzy C Means to achieve fuzzy clustering.\n\n\n## 4.3.2. Disadvantages of Gaussian Mixture Model algorithm\n- Complex algorithm and cannot be applicable to larger data\n- It is hard to find clusters if the data is not Gaussian, hence a lot of data preparation is required.","metadata":{}},{"cell_type":"markdown","source":"## 4.3.3. Code Example :","metadata":{}},{"cell_type":"code","source":"from numpy import unique\nfrom numpy import where\nfrom matplotlib import pyplot\nfrom sklearn.datasets import make_classification\nfrom sklearn.mixture import GaussianMixture\n\n# initialize the data set we'll work with\ntraining_data, _ = make_classification(\n    n_samples=1000,\n    n_features=2,\n    n_informative=2,\n    n_redundant=0,\n    n_clusters_per_class=1,\n    random_state=4\n)\n\n# define the model\ngaussian_model = GaussianMixture(n_components=2)\n\n# train the model\ngaussian_model.fit(training_data)\n\n# assign each data point to a cluster\ngaussian_result = gaussian_model.predict(training_data)\n\n# get all of the unique clusters\ngaussian_clusters = unique(gaussian_result)\n\n# plot Gaussian Mixture the clusters\nfor gaussian_cluster in gaussian_clusters:\n    # get data points that fall in this cluster\n    index = where(gaussian_result == gaussian_clusters)\n    # make the plot\n    pyplot.scatter(training_data[index, 0], training_data[index, 1])\n\n# show the Gaussian Mixture plot\npyplot.show()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-01T18:22:33.552613Z","iopub.status.idle":"2022-07-01T18:22:33.553014Z","shell.execute_reply.started":"2022-07-01T18:22:33.552811Z","shell.execute_reply":"2022-07-01T18:22:33.552842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# > 4.4. BIRCH algorithm\n\n**The Balance Iterative Reducing and Clustering using Hierarchies (BIRCH) algorithm works better on large data sets than the k-means algorithm.**\nIt breaks the data into little summaries that are clustered instead of the original data points. The summaries hold as much distribution information about the data points as possible.\n\n![](https://media.geeksforgeeks.org/wp-content/uploads/20200612004451/BIRCH.png)\n\nThis algorithm is commonly used with other clustering algorithm because the other clustering techniques can be used on the summaries generated by BIRCH.\nThe main downside of the BIRCH algorithm is that it only works on numeric data values. You can't use this for categorical values unless you do some data transformations.","metadata":{}},{"cell_type":"markdown","source":"## 4.3.1. Advantages of BIRCH algorithm\n- Finds a good clustering with a single scan and improves the quality with a few additional scans\n\n## 4.3.2. Disadvantages of BIRCH algorithm\n- Handles only numeric data\n\n## 4.3.3. Applications of BIRCH algorithm\n- Pixel classification in images\n- Image compression\n- Works with very large data sets","metadata":{}},{"cell_type":"markdown","source":"## 4.3.4. Code Example:","metadata":{}},{"cell_type":"code","source":"\n# Import required libraries and modules\nimport matplotlib.pyplot as plt\nfrom sklearn.datasets.samples_generator import make_blobs\nfrom sklearn.cluster import Birch\n \n# Generating 600 samples using make_blobs\ndataset, clusters = make_blobs(n_samples = 600, centers = 8, cluster_std = 0.75, random_state = 0)\n \n# Creating the BIRCH clustering model\nmodel = Birch(branching_factor = 50, n_clusters = None, threshold = 1.5)\n \n# Fit the data (Training)\nmodel.fit(dataset)\n \n# Predict the same data\npred = model.predict(dataset)\n \n# Creating a scatter plot\nplt.scatter(dataset[:, 0], dataset[:, 1], c = pred, cmap = 'rainbow', alpha = 0.7, edgecolors = 'b')\nplt.show()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-01T18:22:33.554561Z","iopub.status.idle":"2022-07-01T18:22:33.555470Z","shell.execute_reply.started":"2022-07-01T18:22:33.555206Z","shell.execute_reply":"2022-07-01T18:22:33.555228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# > 4.5. Affinity Propagation clustering algorithm","metadata":{}},{"cell_type":"markdown","source":"Affinity propagation (AP) is a graph based clustering algorithm similar to k Means or K medoids, which does not require the estimation of the number of clusters before running the algorithm. Affinity propagation finds “exemplars” i.e. members of the input set that are representative of clusters.\n\nEach data point communicates with all of the other data points to let each other know how similar they are and that starts to reveal the clusters in the data. You don't have to tell this algorithm how many clusters to expect in the initialization parameters.\n\n![](https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcRhsCOnmydquBzvTqjIoAj1p6vamOiO5J0Ix7LKrUjzxw&s)","metadata":{}},{"cell_type":"markdown","source":"## 4.5.1. Code Examples","metadata":{}},{"cell_type":"code","source":"from numpy import unique\nfrom numpy import where\nfrom matplotlib import pyplot\nfrom sklearn.datasets import make_classification\nfrom sklearn.cluster import AffinityPropagation\n\n# initialize the data set we'll work with\ntraining_data, _ = make_classification(\n    n_samples=1000,\n    n_features=2,\n    n_informative=2,\n    n_redundant=0,\n    n_clusters_per_class=1,\n    random_state=4\n)\n\n# define the model\nmodel = AffinityPropagation(damping=0.7)\n\n# train the model\nmodel.fit(training_data)\n\n# assign each data point to a cluster\nresult = model.predict(training_data)\n\n# get all of the unique clusters\nclusters = unique(result)\n\n# plot the clusters\nfor cluster in clusters:\n    # get data points that fall in this cluster\n    index = where(result == cluster)\n    # make the plot\n    pyplot.scatter(training_data[index, 0], training_data[index, 1])\n\n# show the plot\npyplot.show()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-01T18:22:33.557380Z","iopub.status.idle":"2022-07-01T18:22:33.557774Z","shell.execute_reply.started":"2022-07-01T18:22:33.557590Z","shell.execute_reply":"2022-07-01T18:22:33.557608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# > 4.6. Mean-Shift clustering algorithm","metadata":{}},{"cell_type":"markdown","source":"Mean shift is a non-parametric feature-space mathematical analysis technique for locating the maxima of a density function, a so-called mode-seeking algorithm. Application domains include cluster analysis in computer vision and image processing.\n\nIt is a density-based clustering algorithm where it firstly, seeks for stationary points in the density function. Then next, the clusters are eventually shifted to a region with higher density by shifting the center of the cluster to the mean of the points present in the current window. The shift if the window is repeated until no more points can be accommodated inside of that window.\n\n![](https://media.geeksforgeeks.org/wp-content/uploads/20190429213154/1354.png)","metadata":{}},{"cell_type":"markdown","source":"## 4.6.1. Advantages of Mean-Shift clustering \n- The following are some advantages of Mean-Shift clustering algorithm −\n- It does not need to make any model assumption as like in K-means or Gaussian mixture.\n- It can also model the complex clusters which have nonconvex shape.\n- It only needs one parameter named bandwidth which automatically determines the number of clusters.\n- There is no issue of local minima as like in K-means.\n- No problem generated from outliers.\n\n## 4.6.2. Disadvantages of Mean-Shift clustering \n- The following are some disadvantages of Mean-Shift clustering algorithm −\n- Mean-shift algorithm does not work well in case of high dimension, where number of clusters changes abruptly.\n- We do not have any direct control on the number of clusters but in some applications, we need a specific number of clusters.\n- It cannot differentiate between meaningful and meaningless modes.","metadata":{}},{"cell_type":"markdown","source":"## 4.6.3. Code Example","metadata":{}},{"cell_type":"code","source":"from numpy import unique\nfrom numpy import where\nfrom matplotlib import pyplot\nfrom sklearn.datasets import make_classification\nfrom sklearn.cluster import MeanShift\n\n# initialize the data set we'll work with\ntraining_data, _ = make_classification(\n    n_samples=1000,\n    n_features=2,\n    n_informative=2,\n    n_redundant=0,\n    n_clusters_per_class=1,\n    random_state=4\n)\n\n# define the model\nmean_model = MeanShift()\n\n# assign each data point to a cluster\nmean_result = mean_model.fit_predict(training_data)\n\n# get all of the unique clusters\nmean_clusters = unique(mean_result)\n\n# plot Mean-Shift the clusters\nfor mean_cluster in mean_clusters:\n    # get data points that fall in this cluster\n    index = where(mean_result == mean_cluster)\n    # make the plot\n    pyplot.scatter(training_data[index, 0], training_data[index, 1])\n\n# show the Mean-Shift plot\npyplot.show()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-01T18:22:33.560070Z","iopub.status.idle":"2022-07-01T18:22:33.560487Z","shell.execute_reply.started":"2022-07-01T18:22:33.560300Z","shell.execute_reply":"2022-07-01T18:22:33.560319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# > 4.7. OPTICS algorithm\n**OPTICS Clustering stands for Ordering Points To Identify Cluster Structure.** It draws inspiration from the DBSCAN clustering algorithm. It adds two more terms to the concepts of DBSCAN clustering.\n\nThey are:-\n- **Core Distance:** It is the minimum value of radius required to classify a given point as a core point. If the given point is not a Core point, then it’s Core Distance is undefined.\n- **Reachability Distance:** It is defined with respect to another data point q(Let). The Reachability distance between a point p and q is the maximum of the Core Distance of p and the Euclidean Distance(or some other distance metric) between p and q. Note that The Reachability Distance is not defined if q is not a Core point.\n\n![](https://media.geeksforgeeks.org/wp-content/uploads/20190711114717/reachability_distance1.png)","metadata":{}},{"cell_type":"markdown","source":"## 4.7.1. OPTICS Clustering v/s DBSCAN Clustering:\n\n- **Memory Cost :** The OPTICS clustering technique requires more memory as it maintains a priority queue (Min Heap) to determine the next data point which is closest to the point currently being processed in terms of Reachability Distance. It also requires more computational power because the nearest neighbour queries are more complicated than radius queries in DBSCAN.\n- **Fewer Parameters :** The OPTICS clustering technique does not need to maintain the epsilon parameter and is only given in the above pseudo-code to reduce the time taken. This leads to the reduction of the analytical process of parameter tuning.\nThis technique does not segregate the given data into clusters. It merely produces a Reachability distance plot and it is upon the interpretation of the programmer to cluster the points accordingly.","metadata":{}},{"cell_type":"markdown","source":"## 4.7.2. Code Example: ","metadata":{}},{"cell_type":"code","source":"from numpy import unique\nfrom numpy import where\nfrom matplotlib import pyplot\nfrom sklearn.datasets import make_classification\nfrom sklearn.cluster import OPTICS\n\n# initialize the data set we'll work with\ntraining_data, _ = make_classification(\n    n_samples=1000,\n    n_features=2,\n    n_informative=2,\n    n_redundant=0,\n    n_clusters_per_class=1,\n    random_state=4\n)\n\n# define the model\noptics_model = OPTICS(eps=0.75, min_samples=10)\n\n# assign each data point to a cluster\noptics_result = optics_model.fit_predict(training_data)\n\n# get all of the unique clusters\noptics_clusters = unique(optics_clusters)\n\n# plot OPTICS the clusters\nfor optics_cluster in optics_clusters:\n    # get data points that fall in this cluster\n    index = where(optics_result == optics_clusters)\n    # make the plot\n    pyplot.scatter(training_data[index, 0], training_data[index, 1])\n\n# show the OPTICS plot\npyplot.show()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-01T18:22:33.562227Z","iopub.status.idle":"2022-07-01T18:22:33.562993Z","shell.execute_reply.started":"2022-07-01T18:22:33.562700Z","shell.execute_reply":"2022-07-01T18:22:33.562720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# > 4.8. Agglomerative Hierarchy clustering algorithm\nThe agglomerative clustering is the most common type of hierarchical clustering used to group objects in clusters based on their similarity. It's also known as AGNES (Agglomerative Nesting). The algorithm starts by treating each object as a singleton cluster.","metadata":{}},{"cell_type":"markdown","source":"## 4.8.1. Advantages of Agglomerative Hierarchy clustering algorithm\n- No prior knowledge about the number of clusters is needed, although the user needs to define a threshold for divisions.\n- Easy to implement across various forms of data and known to provide robust results for data generated via various sources. Hence it has a wide application area.\n\n## 4.8.2. Disadvantages of Agglomerative Hierarchy clustering algorithm\n- The cluster division (DIANA) or combination (AGNES) is really strict and once performed, it cannot be undone and re-assigned in subsequesnt iterations or re-runs.\n- It has a high time complexity, in the order of O(n^2 log n) for all the n data-points, hence cannot be used for larger datasets.\n- Cannot handle outliers and noise","metadata":{}},{"cell_type":"markdown","source":"\n![](https://www.researchgate.net/profile/Satinder-Bal/publication/348137354/figure/fig2/AS:975396417830915@1609564036826/Agglomerative-hierarchical-clustering-algorithm.png)","metadata":{}},{"cell_type":"markdown","source":"## 4.8.3. Code Example: ","metadata":{}},{"cell_type":"code","source":"from numpy import unique\nfrom numpy import where\nfrom matplotlib import pyplot\nfrom sklearn.datasets import make_classification\nfrom sklearn.cluster import AgglomerativeClustering\n\n# initialize the data set we'll work with\ntraining_data, _ = make_classification(\n    n_samples=1000,\n    n_features=2,\n    n_informative=2,\n    n_redundant=0,\n    n_clusters_per_class=1,\n    random_state=4\n)\n\n# define the model\nagglomerative_model = AgglomerativeClustering(n_clusters=2)\n\n# assign each data point to a cluster\nagglomerative_result = agglomerative_model.fit_predict(training_data)\n\n# get all of the unique clusters\nagglomerative_clusters = unique(agglomerative_result)\n\n# plot the clusters\nfor agglomerative_cluster in agglomerative_clusters:\n    # get data points that fall in this cluster\n    index = where(agglomerative_result == agglomerative_clusters)\n    # make the plot\n    pyplot.scatter(training_data[index, 0], training_data[index, 1])\n\n# show the Agglomerative Hierarchy plot\npyplot.show()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-01T18:22:33.564410Z","iopub.status.idle":"2022-07-01T18:22:33.565176Z","shell.execute_reply.started":"2022-07-01T18:22:33.564955Z","shell.execute_reply":"2022-07-01T18:22:33.564980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4.9. DIANA or Divisive Analysis","metadata":{}},{"cell_type":"markdown","source":"DIANA is also known as DIvisie ANAlysis clustering algorithm. It is the top-down approach form of hierarchical clustering where all data points are initially assigned a single cluster. Further, the clusters are split into two least similar clusters. The divisive clustering algorithm is a top-down clustering approach, initially, all the points in the dataset belong to one cluster and split is performed recursively as one moves down the hierarchy. \n![](https://media.geeksforgeeks.org/wp-content/uploads/20190508025314/781ff66c-b380-4a78-af25-80507ed6ff261-300x300.png)","metadata":{}},{"cell_type":"markdown","source":"## 4.9.1. What are the advantages of divisive clustering techniques?\nDivisive clustering is more efficient if we do not generate a complete hierarchy all the way down to individual data leaves. The time complexity of a naive agglomerative clustering is O(n3) because we exhaustively scan the N x N matrix dist_mat for the lowest distance in each of N-1 iterations.","metadata":{}},{"cell_type":"markdown","source":"## 4.9.2. Code Example (R) :","metadata":{}},{"cell_type":"code","source":"\"\"\"\n# Compute diana()\nlibrary(cluster)\nres.diana <- diana(USArrests, stand = TRUE)\n\n# Plot the dendrogram\nlibrary(factoextra)\nfviz_dend(res.diana, cex = 0.5,\n          k = 4, # Cut in four groups\n          palette = \"jco\" # Color palette\n          )\n\n\"\"\"","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-01T18:22:33.566928Z","iopub.status.idle":"2022-07-01T18:22:33.567297Z","shell.execute_reply.started":"2022-07-01T18:22:33.567117Z","shell.execute_reply":"2022-07-01T18:22:33.567134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4.10. Fuzzy Analysis Clustering","metadata":{}},{"cell_type":"markdown","source":"Fuzzy c-means (FCM) is a data clustering technique in which a data set is grouped into N clusters with every data point in the dataset belonging to every cluster to a certain degree. For example, a data point that lies close to the center of a cluster will have a high degree of membership in that cluster, and another data point that lies far away from the center of a cluster will have a low degree of membership to that cluster.\n\nThe fcm function performs FCM clustering. It starts with a random initial guess for the cluster centers; that is the mean location of each cluster. Next, fcm assigns every data point a random membership grade for each cluster. By iteratively updating the cluster centers and the membership grades for each data point, fcm moves the cluster centers to the correct location within a data set and, for each data point, finds the degree of membership in each cluster. This iteration minimizes an objective function that represents the distance from any given data point to a cluster center weighted by the membership of that data point in the cluster.\n\nThis algorithm follows the fuzzy cluster assignment methodology of clustering. The working of FCM Algorithm is almost similar to the k-means – distance-based cluster assignment – however, the major difference is, as mentioned earlier, that according to this algorithm, a data point can be put into more than one cluster. \n\n![](https://es.mathworks.com/help/examples/fuzzy/win64/fcmdemo_codepad_01.png)","metadata":{}},{"cell_type":"markdown","source":"## 4.10.1. Advantages of Fuzzy Analysis Clustering\n- FCM works best for highly correlated and overlapped data, where k-means cannot give any conclusive results.\n- It is an unsupervised algorithm and it has a higher rate of convergence than other partitioning based algorithms.\n\n## 4.10.2. Disadvantages of Fuzzy Analysis Clustering\n- We need to specify the number of clusters “k” prior to the start of the algorithm\n- Although convergence is always guaranteed but the process is very slow and this cannot be used for larger data.\n- Prone to errors if the data has noise and outliers.","metadata":{}},{"cell_type":"markdown","source":"## 4.10.3. Code Example ","metadata":{}},{"cell_type":"code","source":"from __future__ import division, print_function\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport skfuzzy as fuzz\n\ncolors = ['b', 'orange', 'g', 'r', 'c', 'm', 'y', 'k', 'Brown', 'ForestGreen']\n\n# Define three cluster centers\ncenters = [[4, 2],\n           [1, 7],\n           [5, 6]]\n\n# Define three cluster sigmas in x and y, respectively\nsigmas = [[0.8, 0.3],\n          [0.3, 0.5],\n          [1.1, 0.7]]\n\n# Generate test data\nnp.random.seed(42)  # Set seed for reproducibility\nxpts = np.zeros(1)\nypts = np.zeros(1)\nlabels = np.zeros(1)\nfor i, ((xmu, ymu), (xsigma, ysigma)) in enumerate(zip(centers, sigmas)):\n    xpts = np.hstack((xpts, np.random.standard_normal(200) * xsigma + xmu))\n    ypts = np.hstack((ypts, np.random.standard_normal(200) * ysigma + ymu))\n    labels = np.hstack((labels, np.ones(200) * i))\n\n# Visualize the test data\nfig0, ax0 = plt.subplots()\nfor label in range(3):\n    ax0.plot(xpts[labels == label], ypts[labels == label], '.',\n             color=colors[label])\nax0.set_title('Test data: 200 points x3 clusters.')","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-01T18:22:33.569794Z","iopub.status.idle":"2022-07-01T18:22:33.570240Z","shell.execute_reply.started":"2022-07-01T18:22:33.570020Z","shell.execute_reply":"2022-07-01T18:22:33.570039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Comperative Analysis of Algorithms\nCredit: [*Sunit Prasad, Different Types of Clustering Methods and Applications*](https://www.analytixlabs.co.in/blog/types-of-clustering-algorithms/)","metadata":{}},{"cell_type":"markdown","source":"| Clustering Method                              | Description                                                                                                           | Advantages                                                                                                                                                                | Disadvantages                                                                                                                                          | Algorithms                                       |\n| ---------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------ |\n| Hierarchical Clustering                        | Based on top-to-bottom hierarchy of the data points to create clusters.                                               | Easy to implement, the number of clusters need not be specified apriori, dendrograms are easy to interpret.                                                               | Cluster assignment is strict and cannot be undone, high time complexity, cannot work for a larger dataset                                              | DIANA, AGNES, hclust etc.                        |\n| Partitioning methods                           | Based on centroids and data points are assigned into a cluster based on its proximity to the cluster centroid         | Easy to implement, faster processing, can work on larger data, easy to interpret the outputs                                                                              | We need to specify the number of cenrtroids apriori, clusters that get created are of inconsistent sizes and densities, Effected by noise and outliers | k-means, k-medians, k-modes                      |\n| Distribution-based Clustering                  | Based on the probability distribution of the data, clusters are derived from various metrics like mean, variance etc. | Number of clusters need not be specified apriori, works on real-time data, metrics are easy to understand and tune                                                        | Complex algorithm and slow, cannot be scaled to larger data                                                                                            | Gaussian Mixed Models, DBCLASD                   |\n| Density-based Clustering (Model-based methods) | Based on density of the data points, also known as model based clustering                                             | Can handle noise and outliers, need not specify number of clusters in the start, clusters that are created are highly homogenous, no restrictions on cluster shapes.      | Complex algorithm and slow, cannot be scaled to larger data                                                                                            | DENCAST, DBSCAN                                  |\n| Fuzzy Clustering                               | Based on Partitioning Approach but data points can belong to more than one cluster                                    | Can work on highly overlapped data, a higher rate of convergence                                                                                                          | We need to specify the number of centroids apriori, Effected by noise and outliers, Slow algorithm and cannot be scaled                                | Fuzzy C Means, Rough k means                     |\n| Constraint Based (Supervised Clustering)       | Clustering is directed and controlled by user constraints                                                             | Creates a perfect decision boundary, can automatically determine the outcome classes based on constraints, future data can be classified based on the training boundaries | Overfitting, high level of misclassification errors, cannot be trained on larger datasets                                                              | Decision Trees, Random Forest, Gradient Boosting |","metadata":{}},{"cell_type":"markdown","source":"## Appreciate, Upvote, Comment, Share, Enjoy!!","metadata":{}},{"cell_type":"markdown","source":"# References and Credits\n- [Different Types of Clustering Methods and Applications](https://www.analytixlabs.co.in/blog/types-of-clustering-algorithms/)\n- [The 5 Clustering Algorithms Data Scientists Need to Know](https://towardsdatascience.com/the-5-clustering-algorithms-data-scientists-need-to-know-a36d136ef68)\n- [What is Clustering and Different Types of Clustering Methods](https://www.upgrad.com/blog/clustering-and-types-of-clustering-methods/)\n- [8 Clustering Algorithms in Machine Learning that All Data Scientists Should Know](https://www.freecodecamp.org/news/8-clustering-algorithms-in-machine-learning-that-all-data-scientists-should-know/)\n- [Google Developers : Clustering Algorithms ](https://developers.google.com/machine-learning/clustering/clustering-algorithms)\n- [Clustering Technique](https://www.sciencedirect.com/topics/computer-science/clustering-technique)\n- [TYPES OF CLUSTERING METHODS: OVERVIEW AND QUICK START R CODE](https://www.datanovia.com/en/blog/types-of-clustering-methods-overview-and-quick-start-r-code/)\n- [5 Clustering Methods and Applications](https://www.analyticssteps.com/blogs/5-clustering-methods-and-applications)\n- [ML - Clustering Mean Shift Algorithm](https://www.tutorialspoint.com/machine_learning_with_python/machine_learning_with_python_clustering_algorithms_mean_shift.htm#:~:text=Advantages%20and%20Disadvantages&text=It%20can%20also%20model%20the,No%20problem%20generated%20from%20outliers.)\n- [Explain BIRCH algorithm with example](https://www.ques10.com/p/9298/explain-birch-algorithm-with-example/)\n- [Build Better and Accurate Clusters with Gaussian Mixture Models](https://www.analyticsvidhya.com/blog/2019/10/gaussian-mixture-models-clustering/)\n- [ML | BIRCH Clustering](https://www.geeksforgeeks.org/ml-birch-clustering/)\n- [Fig - Centroid Based](https://www.researchgate.net/figure/Centroid-based-clustering-algorithm_fig6_334279038)\n- [Agglomerative hierarchical clustering algorithm. | Download Scientific Diagram](https://www.google.com/search?q=Agglomerative+Hierarchy+clustering+algorithm&rlz=1C1UEAD_enBD996BD996&hl=en&sxsrf=ALiCzsbN-VXLpxg4c8kxBys8tx8VzMsvlw:1656698373967&source=lnms&tbm=isch&sa=X&ved=2ahUKEwiS9NGwotj4AhXo-TgGHTM-DXAQ_AUoAXoECAIQAw&biw=1920&bih=937&dpr=1#imgrc=iROfxWllHmNg0M)\n- [HOW THE HIERARCHICAL CLUSTERING ALGORITHM WORKS](https://dataaspirant.com/hierarchical-clustering-algorithm/)\n- [Agglomerative Hierarchical Clustering - Datanovia](https://www.datanovia.com/en/lessons/agglomerative-hierarchical-clustering/#:~:text=The%20agglomerative%20clustering%20is%20the,object%20as%20a%20singleton%20cluster.)\n- [Divisive Hierarchical Clustering](https://www.datanovia.com/en/lessons/divisive-hierarchical-clustering/)\n- [Fuzzy C-Means Clustering](https://es.mathworks.com/help/fuzzy/fuzzy-c-means-clustering.html;jsessionid=af46b531dd62efa365ff83aa21ed)\n- [Fuzzy C-Means Clustering](https://pythonhosted.org/scikit-fuzzy/auto_examples/plot_cmeans.html)\n- [www.analytixlabs.co.in](www.analytixlabs.co.in)","metadata":{}}]}