{"cells":[{"metadata":{"_cell_guid":"53f49b34-882d-41bb-a2e8-66b1993c9943","_uuid":"833e19a6ad7945421071c89bd131115d03501402"},"cell_type":"markdown","source":"# Visualization: True Tracks VS Clustered Hits (Predicted Tracks)\nI created this kernel because I wanted to know what the tracks look like, what the clusters (made w/ DBSCAN) look like, and if the clustering is working correctly (or working at all) aside from knowing the LB score.\n\nI will show plots of tracks and plots of clustered hits (or predicted tracks) on the same points. The tracks are plotted using the raw hits coordinates and also using the transformed hits coordinates (based on the DBSCAN benchmark kernel). I will also relate the plots to how the clustering is expected to work. Clustered hits are visualized for 4 different settings using DBSCAN.\n\nThe code for the track visualizations are based on Joshua Bonatt's kernel: https://www.kaggle.com/jbonatt/trackml-eda-etc\n\nThe code for the coordinate transformation are based on Mikhail Hushchyn's DBSCAN benchmark kernel: https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark\n\nLearned how to import trackml from Wesam Elshamy's kernel: https://www.kaggle.com/wesamelshamy/trackml-problem-explanation-and-data-exploration"},{"metadata":{"_cell_guid":"3700a5fe-da95-402f-b4ea-686d7f1c0756","_uuid":"3c3a912b2e9e3a4bdc4dc620a39a36d4921e3ab1"},"cell_type":"markdown","source":"# Table of Contents\n- Import Libraries\n- Loading a Single Event\n- Track Visualization\n    - Raw Hits Coordinates\n    - Preprocessesd Hits Data (w/ Coordinate Transformation)\n- Clustered Hits Visualization"},{"metadata":{"_cell_guid":"f42cacce-1cf4-4c7e-9a9f-2b00cf412f3f","_uuid":"2cb38ae99e3c2e6dc4b45a9548c7102c7753915f"},"cell_type":"markdown","source":"# Import Libraries"},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","collapsed":true,"trusted":true},"cell_type":"code","source":"import os\n\nimport numpy as np\nimport pandas as pd\n\nfrom trackml.dataset import load_event\nfrom trackml.score import score_event\n\nfrom sklearn import metrics\nfrom sklearn.cluster import DBSCAN\nfrom sklearn.preprocessing import StandardScaler, MaxAbsScaler, Normalizer\n\nimport matplotlib.pyplot as plt\nfrom mpl_toolkits.mplot3d import Axes3D\n%matplotlib inline\n\nfrom scipy import spatial","execution_count":1,"outputs":[]},{"metadata":{"_cell_guid":"418e9883-75b1-4255-8799-01f4cd1d4c2d","_uuid":"4d621f8ecc04fae3d9e20bca717e3fe4f23195af"},"cell_type":"markdown","source":"## function to be used later\na function for arranging hits in a track"},{"metadata":{"_cell_guid":"b02b5bf2-3a4f-452f-ab89-153e7979abba","_kg_hide-input":true,"_kg_hide-output":true,"_uuid":"dde2dde5cef2dbeee43600235dc640904eedf9db","trusted":true},"cell_type":"code","source":"# function for arranging hits in a track (closest to origin, closest track to previous track, and so on...)\ndef arrange_track(track_points):\n    arranged_track = pd.DataFrame()\n\n    pt = [0, 0, 0]\n    kdtree = spatial.KDTree(track_points)\n    distance, index = kdtree.query(pt)\n\n    arranged_track = arranged_track.append(track_points.iloc[index])\n    track_points = track_points.drop(track_points.index[index]).reset_index(drop=True)\n\n    while not track_points.empty:\n        pt = arranged_track.iloc[-1]\n        kdtree = spatial.KDTree(track_points)\n        distance, index = kdtree.query(pt)\n\n        arranged_track = arranged_track.append(track_points.iloc[index])\n        track_points = track_points.drop(track_points.index[index]).reset_index(drop=True)\n        \n    return arranged_track\n\ntest_points = pd.DataFrame([[0, 0, 5], [0, 0, 1], [0, 0, 3], [0, 0, 2]])\narrange_track(test_points)","execution_count":2,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true},"cell_type":"markdown","source":"# Load a single event from train set"},{"metadata":{"_cell_guid":"086fa4d8-aff2-4771-a248-ca899ddfbf3a","_kg_hide-output":true,"_uuid":"cca10a1397e0e6b13ee8f18d807a4831aead3dc5","trusted":true},"cell_type":"code","source":"hits, cells, particles, truth = load_event('../input/train_1/event000001000')\nhits.head()","execution_count":3,"outputs":[]},{"metadata":{"_cell_guid":"3cbb09e3-9871-44c4-9016-714ef7d01bc3","_kg_hide-output":true,"_uuid":"f1a511e6f4f6e01076f1eb11e1df2479bc231b5d","trusted":true},"cell_type":"code","source":"truth.head()","execution_count":4,"outputs":[]},{"metadata":{"_cell_guid":"57f09e3a-c5c5-4e86-a8b5-b3992e27f478","_uuid":"d98f2c97c8f37dabcbf0e609b678f4a1dbb056c1"},"cell_type":"markdown","source":"# Track Visualization\nvisualization code is based on the kernel: https://www.kaggle.com/jbonatt/trackml-eda-etc\n## Raw Hits Data\nThe following plots show the tracks as they are. The tracks are identified by color using the **truth** data but the coordinates used are from the **hits** data. I use the term \"True Tracks\" to distinguish these plots from the predicted tracks later on.\n### every 500th track"},{"metadata":{"_cell_guid":"cd8c22d6-20df-4631-9a50-a29743c73047","_kg_hide-input":true,"_uuid":"e70f5390962279ecd109e92e3bc43de9ce3e19eb","trusted":true},"cell_type":"code","source":"tracks = truth.particle_id.unique()[1::500]\nfig = plt.figure(figsize=(20,7))\n\nax = fig.add_subplot(121,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x', 'y', 'z']]\n    ax.plot3D(t.z, t.x, t.y, '.', ms=10)\n    \nax.set_xlabel('z (mm)')\nax.set_ylabel('x (mm)')\nax.set_zlabel('y (mm)')\nax.set_title(\"True Tracks (Scatter Plot)\", y=-.15, size=20)\n\nax2 = fig.add_subplot(122,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x', 'y', 'z']]\n    t = arrange_track(t)\n    ax2.plot3D(t.z, t.x, t.y, '.-', ms=10)\n    \nax2.set_xlabel('z (mm)')\nax2.set_ylabel('x (mm)')\nax2.set_zlabel('y (mm)')\nax2.set_title(\"True Tracks (Line Plot)\", y=-.15, size=20)\n\nplt.show()","execution_count":8,"outputs":[]},{"metadata":{"_cell_guid":"02222cc7-b105-4cdc-9f1d-e19740c9ed77","_uuid":"c99449a4f5ef68d55b7222e5dd82a7dc049c895f"},"cell_type":"markdown","source":"### every 100th track"},{"metadata":{"_cell_guid":"6fba11a9-5340-4d9f-9c78-1b7d4a36da92","_kg_hide-input":true,"_uuid":"a3d506a8497dde4f93eb004ede69654f80619b02","trusted":true},"cell_type":"code","source":"tracks = truth.particle_id.unique()[1::100]\nfig = plt.figure(figsize=(20,7))\n\nax = fig.add_subplot(121,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x', 'y', 'z']]\n    ax.plot3D(t.z, t.x, t.y, '.', ms=10)\n    \nax.set_xlabel('z (mm)')\nax.set_ylabel('x (mm)')\nax.set_zlabel('y (mm)')\nax.set_title(\"True Tracks (Scatter Plot)\", y=-.15, size=20)\n\nax2 = fig.add_subplot(122,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x', 'y', 'z']]\n    t = arrange_track(t)\n    ax2.plot3D(t.z, t.x, t.y, '.-', ms=10)\n    \nax2.set_xlabel('z (mm)')\nax2.set_ylabel('x (mm)')\nax2.set_zlabel('y (mm)')\nax2.set_title(\"True Tracks (Line Plot)\", y=-.15, size=20)\n\nplt.show()","execution_count":9,"outputs":[]},{"metadata":{"_cell_guid":"6f931937-84fb-4ef9-b8b9-4cf9d3ad5f49","_uuid":"4b7e2aff3da9296f881144cf0b7494a8a15b9bef"},"cell_type":"markdown","source":"Looking at the plots above, we can see a lot of points near the origin that are very close to each other but belong to different tracks. Because of this, we can expect that using DBSCAN (or other clustering algorithms like K-means) will not be effective because these points will probably be clustered together."},{"metadata":{"_cell_guid":"6e0dc19a-b40b-4612-98e6-ab45c9d4ab63","_uuid":"a339fc73cf0eb229a3d96e5bcab6419aa5951d32"},"cell_type":"markdown","source":"## Preprocessed Hits Data\nThe following plots show the tracks after going through the coordinate transformation performed in the DBSCAN benchmark kernel: https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark\n### Preprocessing / Coordinate Transformation"},{"metadata":{"_cell_guid":"985b1486-f7ac-4edf-8438-1b8abde33add","_kg_hide-input":true,"_uuid":"3297da54a5639862c09b3eb058116415d596f8a6","collapsed":true,"trusted":true},"cell_type":"code","source":"# DBSCAN benchmark preprocessing / coordinate transformation\nx = hits.x.values\ny = hits.y.values\nz = hits.z.values\n\nr = np.sqrt(x**2 + y**2 + z**2)\nhits['x2'] = x/r\nhits['y2'] = y/r\n\nr = np.sqrt(x**2 + y**2)\nhits['z2'] = z/r","execution_count":11,"outputs":[]},{"metadata":{"_cell_guid":"c8d82ace-6c8e-4a8e-af2e-c068c8e8451f","_uuid":"5ffdd3e774bfb4ece2a225f38a1400a9f04e12e1"},"cell_type":"markdown","source":"### every 500th track"},{"metadata":{"_cell_guid":"136375c3-2554-404d-b524-78d0be1fca7f","_kg_hide-input":true,"_uuid":"633bf27da8d2015d8e3163dd3d5dd70dc0942f87","trusted":true},"cell_type":"code","source":"tracks = truth.particle_id.unique()[1::500]\nfig = plt.figure(figsize=(20,7))\n\nax = fig.add_subplot(121,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x2', 'y2', 'z2']]\n    ax.plot3D(t.z2, t.x2, t.y2, '.', ms=10)\n    \nax.set_xlabel('z2')\nax.set_ylabel('x2')\nax.set_zlabel('y2')\nax.set_title(\"True Tracks (Scatter Plot)\", y=-.15, size=20)\n\nax2 = fig.add_subplot(122,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x2', 'y2', 'z2']]\n    t = arrange_track(t)\n    ax2.plot3D(t.z2, t.x2, t.y2, '.-', ms=10)\n    \nax2.set_xlabel('z2')\nax2.set_ylabel('x2')\nax2.set_zlabel('y2')\nax2.set_title(\"True Tracks (Line Plot)\", y=-.15, size=20)\n\nplt.show()","execution_count":12,"outputs":[]},{"metadata":{"_cell_guid":"c69a2ac9-f5cb-457a-9412-87a15250bd45","_uuid":"2cc493143ffafb8dd4e8ad3a180497fcaee9cb8b"},"cell_type":"markdown","source":"Mikhail, the author of the DBSCAN benchmark kernel (https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark) mentioned that \"The transformation just make hits from the same track closer to each other.\" Indeed, that is evident in these plots. Looking at the plots above, DBSCAN or other clustering algorithms might work since the tracks appear separated from each other. However, look at the next plots."},{"metadata":{"_cell_guid":"76a0070b-57be-4238-8533-80bacc64779d","_uuid":"fa27fbd5c163bdcc557390d270b6860de42b3299"},"cell_type":"markdown","source":"### every 100th track"},{"metadata":{"_cell_guid":"562c9a46-7201-44da-a8fb-d5c2cf0de983","_kg_hide-input":true,"_uuid":"0ab9f06b53f0c612593f70bd716fdfaf94369fdb","trusted":true},"cell_type":"code","source":"tracks = truth.particle_id.unique()[1::100]\nfig = plt.figure(figsize=(20,7))\n\nax = fig.add_subplot(121,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x2', 'y2', 'z2']]\n    ax.plot3D(t.z2, t.x2, t.y2, '.', ms=10)\n    \nax.set_xlabel('z2)')\nax.set_ylabel('x2')\nax.set_zlabel('y2')\nax.set_title(\"True Tracks (Scatter Plot)\", y=-.15, size=20)\n\nax2 = fig.add_subplot(122,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x2', 'y2', 'z2']]\n    t = arrange_track(t)\n    ax2.plot3D(t.z2, t.x2, t.y2, '.-', ms=10)\n    \nax2.set_xlabel('z2')\nax2.set_ylabel('x2')\nax2.set_zlabel('y2')\nax2.set_title(\"True Tracks (Line Plot)\", y=-.15, size=20)\n\nplt.show()","execution_count":13,"outputs":[]},{"metadata":{"_cell_guid":"45a12e1f-749d-490c-8153-9af4abbc3c2d","_uuid":"09c115da92bc9ab7f51f95b05087603a9ccac2ea"},"cell_type":"markdown","source":"As more tracks are included in the plot, we can see that many tracks are again close to each other, which means that clustering is still difficult."},{"metadata":{"_cell_guid":"ac46bdaf-c18a-4a18-b902-e1983a331d86","_uuid":"8bd7ad22fed1c973007fb995664ee2111f6e5af7"},"cell_type":"markdown","source":"# Clustered Hits Visualization\nNext, on the same points, let's visualize the resulting clusters (or predicted tracks)."},{"metadata":{"_cell_guid":"86eddb10-bdd5-4149-829f-8d807eb0c9ca","_uuid":"da1e4459a7ce75123dc7655819eacbff10b8363e"},"cell_type":"markdown","source":"## DBSCAN Clustering Setting 1\n## based on the benchmark: coordinate transformation + dbscan"},{"metadata":{"_cell_guid":"9eee5947-62cf-48f7-bb28-5984b5642515","_kg_hide-input":true,"_uuid":"01d4fafaf17b754596c86dec724ad69d08174653","collapsed":true,"trusted":true},"cell_type":"code","source":"X = hits[['x2', 'y2', 'z2']]\nscaler = StandardScaler().fit(X)\nX = scaler.transform(X)\n\neps = 0.008\nmin_samp = 1\ndb = DBSCAN(eps=eps, min_samples=min_samp, metric='euclidean').fit(X)\nlabels = db.labels_\n\nclustering = pd.DataFrame()\nclustering['hit_id'] = truth['hit_id']\nclustering['track_id'] = labels\n\nscore = score_event(truth, clustering)\nprint('track-ml custom metric score:', round(score, 4))\n\nlabels_true = truth['particle_id']\nn_clusters = len(set(labels)) - (1 if -1 in labels else 0)\n\nprint('\\nOTHER CLUSTERING RESULTS:')\nprint('Estimated number of clusters: %d' % n_clusters)\nprint(\"Homogeneity: %0.3f\" % metrics.homogeneity_score(labels_true, labels))\nprint(\"Completeness: %0.3f\" % metrics.completeness_score(labels_true, labels))\nprint(\"Adjusted Rand Index: %0.3f\" % metrics.adjusted_rand_score(labels_true, labels))\nrej_perc = list(labels).count(-1) / float(hits.shape[0]) * 100\nrej_perc = round(rej_perc, 2)\nprint (\"Rejected samples %:\", str(rej_perc) + '%')\nrejected_count = list(labels).count(-1)\nprint (\"Rejected samples:\", rejected_count)\nprint (\"Total samples:\", hits.shape[0])\nprint (\"Clustered samples:\", hits.shape[0] - list(labels).count(-1))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"a3bc5fa9-4ec5-4d1d-aea0-59babf8629ee","_uuid":"bef3bf41e7cfb4a69697c71b63433f22fc028489"},"cell_type":"markdown","source":"## Clustered Hits (Predicted Tracks) vs True Tracks - Transformed Coordinates"},{"metadata":{"_cell_guid":"206764eb-e8a6-4928-8534-4be1a4767fc8","_kg_hide-input":true,"_uuid":"06ac076174538258632b87f2a363da99c3fb95a8","collapsed":true,"scrolled":false,"trusted":true},"cell_type":"code","source":"tracks = truth.particle_id.unique()[1::100]\nfig = plt.figure(figsize=(20,7))\n\nax = fig.add_subplot(121,projection='3d')\n\ntracks_hit_ids = truth[truth['particle_id'].isin(tracks)]['hit_id'] # all hits in tracks\nclusters = clustering[clustering['hit_id'].isin(tracks_hit_ids)].track_id.unique() # all clusters containing the hits in tracks\nfor cluster in clusters:\n    cluster_hit_ids = clustering[clustering['track_id'] == cluster]['hit_id'] # all hits in cluster\n    plot_hit_ids = list(set(tracks_hit_ids) & set(cluster_hit_ids))\n    t = hits[hits['hit_id'].isin(plot_hit_ids)][['x2', 'y2', 'z2']]\n    if cluster == -1:\n        ax.plot3D(t.z2, t.x2, t.y2, '.', ms=10, color='black')\n    else:\n        ax.plot3D(t.z2, t.x2, t.y2, '.-', ms=10)\n    \nax.set_xlabel('z2')\nax.set_ylabel('x2')\nax.set_zlabel('y2')\nax.set_title('Clustered Hits (Predicted Tracks)', y=-.15, size=20)\n\nax2 = fig.add_subplot(122,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x2', 'y2', 'z2']]\n    t = arrange_track(t)\n    ax2.plot3D(t.z2, t.x2, t.y2, '.-', ms=10)\n    \nax2.set_xlabel('z2')\nax2.set_ylabel('x2')\nax2.set_zlabel('y2')\nax2.set_title('True Tracks', y=-.15, size=20)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"df7061e6-724e-4d2b-8bce-52a524659fd4","_uuid":"ca4b70a7916fca4e43f46caa827aa545edfd3da9"},"cell_type":"markdown","source":"Here we see that the clusters are really small and we can't actually observe any tracks."},{"metadata":{"_cell_guid":"7aaee4c0-1663-4b60-ad86-25470a967aba","_uuid":"7bbaaabbd0ea1dc12512c63c68faecdacffc96de"},"cell_type":"markdown","source":"## Clustered Hits (Predicted Tracks) vs True Tracks - Raw Hits Coordinates"},{"metadata":{"_cell_guid":"283e89c7-db40-4f84-b4a5-9387cee1b250","_kg_hide-input":true,"_uuid":"515eaf6eff5a754541fe9109f19c486c4426c238","collapsed":true,"scrolled":false,"trusted":true},"cell_type":"code","source":"tracks = truth.particle_id.unique()[1::100]\nfig = plt.figure(figsize=(20,7))\n\nax = fig.add_subplot(121,projection='3d')\n\ntracks_hit_ids = truth[truth['particle_id'].isin(tracks)]['hit_id'] # all hits in tracks\nclusters = clustering[clustering['hit_id'].isin(tracks_hit_ids)].track_id.unique() # all clusters containing the hits in tracks\nfor cluster in clusters:\n    cluster_hit_ids = clustering[clustering['track_id'] == cluster]['hit_id'] # all hits in cluster\n    plot_hit_ids = list(set(tracks_hit_ids) & set(cluster_hit_ids))\n    t = hits[hits['hit_id'].isin(plot_hit_ids)][['x', 'y', 'z']]\n    if cluster == -1:\n        ax.plot3D(t.z, t.x, t.y, '.', ms=10, color='black')\n    else:\n        ax.plot3D(t.z, t.x, t.y, '.-', ms=10)\n    \nax.set_xlabel('z (mm)')\nax.set_ylabel('x (mm)')\nax.set_zlabel('y (mm)')\nax.set_title('Clustered Hits (Predicted Tracks)', y=-.15, size=20)\n\nax2 = fig.add_subplot(122,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x', 'y', 'z']]\n    t = arrange_track(t)\n    ax2.plot3D(t.z, t.x, t.y, '.-', ms=10)\n    \nax2.set_xlabel('z (mm)')\nax2.set_ylabel('x (mm)')\nax2.set_zlabel('y (mm)')\nax2.set_title('True Tracks', y=-.15, size=20)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"07feb123-7fb5-45f9-ae0e-cc71c884eae3","_uuid":"25f1e2dbf3abe3c51dbb320db96de84e6f6759df"},"cell_type":"markdown","source":"Here, we can see a few formed tracks but still many small clusters. The tracks that were formed are straight and do not follow the shape of the true tracks."},{"metadata":{"_cell_guid":"93874845-125b-4e9f-8a91-971bf26468ee","_uuid":"ce88902874a6240d34464cd67b7c730ada17ed75"},"cell_type":"markdown","source":"## DBSCAN Clustering Setting 2\n## based on the benchmark: coordinate transformation + dbscan\n## high eps (eps=0.018)\nIncreasing the eps means decreasing the density required to form a cluster. eps is the distance that is used to define the neighbors of a sample. Since we increased eps, we expect bigger clusters to be formed."},{"metadata":{"_cell_guid":"c8279b8e-ce69-4198-a412-dac05999fe2d","_kg_hide-input":true,"_uuid":"c9771c3cc97d1d22be1d84ae8b79ed96ce09efef","collapsed":true,"trusted":true},"cell_type":"code","source":"X = hits[['x2', 'y2', 'z2']]\nscaler = StandardScaler().fit(X)\nX = scaler.transform(X)\n\neps = 0.018\nmin_samp = 1\ndb = DBSCAN(eps=eps, min_samples=min_samp, metric='euclidean').fit(X)\nlabels = db.labels_\n\nclustering = pd.DataFrame()\nclustering['hit_id'] = truth['hit_id']\nclustering['track_id'] = labels\n\nscore = score_event(truth, clustering)\nprint('track-ml custom metric score:', round(score, 4))\n\nlabels_true = truth['particle_id']\nn_clusters = len(set(labels)) - (1 if -1 in labels else 0)\n\nprint('\\nOTHER CLUSTERING RESULTS:')\nprint('Estimated number of clusters: %d' % n_clusters)\nprint(\"Homogeneity: %0.3f\" % metrics.homogeneity_score(labels_true, labels))\nprint(\"Completeness: %0.3f\" % metrics.completeness_score(labels_true, labels))\nprint(\"Adjusted Rand Index: %0.3f\" % metrics.adjusted_rand_score(labels_true, labels))\nrej_perc = list(labels).count(-1) / float(hits.shape[0]) * 100\nrej_perc = round(rej_perc, 2)\nprint (\"Rejected samples %:\", str(rej_perc) + '%')\nrejected_count = list(labels).count(-1)\nprint (\"Rejected samples:\", rejected_count)\nprint (\"Total samples:\", hits.shape[0])\nprint (\"Clustered samples:\", hits.shape[0] - list(labels).count(-1))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"668e6449-7ed8-4983-be45-d7278c0f952d","_uuid":"a5468011a72c3887c2f9d8f1b3057b4806b9881f"},"cell_type":"markdown","source":"## Clustered Hits (Predicted Tracks) vs True Tracks - Transformed Coordinates"},{"metadata":{"_cell_guid":"6a38e69c-0db2-4503-baf4-99d861d94d56","_kg_hide-input":true,"_uuid":"a69490b875e75493a5ae1f68a6561d177ab36b9b","collapsed":true,"scrolled":false,"trusted":true},"cell_type":"code","source":"tracks = truth.particle_id.unique()[1::100]\nfig = plt.figure(figsize=(20,7))\n\nax = fig.add_subplot(121,projection='3d')\n\ntracks_hit_ids = truth[truth['particle_id'].isin(tracks)]['hit_id'] # all hits in tracks\nclusters = clustering[clustering['hit_id'].isin(tracks_hit_ids)].track_id.unique() # all clusters containing the hits in tracks\nfor cluster in clusters:\n    cluster_hit_ids = clustering[clustering['track_id'] == cluster]['hit_id'] # all hits in cluster\n    plot_hit_ids = list(set(tracks_hit_ids) & set(cluster_hit_ids))\n    t = hits[hits['hit_id'].isin(plot_hit_ids)][['x2', 'y2', 'z2']]\n    if cluster == -1:\n        ax.plot3D(t.z2, t.x2, t.y2, '.', ms=10, color='black')\n    else:\n        ax.plot3D(t.z2, t.x2, t.y2, '.-', ms=10)\n    \nax.set_xlabel('z2')\nax.set_ylabel('x2')\nax.set_zlabel('y2')\nax.set_title('Clustered Hits (Predicted Tracks)', y=-.15, size=20)\n\nax2 = fig.add_subplot(122,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x2', 'y2', 'z2']]\n    t = arrange_track(t)\n    ax2.plot3D(t.z2, t.x2, t.y2, '.-', ms=10)\n    \nax2.set_xlabel('z2')\nax2.set_ylabel('x2')\nax2.set_zlabel('y2')\nax2.set_title('True Tracks', y=-.15, size=20)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"915bd26f-09af-499a-8eb2-0020f60eb7d4","_uuid":"d560b20a759fb7d6632bff374450eb492649c723"},"cell_type":"markdown","source":"As expected, bigger clusters were formed. We can now observe a few tracks. However, the predicted tracks don't seem to match the true tracks. There is also a strange looking predicted track (orange cluster in the middle)."},{"metadata":{"_cell_guid":"7bf533c3-8f5a-44fb-8581-68a77eaeb0bb","_uuid":"c90c56dd2b40b61c05adcaaca63eacda188b226e"},"cell_type":"markdown","source":"## Clustered Hits (Predicted Tracks) vs True Tracks - Raw Hits Coordinates"},{"metadata":{"_cell_guid":"83cfb96c-23d4-4f31-be48-e0cbb31b7d7e","_kg_hide-input":true,"_uuid":"ad6854a90524ed24649f73e3a603ae6ee9bb3f21","collapsed":true,"trusted":true},"cell_type":"code","source":"tracks = truth.particle_id.unique()[1::100]\nfig = plt.figure(figsize=(20,7))\n\nax = fig.add_subplot(121,projection='3d')\n\ntracks_hit_ids = truth[truth['particle_id'].isin(tracks)]['hit_id'] # all hits in tracks\nclusters = clustering[clustering['hit_id'].isin(tracks_hit_ids)].track_id.unique() # all clusters containing the hits in tracks\nfor cluster in clusters:\n    cluster_hit_ids = clustering[clustering['track_id'] == cluster]['hit_id'] # all hits in cluster\n    plot_hit_ids = list(set(tracks_hit_ids) & set(cluster_hit_ids))\n    t = hits[hits['hit_id'].isin(plot_hit_ids)][['x', 'y', 'z']]\n    if cluster == -1:\n        ax.plot3D(t.z, t.x, t.y, '.', ms=10, color='black')\n    else:\n        ax.plot3D(t.z, t.x, t.y, '.-', ms=10)\n    \nax.set_xlabel('z (mm)')\nax.set_ylabel('x (mm)')\nax.set_zlabel('y (mm)')\nax.set_title('Clustered Hits (Predicted Tracks)', y=-.15, size=20)\n\nax2 = fig.add_subplot(122,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x', 'y', 'z']]\n    t = arrange_track(t)\n    ax2.plot3D(t.z, t.x, t.y, '.-', ms=10)\n    \nax2.set_xlabel('z (mm)')\nax2.set_ylabel('x (mm)')\nax2.set_zlabel('y (mm)')\nax2.set_title('True Tracks', y=-.15, size=20)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"43ec75e5-c508-4539-8cc4-3eb8420d7007","_uuid":"9ab4dd15de7b60f6de8cd5b94866a82479f7dc36"},"cell_type":"markdown","source":"We also observe more tracks in this coordinate system. The strange orange track is seen more clearly here. These are hits that actually belong to different tracks but were clustered together."},{"metadata":{"_cell_guid":"cc413a75-af17-4f13-9ec6-87952310dc8e","_uuid":"1dda933e8e90b91307053a6760ab1ecbb80d41c5"},"cell_type":"markdown","source":"## DBSCAN Clustering Setting 3\n## based on the benchmark: coordinate transformation + dbscan\n## min_samp=3\nIncreasing the minimum samples means decreasing the density required to form a cluster. Let's see the effects of this adjustment to the score and clustering metrics."},{"metadata":{"_cell_guid":"7c721ee7-a2da-48b2-b53f-644d5f0bfe75","_kg_hide-input":true,"_uuid":"87a46bcdefc3021d2f938368e390b09f8738c526","collapsed":true,"trusted":true},"cell_type":"code","source":"X = hits[['x2', 'y2', 'z2']]\nscaler = StandardScaler().fit(X)\nX = scaler.transform(X)\n\neps = 0.008\nmin_samp = 3\ndb = DBSCAN(eps=eps, min_samples=min_samp, metric='euclidean').fit(X)\nlabels = db.labels_\n\nclustering = pd.DataFrame()\nclustering['hit_id'] = truth['hit_id']\nclustering['track_id'] = labels\n\nscore = score_event(truth, clustering)\nprint('track-ml custom metric score:', round(score, 4))\n\nlabels_true = truth['particle_id']\nn_clusters = len(set(labels)) - (1 if -1 in labels else 0)\n\nprint('\\nOTHER CLUSTERING RESULTS:')\nprint('Estimated number of clusters: %d' % n_clusters)\nprint(\"Homogeneity: %0.3f\" % metrics.homogeneity_score(labels_true, labels))\nprint(\"Completeness: %0.3f\" % metrics.completeness_score(labels_true, labels))\nprint(\"Adjusted Rand Index: %0.3f\" % metrics.adjusted_rand_score(labels_true, labels))\nrej_perc = list(labels).count(-1) / float(hits.shape[0]) * 100\nrej_perc = round(rej_perc, 2)\nprint (\"Rejected samples %:\", str(rej_perc) + '%')\nrejected_count = list(labels).count(-1)\nprint (\"Rejected samples:\", rejected_count)\nprint (\"Total samples:\", hits.shape[0])\nprint (\"Clustered samples:\", hits.shape[0] - list(labels).count(-1), '\\n')\n\nprint ('WITHOUT REJECTED SAMPLES:')\nlabels_true_wr = labels_true[labels != -1]\nlabels_wr = labels[labels != -1]\nprint(\"Homogeneity: %0.3f\" % metrics.homogeneity_score(labels_true_wr, labels_wr))\nprint((\"Completeness: %0.3f\" % metrics.completeness_score(labels_true_wr, labels_wr)), '\\n')","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"07d07dd4-d57b-4a68-8f00-c2f75b226b58","_uuid":"dc4c76366225e6b8287e0f505bc1392686be8879"},"cell_type":"markdown","source":"Oddly enough, the score is exactly the same, while the clustering metrics drastically changed compared to Setting 1. In this setting, we observe a lot of rejected samples (samples that did not join any cluster). This is not suprising since the **min_samp** parameter also determines the minimum number of samples in a cluster. This result means that there are hits that are too far from other hits and were unable to join a cluster."},{"metadata":{"_cell_guid":"6d4ab5ef-8c78-4e4c-892d-9558ad723882","_uuid":"e104ee908f5377f4d103c7cf05140a6fe2f7cf2c"},"cell_type":"markdown","source":"## Clustered Hits (Predicted Tracks) vs True Tracks - Transformed Coordinates"},{"metadata":{"_cell_guid":"3843339b-dac2-4455-8e7d-a22701da7722","_kg_hide-input":true,"_uuid":"d0bc93184ef3f6b72345b5bc34da5dc98b4ce525","collapsed":true,"trusted":true},"cell_type":"code","source":"tracks = truth.particle_id.unique()[1::100]\nfig = plt.figure(figsize=(20,7))\n\nax = fig.add_subplot(121,projection='3d')\n\ntracks_hit_ids = truth[truth['particle_id'].isin(tracks)]['hit_id'] # all hits in tracks\nclusters = clustering[clustering['hit_id'].isin(tracks_hit_ids)].track_id.unique() # all clusters containing the hits in tracks\nfor cluster in clusters:\n    cluster_hit_ids = clustering[clustering['track_id'] == cluster]['hit_id'] # all hits in cluster\n    plot_hit_ids = list(set(tracks_hit_ids) & set(cluster_hit_ids))\n    t = hits[hits['hit_id'].isin(plot_hit_ids)][['x2', 'y2', 'z2']]\n    if cluster == -1:\n        ax.plot3D(t.z2, t.x2, t.y2, '.', ms=10, color='black')\n    else:\n        ax.plot3D(t.z2, t.x2, t.y2, '.-', ms=10)\n    \nax.set_xlabel('z2')\nax.set_ylabel('x2')\nax.set_zlabel('y2')\nax.set_title('Clustered Hits (Predicted Tracks)', y=-.15, size=20)\n\nax2 = fig.add_subplot(122,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x2', 'y2', 'z2']]\n    t = arrange_track(t)\n    ax2.plot3D(t.z2, t.x2, t.y2, '.-', ms=10)\n    \nax2.set_xlabel('z2')\nax2.set_ylabel('x2')\nax2.set_zlabel('y2')\nax2.set_title('True Tracks', y=-.15, size=20)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"6346f2f0-dc79-4454-afdd-aff8e8505f9c","_uuid":"390eb61c40aeea9fbf75d03bb2dc26fbd65f9fcd"},"cell_type":"markdown","source":"The rejected hits are plotted as black points."},{"metadata":{"_cell_guid":"2fcd58e8-a88f-46df-aa70-7766352e676c","_uuid":"d5a5ed1518bfdf361fe30ff5880ae48ad1a08d62"},"cell_type":"markdown","source":"## Clustered Hits (Predicted Tracks) vs True Tracks - Raw Hits Coordinates"},{"metadata":{"_cell_guid":"15315125-67a8-4d0c-9765-347ece23c83b","_kg_hide-input":true,"_uuid":"71803a1e3041a7cdd4a96b46c075f08435da4a0a","collapsed":true,"trusted":true},"cell_type":"code","source":"tracks = truth.particle_id.unique()[1::100]\nfig = plt.figure(figsize=(20,7))\n\nax = fig.add_subplot(121,projection='3d')\n\ntracks_hit_ids = truth[truth['particle_id'].isin(tracks)]['hit_id'] # all hits in tracks\nclusters = clustering[clustering['hit_id'].isin(tracks_hit_ids)].track_id.unique() # all clusters containing the hits in tracks\nfor cluster in clusters:\n    cluster_hit_ids = clustering[clustering['track_id'] == cluster]['hit_id'] # all hits in cluster\n    plot_hit_ids = list(set(tracks_hit_ids) & set(cluster_hit_ids))\n    t = hits[hits['hit_id'].isin(plot_hit_ids)][['x', 'y', 'z']]\n    if cluster == -1:\n        ax.plot3D(t.z, t.x, t.y, '.', ms=10, color='black')\n    else:\n        ax.plot3D(t.z, t.x, t.y, '.-', ms=10)\n    \nax.set_xlabel('z (mm)')\nax.set_ylabel('x (mm)')\nax.set_zlabel('y (mm)')\nax.set_title('Clustered Hits (Predicted Tracks)', y=-.15, size=20)\n\nax2 = fig.add_subplot(122,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x', 'y', 'z']]\n    t = arrange_track(t)\n    ax2.plot3D(t.z, t.x, t.y, '.-', ms=10)\n    \nax2.set_xlabel('z (mm)')\nax2.set_ylabel('x (mm)')\nax2.set_zlabel('y (mm)')\nax2.set_title('True Tracks', y=-.15, size=20)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"0d275508-9801-4c24-a1d2-892c5465b8e1","_uuid":"9ad64eb4b2faa5d4d98aa8c314682155223e70c8"},"cell_type":"markdown","source":"## DBSCAN Clustering Setting 4\n## scaling and normalization + dbscan (no coordinate transformation)"},{"metadata":{"_cell_guid":"30c64317-d408-47ba-8cf6-6e8d904cf1c5","_kg_hide-input":true,"_uuid":"b6b16bd427ba22d99a457ccd966fead631277102","collapsed":true,"trusted":true},"cell_type":"code","source":"X = hits[['x', 'y', 'z']]\nscaler = MaxAbsScaler().fit(X)\nX = scaler.transform(X)\nnormalizer = Normalizer(norm='l2').fit(X)\nX = normalizer.transform(X)\n\neps = 0.0022\nmin_samp = 3\ndb = DBSCAN(eps=eps, min_samples=min_samp, metric='euclidean').fit(X)\nlabels = db.labels_\n\nclustering = pd.DataFrame()\nclustering['hit_id'] = truth['hit_id']\nclustering['track_id'] = labels\n\nscore = score_event(truth, clustering)\nprint('track-ml custom metric score:', round(score, 4))\n\nlabels_true = truth['particle_id']\nn_clusters = len(set(labels)) - (1 if -1 in labels else 0)\n\nprint('\\nOTHER CLUSTERING RESULTS:')\nprint('Estimated number of clusters: %d' % n_clusters)\nprint(\"Homogeneity: %0.3f\" % metrics.homogeneity_score(labels_true, labels))\nprint(\"Completeness: %0.3f\" % metrics.completeness_score(labels_true, labels))\nprint(\"Adjusted Rand Index: %0.3f\" % metrics.adjusted_rand_score(labels_true, labels))\nrej_perc = list(labels).count(-1) / float(hits.shape[0]) * 100\nrej_perc = round(rej_perc, 2)\nprint (\"Rejected samples %:\", str(rej_perc) + '%')\nrejected_count = list(labels).count(-1)\nprint (\"Rejected samples:\", rejected_count)\nprint (\"Total samples:\", hits.shape[0])\nprint (\"Clustered samples:\", hits.shape[0] - list(labels).count(-1), '\\n')\n\nprint ('WITHOUT REJECTED SAMPLES:')\nlabels_true_wr = labels_true[labels != -1]\nlabels_wr = labels[labels != -1]\nprint(\"Homogeneity: %0.3f\" % metrics.homogeneity_score(labels_true_wr, labels_wr))\nprint((\"Completeness: %0.3f\" % metrics.completeness_score(labels_true_wr, labels_wr)), '\\n')","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"0a459c6a-879d-4cf8-a133-23ca9132975a","_uuid":"2b5bd6b34da0b02860a84a6f1390c6d91c84b61e"},"cell_type":"markdown","source":"## Clustered Hits (Predicted Tracks) vs True Tracks"},{"metadata":{"_cell_guid":"82550664-1b7b-448d-ac8e-9645e1304c5c","_uuid":"792b91afc7f1822949ad625469154ac6c7a54685","collapsed":true,"trusted":true},"cell_type":"code","source":"tracks = truth.particle_id.unique()[1::100]\nfig = plt.figure(figsize=(20,7))\n\nax = fig.add_subplot(121,projection='3d')\n\ntracks_hit_ids = truth[truth['particle_id'].isin(tracks)]['hit_id'] # all hits in tracks\nclusters = clustering[clustering['hit_id'].isin(tracks_hit_ids)].track_id.unique() # all clusters containing the hits in tracks\nfor cluster in clusters:\n    cluster_hit_ids = clustering[clustering['track_id'] == cluster]['hit_id'] # all hits in cluster\n    plot_hit_ids = list(set(tracks_hit_ids) & set(cluster_hit_ids))\n    t = hits[hits['hit_id'].isin(plot_hit_ids)][['x', 'y', 'z']]\n    if cluster == -1:\n        ax.plot3D(t.z, t.x, t.y, '.', ms=10, color='black')\n    else:\n        ax.plot3D(t.z, t.x, t.y, '.-', ms=10)\n    \nax.set_xlabel('z (mm)')\nax.set_ylabel('x (mm)')\nax.set_zlabel('y (mm)')\nax.set_title('Clustered Hits (Predicted Tracks)', y=-.15, size=20)\n\nax2 = fig.add_subplot(122,projection='3d')\nfor track in tracks:\n    hit_ids = truth[truth['particle_id'] == track]['hit_id']\n    t = hits[hits['hit_id'].isin(hit_ids)][['x', 'y', 'z']]\n    t = arrange_track(t)\n    ax2.plot3D(t.z, t.x, t.y, '.-', ms=10)\n    \nax2.set_xlabel('z (mm)')\nax2.set_ylabel('x (mm)')\nax2.set_zlabel('y (mm)')\nax2.set_title('True Tracks', y=-.15, size=20)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"384eca29-ac83-48c6-a5f4-31149c2229a8","_uuid":"15ae47198126a638f59ed25d44cbe6d3586b522b"},"cell_type":"markdown","source":"In this setting, we observe shorter predicted tracks than when usin a coordinate transformation. Also, the predicted tracks are still straight like in Setting 1."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.5","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}