{"cells":[{"metadata":{"_uuid":"e30a74160a2b809824bc2d39a58e13a17fbd36c5"},"cell_type":"markdown","source":"Uniform Manifold Approximation and Projection (UMAP) is a dimension reduction technique like t-SNE . The algorithm is founded on three assumptions about the data\n\n1.  The data is uniformly distributed on a Riemannian manifold;\n2.  The Riemannian metric is locally constant (or can be approximated as such);\n3.  The manifold is locally connected.\n\n*McInnes, L, Healy, J, UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction, ArXiv e-prints 1802.03426, 2018*\n[https://arxiv.org/abs/1802.03426](http://)\n\nThe original version  is writen in Python and there is also a fast R version (uwot) that uses RcppParallel.\nUMAP scales better than t-SNE and allows for new data points addition as well as semi-supervised  embedding.  \n\n"},{"metadata":{"_uuid":"7c3477192a50949ac7bbf8f46421f2f27c666321","_execution_state":"idle","trusted":true},"cell_type":"code","source":"#devtools::install_github(\"jlmelville/uwot\")\nlibrary(uwot)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1d3f57da85931671b039515bbe22e5d31147e802"},"cell_type":"markdown","source":"Individual tracks can be characterized by the vertex position, and the initial particle moment. Let's take a look how this parametres are distributed for the MonteCarlo events.\n\nFirst we read the true info for some events."},{"metadata":{"trusted":true,"_uuid":"7aa59686e07299039e25d1e3b259c1f493c67ba9"},"cell_type":"code","source":"require(data.table)\npath='../input/train_1/'\ndfp=NULL\n  for(i in 0:15){\n    nev=1000+i\n    dfp=rbind(dfp,cbind(i,fread(paste0(path,'event00000',nev,'-particles.csv'))))\n  }","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a7c25fcb6bc0ef2e3d05942ad429ae5c7b7a5e68"},"cell_type":"markdown","source":"And then perform the embedding for the all the particles using, vertex, momentum and charge information"},{"metadata":{"trusted":true,"_uuid":"99d311347873870090eb25d73427e8cba9a7f5f1"},"cell_type":"code","source":"set.seed(123) \nmlt_umap <- umap(dfp[,.(px,py,pz,vx,vy,vz,q)], n_neighbors = 15, min_dist = 0.01, verbose = TRUE, n_threads = 8)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dfc0b9880752bf4e34ad058a7d8c0877f550de91"},"cell_type":"code","source":"plot(mlt_umap,pch='.',col=rgb(0,0,0,.1))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0635c30a62068307d2f85db895f04973c36ef5b9"},"cell_type":"markdown","source":"The embedding shows a rich and convoluted structure. Let's investigate the meanong of these clusters\n\nFirst, how positive and negative charged particles are distributed "},{"metadata":{"trusted":true,"_uuid":"16ca8c80a9e4c51df7f86060446be06e0ce8bf99"},"cell_type":"code","source":"require(squash)\n  \n  squashgram(mlt_umap[,1],mlt_umap[,2],dfp[,q],FUN=mean,nx=1000)\n  ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"45341a979fb2fc92d6336e9c93c80cc8a01e7bde"},"cell_type":"markdown","source":"Some regions are clearly separating particles by its charge, while others seems to mix them.\n\nAfterward we take a look to the transvers and total momentum:"},{"metadata":{"trusted":true,"_uuid":"40ab3f12e2b0bf5b93168400b62e77594a989cb9"},"cell_type":"code","source":"  squashgram(mlt_umap[,1],mlt_umap[,2],dfp[,log(px^2+py^2)],FUN=mean,nx=1000)\n  squashgram(mlt_umap[,1],mlt_umap[,2],dfp[,log(px^2+py^2+pz^2)],FUN=mean,nx=1000,nz=50)\n  \n ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"583d55da5a5254dcbfabc8dfc4b3c6612f93245d"},"cell_type":"markdown","source":"And finally to the vertex position:"},{"metadata":{"trusted":true,"_uuid":"2ab2f9d8c0b273985416e4c88a18fa8ee7d928f4"},"cell_type":"code","source":"squashgram(mlt_umap[,1],mlt_umap[,2],dfp[,log(vx^2+vy^2+vz^2)],FUN=mean,nx=1000,nz=50)\nsquashgram(mlt_umap[,1],mlt_umap[,2],dfp[,vz],FUN=mean,nx=1000,nz=50)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4377777688ad1e5fdd1e296d4751afc36a7ecc5c"},"cell_type":"markdown","source":"Can be this useful?  My intuition is than each point of this map defines a possible vertex and helix ratio ( and only some combinations seems to appear) and then  and optimal way to transform the (x,y,z) track  parameters in order to cluster together hits from the same track. Moroever, this embedding can be a way to learn this optimal  transformatiom from data  ( the embedding is the label) instead of generate handcrafted attributes and scan the potential track parameters from a priory distributions.  What do you think?"}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}