{"cells":[{"metadata":{"_uuid":"0a9c16484375c8c7500d83023226eefa99534f60"},"cell_type":"markdown","source":"This notebook shows how weights are computed.  First, let's import som packages."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"collapsed":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.\n\nfrom matplotlib import pyplot as plt\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d3c0bbf326d6c876ba3dd5fa74dcf1029e2d6abf"},"cell_type":"markdown","source":"Then let's read one event data, and put it in a single data frame."},{"metadata":{"trusted":true,"_uuid":"7d5b956126a5ac03b2808adbc9bc919244d6d4e3","collapsed":true},"cell_type":"code","source":"event_id = 3\nhits = pd.read_csv('../input/train_1/event00000100%d-hits.csv' % event_id)\nparticles = pd.read_csv('../input/train_1/event00000100%d-particles.csv' % event_id)\ntruth = pd.read_csv('../input/train_1/event00000100%d-truth.csv' % event_id)\n\nhits = hits.merge(truth, how='left', on='hit_id')\nhits = hits.merge(particles, how='left', on='particle_id')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"6c7f8158b32feba3d1ef5cece25111a0f64f2ecc"},"cell_type":"markdown","source":"Then let's sort the data by particle id then by distance to the track origin."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"8b2526d6485b9c860adeefdc2a4d02f11dcd2718"},"cell_type":"code","source":"hits['dv'] = np.sqrt((hits.vx - hits.tx) ** 2 + \\\n                     (hits.vy - hits.ty) ** 2 + \\\n                     (hits.vz - hits.tz) ** 2)\nhits = hits.sort_values(by=['particle_id', 'dv']).reset_index(drop=True)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b4aa9cf7c5ee7072848735b5f32dbc0c7b914ce0"},"cell_type":"markdown","source":"We're almost there.  We compute a rank in percentagfe for each hit along a given particle track."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"248458e0078cc7b205ca303e873c3adc31d3ab66"},"cell_type":"code","source":"hits['rank'] = hits.groupby('particle_id').cumcount()\nhits['len'] = hits.groupby('particle_id').particle_id.transform('count')\nhits['rank'] = (hits['rank']) / (hits['len'] - 1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"492427e731dedfeb2930b8a3c68ce53f6e78d63e"},"cell_type":"markdown","source":"   And we normalize weights so that the largest weight on each track is 1."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"72b804a722c6b2a5d5956f90acec1204aaef1987"},"cell_type":"code","source":"hits['weight'] /= hits.groupby('particle_id').weight.transform('max')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8a9a68a0bacc4a9ac64aaa5a2253e22d8b4d474c"},"cell_type":"markdown","source":"We can now plot the normalized weight along each track.  "},{"metadata":{"trusted":true,"_uuid":"88d9cae71938dd2214a7c24395fc2b268493ffae"},"cell_type":"code","source":"fig, ax = plt.subplots(1, 1, figsize=(15, 15))\n\nax.scatter(hits['rank'], hits['weight'], alpha=0.1, marker='+')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"5c7a0d9604eab55047bacbb32f7b9006c2a6fcee"},"cell_type":"markdown","source":"The pattern is quite clear, isn't it? Note that there are surious points away from the main curve.  Aftert looking at them, they belong to particles with strange tracks, probably errors in the simulation."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.5","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}