{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"UPDATE: [exploration of sensors ftrings](#strings)","metadata":{}},{"cell_type":"markdown","source":"CONTENTS\n\n    0  Intro\n    1  Initial\n        1.1  Imports\n        1.2  Constants\n        1.3  Functions\n        1.4  Settings\n    2  Read and Check Data\n        2.1  Read Data\n        2.2  First Look at Data\n    3  Initial Processing\n        3.1  Alternative Coordinate System for Sensors\n        3.2  Minor Improvements\n    4  Auxiliary\n    5  Charge\n    6  Pulses Quantity in an Event\n    7  Pulses Quantity and Charge Average in an Event\n    8  Zenith\n    9  Azimuth\n    10  Time\n        10.1  Time Slots\n        10.2  Event Timing Assumptions\n        10.3  Time and Charge\n        10.4  Charge Sum in an Event for different Time Slices (or Time Windows)\n        10.5  Some summary for time slices separated by aixiliary.\n    11  Sensors\n        11.1  Repeated Triggering of Sensors Within One Event\n        11.2  Quantity of Unique Sensors Within One Event\n    12  Zenith and Charge\n    13  First triggering of sensor in an event\n    14  Event Duration and Charge\n    15  Coordinates\n    16  Strings of sensors\n        16.1  Common configuration of strings with sensors, plane projection on (x, y)\n        16.2  Broken string about x=444 and y=194\n        16.3  Removing broken string data\n        16.4  Quantity of pulses for sensors strings separated by auxiliary\n        16.5  Total charge of pulses for sensors strings separated by auxiliary\n    17  Assumption to check A\n    18  Assumption to check B\n    19  Summary\n        19.1  Общие замечания\n        19.2  Предположение 1\n        19.3  Предположение 2\n        19.4  Предположение 3","metadata":{"toc":true}},{"cell_type":"markdown","source":"# [IceCube](https://icecube.wisc.edu/) – EDA\n\n[Kaggle](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/)\n\n---\n\n**Introductory analysis** \n\nUsing `batch_1`, 1 piece of data out of almost 700 .\n\n---\n\n**Purpose**  \n\n- understand the physical process;\n- understand the essence of the input data and their correspondence to the physical process;\n- identify as more different areas for analysis as possible, then to choose the most useful ones;\n- make an assumption about the principle of model creating  for trajectory determination.","metadata":{}},{"cell_type":"markdown","source":"**Some explanations**\n\n---\n\nPermanent data tables named like: **data**.  \n\nTemporary data tables named like: **df**.  \n\n---\n\nIntermediate conclusions are highlighted as follows:\n\n> Intermediate conclusion.\n\n---\n\nThe code of the cells are as independent as possible from each other in order to freely manipulate the cells.\n\nAt the end of this analysis it will be clear which calculated fields better to be added at the beginning of the analysis.\nThe same goes for functions.","metadata":{}},{"cell_type":"markdown","source":"![Ice Cube](https://res.cloudinary.com/icecube/images/w_600,h_450/q_auto/v1603431620/icecube_detector_schematic/icecube_detector_schematic.jpg)","metadata":{}},{"cell_type":"markdown","source":"The in-ice component of [IceCube](https://icecube.wisc.edu/science/icecube/) consists of 5,160 digital optical modules (DOMs), each with a ten-inch photomultiplier tube and associated electronics. The DOMs are attached to vertical “strings,” frozen into 86 boreholes, and arrayed over a cubic kilometer from 1,450 meters to 2,450 meters depth. The strings are deployed on a hexagonal grid with 125 meters spacing and hold 60 DOMs each. The vertical separation of the DOMs is 17 meters.","metadata":{}},{"cell_type":"markdown","source":"## Intro\n\nIn short. Can be expanded.\n\n**Section in the process of being translated into English.**","metadata":{}},{"cell_type":"markdown","source":"На нейтрино не действуют сильные ядерные взаимодействия и электромагнетизм. Нейтрино участвуют только в слабом взаимодействии. Из-за этого [сечение рассеяния](https://elementy.ru/LHC/HEP/measures/cross-section) нейтрино очень мало и вероятность их детектирования ничтожна.","metadata":{}},{"cell_type":"markdown","source":"Нейтрино ударяет по атому в рабочем объеме детектора, что приводит к большому энерговыделению – рождению многочисленных фотонов и электрон-позитронных пар, свечение от которых регистрируется чувствительными датчиками света (фотоумножителями, ФЭУ). Чем ярче эта вспышка, тем больше ее пространственный размер. Для энергий нейтрино в сотни ТэВ этот размер составляет сотни метров. Поэтому чтобы отличить нейтринные события с энергией 1 ПэВ и 1 ТэВ — то есть четко различать астрофизические и атмосферные нейтрино, — требуется установка размером порядка километра.\n\nБольшой размер детектора нужен также для того, чтобы повысить вероятность их регистрации. Если теоретические модели процессов рождения астрофизических нейтрино верны, при таких масштабах можно рассчитывать на поимку как минимум нескольких астрофических нейтрино в год.","metadata":{}},{"cell_type":"markdown","source":"Чувствительные фотоэлементы в IceCube организованы в виде гирлянд, которые спускались в шахты на глубину до двух с половиной километров, где они вмерзали в окружающий лед. IceCube состоит из 86 таких цепочек по 60 ФЭУ в каждой, а сами цепочки расположены на расстоянии 125 метров друг от друга. Кроме этого, в центре детектора есть небольшая область с более плотной «застройкой» (установка Deep Core), а сверху над IceCube расположен датчик космических лучей IceTop. Этот датчик позволяет отличить одиночное космическое нейтрино от нейтрино, рожденного в атмосфере: в первом случае мощный нейтринный сигнал внутри IceCube происходит в одиночестве, а во втором случае IceTop зафиксирует совпадающий по времени широкий ливень вторичных частиц.","metadata":{}},{"cell_type":"markdown","source":"На верхние слои атмосферы Земли попадает непрерывный поток космических лучей — релятивистских протонов и ядер с зарегистрированными энергиями от ГэВ до ~10$^{20}$ эВ. При высоких энергиях взаимодействие космической частицы с ядром атома из атмосферы приводит к множественному рождению вторичных частиц, прежде всего π-мезонов. Заряженные π-мезоны распадаются, рождая, в том числе, высокоэнергичные нейтрино. К рождению нейтрино приводят также и распады более тяжёлых вторичных частиц. Вся совокупность рождающихся в этих процессах нейтрино называется атмосферными нейтрино.  \nПотоки космических лучей достаточно низких (ГэВ–ТэВ) энергий очень велики, а источник нейтрино — земная атмосфера — находится в непосредственной близости, поэтому любые астрофизические нейтринные сигналы в этом диапазоне энергий оказываются задавленными атмосферным фоном.","metadata":{}},{"cell_type":"markdown","source":"Детектор IceCube позволяет измерить долю мюонных нейтрино. Это важный показатель для выделения астрофизических нейтрино.  \nАтмосферные нейтрино можно отличать по сопровождающему ливню заряженных частиц. Если одновременно с высокоэнергетическим нейтрино в детекторе есть также сильный сигнал от IceTop, то такое нейтрино скорее всего атмосферное.  \nНо IceTop может и не зафиксировать особого сигнала от вторичных частиц, если первичное столкновение произошло в атмосфере близко к линии горизонта. Тогда высокоэнергетическое нейтрино будет сопровождаться *небольшим* сигналом в IceTop, и выделить его будет трудно, поскольку космические лучи падают постоянно.","metadata":{}},{"cell_type":"markdown","source":"Если событие вызвано мюонным нейтрино, которое, столкнувшись с атомом вещества, превратилось в мюон, то этот мюон большой энергии оставляет четко различимый след — мюонный трек. Направление этого трека позволяет определить направление прилета нейтрино с точностью порядка одного градуса. Если же это было электронное или тау-нейтрино без мюонного трека, тогда направление прилета восстанавливается плохо, с погрешностью около 15 градусов.","metadata":{}},{"cell_type":"markdown","source":"Обычно, подчеркивая огромную проникающую способность нейтрино, говорят, что они могут беспрепятственно пролетать сквозь Землю, однако это верно лишь для нейтрино умеренно больших энергий. При энергии в сотни ТэВ нейтрино уже начинает довольно хорошо поглощаться Землей, поэтому летящие с северного неба нейтрино имеют меньше шансов долететь до детектора.","metadata":{}},{"cell_type":"markdown","source":"Природный лед содержит неоднородности, на которых излучение рассеивается. Это влияет на точность определения тректории нейтрино.","metadata":{}},{"cell_type":"markdown","source":"Два основных типа событий, наблюдаемых в экспериментах, — это «треки» и «каскады».  \nПри взаимодействии мюонного нейтрино с ядром мишени рождается мюон, который является относительно долгоживущей частицей: он пролетает через значительную часть объема детектора до распада, оставляя узкий линейный след — «трек» черенковского излучения.  \nТакие же мюоны рождаются при взаимодействии космических лучей в атмосфере, и, несмотря на достаточно большую глубину, часть таких событий достигает детектора.  \nТрековые события не сопровождают взаимодействия $υ_e$ и $υ_τ$, поскольку рождающиеся в них электроны быстро перерассеиваются, а $τ$-лептоны практически сразу распадаются.\n\nВторой тип событий — «каскады» — связан с лавинообразным развитием многочастичных процессов, вызванных неупругим взаимодействием нейтрино с ядром. При высоких энергиях рождается сразу много частиц, каждая из которых также взаимодействует с окружающим веществом, и в результате формируется лавина релятивистских частиц, распространяющихся от точки взаимодействия. Все эти частицы также излучают черенковские фотоны. Такое событие регистрируется детектором, однако выглядит совершенно иначе, чем мюонный трек: каскадное событие больше напоминает облако, нежели прямую линию. Точность определения направления оказывается хуже, чем у мюонного трека, зато каскады, начавшиеся и развивающиеся в детекторе, позволяют лучше измерить энергию исходного нейтрино.\n\nКаскадные события, вызванные тау-нейтрино наиболее высоких энергий (≳2 ПэВ), могут выглядеть иначе, и их иногда выделяют в отдельный, третий тип событий — «двойной всплеск». Энергичный тау-лептон распадается, лишь отлетев на некоторое расстояние от точки нейтринного взаимодействия, поэтому должны наблюдаться два каскада (от взаимодействия нейтрино с ядром и от распада тау-лептона), соединенные коротким треком (черенковское излучение самого тау-лептона). Экспериментально такие события пока надежно не зафиксированы.","metadata":{}},{"cell_type":"markdown","source":"Для атмосферного фона теоретически известно определенное распределение по зенитным углам. Но какое именно, пока не нашел.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Initial","metadata":{"tags":[]}},{"cell_type":"markdown","source":"### Imports","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport pyarrow as pa\nfrom pyarrow.parquet import ParquetFile\n\nimport os\nimport warnings\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom matplotlib.colors import LinearSegmentedColormap\nimport plotly.graph_objects as go\nimport plotly.io as pio","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:31.164198Z","iopub.execute_input":"2023-02-19T17:47:31.165520Z","iopub.status.idle":"2023-02-19T17:47:31.171865Z","shell.execute_reply.started":"2023-02-19T17:47:31.165433Z","shell.execute_reply":"2023-02-19T17:47:31.170689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Constants","metadata":{}},{"cell_type":"code","source":"PATH_LOCAL = 'datasets/'                                              # local path to data\nPATH_REMOTE = '/kaggle/input/icecube-neutrinos-in-deep-ice/'          # remote path to data\n\nCR = '\\n'                                                             # new line\nRANDOM_STATE = RS = 88                                                # random_state","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:31.190062Z","iopub.execute_input":"2023-02-19T17:47:31.190718Z","iopub.status.idle":"2023-02-19T17:47:31.195723Z","shell.execute_reply.started":"2023-02-19T17:47:31.190681Z","shell.execute_reply":"2023-02-19T17:47:31.194553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Functions","metadata":{}},{"cell_type":"code","source":"def custom_read_csv(file_name, separator=','):\n    \"\"\"\n    чтение датасета в формате CSV:\n      сначала из локального хранилища;\n      при неудаче — из удаленного хранилища Kaggle.\n    \"\"\"\n\n    path_local = f'{PATH_LOCAL}{file_name}'\n    path_remote = f'{PATH_REMOTE}{file_name}'\n    \n    if os.path.exists(path_local):\n        return pd.read_csv(path_local, sep=separator)\n\n    elif os.path.exists(path_remote):\n        return pd.read_csv(path_remote, sep=separator)\n\n    else:\n        print(f'File \"{file_name}\" not found at the specified path ')","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:31.198710Z","iopub.execute_input":"2023-02-19T17:47:31.199193Z","iopub.status.idle":"2023-02-19T17:47:31.209113Z","shell.execute_reply.started":"2023-02-19T17:47:31.199138Z","shell.execute_reply":"2023-02-19T17:47:31.207491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def custom_read_parquet(file_name):\n    \"\"\"\n    чтение датасета в формате Parquet:\n      сначала из локального хранилища;\n      при неудаче — из удаленного хранилища Kaggle.\n    \"\"\"\n\n    path_local = f'{PATH_LOCAL}{file_name}'\n    path_remote = f'{PATH_REMOTE}{file_name}'\n    \n    if os.path.exists(path_local):\n        return pa.parquet.read_table(path_local).to_pandas()\n\n    elif os.path.exists(path_remote):\n        return pa.parquet.read_table(path_remote).to_pandas()\n\n    else:\n        print(f'File \"{file_name}\" not found at the specified path ')","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:31.210886Z","iopub.execute_input":"2023-02-19T17:47:31.212201Z","iopub.status.idle":"2023-02-19T17:47:31.222867Z","shell.execute_reply.started":"2023-02-19T17:47:31.212154Z","shell.execute_reply":"2023-02-19T17:47:31.221353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def custom_read_parquet_group(file_name, group_num=0, columns_list=[0]):\n    \"\"\"\n    чтение определенной группы из датасета в формате Parquet:\n      сначала из локального хранилища;\n      при неудаче — из удаленного хранилища Kaggle.\n      group_num: номер группы в файле\n      columns_list: список полей для чтения\n    \"\"\"\n\n    path_local = f'{PATH_LOCAL}{file_name}'\n    path_remote = f'{PATH_REMOTE}{file_name}'\n    \n    if os.path.exists(path_local):\n        return (ParquetFile(path_local)\n                 .read_row_group(i=group_num, columns=columns_list)\n                 .to_pandas()\n                )\n    elif os.path.exists(path_remote):\n        return (ParquetFile(path_remote)\n                 .read_row_group(i=group_num, columns=columns_list)\n                 .to_pandas()\n                )\n    else:\n        print(f'File \"{file_name}\" not found at the specified path ')","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:31.224982Z","iopub.execute_input":"2023-02-19T17:47:31.225935Z","iopub.status.idle":"2023-02-19T17:47:31.235733Z","shell.execute_reply.started":"2023-02-19T17:47:31.225874Z","shell.execute_reply":"2023-02-19T17:47:31.234712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def df_name(df):\n    \"\"\"\n    table name determination\n    \"\"\"\n    return [name for name in globals() if globals()[name] is df][0]","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:31.238908Z","iopub.execute_input":"2023-02-19T17:47:31.239757Z","iopub.status.idle":"2023-02-19T17:47:31.251029Z","shell.execute_reply.started":"2023-02-19T17:47:31.239706Z","shell.execute_reply":"2023-02-19T17:47:31.249913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def basic_info(df: pd.DataFrame, sample_type='sample', samples=5, describe='all'):\n    \"\"\"\n    first info about dataframe: info(), sample()/head()/tail(), describe()\n    \"\"\"\n    \n    # title (name of dataframe)\n    \n    print(f'\\n\\ndataframe {f.BOLD}{df_name(df)}{f.END}', '≋'*30)\n\n\n    # method info()\n    \n    print('\\n\\n--- method info() ---\\n')\n    print(df.info())\n\n    \n    # several random records\n    \n    print(f'\\n\\n--- method {sample_type}({samples}) ---')\n    \n    if sample_type == 'sample':\n        display(df.sample(samples))\n    elif sample_type == 'head':\n        display(df.head(samples))\n    elif sample_type == 'tail':\n        display(df.tail(samples))\n    else:\n        print(f'{sample_type} – invalid value for parameter \"sample_type\" ')\n    \n    \n    # method describe()\n    \n    print(f'\\n\\n--- method describe({describe}) ---')\n    \n    if describe=='all' or describe=='numeric':\n        try:\n            display(df.describe(include=np.number))\n        except ValueError:\n            pass\n\n    if describe=='all' or describe=='categorical':\n        try:\n            display(df.describe(exclude=np.number).T)\n        except ValueError:\n            pass\n    \n    if describe not in ['numeric','categorical','all']:\n        print(f'{describe} – invalid value for parameter \"describe\" ')","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:31.252822Z","iopub.execute_input":"2023-02-19T17:47:31.253625Z","iopub.status.idle":"2023-02-19T17:47:31.266852Z","shell.execute_reply.started":"2023-02-19T17:47:31.253577Z","shell.execute_reply":"2023-02-19T17:47:31.265458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Settings","metadata":{}},{"cell_type":"code","source":"# text styles\nclass f:\n    BOLD = \"\\033[1m\"\n    ITALIC = \"\\033[3m\"\n    END = \"\\033[0m\"","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:31.269079Z","iopub.execute_input":"2023-02-19T17:47:31.270597Z","iopub.status.idle":"2023-02-19T17:47:31.283565Z","shell.execute_reply.started":"2023-02-19T17:47:31.270551Z","shell.execute_reply":"2023-02-19T17:47:31.282127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# defaults for charts\n\n# Matplotlib, Seaborn\nsns.set_style('whitegrid', {'axes.facecolor': '0.98', 'grid.color': '0.9', 'axes.edgecolor': '1.0'})\nPLOT_DPI = 150  # dpi for charts rendering \nplt.rc(\n       'axes',\n       labelweight='bold',\n       titlesize=16,\n       titlepad=10,\n      )\n\n# Plotly Graph_Objects\npio.templates['my_theme'] = go.layout.Template(\n                                               layout_autosize=True,\n                                               # width=900,\n                                               layout_height=200,\n                                               layout_legend_orientation=\"h\",\n                                               layout_margin=dict(t=40, b=40),         # (l=0, r=0, b=0, t=0, pad=0)\n                                               layout_template='seaborn',\n                                              )\npio.templates.default = 'my_theme'\n\n# colors, color schemes\nCMAP_SYMMETRIC = LinearSegmentedColormap.from_list('', ['steelblue', 'aliceblue', 'steelblue'])","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:31.416747Z","iopub.execute_input":"2023-02-19T17:47:31.417501Z","iopub.status.idle":"2023-02-19T17:47:31.524618Z","shell.execute_reply.started":"2023-02-19T17:47:31.417425Z","shell.execute_reply":"2023-02-19T17:47:31.522757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Pandas defaults\npd.options.display.max_colwidth = 100\npd.options.display.max_rows = 500\npd.options.display.max_columns = 100\npd.options.display.float_format = '{:.3f}'.format\n# pd.options.display.colheader_justify = 'left'","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:31.527060Z","iopub.execute_input":"2023-02-19T17:47:31.527540Z","iopub.status.idle":"2023-02-19T17:47:31.535103Z","shell.execute_reply.started":"2023-02-19T17:47:31.527494Z","shell.execute_reply":"2023-02-19T17:47:31.533011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# others\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:31.537068Z","iopub.execute_input":"2023-02-19T17:47:31.537512Z","iopub.status.idle":"2023-02-19T17:47:31.548298Z","shell.execute_reply.started":"2023-02-19T17:47:31.537451Z","shell.execute_reply":"2023-02-19T17:47:31.546860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Read and Check Data","metadata":{}},{"cell_type":"markdown","source":"### Read Data","metadata":{}},{"cell_type":"code","source":"# pulses data (part 1)\ndata = custom_read_parquet('train/batch_1.parquet')","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:31.550296Z","iopub.execute_input":"2023-02-19T17:47:31.550948Z","iopub.status.idle":"2023-02-19T17:47:33.330853Z","shell.execute_reply.started":"2023-02-19T17:47:31.550902Z","shell.execute_reply":"2023-02-19T17:47:33.329845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# sensors data\ndata_sg = custom_read_csv('sensor_geometry.csv')","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:33.333412Z","iopub.execute_input":"2023-02-19T17:47:33.333968Z","iopub.status.idle":"2023-02-19T17:47:33.343303Z","shell.execute_reply.started":"2023-02-19T17:47:33.333935Z","shell.execute_reply":"2023-02-19T17:47:33.342463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# events meta-data (including trajectory data)\n\nGROUP = 0                                                  # group number in 'batch_1.parquet' (file includes several batches)\nBATCH = 1                                                  # batch number\nCOLUMNS = ['batch_id','event_id','azimuth','zenith']       # fields list to read\n\ndata_meta = custom_read_parquet_group('train_meta.parquet', group_num=GROUP, columns_list=COLUMNS)\ndata_meta = data_meta.query('batch_id == @BATCH')          # one batch of events is enough for first analysis\ndata_meta = data_meta.drop('batch_id', axis=1)             # in this exploration field 'batch_id' is not required","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:33.344940Z","iopub.execute_input":"2023-02-19T17:47:33.345526Z","iopub.status.idle":"2023-02-19T17:47:39.387217Z","shell.execute_reply.started":"2023-02-19T17:47:33.345483Z","shell.execute_reply":"2023-02-19T17:47:39.386193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### First Look at Data","metadata":{}},{"cell_type":"code","source":"basic_info(data)\nbasic_info(data_sg)\nbasic_info(data_meta)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:39.388810Z","iopub.execute_input":"2023-02-19T17:47:39.389546Z","iopub.status.idle":"2023-02-19T17:47:45.434835Z","shell.execute_reply.started":"2023-02-19T17:47:39.389501Z","shell.execute_reply":"2023-02-19T17:47:45.433463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nConverting the index to a regular field `event_id`.","metadata":{}},{"cell_type":"code","source":"data = data.reset_index()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:45.436439Z","iopub.execute_input":"2023-02-19T17:47:45.436947Z","iopub.status.idle":"2023-02-19T17:47:45.769990Z","shell.execute_reply.started":"2023-02-19T17:47:45.436902Z","shell.execute_reply":"2023-02-19T17:47:45.768619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> There are no NaNs.\n>\n> The `event_id` field is represented as an index. Can be converted to a regular field.","metadata":{}},{"cell_type":"markdown","source":"---\nDuplicates in data.","metadata":{}},{"cell_type":"code","source":"data.duplicated().sum(), data_sg.duplicated().sum(), data_meta.duplicated().sum()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:47:45.771318Z","iopub.execute_input":"2023-02-19T17:47:45.771676Z","iopub.status.idle":"2023-02-19T17:48:00.036708Z","shell.execute_reply.started":"2023-02-19T17:47:45.771640Z","shell.execute_reply":"2023-02-19T17:48:00.035367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data[data.duplicated()].sort_values('charge', ascending=False).head()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:48:00.038359Z","iopub.execute_input":"2023-02-19T17:48:00.038755Z","iopub.status.idle":"2023-02-19T17:48:12.665753Z","shell.execute_reply.started":"2023-02-19T17:48:00.038721Z","shell.execute_reply":"2023-02-19T17:48:12.664666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Removing full duplicates.","metadata":{}},{"cell_type":"code","source":"data = data.drop_duplicates()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:48:12.667049Z","iopub.execute_input":"2023-02-19T17:48:12.667931Z","iopub.status.idle":"2023-02-19T17:48:26.083941Z","shell.execute_reply.started":"2023-02-19T17:48:12.667883Z","shell.execute_reply":"2023-02-19T17:48:26.082645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> There are a small number of full duplicates in the data.  \n> It would be safe to remove full duplicates, but there are pulses with a large charge among them.\n\n> Now we'll remove full duplicates, but there is a risk of important data losing (with a large charge).","metadata":{}},{"cell_type":"markdown","source":"---\nA deeper look at duplicates.\n\nIt is assumed the sensor can't detect more than one flash at a time (or sensor can detect only one flash in one event at the same time).  \nLet's check it: count the number of duplicates by a limited list of fields.","metadata":{}},{"cell_type":"code","source":"df = data[['event_id','sensor_id','time']]\n\ndf.duplicated().sum()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:48:26.085691Z","iopub.execute_input":"2023-02-19T17:48:26.086184Z","iopub.status.idle":"2023-02-19T17:48:37.585761Z","shell.execute_reply.started":"2023-02-19T17:48:26.086138Z","shell.execute_reply":"2023-02-19T17:48:37.584507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What does duplicates look like?","metadata":{}},{"cell_type":"code","source":"df = df[df.duplicated(keep=False)].merge(data, on=['event_id','sensor_id','time']).drop_duplicates()\ndf.head(6)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:48:37.587218Z","iopub.execute_input":"2023-02-19T17:48:37.587603Z","iopub.status.idle":"2023-02-19T17:49:06.747457Z","shell.execute_reply.started":"2023-02-19T17:48:37.587569Z","shell.execute_reply":"2023-02-19T17:49:06.746250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Duplicates are changed due to charge or auxiliary?","metadata":{}},{"cell_type":"code","source":"df.auxiliary.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:49:06.752247Z","iopub.execute_input":"2023-02-19T17:49:06.752622Z","iopub.status.idle":"2023-02-19T17:49:06.762309Z","shell.execute_reply.started":"2023-02-19T17:49:06.752591Z","shell.execute_reply":"2023-02-19T17:49:06.760938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How many times can records be duplicated: 2, 3 or more?","metadata":{}},{"cell_type":"code","source":"df.groupby(['event_id','sensor_id','time']).size().value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:49:06.764142Z","iopub.execute_input":"2023-02-19T17:49:06.764733Z","iopub.status.idle":"2023-02-19T17:49:06.792678Z","shell.execute_reply.started":"2023-02-19T17:49:06.764670Z","shell.execute_reply":"2023-02-19T17:49:06.791590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The number of duplicates in group and total charge in group.","metadata":{}},{"cell_type":"code","source":"df['n_diplicate'] = df.groupby(['event_id','sensor_id','time']).charge.transform(lambda x: x.count())\ndf['charge_sum'] = df.groupby(['event_id','sensor_id','time']).charge.transform(lambda x: x.sum())","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:49:06.794074Z","iopub.execute_input":"2023-02-19T17:49:06.794397Z","iopub.status.idle":"2023-02-19T17:49:14.659440Z","shell.execute_reply.started":"2023-02-19T17:49:06.794368Z","shell.execute_reply":"2023-02-19T17:49:14.658435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.sort_values('charge_sum', ascending=False).head(8)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:49:14.661031Z","iopub.execute_input":"2023-02-19T17:49:14.661345Z","iopub.status.idle":"2023-02-19T17:49:14.682948Z","shell.execute_reply.started":"2023-02-19T17:49:14.661317Z","shell.execute_reply":"2023-02-19T17:49:14.681790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Total 22010 duplicates: pulses recorded during the same event at the same time on the same sensor. This is less than 0.07% of all pulses.\n>\n> All duplicates have auxiliary=False, charge differs.\n>\n> Very very rare triple duplicates.\n\n> `Assumption: this is a technical error. It may be necessary to merge the duplicates by summing charge.`\n\n> So far no action has been taken with such duplicates.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Initial Processing","metadata":{}},{"cell_type":"markdown","source":"### Alternative Coordinate System for Sensors","metadata":{}},{"cell_type":"markdown","source":"The alternative coordinate system will simplify some calculations.\n\nThe origin is at the top near corner of the cube. In old coordinates, this point is described as (max(x), min(y), max(z)).  \nAll alternative coordinates are `>= 0`.\n\nX-axis: left-far direction along the edge of the cube from the origin.  \nY-axis: right-far direction along the edge of the cube from the origin.  \nZ-axis: origin at the top, downward direction, z-value is the distance from the top.","metadata":{}},{"cell_type":"code","source":"data_sg['X'] = (data_sg.x - data_sg.x.max()).abs()\ndata_sg['Y'] = (data_sg.y - data_sg.y.min())\ndata_sg['Z'] = (data_sg.z - data_sg.z.max()).abs()\n\ndata_sg.sample(3)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:49:14.684445Z","iopub.execute_input":"2023-02-19T17:49:14.685131Z","iopub.status.idle":"2023-02-19T17:49:14.708135Z","shell.execute_reply.started":"2023-02-19T17:49:14.685087Z","shell.execute_reply":"2023-02-19T17:49:14.707248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Minor Improvements","metadata":{}},{"cell_type":"markdown","source":"Attaching trajectory data (azimuth, zenith).\n\nAttaching sensors coordinates.","metadata":{}},{"cell_type":"code","source":"data = (data\n        .merge(data_meta, on='event_id')\n        .merge(data_sg, on='sensor_id')\n       )\n\ndata.sample()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:49:14.709675Z","iopub.execute_input":"2023-02-19T17:49:14.710817Z","iopub.status.idle":"2023-02-19T17:49:29.999580Z","shell.execute_reply.started":"2023-02-19T17:49:14.710767Z","shell.execute_reply":"2023-02-19T17:49:29.998404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## **Auxiliary**\n\nIf `True` the pulse was not fully digitized, is of lower quality.  \nIf `False` then this pulse was contributed to the trigger decision and the pulse was fully digitized.","metadata":{}},{"cell_type":"markdown","source":"---\nThe quantity of events in the explorated batch.","metadata":{}},{"cell_type":"code","source":"data.event_id.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T17:49:30.001015Z","iopub.execute_input":"2023-02-19T17:49:30.001495Z","iopub.status.idle":"2023-02-19T17:49:30.485443Z","shell.execute_reply.started":"2023-02-19T17:49:30.001433Z","shell.execute_reply":"2023-02-19T17:49:30.484139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nAre there events where all `auxiliary` values are True or False only?","metadata":{}},{"cell_type":"code","source":"data.groupby('event_id').auxiliary.transform(lambda x: x.nunique()).value_counts()","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2023-02-19T17:49:30.487275Z","iopub.execute_input":"2023-02-19T17:49:30.487644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> There is no event where `auxiliary` is exclusively True or False.","metadata":{}},{"cell_type":"markdown","source":"---\nFraction of confirmed pulses (without grouping into events), `auxiliary=False`.","metadata":{}},{"cell_type":"code","source":"data[data.auxiliary == False].shape[0] / data.shape[0]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nFraction of confirmed pulses (grouped by events), `auxiliary=False`.","metadata":{}},{"cell_type":"code","source":"df = (data.groupby(['event_id','auxiliary'])\n          .auxiliary\n          .count()\n          .unstack()\n          .assign(auxiliary_False_fraction = lambda row: row[0]/(row[0]+row[1]))\n     )\n\ndf.sample()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_bins = 50\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\nsns.histplot(x=df.auxiliary_False_fraction, bins=n_bins)\nax.set_title(f'{CR}Fraction of confirmed pulses grouped into events{CR}auxiliary=False')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> The global fraction of confirmed pulses is about 72%.  \n> However when grouped by events the peak of confirmed pulses is about 25%.\n>\n> `Assumption: most of the confirmed impulses are concentrated in a small number of events.`","metadata":{}},{"cell_type":"markdown","source":"---\nDistributions comparison of confirmed and unconfirmed pulses by events.","metadata":{}},{"cell_type":"code","source":"df = (data.groupby(['event_id','auxiliary'])\n          .auxiliary\n          .count()\n          .unstack()\n          .assign(ratio_False_True = lambda row: row[0]/row[1])   # ratio of confirmed pulses to unconfirmed\n          .assign(ratio_True_False = lambda row: row[1]/row[0])   # and inverse ratio\n     )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nMaximum `False/True` and `True/False` ratio values for confirmed and unconfirmed pulses by events.","metadata":{}},{"cell_type":"code","source":"display(df.sort_values('ratio_False_True', ascending=False).head(3),\n        df.sort_values('ratio_True_False', ascending=False).head(3))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n`auxiliary` ratios by events.","metadata":{}},{"cell_type":"code","source":"n_bins = 200\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI, nrows=1, ncols=2)\nsns.histplot(x=df.ratio_False_True, bins=n_bins, ax=ax[0])\nsns.histplot(x=df.ratio_True_False, bins=n_bins, ax=ax[1])\nax[0].set_title(f'{CR}auxiliary: ratio False/True')\nax[1].set_title(f'{CR}auxiliary: ratio True/False')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Confirmed pulses are concentrated in a relatively small quantity of events.\n>\n> False/True ratio can reach much higher values  (about 50 times) than True/False ratio. This explains the fact: with the general predominance of confirmed pulses in most events the proportion of unconfirmed pulses is higher.\n>\n> `Assumption: unconfirmed pulses are noise or have a different origin.`\n>\n> `Assumption: most events are either false (caused by noise) or contain other errors (for example, an incorrectly determined trajectory due to strong noise).`","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## **Charge**","metadata":{}},{"cell_type":"markdown","source":"---\nIs **charge** a continuous value?","metadata":{}},{"cell_type":"code","source":"data.charge.value_counts().sort_index().head(3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.charge.value_counts().sample(10)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **Charge** is discrete with a step of 0.050 starting from 0.025 photoelectrons.","metadata":{}},{"cell_type":"markdown","source":"---\nStatistical describe of pulse **charge** separated by **auxiliary**.","metadata":{}},{"cell_type":"code","source":"data.groupby('auxiliary').charge.describe()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Except for the minimum value the statistics for confirmed and unconfirmed signals differ significantly.","metadata":{}},{"cell_type":"markdown","source":"---\nCharge of pulse across all events.","metadata":{}},{"cell_type":"code","source":"n_bins = 10000\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\nsns.histplot(x=data.charge, bins=n_bins)\nax.set_xlim(0,10)\nax.set_title(f'{CR}Charge of pulse across all events')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Most of the pulses have a charge about 1 photoelectron.","metadata":{}},{"cell_type":"markdown","source":"---\nCharge of pulse across all events separated by auxiliary.","metadata":{}},{"cell_type":"code","source":"n_bins = 10000\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\nsns.histplot(x=data.charge, hue=data.auxiliary, bins=n_bins)\nax.set_xlim(0,10)\nax.set_title(f'{CR}Pulse charge across all events{CR}separated by auxiliary')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> The higher the pulse charge, the higher the proportion of confirmed pulses.","metadata":{}},{"cell_type":"markdown","source":"---\nLogarithm of pulse charge.","metadata":{}},{"cell_type":"code","source":"n_bins = 50\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\nsns.histplot(x=np.log(data.charge), hue=data.auxiliary, bins=n_bins)\nax.set_title(f'{CR}Logarithm of pulse charge across all events separated by auxiliary,{CR}{n_bins} bins')\nax.set_xlabel('log(charge)')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nLogarithm of pulse charge separated by auxiliary.","metadata":{}},{"cell_type":"code","source":"n_bins = 500\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\nsns.histplot(x=np.log(data.charge), hue=data.auxiliary, bins=n_bins)\nax.set_title(f'{CR}Logarithm of pulse charge across all events separated by auxiliary{CR}{n_bins} bins')\nax.set_xlabel('log(charge)')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Charge is the probability of photoelectron detection.\n>\n> `Assumption: the discreteness of the probability of detecting one photoelectron is caused by technical features. Perhaps further study of these features will be useful.`\n\n> The secondary combs on the right side of the graph are a technical error: due to the discreteness of **charge** histogram overlaps during the bins formation. This can be easily verified by changing the number of bins: the number and location of the combs will also change.","metadata":{}},{"cell_type":"markdown","source":"> Следуя описанию, `charge` – это вероятность того, что фотоэлектрон(ы) зафиксирован.\n> \n> `Предположение: дискретность вероятности фиксации одного фотоэлектрона вызвана техническими особенностями. Возможно, дополнительное изучение этих особенностей будет полезно.`\n \n> Вторичные гребни в правой части графика – техническая погрешность: в силу дискретности **charge** происходит наложение при формировании корзин гистограммы. В этом легко убедиться, меняя количество корзин: количество и местоположение гребней также будет меняться.","metadata":{}},{"cell_type":"markdown","source":"---\nTotal charge in event.","metadata":{}},{"cell_type":"code","source":"df = data.groupby(['event_id','auxiliary']).agg(charge_sum = ('charge','sum')).reset_index()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_bins = 100\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\nsns.histplot(x=np.log(df.charge_sum), hue=df.auxiliary, bins=n_bins)\nax.set_title(f'{CR}Logarithm of total charge of pulses in event{CR}separated by auxiliary')\nax.set_xlabel('log(charge_sum)')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> The distributions for the logarithm of the total **charge** are significantly different for confirmed and unconfirmed pulses.\n> \n> At a certain interval a noise signal (or a signal of a different origin) predominates.\n\n> Need to add:  \n> - make similar chart for the first triggering of the sensor;\n> - **charge** depending on **time**.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Pulses Quantity in an Event","metadata":{}},{"cell_type":"code","source":"df = (data.groupby(['event_id','auxiliary'])\n          .size()\n          .unstack()\n          .rename(columns={0:'auxiliary_False', 1:'auxiliary_True'})\n          .assign(total = lambda row: row.auxiliary_False + row.auxiliary_True)\n          .reset_index(drop=True)\n     )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.total.describe()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> In some events the quantity of pulses can exceed the median value by hundreds of times.","metadata":{}},{"cell_type":"code","source":"n_bins = 20000\n\nfig, ax = plt.subplots(figsize=(15,8), dpi=PLOT_DPI, nrows=3, ncols=1, sharex=True, sharey=True)\nfig.subplots_adjust(hspace=0.4)\n\nsns.histplot(x=df.auxiliary_False, bins=n_bins, ax=ax[0])\nax[0].set_title(f'{CR}Pulses quantity in event (auxiliary=False)')\n\nsns.histplot(x=df.auxiliary_True, bins=n_bins//200, ax=ax[1], color='darkorange')\nax[1].set_title(f'{CR}Pulses quantity in event (auxiliary=True)')\n\nsns.histplot(x=df.total, bins=n_bins, ax=ax[2], color='teal')\nax[2].set_title(f'{CR}Pulses quantity in event (total)')\nax[2].set_xlim(0,200)\n\nplt.xlabel('pulses quantity')\n\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Confirmed and unconfirmed pulses have significantly different distributions.\n\n> Unconfirmed pulses and all impulses have similar distributions but proportion of unconfirmed impulses is only about 28%.","metadata":{}},{"cell_type":"markdown","source":"---\nSeveral charts about the same thing in order to better understand how pulses quantity is distributed by events.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,3), dpi=PLOT_DPI)\n\nsns.histplot(data=df.melt(), x='value', y='auxiliary')\nax.set_title(f'{CR}Pulses quantity in event separated by auxiliary')\nax.set_xlabel('pulses quantity')\n\nplt.show()","metadata":{"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> The greater the pulses quantity in an event, the more it is determined by the confirmed pulses quantity.  \n> In other words the proportion of confirmed pulses is growing.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,2), dpi=PLOT_DPI)\n\nsns.boxplot(data=df.melt(), x='value', y='auxiliary', fliersize=2)\nax.set_title(f'{CR}Pulses quantity in event separated by auxiliary')\nax.set_xlabel('pulses quantity')\n\nplt.show()","metadata":{"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,2), dpi=PLOT_DPI)\n\nsns.boxplot(data=df.melt(), x='value', y='auxiliary', palette=['steelblue','darkorange','teal'], flierprops={'marker':'.','markeredgecolor':'grey','markersize':1}, saturation=1)\nax.set_xlim(0,200)\nax.set_title(f'{CR}Pulses quantity in event separated by auxiliary,{CR}zoomed in')\nax.set_xlabel('pulses quantity')\n\nplt.show()","metadata":{"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> The quantity and scale of unconfirmed pulses outliers are large, but outliers of confirmed pulses much larger.\n\n> Confidence intervals of the quantity of confirmed and unconfirmed pulses almost do not intersect.  \n> The median is significantly different.  \n> The minimum and maximum values differ significantly.","metadata":{}},{"cell_type":"markdown","source":"---\nThe most frequent quantity of confirmed pulses in an event.","metadata":{}},{"cell_type":"code","source":"df.value_counts('auxiliary_False').head().to_frame().reset_index().rename(columns={'auxiliary_False':'auxiliary_False_qnty',0:'events_qnty'}).set_index('auxiliary_False_qnty')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nTop 20 events with the most frequent quantity of confirmed pulses in total are:","metadata":{}},{"cell_type":"code","source":"df.value_counts('auxiliary_False').head(20).to_frame().reset_index().rename(columns={0:'counts'}).counts.sum()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nAverage quantity of confirmed pulses for the top 20:","metadata":{}},{"cell_type":"code","source":"df.value_counts('auxiliary_False').head(20).to_frame().reset_index().auxiliary_False.mean()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> In most events, there are about 20-70 unconfirmed pulses (we assume that this is noise).\n\n> The quantity of confirmed pulses in an event is most often about 11. The top 20 in terms of the frequency of confirmed impulses make up about 60% of all events. The average quantity of confirmed pulses in them is 14.5. This may indicate a high level of noise in most events.\n\n> The quantity of unconfirmed pulses is characterized by relatively small outliers.\nQuantity of confirmed pulses are characterized by very huge outliers.\n\n> Taken together this may indicate that:\n>\n> - unconfirmed pulses are noise;  \n> - most of events are strongly noisy;  \n> - the reliability of event increases with the increase of confirmed pulses quantity.","metadata":{}},{"cell_type":"markdown","source":"---\nLogarithm of pulses quantity in event.","metadata":{}},{"cell_type":"code","source":"n_bins = 100\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\nsns.histplot(x=np.log(df.total), bins=n_bins)\nax.set_xlim(1,7)\nax.set_title(f'{CR}Logarithm of pulses quantity in event')\nax.set_xlabel('log(pulses_quantity)')\nplt.show()","metadata":{"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nLogarithm of pulses quantity in event separated by auxiliary.","metadata":{}},{"cell_type":"code","source":"df = df.melt()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_bins = 100\n\nfig, ax = plt.subplots(figsize=(15,5), dpi=PLOT_DPI)\nsns.histplot(x=np.log(df.value), hue=df.auxiliary, bins=n_bins, kde=True)\nax.set_xlim(1,7)\nax.set_title(f'{CR}Logarithm of pulses quantity in event separated by auxiliary,{CR}{n_bins} bins')\nax.set_xlabel('log(pulses_quantity)')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Confirmed pulses dominate in the area of low and high values.\n> \n> Unconfirmed pulses dominate in the area of middle values.","metadata":{}},{"cell_type":"code","source":"n_bins = 1000\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\nsns.histplot(x=np.log(df.value), hue=df.auxiliary, bins=n_bins, kde=True)\nax.set_xlim(1,7)\nax.set_title(f'{CR}Logarithm of pulses quantity in event separated by auxiliary,{CR}{n_bins} bins')\nax.set_xlabel('log(pulses_quantity)')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> You can clearly see the difference between the histograms for 100 and 1000 bins.  \n> Confirmed events are concentrated in a relatively small number of events, so as the number of bins increases, their absolute value decreases slowly. Unlike unconfirmed events, which are scattered across a large number of events, so their histogram declines faster.  \nAs a consequence, the behavior of the summary histogram depends more on unconfirmed events.\n\n> As noted earlier the secondary ridges on the right side are only a technical error in the display of charts.","metadata":{}},{"cell_type":"markdown","source":"---\nThe number of events corresponding to the quantity of pulses confirmed in them.","metadata":{}},{"cell_type":"code","source":"df = data[data['auxiliary'] == False].groupby('event_id').agg(pulse_count = ('charge', 'count'))\n\nprint(f'{CR}The quantity of confirmed pulses in event and the quantity of such events (first 10 in ascending){CR}')\nfor n in range(1, 11):\n    print(f'{n}: {df[df[\"pulse_count\"] == n].shape[0]}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Fewer pulses involved in the trajectory calculation leads to theoretically lower trajectory accuracy. Let's call this the geometric component of the error. On the other hand, a smaller quantity of pulses in calculation leads to a faster algorithm (also theoretically).","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Pulses Quantity and Charge Average in an Event","metadata":{}},{"cell_type":"markdown","source":"---\nAverage charge depending on pulses quantity in event.","metadata":{}},{"cell_type":"code","source":"df = (data.groupby('event_id')\n          .agg(charge_avg = ('charge', 'mean'),\n               pulse_count = ('charge', 'count'))\n     )\n\ndf.sample(3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,7), dpi=PLOT_DPI)\nsns.scatterplot(x=df.pulse_count, y=df.charge_avg)\nax.set_title(f'{CR}Average charge depending on pulses quantity in event')\nplt.axvline(60e3, color='lightgrey')\nplt.axvline(70e3, color='lightgrey')\nplt.axline((0, 0.5), (80e3, 4.5), linewidth=1, color='steelblue')\nplt.axline((20e3, 31), (100e3, 17), linewidth=1, color='steelblue')\n\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> You can trace the wedge converging from left to right (between the blue lines).\n\n> It is possible to determine the average value for any range and the scatter limit of the charge average in an event (i.e. first we average the pulse charge within the event, then we average between events in a given range of the pulses quantity). For example, you can calculate the minimum, maximum, mean and deviation from the predicted mean between the vertical gray lines on a graph.","metadata":{}},{"cell_type":"markdown","source":"---\nLogarith of average charge depending on pulses quantity in event.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,7), dpi=PLOT_DPI)\nsns.scatterplot(x=np.log(df.pulse_count), y=np.log(df.charge_avg), size=np.log(df.charge_avg), sizes=(1,20))\nax.set_title(f'{CR}Logarith of average charge depending on pulses quantity in event')\nax.set_xlabel('log (pulse_count)')\nax.set_ylabel('log (pulse_charge_avg)')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> The lower bound is clearly visible.  \n> Perhaps there is a weak trend.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## **Zenith**\n\nThe direction of the zenith axis is from south to north.","metadata":{}},{"cell_type":"markdown","source":"---\nFraction of events with zenith > $\\frac{\\pi}{2}$ (i.e. above the horizon).","metadata":{}},{"cell_type":"code","source":"data_meta[data_meta.zenith > np.pi/2].shape[0] / data_meta.shape[0]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nFraction of pulses with zenith > $\\frac{\\pi}{2}$ (i.e. above the horizon).","metadata":{}},{"cell_type":"code","source":"data[data.zenith > np.pi/2].shape[0] / data.shape[0]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nFraction of confirmed pulses with zenith > $\\frac{\\pi}{2}$ (i.e. above the horizon).","metadata":{}},{"cell_type":"code","source":"data[(data.zenith > np.pi/2) & (data.auxiliary == False)].shape[0] / data[data.auxiliary == False].shape[0]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Events are almost equaly distributed above and below the horizon line.  \n> The fraction of impulses is shifted down relative to the horizon line.  \n> The proportion of confirmed impulses is even more shifted down relative to the horizon.","metadata":{}},{"cell_type":"markdown","source":"---\nQuantity of events depending on zenith.","metadata":{}},{"cell_type":"code","source":"n_bins = 180\n\nfig, ax = plt.subplots(figsize=(15,3), dpi=PLOT_DPI)\nsns.histplot(x=data_meta.zenith, bins=n_bins)\nax.set_title(f'{CR}Quantity of events depending on zenith{CR}1 degree step')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> There is a clear trend in the quantity of events depending on zenith.  \n> The maximum quantity of events is observed perpendicular to the azimuth line.  \n> When trajectory line approaches the zenith line the quantity of events falls close to zero.\n>\n> `Assumption: this behavior is due to the relative position of the sensors.`  \n> `Assumption: this behavior is due to specific features of atmospheric neutrinos.`","metadata":{}},{"cell_type":"markdown","source":"---\nQuantity of pulses depending on zenith.","metadata":{}},{"cell_type":"code","source":"n_bins = 180\n\nfig, ax = plt.subplots(figsize=(15,3), dpi=PLOT_DPI)\nsns.histplot(x=data.zenith, bins=n_bins)\nax.set_title(f'{CR}Quantity of pulses depending on zenith{CR}1 degree step')\nplt.show()","metadata":{"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nQuantity of pulses depending on zenith, auxiliary=False.","metadata":{}},{"cell_type":"code","source":"n_bins = 180\n\nfig, ax = plt.subplots(figsize=(15,3), dpi=PLOT_DPI)\nsns.histplot(x=data[data.auxiliary == False].zenith, bins=n_bins)\nax.set_title(f'{CR}Quantity of pulses depending on zenith{CR}1 degree step{CR}auxiliary=False')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> The distribution of confirmed impulses is very similar to the distribution of impulses in general, since the proportion of confirmed impulses is 72%.\n> \n> There is a noticeable shift to the left (below the horizon line).\n>\n> The pronounced peaks are probably due to individual events involving a very large number of pulses.","metadata":{}},{"cell_type":"markdown","source":"---\nQuantity of pulses depending on zenith, auxiliary=True.","metadata":{}},{"cell_type":"code","source":"n_bins = 180\n\nfig, ax = plt.subplots(figsize=(15,3), dpi=PLOT_DPI)\nsns.histplot(x=data[data.auxiliary == True].zenith, bins=n_bins, kde=True)\nax.set_title(f'{CR}Quantity of pulses depending on zenith{CR}1 degree step{CR}auxiliary=True')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> The distribution is suspiciously even. Perhaps this suggests that unconfirmed pulses are noise.  \n> This distribution is typical for atmospheric background.\n> \n> The distribution of unconfirmed impulses is almost the same as the distribution of events.  \n> `Assumption: most of the events were due to atmospheric background (noise?).`","metadata":{}},{"cell_type":"markdown","source":"---\nLet's filter out events in which the ratio of confirmed and unconfirmed pulses is greater than a specified value.","metadata":{}},{"cell_type":"code","source":"df = (\n      data\n      .groupby(['event_id','auxiliary'])\n      .size()\n      .unstack()\n      .assign(ratio = lambda row: row[0]/row[1])      # ratio of confirmed and unconfirmed pulses\n     )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def ratio_auxiliary_filter(df: pd.DataFrame, threshold=1):\n    '''\n    В переданном датафрейме оставляет только те события,\n    в которых соотношение подтвержденных и неподтвержденных импульсов выше заданного порога.\n    Добавляет данные о zenith.\n    \n    '''\n    return (\n            df\n            .query('ratio > @threshold')                              # ratio filter\n            .reset_index()\n            .merge(data_meta[['event_id','zenith']], on='event_id')   # zenith data\n            .zenith\n           )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_bins = 180\nthreshould_list = [0.0, 0.5, 1.0, 1.5]\ncolors=['lightgrey','skyblue','steelblue','teal']\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\n\nfor i, threshould in enumerate(threshould_list):\n    sns.histplot(x=ratio_auxiliary_filter(df, threshold=threshould), bins=n_bins, color=colors[i])\n\nax.set_title(f'{CR}Quantity of pulses depending on zenith{CR}with auxiliary False/True ratio {threshould_list}{CR}1 degree step')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_bins = 180\nthreshould_list = [2, 4, 8, 16]\ncolors=['lightgrey','skyblue','steelblue','teal']\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\n\nfor i, threshould in enumerate(threshould_list):\n    sns.histplot(x=ratio_auxiliary_filter(df, threshold=threshould), bins=n_bins, color=colors[i])\n\nax.set_title(f'{CR}Quantity of pulses depending on zenith{CR}with auxiliary False/True ratio {threshould_list}{CR}1 degree step')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> As the ratio of confirmed/unconfirmed pulses in event increases, the distribution becomes less uniform.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## **Azimuth**","metadata":{}},{"cell_type":"markdown","source":"The azimuth analysis is similar to the zenith analysis.  \nThe conclusions are almost the same.","metadata":{}},{"cell_type":"markdown","source":"---\nQuantity of events depending on azimuth.","metadata":{}},{"cell_type":"code","source":"n_bins = 360\n\nfig, ax = plt.subplots(figsize=(15,2), dpi=PLOT_DPI)\nsns.histplot(x=data_meta.azimuth, bins=n_bins)\nax.set_title(f'{CR}Quantity of events depending on azimuth{CR}1 degree step')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nQuantity of pulses depending on azimuth.","metadata":{}},{"cell_type":"code","source":"n_bins = 360\n\nfig, ax = plt.subplots(figsize=(15,2), dpi=PLOT_DPI)\nsns.histplot(x=data.azimuth, bins=n_bins)\nax.set_title(f'{CR}Quantity of pulses depending on azimuth{CR}1 degree step')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nQuantity of pulses depending on azimuth, auxiliary=False.","metadata":{}},{"cell_type":"code","source":"n_bins = 360\n\nfig, ax = plt.subplots(figsize=(15,2), dpi=PLOT_DPI)\nsns.histplot(x=data[data.auxiliary == False].azimuth, bins=n_bins)\nax.set_title(f'{CR}Quantity of pulses depending on azimuth{CR}1 degree step{CR}auxiliary=False')\nplt.show()","metadata":{"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nQuantity of pulses depending on azimuth, auxiliary=True.","metadata":{}},{"cell_type":"code","source":"n_bins = 360\n\nfig, ax = plt.subplots(figsize=(15,2), dpi=PLOT_DPI)\nsns.histplot(x=data[data.auxiliary == True].azimuth, bins=n_bins)\nax.set_title(f'{CR}Quantity of pulses depending on azimuth{CR}1 degree step{CR}auxiliary=True')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Similar with zenith unconfirmed pulses are like noise making most events look noisy.\n\n> The difference between **azimuth** and **zenith** is that unconfirmed pulses are not only evenly but similiar distributed over the entire range of **azimuth** values.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## **Time**\n\nThe time of the pulse in nanoseconds in the current event time window. The absolute time of a pulse has no relevance, and only the relative time with respect to other pulses within an event is of relevance.","metadata":{}},{"cell_type":"markdown","source":"### Time Slots","metadata":{}},{"cell_type":"markdown","source":"Statistical describe of pulse time separated by auxiliary.","metadata":{}},{"cell_type":"code","source":"data.groupby('auxiliary').time.describe().astype(int)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Statistics of pulse time for confirmed and unconfirmed signals differ little.","metadata":{}},{"cell_type":"markdown","source":"---\nTime of pulses.","metadata":{}},{"cell_type":"code","source":"n_bins = 500\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\nsns.histplot(x=data.time, bins=n_bins)\nax.set_xlim(0,40e3)\nax.set_title(f'{CR}Time of pulses')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> There is a sharp rise of pulses quantity when the time reaches about 10,000 ns, as well as a certain constant tail to the left of this rise.","metadata":{}},{"cell_type":"markdown","source":"---\nTime of pulses separated by auxiliary.","metadata":{}},{"cell_type":"code","source":"n_bins = 500\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\nsns.histplot(x=data.time, hue=data.auxiliary, bins=n_bins)\nax.set_xlim(0,40e3)\nax.set_title(f'{CR}Time of pulses{CR}separated by auxiliary')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_bins = 500\n\nfig, ax = plt.subplots(figsize=(15,8), dpi=PLOT_DPI)\nsns.histplot(x=data.time, hue=data.auxiliary, bins=n_bins)\nax.set_xlim(0, 40e3)\nax.set_ylim(0, 1.6e5)\nax.set_title(f'{CR}Time of pulses{CR}separated by auxiliary{CR}zoomed in Y axis')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Several zones can be distinguished on the chart. The reasons for their appearance remain to be explored.\n\n> The quantity of confirmed pulses increases sharply from about 10,000 ns.\n\n> The \"left\" tail is caused by unconfirmed pulses (probably noise).\n> It contains a very small proportion of confirmed signals. It is necessary to find the reason for their appearance.\n\n> Questions about how IceCube works:\n> - how is the beginning of the event determined (point 0 on the timeline)?\n> - what caused clearly defined zones on the chart?\n> - how is it determined whether the event started inside the detector or outside?\n> - how is the end of the event determined?\n\n> **The `time` analysis indicates a lack of information about how IceCube works.**","metadata":{}},{"cell_type":"markdown","source":"### Event Timing Assumptions\n\nLet's estimate what is the maximum and minimum time required for the signal to propagate completely inside the cube.  \nBy full propagation, we mean that all sensors have had enough time to react.","metadata":{}},{"cell_type":"markdown","source":"**Important!**  \nFurther the location of the sensors is considered to be a square when viewed from above, but in fact it is a hexagon. Therefore, the calculations need to be refined.**","metadata":{}},{"cell_type":"markdown","source":"---\nThe maximum flash propagation time in IceCube will be if the sensor in the corner of the cube reacts first. After that, the light should reach the most distant corner of the cube.","metadata":{}},{"cell_type":"code","source":"n_bins = 500\n\nfig, ax = plt.subplots(figsize=(15,6), dpi=PLOT_DPI)\nsns.histplot(x=data.time, hue=data.auxiliary, bins=n_bins)\nax.set_xlim(5000, 12000)\nax.set_ylim(0, 3000)\nax.set_title(f'{CR}Time of pulses{CR}separated by auxiliary{CR}extremely zoomed of the lower left corner')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# approximate speed of light in ice, in nanoseconds (so all results are also approximate)\nICE_LIGHT_SPEED = 229e6 / 1e9\n\n# distance between furthest points / speed of light in ice\nmax_pulse_arrivement = (data[['X','Y','Z']].max()**2).sum()**0.5 / ICE_LIGHT_SPEED\n\nprint(f'maximum possible flash propagation time: {max_pulse_arrivement:0.0f} нс')","metadata":{"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The minimum flash propagation time in IceCube will be if the sensor in the center of the cube face with the longest sides reacts first. After that, the light should reach the most distant corner of the cube.","metadata":{}},{"cell_type":"code","source":"# let's assume that the sensor in the center of one of the faces of the cube reacted first\n# select the side from which the flash propagation path is minimal\n\nmax_pulse_arrivement = min(\n                           ((data.X.max()  )**2 + (data.Y.max()/2)**2 + (data.Z.max()/2)**2),\n                           ((data.X.max()/2)**2 + (data.Y.max()  )**2 + (data.Z.max()/2)**2),\n                           ((data.X.max()/2)**2 + (data.Y.max()/2)**2 + (data.Z.max()  )**2),\n                          ) ** 0.5 / ICE_LIGHT_SPEED\n\nprint(f'minimum possible flash propagation time: {max_pulse_arrivement:0.0f} нс')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> The minimum possible flash propagation time in IceCube almost exactly corresponds to the minimum value in the **time** field.\n>\n> At around 8,000 ns, a gradual increase in the number of pulses begins. This roughly corresponds to the time of the maximum possible propagation of the flash in the cube. It is not yet clear how this is related to the **time** values (i.e. how the timing of events is determined), but the coincidence is unlikely.  \n> It may matter if the event started inside or outside the detector.\n\n> It would be useful to know the principle of determining the end of an event.\n\n> `Assumption: pulses with the small 'time' (up to about 16-20,000 ns) are most useful. Perhaps pulses with larger 'time' values are useless or even harmful.`","metadata":{}},{"cell_type":"markdown","source":"### **Time** and **Charge**","metadata":{}},{"cell_type":"code","source":"# small sample to speed up and readability of visualization\ndf = data.sample(frac=0.01, random_state=RS)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,7), dpi=PLOT_DPI)\nsns.scatterplot(x=df[df.auxiliary == False].time, y=df.charge, size=1, legend=False)\nax.set_ylim(0,400)\nax.set_title(f'{CR}Time and Charge of pulses{CR}auxiliary=False{CR}cropped in Y-axis')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Vertical lines are traced in the rarefied area: flashes of different power with approximately the same time.\n\n> The strongest pulses are recorded around 10,000 ns. It is assumed that the strongest pulses should be caused by astrophysical neutrinos, and such pulses should not be accompanied by a shower of other particles, i.e. the event must last a relatively short time. How short the duration of an event should be depends on the timing principle on IceCube. Presumably not more than `8 + 8 = 16,000 ns`.\n>\n> If the event lasts longer it may be caused by atmospheric neutrinos.\n\n> `Assumption: time above some threshold does not carry useful information (at least for astrophysical neutrinos).`\n> We should try to check it.","metadata":{}},{"cell_type":"markdown","source":"### Charge Sum in an Event for different Time Slices (or Time Windows)","metadata":{}},{"cell_type":"markdown","source":"Based on the previous study let's select the time windows in the first approximation. That's enough for a start.","metadata":{}},{"cell_type":"code","source":"time_bins = [data.time.min(), 8000, 16000, 24000, data.time.max()]\ntime_labels = [0,1,2,3]\n\ndata['time_slice'], bin_eges = pd.cut(data.time, bins=time_bins, labels=time_labels, retbins=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = data.groupby(['event_id', 'time_slice', 'auxiliary']).agg(charge_total = ('charge','sum')).reset_index()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_bins = 100\nn_time_slice = df.time_slice.max() + 1\n\nfig, ax = plt.subplots(figsize=(15,15), dpi=PLOT_DPI, nrows=n_time_slice, ncols=1, sharex=True, sharey=True)\nfig.subplots_adjust(hspace=0.2)\nfig.suptitle(f'{CR}Total Charge in event for different time slices{CR}separated by auxiliary{CR}', fontsize=18)\n\nfor i in range(0, n_time_slice):\n    sns.histplot(x=np.log(df[df.time_slice == i].charge_total), hue=df[df.time_slice == i].auxiliary, bins=n_bins, ax=ax[i])\n    ax[i].set_title(f'time slice {bin_eges[i]}–{bin_eges[i+1]} нс')\n\nfig.tight_layout()\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> As expected after 16,000 ns (a more accurate value can be calculated) the total charge of the confirmed pulses in the event drops significantly. This confirms the assumption that this interval is useless. It is necessary to check this interval on a larger scale in order to avoid unexpected surprises.\n> \n> Up to 8,000 ns there are also no confirmed pulses.\n\n> Distribution of pulse charge over time requires more detailed research.\n\n> This chart is only an assessment of the potential of the feature `time`. Perhaps it makes sense to study the dependence of all signs on time in a separate study.","metadata":{}},{"cell_type":"markdown","source":"### Some summary for time slices separated by aixiliary.","metadata":{}},{"cell_type":"markdown","source":"Pulses total count (accross all events).","metadata":{}},{"cell_type":"code","source":"data.groupby(['time_slice','auxiliary']).size().unstack().assign(ratio = lambda x: x[0]/x[1])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Pulses total charge (accross all events).","metadata":{}},{"cell_type":"code","source":"data.groupby(['time_slice','auxiliary']).charge.sum().unstack().assign(ratio = lambda x: x[0]/x[1])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Sensors","metadata":{}},{"cell_type":"markdown","source":"### Repeated Triggering of Sensors Within One Event\n\nWhen a sensor triggered more than once per event.","metadata":{}},{"cell_type":"code","source":"df = data.groupby(['event_id','sensor_id']).agg(sensor_works = ('x','count')).reset_index()\ndf.sort_values('sensor_works', ascending=False).head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_bins = df.sensor_works.max()\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI, nrows=1, ncols=3)\nsns.histplot(x=df.sensor_works, bins=n_bins, ax=ax[0])\nsns.histplot(x=df.sensor_works, bins=n_bins, ax=ax[1])\nsns.histplot(x=df.sensor_works, bins=n_bins, ax=ax[2])\nfig.suptitle('The quantity of triggers of the same sensor within one event', fontsize=18)\nax[0].set_title(f'{CR}no limt', fontsize=12, pad=5)\nax[1].set_title(f'{CR}limit 1000 by Y-axis', fontsize=12, pad=5)\nax[2].set_title(f'{CR}limit 10 by Y-axis', fontsize=12, pad=5)\nax[1].set_ylim(0,1000)\nax[2].set_ylim(0,10)\nfig.tight_layout()\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Percentage of sensors triggered in the event:')\nfor n in range(1,6):\n    print(f'{n} time(s): {df.query(\"sensor_works == @n\").shape[0] / df.shape[0]:0.2%}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> The triggers quantity of the same sensor within one event — along the X axis.\n \n> In most events the sensors were triggered 1 time, but in a considerable part of the events the sensors were triggered more than 1 time, and in some cases up to 100 times or more.\n>\n> If the event is associated with an astrophysical neutrino, no more than 1 sensor trigger per event is expected.\n\n> We can further explore the chains of repeated sensors triggering within the event.  \n> We can further explore \"clean\" events (all sensors worked no more than 1 time).\n>\n> For \"clean\" events we can compare the dependence of triggered sensors fraction with the pulse charge (averaged, total, or something else).  \n> `Assumption: correlation should show up.`\n\n> We can compare the dependence of repeated triggers of sensors with the sensors quantity in event.  \n> `Assumption: correlation should show up.`","metadata":{}},{"cell_type":"markdown","source":"---\nRepeated triggering of sensors within the same event for confirmed pulses.","metadata":{}},{"cell_type":"code","source":"df = data[data.auxiliary == False].groupby(['event_id','sensor_id']).agg(sensor_works = ('x','count')).reset_index()\n\ndf.sensor_works.value_counts().head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Repeated triggering of sensors within the same event for **un**confirmed pulses.","metadata":{}},{"cell_type":"code","source":"df = data[data.auxiliary == True].groupby(['event_id','sensor_id']).agg(sensor_works = ('x','count')).reset_index()\n\ndf.sensor_works.value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Repeated sensor triggers for unconfirmed pulses are much less common: about 10%, of which 4/5 are double trips.  \n> In fact, it is uncommon for unconfirmed pulses to trigger the same sensor again within the event. Apparently, this is how the noise should look like.","metadata":{}},{"cell_type":"markdown","source":"### Quantity of Unique Sensors Within One Event\n\nAs a first approximation the quantity of unique triggered sensors can be considered as the flash energy equivalent. Or the energy of the detected particle.","metadata":{}},{"cell_type":"markdown","source":"---\nSeries of charts:  \n- zoomed in on the X-axis to view the area with a small quantity of sensors;\n- zoomed in on the Y-axis to view the area of a large quantity of sensors.","metadata":{}},{"cell_type":"code","source":"df = data.groupby(['event_id','auxiliary']).agg(sensor_unique_qnty = ('sensor_id', 'nunique')).reset_index()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_bins_0 = 100\nn_bins_1 = 400\n\nfig, ax = plt.subplots(figsize=(15,7), dpi=PLOT_DPI, nrows=2, ncols=3)\nfig.suptitle(f'{CR}Quantity of unique sensors within Event', fontsize=18)\n\nsns.histplot(x=df.sensor_unique_qnty, hue=df.auxiliary, bins=n_bins_0, ax=ax[0,0])\nsns.histplot(x=df.sensor_unique_qnty, hue=df.auxiliary, bins=n_bins_0, ax=ax[0,1])\nsns.histplot(x=df.sensor_unique_qnty, hue=df.auxiliary, bins=n_bins_0, ax=ax[0,2])\nax[0,0].set_title(f'{CR}no limit', fontsize=12, pad=5)\nax[0,1].set_title(f'{CR}limit 200 by Y-axis', fontsize=12, pad=5)\nax[0,2].set_title(f'{CR}limit 20 by Y-axis', fontsize=12, pad=5)\nax[0,1].set_ylim(0,200)\nax[0,2].set_ylim(0,20)\n\nsns.histplot(x=df.sensor_unique_qnty, hue=df.auxiliary, bins=n_bins_1, ax=ax[1,0])\nsns.histplot(x=df.sensor_unique_qnty, hue=df.auxiliary, bins=n_bins_1, ax=ax[1,1])\nsns.histplot(x=df.sensor_unique_qnty, hue=df.auxiliary, bins=n_bins_1, ax=ax[1,2])\nax[1,0].set_title(f'{CR}no limit', fontsize=12, pad=5)\nax[1,1].set_title(f'{CR}limit 500 by X-axis', fontsize=12, pad=5)\nax[1,2].set_title(f'{CR}limit 100 by X-axis', fontsize=12, pad=5)\nax[1,1].set_xlim(0,500)\nax[1,2].set_xlim(0,100)\n\nfig.tight_layout()\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Histograms of confirmed and unconfirmed pulses are significantly different and easily separated.\n\n> Both in the area of low values (the quantity of unique triggered sensors) and in the area of high values, confirmed impulses significantly prevail. On the one hand, in most cases the number of triggered sensors is no more than 20. On the other hand, the histogram has a long tail to the right.\n>\n> In the region of average values (about 20–70) unconfirmed impulses prevail. The distribution of unconfirmed impulses looks like normal.\n\n> `Assumption: the long right tail is associated with high-energy astrophysical neutrinos.`\n\n> `Assumption: the first triggering of the sensor has the greatest benefit, subsequent triggerings are useless or even harmful (secondary triggerings may be caused by concomitant flows of other particles).`","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## **Zenith** and **Charge**","metadata":{}},{"cell_type":"code","source":"# small sample to speed up and readability of visualization\ndf = data.sample(frac=0.01, random_state=RS)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,7), dpi=PLOT_DPI)\n# sns.scatterplot(x=np.log(df.charge), y=df.zenith, hue=df.auxiliary, size=np.log(df.charge), sizes=(1,10))\nsns.scatterplot(x=np.log(df.charge), y=df.zenith, hue=df.auxiliary, size=2, sizes=(2,2))\nax.set_title(f'{CR}Logarithm of charge depending on zenith')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Noisy data is clearly visible near 0 (pulse charge is about 1 photoelectron).\n\n> Recognizable horizontal bars indicate events with a large total charge.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## First triggering of sensor in an event\n\nThis is the first attempt to evaluate the usefulness of first triggering of the sensors. If the exploration will be useful it can be expanded.\n\nLet's try to consider only the first triggering of the sensor in the event.","metadata":{}},{"cell_type":"code","source":"# поиск первого срабатывания датчика (время=мин)\ndf = data.groupby(['event_id','sensor_id']).time.min().reset_index()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# attaching other pulse features\ndf = df.merge(data, on=['event_id','sensor_id','time'], how='left')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Comparison of the pulses quantity at the first sensor triggering and the total quantity of pulses.","metadata":{}},{"cell_type":"code","source":"print(f'{CR}for first sensor triggering:')\ndf.auxiliary.value_counts(sort=False).to_frame().T.assign(ratio_False_True = lambda x: x[1]/x[0])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'{CR}for all pulses:')\ndata.auxiliary.value_counts(sort=False).to_frame().T.assign(ratio_False_True = lambda x: x[1]/x[0])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Repeated sensor triggering in most cases refers to confirmed pulses.","metadata":{}},{"cell_type":"markdown","source":"---\nPulses charge for first sensor triggering","metadata":{}},{"cell_type":"code","source":"n_bins = 300\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\nsns.histplot(x=np.log(df.charge), hue=df.auxiliary, bins=n_bins)\nax.set_title(f'{CR}Logarithm of pulses charge for first sensor triggering{CR}separated by auxiliary')\nax.set_xlabel('log(charge)')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> After removal of repeated sensors triggering, the charge distributions of  confirmed and unconfirmed pulses become very similar.\n>\n> `Assumption: repeated sensors triggering are important for pulse confirmation (auxiliary=True)`.","metadata":{}},{"cell_type":"markdown","source":"---\nTotal charge in event for first sensor triggering","metadata":{}},{"cell_type":"code","source":"df_1st = df.groupby(['event_id', 'auxiliary']).agg(charge_total = ('charge','sum')).reset_index()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_total = data.groupby(['event_id', 'auxiliary']).agg(charge_total = ('charge','sum')).reset_index().query('charge_total <= @df_1st.charge_total.max()')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_bins = 200\n\nfig, ax = plt.subplots(figsize=(15,4), dpi=PLOT_DPI)\nax_total = sns.histplot(x=np.log(df_total.charge_total), hue=df_total.auxiliary, bins=n_bins, palette=['skyblue','navajowhite'])\nax_1st = sns.histplot(x=np.log(df_1st.charge_total), hue=df_1st.auxiliary, bins=n_bins)\nax.set_title(f'{CR}Logarithm of total charge in event for first sensor triggering only (darker){CR}compared to total charge in event without any filter (lighter){CR}separated by auxiliary')\nax.set_xlabel('log(charge_sum)')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> The general form of distributions has not changed.\n\n> The expected shift to the left has occurred.  \n> For confirmed pulses the shift is greater.  \n> The separation between histograms for confirmed and unconfirmed pulses has become clearer.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Event Duration and Charge","metadata":{}},{"cell_type":"markdown","source":"When calculating the duration of an event either only confirmed or unconfirmed pulses are taken.","metadata":{}},{"cell_type":"code","source":"df_F = (data[data.auxiliary == False]\n      .groupby('event_id')\n      .agg(time_min = ('time','min'), time_max = ('time','max'), charge_total = ('charge','sum'))\n      .assign(duration = lambda row: row.time_max - row.time_min)\n     )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_T = (data[data.auxiliary == True]\n      .groupby('event_id')\n      .agg(time_min = ('time','min'), time_max = ('time','max'), charge_total = ('charge','sum'))\n      .assign(duration = lambda row: row.time_max - row.time_min)\n     )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,5), dpi=PLOT_DPI)\nsns.scatterplot(x=df_F.duration, y=np.log(df_F.charge_total), size=1, sizes=(1,1), legend=False)\nsns.scatterplot(x=df_T.duration, y=np.log(df_T.charge_total), size=1, sizes=(1,1), color='orange', legend=False)\nax.set_xlim(0,40000)\nax.set_title(f'{CR}Event duration and total charge of pulses{CR}separated by auxiliary{CR}{CR}chart for auxiliary=True on top')\nplt.show()\n\nfig, ax = plt.subplots(figsize=(15,5), dpi=PLOT_DPI)\nsns.scatterplot(x=df_T.duration, y=np.log(df_T.charge_total), size=1, sizes=(1,1), color='orange', legend=False)\nsns.scatterplot(x=df_F.duration, y=np.log(df_F.charge_total), size=1, sizes=(1,1), legend=False)\nax.set_xlim(0,40000)\nax.set_title(f'{CR}chart for auxiliary=False on top')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> The shape of the distributions for confirmed and unconfirmed pulses significantly differs.\n>\n> The area of intersection of distributions is relatively small.\n>\n> The intervals between confirmed pulses starts from zero. For unconfirmed pulses intervals starts mostly at about 6000 ns.\n>\n> The distribution shape for confirmed pulses becomes more sparse and bifurcates around 10,000 ns.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Coordinates","metadata":{}},{"cell_type":"markdown","source":"Total charge of pulses depending on the average coordinates of pulses in event","metadata":{}},{"cell_type":"code","source":"df = (data\n      .groupby(['event_id'])\n      .agg(charge_total = ('charge','sum'), x_avg = ('x','mean'), y_avg = ('y','mean'), z_avg = ('z','mean'))\n      .reset_index()\n     )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,3), dpi=PLOT_DPI)\nsns.scatterplot(x=df.x_avg, y=np.log(df.charge_total), size=5, sizes=(5,5), legend=False)\nax.set_title(f'{CR}Total charge of pulses depending on the average X coordinate of pulses in event')\nplt.show()\n\nfig, ax = plt.subplots(figsize=(15,3), dpi=PLOT_DPI)\nsns.scatterplot(x=df.y_avg, y=np.log(df.charge_total), size=5, sizes=(5,5), legend=False)\nax.set_title(f'{CR}Total charge of pulses depending on the average Y coordinate of pulses in event')\nplt.show()\n\nfig, ax = plt.subplots(figsize=(15,3), dpi=PLOT_DPI)\nsns.scatterplot(x=df.z_avg, y=np.log(df.charge_total), size=5, sizes=(5,5), legend=False)\nax.set_title(f'{CR}Total charge of pulses depending on the average Z coordinate of pulses in event')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Charts have a strange rounded shape at the bottom.  \n> Let's try to find the reason.","metadata":{}},{"cell_type":"markdown","source":"---\nTotal charge of pulses depending on the average coordinates of pulses in event **separated by auxiliary**","metadata":{}},{"cell_type":"code","source":"df = (data\n      .groupby(['event_id','auxiliary'])\n      .agg(charge_total = ('charge','sum'), x_avg = ('x','mean'), y_avg = ('y','mean'), z_avg = ('z','mean'))\n      .reset_index()\n     )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,3), dpi=PLOT_DPI)\nsns.scatterplot(x=df.x_avg, y=np.log(df.charge_total), hue=df.auxiliary, size=5, sizes=(5,5), legend=False)\nax.set_title(f'{CR}Total charge of pulses depending on the average X coordinate of pulses in event{CR}separated by auxiliary')\nplt.show()\n\nfig, ax = plt.subplots(figsize=(15,3), dpi=PLOT_DPI)\nsns.scatterplot(x=df.y_avg, y=np.log(df.charge_total), hue=df.auxiliary, size=5, sizes=(5,5), legend=False)\nax.set_title(f'{CR}Total charge of pulses depending on the average Y coordinate of pulses in event{CR}separated by auxiliary')\nplt.show()\n\nfig, ax = plt.subplots(figsize=(15,3), dpi=PLOT_DPI)\nsns.scatterplot(x=df.z_avg, y=np.log(df.charge_total), hue=df.auxiliary, size=5, sizes=(5,5), legend=False)\nax.set_title(f'{CR}Total charge of pulses depending on the average Z coordinate of pulses in event{CR}separated by auxiliary')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Unconfirmed pulses are the reason for the rounded shape at the bottom of the charts in the previous section.\n\n> Graphs look a little cut off at the edges. The probable reason for this is the events that took place outside the IceCube.\n\n> Unconfirmed pulses are noticeably asymmetric.\n\n> For confirmed pulses the boundary between the dense zone and the sparse zone is clearly visible.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"strings\"></a> ","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Strings of sensors","metadata":{"lang":"en"}},{"cell_type":"code","source":"data.groupby(['x','y']).agg(pulse_count = ('charge','count'))","metadata":{"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> In some places we see fragmentation: one coordinate X corresponds to several close coordinates Y (and one Y corresponds several close X).  \n> The fragmentation is accompanied by a multiple drop in the registered pulses quantity.  \n> We will find out the reason later.","metadata":{}},{"cell_type":"markdown","source":"### Common configuration of strings with sensors, plane projection on (x, y)","metadata":{}},{"cell_type":"code","source":"df = data.groupby(['x','y']).agg(pulse_count = ('charge','count')).reset_index()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,15), dpi=PLOT_DPI)\nsns.scatterplot(x=df.x, y=df.y, size=df.pulse_count, sizes=(df.pulse_count.min()/1000, df.pulse_count.max()/1000), legend=False)\nax.set_title(f'{CR}Common configuration of strings with sensors{CR}and the total pulses quantity,{CR}plane projection on (x, y)')\nplt.axis('scaled')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> In the center the strings are denser due to extra strings wich register more pulses.\n>\n> Eight extra strings at the center of the array were deployed more compactly, with a horizontal separation of about 70 meters and a vertical sensors spacing of 7 meters. This denser configuration forms the DeepCore subdetector for lowers the neutrino energy threshold.\n>\n> `Assumption: regular string data and subcore data should be analyzed separately.`\n\n> Regular strings capture approximately the same number of pulses. At the point with coordinates about x=444 and y=194 one string looks extremely unusual. Next we'll explore this broken string.","metadata":{}},{"cell_type":"markdown","source":"### Broken string about x=444 and y=194","metadata":{}},{"cell_type":"code","source":"broken_x_min, broken_x_max = 443.4, 444.1\nbroken_y_min, broken_y_max = -194.6, -193.9\n\nfig, ax = plt.subplots(figsize=(15,15), dpi=PLOT_DPI)\nsns.scatterplot(x=df.x, y=df.y, size=50, sizes=(80,80), legend=False)\nax.set_title(f'{CR}Total pulses quantity for broken string about x=444 and y=194,{CR}plane projection on (x, y)')\nax.set_xlim(broken_x_min, broken_x_max)\nax.set_ylim(broken_y_min, broken_y_max)\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Broken string looks spiral: the coordinates of the sensors vary within +/- 0.35 meters from the center. This should not affect IceCube performance, but is an inconvenience for string analysis.\n>\n> We can either combine this data into a straight string, or temporarily remove it.\n>\n> Even by \"straightening\" the string, the total quantity of pulses registered by it will be radically lower than that of other strings.  \n> Given this, it is possible to recognize this string as indeed broken, and remove the data associated with it.","metadata":{}},{"cell_type":"markdown","source":"### Removing broken string data","metadata":{}},{"cell_type":"code","source":"data = data.loc[~((data.x > broken_x_min) & \n                  (data.x < broken_x_max) & \n                  (data.y > broken_y_min) &\n                  (data.y < broken_y_max))\n               ]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = data.groupby(['x','y']).agg(pulse_count = ('charge','count')).reset_index()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,15), dpi=PLOT_DPI)\nsns.scatterplot(x=df.x, y=df.y, size=df.pulse_count, sizes=(df.pulse_count.min()/1000, df.pulse_count.max()/1000), legend=False)\nax.set_title(f'{CR}Common configuration of strings with sensors{CR}and the total pulses quantity,{CR}plane projection on (x, y){CR}(broken string about x=444 and y=194 removed)')\nplt.axis('scaled')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Broken string data removed.  \n> Now the overall picture looks more diverse.  \n> Special strings in the center register much more pulses than regular strings.\n\n> Now we see that regular strings register very different pulses quantity. Perhaps they also have problems with the health of the sensors. This can be explored sometime.","metadata":{}},{"cell_type":"markdown","source":"### Quantity of pulses for sensors strings separated by `auxiliary`","metadata":{}},{"cell_type":"code","source":"df = data.groupby(['x','y', 'auxiliary']).agg(pulse_count = ('charge','count')).reset_index()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,15), dpi=PLOT_DPI)\nsns.scatterplot(x=df.x, y=df.y, hue=df.auxiliary, size=df.pulse_count, sizes=(df.pulse_count.min()/1000, df.pulse_count.max()/1000))\nax.set_title(f'{CR}Common configuration of strings with sensors{CR}and the total pulses quantity{CR}separated by auxiliary,{CR}plane projection on (x, y)')\nplt.axis('scaled')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Now the pulses are distributed on strings about evenly regardless of the `auxliary`. About IceCore it was mentioned above.","metadata":{}},{"cell_type":"markdown","source":"### Total charge of pulses for sensors strings separated by `auxiliary`","metadata":{}},{"cell_type":"code","source":"df = data.groupby(['x','y', 'auxiliary']).agg(charge_total = ('charge','sum')).reset_index()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,15), dpi=PLOT_DPI)\nsns.scatterplot(x=df.x, y=df.y, hue=df.auxiliary, size=df.charge_total, sizes=(df.charge_total.min()/1000, df.charge_total.max()/1000))\nax.set_title(f'{CR}Total charge of pulses for sensors strings,{CR}separated by auxiliary,{CR}plane projection on (x, y)')\nplt.axis('scaled')\n\nhandles, labels = ax.get_legend_handles_labels()\nfor h in handles[4:]:\n    sizes = [s / 12 for s in h.get_sizes()]\n    h.set_sizes(sizes)\nlabels[4:] = [str(float(s) * 1000) for s in labels[4:]]\nplt.legend(handles, labels)\n\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Nothing out of the ordinary for a total charge.  \n> Events with a large total charge are relatively rare, so there is some unevenness among the strings.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Assumption to check A\n\nThe pulse charge is inversely proportional to the distance$^3$ from the collision point to the sensor.","metadata":{"lang":"en"}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Assumption to check B\n\nThe impulse charge decreases as time increases. The problem with this statement may be related to repeated triggering of the sensors.","metadata":{"lang":"en"}},{"cell_type":"code","source":"(data\n .groupby('event_id')\n .agg(charge_total = ('charge','sum'), pulse_count = ('charge','count'), charge_avg = ('charge','median'))\n .sort_values('charge_avg', ascending=False)\n .query('pulse_count > 100')\n .query('charge_avg > 2')\n#  .head()\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can select any event with certain characteristics.","metadata":{"lang":"en"}},{"cell_type":"code","source":"EVENT_ID = 1767265\n\ndf = (data\n      .query('event_id == @EVENT_ID')\n      .groupby('sensor_id')\n      .time.min()\n      .to_frame()\n      .reset_index()\n      .merge(data, on=['sensor_id','time'], how='left')\n     )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15,6), dpi=PLOT_DPI)\nsns.scatterplot(x=df.time, y=np.log(df.charge), hue=df.auxiliary, size=np.log(df.charge), sizes=(1,20), legend=False)\n# ax.set_title(f'{CR}Total charge of pulses depending on the average X coordinate of pulses in event{CR}separated by auxiliary')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Checked this way and that, with and without repeated triggering: no pattern of charge versus time was found.","metadata":{"lang":"en"}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Summary\n\n**Section in the process of being translated into English.**","metadata":{}},{"cell_type":"markdown","source":"### Общие замечания","metadata":{}},{"cell_type":"markdown","source":"Это черновой вариант исследования, требующий чистки и упорядочивания.  \nРяд исследований можно расширить и углубить.  \nМногие зависимости еще не исследованы.","metadata":{}},{"cell_type":"markdown","source":"Некоторые расчетные поля пригодятся для следующего этапа работы. Например, количество импульсов в событии и т.п.  \nМожно оптимизировать данное исследование, добавив все необходимые данные в самом начале.","metadata":{}},{"cell_type":"markdown","source":"Не окончательной ясности, что представляют собой НЕподтвержденные импульсы (auxiliary=True). Это могут быть атмосферные нейтрино либо шум. В зависимости от контеста задачи атмосферные нейтрино можно считать шумом.","metadata":{}},{"cell_type":"markdown","source":"В постановке задачи нет указания на то, что главная цель – астрофизические нейтрино, хотя по заявлениям это – основная задача IceCube.\n\nЗа весь срок эксплуатации количество обнаруженных астрофизических нейтрино измеряется сотнями (более точного значения не нашел). Но в одной только в первой пачке данных из почти 700 содержится 200 тысяч событий. То есть количество астрофизических нейтрино (а значит, реально интересующих событий) составляет ничтожную часть.\n\nОбучение модели без разделения на астрофизические и прочие нейтрино может привести к оптимизации алгоритма не в пользу астрофизических, поскольку их ничтожно мало, и входные данные не сообщают, к какому классу относится анализируемый нейтрино (на самом деле это известно с высокой вероятностью, благодаря в том числе IceTop).  \n\nВозможно, задачи IceCube связаны не только с астрофизическими нейтрино.","metadata":{}},{"cell_type":"markdown","source":"Насчет оценки качества модели.\n\nПроверка будет по неким \"известным\" результатам. Но сами эти результаты могли быть получены только расчетным путем (если речь не идет о мюонах и черенковском излучении). То есть фактически более качественной окажется не та модель, которая точнее рассчитает траекторию, а та, которая окажется ближе к уже проведенным расчетам. Если за ошибку принимать отклонение от старой модели, то зачем вообще делать новую модель?\n\nСтоит отметить, что достижимая точность при определении траектории мюонных нейтрино и прочих видов нейтрино на данный момент отличается более чем на порядок.\n\nВ описании задания упоминается, что имеющиеся модели при достаточной точности не обладают приемлемой скоростью.  \nТогда имеет место несовпадение целей соревнования (точность) с реальными потребностями (скорость).","metadata":{}},{"cell_type":"markdown","source":"### Предположение 1\n\nИсключительно по вспышкам (импульсам) нельзя восстановить траекторию нейтрино.\n\nНейтрино перестает существовать в момент соударения с веществом. При этом может порождаться каскад частиц, количество которых зависит от начальной энергии нейтрино. ФЭУ улавливают фотоны, образующиеся в результате взаимодействия каскада частиц с веществом. То есть расположение импульсов – это не траектория нейтрино, а фотоны, разлетающиеся в разные стороны.\n\nПо импульсам невозможно (или крайне трудно) восстановить даже точку соударения нейтрино с веществом, поскольку фотоны рождаются не в один момент и не в одной точке пространства. Чем выше энергия нейтрино, тем больший каскад оно способно породить. Световой фронт вовсе не выглядит как шар.  Черенковское излучение, отражающее этот процесс, имеет довольно сложную форму.\n\nВ задании соревнования мельком упомянуто, что заказчик не предоставляет данные об энергии нейтрино, координатах точки столкновения, кинематике процесса, а также данные с IceTop. Все эти данные позволяют определить траекторию нейтрино с определенной точностью (около 1 градуса для одиночных нейтрино и 10-15 градусов для каскадов). Очевидно, что без этих данных не достичь даже такой точности.","metadata":{}},{"cell_type":"markdown","source":"### Предположение 2\n\nЧто делать, если верно Предположение 1?\n\nПомимо данных об импульсах у нас есть данные о траекториях. Видимо, это и есть задача данного соревнования: не рассчитать траекторию нейтрино, а найти связь между импульсами и уже рассчитанной иным способом траекторией. То есть вместо того, чтобы строить модель для каждого события (где наблюдения – это импульсы), а построить модель для всех событий (где наблюдения – это события).\n\nСобытие состоит из некоторого количества наблюдений, но для модели нужно, чтобы событие состояло из одной строки, то есть чтобы событие было наблюдением. Для этого нужно \"свернуть\" все данные об импульсах для каждого события в одно наблюдение (строку), создав таким образом описание события. Новые признаки события должны помочь предсказать таргет – траекторию.","metadata":{}},{"cell_type":"markdown","source":"### Предположение 3\n\nВозможно, нужно делать 2 отдельных модели, каждая для своего таргета: азимут и зенит.\n\nFeature engineering при \"сворачивании\" наблюдений об импульсах в наблюдения о событиях.  \n\nВозможные признаки:\n- размеры облака импульсов: максимальный размер (длина), минимальный размер (ширина).. возможно, также размер по 3-й оси, перпендикулярной первым двум (глубина?).. прочие характеристика формы облака;\n- статистики заряда: суммарный, средний, медиана, минимальный, максимальный, стандартное отклонение и т.п... для заряда, возможно, лучше подойдет логарифмическая шкала;\n- распределение заряда внутри облака;\n- по аналогии с зарядом построить признаки на основе time;\n- если облако содержит крайние датчики, событие могло произойти вне детектора.. можно количественно определить степень касания (например, соотношение площади касания к объему облака – круто, но сложно) .. (либо приблизительная оценка, например, соотношение числа датчиков на боковых слоях ко всему объему).","metadata":{}}]}