{"cells":[{"metadata":{},"cell_type":"markdown","source":"# I made this kernel public after the private LB was revealed. It shows that to match the train mean of TTF, you require a very high TTF on average for your submission."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train_orig = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int32, 'time_to_failure': np.float32})","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Recall an all zero's prediction scores 4.017 on the public LB. This means the public LB has an average TTF of 4.017"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_orig['time_to_failure'].values.mean()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# The mean value in the entire train is 5.678285"},{"metadata":{"trusted":true},"cell_type":"code","source":"# The median value?\nnp.median(train_orig['time_to_failure'].values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# The min value?\nnp.min(train_orig['time_to_failure'].values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# The 1% value?\nnp.quantile(train_orig['time_to_failure'].values, 0.01)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import os\nnum_test_samples = len(os.listdir('../input/test/'))\nprint('There are',num_test_samples,'test samples')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# There are 2624 test samples. Recall the public LB is 13% of the test data."},{"metadata":{"trusted":true},"cell_type":"code","source":"0.13 * 2624","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# This means there are about 341 segments in the public LB. What does the mean of the private LB have to be such that the mean of the entire test equals 5.678285, i.e. the mean of the entire test equals the mean of the entire train?"},{"metadata":{"trusted":true},"cell_type":"code","source":"np.append(np.repeat(4.017, 341), np.repeat(5.926423, 2624-341)).mean()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# So, if the test resembles the train, we expect the private LB to have a mean of 5.926423; thus, I am okay with higher-biased models."},{"metadata":{},"cell_type":"markdown","source":"# So, if we want a hedge away from the predictions made by the pixel measurements from the p4677 data, we could have a submission that scales to average LB prediction of 5.926423. This can be either from scaling, or from a new CV scheme."},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}