{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Google Smartphone Decimeter Challenge by LightGBM\n\nSolution is based on referring to various notebooks posted at Kaggle, along with my changes. Seeing this as a learning experience.\n\n## Notes\n\n If you’re outside, with open sky, the GPS accuracy from your phone is about five meters, and that’s been constant for a while. With raw GNSS measurements from the phones, this can now improve dramatically.\n \n### The GNSS problem description\nhttps://github.com/commaai/laika\n\nGNSS satellites orbit the earth broadcasting signals that allow the receiver to determine the distance to each satellite. These satellites have known orbits and so their positions are known. This makes determining the receiver's position a basic 3-dimensional trilateration problem. In practice observed distances to each satellite will be measured with some offset that is caused by the receiver's clock error. This offset also needs to be determined, making it a 4-dimensional trilateration problem. \n\n<img src= \"https://camo.githubusercontent.com/0d85f5131c63442f8e7b46de7dab8040a7d693effd5e611ebed25be0b7600a32/68747470733a2f2f75706c6f61642e77696b696d656469612e6f72672f77696b6970656469612f636f6d6d6f6e732f7468756d622f632f63332f33737068657265732e7376672f36323270782d33737068657265732e7376672e706e67\"  alt =\"GNSS\" style=\"width:400px;height:400px;\">\n\nSince this problem is generally overdetermined (more than 4 satellites to solve the 4d problem) there is a variety of methods to compute a position estimate from the measurements. One can use a basic weighted least squares solver for experimental purposes. This is far from optimal due to the dynamic nature of the system, this makes a Bayesian estimator like a Kalman filter the preferred estimator.\n\nHowever, the above description is over-simplified. Getting accurate distance estimates to satellites and the satellite's position from the receiver observations is not trivial. This is what we call processing of the GNSS observables and it is this procedure laika is designed to make easy.\n \n ### How positioning works?\n - Send a burst of these transactions and, as a consequence, the system can calculate ranging statistics, such as the mean and the variance.\n \n <img src= \"https://www.gpsworld.com/wp-content/uploads/2018/07/Android-Figure-3.jpg\" alt =\"Wifi distance\">\n \n Wi-Fi RTT principles, basic concept. Image by Frank van Diggelen, Roy Want and Wei Wang\n \n - Take these ranges of separate access points; if those ranges were accurate, they would define four circles that would intersect at a single point. In practice, because of error in each range, a maximum likelihood position is calculated using a least squares multilateration algorithm.\n \n - Further refine this position by repeating the process, particularly as the phone moves, and then calculate trajectory using filtering techniques, such as Kalman filtering, to optimize the estimate.\n \n  <img src= \"https://www.gpsworld.com/wp-content/uploads/2018/07/Android-Figure-4.jpg\" alt =\"Workflow\">\n  \n  Wi-Fi Workflow. Image by: Frank van Diggelen, Roy Want and Wei Wang.\n\n \n \n### Steps for the solution from Sohier Dane\n- Smoothing out the baseline estimates\n- Integrating readings from other phone instruments, like the accelerometer.\n- Satellite triangulation using the *derived.csv files.\n- Building triangulations directly from the raw gnss logs. \n- Incorporating external data for controls like satellite readings from base stations in the area.\n\n### References: \n- https://www.kaggle.com/tensorchoko/google-smartphone-lightgbm\n- https://www.kaggle.com/jeongyoonlee/google-smartphone-decimeter-eda-keras-tpu\n- Discussions topics from Sohier Dane\n- Google I/O https://www.gpsworld.com/how-to-achieve-1-meter-accuracy-in-android/\n- GPS Survey Workshop video https://www.youtube.com/watch?v=vOJ3u7Zd_i0\n- Hardware: Centimeter Positioning with a Smartphone-Quality GNSS Antenna https://www.youtube.com/watch?v=rCOvklUB5vQ","metadata":{}},{"cell_type":"code","source":"!pip install simdkalman","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-07-16T05:29:31.102179Z","iopub.execute_input":"2021-07-16T05:29:31.102738Z","iopub.status.idle":"2021-07-16T05:29:41.192945Z","shell.execute_reply.started":"2021-07-16T05:29:31.102618Z","shell.execute_reply":"2021-07-16T05:29:41.191527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%matplotlib inline\nfrom matplotlib import pyplot as plt\nimport numpy as np # linear algebra\nfrom pathlib import Path\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom scipy import sparse\nfrom sklearn.metrics import mean_squared_error\nfrom sklearn.model_selection import KFold\nfrom sklearn.preprocessing import StandardScaler\nimport seaborn as sns\nimport simdkalman\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tqdm.notebook import tqdm\nfrom warnings import simplefilter\n\nsimplefilter('ignore')\nplt.style.use('fivethirtyeight')\npd.set_option('max_columns', 100)\npd.set_option('max_rows', 100)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2021-07-16T05:29:41.194770Z","iopub.execute_input":"2021-07-16T05:29:41.195105Z","iopub.status.idle":"2021-07-16T05:29:49.088779Z","shell.execute_reply.started":"2021-07-16T05:29:41.195062Z","shell.execute_reply":"2021-07-16T05:29:49.087773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_name = 'nn_v2'\n\ndata_dir = Path('../input/google-smartphone-decimeter-challenge')\ntrain_file = data_dir / 'baseline_locations_train.csv'\ntest_file = data_dir / 'baseline_locations_test.csv'\nsample_file = data_dir / 'sample_submission.csv'\n\nbuild_dir = Path('./build')\nbuild_dir.mkdir(parents=True, exist_ok=True)\npredict_val_file = build_dir / f'{model_name}.val.txt'\npredict_tst_file = build_dir / f'{model_name}.tst.txt'\nsubmission_file = 'submission.csv'\n\ncname_col = 'collectionName'\npname_col = 'phoneName'\nphone_col = 'phone'\nts_col = 'millisSinceGpsEpoch'\ndt_col = 'datetime'\nlat_col = 'latDeg'\nlon_col = 'lngDeg'\n\nlrate = .01 #.001\nbatch_size = 32 #1024\nepochs = 2000#100\nn_stop = 10\nn_fold = 5\nseed = 42#77","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:49.090446Z","iopub.execute_input":"2021-07-16T05:29:49.090873Z","iopub.status.idle":"2021-07-16T05:29:49.099332Z","shell.execute_reply.started":"2021-07-16T05:29:49.090839Z","shell.execute_reply":"2021-07-16T05:29:49.098435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv(train_file)\ntest = pd.read_csv(test_file)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:49.100952Z","iopub.execute_input":"2021-07-16T05:29:49.101427Z","iopub.status.idle":"2021-07-16T05:29:49.689001Z","shell.execute_reply.started":"2021-07-16T05:29:49.101395Z","shell.execute_reply":"2021-07-16T05:29:49.688135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.groupby('collectionName').apply(lambda x: x['phoneName'].unique())","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:49.690201Z","iopub.execute_input":"2021-07-16T05:29:49.690626Z","iopub.status.idle":"2021-07-16T05:29:49.768124Z","shell.execute_reply.started":"2021-07-16T05:29:49.690593Z","shell.execute_reply":"2021-07-16T05:29:49.766896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.groupby('collectionName').apply(lambda x: x['phoneName'].unique())","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:49.769636Z","iopub.execute_input":"2021-07-16T05:29:49.769996Z","iopub.status.idle":"2021-07-16T05:29:49.813919Z","shell.execute_reply.started":"2021-07-16T05:29:49.769950Z","shell.execute_reply":"2021-07-16T05:29:49.812821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.phoneName.unique()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:49.815610Z","iopub.execute_input":"2021-07-16T05:29:49.816058Z","iopub.status.idle":"2021-07-16T05:29:49.835830Z","shell.execute_reply.started":"2021-07-16T05:29:49.816012Z","shell.execute_reply":"2021-07-16T05:29:49.834407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f = open('../input/google-smartphone-decimeter-challenge/train/2020-05-14-US-MTV-1/Pixel4/supplemental/Pixel4_GnssLog.20o', 'r')\ndata = f.readlines()\nf.close()\ndata[:20]","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:49.839140Z","iopub.execute_input":"2021-07-16T05:29:49.839542Z","iopub.status.idle":"2021-07-16T05:29:49.924761Z","shell.execute_reply.started":"2021-07-16T05:29:49.839500Z","shell.execute_reply":"2021-07-16T05:29:49.923500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f = open('../input/google-smartphone-decimeter-challenge/train/2020-05-14-US-MTV-1/Pixel4/supplemental/SPAN_Pixel4_10Hz.nmea', 'r')\ndata = f.readlines()\nf.close()\ndata[:10]","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:49.927168Z","iopub.execute_input":"2021-07-16T05:29:49.927569Z","iopub.status.idle":"2021-07-16T05:29:50.001124Z","shell.execute_reply.started":"2021-07-16T05:29:49.927532Z","shell.execute_reply":"2021-07-16T05:29:49.999698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ground = pd.read_csv('../input/google-smartphone-decimeter-challenge/train/2020-05-14-US-MTV-1/Pixel4/ground_truth.csv')\nground.head(5)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:50.002694Z","iopub.execute_input":"2021-07-16T05:29:50.003081Z","iopub.status.idle":"2021-07-16T05:29:50.055160Z","shell.execute_reply.started":"2021-07-16T05:29:50.003048Z","shell.execute_reply":"2021-07-16T05:29:50.053823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f = open('../input/google-smartphone-decimeter-challenge/train/2020-05-14-US-MTV-1/Pixel4/Pixel4_GnssLog.txt', 'r')\ndata = f.readlines()\nf.close()\ndata[:10]","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:50.056819Z","iopub.execute_input":"2021-07-16T05:29:50.057195Z","iopub.status.idle":"2021-07-16T05:29:52.192374Z","shell.execute_reply.started":"2021-07-16T05:29:50.057158Z","shell.execute_reply":"2021-07-16T05:29:52.190794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived = pd.read_csv('../input/google-smartphone-decimeter-challenge/train/2020-05-14-US-MTV-1/Pixel4/Pixel4_derived.csv')\nderived","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:52.194433Z","iopub.execute_input":"2021-07-16T05:29:52.194866Z","iopub.status.idle":"2021-07-16T05:29:52.669897Z","shell.execute_reply.started":"2021-07-16T05:29:52.194817Z","shell.execute_reply":"2021-07-16T05:29:52.668616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived_unique = list(derived.millisSinceGpsEpoch.unique())\nlen(derived_unique)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:52.671799Z","iopub.execute_input":"2021-07-16T05:29:52.672147Z","iopub.status.idle":"2021-07-16T05:29:52.683017Z","shell.execute_reply.started":"2021-07-16T05:29:52.672112Z","shell.execute_reply":"2021-07-16T05:29:52.681733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ground_unique = list(ground.millisSinceGpsEpoch.unique())\nlen(ground_unique)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:52.684722Z","iopub.execute_input":"2021-07-16T05:29:52.685166Z","iopub.status.idle":"2021-07-16T05:29:52.698323Z","shell.execute_reply.started":"2021-07-16T05:29:52.685126Z","shell.execute_reply":"2021-07-16T05:29:52.697344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import json\njson_open = open('../input/google-smartphone-decimeter-challenge/metadata/accumulated_delta_range_state_bit_map.json', 'r')\njson.load(json_open)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:52.699857Z","iopub.execute_input":"2021-07-16T05:29:52.700263Z","iopub.status.idle":"2021-07-16T05:29:52.725691Z","shell.execute_reply.started":"2021-07-16T05:29:52.700226Z","shell.execute_reply":"2021-07-16T05:29:52.724036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import json\njson_open = open('../input/google-smartphone-decimeter-challenge/metadata/raw_state_bit_map.json', 'r')\njson.load(json_open)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:52.727558Z","iopub.execute_input":"2021-07-16T05:29:52.728024Z","iopub.status.idle":"2021-07-16T05:29:52.745313Z","shell.execute_reply.started":"2021-07-16T05:29:52.727972Z","shell.execute_reply":"2021-07-16T05:29:52.743379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.read_csv('../input/google-smartphone-decimeter-challenge/metadata/constellation_type_mapping.csv').head(5)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:52.747033Z","iopub.execute_input":"2021-07-16T05:29:52.747374Z","iopub.status.idle":"2021-07-16T05:29:52.766154Z","shell.execute_reply.started":"2021-07-16T05:29:52.747341Z","shell.execute_reply":"2021-07-16T05:29:52.764970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from https://www.kaggle.com/sohier/loading-gnss-logs\ndef gnss_log_to_dataframes(path):\n    print('Loading ' + path, flush=True)\n    gnss_section_names = {'Raw','UncalAccel', 'UncalGyro', 'UncalMag', 'Fix', 'Status', 'OrientationDeg'} #これはどこからでてきたのか？\n    with open(path) as f_open:\n        datalines = f_open.readlines()\n\n    datas = {k: [] for k in gnss_section_names}\n    gnss_map = {k: [] for k in gnss_section_names}\n    for dataline in datalines:\n        is_header = dataline.startswith('#')\n        dataline = dataline.strip('#').strip().split(',')\n        # skip over notes, version numbers, etc\n        if is_header and dataline[0] in gnss_section_names:\n            gnss_map[dataline[0]] = dataline[1:]\n        elif not is_header:\n            datas[dataline[0]].append(dataline[1:])\n\n    results = dict()\n    for k, v in datas.items():\n        results[k] = pd.DataFrame(v, columns=gnss_map[k])\n    # pandas doesn't properly infer types from these lists by default\n    for k, df in results.items():\n        for col in df.columns:\n            if col == 'CodeType':\n                continue\n            results[k][col] = pd.to_numeric(results[k][col])\n\n    return results","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-16T05:29:52.767677Z","iopub.execute_input":"2021-07-16T05:29:52.767979Z","iopub.status.idle":"2021-07-16T05:29:52.779628Z","shell.execute_reply.started":"2021-07-16T05:29:52.767949Z","shell.execute_reply":"2021-07-16T05:29:52.778082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gnss_section_names = {'Raw','UncalAccel', 'UncalGyro', 'UncalMag', 'Fix', 'Status', 'OrientationDeg'}\ndatas = {k: [] for k in gnss_section_names}\ngnss_map = {k: [] for k in gnss_section_names}\ndatas","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:52.782648Z","iopub.execute_input":"2021-07-16T05:29:52.783151Z","iopub.status.idle":"2021-07-16T05:29:52.804463Z","shell.execute_reply.started":"2021-07-16T05:29:52.783103Z","shell.execute_reply":"2021-07-16T05:29:52.802897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results = dict()\nfor k, v in datas.items():\n     results[k] = pd.DataFrame(v, columns=gnss_map[k])\nresults","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:52.806432Z","iopub.execute_input":"2021-07-16T05:29:52.807026Z","iopub.status.idle":"2021-07-16T05:29:52.834341Z","shell.execute_reply.started":"2021-07-16T05:29:52.806986Z","shell.execute_reply":"2021-07-16T05:29:52.833445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from https://www.kaggle.com/dannellyz/start-here-simple-folium-heatmap-for-geo-data\nimport folium\nfrom folium import plugins\n\n\ndef simple_folium(df:pd.DataFrame, lat_col:str, lon_col:str):\n    #Preprocess\n    #Drop rows that do not have lat/lon\n    df = df[df[lat_col].notnull() & df[lon_col].notnull()]\n\n    # Convert lat/lon to (n, 2) nd-array format for heatmap\n    # Then send to list\n    df_locs = list(df[[lat_col, lon_col]].values)\n\n    ##folium.Mapオブジェクト作成\n    fol_map = folium.Map([df[lat_col].median(), df[lon_col].median()])\n\n    # plot heatmap\n    heat_map = plugins.HeatMap(df_locs)\n    print(heat_map)\n    fol_map.add_child(heat_map)\n\n    # plot markers\n    markers = plugins.MarkerCluster(locations = df_locs)\n    fol_map.add_child(markers)\n\n    #Add Layer Control\n    folium.LayerControl().add_to(fol_map)\n\n    return fol_map","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-16T05:29:52.835721Z","iopub.execute_input":"2021-07-16T05:29:52.836053Z","iopub.status.idle":"2021-07-16T05:29:53.408136Z","shell.execute_reply.started":"2021-07-16T05:29:52.836021Z","shell.execute_reply":"2021-07-16T05:29:53.407031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from https://www.kaggle.com/jpmiller/baseline-from-host-data\n# simplified haversine distance\ndef calc_haversine(lat1, lon1, lat2, lon2):\n    lat1, lon1, lat2, lon2 = map(np.radians, [lat1, lon1, lat2, lon2])\n    dlat = lat2 - lat1\n    dlon = lon2 - lon1\n    a = np.sin(dlat/2.0)**2 + np.cos(lat1) * np.cos(lat2) * np.sin(dlon/2.0)**2\n\n    c = 2 * np.arcsin(a**0.5)\n    dist = 6_367_000 * c\n    return dist","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-16T05:29:53.409854Z","iopub.execute_input":"2021-07-16T05:29:53.410542Z","iopub.status.idle":"2021-07-16T05:29:53.417363Z","shell.execute_reply.started":"2021-07-16T05:29:53.410500Z","shell.execute_reply":"2021-07-16T05:29:53.416298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from https://www.kaggle.com/emaerthin/demonstration-of-the-kalman-filter\nT = 1.0 \nstate_transition = np.array([[1, 0, T, 0, 0.5 * T ** 2, 0], [0, 1, 0, T, 0, 0.5 * T ** 2], [0, 0, 1, 0, T, 0],\n                             [0, 0, 0, 1, 0, T], [0, 0, 0, 0, 1, 0], [0, 0, 0, 0, 0, 1]])\nprocess_noise = np.diag([1e-5, 1e-5, 5e-6, 5e-6, 1e-6, 1e-6]) + np.ones((6, 6)) * 1e-9\nobservation_model = np.array([[1, 0, 0, 0, 0, 0], [0, 1, 0, 0, 0, 0]])\nobservation_noise = np.diag([5e-5, 5e-5]) + np.ones((2, 2)) * 1e-9\n\nkf = simdkalman.KalmanFilter(\n        state_transition = state_transition,\n        process_noise = process_noise,\n        observation_model = observation_model,\n        observation_noise = observation_noise)\n\ndef apply_kf_smoothing(df, kf_=kf):\n    unique_paths = df[phone_col].unique()\n    for phone in tqdm(unique_paths):\n        data = df.loc[df[phone_col] == phone][[lat_col, lon_col]].values\n        data = data.reshape(1, len(data), 2)\n        smoothed = kf_.smooth(data)\n        df.loc[df[phone_col] == phone, lat_col] = smoothed.states.mean[0, :, 0]\n        df.loc[df[phone_col] == phone, lon_col] = smoothed.states.mean[0, :, 1]\n    return df","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:53.422209Z","iopub.execute_input":"2021-07-16T05:29:53.422899Z","iopub.status.idle":"2021-07-16T05:29:53.438592Z","shell.execute_reply.started":"2021-07-16T05:29:53.422853Z","shell.execute_reply":"2021-07-16T05:29:53.437371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trn = pd.read_csv(train_file)\nprint(trn.shape)\ntrn.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:53.442528Z","iopub.execute_input":"2021-07-16T05:29:53.443361Z","iopub.status.idle":"2021-07-16T05:29:53.679660Z","shell.execute_reply.started":"2021-07-16T05:29:53.443288Z","shell.execute_reply":"2021-07-16T05:29:53.678549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tst = pd.read_csv(test_file)\nprint(tst.shape)\ntst.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:53.681376Z","iopub.execute_input":"2021-07-16T05:29:53.682146Z","iopub.status.idle":"2021-07-16T05:29:53.847463Z","shell.execute_reply.started":"2021-07-16T05:29:53.682087Z","shell.execute_reply":"2021-07-16T05:29:53.846331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = pd.read_csv(sample_file)\nprint(sub.shape)\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:53.849319Z","iopub.execute_input":"2021-07-16T05:29:53.849937Z","iopub.status.idle":"2021-07-16T05:29:54.032470Z","shell.execute_reply.started":"2021-07-16T05:29:53.849887Z","shell.execute_reply":"2021-07-16T05:29:54.031314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cname = trn[cname_col][0]\ncname","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:54.033916Z","iopub.execute_input":"2021-07-16T05:29:54.034358Z","iopub.status.idle":"2021-07-16T05:29:54.041876Z","shell.execute_reply.started":"2021-07-16T05:29:54.034224Z","shell.execute_reply":"2021-07-16T05:29:54.040769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pname = trn[pname_col][0]\npname","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:54.043479Z","iopub.execute_input":"2021-07-16T05:29:54.043815Z","iopub.status.idle":"2021-07-16T05:29:54.058614Z","shell.execute_reply.started":"2021-07-16T05:29:54.043782Z","shell.execute_reply":"2021-07-16T05:29:54.057640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path =str(data_dir / 'train' / cname / pname / f'{pname}_GnssLog.txt')\nwith open(path) as f_open:\n        datalines = f_open.readlines()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:54.060014Z","iopub.execute_input":"2021-07-16T05:29:54.060601Z","iopub.status.idle":"2021-07-16T05:29:54.469765Z","shell.execute_reply.started":"2021-07-16T05:29:54.060557Z","shell.execute_reply":"2021-07-16T05:29:54.468778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"    for dataline in datalines:\n        is_header = dataline.startswith('#')\n        dataline = dataline.strip('#').strip().split(',')\n        break","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:54.471231Z","iopub.execute_input":"2021-07-16T05:29:54.472148Z","iopub.status.idle":"2021-07-16T05:29:54.477237Z","shell.execute_reply.started":"2021-07-16T05:29:54.472079Z","shell.execute_reply":"2021-07-16T05:29:54.476085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"datalines[:10]","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:54.478563Z","iopub.execute_input":"2021-07-16T05:29:54.479075Z","iopub.status.idle":"2021-07-16T05:29:54.500005Z","shell.execute_reply.started":"2021-07-16T05:29:54.479041Z","shell.execute_reply":"2021-07-16T05:29:54.498761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in [cname_col, pname_col]:\n    print(f'# of unique {col:>14s} in training: {trn[col].nunique():4d}')\n    print(f'# of unique {col:>14s}     in test: {tst[col].nunique():4d}')","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:54.501301Z","iopub.execute_input":"2021-07-16T05:29:54.501781Z","iopub.status.idle":"2021-07-16T05:29:54.664360Z","shell.execute_reply.started":"2021-07-16T05:29:54.501749Z","shell.execute_reply":"2021-07-16T05:29:54.662618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trn[pname_col].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:54.665968Z","iopub.execute_input":"2021-07-16T05:29:54.666649Z","iopub.status.idle":"2021-07-16T05:29:54.706914Z","shell.execute_reply.started":"2021-07-16T05:29:54.666604Z","shell.execute_reply":"2021-07-16T05:29:54.705820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tst[pname_col].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:54.708802Z","iopub.execute_input":"2021-07-16T05:29:54.709230Z","iopub.status.idle":"2021-07-16T05:29:54.742105Z","shell.execute_reply.started":"2021-07-16T05:29:54.709194Z","shell.execute_reply":"2021-07-16T05:29:54.741372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'# of unique phone in training: {trn[phone_col].nunique():4d}')\nprint(f'    # of unique phone in test: {tst[phone_col].nunique():4d}')","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:54.743548Z","iopub.execute_input":"2021-07-16T05:29:54.743852Z","iopub.status.idle":"2021-07-16T05:29:54.809803Z","shell.execute_reply.started":"2021-07-16T05:29:54.743823Z","shell.execute_reply":"2021-07-16T05:29:54.808435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trn[phone_col].value_counts()[:10]","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:54.811319Z","iopub.execute_input":"2021-07-16T05:29:54.811610Z","iopub.status.idle":"2021-07-16T05:29:54.853026Z","shell.execute_reply.started":"2021-07-16T05:29:54.811583Z","shell.execute_reply":"2021-07-16T05:29:54.851391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tst[phone_col].value_counts()[:10]","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:54.854880Z","iopub.execute_input":"2021-07-16T05:29:54.855349Z","iopub.status.idle":"2021-07-16T05:29:54.885221Z","shell.execute_reply.started":"2021-07-16T05:29:54.855305Z","shell.execute_reply":"2021-07-16T05:29:54.884408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"overlapping_phones = [x for x in tst[phone_col] if x in trn[phone_col]]\nprint(len(overlapping_phones))","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:54.886704Z","iopub.execute_input":"2021-07-16T05:29:54.887089Z","iopub.status.idle":"2021-07-16T05:29:55.695064Z","shell.execute_reply.started":"2021-07-16T05:29:54.887054Z","shell.execute_reply":"2021-07-16T05:29:55.694003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tst[ts_col].min(), tst[ts_col].max()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:55.696266Z","iopub.execute_input":"2021-07-16T05:29:55.696554Z","iopub.status.idle":"2021-07-16T05:29:55.704093Z","shell.execute_reply.started":"2021-07-16T05:29:55.696527Z","shell.execute_reply":"2021-07-16T05:29:55.703181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dt_offset = pd.to_datetime('1980-01-06 00:00:00')\nprint(dt_offset)\ndt_offset_in_ms = int(dt_offset.value / 1e6)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:55.705597Z","iopub.execute_input":"2021-07-16T05:29:55.705903Z","iopub.status.idle":"2021-07-16T05:29:55.721577Z","shell.execute_reply.started":"2021-07-16T05:29:55.705860Z","shell.execute_reply":"2021-07-16T05:29:55.720731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trn[dt_col] = pd.to_datetime(trn[ts_col] + dt_offset_in_ms, unit='ms')\ntst[dt_col] = pd.to_datetime(tst[ts_col] + dt_offset_in_ms, unit='ms')\nprint(f'Training data range: {trn[dt_col].min()} - {trn[dt_col].max()}')\nprint(f'    Test data range: {tst[dt_col].min()} - {tst[dt_col].max()}')","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:55.722719Z","iopub.execute_input":"2021-07-16T05:29:55.723227Z","iopub.status.idle":"2021-07-16T05:29:56.166202Z","shell.execute_reply.started":"2021-07-16T05:29:55.723194Z","shell.execute_reply":"2021-07-16T05:29:56.165124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"latlon_trn = trn[[lat_col, lon_col]].round(3)\nlatlon_trn['counts'] = 1\nlatlon_trn = latlon_trn.groupby([lat_col, lon_col]).sum().reset_index()\nlatlon_trn.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:56.167937Z","iopub.execute_input":"2021-07-16T05:29:56.168382Z","iopub.status.idle":"2021-07-16T05:29:56.210316Z","shell.execute_reply.started":"2021-07-16T05:29:56.168336Z","shell.execute_reply":"2021-07-16T05:29:56.209145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"    #def simple_folium(df:pd.DataFrame, lat_col:str, lon_col:str):\n    simple_folium(latlon_trn, lat_col, lon_col)\n    df = pd.DataFrame(latlon_trn)\n\n    #Preprocess\n    #Drop rows that do not have lat/lon\n    df = df[df[lat_col].notnull() & df[lon_col].notnull()]\n    df\n","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:56.211788Z","iopub.execute_input":"2021-07-16T05:29:56.212092Z","iopub.status.idle":"2021-07-16T05:29:56.337082Z","shell.execute_reply.started":"2021-07-16T05:29:56.212060Z","shell.execute_reply":"2021-07-16T05:29:56.335944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n    # Convert lat/lon to (n, 2) nd-array format for heatmap\n    # Then send to list\n    df_locs = list(df[[lat_col, lon_col]].values)\n    ","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:56.338649Z","iopub.execute_input":"2021-07-16T05:29:56.338969Z","iopub.status.idle":"2021-07-16T05:29:56.346037Z","shell.execute_reply.started":"2021-07-16T05:29:56.338935Z","shell.execute_reply":"2021-07-16T05:29:56.345232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[lat_col].median() #中央値","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:56.347578Z","iopub.execute_input":"2021-07-16T05:29:56.348192Z","iopub.status.idle":"2021-07-16T05:29:56.363045Z","shell.execute_reply.started":"2021-07-16T05:29:56.348156Z","shell.execute_reply":"2021-07-16T05:29:56.362001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"    fol_map = folium.Map([df[lat_col].median(), df[lon_col].median()])\n    fol_map","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:56.364466Z","iopub.execute_input":"2021-07-16T05:29:56.364788Z","iopub.status.idle":"2021-07-16T05:29:56.389134Z","shell.execute_reply.started":"2021-07-16T05:29:56.364756Z","shell.execute_reply":"2021-07-16T05:29:56.387571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"    # plot heatmap\n    heat_map = plugins.HeatMap(df_locs)\n    print(heat_map)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:56.390913Z","iopub.execute_input":"2021-07-16T05:29:56.391284Z","iopub.status.idle":"2021-07-16T05:29:56.421836Z","shell.execute_reply.started":"2021-07-16T05:29:56.391231Z","shell.execute_reply":"2021-07-16T05:29:56.420729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"    fol_map.add_child(heat_map)\n    fol_map","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:56.423605Z","iopub.execute_input":"2021-07-16T05:29:56.424303Z","iopub.status.idle":"2021-07-16T05:29:56.468863Z","shell.execute_reply.started":"2021-07-16T05:29:56.424233Z","shell.execute_reply":"2021-07-16T05:29:56.467725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n    # plot markers\n    markers = plugins.MarkerCluster(locations = df_locs)\n    ","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:56.470412Z","iopub.execute_input":"2021-07-16T05:29:56.470709Z","iopub.status.idle":"2021-07-16T05:29:56.541716Z","shell.execute_reply.started":"2021-07-16T05:29:56.470680Z","shell.execute_reply":"2021-07-16T05:29:56.540695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fol_map.add_child(markers)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:56.543550Z","iopub.execute_input":"2021-07-16T05:29:56.544196Z","iopub.status.idle":"2021-07-16T05:29:58.063774Z","shell.execute_reply.started":"2021-07-16T05:29:56.544148Z","shell.execute_reply":"2021-07-16T05:29:58.062618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"    #Add Layer Control\n    folium.LayerControl().add_to(fol_map)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:58.065407Z","iopub.execute_input":"2021-07-16T05:29:58.065941Z","iopub.status.idle":"2021-07-16T05:29:58.074188Z","shell.execute_reply.started":"2021-07-16T05:29:58.065900Z","shell.execute_reply":"2021-07-16T05:29:58.072669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"simple_folium(latlon_trn, lat_col, lon_col)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:58.075936Z","iopub.execute_input":"2021-07-16T05:29:58.076608Z","iopub.status.idle":"2021-07-16T05:29:59.694965Z","shell.execute_reply.started":"2021-07-16T05:29:58.076540Z","shell.execute_reply":"2021-07-16T05:29:59.693688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"latlon_tst = tst[[lat_col, lon_col]].round(3)\nlatlon_tst","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:59.696763Z","iopub.execute_input":"2021-07-16T05:29:59.697168Z","iopub.status.idle":"2021-07-16T05:29:59.724218Z","shell.execute_reply.started":"2021-07-16T05:29:59.697128Z","shell.execute_reply":"2021-07-16T05:29:59.722985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"latlon_tst['counts'] = 1\nlatlon_tst = latlon_tst.groupby([lat_col, lon_col]).sum().reset_index()\nlatlon_tst","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:59.726006Z","iopub.execute_input":"2021-07-16T05:29:59.726412Z","iopub.status.idle":"2021-07-16T05:29:59.760311Z","shell.execute_reply.started":"2021-07-16T05:29:59.726370Z","shell.execute_reply":"2021-07-16T05:29:59.759209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"simple_folium(latlon_tst, lat_col, lon_col)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:29:59.761572Z","iopub.execute_input":"2021-07-16T05:29:59.761850Z","iopub.status.idle":"2021-07-16T05:30:01.857674Z","shell.execute_reply.started":"2021-07-16T05:29:59.761822Z","shell.execute_reply":"2021-07-16T05:30:01.856561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cname = trn[cname_col][0]\ncname","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:01.859075Z","iopub.execute_input":"2021-07-16T05:30:01.859411Z","iopub.status.idle":"2021-07-16T05:30:01.866779Z","shell.execute_reply.started":"2021-07-16T05:30:01.859377Z","shell.execute_reply":"2021-07-16T05:30:01.865698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pname = trn[pname_col][0]\npname","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:01.868704Z","iopub.execute_input":"2021-07-16T05:30:01.869217Z","iopub.status.idle":"2021-07-16T05:30:01.890397Z","shell.execute_reply.started":"2021-07-16T05:30:01.869172Z","shell.execute_reply":"2021-07-16T05:30:01.889089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfs = gnss = gnss_log_to_dataframes(str(data_dir / 'train' / cname / pname / f'{pname}_GnssLog.txt'))\nprint(dfs.keys())","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:01.892696Z","iopub.execute_input":"2021-07-16T05:30:01.893420Z","iopub.status.idle":"2021-07-16T05:30:25.108993Z","shell.execute_reply.started":"2021-07-16T05:30:01.893365Z","shell.execute_reply":"2021-07-16T05:30:25.108063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_raw = dfs['Raw']\nprint(df_raw.shape)\ndf_raw.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:25.110560Z","iopub.execute_input":"2021-07-16T05:30:25.111361Z","iopub.status.idle":"2021-07-16T05:30:25.148567Z","shell.execute_reply.started":"2021-07-16T05:30:25.111307Z","shell.execute_reply":"2021-07-16T05:30:25.145861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_raw.info()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:25.150179Z","iopub.execute_input":"2021-07-16T05:30:25.150549Z","iopub.status.idle":"2021-07-16T05:30:25.179224Z","shell.execute_reply.started":"2021-07-16T05:30:25.150512Z","shell.execute_reply":"2021-07-16T05:30:25.177962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_raw['ArrivalTime'] = df_raw['TimeNanos'] - df_raw['FullBiasNanos'] - df_raw['BiasNanos']\nprint(df_raw['ArrivalTime'].describe())\ndf_raw['ArrivalTime'].hist(bins=20)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:25.180769Z","iopub.execute_input":"2021-07-16T05:30:25.181089Z","iopub.status.idle":"2021-07-16T05:30:25.483043Z","shell.execute_reply.started":"2021-07-16T05:30:25.181057Z","shell.execute_reply":"2021-07-16T05:30:25.482031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_raw['BiasUncertaintyNanos'].describe())\ndf_raw['BiasUncertaintyNanos'].hist(bins=20)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:25.484768Z","iopub.execute_input":"2021-07-16T05:30:25.485217Z","iopub.status.idle":"2021-07-16T05:30:25.699392Z","shell.execute_reply.started":"2021-07-16T05:30:25.485170Z","shell.execute_reply":"2021-07-16T05:30:25.698219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_raw['ReceivedSvTimeUncertaintyNanos'].describe())\ndf_raw['ReceivedSvTimeUncertaintyNanos'].hist(bins=20)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:25.700849Z","iopub.execute_input":"2021-07-16T05:30:25.701165Z","iopub.status.idle":"2021-07-16T05:30:25.934756Z","shell.execute_reply.started":"2021-07-16T05:30:25.701133Z","shell.execute_reply":"2021-07-16T05:30:25.933921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_raw.AccumulatedDeltaRangeUncertaintyMeters.describe())\ndf_raw.AccumulatedDeltaRangeUncertaintyMeters.hist(bins=20)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:25.936314Z","iopub.execute_input":"2021-07-16T05:30:25.936678Z","iopub.status.idle":"2021-07-16T05:30:26.166543Z","shell.execute_reply.started":"2021-07-16T05:30:25.936641Z","shell.execute_reply":"2021-07-16T05:30:26.165445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_raw.Cn0DbHz.describe())\ndf_raw.Cn0DbHz.hist(bins=20)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:26.168314Z","iopub.execute_input":"2021-07-16T05:30:26.168668Z","iopub.status.idle":"2021-07-16T05:30:26.387862Z","shell.execute_reply.started":"2021-07-16T05:30:26.168634Z","shell.execute_reply":"2021-07-16T05:30:26.386538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_raw = df_raw.loc[\n    ~pd.isnull(df_raw.FullBiasNanos) &\n    (df_raw.BiasUncertaintyNanos < 100) &\n    (df_raw.ArrivalTime > 0) &\n    (df_raw.ConstellationType != 0) &\n    ~pd.isnull(df_raw.TimeNanos) &\n    (df_raw.State != 3) & (df_raw.State != 14) & (df_raw.State != 7) & (df_raw.State != 15) &\n    (df_raw.ReceivedSvTimeUncertaintyNanos < 100) &\n    (df_raw.AccumulatedDeltaRangeUncertaintyMeters < 0.3) &\n    (df_raw.Cn0DbHz > 20)\n]\nprint(df_raw.shape)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:26.389834Z","iopub.execute_input":"2021-07-16T05:30:26.390199Z","iopub.status.idle":"2021-07-16T05:30:26.428060Z","shell.execute_reply.started":"2021-07-16T05:30:26.390160Z","shell.execute_reply":"2021-07-16T05:30:26.427103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_raw","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:26.429214Z","iopub.execute_input":"2021-07-16T05:30:26.429531Z","iopub.status.idle":"2021-07-16T05:30:26.484317Z","shell.execute_reply.started":"2021-07-16T05:30:26.429500Z","shell.execute_reply":"2021-07-16T05:30:26.483212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived = pd.read_csv(data_dir / 'train' / cname / pname / f'{pname}_derived.csv')\nprint(derived.shape)\nderived.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:26.493180Z","iopub.execute_input":"2021-07-16T05:30:26.493609Z","iopub.status.idle":"2021-07-16T05:30:26.788575Z","shell.execute_reply.started":"2021-07-16T05:30:26.493576Z","shell.execute_reply":"2021-07-16T05:30:26.787533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived.info()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:26.791556Z","iopub.execute_input":"2021-07-16T05:30:26.791863Z","iopub.status.idle":"2021-07-16T05:30:26.828309Z","shell.execute_reply.started":"2021-07-16T05:30:26.791832Z","shell.execute_reply":"2021-07-16T05:30:26.827097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived = derived.loc[derived.constellationType != 0]\nprint(derived.shape)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:26.829701Z","iopub.execute_input":"2021-07-16T05:30:26.830000Z","iopub.status.idle":"2021-07-16T05:30:26.844164Z","shell.execute_reply.started":"2021-07-16T05:30:26.829970Z","shell.execute_reply":"2021-07-16T05:30:26.843108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:26.845632Z","iopub.execute_input":"2021-07-16T05:30:26.846173Z","iopub.status.idle":"2021-07-16T05:30:26.893329Z","shell.execute_reply.started":"2021-07-16T05:30:26.846138Z","shell.execute_reply":"2021-07-16T05:30:26.891916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived['correctedPrM'] = (derived['rawPrM'] + derived['satClkBiasM'] - derived['isrbM'] - \n                           derived['ionoDelayM'] - derived['tropoDelayM'])\nsns.pairplot(data=derived, vars=['correctedPrM', 'rawPrM'], size=3)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:26.894841Z","iopub.execute_input":"2021-07-16T05:30:26.895188Z","iopub.status.idle":"2021-07-16T05:30:28.760576Z","shell.execute_reply.started":"2021-07-16T05:30:26.895155Z","shell.execute_reply":"2021-07-16T05:30:28.759342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived[dt_col] = pd.to_datetime(derived[ts_col] + dt_offset_in_ms, unit='ms')\nprint(f'Data range for {cname}/{pname}: {derived[dt_col].min()} - {derived[dt_col].max()}')","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:28.762132Z","iopub.execute_input":"2021-07-16T05:30:28.762460Z","iopub.status.idle":"2021-07-16T05:30:28.777766Z","shell.execute_reply.started":"2021-07-16T05:30:28.762427Z","shell.execute_reply":"2021-07-16T05:30:28.776623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived[['constellationType', 'svid', 'signalType']].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:28.780154Z","iopub.execute_input":"2021-07-16T05:30:28.780802Z","iopub.status.idle":"2021-07-16T05:30:28.803743Z","shell.execute_reply.started":"2021-07-16T05:30:28.780751Z","shell.execute_reply":"2021-07-16T05:30:28.802431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived[[ts_col, 'constellationType', 'correctedPrM']].groupby([ts_col, 'constellationType']).agg(['mean', 'std', 'count']).describe()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:28.805458Z","iopub.execute_input":"2021-07-16T05:30:28.805808Z","iopub.status.idle":"2021-07-16T05:30:28.858214Z","shell.execute_reply.started":"2021-07-16T05:30:28.805774Z","shell.execute_reply":"2021-07-16T05:30:28.857154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived.loc[derived.constellationType == 1][[ts_col, 'svid', 'correctedPrM']].groupby([ts_col, 'svid']).agg(['mean', 'std', 'count']).describe()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:28.859728Z","iopub.execute_input":"2021-07-16T05:30:28.860029Z","iopub.status.idle":"2021-07-16T05:30:28.912126Z","shell.execute_reply.started":"2021-07-16T05:30:28.860000Z","shell.execute_reply":"2021-07-16T05:30:28.910732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.read_csv('../input/google-smartphone-decimeter-challenge/metadata/constellation_type_mapping.csv')","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:28.913989Z","iopub.execute_input":"2021-07-16T05:30:28.914480Z","iopub.status.idle":"2021-07-16T05:30:28.930501Z","shell.execute_reply.started":"2021-07-16T05:30:28.914429Z","shell.execute_reply":"2021-07-16T05:30:28.929693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived.loc[derived.signalType == 'GPS_L1'][[ts_col, 'svid', 'correctedPrM']].groupby([ts_col, 'svid']).agg(['mean', 'std', 'count'])","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:28.931599Z","iopub.execute_input":"2021-07-16T05:30:28.931909Z","iopub.status.idle":"2021-07-16T05:30:28.984849Z","shell.execute_reply.started":"2021-07-16T05:30:28.931878Z","shell.execute_reply":"2021-07-16T05:30:28.983597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived.loc[derived.signalType == 'GPS_L1'][[ts_col, 'svid', 'correctedPrM']].groupby([ts_col, 'svid']).agg(['mean', 'std', 'count']).describe()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:28.986602Z","iopub.execute_input":"2021-07-16T05:30:28.987003Z","iopub.status.idle":"2021-07-16T05:30:29.046257Z","shell.execute_reply.started":"2021-07-16T05:30:28.986966Z","shell.execute_reply":"2021-07-16T05:30:29.044993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived.loc[derived.signalType == 'GPS_L1'][[ts_col, 'svid']].drop_duplicates().groupby([ts_col]).agg(['mean', 'std', 'count']).describe()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.047983Z","iopub.execute_input":"2021-07-16T05:30:29.048477Z","iopub.status.idle":"2021-07-16T05:30:29.105137Z","shell.execute_reply.started":"2021-07-16T05:30:29.048426Z","shell.execute_reply":"2021-07-16T05:30:29.103976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gps_l1 = derived.loc[derived.signalType == 'GPS_L1'][[ts_col, 'svid', 'correctedPrM']].drop_duplicates([ts_col, 'svid'])\nprint(gps_l1.shape)\ngps_l1.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.106827Z","iopub.execute_input":"2021-07-16T05:30:29.107203Z","iopub.status.idle":"2021-07-16T05:30:29.142590Z","shell.execute_reply.started":"2021-07-16T05:30:29.107171Z","shell.execute_reply":"2021-07-16T05:30:29.141507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"label = pd.read_csv(data_dir / 'train' / cname / pname / 'ground_truth.csv')\nprint(label.shape)\nlabel.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.144203Z","iopub.execute_input":"2021-07-16T05:30:29.144817Z","iopub.status.idle":"2021-07-16T05:30:29.176524Z","shell.execute_reply.started":"2021-07-16T05:30:29.144772Z","shell.execute_reply":"2021-07-16T05:30:29.175156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"label[dt_col] = pd.to_datetime(label[ts_col] + dt_offset_in_ms, unit='ms')\nprint(f'Labels range for {cname}/{pname}: {label[dt_col].min()} - {label[dt_col].max()}')","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.178106Z","iopub.execute_input":"2021-07-16T05:30:29.178476Z","iopub.status.idle":"2021-07-16T05:30:29.187488Z","shell.execute_reply.started":"2021-07-16T05:30:29.178443Z","shell.execute_reply":"2021-07-16T05:30:29.186656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cname = trn[cname_col][10]\npname = trn[pname_col][10]\nderived2 = pd.read_csv(data_dir / 'train' / cname / pname / f'{pname}_derived.csv')\nlabel2 = pd.read_csv(data_dir / 'train' / cname / pname / 'ground_truth.csv')\nprint(f\"Derived data starts at: {pd.to_datetime(derived2[ts_col].min() + dt_offset_in_ms, unit='ms')}\")\nprint(f\"  Label data starts at: {pd.to_datetime(label2[ts_col].min() + dt_offset_in_ms, unit='ms')}\")","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.188783Z","iopub.execute_input":"2021-07-16T05:30:29.189305Z","iopub.status.idle":"2021-07-16T05:30:29.432514Z","shell.execute_reply.started":"2021-07-16T05:30:29.189253Z","shell.execute_reply":"2021-07-16T05:30:29.431237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trn.sort_values([phone_col, ts_col], inplace=True)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.434011Z","iopub.execute_input":"2021-07-16T05:30:29.434347Z","iopub.status.idle":"2021-07-16T05:30:29.516428Z","shell.execute_reply.started":"2021-07-16T05:30:29.434317Z","shell.execute_reply":"2021-07-16T05:30:29.515129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trn[['prev_lat']] = trn[lat_col].shift().where(trn[phone_col].eq(trn[phone_col].shift()))\ntrn[['prev_lat']] ","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.518084Z","iopub.execute_input":"2021-07-16T05:30:29.518422Z","iopub.status.idle":"2021-07-16T05:30:29.562835Z","shell.execute_reply.started":"2021-07-16T05:30:29.518389Z","shell.execute_reply":"2021-07-16T05:30:29.561532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trn[['prev_lon']] = trn[lon_col].shift().where(trn[phone_col].eq(trn[phone_col].shift()))\ntrn[['prev_lon']]","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.564393Z","iopub.execute_input":"2021-07-16T05:30:29.564701Z","iopub.status.idle":"2021-07-16T05:30:29.605868Z","shell.execute_reply.started":"2021-07-16T05:30:29.564672Z","shell.execute_reply":"2021-07-16T05:30:29.604661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tst.sort_values([phone_col, ts_col], inplace=True)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.607635Z","iopub.execute_input":"2021-07-16T05:30:29.607997Z","iopub.status.idle":"2021-07-16T05:30:29.662421Z","shell.execute_reply.started":"2021-07-16T05:30:29.607965Z","shell.execute_reply":"2021-07-16T05:30:29.661235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tst[['prev_lat']] = tst[lat_col].shift().where(tst[phone_col].eq(tst[phone_col].shift()))\ntst[['prev_lat']] ","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.663968Z","iopub.execute_input":"2021-07-16T05:30:29.664310Z","iopub.status.idle":"2021-07-16T05:30:29.695303Z","shell.execute_reply.started":"2021-07-16T05:30:29.664264Z","shell.execute_reply":"2021-07-16T05:30:29.694532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tst[['prev_lon']] = tst[lon_col].shift().where(tst[phone_col].eq(tst[phone_col].shift()))\ntrn.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.696435Z","iopub.execute_input":"2021-07-16T05:30:29.696839Z","iopub.status.idle":"2021-07-16T05:30:29.725593Z","shell.execute_reply.started":"2021-07-16T05:30:29.696807Z","shell.execute_reply":"2021-07-16T05:30:29.724827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from https://www.kaggle.com/jpmiller/baseline-from-host-data\nlabel_files = (data_dir / 'train').rglob('ground_truth.csv')\nlabel_files","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.726664Z","iopub.execute_input":"2021-07-16T05:30:29.727053Z","iopub.status.idle":"2021-07-16T05:30:29.733653Z","shell.execute_reply.started":"2021-07-16T05:30:29.727024Z","shell.execute_reply":"2021-07-16T05:30:29.732721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols = [phone_col, ts_col, lat_col, lon_col]\n\ndf_list = []\nfor t in tqdm(label_files, total=73):\n    label = pd.read_csv(t, usecols=[cname_col, pname_col, ts_col, lat_col, lon_col])\n    df_list.append(label)\n   ","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:29.735010Z","iopub.execute_input":"2021-07-16T05:30:29.735612Z","iopub.status.idle":"2021-07-16T05:30:31.323853Z","shell.execute_reply.started":"2021-07-16T05:30:29.735560Z","shell.execute_reply":"2021-07-16T05:30:31.322498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_label = pd.concat(df_list, ignore_index=True)\ndf_label","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:31.325504Z","iopub.execute_input":"2021-07-16T05:30:31.325819Z","iopub.status.idle":"2021-07-16T05:30:31.371159Z","shell.execute_reply.started":"2021-07-16T05:30:31.325790Z","shell.execute_reply":"2021-07-16T05:30:31.370007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.DataFrame(df_list)[:5]","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:31.372580Z","iopub.execute_input":"2021-07-16T05:30:31.372896Z","iopub.status.idle":"2021-07-16T05:30:31.488095Z","shell.execute_reply.started":"2021-07-16T05:30:31.372868Z","shell.execute_reply":"2021-07-16T05:30:31.486871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_label[phone_col] = df_label[cname_col] + '_' + df_label[pname_col]\ndf_label","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:31.489285Z","iopub.execute_input":"2021-07-16T05:30:31.489589Z","iopub.status.idle":"2021-07-16T05:30:31.560542Z","shell.execute_reply.started":"2021-07-16T05:30:31.489561Z","shell.execute_reply":"2021-07-16T05:30:31.559781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df_label.merge(trn[cols + ['prev_lat', 'prev_lon']], how='inner', on=[phone_col, ts_col], \n                    suffixes=('_gt', '')).drop([cname_col, pname_col], axis=1) #列名が重複している場合のサフィックスを指定: 引数suffixes\ndf","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:31.561622Z","iopub.execute_input":"2021-07-16T05:30:31.562007Z","iopub.status.idle":"2021-07-16T05:30:31.710389Z","shell.execute_reply.started":"2021-07-16T05:30:31.561978Z","shell.execute_reply":"2021-07-16T05:30:31.709066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['sSinceGpsEpoch'] = df[ts_col] // 1000 ## 切り捨て除算\nprint(df.shape)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:31.712220Z","iopub.execute_input":"2021-07-16T05:30:31.712675Z","iopub.status.idle":"2021-07-16T05:30:31.738754Z","shell.execute_reply.started":"2021-07-16T05:30:31.712631Z","shell.execute_reply":"2021-07-16T05:30:31.737684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_tst = sub[[phone_col, ts_col]].merge(tst[[phone_col, ts_col, lat_col, lon_col, 'prev_lat', 'prev_lon']], \n                                        how='left', on=[phone_col, ts_col], suffixes=('', '_basepred'))\ndf_tst","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:31.740507Z","iopub.execute_input":"2021-07-16T05:30:31.741109Z","iopub.status.idle":"2021-07-16T05:30:31.828948Z","shell.execute_reply.started":"2021-07-16T05:30:31.741061Z","shell.execute_reply":"2021-07-16T05:30:31.827540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_tst['sSinceGpsEpoch'] = df_tst[ts_col] // 1000\nprint(df_tst.shape)\ndf_tst.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:31.830556Z","iopub.execute_input":"2021-07-16T05:30:31.830865Z","iopub.status.idle":"2021-07-16T05:30:31.851850Z","shell.execute_reply.started":"2021-07-16T05:30:31.830836Z","shell.execute_reply":"2021-07-16T05:30:31.850788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived_files = (data_dir / 'train').rglob('*_derived.csv')\ncols = [ts_col, 'svid', 'correctedPrM']\nderived_files","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:31.853190Z","iopub.execute_input":"2021-07-16T05:30:31.853507Z","iopub.status.idle":"2021-07-16T05:30:31.861213Z","shell.execute_reply.started":"2021-07-16T05:30:31.853478Z","shell.execute_reply":"2021-07-16T05:30:31.860015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_list = []\nfor t in tqdm(derived_files, total=73):\n    derived = pd.read_csv(t).drop_duplicates([ts_col, 'svid'])\n    derived['correctedPrM'] = (derived['rawPrM'] + derived['satClkBiasM'] - derived['isrbM'] - \n                               derived['ionoDelayM'] - derived['tropoDelayM'])\n    df_list.append(derived[[cname_col, pname_col, ts_col, 'svid', 'correctedPrM']])","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:31.863294Z","iopub.execute_input":"2021-07-16T05:30:31.863781Z","iopub.status.idle":"2021-07-16T05:30:57.587508Z","shell.execute_reply.started":"2021-07-16T05:30:31.863739Z","shell.execute_reply":"2021-07-16T05:30:57.586084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_derived = pd.concat(df_list, ignore_index=True)\ndf_derived","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:57.589164Z","iopub.execute_input":"2021-07-16T05:30:57.589756Z","iopub.status.idle":"2021-07-16T05:30:57.699065Z","shell.execute_reply.started":"2021-07-16T05:30:57.589707Z","shell.execute_reply":"2021-07-16T05:30:57.697814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_derived[phone_col] = df_derived[cname_col] + '_' + df_derived[pname_col]\ndf_derived[phone_col] ","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:57.700834Z","iopub.execute_input":"2021-07-16T05:30:57.701156Z","iopub.status.idle":"2021-07-16T05:30:58.702095Z","shell.execute_reply.started":"2021-07-16T05:30:57.701125Z","shell.execute_reply":"2021-07-16T05:30:58.701017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_derived.drop([cname_col, pname_col], axis=1, inplace=True)\n\nprint(df_derived.shape)\ndf_derived.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:58.703623Z","iopub.execute_input":"2021-07-16T05:30:58.704233Z","iopub.status.idle":"2021-07-16T05:30:59.108238Z","shell.execute_reply.started":"2021-07-16T05:30:58.704188Z","shell.execute_reply":"2021-07-16T05:30:59.107161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_derived_pivot = pd.pivot_table(df_derived, \n                                  values='correctedPrM', \n                                  index=[phone_col, ts_col],\n                                  columns=['svid'],\n                                  aggfunc=np.mean)\ndf_derived_pivot ","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:30:59.109821Z","iopub.execute_input":"2021-07-16T05:30:59.110383Z","iopub.status.idle":"2021-07-16T05:31:02.183376Z","shell.execute_reply.started":"2021-07-16T05:30:59.110337Z","shell.execute_reply":"2021-07-16T05:31:02.182120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_derived_pivot.columns = [f'svid_{x}' for x in df_derived_pivot.columns]\ndf_derived_pivot.columns","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:02.184821Z","iopub.execute_input":"2021-07-16T05:31:02.185127Z","iopub.status.idle":"2021-07-16T05:31:02.192745Z","shell.execute_reply.started":"2021-07-16T05:31:02.185097Z","shell.execute_reply":"2021-07-16T05:31:02.191595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_derived_pivot.reset_index(inplace=True)\ndf_derived_pivot","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:02.194591Z","iopub.execute_input":"2021-07-16T05:31:02.195215Z","iopub.status.idle":"2021-07-16T05:31:02.283402Z","shell.execute_reply.started":"2021-07-16T05:31:02.195168Z","shell.execute_reply":"2021-07-16T05:31:02.282363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_derived_pivot['sSinceGpsEpoch'] = df_derived_pivot[ts_col] // 1000\n\nprint(df_derived_pivot.shape)\ndf_derived_pivot.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:02.285074Z","iopub.execute_input":"2021-07-16T05:31:02.285678Z","iopub.status.idle":"2021-07-16T05:31:02.344223Z","shell.execute_reply.started":"2021-07-16T05:31:02.285632Z","shell.execute_reply":"2021-07-16T05:31:02.342828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.merge(df_derived_pivot, how='left', on=[phone_col, 'sSinceGpsEpoch'], suffixes=['', '_2'])\ndf.drop(['sSinceGpsEpoch', ts_col + '_2'], axis=1, inplace=True)\nprint(df.shape)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:02.345641Z","iopub.execute_input":"2021-07-16T05:31:02.345935Z","iopub.status.idle":"2021-07-16T05:31:02.620327Z","shell.execute_reply.started":"2021-07-16T05:31:02.345906Z","shell.execute_reply":"2021-07-16T05:31:02.619076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['d_lat'] = df['latDeg_gt'] - df[lat_col]\ndf['d_lon'] = df['lngDeg_gt'] - df[lon_col]\ndf[['d_lat', 'd_lon']].describe()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:02.622170Z","iopub.execute_input":"2021-07-16T05:31:02.622619Z","iopub.status.idle":"2021-07-16T05:31:02.686945Z","shell.execute_reply.started":"2021-07-16T05:31:02.622575Z","shell.execute_reply":"2021-07-16T05:31:02.685806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"derived_files = (data_dir / 'test').rglob('*_derived.csv')\ncols = [ts_col, 'svid', 'correctedPrM']\nderived_files ","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:02.688316Z","iopub.execute_input":"2021-07-16T05:31:02.688629Z","iopub.status.idle":"2021-07-16T05:31:02.696017Z","shell.execute_reply.started":"2021-07-16T05:31:02.688575Z","shell.execute_reply":"2021-07-16T05:31:02.694891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_list = []\nfor t in tqdm(derived_files, total=48):\n    derived = pd.read_csv(t)\n    derived['sSinceGpsEpoch'] = derived[ts_col] // 1000\n    derived.drop_duplicates(['sSinceGpsEpoch', 'svid'], inplace=True)\n    derived['correctedPrM'] = (derived['rawPrM'] + derived['satClkBiasM'] - derived['isrbM'] - \n                               derived['ionoDelayM'] - derived['tropoDelayM'])\n    df_list.append(derived[[cname_col, pname_col, 'sSinceGpsEpoch', 'svid', 'correctedPrM']])\n    ","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:02.697239Z","iopub.execute_input":"2021-07-16T05:31:02.697576Z","iopub.status.idle":"2021-07-16T05:31:20.458204Z","shell.execute_reply.started":"2021-07-16T05:31:02.697543Z","shell.execute_reply":"2021-07-16T05:31:20.457305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_derived = pd.concat(df_list, ignore_index=True)\ndf_derived","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:20.459426Z","iopub.execute_input":"2021-07-16T05:31:20.459840Z","iopub.status.idle":"2021-07-16T05:31:20.529989Z","shell.execute_reply.started":"2021-07-16T05:31:20.459808Z","shell.execute_reply":"2021-07-16T05:31:20.528873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_derived[phone_col] = df_derived[cname_col] + '_' + df_derived[pname_col]\ndf_derived.drop([cname_col, pname_col], axis=1, inplace=True)\ndf_derived","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:20.531286Z","iopub.execute_input":"2021-07-16T05:31:20.531777Z","iopub.status.idle":"2021-07-16T05:31:21.465218Z","shell.execute_reply.started":"2021-07-16T05:31:20.531741Z","shell.execute_reply":"2021-07-16T05:31:21.464421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_derived_pivot = pd.pivot_table(df_derived, \n                                  values='correctedPrM', \n                                  index=[phone_col, 'sSinceGpsEpoch'],\n                                  columns=['svid'],\n                                  aggfunc=np.mean)\ndf_derived_pivot","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:21.466425Z","iopub.execute_input":"2021-07-16T05:31:21.466854Z","iopub.status.idle":"2021-07-16T05:31:23.066891Z","shell.execute_reply.started":"2021-07-16T05:31:21.466822Z","shell.execute_reply":"2021-07-16T05:31:23.065593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_derived_pivot.columns = [f'svid_{x}' for x in df_derived_pivot.columns]\ndf_derived_pivot.reset_index(inplace=True)\ndf_derived_pivot","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:23.068252Z","iopub.execute_input":"2021-07-16T05:31:23.068552Z","iopub.status.idle":"2021-07-16T05:31:23.151849Z","shell.execute_reply.started":"2021-07-16T05:31:23.068524Z","shell.execute_reply":"2021-07-16T05:31:23.150650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_tst = df_tst.merge(df_derived_pivot, how='left', \n                      on=[phone_col, 'sSinceGpsEpoch']).drop(['sSinceGpsEpoch'], axis=1)\nprint(df_tst.shape)\ndf_tst.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:23.153463Z","iopub.execute_input":"2021-07-16T05:31:23.153778Z","iopub.status.idle":"2021-07-16T05:31:23.348615Z","shell.execute_reply.started":"2021-07-16T05:31:23.153747Z","shell.execute_reply":"2021-07-16T05:31:23.347374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_tst.describe()","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:23.350433Z","iopub.execute_input":"2021-07-16T05:31:23.350887Z","iopub.status.idle":"2021-07-16T05:31:23.590722Z","shell.execute_reply.started":"2021-07-16T05:31:23.350837Z","shell.execute_reply":"2021-07-16T05:31:23.589910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feature_cols = [x for x in df_tst.columns if x not in [phone_col, ts_col]]\ntarget_cols = ['d_lat', 'd_lon']\ninput_dim = len(feature_cols)\noutput_dim = len(target_cols)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:23.591900Z","iopub.execute_input":"2021-07-16T05:31:23.592313Z","iopub.status.idle":"2021-07-16T05:31:23.597006Z","shell.execute_reply.started":"2021-07-16T05:31:23.592261Z","shell.execute_reply":"2021-07-16T05:31:23.595938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feature_cols ","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:23.598406Z","iopub.execute_input":"2021-07-16T05:31:23.598804Z","iopub.status.idle":"2021-07-16T05:31:23.616057Z","shell.execute_reply.started":"2021-07-16T05:31:23.598774Z","shell.execute_reply":"2021-07-16T05:31:23.614968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scaler = StandardScaler()\nlabel_scaler = StandardScaler()\nscaler.fit(pd.concat([df[feature_cols], df_tst[feature_cols]], axis=0).fillna(0).values)\nX = scaler.transform(df[feature_cols].fillna(0).values)\nX_tst = scaler.transform(df_tst[feature_cols].fillna(0).values)\nY = label_scaler.fit_transform(df[target_cols].values)\nprint(X.shape, Y.shape, X_tst.shape)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:23.618136Z","iopub.execute_input":"2021-07-16T05:31:23.618718Z","iopub.status.idle":"2021-07-16T05:31:24.190181Z","shell.execute_reply.started":"2021-07-16T05:31:23.618669Z","shell.execute_reply":"2021-07-16T05:31:24.188926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def scheduler(epoch, lr, warmup=5):\n    if epoch < warmup:\n        return lr * 1.5\n    else:\n        return lr * tf.math.exp(-.1) #epoch毎に減衰させている。","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:24.191727Z","iopub.execute_input":"2021-07-16T05:31:24.192032Z","iopub.status.idle":"2021-07-16T05:31:24.197205Z","shell.execute_reply.started":"2021-07-16T05:31:24.192003Z","shell.execute_reply":"2021-07-16T05:31:24.196384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import optuna \nimport optuna.integration.lightgbm as lgbo\nimport lightgbm as lgb\n'''\nparams = { 'objective': 'mse', 'metric': 'mse' }\nY = pd.DataFrame(Y,columns={'data','data2'})\n\nlgb_train1 = lgb.Dataset(X, Y.data)\nlgb_valid1 = lgb.Dataset(X, Y.data)\nmodel1 = lgbo.train(params, lgb_train1, valid_sets=[lgb_valid1], verbose_eval=False, num_boost_round=100, early_stopping_rounds=5) \nmodel1.params[\"learning_rate\"] = 0.01\nmodel1.params[\"early_stopping_round\"] = 100\nmodel1.params[\"num_iterations\"] = 8000\nmodel1.params\n'''","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:24.198409Z","iopub.execute_input":"2021-07-16T05:31:24.198815Z","iopub.status.idle":"2021-07-16T05:31:26.293782Z","shell.execute_reply.started":"2021-07-16T05:31:24.198784Z","shell.execute_reply":"2021-07-16T05:31:26.292625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params1= {'objective': 'mse',\n 'metric': 'l2',\n 'feature_pre_filter': False,\n 'lambda_l1': 0.0,\n 'lambda_l2': 0.0,\n 'num_leaves': 253,\n 'feature_fraction': 0.8999999999999999,\n 'bagging_fraction': 0.8540227553324429,\n 'bagging_freq': 2,\n 'min_child_samples': 5,\n 'num_iterations': 8000,\n 'early_stopping_round': 100,\n 'learning_rate': 0.01}","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:26.295363Z","iopub.execute_input":"2021-07-16T05:31:26.295669Z","iopub.status.idle":"2021-07-16T05:31:26.301368Z","shell.execute_reply.started":"2021-07-16T05:31:26.295639Z","shell.execute_reply":"2021-07-16T05:31:26.300579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nlgb_train2 = lgb.Dataset(X, Y.data2)\nlgb_valid2 = lgb.Dataset(X, Y.data2)\nmodel2 = lgbo.train(params, lgb_train2, valid_sets=[lgb_valid2], verbose_eval=False, num_boost_round=100, early_stopping_rounds=5) \nmodel2.params[\"learning_rate\"] = 0.01\nmodel2.params[\"early_stopping_round\"] = 100\nmodel2.params[\"num_iterations\"] = 8000\nmodel2.params\n'''","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:26.302521Z","iopub.execute_input":"2021-07-16T05:31:26.302948Z","iopub.status.idle":"2021-07-16T05:31:26.319236Z","shell.execute_reply.started":"2021-07-16T05:31:26.302918Z","shell.execute_reply":"2021-07-16T05:31:26.318479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params2 = {'objective': 'mse',\n 'metric': 'l2',\n 'feature_pre_filter': False,\n 'lambda_l1': 0.0,\n 'lambda_l2': 0.0,\n 'num_leaves': 242,\n 'feature_fraction': 0.8839999999999999,\n 'bagging_fraction': 0.9480235884535055,\n 'bagging_freq': 3,\n 'min_child_samples': 5,\n 'num_iterations': 8000,\n 'early_stopping_round': 100,\n 'learning_rate': 0.01}","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:26.320404Z","iopub.execute_input":"2021-07-16T05:31:26.320803Z","iopub.status.idle":"2021-07-16T05:31:26.335221Z","shell.execute_reply.started":"2021-07-16T05:31:26.320773Z","shell.execute_reply":"2021-07-16T05:31:26.334229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params = {'objective': 'mse',\n 'metric': 'mse',\n 'num_iterations': 8000,\n 'early_stopping_round': 100,\n 'learning_rate': 0.001}","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:26.336536Z","iopub.execute_input":"2021-07-16T05:31:26.337022Z","iopub.status.idle":"2021-07-16T05:31:26.351037Z","shell.execute_reply.started":"2021-07-16T05:31:26.336987Z","shell.execute_reply":"2021-07-16T05:31:26.350115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfrom sklearn.multioutput import MultiOutputRegressor\nimport lightgbm as lgb\n\nparams={'learning_rate': 0.02, #0.01\n        'objective':'mae', \n        'metric':'mae',\n        'num_leaves': 9, #9 @@\n        'verbose': 0,\n        'bagging_fraction': 0.8, #0.7\n        'feature_fraction': 0.8 #0.7\n       }\nreg = MultiOutputRegressor(lgb.LGBMRegressor(**params, n_estimators=2000))\n\ncv = KFold(n_splits=n_fold, shuffle=True, random_state=seed)\n\nP = np.zeros_like(Y, dtype=float)\nP_tst = np.zeros((X_tst.shape[0], output_dim), dtype=float)\n#Y = pd.DataFrame(Y,columns={'data','data2'})\nfor i, (i_trn, i_val) in enumerate(cv.split(X), 1):\n    print(f'Training for CV #{i}')\n    #lgb_train1 = lgb.Dataset(X[i_trn], Y.data[i_trn])\n    #lgb_valid1 = lgb.Dataset(X[i_val], Y.data[i_val])\n    \n    #lgb_train2 = lgb.Dataset(X[i_trn], Y.data2[i_trn])\n    #lgb_valid2 = lgb.Dataset(X[i_val], Y.data2[i_val])\n    \n    reg.fit(X[i_trn], Y[i_trn])\n    #model1 = lgb.train(params1, lgb_train1, valid_sets=[lgb_valid1], verbose_eval=100)\n    #model2 = lgb.train(params2, lgb_train2, valid_sets=[lgb_valid2], verbose_eval=100)\n    \n    #a =model1.predict(X[i_val])\n    #b =model2.predict(X[i_val])\n    #tt = pd.DataFrame(columns=['a','b'],index=range(len(a)))\n    #tt.a = a\n    #tt.b = b\n   \n    tt = reg.predict(X[i_val])\n    P[i_val] = label_scaler.inverse_transform(tt)\n    \n    \n    #a =model1.predict(X_tst)\n    #b =model2.predict(X_tst)\n    #tt = pd.DataFrame(columns=['a','b'],index=range(len(a)))\n    #tt.a = a\n    #tt.b = b\n    tt = reg.predict(X_tst)\n\n    P_tst += label_scaler.inverse_transform(tt) / n_fold\n    \n    distance_i = calc_haversine(df.latDeg_gt.values[i_val], \n                                df.lngDeg_gt.values[i_val], \n                                P[i_val, 0] + df.latDeg.values[i_val], \n                                P[i_val, 1] + df.lngDeg.values[i_val]).mean()\n    print(f'CV #{i}: {np.percentile(distance_i, [50, 95])}')\n","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:31:26.352652Z","iopub.execute_input":"2021-07-16T05:31:26.353301Z","iopub.status.idle":"2021-07-16T05:34:51.636945Z","shell.execute_reply.started":"2021-07-16T05:31:26.353233Z","shell.execute_reply":"2021-07-16T05:34:51.635897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\ncv = KFold(n_splits=n_fold, shuffle=True, random_state=seed)\n\nP = np.zeros_like(Y, dtype=float)\nP_tst = np.zeros((X_tst.shape[0], output_dim), dtype=float)\nY = pd.DataFrame(Y,columns={'data','data2'})\nfor i, (i_trn, i_val) in enumerate(cv.split(X), 1):\n    print(f'Training for CV #{i}')\n    lgb_train1 = lgb.Dataset(X[i_trn], Y.data[i_trn])\n    lgb_valid1 = lgb.Dataset(X[i_val], Y.data[i_val])\n    \n    lgb_train2 = lgb.Dataset(X[i_trn], Y.data2[i_trn])\n    lgb_valid2 = lgb.Dataset(X[i_val], Y.data2[i_val])\n    \n\n    model1 = lgb.train(params1, lgb_train1, valid_sets=[lgb_valid1], verbose_eval=100)\n    model2 = lgb.train(params2, lgb_train2, valid_sets=[lgb_valid2], verbose_eval=100)\n    \n    a =model1.predict(X[i_val])\n    b =model2.predict(X[i_val])\n    tt = pd.DataFrame(columns=['a','b'],index=range(len(a)))\n    tt.a = a\n    tt.b = b\n   \n    P[i_val] = label_scaler.inverse_transform(tt)\n    \n    \n    a =model1.predict(X_tst)\n    b =model2.predict(X_tst)\n    tt = pd.DataFrame(columns=['a','b'],index=range(len(a)))\n    tt.a = a\n    tt.b = b\n\n    P_tst += label_scaler.inverse_transform(tt) / n_fold\n    \n    distance_i = calc_haversine(df.latDeg_gt.values[i_val], \n                                df.lngDeg_gt.values[i_val], \n                                P[i_val, 0] + df.latDeg.values[i_val], \n                                P[i_val, 1] + df.lngDeg.values[i_val]).mean()\n    print(f'CV #{i}: {np.percentile(distance_i, [50, 95])}')\n    '''","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:34:51.638610Z","iopub.execute_input":"2021-07-16T05:34:51.638983Z","iopub.status.idle":"2021-07-16T05:34:51.648046Z","shell.execute_reply.started":"2021-07-16T05:34:51.638950Z","shell.execute_reply":"2021-07-16T05:34:51.646779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(P.mean(axis=0), P_tst.mean(axis=0))\nnp.savetxt(predict_val_file, P, delimiter=',', fmt='%.6f')\nnp.savetxt(predict_tst_file, P_tst, delimiter=',', fmt='%.6f')","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:34:51.649766Z","iopub.execute_input":"2021-07-16T05:34:51.650134Z","iopub.status.idle":"2021-07-16T05:34:52.423454Z","shell.execute_reply.started":"2021-07-16T05:34:51.650101Z","shell.execute_reply":"2021-07-16T05:34:52.422586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"distance = calc_haversine(df.latDeg_gt, df.lngDeg_gt, P[:, 0] + df.latDeg, P[:, 1] + df.lngDeg)\nprint(f'CV All: {np.percentile(distance, [50, 95])}')","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:34:52.424828Z","iopub.execute_input":"2021-07-16T05:34:52.425240Z","iopub.status.idle":"2021-07-16T05:34:52.452967Z","shell.execute_reply.started":"2021-07-16T05:34:52.425195Z","shell.execute_reply":"2021-07-16T05:34:52.452166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.sort_values([phone_col, ts_col], inplace=True)\ndf_smoothed = df.copy()\ndf_smoothed[lat_col] = df[lat_col] + P[:, 0]\ndf_smoothed[lon_col] = df[lon_col] + P[:, 1]\ndf_smoothed = apply_kf_smoothing(df_smoothed)\ndistance = calc_haversine(df_smoothed.latDeg_gt, df_smoothed.lngDeg_gt, df_smoothed.latDeg, df_smoothed.lngDeg)\nprint(f'CV All (smoothed): {np.percentile(distance, [50, 95])}')","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:34:52.454190Z","iopub.execute_input":"2021-07-16T05:34:52.454664Z","iopub.status.idle":"2021-07-16T05:35:43.350350Z","shell.execute_reply.started":"2021-07-16T05:34:52.454629Z","shell.execute_reply":"2021-07-16T05:35:43.349216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"distance_tst = calc_haversine(df_tst.latDeg, df_tst.lngDeg, P_tst[:, 0] + df_tst.latDeg, P_tst[:, 1] + df_tst.lngDeg)\nprint(f'CV All: {np.percentile(distance_tst, [50, 95])}')","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:35:43.352358Z","iopub.execute_input":"2021-07-16T05:35:43.353050Z","iopub.status.idle":"2021-07-16T05:35:43.377901Z","shell.execute_reply.started":"2021-07-16T05:35:43.353002Z","shell.execute_reply":"2021-07-16T05:35:43.376819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"distance_tst","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:35:43.379532Z","iopub.execute_input":"2021-07-16T05:35:43.379967Z","iopub.status.idle":"2021-07-16T05:35:43.388546Z","shell.execute_reply.started":"2021-07-16T05:35:43.379921Z","shell.execute_reply":"2021-07-16T05:35:43.387740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_tst.sort_values([phone_col, ts_col], inplace=True)\ndf_tst_smoothed = df_tst.copy()\ndf_tst_smoothed[lat_col] = df_tst_smoothed[lat_col] + P_tst[:, 0]\ndf_tst_smoothed[lon_col] = df_tst_smoothed[lon_col] + P_tst[:, 1]\ndf_tst_smoothed","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:35:43.389912Z","iopub.execute_input":"2021-07-16T05:35:43.390435Z","iopub.status.idle":"2021-07-16T05:35:43.542333Z","shell.execute_reply.started":"2021-07-16T05:35:43.390391Z","shell.execute_reply":"2021-07-16T05:35:43.541497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_tst_smoothed = apply_kf_smoothing(df_tst_smoothed)\ndf_tst_smoothed","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:35:43.543705Z","iopub.execute_input":"2021-07-16T05:35:43.544185Z","iopub.status.idle":"2021-07-16T05:36:18.745409Z","shell.execute_reply.started":"2021-07-16T05:35:43.544152Z","shell.execute_reply":"2021-07-16T05:36:18.744486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"distance_tst = calc_haversine(df_tst.latDeg, df_tst.lngDeg, df_tst_smoothed.latDeg, df_tst_smoothed.lngDeg)\nprint(f'CV All (smoothed): {np.percentile(distance_tst, [50, 95])}')","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:36:18.746687Z","iopub.execute_input":"2021-07-16T05:36:18.747115Z","iopub.status.idle":"2021-07-16T05:36:18.769976Z","shell.execute_reply.started":"2021-07-16T05:36:18.747082Z","shell.execute_reply":"2021-07-16T05:36:18.768749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_tst_smoothed[[phone_col, ts_col, lat_col, lon_col]].to_csv(submission_file, index=False)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:36:18.771605Z","iopub.execute_input":"2021-07-16T05:36:18.771988Z","iopub.status.idle":"2021-07-16T05:36:19.468264Z","shell.execute_reply.started":"2021-07-16T05:36:18.771950Z","shell.execute_reply":"2021-07-16T05:36:19.467185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = df_tst_smoothed[[phone_col, ts_col, lat_col, lon_col]]","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:36:19.470328Z","iopub.execute_input":"2021-07-16T05:36:19.470751Z","iopub.status.idle":"2021-07-16T05:36:19.478580Z","shell.execute_reply.started":"2021-07-16T05:36:19.470704Z","shell.execute_reply":"2021-07-16T05:36:19.477546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_removedevice(input_df: pd.DataFrame, divece: str) -> pd.DataFrame:\n    input_df['index'] = input_df.index\n    input_df = input_df.sort_values('millisSinceGpsEpoch')\n    input_df.index = input_df['millisSinceGpsEpoch'].values\n\n    output_df = pd.DataFrame() \n    for _, subdf in input_df.groupby('collectionName'):\n\n        phones = subdf['phoneName'].unique()\n\n        if (len(phones) == 1) or (not divece in phones):\n            output_df = pd.concat([output_df, subdf])\n            continue\n\n        origin_df = subdf.copy()\n        \n        _index = subdf['phoneName']==divece\n        subdf.loc[_index, 'latDeg'] = np.nan\n        subdf.loc[_index, 'lngDeg'] = np.nan\n        subdf = subdf.interpolate(method='index', limit_area='inside')\n\n        _index = subdf['latDeg'].isnull()\n        subdf.loc[_index, 'latDeg'] = origin_df.loc[_index, 'latDeg'].values\n        subdf.loc[_index, 'lngDeg'] = origin_df.loc[_index, 'lngDeg'].values\n\n        output_df = pd.concat([output_df, subdf])\n\n    output_df.index = output_df['index'].values\n    output_df = output_df.sort_index()\n\n    del output_df['index']\n    \n    return output_df","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:36:19.480097Z","iopub.execute_input":"2021-07-16T05:36:19.480480Z","iopub.status.idle":"2021-07-16T05:36:19.493094Z","shell.execute_reply.started":"2021-07-16T05:36:19.480446Z","shell.execute_reply":"2021-07-16T05:36:19.492192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission['collectionName'] = submission['phone'].map(lambda x: x.split('_')[0])\nsubmission['phoneName'] = submission['phone'].map(lambda x: x.split('_')[1])\nsubmission = get_removedevice(submission, 'SamsungS20Ultra')\n\nsubmission = submission.drop(columns=['collectionName', 'phoneName'], axis=1)\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2021-07-16T05:36:19.494555Z","iopub.execute_input":"2021-07-16T05:36:19.494882Z","iopub.status.idle":"2021-07-16T05:36:20.713518Z","shell.execute_reply.started":"2021-07-16T05:36:19.494853Z","shell.execute_reply":"2021-07-16T05:36:20.712240Z"},"trusted":true},"execution_count":null,"outputs":[]}]}