{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":60095,"databundleVersionId":6542333,"sourceType":"competition"}],"dockerImageVersionId":30636,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm, trange\n\nINPUT_PATH = '/kaggle/input/smartphone-decimeter-2023/sdc2023/'\n\ndf_sample_trail_gnss = pd.read_csv(INPUT_PATH + \"train/2020-06-25-00-34-us-ca-mtv-sb-101/pixel4/device_gnss.csv\")\ndf_sample_trail_gt = pd.read_csv(INPUT_PATH + \"train/2020-06-25-00-34-us-ca-mtv-sb-101/pixel4/ground_truth.csv\")\nsample_submission = pd.read_csv(INPUT_PATH + \"sample_submission.csv\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-01-23T01:11:01.261500Z","iopub.execute_input":"2024-01-23T01:11:01.262225Z","iopub.status.idle":"2024-01-23T01:11:01.817485Z","shell.execute_reply.started":"2024-01-23T01:11:01.262176Z","shell.execute_reply":"2024-01-23T01:11:01.816670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this challenge, our only output is a **longitude** and **latitude** for each millisecond-accurate timestep.","metadata":{}},{"cell_type":"code","source":"sample_submission.head()","metadata":{"execution":{"iopub.status.busy":"2024-01-23T01:00:56.497989Z","iopub.execute_input":"2024-01-23T01:00:56.498985Z","iopub.status.idle":"2024-01-23T01:00:56.510292Z","shell.execute_reply.started":"2024-01-23T01:00:56.498934Z","shell.execute_reply":"2024-01-23T01:00:56.509166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Recall that our training data has two components:\n1. device_gnss (Noisy data from phone sensors)\n2. ground_truth (high-quality, accurate data from laser rangefinders + dual-channel GPS sensors).\n\nIn our testing phase, we're only able to make predictions off of device_gnss data. Our goal is thus to predict the longitude and latitude of *ground_truth* given *device_gnss*.\n\nNote that our input data and desired output data is both timeseries. As shown below, the device_gnss data spans the entire duration of a test drive (roughly 20 minutes). Our prediction target has the same duration as well (albeit starting one second later):","metadata":{"execution":{"iopub.status.busy":"2024-01-23T01:05:03.013349Z","iopub.execute_input":"2024-01-23T01:05:03.014261Z","iopub.status.idle":"2024-01-23T01:05:03.021714Z","shell.execute_reply.started":"2024-01-23T01:05:03.014219Z","shell.execute_reply":"2024-01-23T01:05:03.020387Z"}}},{"cell_type":"code","source":"from datetime import datetime\n\ndef utc_to_human_readable(utcTime):\n    utc_datetime_str = datetime.fromtimestamp(utcTime / 1e3)\n    return datetime.strftime(utc_datetime_str, '%Y-%m-%d | %H:%M:%S')\n\nprint(\"Duration of input data (s):\", \n      (df_sample_trail_gnss[\"utcTimeMillis\"].max() - df_sample_trail_gnss[\"utcTimeMillis\"].min()) * 1e-3)\nprint(\"Starting from\", utc_to_human_readable(df_sample_trail_gnss[\"utcTimeMillis\"].min()),\n      \"to\", utc_to_human_readable(df_sample_trail_gnss[\"utcTimeMillis\"].max()))\nlabels = df_sample_trail_gt[[\"LatitudeDegrees\", \"LongitudeDegrees\", \"UnixTimeMillis\"]]\nprint(\"Duration of target data (s):\", \n      (labels[\"UnixTimeMillis\"].max() - labels[\"UnixTimeMillis\"].min()) * 1e-3)\nprint(\"Starting from\", utc_to_human_readable(labels[\"UnixTimeMillis\"].min()),\n      \"to\", utc_to_human_readable(labels[\"UnixTimeMillis\"].max()))","metadata":{"execution":{"iopub.status.busy":"2024-01-23T01:20:12.296021Z","iopub.execute_input":"2024-01-23T01:20:12.296396Z","iopub.status.idle":"2024-01-23T01:20:12.307344Z","shell.execute_reply.started":"2024-01-23T01:20:12.296363Z","shell.execute_reply":"2024-01-23T01:20:12.306366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The frequency at which we collect the data is once per second, or 1Hz. You might be wondering, then, why our input trajectory (gnss data) then has significantly more data points than this low frequency seems to suggest. This is due to the fact that we receive data from many satellites from different GPS constellations at every timestep:\n\nNote that the Svid does not uniquely identify a satellite; distinct satellites between different constellations can have the same Svid. ","metadata":{}},{"cell_type":"code","source":"print(\"Num input data timesteps:\", len(df_sample_trail_gnss[\"utcTimeMillis\"].unique()))\ndf_sample_trail_gnss[df_sample_trail_gnss[\"TimeNanos\"] == 1047929361000000][[\"Svid\", \"State\", \"SvVelocityXEcefMetersPerSecond\", \"ConstellationType\"]]","metadata":{"execution":{"iopub.status.busy":"2024-01-23T01:33:16.849130Z","iopub.execute_input":"2024-01-23T01:33:16.850002Z","iopub.status.idle":"2024-01-23T01:33:16.866652Z","shell.execute_reply.started":"2024-01-23T01:33:16.849960Z","shell.execute_reply":"2024-01-23T01:33:16.865695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our ground truth data, fortunately, has just one row of data per timestep.\n\nTo recap:\n- GNSS input, shape (timesteps, # satellite signals, data_dim). At each timestep, we receive data from a varying number of satellites.\n- Target, shape (timesteps, 2) - Longitude and Latitude for each timestep.\n- Ground truth, shape (timesteps, data_dim) - auxiliary data we can use in train time that is unavailable at test time.","metadata":{}},{"cell_type":"markdown","source":"Let's take a look at the trends within our data:","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nsns.pairplot(df_sample_trail_gnss[[\"SvClockDriftMetersPerSecond\", \"AccumulatedDeltaRangeMeters\",\n                                  \"IonosphericDelayMeters\", \"TroposphericDelayMeters\"]])","metadata":{"execution":{"iopub.status.busy":"2024-01-23T01:42:02.116905Z","iopub.execute_input":"2024-01-23T01:42:02.117889Z","iopub.status.idle":"2024-01-23T01:42:09.971042Z","shell.execute_reply.started":"2024-01-23T01:42:02.117851Z","shell.execute_reply":"2024-01-23T01:42:09.970139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It looks like there are some interesting relationships between the Ionospheric delay and the Tropospheric delay.","metadata":{}},{"cell_type":"markdown","source":"One challenge we will face throughout this competition is the fact that a lot of the data is redundant. Likely only a few columns will contribute to the majority of the accuracy of our model. Let's drop the NaN data first:","metadata":{}},{"cell_type":"code","source":"df_sample_trail_gnss.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2024-01-23T01:44:55.510704Z","iopub.execute_input":"2024-01-23T01:44:55.511856Z","iopub.status.idle":"2024-01-23T01:44:55.542174Z","shell.execute_reply.started":"2024-01-23T01:44:55.511808Z","shell.execute_reply":"2024-01-23T01:44:55.541280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can be comfortable dropping all the data that's completely missing, but the mostly complete data looks relatively important. We can use [imputation techniques](https://www.kaggle.com/code/azminetoushikwasi/all-imputation-techniques-with-pros-and-cons) to best fill in the missing data with synthetic data.","metadata":{}},{"cell_type":"markdown","source":"Now let's take a look at our data at time-of-inference.","metadata":{}},{"cell_type":"code","source":"df_test_trail_gnss = pd.read_csv(INPUT_PATH + \"test/2020-12-11-19-30-us-ca-mtv-e/pixel4xl/device_gnss.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-01-23T01:09:17.927761Z","iopub.execute_input":"2024-01-23T01:09:17.928628Z","iopub.status.idle":"2024-01-23T01:09:18.643127Z","shell.execute_reply.started":"2024-01-23T01:09:17.928591Z","shell.execute_reply":"2024-01-23T01:09:18.642074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test_trail_gnss","metadata":{"execution":{"iopub.status.busy":"2024-01-23T01:09:24.817091Z","iopub.execute_input":"2024-01-23T01:09:24.817983Z","iopub.status.idle":"2024-01-23T01:09:24.858999Z","shell.execute_reply.started":"2024-01-23T01:09:24.817947Z","shell.execute_reply":"2024-01-23T01:09:24.857992Z"},"trusted":true},"execution_count":null,"outputs":[]}]}