{
  "id": 241874,
  "title": "Merging derived files dilemma",
  "url": "/competitions/google-smartphone-decimeter-challenge/discussion/241874",
  "author_name": "",
  "post_date": "2021-05-26T12:50:00.980761100Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am dealing with a dilemma regarding the derived files.</p>\n<p>There are way more rows in the derived files than in the baseline train, and also not every timestamp in the train dataset appears in the derived dataset.</p>\n<p>Possible solutions:</p>\n<ol>\n<li><p>dropping duplicate lines (losing data 😟) from the derived datasets, filling missing rows with some fillna method based on neighbors.</p></li>\n<li><p>melting the rows with the same timestamp and same phone (not losing data 🤓) and filling the missing timestamps.</p></li>\n<li><p>skipping the derived files altogether.</p></li>\n</ol>\n<p>Right now my instinct tells me to let go for now and move on to the next files (like GNSS logs).</p>\n<p>Will be happy to hear about your chosen solution (and other possible solutions) for this problem.</p>",
  "messages": [
    {
      "id": "1323790",
      "postDate": "05/26/2021 12:50:00",
      "content": "<p>I am dealing with a dilemma regarding the derived files.</p>\n<p>There are way more rows in the derived files than in the baseline train, and also not every timestamp in the train dataset appears in the derived dataset.</p>\n<p>Possible solutions:</p>\n<ol>\n<li><p>dropping duplicate lines (losing data 😟) from the derived datasets, filling missing rows with some fillna method based on neighbors.</p></li>\n<li><p>melting the rows with the same timestamp and same phone (not losing data 🤓) and filling the missing timestamps.</p></li>\n<li><p>skipping the derived files altogether.</p></li>\n</ol>\n<p>Right now my instinct tells me to let go for now and move on to the next files (like GNSS logs).</p>\n<p>Will be happy to hear about your chosen solution (and other possible solutions) for this problem.</p>",
      "rawMarkdown": "I am dealing with a dilemma regarding the derived files.\n\nThere are way more rows in the derived files than in the baseline train, and also not every timestamp in the train dataset appears in the derived dataset.\n\nPossible solutions:\n\n1. dropping duplicate lines (losing data 😟) from the derived datasets, filling missing rows with some fillna method based on neighbors.\n\n2. melting the rows with the same timestamp and same phone (not losing data 🤓) and filling the missing timestamps.\n\n3. skipping the derived files altogether.\n\nRight now my instinct tells me to let go for now and move on to the next files (like GNSS logs).\n\nWill be happy to hear about your chosen solution (and other possible solutions) for this problem.",
      "votes": null
    },
    {
      "id": "1324251",
      "postDate": "05/26/2021 18:29:12",
      "content": "<p>Hello, I haven't yet looked into incorporating the derived files nor have I looked too deeply within them but I think the multiple rows are relating to different satellite readings at a given time.</p>",
      "rawMarkdown": "Hello, I haven't yet looked into incorporating the derived files nor have I looked too deeply within them but I think the multiple rows are relating to different satellite readings at a given time.",
      "votes": null
    },
    {
      "id": "1324297",
      "postDate": "05/26/2021 19:23:35",
      "content": "<p>Yes you are right, the question is if and how to incorperate them into a train file</p>",
      "rawMarkdown": "Yes you are right, the question is if and how to incorperate them into a train file",
      "votes": null
    },
    {
      "id": "1324469",
      "postDate": "05/27/2021 01:37:01",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/avivlevi815\" target=\"_blank\">@avivlevi815</a>, currently I'm working on reproducing the baseline result provided by dataset provider from derived files. Not success yet 😓, but I do have some ideas to merge the information from derive file to the baseline/groundtruth file. Basically, we can just merge it base on the nearby timestamp by using merge_asof.</p>\n<pre><code>df_sample_trails_merged_SL = pd.merge_asof(df_sample_trails_gt.sort_values('millisSinceGpsEpoch'), df_sample_trails_estimate.sort_values('millisSinceGpsEpoch'), \n                                           on=\"millisSinceGpsEpoch\", by=[\"collectionName\", \"phoneName\"], direction='nearest',tolerance=100000, suffixes=('_truth', '_pred'))\ndf_sample_trails_merged_SL = df_sample_trails_merged_SL.sort_values(by=[\"collectionName\", \"phoneName\", \"millisSinceGpsEpoch\"], ignore_index=True)\n</code></pre>\n<p>Some relative code can be found in evaluation and submition subsections in <a href=\"https://www.kaggle.com/foreveryoung/least-squares-solution-from-gnss-derived-data\" target=\"_blank\">this nodebook</a>.</p>\n<p>Please let me know if you have any suggestion to reproduce the baseline result. I think it would be super helpful to help us understand the features and build the model on the top of it.</p>",
      "rawMarkdown": "Hi, @avivlevi815, currently I'm working on reproducing the baseline result provided by dataset provider from derived files. Not success yet 😓, but I do have some ideas to merge the information from derive file to the baseline/groundtruth file. Basically, we can just merge it base on the nearby timestamp by using merge_asof.\n\n```python\ndf_sample_trails_merged_SL = pd.merge_asof(df_sample_trails_gt.sort_values('millisSinceGpsEpoch'), df_sample_trails_estimate.sort_values('millisSinceGpsEpoch'), \n                                           on=\"millisSinceGpsEpoch\", by=[\"collectionName\", \"phoneName\"], direction='nearest',tolerance=100000, suffixes=('_truth', '_pred'))\ndf_sample_trails_merged_SL = df_sample_trails_merged_SL.sort_values(by=[\"collectionName\", \"phoneName\", \"millisSinceGpsEpoch\"], ignore_index=True)\n```\n\nSome relative code can be found in evaluation and submition subsections in [this nodebook](https://www.kaggle.com/foreveryoung/least-squares-solution-from-gnss-derived-data).\n\nPlease let me know if you have any suggestion to reproduce the baseline result. I think it would be super helpful to help us understand the features and build the model on the top of it.",
      "votes": null
    },
    {
      "id": "1324580",
      "postDate": "05/27/2021 04:39:57",
      "content": "<p>Thank you YangLiu!<br>\nI will try it.<br>\nAbout the baseline result - try and submit the small antenna's lat&amp;lng features, take them from the test baseline file and put them in sample_submission - submit this, it should get you to ~7 in the LB</p>",
      "rawMarkdown": "Thank you YangLiu!\nI will try it.\nAbout the baseline result - try and submit the small antenna's lat&lng features, take them from the test baseline file and put them in sample_submission - submit this, it should get you to ~7 in the LB",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1324251,
      "author_name": "martinf1",
      "author_url": "",
      "post_date": "05/26/2021 18:29:12",
      "content": "<p>Hello, I haven't yet looked into incorporating the derived files nor have I looked too deeply within them but I think the multiple rows are relating to different satellite readings at a given time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1324297,
          "author_name": "avivlevi815",
          "author_url": "",
          "post_date": "05/26/2021 19:23:35",
          "content": "<p>Yes you are right, the question is if and how to incorperate them into a train file</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1324469,
      "author_name": "foreveryoung",
      "author_url": "",
      "post_date": "05/27/2021 01:37:01",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/avivlevi815\" target=\"_blank\">@avivlevi815</a>, currently I'm working on reproducing the baseline result provided by dataset provider from derived files. Not success yet 😓, but I do have some ideas to merge the information from derive file to the baseline/groundtruth file. Basically, we can just merge it base on the nearby timestamp by using merge_asof.</p>\n<pre><code>df_sample_trails_merged_SL = pd.merge_asof(df_sample_trails_gt.sort_values('millisSinceGpsEpoch'), df_sample_trails_estimate.sort_values('millisSinceGpsEpoch'), \n                                           on=\"millisSinceGpsEpoch\", by=[\"collectionName\", \"phoneName\"], direction='nearest',tolerance=100000, suffixes=('_truth', '_pred'))\ndf_sample_trails_merged_SL = df_sample_trails_merged_SL.sort_values(by=[\"collectionName\", \"phoneName\", \"millisSinceGpsEpoch\"], ignore_index=True)\n</code></pre>\n<p>Some relative code can be found in evaluation and submition subsections in <a href=\"https://www.kaggle.com/foreveryoung/least-squares-solution-from-gnss-derived-data\" target=\"_blank\">this nodebook</a>.</p>\n<p>Please let me know if you have any suggestion to reproduce the baseline result. I think it would be super helpful to help us understand the features and build the model on the top of it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1324580,
          "author_name": "avivlevi815",
          "author_url": "",
          "post_date": "05/27/2021 04:39:57",
          "content": "<p>Thank you YangLiu!<br>\nI will try it.<br>\nAbout the baseline result - try and submit the small antenna's lat&amp;lng features, take them from the test baseline file and put them in sample_submission - submit this, it should get you to ~7 in the LB</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1323790": "I am dealing with a dilemma regarding the derived files.\n\nThere are way more rows in the derived files than in the baseline train, and also not every timestamp in the train dataset appears in the derived dataset.\n\nPossible solutions:\n\n1. dropping duplicate lines (losing data 😟) from the derived datasets, filling missing rows with some fillna method based on neighbors.\n\n2. melting the rows with the same timestamp and same phone (not losing data 🤓) and filling the missing timestamps.\n\n3. skipping the derived files altogether.\n\nRight now my instinct tells me to let go for now and move on to the next files (like GNSS logs).\n\nWill be happy to hear about your chosen solution (and other possible solutions) for this problem.",
    "1324251": "Hello, I haven't yet looked into incorporating the derived files nor have I looked too deeply within them but I think the multiple rows are relating to different satellite readings at a given time.",
    "1324297": "Yes you are right, the question is if and how to incorperate them into a train file",
    "1324469": "Hi, @avivlevi815, currently I'm working on reproducing the baseline result provided by dataset provider from derived files. Not success yet 😓, but I do have some ideas to merge the information from derive file to the baseline/groundtruth file. Basically, we can just merge it base on the nearby timestamp by using merge_asof.\n\n```python\ndf_sample_trails_merged_SL = pd.merge_asof(df_sample_trails_gt.sort_values('millisSinceGpsEpoch'), df_sample_trails_estimate.sort_values('millisSinceGpsEpoch'), \n                                           on=\"millisSinceGpsEpoch\", by=[\"collectionName\", \"phoneName\"], direction='nearest',tolerance=100000, suffixes=('_truth', '_pred'))\ndf_sample_trails_merged_SL = df_sample_trails_merged_SL.sort_values(by=[\"collectionName\", \"phoneName\", \"millisSinceGpsEpoch\"], ignore_index=True)\n```\n\nSome relative code can be found in evaluation and submition subsections in [this nodebook](https://www.kaggle.com/foreveryoung/least-squares-solution-from-gnss-derived-data).\n\nPlease let me know if you have any suggestion to reproduce the baseline result. I think it would be super helpful to help us understand the features and build the model on the top of it.",
    "1324580": "Thank you YangLiu!\nI will try it.\nAbout the baseline result - try and submit the small antenna's lat&lng features, take them from the test baseline file and put them in sample_submission - submit this, it should get you to ~7 in the LB"
  },
  "source": "meta"
}