{
  "id": 228199,
  "title": "What's a good CV strategy in this competition",
  "url": "/competitions/indoor-location-navigation/discussion/228199",
  "author_name": "",
  "post_date": "2021-03-23T19:10:55.703009200Z",
  "votes": 12,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all,<br>\nAfter analysing the data and understanding what is going on in this competition , I was trying to figure out what would be a good cv strategy in this competition . I have got really confused with the question as this is something very new to me .</p>\n<p>What I can think of right now :<br>\nFor every txt file that we have , sort the events by timestamp take the first few timestamp in the training set and the last few in the test set , that way we predict future in time , but I don't know whether this mimics the train set as there we will be given a whole txt file and we have to start predicting from time t=0 .</p>\n<p>It would be great help to know what you guys are using as a cv strategy</p>",
  "messages": [
    {
      "id": "1250087",
      "postDate": "03/23/2021 19:10:55",
      "content": "<p>Hi all,<br>\nAfter analysing the data and understanding what is going on in this competition , I was trying to figure out what would be a good cv strategy in this competition . I have got really confused with the question as this is something very new to me .</p>\n<p>What I can think of right now :<br>\nFor every txt file that we have , sort the events by timestamp take the first few timestamp in the training set and the last few in the test set , that way we predict future in time , but I don't know whether this mimics the train set as there we will be given a whole txt file and we have to start predicting from time t=0 .</p>\n<p>It would be great help to know what you guys are using as a cv strategy</p>",
      "rawMarkdown": "Hi all,\nAfter analysing the data and understanding what is going on in this competition , I was trying to figure out what would be a good cv strategy in this competition . I have got really confused with the question as this is something very new to me .\n\nWhat I can think of right now :\nFor every txt file that we have , sort the events by timestamp take the first few timestamp in the training set and the last few in the test set , that way we predict future in time , but I don't know whether this mimics the train set as there we will be given a whole txt file and we have to start predicting from time t=0 .\n\nIt would be great help to know what you guys are using as a cv strategy",
      "votes": null
    },
    {
      "id": "1250383",
      "postDate": "03/24/2021 01:46:24",
      "content": "<p>I picked certain amount of txt files to be a validation set. So, for each building I used x% of the txt files chosen randomly for validation. I think if you split train and test within a single txt file, your CV score will be better because there will be datas that are very close to each other. This \"improved CV\" won't equate to better model.</p>",
      "rawMarkdown": "I picked certain amount of txt files to be a validation set. So, for each building I used x% of the txt files chosen randomly for validation. I think if you split train and test within a single txt file, your CV score will be better because there will be datas that are very close to each other. This \"improved CV\" won't equate to better model.",
      "votes": null
    },
    {
      "id": "1250787",
      "postDate": "03/24/2021 09:20:01",
      "content": "<p>Thanks for reply <a href=\"https://www.kaggle.com/nooblife\" target=\"_blank\">@nooblife</a> your strategy makes more sense</p>",
      "rawMarkdown": "Thanks for reply @nooblife your strategy makes more sense",
      "votes": null
    },
    {
      "id": "1254202",
      "postDate": "03/27/2021 12:31:19",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/nooblife\" target=\"_blank\">@nooblife</a> for your insight. I haven't got very deep in the competition, but certain aspects of the competition might not be temporally related that much, in which case normal kFold will work fine. But for robust cv, your method sounds the best way</p>",
      "rawMarkdown": "Thanks @nooblife for your insight. I haven't got very deep in the competition, but certain aspects of the competition might not be temporally related that much, in which case normal kFold will work fine. But for robust cv, your method sounds the best way",
      "votes": null
    },
    {
      "id": "1255470",
      "postDate": "03/28/2021 20:43:34",
      "content": "<p><a href=\"https://www.kaggle.com/nooblife\" target=\"_blank\">@nooblife</a> This is really great insight. <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> have you given this a try? I was wondering the exact same thing and have yet to try and implement a strategy. </p>",
      "rawMarkdown": "nooblife This is really great insight. @tanulsingh077 have you given this a try? I was wondering the exact same thing and have yet to try and implement a strategy.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1250383,
      "author_name": "nooblife",
      "author_url": "",
      "post_date": "03/24/2021 01:46:24",
      "content": "<p>I picked certain amount of txt files to be a validation set. So, for each building I used x% of the txt files chosen randomly for validation. I think if you split train and test within a single txt file, your CV score will be better because there will be datas that are very close to each other. This \"improved CV\" won't equate to better model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1250787,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "03/24/2021 09:20:01",
          "content": "<p>Thanks for reply <a href=\"https://www.kaggle.com/nooblife\" target=\"_blank\">@nooblife</a> your strategy makes more sense</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1254202,
          "author_name": "suryajrrafl",
          "author_url": "",
          "post_date": "03/27/2021 12:31:19",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/nooblife\" target=\"_blank\">@nooblife</a> for your insight. I haven't got very deep in the competition, but certain aspects of the competition might not be temporally related that much, in which case normal kFold will work fine. But for robust cv, your method sounds the best way</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255470,
          "author_name": "crained",
          "author_url": "",
          "post_date": "03/28/2021 20:43:34",
          "content": "<p><a href=\"https://www.kaggle.com/nooblife\" target=\"_blank\">@nooblife</a> This is really great insight. <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> have you given this a try? I was wondering the exact same thing and have yet to try and implement a strategy. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1250087": "Hi all,\nAfter analysing the data and understanding what is going on in this competition , I was trying to figure out what would be a good cv strategy in this competition . I have got really confused with the question as this is something very new to me .\n\nWhat I can think of right now :\nFor every txt file that we have , sort the events by timestamp take the first few timestamp in the training set and the last few in the test set , that way we predict future in time , but I don't know whether this mimics the train set as there we will be given a whole txt file and we have to start predicting from time t=0 .\n\nIt would be great help to know what you guys are using as a cv strategy",
    "1250383": "I picked certain amount of txt files to be a validation set. So, for each building I used x% of the txt files chosen randomly for validation. I think if you split train and test within a single txt file, your CV score will be better because there will be datas that are very close to each other. This \"improved CV\" won't equate to better model.",
    "1250787": "Thanks for reply @nooblife your strategy makes more sense",
    "1254202": "Thanks @nooblife for your insight. I haven't got very deep in the competition, but certain aspects of the competition might not be temporally related that much, in which case normal kFold will work fine. But for robust cv, your method sounds the best way",
    "1255470": "nooblife This is really great insight. @tanulsingh077 have you given this a try? I was wondering the exact same thing and have yet to try and implement a strategy."
  },
  "source": "meta"
}