{
  "id": 90535,
  "title": "Is the seg_id hexadecimal expression of time order??",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90535",
  "author_name": "Shinsei66",
  "post_date": "2019-04-24T16:08:05.922000",
  "votes": 8,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Although I've not go through all the seg_id of the test data, I'm wondering whether the id part of the seg_id (seg_XXXXXX ) is a hexadecimal expression of a certain kind of order of the data such as the time of the data taken.</p>\n\n<p>I tried to convert the id part to decimal numbers. Hope this will be of some hint.</p>",
  "messages": [
    {
      "id": 522548,
      "postDate": "2019-04-24T16:08:05.923Z",
      "content": "<p>Although I've not go through all the seg_id of the test data, I'm wondering whether the id part of the seg_id (seg_XXXXXX ) is a hexadecimal expression of a certain kind of order of the data such as the time of the data taken.</p>\n\n<p>I tried to convert the id part to decimal numbers. Hope this will be of some hint.</p>",
      "rawMarkdown": "Although I've not go through all the seg_id of the test data, I'm wondering whether the id part of the seg_id (seg_XXXXXX ) is a hexadecimal expression of a certain kind of order of the data such as the time of the data taken.\n\nI tried to convert the id part to decimal numbers. Hope this will be of some hint.",
      "votes": 8
    },
    {
      "id": 522751,
      "postDate": "2019-04-25T00:09:21.283Z",
      "content": "<p>I hope not, or else Kaggle has learned nothing about preventing leaks over the past N years.</p>",
      "rawMarkdown": "I hope not, or else Kaggle has learned nothing about preventing leaks over the past N years.",
      "votes": 8
    },
    {
      "id": 522961,
      "postDate": "2019-04-25T09:36:12.133Z",
      "content": "<p>The 'finding' is that seg_ids are sorted...</p>",
      "rawMarkdown": "The 'finding' is that seg_ids are sorted...",
      "votes": 4,
      "replies": [
        {
          "id": 522995,
          "postDate": "2019-04-25T10:56:08.867Z",
          "content": "<p>This</p>",
          "rawMarkdown": "This"
        }
      ]
    },
    {
      "id": 522873,
      "postDate": "2019-04-25T06:57:44.660Z",
      "content": "<p>Edit: I should wake up before writing comments...</p>",
      "rawMarkdown": "Edit: I should wake up before writing comments...",
      "votes": 1
    },
    {
      "id": 522898,
      "postDate": "2019-04-25T07:52:19.663Z",
      "content": "<p>Wow.</p>",
      "rawMarkdown": "Wow."
    },
    {
      "id": 522892,
      "postDate": "2019-04-25T07:38:31.667Z",
      "content": "<p>It was the first competition I could ever get to the silver tier but I can smell that I am gonna loose it sooooo fast because of your finding.</p>",
      "rawMarkdown": "It was the first competition I could ever get to the silver tier but I can smell that I am gonna loose it sooooo fast because of your finding.",
      "replies": [
        {
          "id": 522900,
          "postDate": "2019-04-25T07:54:36.973Z",
          "content": "<p>Why is that?</p>",
          "rawMarkdown": "Why is that?",
          "votes": 1
        },
        {
          "id": 522905,
          "postDate": "2019-04-25T08:06:15.420Z",
          "content": "<p>Cuz when you plot what she has found, it's linearly increasing up to 16m. Might be a leak. I am not sure yet and I hope it's not the case.</p>",
          "rawMarkdown": "Cuz when you plot what she has found, it's linearly increasing up to 16m. Might be a leak. I am not sure yet and I hope it's not the case."
        },
        {
          "id": 522922,
          "postDate": "2019-04-25T08:37:58.400Z",
          "content": "<p>I don't understand, if its linear, not much information gain though.</p>",
          "rawMarkdown": "I don't understand, if its linear, not much information gain though.",
          "votes": 1
        },
        {
          "id": 522924,
          "postDate": "2019-04-25T08:39:11.527Z",
          "content": "<p>Or it just uniformly distributed random numbers...\nDoes anything support the hypothesis that id is related to time?</p>",
          "rawMarkdown": "Or it just uniformly distributed random numbers...\nDoes anything support the hypothesis that id is related to time?",
          "votes": 1
        },
        {
          "id": 522925,
          "postDate": "2019-04-25T08:40:31.700Z",
          "content": "<blockquote>\n  <p>it's linearly increasing up to 16m</p>\n</blockquote>\n\n<p>Sure, integer and hexadecimal representation are very well correlated...</p>",
          "rawMarkdown": "&gt; it's linearly increasing up to 16m\n\nSure, integer and hexadecimal representation are very well correlated...",
          "votes": 1
        },
        {
          "id": 522929,
          "postDate": "2019-04-25T08:44:05.443Z",
          "content": "<p>if it's linearly increasing it might be a notion that it belongs to time? so test segments were not shuffled? i.e. nearby segments couldn't be much different from each other and there should be some pattern between groups of them? \nI took the diff() and luckily they weren't equal so that's a good sign. </p>",
          "rawMarkdown": "if it's linearly increasing it might be a notion that it belongs to time? so test segments were not shuffled? i.e. nearby segments couldn't be much different from each other and there should be some pattern between groups of them? \nI took the diff() and luckily they weren't equal so that's a good sign. "
        },
        {
          "id": 522931,
          "postDate": "2019-04-25T08:49:41.337Z",
          "content": "<p>If you had segments ordered by time, you could find high spikes just before quake, so the problem become much simpler as time to quake is decreasing between quakes.</p>",
          "rawMarkdown": "If you had segments ordered by time, you could find high spikes just before quake, so the problem become much simpler as time to quake is decreasing between quakes."
        },
        {
          "id": 522932,
          "postDate": "2019-04-25T08:51:23.800Z",
          "content": "<p><a href=\"/alexfir\">@alexfir</a> that is exactly why I am concerned. If seg_id is somehow related to time and segments were not shuffled it could be a leak and people will take advantage of it soon.</p>",
          "rawMarkdown": "@alexfir that is exactly why I am concerned. If seg_id is somehow related to time and segments were not shuffled it could be a leak and people will take advantage of it soon."
        },
        {
          "id": 522943,
          "postDate": "2019-04-25T09:08:57.873Z",
          "content": "<p>It also would make the idea of this competition useless. As I get original work breakthrough was not only because of using ML, but also because time to quake was predicted without tracking history, only by small chunk of data.</p>",
          "rawMarkdown": "It also would make the idea of this competition useless. As I get original work breakthrough was not only because of using ML, but also because time to quake was predicted without tracking history, only by small chunk of data.",
          "votes": 1
        },
        {
          "id": 522945,
          "postDate": "2019-04-25T09:13:37.090Z",
          "content": "<p><a href=\"/alexfir\">@alexfir</a> correct. It has happened to several Kaggle competitions before and it not only screws the competition but also wastes the efforts so many people put on this competition before a leak was found.</p>",
          "rawMarkdown": "@alexfir correct. It has happened to several Kaggle competitions before and it not only screws the competition but also wastes the efforts so many people put on this competition before a leak was found."
        }
      ]
    },
    {
      "id": 522771,
      "postDate": "2019-04-25T01:52:04.577Z",
      "content": "<p>Thank you for sharing the important findings. If test set is ordered by time, shuffled Kfold modeling will bring big shake into the Leaderboard in this competition.</p>",
      "rawMarkdown": "Thank you for sharing the important findings. If test set is ordered by time, shuffled Kfold modeling will bring big shake into the Leaderboard in this competition."
    },
    {
      "id": 522608,
      "postDate": "2019-04-24T18:02:23.763Z",
      "content": "<p>if this is the case ,then the test segments were not shuffled after taking from a continuous experiment.</p>",
      "rawMarkdown": "if this is the case ,then the test segments were not shuffled after taking from a continuous experiment."
    },
    {
      "id": 523019,
      "postDate": "2019-04-25T11:57:32.737Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 523030,
          "postDate": "2019-04-25T12:30:58.983Z",
          "content": "<blockquote>\n  <p>I think the finding is that the file names are hexadecimal numbers. And that hexadecimal numbers can be converted to integers and that integers can be sorted</p>\n</blockquote>\n\n<p>And sorting the integers is the same as sorting the seg_ids, and they are already sorted.</p>",
          "rawMarkdown": "&gt; I think the finding is that the file names are hexadecimal numbers. And that hexadecimal numbers can be converted to integers and that integers can be sorted\n\nAnd sorting the integers is the same as sorting the seg_ids, and they are already sorted.",
          "votes": 1
        },
        {
          "id": 523057,
          "postDate": "2019-04-25T13:40:12.760Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 523065,
          "postDate": "2019-04-25T14:01:26.053Z",
          "content": "<p>I'm just saying that the sorting your propose is the same as sorting by filename,.  This sorting is not one of many, it is the one used in the sample submission file.</p>\n\n<p>That I need to write it 4 times in 4different comments to get across is puzzling me.</p>",
          "rawMarkdown": "I'm just saying that the sorting your propose is the same as sorting by filename,.  This sorting is not one of many, it is the one used in the sample submission file.\n\nThat I need to write it 4 times in 4different comments to get across is puzzling me."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 522751,
      "author_name": "Branden Murray",
      "author_url": "",
      "post_date": "2019-04-25T00:09:21.283000",
      "content": "<p>I hope not, or else Kaggle has learned nothing about preventing leaks over the past N years.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 522961,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-04-25T09:36:12.133000",
      "content": "<p>The 'finding' is that seg_ids are sorted...</p>",
      "votes": 4,
      "replies": [
        {
          "id": 522995,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-04-25T10:56:08.867000",
          "content": "<p>This</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 522873,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-04-25T06:57:44.660000",
      "content": "<p>Edit: I should wake up before writing comments...</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 522898,
      "author_name": "Timmmmmms",
      "author_url": "",
      "post_date": "2019-04-25T07:52:19.663000",
      "content": "<p>Wow.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 522892,
      "author_name": "Massoud Hosseinali",
      "author_url": "",
      "post_date": "2019-04-25T07:38:31.667000",
      "content": "<p>It was the first competition I could ever get to the silver tier but I can smell that I am gonna loose it sooooo fast because of your finding.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 522900,
          "author_name": "Marcus Lin",
          "author_url": "",
          "post_date": "2019-04-25T07:54:36.973000",
          "content": "<p>Why is that?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 522905,
          "author_name": "Massoud Hosseinali",
          "author_url": "",
          "post_date": "2019-04-25T08:06:15.420000",
          "content": "<p>Cuz when you plot what she has found, it's linearly increasing up to 16m. Might be a leak. I am not sure yet and I hope it's not the case.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 522922,
          "author_name": "Marcus Lin",
          "author_url": "",
          "post_date": "2019-04-25T08:37:58.400000",
          "content": "<p>I don't understand, if its linear, not much information gain though.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 522924,
          "author_name": "Alexander Firsov",
          "author_url": "",
          "post_date": "2019-04-25T08:39:11.527000",
          "content": "<p>Or it just uniformly distributed random numbers...\nDoes anything support the hypothesis that id is related to time?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 522925,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-25T08:40:31.700000",
          "content": "<blockquote>\n  <p>it's linearly increasing up to 16m</p>\n</blockquote>\n\n<p>Sure, integer and hexadecimal representation are very well correlated...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 522929,
          "author_name": "Massoud Hosseinali",
          "author_url": "",
          "post_date": "2019-04-25T08:44:05.443000",
          "content": "<p>if it's linearly increasing it might be a notion that it belongs to time? so test segments were not shuffled? i.e. nearby segments couldn't be much different from each other and there should be some pattern between groups of them? \nI took the diff() and luckily they weren't equal so that's a good sign. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 522931,
          "author_name": "Alexander Firsov",
          "author_url": "",
          "post_date": "2019-04-25T08:49:41.337000",
          "content": "<p>If you had segments ordered by time, you could find high spikes just before quake, so the problem become much simpler as time to quake is decreasing between quakes.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 522932,
          "author_name": "Massoud Hosseinali",
          "author_url": "",
          "post_date": "2019-04-25T08:51:23.800000",
          "content": "<p><a href=\"/alexfir\">@alexfir</a> that is exactly why I am concerned. If seg_id is somehow related to time and segments were not shuffled it could be a leak and people will take advantage of it soon.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 522943,
          "author_name": "Alexander Firsov",
          "author_url": "",
          "post_date": "2019-04-25T09:08:57.873000",
          "content": "<p>It also would make the idea of this competition useless. As I get original work breakthrough was not only because of using ML, but also because time to quake was predicted without tracking history, only by small chunk of data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 522945,
          "author_name": "Massoud Hosseinali",
          "author_url": "",
          "post_date": "2019-04-25T09:13:37.090000",
          "content": "<p><a href=\"/alexfir\">@alexfir</a> correct. It has happened to several Kaggle competitions before and it not only screws the competition but also wastes the efforts so many people put on this competition before a leak was found.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 522771,
      "author_name": "T.Ito",
      "author_url": "",
      "post_date": "2019-04-25T01:52:04.577000",
      "content": "<p>Thank you for sharing the important findings. If test set is ordered by time, shuffled Kfold modeling will bring big shake into the Leaderboard in this competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 522608,
      "author_name": "Ahmed Sleem",
      "author_url": "",
      "post_date": "2019-04-24T18:02:23.763000",
      "content": "<p>if this is the case ,then the test segments were not shuffled after taking from a continuous experiment.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 523019,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-25T11:57:32.737000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 523030,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-25T12:30:58.983000",
          "content": "<blockquote>\n  <p>I think the finding is that the file names are hexadecimal numbers. And that hexadecimal numbers can be converted to integers and that integers can be sorted</p>\n</blockquote>\n\n<p>And sorting the integers is the same as sorting the seg_ids, and they are already sorted.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 523057,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-25T13:40:12.760000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 523065,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-25T14:01:26.053000",
          "content": "<p>I'm just saying that the sorting your propose is the same as sorting by filename,.  This sorting is not one of many, it is the one used in the sample submission file.</p>\n\n<p>That I need to write it 4 times in 4different comments to get across is puzzling me.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "522548": "Although I've not go through all the seg_id of the test data, I'm wondering whether the id part of the seg_id (seg_XXXXXX ) is a hexadecimal expression of a certain kind of order of the data such as the time of the data taken.\n\nI tried to convert the id part to decimal numbers. Hope this will be of some hint.",
    "522751": "I hope not, or else Kaggle has learned nothing about preventing leaks over the past N years.",
    "522961": "The 'finding' is that seg_ids are sorted...",
    "522873": "Edit: I should wake up before writing comments...",
    "522898": "Wow.",
    "522892": "It was the first competition I could ever get to the silver tier but I can smell that I am gonna loose it sooooo fast because of your finding.",
    "522771": "Thank you for sharing the important findings. If test set is ordered by time, shuffled Kfold modeling will bring big shake into the Leaderboard in this competition.",
    "522608": "if this is the case ,then the test segments were not shuffled after taking from a continuous experiment.",
    "523019": ""
  }
}