{
  "id": 584485,
  "title": "Zero is a signal too",
  "url": "/competitions/drw-crypto-market-prediction/discussion/584485",
  "author_name": "Gleb Shanshin",
  "post_date": "2025-06-13T17:29:34.420000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>It is interesting that part of the data contains zeroes. So like around 80% of X864 is zero.</p>\n<h1>1. Different columns has different number of zeros. It means that we can pair some of them to understand that they are connected. For example</h1>\n<p>X613 - 170 zeros<br>\nX619 - 170 zeros</p>\n<p>X614 - 70322 zeros<br>\nX620 - 70322 zeros</p>\n<p>X545 - 184295 zeros<br>\nX546 - 184295 zeros</p>\n<p>You can check that they look correlated</p>\n<h1>2. Some columns are zero in train and not all zero in test. I think it is kinda leak, because it means that all zeroes stand behind all zeroes in test. We can extend this idea in some trick</h1>\n<p>Assume we have i, j, k feature which are zero prefixed, test set is in brackets</p>\n<p>X_i = 000[100]<br>\nX_j = 000[110]<br>\nX_k = 000[111]</p>\n<p>it means that we can shuffle it in according to this prefixes to get right order</p>\n<p>X_i = 000001<br>\nX_j = 000011<br>\nX_k = 000111</p>\n<p>Of course we don't have enough columns to determine all timestamps, however we can understand seasons if we imagine that every 1 in previous example is 3 months for example</p>\n<h1>3. If we know point where columns stopped being zero, we can reverse-engineer coin that was launched this day</h1>",
  "messages": [
    {
      "id": 3223688,
      "postDate": "2025-06-13T17:29:34.420Z",
      "content": "<p>It is interesting that part of the data contains zeroes. So like around 80% of X864 is zero.</p>\n<h1>1. Different columns has different number of zeros. It means that we can pair some of them to understand that they are connected. For example</h1>\n<p>X613 - 170 zeros<br>\nX619 - 170 zeros</p>\n<p>X614 - 70322 zeros<br>\nX620 - 70322 zeros</p>\n<p>X545 - 184295 zeros<br>\nX546 - 184295 zeros</p>\n<p>You can check that they look correlated</p>\n<h1>2. Some columns are zero in train and not all zero in test. I think it is kinda leak, because it means that all zeroes stand behind all zeroes in test. We can extend this idea in some trick</h1>\n<p>Assume we have i, j, k feature which are zero prefixed, test set is in brackets</p>\n<p>X_i = 000[100]<br>\nX_j = 000[110]<br>\nX_k = 000[111]</p>\n<p>it means that we can shuffle it in according to this prefixes to get right order</p>\n<p>X_i = 000001<br>\nX_j = 000011<br>\nX_k = 000111</p>\n<p>Of course we don't have enough columns to determine all timestamps, however we can understand seasons if we imagine that every 1 in previous example is 3 months for example</p>\n<h1>3. If we know point where columns stopped being zero, we can reverse-engineer coin that was launched this day</h1>",
      "rawMarkdown": "It is interesting that part of the data contains zeroes. So like around 80% of X864 is zero.\n\n# 1. Different columns has different number of zeros. It means that we can pair some of them to understand that they are connected. For example\n\nX613 - 170 zeros\nX619 - 170 zeros\n\nX614 - 70322 zeros\nX620 - 70322 zeros\n\nX545 - 184295 zeros\nX546 - 184295 zeros\n\nYou can check that they look correlated\n\n\n# 2. Some columns are zero in train and not all zero in test. I think it is kinda leak, because it means that all zeroes stand behind all zeroes in test. We can extend this idea in some trick\n\nAssume we have i, j, k feature which are zero prefixed, test set is in brackets\n\nX_i = 000[100]\nX_j = 000[110]\nX_k = 000[111]\n\nit means that we can shuffle it in according to this prefixes to get right order\n\nX_i = 000001\nX_j = 000011\nX_k = 000111\n\nOf course we don't have enough columns to determine all timestamps, however we can understand seasons if we imagine that every 1 in previous example is 3 months for example\n\n# 3. If we know point where columns stopped being zero, we can reverse-engineer coin that was launched this day\n",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3223688": "It is interesting that part of the data contains zeroes. So like around 80% of X864 is zero.\n\n# 1. Different columns has different number of zeros. It means that we can pair some of them to understand that they are connected. For example\n\nX613 - 170 zeros\nX619 - 170 zeros\n\nX614 - 70322 zeros\nX620 - 70322 zeros\n\nX545 - 184295 zeros\nX546 - 184295 zeros\n\nYou can check that they look correlated\n\n\n# 2. Some columns are zero in train and not all zero in test. I think it is kinda leak, because it means that all zeroes stand behind all zeroes in test. We can extend this idea in some trick\n\nAssume we have i, j, k feature which are zero prefixed, test set is in brackets\n\nX_i = 000[100]\nX_j = 000[110]\nX_k = 000[111]\n\nit means that we can shuffle it in according to this prefixes to get right order\n\nX_i = 000001\nX_j = 000011\nX_k = 000111\n\nOf course we don't have enough columns to determine all timestamps, however we can understand seasons if we imagine that every 1 in previous example is 3 months for example\n\n# 3. If we know point where columns stopped being zero, we can reverse-engineer coin that was launched this day\n"
  }
}