{
  "id": 680068,
  "title": "Dataset Update: Data Leakage Issues Fixed",
  "url": "/competitions/cyber-physical-anomaly-detection-for-der-systems/discussion/680068",
  "author_name": "",
  "post_date": "2026-03-05T16:29:53.949829700Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thank you to everyone who reported the data leakage issues:</p>\n<ul>\n<li><strong>@αρ</strong> for the original report identifying the perfect score without training</li>\n<li><strong>@Spiritmilk</strong> for the detailed analysis showing Id leak and noting CatBoost behavior</li>\n<li><strong>@Nadiia</strong> for confirming the issue and asking about empty columns</li>\n</ul>\n<p>Your feedback helped us identify and fix several critical problems.</p>\n<h3>What Went Wrong</h3>\n<p>We identified <strong>three main leakage sources</strong> in the old dataset:</p>\n<p><strong>1. Id Column Leakage (Critical)</strong>\nWhen preparing the dataset, we concatenated normal data first, then abnormal data, and assigned sequential IDs afterward. This meant all normal samples had IDs 0-1,659,036 and all abnormal samples had IDs 1,659,037+. A simple threshold could achieve 100% accuracy.</p>\n<p><strong>2. NaN Pattern Leakage (High)</strong>\nOur data comes from simulating different DER system sizes (10kW and 100kW). Due to infrastructure differences in how these were generated, certain columns like <code>DERMeasureAC[0].DERMode</code> had different NaN patterns between normal and abnormal data. This is why CatBoost could reach 1.0 without using the Id column—it treats NaN as a distinct category.</p>\n<p><strong>3. Infrastructure Column Leakage (High)</strong>\nThe <code>DeviceType</code> column was only present in one of our four source datasets, making it a perfect indicator for a subset of abnormal samples.</p>\n<h3>Fixes Applied</h3>\n<table>\n<thead>\n<tr>\n<th>Issue</th>\n<th>Fix</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Id column leakage</td>\n<td>Data is now <strong>shuffled</strong> before Id assignment</td>\n</tr>\n<tr>\n<td>DERMeasureAC[0].DERMode</td>\n<td>Column <strong>removed</strong></td>\n</tr>\n<tr>\n<td>DeviceType column</td>\n<td>Column <strong>removed</strong></td>\n</tr>\n<tr>\n<td>DerSimControls columns (11)</td>\n<td>Columns <strong>removed</strong> (simulator internals, not real DER data)</td>\n</tr>\n<tr>\n<td>MnAlrmInfo column</td>\n<td>Column <strong>removed</strong> (was empty anyway)</td>\n</tr>\n<tr>\n<td>DERTripHF[0].Ena</td>\n<td>Column <strong>removed</strong> (only appeared in normal data)</td>\n</tr>\n</tbody>\n</table>\n<p>The new dataset has <strong>722 feature columns</strong> (down from 739) after removing leakage sources.</p>\n<h3>About the Empty Columns</h3>\n<p>Several competitors asked about empty or duplicate columns. This dataset represents SunSpec protocol data from DER (Distributed Energy Resource) devices. The columns map to the SunSpec standard's data model—some fields are empty because:</p>\n<ul>\n<li>The simulated device state doesn't populate them</li>\n<li>They were intentionally removed to prevent leakage (alarm/status fields)</li>\n<li>They're optional fields in the standard</li>\n</ul>\n<p>We kept the column structure to represent what you'd see scanning a real device.</p>\n<h3>What to Expect</h3>\n<ul>\n<li>Leaderboard scores will reset with the new dataset (Pending Kaggle administrator contact)</li>\n<li>We're committed to fixing any additional issues the community finds</li>\n</ul>\n<p>Thank you again for your patience and for helping us improve this competition. This data generation process is new for us, and your feedback is invaluable.</p>\n<p><strong>Jorge Pineda</strong> - Competition Host</p>",
  "messages": [
    {
      "id": "3417541",
      "postDate": "03/05/2026 16:29:53",
      "content": "<p>Thank you to everyone who reported the data leakage issues:</p>\n<ul>\n<li><strong>@αρ</strong> for the original report identifying the perfect score without training</li>\n<li><strong>@Spiritmilk</strong> for the detailed analysis showing Id leak and noting CatBoost behavior</li>\n<li><strong>@Nadiia</strong> for confirming the issue and asking about empty columns</li>\n</ul>\n<p>Your feedback helped us identify and fix several critical problems.</p>\n<h3>What Went Wrong</h3>\n<p>We identified <strong>three main leakage sources</strong> in the old dataset:</p>\n<p><strong>1. Id Column Leakage (Critical)</strong>\nWhen preparing the dataset, we concatenated normal data first, then abnormal data, and assigned sequential IDs afterward. This meant all normal samples had IDs 0-1,659,036 and all abnormal samples had IDs 1,659,037+. A simple threshold could achieve 100% accuracy.</p>\n<p><strong>2. NaN Pattern Leakage (High)</strong>\nOur data comes from simulating different DER system sizes (10kW and 100kW). Due to infrastructure differences in how these were generated, certain columns like <code>DERMeasureAC[0].DERMode</code> had different NaN patterns between normal and abnormal data. This is why CatBoost could reach 1.0 without using the Id column—it treats NaN as a distinct category.</p>\n<p><strong>3. Infrastructure Column Leakage (High)</strong>\nThe <code>DeviceType</code> column was only present in one of our four source datasets, making it a perfect indicator for a subset of abnormal samples.</p>\n<h3>Fixes Applied</h3>\n<table>\n<thead>\n<tr>\n<th>Issue</th>\n<th>Fix</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Id column leakage</td>\n<td>Data is now <strong>shuffled</strong> before Id assignment</td>\n</tr>\n<tr>\n<td>DERMeasureAC[0].DERMode</td>\n<td>Column <strong>removed</strong></td>\n</tr>\n<tr>\n<td>DeviceType column</td>\n<td>Column <strong>removed</strong></td>\n</tr>\n<tr>\n<td>DerSimControls columns (11)</td>\n<td>Columns <strong>removed</strong> (simulator internals, not real DER data)</td>\n</tr>\n<tr>\n<td>MnAlrmInfo column</td>\n<td>Column <strong>removed</strong> (was empty anyway)</td>\n</tr>\n<tr>\n<td>DERTripHF[0].Ena</td>\n<td>Column <strong>removed</strong> (only appeared in normal data)</td>\n</tr>\n</tbody>\n</table>\n<p>The new dataset has <strong>722 feature columns</strong> (down from 739) after removing leakage sources.</p>\n<h3>About the Empty Columns</h3>\n<p>Several competitors asked about empty or duplicate columns. This dataset represents SunSpec protocol data from DER (Distributed Energy Resource) devices. The columns map to the SunSpec standard's data model—some fields are empty because:</p>\n<ul>\n<li>The simulated device state doesn't populate them</li>\n<li>They were intentionally removed to prevent leakage (alarm/status fields)</li>\n<li>They're optional fields in the standard</li>\n</ul>\n<p>We kept the column structure to represent what you'd see scanning a real device.</p>\n<h3>What to Expect</h3>\n<ul>\n<li>Leaderboard scores will reset with the new dataset (Pending Kaggle administrator contact)</li>\n<li>We're committed to fixing any additional issues the community finds</li>\n</ul>\n<p>Thank you again for your patience and for helping us improve this competition. This data generation process is new for us, and your feedback is invaluable.</p>\n<p><strong>Jorge Pineda</strong> - Competition Host</p>",
      "rawMarkdown": "Thank you to everyone who reported the data leakage issues:\n- **@αρ** for the original report identifying the perfect score without training\n- **@Spiritmilk** for the detailed analysis showing Id leak and noting CatBoost behavior\n- **@Nadiia** for confirming the issue and asking about empty columns\n\nYour feedback helped us identify and fix several critical problems.\n\n### What Went Wrong\n\nWe identified **three main leakage sources** in the old dataset:\n\n**1. Id Column Leakage (Critical)**\nWhen preparing the dataset, we concatenated normal data first, then abnormal data, and assigned sequential IDs afterward. This meant all normal samples had IDs 0-1,659,036 and all abnormal samples had IDs 1,659,037+. A simple threshold could achieve 100% accuracy.\n\n**2. NaN Pattern Leakage (High)**\nOur data comes from simulating different DER system sizes (10kW and 100kW). Due to infrastructure differences in how these were generated, certain columns like `DERMeasureAC[0].DERMode` had different NaN patterns between normal and abnormal data. This is why CatBoost could reach 1.0 without using the Id column—it treats NaN as a distinct category.\n\n**3. Infrastructure Column Leakage (High)**\nThe `DeviceType` column was only present in one of our four source datasets, making it a perfect indicator for a subset of abnormal samples.\n\n### Fixes Applied\n\n| Issue | Fix |\n|-------|-----|\n| Id column leakage | Data is now **shuffled** before Id assignment |\n| DERMeasureAC[0].DERMode | Column **removed** |\n| DeviceType column | Column **removed** |\n| DerSimControls columns (11) | Columns **removed** (simulator internals, not real DER data) |\n| MnAlrmInfo column | Column **removed** (was empty anyway) |\n| DERTripHF[0].Ena | Column **removed** (only appeared in normal data) |\n\nThe new dataset has **722 feature columns** (down from 739) after removing leakage sources.\n\n### About the Empty Columns\n\nSeveral competitors asked about empty or duplicate columns. This dataset represents SunSpec protocol data from DER (Distributed Energy Resource) devices. The columns map to the SunSpec standard's data model—some fields are empty because:\n- The simulated device state doesn't populate them\n- They were intentionally removed to prevent leakage (alarm/status fields)\n- They're optional fields in the standard\n\nWe kept the column structure to represent what you'd see scanning a real device.\n\n### What to Expect\n\n- Leaderboard scores will reset with the new dataset (Pending Kaggle administrator contact)\n- We're committed to fixing any additional issues the community finds\n\nThank you again for your patience and for helping us improve this competition. This data generation process is new for us, and your feedback is invaluable.\n\n**Jorge Pineda** - Competition Host",
      "votes": null
    },
    {
      "id": "3417828",
      "postDate": "03/06/2026 10:37:43",
      "content": "<p>Similarly to the already removed DERTripHF[0].Ena, the analogous HV, LV and LF collumns appear to be of the same type of leak. They all score an AUC = 1.0 on the 10kW sim. </p>",
      "rawMarkdown": "Similarly to the already removed DERTripHF[0].Ena, the analogous HV, LV and LF collumns appear to be of the same type of leak. They all score an AUC = 1.0 on the 10kW sim.",
      "votes": null
    },
    {
      "id": "3417970",
      "postDate": "03/06/2026 18:51:59",
      "content": "<p>I removed these features and updated the dataset. Thanks for your feedback. I will post later with more information about the dataset.</p>",
      "rawMarkdown": "I removed these features and updated the dataset. Thanks for your feedback. I will post later with more information about the dataset.",
      "votes": null
    },
    {
      "id": "3419085",
      "postDate": "03/09/2026 21:02:01",
      "content": "<p>Since the original fix, we completed a root cause investigation and regenerated the source data that had inconsistent settings. The DERTrip Ena columns (reported by <a href=\"https://www.kaggle.com/leo01000111\" target=\"_blank\">@leo01000111</a> ) are back in the dataset now that the underlying issue is resolved.</p>\n<p>The dataset has been updated and the leaderboard reset. Let us know if you find any other issues.</p>",
      "rawMarkdown": "Since the original fix, we completed a root cause investigation and regenerated the source data that had inconsistent settings. The DERTrip Ena columns (reported by @leo01000111 ) are back in the dataset now that the underlying issue is resolved.\n\nThe dataset has been updated and the leaderboard reset. Let us know if you find any other issues.",
      "votes": null
    },
    {
      "id": "3419323",
      "postDate": "03/10/2026 13:29:00",
      "content": "<p>Thanks for the fix!</p>\n<p>Just to understand it better, was the data recreated from 0 or just updated?</p>\n<p>If it was updated, I'm afraid people who had access to the previous data might be able to link the new dataset to the previous one and gather unfair advantage (by essentially having an answer cheat, if they are able to correctly link the old and new data).</p>\n<p>If it was recreated, the only advantage people with the old data might have is data volume, which I don't know if it will be influential, since we already have quite a lot of entries. Also, is there a way to access the old data? Thank you.</p>",
      "rawMarkdown": "Thanks for the fix!\n\nJust to understand it better, was the data recreated from 0 or just updated?\n\nIf it was updated, I'm afraid people who had access to the previous data might be able to link the new dataset to the previous one and gather unfair advantage (by essentially having an answer cheat, if they are able to correctly link the old and new data).\n\nIf it was recreated, the only advantage people with the old data might have is data volume, which I don't know if it will be influential, since we already have quite a lot of entries. Also, is there a way to access the old data? Thank you.",
      "votes": null
    },
    {
      "id": "3419409",
      "postDate": "03/10/2026 16:46:45",
      "content": "<p>Part of the source data was regenerated and we made additional undisclosed changes to the data preparation. Row matching against the old dataset is theoretically possible, though all winners are required to submit their source code and methodology for review — so any submission relying on linking old and new data rather than a trained model would be caught during verification.</p>\n<p>As for extra data volume, we don't think it would provide a meaningful advantage — the old and new data come from the same simulator, so additional samples should be statistically redundant given the size of the current dataset.</p>\n<p>This is our first time hosting a competition like this, so we're learning as we go. We appreciate the vigilance and want to make sure it's fair for everyone — please keep flagging anything that seems off.</p>",
      "rawMarkdown": "Part of the source data was regenerated and we made additional undisclosed changes to the data preparation. Row matching against the old dataset is theoretically possible, though all winners are required to submit their source code and methodology for review — so any submission relying on linking old and new data rather than a trained model would be caught during verification.\n\nAs for extra data volume, we don't think it would provide a meaningful advantage — the old and new data come from the same simulator, so additional samples should be statistically redundant given the size of the current dataset.\n\nThis is our first time hosting a competition like this, so we're learning as we go. We appreciate the vigilance and want to make sure it's fair for everyone — please keep flagging anything that seems off.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3417828,
      "author_name": "leo01000111",
      "author_url": "",
      "post_date": "03/06/2026 10:37:43",
      "content": "<p>Similarly to the already removed DERTripHF[0].Ena, the analogous HV, LV and LF collumns appear to be of the same type of leak. They all score an AUC = 1.0 on the 10kW sim. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3417970,
          "author_name": "jpinedaa",
          "author_url": "",
          "post_date": "03/06/2026 18:51:59",
          "content": "<p>I removed these features and updated the dataset. Thanks for your feedback. I will post later with more information about the dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3419085,
      "author_name": "jpinedaa",
      "author_url": "",
      "post_date": "03/09/2026 21:02:01",
      "content": "<p>Since the original fix, we completed a root cause investigation and regenerated the source data that had inconsistent settings. The DERTrip Ena columns (reported by <a href=\"https://www.kaggle.com/leo01000111\" target=\"_blank\">@leo01000111</a> ) are back in the dataset now that the underlying issue is resolved.</p>\n<p>The dataset has been updated and the leaderboard reset. Let us know if you find any other issues.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3419323,
          "author_name": "araraonline",
          "author_url": "",
          "post_date": "03/10/2026 13:29:00",
          "content": "<p>Thanks for the fix!</p>\n<p>Just to understand it better, was the data recreated from 0 or just updated?</p>\n<p>If it was updated, I'm afraid people who had access to the previous data might be able to link the new dataset to the previous one and gather unfair advantage (by essentially having an answer cheat, if they are able to correctly link the old and new data).</p>\n<p>If it was recreated, the only advantage people with the old data might have is data volume, which I don't know if it will be influential, since we already have quite a lot of entries. Also, is there a way to access the old data? Thank you.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3419409,
              "author_name": "jpinedaa",
              "author_url": "",
              "post_date": "03/10/2026 16:46:45",
              "content": "<p>Part of the source data was regenerated and we made additional undisclosed changes to the data preparation. Row matching against the old dataset is theoretically possible, though all winners are required to submit their source code and methodology for review — so any submission relying on linking old and new data rather than a trained model would be caught during verification.</p>\n<p>As for extra data volume, we don't think it would provide a meaningful advantage — the old and new data come from the same simulator, so additional samples should be statistically redundant given the size of the current dataset.</p>\n<p>This is our first time hosting a competition like this, so we're learning as we go. We appreciate the vigilance and want to make sure it's fair for everyone — please keep flagging anything that seems off.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3417541": "Thank you to everyone who reported the data leakage issues:\n- **@αρ** for the original report identifying the perfect score without training\n- **@Spiritmilk** for the detailed analysis showing Id leak and noting CatBoost behavior\n- **@Nadiia** for confirming the issue and asking about empty columns\n\nYour feedback helped us identify and fix several critical problems.\n\n### What Went Wrong\n\nWe identified **three main leakage sources** in the old dataset:\n\n**1. Id Column Leakage (Critical)**\nWhen preparing the dataset, we concatenated normal data first, then abnormal data, and assigned sequential IDs afterward. This meant all normal samples had IDs 0-1,659,036 and all abnormal samples had IDs 1,659,037+. A simple threshold could achieve 100% accuracy.\n\n**2. NaN Pattern Leakage (High)**\nOur data comes from simulating different DER system sizes (10kW and 100kW). Due to infrastructure differences in how these were generated, certain columns like `DERMeasureAC[0].DERMode` had different NaN patterns between normal and abnormal data. This is why CatBoost could reach 1.0 without using the Id column—it treats NaN as a distinct category.\n\n**3. Infrastructure Column Leakage (High)**\nThe `DeviceType` column was only present in one of our four source datasets, making it a perfect indicator for a subset of abnormal samples.\n\n### Fixes Applied\n\n| Issue | Fix |\n|-------|-----|\n| Id column leakage | Data is now **shuffled** before Id assignment |\n| DERMeasureAC[0].DERMode | Column **removed** |\n| DeviceType column | Column **removed** |\n| DerSimControls columns (11) | Columns **removed** (simulator internals, not real DER data) |\n| MnAlrmInfo column | Column **removed** (was empty anyway) |\n| DERTripHF[0].Ena | Column **removed** (only appeared in normal data) |\n\nThe new dataset has **722 feature columns** (down from 739) after removing leakage sources.\n\n### About the Empty Columns\n\nSeveral competitors asked about empty or duplicate columns. This dataset represents SunSpec protocol data from DER (Distributed Energy Resource) devices. The columns map to the SunSpec standard's data model—some fields are empty because:\n- The simulated device state doesn't populate them\n- They were intentionally removed to prevent leakage (alarm/status fields)\n- They're optional fields in the standard\n\nWe kept the column structure to represent what you'd see scanning a real device.\n\n### What to Expect\n\n- Leaderboard scores will reset with the new dataset (Pending Kaggle administrator contact)\n- We're committed to fixing any additional issues the community finds\n\nThank you again for your patience and for helping us improve this competition. This data generation process is new for us, and your feedback is invaluable.\n\n**Jorge Pineda** - Competition Host",
    "3417828": "Similarly to the already removed DERTripHF[0].Ena, the analogous HV, LV and LF collumns appear to be of the same type of leak. They all score an AUC = 1.0 on the 10kW sim.",
    "3417970": "I removed these features and updated the dataset. Thanks for your feedback. I will post later with more information about the dataset.",
    "3419085": "Since the original fix, we completed a root cause investigation and regenerated the source data that had inconsistent settings. The DERTrip Ena columns (reported by @leo01000111 ) are back in the dataset now that the underlying issue is resolved.\n\nThe dataset has been updated and the leaderboard reset. Let us know if you find any other issues.",
    "3419323": "Thanks for the fix!\n\nJust to understand it better, was the data recreated from 0 or just updated?\n\nIf it was updated, I'm afraid people who had access to the previous data might be able to link the new dataset to the previous one and gather unfair advantage (by essentially having an answer cheat, if they are able to correctly link the old and new data).\n\nIf it was recreated, the only advantage people with the old data might have is data volume, which I don't know if it will be influential, since we already have quite a lot of entries. Also, is there a way to access the old data? Thank you.",
    "3419409": "Part of the source data was regenerated and we made additional undisclosed changes to the data preparation. Row matching against the old dataset is theoretically possible, though all winners are required to submit their source code and methodology for review — so any submission relying on linking old and new data rather than a trained model would be caught during verification.\n\nAs for extra data volume, we don't think it would provide a meaningful advantage — the old and new data come from the same simulator, so additional samples should be statistically redundant given the size of the current dataset.\n\nThis is our first time hosting a competition like this, so we're learning as we go. We appreciate the vigilance and want to make sure it's fair for everyone — please keep flagging anything that seems off."
  },
  "source": "meta"
}