{
  "id": 689286,
  "title": "1st Place solution - Cyber-Physical Anomaly Detection for DER Systems",
  "url": "/competitions/cyber-physical-anomaly-detection-for-der-systems/writeups/1st-place-solution",
  "author_name": "",
  "post_date": "2026-04-08T06:53:36.603Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I would like to thank the organizers for putting together the competition, and for addressing the early dataset issues so quickly.</p>\n<p>Most of the public leaderboard eventually converged around <code>0.9115</code> (public) / <code>0.9119</code> (private), so I was happy to finish a bit above that score.</p>\n<p>It was a fun competition, and I learned a lot from it. Thank you!</p>\n<p><strong>TLDR</strong>: <a href=\"https://www.kaggle.com/code/alvaroborras/anomaly-detection-cyber-der-final-submission\" target=\"_blank\">This notebook</a> reproduces my final submission, and the full implementation is available in my <a href=\"https://github.com/alvaroborras/Anomaly-Detection-DER-Systems\" target=\"_blank\">personal repository</a>.</p>\n<h1>Overview of the solution</h1>\n<p>The dataset is highly structured, and in the original CSV files it is dominated by two exact identity tuples.</p>\n<p>More concretely, I build a 5-field fingerprint from <code>common[0].Mn</code>, <code>common[0].Md</code>, <code>common[0].Opt</code>, <code>common[0].Vr</code>, and <code>common[0].SN</code>. The two dominant tuples correspond to the 10 kW and 100 kW simulator rows, which I refer to as <code>canon10</code> and <code>canon100</code>, and I use that fingerprint to split the data into <code>canon10</code>, <code>canon100</code>, and a small noncanonical bucket.</p>\n<p>This distinction is very important for the rest of the pipeline. The two canonical families are modeled separately, while the noncanonical rows are handled outside the main learned-model path because they already form a very strong anomaly signal in training.</p>\n<p>From <code>665</code> selected source columns, I build semantic features that make the DER structure explicit: missingness and schema-integrity signals, measurement-vs-rating/setting consistency, control-compliance features, protection and curve features, and AC/DC consistency checks.</p>\n<p>For each canonical family, the final predictor is a small ensemble of XGBoost on the semantic numeric table and CatBoost on a narrower raw + categorical table, with the blend and threshold tuned directly for F2.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5517090%2Fa5fd25af7c5dbf46f7b732c1ab815194%2Foverall_pipeline.png?generation=1775626772760115&amp;alt=media\" alt=\"\"></p>\n<p><em>High level view of the pipeline. Rows are first split by device fingerprint; the two canonical families go through family-specific feature engineering and models, while the small noncanonical bucket is handled separately.</em></p>\n<h2>Observations about the dataset</h2>\n<p>The raw dataset has <code>723</code> features, but many are fully empty or not especially useful for the final model. In particular, <code>182</code> features are completely null in both train and test, while there are no fully empty rows. In practice, the useful signal is not in dropping rows, but in being selective about columns and preserving informative missingness.</p>\n<p>A second key observation is that the data is overwhelmingly concentrated around two exact fingerprints from the original CSV identity columns:</p>\n<ul>\n<li><code>('DERSec', 'DER Simulator', '10 kW DER', '1.2.3', 'SN-Three-Phase')</code>  -&gt;  <code>canon10</code></li>\n<li><code>('DERSec', 'DER Simulator 100 kW', '1.2.3.1', '1.0.0', '1100058974')</code>  -&gt;  <code>canon100</code></li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Bucket</th>\n<th>Definition</th>\n<th>Train rows</th>\n<th>Test rows</th>\n<th>How I use it</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>canon10</code></td>\n<td>Exact match to the canonical 10 kW simulator fingerprint</td>\n<td>1,181,604</td>\n<td>505,993</td>\n<td>Family-specific learned models + overrides</td>\n</tr>\n<tr>\n<td><code>canon100</code></td>\n<td>Exact match to the canonical 100 kW simulator fingerprint</td>\n<td>1,168,218</td>\n<td>501,106</td>\n<td>Family-specific learned models + overrides</td>\n</tr>\n<tr>\n<td><code>other</code> / noncanonical</td>\n<td>Anything else</td>\n<td>13,341</td>\n<td>5,686</td>\n<td>Identity-based anomaly bucket</td>\n</tr>\n</tbody>\n</table>\n<p>This matters because all <code>13,341</code> noncanonical training rows are anomalous. So I did not treat <code>other</code> as a third modeling family. Instead, I used it as a strong identity-based anomaly signal.</p>\n<h2>Feature engineering</h2>\n<p>The main feature groups are:</p>\n<ol>\n<li><strong>Identity and missingness</strong>: device fingerprint, missing identity fields, per-block missingness counts, and missingness patterns.</li>\n<li><strong>Schema integrity</strong>: expected model IDs and lengths, used to detect malformed or structurally inconsistent rows.</li>\n<li><strong>Physical consistency</strong>: measurements compared against ratings and limits, plus phase-sum and power-factor consistency checks.</li>\n<li><strong>Control compliance</strong>: whether the reported active or reactive control state matches the observed output.</li>\n<li><strong>Protection and curves</strong>: enter-service logic, trip blocks, Volt-Var / Volt-Watt / Watt-Var behavior, and frequency-droop features.</li>\n<li><strong>AC/DC consistency</strong>: agreement between AC and DC measurements, per-port values, and sign patterns.</li>\n<li><strong>Residual and scenario features</strong>: surrogate-model residuals and smoothed anomaly rates for recurring operating scenarios.</li>\n</ol>\n<h2>Hard overrides</h2>\n<p>I also used a small set of high-confidence hard overrides. They are simply anomaly signatures that were clean enough in training to justify overriding the learned model.</p>\n<p>The most important case is the noncanonical bucket itself. If the 5-field fingerprint does not exactly match one of the two canonical simulator identities, the row is assigned to <code>other</code>, which acts as an identity-based anomaly bucket rather than a learned family.</p>\n<p>Inside the canonical families, the overrides are limited to a few high-precision patterns such as:</p>\n<ul>\n<li>clear electrical envelope violations,</li>\n<li>large mismatches between enabled controls and measured output,</li>\n<li>rare metadata or AC/DC patterns,</li>\n<li>enter-service contradictions,</li>\n<li>producing power while already outside a must-trip region.</li>\n</ul>\n<p>An important implementation detail is that candidate overrides are audited on the training data. Only the most precise ones remain true hard overrides; the rest can still help as features without forcing the final prediction.</p>\n<h2>Family-specific models and ensembling</h2>\n<p>For <code>canon10</code> and <code>canon100</code>, I train separate models. This worked better than a single global classifier because the two canonical devices operate in different regimes and have different local patterns.</p>\n<p>The main model is <strong>XGBoost</strong> on the full semantic numeric feature set. The companion model is <strong>CatBoost</strong> on a narrower raw + categorical table containing selected raw numeric columns, raw identity fields, the full device fingerprint, and additional identity / missingness-engineered features.</p>\n<p>I also add family-specific regressors trained on normal rows to predict quantities such as active power, apparent power, reactive power, power factor, and current. Their residuals become new features for the final classifiers. For validation, I use both a primary 5-fold split based on <code>Id</code> and an audit split based on a hashed representation of the operating scenario.</p>\n<h2>Threshold optimization</h2>\n<p>Since the competition metric is F2, I did not use the default <code>0.5</code> threshold. For each canonical family, I search directly over out-of-fold probabilities to find the threshold that maximizes F2. I do the same for the XGBoost / CatBoost blend weight, and keep the family-specific operating point that performs well on both the primary and audit splits.</p>\n<h2>Tools used</h2>\n<p>To develop the solution, I used <a href=\"https://docs.astral.sh/uv/\" target=\"_blank\">uv</a>, <a href=\"https://github.com/astral-sh/ruff\" target=\"_blank\">ruff</a>, <a href=\"https://www.docker.com/\" target=\"_blank\">Docker</a>, and <a href=\"https://developers.openai.com/codex/cli\" target=\"_blank\">Codex CLI</a>. Most of the work was done on a MacBook Pro, without a GPU.</p>\n<h2>Reproducibility</h2>\n<p>To keep the results reproducible, I seeded the full pipeline and used a <a href=\"https://console.cloud.google.com/artifacts/docker/kaggle-images/us/gcr.io/python/sha256:02c72a7c98e5e0895056901d9c715d181cd30eae392491235dfea93e6d0de3ed\" target=\"_blank\">Kaggle official Docker image</a> for local development. In the GitHub repository, the solution can be reproduced by building the pinned Docker image and running the <code>run_docker.sh</code> script, which executes the code inside the container and writes the final <code>submission.csv</code>.</p>\n<h2>About me</h2>\n<p>I am an applied mathematician turned into software developer, currently based in Madrid. My interests are AI, machine learning, numerical simulation, and optimization.</p>\n<p>If you would like to reach out, you can find me on <a href=\"https://www.linkedin.com/in/alvaro-borras/\" target=\"_blank\">LinkedIn</a>, or via Gmail as <code>alvaroborrasf</code>.</p>\n<p>Thank you again and looking forward to seeing the approaches from other participants!</p>",
  "messages": [
    {
      "id": "3437808",
      "postDate": "04/08/2026 06:03:48",
      "content": "<p>I would like to thank the organizers for putting together the competition, and for addressing the early dataset issues so quickly.</p>\n<p>Most of the public leaderboard eventually converged around <code>0.9115</code> (public) / <code>0.9119</code> (private), so I was happy to finish a bit above that score.</p>\n<p>It was a fun competition, and I learned a lot from it. Thank you!</p>\n<p><strong>TLDR</strong>: <a href=\"https://www.kaggle.com/code/alvaroborras/anomaly-detection-cyber-der-final-submission\" target=\"_blank\">This notebook</a> reproduces my final submission, and the full implementation is available in my <a href=\"https://github.com/alvaroborras/Anomaly-Detection-DER-Systems\" target=\"_blank\">personal repository</a>.</p>\n<h1>Overview of the solution</h1>\n<p>The dataset is highly structured, and in the original CSV files it is dominated by two exact identity tuples.</p>\n<p>More concretely, I build a 5-field fingerprint from <code>common[0].Mn</code>, <code>common[0].Md</code>, <code>common[0].Opt</code>, <code>common[0].Vr</code>, and <code>common[0].SN</code>. The two dominant tuples correspond to the 10 kW and 100 kW simulator rows, which I refer to as <code>canon10</code> and <code>canon100</code>, and I use that fingerprint to split the data into <code>canon10</code>, <code>canon100</code>, and a small noncanonical bucket.</p>\n<p>This distinction is very important for the rest of the pipeline. The two canonical families are modeled separately, while the noncanonical rows are handled outside the main learned-model path because they already form a very strong anomaly signal in training.</p>\n<p>From <code>665</code> selected source columns, I build semantic features that make the DER structure explicit: missingness and schema-integrity signals, measurement-vs-rating/setting consistency, control-compliance features, protection and curve features, and AC/DC consistency checks.</p>\n<p>For each canonical family, the final predictor is a small ensemble of XGBoost on the semantic numeric table and CatBoost on a narrower raw + categorical table, with the blend and threshold tuned directly for F2.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5517090%2Fa5fd25af7c5dbf46f7b732c1ab815194%2Foverall_pipeline.png?generation=1775626772760115&amp;alt=media\" alt=\"\"></p>\n<p><em>High level view of the pipeline. Rows are first split by device fingerprint; the two canonical families go through family-specific feature engineering and models, while the small noncanonical bucket is handled separately.</em></p>\n<h2>Observations about the dataset</h2>\n<p>The raw dataset has <code>723</code> features, but many are fully empty or not especially useful for the final model. In particular, <code>182</code> features are completely null in both train and test, while there are no fully empty rows. In practice, the useful signal is not in dropping rows, but in being selective about columns and preserving informative missingness.</p>\n<p>A second key observation is that the data is overwhelmingly concentrated around two exact fingerprints from the original CSV identity columns:</p>\n<ul>\n<li><code>('DERSec', 'DER Simulator', '10 kW DER', '1.2.3', 'SN-Three-Phase')</code>  -&gt;  <code>canon10</code></li>\n<li><code>('DERSec', 'DER Simulator 100 kW', '1.2.3.1', '1.0.0', '1100058974')</code>  -&gt;  <code>canon100</code></li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Bucket</th>\n<th>Definition</th>\n<th>Train rows</th>\n<th>Test rows</th>\n<th>How I use it</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>canon10</code></td>\n<td>Exact match to the canonical 10 kW simulator fingerprint</td>\n<td>1,181,604</td>\n<td>505,993</td>\n<td>Family-specific learned models + overrides</td>\n</tr>\n<tr>\n<td><code>canon100</code></td>\n<td>Exact match to the canonical 100 kW simulator fingerprint</td>\n<td>1,168,218</td>\n<td>501,106</td>\n<td>Family-specific learned models + overrides</td>\n</tr>\n<tr>\n<td><code>other</code> / noncanonical</td>\n<td>Anything else</td>\n<td>13,341</td>\n<td>5,686</td>\n<td>Identity-based anomaly bucket</td>\n</tr>\n</tbody>\n</table>\n<p>This matters because all <code>13,341</code> noncanonical training rows are anomalous. So I did not treat <code>other</code> as a third modeling family. Instead, I used it as a strong identity-based anomaly signal.</p>\n<h2>Feature engineering</h2>\n<p>The main feature groups are:</p>\n<ol>\n<li><strong>Identity and missingness</strong>: device fingerprint, missing identity fields, per-block missingness counts, and missingness patterns.</li>\n<li><strong>Schema integrity</strong>: expected model IDs and lengths, used to detect malformed or structurally inconsistent rows.</li>\n<li><strong>Physical consistency</strong>: measurements compared against ratings and limits, plus phase-sum and power-factor consistency checks.</li>\n<li><strong>Control compliance</strong>: whether the reported active or reactive control state matches the observed output.</li>\n<li><strong>Protection and curves</strong>: enter-service logic, trip blocks, Volt-Var / Volt-Watt / Watt-Var behavior, and frequency-droop features.</li>\n<li><strong>AC/DC consistency</strong>: agreement between AC and DC measurements, per-port values, and sign patterns.</li>\n<li><strong>Residual and scenario features</strong>: surrogate-model residuals and smoothed anomaly rates for recurring operating scenarios.</li>\n</ol>\n<h2>Hard overrides</h2>\n<p>I also used a small set of high-confidence hard overrides. They are simply anomaly signatures that were clean enough in training to justify overriding the learned model.</p>\n<p>The most important case is the noncanonical bucket itself. If the 5-field fingerprint does not exactly match one of the two canonical simulator identities, the row is assigned to <code>other</code>, which acts as an identity-based anomaly bucket rather than a learned family.</p>\n<p>Inside the canonical families, the overrides are limited to a few high-precision patterns such as:</p>\n<ul>\n<li>clear electrical envelope violations,</li>\n<li>large mismatches between enabled controls and measured output,</li>\n<li>rare metadata or AC/DC patterns,</li>\n<li>enter-service contradictions,</li>\n<li>producing power while already outside a must-trip region.</li>\n</ul>\n<p>An important implementation detail is that candidate overrides are audited on the training data. Only the most precise ones remain true hard overrides; the rest can still help as features without forcing the final prediction.</p>\n<h2>Family-specific models and ensembling</h2>\n<p>For <code>canon10</code> and <code>canon100</code>, I train separate models. This worked better than a single global classifier because the two canonical devices operate in different regimes and have different local patterns.</p>\n<p>The main model is <strong>XGBoost</strong> on the full semantic numeric feature set. The companion model is <strong>CatBoost</strong> on a narrower raw + categorical table containing selected raw numeric columns, raw identity fields, the full device fingerprint, and additional identity / missingness-engineered features.</p>\n<p>I also add family-specific regressors trained on normal rows to predict quantities such as active power, apparent power, reactive power, power factor, and current. Their residuals become new features for the final classifiers. For validation, I use both a primary 5-fold split based on <code>Id</code> and an audit split based on a hashed representation of the operating scenario.</p>\n<h2>Threshold optimization</h2>\n<p>Since the competition metric is F2, I did not use the default <code>0.5</code> threshold. For each canonical family, I search directly over out-of-fold probabilities to find the threshold that maximizes F2. I do the same for the XGBoost / CatBoost blend weight, and keep the family-specific operating point that performs well on both the primary and audit splits.</p>\n<h2>Tools used</h2>\n<p>To develop the solution, I used <a href=\"https://docs.astral.sh/uv/\" target=\"_blank\">uv</a>, <a href=\"https://github.com/astral-sh/ruff\" target=\"_blank\">ruff</a>, <a href=\"https://www.docker.com/\" target=\"_blank\">Docker</a>, and <a href=\"https://developers.openai.com/codex/cli\" target=\"_blank\">Codex CLI</a>. Most of the work was done on a MacBook Pro, without a GPU.</p>\n<h2>Reproducibility</h2>\n<p>To keep the results reproducible, I seeded the full pipeline and used a <a href=\"https://console.cloud.google.com/artifacts/docker/kaggle-images/us/gcr.io/python/sha256:02c72a7c98e5e0895056901d9c715d181cd30eae392491235dfea93e6d0de3ed\" target=\"_blank\">Kaggle official Docker image</a> for local development. In the GitHub repository, the solution can be reproduced by building the pinned Docker image and running the <code>run_docker.sh</code> script, which executes the code inside the container and writes the final <code>submission.csv</code>.</p>\n<h2>About me</h2>\n<p>I am an applied mathematician turned into software developer, currently based in Madrid. My interests are AI, machine learning, numerical simulation, and optimization.</p>\n<p>If you would like to reach out, you can find me on <a href=\"https://www.linkedin.com/in/alvaro-borras/\" target=\"_blank\">LinkedIn</a>, or via Gmail as <code>alvaroborrasf</code>.</p>\n<p>Thank you again and looking forward to seeing the approaches from other participants!</p>",
      "rawMarkdown": "I would like to thank the organizers for putting together the competition, and for addressing the early dataset issues so quickly.\n\nMost of the public leaderboard eventually converged around `0.9115` (public) / `0.9119` (private), so I was happy to finish a bit above that score.\n\nIt was a fun competition, and I learned a lot from it. Thank you!\n\n**TLDR**: [This notebook](https://www.kaggle.com/code/alvaroborras/anomaly-detection-cyber-der-final-submission) reproduces my final submission, and the full implementation is available in my [personal repository](https://github.com/alvaroborras/Anomaly-Detection-DER-Systems).\n\n# Overview of the solution\n\nThe dataset is highly structured, and in the original CSV files it is dominated by two exact identity tuples.\n\nMore concretely, I build a 5-field fingerprint from `common[0].Mn`, `common[0].Md`, `common[0].Opt`, `common[0].Vr`, and `common[0].SN`. The two dominant tuples correspond to the 10 kW and 100 kW simulator rows, which I refer to as `canon10` and `canon100`, and I use that fingerprint to split the data into `canon10`, `canon100`, and a small noncanonical bucket.\n\nThis distinction is very important for the rest of the pipeline. The two canonical families are modeled separately, while the noncanonical rows are handled outside the main learned-model path because they already form a very strong anomaly signal in training.\n\nFrom `665` selected source columns, I build semantic features that make the DER structure explicit: missingness and schema-integrity signals, measurement-vs-rating/setting consistency, control-compliance features, protection and curve features, and AC/DC consistency checks.\n\nFor each canonical family, the final predictor is a small ensemble of XGBoost on the semantic numeric table and CatBoost on a narrower raw + categorical table, with the blend and threshold tuned directly for F2.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5517090%2Fa5fd25af7c5dbf46f7b732c1ab815194%2Foverall_pipeline.png?generation=1775626772760115&alt=media)\n\n*High level view of the pipeline. Rows are first split by device fingerprint; the two canonical families go through family-specific feature engineering and models, while the small noncanonical bucket is handled separately.*\n\n## Observations about the dataset\n\nThe raw dataset has `723` features, but many are fully empty or not especially useful for the final model. In particular, `182` features are completely null in both train and test, while there are no fully empty rows. In practice, the useful signal is not in dropping rows, but in being selective about columns and preserving informative missingness.\n\nA second key observation is that the data is overwhelmingly concentrated around two exact fingerprints from the original CSV identity columns:\n\n- `('DERSec', 'DER Simulator', '10 kW DER', '1.2.3', 'SN-Three-Phase')`  ->  `canon10`\n- `('DERSec', 'DER Simulator 100 kW', '1.2.3.1', '1.0.0', '1100058974')`  ->  `canon100`\n\n| Bucket | Definition | Train rows | Test rows | How I use it |\n|---|---|---:|---:|---|\n| `canon10` | Exact match to the canonical 10 kW simulator fingerprint | 1,181,604 | 505,993 | Family-specific learned models + overrides |\n| `canon100` | Exact match to the canonical 100 kW simulator fingerprint | 1,168,218 | 501,106 | Family-specific learned models + overrides |\n| `other` / noncanonical | Anything else | 13,341 | 5,686 | Identity-based anomaly bucket |\n\nThis matters because all `13,341` noncanonical training rows are anomalous. So I did not treat `other` as a third modeling family. Instead, I used it as a strong identity-based anomaly signal.\n\n## Feature engineering\n\nThe main feature groups are:\n\n1. **Identity and missingness**: device fingerprint, missing identity fields, per-block missingness counts, and missingness patterns.\n2. **Schema integrity**: expected model IDs and lengths, used to detect malformed or structurally inconsistent rows.\n3. **Physical consistency**: measurements compared against ratings and limits, plus phase-sum and power-factor consistency checks.\n4. **Control compliance**: whether the reported active or reactive control state matches the observed output.\n5. **Protection and curves**: enter-service logic, trip blocks, Volt-Var / Volt-Watt / Watt-Var behavior, and frequency-droop features.\n6. **AC/DC consistency**: agreement between AC and DC measurements, per-port values, and sign patterns.\n7. **Residual and scenario features**: surrogate-model residuals and smoothed anomaly rates for recurring operating scenarios.\n\n## Hard overrides\n\nI also used a small set of high-confidence hard overrides. They are simply anomaly signatures that were clean enough in training to justify overriding the learned model.\n\nThe most important case is the noncanonical bucket itself. If the 5-field fingerprint does not exactly match one of the two canonical simulator identities, the row is assigned to `other`, which acts as an identity-based anomaly bucket rather than a learned family.\n\nInside the canonical families, the overrides are limited to a few high-precision patterns such as:\n\n- clear electrical envelope violations,\n- large mismatches between enabled controls and measured output,\n- rare metadata or AC/DC patterns,\n- enter-service contradictions,\n- producing power while already outside a must-trip region.\n\nAn important implementation detail is that candidate overrides are audited on the training data. Only the most precise ones remain true hard overrides; the rest can still help as features without forcing the final prediction.\n\n## Family-specific models and ensembling\n\nFor `canon10` and `canon100`, I train separate models. This worked better than a single global classifier because the two canonical devices operate in different regimes and have different local patterns.\n\nThe main model is **XGBoost** on the full semantic numeric feature set. The companion model is **CatBoost** on a narrower raw + categorical table containing selected raw numeric columns, raw identity fields, the full device fingerprint, and additional identity / missingness-engineered features.\n\nI also add family-specific regressors trained on normal rows to predict quantities such as active power, apparent power, reactive power, power factor, and current. Their residuals become new features for the final classifiers. For validation, I use both a primary 5-fold split based on `Id` and an audit split based on a hashed representation of the operating scenario.\n\n## Threshold optimization\n\nSince the competition metric is F2, I did not use the default `0.5` threshold. For each canonical family, I search directly over out-of-fold probabilities to find the threshold that maximizes F2. I do the same for the XGBoost / CatBoost blend weight, and keep the family-specific operating point that performs well on both the primary and audit splits.\n\n## Tools used\n\nTo develop the solution, I used [uv](https://docs.astral.sh/uv/), [ruff](https://github.com/astral-sh/ruff), [Docker](https://www.docker.com/), and [Codex CLI](https://developers.openai.com/codex/cli). Most of the work was done on a MacBook Pro, without a GPU.\n\n## Reproducibility\n\nTo keep the results reproducible, I seeded the full pipeline and used a [Kaggle official Docker image](https://console.cloud.google.com/artifacts/docker/kaggle-images/us/gcr.io/python/sha256:02c72a7c98e5e0895056901d9c715d181cd30eae392491235dfea93e6d0de3ed) for local development. In the GitHub repository, the solution can be reproduced by building the pinned Docker image and running the `run_docker.sh` script, which executes the code inside the container and writes the final `submission.csv`.\n\n## About me\n\nI am an applied mathematician turned into software developer, currently based in Madrid. My interests are AI, machine learning, numerical simulation, and optimization.\n\nIf you would like to reach out, you can find me on [LinkedIn](https://www.linkedin.com/in/alvaro-borras/), or via Gmail as `alvaroborrasf`.\n\nThank you again and looking forward to seeing the approaches from other participants!",
      "votes": null
    },
    {
      "id": "3438095",
      "postDate": "04/08/2026 16:55:32",
      "content": "<p>Great work!! Also congrats on the win </p>",
      "rawMarkdown": "Great work!! Also congrats on the win",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3438095,
      "author_name": "adwaittagalpallewar",
      "author_url": "",
      "post_date": "04/08/2026 16:55:32",
      "content": "<p>Great work!! Also congrats on the win </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3437808": "I would like to thank the organizers for putting together the competition, and for addressing the early dataset issues so quickly.\n\nMost of the public leaderboard eventually converged around `0.9115` (public) / `0.9119` (private), so I was happy to finish a bit above that score.\n\nIt was a fun competition, and I learned a lot from it. Thank you!\n\n**TLDR**: [This notebook](https://www.kaggle.com/code/alvaroborras/anomaly-detection-cyber-der-final-submission) reproduces my final submission, and the full implementation is available in my [personal repository](https://github.com/alvaroborras/Anomaly-Detection-DER-Systems).\n\n# Overview of the solution\n\nThe dataset is highly structured, and in the original CSV files it is dominated by two exact identity tuples.\n\nMore concretely, I build a 5-field fingerprint from `common[0].Mn`, `common[0].Md`, `common[0].Opt`, `common[0].Vr`, and `common[0].SN`. The two dominant tuples correspond to the 10 kW and 100 kW simulator rows, which I refer to as `canon10` and `canon100`, and I use that fingerprint to split the data into `canon10`, `canon100`, and a small noncanonical bucket.\n\nThis distinction is very important for the rest of the pipeline. The two canonical families are modeled separately, while the noncanonical rows are handled outside the main learned-model path because they already form a very strong anomaly signal in training.\n\nFrom `665` selected source columns, I build semantic features that make the DER structure explicit: missingness and schema-integrity signals, measurement-vs-rating/setting consistency, control-compliance features, protection and curve features, and AC/DC consistency checks.\n\nFor each canonical family, the final predictor is a small ensemble of XGBoost on the semantic numeric table and CatBoost on a narrower raw + categorical table, with the blend and threshold tuned directly for F2.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5517090%2Fa5fd25af7c5dbf46f7b732c1ab815194%2Foverall_pipeline.png?generation=1775626772760115&alt=media)\n\n*High level view of the pipeline. Rows are first split by device fingerprint; the two canonical families go through family-specific feature engineering and models, while the small noncanonical bucket is handled separately.*\n\n## Observations about the dataset\n\nThe raw dataset has `723` features, but many are fully empty or not especially useful for the final model. In particular, `182` features are completely null in both train and test, while there are no fully empty rows. In practice, the useful signal is not in dropping rows, but in being selective about columns and preserving informative missingness.\n\nA second key observation is that the data is overwhelmingly concentrated around two exact fingerprints from the original CSV identity columns:\n\n- `('DERSec', 'DER Simulator', '10 kW DER', '1.2.3', 'SN-Three-Phase')`  ->  `canon10`\n- `('DERSec', 'DER Simulator 100 kW', '1.2.3.1', '1.0.0', '1100058974')`  ->  `canon100`\n\n| Bucket | Definition | Train rows | Test rows | How I use it |\n|---|---|---:|---:|---|\n| `canon10` | Exact match to the canonical 10 kW simulator fingerprint | 1,181,604 | 505,993 | Family-specific learned models + overrides |\n| `canon100` | Exact match to the canonical 100 kW simulator fingerprint | 1,168,218 | 501,106 | Family-specific learned models + overrides |\n| `other` / noncanonical | Anything else | 13,341 | 5,686 | Identity-based anomaly bucket |\n\nThis matters because all `13,341` noncanonical training rows are anomalous. So I did not treat `other` as a third modeling family. Instead, I used it as a strong identity-based anomaly signal.\n\n## Feature engineering\n\nThe main feature groups are:\n\n1. **Identity and missingness**: device fingerprint, missing identity fields, per-block missingness counts, and missingness patterns.\n2. **Schema integrity**: expected model IDs and lengths, used to detect malformed or structurally inconsistent rows.\n3. **Physical consistency**: measurements compared against ratings and limits, plus phase-sum and power-factor consistency checks.\n4. **Control compliance**: whether the reported active or reactive control state matches the observed output.\n5. **Protection and curves**: enter-service logic, trip blocks, Volt-Var / Volt-Watt / Watt-Var behavior, and frequency-droop features.\n6. **AC/DC consistency**: agreement between AC and DC measurements, per-port values, and sign patterns.\n7. **Residual and scenario features**: surrogate-model residuals and smoothed anomaly rates for recurring operating scenarios.\n\n## Hard overrides\n\nI also used a small set of high-confidence hard overrides. They are simply anomaly signatures that were clean enough in training to justify overriding the learned model.\n\nThe most important case is the noncanonical bucket itself. If the 5-field fingerprint does not exactly match one of the two canonical simulator identities, the row is assigned to `other`, which acts as an identity-based anomaly bucket rather than a learned family.\n\nInside the canonical families, the overrides are limited to a few high-precision patterns such as:\n\n- clear electrical envelope violations,\n- large mismatches between enabled controls and measured output,\n- rare metadata or AC/DC patterns,\n- enter-service contradictions,\n- producing power while already outside a must-trip region.\n\nAn important implementation detail is that candidate overrides are audited on the training data. Only the most precise ones remain true hard overrides; the rest can still help as features without forcing the final prediction.\n\n## Family-specific models and ensembling\n\nFor `canon10` and `canon100`, I train separate models. This worked better than a single global classifier because the two canonical devices operate in different regimes and have different local patterns.\n\nThe main model is **XGBoost** on the full semantic numeric feature set. The companion model is **CatBoost** on a narrower raw + categorical table containing selected raw numeric columns, raw identity fields, the full device fingerprint, and additional identity / missingness-engineered features.\n\nI also add family-specific regressors trained on normal rows to predict quantities such as active power, apparent power, reactive power, power factor, and current. Their residuals become new features for the final classifiers. For validation, I use both a primary 5-fold split based on `Id` and an audit split based on a hashed representation of the operating scenario.\n\n## Threshold optimization\n\nSince the competition metric is F2, I did not use the default `0.5` threshold. For each canonical family, I search directly over out-of-fold probabilities to find the threshold that maximizes F2. I do the same for the XGBoost / CatBoost blend weight, and keep the family-specific operating point that performs well on both the primary and audit splits.\n\n## Tools used\n\nTo develop the solution, I used [uv](https://docs.astral.sh/uv/), [ruff](https://github.com/astral-sh/ruff), [Docker](https://www.docker.com/), and [Codex CLI](https://developers.openai.com/codex/cli). Most of the work was done on a MacBook Pro, without a GPU.\n\n## Reproducibility\n\nTo keep the results reproducible, I seeded the full pipeline and used a [Kaggle official Docker image](https://console.cloud.google.com/artifacts/docker/kaggle-images/us/gcr.io/python/sha256:02c72a7c98e5e0895056901d9c715d181cd30eae392491235dfea93e6d0de3ed) for local development. In the GitHub repository, the solution can be reproduced by building the pinned Docker image and running the `run_docker.sh` script, which executes the code inside the container and writes the final `submission.csv`.\n\n## About me\n\nI am an applied mathematician turned into software developer, currently based in Madrid. My interests are AI, machine learning, numerical simulation, and optimization.\n\nIf you would like to reach out, you can find me on [LinkedIn](https://www.linkedin.com/in/alvaro-borras/), or via Gmail as `alvaroborrasf`.\n\nThank you again and looking forward to seeing the approaches from other participants!",
    "3438095": "Great work!! Also congrats on the win"
  },
  "source": "meta"
}