{
  "id": 327649,
  "title": "The data has uniform random noise injected",
  "url": "/competitions/amex-default-prediction/discussion/327649",
  "author_name": "raddar",
  "post_date": "2022-05-28T11:34:42.635000",
  "votes": 82,
  "comment_count": 10,
  "views": 0,
  "content": "<p>The organizers did a good job to properly anonymize and normalize the data. In previous competitions I was able to de-anonymize features. This time I think it is impossible (although not giving up yet!)</p>\n<p>Some findings about that can be found in my notebook:</p>\n<p><a href=\"https://www.kaggle.com/raddar/the-data-has-random-uniform-noise-added\" target=\"_blank\">https://www.kaggle.com/raddar/the-data-has-random-uniform-noise-added</a></p>",
  "messages": [
    {
      "id": 1803959,
      "postDate": "2022-05-28T11:34:42.637Z",
      "content": "<p>The organizers did a good job to properly anonymize and normalize the data. In previous competitions I was able to de-anonymize features. This time I think it is impossible (although not giving up yet!)</p>\n<p>Some findings about that can be found in my notebook:</p>\n<p><a href=\"https://www.kaggle.com/raddar/the-data-has-random-uniform-noise-added\" target=\"_blank\">https://www.kaggle.com/raddar/the-data-has-random-uniform-noise-added</a></p>",
      "rawMarkdown": "The organizers did a good job to properly anonymize and normalize the data. In previous competitions I was able to de-anonymize features. This time I think it is impossible (although not giving up yet!)\n\nSome findings about that can be found in my notebook:\n\nhttps://www.kaggle.com/raddar/the-data-has-random-uniform-noise-added",
      "votes": 81
    },
    {
      "id": 1804040,
      "postDate": "2022-05-28T13:26:39.263Z",
      "content": "<p>It looks like they used the following function to anonymize B_2:</p>\n<pre><code>def anonymize(data):\n    data -= data.min()\n    data /= data.max()\n    data = data.round(2)\n    rng = np.random.default_rng()\n    data += rng.uniform(0, 0.01, len(data))\n    return data\n</code></pre>\n<p>We cannot invert the <code>round(2)</code>, but we can remove the uniform noise. Ideally, we'd remove the noise before converting to <code>float16</code>, and then we should multiply by 100 and convert to <code>int8</code>. </p>",
      "rawMarkdown": "It looks like they used the following function to anonymize B_2:\n\n```\ndef anonymize(data):\n    data -= data.min()\n    data /= data.max()\n    data = data.round(2)\n    rng = np.random.default_rng()\n    data += rng.uniform(0, 0.01, len(data))\n    return data\n```\n\nWe cannot invert the `round(2)`, but we can remove the uniform noise. Ideally, we'd remove the noise before converting to `float16`, and then we should multiply by 100 and convert to `int8`. ",
      "votes": 10,
      "replies": [
        {
          "id": 1804054,
          "postDate": "2022-05-28T13:39:59.700Z",
          "content": "<p>x['B_2'] = x['B_2'].apply(lambda t: np.floor(t*100))</p>",
          "rawMarkdown": "x['B_2'] = x['B_2'].apply(lambda t: np.floor(t*100))",
          "votes": 7
        }
      ]
    },
    {
      "id": 1803977,
      "postDate": "2022-05-28T11:50:48.667Z",
      "content": "<p>Great observation. I posted a similar observation at the same time <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651\" target=\"_blank\">here</a>. Your hypothesis of uniform random noise makes sense. (Especially since columns related to money should show less unique values as you point out).</p>",
      "rawMarkdown": "Great observation. I posted a similar observation at the same time [here][1]. Your hypothesis of uniform random noise makes sense. (Especially since columns related to money should show less unique values as you point out).\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651",
      "votes": 6,
      "replies": [
        {
          "id": 1803984,
          "postDate": "2022-05-28T11:56:47.127Z",
          "content": "<p>I guess a good strategy will be to squeeze this random noise to reduce cardinality of the features. hopefully useful for ML models</p>",
          "rawMarkdown": "I guess a good strategy will be to squeeze this random noise to reduce cardinality of the features. hopefully useful for ML models",
          "votes": 7,
          "replies": [
            {
              "id": 1812752,
              "postDate": "2022-06-06T07:04:24.180Z",
              "content": "<blockquote>\n  <p>I guess a good strategy will be to squeeze this random noise to reduce cardinality of the features. hopefully useful for ML models</p>\n</blockquote>\n<p>I think reducing cardinality of the features is base by restoring noisy caused different data to consistent data，is that correct?👀</p>",
              "rawMarkdown": "> I guess a good strategy will be to squeeze this random noise to reduce cardinality of the features. hopefully useful for ML models\n\nI think reducing cardinality of the features is base by restoring noisy caused different data to consistent data，is that correct?👀"
            }
          ]
        }
      ]
    },
    {
      "id": 1804980,
      "postDate": "2022-05-29T16:13:47.063Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> ,<br>\nIn most big BFSI banks ( including Amex), it is a standard practice to anonymize data before modelling due to Data privacy norms.<br>\nGreat observation though .<br>\nThanks</p>",
      "rawMarkdown": "Hi @raddar ,\nIn most big BFSI banks ( including Amex), it is a standard practice to anonymize data before modelling due to Data privacy norms.\nGreat observation though .\nThanks",
      "replies": [
        {
          "id": 1805827,
          "postDate": "2022-05-30T14:34:55.250Z",
          "content": "<p>Actually what raddar has mentioned here is adding random noise to the normalized dataset.</p>\n<p>Agreed anonymization is a common practice followed by most of the organizations, But adding noise is more of a Kaggle and public data sharing thing.</p>",
          "rawMarkdown": "Actually what raddar has mentioned here is adding random noise to the normalized dataset.\n\nAgreed anonymization is a common practice followed by most of the organizations, But adding noise is more of a Kaggle and public data sharing thing.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1838137,
      "postDate": "2022-06-30T08:43:40.857Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1831994,
      "postDate": "2022-06-24T15:13:48.003Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1832228,
          "postDate": "2022-06-24T20:18:03.260Z",
          "content": "<p>you can just read my other notebooks from other competitions</p>",
          "rawMarkdown": "you can just read my other notebooks from other competitions"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1804040,
      "author_name": "AmbrosM",
      "author_url": "",
      "post_date": "2022-05-28T13:26:39.263000",
      "content": "<p>It looks like they used the following function to anonymize B_2:</p>\n<pre><code>def anonymize(data):\n    data -= data.min()\n    data /= data.max()\n    data = data.round(2)\n    rng = np.random.default_rng()\n    data += rng.uniform(0, 0.01, len(data))\n    return data\n</code></pre>\n<p>We cannot invert the <code>round(2)</code>, but we can remove the uniform noise. Ideally, we'd remove the noise before converting to <code>float16</code>, and then we should multiply by 100 and convert to <code>int8</code>. </p>",
      "votes": 10,
      "replies": [
        {
          "id": 1804054,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-05-28T13:39:59.700000",
          "content": "<p>x['B_2'] = x['B_2'].apply(lambda t: np.floor(t*100))</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1803977,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-05-28T11:50:48.667000",
      "content": "<p>Great observation. I posted a similar observation at the same time <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651\" target=\"_blank\">here</a>. Your hypothesis of uniform random noise makes sense. (Especially since columns related to money should show less unique values as you point out).</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1803984,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-05-28T11:56:47.127000",
          "content": "<p>I guess a good strategy will be to squeeze this random noise to reduce cardinality of the features. hopefully useful for ML models</p>",
          "votes": 7,
          "replies": [
            {
              "id": 1812752,
              "author_name": "Dou Fan",
              "author_url": "",
              "post_date": "2022-06-06T07:04:24.180000",
              "content": "<blockquote>\n  <p>I guess a good strategy will be to squeeze this random noise to reduce cardinality of the features. hopefully useful for ML models</p>\n</blockquote>\n<p>I think reducing cardinality of the features is base by restoring noisy caused different data to consistent data，is that correct?👀</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1804980,
      "author_name": "KritiDoneria",
      "author_url": "",
      "post_date": "2022-05-29T16:13:47.063000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> ,<br>\nIn most big BFSI banks ( including Amex), it is a standard practice to anonymize data before modelling due to Data privacy norms.<br>\nGreat observation though .<br>\nThanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1805827,
          "author_name": "Seeker",
          "author_url": "",
          "post_date": "2022-05-30T14:34:55.250000",
          "content": "<p>Actually what raddar has mentioned here is adding random noise to the normalized dataset.</p>\n<p>Agreed anonymization is a common practice followed by most of the organizations, But adding noise is more of a Kaggle and public data sharing thing.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1838137,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-30T08:43:40.857000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1831994,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-24T15:13:48.003000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1832228,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-24T20:18:03.260000",
          "content": "<p>you can just read my other notebooks from other competitions</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1803959": "The organizers did a good job to properly anonymize and normalize the data. In previous competitions I was able to de-anonymize features. This time I think it is impossible (although not giving up yet!)\n\nSome findings about that can be found in my notebook:\n\nhttps://www.kaggle.com/raddar/the-data-has-random-uniform-noise-added",
    "1804040": "It looks like they used the following function to anonymize B_2:\n\n```\ndef anonymize(data):\n    data -= data.min()\n    data /= data.max()\n    data = data.round(2)\n    rng = np.random.default_rng()\n    data += rng.uniform(0, 0.01, len(data))\n    return data\n```\n\nWe cannot invert the `round(2)`, but we can remove the uniform noise. Ideally, we'd remove the noise before converting to `float16`, and then we should multiply by 100 and convert to `int8`. ",
    "1803977": "Great observation. I posted a similar observation at the same time [here][1]. Your hypothesis of uniform random noise makes sense. (Especially since columns related to money should show less unique values as you point out).\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651",
    "1804980": "Hi @raddar ,\nIn most big BFSI banks ( including Amex), it is a standard practice to anonymize data before modelling due to Data privacy norms.\nGreat observation though .\nThanks",
    "1838137": "",
    "1831994": ""
  }
}