{
  "id": 252249,
  "title": "what is your best input normalization method ?",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/252249",
  "author_name": "",
  "post_date": "2021-07-11T09:58:08.312317800Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I divide integer information and float information<br>\ni. convert int infomation to  one hot format ( if nan, all zero), <br>\nii. make single float information as four types, normalized on each columns of train data.</p>\n<ol>\n<li>percentile normalization</li>\n<li>min = 0 &amp; max = 1 normaliztion</li>\n<li>gaussian cdf normalization</li>\n<li>\"is not null\"</li>\n</ol>\n<p>and I compare with pure input vs normalized input, </p>\n<p>validation result : 1.44(lightgbm) and 1.91(lightgbm)</p>\n<p>validation result on normalized input, 1.55 reached by GRU, 1.65 by transformers</p>\n<p>LB on normalized input was 1.3400(GRU), error (LSTM), 1.3914(transformers)</p>\n<p>I believe transformers is one of the key to reach 1.25<br>\nbut still GRU based model much better than transformers for me.</p>\n<p>is there better way to normalize?</p>",
  "messages": [
    {
      "id": "1383884",
      "postDate": "07/11/2021 09:58:08",
      "content": "<p>I divide integer information and float information<br>\ni. convert int infomation to  one hot format ( if nan, all zero), <br>\nii. make single float information as four types, normalized on each columns of train data.</p>\n<ol>\n<li>percentile normalization</li>\n<li>min = 0 &amp; max = 1 normaliztion</li>\n<li>gaussian cdf normalization</li>\n<li>\"is not null\"</li>\n</ol>\n<p>and I compare with pure input vs normalized input, </p>\n<p>validation result : 1.44(lightgbm) and 1.91(lightgbm)</p>\n<p>validation result on normalized input, 1.55 reached by GRU, 1.65 by transformers</p>\n<p>LB on normalized input was 1.3400(GRU), error (LSTM), 1.3914(transformers)</p>\n<p>I believe transformers is one of the key to reach 1.25<br>\nbut still GRU based model much better than transformers for me.</p>\n<p>is there better way to normalize?</p>",
      "rawMarkdown": "I divide integer information and float information\ni. convert int infomation to  one hot format ( if nan, all zero), \nii. make single float information as four types, normalized on each columns of train data.\n1. percentile normalization\n2. min = 0 & max = 1 normaliztion\n3. gaussian cdf normalization\n4. \"is not null\"\n\nand I compare with pure input vs normalized input, \n\nvalidation result : 1.44(lightgbm) and 1.91(lightgbm)\n\nvalidation result on normalized input, 1.55 reached by GRU, 1.65 by transformers\n\nLB on normalized input was 1.3400(GRU), error (LSTM), 1.3914(transformers)\n\nI believe transformers is one of the key to reach 1.25\nbut still GRU based model much better than transformers for me.\n\nis there better way to normalize?",
      "votes": null
    },
    {
      "id": "1383909",
      "postDate": "07/11/2021 10:21:26",
      "content": "<p>I am usinng \"is not null\" and standard scaling normalitzation. How do you deal with Nans on rosters data and standings data on private set? </p>",
      "rawMarkdown": "I am usinng \"is not null\" and standard scaling normalitzation. How do you deal with Nans on rosters data and standings data on private set?",
      "votes": null
    },
    {
      "id": "1383912",
      "postDate": "07/11/2021 10:26:45",
      "content": "<p><a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> train with dropout, but it makes validation and also LB lower (w/ , w/o 0.001, 0.1, 0.5 dropout in validation phase both increase loss)</p>",
      "rawMarkdown": "enric1296 train with dropout, but it makes validation and also LB lower (w/ , w/o 0.001, 0.1, 0.5 dropout in validation phase both increase loss)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1383909,
      "author_name": "enric1296",
      "author_url": "",
      "post_date": "07/11/2021 10:21:26",
      "content": "<p>I am usinng \"is not null\" and standard scaling normalitzation. How do you deal with Nans on rosters data and standings data on private set? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1383912,
          "author_name": "assign",
          "author_url": "",
          "post_date": "07/11/2021 10:26:45",
          "content": "<p><a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> train with dropout, but it makes validation and also LB lower (w/ , w/o 0.001, 0.1, 0.5 dropout in validation phase both increase loss)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1383884": "I divide integer information and float information\ni. convert int infomation to  one hot format ( if nan, all zero), \nii. make single float information as four types, normalized on each columns of train data.\n1. percentile normalization\n2. min = 0 & max = 1 normaliztion\n3. gaussian cdf normalization\n4. \"is not null\"\n\nand I compare with pure input vs normalized input, \n\nvalidation result : 1.44(lightgbm) and 1.91(lightgbm)\n\nvalidation result on normalized input, 1.55 reached by GRU, 1.65 by transformers\n\nLB on normalized input was 1.3400(GRU), error (LSTM), 1.3914(transformers)\n\nI believe transformers is one of the key to reach 1.25\nbut still GRU based model much better than transformers for me.\n\nis there better way to normalize?",
    "1383909": "I am usinng \"is not null\" and standard scaling normalitzation. How do you deal with Nans on rosters data and standings data on private set?",
    "1383912": "enric1296 train with dropout, but it makes validation and also LB lower (w/ , w/o 0.001, 0.1, 0.5 dropout in validation phase both increase loss)"
  },
  "source": "meta"
}