{
  "id": 467372,
  "title": "There are many outliers in the data",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467372",
  "author_name": "",
  "post_date": "2024-01-12T07:39:23.596693300Z",
  "votes": 12,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello everybody!<br>\nI combined all data from spectrograms and eegs into two complete dataframes, as a result I got the following:</p>\n<ol>\n<li>According to the spectrogram data:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F89d033650a863af4b4b4a8956443baf1%2F3.jpeg?generation=1705044111476480&amp;alt=media\"><br>\n2.According to eegs data<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F36e144a9d7a5ef2dc25b736aeef3afa0%2F4.jpeg?generation=1705044276991148&amp;alt=media\"></li>\n</ol>\n<p>As can be seen from the descriptions, the maximum values for both samples are several orders of magnitude higher than the average and the 75% quantile values. As a result, when trying to normalize the data, these maximum values will unnecessarily compress the underlying data to 0. I tried to solve the problem using IsolationForest. Here are the results</p>\n<ol>\n<li>According to the spectrogram data:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2Fb1ebefa65a204bb6435ef0851c6f8344%2F5.jpeg?generation=1705044740517896&amp;alt=media\"><br>\n2.According to eegs data<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F724c28296235e291e8dd87862ddb790d%2F6.jpeg?generation=1705044753383534&amp;alt=media\"></li>\n</ol>\n<p>This improved the results and made it possible to use MinMaxscaler, but using IsolationForest(contamination=0.1) , namely contamination=0.1 was a mistake, as it turned out, this approach cuts off about 3% of samples completely, it can be used but in auto mode, then significantly fewer samples will be cut off , but MinMaxscaler will also work worse.</p>\n<p>What other options do you have for working with peak values?</p>",
  "messages": [
    {
      "id": "2598193",
      "postDate": "01/12/2024 07:39:23",
      "content": "<p>Hello everybody!<br>\nI combined all data from spectrograms and eegs into two complete dataframes, as a result I got the following:</p>\n<ol>\n<li>According to the spectrogram data:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F89d033650a863af4b4b4a8956443baf1%2F3.jpeg?generation=1705044111476480&amp;alt=media\"><br>\n2.According to eegs data<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F36e144a9d7a5ef2dc25b736aeef3afa0%2F4.jpeg?generation=1705044276991148&amp;alt=media\"></li>\n</ol>\n<p>As can be seen from the descriptions, the maximum values for both samples are several orders of magnitude higher than the average and the 75% quantile values. As a result, when trying to normalize the data, these maximum values will unnecessarily compress the underlying data to 0. I tried to solve the problem using IsolationForest. Here are the results</p>\n<ol>\n<li>According to the spectrogram data:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2Fb1ebefa65a204bb6435ef0851c6f8344%2F5.jpeg?generation=1705044740517896&amp;alt=media\"><br>\n2.According to eegs data<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F724c28296235e291e8dd87862ddb790d%2F6.jpeg?generation=1705044753383534&amp;alt=media\"></li>\n</ol>\n<p>This improved the results and made it possible to use MinMaxscaler, but using IsolationForest(contamination=0.1) , namely contamination=0.1 was a mistake, as it turned out, this approach cuts off about 3% of samples completely, it can be used but in auto mode, then significantly fewer samples will be cut off , but MinMaxscaler will also work worse.</p>\n<p>What other options do you have for working with peak values?</p>",
      "rawMarkdown": "Hello everybody!\nI combined all data from spectrograms and eegs into two complete dataframes, as a result I got the following:\n1. According to the spectrogram data:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F89d033650a863af4b4b4a8956443baf1%2F3.jpeg?generation=1705044111476480&alt=media)\n2.According to eegs data\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F36e144a9d7a5ef2dc25b736aeef3afa0%2F4.jpeg?generation=1705044276991148&alt=media)\n\nAs can be seen from the descriptions, the maximum values for both samples are several orders of magnitude higher than the average and the 75% quantile values. As a result, when trying to normalize the data, these maximum values will unnecessarily compress the underlying data to 0. I tried to solve the problem using IsolationForest. Here are the results\n1. According to the spectrogram data:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2Fb1ebefa65a204bb6435ef0851c6f8344%2F5.jpeg?generation=1705044740517896&alt=media)\n2.According to eegs data\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F724c28296235e291e8dd87862ddb790d%2F6.jpeg?generation=1705044753383534&alt=media)\n\nThis improved the results and made it possible to use MinMaxscaler, but using IsolationForest(contamination=0.1) , namely contamination=0.1 was a mistake, as it turned out, this approach cuts off about 3% of samples completely, it can be used but in auto mode, then significantly fewer samples will be cut off , but MinMaxscaler will also work worse.\n\nWhat other options do you have for working with peak values?",
      "votes": null
    },
    {
      "id": "2598204",
      "postDate": "01/12/2024 07:46:31",
      "content": "<p>Do we know that multiple EEGs are on the same scale? Maybe instance norm makes more sense here.</p>",
      "rawMarkdown": "Do we know that multiple EEGs are on the same scale? Maybe instance norm makes more sense here.",
      "votes": null
    },
    {
      "id": "2598223",
      "postDate": "01/12/2024 08:08:15",
      "content": "<p>It's a little risky, but needs to be tested, but again, sensor recordings are made over a period of time, and the emissions can be (according to the description) random factors like people's movements or some reactions, it needs to be taken into account somehow</p>",
      "rawMarkdown": "It's a little risky, but needs to be tested, but again, sensor recordings are made over a period of time, and the emissions can be (according to the description) random factors like people's movements or some reactions, it needs to be taken into account somehow",
      "votes": null
    },
    {
      "id": "2614588",
      "postDate": "01/22/2024 17:53:34",
      "content": "<p>Hi. Is it possible that you can also provide the relevant notebook ? Or if you have already posted then provide the link. </p>",
      "rawMarkdown": "Hi. Is it possible that you can also provide the relevant notebook ? Or if you have already posted then provide the link.",
      "votes": null
    },
    {
      "id": "2615751",
      "postDate": "01/23/2024 09:44:22",
      "content": "<p>Hi. <br>\nI've already dropped that idea, now all public notebooks use sample normalization. But here is the link</p>\n<p><a href=\"https://www.kaggle.com/datasets/aikhmelnytskyy/hms-scalers\" target=\"_blank\">https://www.kaggle.com/datasets/aikhmelnytskyy/hms-scalers</a> - scalers </p>\n<h1>Specify the path to the CSV file containing training data</h1>\n<p>train_csv_path = '/kaggle/input/hms-harmful-brain-activity-classification/train.csv'</p>\n<h1>Read the training data from the CSV file into a DataFrame</h1>\n<p>train = pd.read_csv(train_csv_path)</p>\n<h1>Specify paths to the directories containing EEG and spectrogram data</h1>\n<p>train_eegs = \"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/\"<br>\ntrain_spectrograms = \"/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/\"</p>\n<h1>Create new columns in the DataFrame to store full paths to EEG and spectrogram files</h1>\n<p>train['train_eegs'] = train_eegs + train['eeg_id'].astype(str) + '.parquet'<br>\ntrain['train_spectrograms'] = train_spectrograms + train['spectrogram_id'].astype(str) + '.parquet'</p>\n<p>train_eegs=train.drop_duplicates(subset='eeg_id').reset_index(drop=True)<br>\ndfs = []<br>\nfor file_path in tqdm( train_eegs['train_eegs'], desc=\"Loading files\"):<br>\n    df = pd.read_parquet(file_path)<br>\n    dfs.append(df)</p>\n<p>merged_df = pd.concat(dfs, axis=0, ignore_index=True)</p>\n<p>This is a sample of how I combine samples. Everything is simple, I don't even have anything to publish there. The rest of the code is posted above.</p>",
      "rawMarkdown": "Hi. \nI've already dropped that idea, now all public notebooks use sample normalization. But here is the link\n\nhttps://www.kaggle.com/datasets/aikhmelnytskyy/hms-scalers - scalers \n\n\n# Specify the path to the CSV file containing training data\ntrain_csv_path = '/kaggle/input/hms-harmful-brain-activity-classification/train.csv'\n\n# Read the training data from the CSV file into a DataFrame\ntrain = pd.read_csv(train_csv_path)\n# Specify paths to the directories containing EEG and spectrogram data\ntrain_eegs = \"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/\"\ntrain_spectrograms = \"/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/\"\n\n# Create new columns in the DataFrame to store full paths to EEG and spectrogram files\ntrain['train_eegs'] = train_eegs + train['eeg_id'].astype(str) + '.parquet'\ntrain['train_spectrograms'] = train_spectrograms + train['spectrogram_id'].astype(str) + '.parquet'\n\ntrain_eegs=train.drop_duplicates(subset='eeg_id').reset_index(drop=True)\ndfs = []\nfor file_path in tqdm( train_eegs['train_eegs'], desc=\"Loading files\"):\n    df = pd.read_parquet(file_path)\n    dfs.append(df)\n\nmerged_df = pd.concat(dfs, axis=0, ignore_index=True)\n\n\nThis is a sample of how I combine samples. Everything is simple, I don't even have anything to publish there. The rest of the code is posted above.",
      "votes": null
    },
    {
      "id": "2615849",
      "postDate": "01/23/2024 11:05:54",
      "content": "<p>Thanks for the explanation.</p>",
      "rawMarkdown": "Thanks for the explanation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2598204,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "01/12/2024 07:46:31",
      "content": "<p>Do we know that multiple EEGs are on the same scale? Maybe instance norm makes more sense here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2598223,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "01/12/2024 08:08:15",
          "content": "<p>It's a little risky, but needs to be tested, but again, sensor recordings are made over a period of time, and the emissions can be (according to the description) random factors like people's movements or some reactions, it needs to be taken into account somehow</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2614588,
      "author_name": "jamshaidsohail5",
      "author_url": "",
      "post_date": "01/22/2024 17:53:34",
      "content": "<p>Hi. Is it possible that you can also provide the relevant notebook ? Or if you have already posted then provide the link. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2615751,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "01/23/2024 09:44:22",
          "content": "<p>Hi. <br>\nI've already dropped that idea, now all public notebooks use sample normalization. But here is the link</p>\n<p><a href=\"https://www.kaggle.com/datasets/aikhmelnytskyy/hms-scalers\" target=\"_blank\">https://www.kaggle.com/datasets/aikhmelnytskyy/hms-scalers</a> - scalers </p>\n<h1>Specify the path to the CSV file containing training data</h1>\n<p>train_csv_path = '/kaggle/input/hms-harmful-brain-activity-classification/train.csv'</p>\n<h1>Read the training data from the CSV file into a DataFrame</h1>\n<p>train = pd.read_csv(train_csv_path)</p>\n<h1>Specify paths to the directories containing EEG and spectrogram data</h1>\n<p>train_eegs = \"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/\"<br>\ntrain_spectrograms = \"/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/\"</p>\n<h1>Create new columns in the DataFrame to store full paths to EEG and spectrogram files</h1>\n<p>train['train_eegs'] = train_eegs + train['eeg_id'].astype(str) + '.parquet'<br>\ntrain['train_spectrograms'] = train_spectrograms + train['spectrogram_id'].astype(str) + '.parquet'</p>\n<p>train_eegs=train.drop_duplicates(subset='eeg_id').reset_index(drop=True)<br>\ndfs = []<br>\nfor file_path in tqdm( train_eegs['train_eegs'], desc=\"Loading files\"):<br>\n    df = pd.read_parquet(file_path)<br>\n    dfs.append(df)</p>\n<p>merged_df = pd.concat(dfs, axis=0, ignore_index=True)</p>\n<p>This is a sample of how I combine samples. Everything is simple, I don't even have anything to publish there. The rest of the code is posted above.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2615849,
              "author_name": "jamshaidsohail5",
              "author_url": "",
              "post_date": "01/23/2024 11:05:54",
              "content": "<p>Thanks for the explanation.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2598193": "Hello everybody!\nI combined all data from spectrograms and eegs into two complete dataframes, as a result I got the following:\n1. According to the spectrogram data:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F89d033650a863af4b4b4a8956443baf1%2F3.jpeg?generation=1705044111476480&alt=media)\n2.According to eegs data\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F36e144a9d7a5ef2dc25b736aeef3afa0%2F4.jpeg?generation=1705044276991148&alt=media)\n\nAs can be seen from the descriptions, the maximum values for both samples are several orders of magnitude higher than the average and the 75% quantile values. As a result, when trying to normalize the data, these maximum values will unnecessarily compress the underlying data to 0. I tried to solve the problem using IsolationForest. Here are the results\n1. According to the spectrogram data:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2Fb1ebefa65a204bb6435ef0851c6f8344%2F5.jpeg?generation=1705044740517896&alt=media)\n2.According to eegs data\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F724c28296235e291e8dd87862ddb790d%2F6.jpeg?generation=1705044753383534&alt=media)\n\nThis improved the results and made it possible to use MinMaxscaler, but using IsolationForest(contamination=0.1) , namely contamination=0.1 was a mistake, as it turned out, this approach cuts off about 3% of samples completely, it can be used but in auto mode, then significantly fewer samples will be cut off , but MinMaxscaler will also work worse.\n\nWhat other options do you have for working with peak values?",
    "2598204": "Do we know that multiple EEGs are on the same scale? Maybe instance norm makes more sense here.",
    "2598223": "It's a little risky, but needs to be tested, but again, sensor recordings are made over a period of time, and the emissions can be (according to the description) random factors like people's movements or some reactions, it needs to be taken into account somehow",
    "2614588": "Hi. Is it possible that you can also provide the relevant notebook ? Or if you have already posted then provide the link.",
    "2615751": "Hi. \nI've already dropped that idea, now all public notebooks use sample normalization. But here is the link\n\nhttps://www.kaggle.com/datasets/aikhmelnytskyy/hms-scalers - scalers \n\n\n# Specify the path to the CSV file containing training data\ntrain_csv_path = '/kaggle/input/hms-harmful-brain-activity-classification/train.csv'\n\n# Read the training data from the CSV file into a DataFrame\ntrain = pd.read_csv(train_csv_path)\n# Specify paths to the directories containing EEG and spectrogram data\ntrain_eegs = \"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/\"\ntrain_spectrograms = \"/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/\"\n\n# Create new columns in the DataFrame to store full paths to EEG and spectrogram files\ntrain['train_eegs'] = train_eegs + train['eeg_id'].astype(str) + '.parquet'\ntrain['train_spectrograms'] = train_spectrograms + train['spectrogram_id'].astype(str) + '.parquet'\n\ntrain_eegs=train.drop_duplicates(subset='eeg_id').reset_index(drop=True)\ndfs = []\nfor file_path in tqdm( train_eegs['train_eegs'], desc=\"Loading files\"):\n    df = pd.read_parquet(file_path)\n    dfs.append(df)\n\nmerged_df = pd.concat(dfs, axis=0, ignore_index=True)\n\n\nThis is a sample of how I combine samples. Everything is simple, I don't even have anything to publish there. The rest of the code is posted above.",
    "2615849": "Thanks for the explanation."
  },
  "source": "meta"
}