{
  "id": 253101,
  "title": "Autoencoder for solving the problem",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/253101",
  "author_name": "",
  "post_date": "2021-07-15T03:09:20.846960Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am trying to use a Autoencoder for anomaly detection but the data is given in such a complicated manner that I find it impossible to create the required databases.   I must admit that my abilities are lacking in data manipulation, for I a beginner at this stuff.   I would probably try a CNN first, but believe that a LSTM Autoencoder is probably the correct choice.   Has anyone else tried an CNN or LSTM Autoencoder?   I believe that this machine learning tool is the correct one to solve the problem.</p>",
  "messages": [
    {
      "id": "1388514",
      "postDate": "07/15/2021 03:09:20",
      "content": "<p>I am trying to use a Autoencoder for anomaly detection but the data is given in such a complicated manner that I find it impossible to create the required databases.   I must admit that my abilities are lacking in data manipulation, for I a beginner at this stuff.   I would probably try a CNN first, but believe that a LSTM Autoencoder is probably the correct choice.   Has anyone else tried an CNN or LSTM Autoencoder?   I believe that this machine learning tool is the correct one to solve the problem.</p>",
      "rawMarkdown": "I am trying to use a Autoencoder for anomaly detection but the data is given in such a complicated manner that I find it impossible to create the required databases.   I must admit that my abilities are lacking in data manipulation, for I a beginner at this stuff.   I would probably try a CNN first, but believe that a LSTM Autoencoder is probably the correct choice.   Has anyone else tried an CNN or LSTM Autoencoder?   I believe that this machine learning tool is the correct one to solve the problem.",
      "votes": null
    },
    {
      "id": "1388804",
      "postDate": "07/15/2021 08:37:17",
      "content": "<blockquote>\n  <p>data is given in such a complicated manner that I find it impossible to create the required databases</p>\n</blockquote>\n<p>What do you find complicated?  Loading one sample is rather easy, you'll find many implementations in public notebooks.  Here is one for instance:</p>\n<pre><code>from pathlib import Path\ninput_path = Path('../input/')\ntrain_path = input_path / 'train'\n\ndef load_data(id, data_path):\n    data = np.load(data_path / id[0] / id[1] / id[2] / (id + '.npy')).astype('float32')\n    return data\n</code></pre>\n<p>I cast in np.float32 but you can cast into anything that matters, for instance to torch tensors:</p>\n<pre><code>def load_data(id, data_path):\n    data = np.load(data_path / id[0] / id[1] / id[2] / (id + '.npy'))\n    data = torch.from_numpy(data).float()\n    return data\n</code></pre>",
      "rawMarkdown": "> data is given in such a complicated manner that I find it impossible to create the required databases\n\nWhat do you find complicated?  Loading one sample is rather easy, you'll find many implementations in public notebooks.  Here is one for instance:\n\n```\nfrom pathlib import Path\ninput_path = Path('../input/')\ntrain_path = input_path / 'train'\n\ndef load_data(id, data_path):\n    data = np.load(data_path / id[0] / id[1] / id[2] / (id + '.npy')).astype('float32')\n    return data\n```\n\nI cast in np.float32 but you can cast into anything that matters, for instance to torch tensors:\n\n```\ndef load_data(id, data_path):\n    data = np.load(data_path / id[0] / id[1] / id[2] / (id + '.npy'))\n    data = torch.from_numpy(data).float()\n    return data\n```",
      "votes": null
    },
    {
      "id": "1389285",
      "postDate": "07/15/2021 15:22:25",
      "content": "<p>Hi Thanks.  I have been able to download the data and use a number of the notebook examples.  I am using TensorFlow, but I know PyTorch also.  I have split the large TRAIN file into these components, but it is making files of \"images\" that I have been unable to do. I have been working with Librosa image creation. I am having problems creating a loop.  The code works for one file at a time. The \"wave\" upload argument is my stumbling block.  For anomaly detection using an autoencoder, I need separate file of \"1's\" and \"0's\" that can be used to create train, validation, and test \"images\" via the various methods demonstrated in the notebooks.    I plan to train with images that are ALL  \"0's\" (no wave) .  The validation file would be both 0's and 1's, and the test file would be all \"1\" (all waves).   The reconstructions of the \"non-wave\" and \"wave\" in autoencoder should be produce thresholds that are different.  This difference should be the discriminator for prediction.  I did read a paper that in fact states that LSTM Autocoders are successful in solving this problem. (<a href=\"https://github.com/eric-moreno/Anomaly-Detection-Autoencoder)\" target=\"_blank\">https://github.com/eric-moreno/Anomaly-Detection-Autoencoder)</a>. However,  I am going to try an image based CNN autoencoder first.  I will be happy just to get results.    The LSTM Autoencoder will present another challenge for the data will not be images. It will be some sort of sequence.  I'll get there, so I will need to struggle a little longer.   At a minimum I will be learning . However, I do appreciate your response and help.  I am new to this and probably should be in a beginner competition, but I am retired so I need something. to do between, my day trading, exercising, reading, banjo playing, and taking a nap.  lol  Our son is a software engineer and I have a \"zoom help call\" scheduled for tonight.  Thanks again.</p>",
      "rawMarkdown": "Hi Thanks.  I have been able to download the data and use a number of the notebook examples.  I am using TensorFlow, but I know PyTorch also.  I have split the large TRAIN file into these components, but it is making files of \"images\" that I have been unable to do. I have been working with Librosa image creation. I am having problems creating a loop.  The code works for one file at a time. The \"wave\" upload argument is my stumbling block.  For anomaly detection using an autoencoder, I need separate file of \"1's\" and \"0's\" that can be used to create train, validation, and test \"images\" via the various methods demonstrated in the notebooks.    I plan to train with images that are ALL  \"0's\" (no wave) .  The validation file would be both 0's and 1's, and the test file would be all \"1\" (all waves).   The reconstructions of the \"non-wave\" and \"wave\" in autoencoder should be produce thresholds that are different.  This difference should be the discriminator for prediction.  I did read a paper that in fact states that LSTM Autocoders are successful in solving this problem. (https://github.com/eric-moreno/Anomaly-Detection-Autoencoder). However,  I am going to try an image based CNN autoencoder first.  I will be happy just to get results.    The LSTM Autoencoder will present another challenge for the data will not be images. It will be some sort of sequence.  I'll get there, so I will need to struggle a little longer.   At a minimum I will be learning . However, I do appreciate your response and help.  I am new to this and probably should be in a beginner competition, but I am retired so I need something. to do between, my day trading, exercising, reading, banjo playing, and taking a nap.  lol  Our son is a software engineer and I have a \"zoom help call\" scheduled for tonight.  Thanks again.",
      "votes": null
    },
    {
      "id": "1484877",
      "postDate": "08/21/2021 16:26:51",
      "content": "<p>Hi, May I ask about the status of your results with the autoencoders?</p>",
      "rawMarkdown": "Hi, May I ask about the status of your results with the autoencoders?",
      "votes": null
    },
    {
      "id": "1508193",
      "postDate": "09/10/2021 01:50:24",
      "content": "<p>Source-Agnostic Gravitational-Wave Detection with Recurrent Autoencoders<br>\n<a href=\"http://cds.cern.ch/record/2777883/files/2107.12698.pdf?version=2\" target=\"_blank\">http://cds.cern.ch/record/2777883/files/2107.12698.pdf?version=2</a></p>\n<p>Abstract. We present an application of anomaly detection techniques based on deep<br>\nrecurrent autoencoders to the problem of detecting gravitational wave signals in laser<br>\ninterferometers. Trained on noise data, this class of algorithms could detect signals using<br>\nan unsupervised strategy, i.e., without targeting a specific kind of source. We develop a<br>\ncustom architecture to analyze the data from two interferometers. We compare the<br>\nobtained performance to that obtained with other autoencoder architectures and with a<br>\nconvolutional classifier.</p>",
      "rawMarkdown": "Source-Agnostic Gravitational-Wave Detection with Recurrent Autoencoders\nhttp://cds.cern.ch/record/2777883/files/2107.12698.pdf?version=2\n\nAbstract. We present an application of anomaly detection techniques based on deep\nrecurrent autoencoders to the problem of detecting gravitational wave signals in laser\ninterferometers. Trained on noise data, this class of algorithms could detect signals using\nan unsupervised strategy, i.e., without targeting a specific kind of source. We develop a\ncustom architecture to analyze the data from two interferometers. We compare the\nobtained performance to that obtained with other autoencoder architectures and with a\nconvolutional classifier.",
      "votes": null
    },
    {
      "id": "1559697",
      "postDate": "10/27/2021 07:07:07",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    },
    {
      "id": "1559984",
      "postDate": "10/27/2021 08:52:47",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1388804,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/15/2021 08:37:17",
      "content": "<blockquote>\n  <p>data is given in such a complicated manner that I find it impossible to create the required databases</p>\n</blockquote>\n<p>What do you find complicated?  Loading one sample is rather easy, you'll find many implementations in public notebooks.  Here is one for instance:</p>\n<pre><code>from pathlib import Path\ninput_path = Path('../input/')\ntrain_path = input_path / 'train'\n\ndef load_data(id, data_path):\n    data = np.load(data_path / id[0] / id[1] / id[2] / (id + '.npy')).astype('float32')\n    return data\n</code></pre>\n<p>I cast in np.float32 but you can cast into anything that matters, for instance to torch tensors:</p>\n<pre><code>def load_data(id, data_path):\n    data = np.load(data_path / id[0] / id[1] / id[2] / (id + '.npy'))\n    data = torch.from_numpy(data).float()\n    return data\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1389285,
          "author_name": "llobos101",
          "author_url": "",
          "post_date": "07/15/2021 15:22:25",
          "content": "<p>Hi Thanks.  I have been able to download the data and use a number of the notebook examples.  I am using TensorFlow, but I know PyTorch also.  I have split the large TRAIN file into these components, but it is making files of \"images\" that I have been unable to do. I have been working with Librosa image creation. I am having problems creating a loop.  The code works for one file at a time. The \"wave\" upload argument is my stumbling block.  For anomaly detection using an autoencoder, I need separate file of \"1's\" and \"0's\" that can be used to create train, validation, and test \"images\" via the various methods demonstrated in the notebooks.    I plan to train with images that are ALL  \"0's\" (no wave) .  The validation file would be both 0's and 1's, and the test file would be all \"1\" (all waves).   The reconstructions of the \"non-wave\" and \"wave\" in autoencoder should be produce thresholds that are different.  This difference should be the discriminator for prediction.  I did read a paper that in fact states that LSTM Autocoders are successful in solving this problem. (<a href=\"https://github.com/eric-moreno/Anomaly-Detection-Autoencoder)\" target=\"_blank\">https://github.com/eric-moreno/Anomaly-Detection-Autoencoder)</a>. However,  I am going to try an image based CNN autoencoder first.  I will be happy just to get results.    The LSTM Autoencoder will present another challenge for the data will not be images. It will be some sort of sequence.  I'll get there, so I will need to struggle a little longer.   At a minimum I will be learning . However, I do appreciate your response and help.  I am new to this and probably should be in a beginner competition, but I am retired so I need something. to do between, my day trading, exercising, reading, banjo playing, and taking a nap.  lol  Our son is a software engineer and I have a \"zoom help call\" scheduled for tonight.  Thanks again.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1484877,
      "author_name": "hareeshthuruthipilly",
      "author_url": "",
      "post_date": "08/21/2021 16:26:51",
      "content": "<p>Hi, May I ask about the status of your results with the autoencoders?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1508193,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/10/2021 01:50:24",
      "content": "<p>Source-Agnostic Gravitational-Wave Detection with Recurrent Autoencoders<br>\n<a href=\"http://cds.cern.ch/record/2777883/files/2107.12698.pdf?version=2\" target=\"_blank\">http://cds.cern.ch/record/2777883/files/2107.12698.pdf?version=2</a></p>\n<p>Abstract. We present an application of anomaly detection techniques based on deep<br>\nrecurrent autoencoders to the problem of detecting gravitational wave signals in laser<br>\ninterferometers. Trained on noise data, this class of algorithms could detect signals using<br>\nan unsupervised strategy, i.e., without targeting a specific kind of source. We develop a<br>\ncustom architecture to analyze the data from two interferometers. We compare the<br>\nobtained performance to that obtained with other autoencoder architectures and with a<br>\nconvolutional classifier.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559697,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 07:07:07",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559984,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:52:47",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1388514": "I am trying to use a Autoencoder for anomaly detection but the data is given in such a complicated manner that I find it impossible to create the required databases.   I must admit that my abilities are lacking in data manipulation, for I a beginner at this stuff.   I would probably try a CNN first, but believe that a LSTM Autoencoder is probably the correct choice.   Has anyone else tried an CNN or LSTM Autoencoder?   I believe that this machine learning tool is the correct one to solve the problem.",
    "1388804": "> data is given in such a complicated manner that I find it impossible to create the required databases\n\nWhat do you find complicated?  Loading one sample is rather easy, you'll find many implementations in public notebooks.  Here is one for instance:\n\n```\nfrom pathlib import Path\ninput_path = Path('../input/')\ntrain_path = input_path / 'train'\n\ndef load_data(id, data_path):\n    data = np.load(data_path / id[0] / id[1] / id[2] / (id + '.npy')).astype('float32')\n    return data\n```\n\nI cast in np.float32 but you can cast into anything that matters, for instance to torch tensors:\n\n```\ndef load_data(id, data_path):\n    data = np.load(data_path / id[0] / id[1] / id[2] / (id + '.npy'))\n    data = torch.from_numpy(data).float()\n    return data\n```",
    "1389285": "Hi Thanks.  I have been able to download the data and use a number of the notebook examples.  I am using TensorFlow, but I know PyTorch also.  I have split the large TRAIN file into these components, but it is making files of \"images\" that I have been unable to do. I have been working with Librosa image creation. I am having problems creating a loop.  The code works for one file at a time. The \"wave\" upload argument is my stumbling block.  For anomaly detection using an autoencoder, I need separate file of \"1's\" and \"0's\" that can be used to create train, validation, and test \"images\" via the various methods demonstrated in the notebooks.    I plan to train with images that are ALL  \"0's\" (no wave) .  The validation file would be both 0's and 1's, and the test file would be all \"1\" (all waves).   The reconstructions of the \"non-wave\" and \"wave\" in autoencoder should be produce thresholds that are different.  This difference should be the discriminator for prediction.  I did read a paper that in fact states that LSTM Autocoders are successful in solving this problem. (https://github.com/eric-moreno/Anomaly-Detection-Autoencoder). However,  I am going to try an image based CNN autoencoder first.  I will be happy just to get results.    The LSTM Autoencoder will present another challenge for the data will not be images. It will be some sort of sequence.  I'll get there, so I will need to struggle a little longer.   At a minimum I will be learning . However, I do appreciate your response and help.  I am new to this and probably should be in a beginner competition, but I am retired so I need something. to do between, my day trading, exercising, reading, banjo playing, and taking a nap.  lol  Our son is a software engineer and I have a \"zoom help call\" scheduled for tonight.  Thanks again.",
    "1484877": "Hi, May I ask about the status of your results with the autoencoders?",
    "1508193": "Source-Agnostic Gravitational-Wave Detection with Recurrent Autoencoders\nhttp://cds.cern.ch/record/2777883/files/2107.12698.pdf?version=2\n\nAbstract. We present an application of anomaly detection techniques based on deep\nrecurrent autoencoders to the problem of detecting gravitational wave signals in laser\ninterferometers. Trained on noise data, this class of algorithms could detect signals using\nan unsupervised strategy, i.e., without targeting a specific kind of source. We develop a\ncustom architecture to analyze the data from two interferometers. We compare the\nobtained performance to that obtained with other autoencoder architectures and with a\nconvolutional classifier.",
    "1559697": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
    "1559984": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}