{
  "id": 90115,
  "title": "Variance related features and intensity distribuition",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90115",
  "author_name": "Roberto Anzaldua",
  "post_date": "2019-04-20T15:56:29.758000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>So, I am, as everyone else is, creating features for around 4100 training inputs. My current best score (1.522) was created using two types of features:</p>\n\n<ol>\n<li><p><strong>Intensity counts features</strong>. There are about 5600 different values in the training set varying from around -5500 to 5500. I have created features for determining <em>how many</em> values for an intensity level the training sample has. For example. I have set a feature that counts how many values between -5500 and -5300 were observed in a given training sample and I do the same across all values.</p></li>\n<li><p><strong>Uniqueness distribution</strong>.  It is a term I am using to describe a type of features that come from using the histogram of value counts. The intuition from this feature comes from the fact that <em>I have observed</em> that when there is a significant time before the next earthquake (&gt; 7 seconds), the number of different intensities varies a lot for the training input. The intuitive explanation to this is that when we are not close to an earthquake, there can be a lot of different intensities, but as we get closer to the earthquake, we observe a much \"stable\" set of recorded intensities.</p></li>\n</ol>\n\n<p>I hope you find this interesting! I am planning to upload a notebook to explain in more detail my observations. </p>\n\n<p>The score I got is without \"much tunning\" of the models. I am using Catboost/XGBoost with similar results.</p>",
  "messages": [
    {
      "id": 520270,
      "postDate": "2019-04-20T15:56:29.757Z",
      "content": "<p>So, I am, as everyone else is, creating features for around 4100 training inputs. My current best score (1.522) was created using two types of features:</p>\n\n<ol>\n<li><p><strong>Intensity counts features</strong>. There are about 5600 different values in the training set varying from around -5500 to 5500. I have created features for determining <em>how many</em> values for an intensity level the training sample has. For example. I have set a feature that counts how many values between -5500 and -5300 were observed in a given training sample and I do the same across all values.</p></li>\n<li><p><strong>Uniqueness distribution</strong>.  It is a term I am using to describe a type of features that come from using the histogram of value counts. The intuition from this feature comes from the fact that <em>I have observed</em> that when there is a significant time before the next earthquake (&gt; 7 seconds), the number of different intensities varies a lot for the training input. The intuitive explanation to this is that when we are not close to an earthquake, there can be a lot of different intensities, but as we get closer to the earthquake, we observe a much \"stable\" set of recorded intensities.</p></li>\n</ol>\n\n<p>I hope you find this interesting! I am planning to upload a notebook to explain in more detail my observations. </p>\n\n<p>The score I got is without \"much tunning\" of the models. I am using Catboost/XGBoost with similar results.</p>",
      "rawMarkdown": "So, I am, as everyone else is, creating features for around 4100 training inputs. My current best score (1.522) was created using two types of features:\n\n1. **Intensity counts features**. There are about 5600 different values in the training set varying from around -5500 to 5500. I have created features for determining *how many* values for an intensity level the training sample has. For example. I have set a feature that counts how many values between -5500 and -5300 were observed in a given training sample and I do the same across all values.\n\n2. **Uniqueness distribution**.  It is a term I am using to describe a type of features that come from using the histogram of value counts. The intuition from this feature comes from the fact that *I have observed* that when there is a significant time before the next earthquake (&gt; 7 seconds), the number of different intensities varies a lot for the training input. The intuitive explanation to this is that when we are not close to an earthquake, there can be a lot of different intensities, but as we get closer to the earthquake, we observe a much \"stable\" set of recorded intensities.\n\nI hope you find this interesting! I am planning to upload a notebook to explain in more detail my observations. \n\nThe score I got is without \"much tunning\" of the models. I am using Catboost/XGBoost with similar results.",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "520270": "So, I am, as everyone else is, creating features for around 4100 training inputs. My current best score (1.522) was created using two types of features:\n\n1. **Intensity counts features**. There are about 5600 different values in the training set varying from around -5500 to 5500. I have created features for determining *how many* values for an intensity level the training sample has. For example. I have set a feature that counts how many values between -5500 and -5300 were observed in a given training sample and I do the same across all values.\n\n2. **Uniqueness distribution**.  It is a term I am using to describe a type of features that come from using the histogram of value counts. The intuition from this feature comes from the fact that *I have observed* that when there is a significant time before the next earthquake (&gt; 7 seconds), the number of different intensities varies a lot for the training input. The intuitive explanation to this is that when we are not close to an earthquake, there can be a lot of different intensities, but as we get closer to the earthquake, we observe a much \"stable\" set of recorded intensities.\n\nI hope you find this interesting! I am planning to upload a notebook to explain in more detail my observations. \n\nThe score I got is without \"much tunning\" of the models. I am using Catboost/XGBoost with similar results."
  }
}