{
  "id": 99635,
  "title": "Filtering using Domain Knowledge",
  "url": "/competitions/bigdata2019-flare-prediction/discussion/99635",
  "author_name": "Dustin Kempton",
  "post_date": "2019-07-12T19:01:51.096000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Based upon information presented in <a href=\"https://iopscience.iop.org/article/10.1086/511857\">Schriver (2007)</a>, participants may find it helpful to use the maximum R_VALUE in a sample as a filtering method on non-flaring regions.</p>\n\n<p>For example, using a threshold of 1.5 on the maximum value of the R_VALUE parameter one can achieve the following filtering statistics when looking at all three training folds:</p>\n\n<p>|Total NF Samples  | # Below Threshold | % Below|\n| --- | --- |---|\n| 164,974 | 76,692 |46.49 |</p>\n\n<p>| Total FL Samples | # Below Threshold  | % Below|\n| --- | --- | --- |\n| 31,286 |190 | 0.607|</p>\n\n<p>The threshold can be played with here and it may be useful to produce two different models or one can simply assume that the 190 samples that fall below the threshold are anomalous readings and should be disregarded for the sake of the competition.</p>\n\n<p>Here is a plot of what the R_VALUE maximum looks like in reference to the USFLUX maximum in the same sample.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2697732%2F7288e25717f1768ea238c9b3e0a0b8fb%2Fbla.png?generation=1562958086925221&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 573799,
      "postDate": "2019-07-12T19:01:51.097Z",
      "content": "<p>Based upon information presented in <a href=\"https://iopscience.iop.org/article/10.1086/511857\">Schriver (2007)</a>, participants may find it helpful to use the maximum R_VALUE in a sample as a filtering method on non-flaring regions.</p>\n\n<p>For example, using a threshold of 1.5 on the maximum value of the R_VALUE parameter one can achieve the following filtering statistics when looking at all three training folds:</p>\n\n<p>|Total NF Samples  | # Below Threshold | % Below|\n| --- | --- |---|\n| 164,974 | 76,692 |46.49 |</p>\n\n<p>| Total FL Samples | # Below Threshold  | % Below|\n| --- | --- | --- |\n| 31,286 |190 | 0.607|</p>\n\n<p>The threshold can be played with here and it may be useful to produce two different models or one can simply assume that the 190 samples that fall below the threshold are anomalous readings and should be disregarded for the sake of the competition.</p>\n\n<p>Here is a plot of what the R_VALUE maximum looks like in reference to the USFLUX maximum in the same sample.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2697732%2F7288e25717f1768ea238c9b3e0a0b8fb%2Fbla.png?generation=1562958086925221&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Based upon information presented in [Schriver (2007)](https://iopscience.iop.org/article/10.1086/511857), participants may find it helpful to use the maximum R_VALUE in a sample as a filtering method on non-flaring regions.\n\nFor example, using a threshold of 1.5 on the maximum value of the R_VALUE parameter one can achieve the following filtering statistics when looking at all three training folds:\n\n|Total NF Samples  | # Below Threshold | % Below|\n| --- | --- |---|\n| 164,974 | 76,692 |46.49 |\n\n| Total FL Samples | # Below Threshold  | % Below|\n| --- | --- | --- |\n| 31,286 |190 | 0.607|\n\nThe threshold can be played with here and it may be useful to produce two different models or one can simply assume that the 190 samples that fall below the threshold are anomalous readings and should be disregarded for the sake of the competition.\n\nHere is a plot of what the R_VALUE maximum looks like in reference to the USFLUX maximum in the same sample.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2697732%2F7288e25717f1768ea238c9b3e0a0b8fb%2Fbla.png?generation=1562958086925221&amp;alt=media)\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "573799": "Based upon information presented in [Schriver (2007)](https://iopscience.iop.org/article/10.1086/511857), participants may find it helpful to use the maximum R_VALUE in a sample as a filtering method on non-flaring regions.\n\nFor example, using a threshold of 1.5 on the maximum value of the R_VALUE parameter one can achieve the following filtering statistics when looking at all three training folds:\n\n|Total NF Samples  | # Below Threshold | % Below|\n| --- | --- |---|\n| 164,974 | 76,692 |46.49 |\n\n| Total FL Samples | # Below Threshold  | % Below|\n| --- | --- | --- |\n| 31,286 |190 | 0.607|\n\nThe threshold can be played with here and it may be useful to produce two different models or one can simply assume that the 190 samples that fall below the threshold are anomalous readings and should be disregarded for the sake of the competition.\n\nHere is a plot of what the R_VALUE maximum looks like in reference to the USFLUX maximum in the same sample.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2697732%2F7288e25717f1768ea238c9b3e0a0b8fb%2Fbla.png?generation=1562958086925221&amp;alt=media)\n"
  }
}