{
  "id": 333351,
  "title": "Time Series Pattern of B_2",
  "url": "/competitions/amex-default-prediction/discussion/333351",
  "author_name": "",
  "post_date": "2022-06-26T04:24:15.628024500Z",
  "votes": 24,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I have been leveraging the Time Series EDA Notebook shared by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> to investigate data patterns. And I am baffled by the behavior of B_2 and how there seems to be two ceilings in the data. Does anyone have a guess why B_2 might show a pattern like this?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3585582%2F03f5be82427b289b41b30e99373f993c%2FB_2.png?generation=1656217269536360&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1833498",
      "postDate": "06/26/2022 04:24:15",
      "content": "<p>Hello,</p>\n<p>I have been leveraging the Time Series EDA Notebook shared by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> to investigate data patterns. And I am baffled by the behavior of B_2 and how there seems to be two ceilings in the data. Does anyone have a guess why B_2 might show a pattern like this?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3585582%2F03f5be82427b289b41b30e99373f993c%2FB_2.png?generation=1656217269536360&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hello,\n\nI have been leveraging the Time Series EDA Notebook shared by @cdeotte to investigate data patterns. And I am baffled by the behavior of B_2 and how there seems to be two ceilings in the data. Does anyone have a guess why B_2 might show a pattern like this?\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3585582%2F03f5be82427b289b41b30e99373f993c%2FB_2.png?generation=1656217269536360&alt=media)",
      "votes": null
    },
    {
      "id": "1834129",
      "postDate": "06/26/2022 16:59:14",
      "content": "<p>That is an interesting pattern. I'm not sure why exactly, but they are basically two different categorical <code>B_38</code> types of rows. The rows with <code>B_2 = 0.8</code> are mainly categorical feature <code>B_38 = 1</code> whereas rows with <code>B_2 = 1.0</code> are mainly categorical feature <code>B_38 = 2</code>.</p>\n<p>If you group all the <code>0.7 &lt; B_2 &lt;0.9</code> rows together. And group all the <code>0.9 &lt; B_2 &lt; 1.1</code> rows together. Then you can plot dual histograms of all the 188 features. If we do this, we see that these two groups of rows differ most in features <code>S_6, S_8, D_60, D_61, S_11, S_13, S_15, S_19, B_38</code>.</p>\n<p>Next, you can find all customers with at least one <code>0.7 &lt; B_2 &lt;0.9</code>. And find all customers with at least one <code>0.9 &lt; B_2 &lt; 1.1</code>. Then compare those histograms. The customer ids have 62% overlap, but enough customers specialize in one or the other so the results are basically the same as the row analysis.</p>",
      "rawMarkdown": "That is an interesting pattern. I'm not sure why exactly, but they are basically two different categorical `B_38` types of rows. The rows with `B_2 = 0.8` are mainly categorical feature `B_38 = 1` whereas rows with `B_2 = 1.0` are mainly categorical feature `B_38 = 2`.\n\nIf you group all the `0.7 < B_2 <0.9` rows together. And group all the `0.9 < B_2 < 1.1` rows together. Then you can plot dual histograms of all the 188 features. If we do this, we see that these two groups of rows differ most in features `S_6, S_8, D_60, D_61, S_11, S_13, S_15, S_19, B_38`.\n\nNext, you can find all customers with at least one `0.7 < B_2 <0.9`. And find all customers with at least one `0.9 < B_2 < 1.1`. Then compare those histograms. The customer ids have 62% overlap, but enough customers specialize in one or the other so the results are basically the same as the row analysis.",
      "votes": null
    },
    {
      "id": "1834211",
      "postDate": "06/26/2022 18:14:04",
      "content": "<p>Thank you Chris as always for sharing your insights! The dual histogram approach sounds great! I will check it out! </p>",
      "rawMarkdown": "Thank you Chris as always for sharing your insights! The dual histogram approach sounds great! I will check it out!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1834129,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/26/2022 16:59:14",
      "content": "<p>That is an interesting pattern. I'm not sure why exactly, but they are basically two different categorical <code>B_38</code> types of rows. The rows with <code>B_2 = 0.8</code> are mainly categorical feature <code>B_38 = 1</code> whereas rows with <code>B_2 = 1.0</code> are mainly categorical feature <code>B_38 = 2</code>.</p>\n<p>If you group all the <code>0.7 &lt; B_2 &lt;0.9</code> rows together. And group all the <code>0.9 &lt; B_2 &lt; 1.1</code> rows together. Then you can plot dual histograms of all the 188 features. If we do this, we see that these two groups of rows differ most in features <code>S_6, S_8, D_60, D_61, S_11, S_13, S_15, S_19, B_38</code>.</p>\n<p>Next, you can find all customers with at least one <code>0.7 &lt; B_2 &lt;0.9</code>. And find all customers with at least one <code>0.9 &lt; B_2 &lt; 1.1</code>. Then compare those histograms. The customer ids have 62% overlap, but enough customers specialize in one or the other so the results are basically the same as the row analysis.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1834211,
          "author_name": "raphael1123",
          "author_url": "",
          "post_date": "06/26/2022 18:14:04",
          "content": "<p>Thank you Chris as always for sharing your insights! The dual histogram approach sounds great! I will check it out! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1833498": "Hello,\n\nI have been leveraging the Time Series EDA Notebook shared by @cdeotte to investigate data patterns. And I am baffled by the behavior of B_2 and how there seems to be two ceilings in the data. Does anyone have a guess why B_2 might show a pattern like this?\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3585582%2F03f5be82427b289b41b30e99373f993c%2FB_2.png?generation=1656217269536360&alt=media)",
    "1834129": "That is an interesting pattern. I'm not sure why exactly, but they are basically two different categorical `B_38` types of rows. The rows with `B_2 = 0.8` are mainly categorical feature `B_38 = 1` whereas rows with `B_2 = 1.0` are mainly categorical feature `B_38 = 2`.\n\nIf you group all the `0.7 < B_2 <0.9` rows together. And group all the `0.9 < B_2 < 1.1` rows together. Then you can plot dual histograms of all the 188 features. If we do this, we see that these two groups of rows differ most in features `S_6, S_8, D_60, D_61, S_11, S_13, S_15, S_19, B_38`.\n\nNext, you can find all customers with at least one `0.7 < B_2 <0.9`. And find all customers with at least one `0.9 < B_2 < 1.1`. Then compare those histograms. The customer ids have 62% overlap, but enough customers specialize in one or the other so the results are basically the same as the row analysis.",
    "1834211": "Thank you Chris as always for sharing your insights! The dual histogram approach sounds great! I will check it out!"
  },
  "source": "meta"
}