{
  "id": 550054,
  "title": "Some Data Visualization and Analysis",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/550054",
  "author_name": "",
  "post_date": "2024-12-05T07:13:37.542851300Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>So Lets begin,<br>\n   finally after loading the data using polars I noticed some features about it which includes what % of training data is just Nan and some of them is just useless and how target feature is distributed.<br>\n<strong>Correlation plot</strong>:<br>\nSo, we can have Multicollinearity problem here which is suggested by the following plot:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2F2386634c7412e42cec54ccc0ab076c07%2FCorrelation%20-%20Copy.jpg?generation=1733382680659886&amp;alt=media\" alt=\"\"></p>\n<p><strong>Target feature distribution:</strong><br>\nThat looks normal(kinda) but it's not normal(suggested by the QQ plot)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2F603e2823c905bca0c2520698b2de8711%2Ftarget%20-%20Copy.jpg?generation=1733382752320798&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2F4883b2227b9a29b1f44dc590ebcb2f46%2FQQ%20-%20Copy.jpg?generation=1733382765762447&amp;alt=media\" alt=\"\"></p>\n<p><strong>Symbols distribution:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2Fbd5121af42394fa38d58fac95dcf74fd%2Fsymbols%20-%20Copy.jpg?generation=1733382797722363&amp;alt=media\" alt=\"\"></p>\n<p><strong>NaN values:</strong><br>\nfor feature_00: 6.752030000081906%<br>\nfor feature_01: 6.752030000081906%<br>\nfor feature_02: 6.752030000081906%<br>\nfor feature_03: 6.752030000081906%<br>\nfor feature_04: 6.752030000081906%<br>\nfor feature_08: 0.637097304328965%<br>\nfor feature_15: 2.566024416655997%<br>\nfor feature_16: 0.0005538186773884831%<br>\nfor feature_17: 0.4282822000258109%<br>\nfor feature_18: 0.00047955180494175164%<br>\nfor feature_19: 0.00047955180494175164%<br>\nfor feature_21: 17.900406341643993%<br>\nfor feature_26: 17.900406341643993%<br>\nfor feature_27: 17.900406341643993%<br>\nfor feature_31: 17.900406341643993%<br>\nfor feature_32: 1.0152429997213082%<br>\nfor feature_33: 1.0152429997213082%<br>\nfor feature_37: 0.0018015021344935714%<br>\nfor feature_39: 9.125592877747518%<br>\nfor feature_40: 0.14398436847844026%<br>\nfor feature_41: 2.3192737939070525%<br>\nfor feature_42: 9.125592877747518%<br>\nfor feature_43: 0.14398436847844026%<br>\nfor feature_44: 2.3192737939070525%<br>\nfor feature_45: 0.672991544737791%<br>\nfor feature_46: 0.672991544737791%<br>\nfor feature_47: 0.00018460622579616104%<br>\nfor feature_50: 9.026815815482724%<br>\nfor feature_51: 0.02929297640363222%<br>\nfor feature_52: 2.2171801853098514%<br>\nfor feature_53: 9.026815815482724%<br>\nfor feature_54: 0.02929297640363222%<br>\nfor feature_55: 2.2171801853098514%<br>\nfor feature_56: 0.00047955180494175164%<br>\nfor feature_57: 0.00047955180494175164%<br>\nfor feature_58: 1.0152323901681015%<br>\nfor feature_62: 0.621352727370258%<br>\nfor feature_63: 0.48287471700608253%<br>\nfor feature_64: 0.5042996487516439%<br>\nfor feature_65: 0.672991544737791%<br>\nfor feature_66: 0.672991544737791%<br>\nfor feature_73: 1.0264933699416674%<br>\nfor feature_74: 1.0264933699416674%<br>\nfor feature_75: 0.12398323877321482%<br>\nfor feature_76: 0.12398323877321482%<br>\nfor feature_77: 0.042529454984281095%<br>\nfor feature_78: 0.042529454984281095%</p>\n<p>Ya… Thanks for reading.. bye :)</p>",
  "messages": [
    {
      "id": "3064054",
      "postDate": "12/05/2024 07:13:37",
      "content": "<p>So Lets begin,<br>\n   finally after loading the data using polars I noticed some features about it which includes what % of training data is just Nan and some of them is just useless and how target feature is distributed.<br>\n<strong>Correlation plot</strong>:<br>\nSo, we can have Multicollinearity problem here which is suggested by the following plot:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2F2386634c7412e42cec54ccc0ab076c07%2FCorrelation%20-%20Copy.jpg?generation=1733382680659886&amp;alt=media\" alt=\"\"></p>\n<p><strong>Target feature distribution:</strong><br>\nThat looks normal(kinda) but it's not normal(suggested by the QQ plot)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2F603e2823c905bca0c2520698b2de8711%2Ftarget%20-%20Copy.jpg?generation=1733382752320798&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2F4883b2227b9a29b1f44dc590ebcb2f46%2FQQ%20-%20Copy.jpg?generation=1733382765762447&amp;alt=media\" alt=\"\"></p>\n<p><strong>Symbols distribution:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2Fbd5121af42394fa38d58fac95dcf74fd%2Fsymbols%20-%20Copy.jpg?generation=1733382797722363&amp;alt=media\" alt=\"\"></p>\n<p><strong>NaN values:</strong><br>\nfor feature_00: 6.752030000081906%<br>\nfor feature_01: 6.752030000081906%<br>\nfor feature_02: 6.752030000081906%<br>\nfor feature_03: 6.752030000081906%<br>\nfor feature_04: 6.752030000081906%<br>\nfor feature_08: 0.637097304328965%<br>\nfor feature_15: 2.566024416655997%<br>\nfor feature_16: 0.0005538186773884831%<br>\nfor feature_17: 0.4282822000258109%<br>\nfor feature_18: 0.00047955180494175164%<br>\nfor feature_19: 0.00047955180494175164%<br>\nfor feature_21: 17.900406341643993%<br>\nfor feature_26: 17.900406341643993%<br>\nfor feature_27: 17.900406341643993%<br>\nfor feature_31: 17.900406341643993%<br>\nfor feature_32: 1.0152429997213082%<br>\nfor feature_33: 1.0152429997213082%<br>\nfor feature_37: 0.0018015021344935714%<br>\nfor feature_39: 9.125592877747518%<br>\nfor feature_40: 0.14398436847844026%<br>\nfor feature_41: 2.3192737939070525%<br>\nfor feature_42: 9.125592877747518%<br>\nfor feature_43: 0.14398436847844026%<br>\nfor feature_44: 2.3192737939070525%<br>\nfor feature_45: 0.672991544737791%<br>\nfor feature_46: 0.672991544737791%<br>\nfor feature_47: 0.00018460622579616104%<br>\nfor feature_50: 9.026815815482724%<br>\nfor feature_51: 0.02929297640363222%<br>\nfor feature_52: 2.2171801853098514%<br>\nfor feature_53: 9.026815815482724%<br>\nfor feature_54: 0.02929297640363222%<br>\nfor feature_55: 2.2171801853098514%<br>\nfor feature_56: 0.00047955180494175164%<br>\nfor feature_57: 0.00047955180494175164%<br>\nfor feature_58: 1.0152323901681015%<br>\nfor feature_62: 0.621352727370258%<br>\nfor feature_63: 0.48287471700608253%<br>\nfor feature_64: 0.5042996487516439%<br>\nfor feature_65: 0.672991544737791%<br>\nfor feature_66: 0.672991544737791%<br>\nfor feature_73: 1.0264933699416674%<br>\nfor feature_74: 1.0264933699416674%<br>\nfor feature_75: 0.12398323877321482%<br>\nfor feature_76: 0.12398323877321482%<br>\nfor feature_77: 0.042529454984281095%<br>\nfor feature_78: 0.042529454984281095%</p>\n<p>Ya… Thanks for reading.. bye :)</p>",
      "rawMarkdown": "So Lets begin,\n   finally after loading the data using polars I noticed some features about it which includes what % of training data is just Nan and some of them is just useless and how target feature is distributed.\n**Correlation plot**:\nSo, we can have Multicollinearity problem here which is suggested by the following plot:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2F2386634c7412e42cec54ccc0ab076c07%2FCorrelation%20-%20Copy.jpg?generation=1733382680659886&alt=media)\n\n**Target feature distribution:**\nThat looks normal(kinda) but it's not normal(suggested by the QQ plot)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2F603e2823c905bca0c2520698b2de8711%2Ftarget%20-%20Copy.jpg?generation=1733382752320798&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2F4883b2227b9a29b1f44dc590ebcb2f46%2FQQ%20-%20Copy.jpg?generation=1733382765762447&alt=media)\n\n**Symbols distribution:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2Fbd5121af42394fa38d58fac95dcf74fd%2Fsymbols%20-%20Copy.jpg?generation=1733382797722363&alt=media)\n\n**NaN values:**\nfor feature_00: 6.752030000081906%\nfor feature_01: 6.752030000081906%\nfor feature_02: 6.752030000081906%\nfor feature_03: 6.752030000081906%\nfor feature_04: 6.752030000081906%\nfor feature_08: 0.637097304328965%\nfor feature_15: 2.566024416655997%\nfor feature_16: 0.0005538186773884831%\nfor feature_17: 0.4282822000258109%\nfor feature_18: 0.00047955180494175164%\nfor feature_19: 0.00047955180494175164%\nfor feature_21: 17.900406341643993%\nfor feature_26: 17.900406341643993%\nfor feature_27: 17.900406341643993%\nfor feature_31: 17.900406341643993%\nfor feature_32: 1.0152429997213082%\nfor feature_33: 1.0152429997213082%\nfor feature_37: 0.0018015021344935714%\nfor feature_39: 9.125592877747518%\nfor feature_40: 0.14398436847844026%\nfor feature_41: 2.3192737939070525%\nfor feature_42: 9.125592877747518%\nfor feature_43: 0.14398436847844026%\nfor feature_44: 2.3192737939070525%\nfor feature_45: 0.672991544737791%\nfor feature_46: 0.672991544737791%\nfor feature_47: 0.00018460622579616104%\nfor feature_50: 9.026815815482724%\nfor feature_51: 0.02929297640363222%\nfor feature_52: 2.2171801853098514%\nfor feature_53: 9.026815815482724%\nfor feature_54: 0.02929297640363222%\nfor feature_55: 2.2171801853098514%\nfor feature_56: 0.00047955180494175164%\nfor feature_57: 0.00047955180494175164%\nfor feature_58: 1.0152323901681015%\nfor feature_62: 0.621352727370258%\nfor feature_63: 0.48287471700608253%\nfor feature_64: 0.5042996487516439%\nfor feature_65: 0.672991544737791%\nfor feature_66: 0.672991544737791%\nfor feature_73: 1.0264933699416674%\nfor feature_74: 1.0264933699416674%\nfor feature_75: 0.12398323877321482%\nfor feature_76: 0.12398323877321482%\nfor feature_77: 0.042529454984281095%\nfor feature_78: 0.042529454984281095%\n\nYa... Thanks for reading.. bye :)",
      "votes": null
    },
    {
      "id": "3064083",
      "postDate": "12/05/2024 08:05:07",
      "content": "<blockquote>\n  <p>That looks normal(kinda) but it's not normal</p>\n</blockquote>\n<p>It's most close to Laplace distribution.</p>",
      "rawMarkdown": ">That looks normal(kinda) but it's not normal\n\nIt's most close to Laplace distribution.",
      "votes": null
    },
    {
      "id": "3064872",
      "postDate": "12/06/2024 04:36:56",
      "content": "<p>yes it looks like that by I've plotted the QQ plot for it as well and it suggests something else</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2Ff96316d68f7c563db355933a8e4f3942%2FScreenshot%202024-12-06%20100157%20-%20Copy.jpg?generation=1733459812698860&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "yes it looks like that by I've plotted the QQ plot for it as well and it suggests something else\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2Ff96316d68f7c563db355933a8e4f3942%2FScreenshot%202024-12-06%20100157%20-%20Copy.jpg?generation=1733459812698860&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3064083,
      "author_name": "shuthdar",
      "author_url": "",
      "post_date": "12/05/2024 08:05:07",
      "content": "<blockquote>\n  <p>That looks normal(kinda) but it's not normal</p>\n</blockquote>\n<p>It's most close to Laplace distribution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3064872,
          "author_name": "nabayansaha",
          "author_url": "",
          "post_date": "12/06/2024 04:36:56",
          "content": "<p>yes it looks like that by I've plotted the QQ plot for it as well and it suggests something else</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2Ff96316d68f7c563db355933a8e4f3942%2FScreenshot%202024-12-06%20100157%20-%20Copy.jpg?generation=1733459812698860&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3064054": "So Lets begin,\n   finally after loading the data using polars I noticed some features about it which includes what % of training data is just Nan and some of them is just useless and how target feature is distributed.\n**Correlation plot**:\nSo, we can have Multicollinearity problem here which is suggested by the following plot:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2F2386634c7412e42cec54ccc0ab076c07%2FCorrelation%20-%20Copy.jpg?generation=1733382680659886&alt=media)\n\n**Target feature distribution:**\nThat looks normal(kinda) but it's not normal(suggested by the QQ plot)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2F603e2823c905bca0c2520698b2de8711%2Ftarget%20-%20Copy.jpg?generation=1733382752320798&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2F4883b2227b9a29b1f44dc590ebcb2f46%2FQQ%20-%20Copy.jpg?generation=1733382765762447&alt=media)\n\n**Symbols distribution:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2Fbd5121af42394fa38d58fac95dcf74fd%2Fsymbols%20-%20Copy.jpg?generation=1733382797722363&alt=media)\n\n**NaN values:**\nfor feature_00: 6.752030000081906%\nfor feature_01: 6.752030000081906%\nfor feature_02: 6.752030000081906%\nfor feature_03: 6.752030000081906%\nfor feature_04: 6.752030000081906%\nfor feature_08: 0.637097304328965%\nfor feature_15: 2.566024416655997%\nfor feature_16: 0.0005538186773884831%\nfor feature_17: 0.4282822000258109%\nfor feature_18: 0.00047955180494175164%\nfor feature_19: 0.00047955180494175164%\nfor feature_21: 17.900406341643993%\nfor feature_26: 17.900406341643993%\nfor feature_27: 17.900406341643993%\nfor feature_31: 17.900406341643993%\nfor feature_32: 1.0152429997213082%\nfor feature_33: 1.0152429997213082%\nfor feature_37: 0.0018015021344935714%\nfor feature_39: 9.125592877747518%\nfor feature_40: 0.14398436847844026%\nfor feature_41: 2.3192737939070525%\nfor feature_42: 9.125592877747518%\nfor feature_43: 0.14398436847844026%\nfor feature_44: 2.3192737939070525%\nfor feature_45: 0.672991544737791%\nfor feature_46: 0.672991544737791%\nfor feature_47: 0.00018460622579616104%\nfor feature_50: 9.026815815482724%\nfor feature_51: 0.02929297640363222%\nfor feature_52: 2.2171801853098514%\nfor feature_53: 9.026815815482724%\nfor feature_54: 0.02929297640363222%\nfor feature_55: 2.2171801853098514%\nfor feature_56: 0.00047955180494175164%\nfor feature_57: 0.00047955180494175164%\nfor feature_58: 1.0152323901681015%\nfor feature_62: 0.621352727370258%\nfor feature_63: 0.48287471700608253%\nfor feature_64: 0.5042996487516439%\nfor feature_65: 0.672991544737791%\nfor feature_66: 0.672991544737791%\nfor feature_73: 1.0264933699416674%\nfor feature_74: 1.0264933699416674%\nfor feature_75: 0.12398323877321482%\nfor feature_76: 0.12398323877321482%\nfor feature_77: 0.042529454984281095%\nfor feature_78: 0.042529454984281095%\n\nYa... Thanks for reading.. bye :)",
    "3064083": ">That looks normal(kinda) but it's not normal\n\nIt's most close to Laplace distribution.",
    "3064872": "yes it looks like that by I've plotted the QQ plot for it as well and it suggests something else\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14232682%2Ff96316d68f7c563db355933a8e4f3942%2FScreenshot%202024-12-06%20100157%20-%20Copy.jpg?generation=1733459812698860&alt=media)"
  },
  "source": "meta"
}