{
  "id": 536694,
  "title": "Have You Noticed Any Anomalous Values? Let’s Share Them Here!",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/536694",
  "author_name": "",
  "post_date": "2024-09-29T08:49:13.183852100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi everyone,<br>\nWhen data is missing, it's usually recorded as NaN. However, I’ve noticed an unusual value in the training data. Specifically, in row 2065 (id = 83525bbe), the \"CGAS-CGAS_Score\" column has a value of 999, even though this column typically contains values between 0 and 100.</p>\n<p>This kind of impossible value should ideally be removed during preprocessing, but I realize it can be time-consuming for each of us to identify these anomalies. If you’ve come across any similar issues, feel free to share them here to help everyone out!</p>\n<p>Looking forward to hearing from you!</p>\n<p><strong>Update:</strong></p>\n<p>As some of the comments have pointed out, many reports on outliers have already been made in Antonina Dolgorukova’s topic (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535354</a>). I recommend referring to the content and notebooks attached there first. <br>\nAlso, outliers and noise in the Actigraphy files are being discussed here as well: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535900</a>.</p>\n<p>If I become aware of other topics discussing outliers, I will update this post accordingly. Please feel free to comment if you notice anything missing!</p>\n<p>While this topic may no longer hold much value, I’ll leave it up for beginners who might encounter similar issues. If this topic causes any confusion in the discussion, please feel free to let me know, and I will delete it.</p>",
  "messages": [
    {
      "id": "3001762",
      "postDate": "09/29/2024 08:49:13",
      "content": "<p>Hi everyone,<br>\nWhen data is missing, it's usually recorded as NaN. However, I’ve noticed an unusual value in the training data. Specifically, in row 2065 (id = 83525bbe), the \"CGAS-CGAS_Score\" column has a value of 999, even though this column typically contains values between 0 and 100.</p>\n<p>This kind of impossible value should ideally be removed during preprocessing, but I realize it can be time-consuming for each of us to identify these anomalies. If you’ve come across any similar issues, feel free to share them here to help everyone out!</p>\n<p>Looking forward to hearing from you!</p>\n<p><strong>Update:</strong></p>\n<p>As some of the comments have pointed out, many reports on outliers have already been made in Antonina Dolgorukova’s topic (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535354</a>). I recommend referring to the content and notebooks attached there first. <br>\nAlso, outliers and noise in the Actigraphy files are being discussed here as well: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535900</a>.</p>\n<p>If I become aware of other topics discussing outliers, I will update this post accordingly. Please feel free to comment if you notice anything missing!</p>\n<p>While this topic may no longer hold much value, I’ll leave it up for beginners who might encounter similar issues. If this topic causes any confusion in the discussion, please feel free to let me know, and I will delete it.</p>",
      "rawMarkdown": "Hi everyone,\nWhen data is missing, it's usually recorded as NaN. However, I’ve noticed an unusual value in the training data. Specifically, in row 2065 (id = 83525bbe), the \"CGAS-CGAS_Score\" column has a value of 999, even though this column typically contains values between 0 and 100.\n\nThis kind of impossible value should ideally be removed during preprocessing, but I realize it can be time-consuming for each of us to identify these anomalies. If you’ve come across any similar issues, feel free to share them here to help everyone out!\n\nLooking forward to hearing from you!\n\n**Update:**\n\nAs some of the comments have pointed out, many reports on outliers have already been made in Antonina Dolgorukova’s topic ([https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535354](url)). I recommend referring to the content and notebooks attached there first. \nAlso, outliers and noise in the Actigraphy files are being discussed here as well: [https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535900](url).\n\nIf I become aware of other topics discussing outliers, I will update this post accordingly. Please feel free to comment if you notice anything missing!\n\nWhile this topic may no longer hold much value, I’ll leave it up for beginners who might encounter similar issues. If this topic causes any confusion in the discussion, please feel free to let me know, and I will delete it.",
      "votes": null
    },
    {
      "id": "3001774",
      "postDate": "09/29/2024 09:00:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/masatoito8823\" target=\"_blank\">@masatoito8823</a>, there are several public EDA notebooks which show many outliers.</p>",
      "rawMarkdown": "Hi @masatoito8823, there are several public EDA notebooks which show many outliers.",
      "votes": null
    },
    {
      "id": "3001796",
      "postDate": "09/29/2024 09:30:36",
      "content": "<p>Hi! I have already shared everything I have found so far, apart from the EDA notebooks, it is summarised in the topic \"Some findings from the features EDA (+ data issues, outliers)\", and also, you can see e.g. \"Worst three time series of the dataset\" - for insights into the time series actigraphy data :)</p>",
      "rawMarkdown": "Hi! I have already shared everything I have found so far, apart from the EDA notebooks, it is summarised in the topic \"Some findings from the features EDA (+ data issues, outliers)\", and also, you can see e.g. \"Worst three time series of the dataset\" - for insights into the time series actigraphy data :)",
      "votes": null
    },
    {
      "id": "3001867",
      "postDate": "09/29/2024 11:29:58",
      "content": "<p>I apologize for overlooking your contributions, and thank you so much for taking the time to comment! I’ve only just started reviewing the data, and it seems you’ve already conducted a very thorough analysis. I will definitely refer to your topic and notebooks for my future work. Thanks again for sharing your insights!</p>",
      "rawMarkdown": "I apologize for overlooking your contributions, and thank you so much for taking the time to comment! I’ve only just started reviewing the data, and it seems you’ve already conducted a very thorough analysis. I will definitely refer to your topic and notebooks for my future work. Thanks again for sharing your insights!",
      "votes": null
    },
    {
      "id": "3001880",
      "postDate": "09/29/2024 11:46:37",
      "content": "<p>Thank you for your comment, and I apologize for posting before reviewing all the available topics. I did check your excellent topics before creating mine. However, I thought there wasn’t a specific discussion summarizing outliers, so I created this topic to share what I had found. That said, it seems that Antonina Dolgorukova has already conducted a comprehensive analysis of outliers, making my topic less necessary. I’m still new to Kaggle, so I apologize for not fully understanding the rules. Thank you again for your valuable feedback!</p>",
      "rawMarkdown": "Thank you for your comment, and I apologize for posting before reviewing all the available topics. I did check your excellent topics before creating mine. However, I thought there wasn’t a specific discussion summarizing outliers, so I created this topic to share what I had found. That said, it seems that Antonina Dolgorukova has already conducted a comprehensive analysis of outliers, making my topic less necessary. I’m still new to Kaggle, so I apologize for not fully understanding the rules. Thank you again for your valuable feedback!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3001774,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "09/29/2024 09:00:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/masatoito8823\" target=\"_blank\">@masatoito8823</a>, there are several public EDA notebooks which show many outliers.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3001880,
          "author_name": "masatoito8823",
          "author_url": "",
          "post_date": "09/29/2024 11:46:37",
          "content": "<p>Thank you for your comment, and I apologize for posting before reviewing all the available topics. I did check your excellent topics before creating mine. However, I thought there wasn’t a specific discussion summarizing outliers, so I created this topic to share what I had found. That said, it seems that Antonina Dolgorukova has already conducted a comprehensive analysis of outliers, making my topic less necessary. I’m still new to Kaggle, so I apologize for not fully understanding the rules. Thank you again for your valuable feedback!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3001796,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "09/29/2024 09:30:36",
      "content": "<p>Hi! I have already shared everything I have found so far, apart from the EDA notebooks, it is summarised in the topic \"Some findings from the features EDA (+ data issues, outliers)\", and also, you can see e.g. \"Worst three time series of the dataset\" - for insights into the time series actigraphy data :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3001867,
          "author_name": "masatoito8823",
          "author_url": "",
          "post_date": "09/29/2024 11:29:58",
          "content": "<p>I apologize for overlooking your contributions, and thank you so much for taking the time to comment! I’ve only just started reviewing the data, and it seems you’ve already conducted a very thorough analysis. I will definitely refer to your topic and notebooks for my future work. Thanks again for sharing your insights!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3001762": "Hi everyone,\nWhen data is missing, it's usually recorded as NaN. However, I’ve noticed an unusual value in the training data. Specifically, in row 2065 (id = 83525bbe), the \"CGAS-CGAS_Score\" column has a value of 999, even though this column typically contains values between 0 and 100.\n\nThis kind of impossible value should ideally be removed during preprocessing, but I realize it can be time-consuming for each of us to identify these anomalies. If you’ve come across any similar issues, feel free to share them here to help everyone out!\n\nLooking forward to hearing from you!\n\n**Update:**\n\nAs some of the comments have pointed out, many reports on outliers have already been made in Antonina Dolgorukova’s topic ([https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535354](url)). I recommend referring to the content and notebooks attached there first. \nAlso, outliers and noise in the Actigraphy files are being discussed here as well: [https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535900](url).\n\nIf I become aware of other topics discussing outliers, I will update this post accordingly. Please feel free to comment if you notice anything missing!\n\nWhile this topic may no longer hold much value, I’ll leave it up for beginners who might encounter similar issues. If this topic causes any confusion in the discussion, please feel free to let me know, and I will delete it.",
    "3001774": "Hi @masatoito8823, there are several public EDA notebooks which show many outliers.",
    "3001796": "Hi! I have already shared everything I have found so far, apart from the EDA notebooks, it is summarised in the topic \"Some findings from the features EDA (+ data issues, outliers)\", and also, you can see e.g. \"Worst three time series of the dataset\" - for insights into the time series actigraphy data :)",
    "3001867": "I apologize for overlooking your contributions, and thank you so much for taking the time to comment! I’ve only just started reviewing the data, and it seems you’ve already conducted a very thorough analysis. I will definitely refer to your topic and notebooks for my future work. Thanks again for sharing your insights!",
    "3001880": "Thank you for your comment, and I apologize for posting before reviewing all the available topics. I did check your excellent topics before creating mine. However, I thought there wasn’t a specific discussion summarizing outliers, so I created this topic to share what I had found. That said, it seems that Antonina Dolgorukova has already conducted a comprehensive analysis of outliers, making my topic less necessary. I’m still new to Kaggle, so I apologize for not fully understanding the rules. Thank you again for your valuable feedback!"
  },
  "source": "meta"
}