{
  "id": 540237,
  "title": "unsupervised learning approach",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/540237",
  "author_name": "",
  "post_date": "2024-10-13T15:39:41.625365800Z",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>as it mentionned in the overview in the competition there is a lost of missing data in the traning data (both features and targets (sii)) , i noticed that most the public notebooks note use unsupervised approaches:<br>\nso we can try:<br>\n1==&gt; clustering <br>\n2==&gt; PCA<br>\n3==&gt; Sampling<br>\n4==&gt; Predicting missing sii instead of filling them or get rid off the row with NANS VALUES</p>",
  "messages": [
    {
      "id": "3016305",
      "postDate": "10/13/2024 15:39:41",
      "content": "<p>as it mentionned in the overview in the competition there is a lost of missing data in the traning data (both features and targets (sii)) , i noticed that most the public notebooks note use unsupervised approaches:<br>\nso we can try:<br>\n1==&gt; clustering <br>\n2==&gt; PCA<br>\n3==&gt; Sampling<br>\n4==&gt; Predicting missing sii instead of filling them or get rid off the row with NANS VALUES</p>",
      "rawMarkdown": "as it mentionned in the overview in the competition there is a lost of missing data in the traning data (both features and targets (sii)) , i noticed that most the public notebooks note use unsupervised approaches:\nso we can try:\n1==> clustering \n2==> PCA\n3==> Sampling\n4==> Predicting missing sii instead of filling them or get rid off the row with NANS VALUES",
      "votes": null
    },
    {
      "id": "3016411",
      "postDate": "10/13/2024 17:36:51",
      "content": "<p>Maybe. However, unsupervised learning will likely just label the missing Sii with mostly 0s or 1s. I’ve tried a few of these techniques already.  </p>\n<p>I feel like we need to look at something bit advanced for tackling this missing value issue. Any ideas from just raw mathematics to extrapolate the distribution of training data?</p>",
      "rawMarkdown": "Maybe. However, unsupervised learning will likely just label the missing Sii with mostly 0s or 1s. I’ve tried a few of these techniques already.  \n\nI feel like we need to look at something bit advanced for tackling this missing value issue. Any ideas from just raw mathematics to extrapolate the distribution of training data?",
      "votes": null
    },
    {
      "id": "3016423",
      "postDate": "10/13/2024 17:54:34",
      "content": "<p>thee approch tha consist build a model to predict the messing value ( associated with high probabilities) , i don't implemented this yet ,</p>",
      "rawMarkdown": "thee approch tha consist build a model to predict the messing value ( associated with high probabilities) , i don't implemented this yet ,",
      "votes": null
    },
    {
      "id": "3016434",
      "postDate": "10/13/2024 18:18:00",
      "content": "<p>Let me give that a try.🙂</p>",
      "rawMarkdown": "Let me give that a try.🙂",
      "votes": null
    },
    {
      "id": "3016442",
      "postDate": "10/13/2024 18:32:39",
      "content": "<p>good luck bro, did you use kaggle notebook to build your models or you are using your local machine for good experiments</p>",
      "rawMarkdown": "good luck bro, did you use kaggle notebook to build your models or you are using your local machine for good experiments",
      "votes": null
    },
    {
      "id": "3016451",
      "postDate": "10/13/2024 18:40:35",
      "content": "<p>kaggle notebooks. If you want to start you can start from one of my notebook.</p>\n<p><a href=\"https://www.kaggle.com/code/chanpreetsingh07/3b1b-jcm\" target=\"_blank\">https://www.kaggle.com/code/chanpreetsingh07/3b1b-jcm</a></p>",
      "rawMarkdown": "kaggle notebooks. If you want to start you can start from one of my notebook.\n\nhttps://www.kaggle.com/code/chanpreetsingh07/3b1b-jcm",
      "votes": null
    },
    {
      "id": "3020648",
      "postDate": "10/17/2024 18:25:32",
      "content": "<p>You can do the coding in your local machine, open kaggle notebook, cut, paste and run. This is because, some people are more comfortable with their local machines.</p>",
      "rawMarkdown": "You can do the coding in your local machine, open kaggle notebook, cut, paste and run. This is because, some people are more comfortable with their local machines.",
      "votes": null
    },
    {
      "id": "3022937",
      "postDate": "10/20/2024 05:16:11",
      "content": "<p>yes especially for complex competitions ( deep learning models) , when you need to create projects in the form of ensemble of .py files python .json files (for experiments ) …etc such the following example :<br>\nmy-deep-learning-project/<br>\n│<br>\n├── README.md               # Overview of the project<br>\n├── requirements.txt        # Dependencies and packages<br>\n├── data/<br>\n│   ├── raw/                # Raw data files<br>\n│   ├── processed/          # Processed data files<br>\n│   └── external/           # External datasets or sources<br>\n│<br>\n├── notebooks/              # Jupyter notebooks for exploration<br>\n│   ├── exploratory_analysis.ipynb<br>\n│   └── model_evaluation.ipynb<br>\n│<br>\n├── src/                   # Source code<br>\n│   ├── data_preprocessing.py  # Data cleaning and preprocessing functions<br>\n│   ├── model.py              # Model architecture definition<br>\n│   ├── train.py              # Training script<br>\n│   └── evaluate.py           # Evaluation script<br>\n│<br>\n├── experiments/            # Experiment tracking and results<br>\n│   ├── experiment_1/        # Folder for a specific experiment<br>\n│   │   ├── logs/            # Training logs<br>\n│   │   ├── results.csv      # Results from this experiment<br>\n│   │   └── model_weights.h5  # Saved model weights (Tensorflow) or .pyth for Pytorch<br>\n│   └── experiment_2/<br>\n│<br>\n└── scripts/                # Utility scripts<br>\n    ├── plot_results.py      # Visualization scripts<br>\n    └── config.py            # Configuration settings</p>",
      "rawMarkdown": "yes especially for complex competitions ( deep learning models) , when you need to create projects in the form of ensemble of .py files python .json files (for experiments ) ...etc such the following example :\nmy-deep-learning-project/\n│\n├── README.md               # Overview of the project\n├── requirements.txt        # Dependencies and packages\n├── data/\n│   ├── raw/                # Raw data files\n│   ├── processed/          # Processed data files\n│   └── external/           # External datasets or sources\n│\n├── notebooks/              # Jupyter notebooks for exploration\n│   ├── exploratory_analysis.ipynb\n│   └── model_evaluation.ipynb\n│\n├── src/                   # Source code\n│   ├── data_preprocessing.py  # Data cleaning and preprocessing functions\n│   ├── model.py              # Model architecture definition\n│   ├── train.py              # Training script\n│   └── evaluate.py           # Evaluation script\n│\n├── experiments/            # Experiment tracking and results\n│   ├── experiment_1/        # Folder for a specific experiment\n│   │   ├── logs/            # Training logs\n│   │   ├── results.csv      # Results from this experiment\n│   │   └── model_weights.h5  # Saved model weights (Tensorflow) or .pyth for Pytorch\n│   └── experiment_2/\n│\n└── scripts/                # Utility scripts\n    ├── plot_results.py      # Visualization scripts\n    └── config.py            # Configuration settings",
      "votes": null
    },
    {
      "id": "3520542",
      "postDate": "09/03/2026 16:12:53",
      "content": "<p>What if the most important patterns in this data are the ones nobody labeled? 👀\n<a href=\"https://mlguidance.blogspot.com/2026/08/what-are-models-of-unsupervised-learning.html\" target=\"_blank\">https://mlguidance.blogspot.com/2026/08/what-are-models-of-unsupervised-learning.html</a></p>",
      "rawMarkdown": "What if the most important patterns in this data are the ones nobody labeled? 👀\nhttps://mlguidance.blogspot.com/2026/08/what-are-models-of-unsupervised-learning.html",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3016411,
      "author_name": "chanpreetsingh07",
      "author_url": "",
      "post_date": "10/13/2024 17:36:51",
      "content": "<p>Maybe. However, unsupervised learning will likely just label the missing Sii with mostly 0s or 1s. I’ve tried a few of these techniques already.  </p>\n<p>I feel like we need to look at something bit advanced for tackling this missing value issue. Any ideas from just raw mathematics to extrapolate the distribution of training data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3016423,
          "author_name": "saidkoussi",
          "author_url": "",
          "post_date": "10/13/2024 17:54:34",
          "content": "<p>thee approch tha consist build a model to predict the messing value ( associated with high probabilities) , i don't implemented this yet ,</p>",
          "votes": null,
          "replies": [
            {
              "id": 3016434,
              "author_name": "chanpreetsingh07",
              "author_url": "",
              "post_date": "10/13/2024 18:18:00",
              "content": "<p>Let me give that a try.🙂</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3016442,
                  "author_name": "saidkoussi",
                  "author_url": "",
                  "post_date": "10/13/2024 18:32:39",
                  "content": "<p>good luck bro, did you use kaggle notebook to build your models or you are using your local machine for good experiments</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3016451,
                      "author_name": "chanpreetsingh07",
                      "author_url": "",
                      "post_date": "10/13/2024 18:40:35",
                      "content": "<p>kaggle notebooks. If you want to start you can start from one of my notebook.</p>\n<p><a href=\"https://www.kaggle.com/code/chanpreetsingh07/3b1b-jcm\" target=\"_blank\">https://www.kaggle.com/code/chanpreetsingh07/3b1b-jcm</a></p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 3020648,
                      "author_name": "sherriffasherriff",
                      "author_url": "",
                      "post_date": "10/17/2024 18:25:32",
                      "content": "<p>You can do the coding in your local machine, open kaggle notebook, cut, paste and run. This is because, some people are more comfortable with their local machines.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3022937,
                          "author_name": "saidkoussi",
                          "author_url": "",
                          "post_date": "10/20/2024 05:16:11",
                          "content": "<p>yes especially for complex competitions ( deep learning models) , when you need to create projects in the form of ensemble of .py files python .json files (for experiments ) …etc such the following example :<br>\nmy-deep-learning-project/<br>\n│<br>\n├── README.md               # Overview of the project<br>\n├── requirements.txt        # Dependencies and packages<br>\n├── data/<br>\n│   ├── raw/                # Raw data files<br>\n│   ├── processed/          # Processed data files<br>\n│   └── external/           # External datasets or sources<br>\n│<br>\n├── notebooks/              # Jupyter notebooks for exploration<br>\n│   ├── exploratory_analysis.ipynb<br>\n│   └── model_evaluation.ipynb<br>\n│<br>\n├── src/                   # Source code<br>\n│   ├── data_preprocessing.py  # Data cleaning and preprocessing functions<br>\n│   ├── model.py              # Model architecture definition<br>\n│   ├── train.py              # Training script<br>\n│   └── evaluate.py           # Evaluation script<br>\n│<br>\n├── experiments/            # Experiment tracking and results<br>\n│   ├── experiment_1/        # Folder for a specific experiment<br>\n│   │   ├── logs/            # Training logs<br>\n│   │   ├── results.csv      # Results from this experiment<br>\n│   │   └── model_weights.h5  # Saved model weights (Tensorflow) or .pyth for Pytorch<br>\n│   └── experiment_2/<br>\n│<br>\n└── scripts/                # Utility scripts<br>\n    ├── plot_results.py      # Visualization scripts<br>\n    └── config.py            # Configuration settings</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3520542,
      "author_name": "kaustubh994",
      "author_url": "",
      "post_date": "09/03/2026 16:12:53",
      "content": "<p>What if the most important patterns in this data are the ones nobody labeled? 👀\n<a href=\"https://mlguidance.blogspot.com/2026/08/what-are-models-of-unsupervised-learning.html\" target=\"_blank\">https://mlguidance.blogspot.com/2026/08/what-are-models-of-unsupervised-learning.html</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3016305": "as it mentionned in the overview in the competition there is a lost of missing data in the traning data (both features and targets (sii)) , i noticed that most the public notebooks note use unsupervised approaches:\nso we can try:\n1==> clustering \n2==> PCA\n3==> Sampling\n4==> Predicting missing sii instead of filling them or get rid off the row with NANS VALUES",
    "3016411": "Maybe. However, unsupervised learning will likely just label the missing Sii with mostly 0s or 1s. I’ve tried a few of these techniques already.  \n\nI feel like we need to look at something bit advanced for tackling this missing value issue. Any ideas from just raw mathematics to extrapolate the distribution of training data?",
    "3016423": "thee approch tha consist build a model to predict the messing value ( associated with high probabilities) , i don't implemented this yet ,",
    "3016434": "Let me give that a try.🙂",
    "3016442": "good luck bro, did you use kaggle notebook to build your models or you are using your local machine for good experiments",
    "3016451": "kaggle notebooks. If you want to start you can start from one of my notebook.\n\nhttps://www.kaggle.com/code/chanpreetsingh07/3b1b-jcm",
    "3020648": "You can do the coding in your local machine, open kaggle notebook, cut, paste and run. This is because, some people are more comfortable with their local machines.",
    "3022937": "yes especially for complex competitions ( deep learning models) , when you need to create projects in the form of ensemble of .py files python .json files (for experiments ) ...etc such the following example :\nmy-deep-learning-project/\n│\n├── README.md               # Overview of the project\n├── requirements.txt        # Dependencies and packages\n├── data/\n│   ├── raw/                # Raw data files\n│   ├── processed/          # Processed data files\n│   └── external/           # External datasets or sources\n│\n├── notebooks/              # Jupyter notebooks for exploration\n│   ├── exploratory_analysis.ipynb\n│   └── model_evaluation.ipynb\n│\n├── src/                   # Source code\n│   ├── data_preprocessing.py  # Data cleaning and preprocessing functions\n│   ├── model.py              # Model architecture definition\n│   ├── train.py              # Training script\n│   └── evaluate.py           # Evaluation script\n│\n├── experiments/            # Experiment tracking and results\n│   ├── experiment_1/        # Folder for a specific experiment\n│   │   ├── logs/            # Training logs\n│   │   ├── results.csv      # Results from this experiment\n│   │   └── model_weights.h5  # Saved model weights (Tensorflow) or .pyth for Pytorch\n│   └── experiment_2/\n│\n└── scripts/                # Utility scripts\n    ├── plot_results.py      # Visualization scripts\n    └── config.py            # Configuration settings",
    "3520542": "What if the most important patterns in this data are the ones nobody labeled? 👀\nhttps://mlguidance.blogspot.com/2026/08/what-are-models-of-unsupervised-learning.html"
  },
  "source": "meta"
}