{
  "id": 516905,
  "title": "I processed competition dataset with separated the positive binding data from the negative data.",
  "url": "/competitions/leash-BELKA/discussion/516905",
  "author_name": "joejeo1",
  "post_date": "2024-07-04T03:50:59.036000",
  "votes": 0,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I processed competition dataset to separate the positive binding data from the negative data. I only included 20% of the negative binding data (randomly sampled) for easy handling of data.  You can easily sample the negative data relative to the positive data size.  <br>\n<a href=\"https://www.kaggle.com/code/joejeo1/belka-mini-dataset\" target=\"_blank\">https://www.kaggle.com/code/joejeo1/belka-mini-dataset</a><br>\n<a href=\"https://www.kaggle.com/datasets/joejeo1/belka-mini-dataset-with-folds\" target=\"_blank\">https://www.kaggle.com/datasets/joejeo1/belka-mini-dataset-with-folds</a></p>",
  "messages": [
    {
      "id": 2903873,
      "postDate": "2024-07-04T03:50:59.037Z",
      "content": "<p>I processed competition dataset to separate the positive binding data from the negative data. I only included 20% of the negative binding data (randomly sampled) for easy handling of data.  You can easily sample the negative data relative to the positive data size.  <br>\n<a href=\"https://www.kaggle.com/code/joejeo1/belka-mini-dataset\" target=\"_blank\">https://www.kaggle.com/code/joejeo1/belka-mini-dataset</a><br>\n<a href=\"https://www.kaggle.com/datasets/joejeo1/belka-mini-dataset-with-folds\" target=\"_blank\">https://www.kaggle.com/datasets/joejeo1/belka-mini-dataset-with-folds</a></p>",
      "rawMarkdown": "I processed competition dataset to separate the positive binding data from the negative data. I only included 20% of the negative binding data (randomly sampled) for easy handling of data.  You can easily sample the negative data relative to the positive data size.  \nhttps://www.kaggle.com/code/joejeo1/belka-mini-dataset\nhttps://www.kaggle.com/datasets/joejeo1/belka-mini-dataset-with-folds"
    },
    {
      "id": 2904932,
      "postDate": "2024-07-04T16:00:10.093Z",
      "content": "<p>Thanks for sharing this processed dataset and code! Processing positive combined data separately from negative data is a very useful step, especially when dealing with unbalanced data sets. Here are some of my thoughts and suggestions:</p>\n<p>Can you further explain why you chose to include only 20% negative binding data? What are the specific benefits of this for data processing and model training?<br>\nThe data sets and code links provided are very helpful, can you briefly explain what each link corresponds to in the post?<br>\nCan you discuss how to resize negative data relative to the size of positive data, and how this affects model training and performance?<br>\nOverall, it was a very valuable share and thank you for your contribution!</p>",
      "rawMarkdown": "Thanks for sharing this processed dataset and code! Processing positive combined data separately from negative data is a very useful step, especially when dealing with unbalanced data sets. Here are some of my thoughts and suggestions:\n\nCan you further explain why you chose to include only 20% negative binding data? What are the specific benefits of this for data processing and model training?\nThe data sets and code links provided are very helpful, can you briefly explain what each link corresponds to in the post?\nCan you discuss how to resize negative data relative to the size of positive data, and how this affects model training and performance?\nOverall, it was a very valuable share and thank you for your contribution!",
      "isDeleted": true,
      "replies": [
        {
          "id": 2904979,
          "postDate": "2024-07-04T16:23:00.503Z",
          "content": "<p>I think the Kaggle notebook has only ~29 GB RAM.   20% negative data should be fine with that.  I use the full negative dataset, that needs at least 64GB RAM.  </p>",
          "rawMarkdown": "I think the Kaggle notebook has only ~29 GB RAM.   20% negative data should be fine with that.  I use the full negative dataset, that needs at least 64GB RAM.  "
        },
        {
          "id": 2905428,
          "postDate": "2024-07-04T23:59:20.683Z",
          "content": "<p>I also uploaded the full negative dataset.  </p>",
          "rawMarkdown": "I also uploaded the full negative dataset.  "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2904932,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-04T16:00:10.093000",
      "content": "<p>Thanks for sharing this processed dataset and code! Processing positive combined data separately from negative data is a very useful step, especially when dealing with unbalanced data sets. Here are some of my thoughts and suggestions:</p>\n<p>Can you further explain why you chose to include only 20% negative binding data? What are the specific benefits of this for data processing and model training?<br>\nThe data sets and code links provided are very helpful, can you briefly explain what each link corresponds to in the post?<br>\nCan you discuss how to resize negative data relative to the size of positive data, and how this affects model training and performance?<br>\nOverall, it was a very valuable share and thank you for your contribution!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2904979,
          "author_name": "joejeo1",
          "author_url": "",
          "post_date": "2024-07-04T16:23:00.503000",
          "content": "<p>I think the Kaggle notebook has only ~29 GB RAM.   20% negative data should be fine with that.  I use the full negative dataset, that needs at least 64GB RAM.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2905428,
          "author_name": "joejeo1",
          "author_url": "",
          "post_date": "2024-07-04T23:59:20.683000",
          "content": "<p>I also uploaded the full negative dataset.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2903873": "I processed competition dataset to separate the positive binding data from the negative data. I only included 20% of the negative binding data (randomly sampled) for easy handling of data.  You can easily sample the negative data relative to the positive data size.  \nhttps://www.kaggle.com/code/joejeo1/belka-mini-dataset\nhttps://www.kaggle.com/datasets/joejeo1/belka-mini-dataset-with-folds",
    "2904932": "Thanks for sharing this processed dataset and code! Processing positive combined data separately from negative data is a very useful step, especially when dealing with unbalanced data sets. Here are some of my thoughts and suggestions:\n\nCan you further explain why you chose to include only 20% negative binding data? What are the specific benefits of this for data processing and model training?\nThe data sets and code links provided are very helpful, can you briefly explain what each link corresponds to in the post?\nCan you discuss how to resize negative data relative to the size of positive data, and how this affects model training and performance?\nOverall, it was a very valuable share and thank you for your contribution!"
  }
}