{
  "id": 151922,
  "title": "New to ML",
  "url": "/competitions/imet-2020-fgvc7/discussion/151922",
  "author_name": "",
  "post_date": "2020-05-17T17:49:30.710263400Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi friends,\nI am completely new to the world of Machine learning. by looking at output column \"attribute_ids \" in train.csv, I am clue less. can any one guide me to recognize the type of problem. As per my understanding it should be \"supervised ML\" because it contains both <strong>feature(id)</strong> and <strong>label(attribute_ids)</strong>. But how to handle space separated values in attribute_ids column. Any help would be highly appreciated.</p>",
  "messages": [
    {
      "id": "851551",
      "postDate": "05/17/2020 17:49:30",
      "content": "<p>Hi friends,\nI am completely new to the world of Machine learning. by looking at output column \"attribute_ids \" in train.csv, I am clue less. can any one guide me to recognize the type of problem. As per my understanding it should be \"supervised ML\" because it contains both <strong>feature(id)</strong> and <strong>label(attribute_ids)</strong>. But how to handle space separated values in attribute_ids column. Any help would be highly appreciated.</p>",
      "rawMarkdown": "Hi friends,\nI am completely new to the world of Machine learning. by looking at output column \"attribute_ids \" in train.csv, I am clue less. can any one guide me to recognize the type of problem. As per my understanding it should be \"supervised ML\" because it contains both **feature(id)** and **label(attribute_ids)**. But how to handle space separated values in attribute_ids column. Any help would be highly appreciated.",
      "votes": null
    },
    {
      "id": "851967",
      "postDate": "05/18/2020 04:31:29",
      "content": "<p>pandas.read_csv(filename, sep = ' ')</p>\n\n<p><strong>EDIT</strong></p>\n\n<p>sorry never mind that's for if a file is space separated instead of some other separator, probably using str.split() should do the trick</p>\n\n<p><strong>EDIT2</strong></p>\n\n<p>also this all assuming you are using python</p>\n\n<p><strong>EDIT3</strong></p>\n\n<p>ask if this is not making sense, I'm not sure how new you are?</p>",
      "rawMarkdown": "pandas.read_csv(filename, sep = ' ')\n\n**EDIT**\n\nsorry never mind that's for if a file is space separated instead of some other separator, probably using str.split() should do the trick\n\n**EDIT2**\n\nalso this all assuming you are using python\n\n**EDIT3**\n\nask if this is not making sense, I'm not sure how new you are?",
      "votes": null
    },
    {
      "id": "852869",
      "postDate": "05/18/2020 18:33:55",
      "content": "<p>Thanks Vishesh for reply. my question is bit different. The output column \"attribute_ids\" contains space separated multiple values. so far I have not faced any problem where output coulmn have space separated values. I have dealt with only single valued output column. which algorithm should i use to solve this type of problem.</p>",
      "rawMarkdown": "Thanks Vishesh for reply. my question is bit different. The output column \"attribute_ids\" contains space separated multiple values. so far I have not faced any problem where output coulmn have space separated values. I have dealt with only single valued output column. which algorithm should i use to solve this type of problem.",
      "votes": null
    },
    {
      "id": "852871",
      "postDate": "05/18/2020 18:34:54",
      "content": "<p>Thanks Vishesh for reply. my question is bit different. The output column \"attribute_ids\" contains space separated multiple values. so far I have not faced any problem where output coulmn have space separated values. I have dealt with only single valued output column. which algorithm should i use to solve this type of problem.</p>",
      "rawMarkdown": "Thanks Vishesh for reply. my question is bit different. The output column \"attribute_ids\" contains space separated multiple values. so far I have not faced any problem where output coulmn have space separated values. I have dealt with only single valued output column. which algorithm should i use to solve this type of problem.",
      "votes": null
    },
    {
      "id": "853558",
      "postDate": "05/19/2020 09:35:21",
      "content": "<p>Hmmm, so your question is how to deal with multi-class prediction? Because each painting could belong to several attributes, and you have to predict those attributes? Well, I couldn't tell you specifically but like it's not any specific model exists particularly built for predicting multiple tags. There is many different ones you could try and see what works? A quick google got me to this link:</p>\n\n<p><a href=\"https://towardsdatascience.com/regression-models-with-multiple-target-variables-8baa75aacd\">https://towardsdatascience.com/regression-models-with-multiple-target-variables-8baa75aacd</a></p>\n\n<p>Basically seems like ensemble decision trees is a good bet. But I'm sure there are others if you keep googling and what not. Furthermore, don't get beholden to a model, always experiment. Sorry I couldn't be of more help. </p>",
      "rawMarkdown": "Hmmm, so your question is how to deal with multi-class prediction? Because each painting could belong to several attributes, and you have to predict those attributes? Well, I couldn't tell you specifically but like it's not any specific model exists particularly built for predicting multiple tags. There is many different ones you could try and see what works? A quick google got me to this link:\n\nhttps://towardsdatascience.com/regression-models-with-multiple-target-variables-8baa75aacd\n\nBasically seems like ensemble decision trees is a good bet. But I'm sure there are others if you keep googling and what not. Furthermore, don't get beholden to a model, always experiment. Sorry I couldn't be of more help.",
      "votes": null
    },
    {
      "id": "853580",
      "postDate": "05/19/2020 09:56:25",
      "content": "<p>Thanks Vishesh!</p>",
      "rawMarkdown": "Thanks Vishesh!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 851967,
      "author_name": "rarebrownpanda",
      "author_url": "",
      "post_date": "05/18/2020 04:31:29",
      "content": "<p>pandas.read_csv(filename, sep = ' ')</p>\n\n<p><strong>EDIT</strong></p>\n\n<p>sorry never mind that's for if a file is space separated instead of some other separator, probably using str.split() should do the trick</p>\n\n<p><strong>EDIT2</strong></p>\n\n<p>also this all assuming you are using python</p>\n\n<p><strong>EDIT3</strong></p>\n\n<p>ask if this is not making sense, I'm not sure how new you are?</p>",
      "votes": null,
      "replies": [
        {
          "id": 852871,
          "author_name": "sujit143",
          "author_url": "",
          "post_date": "05/18/2020 18:34:54",
          "content": "<p>Thanks Vishesh for reply. my question is bit different. The output column \"attribute_ids\" contains space separated multiple values. so far I have not faced any problem where output coulmn have space separated values. I have dealt with only single valued output column. which algorithm should i use to solve this type of problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 853558,
          "author_name": "rarebrownpanda",
          "author_url": "",
          "post_date": "05/19/2020 09:35:21",
          "content": "<p>Hmmm, so your question is how to deal with multi-class prediction? Because each painting could belong to several attributes, and you have to predict those attributes? Well, I couldn't tell you specifically but like it's not any specific model exists particularly built for predicting multiple tags. There is many different ones you could try and see what works? A quick google got me to this link:</p>\n\n<p><a href=\"https://towardsdatascience.com/regression-models-with-multiple-target-variables-8baa75aacd\">https://towardsdatascience.com/regression-models-with-multiple-target-variables-8baa75aacd</a></p>\n\n<p>Basically seems like ensemble decision trees is a good bet. But I'm sure there are others if you keep googling and what not. Furthermore, don't get beholden to a model, always experiment. Sorry I couldn't be of more help. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 853580,
          "author_name": "sujit143",
          "author_url": "",
          "post_date": "05/19/2020 09:56:25",
          "content": "<p>Thanks Vishesh!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 852869,
      "author_name": "sujit143",
      "author_url": "",
      "post_date": "05/18/2020 18:33:55",
      "content": "<p>Thanks Vishesh for reply. my question is bit different. The output column \"attribute_ids\" contains space separated multiple values. so far I have not faced any problem where output coulmn have space separated values. I have dealt with only single valued output column. which algorithm should i use to solve this type of problem.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "851551": "Hi friends,\nI am completely new to the world of Machine learning. by looking at output column \"attribute_ids \" in train.csv, I am clue less. can any one guide me to recognize the type of problem. As per my understanding it should be \"supervised ML\" because it contains both **feature(id)** and **label(attribute_ids)**. But how to handle space separated values in attribute_ids column. Any help would be highly appreciated.",
    "851967": "pandas.read_csv(filename, sep = ' ')\n\n**EDIT**\n\nsorry never mind that's for if a file is space separated instead of some other separator, probably using str.split() should do the trick\n\n**EDIT2**\n\nalso this all assuming you are using python\n\n**EDIT3**\n\nask if this is not making sense, I'm not sure how new you are?",
    "852869": "Thanks Vishesh for reply. my question is bit different. The output column \"attribute_ids\" contains space separated multiple values. so far I have not faced any problem where output coulmn have space separated values. I have dealt with only single valued output column. which algorithm should i use to solve this type of problem.",
    "852871": "Thanks Vishesh for reply. my question is bit different. The output column \"attribute_ids\" contains space separated multiple values. so far I have not faced any problem where output coulmn have space separated values. I have dealt with only single valued output column. which algorithm should i use to solve this type of problem.",
    "853558": "Hmmm, so your question is how to deal with multi-class prediction? Because each painting could belong to several attributes, and you have to predict those attributes? Well, I couldn't tell you specifically but like it's not any specific model exists particularly built for predicting multiple tags. There is many different ones you could try and see what works? A quick google got me to this link:\n\nhttps://towardsdatascience.com/regression-models-with-multiple-target-variables-8baa75aacd\n\nBasically seems like ensemble decision trees is a good bet. But I'm sure there are others if you keep googling and what not. Furthermore, don't get beholden to a model, always experiment. Sorry I couldn't be of more help.",
    "853580": "Thanks Vishesh!"
  },
  "source": "meta"
}