{
  "id": 349091,
  "title": "What are examples and state of art on 200 000 SPARSE features and 20 000 targets ?",
  "url": "/competitions/open-problems-multimodal/discussion/349091",
  "author_name": "Alexander Chervov",
  "post_date": "2022-08-31T08:13:05.838000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I guess one the main interests of the competition is that we have  200 000 SPARSE features and 20 000 targets .<br>\nWhat are the good practices, examples, state of art methods to work with such situation ? <br>\nAnd where they might occur. </p>\n<p>200 000 SPARSE(!) features might occur in NLP - when use tf-idf coding of texts,<br>\nbut seems to me 20 000 targets is not common for  NLP.</p>\n<p>There are some ideas already discussed:<br>\n1) Use dimensional reduction 200 000 targets will be reduced. (Comment: not sure that is way to SOTA, you lose specifics of SPARSE (the same you can do if it not sparce) - but may be I am wrong).<br>\n2) Try to use some Sparse layers in pytorch (never worked with it - as far as I understand pytorch/gpu do not quite like sparce matrices)</p>\n<p>For 20 000 targets - seems to be neural network is natural approach, not gradient boosting, despite the data is tabular. </p>\n<p>Any way I guess there should literature on such tasks - I will look for it by myself later - but may me some one knows of hand…</p>",
  "messages": [
    {
      "id": 1920556,
      "postDate": "2022-08-31T08:13:05.840Z",
      "content": "<p>I guess one the main interests of the competition is that we have  200 000 SPARSE features and 20 000 targets .<br>\nWhat are the good practices, examples, state of art methods to work with such situation ? <br>\nAnd where they might occur. </p>\n<p>200 000 SPARSE(!) features might occur in NLP - when use tf-idf coding of texts,<br>\nbut seems to me 20 000 targets is not common for  NLP.</p>\n<p>There are some ideas already discussed:<br>\n1) Use dimensional reduction 200 000 targets will be reduced. (Comment: not sure that is way to SOTA, you lose specifics of SPARSE (the same you can do if it not sparce) - but may be I am wrong).<br>\n2) Try to use some Sparse layers in pytorch (never worked with it - as far as I understand pytorch/gpu do not quite like sparce matrices)</p>\n<p>For 20 000 targets - seems to be neural network is natural approach, not gradient boosting, despite the data is tabular. </p>\n<p>Any way I guess there should literature on such tasks - I will look for it by myself later - but may me some one knows of hand…</p>",
      "rawMarkdown": "I guess one the main interests of the competition is that we have  200 000 SPARSE features and 20 000 targets .\nWhat are the good practices, examples, state of art methods to work with such situation ? \nAnd where they might occur. \n\n200 000 SPARSE(!) features might occur in NLP - when use tf-idf coding of texts,\nbut seems to me 20 000 targets is not common for  NLP.\n\nThere are some ideas already discussed:\n1) Use dimensional reduction 200 000 targets will be reduced. (Comment: not sure that is way to SOTA, you lose specifics of SPARSE (the same you can do if it not sparce) - but may be I am wrong).\n2) Try to use some Sparse layers in pytorch (never worked with it - as far as I understand pytorch/gpu do not quite like sparce matrices)\n\nFor 20 000 targets - seems to be neural network is natural approach, not gradient boosting, despite the data is tabular. \n\nAny way I guess there should literature on such tasks - I will look for it by myself later - but may me some one knows of hand...\n\n\n\n "
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1920556": "I guess one the main interests of the competition is that we have  200 000 SPARSE features and 20 000 targets .\nWhat are the good practices, examples, state of art methods to work with such situation ? \nAnd where they might occur. \n\n200 000 SPARSE(!) features might occur in NLP - when use tf-idf coding of texts,\nbut seems to me 20 000 targets is not common for  NLP.\n\nThere are some ideas already discussed:\n1) Use dimensional reduction 200 000 targets will be reduced. (Comment: not sure that is way to SOTA, you lose specifics of SPARSE (the same you can do if it not sparce) - but may be I am wrong).\n2) Try to use some Sparse layers in pytorch (never worked with it - as far as I understand pytorch/gpu do not quite like sparce matrices)\n\nFor 20 000 targets - seems to be neural network is natural approach, not gradient boosting, despite the data is tabular. \n\nAny way I guess there should literature on such tasks - I will look for it by myself later - but may me some one knows of hand...\n\n\n\n "
  }
}