{
  "id": 360002,
  "title": "Question about row vs. column normalization",
  "url": "/competitions/open-problems-multimodal/discussion/360002",
  "author_name": "",
  "post_date": "2022-10-14T16:14:41.609435200Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>When I first started on this competition, I used column normalization as usual only to find out a few days later that the Pearson loss and everybody else was using row normalization.</p>\n<p>The proteins in a cell are not uncorrelated whereas the cells for each protein are (up to the fact that they come from the same donor and day). Usually, in machine learning, we try to predict many uncorrelated measurements of the same target feature given some input features.</p>\n<p>Could someone explain to me the rationale for doing it this way?</p>",
  "messages": [
    {
      "id": "1987285",
      "postDate": "10/14/2022 16:14:41",
      "content": "<p>When I first started on this competition, I used column normalization as usual only to find out a few days later that the Pearson loss and everybody else was using row normalization.</p>\n<p>The proteins in a cell are not uncorrelated whereas the cells for each protein are (up to the fact that they come from the same donor and day). Usually, in machine learning, we try to predict many uncorrelated measurements of the same target feature given some input features.</p>\n<p>Could someone explain to me the rationale for doing it this way?</p>",
      "rawMarkdown": "When I first started on this competition, I used column normalization as usual only to find out a few days later that the Pearson loss and everybody else was using row normalization.\n\nThe proteins in a cell are not uncorrelated whereas the cells for each protein are (up to the fact that they come from the same donor and day). Usually, in machine learning, we try to predict many uncorrelated measurements of the same target feature given some input features.\n\nCould someone explain to me the rationale for doing it this way?",
      "votes": null
    },
    {
      "id": "1988065",
      "postDate": "10/15/2022 04:08:36",
      "content": "<p>I guess it's a question of priority: do you care about each kind of protein the same amount or about each cell the same amount.</p>",
      "rawMarkdown": "I guess it's a question of priority: do you care about each kind of protein the same amount or about each cell the same amount.",
      "votes": null
    },
    {
      "id": "1988094",
      "postDate": "10/15/2022 04:41:28",
      "content": "<p>What do you mean by \"everybody else was using row normalization\"? I thought the data was already row-normalized, since the data page says the inputs are \"library-size normalized\". </p>\n<p>Maybe the rational for this is that protein counts vary greatly between cells, but we are more concerned with proteins levels relative to other proteins within the same cell. But I'm not a domain expert. </p>\n<p>For the Pearson correlation choice of metric, I think it is because they didn't like the noisy results of last year's competition, where the metric was some kind of absolute error. </p>",
      "rawMarkdown": "What do you mean by \"everybody else was using row normalization\"? I thought the data was already row-normalized, since the data page says the inputs are \"library-size normalized\". \n\nMaybe the rational for this is that protein counts vary greatly between cells, but we are more concerned with proteins levels relative to other proteins within the same cell. But I'm not a domain expert. \n\nFor the Pearson correlation choice of metric, I think it is because they didn't like the noisy results of last year's competition, where the metric was some kind of absolute error.",
      "votes": null
    },
    {
      "id": "1988102",
      "postDate": "10/15/2022 04:52:25",
      "content": "<p>\"Usually, in machine learning, we try to predict many <strong>uncorrelated</strong> measurements of the same target feature given some input features.\"</p>\n<p>I would not agree - usually we need to predict things which our client/customer wants to be predicted, client does not care about \"uncorrelated\".</p>\n<p>PS<br>\nThere are many details about normalizations in that competition. <br>\nBut I am not sure I fully understand your question. </p>",
      "rawMarkdown": "\"Usually, in machine learning, we try to predict many **uncorrelated** measurements of the same target feature given some input features.\"\n\nI would not agree - usually we need to predict things which our client/customer wants to be predicted, client does not care about \"uncorrelated\".\n\nPS\nThere are many details about normalizations in that competition. \nBut I am not sure I fully understand your question.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1988065,
      "author_name": "alexroberts",
      "author_url": "",
      "post_date": "10/15/2022 04:08:36",
      "content": "<p>I guess it's a question of priority: do you care about each kind of protein the same amount or about each cell the same amount.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1988094,
      "author_name": "chrisrichardmiles",
      "author_url": "",
      "post_date": "10/15/2022 04:41:28",
      "content": "<p>What do you mean by \"everybody else was using row normalization\"? I thought the data was already row-normalized, since the data page says the inputs are \"library-size normalized\". </p>\n<p>Maybe the rational for this is that protein counts vary greatly between cells, but we are more concerned with proteins levels relative to other proteins within the same cell. But I'm not a domain expert. </p>\n<p>For the Pearson correlation choice of metric, I think it is because they didn't like the noisy results of last year's competition, where the metric was some kind of absolute error. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1988102,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "10/15/2022 04:52:25",
      "content": "<p>\"Usually, in machine learning, we try to predict many <strong>uncorrelated</strong> measurements of the same target feature given some input features.\"</p>\n<p>I would not agree - usually we need to predict things which our client/customer wants to be predicted, client does not care about \"uncorrelated\".</p>\n<p>PS<br>\nThere are many details about normalizations in that competition. <br>\nBut I am not sure I fully understand your question. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1987285": "When I first started on this competition, I used column normalization as usual only to find out a few days later that the Pearson loss and everybody else was using row normalization.\n\nThe proteins in a cell are not uncorrelated whereas the cells for each protein are (up to the fact that they come from the same donor and day). Usually, in machine learning, we try to predict many uncorrelated measurements of the same target feature given some input features.\n\nCould someone explain to me the rationale for doing it this way?",
    "1988065": "I guess it's a question of priority: do you care about each kind of protein the same amount or about each cell the same amount.",
    "1988094": "What do you mean by \"everybody else was using row normalization\"? I thought the data was already row-normalized, since the data page says the inputs are \"library-size normalized\". \n\nMaybe the rational for this is that protein counts vary greatly between cells, but we are more concerned with proteins levels relative to other proteins within the same cell. But I'm not a domain expert. \n\nFor the Pearson correlation choice of metric, I think it is because they didn't like the noisy results of last year's competition, where the metric was some kind of absolute error.",
    "1988102": "\"Usually, in machine learning, we try to predict many **uncorrelated** measurements of the same target feature given some input features.\"\n\nI would not agree - usually we need to predict things which our client/customer wants to be predicted, client does not care about \"uncorrelated\".\n\nPS\nThere are many details about normalizations in that competition. \nBut I am not sure I fully understand your question."
  },
  "source": "meta"
}