Three critical failures that undermine machine learning accuracy and fairness.
Training data can fail in three ways. First, it might not be sufficient. The data may lack context or situations the algorithm will encounter in the real world. Second, it might be biased. This means data is systematically missing certain perspectives or groups. For example, if texts predominantly show nurses as women, the algorithm learns that pattern without understanding societal goals like gender equality. Third, the data might be outright wrong. Labels may be inaccurate or based on performances rather than reality, leading algorithms to learn incorrect associations.
— It’s The Data, Stupid! Why AI Might Get It Wrong. · Forbes